The "something something" video database for learning and evaluating visual common sense

Goyal, Raghav; Kahou, Samira Ebrahimi; Michalski, Vincent; Materzyńska, Joanna; Westphal, Susanne; Kim, Heuna; Haenel, Valentin; Fruend, Ingo; Yianilos, Peter; Mueller-Freitag, Moritz; Hoppe, Florian; Thurau, Christian; Bax, Ingo; Memisevic, Roland

Computer Science > Computer Vision and Pattern Recognition

arXiv:1706.04261 (cs)

[Submitted on 13 Jun 2017 (v1), last revised 15 Jun 2017 (this version, v2)]

Title:The "something something" video database for learning and evaluating visual common sense

Authors:Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzyńska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, Roland Memisevic

View PDF

Abstract:Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world. Videos, unlike still images, contain a wealth of detailed information about the physical world. However, most labelled video datasets represent high-level concepts rather than detailed physical aspects about actions and scenes. In this work, we describe our ongoing collection of the "something-something" database of video prediction tasks whose solutions require a common sense understanding of the depicted situation. The database currently contains more than 100,000 videos across 174 classes, which are defined as caption-templates. We also describe the challenges in crowd-sourcing this data at scale.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1706.04261 [cs.CV]
	(or arXiv:1706.04261v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1706.04261

Submission history

From: Raghav Goyal [view email]
[v1] Tue, 13 Jun 2017 21:26:19 UTC (15,311 KB)
[v2] Thu, 15 Jun 2017 21:15:13 UTC (4,726 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:The "something something" video database for learning and evaluating visual common sense

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:The "something something" video database for learning and evaluating visual common sense

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators