Using Sparse Semantic Embeddings Learned from Multimodal Text and Image Data to Model Human Conceptual Knowledge

Derby, Steven; Miller, Paul; Murphy, Brian; Devereux, Barry

Computer Science > Computation and Language

arXiv:1809.02534 (cs)

[Submitted on 7 Sep 2018 (v1), last revised 14 Nov 2018 (this version, v3)]

Title:Using Sparse Semantic Embeddings Learned from Multimodal Text and Image Data to Model Human Conceptual Knowledge

Authors:Steven Derby, Paul Miller, Brian Murphy, Barry Devereux

View PDF

Abstract:Distributional models provide a convenient way to model semantics using dense embedding spaces derived from unsupervised learning algorithms. However, the dimensions of dense embedding spaces are not designed to resemble human semantic knowledge. Moreover, embeddings are often built from a single source of information (typically text data), even though neurocognitive research suggests that semantics is deeply linked to both language and perception. In this paper, we combine multimodal information from both text and image-based representations derived from state-of-the-art distributional models to produce sparse, interpretable vectors using Joint Non-Negative Sparse Embedding. Through in-depth analyses comparing these sparse models to human-derived behavioural and neuroimaging data, we demonstrate their ability to predict interpretable linguistic descriptions of human ground-truth semantic knowledge.

Comments:	Proceedings of the 22nd Conference on Computational Natural Language Learning (CoNLL 2018), pages 260-270. Brussels, Belgium, October 31 - November 1, 2018. Association for Computational Linguistics
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1809.02534 [cs.CL]
	(or arXiv:1809.02534v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1809.02534
Journal reference:	Proceedings of the 22nd Conference on Computational Natural Language Learning (CoNLL 2018), pages 260-270. Brussels, Belgium, October 31 - November 1, 2018. Association for Computational Linguistics

Submission history

From: Barry Devereux [view email]
[v1] Fri, 7 Sep 2018 15:22:04 UTC (1,000 KB)
[v2] Tue, 18 Sep 2018 13:43:50 UTC (1,164 KB)
[v3] Wed, 14 Nov 2018 15:25:30 UTC (1,164 KB)

Computer Science > Computation and Language

Title:Using Sparse Semantic Embeddings Learned from Multimodal Text and Image Data to Model Human Conceptual Knowledge

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Using Sparse Semantic Embeddings Learned from Multimodal Text and Image Data to Model Human Conceptual Knowledge

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators