Listening while Speaking and Visualizing: Improving ASR through Multimodal Chain

Effendi, Johanes; Tjandra, Andros; Sakti, Sakriani; Nakamura, Satoshi

Computer Science > Computation and Language

arXiv:1906.00579 (cs)

[Submitted on 3 Jun 2019 (v1), last revised 14 Nov 2019 (this version, v3)]

Title:Listening while Speaking and Visualizing: Improving ASR through Multimodal Chain

Authors:Johanes Effendi, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura

View PDF

Abstract:Previously, a machine speech chain, which is based on sequence-to-sequence deep learning, was proposed to mimic speech perception and production behavior. Such chains separately processed listening and speaking by automatic speech recognition (ASR) and text-to-speech synthesis (TTS) and simultaneously enabled them to teach each other in semi-supervised learning when they received unpaired data. Unfortunately, this speech chain study is limited to speech and textual modalities. In fact, natural communication is actually multimodal and involves both auditory and visual sensory systems. Although the said speech chain reduces the requirement of having a full amount of paired data, in this case we still need a large amount of unpaired data. In this research, we take a further step and construct a multimodal chain and design a closely knit chain architecture that combines ASR, TTS, image captioning, and image production models into a single framework. The framework allows the training of each component without requiring a large number of parallel multimodal data. Our experimental results also show that an ASR can be further trained without speech and text data and cross-modal data augmentation remains possible through our proposed chain, which improves the ASR performance.

Comments:	Accepted in IEEE ASRU 2019
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:1906.00579 [cs.CL]
	(or arXiv:1906.00579v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1906.00579

Submission history

From: Johanes Effendi [view email]
[v1] Mon, 3 Jun 2019 05:25:42 UTC (281 KB)
[v2] Mon, 7 Oct 2019 05:09:27 UTC (280 KB)
[v3] Thu, 14 Nov 2019 12:05:12 UTC (284 KB)

Computer Science > Computation and Language

Title:Listening while Speaking and Visualizing: Improving ASR through Multimodal Chain

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Listening while Speaking and Visualizing: Improving ASR through Multimodal Chain

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators