End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator

Tjandra, Andros; Sakti, Sakriani; Nakamura, Satoshi

Computer Science > Computation and Language

arXiv:1810.13107 (cs)

[Submitted on 31 Oct 2018]

Title:End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator

Authors:Andros Tjandra, Sakriani Sakti, Satoshi Nakamura

View PDF

Abstract:The speech chain mechanism integrates automatic speech recognition (ASR) and text-to-speech synthesis (TTS) modules into a single cycle during training. In our previous work, we applied a speech chain mechanism as a semi-supervised learning. It provides the ability for ASR and TTS to assist each other when they receive unpaired data and let them infer the missing pair and optimize the model with reconstruction loss. If we only have speech without transcription, ASR generates the most likely transcription from the speech data, and then TTS uses the generated transcription to reconstruct the original speech features. However, in previous papers, we just limited our back-propagation to the closest module, which is the TTS part. One reason is that back-propagating the error through the ASR is challenging due to the output of the ASR are discrete tokens, creating non-differentiability between the TTS and ASR. In this paper, we address this problem and describe how to thoroughly train a speech chain end-to-end for reconstruction loss using a straight-through estimator (ST). Experimental results revealed that, with sampling from ST-Gumbel-Softmax, we were able to update ASR parameters and improve the ASR performances by 11\% relative CER reduction compared to the baseline.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:1810.13107 [cs.CL]
	(or arXiv:1810.13107v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1810.13107

Submission history

From: Andros Tjandra [view email]
[v1] Wed, 31 Oct 2018 05:05:37 UTC (1,462 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2018-10

Change to browse by:

cs
cs.LG
cs.SD
eess
eess.AS

References & Citations

DBLP - CS Bibliography

listing | bibtex

Andros Tjandra
Sakriani Sakti
Satoshi Nakamura

export BibTeX citation

Computer Science > Computation and Language

Title:End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators