Improving noise robustness of automatic speech recognition via parallel data and teacher-student learning

Mošner, Ladislav; Wu, Minhua; Raju, Anirudh; Parthasarathi, Sree Hari Krishnan; Kumatani, Kenichi; Sundaram, Shiva; Maas, Roland; Hoffmeister, Björn

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:1901.02348 (eess)

[Submitted on 5 Jan 2019 (v1), last revised 15 Mar 2019 (this version, v3)]

Title:Improving noise robustness of automatic speech recognition via parallel data and teacher-student learning

Authors:Ladislav Mošner, Minhua Wu, Anirudh Raju, Sree Hari Krishnan Parthasarathi, Kenichi Kumatani, Shiva Sundaram, Roland Maas, Björn Hoffmeister

View PDF

Abstract:For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we apply a logits selection method which only preserves the k highest values to prevent wrong emphasis of knowledge from the teacher and to reduce bandwidth needed for transferring data. We incorporate up to 8000 hours of untranscribed data for training and present our results on sequence trained models apart from cross entropy trained ones. The best sequence trained student model yields relative word error rate (WER) reductions of approximately 10.1%, 28.7% and 19.6% on our clean, simulated noisy and real test sets respectively comparing to a sequence trained teacher.

Comments:	To Appear in ICASSP 2019
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:1901.02348 [eess.AS]
	(or arXiv:1901.02348v3 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.1901.02348

Submission history

From: Minhua Wu [view email]
[v1] Sat, 5 Jan 2019 06:22:40 UTC (60 KB)
[v2] Fri, 11 Jan 2019 06:15:47 UTC (60 KB)
[v3] Fri, 15 Mar 2019 20:16:57 UTC (60 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Improving noise robustness of automatic speech recognition via parallel data and teacher-student learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Improving noise robustness of automatic speech recognition via parallel data and teacher-student learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators