Efficient Personalized Speech Enhancement through Self-Supervised Learning

Sivaraman, Aswin; Kim, Minje

doi:10.1109/JSTSP.2022.3181782

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2104.02017 (eess)

[Submitted on 5 Apr 2021 (v1), last revised 27 Jul 2022 (this version, v2)]

Title:Efficient Personalized Speech Enhancement through Self-Supervised Learning

Authors:Aswin Sivaraman, Minje Kim

View PDF

Abstract:This work presents self-supervised learning methods for developing monaural speaker-specific (i.e., personalized) speech enhancement models. While generalist models must broadly address many speakers, specialist models can adapt their enhancement function towards a particular speaker's voice, expecting to solve a narrower problem. Hence, specialists are capable of achieving more optimal performance in addition to reducing computational complexity. However, naive personalization methods can require clean speech from the target user, which is inconvenient to acquire, e.g., due to subpar recording conditions. To this end, we pose personalization as either a zero-shot task, in which no additional clean speech of the target speaker is used for training, or a few-shot learning task, in which the goal is to minimize the duration of the clean speech used for transfer learning. With this paper, we propose self-supervised learning methods as a solution to both zero- and few-shot personalization tasks. The proposed methods are designed to learn the personalized speech features from unlabeled data (i.e., in-the-wild noisy recordings from the target user) without knowing the corresponding clean sources. Our experiments investigate three different self-supervised learning mechanisms. The results show that self-supervised models achieve zero-shot and few-shot personalization using fewer model parameters and less clean data from the target user, achieving the data efficiency and model compression goals.

Comments:	15 pages, 9 figures, published in IEEE JSTSP 2022
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
Cite as:	arXiv:2104.02017 [eess.AS]
	(or arXiv:2104.02017v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2104.02017
Related DOI:	https://doi.org/10.1109/JSTSP.2022.3181782

Submission history

From: Aswin Sivaraman [view email]
[v1] Mon, 5 Apr 2021 17:12:51 UTC (4,505 KB)
[v2] Wed, 27 Jul 2022 04:48:32 UTC (9,970 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Efficient Personalized Speech Enhancement through Self-Supervised Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Efficient Personalized Speech Enhancement through Self-Supervised Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators