Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain

Gu, Rongzhi; Zhang, Shi-Xiong; Zou, Yuexian; Yu, Dong

doi:10.1109/LSP.2021.3076374

Computer Science > Sound

arXiv:2104.12359 (cs)

[Submitted on 26 Apr 2021]

Title:Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain

Authors:Rongzhi Gu, Shi-Xiong Zhang, Yuexian Zou, Dong Yu

View PDF

Abstract:To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of the target speech in time-frequency domain under supervised deep learning framework. However, the existing deep models for estimating cRM are designed in the way that the real and imaginary parts of the cRM are separately modeled using real-valued training data pairs. The research motivation of this study is to design a deep model that fully exploits the temporal-spectral-spatial information of multi-channel signals for estimating cRM directly and efficiently in complex domain. As a result, a novel TSS network is designed consisting of two modules, a complex neural spatial filter (cNSF) and an MVDR. Essentially, cNSF is a cRM estimation model and an MVDR module is cascaded to the cNSF module to reduce the nonlinear speech distortions introduced by neural network. Specifically, to fit the cRM target, all input features of cNSF are reformulated into complex-valued representations following the supervised learning paradigm. Then, to achieve good hierarchical feature abstraction, a complex deep neural network (cDNN) is delicately designed with U-Net structure. Experiments conducted on simulated multi-channel speech data demonstrate the proposed cNSF outperforms the baseline NSF by 12.1% scale-invariant signal-to-distortion ratio and 33.1% word error rate.

Comments:	5 pages, 3 figures
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2104.12359 [cs.SD]
	(or arXiv:2104.12359v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2104.12359
Related DOI:	https://doi.org/10.1109/LSP.2021.3076374

Submission history

From: Rongzhi Gu [view email]
[v1] Mon, 26 Apr 2021 06:04:27 UTC (1,518 KB)

Computer Science > Sound

Title:Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators