SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection

Zhu, Yi; Koppisetti, Surya; Tran, Trang; Bharaj, Gaurav

Computer Science > Sound

arXiv:2407.18517 (cs)

[Submitted on 26 Jul 2024]

Title:SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection

Authors:Yi Zhu, Surya Koppisetti, Trang Tran, Gaurav Bharaj

View PDF HTML (experimental)

Abstract:Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain data. Moreover, the black-box nature of existing models limits their use in real-world scenarios, where explanations are required for model decisions. To alleviate these issues, we introduce a new ADD model that explicitly uses the StyleLInguistics Mismatch (SLIM) in fake speech to separate them from real speech. SLIM first employs self-supervised pretraining on only real samples to learn the style-linguistics dependency in the real class. The learned features are then used in complement with standard pretrained acoustic features (e.g., Wav2vec) to learn a classifier on the real and fake classes. When the feature encoders are frozen, SLIM outperforms benchmark methods on out-of-domain datasets while achieving competitive results on in-domain data. The features learned by SLIM allow us to quantify the (mis)match between style and linguistic content in a sample, hence facilitating an explanation of the model decision.

Subjects:	Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2407.18517 [cs.SD]
	(or arXiv:2407.18517v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2407.18517

Submission history

From: Gaurav Bharaj [view email]
[v1] Fri, 26 Jul 2024 05:23:41 UTC (4,775 KB)

Computer Science > Sound

Title:SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators