SEPP: Similarity Estimation of Predicted Probabilities for Defending and Detecting Adversarial Text

Nguyen-Son, Hoang-Quoc; Hidano, Seira; Fukushima, Kazuhide; Kiyomoto, Shinsaku

Computer Science > Computation and Language

arXiv:2110.05748 (cs)

[Submitted on 12 Oct 2021 (v1), last revised 13 Oct 2021 (this version, v2)]

Title:SEPP: Similarity Estimation of Predicted Probabilities for Defending and Detecting Adversarial Text

Authors:Hoang-Quoc Nguyen-Son, Seira Hidano, Kazuhide Fukushima, Shinsaku Kiyomoto

View PDF

Abstract:There are two cases describing how a classifier processes input text, namely, misclassification and correct classification. In terms of misclassified texts, a classifier handles the texts with both incorrect predictions and adversarial texts, which are generated to fool the classifier, which is called a victim. Both types are misunderstood by the victim, but they can still be recognized by other classifiers. This induces large gaps in predicted probabilities between the victim and the other classifiers. In contrast, text correctly classified by the victim is often successfully predicted by the others and induces small gaps. In this paper, we propose an ensemble model based on similarity estimation of predicted probabilities (SEPP) to exploit the large gaps in the misclassified predictions in contrast to small gaps in the correct classification. SEPP then corrects the incorrect predictions of the misclassified texts. We demonstrate the resilience of SEPP in defending and detecting adversarial texts through different types of victim classifiers, classification tasks, and adversarial attacks.

Comments:	PACLIC 35 (2021) (Oral)
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2110.05748 [cs.CL]
	(or arXiv:2110.05748v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2110.05748

Submission history

From: Hoang-Quoc Nguyen-Son [view email]
[v1] Tue, 12 Oct 2021 05:36:54 UTC (572 KB)
[v2] Wed, 13 Oct 2021 02:17:45 UTC (571 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2021-10

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Hoang-Quoc Nguyen-Son
Seira Hidano
Shinsaku Kiyomoto

export BibTeX citation

Computer Science > Computation and Language

Title:SEPP: Similarity Estimation of Predicted Probabilities for Defending and Detecting Adversarial Text

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:SEPP: Similarity Estimation of Predicted Probabilities for Defending and Detecting Adversarial Text

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators