Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection

Zhang, Min; He, Jianfeng; Ji, Taoran; Lu, Chang-Tien

Computer Science > Computation and Language

arXiv:2402.11406 (cs)

[Submitted on 18 Feb 2024 (v1), last revised 23 Jul 2024 (this version, v3)]

Title:Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection

Authors:Min Zhang, Jianfeng He, Taoran Ji, Chang-Tien Lu

View PDF

Abstract:The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of practice. However, the extent to which LLMs effectively address this issue remains insufficiently examined. This paper delves into the capability of LLMs to detect implicit hate speech (Classification Task) and express confidence in their responses (Calibration Task). Our evaluation meticulously considers various prompt patterns and mainstream uncertainty estimation methods. Our findings highlight that LLMs exhibit two extremes: (1) LLMs display excessive sensitivity towards groups or topics that may cause fairness issues, resulting in misclassifying benign statements as hate speech. (2) LLMs' confidence scores for each method excessively concentrate on a fixed range, remaining unchanged regardless of the dataset's complexity. Consequently, the calibration performance is heavily reliant on primary classification accuracy. These discoveries unveil new limitations of LLMs, underscoring the need for caution when optimizing models to ensure they do not veer towards extremes. This serves as a reminder to carefully consider sensitivity and confidence in the pursuit of model fairness.

Comments:	ACL 2024 Main Conference
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2402.11406 [cs.CL]
	(or arXiv:2402.11406v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2402.11406

Submission history

From: Min Zhang [view email]
[v1] Sun, 18 Feb 2024 00:04:40 UTC (2,970 KB)
[v2] Mon, 26 Feb 2024 16:43:37 UTC (3,041 KB)
[v3] Tue, 23 Jul 2024 06:20:32 UTC (4,866 KB)

Computer Science > Computation and Language

Title:Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators