Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

Debenedetti, Edoardo; Rando, Javier; Paleka, Daniel; Florin, Silaghi Fineas; Albastroiu, Dragos; Cohen, Niv; Lemberg, Yuval; Ghosh, Reshmi; Wen, Rui; Salem, Ahmed; Cherubin, Giovanni; Zanella-Beguelin, Santiago; Schmid, Robin; Klemm, Victor; Miki, Takahiro; Li, Chenhao; Kraft, Stefan; Fritz, Mario; Tramèr, Florian; Abdelnabi, Sahar; Schönherr, Lea

Computer Science > Cryptography and Security

arXiv:2406.07954 (cs)

[Submitted on 12 Jun 2024]

Title:Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

Abstract:Large language model systems face important security risks from maliciously crafted messages that aim to overwrite the system's original instructions or leak private data. To study this problem, we organized a capture-the-flag competition at IEEE SaTML 2024, where the flag is a secret string in the LLM system prompt. The competition was organized in two phases. In the first phase, teams developed defenses to prevent the model from leaking the secret. During the second phase, teams were challenged to extract the secrets hidden for defenses proposed by the other teams. This report summarizes the main insights from the competition. Notably, we found that all defenses were bypassed at least once, highlighting the difficulty of designing a successful defense and the necessity for additional research to protect LLM systems. To foster future research in this direction, we compiled a dataset with over 137k multi-turn attack chats and open-sourced the platform.

Subjects:	Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2406.07954 [cs.CR]
	(or arXiv:2406.07954v1 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2406.07954

Submission history

From: Javier Rando [view email]
[v1] Wed, 12 Jun 2024 07:27:28 UTC (5,092 KB)

Computer Science > Cryptography and Security

Title:Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators