Counterfactual Memorization in Neural Language Models

Zhang, Chiyuan; Ippolito, Daphne; Lee, Katherine; Jagielski, Matthew; Tramèr, Florian; Carlini, Nicholas

Computer Science > Computation and Language

arXiv:2112.12938 (cs)

[Submitted on 24 Dec 2021 (v1), last revised 13 Oct 2023 (this version, v2)]

Title:Counterfactual Memorization in Neural Language Models

Authors:Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, Nicholas Carlini

View PDF

Abstract:Modern neural language models that are widely used in various NLP tasks risk memorizing sensitive information from their training data. Understanding this memorization is important in real world applications and also from a learning-theoretical perspective. An open question in previous studies of language model memorization is how to filter out "common" memorization. In fact, most memorization criteria strongly correlate with the number of occurrences in the training set, capturing memorized familiar phrases, public knowledge, templated texts, or other repeated data. We formulate a notion of counterfactual memorization which characterizes how a model's predictions change if a particular document is omitted during training. We identify and study counterfactually-memorized training examples in standard text datasets. We estimate the influence of each memorized training example on the validation set and on generated texts, showing how this can provide direct evidence of the source of memorization at test time.

Comments:	NeurIPS 2023; 42 pages, 33 figures
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2112.12938 [cs.CL]
	(or arXiv:2112.12938v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2112.12938

Submission history

From: Chiyuan Zhang [view email]
[v1] Fri, 24 Dec 2021 04:20:57 UTC (3,875 KB)
[v2] Fri, 13 Oct 2023 22:16:41 UTC (6,374 KB)

Computer Science > Computation and Language

Title:Counterfactual Memorization in Neural Language Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Counterfactual Memorization in Neural Language Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators