ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings

Brannon, William; Kang, Wonjune; Fulay, Suyash; Jiang, Hang; Roy, Brandon; Roy, Deb; Kabbara, Jad

Computer Science > Computation and Language

arXiv:2305.14321 (cs)

[Submitted on 23 May 2023 (v1), last revised 9 Jul 2024 (this version, v2)]

Title:ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings

Authors:William Brannon, Wonjune Kang, Suyash Fulay, Hang Jiang, Brandon Roy, Deb Roy, Jad Kabbara

View PDF HTML (experimental)

Abstract:Learning on text-attributed graphs (TAGs), in which nodes are associated with one or more texts, has been the subject of much recent work. However, most approaches tend to make strong assumptions about the downstream task of interest, are reliant on hand-labeled data, or fail to equally balance the importance of both text and graph representations. In this work, we propose Contrastive Graph-Text pretraining (ConGraT), a general, self-supervised approach for jointly learning separate representations of texts and nodes in a TAG. Our method trains a language model (LM) and a graph neural network (GNN) to align their representations in a common latent space using a batch-wise contrastive learning objective inspired by CLIP. We further propose an extension to the CLIP objective that leverages graph structure to incorporate information about inter-node similarity. Extensive experiments demonstrate that ConGraT outperforms baselines on various downstream tasks, including node and text category classification, link prediction, and language modeling. Finally, we present an application of our method to community detection in social graphs, which enables finding more textually grounded communities, rather than purely graph-based ones. Code and certain datasets are available at this https URL.

Comments:	New visualizations, added references, and an application to community detection. To appear at the TextGraphs workshop @ ACL 2024. 21 pages, 5 figures, 13 tables
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2305.14321 [cs.CL]
	(or arXiv:2305.14321v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2305.14321

Submission history

From: William Brannon [view email]
[v1] Tue, 23 May 2023 17:53:30 UTC (953 KB)
[v2] Tue, 9 Jul 2024 23:44:46 UTC (1,051 KB)

Computer Science > Computation and Language

Title:ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators