CTM -- A Model for Large-Scale Multi-View Tweet Topic Classification

Kulkarni, Vivek; Leung, Kenny; Haghighi, Aria

Computer Science > Computation and Language

arXiv:2205.01603 (cs)

[Submitted on 3 May 2022]

Title:CTM -- A Model for Large-Scale Multi-View Tweet Topic Classification

Authors:Vivek Kulkarni, Kenny Leung, Aria Haghighi

View PDF

Abstract:Automatically associating social media posts with topics is an important prerequisite for effective search and recommendation on many social media platforms. However, topic classification of such posts is quite challenging because of (a) a large topic space (b) short text with weak topical cues, and (c) multiple topic associations per post. In contrast to most prior work which only focuses on post classification into a small number of topics ($10$-$20$), we consider the task of large-scale topic classification in the context of Twitter where the topic space is $10$ times larger with potentially multiple topic associations per Tweet. We address the challenges above by proposing a novel neural model, CTM that (a) supports a large topic space of $300$ topics and (b) takes a holistic approach to tweet content modeling -- leveraging multi-modal content, author context, and deeper semantic cues in the Tweet. Our method offers an effective way to classify Tweets into topics at scale by yielding superior performance to other approaches (a relative lift of $\mathbf{20}\%$ in median average precision score) and has been successfully deployed in production at Twitter.

Comments:	12 pages. 1 figure. NAACL Industry Track
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2205.01603 [cs.CL]
	(or arXiv:2205.01603v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2205.01603

Submission history

From: Vivek Kulkarni [view email]
[v1] Tue, 3 May 2022 16:32:09 UTC (250 KB)

Computer Science > Computation and Language

Title:CTM -- A Model for Large-Scale Multi-View Tweet Topic Classification

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:CTM -- A Model for Large-Scale Multi-View Tweet Topic Classification

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators