StructChart: Perception, Structuring, Reasoning for Visual Chart Understanding

Xia, Renqiu; Zhang, Bo; Peng, Haoyang; Ye, Hancheng; Yan, Xiangchao; Ye, Peng; Shi, Botian; Qiao, Yu; Yan, Junchi

Computer Science > Computer Vision and Pattern Recognition

arXiv:2309.11268 (cs)

[Submitted on 20 Sep 2023 (v1), last revised 19 Feb 2024 (this version, v4)]

Title:StructChart: Perception, Structuring, Reasoning for Visual Chart Understanding

Authors:Renqiu Xia, Bo Zhang, Haoyang Peng, Hancheng Ye, Xiangchao Yan, Peng Ye, Botian Shi, Yu Qiao, Junchi Yan

View PDF HTML (experimental)

Abstract:Charts are common in literature across different scientific fields, conveying rich information easily accessible to readers. Current chart-related tasks focus on either chart perception which refers to extracting information from the visual charts, or performing reasoning given the extracted data, e.g. in a tabular form. In this paper, we aim to establish a unified and label-efficient learning paradigm for joint perception and reasoning tasks, which can be generally applicable to different downstream tasks, beyond the question-answering task as specifically studied in peer works. Specifically, StructChart first reformulates the chart information from the popular tubular form (specifically linearized CSV) to the proposed Structured Triplet Representations (STR), which is more friendly for reducing the task gap between chart perception and reasoning due to the employed structured information extraction for charts. We then propose a Structuring Chart-oriented Representation Metric (SCRM) to quantitatively evaluate the performance for the chart perception task. To enrich the dataset for training, we further explore the possibility of leveraging the Large Language Model (LLM), enhancing the chart diversity in terms of both chart visual style and its statistical information. Extensive experiments are conducted on various chart-related tasks, demonstrating the effectiveness and promising potential for a unified chart perception-reasoning paradigm to push the frontier of chart understanding.

Comments:	SimChart9K is available for downloading at: this https URL 26 pages, 15 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2309.11268 [cs.CV]
	(or arXiv:2309.11268v4 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2309.11268

Submission history

From: Bo Zhang [view email]
[v1] Wed, 20 Sep 2023 12:51:13 UTC (2,180 KB)
[v2] Mon, 25 Sep 2023 06:09:36 UTC (2,279 KB)
[v3] Thu, 1 Feb 2024 12:47:15 UTC (2,482 KB)
[v4] Mon, 19 Feb 2024 03:48:55 UTC (2,482 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:StructChart: Perception, Structuring, Reasoning for Visual Chart Understanding

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:StructChart: Perception, Structuring, Reasoning for Visual Chart Understanding

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators