CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval

Jiang, Xintong; Wang, Yaxiong; Li, Mengjian; Wu, Yujiao; Hu, Bingwen; Qian, Xueming

doi:10.1145/3626772.3657823

Computer Science > Computer Vision and Pattern Recognition

arXiv:2405.19149 (cs)

[Submitted on 29 May 2024 (v1), last revised 30 May 2024 (this version, v2)]

Title:CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval

Authors:Xintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu, Bingwen Hu, Xueming Qian

View PDF HTML (experimental)

Abstract:Composed Image Retrieval (CIR) involves searching for target images based on an image-text pair query. While current methods treat this as a query-target matching problem, we argue that CIR triplets contain additional associations beyond this primary relation. In our paper, we identify two new relations within triplets, treating each triplet as a graph node. Firstly, we introduce the concept of text-bridged image alignment, where the query text serves as a bridge between the query image and the target image. We propose a hinge-based cross-attention mechanism to incorporate this relation into network learning. Secondly, we explore complementary text reasoning, considering CIR as a form of cross-modal retrieval where two images compose to reason about complementary text. To integrate these perspectives effectively, we design a twin attention-based compositor. By combining these complementary associations with the explicit query pair-target image relation, we establish a comprehensive set of constraints for CIR. Our framework, CaLa (Complementary Association Learning for Augmenting Composed Image Retrieval), leverages these insights. We evaluate CaLa on CIRR and FashionIQ benchmarks with multiple backbones, demonstrating its superiority in composed image retrieval.

Comments:	To appear at SIGIR 2024. arXiv admin note: text overlap with arXiv:2309.02169
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
Cite as:	arXiv:2405.19149 [cs.CV]
	(or arXiv:2405.19149v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2405.19149
Related DOI:	https://doi.org/10.1145/3626772.3657823

Submission history

From: Xintong Jiang [view email]
[v1] Wed, 29 May 2024 14:52:10 UTC (2,748 KB)
[v2] Thu, 30 May 2024 13:26:43 UTC (2,748 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators