2608.05149v1 Aug 05, 2026 cs.CV

CoCo-IR: 문맥 기반 이미지 검색

CoCo-IR: Contextual Composed Image Retrieval

Zhe Li
Zhe Li
Citations: 213
h-index: 3
Kaifeng Chen
Kaifeng Chen
Citations: 201
h-index: 4
Shengcao Cao
Shengcao Cao
Citations: 1,489
h-index: 8
T. Dabral
T. Dabral
Citations: 126
h-index: 3
Z. Ding
Z. Ding
Citations: 98
h-index: 1
Madhuri Shanbhogue
Madhuri Shanbhogue
Citations: 192
h-index: 2
Mojtaba Seyedhosseini
Mojtaba Seyedhosseini
Citations: 3,550
h-index: 6
Liangyan Gui
Liangyan Gui
Citations: 2,689
h-index: 20
Yu-Xiong Wang
Yu-Xiong Wang
Citations: 74
h-index: 4

현재의 지시 기반 이미지 검색 시스템은 강력하지만, 단일 상호작용에 국한되어 있어 복잡하고 실제적인 시각 검색의 반복적인 특성을 반영하지 못합니다. 이러한 한계를 극복하기 위해, 사용자가 상호작용을 통해 점진적으로 검색 결과를 개선할 수 있는 새로운 작업인 문맥 기반 이미지 검색 (CoCo-IR)을 제안합니다. 우리는 이 새로운 작업을 해결하기 위해, CoCo-IR에 적합한 문맥 인지 추론 기능을 제공하는 대규모 다중 모드 모델(LMM) 기반의 새로운 모델을 개발했습니다. 우리의 모델은 전체 상호작용 기록을 해석하여, 여러 단계에서 진화하는 변환 가능한 이미지 임베딩(TIE)을 생성합니다. 모델 학습에 필요한 데이터셋을 구축하기 위해, 인간의 주석 없이도 고품질의 문맥 기반 검색 데이터를 생성할 수 있는 완전 자동화된 확장 가능 데이터 엔진을 개발했으며, 모델 지침 하에 어려운 부정 샘플을 추출했습니다. 광범위한 실험 결과는 우리의 접근 방식이 새로운 최고 성능을 달성함을 보여줍니다. 우리는 까다로운 단일 상호작용 벤치마크 CIRCO에서 39.4의 mAP@5를 달성했으며, 또한 우리가 새롭게 만든 CoCo-IR 벤치마크에서 모델이 4턴 대화에서 44.1의 R@1을 유지하며 기존 방법(28.2의 4턴 R@1)보다 훨씬 우수한 성능을 보였습니다. 프로젝트 페이지: https://CoCo-IR.github.io.

Original Abstract

Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual searches. To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables users to progressively refine search results through interactions. We address this new task by proposing a new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR. Our model interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns. To fuel the model training without expensive human annotations, we develop a fully autonomous, scalable data engine that leverages LMMs to generate high-quality contextual retrieval data, and uses model-guided verification to mine challenging hard negatives. Extensive experiments demonstrate that our approach establishes new state-of-the-art performance: We achieve 39.4 mAP@5 on the challenging single-turn benchmark CIRCO; furthermore, on our new CoCo-IR benchmark, our model maintains robust performance with 44.1 R@1 on 4-turn dialogues, dramatically outperforming existing methods (28.2 4-turn R@1) that fail to handle multi-turn context. Project page: https://CoCo-IR.github.io.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!