2607.00374v1 Jul 01, 2026 cs.CV

학습을 통한 이미지 합성: 제로샷 복합 이미지 검색을 위한 프록시 태스크 설계 재검토

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

Zheren Fu
Zheren Fu
University of Science and Technology of China
Citations: 5,122
h-index: 6
Zhendong Mao
Zhendong Mao
Citations: 11
h-index: 2
Jingjing Zhang
Jingjing Zhang
Citations: 19
h-index: 1
Lei Zhang
Lei Zhang
Citations: 188
h-index: 8

복합 이미지 검색 (CIR)은 참조 이미지와 텍스트 수정 사항을 기반으로 대상 이미지를 검색하는 기술입니다. 지도 학습 기반의 CIR은 비용이 많이 드는 트리플렛 데이터에 의존하지만, 제로샷 CIR (ZS-CIR)은 이미지-텍스트 쌍으로 학습된 프록시 태스크를 통해 이러한 의존성을 완화합니다. 그러나 기존의 프록시 태스크는 주로 시각적 및 텍스트 표현을 향상시켜 미리 정의된 합성 메커니즘(예: 고정된 텍스트 인코더에 가짜 단어 주입 또는 선형 특징 연산)을 활용하는 데 중점을 둡니다. 그 결과, 합성 함수 자체가 학습되지 않아 모델이 다양한 수준의 의미적 수정을 표현하는 능력이 제한됩니다. 이러한 문제를 해결하기 위해, 우리는 수정과 관련된 시각적 콘텐츠에 집중하고, 이어서 대상의 의미를 완성하는 두 단계로 구성된 합성을 모델링하는 FoCo를 제안합니다. 이를 위해 텍스트 기반 시각적 집계 (text-anchored visual aggregation)라는 프록시 태스크를 통해 지역적인 텍스트 의미에 따라 선택적으로 시각적 콘텐츠를 수집하고, 컨텍스트 조건부 의미 완성 (context-conditioned semantic completion)이라는 다른 프록시 태스크를 통해 수집된 시각 정보를 나머지 장면 맥락과 결합하여 일관성 있는 복합 표현을 생성합니다. 이러한 태스크는 인스턴스 간 대비 학습 목표와 함께 훈련되어 의미적 다양성을 장려하고, 단순한 합성 전략을 방지합니다. 네 가지 ZS-CIR 벤치마크에 대한 광범위한 실험 결과, FoCo가 최첨단 성능과 향상된 일반화 능력을 보여주었습니다.

Original Abstract

Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) alleviates this reliance through proxy tasks trained on image-text pairs. However, existing proxy tasks primarily enhance visual and textual representations to accommodate a predefined composition mechanism such as pseudo-word injection into a frozen text encoder or linear feature arithmetic. As a result, the composition function itself remains unlearned, limiting the model's ability to express diverse and fine-grained semantic modifications. To address this, we propose FoCo, which models composition as two coordinated stages: focusing on modification-relevant visual content, and then completing the target semantics. We realize these through two proxy tasks: text-anchored visual aggregation to selectively gather visual content guided by localized textual semantics, and context-conditioned semantic completion to transform these aggregated visuals with the remaining scene context into a coherent composed representation. The tasks are trained jointly with a cross-instance contrastive objective, encouraging semantic diversity and discouraging shortcut composition strategies. Extensive experiments on four ZS-CIR benchmarks show FoCo's state-of-the-art performance and improved generalization.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!