MELT: 수정 빈도-희소성 균형 네트워크를 이용한 합성 이미지 검색 성능 향상
MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network
합성 이미지 검색(CIR)은 참조 이미지와 수정 텍스트를 쿼리로 사용하여 "텍스트 지침에 따라 참조 이미지를 수정하는" 대상 이미지를 검색하는 기술입니다. 그러나 기존 CIR 방법은 다음과 같은 두 가지 한계점을 가지고 있습니다. (1) 빈도 편향으로 인한 "희소 샘플 무시", (2) 유사도 점수가 어려운 부정 샘플과 노이즈의 영향을 쉽게 받는다는 점입니다. 이러한 한계점을 해결하기 위해, 우리는 비대칭적인 희귀 의미 위치 파악 및 어려운 부정 샘플 하에서의 강력한 유사도 추정이라는 두 가지 핵심 과제를 해결하고자 합니다. 이러한 과제를 해결하기 위해, 수정 빈도-희소성 균형 네트워크 MELT를 제안합니다. MELT는 다중 모달 컨텍스트에서 희귀한 수정 의미에 더 많은 주의를 기울이는 동시에, 높은 유사도 점수를 가진 어려운 부정 샘플에 대해 확산 기반의 노이즈 제거를 적용하여 다중 모달 융합 및 매칭을 향상시킵니다. 두 가지 CIR 벤치마크에 대한 광범위한 실험을 통해 MELT의 우수한 성능이 입증되었습니다. 코드는 다음 주소에서 확인할 수 있습니다: https://github.com/luckylittlezhi/MELT.
Composed Image Retrieval (CIR) uses a reference image and a modification text as a query to retrieve a target image satisfying the requirement of ``modifying the reference image according to the text instructions''. However, existing CIR methods face two limitations: (1) frequency bias leading to ``Rare Sample Neglect'', and (2) susceptibility of similarity scores to interference from hard negative samples and noise. To address these limitations, we confront two key challenges: asymmetric rare semantic localization and robust similarity estimation under hard negative samples. To solve these challenges, we propose the Modification frEquentation-rarity baLance neTwork MELT. MELT assigns increased attention to rare modification semantics in multimodal contexts while applying diffusion-based denoising to hard negative samples with high similarity scores, enhancing multimodal fusion and matching. Extensive experiments on two CIR benchmarks validate the superior performance of MELT. Codes are available at https://github.com/luckylittlezhi/MELT.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.