FiRE: 미세 입력을 활용하여 멀티모달 대규모 언어 모델(MLLM)의 성능을 향상시키는 방법론 - 복잡한 이미지 검색
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
멀티모달 대규모 언어 모델(MLLM)은 강력한 일반화 능력과 추론 능력을 바탕으로 다양한 실제 이미지 검색 작업에 효과적으로 활용될 수 있는 잠재력을 보여줍니다. 그러나 기존 연구들은 MLLM의 검색 성능을 향상시키는 데 있어 미세 입력을 활용한 모델링 및 분리된 학습 목표의 가능성을 간과하고 있습니다. 특히, 긴 텍스트-이미지 검색, 시각적 대화 검색, 그리고 복합 이미지 검색(CIR)과 같은 복잡한 작업에 대한 개선이 필요합니다. 이에 본 연구에서는 자동화된 미세 입력 멀티모달 5튜플 데이터셋 구축 파이프라인과 새로운 두 단계의 미세 입력 멀티모달 학습 전략을 제안합니다. 데이터셋 생성 파이프라인은 세밀하게 작성된 이미지 설명 및 수정 텍스트를 포함하는 포괄적인 CIR 데이터셋을 생성하여 미세 입력 모델링을 용이하게 합니다. 기존의 통합된 학습 방식과 달리, 저희는 학습 과정을 두 가지 단계로 분리했습니다: (1) 미세 입력을 활용한 문맥 추론 중심 학습, 그리고 (2) 검색 성능 향상을 위한 미세 입력 중심 학습. 이러한 단계들은 모델의 문맥 이해 능력과 쿼리와 대상 간의 정렬 능력을 순차적으로 개선하여 검색 성능을 향상시키는 것을 목표로 합니다. 다섯 가지 데이터셋에 대한 광범위한 실험 결과는 저희 방법이 기존 방식보다 뛰어난 성능을 보이며, 특히 더 가벼운 MLLM 기반 모델에서도 우수한 성능을 유지함을 보여줍니다.
Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR). Therefore, in this work, we propose an automated fine-grained multimodal quintuple dataset construction pipeline and a novel two-stage fine-grained multimodal fine-tuning strategy. The dataset generation pipeline produces a comprehensive CIR dataset with fine-grained image captions and modification text, facilitating fine-grained context modeling. Beyond the previously entangled fine-tuning paradigm, our approach separates the fine-tuning process into two distinct stages: (1) fine-grained context reasoning-oriented fine-tuning and (2) fine-grained retrieval-oriented fine-tuning. These stages aim to sequentially enhance the model's context understanding and query-target alignment capabilities, thereby improving retrieval performance. Extensive experiments across five datasets encompassing diverse and complex image retrieval tasks demonstrate the remarkable superiority of our method over existing approaches in zero-shot retrieval settings, even with a more lightweight MLLM backbone compared to those methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.