2606.20280v1 Jun 18, 2026 cs.IR

ELVA: 순위 기반의 범용 다중 모드 검색 연구

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval

Pei Fu
Pei Fu
Citations: 28
h-index: 4
Zhenbo Luo
Zhenbo Luo
Citations: 213
h-index: 7
Jian Luan
Jian Luan
Citations: 220
h-index: 7
Yuhang Liu
Yuhang Liu
Citations: 0
h-index: 0
Hang Li
Hang Li
Citations: 75
h-index: 3
Yukun Qi
Yukun Qi
Citations: 64
h-index: 4
Chao Jiang
Chao Jiang
Citations: 31
h-index: 2
Jingwen Fu
Jingwen Fu
Citations: 58
h-index: 5
Zhen Liu
Zhen Liu
Citations: 0
h-index: 0
Bin Qin
Bin Qin
Citations: 174
h-index: 3
Jingmin Xin
Jingmin Xin
Citations: 0
h-index: 0

다중 모드 대규모 언어 모델(MLLM)을 대비 학습을 통해 활용하는 것은 범용 다중 모드 검색(UMR) 성능 향상을 위한 주류 패러다임이 되었습니다. 그러나 기존 연구에서는 대비 학습 패러다임을 검색 작업에 적용할 때 발생하는 '세부 정보 무시' 문제를 간과했습니다. '세부 정보 무시'는 모델이 쿼리에 포함된 세부 수준의 정보를 제대로 인식하지 못하는 경향을 의미하며, 복잡한 쿼리를 효과적으로 처리하는 데 매우 중요합니다. 이는 대비 학습이 샘플을 이진 분류(긍정/부정)로 취급하면서 각 부정 샘플이 가진 다양한 정보를 무시하기 때문에 발생합니다. 이러한 문제를 해결하기 위해, 우리는 부정 샘플과 긍정 샘플의 유사성에 따라 부정 샘플을 다르게 처리해야 한다고 주장하며, 이를 통해 모델이 각 부정 샘플로부터 서로 다른 세부 정보를 학습할 수 있도록 합니다. 본 논문에서는 'ELVA'라는 간단하면서도 효과적인 프레임워크를 소개합니다. ELVA는 순위 기반의 MLLM을 활용하여 세부 정보 무시 문제를 완화하는 새로운 규칙 기반 강화 학습 프레임워크입니다. 1) 우리는 보상 모델에 의존하지 않고, 검증 가능한 보상을 갖춘 강화 학습(RLVR)을 검색 작업에 확장하여 모델이 명시적인 순위 레이블 없이도 새로운 순위 동작을 탐색할 수 있도록 합니다. 2) 규칙 기반 보상을 활용함으로써, 우리의 접근 방식은 긍정 샘플과 부정 샘플 간의 유사성 차이를 확대하면서 동시에 부정 샘플의 순위를 최적화합니다. 또한, 세부 정보 무시를 보다 정확하게 측정하기 위해, 다중 세부 수준 쿼리 시나리오에 특화된 새로운 벤치마크인 MRBench를 추가적으로 소개합니다. ELVA는 표준 검색 벤치마크에서 최고 성능을 달성했으며, MRBench에서의 13.1%라는 상당한 개선은 세부 정보 무시 문제를 완화하는 데 있어 그 효과성을 더욱 입증합니다.

Original Abstract

Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance of Universal Multimodal Retrieval (UMR). However, previous works have ignored the grain blindness when adapting the contrastive paradigm into retrieval tasks. Grain blindness refers to the tendency of the model to overlook grain-level information contained in the query, which is crucial for effectively handling complex queries. This stems from contrastive learning treating samples as a binary classification (positive/negative), while ignoring the different information carried by each negative sample. To address this, we argue that negatives should be treated differently according to their similarity to the positive sample, enabling the model to learn distinct grain information from each negative. In this paper, we introduce a simple but effective framework, called ELVA, a novel rule-based RL framework that mitigates grain blindness through ranking-driven MLLMs. 1) Instead of relying on reward models, we extend Reinforcement Learning with Verifiable Rewards (RLVR) to retrieval tasks, allowing the model to explore new ranking behaviors without explicit ranking labels. 2) By utilizing rule-based rewards, our approach jointly optimizes the ranking of negative samples while enlarging the similarity gap between positive and negative. To more precisely measure grain blindness, we further introduce MRBench, a new benchmark specifically designed for multi-grain query scenarios. ELVA achieves state-of-the-art results across standard retrieval benchmarks, and its notable 13.1% improvement on MRBench further demonstrates its effectiveness in alleviating grain blindness.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!