CAER: 다중 모드 대규모 언어 모델을 위한 충돌 인식 증거 라우팅 - 이중 프리픽스 전문가 활용
CAER: Conflict-Aware Evidence Routing with Dual Prefix Experts for Multimodal Large Language Models
다중 모드 대규모 언어 모델(MLLM)은 다중 모드 이해 및 생성 분야에서 뛰어난 성능을 보여주었습니다. 하지만 텍스트 입력이 시각적 증거와 충돌할 때, 여전히 환각 현상을 일으키고 시각적 내용과 일치하지 않는 응답을 생성하는 문제가 있습니다. 기존 연구들은 주로 디코딩 전략, 추가적인 학습, 검증 방법 또는 프롬프트 기술에 의존하지만, 미세한 수준의 충돌 위치 파악 및 충돌 상황을 고려한 생성 능력이 부족합니다. 본 논문에서는 시각-언어 간 충돌 탐지 및 충돌 상황을 고려한 생성을 위한 모듈 아키텍처에 독립적인 프레임워크인 CAER를 제안합니다. CAER는 청크 단위의 증거 라우터를 도입하여, 입력 문장을 소프트 텍스트 쿼리로 변환하고, 고정된 시각적 토큰에서 해당 증거를 검색함으로써 미세한 수준의 충돌 추론을 가능하게 합니다. 또한, CAER는 시각적으로 뒷받침되는 입력과 반박되는 입력을 위해 별도의 전문가를 학습하는 이중 프리픽스 전문가 라우팅 메커니즘을 설계하여, 명시적인 전문가 선택을 통해 충돌 상황을 고려한 생성을 지원합니다. 공개된 MMMC 벤치마크 및 저희가 새로 구축한 AgriConflict 데이터셋에 대한 실험 결과는 CAER가 시각-언어 간의 충돌을 효과적으로 탐지하고, 오픈 소스 MLLM의 신뢰도를 향상시키는 데 기여한다는 것을 보여줍니다. 이러한 개선은 모델의 핵심 파라미터를 업데이트하지 않고도 가능합니다.
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in multimodal understanding and generation. However, when textual inputs conflict with visual evidence, they still suffer from hallucinations and produce responses inconsistent with visual content. Existing approaches mainly rely on decoding strategies, additional training, verification methods, or prompting techniques, but often lack fine-grained conflict localization and conflict-aware generation. In this work, we propose CAER, a backbone-agnostic framework for visual-language conflict detection and conflict-aware generation. CAER introduces a span-grounded evidence router that transforms claim representations into soft textual queries and retrieves corresponding evidence from frozen visual tokens, enabling fine-grained conflict estimation. Furthermore, we design a dual-prefix expert routing mechanism that learns separate experts for visually supported and contradicted inputs, enabling conflict-aware generation through explicit expert selection. Experiments on the public MMMC benchmark and our newly curated AgriConflict dataset demonstrate that CAER effectively detects visual-language conflicts and improves the reliability of open-source MLLMs without updating their backbone parameters.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.