2607.13712v1 Jul 15, 2026 cs.CV

Groc-PO: 사실 기반 다중 모드 LLM을 위한 맥락 선호도 최적화

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

Zhi Zheng
Zhi Zheng
Citations: 1,995
h-index: 6
Zhiyuan Yao
Zhiyuan Yao
Citations: 367
h-index: 4
Zheren Fu
Zheren Fu
University of Science and Technology of China
Citations: 5,122
h-index: 6
Zhendong Mao
Zhendong Mao
Citations: 11
h-index: 2
Chunxiao Liu
Chunxiao Liu
Citations: 702
h-index: 6
Dongming Zhang
Dongming Zhang
Citations: 0
h-index: 0

다중 모드 대규모 언어 모델(MLLM)의 빠른 발전에도 불구하고, 시각적 환각, 내용 조작, 부정확한 추론과 같은 진실성 문제로 인해 신뢰성과 실용성이 크게 저하되는 경향이 있습니다. 인간 선호도를 기반으로 하는 정렬 방법인 직접 선호도 최적화(DPO)는 이러한 문제를 해결하기 위해 널리 사용됩니다. 그러나 다중 모드 추론 오류는 종종 여러 단계에 걸쳐 전파되며, 최종 답변 오류는 초기 단계의 잘못에서 비롯되는 경우가 많지만, 일반적인 DPO는 일반적으로 최종 답변 수준에서 선호도 최적화를 적용합니다. 이러한 책임 할당 문제는 초기 단계의 감독이 특정 단계에 대한 직접적인 지시가 아닌 간접적인 형태를 띠게 하여, 접지(grounding) 오류로 인한 오차 전파를 억제하기 어렵게 만듭니다. 이를 해결하기 위해, 우리는 MLLM을 위한 접지 선호도 최적화 프레임워크인 Grounded Context Preference Optimization (Groc-PO)을 제안합니다. 또한, 객체 접지(Object Grounding), 맥락 접지(Contextual Grounding), 접지 추론(Grounded Reasoning)의 세 가지 단계를 중심으로 다단계 선호도 샘플을 구성한 Grounded Context Preference Dataset (GCPD)를 구축하여, 접지된 맥락의 형성, 통합 및 활용을 포착합니다. Groc-PO는 여러 단계에 걸쳐 더 명시적인 선호도 감독을 도입함으로써, 맥락 의존적 추론을 강화하고 단계 간 오류 전파를 완화합니다. 광범위한 실험 결과, 표준 DPO 및 다른 강력한 기준 모델과 비교했을 때 Groc-PO는 환각 감소, 신뢰할 수 있는 추론 및 전체적인 안정성 측면에서 향상된 성능을 달성했으며, 이는 신뢰할 수 있는 다중 모드 추론을 위한 더욱 명시적인 접지 감독의 가치를 뒷받침합니다.

Original Abstract

Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and practical utility. Alignment methods based on human preference, such as Direct Preference Optimization (DPO), have been widely adopted to address these issues. However, multimodal reasoning errors often propagate across stages, and final-answer errors can often be traced to mistakes in early grounding stages, yet standard DPO typically applies preference optimization at the final-answer level. This credit-assignment challenge means that supervision for early grounding stages is indirect rather than stage-specific, making it difficult to suppress error propagation arising from grounding drift and context inconsistency. To address this, we propose Grounded Context Preference Optimization (Groc-PO), a grounded preference optimization framework for MLLMs. We further construct the Grounded Context Preference Dataset (GCPD), organizing multi-stage preference samples around three stages of Object Grounding, Contextual Grounding, and Grounded Reasoning, to capture the formation, integration, and utilization of grounded context. By introducing more explicit preference supervision over multiple grounded stages, Groc-PO strengthens context-dependent reasoning and mitigates cross-stage error propagation. Extensive experiments show that, compared with standard DPO and other strong baselines, Groc-PO achieves improved performance in hallucination mitigation, faithful reasoning, and overall reliability, supporting the value of more explicit grounded supervision for trustworthy multimodal reasoning.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!