신뢰성 있는 GRPO: 제약 조건 최적화를 통한 다중 모드 언어 모델의 시각적 공간 추론 능력 향상
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
검증 가능한 보상을 사용하는 강화 학습으로 훈련된 다중 모드 추론 모델(MRM)은 시각적 추론 벤치마크에서 정확도가 향상되는 것으로 나타났습니다. 그러나 우리는 정확도 향상이 종종 추론 품질 저하를 동반한다는 것을 관찰했습니다. 즉, 생성된 연쇄적 사고(Chain-of-Thought, CoT) 과정이 최종 답변과 일관성이 부족하고 시각적 증거에 제대로 기반하지 않는 경우가 많습니다. 우리는 일곱 가지의 실제 시각적 공간 추론 벤치마크를 통해 이 현상을 체계적으로 연구했으며, ViGoRL-Spatial, TreeVGR 및 표준 그룹 상대 정책 최적화(GRPO)로 훈련된 자체 모델과 같은 현대적인 MRM에 영향을 미친다는 것을 확인했습니다. 우리는 CoT 추론 품질을 '논리적 일관성'(CoT가 최종 답변을 포함하는가?)과 '시각적 기반성'(각 추론 단계가 이미지 내의 객체, 속성 및 공간적 관계를 정확하게 설명하는가?)이라는 두 가지 상호 보완적인 축을 기준으로 분석했습니다. 이러한 문제를 해결하기 위해, 우리는 Lagrangian 이중 상승을 통해 일관성과 기반성을 제약 조건으로 적용하는 GRPO의 변형인 Faithful GRPO (FGRPO)를 제안합니다. FGRPO는 그룹 내의 이점 계산 과정에 배치 수준의 일관성 및 기반성 제약 조건을 통합하고, 최적화 과정에서 제약 조건의 상대적인 중요도를 적응적으로 조정합니다. 우리는 Qwen2.5-VL-7B 및 3B 모델을 기반으로 일곱 가지 공간 데이터 세트에 대해 FGRPO를 평가했습니다. 그 결과, FGRPO는 추론 품질을 크게 향상시켜 불일치율을 24.5%에서 1.7%로 줄이고 시각적 기반성 점수를 +13% 향상시켰습니다. 또한 FGRPO는 단순 GRPO보다 최종 답변 정확도를 향상시켜, 신뢰성 있는 추론이 더 나은 답변을 가능하게 한다는 것을 입증했습니다.
Multimodal reasoning models (MRMs) trained with reinforcement learning with verifiable rewards (RLVR) show improved accuracy on visual reasoning benchmarks. However, we observe that accuracy gains often come at the cost of reasoning quality: generated Chain-of-Thought (CoT) traces are frequently inconsistent with the final answer and poorly grounded in the visual evidence. We systematically study this phenomenon across seven challenging real-world spatial reasoning benchmarks and find that it affects contemporary MRMs such as ViGoRL-Spatial, TreeVGR as well as our own models trained with standard Group Relative Policy Optimization (GRPO). We characterize CoT reasoning quality along two complementary axes: "logical consistency" (does the CoT entail the final answer?) and "visual grounding" (does each reasoning step accurately describe objects, attributes, and spatial relationships in the image?). To address this, we propose Faithful GRPO (FGRPO), a variant of GRPO that enforces consistency and grounding as constraints via Lagrangian dual ascent. FGRPO incorporates batch-level consistency and grounding constraints into the advantage computation within a group, adaptively adjusting the relative importance of constraints during optimization. We evaluate FGRPO on Qwen2.5-VL-7B and 3B backbones across seven spatial datasets. Our results show that FGRPO substantially improves reasoning quality, reducing the inconsistency rate from 24.5% to 1.7% and improving visual grounding scores by +13%. It also improves final answer accuracy over simple GRPO, demonstrating that faithful reasoning enables better answers.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.