2608.04698v1 Aug 05, 2026 cs.CV

MLLM에게 '안 돼'라고 말하게 하기: GRPO 기반 거부 교정을 통한 일반화된 지칭 표현 이해

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

Tao Huang
Tao Huang
Citations: 186
h-index: 4
Caiyan Qin
Caiyan Qin
Citations: 100
h-index: 3
Jun Ling
Jun Ling
Citations: 3
h-index: 1
Peng Wang
Peng Wang
Citations: 40
h-index: 3
Xuzheng Yang
Xuzheng Yang
Citations: 28
h-index: 2

본 연구에서는 아직 충분히 탐구되지 않은, 하지만 매우 어려운 과제인 일반화된 지칭 표현 이해(GREC) 문제를 다룬다. GREC는 모델이 텍스트 표현으로 설명되는 객체가 존재할 때 해당 객체를 정확하게 찾아내고 (긍정 샘플), 존재하지 않을 때는 출력을 거부해야 한다 (부정 샘플). 멀티모달 대규모 언어 모델(MLLM)은 기존 객체 찾기 능력에서는 뛰어난 성능을 보이지만, 훈련 과정에서 부정 샘플이 부족하여 존재하지 않는 객체를 잘못 찾아내는 (환각된 바운딩 박스 생성) 경우가 종종 발생한다. 기존의 추가 학습 방법인 지도 미세 조정(SFT) 및 강화 학습(RL)은 거부 능력을 향상시키지만, 일반적으로 긍정 샘플에서의 위치 정확도를 저하시켜 모델의 핵심 기능을 약화시킨다. 이러한 문제를 해결하기 위해, 본 연구에서는 MLLM의 거부 능력을 강화하면서도 위치 찾기 성능을 유지하는, 교정된 강화 학습 전략인 Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO)를 제안한다. RC-GRPO는 부정 샘플에 대한 정확한 보상 추정을 위해 시뮬레이션 과정에서 '없음(None)' 출력을 강제하고, 긍정 샘플에서의 과도한 거부를 방지하기 위한 패널티를 적용하여 정확성과 신뢰성 간의 균형을 맞춘다. 또한, 두 번째 단계의 추론 강화 학습은 인과적 이해와 해석 가능성을 더욱 강화한다. 세 가지 GREC 벤치마크에 대한 실험 결과, RC-GRPO는 우수한 위치 찾기 정확도를 달성하면서도 강력한 거부 능력을 유지하는 것으로 나타났다.

Original Abstract

We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expression when it exists (positive sample) and to refuse output when it does not (negative sample). Although Multimodal Large Language Models (MLLMs) excel at localizing existing objects, they often fail to reject nonexistent ones due to the absence of negative samples during training, producing hallucinated bounding boxes. Existing post-training approaches such as supervised fine-tuning (SFT) and reinforcement learning (RL) enhance refusal behavior but usually degrade localization accuracy on positive samples, undermining the model's core competence. To address this, we propose Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO), a calibrated RL strategy that strengthens the refusal ability of MLLMs while preserving localization performance. It enforces "None" outputs in rollouts for valid advantage estimation on negative samples and applies a penalty to prevent over-refusal on positives, achieving a balanced trade-off between accuracy and reliability. A second-stage reasoning reinforcement further consolidates causal understanding and interpretability. Experiments on three GREC benchmarks demonstrate that RC-GRPO attains superior localization accuracy while maintaining strong refusal capability.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!