2605.30244v1 May 28, 2026 cs.CV

강건한 채점 기준을 활용한 강화 학습

Reinforcement Learning with Robust Rubric Rewards

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Dandan Tu
Dandan Tu
Citations: 50
h-index: 3
Yong Liao
Yong Liao
Citations: 9
h-index: 2
Ya-Qi Yu
Ya-Qi Yu
Citations: 75
h-index: 3
Fang Hong
Fang Hong
Citations: 2
h-index: 1
Xiangyan Qu
Xiangyan Qu
Citations: 35
h-index: 3
Gaojie Wu
Gaojie Wu
Citations: 250
h-index: 5
Nuo Xu
Nuo Xu
Citations: 168
h-index: 2
Huixin Wang
Huixin Wang
Citations: 11
h-index: 2
Wuheng Xu
Wuheng Xu
Citations: 6
h-index: 2
Haonan Li
Haonan Li
Citations: 112
h-index: 4
Dezhi Peng
Dezhi Peng
Citations: 1,324
h-index: 18
Minghui Liao
Minghui Liao
Citations: 245
h-index: 5
Jihao Wu
Jihao Wu
Citations: 282
h-index: 6
Hao Wang
Hao Wang
Citations: 11
h-index: 2
Ziming Li
Ziming Li
Citations: 491
h-index: 13
Qiaoyu Luo
Qiaoyu Luo
Citations: 2
h-index: 1
Zihao Chen
Zihao Chen
Citations: 2
h-index: 1

검증 가능한 보상을 사용하는 강화 학습(RLVR)은 결정적으로 검증할 수 있는 작업에 효과적이지만, 많은 시각-언어 작업은 부분적으로만 검증 가능하며 다중 기준의 감독이 필요합니다 (예: 인식적인 세부 사항, 추론 단계 및 제약 조건). 채점 기준은 이러한 미세한 수준의 감독을 위한 자연스러운 인터페이스를 제공하지만, 온라인 강화 학습 과정에서의 실행 정확도에 따라 효과가 달라집니다. 본 연구에서는 강건한 채점 기준을 활용한 강화 학습(RLR^3)을 제안합니다. RLR^3은 RLVR을 작업 수준 검증에서 기준 수준 검증으로 확장합니다. RLR^3은 인스턴스별 채점 기준을 두 가지 실행 경로를 통해 처리합니다: LLM을 추출기로 사용하고 결정적인 검증기를 결합하거나, 검증 불가능한 기준에 대해서는 LLM을 판정자로 사용합니다. 정확한 점수 부여를 보장하기 위해 RLR^3은 추출기에게 정답 정보를 숨기고 판정자에게 이미지를 숨기는 최소 노출 전략을 도입했습니다. 또한, RLR^3은 계층적 집계를 사용하여 필수적인 기준에 더 높은 우선순위를 부여하고, rollout 그룹 내의 점수 포화를 완화합니다. Qwen3-VL-30B-A3B 모델을 사용하여 15개의 벤치마크에서 RLR^3을 평가한 결과, RLR^3은 RLVR보다 일관되게 우수한 성능을 보였으며, 기준 모델 대비 4.7점의 성능 향상을 달성하고 공식 instruct-to-thinking 모델과의 격차를 뛰어넘었습니다. 통제된 실험 결과는 우리의 결정적인 검증과 최소 노출 전략이 악용될 수 있는 오탐을 크게 줄이는 것을 확인했습니다.

Original Abstract

While Reinforcement Learning with Verifiable Rewards (RLVR) is effective for deterministically checkable tasks, many vision-language tasks are partially verifiable, demanding multi-criteria supervision (e.g., perceptual details, reasoning steps, and constraints). Rubrics provide a natural interface for this fine-grained supervision, but their effectiveness depends on the execution accuracy during online RL. We propose Reinforcement Learning with Robust Rubric Rewards ($\text{RLR}^3$), extending RLVR from task-level verification to criterion-level verification. $\text{RLR}^3$ routes instance-specific rubrics through two execution paths: an LLM-as-an-extractor paired with a deterministic verifier, or an LLM-as-a-Judge for non-verifiable criteria. To ensure faithful scoring, $\text{RLR}^3$ introduce a minimal exposure strategy that masks ground truths from extractors and images from judges. Furthermore, $\text{RLR}^3$ employs hierarchical aggregation to prioritize essential criteria over additional criteria, and mitigates score saturation within rollout groups. Evaluated on Qwen3-VL-30B-A3B across 15 benchmarks, $\text{RLR}^3$ consistently outperforms RLVR, yielding a 4.7-point improvement over the base model and exceeding the official instruct-to-thinking model gap. Controlled audits confirm our deterministic verification and minimal exposure significantly reduce exploitable false positives.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!