2608.03972v1 Aug 04, 2026 cs.AI

ReflectRL: 반사적-직접 추론을 통한 이상적인 부정 샘플 트레이징 학습

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Tat-Seng Chua
Tat-Seng Chua
Citations: 46
h-index: 3
Jinhe Bi
Jinhe Bi
Citations: 321
h-index: 10
Aniri
Aniri
Citations: 4
h-index: 1
Hu Cao
Hu Cao
Citations: 0
h-index: 0
Volker Tresp
Volker Tresp
Citations: 17
h-index: 3
Fei Shen
Fei Shen
Citations: 37
h-index: 3

온라인 학습은 대규모 언어 모델의 추론 능력을 향상시키는 강력한 후처리 패러다임으로 자리 잡았으며, 종종 더 높은 성능을 가진 전문가 모델에서 생성된 이상적인 트레이징 데이터로 강화됩니다. 그러나 전문가 모델이 더 어려운 문제에 실패하는 경우, 기존의 트레이징 기반 방법은 주요 지도 정보원을 잃게 되며, 이러한 실패한 트레이징 데이터는 일반적으로 부정 샘플로 버려집니다. 본 연구에서는 이러한 실패 사례, 즉 '이상적인 부정 트레이징'이 모방할 수 있는 시연 자료가 아닌, 반성해야 할 오류를 포함하는 경로로서 활용될 경우 여전히 귀중한 추론 신호를 제공할 수 있다고 주장합니다. 우리는 '반성적 이점(Reflection Advantage)'을 제시하며, 어려운 문제의 경우 처음부터 문제를 해결하는 것보다 잘못된 트레이징 경로에 대해 반성하는 것이 더 쉽고 효과적일 수 있습니다. 이러한 동기 부여를 바탕으로, 본 연구에서는 온라인 학습 과정에서 이상적인 부정 트레이징으로부터 학습하는 경량화된 플러그 앤 플레이 프레임워크인 ReflectRL을 제안합니다. ReflectRL은 먼저 이러한 트레이징 데이터를 사용하여 '반성적 추론'을 유도하고, 그런 다음 '반성적-직접 정책 전환(Reflective-to-Direct Policy Transition)'을 적용하여 습득된 추론 행동을 직접적인 추론으로 이전합니다. 9개의 벤치마크, 4가지 LLM 백본 및 4가지 온라인 학습 방법에서의 실험 결과는 ReflectRL이 최소한의 오버헤드로 일관되게 추론 성능을 향상시킴을 보여줍니다.

Original Abstract

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed trajectories to reflect upon. We identify a Reflection Advantage: for hard problems, reflecting on a flawed trajectory can be easier and more effective than solving the problem directly from scratch. Motivated by this, we propose ReflectRL, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training. ReflectRL first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning. Experiments across 9 benchmarks, 4 LLM backbones, and 4 on-policy training methods show that ReflectRL consistently improves reasoning performance with minimal overhead.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!