2608.05080v1 Aug 05, 2026 cs.LG

정책이 학습하는 데이터를 최적화하기 위한 접근 방식: 복원 가능성을 고려한 롤아웃 개입 학습

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yanfang Ye
Yanfang Ye
Citations: 70
h-index: 5
N. Kuang
N. Kuang
Citations: 31
h-index: 4
Manqing Mao
Manqing Mao
Citations: 0
h-index: 0
Samson Koelle
Samson Koelle
Citations: 700
h-index: 10
Jie Yuan
Jie Yuan
Iowa State University
Citations: 108
h-index: 3
James Feng
James Feng
Citations: 0
h-index: 0
Wei Niu
Wei Niu
Citations: 36
h-index: 2

크리틱(Critic) 없이 그룹 기반 강화 학습은 대규모 언어 모델의 사후 훈련에 적합한 확장 가능한 방법으로 자리 잡았습니다. 그러나 대부분의 기존 방법은 모든 작업 및 경로 상태에 동일한 수의 롤아웃을 할당하는데, 이는 일부 롤아웃이 다른 롤아웃보다 훨씬 더 유용한 학습 신호를 제공하는 반면입니다. 최근 연구에서는 롤아웃 생성 과정을 적응적인 의사 결정으로 간주하기 시작했지만, 여전히 두 가지 중요한 제한 사항이 남아 있습니다. 첫째, 개입 전략은 종종 고정된 휴리스틱에 기반하므로 정책이 훈련 중에 변경됨에 따라 조정할 수 없습니다. 둘째, 이러한 방법은 일반적으로 몇 개의 롤아웃을 생성할지 결정하기만 할 뿐, 어디에서, 어떻게 개입할지를 명시적으로 제어하지 않습니다. 이러한 제한 사항을 해결하기 위해, 우리는 각 개입으로 인해 발생하는 개선 정도를 기반으로 롤아웃을 생성하는 방법을 학습하는 훈련 시간 프레임워크인 Recoverability-Aware Intervention Learning (RAIL)을 제안합니다. RAIL은 개입 선택을 온라인 컨텍스추얼 밴딧 문제로 모델링하고, 샤도우-투-라이브(shadow-to-live) 절차를 통해 수집된 개입 추적 데이터를 사용하여 복원 가능성 제어기를 학습시킵니다. 이를 통해 제어기는 기본 정책이 진화하는 동안에도 계속 학습할 수 있습니다. 우리는 RAIL을 효과성, 적응성, 표현력 및 효율성을 기준으로 평가했습니다. 여러 환경에서 RAIL은 제한된 롤아웃 예산 하에서 일관되게 성능을 향상시켰습니다. 이러한 결과는 복원 가능성을 고려한 개입이 더 유익하고 중복되지 않는 롤아웃을 생성하는 원칙적인 방법을 제공하며, 이를 통해 사후 훈련 중에 강력한 학습 신호를 얻을 수 있음을 보여줍니다.

Original Abstract

Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of rollouts to every task and trajectory state, even though some rollouts provide much more useful learning signals than others. Recent work has started to treat rollout generation as an adaptive decision, but two important limitations remain. First, intervention strategies are often based on fixed heuristics and therefore cannot adjust as the policy changes during training. Second, these methods usually decide only how many rollouts to generate, without explicitly controlling where and how to intervene. To address these limitations, we propose Recoverability-Aware Intervention Learning (RAIL), a training-time framework that learns how to generate rollouts based on the improvement produced by each intervention. RAIL models intervention selection as an online contextual-bandit problem and trains a recoverability controller using intervention traces collected through a shadow-to-live procedure. This allows the controller to keep learning while the underlying policy evolves. We evaluate RAIL in terms of effectiveness, adaptivity, expressiveness, and efficiency. Across multiple settings, RAIL consistently improves performance under limited rollout budgets. These results show that recoverability-aware intervention provides a principled way to generate more informative and less redundant rollouts, leading to stronger learning signals during post-training.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!