2607.27610v1 Jul 30, 2026 cs.LG

칼만 필터와 교육 과정의 만남: 적응형 강화 학습 미세 조정(Fine-tuning)을 위한 효율적인 동적 프롬프트 선택 방법

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

Haodong Zhu
Haodong Zhu
Citations: 142
h-index: 3
Yangyang Ren
Yangyang Ren
Citations: 13
h-index: 1
Yanjing Li
Yanjing Li
Beihang University
Citations: 982
h-index: 14
Linlin Yang
Linlin Yang
Citations: 64
h-index: 5
Haiguang Liu
Haiguang Liu
Citations: 42
h-index: 2
Baochang Zhang
Baochang Zhang
Citations: 50
h-index: 5
Sheng Xu
Sheng Xu
Citations: 1,085
h-index: 16

강화 학습(RL) 기반 미세 조정은 거대 언어 모델(LLM)의 추론 능력을 크게 향상시키지만, 그 효과는 현재 정책에 적합한 난이도의 프롬프트를 선택하는 데 달려 있습니다. 이는 프롬프트의 난이도가 훈련 과정 동안 변하기 때문에 어려운 문제입니다. 기존의 온라인 방법은 정확하지만 비용이 많이 드는 평가 기반 방식과 효율적이지만 일반적으로 정적인 난이도를 가정하여 강화 학습의 비정상적인 훈련 동역학에 적합하지 않은 예측 기반 방식 간의 절충점을 갖습니다. 이러한 문제를 해결하기 위해, 우리는 프롬프트 선택을 정적 난이도 예측 대신 동적 상태 추정 문제로 재구성하는 칼만 필터 기반 프롬프트 선택 방법(KGPS)을 제안합니다. KGPS는 각 프롬프트의 잠재적인 성공률을 로짓 공간에서 선형 가우시안 상태-공간 모델로 모델링하며, 과정 잡음은 정책 업데이트의 크기에 연결되어 정책이 더 크게 변경될수록 불확실성이 증가합니다. 칼만 필터는 프롬프트 난이도에 대한 보정된 가우시안 사후 분포를 유지하며, 프롬프트를 선택할 때에는 중간 난이도의 프롬프트를 선호하면서 자연스럽게 불확실한 프롬프트를 재방문하는 사후 기대 훈련 유틸리티를 최대화합니다. 결과적으로 얻어지는 절차는 정책 변화에 적응력이 뛰어나며, 표준 정책 훈련 외에 추가적인 시뮬레이션을 필요로 하지 않습니다. 수학, 계획 및 기하학적 추론 벤치마크와 다양한 강화 학습 알고리즘에서의 광범위한 실험에서 KGPS는 강력한 기준 모델보다 최종 정확도와 롤아웃 효율성 모두에서 꾸준히 우수한 성능을 보이며 온라인 프롬프트 선택 방법 중 최고 수준의 성능을 달성합니다. 예를 들어, DeepSeek-R1-Distill-7B 모델에서 KGPS는 DS 방식에 비해 83% 더 적은 롤아웃으로 사용하면서도 여섯 가지 수학 추론 벤치마크에서 평균적으로 0.12 포인트의 성능 향상을 보였습니다.

Original Abstract

Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy. This is challenging because prompt difficulty evolves throughout training. Existing online methods therefore face a trade-off: evaluation-based approaches are accurate but expensive, while prediction-based approaches are efficient but typically assume stationary difficulty, making them ill-suited to RL's non-stationary training dynamics. To address these issues, we propose a Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction. KGPS models each prompt's latent success rate in logit space using a linear-Gaussian state-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially. A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior-expected training utility that favors intermediate-difficulty prompts while naturally revisiting uncertain ones. The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training. Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods. For example, on DeepSeek-R1-Distill-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0.12 point across six math reasoning benchmarks.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!