2607.19691v1 Jul 22, 2026 cs.CL

SLPO: 서브 정책을 활용한 잠재적 추론 확장

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Yongqi Li
Yongqi Li
Citations: 1,485
h-index: 15
Wenjie Li
Wenjie Li
Citations: 807
h-index: 10
Runyang You
Runyang You
Citations: 38
h-index: 4
Zhiyuan Liu
Zhiyuan Liu
Citations: 1,188
h-index: 11

검증 가능한 보상을 사용하는 강화 학습은 명시적인 연쇄적 사고(Chain-of-Thought) 추론기의 테스트 시간 확장에 가장 효과적인 방법으로 자리 잡았습니다. 그러나 이러한 확장 과정은 계산 비용이 많이 들기 때문에, 각 중간 단계는 언어 토큰으로 디코딩되어야 합니다. 반면, 잠재적 추론은 중간 계산을 연속 벡터로 표현하며, 훨씬 짧은 시간 내에 명시적인 연쇄적 사고를 능가합니다. 그러나 잠재적 추론기는 여전히 모방 학습에 의존하는 경향이 있는 반면, 명시적인 연쇄적 사고는 결과-보상 강화 학습을 통해 이미 모방 학습의 한계를 넘어섰습니다. 잠재적 경로에는 각 단계별 확률 분포를 쉽게 계산할 수 없으며, 고정된 사고 예산 하에서 적응적인 종료 인터페이스가 부족하기 때문에, 결과 보상은 잠재적 테스트 시간 확장을 유도할 수 없습니다. 본 연구에서는 자동 회귀(autoregressive) 잠재적 추론기에 결과-보상 강화 학습을 적용하기 위해 서브 정책 잠재적 최적화(SLPO)를 제안합니다. SLPO는 경로 수준의 신용 할당을 위한 잠재적 전이(transition)에 대한 경험적인 서브 정책 밀도와, 결과-보상 최적화를 통해 가변 길이 정책으로 개선되는 정확성 기반 종료 모듈로 구성됩니다. 연속적 및 부드러운 사고 환경에서 SLPO는 병렬 샘플링 하에서 Pass@$k$ 성능을 향상시키고, 더 높은 결정론적 정확도를 가진 어려운 인스턴스에 더 많은 잠재적 계산 리소스를 할당합니다.

Original Abstract

Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across continuous and soft thinking settings, SLPO improves Pass@$k$ under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!