2606.16620v1 Jun 15, 2026 cs.LG

엔트로피 게이티드 잠재적 반복

Entropy-Gated Latent Recursion

S. Lahlou
S. Lahlou
Citations: 1,730
h-index: 10
M. Takác
M. Takác
Citations: 13
h-index: 3
Soham Bhattacharjee
Soham Bhattacharjee
Citations: 47
h-index: 3
Nils Lukas
Nils Lukas
Citations: 1,108
h-index: 11
Dushyant Singh Chauhan
Dushyant Singh Chauhan
Citations: 1,008
h-index: 14

추론 시간 스케일링은 언어 모델의 추론 능력을 향상시키는 주요 방법으로 자리 잡았지만, 기존 방법들은 주로 확률적인 토큰 레벨 샘플링이라는 단일 요소를 통해 다양성을 확보합니다. 우리는 이러한 단일 축 기반의 샘플링 공간이 근본적으로 제한적이라고 주장하며, 두 번째로 완전히 결정론적이고 상호 보완적인 축을 제시합니다. 이 축은 불확실성이 높은 토큰에 대해 미리 학습된 모델의 최상위 디코더 레이어를 재사용하는 레이어 범위 $L$입니다. $L$의 선택 값에 따라 서로 다른 결과를 도출하며, 이는 다양한 문제 유형을 해결하며 확률적 요소가 전혀 포함되지 않습니다. 우리는 이 축을 엔트로피 게이티드 잠재적 반복(EGLR)이라는 학습이 필요 없는 디코딩 절차를 통해 구현했습니다. EGLR은 다음 토큰의 분포가 수렴될 때까지 최상위 $L$개의 레이어를 최대 $K_{ ext{max}}$번 반복 적용합니다. $T$개의 온도 샘플과 결합된 EGLR은 단일 축 기반의 확률적 rollout 풀을 거의 동일한 비용으로 $L imes T$ 카르테시안 샘플링 공간으로 확장합니다. 우리는 8개의 instruction-tuned 모델과 6개의 수학 추론 벤치마크에서 이 공간을 분석했으며, $L$ 축이 온도를 통해 얻는 효과와 진정으로 상호 보완적임을 확인했습니다. MATH-500 데이터셋에서 Qwen2.5-3B-Instruct 모델을 사용했을 때, $L$과 온도 모두를 활용한 최적의 성능은 91.6%로, 온도만 사용한 경우(83.4%)보다 8.2%p 높고, 레이어만 사용한 경우(81.2%)보다 10.4%p 높은 수치를 보였습니다. 확장된 rollout 풀은 self-consistency, best-of-$N$ with verifiers, 그리고 group-relative RL training (GRPO)과 같은 downstream 절차에 더 풍부한 후보를 제공하며, 확률적 잡음에 의존하지 않는 새로운 추론 시간 스케일링 방향을 제시합니다.

Original Abstract

Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-level sampling. We argue that this single-axis sampling space is fundamentally limiting, and identify a second, fully deterministic and complementary axis: the layer span $L$ at which a frozen model's top decoder layers are recursively re-applied at high-uncertainty tokens. Different choices of $L$ produce distinct rollouts that solve different subsets of problems, with no stochasticity. We instantiate this axis through Entropy-Gated Latent Recursion (EGLR), a training-free decoding procedure that re-applies the top-$L$ layers for at most $K_{\max}$ iterations until the next-token distribution converges. Combined with $T$ temperature samples, EGLR turns a single-axis stochastic rollout pool into an $L\times T$ Cartesian sampling space at almost the same per-rollout cost. We characterize this space across $8$ instruction-tuned models and $6$ math reasoning benchmarks, and show that the $L$-axis is genuinely complementary to temperature: on MATH-500 with Qwen2.5-3B-Instruct, the joint $L\times T$ oracle reaches $91.6\%$, $+8.2$ percentage points beyond the temperature-only oracle ($83.4\%$) and $+10.4$ points beyond the layer-only oracle ($81.2\%$), confirming that the two axes capture genuinely complementary problems. The expanded rollout pool provides richer per-prompt candidates for any downstream procedure that consumes rollouts, including self-consistency, best-of-$N$ with verifiers, and group-relative RL training (GRPO), opening a new direction for inference-time scaling that does not rely on stochastic noise.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!