엔트로피 게이티드 잠재적 반복
Entropy-Gated Latent Recursion
추론 시간 스케일링은 언어 모델의 추론 능력을 향상시키는 주요 방법으로 자리 잡았지만, 기존 방법들은 주로 확률적인 토큰 레벨 샘플링이라는 단일 요소를 통해 다양성을 확보합니다. 우리는 이러한 단일 축 기반의 샘플링 공간이 근본적으로 제한적이라고 주장하며, 두 번째로 완전히 결정론적이고 상호 보완적인 축을 제시합니다. 이 축은 불확실성이 높은 토큰에 대해 미리 학습된 모델의 최상위 디코더 레이어를 재사용하는 레이어 범위 $L$입니다. $L$의 선택 값에 따라 서로 다른 결과를 도출하며, 이는 다양한 문제 유형을 해결하며 확률적 요소가 전혀 포함되지 않습니다. 우리는 이 축을 엔트로피 게이티드 잠재적 반복(EGLR)이라는 학습이 필요 없는 디코딩 절차를 통해 구현했습니다. EGLR은 다음 토큰의 분포가 수렴될 때까지 최상위 $L$개의 레이어를 최대 $K_{ ext{max}}$번 반복 적용합니다. $T$개의 온도 샘플과 결합된 EGLR은 단일 축 기반의 확률적 rollout 풀을 거의 동일한 비용으로 $L imes T$ 카르테시안 샘플링 공간으로 확장합니다. 우리는 8개의 instruction-tuned 모델과 6개의 수학 추론 벤치마크에서 이 공간을 분석했으며, $L$ 축이 온도를 통해 얻는 효과와 진정으로 상호 보완적임을 확인했습니다. MATH-500 데이터셋에서 Qwen2.5-3B-Instruct 모델을 사용했을 때, $L$과 온도 모두를 활용한 최적의 성능은 91.6%로, 온도만 사용한 경우(83.4%)보다 8.2%p 높고, 레이어만 사용한 경우(81.2%)보다 10.4%p 높은 수치를 보였습니다. 확장된 rollout 풀은 self-consistency, best-of-$N$ with verifiers, 그리고 group-relative RL training (GRPO)과 같은 downstream 절차에 더 풍부한 후보를 제공하며, 확률적 잡음에 의존하지 않는 새로운 추론 시간 스케일링 방향을 제시합니다.
Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-level sampling. We argue that this single-axis sampling space is fundamentally limiting, and identify a second, fully deterministic and complementary axis: the layer span $L$ at which a frozen model's top decoder layers are recursively re-applied at high-uncertainty tokens. Different choices of $L$ produce distinct rollouts that solve different subsets of problems, with no stochasticity. We instantiate this axis through Entropy-Gated Latent Recursion (EGLR), a training-free decoding procedure that re-applies the top-$L$ layers for at most $K_{\max}$ iterations until the next-token distribution converges. Combined with $T$ temperature samples, EGLR turns a single-axis stochastic rollout pool into an $L\times T$ Cartesian sampling space at almost the same per-rollout cost. We characterize this space across $8$ instruction-tuned models and $6$ math reasoning benchmarks, and show that the $L$-axis is genuinely complementary to temperature: on MATH-500 with Qwen2.5-3B-Instruct, the joint $L\times T$ oracle reaches $91.6\%$, $+8.2$ percentage points beyond the temperature-only oracle ($83.4\%$) and $+10.4$ points beyond the layer-only oracle ($81.2\%$), confirming that the two axes capture genuinely complementary problems. The expanded rollout pool provides richer per-prompt candidates for any downstream procedure that consumes rollouts, including self-consistency, best-of-$N$ with verifiers, and group-relative RL training (GRPO), opening a new direction for inference-time scaling that does not rely on stochastic noise.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.