2603.21162v1 Mar 22, 2026 cs.AI

LLM을 위한 트리 탐색 재검토: 예산 효율적인 추론을 위한 Gumbel 샘플링 및 순차적 절반 탐색

Revisiting Tree Search for LLMs: Gumbel and Sequential Halving for Budget-Scalable Reasoning

L. Ugadiarov
L. Ugadiarov
Citations: 16
h-index: 3
Aleksandr Panov
Aleksandr Panov
Citations: 4
h-index: 1
Yuri Kuratov
Yuri Kuratov
AIRI
Citations: 1,424
h-index: 13
Alexey Skrynnik
Alexey Skrynnik
Citations: 451
h-index: 13

신경망 트리 탐색은 게임, 모델 기반 강화 학습 등 복잡한 영역에서 널리 사용되는 강력한 의사 결정 알고리즘입니다. 최근 연구에서는 AlphaZero 스타일의 트리 탐색을 사용하여 추론 과정에서 대규모 언어 모델(LLM)의 추론 능력을 향상시키려고 시도했지만, 우리는 이 접근 방식이 확장성 문제를 가지고 있음을 발견했습니다. 즉, GSM8K 및 Game24 데이터셋에서 탐색 예산이 증가함에 따라 정확도가 감소합니다. 본 논문에서는 Gumbel AlphaZero MCTS를 수정하여 ReSCALE이라는 새로운 방법을 제안합니다. ReSCALE은 Dirichlet 노이즈와 PUCT 선택을 Gumbel 샘플링 및 순차적 절반 탐색으로 대체하여, 모델이나 학습 과정에 변경 없이도 정확도가 지속적으로 향상되도록 합니다. ReSCALE은 기존 방법이 성능이 저하되는 예산 수준에서 GSM8K 데이터셋에서 58.4%, Game24 데이터셋에서 85.3%의 정확도를 달성합니다. 추가 분석 결과, 순차적 절반 탐색이 성능 향상의 주요 원인임을 확인했습니다.

Original Abstract

Neural tree search is a powerful decision-making algorithm widely used in complex domains such as game playing and model-based reinforcement learning. Recent work has applied AlphaZero-style tree search to enhance the reasoning capabilities of Large Language Models (LLMs) during inference, but we find that this approach suffers from a scaling failure: on GSM8K and Game24, accuracy drops as the search budget increases. In this paper, we present ReSCALE, an adaptation of Gumbel AlphaZero MCTS that replaces Dirichlet noise and PUCT selection with Gumbel sampling and Sequential Halving, restoring monotonic scaling without changes to the model or its training. ReSCALE reaches 58.4\% on GSM8K and 85.3\% on Game24 at budgets where the baseline degrades. Ablations confirm that Sequential Halving is the primary driver of the improvement.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!