2608.04001v1 Aug 04, 2026 cs.LG

추론 LLM에서의 테스트 시간 스케일링: 추론 방식, 평가 및 재현성

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Debargha Ganguly
Debargha Ganguly
Citations: 75
h-index: 5
Sreehari Sankar
Sreehari Sankar
Citations: 13
h-index: 2
Vipin Chaudhary
Vipin Chaudhary
Citations: 160
h-index: 6
Mohsen Hariri
Mohsen Hariri
Citations: 100
h-index: 6
Nahal Shahini
Nahal Shahini
Citations: 145
h-index: 4
Kai Ye
Kai Ye
Citations: 0
h-index: 0
Amirhossein Samandar
Amirhossein Samandar
Citations: 36
h-index: 4
Michael Hinczewski
Michael Hinczewski
Citations: 9
h-index: 2

대규모 언어 모델은 더 많은 연산 자원을 활용하여 훨씬 복잡한 추론 문제를 해결할 수 있습니다. 그러나 현재 '테스트 시간 스케일링'이라는 용어는 다양한 추론 알고리즘을 포괄하는데, 이들은 단일 경로를 따라 사고 과정을 확장하거나, 완성된 후보들을 샘플링하여 투표 또는 검증을 통해 통합하거나, 불완전한 부분 상태에 대해 탐색하는 방식을 사용합니다. 이러한 알고리즘은 통계적 구조, 연산 처리 방식 및 오류 발생 패턴에서 차이가 있습니다. 이러한 절차를 단일 '예산'이라는 스칼라 값으로 동일하게 취급하거나, 추론 과정을 명시하지 않고 정확도를 보고하면 연구 결과의 비교가 어렵습니다. 본 논문에서는 테스트 시간 스케일링을 세 가지 측면에서 체계적으로 분석합니다. 첫째, 우리는 테스트 시간 스케일링을 자기 회귀 모델의 암묵적인 접두사 트리에서의 제한된 추론으로 공식화하고, 단일 경로 순차적 스케일링, 터미널 감소를 통한 리프 레벨 스케일링, 그리고 접두사 레벨 스케일링이라는 세 가지 구조적 체제를 구분합니다. 둘째, 우리는 평가 대상을 전체 추론 시스템으로 보고, 엔드투엔드 시스템 성능과 후보 풀 진단 간의 분리를 가능하게 하는 평가 원칙을 개발합니다. 우리는 좌표와 간단한 함수를 사용하여 일반적인 반복 샘플링 지표를 복구하거나 제한하며, 연산 및 불확실성에 대한 프로토콜 일치 보고 방식을 제시합니다. 셋째, 우리는 추론 프로토콜에 대한 재현성 요구 사항을 명시하고, 정확한 재생과 분포적 재현성을 구분하며, 각 경우에 필요한 요소들을 식별합니다. 또한, 우리는 오픈 가중치 추론 생태계를 모델 측면 및 인터페이스 메커니즘으로 구성하고, 이러한 원칙을 광범위한 지식, 기호 추론 및 경쟁 수학 벤치마크에 적용하며, 개선된 검증 및 토큰 레벨 신호를 포함하여 20억 개 이상의 완전한 추론 데이터를 공개합니다.

Original Abstract

Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate them through voting or verification, or search over unfinished partial states. These algorithms differ in their statistical structure, compute accounting, and failure modes. Treating these procedures as interchangeable under a single scalar "budget," or reporting accuracy without the inference protocol that produced it, makes results difficult to compare across studies. We develop a systematic account of test-time scaling along three axes. First, we formalize test-time scaling as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish three structural regimes: single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the evaluated object as the entire inference system and develop evaluation principles that separate end-to-end system performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and prescribe protocol-matched reporting of compute and uncertainty. Third, we specify reproducibility requirements for inference protocols, distinguishing exact replay from distributional reproducibility and identifying the artifacts needed to support each. We also organize the open-weight reasoning ecosystem by model-side and interface mechanisms, apply these principles to broad-knowledge, symbolic-reasoning, and competition-mathematics benchmarks, and assemble over 2 billion full reasoning traces for release with progressively richer verifier and token-level signals.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!