2606.25524v1 Jun 24, 2026 cs.AI

절벽 토큰: LLM의 수학적 추론에서 단일 토큰 오류 발생 요인 식별

Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

Pilsung Kang
Pilsung Kang
Citations: 230
h-index: 5
J. Ko
J. Ko
Citations: 0
h-index: 0
Yukyung Lee
Yukyung Lee
Boston University
Citations: 372
h-index: 8

대규모 언어 모델(LLM)은 수학적 추론에서 높은 정확도를 달성하지만, 동일한 문제에 대한 개별적인 추론 과정은 서로 다릅니다. 일부는 정답을 도출하는 반면, 다른 일부는 실패합니다. 기존 연구에서는 단계, 덩어리 또는 문장 수준에서 오류를 분석하거나 이미 오류가 발생한 토큰에서의 오류를 분석했습니다. 하지만, 이러한 방법들은 오류로 이어지는 정확한 토큰을 식별하지 못합니다. 본 연구에서는 '절벽 토큰'이라는 개념을 도입합니다. 절벽 토큰은 일방향 두 비율 Z-검정을 기반으로, 로컬 토큰 잠재력에 따라 조정되는 적응적 임계값 하에서 토큰 잠재력이 현저하게 감소하는 토큰을 의미합니다. 7개의 모델과 세 가지 수학적 추론 벤치마크(GSM1K, MATH500, AIME 2025)를 대상으로 실험한 결과, 절벽 토큰은 오류 발생의 촉매 역할을 합니다. 첫 번째 절벽 토큰을 삭제하고 재샘플링하면 pass@64가 1.0으로 회복되지만, 이를 유지하면 회복률이 0.71에서 1.00 사이로 제한됩니다. 또한, 탐욕적 선택과 토큰 엔트로피를 기준으로 결정론적, 불확실성, 샘플링 오류의 세 가지 유형으로 절벽 토큰을 분류했습니다. 각 유형은 구별되는 확률적 특징을 가지고 있으며, 이러한 분류는 모델 규모에 관계없이 일반화됩니다. 마지막으로, 절벽 위치에서 단일 토큰 선호도 최적화(Cliff-DPO)를 통해 제안된 분류의 유효성을 검증했습니다. GSM8K 데이터셋으로 학습된 Cliff-DPO는 벤치마크 전반에 걸쳐 최대 +6.6까지 정확도를 향상시켰습니다. 불확실성 및 샘플링 오류 절벽에서 최적화를 수행하면 추론 능력이 향상되는 반면, 결정론적 절벽에서는 효과가 미미합니다.

Original Abstract

Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive at the correct answer while others fail. Prior work analyzes failure at the step, chunk, or sentence level, or at tokens where failure has already occurred. Neither identifies the precise token that triggers the shift toward failure. We introduce the cliff token, a token where the token-wise potential drops significantly under an adaptive threshold that scales with the local token-wise potential, based on a one-sided two-proportion z-test. Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to between 0.71 and 1.00. We further introduce a cliff taxonomy of deterministic, uncertain, and sampled-off cliffs, defined by greedy choice and token entropy. Each type has distinct probabilistic characteristics, and the taxonomy generalizes across model scales. Finally, we validate the taxonomy via single-token preference optimization at cliff positions (Cliff-DPO). Trained on GSM8K, Cliff-DPO improves accuracy across benchmarks by up to +6.6. Optimizing at uncertain and sampled-off cliffs improves reasoning, while deterministic cliffs do not.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!