2605.28305v1 May 27, 2026 cs.CL

대규모 언어 모델 추론에서의 인간형 반성 지표 재검토

Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning

Fei Cheng
Fei Cheng
Graduate School of Informatics, Kyoto University
Citations: 909
h-index: 15
Yahan Yu
Yahan Yu
Citations: 529
h-index: 7
Noa Nakanishi
Noa Nakanishi
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 복잡한 추론 과정에서 종종 명시적인 반성 흔적을 생성하며, '잠시 기다려', '음...', '또는'과 같은 인간형 지표를 동반합니다. 이러한 지표들은 반성을 나타내는 시각적 단서로 널리 사용되지만, 그 작동 방식은 여전히 불분명하며, 이는 불필요하고 반복적인 반성 지표로 인해 과도한 사고가 발생할 위험을 초래합니다. 본 연구에서는 인간형 반성 지표의 필요성과 반성 과정에서의 역할을 재검토합니다. 프롬프트 수준 및 토큰 수준의 개입을 통해 이러한 지표를 억제하고, 네 가지 벤치마크와 두 가지 모델 크기에서 작업 성능에 미치는 영향을 분석했습니다. 연구 결과, 인간형 지표는 추론 성능에 일관되게 필요한 요소가 아님을 보여주었습니다. 지표를 억제하면 특정 상황에서 성능을 유지하거나 향상시킬 수 있으며, 특히 더 큰 샘플링 예산 하에서 이러한 효과가 두드러집니다. 동시에, 지표 억제가 반드시 반성 행동을 제거하는 것은 아니며, 모델은 여전히 지표 없이 검증 작업을 수행할 수 있습니다. 이는 인간형 지표가 실제 반성을 나타내는 신뢰할 만한 지표라기보다는 표면적인 단서에 가깝다는 점을 시사하며, 명시적인 지표 패턴을 넘어선 추론 메커니즘에 대한 향후 연구를 촉구합니다.

Original Abstract

Large Language Models (LLMs) often produce explicit reflective traces during complex reasoning, accompanied by anthropomorphic markers such as wait, hmm, and alternatively. Although these markers are commonly used as visible indicators of reflection, their mechanisms remain unclear, which leaves the risk of overthinking associated with redundant and repetitive reflection markers. In this work, we revisit anthropomorphic reflection markers, examining their necessity for reasoning and role in the reflection. We suppress these markers through prompt-level and token-level interventions, and analyze their effects on task performance across four benchmarks and two model scales. Our results show that anthropomorphic markers are not uniformly necessary for reasoning performance: suppressing them can preserve or improve performance in several settings, especially under larger sampling budgets. Meanwhile, marker suppression does not necessarily remove reflection behavior, as models can still perform marker-free verification. These suggest that anthropomorphic markers tend to be surface cues rather than reliable proxies for reflection itself, and motivate future research on reasoning mechanisms beyond explicit marker patterns.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!