2605.26795v1 May 26, 2026 cs.AI

추론 시점에 체인 오브 소트(Chain-of-Thought)가 효과를 내는 이유는 무엇인가? 전역적인 추론보다는 지역적인 공존성

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

Wei Wei
Wei Wei
Citations: 144
h-index: 5
Xiang Wang
Xiang Wang
Citations: 77
h-index: 2

체인 오브 소트(CoT) 프롬프트는 언어 모델의 정확도를 안정적으로 향상시키지만, 어떤 이유 설명 텍스트의 특성이 이러한 성능 향상을 이끄는지에 대한 이해는 부족합니다. 기존 연구들은 주로 생성 시점에서의 동작을 분석했습니다. 본 연구에서는 추론 시점에 초점을 맞춰, 주어진 이유 설명 텍스트 내에서 어떤 요소가 답변에 영향을 미치는지를 조사합니다. 우리는 두 가지 상호 보완적인 요인을 확인했습니다. 첫째, 전역적으로 단어를 재배열한 이유 설명 텍스트조차도 이유 설명이 없는 기준보다 훨씬 뛰어난 성능을 보이며, 이는 강력한 어휘 활성화 효과를 나타냅니다. 더욱 중요한 점은, 구조화된 텍스트가 제공하는 추가적인 성능 향상은 문장 수준의 논리적 순서보다는 짧은 범위의 토큰 인접성에 더 크게 기인한다는 것입니다. 길이가 $n^ extit{*}=2$에서 $3$ 토큰에 불과한 연속적인 윈도우를 유지하는 것만으로도 CoT의 전체 성능 향상 중 대부분을 회복할 수 있습니다. 추가적인 실험 결과는 명시적인 답변 선언 또는 답변 값의 단순 복사, 그리고 완전한 문법적 실현이 주요 원인이 아니라는 것을 배제합니다. 더 나아가, 다양한 모델 계열, 파라미터 규모 및 데이터 세트에 대한 일반화 실험을 통해 이러한 경향성이 안정적으로 유지되는 것을 확인했습니다. 이러한 결과는 추론 시점의 CoT에 대한 지역 공존 활성화(LCA) 가설을 뒷받침하며, 관찰된 성능 향상은 주로 어휘 활성화와 짧은 범위의 토큰 공존에 기인하며 문장 수준의 논리적 추론과는 거리가 멀다는 것을 시사합니다.

Original Abstract

Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly understood. Prior work has largely studied generation-time behavior. We instead ask a probe-time question: given a fixed rationale in context, what in that text changes the answer? We identify two complementary sources of the gain. First, even a globally word-shuffled rationale substantially outperforms the no-rationale baseline, indicating a strong lexical activation effect. More importantly, the additional gain from structured text appears to arise less from sentence-level logical ordering and more from short-range token adjacency. Preserving contiguous windows of just $n^\star{=}2$--$3$ tokens recovers most of the remaining gain toward full CoT performance. Supporting experiments rule out copying of explicit answer declarations or answer values, as well as full grammatical realization, as primary drivers. Further generalization experiments show that the qualitative pattern remains stable across multiple model families, parameter scales, and datasets. These results support a local co-occurrence activation (LCA) account of probe-time CoT, in which the observed gains appear to arise primarily from lexical activation and short-range token co-occurrence rather than sentence-level logical derivation.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!