2603.01437v1 Mar 02, 2026 cs.AI

사고 과정(Chain-of-Thought) 이전의 답변 해독: 사전 사고 과정 탐색(Pre-CoT Probes) 및 활성화 제어(Activation Steering)를 통한 증거

Decoding Answers Before Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering

Adrià Garriga-Alonso
Adrià Garriga-Alonso
Citations: 318
h-index: 6
Kyle Cox
Kyle Cox
Citations: 13
h-index: 2
Darius Kianersi
Darius Kianersi
Citations: 20
h-index: 2

사고 과정(CoT)이 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 핵심 요소가 되면서, 동시에 모델의 의사 결정 과정을 언어적으로 설명함으로써 해석 가능성을 높이는 유망한 도구로 부상했습니다. 하지만 CoT의 해석 가능성은 모델이 제시하는 추론이 실제 의사 결정 과정을 제대로 반영하는지, 즉 '충실성'에 달려 있습니다. 본 연구에서는 지시 학습(instruction-tuning)된 모델들이 종종 CoT를 생성하기 전에 답변을 결정한다는 기계적인 증거를 제시합니다. CoT 생성 직전의 잔류 스트림 활성화(residual stream activations)에 대한 선형 탐색(linear probes)을 통해 학습한 결과, 대부분의 작업에서 모델의 최종 답변을 0.9의 AUC(Area Under the Curve)로 예측할 수 있었습니다. 이러한 활성화 방향은 단순히 예측적일 뿐만 아니라 인과적 관계를 가지는 것으로 나타났습니다. 탐색 방향을 따라 활성화를 제어하면 모델의 답변이 50% 이상의 경우에 반전되며, 이는 수직 방향을 기준으로 하는 기준(baseline)보다 훨씬 높은 수치입니다. 활성화 제어가 잘못된 답변을 유발하는 경우, 두 가지 주요 오류 모드가 관찰되었습니다. 첫째는 '함의 오류(non-entailment)'로, 올바른 전제를 제시하지만 근거 없는 결론을 도출하는 경우입니다. 둘째는 '허구(confabulation)'로, 거짓된 전제를 만들어내는 경우입니다. CoT는 모델이 올바른 사전 지식을 가진 경우 유용할 수 있지만, 위와 같은 오류 모드는 거짓된 지식에서 추론할 때 원치 않는 결과를 초래할 수 있음을 시사합니다.

Original Abstract

As chain-of-thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning. However, the utility of CoT toward interpretability depends upon its faithfulness -- whether the model's stated reasoning reflects the underlying decision process. We provide mechanistic evidence that instruction-tuned models often determine their answer before generating CoT. Training linear probes on residual stream activations at the last token before CoT, we can predict the model's final answer with 0.9 AUC on most tasks. We find that these directions are not only predictive, but also causal: steering activations along the probe direction flips model answers in over 50% of cases, significantly exceeding orthogonal baselines. When steering induces incorrect answers, we observe two distinct failure modes: non-entailment (stating correct premises but drawing unsupported conclusions) and confabulation (fabricating false premises). While post-hoc reasoning may be instrumentally useful when the model has a correct pre-CoT belief, these failure modes suggest it can result in undesirable behaviors when reasoning from a false belief.

10 Citations
1 Influential
3 Altmetric
27.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!