2607.29062v1 Jul 31, 2026 cs.AI

체인 오브 소트(Chain-of-Thought)의 신뢰성 향상을 위한 가이드 벡터 일반화 연구

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

Iv'an Arcuschin
Iv'an Arcuschin
Citations: 282
h-index: 6
Kyle Cox
Kyle Cox
Citations: 13
h-index: 2
Matthew Nguyen
Matthew Nguyen
Citations: 3
h-index: 1
Austin Meek
Austin Meek
Citations: 8
h-index: 1

대규모 언어 모델의 성능 향상은 주로 체인 오브 소트 기술의 발전 덕분입니다. 체인 오브 소트는 AI 안전 측면에서 중요한 역할을 합니다. 모델이 추론 과정을 명시적으로 보여주기 때문에, 이를 통해 모델의 작동 방식을 모니터링할 수 있습니다. 그러나 일부 경우에 모델은 추론 과정에서 중요한 단계를 명시하지 않습니다. 예를 들어, 잘못된 답변을 유도하는 힌트가 주어졌을 때, 모델은 해당 힌트를 무시하는 경우가 있으며, 이 힌트는 모델의 결론에 중요한 영향을 미칠 수 있습니다. 체인 오브 소트가 중요한 추론 단계를 밝히지 못할 때, 우리는 이를 '신뢰성 부족'이라고 표현합니다. 기존 연구에서는 활성화 가이드(activation steering)를 사용하여 체인 오브 소트의 신뢰성을 향상시키는 방법이 효과적임이 입증되었습니다. 본 연구에서는 세 가지 모델(Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B)을 대상으로, 가이드 벡터를 구성하는 다양한 방법, 데이터셋 및 힌트 유형에 따른 신뢰성 향상의 일반화 정도를 분석했습니다. 연구 결과, 가이딩은 가장 큰 모델(Gemma-3 12B)에서만 일관되게 힌트 인지도를 높이는 효과를 보였습니다. 그러나 가이딩이 효과적일 때, 그 효과는 다양한 힌트 유형과 데이터셋에 걸쳐 광범위하게 일반화되는 경향을 보입니다. 즉, 효과 크기는 벡터의 학습 환경보다는 평가 환경에 의해 주로 결정됩니다. 또한, 벡터를 구성하는 방법에 따른 차이는 미미했습니다. 특정 힌트를 언급하지 않는 최적화 방법을 포함한 네 가지 방법이 유사한 효과 크기를 나타냈습니다. 마지막으로, 가이딩이 단순히 힌트의 중요성을 강조하여 힌트 사용을 증가시키는 것인지, 아니면 명시적인 추론 행동을 유도하는 것인지에 대한 가능성을 고려했습니다. 그러나 이러한 주장을 뒷받침할 만한 증거는 발견되지 않았습니다. 가이딩은 힌트 사용률을 거의 변경시키지 않으면서, 명시적으로 언급되지 않는 힌트 사용량을 감소시키는 것으로 나타났습니다.

Original Abstract

Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their conclusion. When chain of thought (CoT) fails to disclose instrumental reasoning steps, we describe it as unfaithful. Prior work has shown that activation steering can be a useful method to improve faithfulness in CoT. We extend this line of work by studying how well steering for faithfulness generalizes across cue types, datasets, and methods of constructing the steering vector for three models (Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B) in a cued question-answering setting. While steering reliably increases cue acknowledgment for only the largest model (Gemma-3 12B), we find that when steering is effective, its effect generalizes broadly across cue types and datasets--in cross-cue and cross-dataset analyses, effect size is determined primarily by the evaluation setting, rather than the vector's train setting. How the vector is built also matters little--four construction methods, including one whose optimization target mentions no specific cue, yield similar effect sizes. Finally, we consider the possibility that steering promotes the salience of the cue and causes greater cue use, rather than targeting verbalization behaviors. However, we find no evidence for this--steering leaves the rate of cue use roughly unchanged while reducing hidden cue use, i.e., cue use that is not acknowledged.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!