2605.25603v1 May 25, 2026 cs.AI

회로 기반 내부-외부 불일치성 분석을 통한 부정확한 연쇄적 사고 탐지

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy

X. Wang
X. Wang
Citations: 247
h-index: 5
Xu Shen
Xu Shen
Citations: 162
h-index: 6
Zhen Tan
Zhen Tan
Citations: 384
h-index: 6
Song Wang
Song Wang
Citations: 57
h-index: 5
Pingjun Hong
Pingjun Hong
Citations: 21
h-index: 4
Rui Miao
Rui Miao
Citations: 213
h-index: 7
Tianlong Chen
Tianlong Chen
Citations: 0
h-index: 0

연쇄적 사고(Chain-of-Thought, CoT) 추론은 대규모 언어 모델(LLM)의 문제 해결 능력을 향상시키지만, 생성된 추론 과정이 실제 모델의 의사 결정 과정을 정확하게 반영하지 않을 수 있습니다. 기존의 CoT 부정확성 탐지 방법은 주로 생성된 설명에서 얻은 외부 신호에 의존하며, 텍스트의 타당성이나 답변과의 일관성과 같은 요소들을 활용하지만, 모델 내부 계산 과정에서 발생하는 증거는 간과합니다. 최근의 회로 추적 방법은 모델 구성 요소를 통해 정보가 흐르는 방식을 추적하여 모델 내부 정보를 얻을 수 있는 방법을 제공하지만, 긴 CoT에 대한 전체적인 회로를 구축하는 것은 비용이 많이 들고 확장하기 어렵습니다. 이러한 문제점을 해결하기 위해, 우리는 CIE-Scorer(Circuit-guided Internal-External Discrepancy Scorer)라는 프레임워크를 제안합니다. 이 프레임워크는 개별 사례 수준에서 CoT의 부정확성을 탐지하는 데 사용됩니다. 핵심 아이디어는 신뢰할 수 있는 추론 과정은 모델의 계산 과정과 일치해야 한다는 것입니다. 반면, 부정확한 추론 과정은 이러한 일관성이 부족할 수 있습니다. CIE-Scorer는 정보가 풍부한 추론 토큰에서 효율적으로 간결한 문장 수준의 회로를 추적하고, 내부 및 외부 추론 그래프를 구축하며, 퓨즈된 그로모프-워테르슈타인 거리를 사용하여 이러한 그래프 간의 불일치를 측정합니다. FaithCoT-Bench에서 제공하는 네 가지 데이터 세트에 대한 실험 결과, CIE-Scorer는 최첨단 성능을 달성하면서 회로 구축 비용을 줄여주었습니다. 이는 메커니즘적 해석 신호와 외부 추론 과정을 결합하여 CoT 부정확성을 탐지하는 데 효과적임을 보여줍니다.

Original Abstract

Chain-of-thought (CoT) reasoning improves the problem-solving ability of large language models (LLMs), but generated reasoning traces may not faithfully reflect the model's actual decision process. Existing CoT unfaithfulness detectors mainly rely on external signals from generated rationales, such as textual plausibility or answer consistency, while overlooking evidence from the model's internal computation. Although recent circuit tracing methods provide a way to obtain model-internal evidence by tracing how information flows through model components during reasoning, constructing full reasoning circuits for long CoTs is costly and difficult to scale. To address these challenges, we propose Circuit-guided Internal-External Discrepancy Scorer (CIE-Scorer), a framework for instance-level CoT unfaithfulness detection. The key idea is that faithful reasoning traces should align with the model's computational process, whereas unfaithful traces may diverge from it. CIE-Scorer efficiently traces compact sentence-level circuits from informative reasoning tokens, constructs internal and external reasoning graphs, and measures their discrepancy using Fused Gromov--Wasserstein distance. Experiments on four datasets from FaithCoT-Bench show that CIE-Scorer achieves state-of-the-art performance while reducing the cost of circuit construction, demonstrating the effectiveness of combining mechanistic interpretability signals with external reasoning traces for CoT unfaithfulness detection.

3 Citations
0 Influential
3.5 Altmetric
20.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!