2608.03291v1 Aug 04, 2026 cs.LG

결정적인 단서: 체인 오브 소트(Chain-of-Thought) 역학을 활용하여 LLM의 추론 오류 탐지

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

Shashwat Sourav
Shashwat Sourav
Citations: 48
h-index: 4
Aishwarya H. Balwani
Aishwarya H. Balwani
Citations: 196
h-index: 5

체인 오브 소트(CoT) 추론은 대규모 언어 모델(LLM)의 성능을 향상시키는 동시에 모델의 추론 과정을 관찰할 수 있는 인터페이스를 제공합니다. 기존 연구에서는 구체적으로 명시된 CoT를 사용하여 추론의 정확성을 모니터링하는 경우가 많지만, 이러한 접근 방식은 주로 개별적인 중간 단계의 의미적 정확성 또는 일관성을 평가하며, 전체 추론 과정이 어떻게 전개되는지에 대한 측면은 상대적으로 간과됩니다. 결과적으로, 단일 오류 단계에 국한되지 않고 추론 경로 전체에 걸쳐 발생하는 오류는 아직 충분히 연구되지 않았습니다. 또한, 명시된 CoT가 모델의 내부적인 추론 과정을 정확하게 반영하지 않을 수도 있으므로, 개별 문장을 내부 계산의 문자 그대로의 기록으로 간주하지 않는 분석이 필요합니다. 본 연구에서는 가시적인 CoT의 역학적 특성이 그러한 의미적 충실성을 가정하지 않고도 성공적인 추론과 실패한 추론을 체계적으로 구별하는 데 활용될 수 있는지 질문합니다. 우리는 다양한 LLM 모델을 사용하여 변동 가능한 복잡도를 가진 검증 가능한 불리언 만족성(SAT) 문제에 대한 연구를 수행하여, 각 모델의 성능 한계 근처에서 통제된 비교를 가능하게 했습니다. CoT 문장을 추론 기능으로 분류한 결과, SAT 문제에서 조기에 검증이 실패하는 현상이 나타났습니다: 잘못된 추론 과정은 더 일찍 절(clause) 검사를 시작하고, 유사한 연산을 반복하며, 더 빨리 완료됩니다. UNSAT 문제에서는 모델들이 성급하게 잘못된 SAT 결론으로 이동하며, 후보 할당을 확인하지만, 구성된 사례에서 모순을 도출하지 않습니다. 그 후, 특정 증명 검색 프롬프트를 통해 Llama3-70B의 정확도를 13.3%에서 85%로 향상시켰으며, 이러한 오류의 84.6%를 수정했습니다. 이러한 결과는 성능 실패가 눈에 보이는 추론 구조의 분산되고 작업 의존적인 변화로 나타날 수 있으며, 모델의 내부 계산을 반영하는지 여부에 관계없이 CoT 역학을 통해 이러한 실패를 진단하고 수정할 수 있음을 보여줍니다.

Original Abstract

Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning correctness, however, largely evaluate the semantic correctness or consistency of individual intermediate steps, rather than how the reasoning process evolves across the trace. As a result, failures distributed across the reasoning trajectory, rather than those localized to a single incorrect step, remain comparatively underexplored. Furthermore, verbalized CoTs need not faithfully reflect the model's internal reasoning, motivating analyses that do not treat individual statements as literal accounts of internal computation. In this work, we therefore ask whether the dynamics of visible CoT can be leveraged to systematically distinguish successful from failed reasoning without assuming such semantic faithfulness. We study a range of LLMs on verifiable Boolean satisfiability tasks with variable complexity, enabling controlled comparisons near each model's capability frontier. Tagging CoT sentences by reasoning function reveals premature verification collapse on SAT problems: incorrect traces enter clause checking earlier, repeat similar operations, and finalize sooner. On UNSAT problems, models presumptuously move towards incorrect SAT conclusions, checking candidate assignments rather than deriving contradictions across constructed cases. Subsequently, a targeted proof-search prompt intervention raises Llama3-70B accuracy from 13.3% to 85%, correcting 84.6% of these errors. These results show that capability failures can manifest as distributed, task-dependent changes in the structure of visible reasoning, and that CoT dynamics agnostic to whether the verbalized trace reflects the model's internal computations can help diagnose and correct failures.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!