2605.26893v1 May 26, 2026 cs.CL

GeoFaith: 신뢰성 있는 연쇄적 사고 과정의 시공간적 이중 관점

GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought

Weijian Lv
Weijian Lv
Citations: 68
h-index: 2
Xiaobo Xia
Xiaobo Xia
Citations: 348
h-index: 7
Wen Zhao
Wen Zhao
Citations: 1
h-index: 1
Jiayu Wang
Jiayu Wang
Citations: 54
h-index: 2
Yuhao Wu
Yuhao Wu
Citations: 1
h-index: 1
Jiaheng Wei
Jiaheng Wei
Citations: 22
h-index: 2

연쇄적 사고(Chain-of-Thought, CoT) 추론은 대규모 언어 모델(LLM) 발전에 기여했지만, 결과 기반 감독 방식은 널리 퍼진 사후 정당화 현상을 야기하여, 그럴듯하지만 신뢰성이 떨어지는 연쇄적 추론 과정을 생성합니다. 기존의 대부분 신뢰성 평가 방법은 확장성이 낮거나, 비용이 많이 들거나, 또는 신뢰할 수 없는 문제가 있었습니다. 본 연구에서는 잠재적인 기하학적 구조와 엔트로피 동역학을 활용하여 신뢰성 있는 추론을 진단하고 강화하는 시공간 프레임워크인 GeoFaith를 제안합니다. 우리는 1천 개에서 2만 개로 확장되는 단계별 주석을 포함하는 확장 가능한 부트스트랩 파이프라인을 개발하고, 표준 벤치마크에서 GPT-5보다 뛰어난 성능을 보이는 80억 개의 매개변수를 가진 신뢰성 검출기를 학습했으며, 결과 정확도, 과정의 신뢰성 및 경로 일관성을 동시에 최적화하는 신뢰성을 고려한 강화 학습 프레임워크를 설계했습니다. 실험 결과, 제안된 방법은 신뢰성 검출과 후속 추론 모두에서 우수한 성능을 달성하며, 정확도를 희생하지 않고 더 짧고 해석 가능성이 높은 연쇄를 생성합니다. 본 연구의 코드는 공개될 예정입니다.

Original Abstract

Chain-of-Thought (CoT) reasoning has advanced large language models (LLMs), but outcome-based supervision leads to pervasive post-hoc rationalization, producing plausible yet unfaithful reasoning chains. Most prior faithfulness assessment methods are either unscalable, expensive, or unreliable. We propose GeoFaith, a spatio-temporal framework that leverages latent geometric structure and entropy dynamics to diagnose and enforce faithful reasoning. We develop a scalable bootstrapping pipeline expanding step-level annotations from 1k to 20k samples across four domains, train an 8B faithfulness detector outperforming GPT-5 on standard benchmarks, and design a faithfulness-aware reinforcement learning framework jointly optimizing outcome correctness, process faithfulness, and trajectory consistency. Experiments show the proposed method achieves superior performance on both faithfulness detection and downstream reasoning, producing shorter, more interpretable chains without sacrificing accuracy. Our code will be made available publicly.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!