VeryTrace: 컴파일 가능한 형식과 구조화된 검증을 통한 추론 과정 검증
VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification
체인 오브 소트(Chain-of-Thought, CoT) 프롬프팅을 사용한 다단계 추론은 여전히 취약합니다. 초기 단계에서 발생하는 논리적 오류나 환각 현상은 조용히 전파되어, 확신에 찬 동시에 부정확한 결론을 야기합니다. 본 논문에서는 VeryTrace라는 검증 및 수정 프레임워크를 제안하며, 이는 자연어 추론 과정을 구조화되고 컴파일 가능한 형태로 표현합니다. VeryTrace는 다음과 같은 특징을 가진 도메인 특화 언어(DSL)를 도입합니다: (i) 단계 간의 의존성을 명시적으로 나타내고, (ii) 양적 내용을 실행 가능한 표현으로 변환하며, (iii) 추론 과정을 연역 스키마를 통해 구조화합니다. VeryTrace는 계산 정확성, 의존성 해결 및 제약 조건 만족에 대한 결정적인 검사와 함께, 기계적으로 처리하기 어려운 의미적 판단을 위한 LLM 감사를 결합하는 하이브리드 검증기를 사용하며, 이를 통해 단계별 오류 위치 파악 및 수정을 가능하게 합니다. 세 가지 다양한 도메인(수학 경시대회 문제 - AIME 2025, 로봇 계획 - LLM-BabyBench, 친족 관계 추론 - CLUTRR)에서 VeryTrace는 최첨단 LLM 모델에 대한 제로샷 기준 성능을 향상시키며, 특정 도메인에 대한 추가 학습이나 컨텍스트 예제 없이도 정확성과 일반화 능력을 달성함을 보여줍니다.
Multi-step reasoning with Chain-of-Thought (CoT) prompting remains fragile: logical errors or hallucinations in early steps silently propagate, producing confident but incorrect conclusions. This paper presents VeryTrace, a zero-shot verification-and-repair framework that formalizes natural-language reasoning traces into a structured, compilable representation. VeryTrace introduces a Domain-Specific Language (DSL) that (i) makes step dependencies explicit, (ii) mechanizes quantitative content as executable expressions, and (iii) structures semantic inferences via deduction schemas. Our hybrid verifier combines deterministic checks for computational correctness, dependency resolution, and constraint satisfaction with targeted LLM audits for non-mechanizable semantic judgments, enabling step-level error localization and repair. Across three diverse domains-competition mathematics (AIME 2025), robotics planning (LLM-BabyBench), and kinship reasoning (CLUTRR), VeryTrace improves accuracy over zero-shot baselines on state-of-the-art LLMs without requiring domain-specific training or in-context examples, demonstrating that formalized trace verification achieves both precision and generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.