대규모 언어 모델에서 발생하는 잠재적인 추론 오류를 감지하기 위한 참조 불필요 평가 지표
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
수학적 연쇄 추론(Chain of Thought, CoT) 평가는 일반적으로 최종 답변이 주어진 정답과 일치하는지에 대한 여부로 단순화됩니다. 이는 정확한 결론을 도출하는 것과 유효한 과정을 생성하는 것을 혼동합니다. 잘못된 과정이라도 우연히 올바른 답에 도달할 수 있으며, 유효한 계산 후에도 오기 오류가 발생할 수 있습니다. 우리는 이러한 불일치를 '추론 답변 일관성 격차(Reasoning Answer Consistency Gap)'라고 부릅니다. 본 논문에서는 Reasoning Answer Faithfulness Score (RAFS)라는 참조 불필요 평가 지표를 소개합니다. RAFS는 모델이 생성한 수학적 과정을 개별적으로 분석하여, 해당 과정이 현지적으로 신뢰할 수 있는지, 답변을 뒷받침하는지, 그리고 재샘플링 및 의도적인 반사실적 개입에 대해 안정적인지를 진단합니다. RAFS는 단계의 유효성, 추론-답변 함축 관계, 반사실적 민감성, 답변 일관성, 조건부 추론 안정성을 종합적으로 평가합니다. RAFS는 모델의 사설 계산이나 테스트된 수학적 환경 외부의 사실 정확도를 평가하는 것이 아니라, 트랜스크립트 수준의 일치 여부를 평가합니다. 본 연구에서는 GSM8K 및 MATH 데이터셋에 대한 사전 등록된 결과를 검증하는 실험을 진행하며, 가설, 허용 규칙, 보정 방법 및 테스트는 결과가 확인되기 전에 미리 설정됩니다. 전체 실행 가능성을 확인하고 개입 범위를 추정하기 위한 별도의 파일럿 연구도 수행되며, 숫자 기반의 파일럿 주장은 추적 정보가 확보된 경우에만 보고됩니다. 본 논문에서는 네 가지 추론 답변 결과를 형식화하고, 비보상 집계 방식을 정당화하며, 의미적 추적 거리를 구체적으로 정의하고, 계산 및 회피(abstention) 사이의 균형을 정량화하며, 검증기의 독립성과 성능 분석을 정의합니다. RAFS는 수학적 답변의 정확도에 더하여 잠재적인 추론 오류와 답변 추출 오류를 감지하기 위한 감사 가능한 경고 신호를 제공하는 것을 목표로 합니다.
Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answer, while a valid calculation can be followed by a transcription error. We call this mismatch the reasoning answer consistency gap. This framework paper introduces the Reasoning Answer Faithfulness Score (RAFS), a reference free, instance level diagnostic of whether an emitted mathematical trace is locally credible, supports its answer, and is stable under resampling and targeted counterfactual interventions. RAFS combines step validity, reasoning to answer entailment and counterfactual sensitivity, answer consensus, and conditional reasoning stability. It evaluates transcript level agreement, not a models private computation and not factual correctness outside the tested mathematical setting. We retain a preregistered, results blind confirmatory study on GSM8K and MATH, with hypotheses, admissibility rules, calibration, and tests fixed before confirmatory outcomes are inspected. A separate feasibility pilot is specified to verify end to end execution and estimate interven tion coverage before that freeze numerical pilot claims are re ported only when trace level artifacts are available. We formalize four reasoning answer outcomes, justify the non compensatory aggregator, instantiate semantic trace distance, quantify compute and abstention tradeoffs, and define verifier independence and power analyses. RAFS is intended to complement mathematical answer accuracy with an auditable warning signal for silent reasoning failures and answer extraction errors
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.