2607.19678v1 Jul 22, 2026 cs.CL

개방형 질문 응답에서 추론 능력에 대한 참조 불필요 평가: 논리적 근거 기반 접근 방식

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

Yuxiang Zhou
Yuxiang Zhou
Citations: 120
h-index: 7
Guneet Singh Kohli
Guneet Singh Kohli
Citations: 17
h-index: 3
M. Schlichtkrull
M. Schlichtkrull
Citations: 8,227
h-index: 16
Gregory E Dean
Gregory E Dean
Citations: 0
h-index: 0
Maria Liakata
Maria Liakata
Citations: 0
h-index: 0

고위험 분야에서 생성된 AI 답변은 종종 유창하지만, 특히 단일 최종 답변이 아닌 다단계 추론을 포함하는 경우 검증하기 어렵습니다. 본 연구에서는 LLM(대규모 언어 모델)이 생성한 결과물을 감사하기 위한 참조 불필요 프레임워크를 제안하며, 이는 추론에 기반합니다. 이 방법은 생성된 추론 과정을 세그먼트로 분해하고, 자연어 추론(NLI)을 사용하여 각 세그먼트 내의 지역적인 전제-목표 관계를 라벨링하여 이러한 관계를 하이퍼 그래프로 구성합니다. 결정적이며 역방향 AND-OR 검색 방식을 통해 각 세그먼트가 생성된 답변 내에서 어떻게 근거화되는지를 나타내는 세그먼트 수준의 감사 레이블을 할당합니다. 본 프레임워크는 두 가지 환경에서 평가되었습니다. 첫째, Hard2Verify를 사용한 귀납적 수학적 추론이고, 둘째는 새로운 의사 기반 벤치마크인 UroReason을 사용한 개방형 의료 분야 추론입니다. 이러한 환경에서, 우리의 NLI-하이퍼 그래프 감사 방식은 LLM 자체를 판별기로 사용하는 기존 방법보다 더 신뢰할 수 있는 참조 불필요 평가 지표를 제공합니다. 특히 임상 환경에서는 최첨단 LLM 판단기가 문제가 있는 추론 세그먼트를 식별하는 데 자주 실패하여, 유창하지만 근거가 약한 답변을 과도하게 허용하는 경향이 있습니다. 우리의 결과는 질문 응답 평가가 최종 답변이나 LLM 검증기에만 의존하는 것이 아니라, 추론 과정 전반에 걸쳐 추론적 관계가 어떻게 구성되는지를 고려해야 함을 보여줍니다. UroReason은 API를 통해 제공될 예정이며, 저희 코드는 오픈 소스로 공개될 것입니다.

Original Abstract

AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a hypergraph. A deterministic backward AND-OR search then assigns segment-level audit labels that indicate how each segment is grounded within the generated response. We evaluate the framework in two settings: deductive mathematical reasoning with Hard2Verify, and open-ended medical reasoning with UroReason, a new physician-annotated benchmark of LLM reasoning traces from real clinical cases. Across these settings, our NLI-hypergraph audit provides a more reliable reference-free evaluation signal than direct LLM-as-judge baselines. In the clinical setting, state-of-the-art LLM judges often fail to identify problematic reasoning segments, over-accepting fluent but weakly grounded responses. Our results show that QA evaluation should account for how inferential relations compose across a reasoning trace, rather than relying only on final answers or LLMs as verifiers. UroReason will be made available through an API, and our code will be released as open source.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!