2602.18446v1 Jan 27, 2026 cs.CL

ReportLogic: 심층 연구 보고서의 논리적 품질 평가

ReportLogic: Evaluating Logical Quality in Deep Research Reports

Jujia Zhao
Jujia Zhao
Citations: 181
h-index: 6
Zhaoxin Huan
Zhaoxin Huan
Citations: 407
h-index: 10
Zihan Wang
Zihan Wang
Citations: 131
h-index: 2
Xiaolu Zhang
Xiaolu Zhang
Citations: 52
h-index: 4
Jun Zhou
Jun Zhou
Citations: 9
h-index: 2
Suzan Verberne
Suzan Verberne
Citations: 249
h-index: 6
Zhaochun Ren
Zhaochun Ren
Citations: 172
h-index: 5

사용자들은 점점 더 많은 양의 정보를 심층 연구에 활용하기 위해 대규모 언어 모델(LLM)에 의존하며, 이를 통해 다양한 출처를 구조화된 보고서로 통합하여 이해와 실행을 지원합니다. 이러한 보고서의 실질적인 신뢰성은 논리적 품질에 달려 있습니다. 즉, 보고서의 주장과 논리가 명시적으로 뒷받침되어야 하며, 단순히 유창하거나 유익해 보이는 것 이상으로, 하위 작업의 기반으로 신뢰할 수 있어야 합니다. 그러나 현재의 평가 프레임워크는 이러한 요구 사항을 대부분 간과합니다. 이러한 격차를 해소하기 위해, 우리는 ReportLogic이라는 벤치마크를 소개합니다. ReportLogic은 독자 중심의 감사 가능성 관점에서 보고서 수준의 논리적 품질을 정량화합니다. 구체적으로, ReportLogic은 계층적 분류 체계를 채택하여 독자들이 (1) 일관된 분석 흐름을 갖춘 주제에 맞는 보고서 구조를 파악할 수 있는지 (매크로 논리), (2) 필요한 맥락을 통해 진행 과정을 이해할 수 있는지 (설명적 논리), 그리고 (3) 명시적인 주장-근거 관계를 통해 결론을 검증할 수 있는지 (구조적 논리)를 평가합니다. 이 분류 체계를 기반으로, 우리는 인간이 주석을 달고 지침이 포함된 데이터 세트를 구축하고, 확장 가능한 평가를 위한 오픈 소스 LogicJudge를 훈련했습니다. 또한, 우리는 적대적 공격을 통해 평가자의 견고성을 평가했으며, 상용 LLM 평가기가 표면적인 단서(예: 장황함)에 자주 영향을 받고, 추론 방식이 잘못된 근거 관계를 가릴 수 있음을 보여주었습니다. 전반적으로, 우리의 결과는 보다 강력한 논리 평가기를 구축하고 LLM이 생성하는 보고서의 논리적 신뢰성을 향상시키는 데 필요한 실질적인 지침을 제공합니다.

Original Abstract

Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports that support understanding and action. In this context, the practical reliability of such reports hinges on logical quality: whether the report's claims and arguments are explicitly supported and can be trusted as a basis for downstream use, rather than merely appearing fluent or informative. However, current evaluation frameworks largely overlook this requirement. To bridge this gap, we introduce ReportLogic, a benchmark that quantifies report-level logical quality through a reader-centric lens of auditability. Specifically, ReportLogic adopts a hierarchical taxonomy that evaluates whether readers can (1) trace an on-topic report structure with a unified analytical arc (Macro-Logic), (2) understand the progression with necessary context (Expositional-Logic), and (3) verify conclusions via explicit claim--support (Structural-Logic). Based on this taxonomy, we construct a human-annotated rubric-guided dataset and train an open-source LogicJudge for scalable evaluation. We further evaluate judge robustness via adversarial attacks, showing that off-the-shelf LLM judges are frequently influenced by superficial cues (e.g., verbosity), and reasoning modes can mask broken support relations. Overall, our results provide actionable guidance for building more robust logic evaluators and improving the logical reliability of LLM-generated reports.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!