DAGverse: 과학 논문으로부터 문서 기반 의미론적 DAG 구축
DAGverse: Building Document-Grounded Semantic DAGs from Scientific Papers
방향성 비순환 그래프(DAG)는 과학 및 기술 분야에서 구조화된 지식을 표현하는 데 널리 사용됩니다. 그러나 실제 DAG에 대한 데이터셋은 도메인 문서에 대한 전문가의 해석이 일반적으로 필요하기 때문에 여전히 부족합니다. 본 연구에서는 Doc2SemDAG 구축 문제를 다룹니다. 이는 문서와 함께, 이를 설명하는 인용 증거 및 맥락을 기반으로 선호하는 의미론적 DAG를 복원하는 문제입니다. 이 문제는 문서가 여러 개의 가능한 추상화를 가질 수 있고, 의도된 구조가 종종 암시되어 있으며, 뒷받침하는 증거가 텍스트, 방정식, 캡션 및 그림에 흩어져 있기 때문에 어렵습니다. 이러한 문제점을 해결하기 위해, 명시적인 DAG 그림을 포함하는 과학 논문을 자연스러운 지도 학습 데이터 소스로 활용합니다. 이 설정에서, DAG 그림은 DAG 구조를 제공하고, 동반 텍스트는 맥락과 설명을 제공합니다. 본 연구에서는 온라인 과학 논문에서 문서 기반 의미론적 DAG를 구축하기 위한 프레임워크인 DAGverse를 소개합니다. DAGverse의 핵심 구성 요소인 DAGverse-Pipeline은 그림 분류, 그래프 재구성, 의미론적 연결 및 검증을 통해 고정밀의 의미론적 DAG 예제를 생성하도록 설계된 준자동 시스템입니다. 사례 연구로서, 본 프레임워크를 인과적 DAG에 적용하고, 그래프 수준, 노드 수준 및 엣지 수준의 증거를 포함하는 108개의 전문가 검증된 의미론적 DAG 데이터셋인 DAGverse-1을 공개합니다. 실험 결과, DAGverse-Pipeline은 DAG 분류 및 어노테이션에서 기존의 Vision-Language 모델보다 우수한 성능을 보였습니다. DAGverse는 문서 기반 DAG 벤치마킹을 위한 기반을 제공하며, 실제 증거에 근거한 구조화된 추론 연구에 대한 새로운 방향을 제시합니다.
Directed Acyclic Graphs (DAGs) are widely used to represent structured knowledge in scientific and technical domains. However, datasets for real-world DAGs remain scarce because constructing them typically requires expert interpretation of domain documents. We study Doc2SemDAG construction: recovering a preferred semantic DAG from a document together with the cited evidence and context that explain it. This problem is challenging because a document may admit multiple plausible abstractions, the intended structure is often implicit, and the supporting evidence is scattered across prose, equations, captions, and figures. To address these challenges, we leverage scientific papers containing explicit DAG figures as a natural source of supervision. In this setting, the DAG figure provides the DAG structure, while the accompanying text provides context and explanation. We introduce DAGverse, a framework for constructing document-grounded semantic DAGs from online scientific papers. Its core component, DAGverse-Pipeline, is a semi-automatic system designed to produce high-precision semantic DAG examples through figure classification, graph reconstruction, semantic grounding, and validation. As a case study, we test the framework for causal DAGs and release DAGverse-1, a dataset of 108 expert-validated semantic DAGs with graph-level, node-level, and edge-level evidence. Experiments show that DAGverse-Pipeline outperforms existing Vision-Language Models on DAG classification and annotation. DAGverse provides a foundation for document-grounded DAG benchmarks and opens new directions for studying structured reasoning grounded in real-world evidence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.