2608.04738v1 Aug 05, 2026 cs.AI

EviGraph: 증거 기반의 자율 연구 에이전트

EviGraph: Evidence-Guided Autonomous Research Agents

Shuo Ren
Shuo Ren
Citations: 164
h-index: 5
Zhenjiang Ren
Zhenjiang Ren
Citations: 81
h-index: 1
Ziliang Pang
Ziliang Pang
Citations: 0
h-index: 0
Ruiji Li
Ruiji Li
Citations: 0
h-index: 0
Xujin Zhang
Xujin Zhang
Citations: 0
h-index: 0
Jiajun Zhang
Jiajun Zhang
Citations: 178
h-index: 5

자율 연구 에이전트는 가설을 생성하고, 실험을 수행하며, 논문을 작성할 수 있지만, 종종 근거 없는 주장과 연구 질문, 실험, 결과 및 결론 간의 불일치를 포함하는 결과를 생성합니다. 우리는 이러한 문제가 부분적으로 시스템 아키텍처에 기인한다고 주장합니다. 기존 시스템은 연구를 순차적 파이프라인으로 구성하지만, 연구 단계 전체에서 진화하는 주장-증거 구조를 명시적으로 유지하거나 검증하지 않습니다. 본 논문에서는 EviGraph라는 자율 연구 프레임워크를 소개합니다. EviGraph는 문제(Problem), 간극(Gap), 가설(Hypothesis), 실험(Experiment), 발견(Finding) 및 주장의 노드로 구성된 유형화된 증거 그래프로 연구 과정을 표현합니다. 이 그래프는 에이전트의 운영 상태 역할을 하며, 사후 기록으로 사용되지 않습니다. EviGraph는 누락된 종속성, 의미적 불일치 및 결과-주장 불일치를 포함한 증거 체인을 검사하고, 가장 취약한 노드를 찾아 해당 하위 그래프를 다시 생성합니다. 그래프 체크포인팅은 실패한 복구 시도가 이전에 검증된 증거를 손상시키지 않도록 합니다. 모든 유지된 주장이 검증된 증거 체인에 기반해야 하는 경우에만 논문이 생성됩니다. ARC-Bench-ML 및 NanoResearch-20 데이터셋에서의 실험 결과, EviGraph는 다른 엔드-투-엔드 연구 에이전트 기준 모델보다 전반적인 연구 성능이 우수하며, 가장 강력한 기준 모델 대비 주장 지지율을 40.19% 향상시키고, 실험 데이터 일관성을 87.73% 달성했습니다. 이러한 결과는 신뢰할 수 있는 자율 연구를 위한 명시적인 증거 상태 유지의 가치를 입증합니다.

Original Abstract

Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions. We argue that this problem is partly architectural: existing systems organize research as sequential pipelines but do not explicitly maintain or validate the evolving claim-evidence structure across stages. In this paper, we introduce EviGraph, an autonomous research framework that represents the research process as a typed evidence graph containing Problem, Gap, Hypothesis, Experiment, Finding, and Claim nodes. The graph serves as the operational state of the agent rather than a post-hoc record. EviGraph inspects evidence chains for missing dependencies, semantic misalignment, and result-claim inconsistencies, localizes the earliest weak node, and regenerates its affected downstream subgraph. Graph checkpointing prevents unsuccessful repairs from corrupting previously validated evidence. Manuscripts are generated only after every retained claim is grounded in a validated evidence chain. Experiments on ARC-Bench-ML and NanoResearch-20 show that EviGraph outperforms the compared end-to-end research-agent baselines in overall research performance, improves Claim Support Rate by 40.19% over the strongest baseline, and achieves 87.73% Experimental Data Consistency. These results demonstrate the value of explicit evidence-state maintenance for reliable autonomous research.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!