2606.18874v1 Jun 17, 2026 cs.AI

연구 하니스를 활용한 인공지능 과학자의 연구 종합 및 검증: 외부화 접근 방식

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

Zichen Zhu
Zichen Zhu
Citations: 411
h-index: 11
Da Ma
Da Ma
Citations: 469
h-index: 10
Lu Chen
Lu Chen
Citations: 460
h-index: 12
Kai Yu
Kai Yu
Citations: 542
h-index: 13
Jing Peng
Jing Peng
Citations: 127
h-index: 4
Xin Chen
Xin Chen
Citations: 6
h-index: 1
Hanqi Li
Hanqi Li
Citations: 44
h-index: 4
Yunzhe Zhang
Yunzhe Zhang
Citations: 41
h-index: 2
Ziyue Yang
Ziyue Yang
Citations: 23
h-index: 2
Zijian Hu
Zijian Hu
Scale AI
Citations: 345
h-index: 6
Tiancheng Huang
Tiancheng Huang
Citations: 50
h-index: 4
Chenrun Wang
Chenrun Wang
Citations: 113
h-index: 1
Zijian Wang
Zijian Wang
Citations: 4,581
h-index: 12
S. Zuo
S. Zuo
Citations: 5
h-index: 1
Danyu Luo
Danyu Luo
Citations: 8
h-index: 2
Sijia Guo
Sijia Guo
Citations: 42
h-index: 2
Huayang Wang
Huayang Wang
Citations: 16
h-index: 2
Senyu Han
Senyu Han
Citations: 71
h-index: 3
Yilu Cao
Yilu Cao
Citations: 1
h-index: 1
Bo Chen
Bo Chen
Citations: 113
h-index: 5

인공지능 시스템은 점점 더 많은 과학적 워크플로우를 자동화할 수 있지만, 이전 증거, 생성된 아이디어, 실험, 그리고 최종 주장을 연결하는 추론 과정은 종종 모델 추론 내부에 암묵적으로 남아 있습니다. 본 연구에서는 Xcientist라는 연구 하니스를 소개합니다. Xcientist는 연구 종합 및 실험 검증 과정을 외부화하여 검토 가능하고 계약 기반의 프로세스로 만듭니다. Xcientist는 문헌 증거, 아이디어 상태, 구현 계획, ablation 기록 및 수정 추적 정보를 지속적인 연구 결과물로 구성하여 생성된 메커니즘이 그 근거를 잃지 않고 실행, 테스트 및 수정될 수 있도록 합니다. 우리는 자동화된 연구의 실패 모드 중 하나인 '주장 변화(claim drift)' 현상을 분석했습니다. 이 현상은 실행 가능한 결과물이 원래 주장했던 메커니즘을 더 이상 뒷받침하지 못하는 경우를 의미합니다. Xcientist는 학습이 필요 없는 메모리 시스템, 그래프 기반 교통 예측 모델 및 다중 스케일 물리 기반 신경망에 적용하여 문제 정의부터 메커니즘 설계, 검증 및 제한적인 수정까지 추적 가능한 경로를 유지합니다. 이러한 결과는 인공지능 과학자를 최종 결과물뿐만 아니라, 그들의 종합 및 검증 과정이 책임감 있게 설명 가능하고 검토 가능하며 과학적으로 타당한 방식으로 수행되는지를 기준으로 평가해야 함을 시사합니다.

Original Abstract

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, idea states, implementation plans, ablation records and repair traces as persistent research artifacts, so that generated mechanisms can be grounded, executed, tested and revised without losing their evidential basis. We identify claim drift as a failure mode of automated research, where runnable artifacts no longer support the mechanism originally claimed. Across training-free memory systems, graph-structured traffic forecasting and multi-scale physics-informed neural networks, Xcientist preserves traceable trajectories from problem formulation to mechanism design, validation and bounded revision. These results suggest that AI scientists should be evaluated not only by their final artifacts, but by whether their synthesis and validation processes remain attributable, inspectable and scientifically accountable.

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!