2608.03501v1 Aug 04, 2026 cs.AI

LLM은 고품질 실험을 설계할 수 있는가? 자율적 실험 설계에 대한 종합적이고 체계적인 성능 평가

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Renhe Jiang
Renhe Jiang
Citations: 391
h-index: 10
Ru Peng
Ru Peng
Zhejiang University
Citations: 104
h-index: 6

AI for Research (AI4Research)는 AI를 활용하여 과학 연구 워크플로우를 자동화하고 개선하는 것을 목표로 합니다. 기존 연구에서 실험 설계는 중요한 단계임에도 불구하고, 주로 코드 구현 및 실행에 초점이 맞춰져 왔으며, 이 단계의 중요성이 간과되어 왔습니다. 또한, AI가 체계적인 실험 설계를 수행할 수 있는 능력을 평가할 수 있는 벤치마크는 존재하지 않았습니다. 이러한 격차를 해소하기 위해, 우리는 최고 수준의 학술지(예: ICML, NeurIPS, ICLR)에서 발표된 19개 연구 분야에 걸쳐 엄선된 300편의 최신 논문을 기반으로 구성된 종합적인 실험 설계 평가 벤치마크인 SCOPE를 제안합니다. SCOPE는 LLM을 다음 두 가지 측면에서 평가합니다. 첫째, 고수준 계획의 완전성(주요 실험, 변형 실험, 분석 실험); 둘째, 저수준 설정의 정확성과 합리성(데이터셋, 기준 모델, 지표). 벤치마킹 결과, 세 가지 주요 결과를 얻었습니다: (1) 대부분의 LLM은 직접적으로 고품질 실험을 설계할 수 없습니다; (2) 모든 LLM은 저수준 설정 단계에서 성능 병목 현상을 보입니다; (3) 검색 모드는 실험 설계 품질을 향상시키지 못합니다. 이러한 문제점을 해결하기 위해, 우리는 LLM 기반의 실험 설계를 최적화하는 새로운 에이전트 기반 워크플로우인 OptED를 제안합니다. OptED는 단계별 분리, 도구 보강, 규칙 기반 제약 등을 통해 LLM 기반의 실험 계획을 향상시켜 저수준 설정 병목 현상을 효과적으로 완화합니다.

Original Abstract

AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design. To bridge this gap, we propose SCOPE, a Scientific COmprehensive Planning Evaluation Benchmark constructed from 300 high-quality latest papers across 19 research domains from top-tier venues (e.g., ICML, NeurIPS, and ICLR),evaluating LLMs on two dimensions: High-Level planning completeness (main, ablation, and analysis experiments) and Low-Level configuration accuracy and rationality (datasets, baselines, and metrics). Benchmarking reveals three findings: (1) most LLMs cannot directly design high-quality experiments; (2) all LLMs exhibit a performance bottleneck in low-level configuration; and (3) search mode does not improve design quality. Furthermore, to address these challenges, we propose OptED, a novel agentic workflow to optimize LLM-based experimental design, that enhances LLM-based experimental planning through stage isolation, tool augmentation, and rule-based constraints, effectively alleviating the configuration bottleneck.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!