EpiBench: 에피지노믹스 분석 AI 에이전트의 검증 가능한 평가
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis
본 연구에서는 단기 에피지노믹스 분석을 위한 검증 가능한 벤치마크인 EpiBench를 소개합니다. EpiBench는 에이전트가 현실적인 작업 흐름 상태에서 명확하게 정의된 분석 결정을 내리고, 결정적으로 평가 가능한 답변을 제공할 수 있는지 평가합니다. 이 벤치마크는 CUT&Tag/CUT&RUN, ATAC-seq, ChIP-seq 및 DNA 메틸화 워크플로우를 포함한 총 106개의 평가 항목으로 구성됩니다. 16개의 모델-프레임워크 조합에서 도출된 5,088개의 유효한 실행 경로에 대해, 어떤 시스템도 대부분의 시도에서 성공하지 못했습니다. GPT-5.5 / Pi가 45.0% (318번 시도 중 143회 성공; 95% 신뢰 구간: 36.3--53.7)로 가장 높은 성공률을 보였으며, 그 뒤를 GPT-5.5 / OpenAI Codex가 39.9% (318번 시도 중 127회 성공; 95% 신뢰 구간: 31.6--48.3)로 이어졌습니다. Claude Opus 4.8 Max / Pi와 GPT-5.4 / Pi는 각각 39.0% (318번 시도 중 124회 성공; 95% 신뢰 구간: 30.2--47.8 및 31.0--47.0)의 성공률을 기록했습니다. 성능은 분석 유형에 따라 다르게 나타났으며, 많은 실패한 실행 결과에서도 정답의 일부가 포함되어 있었습니다. 에이전트는 종종 올바른 파일을 찾고 유용한 중간 결과를 계산했지만, 과제가 더 깊은 수준의 분석이나 특정 분석 분야의 과학적 판단을 요구할 때 실패하는 경우가 많았습니다.
We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of attempts: GPT-5.5 / Pi led at 45.0\% (143/318 attempts; 95\% confidence interval (CI), 36.3--53.7), followed by GPT-5.5 / OpenAI Codex at 39.9\% (127/318 attempts; 95\% CI, 31.6--48.3). Claude Opus 4.8 Max / Pi and GPT-5.4 / Pi each passed 39.0\% (124/318 attempts; 95\% CI, 30.2--47.8 and 31.0--47.0, respectively). Performance varies across assay types, and many failed runs still contain parts of the correct answer. Agents often found the right files and computed useful intermediate results, but failed when the task required deeper, assay-specific scientific judgment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.