2606.13602v1 Jun 11, 2026 cs.AI

EpiBench: 에피지노믹스 분석 AI 에이전트의 검증 가능한 평가

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

Kenny Workman
Kenny Workman
Citations: 5
h-index: 2
H. Muralidharan
H. Muralidharan
Citations: 99
h-index: 6
Timothy Proctor
Timothy Proctor
Citations: 11
h-index: 2
Reema Baskar
Reema Baskar
Citations: 570
h-index: 9
Soo-Yeon Lee
Soo-Yeon Lee
Citations: 17
h-index: 2

본 연구에서는 단기 에피지노믹스 분석을 위한 검증 가능한 벤치마크인 EpiBench를 소개합니다. EpiBench는 에이전트가 현실적인 작업 흐름 상태에서 명확하게 정의된 분석 결정을 내리고, 결정적으로 평가 가능한 답변을 제공할 수 있는지 평가합니다. 이 벤치마크는 CUT&Tag/CUT&RUN, ATAC-seq, ChIP-seq 및 DNA 메틸화 워크플로우를 포함한 총 106개의 평가 항목으로 구성됩니다. 16개의 모델-프레임워크 조합에서 도출된 5,088개의 유효한 실행 경로에 대해, 어떤 시스템도 대부분의 시도에서 성공하지 못했습니다. GPT-5.5 / Pi가 45.0% (318번 시도 중 143회 성공; 95% 신뢰 구간: 36.3--53.7)로 가장 높은 성공률을 보였으며, 그 뒤를 GPT-5.5 / OpenAI Codex가 39.9% (318번 시도 중 127회 성공; 95% 신뢰 구간: 31.6--48.3)로 이어졌습니다. Claude Opus 4.8 Max / Pi와 GPT-5.4 / Pi는 각각 39.0% (318번 시도 중 124회 성공; 95% 신뢰 구간: 30.2--47.8 및 31.0--47.0)의 성공률을 기록했습니다. 성능은 분석 유형에 따라 다르게 나타났으며, 많은 실패한 실행 결과에서도 정답의 일부가 포함되어 있었습니다. 에이전트는 종종 올바른 파일을 찾고 유용한 중간 결과를 계산했지만, 과제가 더 깊은 수준의 분석이나 특정 분석 분야의 과학적 판단을 요구할 때 실패하는 경우가 많았습니다.

Original Abstract

We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of attempts: GPT-5.5 / Pi led at 45.0\% (143/318 attempts; 95\% confidence interval (CI), 36.3--53.7), followed by GPT-5.5 / OpenAI Codex at 39.9\% (127/318 attempts; 95\% CI, 31.6--48.3). Claude Opus 4.8 Max / Pi and GPT-5.4 / Pi each passed 39.0\% (124/318 attempts; 95\% CI, 30.2--47.8 and 31.0--47.0, respectively). Performance varies across assay types, and many failed runs still contain parts of the correct answer. Agents often found the right files and computed useful intermediate results, but failed when the task required deeper, assay-specific scientific judgment.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!