2605.28065v1 May 27, 2026 cs.AI

장기 예측 공간 생물학의 검증 가능한 성능 평가

Verifiable Benchmarking of Long-Horizon Spatial Biology

Kenny Workman
Kenny Workman
Citations: 5
h-index: 2
H. Muralidharan
H. Muralidharan
Citations: 99
h-index: 6
Ian F. Diks
Ian F. Diks
Citations: 10
h-index: 2
Timothy Proctor
Timothy Proctor
Citations: 11
h-index: 2

인공지능 에이전트는 생물학 데이터 분석에 점점 더 유용하게 활용되고 있지만, 기존의 성능 평가는 주로 광범위한 생물학적 지식, 실행 가능한 워크플로우 또는 특정 분석 단계만을 테스트하는 경향이 있습니다. 본 연구에서는 공간 정보를 활용한 장기 예측 생물학 분석을 위한 벤치마크인 SpatialBench-Long을 소개합니다. SpatialBench-Long은 에이전트가 지정된 방법 없이 원시 데이터 또는 준원시 데이터를 기반으로 실험적 맥락과 함께 생물학적 가설을 추론하도록 설계되었습니다. 이 벤치마크는 주요 췌장 선암(PDAC), 인공 글리오블라스 오가노이드 및 생체 종양, Cas9 유전자 편집 기술을 이용한 폐 선암 모델, 그리고 마우스 시신경 노화/개입 시스템 등 다양한 데이터를 포함합니다. 데이터에는 CosMx, Visium, Xenium, 다중 오류 방지 형광 현미경(MERFISH), 단일 세포 RNA 시퀀싱(scRNA-seq), Slide-seq, Slide-tags, 조직학 및 계통 발생 기록 데이터가 포함됩니다. 제시된 가설은 재현성 검증, 독립적인 과학자 평가 및 경로 분석을 통해 검증됩니다. 최종 답변은 정해진 어휘와 기호를 사용하여 객관적으로 평가되며, 주요 분석 단계별 진행 상황을 반영하는 평가 기준이 함께 제공됩니다. SpatialBench-Long 벤치마크 결과, Gemini 3.5 Flash / Pi terminal coding harness, GPT-5.5 / Pi, 그리고 GPT-5.5 / OpenAI Codex 모델-시스템 조합이 각각 8/72회(11.1%)로 동률을 기록했습니다. SpatialBench-Long은 에이전트가 절차적 분석 실행 단계를 넘어 복잡한 공간 정보로부터 정확한 과학적 결론을 도출할 수 있는지 평가하는 데 목적을 두고 있습니다.

Original Abstract

AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps rather than end-to-end scientific reasoning over spatial measurements. We introduce SpatialBench-Long, a benchmark for long-horizon spatial biology in which agents must recover biological claims from raw or near-raw data and calibrated experimental context without prescribed methods. SpatialBench-Long contains 24 evaluations across primary pancreatic ductal adenocarcinoma (PDAC), engineered glioblastoma organoids and in vivo tumors, Cas9 lineage-traced lung adenocarcinoma, and mouse optic nerve aging/intervention systems, spanning CosMx, Visium, Xenium, multiplexed error-robust fluorescence in situ hybridization (MERFISH), single-cell RNA sequencing (scRNA-seq), Slide-seq, Slide-tags, histology, and lineage-recording data. Candidate claims are hardened through reproduction, independent scientist review, and trajectory inspection. Final answers are graded deterministically over controlled vocabularies and symbols with companion rubrics capturing progress through key analysis chokepoints. Across the SpatialBench-Long benchmark, three model-harness pairs tie at 8/72 runs (11.1\%): Gemini 3.5 Flash / Pi terminal coding harness, GPT-5.5 / Pi, and GPT-5.5 / OpenAI Codex. SpatialBench-Long tests whether agents can move beyond executing procedural analysis to deriving accurate scientific conclusions from complex spatial measurements.

3 Citations
0 Influential
3 Altmetric
18.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!