2606.26563v1 Jun 25, 2026 q-bio.GN

scBench-Long: 장기 예측을 위한 단일 세포 생물학의 검증 가능한 벤치마킹

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

Kenny Workman
Kenny Workman
Citations: 5
h-index: 2
Zhenning Yang
Zhenning Yang
University of Michigan
Citations: 163
h-index: 5
Ian F. Diks
Ian F. Diks
Citations: 10
h-index: 2
Timothy Proctor
Timothy Proctor
Citations: 11
h-index: 2
Arjun Banerjee
Arjun Banerjee
Citations: 1
h-index: 1

단일 세포 연구는 분석가들이 다양한 단계의 워크플로우와 메타데이터, 실험 환경 및 보조 증거를 통합하여 원시 데이터를 특정 생물학적 결론으로 변환하도록 요구합니다. 기존의 AI-생물학 벤치마크는 주로 광범위한 지식, 실행 가능한 워크플로우 또는 로컬 분석 단계를 측정합니다. 본 연구에서는 scBench-Long을 소개하며, 이는 에이전트가 미리 정해진 방법 없이 원시 데이터 또는 거의 원시 상태의 데이터로부터 과학적 결론을 도출해야 하는 장기 예측 단일 세포 생물학 벤치마크입니다. 이 벤치마크는 흑색종 CD8 T 세포 반응성, CD8 RNA+ATAC 조절 추론, 인간-원숭이 키메라 발달, KRAS 유발 폐 종양 노화, 그리고 치명적인 COVID-19 폐 병변을 포함한 21개의 평가 항목으로 구성됩니다. 작업에는 paired scRNA/TCR 시퀀싱, RNA 및 염색질 프로파일링, 종 간 트랜스크립토믹스, 조합형 scRNA-seq, 단일 핵 RNA-seq, 면역 체계, 정사상 지도, 리간드-수용체 자원 및 검증 증거가 포함됩니다. 후보 주장은 재현되고 검토되어 결정적인 평가 기준과 경향성 척도를 가진 통제된 답변 어휘로 변환됩니다. 1,068개의 완료된 경로에서 가장 강력한 모델-활성화 쌍은 63회 실행 중 16회(25.4%)를 성공적으로 수행했습니다. scBench-Long은 에이전트가 로컬 분석 단계를 넘어 복잡한 과학적 주장을 얼마나 잘 할 수 있는지, 그리고 이러한 주장들이 단일 세포 데이터에 의해 뒷받침되는지를 평가합니다.

Original Abstract

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence. Existing AI-biology benchmarks largely measure broad knowledge, executable workflows, or local analysis steps. We introduce scBench-Long, a benchmark for long-horizon single-cell biology in which agents must recover scientific conclusions from raw or near-raw data without prescribed methods. The benchmark contains 21 evaluations spanning melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, human--monkey chimera development, KRAS-driven lung tumor aging, and lethal COVID-19 lung pathology. Tasks cover paired scRNA/TCR sequencing, RNA and chromatin profiling, cross-species transcriptomics, combinatorial scRNA-seq, single-nucleus RNA-seq, immune repertoires, ortholog maps, ligand--receptor resources, and validation evidence. Candidate claims are reproduced, reviewed, and converted into controlled answer vocabularies with deterministic grading and trajectory rubrics. Across 1,068 completed trajectories, the strongest model--harness pair passes 16/63 runs (25.4\%). scBench-Long evaluates whether agents can move beyond local analysis steps and make complex scientific claims that are supported by single-cell data.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!