scBench-Long: 장기 예측을 위한 단일 세포 생물학의 검증 가능한 벤치마킹
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
단일 세포 연구는 분석가들이 다양한 단계의 워크플로우와 메타데이터, 실험 환경 및 보조 증거를 통합하여 원시 데이터를 특정 생물학적 결론으로 변환하도록 요구합니다. 기존의 AI-생물학 벤치마크는 주로 광범위한 지식, 실행 가능한 워크플로우 또는 로컬 분석 단계를 측정합니다. 본 연구에서는 scBench-Long을 소개하며, 이는 에이전트가 미리 정해진 방법 없이 원시 데이터 또는 거의 원시 상태의 데이터로부터 과학적 결론을 도출해야 하는 장기 예측 단일 세포 생물학 벤치마크입니다. 이 벤치마크는 흑색종 CD8 T 세포 반응성, CD8 RNA+ATAC 조절 추론, 인간-원숭이 키메라 발달, KRAS 유발 폐 종양 노화, 그리고 치명적인 COVID-19 폐 병변을 포함한 21개의 평가 항목으로 구성됩니다. 작업에는 paired scRNA/TCR 시퀀싱, RNA 및 염색질 프로파일링, 종 간 트랜스크립토믹스, 조합형 scRNA-seq, 단일 핵 RNA-seq, 면역 체계, 정사상 지도, 리간드-수용체 자원 및 검증 증거가 포함됩니다. 후보 주장은 재현되고 검토되어 결정적인 평가 기준과 경향성 척도를 가진 통제된 답변 어휘로 변환됩니다. 1,068개의 완료된 경로에서 가장 강력한 모델-활성화 쌍은 63회 실행 중 16회(25.4%)를 성공적으로 수행했습니다. scBench-Long은 에이전트가 로컬 분석 단계를 넘어 복잡한 과학적 주장을 얼마나 잘 할 수 있는지, 그리고 이러한 주장들이 단일 세포 데이터에 의해 뒷받침되는지를 평가합니다.
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence. Existing AI-biology benchmarks largely measure broad knowledge, executable workflows, or local analysis steps. We introduce scBench-Long, a benchmark for long-horizon single-cell biology in which agents must recover scientific conclusions from raw or near-raw data without prescribed methods. The benchmark contains 21 evaluations spanning melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, human--monkey chimera development, KRAS-driven lung tumor aging, and lethal COVID-19 lung pathology. Tasks cover paired scRNA/TCR sequencing, RNA and chromatin profiling, cross-species transcriptomics, combinatorial scRNA-seq, single-nucleus RNA-seq, immune repertoires, ortholog maps, ligand--receptor resources, and validation evidence. Candidate claims are reproduced, reviewed, and converted into controlled answer vocabularies with deterministic grading and trajectory rubrics. Across 1,068 completed trajectories, the strongest model--harness pair passes 16/63 runs (25.4\%). scBench-Long evaluates whether agents can move beyond local analysis steps and make complex scientific claims that are supported by single-cell data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.