2602.09063v1 Feb 09, 2026 q-bio.GN

scBench: 단일 세포 RNA 시퀀싱 분석에 대한 AI 에이전트 평가

scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis

Kenny Workman
Kenny Workman
Citations: 5
h-index: 2
Zhen Yang
Zhen Yang
Citations: 14
h-index: 3
H. Muralidharan
H. Muralidharan
Citations: 99
h-index: 6
Aidan Abdulali
Aidan Abdulali
Citations: 2
h-index: 1
Hannah Le
Hannah Le
Citations: 5
h-index: 2

단일 세포 RNA 시퀀싱 데이터셋이 채택, 규모, 복잡성 측면에서 증가함에 따라, 데이터 분석은 많은 연구 그룹에게 여전히 병목 현상입니다. 최첨단 AI 에이전트가 소프트웨어 엔지니어링 및 일반 데이터 분석 분야에서 괄목할 만한 발전을 이루었지만, 이러한 에이전트가 실제 복잡한 단일 세포 데이터셋에서 생물학적 통찰력을 추출할 수 있는지 여부는 불분명합니다. 본 연구에서는 scBench를 소개합니다. scBench는 6개의 시퀀싱 플랫폼과 7가지 작업 범주에 걸친 실제 scRNA-seq 워크플로우에서 파생된 394개의 검증 가능한 문제들로 구성된 벤치마크입니다. 각 문제는 분석 단계 직전의 실험 데이터 스냅샷과 핵심 생물학적 결과를 평가하는 결정론적 평가 도구를 제공합니다. 8개의 최첨단 모델에 대한 벤치마크 데이터는 정확도가 29%에서 53% 사이임을 보여주며, 모델-작업 및 모델-플랫폼 간에 상당한 상호 작용이 존재합니다. 플랫폼 선택은 모델 선택만큼 정확도에 영향을 미치며, 문서화가 덜 된 기술의 경우 정확도가 40% 이상 감소하는 경향이 있습니다. scBench는 SpatialBench와 함께 두 가지 주요 단일 세포 분석 방식을 포괄하며, 실제 scRNA-seq 데이터셋을 정확하고 재현 가능하게 분석할 수 있는 에이전트를 개발하기 위한 측정 도구이자 진단 도구 역할을 수행합니다.

Original Abstract

As single-cell RNA sequencing datasets grow in adoption, scale, and complexity, data analysis remains a bottleneck for many research groups. Although frontier AI agents have improved dramatically at software engineering and general data analysis, it remains unclear whether they can extract biological insight from messy, real-world single-cell datasets. We introduce scBench, a benchmark of 394 verifiable problems derived from practical scRNA-seq workflows spanning six sequencing platforms and seven task categories. Each problem provides a snapshot of experimental data immediately prior to an analysis step and a deterministic grader that evaluates recovery of a key biological result. Benchmark data on eight frontier models shows that accuracy ranges from 29-53%, with strong model-task and model-platform interactions. Platform choice affects accuracy as much as model choice, with 40+ percentage point drops on less-documented technologies. scBench complements SpatialBench to cover the two dominant single-cell modalities, serving both as a measurement tool and a diagnostic lens for developing agents that can analyze real scRNA-seq datasets faithfully and reproducibly.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!