2607.26155v1 Jul 28, 2026 cs.AI

ClinLens: 장기적인 시야를 가진 코딩 에이전트를 위한 다중 모드 임상 데이터 분석 벤치마크

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

Jindong Han
Jindong Han
Citations: 990
h-index: 19
Yuan Zhu
Yuan Zhu
Citations: 0
h-index: 0
Ethan B. Liu
Ethan B. Liu
Citations: 0
h-index: 0
Frank Nie
Frank Nie
Citations: 0
h-index: 0

임상 데이터 과학 에이전트는 다양한 형태의 연장 기록을 감사 가능한 분석으로 변환해야 하지만, 기존의 벤치마크는 주로 의료 질문 응답, 구조화된 테이블 추론 또는 일반적인 과학 저장소에 국한되어 있습니다. 본 논문에서는 구조화된 전자 건강 기록, 의무 기록, 심전도, 흉부 X선 촬영 및 심장 초음파 이미지 등 다섯 가지 연결된 MIMIC 리소스를 포괄하는 200개의 실행 가능한 작업으로 구성된 벤치마크인 CLINLENS를 소개합니다. 본 벤치마크는 환자-시간 범위의 네 가지 측면과 분석 능력의 다섯 가지 측면을 조합하여 분류됩니다. 프로그램 기반 역합성 방법론은 각 제한된 반정형 데이터 패키지에 대해 평가자가 보유하는 참조 워크플로우를 매칭하고, 필수적인 산출물, 코호트 및 시간적 의미, 그리고 최종 답변을 검증합니다. 고정된 126개의 작업 세트에서, 가장 성능이 좋은 24가지 표준화된 모델 구조는 100%의 실행 성공률(EXECSUCCESS)에도 불구하고 56.3%의 범위-전체 엄격 통과율(scope-macro STRICTPASS)을 달성합니다. 참고로, 별도로 구성된 코딩 에이전트는 126개의 작업 중 83개를 해결하며, GPT-4o-mini에 맞게 조정된 다섯 가지 생물 의학 시스템은 최대 2.9%의 범위-전체 엄격 통과율을 기록합니다. 이러한 결과는 실행 가능한 제출물과 정확한 임상 분석 간의 상당한 격차를 보여줍니다.

Original Abstract

Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories. We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms. A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities. Program-first reverse synthesis pairs each bounded semi-raw package with an evaluator-private reference workflow and checks required artifacts, cohort and temporal semantics, and the final answer. On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieves 56.3% scope-macro STRICTPASS despite 100% EXECSUCCESS. For reference, a separately configured coding agent solves 83 of 126 tasks, while five biomedical systems adapted to GPT-4o-mini reach at most 2.9% scope-macro STRICTPASS. These results expose a substantial gap between runnable submissions and correct clinical analyses.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!