2606.11830v1 Jun 10, 2026 cs.AI

의료 연구 분석을 위한 기술 기반 인공지능 에이전트: 비소세포 폐암 전사체 바이오마커 작업에서의 다중 모델 인간 평가 연구

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

Fei Sun
Fei Sun
Citations: 2
h-index: 1
Bo-Sheng Huang
Bo-Sheng Huang
Citations: 2
h-index: 1
Qianyun Yao
Qianyun Yao
Citations: 24
h-index: 3
Wei Chen
Wei Chen
Citations: 53
h-index: 2
Jiarui Jiang
Jiarui Jiang
Citations: 8
h-index: 1
Shu Quan
Shu Quan
Citations: 141
h-index: 2
Yifei Chen
Yifei Chen
Citations: 9
h-index: 1
Wenjie Xu
Wenjie Xu
Citations: 6
h-index: 2
Bo Li
Bo Li
Citations: 20
h-index: 2
Li Su
Li Su
Citations: 11
h-index: 1
R. Wu
R. Wu
Citations: 8
h-index: 2
Huhai Hong
Huhai Hong
Citations: 41
h-index: 3
Huimei Wang
Huimei Wang
Citations: 8
h-index: 1

배경. 대규모 언어 모델과 인공지능 에이전트는 생의학 연구를 지원하는 데 점점 더 많이 사용되고 있지만, 기본 모델의 출력은 중요한 분석 단계를 누락하거나, 방법을 오용하거나, 결론을 과장할 수 있습니다. 본 연구에서는 기술 패키지에 대한 자율적인 접근성이, 기술 없이 기본 AI보다 고품질의 AI 생성 전사체 연구 분석 결과를 가져다주는지 평가했습니다. 방법. 비소세포 폐암 면역 요법 바이오마커 작업을 사용하여 탐색적 다중 모델 인간 평가를 수행했습니다. 6개의 모델 백본을 테스트했으며, 평가에는 익명화된 21개의 출력 결과가 포함되었습니다 (기본 AI 출력 9개, 기술 기반 에이전트(OpenClaw) 구현을 통해 생성된 기술 강화 출력 12개). 4명의 생의학 분야 비전문가 검토자와 2명의 익명 전문가가 각 출력 결과를 평가했으며, 각 검토자 유형별로 2개의 등급을 부여했습니다. 주요 결과는 전문가가 평가한 전반적인 품질이었습니다. 결과. 기술 강화된 출력은 기본 AI 출력보다 전문가의 전반적인 품질 측면에서 더 높은 경향을 보였습니다 (평균 5.50 vs 5.11; 차이 = 0.39; 부트스트랩 95% CI, -0.04 to 0.90; Welch p = 0.156). 비전문가 검토자의 품질 평가에서도 동일한 경향을 보였습니다 (평균 4.72 vs 4.47; 차이 = 0.26; 부트스트랩 95% CI, -0.25 to 0.80; Welch p = 0.373). 전문가 간의 일치도는 제한적이었으며 (단일 등급 ICC = -0.15), 모델별 효과는 기술적으로 의미가 있었지만 이질적이었습니다. 결론. 본 연구에서는 자율적인 기술 접근이 탐색적 샘플에서 품질 향상에 대한 긍정적인 경향을 보였지만, 이러한 경향은 전문가 평가의 노이즈보다 작으므로 확증적인 증거로 해석되어서는 안 됩니다. 본 연구 결과는 더 강력한 신뢰성 제어, 플랫폼 복제 및 생물학적 타당성 평가를 포함하는 기술 기반 AI 에이전트에 대한 더 큰 규모의 평가를 촉구합니다.

Original Abstract

Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or overstate conclusions. We evaluated whether autonomous access to a medical research skill package was associated with higher-quality AI-generated transcriptomic research-analysis outputs compared with native AI without skills. Methods. We conducted an exploratory multi-model human evaluation using a non-small cell lung cancer immunotherapy biomarker task. Six model backbones were tested. The evaluation included 21 anonymized outputs: 9 native-AI outputs and 12 skill-augmented outputs generated through an AI agent implementation represented by OpenClaw. Four non-expert biomedical reviewers and two blinded experts evaluated each output, with two ratings from each reviewer type. The primary outcome was expert-rated overall quality. Results. Skill-augmented outputs showed directionally higher expert overall quality than native-AI outputs (mean 5.50 vs 5.11; difference=0.39; bootstrap 95\% CI, -0.04 to 0.90; Welch p=0.156). Non-expert reviewer quality showed the same direction (mean 4.72 vs 4.47; difference=0.26; bootstrap 95\% CI, -0.25 to 0.80; Welch p=0.373). Expert agreement was limited (single-rating ICC=-0.15), and model-specific effects were descriptive and heterogeneous. Conclusions. Autonomous skill access showed a directional quality signal in this exploratory sample, but the signal was smaller than expert-rating noise and should not be interpreted as confirmatory evidence. The findings primarily motivate larger evaluations of skill-augmented AI agents with stronger reliability controls, platform replication, and biological-validity assessment.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!