2606.16149v1 Jun 15, 2026 cs.AI

LiteOdyssey: 해석 가능성을 갖춘 희귀 질환 진단을 위한 경량화된 추론 인공지능 에이전트

Teaching agentic AI to learn expert reasoning for rare disease diagnosis

Minh-Ha Nguyen
Minh-Ha Nguyen
Citations: 3
h-index: 1
Erica Gray
Erica Gray
Citations: 44
h-index: 4
Chih-Ting Yang
Chih-Ting Yang
Citations: 98
h-index: 4
R. Hamid
R. Hamid
Citations: 127
h-index: 6
Lingyao Li
Lingyao Li
Citations: 12
h-index: 3
Siyuan Ma
Siyuan Ma
Citations: 159
h-index: 5
T. Cassini
T. Cassini
Citations: 285
h-index: 10
Cathy Shyr
Cathy Shyr
Citations: 31
h-index: 3
B. Schuler
B. Schuler
Citations: 9
h-index: 2

대부분의 의료 AI 시스템은 추가적인 자원(더 많은 미세 조정 데이터, 더 많은 에이전트 및/또는 더 큰 검색 데이터베이스)을 확장하여 성능을 향상시킵니다. 그러나 희귀 질환 진단에서 이러한 확장은 배포, 감사 및 유지 관리가 어려운 시스템을 만들 수 있습니다. 본 연구에서는 최첨단 진단 성능이 단일 AI 에이전트의 추론 체인을 확장함으로써 달성될 수 있는지 조사했습니다. 즉, 인간-AI 협력을 통해 개발된 진단 정책으로 에이전트를 안내하고, 공개적으로 이용 가능한 생물 의학 도구를 활용하는 것입니다. 본 연구에서는 임상 유전 공학 워크플로우를 통해 추론 언어 모델을 안내하는 경량화된 희귀 질환 진단 프레임워크인 LiteOdyssey를 소개합니다. 이 프레임워크는 인간 피드백 기반 정책 반복(PIHF)을 통해 개발되었으며, 공개 생물 의학 도구에 대한 동적 접근 방식을 사용합니다. 두 가지 어려운 벤치마크에서, 환자의 임상 특징만을 제공하는 환경에서 LiteOdyssey는 최첨단 성능을 달성했습니다. LIRICAL (n = 370) 및 PhenoPacket Store (n = 873)의 총 1,243건의 사례에 대해 전체 질병 Recall@1이 59.3%였습니다. 두 벤치마크 모두 극히 드문 질환의 비율이 높으며(각각 약 45% 및 52.8%), 발생률이 100만 명당 1명 미만입니다. 더 어려운 PhenoPacket 데이터 세트에서, 원인 질환이 당사 희귀 질환 매핑 파이프라인에서 Orphanet에 매핑되지 않은 경우, LiteOdyssey는 Recall@1이 60.7%로, 동일한 기준 모델(GPT-5.4)의 10.7%보다 훨씬 높은 성능을 보였습니다. 이러한 성능은 미세 조정, 다중 에이전트 앙상블 또는 대규모 사례 검색 데이터베이스 없이 달성되었습니다. 또한 개발 중에 한 번도 관찰되지 않은 경우, 실제 희귀 질환 환자 코호트에서, 그리고 더 작은 오픈 액세스 모델에서도 성능 향상이 관찰되었습니다. LiteOdyssey는 정확하고, 배포가 용이하며, 의료 전문가의 검토를 위한 투명성이 높은 희귀 질환 AI 시스템 개발을 위한 가능성을 제시합니다.

Original Abstract

Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be converted into a scalable AI capability through a governed learning process rather than model training alone. We developed liteOdyssey through Policy Iteration with Human Feedback (PIHF), an in-context policy-learning method adapted from generalized policy iteration in reinforcement learning, in which model failures and expert corrections consolidate into an explicit, clinician-gated policy that turns an off-the-shelf LLM into an agentic diagnostic system. We demonstrated that such a policy improved diagnostic accuracy to match the best published systems at a fraction of their deployment footprint, generalized to unseen diseases, transferred across models, and remained under clinician control. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with nearly identical gains on the 1,193 cases and 679 diseases excluded from policy development. Ablations showed that gains exceeded automated prompting improvement or source access alone, and the policy transferred without modification across closed- and open-weight models. In 515 Undiagnosed Diseases Network patients, liteOdyssey again improved accuracy, and blinded physicians rated its differentials more often exact and less often unhelpful. Through PIHF, expert reasoning becomes an LLM capability that experts can inspect, revise, and transfer across models.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!