2607.23290v2 Jul 25, 2026 cs.AI

RareLens: 다양한 대규모 언어 모델 추론을 활용하여 희귀 질환 진료의 전 과정 혁신

RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning

Xi Chen
Xi Chen
Citations: 11
h-index: 1
Huahui Yi
Huahui Yi
Citations: 80
h-index: 5
Hanyu Zhou
Hanyu Zhou
Citations: 35
h-index: 3
W. Fu
W. Fu
Citations: 1,068
h-index: 21
Kang Li
Kang Li
Citations: 5
h-index: 1
Jian Li
Jian Li
Citations: 537
h-index: 6
Shaoting Zhang
Shaoting Zhang
Citations: 1,256
h-index: 14
Shiyu Feng
Shiyu Feng
Citations: 1
h-index: 1
Kun Wang
Kun Wang
Citations: 93
h-index: 4
Benyou Wang
Benyou Wang
Citations: 3
h-index: 1
T. He
T. He
Citations: 615
h-index: 14
Qiankun Li
Qiankun Li
Citations: 517
h-index: 9
Xiaohong Zheng
Xiaohong Zheng
Citations: 0
h-index: 0
Hongru Zhou
Hongru Zhou
Citations: 19
h-index: 2
Rongsheng Wang
Rongsheng Wang
Citations: 33
h-index: 2
Ping Liu
Ping Liu
Citations: 0
h-index: 0
Sicheng Lin
Sicheng Lin
Citations: 35
h-index: 1
Huiying Ou
Huiying Ou
Citations: 0
h-index: 0
Tianying Zang
Tianying Zang
Citations: 35
h-index: 3
Zhuohang Wu
Zhuohang Wu
Citations: 6
h-index: 1
Leheng Jiang
Leheng Jiang
Citations: 14
h-index: 2
Ke-jun Cao
Ke-jun Cao
Citations: 0
h-index: 0
Wen-Hui Zhang
Wen-Hui Zhang
Citations: 0
h-index: 0
Cheng-Yi Li
Cheng-Yi Li
Citations: 128
h-index: 5
Zhiyang Wang
Zhiyang Wang
Citations: 46
h-index: 4
Songlin Li
Songlin Li
Citations: 162
h-index: 8
Ningbei Yin
Ningbei Yin
Citations: 36
h-index: 3

희귀 질환은 이질적인 증상, 부족한 근거 자료 및 제한된 전문 지식으로 인해 임상 의사 결정에서 가장 어려운 영역 중 하나이며, 치료 과정 전체에 걸쳐 지속적인 불확실성을 야기합니다. 인공지능이 도움이 될 수 있지만, 기존 시스템은 주로 진단과 같은 개별적인 작업에 초점을 맞추며, 초기 검사 시 얻을 수 있는 정보보다는 후속 검사에 의존하는 경향이 있습니다. 본 연구에서는 단일 모델의 확장을 통해 개선되는 것이 아니라, 여러 불완전한 추론 시스템의 다양성을 활용함으로써 불확실성 하에서의 임상 인공지능 성능을 향상시킬 수 있음을 보여줍니다. 다양한 대규모 언어 모델에서 상호 보완적인 오류 패턴을 가진 서로 다른 추론 경로를 파악하고, RareLens를 개발하여 이러한 관점을 조화시켜 희귀 질환 진료의 네 가지 단계(위험성 선별, 진단, 치료 계획 및 예후 예측)에서 실행 가능한 의사 결정을 내리도록 합니다. 실제 데이터를 기반으로 구축된 RareLensBench 데이터셋은 33개의 Orphanet 범주와 7,000개 이상의 질환을 포함하는 총 157,525건의 사례로 구성되어 있습니다. RareLens는 GPT-5, DeepSeek-R1, Claude-3.7-Sonnet 및 Gemini-2.5-Pro를 포함한 모든 테스트 대상 최첨단 모델보다 우수한 성능을 보였습니다. 선별 단계에서 0.917의 AUC(Area Under the Curve) 값을, 진단 및 치료 단계에서는 각각 65.5%와 89.8%의 최고 정확도를 달성했습니다. 외부 평가에서는 1,287건의 사례를 대상으로 23명의 의사들이 참여했으며, RareLens를 활용한 자율 시스템과 RareLens의 도움을 받은 의료진 모두 독자적인 판단을 내린 의료진보다 우수한 성능을 보였습니다. 이는 효과적인 인간-AI 협업이 단순히 모델 출력을 제공하는 것 이상임을 보여줍니다. 이러한 결과는 서로 다른 모델의 추론 방식을 유용한 정보원으로 활용할 수 있음을 입증하며, 높은 임상적 불확실성 하에서 안정적으로 작동하는 인공지능 시스템을 구축하기 위한 일반적인 전략을 제시합니다.

Original Abstract

Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse evidence and limited expertise create persistent uncertainty throughout the care pathway. Although artificial intelligence could help, existing systems largely address isolated tasks, particularly diagnosis, and usually rely on downstream investigations rather than information available at initial presentation. Here we show that clinical AI performance under uncertainty can be improved not by scaling a single model, but by exploiting the diversity of multiple imperfect reasoning systems. Across heterogeneous large language models, we identify divergent reasoning trajectories with complementary error patterns and develop RareLens, which learns to reconcile these perspectives into actionable decisions across four stages of rare disease care: risk screening, diagnosis, treatment planning and prognosis prediction. Built on RarelensBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, across all stages. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external evaluation involving 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both outperformed unaided physicians, while demonstrating that effective human-AI collaboration requires more than simply providing model outputs. These findings establish divergent model reasoning as an exploitable source of information and suggest a general strategy for building AI systems that operate reliably under high clinical uncertainty.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!