2607.08038v1 Jul 09, 2026 cs.AI

AI 기반 감별 진단 시스템을 위한 안전 중심의 가설 연역 프레임워크

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

Lingfei Qian
Lingfei Qian
Citations: 359
h-index: 10
Qingyu Chen
Qingyu Chen
Citations: 630
h-index: 11
Mauro Giuffrè
Mauro Giuffrè
Citations: 2
h-index: 1
F. Ma
F. Ma
Citations: 5
h-index: 1
L. Ohno-Machado
L. Ohno-Machado
Citations: 576
h-index: 7
Huan He
Huan He
Citations: 229
h-index: 5
M. Jiang
M. Jiang
Citations: 157
h-index: 3
Cathy Shyr
Cathy Shyr
Citations: 31
h-index: 3
D. Wright
D. Wright
Citations: 48
h-index: 2
K. McCann
K. McCann
Citations: 0
h-index: 0
M. Iscoe
M. Iscoe
Citations: 184
h-index: 8
Chi Wing Ng
Chi Wing Ng
Citations: 55
h-index: 1
Na Hong
Na Hong
Citations: 65
h-index: 4
Lee H Schwamm
Lee H Schwamm
Citations: 112
h-index: 3
Hua Xu
Hua Xu
Citations: 25
h-index: 2

진단 오류는 환자 안전에 대한 주요 위협이지만, 현재의 대규모 언어 모델(LLM) 시스템은 종종 진단을 일회성 예측 작업으로 처리하며, 놓칠 수 있는 고위험 옵션에 대한 안전장치가 부족하고 추론 과정을 엄격하게 검증하지 못하는 경우가 많습니다. 본 연구에서는 안전 중심의 가설 연역적 임상 추론 프레임워크인 AegisDx를 제안합니다. AegisDx는 역할별 계약, 구조화된 중간 결과물, 증거 검색 인터페이스 및 검증 게이트를 통해 특수 LLM 구성 요소를 조정하여 광범위한 감별 진단을 생성하고, 위험한 '절대 놓쳐서는 안 되는' 질환에 대한 명시적인 선별을 수행하며, 근거 기반의 의학적 증거에 따른 추론 과정을 검증하고, 실행 가능한 다음 단계를 구조화합니다. AegisDx는 세 가지 측면에서 평가되었습니다. NEJM 및 JAMA에서 추출한 문헌 기반 사례 보고서에서는 GPT-oss-120B를 백본으로 사용하여 상위 3개 진단의 정확도가 JAMA 사례의 경우 52.1% 대비 59.9%, NEJM 사례의 경우 51.4% 대비 62.7%로 향상되었습니다. Annals of Emergency Medicine의 사례에서는 상위 3개 진단 정확도가 68.6% 대비 85.7%였습니다. 의사 합의하에 '절대 놓쳐서는 안 되는' 진단 목록에 대해, AegisDx는 전체 사례 중 78.0%에서 해당 조건을 상위 3개의 진단 항목 중 하나로 포함시킨 반면, 독립 실행형 LLM은 52.0%였습니다. Yale New Haven Health System의 실제 응급실 기록 43건을 GPT-5와 비교한 검토 결과, AegisDx는 의사가 평가한 종합 안전 점수를 4.31에서 4.55로 향상시켰습니다 (조정된 p = 2.1x10^-4). 또한 '절대 놓쳐서는 안 되는' 질환 식별 및 추론 안전성 측면에서도 질적 개선이 있었습니다. 본 연구 결과는 진단 AI를 단순한 예측 정확도를 최적화하는 것이 아니라, 안전 중심의 추론 프레임워크로 설계함으로써 급성 치료 워크플로우에 대한 더욱 안전하고 투명하며 임상적으로 의미 있는 의사 결정 지원 시스템을 제공할 수 있음을 시사합니다.

Original Abstract

Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning. Here, we present AegisDx, a safety-oriented framework for hypothetico-deductive clinical reasoning. AegisDx coordinates specialized LLM components through role-specific contracts, structured intermediate outputs, evidence-retrieval interfaces, and verification gates to generate broad differential diagnoses, enforce explicit screening for dangerous "must-not-miss" conditions, verify reasoning against grounded medical evidence, and structure actionable next steps. We evaluated AegisDx across three layers. On literature-derived case reports from NEJM and JAMA, with GPT-oss-120B as the shared backbone, Top-3 diagnostic accuracy was 59.9% versus 52.1% for the standalone LLM on JAMA cases and 62.7% versus 51.4% on NEJM cases. On cases from Annals of Emergency Medicine, Top-3 accuracy was 85.7% versus 68.6%; against physician-consensus must-not-miss diagnosis sets, AegisDx captured at least one such condition among its top three diagnoses in 78.0% of cases versus 52.0%. In a blinded physician evaluation of 43 real-world emergency department notes from the Yale New Haven Health System compared against GPT-5, AegisDx improved the physician-rated composite safety score from 4.31 to 4.55 on a 5-point scale (adjusted p = 2.1x10^-4), with qualitative gains in must-not-miss identification and reasoning safety. Our findings suggest that engineering diagnostic AI as a safety-oriented reasoning framework, rather than optimizing raw predictive accuracy alone, can provide a safer, more transparent, and clinically meaningful layer of bedside decision support for acute care workflows.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!