오류 없는 전자 건강 기록(EHR)을 향하여: 전자 건강 기록 시스템 내 임상 노트와 정형화된 테이블 간의 추론 기반 일관성 검증
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records
전자 건강 기록(EHR)에서 비정형 임상 노트와 정형화된 테이블 간의 데이터 일관성은 환자 안전 및 임상 의사 결정에 필수적입니다. 그러나 기존의 노트-테이블 일관성 검증 연구는 주로 숫자 값 또는 단순 이벤트의 표면 수준 매칭에 의존합니다. 이러한 접근 방식은 임상 해석, 사건 관계 및 시간적 변화를 포함한 실제 EHR 문서의 근본적인 추론을 포착하지 못합니다. 이러한 격차를 해결하기 위해 우리는 노트-테이블 일관성 검증을 위한 추론 중심 벤치마크인 EHR-ReasonCon을 소개합니다. 전문가의 지도를 받아 구축되었으며, MIMIC-III 데이터셋을 기반으로 하며, 임상 노트에서 추출된 8,048개의 개체에 대한 고품질의 정답 레이블을 제공합니다. 이 주석 프로토콜은 체계적인 증거 검색 및 신뢰할 수 있는 일관성 평가를 보장하기 위해 특수 테이블 탐색 도구를 지원합니다. 또한 우리는 노트를 분할하고, 핵심 개체 및 시간 참조를 추출하며, 테이블 탐색 도구를 사용하여 정형화된 테이블과의 일관성을 검증하는 LLM 기반 프레임워크인 EHR-Inspector를 제안합니다. 전문가가 검증한 LLM을 사용하여 엄격하고 관대한 기준 하에서 평가한 결과, EHR-Inspector는 다양한 모델 아키텍처에서 최첨단 성능을 달성했습니다. 추가 분석은 구성 요소의 효과성을 입증하고 인간 검증과의 차이점을 강조합니다.
Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and clinical decision-making. However, existing work on note-table consistency verification mainly relies on surface-level matching of numeric values or simple events. Such approaches fail to capture the reasoning underlying real-world EHR documentation, including clinical interpretation, event relations, and temporal changes. To address this gap, we introduce EHR-ReasonCon, a reasoning-intensive benchmark for note-table consistency verification. Built on MIMIC-III with expert-guided annotations, it comprises 8,048 entities derived from clinical notes and provides high-quality ground-truth labels. The annotation protocol is supported by specialized table-exploration tools to ensure systematic evidence retrieval and reliable consistency assessment. We also propose EHR-Inspector, an LLM-based framework that segments notes, extracts anchor entities and temporal references, and uses table-exploration tools to verify consistency against structured tables. Evaluated using expert-validated LLM-as-a-judge metrics under harsh and lenient criteria, EHR-Inspector achieves state-of-the-art performance across multiple model backbones. Analyses further demonstrate the effectiveness of its components and highlight differences from human verification.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.