RIHA: 보고서-이미지 계층적 정렬을 통한 방사선 보고서 생성
RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation
방사선 보고서 생성(RRG)은 의료 영상으로부터 자동으로 진단 보고서를 생성하여 방사선 전문의의 업무 부담을 줄이고 인적 오류를 감소시키는 유망한 접근 방식입니다. RRG의 핵심 과제는 복잡한 시각적 특징과 긴 형식의 방사선 보고서의 계층적 구조 간의 정밀한 정렬을 달성하는 것입니다. 최근 방법들은 이미지-텍스트 표현 학습을 개선했지만, 종종 보고서를 단순한 시퀀스로 취급하여 구조화된 섹션과 의미적 계층 구조를 간과합니다. 이러한 단순화는 정확한 모달 간 정렬을 방해하고 RRG의 정확도를 약화시킵니다. 이러한 과제를 해결하기 위해, 우리는 RIHA(Report-Image Hierarchical Alignment Transformer)를 제안합니다. RIHA는 단락, 문장, 단어 수준에서 방사선 이미지와 해당 보고서 간의 다단계 정렬을 수행하는 새로운 엔드-투-엔드 프레임워크입니다. 이러한 계층적 정렬은 임상 내러티브에 내재된 미묘한 의미를 포착하는 데 필수적인 보다 정확한 모달 간 매핑을 가능하게 합니다. 구체적으로, RIHA는 다중 스케일 시각적 특징을 추출하기 위한 시각적 특징 피라미드(VFP)와 다중 수준 텍스트 구조를 표현하기 위한 텍스트 특징 피라미드(TFP)를 도입합니다. 이러한 구성 요소는 최적 수송을 활용하여 다양한 수준에서 시각적 및 텍스트 특징을 효과적으로 정렬하는 크로스-모달 계층적 정렬(CHA) 모듈을 통해 통합됩니다. 또한, RIHA는 디코더에 상대적 위치 인코딩(RPE)을 통합하여 토큰 간의 공간적 및 의미적 관계를 모델링하여 시각적 특징과 생성된 텍스트 간의 토큰 수준 정렬을 향상시킵니다. IU-Xray 및 MIMIC-CXR이라는 두 개의 벤치마크 흉부 X-ray 데이터 세트에 대한 광범위한 실험 결과, RIHA가 자연어 생성 및 임상 효능 지표 모두에서 기존의 최첨단 모델보다 우수한 성능을 발휘하는 것으로 나타났습니다.
Radiology report generation (RRG) has emerged as a promising approach to alleviate radiologists' workload and reduce human errors by automatically generating diagnostic reports from medical images. A key challenge in RRG is achieving fine-grained alignment between complex visual features and the hierarchical structure of long-form radiology reports. Although recent methods have improved image-text representation learning, they often treat reports as flat sequences, overlooking their structured sections and semantic hierarchies. This simplification hinders precise cross-modal alignment and weakens RRG accuracy. To address this challenge, we propose RIHA (Report-Image Hierarchical Alignment Transformer), a novel end-to-end framework that performs multi-level alignment between radiological images and their corresponding reports across paragraph, sentence, and word levels. This hierarchical alignment enables more precise cross-modal mapping, essential for capturing the nuanced semantics embedded in clinical narratives. Specifically, RIHA introduces a Visual Feature Pyramid (VFP) to extract multi-scale visual features and a Text Feature Pyramid (TFP) to represent multi-granularity textual structures. These components are integrated through a Cross-modal Hierarchical Alignment (CHA) module, leveraging optimal transport to effectively align visual and textual features across various levels. Furthermore, we incorporate Relative Positional Encoding (RPE) into the decoder to model spatial and semantic relationships among tokens, enhancing the token-level alignment between visual features and generated text. Extensive experiments on two benchmark chest X-ray datasets, IU-Xray and MIMIC-CXR, demonstrate that RIHA outperforms existing state-of-the-art models in both natural language generation and clinical efficacy metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.