CURE: 교육 과정을 활용한 다중 작업 훈련을 통한 신뢰성 있는 해부학적 기반 보고서 생성
CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation
의료 영상-언어 모델은 방사선 보고서 생성을 자동화할 수 있지만, 정확한 시각적 기반 설정 및 사실 일관성 측면에서 어려움을 겪습니다. 기존 모델은 종종 텍스트 정보와 시각적 증거 간의 불일치를 보여, 신뢰성이 낮거나 기반이 약한 예측을 생성합니다. 본 연구에서는 CURE라는 오류 인지 교육(curriculum learning) 프레임워크를 제시합니다. CURE는 추가 데이터 없이 시각적 기반 설정 및 보고서 품질을 향상시킵니다. CURE는 공개 데이터 세트를 사용하여 구문 기반 설정, 기반 설정된 보고서 생성 및 해부학적 기반 보고서 생성에 대해 다중 모드 모델을 미세 조정합니다. 이 방법은 모델 성능에 따라 동적으로 샘플링을 조정하여 공간적 및 텍스트 정렬을 개선하기 위해 더 어려운 샘플에 중점을 둡니다. CURE는 IoU를 +0.37만큼 향상시켜 시각적 기반 설정 정확도를 높이고, CXRFEScore를 +0.188만큼 향상시켜 보고서 품질을 높이며, 환각 현상을 18.6% 줄입니다. CURE는 시각적 기반 설정 정확도와 보고서 신뢰성을 모두 향상시키는 데이터 효율적인 프레임워크입니다. 코드는 https://github.com/PabloMessina/CURE에서, 모델 가중치는 https://huggingface.co/pamessina/medgemma-4b-it-cure에서 확인할 수 있습니다.
Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign textual findings with visual evidence, leading to unreliable or weakly grounded predictions. We present CURE, an error-aware curriculum learning framework that improves grounding and report quality without any additional data. CURE fine-tunes a multimodal instructional model on phrase grounding, grounded report generation, and anatomy-grounded report generation using public datasets. The method dynamically adjusts sampling based on model performance, emphasizing harder samples to improve spatial and textual alignment. CURE improves grounding accuracy by +0.37 IoU, boosts report quality by +0.188 CXRFEScore, and reduces hallucinations by 18.6%. CURE is a data-efficient framework that enhances both grounding accuracy and report reliability. Code is available at https://github.com/PabloMessina/CURE and model weights at https://huggingface.co/pamessina/medgemma-4b-it-cure
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.