CWCD: 범주별 대비 디코딩을 통한 구조화된 의료 보고서 생성
CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation
흉부 X-선 판독은 해부학적 구조 간의 중첩과 임상적으로 중요한 많은 병변의 미묘한 표현으로 인해 본질적으로 어렵기 때문에, 숙련된 방사선과 의사조차도 정확한 진단을 내리는 데 상당한 시간이 소요됩니다. 최근의 방사선학 중심 기초 모델인 LLaVA-Rad 및 Maira-2와 같은 모델들은 다중 모달 대규모 언어 모델(MLLM)을 자동 방사선 보고서 생성(RRG)의 최전선으로 이끌었습니다. 그러나 이러한 발전에도 불구하고, 현재의 기초 모델들은 단일 순방향 패스 방식으로 보고서를 생성합니다. 이러한 디코딩 전략은 시각적 토큰에 대한 주의를 감소시키고 생성 과정에서 언어적 사전 지식에 대한 의존도를 증가시켜, 결과적으로 생성된 보고서에 잘못된 병변 동반 발생을 유발합니다. 이러한 한계점을 극복하기 위해, 우리는 구조화된 방사선 보고서 생성(SRRG)을 향상시키기 위한 새로운 모듈형 프레임워크인 범주별 대비 디코딩(CWCD)을 제안합니다. 우리의 접근 방식은 범주별 매개변수화를 도입하고, 범주별 시각적 프롬프트를 사용하여 정상 X-선과 마스킹된 X-선을 대비시켜 범주별 보고서를 생성합니다. 실험 결과는 CWCD가 임상 효능 및 자연어 생성 지표 모두에서 기존 방법보다 일관되게 우수한 성능을 보인다는 것을 보여줍니다. 또한, 추가적인 분석을 통해 전체 성능에 대한 각 아키텍처 구성 요소의 기여도를 명확히 밝힙니다.
Interpreting chest X-rays is inherently challenging due to the overlap between anatomical structures and the subtle presentation of many clinically significant pathologies, making accurate diagnosis time-consuming even for experienced radiologists. Recent radiology-focused foundation models, such as LLaVA-Rad and Maira-2, have positioned multi-modal large language models (MLLMs) at the forefront of automated radiology report generation (RRG). However, despite these advances, current foundation models generate reports in a single forward pass. This decoding strategy diminishes attention to visual tokens and increases reliance on language priors as generation proceeds, which in turn introduces spurious pathology co-occurrences in the generated reports. To mitigate these limitations, we propose Category-Wise Contrastive Decoding (CWCD), a novel and modular framework designed to enhance structured radiology report generation (SRRG). Our approach introduces category-specific parameterization and generates category-wise reports by contrasting normal X-rays with masked X-rays using category-specific visual prompts. Experimental results demonstrate that CWCD consistently outperforms baseline methods across both clinical efficacy and natural language generation metrics. An ablation study further elucidates the contribution of each architectural component to overall performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.