CARE-X: 보조적 지도 학습, 보상 정렬 학습 및 도구 기반 측정 방식을 활용한 임상적으로 유용한 영상 언어 모델 개발
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
임상적으로 유용한 흉부 X선 시스템은 단순히 보고서 생성을 넘어, 조정 가능한 판단 기준을 가진 질병 분류, 공간적 위치 파악, 그리고 많은 진단에 필수적인 해부학적 측정값을 도출해야 합니다. 현재의 영상 언어 모델(VLMs)들은 이러한 기능을 개별적인 문제로 취급하거나 아예 고려하지 않아, 방사선과 의사가 필요로 하는 것과 생성 모델이 제공하는 것 사이의 간극이 존재합니다. 본 논문에서는 보조적 판별 학습과 보상 정렬 생성을 통합하여 이 간극을 줄이는 흉부 X선 VLM인 CARE-X를 소개합니다. CARE-X는 생성 모델의 기반에 focal-loss 분류 및 composite-loss 기반 위치 파악 모듈을 추가하고, 언어 모델링 목표와 함께 공동으로 학습시킵니다. 이러한 보조적 학습은 조정 가능한 판단 기준을 가진 판별 진단 예측과 정확한 공간적 위치 파악을 가능하게 하며, 보고서 품질 향상에도 기여합니다. 이는 구조화된 예측과 생성이 서로를 강화한다는 증거입니다. 이 기반 위에 Decoupled Clip과 Dynamic Sampling Policy Optimization (DAPO)을 활용하여, 보고서 생성, 시각 질의 응답(VQA), 그리고 공간적 위치 파악에 대한 작업별 보상 신호를 사용하여 실제 임상에서 중요한 품질 지표를 직접 최적화합니다. 그 결과, 네 가지 보고서 생성 벤치마크의 대부분 지표에서 최고 수준의 성능을 달성했으며, ReXVQA 데이터셋에서 94.0%의 VQA 정확도를 기록했습니다 (최고 성능 모델 대비 +6.0 pp 향상). 또한, 생성된 공간 정보는 전용 검출 모듈과 거의 동등한 성능을 보입니다. 별도로, 측정값에 의존하는 진단을 해결하기 위해, Qwen3-VL-4B-Instruct 모델과 원산(native) 도구 호출 기능을 결합하여, 결정적인 측정 도구를 사용하고 동시에 이미지 전체를 시각적으로 분석할 수 있도록 했습니다. 이러한 하이브리드 추론은 5가지 측정값에 의존하는 조건에서 기존의 시각 정보만 사용하는 모델 대비 평균 F1 점수가 +43.6 pp 향상되었습니다.
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained alongside the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality, providing evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, visual question answering (VQA), and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report-generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct with native tool-calling capabilities for invoking deterministic measurement tools, while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.