2608.03890v1 Aug 04, 2026 cs.CV

CARE-X: 보조적 지도 학습, 보상 정렬 학습 및 도구 기반 측정 방식을 활용한 임상적으로 유용한 영상 언어 모델 개발

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Tanuja Ganu
Tanuja Ganu
Citations: 982
h-index: 13
M. Ranjit
M. Ranjit
Citations: 511
h-index: 9
Anirban Porya
Anirban Porya
Citations: 2
h-index: 1
Sathvik Joel
Sathvik Joel
Citations: 125
h-index: 4
Niharika Vadlamudi
Niharika Vadlamudi
Citations: 9
h-index: 2
Nikhilesh Chowdary Eathamukkala
Nikhilesh Chowdary Eathamukkala
Citations: 0
h-index: 0
V. PrasanthV
V. PrasanthV
Citations: 0
h-index: 0
A. Swamy
A. Swamy
Citations: 24
h-index: 2
P. Umredkar
P. Umredkar
Citations: 0
h-index: 0
Pradeep Narayan
Pradeep Narayan
Citations: 0
h-index: 0
Vivek Rajagopal
Vivek Rajagopal
Citations: 4
h-index: 1

임상적으로 유용한 흉부 X선 시스템은 단순히 보고서 생성을 넘어, 조정 가능한 판단 기준을 가진 질병 분류, 공간적 위치 파악, 그리고 많은 진단에 필수적인 해부학적 측정값을 도출해야 합니다. 현재의 영상 언어 모델(VLMs)들은 이러한 기능을 개별적인 문제로 취급하거나 아예 고려하지 않아, 방사선과 의사가 필요로 하는 것과 생성 모델이 제공하는 것 사이의 간극이 존재합니다. 본 논문에서는 보조적 판별 학습과 보상 정렬 생성을 통합하여 이 간극을 줄이는 흉부 X선 VLM인 CARE-X를 소개합니다. CARE-X는 생성 모델의 기반에 focal-loss 분류 및 composite-loss 기반 위치 파악 모듈을 추가하고, 언어 모델링 목표와 함께 공동으로 학습시킵니다. 이러한 보조적 학습은 조정 가능한 판단 기준을 가진 판별 진단 예측과 정확한 공간적 위치 파악을 가능하게 하며, 보고서 품질 향상에도 기여합니다. 이는 구조화된 예측과 생성이 서로를 강화한다는 증거입니다. 이 기반 위에 Decoupled Clip과 Dynamic Sampling Policy Optimization (DAPO)을 활용하여, 보고서 생성, 시각 질의 응답(VQA), 그리고 공간적 위치 파악에 대한 작업별 보상 신호를 사용하여 실제 임상에서 중요한 품질 지표를 직접 최적화합니다. 그 결과, 네 가지 보고서 생성 벤치마크의 대부분 지표에서 최고 수준의 성능을 달성했으며, ReXVQA 데이터셋에서 94.0%의 VQA 정확도를 기록했습니다 (최고 성능 모델 대비 +6.0 pp 향상). 또한, 생성된 공간 정보는 전용 검출 모듈과 거의 동등한 성능을 보입니다. 별도로, 측정값에 의존하는 진단을 해결하기 위해, Qwen3-VL-4B-Instruct 모델과 원산(native) 도구 호출 기능을 결합하여, 결정적인 측정 도구를 사용하고 동시에 이미지 전체를 시각적으로 분석할 수 있도록 했습니다. 이러한 하이브리드 추론은 5가지 측정값에 의존하는 조건에서 기존의 시각 정보만 사용하는 모델 대비 평균 F1 점수가 +43.6 pp 향상되었습니다.

Original Abstract

A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained alongside the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality, providing evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, visual question answering (VQA), and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report-generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct with native tool-calling capabilities for invoking deterministic measurement tools, while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!