2605.26621v1 May 26, 2026 cs.CV

MedVol-R1: 보상 기반 증거 연계를 통한 입체 추론 분할

MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

Zihua Wang
Zihua Wang
Citations: 69
h-index: 3
Zichu Wang
Zichu Wang
Citations: 25
h-index: 2
Hairong Shi
Hairong Shi
Citations: 98
h-index: 2
Bingzheng Wei
Bingzheng Wei
Citations: 555
h-index: 11
Yan Xu
Yan Xu
Citations: 9
h-index: 2

입체 추론 분할(VRS)은 3차원 의료 영상에서 자유 형식의 임상 질문에 따라 특정 영역을 분할하는 것을 목표로 하며, 이때 참조 대상은 종종 명시적으로 제시되지 않고 의학적 지식과 입체 기반 추론이 모두 필요합니다. 기존 방법들은 일반적으로 언어 정보와 마스크 디코딩을 연결하기 위해 특수 분할 토큰에 의존하지만, 이러한 결합 방식은 의사 결정 과정을 불투명한 잠재 표현으로 만들고 해석 가능성과 다양한 서술 표현에 대한 일반화 능력을 제한합니다. 본 논문에서는 VRS를 위한 강화 학습 기반 프레임워크인 MedVol-R1을 제안합니다. MedVol-R1은 증거 연계를 입체 영역 규정으로부터 명시적으로 분리합니다. LVLM 모듈은 임상적 추론을 검증 가능한 2차원 증거 지점(주요 축 단면 및 2차원 경계 상자)에 연결하고, 이 정보는 고정된 MedSAM2 모듈에 의해 일관성 있는 3차원 마스크로 변환됩니다. MedVol-R1은 초기 지도 학습 미세 조정 후 GRPO를 사용하여 학습되며, 비용이 많이 드는 사고 과정 주석 없이도 유용한 증거 선택, 정확한 2차원 공간 정렬 및 슬라이스 간 입체 일관성을 장려하는 다중 구성 보상을 사용합니다. M3D-Seg 벤치마크의 CT-ORG, AbdomenCT-1K 및 KiTS23 데이터셋에 대한 실험 결과, MedVol-R1은 강력한 기준 모델보다 우수한 성능을 지속적으로 보여주며 최고 수준의 성능을 달성했습니다. 강화 학습은 순수 지도 학습 미세 조정에 비해 명확한 성능 향상을 제공합니다.

Original Abstract

Volumetric Reasoning Segmentation (VRS) aims to segment a target region in a 3D medical scan from a free-form clinical query, where the referent is often implicit and requires both medical knowledge and volume-grounded reasoning. Existing methods typically rely on specialized segmentation tokens to connect language with mask decoding, but this coupling collapses the decision process into opaque latent representations, limiting interpretability and generalization to diverse narrative expressions. In this paper, we present MedVol-R1, a reinforcement learning-based framework for VRS that explicitly decouples evidence grounding from volumetric delineation: the LVLM grounds clinical reasoning to a verifiable 2D evidence anchor (key axial slice and 2D bounding boxes), which is then propagated into a coherent 3D mask by a frozen MedSAM2 module. We train MedVol-R1 with cold-start supervised fine-tuning followed by GRPO, guided by a multi-component reward that encourages informative evidence selection, accurate 2D spatial grounding, and cross-slice volumetric coherence, without requiring costly chain-of-thought annotations. Experiments on CT-ORG, AbdomenCT-1K, and KiTS23 from the M3D-Seg benchmark demonstrate that MedVol-R1 consistently outperforms strong baselines and achieves state-of-the-art performance, with reinforcement learning providing clear gains over pure supervised fine-tuning.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!