2606.31599v1 Jun 30, 2026 cs.CV

이중 스트림 강화 학습을 통한 토큰 희소 의료 다중 모드 추론

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

Qihao Zheng
Qihao Zheng
Citations: 176
h-index: 9
Mianxin Liu
Mianxin Liu
Citations: 147
h-index: 8
Shangquan Sun
Shangquan Sun
Citations: 26
h-index: 1
Chunfeng Song
Chunfeng Song
Citations: 192
h-index: 9
Xiaosong Wang
Xiaosong Wang
Citations: 37
h-index: 3
Kaitao Chen
Kaitao Chen
Citations: 45
h-index: 4
Jiamin Wu
Jiamin Wu
Citations: 123
h-index: 7
Mu Zhou
Mu Zhou
Citations: 764
h-index: 15
Wei Zhao
Wei Zhao
Citations: 85
h-index: 5

강화 학습(RL)을 결합한 시각-언어 모델(VLMs)은 다중 모드 추론 분야에서 놀라운 발전을 이루었지만, 임상 의사 결정에 필요한 극도로 희소한 시각적 증거를 보이는 의료 영상에서는 여전히 어려움을 겪습니다. 본 연구에서는 시각적 토큰 중 중요한 영역을 벗어난 불필요한 토큰을 제거하는 것이 의료 추론 성능 향상에 크게 기여한다는 점을 확인했습니다. 그러나 능동적인 시각적 토큰 가지치기(VTP)와 의료 다중 모드 추론을 위한 통합된 강화 학습 프레임워크는 아직 확립되지 않았습니다. 이에 본 연구에서는 토큰 가지치기와 질의 응답을 수행하는 이중 스트림 강화 학습 프레임워크인 ViToS를 제안합니다. ViToS는 두 개의 작업 분기로 구성된 정책 모델을 훈련하며, 한 분기는 지역화(grounding)에 집중하고 다른 분기는 VTP 후 토큰 희소 추론을 수행합니다. 또한, 교차 피드백 순차 최적화를 도입하여 결합된 정책 학습 문제를 해결함으로써, 기울기 충돌을 방지하고 공유 정책 모델의 수렴을 촉진합니다. 7개의 의료 데이터셋에 대한 실험 결과, ViToS는 시각적 토큰 수를 원래 길이의 77% 수준으로 줄이면서 Lingshu-7B에서 108.27%, HuatuoGPT-Vision-7B에서 104.16%의 상대적인 성능 향상을 달성했습니다. 전반적으로 ViToS는 우수한 성능과 추론 속도 향상을 제공하며, 의료 다중 모드 추론을 위한 효율적인 패러다임을 제시합니다.

Original Abstract

Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medical images, which typically exhibit extremely sparse visual evidence to inform clinical decision-making. We recognize that pruning visual tokens outside the grounding region greatly enhances medical reasoning. However, a united RL framework for active visual token pruning (VTP) and medical multimodal reasoning remains unestablished. Here, we propose a dual-stream RL framework, ViToS, to fulfill token pruning and question answering. ViToS trains one policy model with two task branches, where one focuses on grounding while the other conducts token-sparse reasoning after VTP. Furthermore, we solve the coupled policy learning problem by introducing the cross-feedback sequential optimization, avoiding gradient conflict and facilitating convergence of the shared policy model. Evaluated on seven medical benchmarks, our method reduces visual tokens to 77% of the original sequence length while achieving a 108.27% relative performance on Lingshu-7B and 104.16% relative performance on HuatuoGPT-Vision-7B. Overall, ViToS delivers superior performance and inference speedup, establishing an efficient paradigm for medical multimodal reasoning.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!