2608.10635v1 Aug 11, 2026 cs.CV

MedUP: 의료 영상-언어 모델에서 통합적인 이해와 인지 능력 향상

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shujian Gao
Shujian Gao
Citations: 24
h-index: 2
Songtao Jiang
Songtao Jiang
Citations: 239
h-index: 7
Jian Wu
Jian Wu
Citations: 146
h-index: 5
Siming Fu
Siming Fu
Citations: 343
h-index: 8
Yixin Chen
Yixin Chen
Citations: 0
h-index: 0
Jiaming Lin
Jiaming Lin
Citations: 0
h-index: 0

의료 영상-언어 모델(Med-VLM)은 시각적 내용을 설명하는 데 뛰어난 성능을 보이지만, 정확한 시각적 인식, 분할 및 위치 지정을 하는 것은 여전히 어려운 과제입니다. 기존 방법들은 영역을 좌표 문자열로 표현하거나, 인지 능력과 이해 능력을 분리시키는 외부 모듈에 의존하여 영역-언어 정렬 과정에서 정보 손실이 발생합니다. 본 연구에서는 MedUP이라는 Med-VLM을 제안하며, 이는 하나의 토큰 공간 내에서 시각적 인식과 이해 능력을 통합하는 모델입니다. 핵심 기술인 UniMedTok은 마스크를 LLM 어휘의 개별 토큰으로 인코딩하여, 모델이 마스크 토큰과 텍스트를 원활하게 결합할 수 있도록 합니다. 또한, 본 연구에서는 텍스트 기반 분할, 영역 기반 이해, 의료 질의 응답 및 추론 기반 분할을 포함하는 184만 개의 데이터셋인 UniMed-Train을 구축하고, 통합적인 평가를 위한 UniMed-Bench를 소개합니다. 광범위한 실험 결과는 MedUP이 기존의 모델, 에이전트 기반 모델 및 이중 디코더 모델보다 모든 작업에서 우수한 성능을 보이며, 전문 분할 모델과 경쟁력 있는 수준임을 보여주므로, 통합적인 이해와 인지 능력 모델링의 잠재력이 매우 크다는 것을 입증합니다.

Original Abstract

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!