VITAL: 향상된 설명 가능성을 위한 시각-의미 이중 감독을 통한 의료용 대규모 언어 모델의 잠재적 추론
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
잠재적 추론은 명시적인 토큰 대신 연속적인 숨겨진 상태에 기반하여 추론을 수행하며, 이를 통해 의료 분야의 시각 질문 답변(VQA)에서 발생하는 언어 장벽과 추론 비용을 줄일 수 있습니다. 그러나 기존 방법들은 모달리티 붕괴, 불충분한 시각적 감독, 학습-추론 불일치 등의 문제를 가지고 있으며, 또한 투명하지 않은 잠재 상태는 임상 응용 분야에서 매우 중요한 해석 가능성을 제공하지 못합니다. 본 논문에서는 VITAL이라는 프레임워크를 제안하며, 이는 의료용 대규모 언어 모델(MLLM)의 잠재 공간 추론을 위한 시각-의미 이중 감독 방법을 사용합니다. 보조 텍스트 디코더는 잠재 상태로부터 추론 과정을 재구성하고, 시각적 투영기는 독립적인 의료 영상 인코더에서 추출된 관심 영역(ROI) 특징을 회귀 예측합니다. 두 모듈 모두 추론 과정에서는 제거되어 오버헤드가 없지만, 필요에 따라 후처리하여 텍스트 및 시각적 설명을 제공함으로써 효율성을 저해하지 않고 이중 해석 가능성을 확보할 수 있습니다. 우리는 9가지 영상 모달리티를 포함하는 61K 데이터셋을 구축했으며, 이는 기존의 의료 분야 시각 잠재 추론 데이터셋보다 훨씬 큰 규모입니다. 7개의 벤치마크 실험 결과, VITAL은 기반 모델 및 다른 잠재적 추론 방법들을 지속적으로 크게 능가하며, 방대한 데이터를 사용하여 학습된 의료 MLLM과 비교해도 경쟁력 있는 최첨단 결과를 달성했습니다.
Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead of chain-of-thought for medical VQA. However, existing methods suffer from modality collapse, insufficient visual supervision, and train-inference mismatch. Moreover, their opaque latent states offer no interpretability, which is critical in clinical applications. We propose VITAL, a latent-space reasoning framework for medical MLLMs with visual-semantic dual supervision: an auxiliary text decoder reconstructs reasoning chains from latent states, while a visual projector regresses ROI features from a frozen, independent medical vision encoder. Both modules are discarded at inference with zero overhead, yet can be re-attached post-hoc for dual interpretability, providing textual and visual explanations of the reasoning process without sacrificing efficiency. We construct a 61K dataset spanning 9 imaging modalities, exceeding prior medical visual latent reasoning datasets by an order of magnitude. Experiments on 7 benchmarks show that VITAL consistently and substantially outperforms the backbone, all latent reasoning baselines, and medical MLLMs trained on far larger data, achieving state-of-the-art results competitive with trillion-parameter proprietary models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.