MultiToP: 시각적 토큰 패치를 통한 환각 현상 완화를 위한 비디오 대규모 다중 모드 모델 학습
MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models
비디오 대규모 다중 모드 모델은 비디오 이해 분야에서 상당한 발전을 이루었지만, 입력 비디오에 의해 충실하게 뒷받침되지 않는 응답을 생성하는 환각 현상에 취약합니다. 본 논문에서는 시각적 토큰 패치를 통해 언어 생성이 이루어지기 전에 신뢰할 수 없는 시각적 토큰을 개선하여 환각 현상을 완화하는 다중 모드 컨텍스트 인식 프레임워크인 MultiToP를 제안합니다. MultiToP는 토큰 수준의 대체 분포를 예측하고 신뢰할 수 없는 시각적 토큰을 동적으로 생성된 전역 패치 토큰으로 선택적으로 대체하는 경량의 시각적 토큰 패처를 도입합니다. 패처를 효과적으로 학습하기 위해, 우리는 답변에 조건부로 프레임 수준 정보 힌트를 사용하여 토큰 대체 지침을 제공하는 정보 기반 순위 보정을 추가로 제안합니다. MultiToP는 정답 데이터 감독과 희소성 규제를 결합하여 원래 모델을 수정하지 않고도 지역적인 시각적 증거 개선을 가능하게 합니다. 광범위한 실험 결과, MultiToP는 Vript-HAL에서 환각 현상을 효과적으로 줄이며, 추론 오버헤드는 미미합니다. 또한 Qwen3-VL-4B-Instruct 모델의 F1 점수를 50.60% 향상시켰습니다. 동시에, MultiToP는 일반적인 비디오 이해 능력을 유지하여 Video-LLaVA-7B 모델에서 ActivityNet-QA 데이터셋에 대한 상대적 정확도를 18.58% 향상시켰습니다.
Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully supported by the input video. In this paper, we propose MultiToP, a multimodal-context-aware visual token patching framework that mitigates hallucinations by refining unreliable visual tokens before language generation. MultiToP introduces a lightweight Visual Token Patcher to predict token-level replacement distributions and selectively substitute unreliable visual tokens with a dynamic global patch token. To train the patcher effectively, we further propose information-guided rank calibration, which uses answer-conditioned frame-level information cues derived from the backbone to guide token replacement. Combined with ground-truth answer supervision and sparsity regularization, MultiToP enables localized visual evidence refinement without modifying the original model. Extensive experiments demonstrate that MultiToP effectively reduces hallucinations on Vript-HAL with negligible inference overhead, improving the F1 scores of Qwen3-VL-4B-Instruct by 50.60% over the vanilla model. Meanwhile, MultiToP preserves general video understanding ability, yielding an 18.58% relative accuracy gain on ActivityNet-QA for Video-LLaVA-7B.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.