2607.05978v1 Jul 07, 2026 cs.CV

제안 및 참여: 다중 토큰 로컬화된 어텐션을 이용한 학습 불필요한 MLLM 근거 신뢰도 측정

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention

Tal Remez
Tal Remez
Citations: 22,001
h-index: 23
Daniel Shalam
Daniel Shalam
Citations: 31
h-index: 2
Emanuel Ben Baruch
Emanuel Ben Baruch
Citations: 886
h-index: 4
Avi Ben Cohen
Avi Ben Cohen
Citations: 0
h-index: 0

다중 모드 대규모 언어 모델(MLLM)은 객체에 대한 경계 상자, 비디오 및 오디오 이벤트에 대한 시간 창과 같은 지역화된 예측을 수행할 수 있지만, 이러한 영역을 과도하게 생성하는 현상이 발생합니다. 모델 자체의 토큰 로그 확률은 거의 유용한 정보를 제공하지 못하며, 이는 근거 품질과 입력 모호성을 혼동하거나, 모델이 특정 결정을 내리면 좌표 토큰이 거의 결정론적으로 변하기 때문입니다. 본 연구에서는 다중 토큰 로컬화된 어텐션(MTLA)을 제안합니다. MTLA는 학습 없이 사후적으로 예측의 토큰들이 해당 영역에 얼마나 집중하는지를 측정하는 지표입니다. 기존의 어텐션을 기반으로 하는 검출기들은 전체 입력 모달리티에 대한 어텐션을 합산하고 단일 응답 토큰을 읽는데, 이는 더 약한 형태입니다. 우리는 제안된 영역 내에서만 어텐션을 합산하고 모든 예측 토큰에 대해 집계하면 더욱 강력한 근거 신호를 얻을 수 있음을 보여줍니다. 동일한 방법은 이미지의 객체 검출 및 비디오/오디오에서의 시간 지역화와 같은 다른 모달리티 및 작업에도 거의 적용될 수 있습니다. 여러 MLLM 계열과 세 가지 모달리티에 걸쳐 MTLA는 기존의 학습 불필요한 최적 기준보다 환각 AUROC를 +7에서 +38% 향상시켰습니다. 재순위를 위한 신뢰도 점수로 사용하면 오픈 소스 8B 범용 모델의 제로샷 COCO 검출 AP가 거의 두 배(20.4에서 37.0으로) 증가하여, 특정 작업에 대한 학습 없이 지도 학습 기반 검출기와 격차를 줄였습니다.

Original Abstract

Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prolifically. The model's own token log-probabilities are nearly uninformative: they conflate grounding quality with input ambiguity, and coordinate tokens become near-deterministic once the model commits. We propose Multi-Token Localized Attention (MTLA): a training-free, post-hoc score that measures how strongly a prediction's tokens attend to the region they claim. Prior attention-based detectors, which sum attention over the entire input modality and read a single response token, are weaker special cases; we show that summing only within the claimed region and aggregating across all prediction tokens recovers a stronger grounding signal. The same recipe applies almost trivially to other modalities and tasks: object detection in images and temporal localization in video and audio. Across multiple MLLM families and three modalities, MTLA improves hallucination AUROC by +7 to +38 over the best prior training-free baseline. Used as a confidence score for re-ranking, it nearly doubles the zero-shot COCO detection AP of an open-source 8B generalist (from 20.4 to 37.0), narrowing the gap to supervised detectors without any task-specific training.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!