2607.27667v1 Jul 30, 2026 cs.CV

증거 포트폴리오: 단일 프리필 위험 감지를 통한 폐쇄형 다중 모드 답변 평가

Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers

Fexiang Liu
Fexiang Liu
Citations: 0
h-index: 0
Shiye Wang
Shiye Wang
Citations: 0
h-index: 0
Qiang Qiu
Qiang Qiu
Citations: 115
h-index: 4
Zheng Wang
Zheng Wang
Citations: 0
h-index: 0

다중 모드 대규모 언어 모델(MLLM)의 안정적인 활용을 위해서는, 신뢰도가 높은 시각적 답변이 실제로 신뢰할 만한지, 검토가 필요한지, 아니면 더 강력한 시스템으로 연결되어야 하는지를 결정해야 합니다. 신뢰도 점수는 후보 답변의 여백을 나타내지만, 해당 여백과 관련된 추정된 부호화된 시각적 정보가 어디에서 비롯되었는지 또는 어떻게 분포하는지는 알려주지 않습니다. 본 연구에서는 동일한 화이트박스 프리필 경로를 사용하여 폐쇄형 시각적 답변에 대한 추론 시간 위험 감지를 수행합니다. 증거 포트폴리오(WEP)는 먼저, 예측된 후보 답변을 지지하거나 반박하는 시각적 요소를 계층별로 추정합니다. 이러한 요소들은 두 가지 해석 가능한 경로 그룹으로 요약됩니다. 즉, 질문 관련 증거 출처와 부호화된 증거 집중도입니다. 중첩된 그룹 검증은 더 신뢰할 수 있는 경로 그룹과 희소한 상위 k개 경로 포트폴리오를 선택하며, 이는 후보 답변의 신뢰도 점수와 결합됩니다. WEP는 이미지 변조, 디코딩 변경, 역방향 패스 또는 외부 검증기가 필요하지 않습니다. 세 가지 MLLM 모델 및 네 가지 이진 답변 벤치마크에서, WEP는 평균 오류 AP를 0.134만큼 향상시켰습니다. 모든 모델-데이터셋 조합에서 성능이 향상되었으며, 이미지 클러스터 부트스트랩 구간은 10쌍에서 엄격하게 양수 값을 보였습니다. WEP는 화이트박스 폐쇄형 답변 시스템을 대상으로 하며, 레이블된 교정 데이터 세트를 사용합니다.

Original Abstract

Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or routed to a stronger system. Confidence scores capture candidate margins, but not where the estimated signed visual readouts associated with those margins come from or how they are distributed. We study inference-time risk detection for closed visual answers using the same white-box prefill path that produces the answer. Witness Evidence Portfolios (WEP) first estimates, layer by layer, which visual contributions support or contradict the predicted candidate. It summarizes these contributions through two interpretable route families: question-related evidence provenance and signed evidence concentration. Nested grouped validation chooses the more reliable family and a sparse top-k route portfolio, which is fused with candidate confidence. WEP needs no image perturbation, decoding change, backward pass, or external verifier. Across three MLLMs and four binary-answer benchmarks, WEP improves mean error AP by 0.134. All 12 model--dataset gains are positive, and image-cluster bootstrap intervals are strictly positive on 10 pairs. WEP targets white-box closed-answer systems and uses a labeled calibration slice.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!