정확성 너머: OCR 성능이 중요한 멀티모달 대규모 언어 모델 추론에서의 시각적 토큰 가지치기 과정에 대한 공간적 출처 감사
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
시각적 토큰 가지치기는 일반적으로 고정된 유지 비율에서 답변 품질을 기준으로 평가됩니다. 그러나 텍스트가 풍부한 멀티모달 대규모 언어 모델(MLLM)의 경우, 이 방법은 다음과 같은 중요한 실패 요소를 간과할 수 있습니다. 바로 답변이 정확하더라도, 해당 답변에 대한 근거가 되는 작은 OCR 영역과의 지역적 연관성이 전혀 없을 수 있다는 점입니다. 본 연구에서는 이러한 사각지대를 제거하기 위해, 답변 행동과 기하학적인 토큰 출처 정보, 개입 방법 및 실제 비용을 결합한 증거 기반 위험 감사 시스템을 제안합니다. 이 시스템은 투명하고 학습이 필요 없는 선택기를 사용하여, 제어 가능한 운영 지점을 식별합니다. 이미지 데이터가 서로 독립적인 환경에서 30% 유지 비율로 설정된 Qwen Target 모델의 정확도는 0.786으로 나타났습니다. 이는 Full 모델(0.783)보다 약간 높지만 (짝지어진 이미지 클러스터 차이: +0.003, 95% 신뢰 구간 [-0.014, +0.020]), 동일한 유지 비율을 가진 Target, Random 및 Grid 방식은 양성 지원 범위(positive-support coverage)에서 현저하게 다른 결과를 보입니다 (각각 0.620, 0.270, 0.318). Qwen3-VL-8B, LLaVA-1.5-7B 및 InternVL3.5-8B 모델에 대한 실험 결과, 제어 그룹, 개입 실험, 검출기 테스트 및 외부 방법을 통해 각 모델별 품질-위험-추적 가능성 경계를 파악할 수 있었으며, 이는 정확도만으로는 확인할 수 없는 중요한 정보입니다. 구체화된 접두사(materialized prefixes)를 사용하면 배치 전처리 속도가 최대 4.32배 향상되고, 최대 메모리 사용량이 76.4% 감소했습니다. 또한, 전체 검증을 거친 TextVQA 및 DocVQA 데이터셋을 사용하여, 특정 모델에서 효과적인 대상 검증 지점이 항상 일반적인 작업 성능 향상을 의미하지는 않음을 확인했습니다. 따라서 시각적 토큰 가지치기 과정에서는 품질과 압축 효율성 외에도, 남아있는 공간적 출처 정보와 실제 비용에 대한 보고가 반드시 포함되어야 합니다.
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.