2607.11862v1 Jul 13, 2026 cs.CV

증거 기반 비디오 질의 응답

Evidence-Backed Video Question Answering

Juan Carlos Niebles
Juan Carlos Niebles
Salesforce
Citations: 26,446
h-index: 66
Ziyang Wang
Ziyang Wang
Citations: 636
h-index: 10
Silvio Savarese
Silvio Savarese
Citations: 4,449
h-index: 31
Caiming Xiong
Caiming Xiong
Citations: 3,801
h-index: 27
Shijie Wang
Shijie Wang
Brown University
Citations: 487
h-index: 7
Honglu Zhou
Honglu Zhou
Citations: 290
h-index: 6
Ran Xu
Ran Xu
Citations: 161
h-index: 3
Chen Sun
Chen Sun
Citations: 25,502
h-index: 45

현재 비디오 거대 언어 모델(Video LLM)은 질의 응답(QA)에서 뛰어난 성능을 보이지만, 대부분 검증 가능한 시각적 근거 없이 텍스트 답변만 제공하는 블랙박스 형태로 운영됩니다. 기존의 설명 가능성 연구는 주로 텍스트 기반 추론이나 제한적인 바운딩 박스를 사용하는데, 이는 가려짐 현상(occlusion)이나 비정형 변형과 같은 복잡한 비디오 동적 특성을 제대로 반영하기 어렵습니다. 본 논문에서는 모델이 의미 있는 답변과 함께 정확한 시공간적 증거(시간 구간 및 밀집된 객체 분할 마스크)를 동시에 출력하도록 요구하는 새로운 작업인 '증거 기반 비디오 질의 응답(E-VQA)'을 제안합니다. 이를 지원하기 위해, 우리는 판별 모델과 생성 모델 모두에 대한 픽셀 수준의 정확한 근거를 제공하는 최초의 인간 검증 벤치마크인 ST-Evidence를 소개합니다. 최첨단 모델의 평가 결과, QA 정확도와 실제 시각적 인식 능력 간에 중요한 격차가 존재하며, 단순히 규모를 키워서는 이 격차를 해소할 수 없다는 것을 확인했습니다. 이러한 문제를 해결하기 위해, 우리는 고수준 추론과 세밀한 근거 사이를 연결하는 16만 크기의 데이터셋인 ST-Evidence-Instruct를 생성하는 확장 가능한 자동화 파이프라인을 개발했습니다. 이 데이터를 사용하여 학습된 근거 기반 Video LLM은 동일 규모의 UniPixel 모델보다 상당한 성능 향상(예: 7B 모델에서 t-mean +27.2, J&F +13.8)을 보이며, 설명 가능하고 증거 기반 비디오 이해를 위한 견고한 기준점을 제시합니다. 코드와 데이터는 https://github.com/SalesforceAIResearch/EVQA 에서 확인할 수 있습니다.

Original Abstract

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. To support this, we introduce ST-Evidence, the first human-verified benchmark for both discriminative and generative pixel-level grounding. Evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and true visual perception that scaling alone fails to bridge. To address this, we develop scalable, automated generation pipelines to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding. Fine-tuning grounded Video LLMs on this data yields substantial gains over the corresponding size-matched UniPixel baselines (e.g., +27.2 t-mean and +13.8 J&F on a 7B model), establishing a robust baseline for explainable, evidence-backed video understanding. Code and data are available at https://github.com/SalesforceAIResearch/EVQA.

0 Citations
0 Influential
50 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!