강력한 비디오 이해를 위한 신뢰도 기반 도구 조율
Confidence-Aware Tool Orchestration for Robust Video Understanding
비디오 추론 언어 모델은 모든 입력 프레임이 동일하게 신뢰할 수 있다고 암묵적으로 가정합니다. 이는 우리가 '맹목적 신뢰 문제'라고 부르는 현상으로 이어집니다. 실제 환경에서 발생하는 모션 블러, 눈부심 또는 가려짐과 같은 교란에 의해 최첨단 비디오 추론 모델은 실제 로봇 시스템 벤치마크에서 최대 15-30%p의 정확도 저하를 경험하는 반면, 시각적 증거가 손상되었다는 사실을 인지하지 못합니다. 이러한 문제를 해결하기 위해, 우리는 각 추론 단계에 프레임별 신뢰도를 명시적으로 통합하는 에이전트 기반 비디오 이해 프레임워크인 Robust-TO를 제안합니다. Robust-TO는 다양한 시각 인식 도구를 통일된 증거 인터페이스 하에서 구성합니다. 각 도구는 원래 질문에서 파생된 부분적인 쿼리와 신뢰도-관련성 점수로 선택된 일련의 신뢰할 수 있는 프레임을 받습니다. 그리고 공유 형식을 갖춘 증거를 반환하는데, 이는 구체적인 예측(예: 바운딩 박스, 움직임 경로, 인식된 텍스트 또는 동작 레이블), 시간적 위치 정보 및 보정된 신뢰도 점수를 포함합니다. 추론 과정에서 이러한 보정된 점수는 세 단계의 합성 프로세스(높음/중간/낮음)에서 증거 가중치를 안내하며, 정확성, 증거 신뢰도 및 효율성을 동시에 최적화하는 신뢰도-비용 GRPO (Gain, Risk, Penalty, Optimization) 보상을 정의합니다. 두 가지 비디오 추론 벤치마크와 관련된 여덟 가지 작업에서 Robust-TO는 깨끗한 입력에 대해 평균 56.4%의 정확도를 달성하여 가장 강력한 오픈 소스 기준 모델보다 10.6%p 높고, Gemini-2.5-Pro (46.2%)를 능가합니다. 다섯 가지 현실적인 왜곡 유형 하에서 Robust-TO는 평균 54.3%의 정확도를 유지하며, 이는 가장 강력한 오픈 소스 기준 모델보다 5.8%p 높은 수치입니다. 또한 비교된 모든 방법 중에서 깨끗한 입력과 손상된 입력 간의 정확도 감소폭이 가장 작습니다.
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: under realistic perturbations such as motion blur, glare, or occlusion, frontier video reasoning models can suffer 15-30%p accuracy drops on real-world embodied benchmarks, while remaining unaware that their visual evidence has been degraded. To address this challenge, we propose Robust-TO, an agentic video understanding framework that explicitly integrates per-frame trustworthiness into every stage of reasoning. Robust-TO organizes heterogeneous visual perception tools under a unified evidence interface. Each tool receives a sub-query derived from the original question and a set of trustworthy frames selected by the reliability-relevance score. It returns evidence in a shared format: a concrete prediction (e.g., a bounding box, motion trajectory, recognized text, or action label), temporal grounding, and a calibrated reliability score. During reasoning, these calibrated scores guide evidence weighting in a three-tier synthesis process (high/medium/low) and define a confidence-cost GRPO reward that jointly optimizes correctness, evidence reliability, and efficiency. On two video reasoning benchmarks spanning eight tasks, Robust-TO achieves 56.4% average accuracy on clean inputs, surpassing the strongest open-source baseline by 10.6%p and outperforming Gemini-2.5-Pro (46.2%). Under five realistic corruption types, Robust-TO maintains 54.3% average accuracy, 5.8%p above the strongest open-source baseline, while exhibiting the smallest clean-to-corrupted accuracy drop among all compared methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.