CADER: 신뢰도 기반 동적 증거 추론을 통한 장편 영상 이해
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
장편 영상 이해는 점점 더 대규모의 시각-언어 모델과 도구 지원 추론에 의존하고 있지만, 대부분의 시스템은 난이도와 관계없이 모든 예시에 동일한 추론 절차를 적용합니다. 이러한 균일적인 전략은 쉬운 질문에 불필요한 도구 기반 처리를 유발하며, 어려운 질문에 대해 세밀한 시간적 증거가 필요할 때 제한적인 제어만 제공합니다. 우리는 CADER (Confidence-Aware Dynamic Evidence Reasoning), 즉 신뢰도 기반 동적 증거 추론이라는 학습이 필요 없는 프레임워크를 제안합니다. CADER은 먼저 균일하게 추출된 프레임을 사용하여 전반적인 추론을 수행하고, 로짓 마진 신호를 통해 답변의 신뢰도를 추정하여, 높은 신뢰도를 가진 예제는 초기에 종료될 수 있도록 합니다. 불확실한 예제의 경우, CADER은 시간적 잘라내기, 경량 의미 검증 및 관련성 기반 재샘플링을 결합하는 두 번째 단계의 도구 지원 루프를 활성화하여 질문과 관련된 증거를 점진적으로 찾습니다. 이러한 설계는 도구 사용을 샘플 수준의 결정으로 취급합니다. 즉, 하나의 전반적인 추론은 쉬운 경우를 처리하고, 추가적인 추론은 불확실성이 더 많은 증거가 필요하다는 것을 나타내는 예제에만 적용됩니다. 여러 VideoQA 벤치마크에서의 실험 결과, CADER은 장편 영상 추론 성능을 향상시키면서도 신뢰도가 높은 샘플의 경우 두 번째 단계를 건너뛸 수 있음을 보여줍니다. 또한, 도구 없이 체인-오브-생트(Chain-of-Thought) 방식으로만 학습된 기본 모델에 CADER을 적용했을 때, 특수하게 설계된 도구 지원 프레임워크와 경쟁력 있는 성능을 달성했습니다. 이는 적응적인 장편 영상 추론을 위한 실용적인 추론 시간 경로를 제시합니다.
Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides limited control when difficult questions require fine-grained temporal evidence. We propose CADER (Confidence-Aware Dynamic Evidence Reasoning), a training-free framework for adaptive and reliable long-video reasoning. CADER first performs global reasoning over uniformly sampled frames and estimates answer confidence with a logit-margin signal, allowing high-confidence examples to exit early. For uncertain examples, CADER activates a second-stage tool-augmented loop that combines temporal cropping, lightweight semantic verification, and Relevance-Guided Resampling to progressively localize question-relevant evidence. This design treats tool use as a sample-level decision: a single global pass handles easy cases, while additional reasoning is reserved for examples where uncertainty suggests that more evidence is needed. Experiments on multiple VideoQA benchmarks show that CADER improves long-video reasoning while bypassing Stage~2 for high-confidence samples. Moreover, when applied to a backbone trained only with tool-free chain-of-thought supervision, CADER achieves competitive performance against specialized tool-augmented frameworks, suggesting a practical inference-time route for adaptive long-video reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.