2608.04587v1 Aug 05, 2026 cs.CV

MetaVideoAgent: 장편 비디오 이해를 위한 자동화된 비디오 에이전트 진화

MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding

Haiwen Hong
Haiwen Hong
Citations: 238
h-index: 5
Longtao Huang
Longtao Huang
Citations: 179
h-index: 7
Junjie Li
Junjie Li
Citations: 286
h-index: 4
Yuwen Zhai
Yuwen Zhai
Citations: 16
h-index: 3
Benlei Cui
Benlei Cui
Citations: 225
h-index: 5
Ruijian Jia
Ruijian Jia
Citations: 0
h-index: 0
Pengfei Sun
Pengfei Sun
Citations: 8
h-index: 2
Ruize Wang
Ruize Wang
Citations: 0
h-index: 0
Jinhao Chen
Jinhao Chen
Citations: 33
h-index: 2
Ying-Hao Chen
Ying-Hao Chen
Citations: 0
h-index: 0
Jingqun Tang
Jingqun Tang
Citations: 52
h-index: 4
Weiwei Wu
Weiwei Wu
Citations: 24
h-index: 1

장편 비디오 이해는 긴, 다중 모달 비디오에서 질문과 관련된 희소한 증거를 찾는 것을 필요로 합니다. 실제 비디오 데이터의 분포는 모드별 정보 밀도, 콘텐츠 구조 및 증거 패턴 측면에서 다르기 때문에, 고정된 비디오 에이전트 설계는 불필요한 처리량을 발생시키거나 일치하지 않는 경우 실패할 수 있습니다. 텍스트 기반 자동 에이전트 진화를 비디오로 확장하는 것은 전체 장편 비디오 실행으로 인해 후보 검증 비용이 많이 들고, 오류가 연결된 증거 처리 단계에 걸쳐 전파되며, 복잡한 전처리, 인식 도구 및 지역화 전략으로 인해 코드 수준 업데이트를 안정적으로 구현하기 어렵기 때문에 어려운 문제입니다. 저희는 대상 분포에 대한 비디오 에이전트를 자동으로 진화시키는 프레임워크인 MetaVideoAgent를 소개합니다. 이 프레임워크는 희소하게 샘플링된 프레임과 관련된 질문에서 정보 밀도 및 증거 요구 사항을 파악하여 초기 설계를 안내하고, 지역화된 오류를 독립적으로 실행 가능한 최소 검증 작업으로 압축합니다. 이 프레임워크는 증거 기반의 Gold Path를 구성하고, Student 트랙션을 감사하며, 샘플 전체에 걸쳐 반복되는 실패를 집계하고, 해당 모듈에 귀속시킵니다. 모듈식 에이전트 표현은 각 업데이트를 주요 책임 모듈과 필요한 종속성으로 제한합니다. 저희는 또한 8개의 비디오 분포를 포함하는 VA-EvoBench 데이터셋을 소개하며, 각 분포별로 진화 및 검증 세트를 분리했습니다. 각 분포에 대해 4번의 진화 단계를 거친 MetaVideoAgent는 모든 초기 에이전트의 성능을 향상시키고, 평균 정확도를 38.44%에서 51.47%로 높였습니다. 또한, 진화 비용은 분포당 평균 354만 개의 토큰으로 매우 효율적입니다. 진화된 에이전트는 비교 대상이었던 가장 강력한 기존 고정 설계 비디오 에이전트보다 6.39% 더 높은 성능을 보였으며, 질문 당 사용되는 토큰 및 비디오 프레임 수가 가장 적었습니다. 저희는 재현 가능한 연구를 지원하기 위해 모든 코드와 데이터를 공개할 예정입니다.

Original Abstract

Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific information density, content structure, and evidence patterns, causing fixed video-agent designs to incur redundant processing or fail when mismatched. Extending automated agent evolution from text to video is challenging because full long-video execution makes candidate validation expensive, failures propagate across coupled evidence-processing stages, and complex preprocessing, perception tools, and localization strategies make code-level updates difficult to implement reliably. We introduce MetaVideoAgent, a framework that automatically evolves a video agent for a target distribution. It profiles information density and evidence requirements from sparsely sampled frames and associated queries to guide initial design, then compresses localized failures into independently executable minimal validation tasks. It constructs evidence-grounded Gold Paths, audits Student trajectories, aggregates recurring failures across samples, and attributes them to responsible modules. A modular agent representation constrains each update to the primary responsible module and its necessary dependencies. We further introduce VA-EvoBench, covering eight video distributions with separate evolution and held-out splits. With four evolution iterations per distribution, MetaVideoAgent improves every initial agent and raises macro-average accuracy from 38.44% to 51.47%, at an average evolution cost of 3.54M tokens per distribution. The evolved agents outperform the strongest prior fixed-design video agent by 6.39 percentage points while using the fewest tokens and video frames per question among the compared video agents. We will release all code and data to support reproducible research.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!