2608.03979v1 Aug 04, 2026 cs.CV

Video-DeepResearch: 차세대 다중 모드 심층 연구 에이전트에 대한 제안

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Shiting Huang
Shiting Huang
Citations: 76
h-index: 5
Yu Zeng
Yu Zeng
Citations: 221
h-index: 9
Qingnan Ren
Qingnan Ren
Citations: 10
h-index: 2
Zhen Fang
Zhen Fang
Citations: 89
h-index: 5
Qisheng Su
Qisheng Su
Citations: 50
h-index: 3
Zehui Chen
Zehui Chen
Citations: 1,988
h-index: 11
Feng Zhao
Feng Zhao
Citations: 638
h-index: 9
Lionel Z. Wang
Lionel Z. Wang
Citations: 27
h-index: 2
Wenxuan Huang
Wenxuan Huang
Citations: 102
h-index: 6
Shuang Chen
Shuang Chen
Citations: 49
h-index: 4
Zhenfei Yin
Zhenfei Yin
Citations: 301
h-index: 7
Lin Chen
Lin Chen
Citations: 55
h-index: 3
Yiming Zhao
Yiming Zhao
Citations: 84
h-index: 5
Yao Hu
Yao Hu
Citations: 725
h-index: 7
Wanli Ouyang
Wanli Ouyang
Citations: 62
h-index: 5
Shaosheng Cao
Shaosheng Cao
Citations: 46
h-index: 4
Qingyu Yin
Qingyu Yin
Citations: 61
h-index: 3
Tianfei Ren
Tianfei Ren
Citations: 14
h-index: 2
Qiutinglu Lu
Qiutinglu Lu
Citations: 0
h-index: 0
Shaohui Lin
Shaohui Lin
Citations: 68
h-index: 4

본 논문에서는 정적인 이미지에서 연속적인 비디오 스트림으로 확장된 다중 모드 에이전트인 Video-DeepResearch (Video-DR)를 소개합니다. 이러한 설정은 밀집된 시공간적 연관성과 웹 검색을 결합해야 합니다. 초기 실험 결과, 현재 모델에는 두 가지 중요한 병목 현상이 존재함을 보여줍니다. 첫째는 모달리티 편향으로, 에이전트가 시각 도구 대신 텍스트 검색을 선호하는 경향입니다. 둘째는 매개변수 지식 유출로, 모델이 실제 도구를 활용한 실행보다는 내부 메모리에 의존합니다. 이러한 문제점을 해결하기 위해, 저희는 단계별 도구 사용 해제를 통해 프레임 간의 철저한 시각적 연관성을 확보하고 웹 검색을 수행하는 분리된 인지-탐색 파이프라인을 갖춘 Video-DR을 제안합니다. 저희의 프레임워크는 지도 학습 미세 조정과 Group Relative Policy Optimization (GRPO)이라는 두 단계로 구성된 학습 방식을 채택하여, 모방 학습의 한계를 뛰어넘는 자율적인 탐색을 가능하게 합니다. 또한, 200개의 복잡하고 다단계 질문 답변(VQA) 예제를 포함하는 인간-AI 협업 벤치마크인 Video-DR-Bench를 구축했습니다. 실험 결과, 저희가 개발한 Video-DeepResearch-35B-A3B 모델은 평균 정확도 64.0%라는 새로운 최고 성능을 달성했으며, 이는 독점 모델인 Claude-4.5-Sonnet (59.0%)보다 5.0% 포인트 높고 GPT-5 (52.5%) 및 Gemini 2.5 Pro (57.5%)보다도 훨씬 뛰어난 성능입니다. 30B-A3B 변형 모델은 59.3%의 정확도를 달성하여 Claude-4.5-Sonnet과 경쟁하며, 더 작은 규모에서도 저희의 학습 방식이 효과적임을 보여줍니다. 코드: https://github.com/Osilly/Vision-DeepResearch.

Original Abstract

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!