2608.07959v1 Aug 08, 2026 cs.AI

SCOUT: 자체 점검 및 복구 기능을 갖춘 도구 기반 추론 에이전트를 활용한 초장기 일인칭 영상 이해

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

Yanhao Zhang
Yanhao Zhang
Citations: 119
h-index: 5
Keyang Zhong
Keyang Zhong
Citations: 84
h-index: 4
Junlin Xie
Junlin Xie
Citations: 129
h-index: 4
Guanbin Li
Guanbin Li
Citations: 3,856
h-index: 32
Kuo Wang
Kuo Wang
Citations: 3
h-index: 1
Peng Liu
Peng Liu
Citations: 27
h-index: 3
Quanlong Zheng
Quanlong Zheng
Citations: 12
h-index: 2
Zhijia Liang
Zhijia Liang
Citations: 3
h-index: 1

초장기 일인칭 영상 이해는 시간적으로 분산된 정보를 바탕으로 수 시간 또는 수일에 걸쳐 추론해야 하며, 이는 제한적인 맥락과 핵심 영상 구간의 연결을 가진 기존 다중 모드 모델에 어려움을 야기합니다. Chain-of-Tool-Thought (CoTT) 에이전트 시스템은 반복적인 검색 및 검사를 가능하게 하지만, 경직된 확대 전략으로 인해 오류가 발생하기 쉽고 복구 메커니즘이 부족합니다. 본 연구에서는 SCOUT (Self-Checking Chain-Of-Tool-thought), 즉 복구 기능을 갖춘 에이전트 기반 프레임워크를 통해 이러한 문제점을 해결합니다. SCOUT는 중간 단계의 도구 사용 결과를 평가하는 적응형 정책을 도입하여, 탐색(지역 전환)과 활용(확대) 간 균형을 동적으로 조절함으로써 매우 긴 시간 범위에 걸쳐 안정적인 다중 단계 추론을 가능하게 합니다. 그러나 이러한 다단계 도구 사용 에이전트를 훈련하는 것은 여전히 어렵습니다. 기존 강화 학습 방법은 희소한 결과 수준의 보상에 의존하며, 확장된 의사 결정 경로에 대한 감독이 부족하여 장기적 추론에 대한 최적화된 보상 할당을 달성하기 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 불확실성을 우선시하는 정책 최적화 방법인 UPS-GRPO를 개발했습니다. 이 방법은 높은 불확실성을 가진 도구 사용 후 상태에 탐색을 집중하면서도 샘플 효율성을 유지합니다. 또한, 결과 보상과 도구 기반 시간 정렬 보상을 통합하여 단계별 성능 향상을 위한 보상 할당 방식을 개선했습니다. 실험 결과, SCOUT는 초장기 일인칭 영상 이해 벤치마크에서 최첨단 결과를 달성했으며, 더 짧은 시간 범위의 장편 영상에서도 경쟁력 있는 성능을 보여주었습니다.

Original Abstract

Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent systems enable iterative retrieval and inspection, they suffer from error propagation due to rigid zoom-in strategies that lack recovery mechanisms. In this work, we address these challenges through SCOUT (Self-Checking Chain-Of-Tool-thought), a recovery-aware agentic framework introducing an adaptive policy that evaluates intermediate tool observations and dynamically trades off exploitation (zoom-in) and exploration (region switching), enabling robust multi-hop reasoning over extremely long horizons. However, training such multi-turn tool-using agents remains challenging, as existing RL methods rely on sparse outcome-level rewards and lack supervision over extended decision trajectories, resulting in suboptimal credit assignment for long-horizon reasoning. To address this, we develop UPS-GRPO, an uncertainty-prioritized policy optimization method that concentrates exploration on high-uncertainty post-tool states while preserving sample efficiency. We further introduce a turn-level advantage decomposition that integrates outcome rewards with tool-grounded temporal alignment rewards for improved credit assignment. Experiments show that SCOUT achieves state-of-the-art results on ultra-long egocentric benchmarks, while remaining competitive on shorter-horizon long-video settings.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!