2605.28192v1 May 27, 2026 cs.AI

에이전트 기반 능동적 다중 모달 인식을 통한 멀티홉 오디오-비주얼 추론

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

Yu Wang
Yu Wang
Citations: 58
h-index: 2
Yanfeng Wang
Yanfeng Wang
Citations: 366
h-index: 10
Ke Xu
Ke Xu
Citations: 13
h-index: 2
Ziyang Cheng
Ziyang Cheng
Citations: 64
h-index: 4
Hongcheng Liu
Hongcheng Liu
Citations: 171
h-index: 6
Yuhao Wang
Yuhao Wang
Citations: 4
h-index: 1

멀티홉 오디오-비주얼 추론은 관련 증거가 종종 희소하고, 시간적으로 분산되어 있으며, 오디오 및 비주얼 스트림 전체에 걸쳐 분포하기 때문에 Omni-LLM에게 여전히 어려운 과제입니다. 기존 벤치마크는 이러한 설정을 제한적으로 조사하며, 일반적으로 제한된 수의 모달리티, 관련 시간 구간 또는 추론 단계를 포함합니다. 본 연구에서는 멀티홉 추론을 요구하는 시간적으로 분산된 오디오-비주얼 증거에 대한 신중하게 선별된 519개의 질문으로 구성된 벤치마크인 MOV-Bench를 소개합니다. MOV-Bench에서의 평가 결과, 현재의 Omni-LLM은 여전히 멀티홉 교차 모달 추론에 어려움을 겪고 있음을 보여줍니다. 이러한 과제를 해결하기 위해, 우리는 오픈 소스 Omni-LLM을 기반으로 구축된 효율적인 에이전트 프레임워크인 AOP-Agent를 제안합니다. 계층적 다중 모달 메모리와 협력적인 관찰-반성-재계획 루프를 결합하여, AOP-Agent는 추가 훈련이나 독점 모델 없이 오픈 소스 Omni-LLM이 능동적인 인식을 수행할 수 있도록 합니다. MOV-Bench 및 OmniVideoBench에서의 실험 결과, AOP-Agent가 일관되게 추론 성능을 향상시키며, 특히 긴 비디오와 추론 집약적인 질문에 대해 뚜렷한 개선 효과를 보임을 확인했습니다.

Original Abstract

Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited investigation of this setting, typically involving only a limited number of modalities, relevant temporal segments, or reasoning steps. In this work, we introduce MOV-Bench, a benchmark containing 519 carefully curated questions that require multi-hop reasoning over temporally dispersed audio-visual evidence. Evaluations on MOV-Bench reveal that current Omni-LLMs still struggle with multi-hop cross-modal reasoning. To address this challenge, we further propose AOP-Agent, an efficient agentic framework built on open-source Omni-LLMs for active omni-modal perception. By combining a hierarchical omni-modal memory with a collaborative observe-reflect-replan loop, AOP-Agent enables open-source Omni-LLMs to perform active perception without additional training or proprietary models. Experiments on MOV-Bench and OmniVideoBench demonstrate that AOP-Agent consistently improves reasoning performance, with particularly notable gains on long videos and reasoning-intensive questions.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!