에이전트 기반 능동적 다중 모달 인식을 통한 멀티홉 오디오-비주얼 추론
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
멀티홉 오디오-비주얼 추론은 관련 증거가 종종 희소하고, 시간적으로 분산되어 있으며, 오디오 및 비주얼 스트림 전체에 걸쳐 분포하기 때문에 Omni-LLM에게 여전히 어려운 과제입니다. 기존 벤치마크는 이러한 설정을 제한적으로 조사하며, 일반적으로 제한된 수의 모달리티, 관련 시간 구간 또는 추론 단계를 포함합니다. 본 연구에서는 멀티홉 추론을 요구하는 시간적으로 분산된 오디오-비주얼 증거에 대한 신중하게 선별된 519개의 질문으로 구성된 벤치마크인 MOV-Bench를 소개합니다. MOV-Bench에서의 평가 결과, 현재의 Omni-LLM은 여전히 멀티홉 교차 모달 추론에 어려움을 겪고 있음을 보여줍니다. 이러한 과제를 해결하기 위해, 우리는 오픈 소스 Omni-LLM을 기반으로 구축된 효율적인 에이전트 프레임워크인 AOP-Agent를 제안합니다. 계층적 다중 모달 메모리와 협력적인 관찰-반성-재계획 루프를 결합하여, AOP-Agent는 추가 훈련이나 독점 모델 없이 오픈 소스 Omni-LLM이 능동적인 인식을 수행할 수 있도록 합니다. MOV-Bench 및 OmniVideoBench에서의 실험 결과, AOP-Agent가 일관되게 추론 성능을 향상시키며, 특히 긴 비디오와 추론 집약적인 질문에 대해 뚜렷한 개선 효과를 보임을 확인했습니다.
Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited investigation of this setting, typically involving only a limited number of modalities, relevant temporal segments, or reasoning steps. In this work, we introduce MOV-Bench, a benchmark containing 519 carefully curated questions that require multi-hop reasoning over temporally dispersed audio-visual evidence. Evaluations on MOV-Bench reveal that current Omni-LLMs still struggle with multi-hop cross-modal reasoning. To address this challenge, we further propose AOP-Agent, an efficient agentic framework built on open-source Omni-LLMs for active omni-modal perception. By combining a hierarchical omni-modal memory with a collaborative observe-reflect-replan loop, AOP-Agent enables open-source Omni-LLMs to perform active perception without additional training or proprietary models. Experiments on MOV-Bench and OmniVideoBench demonstrate that AOP-Agent consistently improves reasoning performance, with particularly notable gains on long videos and reasoning-intensive questions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.