EgoBench: 도구 사용 에이전트를 위한 인터랙티브한 자아 중심 다중 모드 벤치마크
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
인공지능 에이전트가 점점 더 광범위하고 실제 환경에서 운영됨에 따라, 다중 모드 인식, 다단계 추론을 통한 도구 사용 및 사용자 상호 작용 간의 깊은 조화가 필요합니다. 그러나 기존 벤치마크는 엄격하게 결합된 다중 기능 작업을 설계하는 데 따르는 어려움, 자연스럽고 작업 제약 조건이 있는 사용자 피드백 시뮬레이션 및 동적 상호 작용에 대한 객관적인 평가 보장 때문에 이러한 기능을 함께 평가하는 데 실패합니다. 이러한 격차를 해소하기 위해 도구 사용 에이전트를 위한 최초의 인터랙티브 다중 모드 벤치마크인 EgoBench를 소개합니다. EgoBench는 네 가지 일상 시나리오를 포괄하는 1,045개의 자아 중심 비디오 기반 작업과 평가를 위한 사용자-에이전트-도구 상호 작용 환경으로 구성됩니다. 우리는 세 단계로 이루어진 시너지 파이프라인을 구현하여 각 작업을 설계함으로써 시각적 인식과 도구를 활용한 다단계 추론의 공동 적용을 장려합니다. 또한, EgoBench 내에서 에이전트의 상호 작용 능력을 평가하기 위한 멀티 에이전트 시뮬레이션 사용자를 개발하여 에이전트에 대해 높은 충실도와 작업에 맞는 응답을 생성합니다. 더욱이, 우리는 프로세스 기반 및 결과 기반 동등성을 통해 객관적인 평가를 보장하는 결정론적 통합 검증 프레임워크를 구축했습니다. EgoBench에서 8개의 최첨단 비디오-MLLM 에이전트를 벤치마킹한 결과 심각한 성능 한계가 드러났습니다. 가장 우수한 모델조차도 가장 성능이 좋은 시나리오에서 30.62%의 정확도를 달성했으며, 모든 네 가지 시나리오에서 평균 19.43%를 기록했습니다. 마지막으로, 우리는 다양한 오류 분석을 수행하여 실패 원인을 파악하고 향후 인공지능 에이전트 발전을 위한 능력적 제약 요소를 밝혀냈습니다.
As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to jointly evaluate these capabilities due to challenges in designing strictly coupled multi-capability tasks, simulating natural and task-constrained user feedback, and ensuring objective evaluation of dynamic interaction. To bridge this gap, we introduce EgoBench, the first interactive multimodal benchmark for tool-using agents. EgoBench comprises 1,045 egocentric-video-grounded tasks covering four daily scenarios, along with a user-agent-tool interactive environment for evaluation. We implement a three-stage synergistic pipeline through which each task is designed to enforce the joint application of visual perception and tool-augmented multi-hop reasoning. We additionally develop a multi-agent simulated user within EgoBench to evaluate agents' interaction capabilities, which generates high-fidelity, task-aligned responses to agents. Furthermore, we establish a deterministic joint validation framework that guarantees objective assessment through process-based and result-based equivalence. Benchmarking eight SOTA video-MLLM agents on EgoBench reveals a severe performance ceiling: the best model achieves only 30.62% accuracy in the best-performing scenario, averaging 19.43% across all four scenarios. Finally, we conduct a multi-dimensional error analysis to disentangle failure modes, exposing capability bottlenecks for advancing future AI agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.