OmniGameArena: VLM 게임 에이전트를 위한 통일된 UE5 벤치마크 및 성능 개선 동역학
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
비전-언어 모델(VLM) 에이전트가 인터랙티브한 게임 환경에서 점점 더 많이 활용되고 있습니다. 그러나 VLM 에이전트를 위한 게임 벤치마크는 일반적으로 (에이전트, 게임) 쌍마다 단일의 초기 시도 점수를 보고하고, 단일 에이전트의 솔로 플레이에 초점을 맞추며, 상용 VLM, 오픈 가중 VLM 및 특수화된 게임 정책 등 다양한 에이전트 클래스를 동일한 기준으로 평가할 수 있는 통일된 프로토콜이 부족합니다. 우리는 이러한 격차를 OmniGameArena라는 실시간 벤치마크로 해결하고자 합니다. OmniGameArena는 새로운 Unreal Engine 5 기반의 12개의 게임(솔로 모드 7개, PvP 3개, 협동 모드 2개)으로 구성되어 있으며, 통일된 액션 인터페이스를 제공합니다. 또한 Improvement Dynamics Curve (IDC)라는 에이전트 기반 피드백 시스템을 도입하여, 도구를 사용하는 LLM이 여러 라운드를 거치면서 제한된 스킬 프롬프트를 자율적으로 개선하도록 합니다. IDC는 초기 점수 외에도 각 (에이전트, 게임) 쌍에 대해 두 가지 추가적인 정보를 제공합니다. 첫째, 피드백 라운드에 따른 점수의 변화 추세이고, 둘째, 학습된 스킬이 독립적인 작업 변형에서 어떻게 동작하는지를 보여줍니다. 우리는 OmniGameArena의 초기 점수 리더보드에 등록된 12개의 VLM 에이전트와 IDC를 통해 평가된 상위 4개 에이전트에 대한 이러한 정보들을 제공합니다.
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair, focus on single-agent Solo play, and lack unified protocols for evaluating heterogeneous agent classes (commercial VLMs, open-weight VLMs, and specialized game policies) on the same footing. We address these gaps with OmniGameArena, a real-time benchmark of twelve newly built Unreal Engine 5 games spanning Solo (7), PvP (3), and Coop (2) with unified action interfaces, and the Improvement Dynamics Curve (IDC), an agentic-reflection harness in which a tool-using reflector LLM autonomously refines a bounded skill prompt across multiple rounds. Beyond cold-start leaderboard scores, IDC exposes two additional observables for each (agent, game) pair: how the score evolves across reflection rounds, and how the learned skill behaves on held-out task variants. We report these observables for twelve VLM agents on the cold-start leaderboard and four top agents under IDC.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.