2607.08233v1 Jul 09, 2026 cs.AI

ZendoWorld 활용: 능동적인 시각적 개념 추론을 통한 인공지능 에이전트 도전

Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

Kevin Ellis
Kevin Ellis
Citations: 0
h-index: 0
Wasu Top Piriyakulkij
Wasu Top Piriyakulkij
Citations: 0
h-index: 0
Wolfgang Stammer
Wolfgang Stammer
Citations: 188
h-index: 6
Sophia Koehler
Sophia Koehler
Citations: 0
h-index: 0
Antonia Wust
Antonia Wust
Citations: 38
h-index: 3
Inga Ibs
Inga Ibs
Citations: 10
h-index: 2
Constantin A. Rothkopf
Constantin A. Rothkopf
Citations: 43
h-index: 3
Kristian Kersting
Kristian Kersting
Citations: 34
h-index: 3

지능형 시스템 구축의 핵심 과제는 에이전트가 복잡한 입력을 동시에 인식하고, 숨겨진 패턴에 대한 가설을 설정하며, 이를 검증하기 위한 유용한 실험을 설계할 수 있도록 하는 것입니다. 이러한 문제를 연구하기 위해, 우리는 ZendoWorld를 제안합니다. ZendoWorld는 에이전트가 시각적인 게임 관찰로부터 논리적 규칙을 추론하고, 새로운 장면을 제시하여 정보를 획득하며, 게임 환경으로부터 받은 피드백을 바탕으로 가설을 개선하는 제어된 상호작용 환경입니다. 우리는 순수 VLM 추론, 베이지안 파티클 필터링, 동적 개념 발견 및 신경-기호 방식 등 다양한 에이전트를 평가했습니다. 주요 결과는 다음과 같습니다 (1) 관찰된 예시에 대한 레이블 예측의 높은 정확도가 반드시 기본 규칙을 복구하는 것을 의미하지 않습니다; (2) 인식과 유도는 서로 다른 에이전트 클래스에 대해 명확한 병목 현상으로 작용합니다; (3) VLM 기반 에이전트는 거의 정보가 없는 실험을 제안하여, 적극적으로 가설의 불확실성을 줄이는 데 실패합니다. 이러한 결과를 비교하기 위해, 우리는 이 작업에 대한 인간 데이터를 수집했으며, 이는 유도적 추론, 특히 더 복잡한 규칙에 대한 격차를 보여줍니다. 전반적으로, ZENDOWORLD는 지능형 에이전트 평가를 위한 중요한 발판을 마련하며, 과학적 발견과 같은 영역에서 개선될 수 있는 구체적인 방법을 제시합니다.

Original Abstract

A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments to test them. To study this problem, we propose ZendoWorld, a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment. We evaluate several agents spanning pure VLM reasoning, Bayesian particle filtering, dynamic concept discovery, and neuro-symbolic methods. Our main findings are: (1) high accuracy in predicting labels for observed examples does not imply recovery of the underlying rule; (2) perception and induction are distinct bottlenecks for different agent classes; and (3) VLM-based agents propose near-uninformative experiments, failing to actively reduce hypothesis uncertainty. To compare these results, we collect human data on the task, which reveals a gap in inductive reasoning, particularly for more complex rules. Overall, ZENDOWORLD takes an important step toward evaluating intelligent agents and identifies concrete avenues for improvement, particularly in domains like scientific discovery.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!