2602.05249v1 Feb 05, 2026 cs.AI

엠바디드 에이전트의 현장(In-Situ) 평가를 위한 자동 인지 과제 생성

Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents

Song Zhu
Song Zhu
Citations: 1
h-index: 1
Zhenliang Zhang
Zhenliang Zhang
Citations: 9
h-index: 2
Yujia Peng
Yujia Peng
Citations: 5
h-index: 1
Xinyi He
Xinyi He
Citations: 3
h-index: 1
Chuanjian Fu
Chuanjian Fu
Citations: 1
h-index: 1
Ying Yang
Ying Yang
Citations: 74
h-index: 3
Sihan Guo
Sihan Guo
Citations: 40
h-index: 3
Lifeng Fan
Lifeng Fan
Citations: 53
h-index: 2

범용 지능 에이전트가 다양한 가정 환경에 널리 배포될 준비가 됨에 따라, 각각의 고유하고 미지인 3D 환경에 맞춤화된 평가가 필수적인 전제 조건이 되었다. 그러나 기존 벤치마크들은 심각한 데이터 오염과 장면 특이성 부족 문제를 겪고 있어, 미지의 환경에서 에이전트의 능력을 평가하기에는 부적절하다. 이를 해결하기 위해, 우리는 인간의 인지 능력에서 영감을 받아 미지의 환경을 위한 동적 현장(in-situ) 과제 생성 방법을 제안한다. 우리는 구조화된 그래프 표현을 통해 과제를 정의하고, 엠바디드 에이전트를 위한 2단계 상호작용-진화 과제 생성 시스템(TEA)을 구축했다. 상호작용 단계에서 에이전트는 환경과 능동적으로 상호작용하며 과제 실행과 생성 간의 루프를 형성해 지속적인 과제 생성을 가능하게 한다. 진화 단계에서는 과제 그래프 모델링을 통해 외부 데이터 없이 기존 과제를 재조합하고 재사용하여 새로운 과제를 생성한다. 10개의 미지의 장면에서 진행된 실험 결과, TEA는 두 번의 사이클 동안 87,876개의 과제를 자동으로 생성했으며, 인간 검증을 통해 해당 과제들이 물리적으로 합당하고 필수적인 일상 인지 능력을 포괄하고 있음이 확인되었다. 우리가 생성한 현장 과제에 대해 SOTA 모델들을 인간과 비교 벤치마킹한 결과, 모델들은 공개 벤치마크에서는 뛰어난 성능을 보였음에도 불구하고, 기초적인 지각 과제에서는 놀라울 정도로 저조한 성능을 보였으며, 3D 상호작용 인식이 심각하게 결여되어 있고, 추론 시 과제 유형에 매우 민감한 것으로 드러났다. 이러한 엄중한 결과는 에이전트를 실제 인간 환경에 배포하기 전 현장 평가의 필요성을 강조한다.

Original Abstract

As general intelligent agents are poised for widespread deployment in diverse households, evaluation tailored to each unique unseen 3D environment has become a critical prerequisite. However, existing benchmarks suffer from severe data contamination and a lack of scene specificity, inadequate for assessing agent capabilities in unseen settings. To address this, we propose a dynamic in-situ task generation method for unseen environments inspired by human cognition. We define tasks through a structured graph representation and construct a two-stage interaction-evolution task generation system for embodied agents (TEA). In the interaction stage, the agent actively interacts with the environment, creating a loop between task execution and generation that allows for continuous task generation. In the evolution stage, task graph modeling allows us to recombine and reuse existing tasks to generate new ones without external data. Experiments across 10 unseen scenes demonstrate that TEA automatically generated 87,876 tasks in two cycles, which human verification confirmed to be physically reasonable and encompassing essential daily cognitive capabilities. Benchmarking SOTA models against humans on our in-situ tasks reveals that models, despite excelling on public benchmarks, perform surprisingly poorly on basic perception tasks, severely lack 3D interaction awareness and show high sensitivity to task types in reasoning. These sobering findings highlight the necessity of in-situ evaluation before deploying agents into real-world human environments.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!