$\mathcal{P}^3$: 다재다능한 로봇 에이전트를 향하여
$\mathcal{P}^3$: Toward Versatile Embodied Agents
로봇 에이전트는 물리적 환경과의 상호작용에서 유망한 잠재력을 보여주었습니다. 그러나 다재다능한 로봇 에이전트는 동적인 환경 인식, 개방형 도구 활용 및 복잡한 멀티태스킹 계획이라는 세 가지 핵심적인 문제에 직면합니다. 기존 방법들은 장면 변화 및 작업 진행 상황을 추적하기 위해 도구 피드백에 전적으로 의존하여 실시간 적응력이 떨어지고 오류가 누적되며, 도구 호환성이 제한됩니다. 또한, 멀티태스킹 계획은 작업 종속성과 충돌하는 우선순위를 처리하기 어렵기 때문에 연구가 부족합니다. 이러한 한계를 극복하기 위해, 우리는 실시간 인식과 동적 스케줄링을 통합한 통일된 프레임워크인 $\mathcal{P}^3$를 제안합니다. $\mathcal{P}^3$는 환경으로부터 관련 정보를 능동적으로 인지하고, 피드백 요구 사항 없이 도구를 활용하며, 긴급 작업을 우선시하고 종속성에 따라 작업 순서를 동적으로 조정하여 멀티태스킹 실행을 계획합니다. 또한, VLM의 능동적인 장면 이해 및 작업 제안 능력을 정량적으로 측정하기 위한 Active Task Perception (ATP) 벤치마크를 구축했습니다. ATP 벤치마크에 대한 평가 결과, 여러 VLM이 능동적인 작업을 감지하고 제안할 수 있으며, 실제 로봇 실험을 통해 본 방법이 벤치마크와 실제 적용 간의 격차를 해소하며, 일반적인 용도의 로봇 에이전트를 구현할 수 있음을 입증했습니다. 코드 및 데이터는 https://github.com/fz-zsl/P3 에서 확인할 수 있습니다.
Embodied agents have demonstrated promising capabilities in interacting with physical environments. Yet, versatile embodied agents face three core bottlenecks: dynamic environmental perception, open tool access, and complex multi-task planning. Prior methods depend entirely on tool feedback to track scene changes and task progress, leading to poor real-time adaptability, error accumulation, and limited tool compatibility; multi-task scheduling is also understudied due to the difficulty of handling task dependencies and conflicting priorities. To address these limitations, we propose $\mathcal P^3$, a unified framework integrating real-time perception and dynamic scheduling, which perceives task-relevant information actively from the environment, plugs and utilizes tools without feedback requirements, and plans multi-task execution by prioritizing urgent tasks and dynamically adjusting task order based on dependencies. We additionally build the Active Task Perception (ATP) benchmark to quantitatively measure VLMs' capacity for active scene understanding and task proposal. Evaluations on the ATP benchmark verify that multiple VLMs can detect and propose active tasks, and comprehensive real-world robot experiments prove our method bridges the gap between benchmarks and practical deployment, yielding transferable general-purpose embodied agents. Code and data are available at https://github.com/fz-zsl/P3.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.