REMAC: 자기 성찰 및 자기 진화 기반의 다중 에이전트 협업 시스템을 이용한 장기 로봇 조작
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
비전-언어 모델(VLMs)은 특히 환경에 대한 전체적인 이해를 바탕으로 작업을 분해해야 하는 장기 작업에서 로봇 계획 분야에서 놀라운 성능을 보여주었습니다. 기존 방법들은 일반적으로 사전 정의된 환경 정보나 신중하게 설계된 작업별 프롬프트에 의존하기 때문에, 동적인 장면 변화나 예상치 못한 작업 조건(예: 로봇이 당근을 전자레인지에 넣으려 하지만 문이 닫혀 있는 경우)에서 어려움을 겪습니다. 이러한 과제는 적응성 및 효율성과 같은 두 가지 중요한 문제를 강조합니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 지속적인 성찰과 자기 진화를 통해 효율적이고 장면 무관한 다중 로봇 장기 작업 계획 및 실행을 가능하게 하는 적응형 다중 에이전트 계획 프레임워크인 REMAC을 제안합니다. REMAC은 두 가지 핵심 모듈로 구성됩니다. 첫 번째는 진행 상황을 평가하고 계획을 개선하기 위해 전제 조건과 후위 조건을 반복적으로 확인하는 자기 성찰 모듈이고, 두 번째는 장면별 추론에 따라 동적으로 계획을 조정하는 자기 진화 모듈입니다. REMAC은 다음과 같은 장점을 제공합니다. 1) 로봇은 복잡한 프롬프트 설계 없이 환경을 탐색하고 추론할 수 있습니다. 2) 로봇은 잠재적인 계획 오류를 지속적으로 평가하고 작업별 통찰력을 기반으로 계획을 조정할 수 있습니다. 3) 반복을 통해, 하나의 로봇이 다른 로봇에게 작업을 요청하여 병렬로 수행함으로써 작업 실행 효율성을 극대화할 수 있습니다. REMAC의 효과성을 검증하기 위해, RoboCasa를 기반으로 장기 로봇 조작 및 탐색을 위한 다중 에이전트 환경을 구축했으며, 4가지 작업 범주와 27가지 스타일, 50개 이상의 다양한 객체를 포함합니다. 이를 바탕으로 DeepSeek-R1, o3-mini, QwQ, Grok3과 같은 최첨단 추론 모델들을 평가하여, REMAC이 평균 성공률을 40% 향상시키고 실행 효율성을 52.7% 향상시켜 단일 로봇 기준 성능보다 우수함을 입증했습니다.
Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition. Existing methods typically rely on prior environmental knowledge or carefully designed task-specific prompts, making them struggle with dynamic scene changes or unexpected task conditions, e.g., a robot attempting to put a carrot in the microwave but finds the door was closed. Such challenges underscore two critical issues: adaptability and efficiency. To address them, in this work, we propose an adaptive multi-agent planning framework, termed REMAC, that enables efficient, scene-agnostic multi-robot long-horizon task planning and execution through continuous reflection and self-evolution. REMAC incorporates two key modules: a self-reflection module performing pre-condition and post-condition checks in the loop to evaluate progress and refine plans, and a self-evolvement module dynamically adapting plans based on scene-specific reasoning. It offers several appealing benefits: 1) Robots can initially explore and reason about the environment without complex prompt design. 2) Robots can keep reflecting on potential planning errors and adapting the plan based on task-specific insights. 3) After iterations, a robot can call another one to coordinate tasks in parallel, maximizing the task execution efficiency. To validate REMAC's effectiveness, we build a multi-agent environment for long-horizon robot manipulation and navigation based on RoboCasa, featuring 4 task categories with 27 task styles and 50+ different objects. Based on it, we further benchmark state-of-the-art reasoning models, including DeepSeek-R1, o3-mini, QwQ, and Grok3, demonstrating REMAC's superiority by boosting average success rates by 40% and execution efficiency by 52.7% over the single robot baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.