ELITE: 경험 기반 학습 및 의도 인식 전이 기술을 활용한 자가 개선 능력을 갖춘 로봇 에이전트
ELITE: Experiential Learning and Intent-Aware Transfer for Self-improving Embodied Agents
시각-언어 모델(VLM)은 뛰어난 일반적인 능력을 보여주지만, 이를 기반으로 구축된 로봇 에이전트는 복잡한 작업을 수행하는 데 어려움을 겪으며, 종종 중요한 단계를 건너뛰거나, 잘못된 행동을 제안하거나, 실수를 반복합니다. 이러한 실패는 VLM의 정적 학습 데이터와 로봇 작업에 필요한 물리적 상호 작용 간의 근본적인 차이에서 비롯됩니다. VLM은 정적 데이터로부터 풍부한 의미론적 지식을 학습할 수 있지만, 세상과 상호 작용하는 능력은 부족합니다. 이러한 문제를 해결하기 위해, 저희는 경험 기반 학습(Experiential Learning)과 의도 인식 전이(Intent-aware Transfer) 기술을 통합한 로봇 에이전트 프레임워크인 ELITE를 제안합니다. ELITE는 에이전트가 자신의 환경과의 상호 작용 경험으로부터 지속적으로 학습하고, 획득한 지식을 절차적으로 유사한 작업에 적용할 수 있도록 합니다. ELITE는 자기 성찰적 지식 구축(self-reflective knowledge construction)과 의도 인식 검색(intent-aware retrieval)이라는 두 가지 상호 보완적인 메커니즘을 통해 작동합니다. 구체적으로, 자기 성찰적 지식 구축은 실행 경로에서 재사용 가능한 전략을 추출하고, 구조화된 개선 작업을 통해 진화하는 전략 풀을 유지합니다. 그런 다음, 의도 인식 검색은 풀에서 관련 전략을 식별하고, 이를 현재 작업에 적용합니다. EB-ALFRED 및 EB-Habitat 벤치마크에서의 실험 결과, ELITE는 지도 학습 없이 온라인 환경에서 기본 VLM보다 각각 9% 및 5%의 성능 향상을 달성했습니다. 지도 학습 환경에서는 ELITE가 최첨단 훈련 기반 방법보다 우수한 성능으로 새로운 작업 범주에 효과적으로 일반화됩니다. 이러한 결과는 ELITE가 의미론적 이해와 안정적인 행동 실행 간의 격차를 해소하는 데 효과적임을 보여줍니다.
Vision-language models (VLMs) have shown remarkable general capabilities, yet embodied agents built on them fail at complex tasks, often skipping critical steps, proposing invalid actions, and repeating mistakes. These failures arise from a fundamental gap between the static training data of VLMs and the physical interaction for embodied tasks. VLMs can learn rich semantic knowledge from static data but lack the ability to interact with the world. To address this issue, we introduce ELITE, an embodied agent framework with {E}xperiential {L}earning and {I}ntent-aware {T}ransfer that enables agents to continuously learn from their own environment interaction experiences, and transfer acquired knowledge to procedurally similar tasks. ELITE operates through two synergistic mechanisms, \textit{i.e.,} self-reflective knowledge construction and intent-aware retrieval. Specifically, self-reflective knowledge construction extracts reusable strategies from execution trajectories and maintains an evolving strategy pool through structured refinement operations. Then, intent-aware retrieval identifies relevant strategies from the pool and applies them to current tasks. Experiments on the EB-ALFRED and EB-Habitat benchmarks show that ELITE achieves 9\% and 5\% performance improvement over base VLMs in the online setting without any supervision. In the supervised setting, ELITE generalizes effectively to unseen task categories, achieving better performance compared to state-of-the-art training-based methods. These results demonstrate the effectiveness of ELITE for bridging the gap between semantic understanding and reliable action execution.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.