인형 로봇을 위한 확장 가능한 행동 기반 모델
Scaling Behavior Foundation Model for Humanoid Robots
인형 로봇 제어는 자연스러운 전신 조화, 정밀한 실시간 반응, 그리고 다양한 환경 조건에서의 강력한 일반화 능력을 요구하며, 이는 범용 에이전트의 핵심 요소입니다. 최근, 행동 기반 모델(BFM)은 대규모 행동 데이터를 활용하여 뛰어난 표현력, 다양성 및 일반화 성능을 달성함으로써 이러한 과제를 해결할 수 있는 유망한 솔루션으로 부상했습니다. 그러나 BFM의 기능을 더욱 향상시키기 위한 확장 연구에 대한 관심이 높아지고 있음에도 불구하고, 학습 패러다임, 행동 데이터 및 모델 아키텍처와 같은 주요 요소들이 효과적인 확장을 위해 어떻게 조율되어야 하는지에 대한 명확성은 여전히 부족합니다. 본 연구에서는 BFM의 확장 전략을 재검토하고 세 가지 핵심 구성 요소를 조정함으로써 상당한 성능 향상을 얻을 수 있음을 보여줍니다. 첫째, 모션 추적 학습 패러다임을 통해 다양한 인형 로봇 제어 문제를 글로벌 좌표계에서 통합된 전신 행동 재생으로 재구성합니다. 둘째, 온-정책 롤아웃 양과 참조 동작 다양성 간의 전략적인 시너지 효과를 활용합니다. 셋째, 자연스러운 구조화된 행동 표현을 가능하게 하는 '인형 트랜스포머'라는 표현력 있고 확장 가능한 모델 아키텍처를 사용합니다. 시뮬레이션 및 실제 환경에서의 광범위한 실험을 통해 본 연구의 접근 방식이 제어 정확도와 작업 일반화 성능을 크게 향상시키며, 기존 인형 로봇 제어 시스템에 비해 테스트 세트에서 평균 키포인트 위치 오차(MPKPE)를 로컬 모드에서는 10% 이상, 글로벌 모드에서는 82% 이상 감소시킴을 입증했습니다. 이러한 결과는 BFM이 확장 가능하고 범용적인 인형 로봇 제어를 위한 효과적이고 체계적인 기반임을 보여줍니다.
Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.