2607.11689v1 Jul 13, 2026 cs.RO

세계 행동 모델에서 구현된 두뇌로: 개방형 환경에서의 물리적 지능을 위한 로드맵

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Yuanzhi Liang
Yuanzhi Liang
Citations: 300
h-index: 6
Haibin Huang
Haibin Huang
Citations: 65
h-index: 4
Chi Zhang
Chi Zhang
Citations: 52
h-index: 4
Xuelong Li
Xuelong Li
Citations: 124
h-index: 7
Xufeng Zhan
Xufeng Zhan
Citations: 11
h-index: 2

인공 일반 지능은 궁극적으로 실제 세계에서 추론하고 행동할 수 있는 에이전트를 필요로 합니다. 행동 모델, 시각-언어-행동 정책 및 세상 모델은 이러한 목표를 발전시켜 왔으며, 특히 '세계 행동 모델(WAM)'은 잠재적인 개입과 예측된 결과 간의 연결을 통해 유망합니다. 그러나 현재 진행 상황은 여전히 단편적입니다. 모델들은 호환되지 않는 행동 공간과 예측 대상을 사용하며, 데이터셋 및 작업은 서로 다른 규칙을 따르고, 런타임 시스템은 재사용 및 평가를 위한 제한적인 인터페이스를 제공합니다. 우리는 WAM으로의 진화를 검토하고 이러한 한계점을 세 가지 상호 관련된 격차로 분류했습니다: 모델 역할 및 표현, 목표 및 표준화, 그리고 시스템 구성. 본 분석을 바탕으로, 우리는 다중 모달 컨텍스트 통합, 후보 개입 비교, 직접적인 액추에이터 명령 대신 상태 전환 또는 기능 요청을 수행하는 장기적인 모델 목표인 '구현된 두뇌'를 중심으로 하는 물리적 지능의 공동 진화 로드맵을 제안합니다. WAM은 예측 기능을 위한 유망한 프로토타입을 제공하며, 물리적 하드웨어는 도구, 컨트롤러, 검증 및 추적 로깅을 통해 모델 출력을 실제 세계에 연결합니다. 공유된 계약은 이기종 모델, 데이터, 작업 및 구현체를 통합하고, 폐쇄 루프 후훈련은 검증된 상호 작용을 재사용 가능한 경험으로 변환합니다. 이러한 구성 요소들은 적응적이고 자기 개선이 가능한 구현 에이전트를 위한 모듈형 물리 지능 스택을 정의합니다.

Original Abstract

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems expose limited interfaces for reuse and evaluation. We review the evolution toward WAMs and organize these limitations into three coupled gaps: model roles and representations, objectives and standardization, and system composition. Building on this analysis, we propose a co-evolution roadmap for physical intelligence centered on the \emph{embodied brain}, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands. WAMs provide promising prototypes for its predictive functions, while a physical harness grounds model outputs through tools, controllers, verification, and trace logging. Shared contracts align heterogeneous models, data, tasks, and embodiments, and closed-loop post-training converts verified interaction into reusable experience. Together, these components define a modular physical-intelligence stack for adaptive and self-improving embodied agents.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!