2606.27251v1 Jun 25, 2026 cs.RO

분리된 기술에서 일상적인 물리적 자율성으로 나아가는 다중 모드 에이전트

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

Xipeng Qiu
Xipeng Qiu
Citations: 28
h-index: 2
Jingjing Gong
Jingjing Gong
Citations: 174
h-index: 6
Zhaoye Fei
Zhaoye Fei
Citations: 1,048
h-index: 12
Yu-Gang Jiang
Yu-Gang Jiang
Citations: 41
h-index: 3
Hechang Chen
Hechang Chen
Citations: 256
h-index: 6
Siyin Wang
Siyin Wang
Fudan University
Citations: 343
h-index: 9
Junhao Shi
Junhao Shi
Fudan University
Citations: 160
h-index: 5
Zezheng Huai
Zezheng Huai
Citations: 1
h-index: 1
Jia Chen
Jia Chen
Citations: 0
h-index: 0
Yubang Wang
Yubang Wang
Citations: 47
h-index: 2

비정형 환경에서 지속 가능한 에이전트를 구축하려면 사이버(API, IoT) 및 물리(조작, 내비게이션) 도메인에 걸쳐 이질적인 도구를 통합적으로 제어해야 하며, 장기간 작동 중에 발생하는 물리적 오류로부터 자율적으로 복구할 수 있어야 합니다. 기존 시스템은 이러한 요소들을 별개의 문제로 취급합니다. VLM 기반 계획기는 통일된 사이버-물리 공간을 갖추지 못하고, 에이전트 프레임워크는 무한대의 컨텍스트를 축적하여 시간적 일관성을 저하시키며, VLA 정책은 자체 오류를 감지하지 않고 개방 루프 방식으로 실행됩니다. 우리는 지속적인 자율성이 단일 모델이 아닌 계획, 메모리 및 검증을 명시적으로 분리하는 계층적 비동기 아키텍처를 필요로 한다고 주장합니다. 이를 위해 다중 모드 의미론적 플래너를 사용하여 통일된 액션 공간에서 기술 라우팅을 수행하고, 이벤트 경계를 기반으로 적응적인 계층적 메모리를 통해 컨텍스트 증가를 줄이며, 물리적 실행 중에 의미론적 루프를 닫는 비동기 시각적 사전 중단 엔진을 통합한 프레임워크인 OmniAct을 제시합니다. 두 대의 로봇 플랫폼에서 4개의 IoT 장치를 조정하여 수행된 40가지 실제 장기 작업에서 OmniAct은 모든 복잡도 수준에서 전체 성공률을 꾸준히 향상시키고, 10만 건 이상의 누적 상호 작용 토큰 동안 거의 일정하게 토큰 소비량을 유지하며, 중간 규모의 오픈 웨이트 모델을 독점적인 수준의 성능으로 끌어올립니다.

Original Abstract

Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physical action space, agent frameworks accumulate unbounded context that degrades temporal coherence, and VLA policies execute open-loop without detecting their own failures. We argue that persistent autonomy requires not a monolithic model but a hierarchical asynchronous architecture with explicit separation of planning, memory, and verification. To this end, we present OmniAct, a framework integrating a multimodal semantic planner for skill routing across unified action spaces, an adaptive hierarchical memory with event-boundary-driven compression for sub-linear context growth, and an asynchronous visual preemption engine that closes the semantic loop during physical execution. Across 40 real-world long-horizon tasks on two robotic platforms coordinating four IoT devices, OmniAct achieves consistent improvements in end-to-end success across all complexity levels, maintains near-flat token consumption over under 100k+ accumulated interaction tokens, and elevates mid-scale open-weight models to proprietary-level performance.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!