2607.26657v1 Jul 29, 2026 cs.RO

Enfold: 예측 표현을 위한 월드 제너레이터 연산 통합 - 효율적인 Embodied 제어를 위한 방법

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

Jingnan Gao
Jingnan Gao
Citations: 74
h-index: 5
Weili Zeng
Weili Zeng
Citations: 27
h-index: 2
Yichao Yan
Yichao Yan
Citations: 28
h-index: 3
Jisong Cai
Jisong Cai
Citations: 710
h-index: 4
Yitong Xing
Yitong Xing
Citations: 7
h-index: 1
Fulong Liu
Fulong Liu
Citations: 0
h-index: 0
Chengqun Yang
Chengqun Yang
Citations: 8
h-index: 1
Antao Xiang
Antao Xiang
Citations: 0
h-index: 0
Feng Tian
Feng Tian
Citations: 0
h-index: 0
Xin Wang
Xin Wang
Citations: 0
h-index: 0
Xiaomin Wu
Xiaomin Wu
Citations: 0
h-index: 0
Yao Mu
Yao Mu
Citations: 253
h-index: 2

월드 생성 모델은 일반적으로 생성된 결과물, 즉 렌더링된 미래 시뮬레이션, 비디오 기반 액션 또는 비용이 많이 드는 생성 브랜치에 의해 계산된 잠재적 컨텍스트을 통해 사용됩니다. 우리는 이러한 모델의 가장 유용한 자산은 실제로 미래를 구성하는 연산 자체라고 주장합니다. 제너레이터가 손상된 미래를 일관성 있는 궤적으로 변환할 때, 중간 상태는 다양한 추상화 수준에서 외형, 공간 배치 및 상호 작용을 조직합니다. 이러한 미래 생성 연산을 현재의 정보만으로 추론되는 표현 내부에 통합할 수 있을까요? 우리는 Enfold를 제안하며, 이는 시각적 컨텍스트와 언어 지침으로부터 예측된 표현에 이 연산을 통합하는 방법입니다. 훈련 과정에서, 관찰된 미래를 처리하면서 노출되는 다층 상태가 현재의 정보만을 사용하는 인코더를 학습하도록 지도합니다. 학습된 표현은 미래 생성에 영향을 미치도록 피드백되고, 작업 헤드가 이를 활용하지만, 작업 관련 기울기가 인코더를 변경하지 못하도록 합니다. 배포 시, 액션 예측은 더 이상 제너레이터를 실행하지 않습니다. LIBERO, RoboTwin2.0 및 실제 로봇 작업을 통해 Enfold는 강력한 제어 성능을 제공하면서 Fast--WAM에 비해 액션 지연 시간을 3.7배 줄이고, Enfold-Flash는 10.1배까지 개선합니다. 표현 분석 결과, 불필요한 변동성을 억제하고 더 긴 시간 규모에서 발생하는 변화를 우선적으로 포착하는 것을 확인했습니다. 현재 장면이 인간의 개입으로 변경될 때, 생성된 연속적인 부분과 실행되는 액션 모두 적응하는데, 이는 고정된 궤적 재생과는 일치하지 않습니다. 이러한 결과는 월드 제너레이터를 예측 제어 표현의 원천으로 재해석합니다. 미래를 매 단계마다 물질화할 필요 없이, 내부 구조를 현재에 통합(enfold)할 수 있다면 충분합니다.

Original Abstract

World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by $3.7\times$ relative to Fast--WAM, Enfold-Flash reaches $10.1\times$. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!