2607.26056v1 Jul 28, 2026 cs.RO

INTACT: 등방성 의도-행동 학습을 통한 검색 없는 세계 모델

INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

Hao Zhao
Hao Zhao
Citations: 60
h-index: 5
Jun Sun
Jun Sun
Citations: 3
h-index: 1
Guofeng Zhang
Guofeng Zhang
Citations: 1
h-index: 1

기존의 잠재 변수 기반 세계 모델은 행동이 장면을 어떻게 변화시키는지 예측하지만, 원하는 변화를 달성하기 위해서는 비용이 많이 드는 테스트 시간 동안의 탐색 과정이 필요합니다. 본 논문에서는 INTACT (INtent-To-ACTion)라는 엔드투엔드 JEPA 모델을 소개합니다. 이 모델은 액션 레이블이 부여된, 보상 정보가 없는 trajectory 데이터를 활용하여 실제 환경에 적용 가능한 의도-행동 인터페이스를 구축합니다. 각 transition은 물리적 의도를 나타내는 $z_{t+1}-z_t$ 값을 제공하며, 미래의 목표는 deployment 의도를 나타내는 $ rac{sg(z_g)-z_t}{}$ 값을 제공합니다. 제안하는 모델 아키텍처는 로컬 및 목표 운동-의도 백본-입력 그래프 간에 동일한 4개의 슬롯을 가진 문법과 공유된 파라미터를 통해 동형 구조를 갖습니다. 또한, 액션-법(action-law) 의미론을 사용하여 동일한 예측기를 통해 지원되는 로컬 및 목표 운동-의도 패밀리 간에도 연결성을 제공하며, 잠재 변수의 pointwise equality에 의존하지 않습니다. INTACT는 RGB 정보에서 액션 효과적인 잠재 의도 좌표로의 원활한 전이와 의도 패밀리에서 해당 액션-법 패밀리로의 전이를 가능하게 합니다. 비대칭 endpoint gradient를 통해 물리적 결과를 학습하고 미래 목표를 고정하여, pointwise latent matching이나 글로벌 선형 동역학 없이 representation learning과 control을 결합합니다. 결과적으로 얻어지는 좌표는 강력한 분포 기반 액션-법을 지원하며, 조건부 평균은 검색이 필요 없는 정책으로 직접 활용될 수 있으며, 다양성을 확보하거나 선택적인 검증을 위해 sampling도 가능합니다. 4개의 공식 LeWM task에서 단일 epoch 학습과 zero-search 모델로 각각 85.78%, 100.00%, 97.67%, 97.89%의 성공률을 달성했습니다. Direct 계획을 중심으로 하는 로컬 CEM (Cross Entropy Method)은 384개의 후보 sequence만을 사용하여 96.86%의 macro success를 달성했으며, 이는 기존 CEM 방식보다 16.00 포인트 향상된 결과입니다. 또한, 하나의 공유된 4-task encoder는 89.39%의 E5 Direct macro 성공률을 달성했으며, Jointly trained LeWM 모델에 비해 모든 task에서 성능이 향상되었습니다. 예측된-전문가 액션 패밀리 kNN은 Direct success를 $r=0.954$의 정확도로 추적합니다. Direct inference는 2.9~5.5ms의 시간 내에 완료됩니다.

Original Abstract

Forward latent world models predict how actions change a scene, but recover actions for a desired change only through expensive test-time search. We introduce INTACT (INtent-To-ACTion), an end-to-end JEPA that turns action-labeled, reward-free trajectories into a deployable intent-to-action interface. Each transition supplies physical intent $z_{t+1}-z_t$, while a future goal supplies deployment intent $\operatorname{sg}(z_g)-z_t$. The architecture is isomorphic between the local and goal motion-intent backbone-input graphs through an identical four-slot grammar and shared parameters, and between supported local and goal motion-intent families through action-law semantics induced by the same predictor rather than pointwise latent equality. INTACT also provides intact transfer from RGB evidence to action-effective latent intent coordinates and from intent families to their corresponding action-law families. Asymmetric endpoint gradients ground physical successors and fix future goals as anchors, joining representation learning and control without pointwise latent matching or globally linear dynamics. The resulting coordinates support a robust distributional action law: its conditional mean serves directly as a search-free policy, while sampling remains available for diversity or optional verification. On the four official LeWM tasks, one-epoch, zero-search models reach 85.78\%, 100.00\%, 97.67\%, and 97.89\% success. Optional local CEM centered on the Direct plan reaches 96.86\% macro success using 384 instead of 9,000 candidate sequences, reducing sampling by $23.44\times$ while improving pure CEM by 16.00 points. One shared four-task encoder reaches 89.39\% E5 Direct macro and improves every task over jointly trained LeWM, while predicted--expert action-family kNN tracks Direct success at $r=0.954$. Direct inference takes 2.9--5.5 ms.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!