2606.12217v1 Jun 10, 2026 cs.CV

미래 예측을 실용적으로 활용하기: 월드 액션 모델에서의 표현 정렬 재활용

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

Yi Chen
Yi Chen
Citations: 418
h-index: 9
Yuying Ge
Yuying Ge
Citations: 4,788
h-index: 29
Yixiao Ge
Yixiao Ge
Citations: 599
h-index: 13
Lu Qiu
Lu Qiu
Citations: 91
h-index: 4
Xihui Liu
Xihui Liu
Citations: 527
h-index: 11
Yizhuo Li
Yizhuo Li
Citations: 202
h-index: 6

월드 액션 모델(WAM)은 비디오 생성 모델을 사용하여 제어 동작을 수행하기 전에 미래 장면의 변화를 모델링함으로써 로봇 조작에 유망한 방법을 제공합니다. 그러나 우리의 실험적 관찰 결과, 설득력 있는 시각적 미래를 생성하는 것이 항상 정확한 동작 추출을 보장하지는 않는다는 현상이 밝혀졌습니다. 이러한 실패 원인을 진단하기 위해, 액션 헤드 어텐션 분석 및 인과적 개입을 수행했습니다. 그 결과, 액션 디코더가 작업 관련 상호 작용 영역에 집중하지 못하고 작업과 관련 없는 영역의 변화에 민감하게 반응한다는 것을 발견했습니다. 이는 표현 불일치(representation mismatch)를 보여주며, 시각 재구성에 최적화된 숨겨진 상태는 저수준 동작 제어에 유용한 형태로 구성되어 있지 않음을 의미합니다. 본 논문에서는 AGRA라는 액션 기반 표현 정렬 객관 함수를 제안합니다. AGRA는 월드-액션 인터페이스를 정규화하여 중간 비디오 확산 특징을 기본 시각 인코더에서 얻은 공간적으로 일관된 의미론적 표현과 정렬합니다. 우리는 실제 조작 작업에서 AGRA를 평가했습니다. 실험 결과, AGRA가 월드 모델 표현을 더욱 액션 지향적으로 만든다는 것을 보여줍니다. AGRA는 액션 디코더가 올바른 상호 작용 영역에 집중하도록 하여 객체 위치 정확도와 활용 가능성 이해도를 향상시키고, 작업과 관련 없는 영역의 변화에 대한 정책의 강건성을 높입니다. 그 결과, AGRA는 기존 월드 액션 모델보다 일관되게 데이터 분포 내 성능과 외부 데이터 일반화 능력을 향상시킵니다.

Original Abstract

World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions. To diagnose this failure, we conduct action-head attention analysis and causal interventions. We find that the action decoder fails to focus on task-relevant interaction regions and remains sensitive to perturbations in task-irrelevant areas. This reveals a representation mismatch: hidden states optimized for visual reconstruction are not inherently organized in a form useful for low-level action control. In this paper, we propose AGRA, an Action-Grounded Representation Alignment objective that regularizes the world-action interface by aligning intermediate video diffusion features with spatially coherent semantic representations from a foundation visual encoder. We evaluate AGRA on real-world manipulation tasks. Experiments show that AGRA makes world model representations more action-grounded: by focusing the action decoder on the correct interaction regions, it improves object localization accuracy and affordance understanding, and makes the policy more robust to perturbations in task-irrelevant regions. As a result, AGRA consistently improves both in-distribution performance and out-of-distribution generalization over the baseline world action model.

1 Citations
0 Influential
14.5 Altmetric
73.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!