에이전트 불확실성 하에서의 요인화된 전환 효과로부터의 잠재적인 행동
Latent Actions from Factorized Transition Effects under Agent Ambiguity
잠재적 행동 모델(LAM)은 관찰 데이터를 기반으로 행동과 유사한 표현을 학습합니다. 그러나 다중 객체 또는 많은 배경 요소가 포함된 환경에서는 이러한 시각적 효과가 에이전트의 움직임, 카메라 동역학 및 배경 변화와 혼합되어, 지도 없이도 실제 행동의 원인을 명확하게 파악하기 어렵습니다. 본 연구에서는 이러한 복잡성을 재사용 가능한 전환 효과로 구조화하여, 더욱 안정적으로 행동과 유사한 잠재 변수를 생성할 수 있는 중간 표현을 제공합니다. 우리는 관찰된 전환 요인화(OTF)라는 방법을 제안하며, 이는 각 전환 과정을 일련의 희소한 관찰 가능 원시 요소로 분해합니다. 이러한 원시 요소를 전환 인터페이스로 사용하여, 표준 역/정방향 동역학 프레임워크 내에서 움직임 원시 요소를 행동과 유사한 잠재 변수로 추상화하는 OTF-LAM 모델을 제안합니다. 또한, 디코더가 없는 변형인 OTF-LAM-Dino는 DINOv2 표현 공간에서 미래 상태를 예측합니다. 실험 결과, OTF 원시 요소들은 다양한 환경 및 형태 변화에서도 제로샷으로 이전 가능성을 보여주며, 재사용성이 입증되었습니다. 더욱이, 복잡한 전환 불확실성 하에서 다운스트림 정책 학습 결과는 기존 방법보다 우수하거나 동등한 성능을 보였습니다.
Latent Action Models (LAMs) learn action-like proxies from observation transitions. However, in multi-object or distractor-rich scenes, these visual effects mix agent motion with distractors, camera dynamics, and background changes, making the underlying action source ambiguous without supervision. Structuring this mixture as reusable transition effects provides an intermediate representation from which action-like latents can be more robustly formed. We introduce Observed Transition Factorization (OTF), which decomposes each transition into a sparse set of observed transition primitives. Using these primitives as the transition interface, we propose OTF-LAM, which abstracts motion primitives into action-like latents within the standard inverse-forward dynamics framework, and OTF-LAM-Dino, a decoder-free variant that predicts future states in a frozen DINOv2 representation space. Empirically, OTF primitives transfer zeroshot across controlled carrier and morphology shifts, showing reusability. Furthermore, downstream policy learning results match or outperform baselines under complex transition ambiguity.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.