DLAM: 시간 제약 조건이 적용된 분포형 잠재 액션
DLAM: Distributional Latent Actions with Temporal Constraints
비전-언어-액션(VLA) 모델은 여전히 제한적인 액션 레이블이 부착된 로봇 데이터에 의존하는 반면, 액션 정보가 없는 비디오는 물리적 변화에 대한 풍부한 정보를 제공합니다. 잠재 액션 모델은 이러한 사전 지식을 추출할 수 있지만, 재구성을 통해 학습된 코드는 로봇 액션과의 결합 생성을 위해 필요한 구조 없이 미래 관찰 값을 예측할 수 있습니다. 기존의 구조화된 방법들은 시간 제약 조건을 추가하지만, 여전히 결정적인 전환 지점을 유지하기 때문에, 지역적으로 추론된 전환 과정에서의 잔여 오류가 재귀적 조합을 통해 전파되고 증폭될 수 있습니다. 본 연구에서는 각 전환 과정을 대각 가우스 분포로 표현하는 분포형 잠재 액션 모델인 DLAM을 제안합니다. 참조 프레임에 기반한 재구성은 관찰된 시각적 변화를 기준으로 평균값을 설정하며, 정규화된 조합과 동일 간격의 3개 요소에 대한 역전은 평균값뿐만 아니라 차원별 분산도 제약합니다. 분산 조합은 인접한 전환 과정에서 공유되는 중간 프레임과의 의존성을 고려하기 위해 가벼운 공유 상관 계수를 사용하며, 역전은 평균값을 반전시키고 분산을 유지합니다. 다운스트림 정책 학습을 위해, 인코더는 고정하고 플로우 매칭 정책을 학습하여 평균 전환 시퀀스와 로봇 액션을 동시에 생성합니다. DLAM은 기존의 잠재 액션 모델에 비해 더 일관된 시간적 동역학을 학습하며, 보이지 않는 비디오에서 더 강력한 직접 및 누적 재구성을 달성합니다. 동일한 제어된 $π_0$ 전이 프로토콜 하에서, MetaWorld MT50, LIBERO 및 실제 로봇 조작 작업에서 정책 성능이 향상됩니다. 통제된 실험 결과는 정규화된 평균 제약 조건이 대부분의 재구성 개선에 기여하는 반면, 학습된 분산과 상관 관계를 고려한 조합은 다운스트림 제어 측면에서 추가적인 개선을 제공한다는 것을 보여줍니다.
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled $π_0$ transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.