2607.26924v1 Jul 29, 2026 cs.LG

시간 중심 SIGReg을 활용한 다중 작업 LeWorldModel 학습 향상: 분석에서 방법론으로

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

Yusuke Iwasawa
Yusuke Iwasawa
Citations: 10,847
h-index: 23
Chang Liu
Chang Liu
Citations: 0
h-index: 0
Feiya Suo
Feiya Suo
Citations: 0
h-index: 0
Yanzhou Jin
Yanzhou Jin
Citations: 0
h-index: 0
Yutaka Matsuo
Yutaka Matsuo
Citations: 8,669
h-index: 14
Yaonan Zhu
Yaonan Zhu
Citations: 10
h-index: 2

최근 LeWorldModel (LeWM) 연구에서는 Sketched Isotropic Gaussian Regularizer (SIGReg)가 잠재 분포를 등방성 가우시안 분포로 정규화하여 표현 붕괴를 방지함으로써, 픽셀로부터 안정적인 엔드-투-엔드 세계 모델 학습을 가능하게 한다는 것이 밝혀졌습니다. 단일 작업 환경에서는 효과적이고 우아한 방법이지만, 이 방식은 다중 작업 학습으로 확장될 때 신뢰성이 떨어져 다운스트림 행동 복제 성능이 크게 저하됩니다. 본 논문에서는 marginal Gaussianization이 작업 의존적인 잠재 클러스터 간의 분리를 클러스터 내부 변동에 비해 압축한다는 것을 보여줍니다. 이러한 압축은 작업 및 상태 간의 표현 왜곡을 초래하며, 학습된 표현을 작은 시각적 변화에 매우 민감하게 만듭니다. 이 문제를 해결하기 위해, 우리는 잠재 marginal 분포가 아닌 시간 중심 잔차에 SIGReg를 적용합니다. 이 대체 목표는 클러스터 중심 간의 분리에 직접적인 정규화 압력을 가하지 않으며, 전체 잠재 변수가 단일 등방성 가우시안을 따르도록 요구하는 제약을 제거하고, SIGReg의 붕괴 방지 효과를 유지합니다. LIBERO 벤치마크에서, 우리의 방법은 장기 예측 작업에서 성공률을 1.7배 향상시키고, 네 가지 테스트 세트 전체의 평균 성공률을 53.2%에서 73.6%로 높입니다. 외부 사전 학습 없이도, 우리의 방법은 처음부터 학습된 Diffusion Policy보다 약간 더 좋은 성능을 보이며, 대규모 사전 학습 정책 모델 수준에 근접하는 성능을 달성합니다. 이러한 결과는 marginal Gaussian prior와 다중 작업 잠재 구조 간의 구조적 불일치를 보여주며, 안정적이고 확장 가능한 엔드-투-엔드 다중 작업 세계 모델 학습을 위한 간단한 방법을 제시합니다.

Original Abstract

Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. While effective and elegant in single-task settings, this recipe does not extend reliably to multi-task training, leading to substantially worse downstream behavior-cloning performance. In this paper, we show that marginal Gaussianization compresses the separation between task-dependent latent clusters relative to within-cluster variation. This compression introduces representation aliasing across tasks and states, and makes the learned representations highly sensitive to small visual perturbations. To address this problem, we apply SIGReg to temporally centered residuals rather than to the latent marginal distribution. This surrogate target places no direct regularization pressure on the separation among cluster centers, removes the requirement that the full latent follow a single isotropic Gaussian, and retains the anti-collapse effect of SIGReg. On the LIBERO benchmark, our method improves downstream success on the long-horizon suite by 1.7x and raises the average success rate across four suites from 53.2% to 73.6%. Without external pretraining, it slightly outperforms Diffusion Policy trained from scratch and approaches the performance of large-scale pretrained policy baselines. These results reveal a structural incompatibility between marginal Gaussian priors and multi-task latent structure, and provide a simple route toward stable and scalable end-to-end multi-task world-model learning.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!