2605.25620v1 May 25, 2026 cs.AI

파라미터 수를 줄인 잠재 변수를 활용: 시각 기반 학습을 통해 작업 중심의 세계 모델 구축

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations

Biwei Huang
Biwei Huang
Citations: 24
h-index: 2
Fan Feng
Fan Feng
Citations: 21
h-index: 3
Minghao Fu
Minghao Fu
Citations: 19
h-index: 2
Nicklas Hansen
Nicklas Hansen
Citations: 3,740
h-index: 19

세계 모델은 에이전트가 행동에 따라 미래의 동역학을 예측할 수 있도록 하며, 계획 및 제어에 있어 중요한 역할을 하는 잠재 표현(latent representation)의 선택이 핵심입니다. 이러한 표현은 종종 제한적인 의미 구조를 가진 픽셀로부터 직접 학습되거나, 과도한 작업과 관련 없는 세부 정보를 포함하는 고정된 시각 기반 모델로부터 상속되는데, 이는 다운스트림 계획 및 제어에 적합하지 않은 상태 공간을 초래합니다. 특히 보상 신호가 없는 오프라인 환경에서 이러한 문제는 더욱 심각합니다. 모델은 보상 감독이나 실시간 상호 작용 없이 고정된 경로에서 학습해야 하기 때문입니다. 이러한 문제를 해결하기 위해, 우리는 TC-WM이라는 프레임워크를 제안합니다. 이 프레임워크는 사전 훈련된 임베딩 공간을 최종 상태 공간이 아닌 의미론적 기반으로 활용하여, 작업에 필요한 정보만을 담은 간결한 세계 표현을 구축합니다. 핵심 설계 요소는 다음과 같습니다. 첫째, 고차원 시각 임베딩을 선형적으로 투영하여 동적인 상태 공간을 구성하고, 둘째, 콘트라스트 학습을 통해 에이전트의 물리적 상태와 일치하는 부분 공간을 정렬하며, 셋째, 유용한 시각적 구조를 유지하기 위해 임베딩을 재구성합니다. 이를 통해 기반 모델의 일반성을 유지하면서도 작업 중심적인 동역학을 제어할 수 있습니다. 이론적으로 TC-WM은 기본적인 변환을 통해 근본적인 작업 중심 잠재 요인을 식별하는 데 충분하다는 것을 보여줍니다. 실험적으로 TC-WM은 다양한 환경(예: Robomimic 및 D4RL)에서 테스트 시간 계획을 가능하게 하며, 최첨단 접근 방식보다 더 높은 세계 모델링 품질과 정확한 제어를 달성합니다.

Original Abstract

World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are often either learned directly from pixels with limited semantic structure or inherited from frozen visual foundation models with excessive task-irrelevant detail, yielding state spaces that are poorly matched to downstream planning and control. This is especially challenging in reward-free offline settings, where the model must learn from fixed trajectories without reward supervision or online interaction. To address this, we propose TC-WM, a framework for turning foundation-model embeddings into compact, task-sufficient world representations. The key design is to treat the pretrained embedding space as a semantic scaffold rather than as the final state space: TC-WM linearly projects high-dimensional visual embeddings into a compact latent as the dynamic space, aligns a subspace with the agent's physical state via contrastive learning, and reconstructs embeddings to preserve useful visual structure. This combines the generality of foundation features with the controllability of task-centric dynamics. Theoretically, we show that TC-WM suffices to identify the underlying task-centric latent factors up to a simple transformation. Empirically, TC-WM enables test-time planning across diverse environments (e.g., Robomimic and D4RL), achieving better world-modeling quality and more precise control than state-of-the-art approaches.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!