2607.27017v2 Jul 29, 2026 cs.LG

잠재 세계 모델이 무엇을 알 수 있는가? 다중 모드 예측 표현에서의 물리적 매개변수 식별 가능성

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yang Feng
Yang Feng
Citations: 795
h-index: 14
Xin Xu
Xin Xu
Citations: 0
h-index: 0
Kaizhen Tan
Kaizhen Tan
Citations: 1
h-index: 1
Siru Tao
Siru Tao
Citations: 18
h-index: 1
Han Hong
Han Hong
Citations: 60
h-index: 1
Heqing Du
Heqing Du
Citations: 0
h-index: 0

잠재 세계 모델의 핵심 가정은 미래를 예측하는 과정에서 표현이 환경의 물리학적 원리를 내부에 포함하게 된다는 것이다. 훈련된 잠재 변수는 실제로 어떤 물리적 양을 포함하고 있으며, 이는 어떻게 결정되는가? 우리는 POKEWORLD라는 상호 작용 환경에서 통제된 개입을 통해 이 질문에 답한다. POKEWORLD는 시각적으로 동일한 객체들이 질량, 저항 및 접촉 강성이라는 다른 물리적 특성을 숨기고 있다. 인증 게이트 프로토콜은 먼저 각 매개변수가 원시 관측 데이터로부터 복구 가능한지를 확인하고, 그런 다음 해당 매개변수가 잠재 변수에 포함되는지 측정한다. 따라서 결과가 없을 경우, 이는 환경보다는 학습 목표 때문이라고 판단할 수 있다. 생성된 식별 가능성 지도는 두 가지 조직 메커니즘과 하나의 경계를 가진다. 입력 데이터는 알 수 있는 정보의 범위를 제한하며, 예측 대상은 유지되는 정보를 결정한다. 접촉 강성은 접촉을 예측해야 할 때만 잠재 변수에 포함되며 (R^2 = 0.50), 동일한 신호가 단순히 입력으로 통합될 때는 -0.02에 불과하다. 단일 단계 예측에서 시각 정보만 사용하는 경우, 심지어 완벽하게 보이는 객체의 상태조차도 잠재 변수에서 제거된다. 저항은 경계를 나타낸다. 이는 0.89의 복구 가능성 인증을 갖지만, 테스트한 모든 결정론적 예측 목표에서는 약 0.13으로 정체되는 반면, 동일한 구조를 가진 지도 학습 모델에서는 0.45에 도달한다. 감지된 좌표 하에서 읽기 속도가 느리고 비율 형태인 매개변수는 이러한 학습 목표가 습득하지 못하는 범주에 속한다. RH20T라는 입력-대상 요인 실험을 통해 확장 곡선에서 두 로봇과 4,258개의 에피소드에 걸쳐 위와 같은 메커니즘이 재현된다. 정보가 부족하거나 예측 압력이 없는 모든 설정은 데이터 범위가 다섯 배 증가해도 변하지 않으며, 완전한 다중 모드 목표만 예측력을 높여서 기본 수준을 넘어선다. 이때 숨겨진 데이터를 사용했을 때 얻는 이점은 확장 정도에 따라 증가한다. 학습 목표의 구조는 잠재 변수가 어떤 물리적 매개변수를 습득하는지를 결정하며, 추가 데이터는 이미 습득하고 있는 매개변수의 성능만 향상시킨다.

Original Abstract

A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose visually identical objects hide mass, drag, and contact stiffness. A certificate-gated protocol first certifies each parameter as recoverable from raw observations, then measures whether it enters the latent, so a null result can be attributed to the objective rather than to the environment. The resulting identifiability map has two organizing mechanisms and one frontier. Inputs limit what can be known, while prediction targets decide what is retained. Stiffness enters the latent only when touch is forecast ($R^2=0.50$, compared with $-0.02$ when the same signal is merely fused into the input), and under single-step prediction a vision-only latent discards even perfectly visible object state. Drag marks the frontier. It carries a recoverability certificate of 0.89 yet plateaus near 0.13 under every deterministic prediction objective we test, while a supervised head on the same trunk reaches 0.45. Parameters whose readout is slow and ratio-type under the sensed coordinates fall outside what these objectives acquire. On RH20T, an input-target factorial across scaling curves reproduces both mechanisms across two robots and 4,258 episodes. Every arm missing information or prediction pressure stays flat over a fivefold data range, and only the full multimodal objective forecasts force beyond a persistence baseline, with held-out gains that grow with scale. Objective structure determines which physical parameters a latent acquires, and additional data improves only the parameters it already acquires.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!