비디오 표현 정규화를 통한 복합 오류 완화
Mitigating Compounding Error via Video Representation Regularization
비디오 확산 기반 세계 모델은 로봇 공학, 자율 주행 및 시뮬레이션 작업에 대한 장기적인 자기 회귀 비디오 생성을 가능하게 하지만, 슬라이딩 윈도우 자기 회귀 추론은 심각한 오류 누적 현상으로 인해 시간이 지남에 따라 프레임 품질을 저하시킵니다. 이 현상은 광범위하게 관찰되었지만, 복합 오류의 근본적인 메커니즘과 안정적인 장기 생성 방법을 달성하는 문제는 여전히 해결되지 않았습니다. 본 논문에서는 비디오 세계 모델의 내부 표현 역학을 조사하고 복합 오류가 숨겨진 표현의 차원 축소와 밀접하게 관련되어 있음을 밝혀냈습니다. 특히, 모델 표현의 유효 순위는 생성 드리프트가 시작되는 시점에 급격하게 감소하며, 이는 표현 저하와 장기 롤아웃 불안정성 간의 강력한 연관성을 보여줍니다. 또한, 순수한 학습 데이터 확장만으로는 모델이 오류 드리프트에 대한 저항력을 높이는 데 효과적이지 않으며, 이는 주류 확장 패러다임과 상반됩니다. 이러한 문제를 해결하기 위해, 본 연구에서는 잠재 표현을 안정화하고 반복적인 오류 누적을 억제하는 경량 학습 제약 조건인 비디오 표현 정규화를 제안합니다. Diffusion Forcing 방법에 비해, 제안하는 방법은 VBench의 미학 품질 및 이미지 품질 지표에서 각각 38.65에서 55.56, 44.37에서 72.08로 성능 향상을 달성했습니다. 본 연구는 자기 회귀 비디오 드리프트와 모델 내부 표현 간의 첫 번째 연관성을 확립하고, 오류 누적을 정량적으로 측정하기 위해 erank를 사용하며, 비디오 세계 모델에 대한 직관에 어긋나는 확장 한계를 밝히고, 장기 비디오 생성의 견고성을 향상시키는 간단하면서도 효과적인 정규화 전략을 제시합니다.
Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.