2607.29180v1 Jul 31, 2026 cs.CV

MoRAE: 텍스트-모션 생성을 위한 흐름 친화적인 자기 지도 학습 기반 잠재 공간

MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

Mingyi Shi
Mingyi Shi
Citations: 129
h-index: 4
Yifei Zhu
Yifei Zhu
Citations: 1
h-index: 1
Yangyang Cai
Yangyang Cai
Citations: 1
h-index: 1
Miao Cheng
Miao Cheng
Citations: 6
h-index: 2
Y. Kitamura
Y. Kitamura
Citations: 3,381
h-index: 28
Taku Komura
Taku Komura
Citations: 141
h-index: 5

텍스트-모션 생성은 의미적으로 정확하고, 시간적으로 일관되며, 물리적으로 타당한 움직임을 만들어내야 합니다. 자연스러운 접근 방식은 먼저 동작 데이터를 구조화된 의미 공간으로 투영한 다음, 해당 공간 내에서 생성 모델을 학습하는 것입니다. 이러한 패러다임은 Representation Autoencoders (RAEs)를 통해 이미지 생성 분야에서 큰 성공을 거두었으며, 여기서는 고정된 자기 지도 학습 인코더가 확산 또는 흐름 모델이 학습할 수 있는 의미적 특징을 제공합니다. 그러나 Motion-JEPA를 고정된 인코더로 사용하는 이러한 패러다임을 동작 공간에 직접 적용하면 심각한 실패를 초래합니다. 우리는 이러한 실패의 원인을 기하학적으로 분석하고, 두 가지 동작 특유의 문제점을 파악했습니다. (1) JEPA 특징 공간은 스펙트럼이 불안정하여 가우시안 분포에서 실제 데이터로의 변환을 어렵게 만듭니다. (2) 심지어 잘 조정된 스펙트럼이라 하더라도, 흐름 잔차가 디코더에 민감한 방향과 일치하는 경향이 있으며, 이로 인해 작은 잠재 오류가 디코딩 후 큰 동작 왜곡으로 증폭됩니다. 이러한 통찰력을 바탕으로 우리는 MoRAE를 제안합니다. MoRAE는 위에서 언급된 두 가지 문제점에 대해 각각 해결책을 제시합니다. 컴팩트한 병목 구조는 구조화된 JEPA 표현을 추출하면서 약하고 중복되는 방향을 제거하여 잠재 스펙트럼을 변환 안정적인 상태로 만듭니다. 동작 결합 학습은 유지된 잠재 공간의 기하학적 구조를 디코더와 일치시켜, 디코딩 후 발생하는 특징적인 흐름 오류를 줄입니다. 이러한 흐름 친화적인 잠재 공간을 통해 표준 비자동 회귀 Flow-Matching DiT 모델이 최첨단 성능을 달성했습니다.

Original Abstract

Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into a structured semantic space and then train a generative model within that space. Such a paradigm has been highly successful in image generation through Representation Autoencoders (RAEs), where a frozen self-supervised encoder provides semantic features for diffusion or flow models to learn from. However, direct transfer of such a paradigm to motion space using Motion-JEPA as the frozen encoder fails dramatically. We diagnose this failure geometrically and identify two motion-specific bottlenecks: (1) the JEPA feature space is spectrally ill-conditioned, making the Gaussian-to-data transport unstable; and (2) even with a well-conditioned spectrum, flow residuals tend to align with decoder-sensitive directions, where small latent errors are amplified into large motion artifacts after decoding. Based on these insights, we propose MoRAE. MoRAE addresses the two bottlenecks separately. A compact bottleneck distills the structured JEPA representation while removing weak and redundant directions, bringing the latent spectrum into a transport-stable regime. Motion-coupled training then aligns the retained latent geometry with the decoder, making characteristic flow errors less costly after decoding. With this flow-friendly latent, a standard non-autoregressive Flow-Matching DiT achieves state-of-the-art performance.

0 Citations
0 Influential
14 Altmetric
70.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!