2607.27581v1 Jul 30, 2026 cs.LG

MUGEN: 효율적인 동작 이해 및 생성을 위한 통합 프레임워크

MUGEN: A Unified Framework for Efficient Motion Understanding and Generation

Zhankai Ye
Zhankai Ye
Citations: 1
h-index: 1
Yukai Jin
Yukai Jin
Citations: 164
h-index: 3
Shangqian Gao
Shangqian Gao
Citations: 13
h-index: 1
Xin Liu
Xin Liu
Citations: 1
h-index: 1
Bofan Li
Bofan Li
Citations: 1,623
h-index: 22
Yusen Wu
Yusen Wu
Citations: 65
h-index: 3
Bingyang Wei
Bingyang Wei
Citations: 0
h-index: 0
Fangyi Li
Fangyi Li
Citations: 15
h-index: 2

인간의 동작을 언어로, 그리고 언어를 동작으로 연결하는 것은 인간 행동을 이해하고 생성하며 소통할 수 있는 물리적 AI 시스템 개발에 있어 핵심적인 단계입니다. 기존의 통합된 동작-언어 시스템은 공유된 이산적인 동작 코드북을 사용하여 두 방향을 연결했지만, 양자화는 생성 품질을 제한합니다. 성능이 뛰어난 생성 모델들은 이러한 품질 저하를 해결하기 위해 다양한 방법을 사용하지만, 이는 비용 증가로 이어집니다. 예를 들어, 스택된 잔여 코드북은 표현력을 확장하고, 마스킹된 디코딩 단계, 긴 오토 회귀 과정, 그리고 수십에서 수백 단계를 거치는 노이즈 제거 체인은 추론 시간을 늘립니다. 또한, 연속적인 잠재 변수를 사용하는 모델조차도 반복적인 확산 과정을 통해 잠재 변수를 얻습니다. 하지만 이러한 모든 디코딩 메커니즘은 동작 이해에는 기여하지 못합니다. 따라서 우리는 비용이 전혀 들지 않는 통합된 동작-언어 프레임워크인 MUGEN을 제안합니다. 단일 어댑티브 길이 오토인코더는 임의 길이의 동작을 몇 개의 연속적인 잠재 변수로 압축하며, 이는 시스템의 유일한 동작 표현입니다. 언어 모델은 텍스트-투-모션 생성 시 이러한 잠재 변수를 생성하고, 동작 이해 시에는 이를 읽어들입니다. 깊이 라우팅된 숨겨진 상태는 각 잠재 변수가 필요한 트랜스포머 깊이에서 정보를 읽을 수 있도록 하며, 보정된 헤드는 전체 잠재 변수 집합에 대한 확률 분포를 예측하여 단일 샘플링으로 텍스트 조건부의, 그리고 슬롯 간의 다양한 표현을 생성합니다. MUGEN은 K번의 언어 모델 단계와 하나의 디코더 패스를 통해 동작-언어 기반 모델보다 HumanML3D 데이터셋에서 더 나은 FID 점수를 달성하고, 표준 평가 기준으로 실제 동작 참조를 능가하는 높은 검색 정확도를 제공하며, SnapMoGen 데이터셋에서 모든 검색 및 정렬 지표에서 기존의 이산 토큰 기반 최고 성능을 뛰어넘습니다. 또한, CIDEr 및 BLEU@4 점수에서도 최적의 결과를 보입니다.

Original Abstract

Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two directions through a shared discrete motion codebook, but quantization limits generation quality. The strongest generators buy quality back at growing cost: stacked residual codebooks enlarge the representation; masked decoding stages, long autoregressive rollouts, and denoising chains of tens to hundreds of steps stretch inference; even the continuous-latent designs among them reach their latent only through an iterative diffusion head; and none of this decoding machinery serves understanding. We therefore propose MUGEN, a unified motion--language framework that pays neither cost: no codebook, one draw. A single adaptive-length autoencoder compresses any-length motion into a few continuous latent slots, the system's only motion representation: the language model generates them for text-to-motion and reads them back for motion understanding. Depth-routed hidden states let each slot read from the transformer depth it needs, and a calibrated head predicts a joint distribution over the full latent set, so a single draw carries the text-conditional, cross-slot variation a description permits. At a decoding cost of K language-model steps, one draw, and one decoder pass, MUGEN leads language-model baselines on FID on HumanML3D while raising retrieval precision above the real-motion reference under the standard evaluator, achieves the best CIDEr and BLEU@4 scores, and surpasses the discrete-token state of the art on every retrieval and alignment metric on SnapMoGen.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!