2605.29488v1 May 28, 2026 cs.CV

AnyMo: 마스크 모델링을 이용한 다중 모달 조건부 동작 생성의 확장

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

Ruibing Hou
Ruibing Hou
Citations: 144
h-index: 6
Hong Chang
Hong Chang
Citations: 123
h-index: 5
Zhuo Li
Zhuo Li
Citations: 2
h-index: 1
Shiguang Shan
Shiguang Shan
Citations: 27
h-index: 2
Yiheng Li
Yiheng Li
Citations: 35
h-index: 2
Yingjie Chen
Yingjie Chen
Citations: 22
h-index: 1
Hao Liu
Hao Liu
Citations: 33
h-index: 2

조건부 인간 동작 생성은 컴퓨터 비전 및 로봇 공학 분야에서 중요한 과제입니다. 상당한 발전에도 불구하고, 현재 방법들은 종종 고정된 모달 구성과 작업별 아키텍처에 의해 제약되며, 다중 모달 상호 작용 및 다중 모달 조건부 합성의 확장 법칙은 충분히 연구되지 않았습니다. 주요 병목 현상은 대규모로 정렬된 동작 데이터의 부족으로, 이는 다양한 제어 신호에 대한 일반화 능력을 제한합니다. 본 연구에서는 5,000시간 이상의 동작 데이터와 320만 개의 시퀀스로 구성된 대규모 고품질 데이터셋인 OmniHuMo를 소개합니다. OmniHuMo는 텍스트, 음성, 음악 및 경로 등 정확하게 정렬된 다중 모달 주석을 포함하고 있습니다. 우리는 OmniHuMo를 활용하여 AnyMo라는 통일된 다중 모달 프레임워크를 제안합니다. AnyMo는 Residual FSQ 기반의 동작 토크나이저와 확장 가능한 마스크 모델링 트랜스포머를 결합하여, 임의의 모달 조합 하에서 고품질 동작 생성을 가능하게 합니다. 광범위한 실험 결과, AnyMo는 높은 충실도의 합성을 달성하는 동시에 공간적 및 스타일적 속성에 대한 유연한 제어를 제공함을 보여줍니다.

Original Abstract

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce OmniHuMo, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose AnyMo, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!