2607.13978v1 Jul 15, 2026 cs.CV

원자적 동작을 이용한 음악 기반 댄스 생성

Music-to-Dance Generation via Atomic Movements

Minghang Zheng
Minghang Zheng
Citations: 95
h-index: 4
Xin Jin
Xin Jin
Citations: 9
h-index: 2
Xinhao Cai
Xinhao Cai
Citations: 35
h-index: 2
Yixuan Sun
Yixuan Sun
Citations: 16
h-index: 2
Qingchao Chen
Qingchao Chen
Citations: 1,507
h-index: 22
Song-Chun Zhu
Song-Chun Zhu
Citations: 35,398
h-index: 92
Yang Liu
Yang Liu
Citations: 18
h-index: 3

음악 기반 댄스 생성이란, 리듬적으로 동기화되고 음악과 의미적으로 일관된 인간의 움직임을 생성하는 것을 목표로 합니다. 최근의 신경망 기반 접근 방식은 뛰어난 시각적 사실감을 달성했지만, 일반적으로 움직임을 연속적인 신호로 모델링하며 그 구성적 특성을 간과하여 생성된 댄스가 구조적으로 일관성이 없고 제어하기 어렵습니다. 본 연구에서는 구조를 고려한 프레임워크를 소개합니다. 이 프레임워크는 안무를 '원자적 동작'의 시퀀스로 모델링하는데, 원자적 동작은 의미적으로 해석 가능한 움직임 이벤트이며 댄스의 구성 요소 역할을 합니다. 이러한 원자적 동작 어휘를 구축하기 위해 먼저 대규모 댄스 데이터를 분할하고 클러스터링하여 원자적 동작 그룹으로 나눕니다. 그런 다음 대규모 언어 모델을 사용하여 클러스터를 의미적으로 재표시하고 개선하여 해석 가능하고 재사용 가능한 원자적 동작 집합을 얻습니다. 이러한 원자적 동작 주석을 기반으로, 인간의 안무 과정을 반영하는 두 단계 생성 프레임워크를 설계했습니다. 첫 번째 단계인 원자적 동작 계획 단계에서는 모델이 입력 음악에 따라 원자적 동작의 유형, 지속 시간 및 타이밍을 예측하여 상징적인 댄스 할당을 형성합니다. 두 번째 단계인 완성 단계에서는 전환(transition)을 고려한 생성기가 계획된 구조를 기반으로 부드럽고 스타일적으로 일관성 있는 움직임을 합성합니다. 광범위한 실험 결과, 제안하는 방법은 기존의 기준 모델보다 구조적 일관성, 리듬 정렬 및 인지적 자연스러움이 크게 향상된 댄스를 생성하며, 명시적인 구조 표현을 통해 해석 가능성과 제어 가능한 편집 기능을 제공한다는 것을 보여줍니다.

Original Abstract

Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continuous signal and neglect its compositional nature, making generated dances structurally incoherent and difficult to control. In this work, we introduce a structure-aware framework that models choreography as a sequence of atomic movements-semantically interpretable motion events that serve as the building blocks of dance. To construct this atomic movement vocabulary, we first segment large-scale dance data and cluster them into atomic movement groups. We then employ a large language model to semantically relabel and refine the clusters, yielding a set of interpretable and reusable atomic movements. Based on these atomic movement annotations, we design a two-stage generation framework that mirrors the human choreography process. In the atomic movement planning stage, the model predicts the type, duration, and timing of atomic movements conditioned on the input music, forming a symbolic dance allocation. In the completion stage, a transition-aware generator synthesizes smooth and stylistically coherent motion conditioned on the planned structure. Extensive experiments demonstrate that our method produces dances with significantly improved structural coherence, rhythmic alignment, and perceptual naturalness compared to existing baselines, while providing enhanced interpretability and controllable editing through explicit structural representation.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!