2606.31981v1 Jun 30, 2026 cs.CV

LUNA: 스키닝을 넘어선 범용 3D 인간 애니메이션 학습

LUNA: Learning Universal 3D Human Animation Beyond Skinning

Peng Li
Peng Li
Citations: 258
h-index: 6
Rawal Khirodkar
Rawal Khirodkar
Citations: 2,188
h-index: 13
Yuan Dong
Yuan Dong
Citations: 21
h-index: 1
Chen Cao
Chen Cao
Citations: 116
h-index: 5
Wenhan Luo
Wenhan Luo
Citations: 517
h-index: 14
Yi-Ting Guo
Yi-Ting Guo
Citations: 2,143
h-index: 26
Shunsuke Saito
Shunsuke Saito
Citations: 504
h-index: 9
Junxuan Li
Junxuan Li
Citations: 280
h-index: 7
Yuan Liu
Yuan Liu
Citations: 263
h-index: 6

단안 이미지를 기반으로 사실적인 움직임을 갖는 3D 인간 아바타를 생성하는 것은 여전히 주로 선형 블렌드 스키닝(Linear Blend Skinning, LBS) 및 파라메트릭 신체 모델에 의존하며, 이는 표현력을 제한하고 종종 불완전한 적합으로 인해 왜곡을 발생시킨다. 본 논문에서는 LBS 없이 다양한 2D 제어 입력(이미지, 주요 지점, 스케치 등)과 새로운 캐릭터를 직접 3D 가우시안 변형으로 매핑하는 범용 신경망 애니메이션 모델인 LUNA를 제안한다. LUNA의 핵심은 트랜스포머 기반 모션 회귀기로, 전역적인 강체 운동과 미세한 국소적 움직임을 분리하여 일관성 있는 움직임과 미묘한 비강성 효과를 모두 포착한다. 2D-3D 매핑 과정에서 발생하는 고유한 불확실성을 해결하고, 기존 데이터셋에 국한되지 않고 확장하기 위해, LBS 기반 모델을 활용한 지식 전달(knowledge distillation) 및 제한된 적합 데이터와 대규모의 레이블이 없는 비디오 데이터를 모두 활용할 수 있도록 하는 하이브리드 지도 학습 방법을 도입했다. 광범위한 실험 결과, LUNA는 LBS 기반 접근 방식과 비교하여 경쟁력 있는 시각적 충실도를 달성했으며, 사실적인 인간 움직임을 제공하고 다양한 입력 방식에 대한 제로샷(zero-shot) 교차 아이덴티티 일반화 능력을 보여주었다. 현재까지 알려진 바로는, LUNA는 암묵적인 2D 입력을 지원하는 최초의 엔드투엔드 3D 애니메이션 모델이다.

Original Abstract

Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting. We propose LUNA, an LBS-free universal neural animation model that directly maps multiple 2D controls like images, keypoints, sketches, and unseen characters into 3D Gaussian deformations, bypassing explicit body fitting. At its core, a transformer-based motion regressor disentangles global rigid motion from fine-grained local dynamics to capture both coherent movement and subtle non-rigid effects. To resolve the inherent ambiguity of 2D-to-3D lifting while scaling beyond fitted datasets, we introduce hybrid supervision that distills soft structural priors from an LBS teacher and a loss that supports training on both limited fitted data and large in-the-wild unlabeled videos. Extensive experiments show LUNA achieves competitive visual fidelity compared to LBS-based approaches, while delivering realistic human motion and zero-shot cross-identity generalization across diverse driving modalities. To the best of our knowledge, LUNA is the first end-to-end 3D animatable model that supports implicit 2D driving.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!