ViP-Rig: 시각적 프롬프트 기반 제어 가능한 리깅
ViP-Rig: Visual-Prompted Controllable Rigging
리깅은 본질적으로 작업 의존적인 과정이며, 동일한 메시에 대해서도 애니메이션 작업에 따라 다른 골격 구조와 변형 동작이 필요할 수 있습니다. 실제로 아티스트는 초기 리깅을 검토하고, 특정 애니메이션 요구 사항을 충족하기 위해 골격 구조와 변형 동작을 반복적으로 수정합니다. 기존의 자동화된 방법은 주로 기하학적 정보를 기반으로 실행 가능한 리깅을 생성하지만, 결과적인 골격 및 변형 동작에 대한 명시적인 제어는 제한적입니다. 본 연구에서는 사용자가 그린 또는 편집한 2차원 골격 및 강성 프롬프트에서 추출한 특징을 활용하여 기존의 사전 학습된 모델에 주입함으로써, 프롬프트 기반 리깅과 결과 지향적 편집을 모두 지원하는 시각적 프롬프트 기반 프레임워크인 ViP-Rig를 제안합니다. 구체적으로, ViP-Rig는 골격 생성 단계와 스키닝 예측 단계의 두 부분으로 구성됩니다. 첫 번째 단계에서는 덴스-투-컴팩트 비주얼 프롬프트 인코딩을 사용하여 골격 스케치를 처리하여 컴팩트하고 고정된 길이의 조건부 토큰을 생성합니다. 생성된 토큰은 게이티드 어댑터를 통해 사전 학습된 자기 회귀 생성기에 주입되어, 생성기의 기하학적 사전 지식을 유지하면서 관절 배치 및 분기 구조를 제어합니다. 두 번째 단계에서는 동일한 시각적 인코딩 방식을 사용하여 강성 맵을 처리하며, 사전 학습된 스키닝 백본은 고정 상태로 유지됩니다. 결과적으로 생성된 토큰은 포인트 스트림과 관절 스트림에 대칭적으로 주입되어, 포인트-관절 호환성과 결과적인 스키닝 가중치를 조절합니다. Articulation-XL2.0 데이터셋에서의 실험 및 ModelsResource 데이터셋에서의 제로샷 평가를 통해, ViP-Rig가 프롬프트 기반 평가 하에서 기하학적 정보에 의존하는 기존 방법보다 대상 골격과 스키닝 가중치를 더 정확하게 복원한다는 것을 확인했습니다. 또한, 정성적인 결과는 프롬프트 우선 리깅 및 결과 지향적 편집 모두에서 명시적이고 국소적인 제어가 가능하다는 것을 보여줍니다.
Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an initial rig and repeatedly edit its skeletal structure and deformation behavior to meet specific animation requirements. Existing automatic methods primarily generate a plausible rig from geometry, offering limited explicit control over the resulting skeleton and deformation behavior. In this work, we present ViP-Rig, a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones. Specifically, ViP-Rig consists of two stages, Skeleton Generation and Skinning Prediction. In the first stage, the skeletal sketch is processed by the Dense-to-Compact Visual Prompt Encoding to produce compact, fixed-length conditioning tokens. The resulting tokens are injected into a frozen pretrained autoregressive generator through gated adapters to control joint placement and branching structure while preserving the generator's geometric prior. In the second stage, the rigidity map is processed using the same visual encoding design, while the pretrained skinning backbone remains frozen. The resulting tokens are symmetrically injected into the point and joint streams to modulate point-joint compatibility and the resulting skinning weights. Experiments on Articulation-XL2.0 and zero-shot evaluation on ModelsResource show that ViP-Rig more accurately recovers target skeletons and skinning weights than geometry-conditioned baselines under prompt-guided evaluation. Qualitative results further demonstrate explicit and localized control in both prompt-first rigging and result-guided editing.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.