2608.02214v1 Aug 03, 2026 cs.CV

VARPose: 시각적 자기회귀 모델링을 통한 유연한 2차원 자세 밀도화 방법 - 향상된 3차원 정보 추출

VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting

Danel Zeng
Danel Zeng
Citations: 3
h-index: 1
Kai Pu
Kai Pu
Citations: 1
h-index: 1
Tiantian Yang
Tiantian Yang
Citations: 0
h-index: 0

시각적 자기회귀 모델링(VAR)은 다음 단계 예측을 통해 자연 이미지 생성 분야에서 뛰어난 성능을 보여주었지만, 인간 골격과 같은 구조화된 데이터에 적용된 사례는 아직 미미합니다. 본 논문에서는 VARPose를 제안하며, 이는 2차원 희소 자세 데이터를 적응적으로 밀도화하여 3차원 정보 추출 모델에 활용될 수 있는 해부학적 정보를 풍부하게 제공하는 방법입니다. 저희 연구의 핵심 기여는 두 가지입니다. 첫째, 다양한 밀도의 자세를 단일 하이브리드 코드북과 잔차 양자화 전략을 사용하여 통합된 다중 스케일 이산 표현으로 인코딩하는 '밀도-불변 자세 토크나이저(GPT)'를 제안합니다. 실험 결과는 이 표현 방식의 뛰어난 일반화 성능을 입증합니다. 투영 과정으로부터 표현 방식을 분리함으로써, 재학습된 디코더와 고정된 코드북을 사용하여 새로운 밀도의 자세를 성공적으로 디코딩할 수 있습니다. 둘째, 'joint density(관절 밀도)'를 'scale(스케일)'로 간주하는 통합 자기회귀 모델인 UniSkelar를 제안합니다. UniSkelar는 가장 희소한 자세를 조건으로 하여, 거칠기에서 세밀함으로 점진적으로 다음 밀도 수준의 토큰 시퀀스를 예측하도록 학습됩니다. VARPose는 최첨단 방법보다 뛰어난 성능을 보일 뿐만 아니라, 아직 관찰되지 않은 밀도로 일반화될 수 있으며, 2차원 자세 밀도화를 통해 3차원 자세 추정 및 인간 메시 복구와 같은 후속 작업에서 실질적인 성능 향상을 가져다줍니다. 저희의 코드와 모델은 다음 링크에서 확인할 수 있습니다: https://github.com/BRL-SYSU/VARPose.git.

Original Abstract

Visual AutoRegressive Modeling (VAR) has excelled in natural image generation via next-scale prediction, but its use on topology-structured data like human skeletons is still unexplored. VARPose is proposed to adaptively densify 2D sparse poses, thereby enriching the anatomical information available for 3D lifting models. Our core contributions are twofold. First, we introduce a Granularity-agnostic Pose Tokenizer (GPT), which employs a single hybrid codebook and a residual quantization strategy to encode poses of varying densities into a unified, multi-scale discrete representation. Our results demonstrate the strong generalizability of this representation. By decoupling the representation from the projection, we can successfully decode novel pose granularities using a frozen codebook with a retrained decoder. Second, we propose UniSkelar, a unified autoregressive model that treats "joint density" as "scale". UniSkelar learns to predict the token sequence for the next density level in a coarse-to-fine manner, conditioned on the sparsest pose. VARPose not only outperforms state-of-the-art methods and generalizes to unseen granularities, but also confers tangible performance gains on downstream tasks, such as 3D Pose Estimation and Human Mesh Recovery, through 2D pose densification. Our code and model are available at https://github.com/BRL-SYSU/VARPose.git.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!