포즈-ICL: 3차원 인지 In-Context Learning을 활용한 포즈 제어 가능 객체 맞춤 설정
Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization
객체 맞춤 설정은 현대 이미지 생성의 핵심적인 과제입니다. 사용자는 몇 개의 참조 이미지와 텍스트 프롬프트를 제공하여 특정 객체를 원하는 장면에서 생성할 수 있습니다. 그러나 기존 방법들은 여전히 맞춤 설정된 객체의 효과적인 포즈 제어에 어려움을 겪습니다. 실제로, 이들은 종종 부정확한 포즈를 보이거나, 서로 다른 포즈 간의 일관성이 부족한 모습을 나타냅니다. 이러한 한계는 2차원 기반 모델이 객체를 입체적으로 이해하는 데 여전히 중요한 과제가 있음을 시사합니다. 이러한 문제를 해결하기 위해, 저희는 3차원 인지 In-Context Learning (ICL)을 활용하여 여러 쌍의 이미지-포즈 참조를 통해 새로운 객체에 직접적으로 적용할 수 있는, 튜닝이 필요 없는 프레임워크인 Pose-ICL을 제안합니다. 핵심 메커니즘인 Surface-Anchored Position Embedding (SAPE)은 이미지 토큰을 볼륨 바운딩 박스의 표면 좌표에 연결하여 모델에게 명시적인 3차원 인식을 부여합니다. 특별히 설계된 개선 사항들은 기존의 DiT 모델과의 원활한 호환성을 보장합니다. 3D 자산 및 실제 객체에 대한 광범위한 실험 결과는 Pose-ICL이 현재 방법들보다 포즈 정확도와 객체 일관성 모두에서 현저하게 우수한 성능을 보임을 입증합니다.
Subject Customization is a foundational task in modern image generation. By providing a few reference images and a text prompt, users can generate images of a specific object in any desired scene. However, existing methods still struggle to achieve effective pose control for customized subjects. In practice, they often exhibit inaccurate poses or inconsistent cross-pose appearances. These limitations suggest that understanding objects in a volumetric manner remains a significant challenge for 2D-native backbones. To address this challenge, we propose Pose-ICL, a tuning-free framework that leverages 3D-aware In-Context Learning (ICL) to directly adapt to new subjects through multiple paired image-pose references. Its core mechanism,Surface-Anchored Position Embedding (SAPE), equips the model with explicit 3D awareness by anchoring image tokens to the surface coordinates of a volumetric bounding box. Dedicated refinements ensure its seamless compatibility with existing DiT models. Extensive evaluations on both 3D assets and real-world subjects demonstrate that Pose-ICL significantly outperforms current methods in both pose accuracy and identity consistency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.