FashionPose: 통합된 텍스트 기반 패션 합성 모델 - 합동적인 기하학적 및 광도학적 제어
FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control
현실적이고 정밀하게 제어 가능한 의상 합성은 패션 전자 상거래에 필수적이지만, 이는 인간 자세의 기하학과 환경 광도의 정확한 조화를 요구합니다. 기존의 자세 기반 프레임워크는 두 가지 근본적인 한계점을 가지고 있습니다. 첫째, 이들은 시판되는 추정기에서 얻은 미리 정의된 골격에 크게 의존하여 의미적 유연성을 제한하며, 둘째, 대부분 중립적인 조명 조건 하에서 스튜디오와 유사한 이미지를 생성하는 데 집중하여 자연어 묘사로 표현되는 복잡하고 장면 특유의 조명과 기하학적 구성을 일치시키지 못합니다. 이러한 격차를 해소하기 위해, 우리는 통합된 언어 기반 인터페이스 내에서 기하학과 광도를 조정하는 순차적 아키텍처인 FashionPose를 제안합니다. 기존 프레임워크와 달리, 우리의 프레임워크는 분리되지만 시너지 효과가 있는 전략을 사용합니다: (1) 텍스트 의미를 명시적인 기하학적 공간에 연결하는 양방향 대비 정렬 메커니즘으로, 템플릿이 필요 없는 자세 생성이 가능하며; (2) 이러한 기하학적 사전 지식을 고해상도 이미지로 변환하면서 미세한 외관을 유지하는 ID 기반 합성 모듈이며; (3) 생성된 자세를 공간 기준점으로 활용하여 환경 인지 쉐이딩을 달성하는 프롬프트 기반 재조명 모듈입니다. 이러한 계층적 설계는 고수준 지시 사항을 일관된 시각적 표현으로 변환하여 구조적인 정확성과 분위기 조화를 모두 보장합니다. 이 새로운 패러다임을 지원하기 위해, 우리는 4만 개 이상의 캡션-키포인트 쌍으로 구성된 데이터셋인 PoseCap을 구축했습니다. 광범위한 실험 결과는 FashionPose가 기존의 성능 지표에서 자세 정확도와 물리적 현실성 측면에서 뛰어난 성능을 보이며, 개인 맞춤형 및 장면 인지 가상 패션 디스플레이를 위한 강력한 솔루션을 제공한다는 것을 보여줍니다.
Realistic and controllable garment synthesis is essential for fashion e-commerce, yet it demands precise coordination between human pose geometry and environmental photometry. Conventional pose-guided frameworks suffer from two fundamental limitations: they rely heavily on predefined skeletons from off-the-shelf estimators, restricting semantic flexibility; and they predominantly focus on studio-like generation under neutral lighting, failing to reconcile geometric configurations with complex, scene-specific illumination described in natural language. To bridge this gap, we propose FashionPose, a cascaded architecture that reconciles geometric and photometric control within a unified language-driven interface. Unlike conventional frameworks, our framework employs a decoupled yet synergistic strategy: (1) a bidirectional contrastive alignment mechanism that grounds textual semantics into an explicit geometric manifold, enabling template-free pose generation; (2) an identity-anchored synthesis module that translates these geometric priors into high-fidelity imagery while preserving fine-grained appearance; and (3) a prompt-conditioned relighting module that leverages the generated pose as a spatial anchor to achieve environment-aware shading. This hierarchical design effectively transforms high-level instructions into consistent visual representations, ensuring both structural precision and atmospheric harmony. To facilitate this paradigm, we construct PoseCap, a dataset with over 40,000 caption-keypoint pairs. Extensive experiments demonstrate that FashionPose outperforms existing benchmarks in pose accuracy and physical realism, providing a robust solution for personalized, scene-aware virtual fashion displays.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.