2603.20192v1 Mar 20, 2026 cs.CV

LumosX: 개인 맞춤형 비디오 생성 시스템: 모든 개체와 그 속성을 연결하여 개인화된 비디오 생성

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

Jiazheng Xing
Jiazheng Xing
Citations: 35
h-index: 3
Fei Du
Fei Du
Citations: 29
h-index: 3
Hangjie Yuan
Hangjie Yuan
Citations: 23
h-index: 3
Peng Liu
Peng Liu
Citations: 36
h-index: 3
Hongbin Xu
Hongbin Xu
Citations: 29
h-index: 3
Ruigang Niu
Ruigang Niu
Citations: 396
h-index: 7
Weihua Chen
Weihua Chen
Citations: 65
h-index: 4
Fan Wang
Fan Wang
Citations: 45
h-index: 3
Yong Liu
Yong Liu
Citations: 227
h-index: 6
Hai Ci
Hai Ci
Citations: 404
h-index: 11

최근 딥러닝 분야에서 발전된 확산 모델은 텍스트-비디오 생성 능력을 크게 향상시켜, 전경 및 배경 요소를 정밀하게 제어하면서 개인 맞춤형 콘텐츠 제작을 가능하게 했습니다. 그러나 기존 방법은 명시적인 일관성 확보 메커니즘이 부족하여, 다양한 피사체 간의 정확한 얼굴 속성 정합을 달성하는 데 어려움이 있습니다. 이러한 문제를 해결하기 위해서는 명시적인 모델링 전략과 얼굴-속성 정보를 활용한 데이터 리소스가 필요합니다. 따라서 본 논문에서는 데이터와 모델 설계 모두를 발전시키는 프레임워크인 LumosX를 제안합니다. 데이터 측면에서는, 독립적인 비디오에서 캡션과 시각적 단서를 수집하는 맞춤형 파이프라인을 구축하고, 멀티모달 대규모 언어 모델(MLLM)을 사용하여 피사체별 의존성을 추론하고 할당합니다. 이렇게 추출된 관계 기반 정보는 개인 맞춤형 비디오 생성의 표현력을 향상시키고, 포괄적인 벤치마크 구축을 가능하게 합니다. 모델 측면에서는, 위치 정보를 고려한 임베딩과 개선된 어텐션 메커니즘을 결합하여 Relational Self-Attention 및 Relational Cross-Attention을 통해 명시적인 피사체-속성 의존성을 표현하고, 엄격한 그룹 내 일관성을 유지하며, 서로 다른 피사체 그룹 간의 구분을 강화합니다. 제안하는 벤치마크에 대한 종합적인 평가 결과, LumosX는 정밀한 제어, 피사체 일관성, 의미적 정렬을 갖춘 개인 맞춤형 다중 피사체 비디오 생성에서 최첨단 성능을 달성함을 보여줍니다. 코드 및 모델은 https://jiazheng-xing.github.io/lumosx-home/ 에서 확인할 수 있습니다.

Original Abstract

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face-attribute alignment across subjects remains challenging, as existing methods lack explicit mechanisms to ensure intra-group consistency. Addressing this gap requires both explicit modeling strategies and face-attribute-aware data resources. We therefore propose LumosX, a framework that advances both data and model design. On the data side, a tailored collection pipeline orchestrates captions and visual cues from independent videos, while multimodal large language models (MLLMs) infer and assign subject-specific dependencies. These extracted relational priors impose a finer-grained structure that amplifies the expressive control of personalized video generation and enables the construction of a comprehensive benchmark. On the modeling side, Relational Self-Attention and Relational Cross-Attention intertwine position-aware embeddings with refined attention dynamics to inscribe explicit subject-attribute dependencies, enforcing disciplined intra-group cohesion and amplifying the separation between distinct subject clusters. Comprehensive evaluations on our benchmark demonstrate that LumosX achieves state-of-the-art performance in fine-grained, identity-consistent, and semantically aligned personalized multi-subject video generation. Code and models are available at https://jiazheng-xing.github.io/lumosx-home/.

6 Citations
0 Influential
5.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!