2606.11670v1 Jun 10, 2026 cs.CV

ARGUS: 피사체 보존 비디오 생성을 위한 다중 시점 아이덴티티 모자이크 주입 기술

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

Chengzhuo Tong
Chengzhuo Tong
Citations: 48
h-index: 3
Yuanxing Zhang
Yuanxing Zhang
Citations: 951
h-index: 19
Zijie Meng
Zijie Meng
Citations: 115
h-index: 6
Xiaoqiang Liu
Xiaoqiang Liu
Citations: 333
h-index: 7
Jiwen Liu
Jiwen Liu
Citations: 90
h-index: 5
Yulong Xu
Yulong Xu
Citations: 40
h-index: 2
Pengfei Wan
Pengfei Wan
Citations: 71
h-index: 3
Yufei Liu
Yufei Liu
Citations: 49
h-index: 3

피사체를 유지하는 비디오 생성은 정면 얼굴의 유사성만으로는 해결되지 않습니다. 생성된 인물은 움직임, 큰 시야각 변화, 표정 변화, 가려짐, 크기 변화 및 텍스트, 첫 번째 프레임, 그리고 아이덴티티 참조 간의 충돌 등 다양한 요인에서도 인식 가능해야 합니다. 우리는 이러한 주요 병목 현상이 포인트 참조 패러다임에 있다고 주장합니다. 이 방식은 아이덴티티를 단일 정적 관찰로 축소시키는데, 이는 자세, 액세서리, 조명, 배경 및 카메라 통계와 얽혀 있습니다. 본 논문에서는 Stacked Multi-View Identity Mosaic Injection (SMII)을 중심으로 하는 Wan 기반 프레임워크인 Argus를 소개합니다. SMII는 MLLM(Multi-modal Large Language Model)에서 선택한 이미지/비디오 아이덴티티 정보를 3x3으로 구성된 모자이크로 변환하고, 이 모자이크를 현재 디퓨전 시간과 동기화하여 Wan의 기본 토큰 공간에 읽기 전용 메모리로 주입합니다. 이를 통해 아이덴티티는 외부 어댑터나 단일 참조 이미지 대신 소형의 동적 분포로 표현됩니다. SMII 주변에는 MLLM Identity Director가 정보적인 아이덴티티 순간을 선택하고 조건 충돌을 해결하며, no-cross-pair counterfactual training, Temporal Identity Annealing 및 Adaptive Self-Likeness Guidance를 통해 페어링된 피사체-비디오 감독 없이도 견고성을 향상시킵니다. 또한, 공개 인물 아이덴티티 스트레스 벤치마크인 HardID-Celeb를 공개하고, 큰 시야각과 첫 번째 프레임 가려짐에 대한 강건성을 평가하기 위한 YawScore 및 OccScore를 소개합니다. Argus는 OpenS2V-Eval Human-Domain에서 최첨단 결과를 달성하여 총 64.38점, FaceSim 71.86점, NexusScore 51.62점, NaturalScore 79.14점을 기록했습니다. HardID-Celeb 데이터셋에서는 Argus가 FaceSim 점수를 76.80점으로 얻었으며, 가장 강력한 기본 모델보다 YawScore와 OccScore를 각각 12.60점과 15.10점 향상시켰습니다. 이는 동적 아이덴티티 메모리와 대규모 반사실적 자기 지도 학습이 피사체 보존 비디오 생성에 매우 효과적임을 보여줍니다.

Original Abstract

Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references. We argue that the central bottleneck is the point-reference paradigm, which collapses identity into a single static observation entangled with pose, accessories, lighting, background, and camera statistics. We introduce Argus, a Wan-based framework centered on Stacked Multi-View Identity Mosaic Injection (SMII). SMII converts MLLM-selected image/video identity evidence into a 3*3 stacked mosaic, synchronizes the mosaic with the current diffusion time, and injects it as negative-time read-only memory in Wan's native token space. This turns identity from an external clean adapter or a single reference image into a compact dynamic distribution. Around SMII, an MLLM Identity Director selects informative identity moments and resolves condition conflicts, while no-cross-pair counterfactual training, Temporal Identity Annealing, and Adaptive Self-Likeness Guidance improve robustness without paired subject-video supervision. We further release HardID-Celeb, a public-figure identity-stress benchmark, and introduce YawScore and OccScore to probe large-yaw and first-frame-occlusion robustness. Argus achieves state-of-the-art results on OpenS2V-Eval Human-Domain, reaching 64.38 Total Score, 71.86 FaceSim, 51.62 NexusScore, and 79.14 NaturalScore. On HardID-Celeb, Argus obtains 76.80 FaceSim and improves YawScore and OccScore by 12.60 and 15.10 points over the strongest baselines, demonstrating that dynamic identity memory and large-scale counterfactual self-supervision are highly effective for subject-preserving video generation.

6 Citations
0 Influential
9.5 Altmetric
53.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!