2607.28611v1 Jul 30, 2026 cs.CV

Chimera: 친칠라 스케일링을 위한 하이브리드 시각적 디퓨전 트랜스포머 설계

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Yiran Xu
Yiran Xu
Citations: 235
h-index: 4
Hailin Jin
Hailin Jin
Citations: 11,990
h-index: 55
Chongjian Ge
Chongjian Ge
Citations: 47
h-index: 2
Hanwen Jiang
Hanwen Jiang
Citations: 1,702
h-index: 14
Tianyu Wang
Tianyu Wang
Citations: 0
h-index: 0
Jiuxiang Gu
Jiuxiang Gu
Citations: 128
h-index: 6
Ziwen Chen
Ziwen Chen
Citations: 135
h-index: 4
Shaoteng Liu
Shaoteng Liu
Citations: 383
h-index: 5
Jing Shi
Jing Shi
Citations: 74
h-index: 3
Yicong Hong
Yicong Hong
Citations: 1,602
h-index: 6
Zefan Cai
Zefan Cai
Citations: 148
h-index: 7
Hao Tan
Hao Tan
Citations: 143
h-index: 3

시각적 생성은 점차 고해상도 이미지, 긴 비디오 및 다중 모드 컨텍스트를 요구하며, 이는 전체 어텐션의 2차원 비용으로 인해 어려움을 야기합니다. 본 논문에서는 체계적인 스케일링 방식을 갖춘 하이브리드 시각적 디퓨전 백본인 Chimera를 소개합니다. Chimera는 위치 임베딩 없이 텍스트, 이미지 및 비디오 토큰을 하나의 순서대로 처리하며, 긴 컨텍스트 상태 추적을 위한 Kimi Delta Attention (KDA) (O(N) 복잡도), 직접적인 글로벌 상호 작용을 위한 인터리브된 Multi-head Latent Attention (MLA), 그리고 로컬 시공간 컨텍스트를 위한 모달리티 인지형 Short Convolutions을 결합합니다. Sparse Mixture-of-Experts (MoE) 레이어는 활성화된 연산량을 제어하면서 모델 용량을 확장합니다. 이 이질적인 아키텍처를 스케일링하기 위해, 우리는 각 텐서의 기능적 팬-인과 모델 깊이에 따라 너비와 깊이를 통해 하이퍼파라미터를 전송하는 모듈 단위 방식인 HeteroP를 도입했습니다. HeteroP는 일관되게 조정된 패밀리를 생성하며, 이를 통해 활성화된 모델 크기, 학습 토큰 수 및 이미지-비디오 데이터 비율에 대한 친칠라 스타일의 최적화된 법칙을 적용합니다. 이러한 법칙에 따라, 우리는 110억 개의 파라미터를 가진 Chimera를 20억 개의 활성화된 파라미터로 학습했습니다. 실험 결과는 세 가지 중요한 결과를 보여줍니다. 첫째, 사전 학습 디퓨전 손실 측면에서, 본 백본은 동일한 전체 어텐션 Wan-2.1 2B 기준 모델보다 연산 효율성이 1.7배 더 높으며, 전체 시스템은 7.3배 더 높은 효율성을 달성합니다. 둘째, 길이별 미세 조정 없이도 Chimera는 5초의 학습 클립에서 30초 비디오로 확장할 수 있으며, 마지막 5초 동안 FID 점수가 6.5%만 감소합니다. 셋째, 도출된 법칙은 연산 최적화된 이미지 사전 학습이 활성화된 모델 크기와 학습 토큰 수를 거의 균등하게 분배하는 반면, 비디오 사전 학습은 더 큰 예산에서 약간 더 많은 모델 크기를 선호한다는 것을 보여줍니다. 이러한 결과는 효율적인 장기 컨텍스트 디퓨전 아키텍처를 설계하고 확장하기 위한 기반을 제공합니다.

Original Abstract

Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.

0 Citations
0 Influential
27.5 Altmetric
137.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!