Chimera: 친칠라 스케일링을 위한 하이브리드 시각적 디퓨전 트랜스포머 설계
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
시각적 생성은 점차 고해상도 이미지, 긴 비디오 및 다중 모드 컨텍스트를 요구하며, 이는 전체 어텐션의 2차원 비용으로 인해 어려움을 야기합니다. 본 논문에서는 체계적인 스케일링 방식을 갖춘 하이브리드 시각적 디퓨전 백본인 Chimera를 소개합니다. Chimera는 위치 임베딩 없이 텍스트, 이미지 및 비디오 토큰을 하나의 순서대로 처리하며, 긴 컨텍스트 상태 추적을 위한 Kimi Delta Attention (KDA) (O(N) 복잡도), 직접적인 글로벌 상호 작용을 위한 인터리브된 Multi-head Latent Attention (MLA), 그리고 로컬 시공간 컨텍스트를 위한 모달리티 인지형 Short Convolutions을 결합합니다. Sparse Mixture-of-Experts (MoE) 레이어는 활성화된 연산량을 제어하면서 모델 용량을 확장합니다. 이 이질적인 아키텍처를 스케일링하기 위해, 우리는 각 텐서의 기능적 팬-인과 모델 깊이에 따라 너비와 깊이를 통해 하이퍼파라미터를 전송하는 모듈 단위 방식인 HeteroP를 도입했습니다. HeteroP는 일관되게 조정된 패밀리를 생성하며, 이를 통해 활성화된 모델 크기, 학습 토큰 수 및 이미지-비디오 데이터 비율에 대한 친칠라 스타일의 최적화된 법칙을 적용합니다. 이러한 법칙에 따라, 우리는 110억 개의 파라미터를 가진 Chimera를 20억 개의 활성화된 파라미터로 학습했습니다. 실험 결과는 세 가지 중요한 결과를 보여줍니다. 첫째, 사전 학습 디퓨전 손실 측면에서, 본 백본은 동일한 전체 어텐션 Wan-2.1 2B 기준 모델보다 연산 효율성이 1.7배 더 높으며, 전체 시스템은 7.3배 더 높은 효율성을 달성합니다. 둘째, 길이별 미세 조정 없이도 Chimera는 5초의 학습 클립에서 30초 비디오로 확장할 수 있으며, 마지막 5초 동안 FID 점수가 6.5%만 감소합니다. 셋째, 도출된 법칙은 연산 최적화된 이미지 사전 학습이 활성화된 모델 크기와 학습 토큰 수를 거의 균등하게 분배하는 반면, 비디오 사전 학습은 더 큰 예산에서 약간 더 많은 모델 크기를 선호한다는 것을 보여줍니다. 이러한 결과는 효율적인 장기 컨텍스트 디퓨전 아키텍처를 설계하고 확장하기 위한 기반을 제공합니다.
Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.