2608.03082v1 Aug 04, 2026 cs.CV

DiverseDiT++: 확산 트랜스포머(Diffusion Transformer)에서 표현 다변성을 정량화, 분석 및 촉진하는 연구

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers

Xiaomeng Yang
Xiaomeng Yang
Citations: 196
h-index: 6
Zhiyu Tan
Zhiyu Tan
Citations: 958
h-index: 14
Mengping Yang
Mengping Yang
Citations: 112
h-index: 5
Hao Li
Hao Li
Citations: 306
h-index: 10
Binglei Li
Binglei Li
Citations: 16
h-index: 3
Zhizhong Huang
Zhizhong Huang
Fudan University
Citations: 1,194
h-index: 16
Junping Zhang
Junping Zhang
Citations: 9
h-index: 2

최근 확산 트랜스포머(DiTs) 기술의 발전은 뛰어난 확장성 덕분에 시각적 합성 분야에서 괄목할 만한 성과를 가져왔습니다. DiTs가 의미 있는 내부 표현을 효과적으로 학습하도록 돕기 위해, REPA와 같은 최근 연구에서는 표현 정렬을 위해 사전 학습된 외부 인코더를 활용합니다. 그러나 DiTs 내에서의 표현 학습 메커니즘은 아직까지 명확하게 이해되지 않고 있습니다. 이에 본 논문에서는 블록 단위의 표현 다변성을 정량화하여 DiTs의 표현 동역학에 대한 체계적인 분석을 수행합니다. 특히, 다양한 블록 간의 표현 차이를 측정하기 위한 새로운 지표인 가중 다양성 점수(Weighted Diversity Score, WDS)를 제안합니다. 다양한 설정에서 내부 표현의 진화 및 영향에 대한 광범위한 조사를 통해, 블록 간의 표현 다변성이 DiTs에서의 효과적인 표현 학습에 중요한 요소임을 밝혀냅니다. 더욱 중요하게는, WDS가 다양한 설정, 모델 크기 및 훈련 단계에서 합성 품질과 강한 상관관계를 보입니다(Pearson's $r=-0.869$ with $ ext{FID의 로그 값)), 이는 모델 성능을 나타내는 지표로서 잠재력을 가지며, 모델 최적화를 위한 원칙적인 가이드 역할을 할 수 있음을 시사합니다. 이러한 핵심 발견을 바탕으로, 우리는 표현 다변성을 명시적으로 촉진하는 새로운 프레임워크인 DiverseDiT++를 제안합니다. 구체적으로, 우리의 방법은 블록 간의 입력 표현을 다양화하기 위해 긴 잔차 연결을 사용하고, 블록들이 뚜렷한 특징을 학습하도록 유도하기 위한 표현 다변성 손실을 도입합니다. ImageNet $256 imes256$ 및 $512 imes512$ 데이터셋에 대한 광범위한 실험 결과는 DiverseDiT++가 다양한 크기의 백본 모델에 적용될 때 일관된 성능 향상과 수렴 속도 가속화를 가져온다는 것을 보여줍니다.

Original Abstract

Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. To this end, this paper first presents a systematic analysis of the representation dynamics of DiTs via quantifying the diversity of block-wise representations. Specifically, we introduce a novel metric, termed the Weighted Diversity Score (WDS), to measure the representational discrepancies across different blocks. Through extensive investigations on the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a critical factor for effective representation learning in DiTs. More importantly, WDS exhibits a strong correlation with synthesis quality across diverse settings, model scales, and training stages (Pearson's $r=-0.869$ with $\log(\text{FID})$), suggesting its potential as an indicator to reflect model performance and a principled guide for model optimization. Based on this key finding, we propose DiverseDiT++, a novel framework that explicitly promotes diverse representation learning. Concretely, our method incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet $256\times256$ and $512\times512$ demonstrate that our DiverseDiT++ yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes,...

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!