MEPA: 시각적 오토리그레시브 모델링을 위한 다중 스케일 표현 정렬 방법 - Mixture of Experts 활용
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
시각적 오토리그레시브 모델링(VAR)은 거칠기에서 세밀함으로 이어지는 다중 스케일 오토리그레시브 생성 패러다임을 선도하며, 이미지 생성 분야에서 강력한 성능을 보여왔습니다. 그러나 VAR은 여전히 다중 스케일 표현 학습에 내재된 문제점을 가지고 있습니다. 구체적으로, 낮은 스케일은 주로 전반적인 의미(global semantics)를 포착하는 반면, 높은 스케일은 세부적인 디테일에 집중합니다. 이러한 스케일 간의 차이에도 불구하고 공유된 아키텍처를 사용하는 것은 최적화 충돌을 야기할 수 있습니다. 또한, 인과 관계에 기반한 오토리그레시브 프로세스 때문에 초기 스케일에서의 부정확한 의미 정보는 전파되어 최종 결과물의 품질을 크게 저하시킬 수 있습니다. 이러한 문제점을 해결하기 위해, 우리는 각 스케일에 적합한 전문가를 선택할 수 있도록 하는 스케일 인식 토큰 기반 Mixture of Experts (MoE) 아키텍처를 제안합니다. 이를 통해 스케일 간의 독립적인 표현 학습이 가능해집니다. 또한, 초기 스케일에서의 의미 모델링을 강화하기 위해 외부에서 얻은 자기 지도(self-supervised) 특징을 통합합니다. 기존의 단순한 정렬 방식과는 달리, 우리는 VAR 패러다임에 특화된 잔차 특징 집계 방식을 분석하고 설계했습니다. 광범위한 실험 결과, 제안하는 방법이 학습 효율성과 생성 품질 모두에서 상당한 개선을 가져옴을 확인했습니다. ImageNet 256*256 데이터셋에 대한 평가에서는, 우리의 모델이 기존 방식보다 우수한 FID 값을 달성했으며, 이는 기본 설정 대비 절반의 학습 에포크와 더 작은 파라미터 수로 얻어진 결과입니다. 또한, 학습 에포크가 증가함에 따라 성능 차이는 더욱 벌어집니다.
Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization conflicts. Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. In addition, we enhance semantic modeling at early scales by incorporating external self-supervised features. Unlike naive alignment, we analyse and design a residual feature aggregation scheme tailored to the VAR paradigm. Extensive experiments show that our method significantly improves both training efficiency and generation quality. On the ImageNet 256*256 benchmark, our model achieves a superior FID compared to the dense baseline while requiring only half of the default training epochs and a smaller parameter budget, with a merely marginal increase in training cost. Moreover, the performance gap further widens with larger training epochs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.