MCMC 수정을 통한 다중 모드 변분 오토인코더를 이용한 다중 모드 에너지 기반 모델 학습
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision
에너지 기반 모델(EBM)은 유연한 심층 생성 모델의 한 종류이며, 다중 모드 데이터의 복잡한 의존성을 포착하는 데 적합합니다. 그러나 최대 우도 추정(MLE)을 통해 다중 모드 EBM을 학습하려면, 공동 데이터 공간에서 마르코프 체인 몬테카를로(MCMC) 샘플링이 필요하며, 이때 노이즈 초기화된 Langevin 동역학은 종종 제대로 작동하지 않아 일관성 있는 모드 간 관계를 발견하지 못합니다. 다중 모드 VAE는 공유 잠재 생성기와 공동 추론 모델을 도입하여 이러한 모드 간 의존성을 포착하는 데 진전을 이루었습니다. 그러나 공유 잠재 생성기와 공동 추론 모델 모두 단일 모드 가우스(또는 Laplace) 분포로 매개변수화되어 있으며, 이는 다중 모드 데이터에 의해 유도되는 복잡한 구조를 근사하는 능력을 심각하게 제한합니다. 본 연구에서는 다중 모드 EBM, 공유 잠재 생성기, 그리고 공동 추론 모델의 학습 문제를 다룹니다. 데이터 공간과 잠재 공간 모두에서 MLE 업데이트와 해당 MCMC 개선을 효과적으로 결합하는 학습 프레임워크를 제시합니다. 구체적으로, 생성기는 EBM 샘플링을 위한 강력한 초기 상태 역할을 하는 일관성 있는 다중 모드 샘플을 생성하도록 학습되며, 추론 모델은 생성기 사후 샘플링을 위한 유용한 잠재 초기화를 제공하도록 학습됩니다. 이 두 모델은 효과적인 EBM 샘플링 및 학습을 가능하게 하는 상호 보완적인 모델로서, 현실적이고 일관성 있는 다중 모드 EBM 샘플을 생성합니다. 광범위한 실험을 통해 다양한 기준 모델과 비교했을 때, 다중 모드 합성 품질 및 일관성 측면에서 우수한 성능을 보임을 입증합니다. 제안된 다중 모드 프레임워크의 효과성과 확장성을 검증하기 위해 다양한 분석 및 ablation 연구를 수행했습니다.
Energy-based models (EBMs) are a flexible class of deep generative models and are well-suited to capture complex dependencies in multimodal data. However, learning multimodal EBM by maximum likelihood requires Markov Chain Monte Carlo (MCMC) sampling in the joint data space, where noise-initialized Langevin dynamics often mixes poorly and fails to discover coherent inter-modal relationships. Multimodal VAEs have made progress in capturing such inter-modal dependencies by introducing a shared latent generator and a joint inference model. However, both the shared latent generator and joint inference model are parameterized as unimodal Gaussian (or Laplace), which severely limits their ability to approximate the complex structure induced by multimodal data. In this work, we study the learning problem of the multimodal EBM, shared latent generator, and joint inference model. We present a learning framework that effectively interweaves their MLE updates with corresponding MCMC refinements in both the data and latent spaces. Specifically, the generator is learned to produce coherent multimodal samples that serve as strong initial states for EBM sampling, while the inference model is learned to provide informative latent initializations for generator posterior sampling. Together, these two models serve as complementary models that enable effective EBM sampling and learning, yielding realistic and coherent multimodal EBM samples. Extensive experiments demonstrate superior performance for multimodal synthesis quality and coherence compared to various baselines. We conduct various analyses and ablation studies to validate the effectiveness and scalability of the proposed multimodal framework.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.