다중 모드 마스크 확산 모델의 생성 순서 강화
Reinforcing the Generation Order of Multimodal Masked Diffusion Models
최근 디퓨전 언어 모델(DLM)은 자연어 생성 작업에서 상당한 발전을 이루었습니다. 최근 연구에 따르면, 적응적인 토큰 생성 순서 제어를 통해 수학적 추론 및 코드 합성 애플리케이션에서 성능을 크게 향상시킬 수 있습니다. 본 연구에서는 텍스트-이미지 합성 및 다중 모드 이해를 위한 생성 순서 최적화 문제를 다룹니다. 먼저, 스도쿠 퍼즐과 같은 언어 생성의 구조화된 문제와 달리, 텍스트-이미지 생성 및 다중 모드 이해에서 모델 로짓만으로는 최적의 생성 시퀀스를 결정하기에 충분하지 않음을 확인합니다. 이 과제를 해결하기 위해, 그룹 상대 정책 최적화(GRPO)를 통해 학습되는 제어 모듈을 도입하여 생성 순서를 결정합니다. 실험 결과는 이러한 제어 블록 학습이 DLM에서 텍스트-이미지 정렬 및 다중 모드 이해 능력을 크게 향상시킨다는 것을 보여줍니다. 특히, 모델은 생성된 이미지 내의 미세한 공간적 관계를 더 잘 파악하는 동시에 다중 모드 추론 및 이해 작업에서의 성능을 강화합니다. 제안하는 프레임워크는 텍스트-이미지 정렬을 위한 객체 중심 벤치마크인 GenEval에서 4.08%의 상대적인 성능 향상을 달성했습니다. 또한, VLMEvalKit 실험 결과는 다중 모드 이해 능력에서 4.85%의 상대적인 성능 향상을 확인하여, 제안하는 접근 방식의 광범위한 효과성을 강조합니다.
Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathematical reasoning and code synthesis applications. In this work, we investigate the optimization of generation order for both text-to-image synthesis and multimodal understanding. We first establish that, unlike structured problems in language generation such as Sudoku puzzles, model logits alone are insufficient for determining optimal generation sequences in text-to-image generation and multimodal understanding. To address this challenge, we introduce a learnable control module trained via Group Relative Policy Optimization (GRPO) to determine the generation order. Our results demonstrate that learning this control block substantially improves both text-to-image alignment and multimodal understanding in DLMs. In particular, it enhances the model's ability to capture fine-grained spatial relationships in generated images while also strengthening performance on multimodal reasoning and comprehension tasks. We evaluate our framework on GenEval, an object-focused benchmark for text-to-image alignment, where it achieves 4.08% relative improvements. In addition, experiments on VLMEvalKit confirm 4.85% relative improvements in multimodal understanding, highlighting the broad effectiveness of our approach.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.