Mage-Flow: 효율적인 고해상도 기반 모델을 활용한 이미지 생성 및 편집
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
대규모 시각적 생성 모델은 점점 더 강력해지고 있지만, 학습, 미세 조정 및 배포에 많은 비용이 소요됩니다. 본 논문에서는 효율적인 텍스트-이미지 생성 및 지시 기반 이미지 편집을 위한 작고 가벼운 40억 파라미터 규모의 생성 스택인 Mage-Flow를 소개합니다. 이 스택은 두 가지 상호 설계된 구성 요소로 구성됩니다. 첫째, 고품질 잠재 토큰화를 수행하는 경량 모델인 Mage-VAE입니다. 둘째, 수정된 플로우 매칭을 통해 학습된 네이티브 해상도 다중 모드 디퓨전 트랜스포머입니다. Mage-VAE는 앵커 잠재 규제와 함께 단일 단계의 디퓨전 스타일 인코딩 및 디코딩을 사용하여 강력한 공개 VAE의 재구성 품질을 유지하면서 토큰화 비용을 수십 배 줄입니다. 또한 네이티브 해상도 패킹과 스택 수준의 CUDA 커널 융합을 통해 유연한 해상도의 학습을 지원하고 엔드 투 엔드 학습 처리 속도를 약 2.5배 향상시킵니다. 이러한 기반 위에, 생성 및 편집 모두에 대해 Base 모델, RL(강화 학습) 정렬 모델, 그리고 Turbo 모델 등 다양한 모델 패밀리를 개발했습니다. Diffusion-NFT는 프롬프트 준수성, 텍스트 렌더링, 미적 품질 및 편집 정확도를 향상시키고, 적은 단계의 증류와 적대적인 지각 가이드를 통해 4단계 Turbo 모델을 생성하여 낮은 지연 시간으로 추론할 수 있도록 합니다. Mage-Flow와 Mage-Flow-Edit는 작지만 강력한 규모에도 불구하고 표준 생성 및 편집 벤치마크에서 경쟁력 있는 성능을 보입니다. 더욱 중요한 점은, Turbo 모델 덕분에 고해상도 이미지 생성 및 편집이 실시간 사용에 적합하게 되었습니다. 단일 NVIDIA A100 GPU에서 1024x1024 해상도로 Mage-Flow-Turbo는 이미지를 0.59초 만에 생성하고, Mage-Flow-Edit-Turbo는 이미지를 1.02초 만에 편집하며, 작은 메모리 공간을 차지합니다. 이러한 결과는 토큰화, 백본 및 시스템 간의 신중한 공동 설계를 통해 효율적인 40억 파라미터 규모의 모델 패밀리 내에서 강력한 고해상도 생성 및 편집이 가능하다는 것을 보여줍니다.
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about $2.5\times$. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at $1024^2$ resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.