2607.19064v1 Jul 21, 2026 cs.CV

Mage-Flow: 효율적인 고해상도 기반 모델을 활용한 이미지 생성 및 편집

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Xiaoyi Zhang
Xiaoyi Zhang
Citations: 136
h-index: 6
Dongnan Gui
Dongnan Gui
Citations: 195
h-index: 4
Zhening Liu
Zhening Liu
Citations: 443
h-index: 10
Zongyu Guo
Zongyu Guo
Citations: 68
h-index: 4
Jiahao Li
Jiahao Li
Citations: 2,495
h-index: 16
Bin Li
Bin Li
Citations: 1,029
h-index: 13
Zimo Wen
Zimo Wen
Citations: 5
h-index: 1
Yifei Shen
Yifei Shen
Citations: 27
h-index: 3
Fanyi Pu
Fanyi Pu
Nanyang Technological University
Citations: 958
h-index: 6
Senqiao Yang
Senqiao Yang
Citations: 722
h-index: 9
Peng Zhang
Peng Zhang
Citations: 22
h-index: 3
Xinjie Zhang
Xinjie Zhang
Citations: 283
h-index: 9
Shicheng Zheng
Shicheng Zheng
Citations: 67
h-index: 3
J. Guo
J. Guo
Citations: 0
h-index: 0
Zhaoyang Jia
Zhaoyang Jia
Citations: 328
h-index: 8
Xun Guo
Xun Guo
Citations: 92
h-index: 3
Yuxuan Luo
Yuxuan Luo
Peking University
Citations: 14
h-index: 2
Wenxuan Xie
Wenxuan Xie
Citations: 140
h-index: 5
Kaichen Zhang
Kaichen Zhang
Citations: 98
h-index: 3
Tianci Bi
Tianci Bi
Citations: 18
h-index: 2
Zihan Zheng
Zihan Zheng
Citations: 19
h-index: 3
Xiao Li
Xiao Li
Citations: 367
h-index: 11
Jinglu Wang
Jinglu Wang
Citations: 1,515
h-index: 20
Yan Lu
Yan Lu
Citations: 2,157
h-index: 16

대규모 시각적 생성 모델은 점점 더 강력해지고 있지만, 학습, 미세 조정 및 배포에 많은 비용이 소요됩니다. 본 논문에서는 효율적인 텍스트-이미지 생성 및 지시 기반 이미지 편집을 위한 작고 가벼운 40억 파라미터 규모의 생성 스택인 Mage-Flow를 소개합니다. 이 스택은 두 가지 상호 설계된 구성 요소로 구성됩니다. 첫째, 고품질 잠재 토큰화를 수행하는 경량 모델인 Mage-VAE입니다. 둘째, 수정된 플로우 매칭을 통해 학습된 네이티브 해상도 다중 모드 디퓨전 트랜스포머입니다. Mage-VAE는 앵커 잠재 규제와 함께 단일 단계의 디퓨전 스타일 인코딩 및 디코딩을 사용하여 강력한 공개 VAE의 재구성 품질을 유지하면서 토큰화 비용을 수십 배 줄입니다. 또한 네이티브 해상도 패킹과 스택 수준의 CUDA 커널 융합을 통해 유연한 해상도의 학습을 지원하고 엔드 투 엔드 학습 처리 속도를 약 2.5배 향상시킵니다. 이러한 기반 위에, 생성 및 편집 모두에 대해 Base 모델, RL(강화 학습) 정렬 모델, 그리고 Turbo 모델 등 다양한 모델 패밀리를 개발했습니다. Diffusion-NFT는 프롬프트 준수성, 텍스트 렌더링, 미적 품질 및 편집 정확도를 향상시키고, 적은 단계의 증류와 적대적인 지각 가이드를 통해 4단계 Turbo 모델을 생성하여 낮은 지연 시간으로 추론할 수 있도록 합니다. Mage-Flow와 Mage-Flow-Edit는 작지만 강력한 규모에도 불구하고 표준 생성 및 편집 벤치마크에서 경쟁력 있는 성능을 보입니다. 더욱 중요한 점은, Turbo 모델 덕분에 고해상도 이미지 생성 및 편집이 실시간 사용에 적합하게 되었습니다. 단일 NVIDIA A100 GPU에서 1024x1024 해상도로 Mage-Flow-Turbo는 이미지를 0.59초 만에 생성하고, Mage-Flow-Edit-Turbo는 이미지를 1.02초 만에 편집하며, 작은 메모리 공간을 차지합니다. 이러한 결과는 토큰화, 백본 및 시스템 간의 신중한 공동 설계를 통해 효율적인 40억 파라미터 규모의 모델 패밀리 내에서 강력한 고해상도 생성 및 편집이 가능하다는 것을 보여줍니다.

Original Abstract

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about $2.5\times$. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at $1024^2$ resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.

1 Citations
0 Influential
10 Altmetric
51.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!