2604.13540v1 Apr 15, 2026 cs.CV

통합 다중 모달 모델을 위한 무료 점심: 내재적 이해를 활용한 반사적 교정을 통한 생성 능력 향상

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

Zequn Qin
Zequn Qin
Citations: 13
h-index: 3
Xi Li
Xi Li
Citations: 18
h-index: 3
Tao Wu
Tao Wu
Citations: 9
h-index: 2
Yehao Lu
Yehao Lu
Citations: 195
h-index: 5
Yibo Jiang
Yibo Jiang
Citations: 8
h-index: 2
Rui Jiang
Rui Jiang
Citations: 195
h-index: 6
Chaoxiang Cai
Chaoxiang Cai
Citations: 23
h-index: 3

통합 다중 모달 모델(UMM)은 시각적 이해와 생성을 단일 구조 내에서 통합하는 것을 목표로 합니다. 그러나 이러한 모델은 상당한 능력 불균형을 보이며, 특히 이해 능력은 생성 능력보다 훨씬 뛰어납니다. 이러한 불균형은 모델의 풍부한 내부 지식이 이해 작업에는 효과적이지만 생성 과정에서는 충분히 활용되지 않고 있다는 것을 의미합니다. 이러한 문제를 해결하기 위해, 우리는 인간의 '생각하면서 그리기' 방식에서 영감을 얻어, 인간이 지속적으로 자신의 지식을 활성화하고 중간 결과를 교정하는 방식을 모방했습니다. 본 논문에서는 학습이 필요 없는 통합 교정 체인 오브 씽크(CoT) 프레임워크인 UniRect-CoT를 제안합니다. 우리의 접근 방식은 UMM의 강력한 내재적 이해 능력에 숨겨진 '무료 점심'을 활용하여, 지속적으로 내부 지식을 활성화하고 생성 과정에서 중간 결과를 교정합니다. 우리는 UMM의 확산 노이즈 제거 과정을 내재적인 시각적 추론 과정으로 간주하고, 모델이 이해하는 대상 지침과 중간 결과를 일치시켜 UMM의 생성 과정을 자체 감독 방식으로 교정하는 신호를 제공합니다. 광범위한 실험 결과, UniRect-CoT는 기존 UMM에 쉽게 통합될 수 있으며, 다양한 복잡한 작업에서 생성 품질을 크게 향상시킬 수 있음을 보여줍니다.

Original Abstract

Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability significantly outperforms their generation. This mismatch indicates that the model's rich internal knowledge, while effective for understanding tasks, remains underactivated during generation. To address this, we draw inspiration from the human ``Thinking-While-Drawing'' paradigm, where humans continuously reflect to activate their knowledge and rectify intermediate results. In this paper, we propose UniRect-CoT, a training-free unified rectification chain-of-thought framework. Our approach unlocks the ``free lunch'' hidden in the UMM's powerful inherent understanding to continuously reflect, activating its internal knowledge and rectifying intermediate results during generation.We regard the diffusion denoising process in UMMs as an intrinsic visual reasoning process and align the intermediate results with the target instruction understood by the model, serving as a self-supervisory signal to rectify UMM generation.Extensive experiments demonstrate that UniRect-CoT can be easily integrated into existing UMMs, significantly enhancing generation quality across diverse complex tasks.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!