2606.09169v1 Jun 08, 2026 cs.AI

IMUG-Bench: 인터리브된 이해 및 생성에 대한 통합 다중 모드 모델의 성능 평가

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

Weitong Lian
Weitong Lian
Citations: 2
h-index: 1
Zecong Tang
Zecong Tang
Citations: 10
h-index: 2
L. Meng
L. Meng
Citations: 3
h-index: 1
Tengju Ru
Tengju Ru
Citations: 2
h-index: 1
Zhejun Cui
Zhejun Cui
Citations: 2
h-index: 1
Qi Kang
Qi Kang
Citations: 2
h-index: 1
Yu Zhang
Yu Zhang
Citations: 2
h-index: 1
Kaixuan Wang
Kaixuan Wang
Citations: 36
h-index: 3
Yechi Liu
Yechi Liu
Citations: 3
h-index: 1
Haoran Li
Haoran Li
Citations: 7
h-index: 2
Hang Cao
Hang Cao
Citations: 30
h-index: 3
Yichen Zhu
Yichen Zhu
Citations: 407
h-index: 3
Yutao Yuan
Yutao Yuan
Citations: 0
h-index: 0
Chunwei Wang
Chunwei Wang
Citations: 236
h-index: 9
Bo Dai
Bo Dai
Citations: 148
h-index: 6

최근 몇 년 동안, 단일 프레임워크 내에서 이해와 생성을 모두 지원하는 통합 다중 모드 모델(Unified Multimodal Models, UMMs)이 등장했습니다. 역동적이고 다단계 대화형 이미지-텍스트 상호 작용은 실제 응용 분야에서 UMM의 중요한 과제입니다. 그러나 기존 벤치마크는 이 중요한 과제를 평가하는 데 부족하며, 종종 단일 단계 또는 정적인 설정에 제한되거나 다단계 상호 작용에서의 노출 편향(exposure bias)을 간과합니다. 이러한 격차를 해소하기 위해, 우리는 UMM의 이해 및 생성 능력을 종합적으로 평가하는 다단계 대화형 이미지-텍스트 벤치마크인 IMUG-Bench를 제안합니다. IMUG-Bench는 정적 공간(Static Spatial), 시간적 인과 관계(Temporal Causal), 하이브리드(Hybrid)의 세 가지 클래스로 구성되어 있으며, 총 3,113개의 샘플과 12,034번의 상호 작용을 포함합니다. 또한, 동적인 이해 질문을 포함하여 실제 다단계 상호 작용 시나리오를 더 잘 반영하는 평가를 지원합니다. IMUG-Bench에 대한 대규모 실험을 통해 주류 오픈 소스 및 클로즈드 소스 UMM의 성능 한계와 실패 모드를 체계적으로 평가하고, 다단계 상호 작용에서 생성 측면에서의 두드러진 노출 편향을 발견했습니다. 또한, Chain-of-Thought, Self-Verification, Best-of-N Sampling과 같은 다양한 테스트 시간 스케일링 전략을 탐색하여 생성 정확도를 향상시키고 생성 작업에서의 노출 편향을 완화하는 효과를 확인했습니다. 이러한 결과는 미래의 UMM의 견고성과 다단계 상호 작용 능력을 향상시키는 데 도움이 되는 통찰력을 제공합니다.

Original Abstract

In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework. Mastering dynamic, multi-turn interleaved image-text dialogues is a crucial task for UMMs in real-world applications. However, existing benchmarks fail to evaluate this important task, as they are often limited to single-turn or static settings, and typically overlook exposure bias in multi-turn interactions. To bridge this gap, we propose IMUG-Bench, a comprehensive benchmark for multi-turn interleaved image-text dialogue of UMMs that jointly evaluates their understanding and generation capabilities. Our IMUG-Bench comprises three classes: Static Spatial, Temporal Causal, and Hybrid, covering 3,113 samples and 12,034 interaction turns. It also includes dynamic understanding questions, thereby supporting evaluation that better reflects real-world multi-turn interaction scenarios. Large-scale experiments on IMUG-Bench systematically evaluate mainstream open-source and closed-source UMMs, revealing their capability boundaries and failure modes, and uncovering pronounced exposure bias on the generation side in multi-turn interactions. We further explore several test-time scaling strategies, including Chain-of-Thought, Self-Verification, and Best-of-N Sampling, which effectively improve generation accuracy and mitigate exposure bias in generation tasks. These findings provide insights into enhancing the robustness and multi-turn interaction capability of future UMMs.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!