IMUG-Bench: 인터리브된 이해 및 생성에 대한 통합 다중 모드 모델의 성능 평가
IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation
최근 몇 년 동안, 단일 프레임워크 내에서 이해와 생성을 모두 지원하는 통합 다중 모드 모델(Unified Multimodal Models, UMMs)이 등장했습니다. 역동적이고 다단계 대화형 이미지-텍스트 상호 작용은 실제 응용 분야에서 UMM의 중요한 과제입니다. 그러나 기존 벤치마크는 이 중요한 과제를 평가하는 데 부족하며, 종종 단일 단계 또는 정적인 설정에 제한되거나 다단계 상호 작용에서의 노출 편향(exposure bias)을 간과합니다. 이러한 격차를 해소하기 위해, 우리는 UMM의 이해 및 생성 능력을 종합적으로 평가하는 다단계 대화형 이미지-텍스트 벤치마크인 IMUG-Bench를 제안합니다. IMUG-Bench는 정적 공간(Static Spatial), 시간적 인과 관계(Temporal Causal), 하이브리드(Hybrid)의 세 가지 클래스로 구성되어 있으며, 총 3,113개의 샘플과 12,034번의 상호 작용을 포함합니다. 또한, 동적인 이해 질문을 포함하여 실제 다단계 상호 작용 시나리오를 더 잘 반영하는 평가를 지원합니다. IMUG-Bench에 대한 대규모 실험을 통해 주류 오픈 소스 및 클로즈드 소스 UMM의 성능 한계와 실패 모드를 체계적으로 평가하고, 다단계 상호 작용에서 생성 측면에서의 두드러진 노출 편향을 발견했습니다. 또한, Chain-of-Thought, Self-Verification, Best-of-N Sampling과 같은 다양한 테스트 시간 스케일링 전략을 탐색하여 생성 정확도를 향상시키고 생성 작업에서의 노출 편향을 완화하는 효과를 확인했습니다. 이러한 결과는 미래의 UMM의 견고성과 다단계 상호 작용 능력을 향상시키는 데 도움이 되는 통찰력을 제공합니다.
In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework. Mastering dynamic, multi-turn interleaved image-text dialogues is a crucial task for UMMs in real-world applications. However, existing benchmarks fail to evaluate this important task, as they are often limited to single-turn or static settings, and typically overlook exposure bias in multi-turn interactions. To bridge this gap, we propose IMUG-Bench, a comprehensive benchmark for multi-turn interleaved image-text dialogue of UMMs that jointly evaluates their understanding and generation capabilities. Our IMUG-Bench comprises three classes: Static Spatial, Temporal Causal, and Hybrid, covering 3,113 samples and 12,034 interaction turns. It also includes dynamic understanding questions, thereby supporting evaluation that better reflects real-world multi-turn interaction scenarios. Large-scale experiments on IMUG-Bench systematically evaluate mainstream open-source and closed-source UMMs, revealing their capability boundaries and failure modes, and uncovering pronounced exposure bias on the generation side in multi-turn interactions. We further explore several test-time scaling strategies, including Chain-of-Thought, Self-Verification, and Best-of-N Sampling, which effectively improve generation accuracy and mitigate exposure bias in generation tasks. These findings provide insights into enhancing the robustness and multi-turn interaction capability of future UMMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.