2608.02059v1 Aug 03, 2026 cs.CV

MIEScore: 인간 선호도에 기반한 다중 소스 이미지 편집 평가

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Zhengxue Cheng
Zhengxue Cheng
Citations: 263
h-index: 9
Xiongkuo Min
Xiongkuo Min
Citations: 12,533
h-index: 51
Guangtao Zhai
Guangtao Zhai
Citations: 298
h-index: 9
Tianyi Zheng
Tianyi Zheng
Citations: 10
h-index: 2
Huiyu Duan
Huiyu Duan
Citations: 2,247
h-index: 25
Xinyu Zhang
Xinyu Zhang
Citations: 16
h-index: 2
Bo Li
Bo Li
Citations: 427
h-index: 10
Qiang Hu
Qiang Hu
Citations: 4
h-index: 1
Weifei Xiong
Weifei Xiong
Citations: 0
h-index: 0
Zitong Xu
Zitong Xu
Citations: 94
h-index: 4

최근 통합 멀티모달 모델의 발전은 텍스트 기반 이미지 편집 능력을 크게 향상시켰습니다. 특히 Nano-Banana-Pro 및 GPT-Image-2와 같은 모델들은 객체 합성, 인물-배경 구성, 그리고 이미지 간 스타일 결합과 같은 다중 소스 이미지 편집(MIE) 분야에서 뛰어난 성능을 보여줍니다. 그러나 기존의 벤치마크 및 이미지 편집 평가(IEQA) 방법은 주로 단일 이미지 편집 작업에 초점을 맞추고 있으며, 보다 어려운 MIE 환경은 대부분 간과되고 있습니다. 이는 MIE를 위한 포괄적이고 인간 선호도를 반영하는 벤치마크의 필요성을 강조합니다. 이러한 문제를 해결하기 위해, 우리는 미세 조정된 인간 선호도 주석을 포함하는 최초의 대규모 다중 이미지 편집 벤치마크인 MIE-Bench를 소개합니다. 구체적으로, MIE-Bench는 16개의 작업에 걸쳐 3,000개의 편집 인스턴스를 포함하며, 각 인스턴스는 두 개 이상의 소스 이미지와 편집 프롬프트를 포함합니다. 또한, 12개의 최첨단 편집 모델이 생성한 36,000장의 편집된 이미지와 시각적 품질, 지시사항 준수 및 속성 보존을 포괄하는 108,000개 이상의 평균 의견 점수(MOS)를 포함합니다. MIE-Bench를 기반으로, 우리는 기술 최적화 및 다차원 지도 학습을 통해 강화된 멀티모달 대규모 언어 모델(MLLM) 기반 평가 모델인 MIEScore를 제안하여, MIE에 대한 인간 선호도에 부합하는 피드백을 제공합니다. 광범위한 실험 결과는 MIEScore가 인간 선호도와 일치하는 최첨단 성능을 달성하며 다른 IEQA 데이터셋에서도 잘 일반화됨을 보여줍니다. 데이터셋과 모델은 https://github.com/IntMeGroup/MIEScore에서 이용 가능합니다.

Original Abstract

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion. However, existing benchmarks and image editing assessment (IEQA) methods remain primarily focused on single-image editing tasks and largely overlook the more challenging setting of MIE. This highlights the urgent need for a comprehensive and human-aligned benchmark for MIE. To this end, we introduce MIE-Bench, the first large-scale multiple image editing benchmark with fine-grained human preference annotations. Specifically, MIE-Bench includes 3,000 editing instances across 16 tasks, each involving more than two source images and an editing prompt, together with 36K edited images produced by 12 state-of-the-art editing models and over 108K mean opinion scores (MOSs) covering visual quality, instruction following, and attribute preservation. Based on MIE-Bench, we propose MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback for MIE. Extensive experiments show that MIEScore achieves state-of-the-art performance in aligning with human preferences and generalizes well across other IEQA datasets. Both the dataset and the model are available at https://github.com/IntMeGroup/MIEScore.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!