일관성 있는 다중 참조 이미지 편집을 위한 평가-검증 보상 기법
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
최근 이미지 편집 모델은 빠른 발전을 이루었지만, 특히 여러 참조 이미지를 사용하는 경우 시각적 일관성을 유지하고 전체적인 조화를 확보하는 것은 여전히 어려운 과제입니다. 강화 학습은 텍스트-이미지 생성 및 단일 이미지 편집에 효과적임이 입증되었지만, 다중 참조 편집으로 확장할 때에는 여러 이미지 간의 관계 제약을 포착할 수 있는 적절한 보상 모델이 부족하여 어려움이 있습니다. 또한, 멀티모달 대규모 언어 모델(MLLM)을 초기 상태 평가기로 사용하는 경우, 환각 가능성이 높은 장문 추론과 짧은 형식 판단의 제한적인 추론 능력 사이의 중요한 긴장이 존재합니다. 우리는 이러한 문제를 해결하기 위해 다차원 평가-검증 보상 기법(EVR)을 제안합니다. EVR은 평가를 명확한 시각적 기준으로 분해하고, 각 기준에 대해 MLLM 평가기가 여러 후보 가설을 생성하며, 검증기는 각 주장을 구체적인 시각적 증거를 바탕으로 수용하거나 거부하여 신뢰할 수 있고 세분화된 보상 신호를 생성합니다. 또한 확장 가능한 데이터 파이프라인과 함께, 우리의 방법은 아키텍처 변경 없이 기존 편집기를 강화 학습 방식으로 미세 조정할 수 있도록 합니다. 광범위한 실험 결과는 Qwen-Image-Edit의 기본 모델보다 상당한 성능 향상을 보여주었으며, 일관성과 조화성을 NanoBanana 수준 또는 그 이상으로 개선했습니다.
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.