Edit-R2: 문맥 인지 강화 학습 기반 다중 회차 이미지 편집
Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing
텍스트 기반 이미지 편집은 확산 모델과 통합형 멀티모달 기초 모델의 발전으로 빠르게 진전해 왔습니다. 그러나 대부분의 기존 방법은 단일 회차 환경에 국한되어 있으며, 사용자가 일련의 지침을 통해 이미지를 반복적으로 개선하는 보다 현실적인 다중 회차 환경을 고려하지 못합니다. 이러한 환경에서 모델은 새로운 지침을 따르면서 동시에 누적된 세션 수준의 제약을 유지해야 하며, 이는 '장문맥 희석'과 '상태 오염'이라는 두 가지 문제점에 직면합니다. '장문맥 희석'은 이미지와 텍스트가 번갈아 나타나는 기록이 증가함에 따라 드물게 나타나는 텍스트 기반 제약 조건을 복구하기 어렵게 만들고, '상태 오염'은 이전 편집 과정에서의 오류가 이후 생성 결과에 영향을 미쳐 성능을 저하시킵니다. 본 논문에서는 통합형 멀티모달 모델을 위한 새로운 강화 학습 후속 학습 프레임워크인 Edit-R2를 소개합니다. Edit-R2는 실제 세션 의도를 재구성하여, 각 편집 단계 전에 분산된 과거 제약 조건을 명시적인 추론 경로로 통합합니다. 또한, Edit-R2는 추론과 생성 과정을 모두 아우르는 통합 목표 함수를 통해 강화 학습을 수행하며, 이 목표 함수는 이산 텍스트 공간에서의 의도 재구성 생성과 연속 잠재 공간에서의 흐름 일치 이미지 생성을 동시에 최적화합니다. 또한, 트래jectory 필터링 메커니즘은 '상태 오염'으로 인한 불안정한 학습 문제를 완화하기 위해 잘못된 결과를 걸러냅니다. 체계적인 평가를 지원하기 위해, 본 논문에서는 다중 회차 환경에서의 이미지 편집 성능을 자동으로 측정할 수 있는 MICE-Bench라는 대규모 벤치마크를 소개합니다. 실험 결과, Edit-R2는 다중 회차 환경에서의 이미지 편집 성능을 크게 향상시키며, 기존의 강력한 모델과 경쟁력 있는 성능을 보여줍니다.
Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn settings, overlooking the more realistic scenario of multi-turn in-context editing, where users iteratively refine an image through a sequence of instructions. In this setting, a model must follow each new instruction while preserving accumulated session-level constraints, challenged by two coupled failure modes: long-context dilution, where sparse textual constraints become difficult to recover from growing interleaved image-text histories, and state contamination, where earlier editing mistakes degrade subsequent generations. We introduce Edit-R2, a novel reinforcement learning post-training framework for unified multimodal models. Edit-R2 reconstructs the operative session intent, which effectively consolidates scattered historical constraints into an explicit reasoning trace before each editing turn. It further enables multi-turn RL over both reasoning and generation through a unified objective that jointly optimizes intent reconstruction generation in discrete text space and flow-matching image generation in continuous latent space, while a trajectory filtering mechanism suppresses corrupted rollouts to stabilize training under state contamination. To support systematic evaluation, we introduce MICE-Bench, a large-scale benchmark for multi-turn in-context editing with automated metrics for instruction following (IF), content consistency (CC), and global awareness (GA) over accumulated session constraints. Experiments show that Edit-R2 substantially improves multi-turn in-context editing and achieves competitive performance compared against strong baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.