ProductConsistency: SFT 및 강화 학습을 활용한 지시 기반 이미지 편집에서 제품 식별 보존 성능 향상
ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL
최근 지시 기반 이미지 편집 기술의 발전으로 모델은 자연어 지침에 따라 복잡한 시각적 편집 작업을 수행할 수 있게 되었습니다. 그러나 제품 관련 시나리오에서는 제품의 특징, 브랜드 및 텍스트 요소를 유지하는 것이 매우 중요하며, 현재 공개 및 비공개 모델들은 이러한 미세한 객체 식별을 유지하는 데 어려움을 겪는 경우가 많습니다. 이 문제는 또한 지시 기반 제품 이미지 편집에 대한 데이터셋이 부족하다는 점 때문에 더욱 심화됩니다. 기존의 연구에서는 이러한 문제를 지시 기반 이미지 편집 모델의 암묵적인 능력으로 간주하는 경향이 있습니다. 본 논문에서는 제품 중심 이미지 편집 성능을 향상시키기 위해 설계된 ProductConsistency 데이터셋을 소개합니다. 저희의 접근 방식은 다음과 같습니다: 87,000개의 샘플로 구성된 제품 편집을 위한 지도 학습(SFT) 데이터셋, 869개의 고유한 제품 이미지를 포함하는 강화 학습(RL) 데이터셋, 그리고 편집 모델에 대한 엄격하고 표준화된 평가를 가능하게 하는 새로운 벤치마크 데이터셋인 ProductConsistency Benchmark입니다. RL 학습을 안내하기 위해, 저희는 원본 제품 설명과 편집된 이미지에서 생성된 캡션 간의 유사성을 활용하여 제품 식별의 의미적 일관성을 강화하는 Cyclic Consistency 보상 함수를 제안합니다. 저희는 Qwen-Image-Edit-2511 및 Flux.1-Kontext-dev 모델을 저희 데이터셋으로 미세 조정하고, OCR 및 지각적 메트릭뿐만 아니라 MLLM 기반 평가에서도 기준 모델보다 일관된 성능 향상을 보여주었습니다. 이는 제품의 일관성, 텍스트 렌더링 및 전반적인 시각 품질이 향상되었음을 나타냅니다. 특히 Qwen-Image-Edit-2511 모델은 문자 오류율을 5배 감소시켰습니다. 코드 및 파이프라인은 https://anonymous.4open.science/r/ProductConsistency-6FCC/README.md 에서 확인할 수 있습니다.
Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions. However, in product-centric scenarios where preserving product features, branding, and textual elements are critical, current open and closed source models often struggle to maintain this fine-grained object identity. This issue is further compounded by the lack of datasets for instruction-based product image editing with text fidelity constraints, leaving it largely treated as an implicit capability of instruction-based image editing models. In this work, we introduce the ProductConsistency dataset which is designed to improve product-centric image editing. Our approach includes a supervised fine-tuning (SFT) dataset of 87k samples for product editing, a reinforcement learning (RL) dataset with 869 unique product images, and a new benchmark dataset, the ProductConsistency Benchmark, to allow rigorous and standardized evaluation of editing models. To guide RL training, we propose a Cyclic Consistency reward that enforces semantic preservation of product identity by using caption similarity between the original product description and captions generated from the edited image. We fine-tune both Qwen-Image-Edit-2511 and Flux.1-Kontext-dev using our dataset and demonstrate consistent improvements over baseline models in OCR and Perceptual metrics, and MLLM-based evaluations as well, indicating stronger product consistency, text rendering, and overall visual quality; with the Qwen-Image-Edit-2511 model achieving a 5x reduction in the character error rate. The code and pipeline is available at https://anonymous.4open.science/r/ProductConsistency-6FCC/README.md
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.