지시 기반 이미지 편집: 데이터, 모델, 평가 및 응용에 대한 개관
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
지시 기반 이미지 편집(IIE)은 주어진 이미지를 텍스트 지시에 따라 새로운 이미지로 변환하는 것을 목표로 합니다. 대규모 언어 모델(LLM)과 시각-언어 모델(VLM)의 발전으로 인해 실용적인 "한 문장 이미지 편집" 시스템 개발이 가속화되었습니다. 본 논문에서는 IIE 연구에 대한 체계적인 분류와 종합적인 검토를 제시하며, 다음 다섯 가지 핵심 측면을 중심으로 구성됩니다: (1) 편집 작업의 정의 및 계층적 분류, (2) 학습 데이터 구축 방법론, (3) GAN 기반에서 확산 모델 및 자기 회귀 모델로의 아키텍처 진화, (4) 표준화된 평가 지표 및 벤치마크 개발, 그리고 (5) 상용 솔루션 소개. 본 논문의 분석은 다양한 모델 세대를 거치면서 나타난 중요한 기술적 이정표를 보여줍니다. 또한, IIE 작업의 종합적이고 심층적인 성능 평가를 위한 벤치마크(CDD-IIE Bench)를 제안하며, 이를 통해 모델의 다양한 측면을 엄격하게 평가할 수 있습니다. 공개된 솔루션들의 경험적 비교 분석을 통해 각 솔루션의 강점과 한계를 강조합니다. 마지막으로, 본 연구 분야의 발전을 위한 미래 연구 방향에 대해 논의합니다.
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical ``one-sentence image editing" systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of model performance. Through empirical comparisons of open-source solutions, we highlight their respective capabilities and limitations. Finally, we discuss future research directions to advance the field.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.