프롬프트에서 프로덕션까지: 텍스트-이미지 변환 모델을 활용한 브랜드 안전 마케팅 이미지 자동화
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
텍스트-이미지 변환 모델은 텍스트 설명으로부터 이미지를 생성하는 데 있어 인상적인 결과를 내며 괄목할 만한 발전을 이루었다. 그러나 이러한 모델을 실제 프로덕션 환경에 배포하기 위한 확장 가능한 파이프라인을 구축하는 것은 여전히 과제로 남아 있다. 확장성과 품질을 모두 유지하려면 자동화와 인간 피드백 간의 적절한 균형을 맞추는 것이 중요하다. 자동화를 통해 대용량 처리가 가능하지만, 생성된 이미지가 요구되는 기준을 충족하고 창의적인 비전에 부합하는지 확인하려면 인간의 감독이 여전히 필수적이다. 본 논문은 텍스트-이미지 변환 모델을 활용해 상업용 제품의 마케팅 이미지를 생성하는, 완전 자동화되고 확장 가능한 솔루션을 제공하는 새로운 파이프라인을 제시한다. 제안하는 시스템은 이미지의 품질과 충실도(fidelity)를 유지하면서도, 마케팅 가이드라인을 준수하기 위한 충분한 창의적 변형을 도입한다. 이 과정을 간소화함으로써 효율성과 인간 감독의 매끄러운 조화를 보장하며, 결과적으로 DINOV2를 사용한 마케팅 객체 충실도에서 30.77%, 생성된 결과물에 대한 인간 선호도에서 52.00%의 향상을 달성하였다.
Text-to-image models have made significant strides, producing impressive results in generating images from textual descriptions. However, creating a scalable pipeline for deploying these models in production remains a challenge. Achieving the right balance between automation and human feedback is critical to maintain both scale and quality. While automation can handle large volumes, human oversight is still an essential component to ensure that the generated images meet the desired standards and are aligned with the creative vision. This paper presents a new pipeline that offers a fully automated, scalable solution for generating marketing images of commercial products using text-to-image models. The proposed system maintains the quality and fidelity of images, while also introducing sufficient creative variation to adhere to marketing guidelines. By streamlining this process, we ensure a seamless blend of efficiency and human oversight, achieving a $30.77\%$ increase in marketing object fidelity using DINOV2 and a $52.00\%$ increase in human preference over the generated outcome.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.