Creo: 단일 이미지 생성에서 점진적이고 협력적인 아이디어 구상까지
Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation
텍스트-이미지(T2I) 시스템은 고품질 이미지를 빠르게 생성할 수 있지만, 시각적 아이디어가 발전하는 방식과는 일치하지 않습니다. T2I 시스템은 사용자를 대신하여 암묵적인 시각적 결정을 내리고, 종종 사용자가 너무 일찍 특정 세부 사항에 얽매이도록 유도하여 다양한 선택지를 탐색할 수 있는 능력을 제한하며, 편집 과정에서 예기치 않은 변경을 발생시켜 수정하기 어렵게 만들고 사용자의 통제력을 감소시킵니다. 이러한 문제점을 해결하기 위해, 우리는 Creo라는 다단계 T2I 시스템을 제안합니다. Creo는 초기 스케치에서 고해상도 결과물로 점진적으로 발전하며, 중간 단계에서 사용자가 단계적으로 변경할 수 있도록 추상화된 표현을 제공합니다. 스케치와 유사한 표현은 사용자의 편집을 유도하고, 아이디어가 형성되는 초기 단계에서 사용자가 다양한 디자인 옵션을 유지할 수 있도록 돕습니다. Creo의 각 단계는 수동 수정과 AI 기반 작업 모두를 지원하며, 잠금 기능을 통해 이전 결정을 유지하여 후속 편집이 특정 영역이나 속성에만 영향을 미치도록 하여 세밀하고 단계적인 제어를 가능하게 합니다. 사용자는 각 단계에서 결정을 내리고 검증하며, 시스템은 전체 이미지를 다시 생성하는 대신 변경 사항(diff)을 적용하여, 이미지 품질이 향상되는 과정에서 발생할 수 있는 오차를 줄입니다. 단일 단계 시스템과의 비교 연구 결과, 참가자들은 Creo를 통해 생성된 결과물에 대해 더 큰 주인의식을 느꼈으며, 이미지 구축 과정에서 자신의 결정을 추적할 수 있었기 때문입니다. 또한, 임베딩 기반 분석 결과, Creo의 결과물은 단일 단계 결과물보다 다양성이 더 높았습니다. 이러한 결과는 다단계 생성, 중간 단계에서의 제어, 그리고 결정 잠금 기능이 생성 시스템의 제어 가능성, 사용자 주도성, 창의성, 그리고 결과물의 다양성을 향상시키는 핵심적인 설계 원칙임을 시사합니다.
Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I systems generate outputs that make implicit visual decisions on behalf of the user, often introduce fine-grained details that can anchor users prematurely and limit their ability to keep options open early on, and cause unintended changes during editing that are difficult to correct and reduce users' sense of control. To address these concerns, we present Creo, a multi-stage T2I system that scaffolds image generation by progressing from rough sketches to high-resolution outputs, exposing intermediary abstractions where users can make incremental changes. Sketch-like abstractions invite user editing and allow users to keep design options open when ideas are still forming due to their provisional nature. Each stage in Creo can be modified with manual changes and AI-assisted operations, enabling fine-grained, step-wise control through a locking mechanism that preserves prior decisions so subsequent edits affect only specified regions or attributes. Users remain in the loop, making and verifying decisions across stages, while the system applies diffs instead of regenerating full images, reducing drift as fidelity increases. A comparative study with a one-shot baseline shows that participants felt stronger ownership over Creo outputs, as they were able to trace their decisions in building up the image. Furthermore, embedding-based analysis indicates that Creo outputs are less homogeneous than one-shot results. These findings suggest that multi-stage generation, combined with intermediate control and decision locking, is a key design principle for improving controllability, user agency, creativity, and output diversity in generative systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.