ToolArtist: 도구 사용 기반의 통합 멀티모달 모델을 활용한 자율적 이미지 생성
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
텍스트-이미지(T2I) 모델은 시각적으로 매력적인 이미지를 생성할 수 있지만, 복잡한 의미 이해, 다단계 추론 및 외부 세계 지식 통합이 필요한 실제 환경에서의 작업에는 한계가 있습니다. 기존 연구에서는 이미지 생성에 에이전트 기능을 도입하려는 노력이 있었지만, 이는 고정된 워크플로우를 따르거나 전체 이미지 생성 과정 중 일부만 에이전트의 제어 하에 두는 경우가 많았습니다. 결과적으로, 추론, 도구 사용 및 이미지 생성이 단일 정책으로 조정되지 않습니다. 본 논문에서는 통합 멀티모달 모델(UMM)을 추가 학습하여 개발된 완전한 자율적 이미지 생성 모델인 ToolArtist를 제안합니다. ToolArtist는 하나의 통합 정책 내에서 추론, 외부 도구 사용 및 기본 이미지 생성을 동적으로 조율합니다. 지도 미세 조정(SFT) 단계에서는 교사 에이전트에게 검색 도구와 이미지 생성 도구를 함께 제공합니다. 수집된 시퀀스를 UMM과 호환되는 형식으로 변환하며, 이 때 이미지 생성 도구는 숨겨지지만, 결과적으로 생성된 이미지는 유지됩니다. 강화 학습(RL) 단계에서는 UMM을 위한 에이전트 기반 RL 인프라를 개발하고, 모델을 공동으로 최적화하기 위해 상호 보완적인 의도 및 품질 보상을 사용하는 RAD-GRPO라는 알고리즘을 소개합니다. 실험 결과는 전체 실제 환경 이미지 생성 과정을 에이전트 정책에 맡기는 것이 고정된 파이프라인이나 부분적으로만 에이전트 제어를 받는 접근 방식보다 일관되게 우수한 성능을 나타냅니다. 본 연구에서는 학습 데이터와 전체 추가 학습 인프라를 공개합니다.
Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescribe a fixed workflow or place only a subset of the open-world image generation process under agent control. Consequently, reasoning, tool invocation, and image generation are not coordinated by a single policy. We propose ToolArtist, a fully agentic image generation model obtained by post-training a Unified Multimodal Model (UMM). ToolArtist dynamically orchestrates reasoning, external tool use, and native image generation within one unified policy. During Supervised Fine-Tuning (SFT), we equip a teacher agent with search tools alongside an image-generation tool. We then convert the collected trajectories into a UMM compatible format, where the image-generation tool is concealed while the resulting generated images are retained. During Reinforcement Learning (RL), we develop an agentic RL infrastructure for UMMs and introduce Reason-Act-Draw GRPO (RAD-GRPO), which uses complementary intent and quality rewards to jointly optimize the model. Experiments show that placing the entire open-world image-generation process under an agent policy consistently outperforms approaches with fixed pipelines or only partially agent-controlled components. We release the training data and the complete post-training infrastructure.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.