2608.04436v1 Aug 05, 2026 cs.CV

ToolArtist: 도구 사용 기반의 통합 멀티모달 모델을 활용한 자율적 이미지 생성

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Shuicheng Yan
Shuicheng Yan
Citations: 113
h-index: 6
Jiahao Zhao
Jiahao Zhao
Citations: 424
h-index: 8
ZhongXiang Sun
ZhongXiang Sun
Renmin University of China
Citations: 1,266
h-index: 15
Jun Xu
Jun Xu
Citations: 361
h-index: 9
Xiaomin Yu
Xiaomin Yu
Citations: 21
h-index: 3
Xiaobin Hu
Xiaobin Hu
Citations: 32
h-index: 3
Chengwei Qin
Chengwei Qin
Citations: 16
h-index: 3
Fengwei Teng
Fengwei Teng
Citations: 375
h-index: 5

텍스트-이미지(T2I) 모델은 시각적으로 매력적인 이미지를 생성할 수 있지만, 복잡한 의미 이해, 다단계 추론 및 외부 세계 지식 통합이 필요한 실제 환경에서의 작업에는 한계가 있습니다. 기존 연구에서는 이미지 생성에 에이전트 기능을 도입하려는 노력이 있었지만, 이는 고정된 워크플로우를 따르거나 전체 이미지 생성 과정 중 일부만 에이전트의 제어 하에 두는 경우가 많았습니다. 결과적으로, 추론, 도구 사용 및 이미지 생성이 단일 정책으로 조정되지 않습니다. 본 논문에서는 통합 멀티모달 모델(UMM)을 추가 학습하여 개발된 완전한 자율적 이미지 생성 모델인 ToolArtist를 제안합니다. ToolArtist는 하나의 통합 정책 내에서 추론, 외부 도구 사용 및 기본 이미지 생성을 동적으로 조율합니다. 지도 미세 조정(SFT) 단계에서는 교사 에이전트에게 검색 도구와 이미지 생성 도구를 함께 제공합니다. 수집된 시퀀스를 UMM과 호환되는 형식으로 변환하며, 이 때 이미지 생성 도구는 숨겨지지만, 결과적으로 생성된 이미지는 유지됩니다. 강화 학습(RL) 단계에서는 UMM을 위한 에이전트 기반 RL 인프라를 개발하고, 모델을 공동으로 최적화하기 위해 상호 보완적인 의도 및 품질 보상을 사용하는 RAD-GRPO라는 알고리즘을 소개합니다. 실험 결과는 전체 실제 환경 이미지 생성 과정을 에이전트 정책에 맡기는 것이 고정된 파이프라인이나 부분적으로만 에이전트 제어를 받는 접근 방식보다 일관되게 우수한 성능을 나타냅니다. 본 연구에서는 학습 데이터와 전체 추가 학습 인프라를 공개합니다.

Original Abstract

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescribe a fixed workflow or place only a subset of the open-world image generation process under agent control. Consequently, reasoning, tool invocation, and image generation are not coordinated by a single policy. We propose ToolArtist, a fully agentic image generation model obtained by post-training a Unified Multimodal Model (UMM). ToolArtist dynamically orchestrates reasoning, external tool use, and native image generation within one unified policy. During Supervised Fine-Tuning (SFT), we equip a teacher agent with search tools alongside an image-generation tool. We then convert the collected trajectories into a UMM compatible format, where the image-generation tool is concealed while the resulting generated images are retained. During Reinforcement Learning (RL), we develop an agentic RL infrastructure for UMMs and introduce Reason-Act-Draw GRPO (RAD-GRPO), which uses complementary intent and quality rewards to jointly optimize the model. Experiments show that placing the entire open-world image-generation process under an agent policy consistently outperforms approaches with fixed pipelines or only partially agent-controlled components. We release the training data and the complete post-training infrastructure.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!