2606.23221v1 Jun 22, 2026 cs.CV

RS-Gen: 추론 및 검색 기반 이미지 생성을 위한 다단계 에이전트 프레임워크

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

Jian Luan
Jian Luan
Citations: 220
h-index: 7
Daiguo Zhou
Daiguo Zhou
Citations: 44
h-index: 4
Weiwei Deng
Weiwei Deng
Citations: 1,291
h-index: 17
Feifei Bian
Feifei Bian
Citations: 72
h-index: 5
Zhi Zheng
Zhi Zheng
University of Chinese Academy of Sciences
Citations: 1,254
h-index: 7

최근 몇 년 동안 이미지 생성 및 편집 분야에서, 특히 지시 사항 준수 및 시각적 충실도 측면에서 놀라운 발전이 있었습니다. 그러나 모호한 의도 처리, 논리적 추론 및 외부(OOD) 지식에 대한 작업에서는 기존의 이미지 모델이 심층적인 추론 능력과 실시간 외부 정보 부족으로 인해 최적이 아닌 결과를 초래하는 경우가 많습니다. 통합된 이해-생성 모델이 이러한 격차를 해소하려고 시도하지만, 여전히 고유한 파라미터 규모와 정적인 지식 격자에 의해 제약됩니다. 에이전트 패러다임에서 영감을 받아, 우리는 RS-Gen을 제안합니다. RS-Gen은 플러그 앤 플레이 방식으로 작동하며 학습 과정 없이 다단계 이미지 에이전트 프레임워크를 제공합니다. RS-Gen은 혁신적으로 "질문 및 해결"이라는 폐쇄 루프 메커니즘을 도입하여 논리적 문제와 지식 격차를 정확하게 식별하고, 정보 부족을 해소하기 위한 작업을 자율적으로 계획하고 심층적인 논리적 추론을 수행합니다. 광범위한 실험 결과, RS-Gen은 기존의 이미지 생성 및 편집 모델의 성능 한계를 크게 확장한다는 것을 보여줍니다. 특히 WISE Verified 및 RISEBench 벤치마크에서 RS-Gen은 각각 Qwen-Image에서 0.313, Qwen-Image-Edit-2511에서 19.70이라는 상당한 성능 향상을 가져왔으며, 이를 통해 두 모델 모두 오픈 소스 모델 중에서 최첨단(SOTA) 수준에 도달했습니다.

Original Abstract

Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous intentions, logical reasoning, and Out-of-Distribution (OOD) knowledge, existing image models often yield sub-optimal results due to a lack of deep reasoning capabilities and real-time external information. Although emerging unified understanding-and-generation models attempt to bridge this gap, they remain constrained by their intrinsic parameter scales and static knowledge gaps. Inspired by agentic paradigms, we propose RS-Gen: a plug-and-play, training-free, multi-stage image agentic framework. RS-Gen innovatively introduces a "Questioning-and-Solving" closed-loop mechanism to accurately identify logical issues and knowledge gaps, autonomously planning actions to bridge information deficits and execute deep logical reasoning. Extensive experiments demonstrate that RS-Gen significantly expands the capability boundaries of foundational image generation and editing models. Specifically, on the WISE Verified and RISEBench benchmarks, RS-Gen yields substantial absolute performance gains of 0.313 for Qwen-Image and 19.70 for Qwen-Image-Edit-2511, respectively, successfully elevating both to the state-of-the-art (SOTA) level among open-source models.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!