2607.01709v1 Jul 02, 2026 cs.AI

COMFYCLAW: 스스로 발전하는 기술 활용 시스템 - 이미지 생성 워크플로우를 위한

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

Xiao-Ming Wu
Xiao-Ming Wu
Citations: 171
h-index: 8
Jingjing Xie
Jingjing Xie
Citations: 1,036
h-index: 5
Zongxia Li
Zongxia Li
Citations: 912
h-index: 7
Xiyang Wu
Xiyang Wu
Citations: 322
h-index: 8
Lichao Sun
Lichao Sun
Citations: 75
h-index: 5
Dawei Liu
Dawei Liu
Citations: 43
h-index: 2
Fuxiao Liu
Fuxiao Liu
Citations: 2,745
h-index: 19
Jingxi Chen
Jingxi Chen
Citations: 70
h-index: 6
Yuhang Zhou
Yuhang Zhou
Citations: 14
h-index: 2

에이전트는 점점 더 많은 분야에서 워크플로우를 구축하고 인간이 반복적인 작업을 더욱 효율적으로 수행하도록 지원하는 데 사용됩니다. 이러한 워크플로우가 반복되고 특정 도메인에 특화됨에 따라, 에이전트의 기억 기능과 재사용 가능한 기술은 매우 중요해집니다. 에이전트는 이전 실행 과정에서 워크플로우 패턴, 실행 제약 조건 및 사용자 선호도를 기억할 수 있어야 합니다. 본 연구는 워크플로우 기반 이미지 생성 분야에서 이러한 문제를 다루며, ComfyUI 워크플로우를 제어하기 위한 에이전트 기술 발전 시스템인 COMFYCLAW를 소개합니다. COMFYCLAW는 워크플로우 구축을 타입화된 그래프 편집으로 정의하고, 구축 단계별로 구성된 도구를 제공하며, 잘못된 편집 사항은 자동으로 되돌립니다. 또한, 시각적 오류를 실행 가능한 수정 제안으로 변환하기 위해 영역 수준의 비전-언어 모델(VLM) 검증기를 사용합니다. 이 프레임워크는 이전 실행 과정에서 얻은 데이터(경로, 실행 오류 및 검증기 피드백)를 재사용 가능한 에이전트 기술 라이브러리로 발전시킵니다. 네 가지 벤치마크 세트, 세 가지 에이전트 모델 및 두 개의 이미지 백본을 사용하여 평가한 결과, COMFYCLAW는 모든 여섯 가지 에이전트 구성에서 가장 높은 평균 이미지 생성 평가 점수를 달성했으며, 기술 발전을 사용하지 않은 검증기 기반 시스템보다 우수한 성능을 보였습니다. 인간 평가 결과에서도, 피험자들은 기술 발전 기능을 포함하지 않은 COMFYCLAW의 변종보다 COMFYCLAW를 선호했습니다. 이러한 결과는 기술 발전이 반복적인 시각적 워크플로우 구축에서 에이전트의 신뢰성과 성능을 향상시키는 효과적인 방법임을 시사합니다.

Original Abstract

Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently. As these workflows become repeated and domain-specific, agent memory and reusable skills become increasingly important: agents should be able to recall workflow patterns, execution constraints, and user preferences from previous runs. We study this problem in workflow-based image generation and introduce COMFYCLAW, an agentic skill evolution harness for controlling ComfyUI workflows. COMFYCLAW formulates workflow construction as typed graph editing, exposes tools organized by construction stage, automatically reverts invalid edits, and uses a region-level vision-language model (VLM) verifier to translate visual failures into actionable repair suggestions. The framework further evolves a progressively disclosed skill library, where trajectories, execution errors, and verifier feedback from previous runs are distilled into reusable Agent Skills. Across four benchmark splits, three agent models, and two image backbones, COMFYCLAW achieves the best average image-generation evaluation score across all six agent configurations, outperforming a verifier-only baseline without skill evolution. Human annotations further show that annotators prefer COMFYCLAW over variants without skill evolution. Our results suggest that skill evolution is an effective mechanism for improving agent reliability and performance in recurring visual workflow construction.

1 Citations
0 Influential
9.5 Altmetric
48.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!