2607.14456v1 Jul 16, 2026 cs.SE

일반적인 LLM을 넘어서: 구조화된 코드 워크플로우 실행을 위한 전문 에이전트 시스템

Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution

A. Leontjeva
A. Leontjeva
Citations: 414
h-index: 8
Harris Borman
Harris Borman
Citations: 6
h-index: 1
Herman Wandabwa
Herman Wandabwa
Citations: 3
h-index: 1
Fusun Yu
Fusun Yu
Citations: 0
h-index: 0
Sandeepa Kannangara
Sandeepa Kannangara
Citations: 47
h-index: 2
Ritchie Ng
Ritchie Ng
Citations: 0
h-index: 0
Justin Liu
Justin Liu
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 소프트웨어 개발 에이전트의 보급을 가속화했으며, 이러한 에이전트는 현재 통합 개발 환경(IDE) 확장 프로그램 및 독립형 애플리케이션으로 널리 사용되고 있습니다. 이러한 에이전트는 일반적으로 범용적이지만, 전문 에이전트가 추가적인 개발 노력을 정당화하는지 여부는 불분명합니다. 본 연구에서는 비즈니스 프로세스 자동화의 맥락에서 이 질문을 조사하며, 특히 비즈니스 프로세스 모델 및 표기법(BPMN) 다이어그램을 실행 가능한 에이전트 워크플로우로 변환하는 데 초점을 맞춥니다. BPMN은 명시적인 제어 흐름 의미를 지정하므로, 고정된 프로세스 모델과 입력이 실행 경로를 결정하는 결정론적 워크플로우에 중점을 둡니다. 본 연구에서는 이 작업에 대한 전문 워크플로우를 소개하고, Roo 및 Cline과 같은 범용 에이전트와 비교합니다. 결과는 전문 솔루션이 도구 사용 정확도에서 약 9~20%p, 페널티 조정된 지연 시간에서 2~4배, 도구 호출 오류 수를 3배 줄이는 성능을 보이는 에이전트를 생성하며, 동시에 생성 토큰 비용을 95% 이상 절감하고 수정 반복을 없앤다는 것을 보여줍니다. 또한 범용 에이전트는 기능 및 품질 측면에서 일관성 없는 코드를 생성하므로, 안정성과 유지 관리성이 필수적인 산업 환경에 적합하지 않음을 확인했습니다.

Original Abstract

Large Language Models (LLMs) have accelerated the adoption of software development agents, now widely available as Integrated Development Environment (IDE) extensions and standalone applications. While these agents are typically general-purpose, it remains unclear whether specialist agents justify their additional development effort. We investigate this question in the context of business process automation, focusing on the transformation of Business Process Model and Notation (BPMN) diagrams into executable agentic workflows. Since BPMN specifies explicit control-flow semantics, we focus on deterministic workflows in which a fixed process model and inputs uniquely determine the executed path. We introduce a specialist workflow for this task and compare it against generalist agents such as Roo and Cline. Our results show that the specialist solution produces agents that outperform generalist baselines by approximately 9-20 percentage points in tool-use exactness, 2-4x in penalty-adjusted latency, and 3x fewer tool-call errors, while reducing generation token cost by over 95% and eliminating repair iterations. We also find that generalist agents generate code inconsistently in both functionality and quality, limiting their suitability for industrial settings where reliability and maintainability are essential.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!