2605.25430v1 May 25, 2026 cs.AI

CODESKILL: 코딩 에이전트를 위한 자기 진화 기술 학습

CODESKILL: Learning Self-Evolving Skills for Coding Agents

Yang Liu
Yang Liu
Citations: 54
h-index: 3
Yanzhou Li
Yanzhou Li
Nanyang Technological University
Citations: 256
h-index: 7
Yiran Zhang
Yiran Zhang
Citations: 37
h-index: 3
Xiaoyu Zhang
Xiaoyu Zhang
Citations: 6
h-index: 1
Xiaoxia Liu
Xiaoxia Liu
Citations: 1
h-index: 1

코딩 에이전트는 소프트웨어 엔지니어링 작업을 수행하면서 다양한 실행 경로를 생성합니다. 이러한 경로를 활용하여 에이전트의 자기 진화를 가능하게 하는 재사용 가능한 절차적 기술을 추출할 수 있으며, 이를 통해 경험을 압축적으로 표현하고 향후 행동을 안내할 수 있습니다. 그러나 기존의 기술 구성 및 유지 관리 방법은 종종 고정된 프롬프트와 휴리스틱 업데이트 규칙에 의존하며, 지식을 어떻게 선택, 추상화 및 유지해야 하는지에 대한 명확성이 부족합니다. 본 논문에서는 CODESKILL이라는 LLM 기반 프레임워크를 제안합니다. CODESKILL은 기술 추출과 기술 저장소 유지를 학습 가능한 관리 정책으로 재구성합니다. CODESKILL은 코딩 에이전트의 실행 경로에서 다양한 수준의 절차적 기술을 추출하고, 새로운 경험을 통해 기술을 발전시키며, 향후 작업 해결을 위한 간결한 기술 저장소를 유지합니다. 강화 학습을 사용하여 CODESKILL을 훈련했으며, 밀집된 채점 기준 기반의 기술 품질 피드백과 고정된 하위 에이전트로부터 얻은 희소한 실행 가능성 피드백을 결합한 혼합 보상을 사용했습니다. EnvBench, SWE-Bench Verified 및 Terminal-Bench 2에 대한 실험 결과, CODESKILL은 기술을 사용하지 않는 기준 대비 평균 성공률을 9.69% 향상시키고, 가장 강력한 프롬프트 기반 또는 메모리 기반 기준 대비 4.01% 더 높은 성능을 보였으며, 반복적인 구성 과정에서 기술 저장소의 크기를 안정적으로 유지했습니다.

Original Abstract

Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled into reusable procedural skills that compactly encode experience to guide future behavior. However, existing skill construction and maintenance methods often rely on fixed prompts and heuristic update rules, leaving it unclear how knowledge should be selected, abstracted, and maintained to best serve downstream agents. We propose CODESKILL, an LLM-based framework that reformulates skill extraction and skill-bank maintenance as a learnable management policy. CODESKILL extracts multi-granularity procedural skills from coding-agent trajectories, evolves skills with new experience, and maintains a compact skill bank for future task solving. We train CODESKILL with reinforcement learning, using a hybrid reward that combines dense rubric-based skill-quality feedback with sparse verifiable execution feedback from the frozen downstream agent. Experiments on EnvBench, SWE-Bench Verified, and Terminal-Bench 2 show that CODESKILL improves average pass rate by 9.69 over the no-skill baseline and by 4.01 over the strongest prompt-based or memory baseline, while maintaining the skill bank at a stable size during iterative construction.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!