2608.03874v1 Aug 04, 2026 cs.AI

ContinualSkillBench: LLM 에이전트는 정말로 능력을 발전시킬 수 있는가?

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Haotong Yang
Haotong Yang
Citations: 170
h-index: 7
Yi Hu
Yi Hu
Citations: 225
h-index: 6
Tianyi Guan
Tianyi Guan
Citations: 0
h-index: 0
Siyuan Cao
Siyuan Cao
Citations: 0
h-index: 0

최신 에이전트 프레임워크는 대규모 언어 모델에 외부 기술 라이브러리를 제공하여 복잡한 작업을 해결합니다. 그러나 이러한 시스템이 기술을 효과적으로 발전시키고, 그 결과 얻어지는 기술이 작업 수행 능력을 향상시키는지 여부는 아직 불확실합니다. 이 격차를 해소하기 위해, 우리는 문맥 내 지속적인 기술 학습을 위한 동적 평가 프레임워크인 ContinualSkillBench를 소개합니다. 이는 다섯 가지 대표적인 영역으로 구성되어 있으며, 각 영역은 난이도가 점진적으로 증가하고 작업 간 기술 재사용 기회가 있는 100개의 상호 연결된 하위 작업으로 구성됩니다. 우리의 실험 결과는 순차적 실행이 일반적으로 성능을 향상시키지만, 모델 및 영역에 따라 이익의 정도가 크게 다르다는 것을 보여줍니다. 또한, 평균적으로 문맥 내 학습은 명시적인 기술 유지와 유사한 성능을 보이며, 이는 상당 부분 개선 사항이 재사용 가능한 기술 추상화뿐만 아니라 이전 문맥과 피드백에 대한 적응에서 비롯된다는 것을 시사합니다. 그러나 명시적인 기술은 재사용 가능한 절차 또는 정확한 출력을 요구하는 작업에 대해 선택적으로 이점을 제공합니다. 또한, 성능이 낮은 모델은 경향적으로 더 크고 단편화된 작업별 기술 모음을 축적한다는 것을 발견했습니다. 이러한 결과는 현재 문맥 내 기술 발전 메커니즘이 지속적인 적응을 지원할 수 있지만, 여전히 경험을 견고하고 전송 가능한 기술로 일관되게 통합하는 데 어려움을 겪고 있음을 보여줍니다.

Original Abstract

Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!