OpenClaw-Skill: 자율적 대규모 언어 모델을 위한 집단 스킬 트리 탐색
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
대규모 언어 모델(LLM) 에이전트에 효과적인 능력을 부여하는 것은 OpenClaw와 같은 실제 시스템에서 복잡한 문제를 해결하는 데 매우 중요합니다. 본 연구에서는 LLM의 도구 사용 능력, 다단계 추론 능력 및 동적 환경과의 상호 작용을 향상시키기 위해 재사용 가능한 기술을 자동으로 구축하는 프레임워크를 개발하고자 합니다. 이를 위해 우리는 집단 스킬 트리 탐색(CSTS)이라는 새로운 트리-탐색 기반 기술 구축 프레임워크를 제안합니다. CSTS는 구조화되고, 다양하며, 일반화 가능성이 높은 기술 트리를 구축합니다. CSTS의 핵심 아이디어는 집단 지능을 활용하여 두 가지 반복적인 단계인 '집단 스킬 노드 생성(CSN-Gen)' 및 '집단 스킬 노드 평가(CSN-Assess)'를 통해 효과적인 기술을 공동으로 탐색, 식별하고 조합하는 것입니다. CSN-Gen은 여러 모델의 집단 지식을 활용하여 각 하위 작업에 대한 다양한 후보 기술을 탐색함으로써 포괄적인 기술 탐색을 가능하게 합니다. CSN-Assess는 여러 모델을 평가자로 사용하여 기술 노드를 평가하고 선택하며, 두 가지 평가 메커니즘을 사용합니다: (1) 독립적인 평가를 통합하여 기술의 효과성을 견고하게 추정하는 '집단 품질 점수' 및 (2) 기술이 다양한 모델에서 얼마나 잘 일반화되는지 명시적으로 확인하는 '집단 전송성 점수'. CSTS를 통해 우리는 포괄적인 기술 트리와 함께 기술을 추가한 학습 데이터를 구축하여 모델이 효과적으로 기술을 학습하고 활용할 수 있도록 합니다. 또한, 우리는 집단 스킬 강화 학습을 도입하여 트리에 있는 여러 관련 기술을 적극적으로 선택함으로써 해결 공간 탐색 범위를 넓히고 단일 기술과 그 결과로 발생하는 균질하거나 최적이 아닌 솔루션에 갇히는 것을 방지합니다. 그 결과, 우리의 학습된 모델인 OpenClaw-Skill은 장기 계획 수립, 도구 사용 및 어려운 벤치마크에서의 일반화 능력에서 뛰어난 에이전트 기능을 보여줍니다.
Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framework that automatically constructs such reusable skills to enhance LLMs in tool use, multi-step reasoning, and dynamic environment interaction. To this end, we propose Collective Skill Tree Search (CSTS), a novel tree-search-based skill construction framework that constructs structured, diverse and generalizable tree of skills. The core idea of CSTS is to leverage collective intelligence to jointly search, identify and compose effective skills via two iterative phases: Collective Skill Node Generation (CSN-Gen) and Collective Skill Node Assessment (CSN-Assess). CSN-Gen exploits collective knowledge from multiple models to explore diverse candidate skills for each subtask, enabling comprehensive skill exploration. CSN-Assess employs multiple models as judges to evaluate and select skill nodes with two scoring mechanisms: (1) collective quality scoring that aggregates independent evaluations to produce a robust estimate of skill effectiveness, and (2) collective transferability scoring that explicitly verifies whether a skill generalizes well across different models. With CSTS, we construct a set of comprehensive tree of skills along with skill-augmented training data, enabling models to effectively learn and utilize skills. Besides, we introduce Collective Skill Reinforcement Learning, which actively selects multiple relevant skills from the tree to broaden solution-space exploration, avoid being trapped by a single skill and its resulting homogeneous or suboptimal solutions. As a result, our trained model, OpenClaw-Skill, exhibits outstanding agentic capabilities in long-horizon planning, tool use and generalization over challenging benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.