Search2Skill: 기준 기반 강화 학습을 통한 지식 경계를 초월하는 기술 증류
Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning
실제 업무 문제를 해결하는 데 필요한 절차적 지식을 담고 있는 재사용 가능한 기술은 LLM 기반 에이전트가 전문 분야에서 스스로 발전할 수 있도록 돕습니다. 기존의 자체 진화형 기술 학습 방법은 모델의 파라미터 지식 또는 경로를 통해 내부적으로 기술을 구축하며, 따라서 모델이 이미 알고 있는 내용에 한계가 있습니다. 그러나 전문적인 기술의 기반이 되는 도메인 규칙 및 표준 절차는 종종 이러한 경계를 벗어나며, 에이전트만으로는 이를 추출하기 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 Search2Skill이라는 새로운 프레임워크를 제안합니다. Search2Skill은 에이전트의 능력 부족을 자동으로 식별하고, 외부 소스를 검색하여 이를 보완하며, 검색된 정보를 구조화된 재사용 가능한 기술로 증류합니다. 구체적으로, Search2Skill은 기준 기반 강화 학습 방식을 사용하여 검색 시점, 검색 방법 및 기술 생성 방법을 동시에 개선하도록 최적화됩니다. 세 가지 벤치마크에서 추출한 여덟 개의 전문가 수준 도메인에 대한 실험 결과, Search2Skill은 스트리밍 및 별도 평가 프로토콜 모두에서 검색 기능을 추가하거나 경로 기반 기술 학습을 사용하는 기존 방식보다 일관되게 우수한 성능을 보였습니다. 추가 분석 결과, 성능 향상은 단순히 검색된 정보가 아닌 기술 추상화에서 비롯되며, 획득한 기술이 다양한 모델 규모에 걸쳐 전송될 수 있음을 확인했습니다.
Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from the model's parametric knowledge or trajectories, and are therefore bounded by what the model already knows. However, the domain conventions and standard procedures underlying professional skills often lie beyond this boundary and are hard to elicit from the agent alone. To address this issue, we therefore propose a novel framework, Search2Skill, that automatically identifies the agent's capability gaps, searches external sources to address them, and distills the retrieved evidence into structured, reusable skills. Specifically, Search2Skill is optimized by a rubric-based reinforcement learning scheme that jointly improves when to search, how to search, and how to generate skills. Experiments on eight expert-level domains from three benchmarks show that Search2Skill consistently outperforms both search-augmented and trajectory-based skill-learning baselines under both streaming and held-out evaluation protocols. Further analyses show that the gains arise from skill abstraction rather than raw retrieved evidence, and that the acquired skills transfer across model scales.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.