2606.09371v1 Jun 08, 2026 cs.AI

도구 활용 LLM을 위한 역량 정렬 계층적 학습

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

Haotong Yang
Haotong Yang
Citations: 170
h-index: 7
Ting Long
Ting Long
Citations: 48
h-index: 2
Yi Chang
Yi Chang
Citations: 4
h-index: 1

도구 학습은 LLM이 외부 도구를 사용하여 작업을 수행하도록 합니다. 이전 연구에서는 계층적 구조가 효과적인 것으로 입증되었습니다. 고수준 정책은 전반적인 계획을 관리하고 작업을 해결 가능한 하위 작업으로 분해하며, 저수준 정책은 이러한 하위 작업을 해결하기 위해 도구를 호출하는 데 중점을 둡니다. 그러나 기존 연구는 일반적으로 고수준 및 저수준 정책을 개별적으로 최적화하여 계획자와 실행자 간의 불일치를 초래하고 LLM의 도구 활용 작업 성능을 제한합니다. 본 논문에서는 RLVR을 활용하여 두 정책을 동시에 최적화하여 고수준 계획자와 저수준 실행자 간의 더 나은 조화를 이루는 방법인 Capability-Aligned Hierarchical Learning (CAHL)을 제안합니다. 제한된 도구 활용 벤치마크(API-Bank 및 BFCL)와 개방형 환경(Bamboogle)에서의 실험 결과는 CAHL의 효과를 입증합니다.

Original Abstract

Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure: a high-level policy handles global planning and decomposes tasks into manageable sub-tasks, and a low-level policy focuses on invoking tools to solve these sub-tasks. However, these works typically optimize the high-level and low-level policies separately, leading to planner-executor misalignment and limiting LLM performance on tool-use tasks. In this paper, we propose a method called Capability-Aligned Hierarchical Learning (CAHL), which leverages RLVR to jointly optimize both policies, enabling better alignment between the high-level planner and the low-level executor. Experiments on constrained tool-use benchmarks (API-Bank and BFCL) and an open-ended environment (Bamboogle) demonstrate the effectiveness of CAHL.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!