2608.04828v1 Aug 05, 2026 cs.CL

기술 활용: LLM이 실제로 에이전트 시스템에서 기술을 사용할 수 있는가?

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Zixiang Di
Zixiang Di
Citations: 2
h-index: 1
Jinyi Han
Jinyi Han
Citations: 37
h-index: 4
Ying Liao
Ying Liao
Citations: 16
h-index: 3
Zhichao Hu
Zhichao Hu
Citations: 50
h-index: 5
Yuanjian Xu
Yuanjian Xu
Citations: 14
h-index: 2
Xinyi Wang
Xinyi Wang
Citations: 23
h-index: 3
Fan Lu
Fan Lu
Citations: 195
h-index: 5
Zishang Jiang
Zishang Jiang
Citations: 25
h-index: 3
Yanghua Xiao
Yanghua Xiao
Citations: 41
h-index: 4

최근 대규모 언어 모델(LLM) 기반 에이전트는 점점 더 '기술'에 의존하고 있습니다. 여기서 '기술'은 언제 행동해야 하는지, 어떤 절차를 따라야 하는지, 그리고 어떤 도구를 사용할 수 있는지 등을 명시하는 구조화된 문서입니다. 기존의 평가는 주로 기술 자체의 품질이나 작업 성공 기여도에 초점을 맞추고 있으며, 에이전트가 관련 기술을 인식하고 스스로 적용할 수 있는지를 평가하지 못합니다. 본 연구에서는 'Skill-Use'라는 벤치마크를 소개하며, 이는 점진적인 정보 공개 방식을 통해 기술 활용 능력을 평가합니다. 여기서 에이전트는 기술의 이름과 간단한 설명만 보고 전체 절차를 검색하여 따라야 합니다. Skill-Use는 기술 활용의 세 가지 측면을 분리합니다. 'Trigger'는 에이전트가 관련 기술을 호출하는지 측정하고, 'Compliance'는 지정된 절차를 얼마나 정확히 따르는지 측정하며, 'Boundary'는 금지된 작업을 회피하는지 측정합니다. Skill-Use(SU) 점수는 이 세 가지 요소를 결합하며, 기술이 호출된 후에만 실행에 대한 점수를 부여합니다. 본 연구에서는 79개의 실제 기술과 177개의 실행 가능한 작업을 9개 도메인에 걸쳐 조합하고, 각 작업은 실제 파일 기반이며, 격리된 Docker 환경에서 실행되고, 경로 기반의 평가 기준을 사용하여 점수가 매겨집니다. 두 가지 유형의 에이전트 시스템 하에서 8개의 LLM을 평가한 결과, 신뢰할 수 있는 기술 활용 능력은 여전히 달성하기 어렵다는 것을 확인했습니다. 가장 성능이 좋은 구성에서도 SU 점수는 0.613에 불과했습니다. 'Trigger' 및 절차 준수 측면 모두 독립적인 문제점으로 작용하며, 점수와 모델 순위는 사용된 에이전트 시스템에 따라 달라지는 경향을 보여줍니다. 이는 기술 활용 능력이 모델의 고정적인 속성이 아니라, 에이전트 시스템에 의해 조건화되는 능력이라는 것을 시사합니다.

Original Abstract

Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its contribution to task success, leaving unexamined whether an agent can recognize a relevant skill and apply it on its own. We introduce Skill-Use, a benchmark that evaluates skill use under progressive disclosure, where an agent sees only a skill's name and short description and must retrieve the full procedure before following it. Skill-Use separates three facets of skill use. Trigger measures whether the agent invokes the relevant skill, Compliance measures how faithfully it follows the prescribed procedure, and Boundary measures whether it avoids forbidden operations. A Skill-Use (SU) score combines the three and credits execution only after the skill is triggered. Skill-Use pairs 79 real skills with 177 executable tasks across nine domains, each grounded in real files, run in an isolated Docker sandbox, and scored by a trajectory-based rubric. Evaluating eight LLMs under two agent harnesses, we find that reliable skill use remains out of reach, as the strongest configuration reaches an SU of only 0.613. Triggering and procedural compliance fail as independent bottlenecks, and both scores and model rankings shift with the harness, so skill use behaves as a capability conditioned on the harness rather than a fixed property of the model.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!