2606.11543v1 Jun 10, 2026 cs.AI

SkillJuror: 에이전트의 기술 구성 방식이 런타임 동작에 미치는 영향 측정

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

Jianghao Lin
Jianghao Lin
Shanghai Jiao Tong University
Citations: 1,555
h-index: 20
Bingwei Lu
Bingwei Lu
Citations: 3
h-index: 1
Bo-Sheng Huang
Bo-Sheng Huang
Citations: 2
h-index: 1
Yuanjian Zhou
Yuanjian Zhou
Citations: 68
h-index: 3
Weinan Zhang
Weinan Zhang
Citations: 992
h-index: 18
Zhiyu Chen
Zhiyu Chen
Citations: 20
h-index: 2
Zihan Guo
Zihan Guo
Citations: 42
h-index: 2

에이전트 스킬(Agent Skills)은 추론 시점에 절차적 지식을 대규모 언어 모델(LLM) 에이전트에 제공하지만, 현재 벤치마크는 스킬의 내용과 구성 방식 간의 차이를 제대로 구분하지 못합니다. 본 연구에서는 프로그레시브 디스클로저(Progressive Disclosure), 즉 간결한 루트 파일이 필요에 따라 에이전트를 관련 리소스로 안내하는 방식을 활용하여 이러한 차이점을 분석하고, 정규화된 플랫(flat) 방식과 비교했습니다. SkillJuror는 동일한 작업 지식 하에서 의미적으로 제어된 다양한 스킬 작성 방식을 평가하기 위한 프레임워크로, 여러 번의 실험을 통해 결과를 비교하고 경로 데이터를 활용합니다. 82개의 작업으로 구성된 SkillsBench 연구에서 프로그레시브 디스클로저 방식은 전체 결과에 영향을 미치기 전에 런타임 동작에 변화를 가져옵니다. 구체적으로, 각 경로에서 사용되는 스킬 리소스의 수는 1.18개에서 3.85개로 증가하고, 효과적인 활용 이벤트는 1.33개에서 3.92개로 증가합니다. 또한, 정규화된 플랫 방식에 비해 410개의 일치된 실험에서 17번 더 검증 단계를 통과했습니다 (+4.1%). 이러한 효과는 작업에 따라 다릅니다. 프로그레시브 디스클로저 방식은 지원 리소스가 구현, 확인 또는 수리에 도움이 될 때 효과적이지만, 정확한 출력 규칙, 수치 임계값 또는 긴 아티팩트 생성 파이프라인에 의존하는 경우에는 효과가 미미합니다. 이러한 결과는 스킬 구성 방식이 단순한 표현 이상의 의미를 지니며, 에이전트의 절차적 지식 검색 및 적용 방식을 변화시킬 수 있다는 것을 보여줍니다. 또한, 최종 결과 개선은 노출된 리소스가 해당 작업에 실제로 활용될 수 있는지 여부에 따라 달라집니다. 코드 및 관련 자료는 https://github.com/zhiyuchen-ai/skill-juror 에서 확인할 수 있습니다.

Original Abstract

Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organized. We study this distinction through Progressive Disclosure, where a concise root file points agents to supporting resources on demand, and compare it with a normalized flat baseline. We present SkillJuror, a framework for evaluating Skill writing paradigms through semantically controlled variants, matched multi-trial evaluations, and trajectory evidence while holding task knowledge fixed. In an 82-task SkillsBench study, Progressive Disclosure changes runtime behavior before aggregate outcomes: distinct Skill resources touched per trajectory rise from 1.18 to 3.85, and effective uptake events rise from 1.33 to 3.92. It also yields 17 additional verifier-passing trials out of 410 matched trials (+4.1%) over the normalized flat baseline. The benefit is task-dependent. Progressive Disclosure helps when supporting resources guide implementation, checking, or repair, but is weaker when success hinges on exact output conventions, numerical thresholds, or long artifact-generation pipelines. These results show that Skill organization is not mere presentation: it can change how agents search and apply procedural knowledge, while outcome gains depend on whether the exposed resources are actionable for the task. Code is available at https://github.com/zhiyuchen-ai/skill-juror.

3 Citations
0 Influential
33.4657359028 Altmetric
13.9 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!