SkillEval: 에이전트 기술 품질을 해석 가능한 신호로 분해하기
SkillEval: Decomposing Agent Skill Quality into Interpretable Signals
에이전트 기술은 에이전트가 특정 작업을 해결하는 데 도움이 되는 재사용 가능한 절차적 지식을 제공합니다. 이러한 기술의 활용이 증가함에 따라, 기술 품질 평가의 중요성이 더욱 커지고 있습니다. 기존의 평가는 주로 특정 하위 작업에서의 성능 향상을 통해 기술 품질을 측정합니다. 그러나 재사용 가능한 기술은 여러 가지 작업 시나리오에 적용될 수 있습니다. 하위 작업 평가는 기술과 평가 대상 작업 간의 호환성을 주로 반영하며, 기술 품질의 일부 측면만을 보여주고, 어떤 부분이 개선되어야 하는지 식별하지 못합니다. 우리는 `SKILL.md` 문서의 일반적인 속성이 기술 품질에서 중요한 역할을 한다는 것을 발견했습니다. 이러한 속성을 평가하기 위해, 문서 수준의 기술 평가를 위한 해석 가능한 프레임워크인 **SkillEval**을 제안합니다. SkillEval은 각 속성에 대해 고정되고 검사 가능한 평가 방향을 사용하여 해석 가능한 점수를 생성합니다. 또한, 길이 및 서식과 같은 관련 없는 문서 특징의 영향을 측정하고 줄여서, 각 점수가 의도된 의미론적 속성을 보다 구체적으로 반영하도록 합니다. 특히, SkillEval은 모델의 잠재 표현 공간에서 제어된 긍정-부정 기술 쌍으로부터 각 품질 속성에 대한 해석 가능한 방향을 학습하고, 새로운 기술의 표현을 이러한 고정된 방향으로 투영하여 점수를 매깁니다. 우리는 SkillEval을 사용하여 통제된 품질 테스트에서 기술을 평가하고, SkillEval이 다양한 품질의 기술을 신뢰성 있게 구별한다는 것을 보여줍니다. 또한, SkillEval 점수는 하위 작업 성능과 밀접하게 관련되어 있으며, 에이전트가 작업을 완료하는 데 기술이 도움이 될 가능성이 있는지 여부에 대한 초기 지표를 제공합니다. 우리는 또한 SkillEval을 사용하여 기술 문서의 약점을 진단하고, 표적 수정 사항을 안내하기 위해 활용했습니다. 수정된 기술은 표적 속성을 개선하고, 하위 작업에서 더 높은 성공률을 달성했습니다.
Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing whether a skill improves performance on specific downstream tasks. However, a reusable skill may apply to multiple task scenarios. Downstream evaluation mainly reflects the compatibility between a skill and the evaluated task, provides only a partial view of skill quality, and does not identify which aspect of the skill should be improved. We find that general properties of the \texttt{SKILL.md} document play an important role in skill quality. To evaluate these properties, we propose \textbf{SkillEval}, an interpretable framework for document-level skill evaluation. SkillEval evaluates each property using a fixed and inspectable scoring direction, producing interpretable scores. It further measures and reduces the influence of unrelated document features, such as length and formatting, so that each score captures its intended semantic property more specifically. Specifically, SkillEval learns an interpretable direction for each quality property from controlled positive--negative skill pairs in the hidden representation space of the model, and scores a new skill by projecting its representation onto these fixed directions. We use SkillEval to evaluate skills in controlled quality tests and show that SkillEval reliably distinguishes skills of different quality. In addition, SkillEval scores closely reflect downstream task performance, providing an early indication of whether a skill is likely to help an agent complete a task. We further explore SkillEval for diagnosing weaknesses in skill documents and guiding targeted revisions. The revised skills improve the targeted properties and achieve higher pass rates on downstream tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.