2608.09253v1 Aug 10, 2026 cs.AI

SkillSentry: 런타임 보증을 통한 LLM 에이전트의 안정적인 기술 실행

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

Xinyu Huang
Xinyu Huang
Citations: 24
h-index: 2
Bihuan Chen
Bihuan Chen
Citations: 3,774
h-index: 27
You Lu
You Lu
Citations: 34
h-index: 3
Xin Peng
Xin Peng
Citations: 8
h-index: 2

LLM(Large Language Model) 에이전트는 복잡한 작업을 수행하기 위해 다단계 추론과 도구 사용 기능을 갖추고 점점 더 많이 활용되고 있습니다. 기술은 재사용 가능한 절차적 지식을 제공하지만, 에이전트가 여전히 불안정하게 실행할 수 있습니다. 에이전트가 특정 기술의 안내에 따라 작업을 완료할 수 있는 능력을 보여주더라도, 기술 절차에서 벗어나거나 개별 단계를 잘못 수행하면 유사한 작업이나 반복적인 실행에서 일관성을 유지하지 못할 수 있습니다. 이러한 불안정성은 LLM 에이전트의 실질적인 신뢰성을 제한합니다. 이 문제를 해결하기 위해, 우리는 기술 중심의 런타임 보증 프레임워크인 SkillSentry를 제안합니다. SkillSentry는 기술 실행에 대한 런타임 지침을 표현하기 위한 새로운 도메인 특화 언어(DSL)를 기반으로 구축되었습니다. SkillSentry는 해당 기술 문서에서 추출한 기술 사양과 과거 성공 및 실패 사례에서 얻은 실행 경험을 결합하여 런타임 지침을 초기화합니다. 그런 다음, SkillSentry는 에이전트 실행 루프에 통합되어 현재 지침에 따라 기술 실행을 모니터링하고 안내하며, 수집된 새로운 데이터를 사용하여 지침을 반복적으로 개선합니다. 우리는 Claude Code (Claude-Haiku-4.5 및 Claude-Opus-4.6)와 Codex (GPT-5.2 및 GPT-5.4)를 사용하는 두 가지 LLM 에이전트, 각 에이전트에 대해 두 개의 백본 모델을 쌍으로 묶어 총 15개의 기술에 대해 SkillSentry를 평가했습니다. 그 결과, SkillSentry는 평균적으로 기술별로 LLM 에이전트의 작업 성공률을 24.1% 향상시켰으며, 반복 실행 간의 변동성을 줄이는 효과가 있음을 확인했습니다.

Original Abstract

LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent has demonstrated the capability to complete tasks under the guidance of a skill, it may fail to do so consistently across similar tasks or repeated runs due to deviations from the skill procedure or incorrect execution of individual steps. Such instability limits the practical reliability of LLM agents. To address this problem, we propose SkillSentry, a skill-oriented runtime assurance framework built upon a new domain-specific language (DSL) for representing runtime guidance for skill execution. SkillSentry initializes the runtime guidance by combining a skill specification extracted from the corresponding skill document with execution experience mined from historical successful and failed traces. It then wraps around the agent execution loop to monitor and guide skill execution under the current guidance, while iteratively refining the guidance using newly collected traces. We evaluate SkillSentry on 15 skills across two LLM agents, each paired with two backbone models, i.e., Claude Code with Claude-Haiku-4.5 and Claude-Opus-4.6, and Codex with GPT-5.2 and GPT-5.4. Our results show that SkillSentry improves the task success rate of LLM agents by 24.1% across skills, on average, while exhibiting lower variability across repeated runs.

0 Citations
0 Influential
13.5 Altmetric
67.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!