신뢰할 수 없는 에이전트 기술에 대한 체계적인 보안 감사 및 견고성 향상
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
에이전트 기술 패키지는 SKILL.md 파일, 스크립트, 참조 문서 및 저장소 컨텍스트를 재사용 가능한 기능 단위로 묶어, 미리 로드되는 기술에 대한 감사를 단일 프롬프트 필터링에서 파일 간의 보안 검토로 확장합니다. 기존의 안전 장치는 종종 위험을 식별하지만, 의미를 보존하는 재작성 시 악성 의도를 일관성 없이 회피하는 경우가 있습니다. 본 논문에서는 신뢰할 수 없는 에이전트 기술에 대한 사전 로드 감사를 견고한 세 가지 분류 문제로 정의하고, 역할 인지 증거 추출, 선택적 의미 검증 및 일관성을 유지하는 판단을 결합한 SkillGuard-Robust를 소개합니다. SkillGuard-Robust를 SkillGuardBench 및 두 개의 공개 생태계 확장에 대해 254개에서 404개의 패키지를 포함하는 다섯 가지 주요 평가 관점에서 평가했습니다. 404개의 패키지로 구성된 독립적인 데이터 세트에서 SkillGuard-Robust는 97.30%의 전체 정확 일치, 98.33%의 악성 위험 재현율, 98.89%의 공격 정확 일관성을 달성했습니다. 외부 생태계 관점에서 99.66%, 100.00%, 100.00%의 결과를 보였습니다. 이러한 결과는 다음과 같은 결론을 뒷받침합니다. 패키지 감사 기능 분리는 프레임워크 및 공개 생태계의 견고성을 크게 향상시키지만, 외부 소스에서 가져온 콘텐츠의 악용은 여전히 해결해야 할 과제입니다.
Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering into cross-file security review. Existing guardrails often flag risk but recover malicious intent inconsistently under semantics-preserving rewrites. This paper formulates pre-load auditing for untrusted Agent Skills as a robust three-way classification task and introduces SkillGuard-Robust, which combines role-aware evidence extraction, selective semantic verification, and consistency-preserving adjudication. We evaluate SkillGuard-Robust on SkillGuardBench and two public-ecosystem extensions through five large evaluation views ranging from 254 to 404 packages. On the 404-package held-out aggregate, SkillGuard-Robust reaches 97.30% overall exact match, 98.33% malicious-risk recall, and 98.89% attack exact consistency. On the 254-package external-ecosystem view, it reaches 99.66%, 100.00%, and 100.00%, respectively. These results support a bounded conclusion: factorized package auditing materially improves frozen and public-ecosystem robustness, while harsher external-source transfer remains an open challenge.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.