2603.28166v1 Mar 30, 2026 cs.CR

실제 도구에서 에이전트의 권한 사용 평가

Evaluating Privilege Usage of Agents on Real-World Tools

Quan Zhang
Quan Zhang
Citations: 360
h-index: 10
Li Fu
Li Fu
Citations: 5
h-index: 1
Lvsi Lian
Lvsi Lian
Citations: 5
h-index: 1
Gwihwan Go
Gwihwan Go
Citations: 103
h-index: 5
Yujue Wang
Yujue Wang
Citations: 42
h-index: 4
Chijin Zhou
Chijin Zhou
Citations: 781
h-index: 15
Yu Jiang
Yu Jiang
Citations: 115
h-index: 5
G. Pu
G. Pu
Citations: 3,531
h-index: 30

LLM 에이전트에게 실제 도구를 제공하면 생산성을 크게 향상시킬 수 있습니다. 그러나 에이전트에게 도구 사용의 자율성을 부여하는 것은 관련 권한을 에이전트와 기본 LLM 모두에게 이전하는 것을 의미합니다. 부적절한 권한 사용은 정보 유출 및 인프라 손상과 같은 심각한 결과를 초래할 수 있습니다. 에이전트의 보안을 연구하기 위한 여러 벤치마크가 개발되었지만, 이들은 종종 미리 코딩된 도구와 제한적인 상호 작용 패턴에 의존합니다. 이러한 인위적인 환경은 실제 환경과 크게 다르기 때문에, 에이전트의 중요한 권한 제어 및 사용 능력에 대한 보안성을 평가하기 어렵습니다. 따라서, 본 연구에서는 에이전트의 권한 사용을 분석하기 위한 보안 평가 샌드박스인 GrantBox를 제안합니다. GrantBox는 실제 도구를 자동으로 통합하고 LLM 에이전트가 실제 권한을 호출할 수 있도록 하여, 프롬프트 주입 공격에 대한 권한 사용을 평가할 수 있도록 합니다. 우리의 결과는 LLM이 기본적인 보안 인식을 가지고 일부 직접적인 공격을 차단할 수 있지만, 더 정교한 공격에 취약하며, 신중하게 설계된 시나리오에서 평균 공격 성공률이 84.80%라는 것을 나타냅니다.

Original Abstract

Equipping LLM agents with real-world tools can substantially improve productivity. However, granting agents autonomy over tool use also transfers the associated privileges to both the agent and the underlying LLM. Improper privilege usage may lead to serious consequences, including information leakage and infrastructure damage. While several benchmarks have been built to study agents' security, they often rely on pre-coded tools and restricted interaction patterns. Such crafted environments differ substantially from the real-world, making it hard to assess agents' security capabilities in critical privilege control and usage. Therefore, we propose GrantBox, a security evaluation sandbox for analyzing agent privilege usage. GrantBox automatically integrates real-world tools and allows LLM agents to invoke genuine privileges, enabling the evaluation of privilege usage under prompt injection attacks. Our results indicate that while LLMs exhibit basic security awareness and can block some direct attacks, they remain vulnerable to more sophisticated attacks, resulting in an average attack success rate of 84.80% in carefully crafted scenarios.

5 Citations
0 Influential
15 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!