2601.12449v1 Jan 18, 2026 cs.CR

AgenTRIM: 에이전트 기반 AI를 위한 도구 위험 완화 도구

AgenTRIM: Tool Risk Mitigation for Agentic AI

Roy Betser
Roy Betser
Citations: 49
h-index: 5
Shamik Bose
Shamik Bose
Citations: 14
h-index: 2
Amit Giloni
Amit Giloni
Citations: 81
h-index: 5
Chiara Picardi
Chiara Picardi
Citations: 13
h-index: 2
Sindhu Padakandla
Sindhu Padakandla
Citations: 710
h-index: 7
R. Vainshtein
R. Vainshtein
Citations: 65
h-index: 5

AI 에이전트는 LLM과 외부 도구를 결합하여 복잡한 작업을 해결하는 자율 시스템입니다. 이러한 도구는 기능을 확장하지만, 부적절한 도구 권한은 간접적인 프롬프트 주입 및 도구 오용과 같은 보안 위험을 초래합니다. 우리는 이러한 실패를 불균형적인 도구 중심 에이전시로 규정합니다. 에이전트는 불필요한 권한을 유지(과도한 에이전시)하거나 필요한 도구를 호출하지 못할 수 있으며(불충분한 에이전시), 이는 공격 표면을 확대하고 성능을 저하시킵니다. 우리는 에이전트의 내부 추론을 변경하지 않고 도구 중심 에이전시 위험을 탐지하고 완화하는 프레임워크인 AgenTRIM을 소개합니다. AgenTRIM은 상호 보완적인 오프라인 및 온라인 단계를 통해 이러한 위험을 해결합니다. 오프라인에서 AgenTRIM은 코드 및 실행 추적으로부터 에이전트의 도구 인터페이스를 재구성하고 검증합니다. 런타임 시에는 적응적 필터링 및 도구 호출의 상태 인지 유효성 검사를 통해 단계별 최소 권한 도구 접근을 적용합니다. AgentDojo 벤치마크에서 AgenTRIM은 공격 성공률을 크게 줄이면서 높은 작업 성능을 유지합니다. 추가 실험에서 설명 기반 공격에 대한 견고성 및 명시적인 안전 정책의 효과적인 적용을 확인했습니다. 이러한 결과는 AgenTRIM이 LLM 기반 에이전트에서 더 안전한 도구 사용을 위한 실용적이고 기능 보존적인 접근 방식을 제공한다는 것을 보여줍니다.

Original Abstract

AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks. While such tools extend capability, improper tool permissions introduce security risks such as indirect prompt injection and tool misuse. We characterize these failures as unbalanced tool-driven agency. Agents may retain unnecessary permissions (excessive agency) or fail to invoke required tools (insufficient agency), amplifying the attack surface and reducing performance. We introduce AgenTRIM, a framework for detecting and mitigating tool-driven agency risks without altering an agent's internal reasoning. AgenTRIM addresses these risks through complementary offline and online phases. Offline, AgenTRIM reconstructs and verifies the agent's tool interface from code and execution traces. At runtime, it enforces per-step least-privilege tool access through adaptive filtering and status-aware validation of tool calls. Evaluating on the AgentDojo benchmark, AgenTRIM substantially reduces attack success while maintaining high task performance. Additional experiments show robustness to description-based attacks and effective enforcement of explicit safety policies. Together, these results demonstrate that AgenTRIM provides a practical, capability-preserving approach to safer tool use in LLM-based agents.

11 Citations
2 Influential
3.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!