2606.10749v1 Jun 09, 2026 cs.CR

안전한 LLM 에이전트 개발을 위한 연구: 위협 요소, 공격 방법, 방어 기술 및 평가

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

Chunrong Fang
Chunrong Fang
Citations: 882
h-index: 15
Zhenyu Chen
Zhenyu Chen
Citations: 1,336
h-index: 20
Yuchen Ling
Yuchen Ling
Citations: 38
h-index: 4
Shengcheng Yu
Shengcheng Yu
Citations: 603
h-index: 12

대규모 언어 모델(LLM) 기반 에이전트는 단순히 대화형 인터페이스를 넘어 계획 수립, 도구 실행, 메모리 관리 및 외부 환경과의 상호 작용을 수행하는 소프트웨어 구성 요소로 빠르게 발전하고 있습니다. 이러한 변화는 보안 위험의 성격을 바꿉니다. 에이전트 환경에서 오류는 안전하지 않은 텍스트 생성으로만 제한되지 않습니다. 신뢰할 수 없는 콘텐츠가 제어 흐름을 변경하거나, 도구 권한을 남용하거나, 지속적인 상태를 손상시키거나, 민감한 정보를 유출하거나, 악의적인 외부 작업을 트리거할 수 있습니다. 동시에 LLM 에이전트 보안에 대한 연구는 빠르게 확장되고 있지만 공격 유형, 방어 계층, 응용 분야 및 평가 환경에 따라 여전히 분산되어 있습니다. 본 논문에서는 정보 흐름, 위임된 권한 및 지속적인 상태 간의 상호 작용을 중심으로 에이전트 보안을 모델링하는 수명 주기 기반 시스템 지향 프레임워크를 사용하여 247편의 연구를 종합 분석했습니다. 우리는 문헌을 다음 네 가지 질문에 따라 분류했습니다: LLM 에이전트 보안은 어떻게 모델링되어야 하는가, 어떤 위협 요소 및 공격 유형이 가장 일반적인가, 어떤 방어 기술이 제안되었으며 그 장단점은 무엇인가, 그리고 보안 주장은 어떻게 평가되는가. 분석 결과, 프롬프트 인젝션과 도구 기반의 제어 흐름 탈취 공격이 여전히 주요 문제이며, 지속적인 상태 손상 및 다중 에이전트 전파는 점점 더 중요한 우려 사항으로 부상하고 있음을 확인했습니다. 또한 현재 방어 기술은 유용한 구성 요소를 제공하지만, 아직 통합성이 부족하며, 기존 벤치마크는 장기적인 관점, 상태 관리 및 실제 배포 환경에서 발생할 수 있는 위험을 충분히 반영하지 못한다는 것을 발견했습니다. 우리는 안전한 LLM 에이전트를 개발하기 위해서는 명확한 신뢰 경계, 체계적인 권한 제어, 출처 추적 기능을 갖춘 상태 관리, 그리고 현실적인 운영 환경과 일치하는 평가 방법이 필요하다고 주장합니다.

Original Abstract

Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments. This transition changes the nature of security risk. In agentic settings, failures are no longer limited to unsafe text generation. Untrusted content may redirect control flow, misuse tool privileges, corrupt persistent state, leak sensitive information, or trigger harmful external actions. At the same time, research on LLM agent security is expanding quickly but remains fragmented across attack families, defense layers, application domains, and evaluation settings. This paper synthesizes 247 papers through a lifecycle-based, systems-oriented framework that models agent security around the interaction of information flow, delegated authority, and persistent state. We organize the literature around four questions: how LLM agent security should be modeled, which threat surfaces and attack families dominate, what defenses have been proposed and with what tradeoffs, and how security claims are evaluated. We find that prompt injection and tool-mediated control-flow hijacking still dominate the field, while persistent state corruption and multi-agent propagation are becoming central emerging concerns. We further find that current defenses provide useful building blocks but remain weakly compositional, and that existing benchmarks still underrepresent long-horizon, stateful, and deployment-sensitive risks. We argue that secure LLM agents require explicit trust boundaries, principled privilege control, provenance-aware state management, and evaluation practices aligned with realistic operational settings.

2 Citations
0 Influential
10 Altmetric
52.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!