2607.12406v1 Jul 14, 2026 cs.AI

LLM 에이전트 시스템 안전을 위한 격리: 개념, 분류 체계, 과제 및 향후 연구 방향

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

Huihao Jing
Huihao Jing
Citations: 102
h-index: 5
Haoran Li
Haoran Li
Citations: 520
h-index: 12
Shaojin Chen
Shaojin Chen
Citations: 1
h-index: 1
W. Y. Chan
W. Y. Chan
Citations: 5
h-index: 2
Zhongwei Xie
Zhongwei Xie
Citations: 32
h-index: 3
Wei Fan
Wei Fan
Citations: 761
h-index: 8
Wenbin Hu
Wenbin Hu
Citations: 84
h-index: 4
Changxuan Fan
Changxuan Fan
Citations: 1
h-index: 1
Haochen Shi
Haochen Shi
Citations: 214
h-index: 8
Sirui Zhang
Sirui Zhang
Citations: 0
h-index: 0
Hanyu Yang
Hanyu Yang
Citations: 2
h-index: 1
Hongyu Luo
Hongyu Luo
Citations: 53
h-index: 2
Yangqiu Song
Yangqiu Song
Citations: 3,059
h-index: 15

LLM 에이전트가 시스템의 "두뇌" 역할을 수행할 수 있다는 점은 분석 범위를 독립적인 모델 수준에서 훨씬 확장합니다. 따라서 안전 문제는 단순히 입력-출력 내용 일치 여부에만 국한되지 않고, 시스템 동작 및 실제 실행 결과에도 영향을 미칩니다. 그러나 현재 연구는 공격 유형, 응용 분야 및 벤치마크에 따라 분산되어 있어, 프롬프트 주입, 도구 오용, 메모리 독살과 같은 오류들이 종종 동일한 구조적 원인을 공유하고 에이전트 워크플로우를 통해 어떻게 전파되는지 설명하기 어렵습니다. 본 논문에서는 LLM-에이전트 시스템 안전을 위한 핵심 원칙으로 "격리"를 제시합니다. 여기서 격리는 사용자 입력, 도구 접근, 실행 채널, 에이전트 간 통신 및 환경에서 비롯된 컨텍스트의 분리를 의미합니다. 우리는 사용자-에이전트, 에이전트-도구, 에이전트-실행, 에이전트-에이전트, 시스템-환경이라는 다섯 가지 경계를 중심으로 문헌을 분류했습니다. 이러한 관점을 통해 격리가 처음으로 손상되는 지점과, 침해가 경계를 넘어 어떻게 전파되는지, 그리고 각 인터페이스에서 어떤 방어 기술이 가장 관련성이 있는지 파악하는 데 도움이 됩니다. 또한, 경계를 넘나드는 오류 경로를 요약하고, 해결해야 할 과제를 논의하며, 향후 에이전트 시스템에서 격리를 통해 안전성을 확보하기 위한 연구 방향을 제시합니다.

Original Abstract

The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior and real-world execution outcomes. However, the current literature is fragmented across attack types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they spread through an agent workflow. In this survey, we treat isolation as a first-class principle for LLM-agent system safety. By isolation, we refer to the separation of user inputs, tool access, execution channels, inter-agent communication, and environment-originated context. We organize the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment. This view helps identify where the loss of isolation first occurs, how compromise propagates across boundaries, and which defenses are most relevant at each interface. We also summarize cross-boundary failure paths, discuss open challenges, and outline a research agenda for isolation-by-construction in future agent systems.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!