자기 진화형 LLM 에이전트 시스템의 안전성: 위협, 증폭 효과 및 사례 연구
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
자체적으로 모델 파라미터, 메모리, 도구 및 아키텍처를 업데이트하는 자기 진화형 LLM 에이전트 시스템은 적대적인 영향이 영구적으로 인코딩되고, 세대를 거듭하며 스스로 증폭되며, 지속적인 공격자의 접근 없이도 전체 시스템에 전파되는 새로운 유형의 위협 환경을 야기합니다. 본 연구에서는 모듈 라이프사이클 공격 표면(MLAS) 매트릭스를 중심으로 체계적인 보안 및 개인 정보 보호 분석을 수행했습니다. MLAS는 5가지 기능 모듈(Brain, Cognitive Resource, Execution, Self-Design, Collective)과 5가지 라이프사이클 단계(Bootstrap, Propose, Evaluate, Commit, Serve)로 공격 표면을 세분화합니다. 결과적으로 도출된 25개의 셀에 대한 분석 결과, 17개 셀이 심각한 위협에 노출되어 있으며, 효과적인 부분적인 완화 방안은 존재하지 않습니다. 우리는 시너지 효과를 내는 7가지 상호 관련된 증폭 효과를 식별했으며, 이러한 효과는 개별 모듈을 독립적으로 보호하는 것만으로는 해결할 수 없습니다. 두 개의 오픈 소스 프레임워크에 대한 비교 사례 연구 결과, 진화 기능을 기본으로 설계된 시스템은 기존 방식보다 3.5배 더 많은 공격 표면 셀을 활성화하며, 모든 CIA+개인 정보 보호 범주에서 100%의 공격 지속률(40/40 페이로드)을 보입니다. 반면, 함께 설치된 보안 스캐너는 공격의 2.5%만을 차단합니다. 본 연구 결과는 자기 진화가 알려진 모든 공격 유형을 일시적인 세션 기반 공격에서 영구적인 계보 기반 공격으로 전환시키고, 완전히 새로운 유형의 공격을 야기하며, 기존의 정적 방어 체계를 근본적으로 무력화한다는 것을 보여줍니다. 이러한 발견은 자기 수정 시스템에 대한 진화 인식 보안 프레임워크 및 형식적 검증의 필요성을 강조합니다.
Self-evolving LLM agent systems, which autonomously update their model parameters, memory, tools, and architectures, introduce a qualitatively new threat landscape in which adversarial influences become permanently encoded, self-amplify across generations, and propagate through populations without sustained attacker access. We present a systematic security and privacy analysis organized around the Module-Lifecycle Attack Surface (MLAS) matrix, which decomposes the attack surface into five functional modules (Brain, Cognitive Resource, Execution, Self-Design, Collective) $\times$ five lifecycle stages (Bootstrap, Propose, Evaluate, Commit, Serve). Analysis of the resulting 25 cells reveals that 17 face critical threats for which no effective partial mitigation. We identify seven cross-cutting amplification effects that interact synergistically and cannot be addressed by securing individual modules in isolation. Comparative case studies of two open-source frameworks demonstrate that evolution-native design activates $3.5\times$ more attack surface cells and achieves a 100% attack persistence rate (40/40 payloads across all CIA+Privacy categories), while co-located security scanners block only 2.5% of attacks. Our findings establish that self-evolution converts every known attack category from session-bounded to lineage-persistent, gives rise to entirely new attack classes, and renders static defenses structurally inadequate, motivating evolution-aware security frameworks and formal verification for self-modifying systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.