발톱이 기억하지만 말하지 않을 때: 지속적인 개인 에이전트 시스템에서의 은밀한 메모리 주입
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
지속적인 개인 에이전트는 장기 기억과 사용자의 외부 환경에 대한 접근성을 결합하여 맞춤형 프런트엔드 지원 및 능동적인 백그라운드 실행을 가능하게 합니다. 이러한 통합은 새로운 보안 취약점을 야기합니다. 신뢰할 수 없는 외부 콘텐츠가 지속적인 메모리에 조용히 기록되어 이후 신뢰된 상태로 재사용될 수 있습니다. 본 연구에서는 이를 '은밀한 메모리 주입'이라는 위협으로 정의하고, 원격의 블랙박스 공격자가 단일 이메일 페이로드를 통해 에이전트가 악성 메모리를 기록하도록 유도하고, 사용자의 응답에 숨겨져 있으며, 향후 행동에 영향을 미치도록 하는 과정을 분석합니다. 본 연구에서는 5가지 위험 범주와 사실 조작 및 선호도 조작을 포괄하는 108개의 테스트 케이스로 구성된 WhisperBench 벤치마크를 소개합니다. 실제 IMAP/SMTP 워크플로우와 정통 이메일 에이전트 기능을 기반으로 구축되었으며, 은밀한 메모리 주입 공격에 대한 전체 주기 평가를 가능하게 합니다. 단일 이메일 전송 및 런타임 피드백 없이 블랙박스 공격을 수행하기 위해, 우리는 MemGhost라는 일회성 페이로드 생성 프레임워크를 제안합니다. MemGhost는 환경 프록시를 사용하여 지속적인 에이전트 실행을 에뮬레이션하고, 객체 프록시를 사용하여 메모리 채택 및 대화형 은밀성을 밀집된 규칙 기반 보상으로 변환한 후, 지도 학습 미세 조정과 강화 학습을 통해 공격자 정책을 훈련합니다. 56개의 테스트 케이스에서 MemGhost는 GPT-5.4를 사용한 OpenClaw에서 전체 성공률 87.5%, Claude Code SDK와 Sonnet 4.6을 사용한 경우 71.4%의 성능을 달성했습니다. 또한, MemGhost는 다양한 개인 에이전트 아키텍처(NanoClaw 및 Hermes Agent) 및 메모리 백엔드(파일 시스템 및 벡터 기반 Mem0)에서 작동하며, 입력 레벨, 모델 레벨 및 시스템 레벨 방어에 대해서도 효과적인 것으로 나타났습니다. 이러한 결과는 지속적인 메모리가 일반적인 외부 처리를 장기 에이전트 침해를 위한 실질적인 경로로 만들 수 있음을 시사합니다.
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution. This integration also creates a new path to compromise: untrusted external content can be silently written into persistent memory and later reused as trusted state. We study this threat as stealth memory injection, in which a remote black-box adversary delivers a single email payload that must induce the agent to write poisoned memory, stay hidden in the agent's response to the user, and affect future behavior. We introduce WhisperBench, a 108-case benchmark spanning five risk categories and both fact and preference poisoning. Built on a real IMAP/SMTP workflow and an authentic email agent skill, it enables full-cycle evaluation of stealth memory injection attacks. To enable this black-box attack under single-email delivery and without runtime feedback, we propose MemGhost, a one-shot payload generation framework. MemGhost uses an environment proxy to emulate persistent-agent execution and an objective proxy to convert memory adoption and conversational stealth into dense rubric-based rewards, then trains the attacker policy with supervised fine-tuning and reinforcement learning. Across 56 held-out test cases, MemGhost achieves 87.5% end-to-end success on OpenClaw with GPT-5.4 and 71.4% on Claude Code SDK with Sonnet 4.6. It also transfers across personal-agent architectures (NanoClaw and Hermes Agent) and memory backends (filesystem and vector-based Mem0), and remains effective against input-level, model-level, and system-level defenses. These results suggest that persistent memory can turn ordinary external processing into a practical pathway for long-term agent compromise.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.