2607.05189v1 Jul 06, 2026 cs.CR

발톱이 기억하지만 말하지 않을 때: 지속적인 개인 에이전트 시스템에서의 은밀한 메모리 주입

When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents

Gelei Deng
Gelei Deng
Citations: 103
h-index: 6
Yechao Zhang
Yechao Zhang
Citations: 20
h-index: 2
Shiqian Zhao
Shiqian Zhao
Citations: 206
h-index: 6
Xiaogeng Liu
Xiaogeng Liu
Citations: 2,503
h-index: 19
Jie Zhang
Jie Zhang
Citations: 4
h-index: 1
Jiawen Zhang
Jiawen Zhang
Citations: 1,494
h-index: 20
Tianwei Zhang
Tianwei Zhang
Citations: 787
h-index: 11
Chaowei Xiao
Chaowei Xiao
Citations: 1,201
h-index: 13

지속적인 개인 에이전트는 장기 기억과 사용자의 외부 환경에 대한 접근성을 결합하여 맞춤형 프런트엔드 지원 및 능동적인 백그라운드 실행을 가능하게 합니다. 이러한 통합은 새로운 보안 취약점을 야기합니다. 신뢰할 수 없는 외부 콘텐츠가 지속적인 메모리에 조용히 기록되어 이후 신뢰된 상태로 재사용될 수 있습니다. 본 연구에서는 이를 '은밀한 메모리 주입'이라는 위협으로 정의하고, 원격의 블랙박스 공격자가 단일 이메일 페이로드를 통해 에이전트가 악성 메모리를 기록하도록 유도하고, 사용자의 응답에 숨겨져 있으며, 향후 행동에 영향을 미치도록 하는 과정을 분석합니다. 본 연구에서는 5가지 위험 범주와 사실 조작 및 선호도 조작을 포괄하는 108개의 테스트 케이스로 구성된 WhisperBench 벤치마크를 소개합니다. 실제 IMAP/SMTP 워크플로우와 정통 이메일 에이전트 기능을 기반으로 구축되었으며, 은밀한 메모리 주입 공격에 대한 전체 주기 평가를 가능하게 합니다. 단일 이메일 전송 및 런타임 피드백 없이 블랙박스 공격을 수행하기 위해, 우리는 MemGhost라는 일회성 페이로드 생성 프레임워크를 제안합니다. MemGhost는 환경 프록시를 사용하여 지속적인 에이전트 실행을 에뮬레이션하고, 객체 프록시를 사용하여 메모리 채택 및 대화형 은밀성을 밀집된 규칙 기반 보상으로 변환한 후, 지도 학습 미세 조정과 강화 학습을 통해 공격자 정책을 훈련합니다. 56개의 테스트 케이스에서 MemGhost는 GPT-5.4를 사용한 OpenClaw에서 전체 성공률 87.5%, Claude Code SDK와 Sonnet 4.6을 사용한 경우 71.4%의 성능을 달성했습니다. 또한, MemGhost는 다양한 개인 에이전트 아키텍처(NanoClaw 및 Hermes Agent) 및 메모리 백엔드(파일 시스템 및 벡터 기반 Mem0)에서 작동하며, 입력 레벨, 모델 레벨 및 시스템 레벨 방어에 대해서도 효과적인 것으로 나타났습니다. 이러한 결과는 지속적인 메모리가 일반적인 외부 처리를 장기 에이전트 침해를 위한 실질적인 경로로 만들 수 있음을 시사합니다.

Original Abstract

Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution. This integration also creates a new path to compromise: untrusted external content can be silently written into persistent memory and later reused as trusted state. We study this threat as stealth memory injection, in which a remote black-box adversary delivers a single email payload that must induce the agent to write poisoned memory, stay hidden in the agent's response to the user, and affect future behavior. We introduce WhisperBench, a 108-case benchmark spanning five risk categories and both fact and preference poisoning. Built on a real IMAP/SMTP workflow and an authentic email agent skill, it enables full-cycle evaluation of stealth memory injection attacks. To enable this black-box attack under single-email delivery and without runtime feedback, we propose MemGhost, a one-shot payload generation framework. MemGhost uses an environment proxy to emulate persistent-agent execution and an objective proxy to convert memory adoption and conversational stealth into dense rubric-based rewards, then trains the attacker policy with supervised fine-tuning and reinforcement learning. Across 56 held-out test cases, MemGhost achieves 87.5% end-to-end success on OpenClaw with GPT-5.4 and 71.4% on Claude Code SDK with Sonnet 4.6. It also transfers across personal-agent architectures (NanoClaw and Hermes Agent) and memory backends (filesystem and vector-based Mem0), and remains effective against input-level, model-level, and system-level defenses. These results suggest that persistent memory can turn ordinary external processing into a practical pathway for long-term agent compromise.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!