2608.01637v1 Aug 03, 2026 cs.AI

살라미 공격: OpenClaw에 대한 은밀한 공모형 메모리 오염

Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw

Zheng Lin
Zheng Lin
Citations: 46
h-index: 3
Haichang Gao
Haichang Gao
Citations: 27
h-index: 2
Zhenxing Niu
Zhenxing Niu
Citations: 46
h-index: 3
Yuzhen Huang
Yuzhen Huang
Hong Kong University of Science and Technology
Citations: 1,595
h-index: 9
Xianmin Ye
Xianmin Ye
Citations: 0
h-index: 0

장기 메모리는 LLM 에이전트가 여러 세션 동안 유용한 정보를 유지할 수 있도록 하지만, 동시에 적대자가 에이전트의 지속적인 메모리를 오염시켜 행동을 조작할 수 있는 취약점을 생성합니다. 기존의 메모리 오염 공격은 주로 개별적으로 악성인 레코드에 의존하며, 여러 개의 겉보기에는 안전한 메모리가 함께 작용하여 위험한 동작을 유발하는 복합적인 위협을 간과합니다. 본 논문에서는 공모형 메모리 오염 공격을 구성하기 위한 자동화된 레드 팀 프레임워크인 MemCollusion을 소개합니다. MemCollusion은 '살라미 전술'이라는 전략을 사용하며, 이는 적대적 목표를 작고 개별적으로 무해한 조각으로 나누는 방식입니다. 이를 통해 겉보기에는 안전하지만 전체적으로 해로운 메모리 단편들을 생성합니다. MemCollusion은 네 가지 설계 제약 조건, 다섯 가지 이론 기반 전략 및 미세 조정된 생성기를 사용하여 메모리 연합을 구성합니다. 현실적인 크로스 세션 환경에서 공모형 메모리 오염을 평가하기 위해, 우리는 제작된 플랫폼 콘텐츠가 먼저 지속적인 메모리에 기록되고, 이후 별도의 세션에서 에이전트의 행동에 영향을 미치도록 설계된 제어된 연구 재현 시스템인 MoltLab을 개발했습니다. 우리는 두 가지 기본 모델을 사용하여 48개의 시나리오에서 MemCollusion을 OpenClaw에서 평가했습니다. 가장 강력한 메모리 절약 설정에서, MemCollusion은 평균적으로 81.3%의 메모리 절감률과 75.0%의 공격 성공률을 달성했으며, 양호한 메모리 희석 및 메모리 레벨 방어 모두에 효과적임이 입증되었습니다.

Original Abstract

Long-term memory enables LLM agents to retain useful information across sessions, but also creates an attack surface through which adversaries may poison an agent's persistent memory to steer its behavior. Existing memory poisoning attacks mainly rely on individually malicious records, overlooking a compositional threat: multiple benign-looking memories may jointly induce unsafe behavior. In this paper, we introduce MemCollusion, an automated red-teaming framework for constructing collusive memory poisoning attacks. MemCollusion applies salami tactics---a strategy that slices an adversarial objective into small, individually innocuous pieces---to generate memory fragments that are individually benign looking but collectively harmful. It constructs memory coalitions using four design constraints, five theory-informed strategies, and a fine-tuned generator. To assess collusive memory poisoning in a realistic cross-session setting, we develop MoltLab, a controlled research reproduction of Moltbook, in which crafted platform content must first be observed and distilled into persistent memory before influencing the agent's behavior in a separate session. We evaluate MemCollusion on OpenClaw using two backbone models across 48 scenarios. Under the strongest memory-saving setting, MemCollusion achieves an average Memory Save Rate of 81.3% and an Attack Success Rate of 75.0%, and remains effective under both benign memory dilution and memory-level defenses.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!