2607.28103v1 Jul 30, 2026 cs.AI

MIND: LLM 에이전트를 위한 가볍고 효과적인 메모리 주입 방어 기술 - 의도 기반 정보 병목 현상 활용

MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck

Xiaobao Wu
Xiaobao Wu
Citations: 83
h-index: 4
Dongyi Liu
Dongyi Liu
Citations: 0
h-index: 0
Haixing He
Haixing He
Citations: 0
h-index: 0
Jia Li
Jia Li
Citations: 289
h-index: 2

메모리를 활용하는 LLM 기반 에이전트는 메모리 주입 공격에 취약합니다. 공격자가 악성 메모리를 삽입하여, 에이전트의 초기 사용자 의도를 벗어나게 만들고 결국 작업 실패를 초래할 수 있습니다. 기존 방어 기법들은 높은 계산 비용을 요구하거나 다중 턴 환경에서 정보 중복 문제를 가지고 있습니다. 이러한 문제점을 해결하기 위해, 우리는 가벼운 방어 프레임워크인 Memory Intent-Aware Neural Denoising (MIND)를 제안합니다. 초기 분석 결과, 정상적인 실행 경로와 공격에 의해 오염된 경로 간에는 초기 사용자 의도와 후속 행동 사이의 관계에서 구별되는 특징이 존재합니다. 이러한 관찰을 바탕으로, MIND는 의도 기반 정보 병목(Information Bottleneck, IB)을 사용하여 초기 의도와 각 턴에서의 행동으로부터 핵심적인 의도-행동 표현을 추출합니다. IB는 의도 관련 정보를 유지하면서 동시에 작업과 무관하거나 반복적인 정보를 제거하며, 경량화된 감지기는 이러한 표현으로부터 악성 메모리를 식별합니다. 결과적으로, MIND는 다중 턴 환경에서의 정보 중복 문제를 완화하고, LLM 감사에 필요한 추가 비용을 줄입니다. 광범위한 실험 결과, MIND는 공격 성공률을 감소시키면서도 작업 정확도와 추론 효율성을 유지합니다. 특히, ReAct-StrategyQA 데이터셋에서 MIND는 평균 공격 성공률(ASR-r 및 ASR-a)을 각각 55.4% 및 55.3% 감소시켰으며, 방어되지 않은 에이전트와 동일한 수준의 정확도와 지연 시간을 보였습니다.

Original Abstract

Memory-augmented LLM-based agents are vulnerable to memory injection attacks: Agents may retrieve poisoned memory from attackers, which diverts their behavior from initial user intent and finally causes task failure. However, existing defense mechanisms either incur high computational cost or suffer from information redundancy in multi-turn contexts. To address these challenges, we propose Memory Intent-Aware Neural Denoising(MIND), a lightweight defense framework for memory injection attack. Our preliminary analysis reveals that benign and poisoned trajectories exhibit distinguishable relationships between the initial user intent and subsequent behavior. Building on this observation, MIND employs an intent-aware Information Bottleneck(IB) to extract compact intent--behavior representations from the initial intent and turn-level behavior. The IB preserves intent-relevant cross-turn attack signals while filtering task-irrelevant and repetitive information, and a lightweight detector identifies malicious memories from the resulting representations. As such, MIND mitigates information redundancy in multi-turn contexts while avoiding the overhead of repeated LLM auditing. Extensive experiments show that MIND reduces attack success rates while preserving task accuracy and inference efficiency. Notably, on ReAct-StrategyQA, MIND reduces mean ASR-r and ASR-a by 55.4% and 55.3%, respectively, while matching the undefended agent in average accuracy and latency.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!