도구 사용 LLM 에이전트에 대한 컨텍스트 분열 분해 공격: 아티팩트 출처 정보 누락을 이용한 공격
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps
도구를 사용하는 LLM 에이전트는 작업 공간 파일이나 로그와 같은 아티팩트에 상태를 저장하면서 외부 환경과 상호 작용합니다. 따라서, 제어 우회 방어는 개별 텍스트가 아닌 여러 단계에 걸친 상호 작용을 고려해야 합니다. 하지만 대부분의 기존 공격 및 방어 시스템, 예를 들어 Crescendo나 Tree of Attacks와 같은 다중 라운드 제어 우회 기법은 여전히 방어자에게 단일하고 연속적인 대화만 보이는 것을 전제로 작동합니다. 이러한 가정은 실제 에이전트 파이프라인에서는 유효하지 않습니다. 왜냐하면, 제어가 다양한 도구, 모듈 및 시간에 걸쳐 분산되어 있으며, 아티팩트의 출처 정보가 제대로 추적되지 않는 경우가 많기 때문입니다. 본 연구에서는 도구를 사용하는 LLM 에이전트의 배포 실패 요인인 '출처 정보 누락'을 정의하고, 이를 유발하는 재현 가능한 공격 기법인 '컨텍스트 분열 분해(CFD)'를 제안합니다. CFD는 초기 상호 작용에서 무해하게 보이는 중간 아티팩트를 보존하고, 이후에 개별적으로는 무해한 도구 사용을 통해 악의적인 행동을 유발하는 다단계 제어 우회 기법입니다. 이러한 공격 메커니즘을 상세히 분석하고, 출처 정보 추적 시스템을 도입하여 완화 전략을 제시합니다. 다양한 에이전트 시스템에 대한 제어 우회 벤치마크에서 CFD는 최첨단 모델 대비 최대 28.3%p의 성공률 향상을 보였으며, 강력한 단일 라운드 평가 시스템에서도 효과를 입증했습니다. (주의: 본 논문에는 유해하거나 불쾌할 수 있는 표현이 포함되어 있습니다.)
Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak defenses must reason about cross-step composition rather than isolated text. Yet most existing attacks and defenses, including ``multi-turn'' jailbreaks such as Crescendo and Tree of Attacks,still assume a single contiguous conversation visible to the defender. This assumption breaks down in real agent pipelines, where enforcement is fragmented across tools, modules, and time, and where artifact provenance is often not tracked. We operationalize a deployment failure mode for tool-using LLM agents, the \emph{provenance gap}, and study reproducible triggers for it: \emph{Context-Fractured Decomposition} (CFD), a family of cross-context multi-step jailbreaks that preserve benign-looking intermediate artifacts from an early interaction and elicit harmful behavior much later, potentially in a different agent instance or workflow stage, via individually innocuous tool actions whose risk emerges only under delayed artifact-mediated composition. We instrument the failure mode with trace-level diagnostics and outline a verifiable mitigation direction (provenance lineage tagging). Across agent-system jailbreak benchmarks, CFD improves success rates by up to 28.3 percentage points over state-of-the-art baselines, even against strong single-turn judges. Disclaimer: This paper contains examples of harmful or offensive language.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.