경험이 지침으로 변모될 때: 자가 진화 에이전트 기술 시스템에서의 경로 독성화
When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
자가 진화 기술(SES) 시스템은 에이전트의 행동 패턴을 지속적인 기술로 추출하여, 신뢰할 수 없는 경험을 신뢰할 수 있는 지침으로 변환합니다. 본 연구에서는 이러한 과정에 대한 공격 방법인 'PoisonedEvolution'이라는 경로 독성화 공격을 소개합니다. 우리의 관찰 가능한 블랙박스 공격자는 대상 기술을 검사하고 제한된 수준의 정보를 제공할 수 있지만, 비공개 데이터 풀이나 진화 로직을 관찰하거나 기술 저장소를 수정할 수는 없습니다. 악성 코드 주입에는 포함(Inclusion), 진화 속성 부여(Evolution Attribution), 그리고 실현(Realization)이 필요합니다. 진화 속성 부여는 가장 중요한 제약 요소입니다. 대상 행동은 승격되기 전에 인과적으로 유용하고, 반복적이며, 일반화 가능해야 합니다. 우리는 비활성 '카나리아' 명세를 사용하여 대표적인 보안 효과 그룹 네 가지를 평가했습니다. 공격자가 10%의 지원을 제공할 경우, SkillClaw에서 사용되는 여섯 가지 주요 LLM 진화 시스템에서 PoisonedEvolution은 총 600번의 시도 중 546번(91.0% 성공률)에 대상 행동을 삽입하는 데 성공했습니다. 구조적으로 다른 Trace2Skill 파이프라인에서도 동일한 비율로, 600번의 시도 중 369번(61.5% 성공률)에 대상 행동을 삽입하여, 다양한 진화 아키텍처 간의 전이성을 보여줍니다. 제어된 실험 결과, 세 개의 일관된 공격자 레코드가 30개의 레코드 배치에서 충분한 효과를 나타내는 반면, 단일 레코드는 훨씬 약한 성능을 보였습니다. 추가 분석 결과, 반복적인 지원 제공, 인과적 프레임 구성, 그리고 도메인에 맞는 인코딩이 성공의 주요 요인임을 확인했습니다. 이러한 연구 결과는 증거 기반 홍보가 자가 진화 에이전트의 보안 경계로 작용함을 보여줍니다.
Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounded evidence, but cannot observe private pools or evolution logic or edit the skill bank. Artifact poisoning requires Inclusion, Evolution Attribution, and Realization. Attribution is the distinctive bottleneck: the target behavior must appear causally useful, recurrent, and generalizable before promotion. We evaluate four representative security-effect families using inert canary specifications. At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER). On the structurally different Trace2Skill pipeline at the same ratio, it embeds target behaviors in 369/600 trials (61.5% SER), demonstrating transfer across evolution architectures. In a representative controlled study, three consistent attacker records suffice in a 30-record batch, whereas a single record is much weaker. Ablations identify recurring support, causal framing, and domain-aligned encoding as the main determinants of success. These findings expose evidence promotion as a security boundary for self-evolving agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.