2607.17986v1 Jul 20, 2026 cs.CR

자가 호스팅 AI 에이전트에 대한 자기 상태 공격: 운영체제 방어 기술은 어디까지 나아갈 수 있는가?

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?

Roberto Di Pietro
Roberto Di Pietro
Citations: 316
h-index: 4
Jurgen Schmidhuber
Jurgen Schmidhuber
Citations: 211
h-index: 5
Yimeng Chen
Yimeng Chen
Citations: 47
h-index: 4
Nathanaël Denis
Nathanaël Denis
Citations: 22
h-index: 3

자가 호스팅 AI 에이전트는 작동을 위해 자체 메모리와 구성 파일을 읽고 씁니다. 에이전트의 상태가 손상되면, 이는 합법적인 운영체제 시스템 호출을 통해 발생하는 보안 침해로 이어질 수 있습니다. 우리는 이러한 유형의 위협을 '자기 상태 공격'이라고 부릅니다. 본 논문에서는 운영체제가 이러한 공격에 대한 탄력성을 얼마나 가지는지 조사합니다. 공식적으로, 우리는 네 가지 축(대상, 메커니즘, 세분성, 시간)으로 구성된 공격 공간을 정의하고, 방어, 탐지 및 복구의 구조적 한계를 분석하며, 워크로드에 따른 탐지 가능성에 대한 관점을 제시합니다. 프레임워크를 구체화하기 위해, 다양한 워크로드 프로필에서 실행되는 대표적인 자가 호스팅 에이전트로부터 실시간 활동 로그를 수집하고, 이 로그에 공격 공간을 23개의 셀로 구성된 행렬, 실제 자기 상태 파일에 대한 43가지의 구체적인 작업으로 구현하여 삽입했습니다. 그런 다음, 표준적인 방어 전략과 워크로드 기반 방어 전략 모두를 평가했습니다. 실험 결과, 계층화된 방어 체계(명령 및 구성 레이어에서의 접근 제어를 통한 방지, 메모리 레이어에서의 워크로드 기반 탐지, 그리고 주기적인 백업을 통한 복구)가 대부분의 공격 셀에 효과적이지만, 여전히 운영체제 수준에서 구조적으로 구별하기 어려운 작은 범위의 공격 표면이 남아 있음을 보여줍니다. 이러한 결과는 새로 정의된 '자기 상태 공격'에 맞서 운영체제 수준의 방어가 재고되어야 하며, 이는 해당 분야에서 새로운 연구 방향을 제시할 수 있음을 시사합니다.

Original Abstract

Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. We refer to this class of threats as self-state attacks. In this paper, we investigate the OS resilience to this class of attacks. Formally, we characterize a four-axis attack space (Target, Mechanism, Granularity, Temporal); investigate the structural limits of prevention, detection, and recovery; and introduce a workload-conditioned view of detectability. To instantiate the framework, we collect live activity traces from a representative self-hosted agent running across distinct workload profiles, and realize the attack space as a 23-cell matrix, 43 concrete operations on real self-state files, and injected into those traces. We then evaluate both canonical and workload-conditioned defense strategies. The empirical results show that a layered defense stack (access-control prevention on the instruction and configuration layers, workload-conditioned detection on the memory layer, and periodic backup for recovery) is effective on most attack cells while a small residual attack surface remains structurally indistinguishable at the OS level. These findings suggest that against the newly established class of self-state attacks, OS-level defense needs to be reconsidered, potentially opening new research directions in the field.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!