2607.26099v1 Jul 28, 2026 cs.CR

릴리트: 학습-추론 트리거 변화 하에서의 백도어 일반화

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

Jianhai Chen
Jianhai Chen
Citations: 17
h-index: 2
Chunyi Zhou
Chunyi Zhou
Citations: 184
h-index: 7
JinBao Li
JinBao Li
Citations: 1
h-index: 1
Jiahao Chen
Jiahao Chen
Citations: 51
h-index: 4
Yuan Su
Yuan Su
Citations: 16
h-index: 1
Shouling Ji
Shouling Ji
Citations: 174
h-index: 6
Zhou Feng
Zhou Feng
Citations: 26
h-index: 3
Tianyu Du
Tianyu Du
Citations: 0
h-index: 0
Yuwen Pu
Yuwen Pu
Citations: 258
h-index: 11

머신러닝 서비스는 점점 더 많은 공용 데이터, 제3자 제공업체 및 아웃소싱된 학습에 의존하고 있으며, 이는 지속적인 악성 동작을 심으면서도 정상적인 유용성을 유지하는 데이터 포이즈닝 공격의 기회를 제공합니다. 그러나 기존의 백도어 연구는 주로 정확한 트리거 재사용, 학습 환경에 노출된 다양한 트리거 또는 미리 정의된 변환 축을 중심으로 평가됩니다. 따라서 중요한 한계점이 존재합니다. 즉, 특정 학습 시점의 트리거로 학습된 백도어가 피해 시스템의 학습 과정에서 나타나지 않는 추론 시점의 트리거 집단으로 일반화될 수 있는지 여부에 대한 문제입니다. 본 연구에서는 이 문제를 학습-추론 트리거 변화 하에서의 백도어 일반화 문제로 정의하고, 릴리트(Lilith)라는 블랙박스 기반의 앵커-집단 프레임워크를 제안합니다. 릴리트는 분리된 가짜 데이터 자원을 사용하여 먼저 단일 학습 앵커를 통해 대상 시스템에 제한적인 취약점을 유도한 다음, 앵커로 유도된 표현 형상을 유지하는 추론 전용의 제한적인 트리거 집단을 구성합니다. 본 연구에서는 앵커 제거 및 집단 도달 범위라는 개념을 사용하여 이 메커니즘을 설명하고, 국소적 규칙성과 제한된 가짜 데이터-대상 시스템 간의 불일치 하에서 집단 전체의 대상 유지에 필요한 조건을 도출합니다. 다양한 데이터셋, 아키텍처, 포이즈닝 비율 및 방어 기법에서의 실험 결과, 릴리트는 제한적인 유용성 저하와 작은 트리거 일반화 격차로 높은 수준의 집단별 공격 성공률을 달성하는 것으로 나타났습니다. 추가 분석 결과, 집단 활성화는 제안 메커니즘이 아닌 표현 정렬에 따라 달라지며, 이는 정확한 트리거 평가에서 간과되는 더 광범위한 위협을 드러냅니다.

Original Abstract

Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!