2608.09047v1 Aug 10, 2026 cs.CR

다양성이 중요합니다: 데이터 효율적인 백도어 공격을 위한 분포 기반 특징 커버리지 샘플 선택

Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xiaokang Zhou
Xiaokang Zhou
Citations: 10,438
h-index: 53
Feng-Qi Cui
Feng-Qi Cui
Citations: 34
h-index: 2
Meng Li
Meng Li
Citations: 22
h-index: 3
Xiaoke Chen
Xiaoke Chen
Citations: 0
h-index: 0
Yu-Tong Guo
Yu-Tong Guo
Citations: 3
h-index: 1

백도어 공격은 학습 데이터를 손상시켜 모델이 정상적인 정확도를 유지하면서 특정 입력에 대해 공격자가 지정한 목표를 예측하도록 만듭니다. 매우 낮은 오염 비율에서, 몇 개의 샘플만이 트리거-목표 연관성을 전달하므로, 악성 샘플 선택이 중요합니다. 기존 방법들은 일반적으로 개별 샘플 점수를 사용하여 후보를 순위화하는데, 이는 유사한 의미 영역의 중복된 샘플을 선택할 수 있으며, 많은 방법들이 특정 작업에 대한 추가적인 학습을 필요로 합니다. 우리는 분포 기반 특징 커버리지 샘플 선택 (DFCS)이라는, 학습이 필요 없고 트리거 방식에 독립적인 방법을 제안합니다. DFCS는 고정된 사전 훈련된 특징을 사용하여 각 오염 슬롯당 하나의 영역으로 클러스터링하고, 각 영역에서 중심점과 가장 가까운 샘플을 선택합니다. 지역적 1차 분석은 이러한 할당 방식을 특징 커버리지와 대표성 질량 용어로 연결합니다. CIFAR-10, Tiny-ImageNet 및 Imagenette 데이터셋에 대한 BadNets 및 Blended 공격에서 DFCS는 7개의 선택 방법 중 가장 높은 평균 공격 성공률을 달성했으며, 모든 6가지 데이터셋-공격 설정에서 평균적으로 96.30%의 성공률을 보였습니다. 이는 가장 강력한 비교 방법에 비해 평균적으로 4.60% 포인트 높았습니다. 이러한 결과는 분포 기반 특징 커버리지가 저가형 오염된 레이블 백도어 공격에 대한 효과적인 선택 원칙임을 뒷받침합니다.

Original Abstract

Backdoor attacks compromise training data so that a model retains clean accuracy but predicts an attacker-chosen target on triggered inputs. At very low poisoning rates, only a few samples convey the trigger--target association, making poison-sample selection critical. Existing methods typically rank candidates using per-sample scores, which can select redundant samples from similar semantic regions, and many require task-specific surrogate training. We propose Distributional Feature Coverage Sample Selection (DFCS), a training-free, trigger-agnostic method that clusters fixed pretrained features into one region per poisoning slot and selects the centroid-nearest sample from each region. A local first-order analysis relates this allocation to feature-coverage and representative-mass terms. Across BadNets and Blended attacks on CIFAR-10, Tiny-ImageNet, and Imagenette, DFCS achieves the highest mean attack success rate among seven selectors in all six dataset--attack settings, averaging $96.30\%$ and exceeding the strongest comparator in each setting by 4.60 percentage points on average while preserving clean accuracy. These results support distributional feature coverage as an effective selection principle for low-budget dirty-label backdoor attacks.

0 Citations
0 Influential
26.5 Altmetric
132.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!