2604.21700v1 Apr 23, 2026 cs.CR

자연스러운 스타일 트리거 기반의 LLM에 대한 은밀한 백도어 공격

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

Jiali Wei
Jiali Wei
Citations: 50
h-index: 2
Ming Fan
Ming Fan
Citations: 144
h-index: 6
Guoheng Sun
Guoheng Sun
Citations: 1
h-index: 1
Xicheng Zhang
Xicheng Zhang
Citations: 29
h-index: 4
Haijun Wang
Haijun Wang
Citations: 154
h-index: 5
Ting-Chun Liu
Ting-Chun Liu
Citations: 88
h-index: 2

대규모 언어 모델(LLM)이 안전이 중요한 영역에서 광범위하게 사용됨에 따라, 그 보안에 대한 우려가 커지고 있습니다. 최근 연구들은 LLM에 대한 백도어 공격의 가능성을 보여주었습니다. 하지만 기존 방법들은 명시적인 트리거 패턴으로 인해 자연스러움을 저해하고, 장문 생성 시 공격자가 지정한 내용을 안정적으로 주입하기 어렵으며, 백도어가 실제로 어떻게 전달되고 활성화되는지에 대한 불완전한 위협 모델을 가지고 있다는 세 가지 주요 단점을 가지고 있습니다. 이러한 문제점을 해결하기 위해, 우리는 BadStyle이라는 완전한 백도어 공격 프레임워크 및 파이프라인을 제시합니다. BadStyle은 LLM을 활용하여 자연스럽고 은밀한 악성 샘플을 생성하며, 의미와 유창성을 유지하면서 인지하기 어려운 스타일 레벨의 트리거를 포함합니다. 미세 조정 과정에서 공격자가 지정한 콘텐츠가 제대로 주입되도록 하기 위해, 악성 입력에 대한 응답에서 공격자가 지정한 목표 콘텐츠를 강화하고, 정상적인 응답에서 해당 콘텐츠가 나타나는 것을 억제하는 보조 목표 손실 함수를 설계했습니다. 또한, BadStyle을 현실적인 위협 모델에 기반하여 설계하고, 프롬프트 기반 및 PEFT 기반 주입 전략 모두에서 체계적으로 평가했습니다. LLaMA, Phi, DeepSeek 및 GPT 시리즈를 포함한 7개의 LLM을 대상으로 한 광범위한 실험 결과, BadStyle은 높은 공격 성공률(ASR)을 달성하면서도 강력한 은밀성을 유지하는 것으로 나타났습니다. 제안된 보조 목표 손실 함수는 백도어 활성화의 안정성을 크게 향상시켜, 스타일 레벨 트리거에 대한 평균 ASR을 약 30% 향상시켰습니다. 또한, 주입 과정에서 알 수 없는 다운스트림 배포 시나리오에서도 백도어는 효과적으로 작동합니다. 더욱이, BadStyle은 대표적인 입력 레벨 방어 기법을 회피하고, 간단한 위장 기술을 통해 출력 레벨 방어 기법을 우회합니다.

Original Abstract

The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studies have demonstrated the feasibility of backdoor attacks against LLMs. However, existing methods suffer from three key shortcomings: explicit trigger patterns that compromise naturalness, unreliable injection of attacker-specified payloads in long-form generation, and incompletely specified threat models that obscure how backdoors are delivered and activated in practice. To address these gaps, we present BadStyle, a complete backdoor attack framework and pipeline. BadStyle leverages an LLM as a poisoned sample generator to construct natural and stealthy poisoned samples that carry imperceptible style-level triggers while preserving semantics and fluency. To stabilize payload injection during fine-tuning, we design an auxiliary target loss that reinforces the attacker-specified target content in responses to poisoned inputs and penalizes its emergence in benign responses. We further ground the attack in a realistic threat model and systematically evaluate BadStyle under both prompt-induced and PEFT-based injection strategies. Extensive experiments across seven victim LLMs, including LLaMA, Phi, DeepSeek, and GPT series, demonstrate that BadStyle achieves high attack success rates (ASRs) while maintaining strong stealthiness. The proposed auxiliary target loss substantially improves the stability of backdoor activation, yielding an average ASR improvement of around 30% across style-level triggers. Even in downstream deployment scenarios unknown during injection, the implanted backdoor remains effective. Moreover, BadStyle consistently evades representative input-level defenses and bypasses output-level defenses through simple camouflage.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!