2604.18663v1 Apr 20, 2026 cs.CR

명시적인 거부 너머: 검색 증강 생성 시스템에 대한 소프트 실패 공격

Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation

Wentao Zhang
Wentao Zhang
Citations: 7
h-index: 1
Zhuang Yan
Zhuang Yan
Citations: 1
h-index: 1
ZhuHang Zheng
ZhuHang Zheng
Citations: 0
h-index: 0
Mingfei Zhang
Mingfei Zhang
Citations: 26
h-index: 2
Jiawen Deng
Jiawen Deng
Citations: 117
h-index: 5
Fuji Ren
Fuji Ren
Citations: 11
h-index: 2

기존의 검색 증강 생성(RAG) 시스템에 대한 공격은 주로 명시적인 거부 또는 서비스 거부(DoS)를 유발하는데, 이는 눈에 띄고 쉽게 탐지될 수 있습니다. 본 연구에서는 시스템 유용성을 저하시키는 보다 미묘한 가용성 위협인 '소프트 실패'를 정의합니다. 소프트 실패는 명백한 오류 대신 유창하고 일관적이지만 정보가 없는 응답을 유발합니다. 우리는 대규모 언어 모델의 안전 관련 동작을 악용하여 이러한 소프트 실패를 유발하는 적대적인 문서를 생성하는 자동화된 블랙박스 공격 프레임워크인 '기만적 진화적 공격(DEJA)'을 제안합니다. DEJA는 LLM 기반 평가기를 통해 계산된 세분화된 응답 유용성 점수(AUS)에 의해 안내되는 진화적 최적화 프로세스를 사용하여 응답의 확실성을 체계적으로 저하시키면서 높은 검색 성공률을 유지합니다. 다양한 RAG 구성 및 벤치마크 데이터 세트에 대한 광범위한 실험 결과, DEJA는 일관되게 응답을 낮은 유용성의 소프트 실패로 유도하며, SASR(성공률)이 79% 이상인 반면 하드 실패율은 15% 미만으로, 기존 공격보다 훨씬 뛰어난 성능을 보입니다. 생성된 적대적인 문서는 높은 은폐성을 가지며, 퍼플렉시티 기반 탐지를 회피하고, 쿼리 재작성을 견디며, 재타겟팅 없이도 모델 패밀리 간에 전송되어 독점 시스템에서도 작동합니다.

Original Abstract

Existing jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-of-service behaviors, which are conspicuous and easy to detect. In this work, we formalize a subtler availability threat, termed soft failure, which degrades system utility by inducing fluent and coherent yet non-informative responses rather than overt failures. We propose Deceptive Evolutionary Jamming Attack (DEJA), an automated black-box attack framework that generates adversarial documents to trigger such soft failures by exploiting safety-aligned behaviors of large language models. DEJA employs an evolutionary optimization process guided by a fine-grained Answer Utility Score (AUS), computed via an LLM-based evaluator, to systematically degrade the certainty of answers while maintaining high retrieval success. Extensive experiments across multiple RAG configurations and benchmark datasets show that DEJA consistently drives responses toward low-utility soft failures, achieving SASR above 79\% while keeping hard-failure rates below 15\%, significantly outperforming prior attacks. The resulting adversarial documents exhibit high stealth, evading perplexity-based detection and resisting query paraphrasing, and transfer across model families to proprietary systems without retargeting.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!