2605.28632v1 May 27, 2026 cs.CR

블라인드 PRNG 하이재킹: LLM 워터마킹에 대한 검증 가능한 무결성을 유지하는 탐지 불가능한 공격

Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking

Ziyang You
Ziyang You
Citations: 6
h-index: 1
Huilong He
Huilong He
Citations: 0
h-index: 0
Xiaoke Yang
Xiaoke Yang
Citations: 7
h-index: 1
Xuxing Lu
Xuxing Lu
Citations: 1
h-index: 1

암호화된 워터마킹은 대규모 언어 모델(LLM)이 생성한 텍스트의 출처를 밝히는 데 있어 핵심적인 방어 기술입니다. 기존 방식인 KGW, Unigram 및 DipMark를 포함한 대부분의 워터마킹 시스템은 보안성을 보장하기 위해 기본적으로 사용되는 의사 난수 생성기(PRNG)가 신뢰할 수 있다는 가정에 기반합니다. 본 연구에서는 LLM 워터마킹에 대한 최초의 공급망 공격인 SeedHijack을 소개합니다. SeedHijack은 다음과 같은 특징을 갖습니다: (i) 블라인드 방식 – 워터마크 키, 검출기 또는 모델 로짓에 대한 어떠한 정보도 필요하지 않음, (ii) 무결성 유지 – 워터마크 신호를 제거하는 대신 증폭시킴, (iii) 검출과 독립적 – 공격으로 인한 편향은 콘텐츠 측면의 모든 검출 통계와 통계적으로 독립적이므로, 증폭과 회피가 절충 없이 공존함. SeedHijack은 생성된 텍스트를 변경하는 대신, 공급망 계층에서 PRNG를 교체하여, 출력 토큰을 변경하거나 텍스트 품질을 저하시키지 않으면서 화이트리스트 선택에 편향을 만듭니다. 세 가지 워터마킹 방식과 세 개의 오픈 소스 LLM을 대상으로 실험한 결과, SeedHijack 공격은 최첨단 콘텐츠 측면 통계 검출기 6개 모두를 탐지하지 못했으며, 동시에 워터마크 Z-점수를 최대 2.42배까지 증가시켰습니다 (엔트로피 소스 증명과 같은 시스템 수준 방어는 여전히 독립적이며 상호 보완적임). 양자 난수 생성기(QRNG)에 의한 대응책은 공격을 완전히 무력화하면서도 정상적인 워터마킹 기능을 유지하는 것으로 나타났습니다. 이러한 결과는 PRNG의 무결성을 암호화된 콘텐츠 출처 추적 시스템의 중요한 보안 요구 사항으로 확립합니다.

Original Abstract

Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigram, and DipMark, derive their security guarantees from the assumption that the underlying pseudo-random number generator (PRNG) is trustworthy. This work introduces SeedHijack, the first supply-chain attack on LLM watermarking that is simultaneously (i) blind -- requiring no knowledge of the watermark key, detector, or model logits, (ii) integrity-preserving -- amplifying rather than erasing the watermark signal, and (iii) orthogonal to detection -- the attack-induced bias is statistically independent of all content-side detector statistics, ensuring that amplification and evasion coexist without trade-off. Rather than perturbing generated text, SeedHijack replaces the PRNG at the supply-chain layer, biasing green-list selection without altering output tokens or degrading text quality. Across three watermarking schemes and three open-source LLMs, the attack triggers 0/6 state-of-the-art content-side statistical detectors while inflating the watermark z-score up to 2.42x (system-level defenses such as entropy-source attestation remain orthogonal and complementary). A quantum random number generator (QRNG) countermeasure is shown to fully neutralize the attack while preserving benign watermarking utility. These findings establish PRNG integrity as a first-class security requirement for cryptographic content-provenance systems.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!