SAMark: 문단 수준의 패러프레이징에 강건한 자체 고정 방식 텍스트 워터마킹
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
시맨틱 레벨 워터마킹(SWM)은 문장을 기본 단위로 사용하여 텍스트 수정에 대한 강건성을 향상시키지만, 문단 수준의 패러프레이징 공격에 대한 강건성은 여전히 어렵습니다. 이러한 공격은 문장 순서를 변경하여 워터마크 신호를 전반적으로 방해하기 때문입니다. 본 연구에서는 문장 순서에 대한 의존성을 제거하고 의미 공간에서 독립적인 '안전 영역'을 구축하는 자체 고정 방식 워터마킹 프레임워크인 SAMark를 제안합니다. 탐지 가능성을 향상시키기 위해, 약하게 정렬된 후보에서 발생하는 노이즈를 억제하면서 워터마크 신호를 증폭시키는 다중 채널 하이퍼볼릭 스코어링 메커니즘을 도입했습니다. 또한, 단순한 n-gram 반복 필터를 넘어 의미 중복 문제를 해결하기 위해 강력한 필터링과 소프트 정규화를 결합하는 다양성 기반 필터링 전략을 제안합니다. 실험 결과는 SAMark가 일반적인 문단 수준의 패러프레이징 공격 하에서 최대 90.2%의 TP@FP1%를 달성하며, 기존 최고 성능 모델보다 평균 30% 이상 우수한 성능을 보인다는 것을 보여줍니다. 또한, SAMark는 워터마크가 적용되지 않은 텍스트와 경쟁력 있는 생성 품질을 유지하면서 기존 방법이 가진 강건성과 품질 간의 균형 문제를 해결합니다.
Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.