IRIS: 확산 이미지의 위조 방지 워터마킹을 위한 시각-의미 연결
IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Images
대부분의 생성 과정에서 사용되는 워터마크는 해당 이미지를 포함하지 않는 독립적인 패턴을 내장하며, 공격자는 이러한 패턴을 생성기가 만들지 않은 이미지에 이식하여 위조를 발생시킵니다. 워터마크를 시각적 의미와 연결하면 이러한 이식을 방지할 수 있지만, 기존 방식은 워터마크가 대상 이미지가 아닌 프록시 이미지를 기준으로 합니다. 생성 과정 내에서 시각-의미 연결을 구현하는 것은 두 가지 과제를 안고 있습니다. 워터마크는 이미지 자체에서 파생되지만 해당 이미지가 존재하기 전에 샘플링 경로에 입력되며, 이는 워터마크 자체가 연결되는 의미를 변경할 수 있습니다. 또한, 워터마크는 상반된 민감도 요구 사항을 충족해야 합니다. 즉, 의미 변화에는 깨지기 쉬워야 하지만 일반적인 처리 과정에서는 유지되어야 합니다. 본 논문에서는 훈련이 필요 없는 워터마킹 기법인 IRIS를 제시합니다. IRIS는 이미지의 고유한 의미에서 파생된 '내재적 링 식별자'를 삽입합니다. IRIS는 워터마크가 없는 생성된 이미지로부터 콘텐츠 코드를 읽고, 해당 코드와 비밀 키를 사용하여 일회성 링을 생성하며, 생성 과정의 최종 단계로 돌아가서 이미지가 가진 의미가 확정된 후 해당 링을 병합합니다. 상반된 민감도 요구 사항을 충족하기 위해, 콘텐츠 코드는 삽입 및 검출 과정에서 공유되는 정규화 방식을 통해 읽히므로 일반적인 왜곡 및 경미한 재생에도 견고하며, 동시에 의미 변화에 의해 깨지도록 설계되었습니다. 검출은 쿼리 이미지와 키만으로 링을 재계산하며, 따라서 워터마크는 외부 이미지나 조작된 이미지에서는 실패합니다. 또한, 검출 과정은 워터마크의 허용 여부를 결정할 때 의미의 변화를 추적합니다. 세 가지 프롬프트 데이터셋에 대한 실험 결과, IRIS는 안정적으로 동작하며 동일한 시드 값으로 생성된 워터마크 없는 이미지와 유사한 품질을 유지합니다. 이는 기존 생성 과정에서 사용되는 워터마크가 달성하지 못하는 수준의 충실도입니다. 위조 공격은 고정 패턴 워터마크를 이식하고, 재생은 사후 삽입 워터마크를 제거하지만, 비교된 다른 워터마킹 방식들과 달리 IRIS는 이러한 모든 유형의 공격에 대한 저항성을 보입니다.
Most in-generation diffusion watermarks embed patterns independent of the image that carries them, and attackers transplant the marks onto images the generator did not produce, resulting in forgery. Binding the mark to visual semantics prevents such transplantation, yet existing bindings anchor to a proxy image rather than the image they mark. Realizing visual-semantic binding inside generation faces two challenges. The mark derives from the image itself yet enters the sampling trajectory before that image exists, and may itself shift the semantics it binds. The binding also meets opposite sensitivity demands, breaking under semantic change while holding through common processing. We present IRIS, a training-free watermarking scheme that embeds an Intrinsic Ring Identifier from Semantics. IRIS reads a content code from the non-watermarked generated image, derives a one-time ring from the code and a secret key, returns to the final low-noise steps of the same trajectory and blends the ring in, after the semantics it binds are settled. To meet the opposite sensitivity demands, the code is read through a canonicalization shared between embedding and detection, holding through common distortions and mild regeneration while flipping under semantic change. Detection recomputes the ring from the query image and the key alone, and the mark therefore fails on a foreign or spliced image, with acceptance tracking semantic displacement. On three prompt datasets IRIS detects reliably and stays close to its same-seed non-watermarked counterpart, a fidelity prior in-generation marks do not reach. While forgeries transfer fixed-pattern marks and regeneration strips post-hoc marks, IRIS alone among the compared marks withstands both.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.