Weights에서 CSAM 텍스트-이미지 LoRA 모델 탐지
Detecting CSAM Text-to-Image LoRAs From Weights
저랭크 적응(LoRA) 미세 조정은 특정 작업, 특히 아동 성 학대 자료(CSAM) 생성과 같은 작업을 위해 오픈 웨이트 이미지 생성 모델을 저렴하고 쉽게 사용자 정의할 수 있도록 합니다. 기존의 검열 방식은 메타데이터 또는 생성된 결과에 의존하지만, 메타데이터는 오해를 불러일으킬 수 있으며, 결과를 직접 생성하는 것 자체가 용납될 수 없거나 불법일 수 있습니다. 우리는 더 안전한 신호가 모델 가중치 자체에 존재한다는 것을 보여줍니다. LoRA의 업데이트에서 가장 중요한 왼쪽 고유 벡터들은 해당 모델이 학습한 가장 큰 변화를 나타내는 간결하고 추론 없이 사용할 수 있는 지문($u_1$)을 형성합니다. CSAM을 판단하는 데 사용될 수 있는 무해한 대리 변수인 인간 주체의 나이를 사용하여, $u_1$이 LoRA가 어떤 데이터로 훈련되었는지 식별하고, 다양한 기본 모델에 적용 가능하며, 관련 없는 무해한 콘텐츠에는 적용되지 않음을 확인했습니다. 이 신호는 추가적인 가중치 노이즈, 크기 조정 및 정밀도 감소에도 강합니다. 이러한 결과는 유해한 LoRA 모델을 메타데이터나 유해한 결과 생성을 의존하지 않고, 직접적으로 해당 모델의 가중치에서 검사할 수 있음을 시사합니다.
Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's updates form a compact, inference-free fingerprint ($u_1$) of its strongest learned change. Using human-subject age as a benign proxy for CSAM, we find that $u_1$ identifies what a LoRA was trained on, generalizes across base models, and abstains on unrelated benign content. The signal is robust to additive weight noise, rescaling, and precision reduction. These results indicate that harmful LoRAs could be screened directly from their weights without relying on metadata or generating harmful outputs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.