2607.28187v1 Jul 30, 2026 cs.AI

오래된 기술, 새로운 모델: 간단한 이미지 변환이 최신 AI 기반 콘텐츠 검열 시스템을 어떻게 무력화하는가

Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation

Iyiola E. Olatunji
Iyiola E. Olatunji
Citations: 474
h-index: 10
Tegawendé F. Bissyandé
Tegawendé F. Bissyandé
Citations: 2
h-index: 1
Jacques Klein
Jacques Klein
Citations: 50
h-index: 4
Marco Alecci
Marco Alecci
Citations: 77
h-index: 6
Francesco Marchiori
Francesco Marchiori
Citations: 180
h-index: 8

자동 콘텐츠 검열 시스템은 대규모 유해 콘텐츠 필터링에 필수적이지만, 기존의 특정 작업 전용 분류기는 종종 제한적인 정책 적용 범위와 맥락 이해를 제공합니다. 최근에는 대규모 기반 모델을 활용하여 더 넓고 강력한 안전 기능을 제공할 것으로 기대되는 상업용 멀티모달 검열 API가 출시되었습니다. 본 연구에서는 이러한 변화가 이미지 검열의 견고성 향상으로 이어지는지 분석합니다. 우리는 세 가지 주요 상업용 이미지 검열 서비스에 대한 대규모 블랙박스 평가를 수행하고, 이들의 견고성을 비교했습니다. 모델에 독립적인 7가지 간단한 이미지 변환을 다양한 제공 업체, 데이터 세트, 유해 콘텐츠 유형, 인식적 유사성 제약 조건 및 변환 강도를 기준으로 평가한 결과, 다음과 같은 사실을 확인했습니다: (1) 세 가지 상업용 서비스 모두 저렴하고 간단한 이미지 변환을 통해 우회될 수 있으며, 이는 기울기 계산, 대체 모델 또는 대상 시스템에 대한 지식을 필요로 하지 않습니다. (2) 색상 반전 및 회색조 변환과 같은 고정된 변환이라도 안전하지 않은 콘텐츠를 안전한 콘텐츠로 변경시키면서 인간이 여전히 인지할 수 있는 수준의 내용을 유지합니다. (3) 이들의 견고성은 데이터 세트 및 유해 콘텐츠 유형에 따라 크게 다르며, 특히 멀티모달 콘텐츠와 자해 관련 콘텐츠에서 취약성이 두드러집니다. 이러한 결과는 기존 검열 분류기를 기반 모델 기반 API로 대체하는 것만으로는 신뢰할 수 있는 보안 경계를 제공하지 않는다는 결론을 내릴 수 있게 합니다. 이러한 시스템은 실제 변환 환경에서 평가되어야 하며, 독립적인 안전 필터가 아닌 다층 검열 파이프라인의 구성 요소로서 배포되어야 합니다.

Original Abstract

While automated content-moderation systems have become essential for screening harmful content at scale, conventional task-specific classifiers often provide limited policy cov- erage and contextual understanding. Recently, commercial multimodal moderation APIs built on large foundation models have been introduced with the promise of providing broader and more capable safety filters. In this work, we analyze whether this shift also yields more robust image moderation. We conduct a large-scale black-box evaluation on three established commercial image-moderation services and compare their robustness. By evaluating seven simple, model-agnostic image transformations across multiple providers, datasets, harm categories, perceptual-similarity constraints, and transformation intensities, we find that: (1) all three commercial services can be bypassed using inexpensive image transformations that require no gradients, surrogate models, or knowledge of the target system; (2) even fixed transformations such as color inversion and grayscale conversion induce unsafe-to-safe decision changes while preserving content that remains recognizable to humans; (3) their robustness varies substantially across datasets and harm categories, with multimodal content and self-harm exhibiting pronounced vulnerabilities. This yields the conclusion that replacing conventional moderation classifiers with foundation-model-based APIs does not, by itself, provide a reliable security boundary. Such systems must be evaluated under realistic transformations and deployed as one component of a layered moderation pipeline rather than as standalone safety filters.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!