2607.24016v2 Jul 27, 2026 cs.CV

DailyBench: 현대적인 생성 모델에서 생성 및 조작된 이미지에 대한 통합 벤치마크

DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

Dongming Zhang
Dongming Zhang
Citations: 49
h-index: 3
Junyao Gao
Junyao Gao
Citations: 348
h-index: 7
Xin Jiang
Xin Jiang
Citations: 306
h-index: 5
Hao Tang
Hao Tang
Citations: 194
h-index: 2
Meiqi Cao
Meiqi Cao
Citations: 24
h-index: 3
Fei Shen
Fei Shen
Citations: 28
h-index: 2
Yongdong Zhang
Yongdong Zhang
Citations: 0
h-index: 0

최근 생성 모델의 발전으로 인해 AI가 생성한 이미지 감지는 쉽게 구별되는 완전 합성 이미지를 식별하는 것에서, 최신 생성 및 조작 파이프라인을 통해 생성된 매우 현실적인 콘텐츠를 식별하는 것으로 전환되었습니다. 그러나 기존의 감지 벤치마크는 종종 오래된 생성 모델로 구축되며, 주로 전체 이미지 합성에 중점을 두기 때문에 벤치마크 데이터와 실제 생성 및 편집 시나리오에서 접하는 이미지 간의 불일치가 점점 커지고 있습니다. 이러한 격차를 해소하기 위해, 우리는 AI가 생성한 이미지 감지기가 최신 형태의 전체 이미지 합성 및 객체 수준 조작에 걸쳐 일반화할 수 있는지 평가하기 위한 고품질 통합 벤치마크인 DailyBench를 소개합니다. DailyBench는 FakeBench와 ManipulationBench라는 두 가지 상호 보완적인 하위 집합으로 구성됩니다. FakeBench는 최신 오픈 소스 및 상용 생성 모델을 통해 합성된 고품질 이미지를 포함하며, ManipulationBench는 고급 이미지 조건부 모델을 사용하여 실제 이미지에 적용된 어려운 객체 수준 편집을 소개합니다. 이러한 설계로 인해 DailyBench는 생성기 수준의 일반화와 미묘한 로컬 편집 하에서의 조작 인지 감지를 연구하기 위한 현실적인 테스트 환경이 됩니다. DailyBench에 대한 실험 결과, 현재 감지기의 상당한 안정성 격차가 드러났습니다. GenImage에서 91-96%의 균형 잡힌 정확도를 보고하는 방법은 FakeBench에서는 60-76%, ManipulationBench에서는 54-66%로 감소합니다. 이러한 결과는 기존 감지기가 현실적인 합성 및 조작에 여전히 제대로 일반화되지 않았음을 보여주며, DailyBench가 견고하고 조작 인지 AI 생성 이미지 감지 방법을 개발하기 위한 엄격한 테스트 환경임을 강조합니다. 프로젝트는 https://dailybench.github.io/ 에서 확인할 수 있습니다.

Original Abstract

Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91-96% balanced accuracy on GenImage drop to 60-76% on FakeBench and 54-66% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods. The project is available at https://dailybench.github.io/

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!