2606.26552v1 Jun 25, 2026 cs.CV

인식, 판단 및 진화: 후향적 자기 개선 기반의 인공지능 생성 이미지 탐지 시스템

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

Jingren Zhou
Jingren Zhou
Citations: 26,959
h-index: 29
Zhou Zhao
Zhou Zhao
Citations: 10
h-index: 2
Fei Wu
Fei Wu
Citations: 8
h-index: 2
Yang Wu
Yang Wu
Citations: 0
h-index: 0
Keyu Yan
Keyu Yan
Citations: 1,880
h-index: 4
Yu Liu
Yu Liu
Citations: 198
h-index: 4
Fei Huang
Fei Huang
Citations: 417
h-index: 4
Rong Zhang
Rong Zhang
Citations: 0
h-index: 0

생성 모델의 급속한 발전은 기존의 딥페이크 탐지 방법론에 상당한 어려움을 야기하며, 특히 매우 사실적인 인공지능이 생성한 이미지의 광범위한 확산은 더욱 심각한 문제입니다. 다중 모드 대규모 언어 모델(MLLM)이 이 분야에서 강력한 잠재력을 보여주지만, 기존 접근 방식은 미세한 법의학적 단서에 대한 민감성 부족과 최첨단 모델로부터 얻은 정적인 합성 감독 데이터 의존이라는 두 가지 주요 한계를 가지고 있으며, 이는 유연성과 높은 비용으로 이어집니다. 이러한 문제점을 해결하기 위해, 우리는 반복적인 자기 진화를 통해 인공지능 생성 이미지 탐지를 위한 에이전트 기반 법의학 프레임워크인 ForeAgent를 제안합니다. 첫째, ForeAgent는 의미, 공간 및 주파수 영역 특징을 포괄하는 다중 시점 정보를 통합하고, MLLM을 판단 모듈로 활용하여 이러한 신호들을 융합하여 논리적으로 타당한 판단을 내리는 인식-판단 아키텍처를 채택합니다. 둘째, 지속적인 자기 개선을 가능하게 하기 위해, 우리는 샘플링-반성-진화 패러다임을 따르는 후향적 자기 개선 전략을 도입했습니다. 에이전트는 훈련 데이터에 대한 추론 과정을 수행하고, 정답 레이블을 후향 정보로 활용하여 실패 사례 및 품질이 낮은 추론 경로를 분석하고, 더 높은 품질의 추론 과정을 재생성합니다. 이러한 합성 샘플은 엄격하게 필터링된 이중 전문가 품질 게이트 모듈을 통해 처리됩니다. ForeAgent는 자체적으로 선별한 고품질 샘플에 대한 미세 조정을 통해 지속적으로 진화합니다. 광범위한 실험 결과, ForeAgent는 Chameleon 벤치마크에서 최고 성능을 달성하여 정확도가 82.18%로 AIDE보다 16.41% 향상되었으며, 16개의 생성 모델에 대한 AIGCDetect-Benchmark에서 평균 정확도가 93.3%를 달성했습니다. 또한, 외부 평가 결과 ForeAgent가 GPT-5 및 GPT-5-mini와 비교하여 더 일관되고 인과적으로 타당한 추론 결과를 제공하는 것으로 나타났습니다.

Original Abstract

The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images. Although Multimodal Large Language Models (MLLMs) show strong potential for this task, existing approaches suffer from two key limitations: insufficient sensitivity to fine-grained forensic artifacts and reliance on static synthetic supervision from frontier models, leading to limited flexibility and high-cost. To address these issues, we propose ForeAgent, an agentic forensics framework for AI-generated image detection with iterative self-evolution. First, ForeAgent adopts a Perception-Verdict architecture that aggregates multi-view cues spanning semantic, spatial, and frequency-domain features, and leverages an MLLM as a verdict module to fuse these signals for a logical-grounded verdict. Second, to enable continual self-improvement, we introduce a Hindsight-Driven Self-Refining strategy following a Sampling-Reflection-Evolution paradigm. The agent performs inference rollouts on training instances. Guided by ground-truth labels as hindsight, it reflects on failure cases and low-quality reasoning trajectories to regenerate higher-quality reasoning traces. These synthesized samples are then strictly filtered through a dual-expert quality gating module. ForeAgent continuously evolves via fine-tuning on self-curated high-quality samples. Extensive experiments demonstrate that ForeAgent achieves state-of-the-art performance on the Chameleon benchmark, reaching 82.18% accuracy (+16.41% over AIDE), and achieves 93.3% mean accuracy on AIGCDetect-Benchmark across 16 generators. In addition, external evaluation shows that ForeAgent produces more consistent and causally grounded reasoning compared to GPT-5 and GPT-5-mini.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!