Veritas++: 인지 능력 향상을 위한 가치 기반 온라인 증류를 통한 AI 생성 이미지 탐지
Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection
이미지 생성 모델의 발전으로 인해 합성 이미지가 공개 매체에서 흔하게 등장하면서, 강력하고 일반화된 AI 생성 이미지(AIGI) 탐지의 중요성이 더욱 커지고 있습니다. 멀티모달 대규모 언어 모델(MLLM)은 블랙박스 이진 판별 방식에 대한 투명한 대안을 제공하지만, 현재의 MLLM 기반 탐지기는 여전히 미세한 이상 징후를 감지하는 데 상당한 인지적 한계를 보입니다. 이러한 모델들은 주로 시각적 증거가 어떻게 구성되고 합성되는지에 초점을 맞추며, 근본적인 인지 능력은 충분히 최적화되지 않은 경우가 많습니다. 이러한 격차를 해소하기 위해, 우리는 신뢰할 수 있는 인지를 진위 판단의 기반으로 하는 인지 능력 향상 추론 프레임워크인 Veritas++를 제안합니다. 모델의 설명 능력을 직접적으로 최적화하는 대신, AIGI 탐지를 세 가지 기본적인 인지 능력, 즉 미세한 시각적 디테일 캡처, 의미적 이상 및 픽셀 수준의 차이점을 감지하는 데 기반을 두고 있습니다. 이러한 통찰력을 바탕으로, 우리는 개방형 설명 감독을 검증 가능한 보상으로 대체하여 이러한 능력을 명시적으로 강화하는 인지 중심 학습(PoRL) 방식을 도입했습니다. 또한, 향상된 인지를 추론과 통합하기 위해, 가치 기반 온라인 증류(VaOPD)라는 적응형 증류 메커니즘을 소개합니다. VaOPD는 균일한 감독 대신 고가치의 증류 신호에 우선순위를 부여하여, 특권적인 자기 지도자를 통해 인지 능력을 고려한 추론을 내재화합니다. 표준 벤치마크, 실제 환경 데이터 및 새로운 벤치마크를 대상으로 실시한 광범위한 실험 결과, Veritas++는 뛰어난 일반화 성능을 보여줍니다. 인지 학습은 인지적 격차를 효과적으로 해소하고 탐지 성능을 향상시키며, VaOPD는 기존 성능을 희생하지 않고 효율적인 능력 발전을 가능하게 합니다. 코드 및 체크포인트는 https://github.com/EricTan7/VeritasPP 에서 이용하실 수 있습니다.
The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generated Image (AIGI) detection increasingly essential. While multi-modal large language models (MLLMs) offer a transparent alternative to black-box binary scoring, we observe that current MLLM-based detectors still exhibit notable perception bottlenecks in capturing fine-grained anomalies. They primarily focus on how visual evidence is organized and synthesized, leaving the intrinsic perception less optimized. To mitigate this gap, we present Veritas++, a perception-enhanced reasoning framework that establishes reliable perception as the foundation of authenticity reasoning. Rather than directly optimizing the model's explanatory ability, we ground AIGI detection on three basic perception abilities, i.e., capturing fine-grained visual details, semantic anomalies and pixel-level differences. Building on this insight, we introduce Perception-oriented Learning (PoRL), which replaces open-ended description supervision with verifiable rewards to explicitly strengthen these capacities. To further integrate enhanced perception with reasoning, we introduce Value-aware On-Policy Distillation (VaOPD), an adaptive distillation mechanism that prioritizes high-value distillation signals over uniform supervision, internalizing perception-aware reasoning through a privileged self-teacher. Extensive experiments across standard, in-the-wild and emerging benchmarks demonstrate that Veritas++ achieves promising generalization. The perception learning effectively bridges the perception gap and yields seamless gains on detection, while VaOPD further enables efficient capability evolvement without sacrificing existing performance. Code and checkpoints are available at https://github.com/EricTan7/VeritasPP.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.