SEED: 간단한 ViT 모델과 진화적 프레임워크를 활용한 설명 가능한 텍스트 위조 탐지
SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection
AI 기반 이미지 편집 기술은 금융, 법률 및 신원 정보 기록에 대한 신뢰를 위협합니다. ACM MM 2026의 GenText-Forensics Challenge는 이러한 문제를 해결하기 위해 구조화된 포렌식 보고서를 요구하며, 다국어 중심의 위조 이미지를 대상으로 탐지, 픽셀 단위 위치 파악 및 자연어 설명을 통합해야 합니다. 본 논문에서는 SEED라는 모듈형 시스템을 제안합니다. 첫째, 유사성 기반 파이프라인은 다양한 합성 위조 데이터를 활용하여 학습을 강화합니다. 둘째, DINOv3를 기반으로 LoRA 적응을 통해 구축된 단일 ViT 모델은 탐지 및 픽셀 단위 위치 파악을 동시에 수행하며, 사전 학습된 정보를 유지하면서도 최소한의 학습 가능한 파라미터로 작동합니다. 셋째, 진화적 프레임워크는 탐지기의 예측 결과를 활용하여 MLLM(대규모 언어 모델)을 통해 완전한 포렌식 보고서를 생성하며, 제안-평가 루프를 통해 보고서 품질을 반복적으로 개선합니다. SEED는 GenText-Forensics Challenge에서 3위를 차지했습니다. 코드 및 데이터는 https://github.com/KahimWong/GenText-Forensics-3rd-Place 에서 확인할 수 있습니다.
AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by requiring structured forensic reports, in which integrating detection, pixel-level localization, and natural language explanation for multilingual text-centric forgery images. We present SEED, a modular system with three components. First, a similarity-guided pipeline augments training with diverse synthetic forgeries. Second, a single ViT, built on DINOv3 with LoRA adaptation, jointly performs detection and pixel-level localization while preserving pre-trained priors with minimal trainable parameters. Third, an evolving harness takes the detector's predictions and generates a complete forensic report via an MLLM, iteratively improved through a proposer-evaluator loop optimizing report quality. SEED ranked 3rd in the GenText-Forensics Challenge. Code and data are available at https://github.com/KahimWong/GenText-Forensics-3rd-Place.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.