2605.27348v1 May 26, 2026 cs.CV

AI가 오판할 때: 사회적 시선 일관성이 AI 생성 이미지 탐지를 위한 의미론적 단서로 작용하는 경우

When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

J. Rehg
J. Rehg
Citations: 1,468
h-index: 21
Hyesong Choi
Hyesong Choi
Citations: 202
h-index: 7
Jihyeon Kim
Jihyeon Kim
Citations: 218
h-index: 5
Sohee Kim
Sohee Kim
Citations: 28
h-index: 2
Soosang Lee
Soosang Lee
Citations: 84
h-index: 5
Souhwan Jung
Souhwan Jung
Citations: 72
h-index: 5

최근의 생성 모델은 낮은 수준의 특징(픽셀 지문, 주파수 이상 현상, 업샘플링 흔적 등)을 상당 부분 개선했으며, 특히 조작된 영역이 작고 주변에 실제와 유사한 내용이 있는 인물 중심 및 부분 편집 환경에서 이러한 경향이 두드러집니다. 본 연구에서는 사회적 시선 일관성이라는 고수준의 의미론적 단서를 제시합니다. 이는 상호작용하는 개인 간의 시선 방향, 머리와 눈의 정렬, 동공 위치 사이의 상호 연관성을 의미하며, 기존의 낮은 수준의 방식과는 다른 새로운 탐지 방법을 제공합니다. 이러한 아이디어를 세 가지 주요 메커니즘을 통해 구현했습니다: (i) 특정 영역에 대한 시선 일관성 이미지의 제어된 변형을 포함하는 진단 데이터셋으로, 엄격한 쌍 기반 그룹화를 통해 학습 과정에서 생성 모델의 특징을 암기하는 것을 방지하고 증강 기술에 의존하지 않도록 합니다; (ii) 블록 구성 캡션 감독(Block-Compositional Caption Supervision), 이는 1,250개의 거시 수준으로 결합된 캡션을 통틀어 일관된 5개 블록의 추론 구조를 유지하여 추론의 일관성을 표면적인 다양성과 분리합니다; (iii) 다양한 아키텍처에 대한 검증 결과, 제안하는 감독 방식이 비전-언어 모델(FakeVLM)의 COCOAI Interaction 데이터셋에서 정확도가 3.7%p 향상(67.8 -> 71.5), COCOAI Person 데이터셋에서 1.3%p 향상(83.0 -> 84.3)되는 것을 확인했으며, 비전만 사용하는 모델(Effort)에서도 일관된 성능 향상을 보였습니다. 이는 제안하는 방법이 특정 모델에 의존적이지 않음을 보여줍니다. 실제 이미지와 가짜 이미지의 탐지율 모두 동시에 상승하여 모든 이미지를 가짜로 예측하는 현상은 나타나지 않습니다. 네 단계로 구성된 메커니즘 분석(쌍 편집 단축 방지, 쉬운 문제에서 어려운 문제로의 난이도 이전, CLIP 사전 지식 유지, 디퓨전 모델 계열의 주변 구조에 대한 공통적인 약점)은 단일 인페인터(FLUX.1-Fill)로 학습한 결과가 여러 생성 모델 환경에도 적용될 수 있는 이유를 설명합니다. 본 연구의 코드는 논문 게재 후 공개하여 재현성을 높이도록 하겠습니다.

Original Abstract

Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularly in person-centric and partial-edit settings where the manipulated region is small and surrounded by photometrically authentic content. We introduce Social Gaze Consistency, a high-level semantic cue defined as the mutual coherence of gaze direction, head-eye alignment, and pupil placement between interacting individuals, and show that it constitutes a previously underutilized detection axis orthogonal to existing low-level paradigms. We instantiate this insight through three coupled mechanisms: (i) a controlled diagnostic dataset with region-specific perturbations of gaze-consistent imagery, where strict pair-level grouping forecloses generator-fingerprint memorization as an optimization-time shortcut rather than relying on augmentation; (ii) Block-Compositional Caption Supervision, which holds a single 5-block reasoning skeleton invariant across 1,250 macro-combined captions, decoupling reasoning consistency from surface diversity; (iii) Cross-architecture validation showing the same supervision improves a vision-language backbone (FakeVLM) by +3.7 pp on the COCOAI Interaction subset (balanced accuracy 67.8 -> 71.5) and +1.3 pp on the COCOAI Person subset (83.0 -> 84.3), with consistent gains on a vision-only backbone (Effort), evidencing a backbone-agnostic cue. Real- and fake-class recalls rise simultaneously, ruling out a "predict-all-fake" artifact. A four-step mechanistic account - paired-edit shortcut blocking, hard-to-easy difficulty transfer, CLIP prior preservation, and diffusion-family shared spectral weakness in periocular structure - explains why training on a single inpainter (FLUX.1-Fill) transfers to multi-generator suites. We will release the code upon acceptance to facilitate reproducibility.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!