2607.26432v1 Jul 29, 2026 cs.CV

FAS-R1: 추론 기반 얼굴 위조 탐지를 위한 통합 다중 작업 거대 언어 모델

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

Jun Feng
Jun Feng
Citations: 52
h-index: 2
Zitong Yu
Zitong Yu
Citations: 23
h-index: 3
Yichen Shi
Yichen Shi
Citations: 111
h-index: 5
Yiru Huo
Yiru Huo
Citations: 31
h-index: 1
Hongyang Wang
Hongyang Wang
Citations: 95
h-index: 5
Hongrui Li
Hongrui Li
Citations: 7
h-index: 2

얼굴 위조 탐지(FAS)는 단순히 진본/위조 여부를 판단하는 것뿐만 아니라, 공격의 의미와 이미지 기반 증거를 제공하여 사용자의 검토를 돕는 역할이 점점 중요해지고 있습니다. 기존의 판별 모델은 주로 레이블 중심적인 방식으로 작동하며, 최근 등장한 다중 작업 언어 모델(MLLM)은 구조화된 결과를 제공하지만 여전히 감독 학습에 의존하는 경향이 있으며, 종종 획일화된 설명과 어려운 공격 유형에 대한 취약한 최적화를 보입니다. 본 논문에서는 통합적인 FAS 예측을 위한 추론 기반 MLLM 프레임워크인 FAS-R1을 제안합니다. FAS-R1은 진본/위조 판별, 공격 유형 식별 및 위조 영역 특정화의 세 가지 작업을 수행합니다. 먼저, 고품질의 긴 Chain-of-Thought(CoT) 데이터셋인 FAS-R1-23K를 사용하여 초기 감독 학습을 진행하고, 이후 FAS에 특화된 강화 학습 기반 정책 최적화(GRPO) 후속 훈련을 수행합니다. Degradation-Simulated Augmentation (DSA) 기법은 시각 품질 변화에 따른 안정적인 위조 단서 추론을 유도하며, Difficulty-Aware GRPO (DA-GRPO)는 간단한 샘플의 지배를 완화하여 메이크업 또는 마스크와 같이 미묘하거나 모호한 공격 유형과 같은 어려운 작업에 대한 최적화를 개선합니다. 제안하는 3B 파라미터 FAS-R1 모델은 동일 환경에서 98.75%의 진본/위조 판별 정확도, 93.33%의 공격 유형 식별 정확도를 달성했으며, AP@40 및 AP@50 지표에서는 각각 96.30%와 94.73%를 기록했습니다. 또한, 제안하는 모델은 다른 시스템보다 뛰어난 성능을 보이며, 특히 다양한 도메인에서의 진본/위조 판별 정확도와 답변 및 설명의 품질 측면에서 우수합니다. 다양한 기본 모델을 사용한 실험 결과에서도 확장성에 대한 긍정적인 결과를 확인했습니다. 관련 코드는 곧 공개될 예정입니다.

Original Abstract

Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but still rely mainly on supervised fine-tuning, often producing template-like rationales and weak optimization for difficult attacks. We propose FAS-R1, a two-stage reasoning-oriented MLLM framework for unified FAS prediction, covering authenticity classification, attack-type recognition and spoof-region localization. FAS-R1 first uses FAS-R1-23K, a high-quality long-CoT dataset, for cold-start supervised fine-tuning, and then performs FAS-specific GRPO post-training. Degradation-Simulated Augmentation (DSA) encourages stable spoof-cue reasoning across visual-quality shifts, while Difficulty-Aware GRPO (DA-GRPO) mitigates easy-sample dominance that may leave difficult task--attack groups under-optimized, especially for subtle or ambiguous attacks such as makeup and mask attacks. The main 3B FAS-R1 model achieves 98.75\% authenticity accuracy, 93.33\% attack-type accuracy, and 96.30/94.73\% AP@40/AP@50 in-domain. It also outperforms the compared systems in cross-domain authenticity generalization and answer-and-rationale quality. Experiments with different base models further show favorable scaling behavior. The code will be released soon.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!