2603.01038v1 Mar 01, 2026 cs.CV

직관에서 조사로: 일반화된 얼굴 반가짜 탐지를 위한 도구 기반 추론 MLLM 프레임워크

From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing

Haoyuan Zhang
Haoyuan Zhang
Citations: 32
h-index: 2
Keyao Wang
Keyao Wang
Citations: 157
h-index: 6
Haixiao Yue
Haixiao Yue
Citations: 252
h-index: 8
Xiao Tan
Xiao Tan
Citations: 173
h-index: 7
Wei He
Wei He
Citations: 10
h-index: 2
Jingdong Wang
Jingdong Wang
Citations: 19
h-index: 2
Ajian Liu
Ajian Liu
Citations: 49
h-index: 2
Xiangyu Zhu
Xiangyu Zhu
Citations: 7,984
h-index: 33
Zhen Lei
Zhen Lei
Citations: 214
h-index: 8
Guosheng Zhang
Guosheng Zhang
Citations: 135
h-index: 7
Zhiwen Tan
Zhiwen Tan
Citations: 38
h-index: 2
Siran Peng
Siran Peng
Citations: 90
h-index: 4
Tianshuo Zhang
Tianshuo Zhang
Citations: 261
h-index: 6
Kunbin Chen
Kunbin Chen
Citations: 16
h-index: 2

얼굴 인식 기술은 여전히 위조 공격에 취약하며, 이를 해결하기 위한 강력한 얼굴 반가짜 탐지(FAS) 솔루션이 필요합니다. 최근 MLLM 기반 FAS 방법은 이진 분류 작업을 간략한 텍스트 설명 생성으로 재구성하여 도메인 간 일반화 성능을 향상시킵니다. 그러나 이러한 방법의 일반화 능력은 여전히 제한적이며, 텍스트 설명이 주로 직관적인 의미적 단서(예: 마스크 윤곽선)를 포착하는 반면, 미세한 시각적 패턴을 인식하는 데 어려움을 겪습니다. 이러한 제한점을 해결하기 위해, 우리는 MLLM에 외부 시각 도구를 통합하여 미묘한 위조 단서를 더 깊이 있게 조사하도록 유도합니다. 구체적으로, 우리는 도구 기반 추론 FAS(TAR-FAS) 프레임워크를 제안하며, 이는 FAS 작업을 시각 도구를 활용한 체인 오브 소트(CoT-VT) 방식으로 재구성하여, MLLM이 직관적인 관찰부터 시작하여 외부 시각 도구를 적응적으로 호출하여 미세한 조사를 수행할 수 있도록 합니다. 이를 위해, 우리는 도구 기반 데이터 어노테이션 파이프라인을 설계하고, 다단계 도구 사용 추론 경로를 포함하는 ToolFAS-16K 데이터셋을 구축했습니다. 또한, 모델이 효율적인 도구 사용을 자율적으로 학습할 수 있도록 다양한 도구 그룹 상대 정책 최적화(DT-GRPO)를 활용한 도구 인식 FAS 학습 파이프라인을 도입했습니다. 어려운 one-to-eleven 도메인 간 프로토콜에서 수행된 광범위한 실험 결과, TAR-FAS는 최첨단 성능을 달성하며, 신뢰할 수 있는 위조 탐지를 위한 미세한 시각적 조사를 제공합니다.

Original Abstract

Face recognition remains vulnerable to presentation attacks, calling for robust Face Anti-Spoofing (FAS) solutions. Recent MLLM-based FAS methods reformulate the binary classification task as the generation of brief textual descriptions to improve cross-domain generalization. However, their generalizability is still limited, as such descriptions mainly capture intuitive semantic cues (e.g., mask contours) while struggling to perceive fine-grained visual patterns. To address this limitation, we incorporate external visual tools into MLLMs to encourage deeper investigation of subtle spoof clues. Specifically, we propose the Tool-Augmented Reasoning FAS (TAR-FAS) framework, which reformulates the FAS task as a Chain-of-Thought with Visual Tools (CoT-VT) paradigm, allowing MLLMs to begin with intuitive observations and adaptively invoke external visual tools for fine-grained investigation. To this end, we design a tool-augmented data annotation pipeline and construct the ToolFAS-16K dataset, which contains multi-turn tool-use reasoning trajectories. Furthermore, we introduce a tool-aware FAS training pipeline, where Diverse-Tool Group Relative Policy Optimization (DT-GRPO) enables the model to autonomously learn efficient tool use. Extensive experiments under a challenging one-to-eleven cross-domain protocol demonstrate that TAR-FAS achieves SOTA performance while providing fine-grained visual investigation for trustworthy spoof detection.

1 Citations
0 Influential
16.5 Altmetric
83.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!