2602.10687v2 Feb 11, 2026 cs.CV

OmniVL-Guard: 균형 잡힌 강화 학습을 통한 통합된 시각-언어 위조 탐지 및 근거 제시

OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL

Jinjie Shen
Jinjie Shen
Citations: 8
h-index: 2
Jing Wu
Jing Wu
Citations: 177
h-index: 5
Yaxiong Wang
Yaxiong Wang
Citations: 86
h-index: 6
Lechao Cheng
Lechao Cheng
Citations: 60
h-index: 4
Shengeng Tang
Shengeng Tang
Citations: 5
h-index: 2
Tianrui Hui
Tianrui Hui
Citations: 7
h-index: 2
Zhun Zhong
Zhun Zhong
Citations: 104
h-index: 5
Nan Pu
Nan Pu
Citations: 198
h-index: 7

기존의 위조 탐지 방법은 종종 단일 모드 또는 이중 모드 환경에 제한되어 있으며, 실제 오정보에서 흔히 나타나는 텍스트, 이미지 및 비디오가 혼합된 형태를 처리하는 데 어려움을 겪습니다. 이러한 격차를 해소하기 위해, 본 논문에서는 통합된 시각-언어 위조 탐지 및 근거 제시 프레임워크를 개발하는 것을 목표로 합니다. 이러한 통합 환경에서, 다양한 모드 간의 상호 작용과 동시에 탐지 및 위치 지정을 요구하는 이중적인 요구 사항은 중요한 "어려움 편향(difficulty bias)" 문제를 야기합니다. 즉, 비교적 간단한 진실 여부 분류 작업이 그래디언트에 지배적인 영향을 미쳐, 다중 작업 최적화 과정에서 세밀한 근거 제시 성능이 저하될 수 있습니다. 이러한 과제를 해결하기 위해, 우리는 통합된 시각-언어 위조 탐지 및 근거 제시를 위한 균형 잡힌 강화 학습 프레임워크인 extbf{OmniVL-Guard}를 제안합니다. 특히, OmniVL-Guard는 두 가지 핵심 구성 요소로 이루어져 있습니다. 첫째, {자기 진화적 추론 경로 생성(Self-Evolving CoT Generation)}은 고품질의 추론 경로를 생성하여 초기 학습 문제를 효과적으로 해결합니다. 둘째, {적응형 보상 스케일링 정책 최적화(Adaptive Reward Scaling Policy Optimization, ARSPO)}는 보상 스케일과 작업 가중치를 동적으로 조절하여 균형 잡힌 공동 최적화를 보장합니다. 광범위한 실험 결과는 OmniVL-Guard가 최첨단 방법보다 훨씬 우수한 성능을 보이며, 도메인 외부 시나리오에서도 뛰어난 일반화 성능을 보이는 것을 입증합니다.

Original Abstract

Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinformation. To bridge this gap, this paper targets to develop a unified framework for omnibus vision-language forgery detection and grounding. In this unified setting, the {interplay} between diverse modalities and the dual requirements of simultaneous detection and localization pose a critical ``difficulty bias`` problem: the simpler veracity classification task tends to dominate the gradients, leading to suboptimal performance in fine-grained grounding during multi-task optimization. To address this challenge, we propose \textbf{OmniVL-Guard}, a balanced reinforcement learning framework for omnibus vision-language forgery detection and grounding. Particularly, OmniVL-Guard comprises two core designs: Self-Evolving CoT Generatio and Adaptive Reward Scaling Policy Optimization (ARSPO). {Self-Evolving CoT Generation} synthesizes high-quality reasoning paths, effectively overcoming the cold-start challenge. Building upon this, {Adaptive Reward Scaling Policy Optimization (ARSPO)} dynamically modulates reward scales and task weights, ensuring a balanced joint optimization. Extensive experiments demonstrate that OmniVL-Guard significantly outperforms state-of-the-art methods and exhibits zero-shot robust generalization across out-of-domain scenarios.

2 Citations
0 Influential
3.5 Altmetric
19.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!