OmniVerifier-M1: 명시적인 구조적 재보정을 활용한 다중 모드 메타 검증기
OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
다중 모드 대규모 언어 모델에서 시각적 결과의 중요성이 점점 커지고 있으며, 이는 범용 기초 모델을 확장하기 위한 신뢰성 있고 세밀한 검증의 필요성을 강조합니다. 본 연구에서는 검증기가 생성하는 설명(rationale)을 활용하여 의사 결정 정보만을 사용하는 것보다 더 효과적인 다중 모드 메타 검증 방법을 조사하고, 이러한 메타 검증 피드백을 다중 모드 검증기 훈련에 어떻게 효율적으로 통합할 수 있는지 탐구합니다. 우리는 두 가지 주요 결과를 발견했습니다. 첫째, 텍스트 설명보다 기호 기반 검증기의 출력(예: 바운딩 박스)이 메타 검증의 근거로 더 우수하며, 이는 모델 기반 보상 없이 규칙 기반 강화 학습을 통해 효율적인 보상을 제공합니다. 둘째, 이진 판단과 메타 검증에 대한 강화 학습 목표를 분리하면 공동 최적화보다 성능이 훨씬 뛰어나는데, 이는 출력 구조와 학습 역학의 근본적인 차이 때문입니다. 이러한 통찰력을 바탕으로, 본 연구에서는 기호 기반 메타 검증과 분리된 강화 학습을 활용하는 범용 시각 검증기인 OmniVerifier-M1을 훈련했습니다. OmniVerifier-M1은 강력한 검증 능력과 세밀한 오류 위치 파악 기능을 제공하며, 또한 검증기를 통해 구동되는 능동적 생성 시스템인 M1-TTS를 가능하게 하여 동적인 영역 수준의 자체 수정 기능을 구현합니다. 이러한 접근 방식은 더욱 신뢰성 있고 해석 가능하며 세밀한 다중 모드 검증을 가능하게 하여, 더 안전하고 제어 가능한 기초 모델 배포를 지원합니다.
Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In this work, we investigate multimodal meta-verification, which leverages verifier-generated rationales rather than decision-only signals, and explore how to effectively incorporate meta-verification feedback into multimodal verifier training. We identify two key findings. First, symbolic verifier outputs (e.g., bounding boxes) outperform textual explanations as meta-verification rationales, enabling efficient rule-based reinforcement learning rewards while avoiding reliance on model-based rewards from auxiliary judge models. Second, decoupling reinforcement learning objectives for binary judgment and meta-verification substantially outperforms joint reward optimization, due to intrinsic differences in output structure and learning dynamics. Based on these insights, we train OmniVerifier-M1, a generalist visual verifier leveraging symbolic meta-verification and decoupled reinforcement learning. OmniVerifier-M1 provides robust verification and fine-grained error localization, and further enables M1-TTS, a verifier-driven agentic generation system achieving dynamic region-level self-correction. This approach paves the way for more reliable, interpretable, and fine-grained multimodal verification, supporting safer and more controllable foundation model deployment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.