ReasonCLIP-58M: 시각적으로 기반된 상식 추론 감독을 통한 CLIP 모델 학습
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
CLIP 및 그 변형체는 다양한 멀티모달 시스템에서 널리 사용되는 시각적 백본이지만, 사전 학습은 여전히 설명적인 이미지-텍스트 정렬에 의존하고 있습니다. 하위 작업이 점차 시각적으로 기반된 상식 추론과 조립적 추론을 요구함에 따라, CLIP 스타일 인코더가 구조적 변경 없이 이러한 추론을 지원할 수 있는지 불확실합니다. 이를 해결하기 위해, 우리는 두 단계 전략을 통해 대규모 추론 감독 신호를 CLIP 스타일 모델에 통합하는 지속적인 사전 학습 프레임워크인 ReasonCLIP-58M을 제안합니다. 이 과정에서 설명적 정렬을 유지하면서 점진적으로 추론 신호를 통합하고, 범주 구조화된 추론 감독을 수행합니다. 본 프레임워크를 지원하기 위해, 우리는 두 가지 상호 보완적인 데이터셋과 벤치마크를 구축했습니다. 첫 번째는 시각적으로 검증 가능한 개방형 추론 설명을 포함하는 ReasonLite-42M이고, 두 번째는 범주별 추론 감독을 제공하는 ReasonPro-16M입니다. 또한, 시각적으로 기반된 추론 능력을 진단 평가하기 위한 RCLIP-Bench를 개발했습니다. 우리는 ReasonCLIP 모델 패밀리를 학습하여 시각적으로 기반된 상식 및 조립적 추론 능력을 향상시키고, 동시에 제로샷 검색 성능도 향상시켰습니다. LLaVA-NeXT와 같은 멀티모달 대규모 언어 모델의 대체 가능한 시각 인코더로서, ReasonCLIP은 추가적인 추론 비용 없이 일관된 성능 향상을 제공하며, 이는 구조화된 추론 감독이 CLIP 스타일의 시각적 표현 능력 향상에 기여한다는 것을 보여줍니다. 모든 데이터셋, 모델 및 학습 코드는 https://github.com/RISys-Lab/ReasonCLIP 에서 확인할 수 있습니다.
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually grounded commonsense inference and compositional reasoning, it remains unclear whether CLIP-style encoders can support such reasoning without architectural changes. To address this, we present ReasonCLIP-58M, a continual pretraining framework that integrates large-scale reasoning supervision into CLIP-style models through our two-stage strategy, which progressively integrates reasoning signals while preserving descriptive alignment, followed by category-structured reasoning supervision. To support this framework, we construct two complementary datasets and a benchmark: ReasonLite-42M, with open-form, visually verifiable reasoning captions; ReasonPro-16M, with category-specific reasoning supervision; and RCLIP-Bench for diagnostic evaluation of visually grounded reasoning. We train a family of ReasonCLIP that improves visually grounded commonsense and compositional reasoning while also enhancing zero-shot retrieval performance. As a drop-in visual encoder for multimodal large language models such as LLaVA-NeXT, ReasonCLIP delivers consistent gains without additional inference cost, demonstrating that structured reasoning supervision enhances the expressive capacity of CLIP-style visual representations. All datasets, models, and training code are available at https://github.com/RISys-Lab/ReasonCLIP.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.