VirPro: 시각 정보 기반 확률적 프롬프트 학습을 활용한 약지도 학습 기반 단안 3차원 객체 인식
VirPro: Visual-referred Probabilistic Prompt Learning for Weakly-Supervised Monocular 3D Detection
단안 3차원 객체 인식은 일반적으로 실제 어노테이션에 대한 의존성을 줄이기 위해 유사 레이블링 기술에 의존합니다. 최근 연구에서는 결정적인 언어적 단서가 효과적인 보조 약지도 학습 신호 역할을 할 수 있으며, 상호 보완적인 의미론적 컨텍스트를 제공한다는 사실이 밝혀졌습니다. 그러나 수동으로 제작된 텍스트 설명은 장면 전체의 개별 객체의 고유한 시각적 다양성을 포착하는 데 어려움을 겪으며, 이는 모델의 장면 인지 표현 학습 능력을 제한합니다. 이러한 문제를 해결하기 위해, 우리는 시각 정보 기반 확률적 프롬프트 학습(Visual-referred Probabilistic Prompt Learning, VirPro)이라는 적응형 멀티모달 사전 훈련 패러다임을 제안합니다. 이는 다양한 약지도 학습 기반 단안 3차원 객체 인식 프레임워크에 원활하게 통합될 수 있습니다. 구체적으로, 우리는 장면 전반에 걸쳐 학습 가능한, 인스턴스 기반의 다양한 프롬프트를 생성하고, 이를 적응형 프롬프트 뱅크(Adaptive Prompt Bank, APB)에 저장합니다. 그 후, 우리는 장면 기반의 시각적 특징을 해당 텍스트 임베딩에 통합하여 텍스트 프롬프트가 시각적 불확실성을 표현할 수 있도록 하는 멀티 가우시안 프롬프트 모델링(Multi-Gaussian Prompt Modeling, MGPM)을 도입합니다. 그런 다음, 융합된 시각-언어 임베딩으로부터 프롬프트 기반 가우시안 분포를 디코딩하여 각 인스턴스에 대한 통일된 객체 수준의 프롬프트 임베딩을 파생합니다. ROI 수준의 대조 학습 매칭을 사용하여 모달리티 정렬을 강화하고, 동일 장면 내에서 동시에 나타나는 객체의 임베딩을 잠재 공간에서 가깝게 배치하여 의미론적 일관성을 향상시킵니다. KITTI 벤치마크에 대한 광범위한 실험 결과, 우리의 사전 훈련 패러다임을 통합하면 일관되게 상당한 성능 향상을 가져오며, 기준 모델 대비 평균 정밀도 4.8% 향상을 달성했습니다.
Monocular 3D object detection typically relies on pseudo-labeling techniques to reduce dependency on real-world annotations. Recent advances demonstrate that deterministic linguistic cues can serve as effective auxiliary weak supervision signals, providing complementary semantic context. However, hand-crafted textual descriptions struggle to capture the inherent visual diversity of individuals across scenes, limiting the model's ability to learn scene-aware representations. To address this challenge, we propose Visual-referred Probabilistic Prompt Learning (VirPro), an adaptive multi-modal pretraining paradigm that can be seamlessly integrated into diverse weakly supervised monocular 3D detection frameworks. Specifically, we generate a diverse set of learnable, instance-conditioned prompts across scenes and store them in an Adaptive Prompt Bank (APB). Subsequently, we introduce Multi-Gaussian Prompt Modeling (MGPM), which incorporates scene-based visual features into the corresponding textual embeddings, allowing the text prompts to express visual uncertainties. Then, from the fused vision-language embeddings, we decode a prompt-targeted Gaussian, from which we derive a unified object-level prompt embedding for each instance. RoI-level contrastive matching is employed to enforce modality alignment, bringing embeddings of co-occurring objects within the same scene closer in the latent space, thus enhancing semantic coherence. Extensive experiments on the KITTI benchmark demonstrate that integrating our pretraining paradigm consistently yields substantial performance gains, achieving up to a 4.8% average precision improvement than the baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.