불확실성 기반 구성 요소 정렬: 하이퍼볼릭 비전-언어 모델에서 부분-전체 의미적 대표성을 활용
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
비전-언어 모델(VLMs)은 뛰어난 성능을 보이지만, 유클리디안 임베딩은 부분-전체 또는 부모-자식 구조와 같은 계층적 관계를 포착하는 데 한계가 있으며, 종종 다중 객체 구성 시나리오에서 어려움을 겪습니다. 하이퍼볼릭 VLM은 계층적 구조를 더 잘 보존하고, 함의 관계를 통해 부분-전체 관계(즉, 전체 장면과 해당 부분 이미지)를 모델링하여 이러한 문제를 완화합니다. 그러나 기존 접근 방식은 각 부분이 전체에 대한 서로 다른 수준의 의미적 대표성을 갖는다는 점을 고려하지 않습니다. 본 연구에서는 하이퍼볼릭 VLM을 향상시키기 위한 불확실성 기반 구성 요소 하이퍼볼릭 정렬(UNCHA)을 제안합니다. UNCHA는 하이퍼볼릭 불확실성을 사용하여 부분-전체 의미적 대표성을 모델링하며, 전체 장면에서 더 대표적인 부분에는 낮은 불확실성을, 덜 대표적인 부분에는 높은 불확실성을 할당합니다. 이러한 대표성은 불확실성 기반 가중치를 사용하여 대비 학습 목표에 통합됩니다. 마지막으로, 불확실성은 엔트로피 기반 항으로 정규화된 함의 손실을 통해 추가적으로 보정됩니다. 제안된 손실 함수를 통해 UNCHA는 더 정확한 부분-전체 순서를 학습하고, 이미지의 기본 구성 구조를 파악하여 복잡한 다중 객체 장면을 더 잘 이해할 수 있습니다. UNCHA는 제로샷 분류, 검색 및 다중 레이블 분류 벤치마크에서 최첨단 성능을 달성합니다. 본 연구의 코드와 모델은 다음 링크에서 확인할 수 있습니다: https://github.com/jeeit17/UNCHA.git.
While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face challenges in multi-object compositional scenarios. Hyperbolic VLMs mitigate this issue by better preserving hierarchical structures and modeling part-whole relations (i.e., whole scene and its part images) through entailment. However, existing approaches do not model that each part has a different level of semantic representativeness to the whole. We propose UNcertainty-guided Compositional Hyperbolic Alignment (UNCHA) for enhancing hyperbolic VLMs. UNCHA models part-to-whole semantic representativeness with hyperbolic uncertainty, by assigning lower uncertainty to more representative parts and higher uncertainty to less representative ones for the whole scene. This representativeness is then incorporated into the contrastive objective with uncertainty-guided weights. Finally, the uncertainty is further calibrated with an entailment loss regularized by entropy-based term. With the proposed losses, UNCHA learns hyperbolic embeddings with more accurate part-whole ordering, capturing the underlying compositional structure in an image and improving its understanding of complex multi-object scenes. UNCHA achieves state-of-the-art performance on zero-shot classification, retrieval, and multi-label classification benchmarks. Our code and models are available at: https://github.com/jeeit17/UNCHA.git.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.