범용 멀티모달 임베딩에서의 시각적 식별력 향상
Illuminating Visual Identity in Universal Multimodal Embeddings
범용 멀티모달 임베딩(UMEs)은 다양한 모드와 작업을 하나의 공유 표현 공간으로 통합하는 것을 목표로 합니다. 최근 몇 년 동안, 멀티모달 대규모 언어 모델(MLLMs)의 발전으로 인해 이 분야는 상당한 발전을 이루었습니다. 그러나 기존 UME 방법에서는 중요한 능력인 시각적 식별력 구분이 충분히 연구되지 않았습니다. 이는 인스턴스 검색, 재식별 및 AI 생성 콘텐츠에서의 정체성 보존 등 다양한 작업에서 핵심적인 역할을 합니다. 이러한 격차를 해소하기 위해, 우리는 시각적 식별력 구분(VisID)에 대한 통합 프레임워크를 제안하고, 실제 데이터와 합성 데이터를 모두 활용하여 구성된 대규모 벤치마크인 $ extbf{MVEB}$ ($ extbf{M}$ultimodal $ extbf{V}$isual $ extbf{I}$dentity $ extbf{E}$mbedding $ extbf{B}$enchmark)를 소개합니다. 또한, 신중하게 설계된 식별 정보 기반 샘플링 메커니즘을 통해 일반적인 멀티모달 표현과 시각적 식별 표현을 동시에 최적화하는 간단하면서도 효과적인 학습 프레임워크를 제시합니다. 광범위한 실험 결과는 우리의 접근 방식이 UME에 강력한 식별력 구분 능력을 부여하고 경쟁력 있는 일반적인 멀티모달 성능을 유지한다는 것을 보여줍니다. 우리는 이 연구가 중요한 그러나 간과되었던 기능을 강조할 뿐만 아니라, 더욱 포괄적인 범용 멀티모달 임베딩으로 나아가는 한 걸음이 될 것이라고 믿습니다. 코드와 데이터는 [https://chrisclear3.github.io/MVEB](https://chrisclear3.github.io/MVEB) 에서 이용 가능합니다.
Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs). However, a crucial capability, visual identity discrimination, remains underexplored in existing UME methods, despite its critical role in a wide range of tasks, including instance retrieval, re-identification, and identity preservation in AI-generated content. To bridge this gap, we propose a unified formulation for visual identity discrimination~(VisID) and introduce $\textbf{MVEB}$ ($\textbf{M}$ultimodal $\textbf{V}$isual Identity $\textbf{E}$mbedding $\textbf{B}$enchmark), a large-scale benchmark curated from both real-world and synthetic datasets to support evaluation and training. Furthermore, we present a simple yet effective learning framework that jointly optimizes general multimodal and visual identity representations through a carefully designed identity-aware sampling mechanism. Extensive experiments demonstrate that our approach successfully endows UMEs with strong identity discrimination capability and maintains competitive general multimodal performance. We believe this work not only illuminates a critical yet neglected capability, but also takes a step toward more holistic universal multimodal embeddings. Code and data are available at \href{https://chrisclear3.github.io/MVEB}{MVEB}.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.