2606.13288v1 Jun 11, 2026 cs.CV

시각-언어 조화성을 향상시키기 위한 교차 모달 마스킹 기반 구성 개념 모델링

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

Xinmei Tian
Xinmei Tian
Citations: 157
h-index: 4
Wei Li
Wei Li
Citations: 9
h-index: 2
Zhenpeng Huang
Zhenpeng Huang
Citations: 117
h-index: 7

CLIP과 같은 대조 학습 기반의 시각-언어 모델은 이미지-텍스트 표현 학습에서 상당한 발전을 이루었지만, 여전히 구성적 이해에 어려움을 겪고 있습니다. 이러한 모델들은 종종 '단어 가방' 현상을 보이는데, 이는 객체 간 관계, 속성-객체 결합 및 단어 순서 의존성을 파악하는 데 어려움이 있음을 의미합니다. 이러한 제한은 최적화를 위한 전역적인 단일 벡터 표현에 대한 의존뿐만 아니라, 쌍으로 연결된 이미지-텍스트 데이터에 내재된 풍부한 구성 정보를 충분히 활용하고 모델링하지 못하기 때문에 발생합니다. 본 연구에서는 MACCO (MAsked Compositional Concept MOdeling)라는 프레임워크를 제안합니다. MACCO는 한 모달리티에서 구성 개념을 마스킹하고, 다른 모달리티로부터 얻은 완전한 문맥 정보를 바탕으로 이를 재구성하여 모델이 교차 모달 구성 구조를 보다 효과적으로 파악하고 정렬할 수 있도록 합니다. 이 과정을 지원하기 위해, 본 연구에서는 마스킹된 특징들을 교차 및 내 모달 방식으로 함께 정렬하고 규제하는 두 가지 보조 목표를 도입했습니다. 5가지 구성 벤치마크에 대한 광범위한 실험과 심층적인 분석 결과는 제안하는 접근 방식이 VLM의 구성성을 크게 향상시킬 뿐만 아니라, 구문 구조 및 언어 정보를 파악하는 능력도 향상시킨다는 것을 보여줍니다. 또한, 개선된 구성성은 텍스트-이미지 생성 및 멀티모달 대규모 언어 모델에도 긍정적인 영향을 미칩니다. 코드: https://github.com/hiker-lw/MACCO

Original Abstract

Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding. They often exhibit a "bag-of-words" behavior--struggling to capture the object relations, attribute-object bindings, and word order dependencies. This limitation arises not only from the reliance on global, single-vector representations for optimization, but also from the insufficient exploitation and modeling of the rich compositional information inherently present in paired image text data. In this work, we propose MACCO (MAsked Compositional Concept MOdeling), a framework that masks compositional concepts in one modality and reconstructs them conditioned on the full contextual information from the other, enabling the model to capture and align cross-modal compositional structures more effectively. To facilitate this process, we introduce two auxiliary objectives that jointly align and regularize masked features both inter-modally and intra-modally. Extensive experiments on five compositional benchmarks, along with in-depth analyses, demonstrate that our approach not only significantly enhances compositionality in VLMs but also improves their ability to capture syntactic structure and linguistic information. Additionally, the improved compositionality also benefits text-to-image generation and multimodal large language model. Code is available at https://github.com/hiker-lw/MACCO.

0 Citations
0 Influential
23.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!