2607.27798v1 Jul 30, 2026 cs.AI

MemeBench: 문화 의존형 밈 해석 시 대규모 비전-언어 모델(LVLMs)이 놓치는 점

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yuchen He
Yuchen He
Citations: 20
h-index: 2
Weihang Wang
Weihang Wang
Fudan University
Citations: 39
h-index: 3
Kainan Tu
Kainan Tu
Citations: 0
h-index: 0
Boheng Sheng
Boheng Sheng
Citations: 9
h-index: 1
Peiyi Li
Peiyi Li
Citations: 12
h-index: 1
Longwen Gao
Longwen Gao
Citations: 16
h-index: 2
Zhouhui Lian
Zhouhui Lian
Citations: 135
h-index: 6

대규모 비전-언어 모델은 시각적 콘텐츠 설명을 개선하는 데 상당한 발전을 이루었지만, 정확한 설명만으로는 픽셀 너머의 지식이 필요할 때 의미를 제대로 해석한다고 보장할 수 없습니다. 밈은 문화적 요소, 배경 지식 및 커뮤니티 관습에 의존하기 때문에 이러한 한계를 드러냅니다. 대부분의 밈 벤치마크는 해석을 레이블 또는 전체 점수로 줄여, 설명이 실패하는 부분을 가립니다. 본 연구에서는 애니메이션, 만화, 게임 및 관련 온라인 하위 문화를 중심으로 작성된 인간 참조 자료와 품질 관리된 VIKR 주석을 포함하는 1,253개의 중국어 및 영어 밈으로 구성된 진단 벤치마크인 MemeBench를 소개합니다. MemeBench의 VIKR 체계는 설명을 시각적 단서, 정체성 연결, 지식 단위 및 추론 메커니즘으로 분해합니다. 26개의 LVLMs를 분석한 결과, 모든 모델이 보이는 콘텐츠에 대해서는 더 안정적으로 설명하는 반면, 해석에 필요한 지식에 대해서는 신뢰성이 낮으며, 가장 강력한 모델조차도 여전히 22.6%의 시각-지식 격차가 존재합니다. 이러한 진단이 개선을 위한 가이드라인이 될 수 있는지 테스트하기 위해, CultureBase를 기반으로 한 개체 중심 검색 시스템인 KAR를 소개합니다. 제어된 4개의 모델에서 KAR는 VIKR 성공률을 3.6%~7.4% 향상시켰으며, 일반적인 검색 방식과 비교했을 때 더 많은 답변을 개선하고 오류를 줄였습니다. 그러나 두 가지 검색 조건 모두 정체성과 지식 이해도를 높이는 동시에 시각적 정보의 적절성을 감소시키는 경향을 보입니다. MemeBench는 해석이 성공하는지, 무엇이 부족한지, 그리고 특정 증거가 진단된 격차를 채울 수 있는지 여부를 보여줍니다.

Original Abstract

Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!