2607.24957v1 Jul 27, 2026 cs.CV

PerceptionBench: 다중 모드 대규모 언어 모델의 기본 시각 인지 능력 평가

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Chenzhuang Du
Chenzhuang Du
Citations: 2,457
h-index: 8
Haotian Yao
Haotian Yao
Citations: 1,839
h-index: 5
Xinxing Zu
Xinxing Zu
Citations: 1,773
h-index: 6
Y. Charles
Y. Charles
Citations: 1,708
h-index: 7
Yangyang Liu
Yangyang Liu
Citations: 1,678
h-index: 5
Yiping Bao
Yiping Bao
Citations: 1,776
h-index: 6
Zaida Zhou
Zaida Zhou
Citations: 2,519
h-index: 9
Zhiqi Huang
Zhiqi Huang
Peking University
Citations: 2,158
h-index: 15
Hongcheng Gao
Hongcheng Gao
Citations: 82
h-index: 4
Mengfan Dong
Mengfan Dong
Citations: 495
h-index: 3
Jia Li
Jia Li
Citations: 233
h-index: 4
Weihong Li
Weihong Li
Citations: 214
h-index: 1
Bowen Qu
Bowen Qu
Citations: 1,005
h-index: 7
Haiming Wang
Haiming Wang
Citations: 313
h-index: 4
Hao Yang
Hao Yang
Citations: 1,397
h-index: 6
Junwei Yang
Junwei Yang
Citations: 400
h-index: 5
Zijia Zhao
Zijia Zhao
Citations: 573
h-index: 7
Xinyu Zhou
Xinyu Zhou
Citations: 977
h-index: 6
Yifeng Xie
Yifeng Xie
Citations: 9
h-index: 1
Zichao Lin
Zichao Lin
Citations: 146
h-index: 4
Yuhao Dong
Yuhao Dong
Citations: 1,279
h-index: 15
Haoning Wu
Haoning Wu
Citations: 483
h-index: 5
Zuhao Yang
Zuhao Yang
Citations: 67
h-index: 3
Jinguo Zhu
Jinguo Zhu
Citations: 3,633
h-index: 9
Haoyu Lu
Haoyu Lu
Citations: 396
h-index: 3
Tongtian Yue
Tongtian Yue
Citations: 280
h-index: 10
Zhangyang Qi
Zhangyang Qi
Citations: 309
h-index: 6
Peizhou Cao
Peizhou Cao
Citations: 96
h-index: 3
Lin Sui
Lin Sui
Citations: 329
h-index: 3
Jia Chen
Jia Chen
Citations: 24
h-index: 3
Yao Wang
Yao Wang
Citations: 0
h-index: 0
Xiaoxue Wu
Xiaoxue Wu
Citations: 35
h-index: 3
Yaling Wang
Yaling Wang
Citations: 0
h-index: 0

본 논문에서는 다중 모드 대규모 언어 모델(MLLM)의 기본적인 시각 인지 능력을 평가하기 위해 특별히 설계된 벤치마크인 PerceptionBench를 소개합니다. 기존 벤치마크는 종종 인지 능력 평가에 한계를 보입니다. 종합적인 평가는 인지 오류와 추론 또는 도메인 지식 부족을 혼동시키며, 응용 프로그램 중심 벤치마크는 휴리스틱 설계로 인해 제한적이고 단편적인 영역만 다룹니다. 이러한 한계점을 해결하기 위해 PerceptionBench는 하향식 접근 방식을 채택합니다. 최첨단 MLLM의 응답에서 가장 초기에 발생하는 오류 지점을 분석하여 총 42개의 기존 벤치마크를 기반으로 오류 분류 체계를 구축하고, 이 체계 내에서 인지 분야에 해당하는 10가지 기본적인 인지 능력을 정의했습니다. 이 분류 체계를 바탕으로, 추론이나 지식이 아닌 인지에 의해 발생하는 난이도를 가진, 짧고 명확한 답변을 요구하는 3,000개의 검증된 질문 세트를 구축했습니다. 16개의 최첨단 MLLM에 대한 벤치마크 결과는 기본적인 인지 능력이 여전히 해결되지 않은 문제임을 보여줍니다. 어떤 모델도 60%의 정확도를 달성하지 못했으며, 평균적으로 시각 관련 환각이 가장 취약한 능력이며, 유사한 전반적인 점수는 뚜렷하게 다른 능력 프로필을 숨깁니다. 따라서 PerceptionBench는 MLLM의 시각 인지 경계를 측정하고 진단하기 위한 능력 수준의 표준을 제공합니다.

Original Abstract

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!