PerceptionBench: 다중 모드 대규모 언어 모델의 기본 시각 인지 능력 평가
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
본 논문에서는 다중 모드 대규모 언어 모델(MLLM)의 기본적인 시각 인지 능력을 평가하기 위해 특별히 설계된 벤치마크인 PerceptionBench를 소개합니다. 기존 벤치마크는 종종 인지 능력 평가에 한계를 보입니다. 종합적인 평가는 인지 오류와 추론 또는 도메인 지식 부족을 혼동시키며, 응용 프로그램 중심 벤치마크는 휴리스틱 설계로 인해 제한적이고 단편적인 영역만 다룹니다. 이러한 한계점을 해결하기 위해 PerceptionBench는 하향식 접근 방식을 채택합니다. 최첨단 MLLM의 응답에서 가장 초기에 발생하는 오류 지점을 분석하여 총 42개의 기존 벤치마크를 기반으로 오류 분류 체계를 구축하고, 이 체계 내에서 인지 분야에 해당하는 10가지 기본적인 인지 능력을 정의했습니다. 이 분류 체계를 바탕으로, 추론이나 지식이 아닌 인지에 의해 발생하는 난이도를 가진, 짧고 명확한 답변을 요구하는 3,000개의 검증된 질문 세트를 구축했습니다. 16개의 최첨단 MLLM에 대한 벤치마크 결과는 기본적인 인지 능력이 여전히 해결되지 않은 문제임을 보여줍니다. 어떤 모델도 60%의 정확도를 달성하지 못했으며, 평균적으로 시각 관련 환각이 가장 취약한 능력이며, 유사한 전반적인 점수는 뚜렷하게 다른 능력 프로필을 숨깁니다. 따라서 PerceptionBench는 MLLM의 시각 인지 경계를 측정하고 진단하기 위한 능력 수준의 표준을 제공합니다.
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.