Almieyar-Oryx-BloomBench: 인지적 기반 평가를 위한 양방향 다중 모드 벤치마크
Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
비전-언어 모델(VLMs)의 빠른 발전에도 불구하고, 해당 분야는 VLM의 실제 추론 능력을 정확하게 진단하고 인간과 유사한 다중 모드 지능으로 나아가는 데 필요한 벤치마크가 부족합니다. 대부분의 기존 평가는 단편적이고 독립적인 작업에 초점을 맞추어 중요한 인지적 약점을 가리고, 표적 개선을 위한 통찰력을 제공하지 못합니다. 이러한 격차를 해소하기 위해, 우리는 Almieyar 벤치마킹 시리즈의 일부인 BloomBench를 소개합니다. BloomBench는 VLM을 위한 최초의 인지적으로 인간 기반, 양방향(영어-아랍어) 다중 모드 벤치마크입니다. Bloom의 분류 체계를 바탕으로, BloomBench는 신중하게 설계된 이미지-질문-답변 작업을 통해 인지의 여섯 가지 수준(기억, 이해, 적용, 분석, 평가, 창조)을 체계적으로 평가합니다. 반자동 파이프라인으로 구축되고 계층화된 하이브리드 품질 보증 프로토콜을 통해 검증되었으며, 확장성, 문화적 포괄성 및 언어적 정확성을 보장합니다. 이 프레임워크를 활용하여 최첨단 VLM에 대한 종합적인 연구를 수행하고 그들의 인지적 특성을 진단했습니다. 분석 결과, 상당한 인지적 비대칭성이 나타났습니다. 즉, 최첨단 모델은 의미 이해 측면에서 뛰어난 성능을 보이지만, 사실 정보 회상 및 창의적 합성에는 어려움을 겪는 것으로 나타났습니다. 이는 현재의 일반적인 다중 모드 능력이 특정 인지 계층에 존재하는 심오한 한계를 가리고 있음을 보여줍니다. 또한, 본 연구는 아랍어와 영어 간의 중요한 성능 격차를 강조하여 현재 교차 언어 다중 모드 추론의 한계를 드러냅니다. 이러한 결과는 보다 인지적으로 정렬되고 포괄적인 VLM을 개발하는 데 필요한 기반을 제공합니다. 벤치마크 프레임워크 및 데이터 세트는 다음 주소에서 확인할 수 있습니다: https://github.com/qcri/Almieyar-Oryx-BloomBench.
Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence. Most existing evaluations focus on piecemeal or disconnected tasks, obscuring critical cognitive weaknesses and providing little insight for targeted improvement. To address this gap, we introduce BloomBench, part of the Almieyar benchmarking series, the first cognitively human-grounded, bilingual (English-Arabic) multimodal benchmark for VLMs. Grounded in Bloom's Taxonomy, BloomBench systematically evaluates six levels of cognition (Remember, Understand, Apply, Analyze, Evaluate, Create) through carefully designed image-question-answer tasks. Built with a semi-automated pipeline and validated through a stratified hybrid quality assurance protocol, it ensures scalability, cultural inclusivity, and linguistic fidelity. Leveraging this framework, we conduct a comprehensive study of state-of-the-art VLMs to diagnose their cognitive profiles. Our analysis reveals a sharp cognitive asymmetry: while state-of-the-art models achieve strong performance ceilings in semantic understanding, they struggle substantially with factual recall and creative synthesis. This demonstrates that current general multimodal proficiency masks deeper limitations in specific cognitive layers. Furthermore, our study highlights a critical performance gap between Arabic and English, exposing limitations in current cross-lingual multimodal reasoning. These findings establish a foundation for developing more cognitively aligned and inclusive VLMs. The benchmark framework and dataset is available at: https://github.com/qcri/Almieyar-Oryx-BloomBench.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.