2604.16054v1 Apr 17, 2026 cs.CV

마음의 눈: 다중 모드 대규모 언어 모델을 위한 시각적 추상화, 변환 및 조합 벤치마크

Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs

Vineeth N. Balasubramanian
Vineeth N. Balasubramanian
Citations: 390
h-index: 8
Sai Srinivas Kancheti
Sai Srinivas Kancheti
Citations: 50
h-index: 2
Aditya Kanade
Aditya Kanade
Citations: 61
h-index: 3
Rohit Sinha
Rohit Sinha
Citations: 12
h-index: 2
Tanuja Ganu
Tanuja Ganu
Citations: 982
h-index: 13

다중 모드 대규모 언어 모델(MLLM)은 시각-언어 벤치마크에서 상당한 발전을 이루었지만, 시각적 인지 능력과 시공간적 추론 능력은 여전히 명확하게 이해되지 않습니다. 본 연구에서는 고전적인 인간 지능 테스트에서 영감을 받아, '추상화(Abstraction)', '관계(Relation)', '변환(Transformation)'이라는 새로운 분류 체계인 'A-R-T'로 구성된 다지선다형 벤치마크인 "마음의 눈"을 소개합니다. 이 벤치마크는 패턴 유도, 유추 관계 매핑, 정신적 변환과 같은 핵심적인 유동 지능 과정을 평가합니다. 폐쇄형 및 오픈 소스 MLLM을 포함한 다양한 모델을 평가하고, 인간 참가자들과의 성능을 비교했습니다. 인간은 80%의 정확도를 달성하는 반면, 가장 뛰어난 성능을 보이는 MLLM은 50% 미만의 정확도를 보였습니다. 오차 분석 결과, (i) 시각적 주의 집중, (ii) 내부적인 지각적 조작, (iii) 기본적인 시각적 개념의 추상화 부족 등의 실패 원인이 발견되었습니다. 이러한 결과는 현재 MLLM이 인간 참가자와 비교했을 때 제한적인 시공간적 추론 능력을 가지고 있음을 시사하며, 보다 인지적으로 기반한 평가 프레임워크의 필요성을 강조합니다.

Original Abstract

Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a multiple-choice benchmark of eight visuo-cognitive tasks inspired by classic human intelligence tests and organized under a novel "A-R-T" taxonomy: Abstraction, Relation, and Transformation. The tasks probe core processes of fluid intelligence such as pattern induction, analogical relation mapping, and mental transformation. We evaluate a diverse suite of closed-source and open-source MLLMs and compare their performance with human participants. Humans achieve 80% accuracy, while top performing MLLMs remain below 50%. Error analysis reveals failures in: (i) visual attention allocation, (ii) internal perceptual manipulation, and (iii) weak abstraction of underlying visual concepts. Our findings suggest that current MLLMs exhibit limited visuospatial reasoning capabilities, when compared with human participants, highlighting the need for more cognitively grounded evaluation frameworks.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!