2607.08317v1 Jul 09, 2026 cs.AI

Blind-Spots-Bench: 다중 모드 모델의 약점 평가

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Emmanuel Abbe
Emmanuel Abbe
Citations: 134
h-index: 4
Matteo Santelmo
Matteo Santelmo
Citations: 0
h-index: 0
Xiuying Wei
Xiuying Wei
Citations: 53
h-index: 4
Israa Fakih
Israa Fakih
Citations: 1
h-index: 1
Felix Bauer
Felix Bauer
Citations: 0
h-index: 0
Juan Garc'ia Giraldo
Juan Garc'ia Giraldo
Citations: 46
h-index: 1
Etienne Bamas
Etienne Bamas
Citations: 0
h-index: 0
Chengkun Li
Chengkun Li
EPFL
Citations: 374
h-index: 4

최신 AI 모델은 많은 기존 벤치마크에서 뛰어난 성능을 보이지만, 여전히 인간에게는 매우 쉬운 작업들, 예를 들어 문자열 조작이나 다섯 개의 다리를 가진 강아지 그림 그리기와 같은 작업에서는 실패합니다. 이러한 예시는 현재 시스템의 지속적인 약점을 제대로 측정하지 못하는 기존 벤치마크가 존재할 수 있음을 시사합니다. 본 연구에서는 인간에게는 간단해 보이지만 현대 AI에게는 여전히 어려운 작업을 통해 그러한 약점을 드러내기 위한 벤치마크인 $ exttt{blind-spots-bench}$를 소개합니다. 우리는 AI 강좌 학생들로부터 질문을 수집하고, 정제 및 구조화된 참고 해답으로 주석을 달아, 총 235개의 샘플로 구성된 데이터셋에 적합한 작업 분류 체계를 제안했습니다. 또한 다양한 모델(오픈 소스 및 비공개 언어 모델, 시각-언어 모델, 이미지 생성 모델 포함)을 평가하기 위한 자동 채점 파이프라인을 개발했습니다. $ exttt{blind-spots-bench}$에 대한 분석 결과, 기존 벤치마크에서 유사한 성능을 보이지만 비공개 최첨단 모델이 오픈 소스 모델보다 현저히 뛰어난 성능(약 10% 차이)을 보이는 것으로 나타났습니다. 더 자세한 분석 결과, 어떤 단일 모델도 모든 작업 유형에서 우위를 점하지 않으며, 일부 작업은 평가된 모든 모델에게 여전히 어려운 과제임을 확인했습니다. 이러한 결과는 $ exttt{blind-spots-bench}$가 현재의 현대 AI 모델의 구체적인 약점을 식별하기 위한 진단적 스트레스 테스트로서 가치가 있음을 강조합니다.

Original Abstract

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce $\texttt{blind-spots-bench}$, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on $\texttt{blind-spots-bench}$ reveals that closed-source frontier models can substantially outperform open-weight models with even $\approx10\%$ gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of $\texttt{blind-spots-bench}$ as a diagnostic stress test for identifying concrete weaknesses in current modern models.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!