SABRE: 스트레스 상황에서의 거대 시각언어 모델(VLM)의 확장 가능하고 자동화된 성능 평가
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
시각-언어 모델(VLM)은 빠르게 발전하고 있지만, 이에 상응하는 벤치마크 개발은 뒤쳐져 있어 VLM의 약점을 파악하기 어렵습니다. 스트레스 테스트를 구축하는 것은 비용이 많이 들기 때문에, 샘플은 제어된 조건을 만족해야 하며, 답변 가능해야 하고, 현재 모델에 도전적이어야 합니다. 본 논문에서는 확장 가능하고 자동화된 파이프라인인 SABRE를 소개합니다. SABRE는 Test Primer (마크다운 형식의 작업 설계 및 데이터 스키마)를 구조화된 사양, 생성 또는 편집된 이미지, 질의응답 쌍으로 변환합니다. 자동 필터링은 특정 VLM에 의해 해결된 후보를 제거하고, 인간 검토자는 후보의 유효성을 확인하며, 주석 수정 및 이미지 보정을 지원합니다. 우리는 SABRE-Prior를 구현하여 VLM이 시각적 증거를 따르는지, 아니면 세계 사전 지식(익숙한 객체와 장면)에 의존하는지를 테스트했습니다. 이 데이터셋은 600개의 이미지와 1,000개의 질문으로 구성되며, Context (익숙한 장면 내의 예상치 못한 개체), Texture (가상 재질), Attribute (비정규 부품 수), 그리고 Language Elicitation (이미지에 의해 뒷받침되지 않는 언어적 단서) 영역을 포함합니다. 6개의 VLM 모델에 대한 평균 정확도는 17.8%에서 31.3% (평균 22.6%) 범위였습니다. 실제 이미지 기반의 Attribute 검증은 필터링 VLM에게 특히 어려운 것으로 나타났습니다. SABRE-Counting 및 SABRE-Spatial 파일럿 테스트는 워크플로우가 다른 스트레스 테스트 설정도 지원한다는 것을 보여줍니다. 이러한 결과는 SABRE를 단일 고정 벤치마크가 아닌, VLM 스트레스 테스트 구축 및 업데이트를 위한 재사용 가능한 프레임워크로 확립합니다.
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question-answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE-Prior to test whether VLMs follow visual evidence instead of relying on world priors -- learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (noncanonical component counts), and Language Elicitation (answers suggested by language but unsupported by the image). Across six VLMs, macro-average accuracy ranges from 17.8% to 31.3% (22.6% mean). A real-image Attribute control is comparably difficult for the Filtering VLM. SABRE-Counting and SABRE-Spatial pilots show that the workflow supports other stress-test settings. These results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.