MMBench-Live: 다중 모드 모델을 위한 지속적으로 진화하는 벤치마크
MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models
시각-언어 모델(VLM)을 평가하기 위한 벤치마크는 필수적이지만, 대부분의 다중 모드 벤치마크는 정적인 특성을 가지므로 시간 경과에 따른 노후화, 데이터 오염 및 비용이 많이 드는 유지 관리에 취약합니다. 본 논문에서는 다중 에이전트 기반 자동화 파이프라인으로 구축된 지속적으로 진화하는 다중 모드 벤치마크인 MMBench-Live를 소개합니다. 저희의 프레임워크는 벤치마크 진화를 작업 지향적인 데이터셋 구성으로 간주하며, 구조화된 벤치마크 사양, 피드백 기반 실시간 데이터 수집 및 실행 가능한 추론을 갖춘 검증 가능 QA 생성 기능을 통합합니다. 버전 간 비교성을 유지하기 위해, 원래 벤치마크에서 작업 관련 시각적 패턴을 추출하여 데이터 수집 및 필터링을 안내하는 분포 일관성 업데이트 전략을 도입했습니다. MMBench 기반으로 구축된 MMBench-Live는 높은 정답 정확도를 가진 5,900개의 새로운 평가 인스턴스를 포함하며, 각 업데이트에는 약 30달러의 비용이 들고 1~2시간이 소요됩니다. 광범위한 실험 결과, MMBench-Live는 모델 순위를 안정적으로 유지하고, 원래 벤치마크와의 의미론적 일관성을 유지하며, 데이터 오염과 관련된 암기 현상이 약하다는 것을 보여주었습니다. 이는 지속 가능한 다중 모드 벤치마크 진화를 위한 실용적이고 확장 가능한 패러다임을 제시합니다. 프로젝트는 https://github.com/PRIS-CV/MMBench-Live 에서 확인할 수 있습니다.
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We present MMBench-Live, a continuously evolving multimodal benchmark built by a multi-agent-driven automated pipeline. Our framework treats benchmark evolution as task-guided dataset construction, integrating structured benchmark specification, feedback-controlled real-time data acquisition, and verifiable QA generation with executable reasoning. To maintain cross-version comparability, we introduce a distribution-consistent update strategy that extracts task-related visual patterns from the original benchmark to guide data collection and filtering. Instantiated from MMBench, MMBench-Live contains 5.9K newly generated evaluation instances with a high answer correctness rate, while each update costs about USD 30 and takes 1-2 hours. Extensive evaluations show that MMBench-Live preserves stable model rankings, maintains semantic alignment with the original benchmark, and exhibits weaker contamination-related memorization signals, suggesting a practical and scalable paradigm for sustainable multimodal benchmark evolution. The project is available at https://github.com/PRIS-CV/MMBench-Live.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.