2606.06462v1 Jun 04, 2026 cs.AI

모든 것, 모든 곳의 벤치마크: 종합적인 평가 시스템

Benchmark Everything Everywhere All at Once

Yuang Ai
Yuang Ai
Citations: 416
h-index: 10
Xiaohui Li
Xiaohui Li
Citations: 20
h-index: 2
Dongming Wu
Dongming Wu
Citations: 50
h-index: 2
Shiyun Xiong
Shiyun Xiong
Citations: 8
h-index: 1
Wencheng Han
Wencheng Han
Citations: 236
h-index: 9
Xiangyu Yue
Xiangyu Yue
Citations: 209
h-index: 3
Peiwen Sun
Peiwen Sun
Citations: 94
h-index: 4
Bo Yang
Bo Yang
Citations: 11
h-index: 1

벤치마크는 LLM(대규모 언어 모델) 및 MLLM(멀티 모달 대규모 언어 모델)을 평가하고 발전시키는 데 필수적이며, 성능에 대한 표준화되고 명확한 지표를 제공합니다. 그러나 벤치마크 구축은 많은 노동력을 필요로 하며 재사용이 어렵기 때문에 지속 가능성과 확장성에 대한 우려가 제기됩니다. 또한 기존 벤치마크는 출시 후 빠르게 성능 포화 상태에 도달하여 최첨단 모델 간의 차별성을 제대로 반영하지 못하는 경우가 많습니다. 이러한 문제점을 해결하기 위해, 우리는 벤치마크 구축을 위한 완전 자동 에이전트 시스템인 Benchmark Agent를 소개합니다. 우리의 프레임워크는 사용자 쿼리 분석 및 하위 작업 설계부터 데이터 어노테이션 및 품질 관리까지 전체 벤치마크 구축 파이프라인을 조정합니다. Benchmark Agent의 성능을 평가하기 위해, 우리는 다양한 평가 시나리오(텍스트 이해, 멀티모달 이해, 특정 도메인 추론 등)를 포괄하는 15개의 대표적인 벤치마크를 생성했습니다. 인간 평가, LLM-as-a-judge 평가 및 일관성 검사와 같은 광범위한 실험을 통해 Benchmark Agent가 최소한의 인간 개입으로 고품질의 벤치마크 샘플을 생성할 수 있음을 입증했습니다. 더욱 중요한 점은 지속적인 평가를 통해 현재 모델이 특정 도메인별 추론 작업에 어려움을 겪는다는 등 흥미로운 사실들을 발견했습니다. 우리는 빠르게 진화하는 벤치마크가 연구 커뮤니티에 크게 기여할 수 있다고 믿습니다. 데모 페이지 및 코드 저장소에서 미리 보기와 코드를 공개할 예정입니다.

Original Abstract

Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensive and hard to reuse, raising concerns about sustainability and scalability. Moreover, existing benchmarks often quickly reach performance saturation after their release, resulting in insufficient discrimination among state-of-the-art models. To address these challenges, we introduce Benchmark Agent, a fully autonomous agentic system designed for benchmark building. Our framework orchestrates the complete benchmark construction pipeline, from user query analysis and subtask design to data annotation and quality control. To assess Benchmark Agent, we implement it to produce 15 representative benchmarks, spanning diverse evaluation scenarios, including text understanding, multimodal understanding, and domain-specific reasoning. Extensive experiments, including human evaluation, LLM-as-a-judge assessment, and consistency checks, demonstrate Benchmark Agent can generate high-quality benchmark samples with minimal human involvement. More importantly, through continual evaluation, we observe several insightful findings, including that current models struggle with certain domain-specific reasoning tasks. We believe that rapidly evolving benchmarks can contribute significantly to the research community. The preview and code will be publicly available at the demo page and code repository.

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!