2606.22826v1 Jun 22, 2026 cs.AI

MINCE: 소수의 모델 기반 몬테카를로 교정을 통한 LLM 평가 데이터셋 축소

MINCE: Shrinking LLM Evaluation Datasets via Few-Model Monte Carlo Calibration

Ashish Sirasao
Ashish Sirasao
Citations: 896
h-index: 9
Devleena Das
Devleena Das
Citations: 26
h-index: 3
Rajeev Patwari
Rajeev Patwari
Citations: 16
h-index: 3
Vikram Kumar Bukka
Vikram Kumar Bukka
Citations: 0
h-index: 0
Nithin Kumar Guggilla
Nithin Kumar Guggilla
Citations: 0
h-index: 0
Elliott Delaye
Elliott Delaye
Citations: 101
h-index: 6

다양한 LLM 변형(양자화, 미세 조정 또는 배포 환경에 특화된 모델 등)을 평가하려면 대규모 벤치마크를 반복적으로 실행해야 하는데, 이는 NPU와 같은 엣지 하드웨어에서 모델당 수십 시간을 소요할 수 있습니다. 기존의 부분 집합 선택 방법은 이러한 비용을 줄이지만, 대규모 교정 데이터 풀 또는 학습된 예측 레이어를 필요로 합니다. 본 연구에서는 MINCE (Monte Carlo Informed N-sizing for Compact Evaluation)를 소개합니다. MINCE는 소수의 교정 모델에서 얻은 개별 항목 로그에 대한 몬테카를로 시뮬레이션을 사용하여 정확도 변화를 제한하는 최소 부분 집합 크기를 찾고, 예측 레이어 없이 해당 크기로 무작위로 샘플링된 부분 집합을 고정합니다. MINCE는 최대 드리프트가 ≤2.62%p인 BF16 모델에서 IFEVAL을 54%, MMLU를 89%, GSM8K를 70% 줄입니다. 또한, 독립적인 NPU 모델에서의 평균 드리프트는 0.77–3.59%p이며, GPU 평가 속도는 중앙값 기준 2.7~8.1배, NPU 평가 속도는 1.7~2.0배 향상됩니다. 본 방법은 교정 데이터 풀의 크기에 강인하며, tinyBenchmarks보다 낮은 드리프트를 달성합니다 (MMLU에서 12배 낮음, GSM8K에서 3.3배 낮음) 동시에 교정 모델 수를 57배 줄입니다.

Original Abstract

Evaluating LLMs across many model variants -- quantized, fine-tuned, or deployment-specific -- requires running large benchmarks repeatedly, a process that can take tens of hours per model on edge hardware such as NPUs. Existing subset selection methods reduce this cost but depend on large calibration pools or learned prediction layers. We introduce MINCE (Monte Carlo Informed N-sizing for Compact Evaluation), which uses Monte Carlo simulation over per-item logs from a small set of calibration models to find the minimum subset size that bounds accuracy drift and then fixes a randomly sampled subset at that size, with no prediction layer needed. MINCE reduces IFEVAL by 54\%, MMLU by 89\%, and GSM8K by 70\% with maximum drift $\leq$2.62\,pp on BF16 models and mean drift of 0.77--3.59\,pp on held-out NPU models, while delivering median GPU evaluation speedups of 2.7--8.1$\times$ and NPU evaluation speedups of 1.7--2.0$\times$. The method is robust to calibration pool size and achieves lower drift than tinyBenchmarks (12$\times$ lower on MMLU, 3.3$\times$ on GSM8K) while using 57$\times$ fewer calibration models.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!