ProEval: 생성형 AI 평가를 위한 사전 예방적 오류 탐지 및 효율적인 성능 추정
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
생성형 AI 모델 평가는 느린 추론 속도, 높은 평가 비용, 그리고 빠르게 확장되는 모델 및 벤치마크의 다양성으로 인해 점점 더 많은 자원을 필요로 합니다. 본 논문에서는 전이 학습을 활용하여 성능을 효율적으로 추정하고 오류 사례를 식별하는 사전 예방적 평가 프레임워크인 ProEval을 제안합니다. ProEval은 사전 학습된 가우시안 프로세스(GP)를 성능 점수 함수의 대리 모델로 사용하여 모델 입력을 오류 심각도 또는 안전 위반과 같은 지표에 매핑합니다. 성능 추정을 베이지안 사분법(BQ)으로, 오류 탐지를 초레벨 집합 샘플링으로 정의함으로써, 불확실성을 고려한 의사 결정 전략을 개발하여 테스트를 위한 유용한 입력 데이터를 능동적으로 선택하거나 합성합니다. 이론적으로, 사전 학습된 GP 기반 BQ 추정기가 편향되지 않고 경계 내에 있음을 증명합니다. 광범위한 실험 결과, 추론, 안전 정렬 및 분류 벤치마크에서 ProEval이 경쟁적인 기본 모델보다 훨씬 효율적임을 보여줍니다. ProEval은 실제 값과 1% 이내의 추정치를 얻는 데 8-65배 적은 샘플만 필요하며, 동시에 더 엄격한 평가 예산 하에서 더 다양한 오류 사례를 드러냅니다.
Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks. We propose ProEval, a proactive evaluation framework that leverages transfer learning to efficiently estimate performance and identify failure cases. ProEval employs pre-trained Gaussian Processes (GPs) as surrogates for the performance score function, mapping model inputs to metrics such as the severity of errors or safety violations. By framing performance estimation as Bayesian quadrature (BQ) and failure discovery as superlevel set sampling, we develop uncertainty-aware decision strategies that actively select or synthesize highly informative inputs for testing. Theoretically, we prove that our pre-trained GP-based BQ estimator is unbiased and bounded. Empirically, extensive experiments on reasoning, safety alignment, and classification benchmarks demonstrate that ProEval is significantly more efficient than competitive baselines. It requires 8-65x fewer samples to achieve estimates within 1% of the ground truth, while simultaneously revealing more diverse failure cases under a stricter evaluation budget.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.