2606.12117v1 Jun 10, 2026 cs.CL

공정하고 효율적인 LLM 성능 평가를 위한 소프트 프롬프트 튜닝

Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation

Kristian Kersting
Kristian Kersting
Citations: 95
h-index: 5
Bjorn Deiseroth
Bjorn Deiseroth
Citations: 860
h-index: 9
Letitia Parcalabescu
Letitia Parcalabescu
Citations: 1
h-index: 1
Selen Erkan
Selen Erkan
Citations: 28
h-index: 3
Bastian Boll
Bastian Boll
Citations: 41
h-index: 4

벤치마크 점수는 대규모 언어 모델(LLM)의 지식을 종종 왜곡합니다. 이는 벤치마크가 특정 형식 요구 사항을 따르는 모델의 능력에 의존하기 때문입니다. 특히, 올바른 답을 알고 있지만, 일반적으로 추가 학습을 통해 얻는 방식으로 답변을 구성하는 능력이 부족한 기본 모델이 불리하게 평가될 수 있습니다. 이를 해결하기 위해, 우리는 효율적이고 공정하며 아키텍처에 구애받지 않는 모델 평가 방법인 소프트 프롬프트 튜닝을 제안합니다. 짧은 튜닝 기간 동안 약 10개의 소프트 프롬프트 벡터(7B 모델의 경우 대략 0.0006%의 파라미터)만 최적화함으로써, 모델을 특정 벤치마크 형식에 맞게 조정하여 형식 준수 능력 부족 문제를 해결하고, 기본 지식이 벤치마크 점수에 정확하게 반영되도록 합니다. 이를 통해 다양한 사전 학습 방식을 사용한 서로 다른 기본 모델을 전체 추가 학습 없이 공정하게 벤치마크할 수 있습니다. 우리는 7개의 모델과 7개의 데이터 세트에서 소프트 프롬프트 튜닝을 평가했습니다. 결과는 (a) 소프트 프롬프트 튜닝이 80 단계(약 640개 샘플) 내에 형식 준수를 최대한 활용하여 매우 효율적임을 보여줍니다. (b) 소프트 프롬프트 튜닝은 제로-샷 및 퓨-샷 프롬프팅보다 훨씬 뛰어난 성능을 보이며, 표준 프롬프팅으로는 놓치는 기본 모델의 지식을 드러냅니다. (c) 심지어 추가 학습된 모델도 형식 준수를 극대화하기 위해 소프트 프롬프트로부터 이점을 얻을 수 있습니다. (d) 소프트 프롬프트가 적용된 기본 모델의 성능은 제로-샷 및 퓨-샷 기준선보다 후속 모델 순위를 더 안정적으로 예측하여, 다운스트림 모델 품질에 대한 저렴하고 효과적인 지표를 제공합니다. 우리의 기여는 다음과 같습니다 (1) 형식 준수와 지식 정확성을 분리하는 메트릭, (2) LLM 지식을 평가하기 위한 더욱 공정한 벤치마킹 프로토콜, 그리고 (3) LLM 개발 초기에 최적의 사전 학습 전략을 식별할 수 있는 비용 효율적인 방법입니다.

Original Abstract

Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific formatting requirements. This especially penalizes base models that may know the correct answers but lack the ability -- typically introduced in post-training -- to structure them as instructed. To overcome this, we propose soft-prompt tuning, an efficient, fair, and architecture-agnostic model evaluation. By optimizing only 10 soft-prompt vectors (roughly 0.0006% parameters for a 7B model) over a short tuning period, we adapt models to specific benchmark formats, closing gaps in format-following and ensuring that underlying knowledge is accurately reflected in benchmark scores. This allows one to fairly compare different base models -- trained with various pre-training recipes -- on benchmarks without the need for full post-training. We evaluated soft-prompt tuning across 7 models and 7 datasets. The results show that (a) soft-prompt tuning saturates format-following within 80 steps (~640 samples) making it highly efficient, (b) soft-prompt tuning significantly outperforms zero- and few-shot prompting, surfacing base model knowledge that standard prompting misses, that (c) even post-trained models can benefit from soft-prompts to maximize format compliance, and that (d) soft-prompted base model performance predicts post-trained model rankings more reliably than zero- and few-shot baselines, offering a low-cost proxy for downstream model quality. Our contributions include (1) metrics which disentangle format-following and knowledge accuracy, (2) a fairer benchmarking protocol of LLM knowledge, and (3) a cost- and memory-effective recipe to identify optimal pre-training strategies early in LLM development.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!