2602.09492v1 Feb 10, 2026 cs.LG

배치 크기에 주의: LoRA 평가 시의 하이퍼파라미터 편향

Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA

Sangyoon Lee
Sangyoon Lee
Citations: 37
h-index: 4
Jaeho Lee
Jaeho Lee
Citations: 22
h-index: 3

로우 랭크 어댑테이션(LoRA)은 대규모 언어 모델을 미세 조정하는 표준적인 방법이지만, 다양한 LoRA 변형은 동일한 벤치마크에서 종종 상반되는 성능 향상을 보고합니다. 우리는 이러한 모순이 단일하고 간과된 요인, 즉 배치 크기에서 비롯된다는 것을 보여줍니다. 적절하게 조정된 경우, 기본적인 LoRA는 더 복잡한 변형과 비슷한 성능을 나타내는 경우가 많습니다. 우리는 또한 배치 크기 조정을 위한 비용 효율적인 프록시 기반 전략을 제안하여, 랭크, 데이터 세트 크기 및 모델 용량이 최적의 배치 크기에 미치는 영향을 밝혀냅니다. 우리의 연구 결과는 배치 크기를 사소한 구현 세부 사항이 아닌, 최우선적인 설계 파라미터로 격상시키고, 기존의 불일치를 해소하며, LoRA 변형에 대한 더욱 신뢰할 수 있는 평가를 가능하게 합니다.

Original Abstract

Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical gains, often on the same benchmarks. We show that these contradictions arise from a single overlooked factor: the batch size. When properly tuned, vanilla LoRA often matches the performance of more complex variants. We further propose a proxy-based, cost-efficient strategy for batch size tuning, revealing the impact of rank, dataset size, and model capacity on the optimal batch size. Our findings elevate batch size from a minor implementation detail to a first-order design parameter, reconciling prior inconsistencies and enabling more reliable evaluations of LoRA variants.

4 Citations
0 Influential
2 Altmetric
14.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!