점진적 적응에서의 최적 학습 시간 스케일링
Optimal Training-Time Scaling in Gradual Adaptation
점진적 적응에서, 중간 작업의 수가 증가함에 따라 각 작업에 대한 학습 시간을 어떻게 조정해야 할까요? 본 연구에서는 부드럽게 변화하고 공통적인 손실 최소 해를 가지는 과매개변수 선형 회귀 문제에 대해 이 질문을 다룹니다. $N$개의 작업과 각 작업당 $s_N$의 학습 시간을 갖는 경우, 최종 학습 진행은 $Ns_N oτ$일 때 연속적인 곡선으로 수렴합니다. 작은 $ au$ 값에서는 제한된 진행률이 $Θ(τ)$이고, 큰 $ au$ 값에서는 $Θ(τ^{-1})$입니다. 따라서 매우 짧거나 매우 긴 학습은 모두 큰 진전을 가져오지 않습니다. 결과적으로 각 작업에 대한 최적의 학습 시간은 $s_N^ ext{*} = Θ(N^{-1})$으로 스케일링되며, 이는 $Ns_N^ ext{*}=Θ(1)$과 동일합니다. 점진적으로 회전된 MNIST 데이터셋 및 실제 Yearbook 시간 변화에 대한 실험 결과는 경로가 더 세분화될수록 각 작업에 대한 학습 시간을 줄이는 것이 효과적임을 뒷받침합니다.
In gradual adaptation, how should the training time on each task change as the number of intermediate tasks increases? We study this question for overparameterized linear regression tasks that change smoothly and share a zero-loss solution. With $N$ tasks and training time $s_N$ on each, the final learning progress converges to a continuum curve when $Ns_N\toτ$. The limiting progress is $Θ(τ)$ for small $τ$ and $Θ(τ^{-1})$ for large $τ$, so both very short and very long training produce little progress. It follows that optimal per-task training times scale as $s_N^\star=Θ(N^{-1})$, equivalently $Ns_N^\star=Θ(1)$. Experiments on gradually rotated MNIST and a natural Yearbook time shift are consistent with less per-task training as the path is divided more finely.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.