언어 모델 사전 학습 과정에서 작업(task)에 독립적인 학습 데이터의 영향력 측정
Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
언어 모델 사전 학습 과정 전반에 걸쳐 학습 데이터의 영향력을 일관성 있게 측정하는 것은 어려운 과제입니다. 모델의 일반적인 능력과 대표성을 갖는 다운스트림 태스크 또는 검증 세트를 선택하기 어렵고, 중간 체크포인트에서의 성능 지표 의존성은 학습 과정 전반의 비교를 복잡하게 만듭니다. 본 연구에서는 다운스트림 태스크나 검증 세트를 선택하지 않고도 학습 데이터의 영향력을 측정할 수 있는 방법을 제안합니다. 구체적으로, 우리는 특정 사전 학습 실행에서 최종 파라미터에 대한 제곱 거리 감소량을 기준으로 예제의 영향력을 정의하고, 이 값을 모델을 재학습하지 않고 중간 체크포인트로부터 추정합니다. Pythia 및 PolyPythia 제품군에서 제공하는 18개의 구성에 본 방법을 적용한 결과, 학습 과정 동안 영향을 미치는 데이터의 체계적인 시간적 변화를 확인했습니다. 학습 초기 단계에서는 문헌 관련 데이터가 최종 파라미터 방향으로 더 강하게 정렬되는 반면, STEM(과학, 기술, 공학 및 수학) 관련 데이터는 후반 단계에서 더 강한 정렬을 보입니다. 이러한 질적 변화는 다양한 모델 구성 전반에 걸쳐 광범위하게 나타났습니다. 본 연구 결과는 사전 학습 과정 동안 영향력 있는 데이터가 어떻게 변화하는지에 대한 추세 수준의 이해를 제공하며, 특정 다운스트림 태스크 또는 검증 세트에 기반한 영향 분석을 보완합니다.
Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. We propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, we define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and estimate this quantity from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, we find systematic temporal changes in influential data. Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters, whereas STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. Our results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.