2608.13515v1 Aug 13, 2026 cs.CL

언어 모델 사전 학습 과정에서 작업(task)에 독립적인 학습 데이터의 영향력 측정

Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

Chaoran Liu
Chaoran Liu
Citations: 24
h-index: 2
Yuto Nishida
Yuto Nishida
Citations: 32
h-index: 3
Hirokazu Kiyomaru
Hirokazu Kiyomaru
Citations: 160
h-index: 7
Yusuke Oda
Yusuke Oda
Citations: 184
h-index: 5
Takashi Kodama
Takashi Kodama
Citations: 15
h-index: 2
Daisuke Kawahara
Daisuke Kawahara
Citations: 56
h-index: 3
Yusuke Miyao
Yusuke Miyao
Citations: 75
h-index: 4
Max Müller-Eberstein
Max Müller-Eberstein
Citations: 0
h-index: 0
Masaru Isonuma
Masaru Isonuma
Citations: 4
h-index: 1

언어 모델 사전 학습 과정 전반에 걸쳐 학습 데이터의 영향력을 일관성 있게 측정하는 것은 어려운 과제입니다. 모델의 일반적인 능력과 대표성을 갖는 다운스트림 태스크 또는 검증 세트를 선택하기 어렵고, 중간 체크포인트에서의 성능 지표 의존성은 학습 과정 전반의 비교를 복잡하게 만듭니다. 본 연구에서는 다운스트림 태스크나 검증 세트를 선택하지 않고도 학습 데이터의 영향력을 측정할 수 있는 방법을 제안합니다. 구체적으로, 우리는 특정 사전 학습 실행에서 최종 파라미터에 대한 제곱 거리 감소량을 기준으로 예제의 영향력을 정의하고, 이 값을 모델을 재학습하지 않고 중간 체크포인트로부터 추정합니다. Pythia 및 PolyPythia 제품군에서 제공하는 18개의 구성에 본 방법을 적용한 결과, 학습 과정 동안 영향을 미치는 데이터의 체계적인 시간적 변화를 확인했습니다. 학습 초기 단계에서는 문헌 관련 데이터가 최종 파라미터 방향으로 더 강하게 정렬되는 반면, STEM(과학, 기술, 공학 및 수학) 관련 데이터는 후반 단계에서 더 강한 정렬을 보입니다. 이러한 질적 변화는 다양한 모델 구성 전반에 걸쳐 광범위하게 나타났습니다. 본 연구 결과는 사전 학습 과정 동안 영향력 있는 데이터가 어떻게 변화하는지에 대한 추세 수준의 이해를 제공하며, 특정 다운스트림 태스크 또는 검증 세트에 기반한 영향 분석을 보완합니다.

Original Abstract

Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. We propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, we define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and estimate this quantity from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, we find systematic temporal changes in influential data. Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters, whereas STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. Our results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!