2607.24717v1 Jul 27, 2026 cs.CL

DataOrchestra: 사전 학습 데이터의 개별 샘플에 따른 맞춤형 관리 체계 학습

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Shijie Xia
Shijie Xia
Shanghai Jiao Tong University
Citations: 967
h-index: 9
Zhen Huang
Zhen Huang
Citations: 870
h-index: 6
Pengfei Liu
Pengfei Liu
Citations: 528
h-index: 7
Yikun Wang
Yikun Wang
Citations: 112
h-index: 5

사전 학습 데이터 처리는 대규모 언어 모델(LLM)의 성능에 매우 중요한 영향을 미칩니다. 그러나 기존 방식들은 대부분 코퍼스 또는 도메인 수준에서 고정된 처리 전략을 정의하고, 이를 모든 샘플에 동일하게 적용하여 각 샘플의 특성에 맞게 조정하지 못하는 경우가 많습니다. 본 논문에서는 DataOrchestra라는 프레임워크를 제안합니다. DataOrchestra는 다양한 처리 작업을 통합하고, 각 샘플에 대한 개별적인 파이프라인을 구성합니다. 사전 학습 데이터 묶음이 주어지면, 관리자는 해당 데이터를 삭제하거나, 그대로 유지하거나, 또는 정제할지 여부를 결정합니다. 데이터가 정제될 경우, 관리자는 프로그래밍 기반 편집부터 다양한 형태의 LLM 기반 재작성까지, 여러 하위 작업을 선택합니다. 각 재작성 단계마다 구체적인 지시사항을 생성하며, 이는 해당 하위 작업 모델에 의해 실행됩니다. DataOrchestra를 통해 웹 데이터를 처리하여 0.5B에서 7B 파라미터 규모의 모델을 처음부터 학습했으며, 11개의 벤치마크에서 개별 데이터 처리 방식보다 평균적으로 성능 향상이 있음을 확인했습니다. 또한, DataOrchestra는 수학 관련 사전 학습에도 효과적이며, 더 강력한 처리 방식을 사용하는 기존 방식보다 뛰어난 성능을 보입니다. 동시에 불필요한 하위 작업을 건너뛰어 전체적인 처리 비용을 절감합니다.

Original Abstract

Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example. Given a chunk of pretraining data, an orchestrator decides whether to drop, untouch, or clean it. For a chunk to be cleaned, it selects one or more downstream operations, ranging from programmatic editing to different forms of LLM-based rewriting. For each rewriting step, it further generates a concrete instruction, which is executed by the corresponding downstream tool model. We pretrain models from 0.5B to 7B from scratch on web data processed by DataOrchestra and observe stable average gains over individual data-processing methods across 11 benchmarks. DataOrchestra is also effective for math continued pretraining and outperforms stronger processing baselines, while reducing processing compute by skipping unnecessary downstream operations.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!