2607.08646v1 Jul 09, 2026 cs.CL

UltraX: 적응적 프로그래밍 편집을 통한 대규모 사전 학습 데이터 정제

UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

Yudong Wang
Yudong Wang
Citations: 93
h-index: 4
Hengyu Zhao
Hengyu Zhao
Citations: 88
h-index: 4
Jie Zhou
Jie Zhou
Citations: 146
h-index: 3
Zixuan Fu
Zixuan Fu
Citations: 92
h-index: 4
Xuanhe Zhou
Xuanhe Zhou
Citations: 1,754
h-index: 22
Zhiyuan Liu
Zhiyuan Liu
Citations: 84
h-index: 4
Xinlong Zhao
Xinlong Zhao
Citations: 179
h-index: 8
Xu Han
Xu Han
Citations: 106
h-index: 3
Zheng Wang
Zheng Wang
Citations: 0
h-index: 0
Jie Cai
Jie Cai
Citations: 2,135
h-index: 8
Qian Ma
Qian Ma
Citations: 0
h-index: 0
Dongsheng Liu
Dongsheng Liu
Citations: 0
h-index: 0

현재 사용 가능한 학습 데이터가 물리적인 한계에 가까워짐에 따라, 스케일링 법칙의 효과는 점차 감소하고 있습니다. 따라서, 현재 거대 언어 모델(LLM)의 성능 향상은 데이터 확장에 의존하기보다는 고품질 데이터 활용에 더 크게 좌우됩니다. 하지만, 대규모 코퍼스 환경에서 기존의 정제 방법론은 품질, 효율성 및 신뢰성 측면에서 상당한 한계를 가지고 있습니다. 규칙 기반 접근 방식은 고정된 휴리스틱으로 인해 제약을 받으며, 인스턴스 수준의 변동에 어려움을 겪습니다. LLM 기반 접근 방식은 품질을 향상시키지만, 대규모 데이터 처리에 필요한 효율성과 신뢰성 요구 사항을 충족하지 못합니다. 이러한 문제점을 해결하기 위해, 저희는 삽입 기능을 추가하여 삭제 및 수정 기능 외에도 편집 기능 공간을 확장하는 대규모 사전 학습 데이터를 위한 함수 호출 기반 정제 프레임워크인 UltraX를 제안합니다. 구체적으로, UltraX는 신뢰성 있는 프로그램 감독 생성 파이프라인을 구축합니다. 이 파이프라인에서 데이터셋에 적응된 프롬프트 최적화는 먼저 전문가 LLM을 활용하여 고품질의 전체 텍스트 정제를 수행하도록 유도하고, Line Alignment Mapping 및 Dynamic Context Replacement를 통해 원본-정제 텍스트 쌍을 구조화된 프로그램 감독으로 변환합니다. 또한, UltraX는 저신뢰 예제 필터링과 연산 조합에 의한 비율 제어 샘플링을 통해 감독 품질을 향상시키고 학습 분포를 안정화시킵니다. 추론 및 실행 과정에서, UltraX는 슬라이딩 윈도우 예측, 전역 연산 집계 및 체계적인 후처리 과정을 통해 모델 출력을 정규화하고 검증하여 대규모 실행의 안정성과 신뢰성을 향상시킵니다. 실험 결과, UltraX는 모든 코퍼스에서 가장 높은 평균 성능을 달성했으며, 더 적은 학습 토큰으로도 기존 방법과 동등하거나 뛰어넘는 성능을 보여주어 데이터 효율성과 정제 신뢰성이 우수함을 입증했습니다.

Original Abstract

As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Language Models (LLMs) now depends less on data expansion and more on higher-quality data utilization. However, in the context of large-scale corpora, existing refinement methodologies face significant limitations in quality, efficiency, and reliability: Rule-based approaches are constrained by fixed heuristics and struggle with instance-level variations; LLM-based approaches improve quality but fail to meet the efficiency and reliability requirements of large-scale data processing. To address these challenges, we propose UltraX, a function-calling refinement framework for large-scale pre-training data that completes the editing function space by introducing insertion in addition to deletion and modification, enabling fine-grained instance-level editing. Specifically, UltraX builds a reliable program-supervision generation pipeline. In this pipeline, dataset-adaptive prompt optimization first guides an expert LLM to produce high-quality end-to-end refined texts, and Line Alignment Mapping and Dynamic Context Replacement then convert original-refined text pairs into structured program supervision. Meanwhile, UltraX improves supervision quality and stabilizes the training distribution with low-confidence example filtering and ratio-controlled sampling by operation combination. During inference and execution, it normalizes and validates model outputs through sliding-window prediction, global operation aggregation, and systematic post-processing, improving the stability and reliability of large-scale execution. Experiments show that UltraX achieves the highest average performance across all corpora and also matches or surpasses baselines with fewer training tokens, demonstrating stronger data efficiency and refinement reliability.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!