2607.28109v1 Jul 30, 2026 cs.AI

재구성 이상의 의미: 책 수준의 구성이 중간 단계 학습을 위한 합성 교과서 데이터 개선에 미치는 영향

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

Wenhan Yu
Wenhan Yu
Citations: 10
h-index: 2
Xiaokun Yuan
Xiaokun Yuan
Citations: 3
h-index: 1
Tong Yang
Tong Yang
Citations: 43
h-index: 3
Nuo Chen
Nuo Chen
Citations: 582
h-index: 12
Mengzhou Wu
Mengzhou Wu
Citations: 79
h-index: 4
Ji-Hua Tao
Ji-Hua Tao
Citations: 599
h-index: 4
Guoan Wang
Guoan Wang
Citations: 277
h-index: 4
Miao Peng
Miao Peng
Citations: 15
h-index: 1
Yaoming Li
Yaoming Li
Citations: 19
h-index: 3
Maxm Pan
Maxm Pan
Citations: 0
h-index: 0

합성 교과서 데이터는 언어 모델 사전 학습 성능 향상에 기여했지만, 기존 연구에서는 이러한 효과를 주로 생성된 콘텐츠 자체 또는 로컬 재작성 스타일의 특성으로 간주했습니다. 본 연구에서는 관련 콘텐츠가 일관된 책 수준의 문서로 구성되는지 여부라는 다른 요인을 탐구합니다. 우리는 확장 가능한 합성 파이프라인과 함께 이 구성이 중요하다는 통제된 증거를 제시합니다. 파이프라인은 사전 학습 코퍼스에서 소스 자료를 검색하고, 주제별 단위로 클러스터링하며, 계층적 목차를 계획하고, 소스 기반 섹션을 완전한 책으로 조립합니다(Full 설정). 이를 통해 15,000개 이상의 분야에 걸쳐 686,000개의 교과서(32B 토큰)가 생성되었습니다. 중간 단계 학습 데이터셋에서 자연 교과서를 이 코퍼스로 대체하면 평균 +1.09의 성능 향상을 보입니다. 통제된 비교 실험을 통해 관련된 설계 요인을 분리합니다. 콘텐츠 일치 조건을 사용하여 생성된 텍스트와 토큰은 고정하고 각 섹션을 독립적인 문서로 처리하면, Full 설정의 +1.02만큼의 평균 성능 향상은 문서 패키징의 효과를 나타냅니다. 길이 일치 조건에서 서로 다른 책의 섹션을 무작위로 연결하는 방식은 Full 설정을 능가하지 못하며, 이는 문서 길이만으로는 설명할 수 없음을 시사합니다. 검색 풀 일치 조건을 사용하여 개별적으로 검색된 문서를 동일한 청중 및 스타일 체계를 적용하여 재작성하지만 클러스터링, 목차 계획 또는 책 조립은 수행하지 않습니다. Full 설정의 +1.17만큼의 성능 향상은 구조화된 합성 방식의 가치를 입증합니다. Llama3-8B 모델에서 Full 설정은 무작위 연결 방식과 자연 교과서 모두보다 우수한 성능을 보이며, 이는 책 수준의 구성이 합성 사전 학습 데이터 설계에 유용한 요소임을 뒷받침합니다.

Original Abstract

Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a different factor: whether related content is organized into coherent book-level documents. We contribute both a scalable synthesis pipeline and controlled evidence that this organization matters. The pipeline retrieves source material from a pre-training corpus, clusters it into topical units, plans hierarchical tables of contents, and assembles source-grounded sections into complete books (our Full setting), yielding 686K textbooks (32B tokens) across 15,000+ disciplines. Replacing natural books in a mid-training mix with this corpus improves downstream performance by +1.09 on average. Controlled comparisons then disentangle the relevant design factors. A content-matched Split condition holds generated text and tokens fixed but treats each section as an independent document; Full's +1.02 mean gain isolates document packaging. A length-matched RandomConcat control that joins sections from different books remains below Full, ruling out document length alone. A retrieval-pool-matched Rephrase condition independently rewrites individual retrieved documents under the same audience-by-style scheme, without clustering, TOC planning, or book assembly; Full's +1.17 gain demonstrates the value of structured synthesis. On Llama3-8B, Full likewise outperforms both RandomConcat and Natural Books, supporting book-level organization as a useful axis for synthetic pre-training data design.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!