LLMSurgeon: 대규모 언어 모델의 데이터 혼합 진단
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
대규모 언어 모델(LLM)의 사전 학습 데이터 혼합은 모델의 '디지털 DNA'를 구성하며, 모델의 행동, 능력 및 오류 패턴을 결정합니다. 그러나 이러한 구성 정보는 거의 공개되지 않아, 사후적으로 데이터 조합 또는 출처를 감사하는 것이 어렵습니다. 본 연구에서는 **데이터 혼합 수술(Data Mixture Surgery, DMS)**이라는 개념을 정의하고, 목표 LLM에서 생성된 텍스트만으로 사전 학습 코퍼스의 도메인 수준 분포를 미리 정의된 분류 체계에 따라 추정하는 방법을 제시합니다. 우리는 **LLMSurgeon**이라는 강력한 프레임워크를 제안하며, 이는 DMS를 레이블 이동(label-shift) 가정 하에서 역문제로 간주합니다. LLMSurgeon은 기존의 분류기 출력 단순 합산 방식이 아닌, 보정된 '소프트' 혼동 행렬을 추정하고, 체계적인 도메인 혼동을 수정하고 잠재적인 데이터 혼합 우선순위를 복원하기 위해 제약 조건이 있는 역문제 해결 방식을 사용합니다. 성능 평가를 위해, 투명한 사전 학습 데이터를 갖는 오픈 소스 LLM으로 구성된 검증 스위트인 **LLMScan**을 개발했습니다. LLMScan 환경에서, LLMSurgeon은 고정된 프로토콜 하에 높은 정확도로 도메인 혼합을 복원합니다. 본 연구는 훈련 데이터에 대한 접근 권한 없이도 기반 모델의 디지털 DNA를 감사할 수 있는 실용적인 사후 분석 방법을 제시합니다.
The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{Data Mixture Surgery (DMS)}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{LLMSurgeon}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{LLMScan}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.