2605.30039v1 May 28, 2026 cs.AI

최소 충분 표현 학습을 통한 LLM의 도메인 특화 데이터 합성

Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning

Peng Di
Peng Di
Citations: 55
h-index: 5
Jianwei Yin
Jianwei Yin
Citations: 312
h-index: 10
Tong Ye
Tong Ye
Citations: 71
h-index: 4
Hang Yu
Hang Yu
Citations: 437
h-index: 10
Tengfei Ma
Tengfei Ma
Citations: 67
h-index: 4
Pei Liu
Pei Liu
Citations: 6
h-index: 2
Wenhai Wang
Wenhai Wang
Citations: 272
h-index: 11
Xuhong Zhang
Xuhong Zhang
Citations: 39
h-index: 1
Jianguo Li
Jianguo Li
Citations: 518
h-index: 13

대규모 언어 모델(LLM)은 일반적인 능력에서 뛰어난 발전을 보여왔으며, 도메인 특화 데이터를 활용한 미세 조정(fine-tuning)을 통해 특정 분야에서도 강력한 성능을 달성할 수 있습니다. 그러나 목표 도메인에 대한 고품질 데이터 확보는 여전히 중요한 과제입니다. 기존의 데이터 합성 방법은 주로 연역적 패러다임을 따르며, 자연어 표현으로 명시적으로 설명된 도메인 정보와 신중한 프롬프트 엔지니어링에 크게 의존하는데, 이는 실제 시나리오에서 도메인을 설명하거나 공식적으로 정의하기 어려운 경우 적용 가능성이 제한됩니다. 본 연구에서는 자연어 설명이 어렵거나 형식적인 정의가 불가능한 경우에도 참조 예제를 통해 정의되는 유도적 패러다임을 사용하여 아직 탐구되지 않은 도메인 특화 데이터 합성 문제를 다룹니다. 우리는 DOMINO라는 새로운 프레임워크를 제안합니다. DOMINO는 참조 샘플로부터 최소 충분한 도메인 표현을 학습하고, 이를 활용하여 도메인과 일치하는 합성 데이터를 생성하도록 안내합니다. DOMINO는 프롬프트 튜닝과 대조적인 분리 목표(contrastive disentanglement objective)를 통합하여 도메인 수준의 패턴을 샘플별 노이즈로부터 분리함으로써 과적합을 방지하면서 핵심 도메인 특성을 유지합니다. 이론적으로, DOMINO는 합성 데이터 분포의 지지를 확장하여 더 큰 다양성을 보장한다는 것을 증명했습니다. 실험적으로, 도메인 정의가 암시적인 어려운 코딩 벤치마크에서, DOMINO에 의해 생성된 데이터를 사용하여 미세 조정을 수행했을 때 Pass@1 정확도가 강력하고 명령형으로 조정된 기본 모델(backbone)을 사용하는 경우보다 최대 4.63% 향상되었습니다. 이는 DOMINO의 효과성과 견고성을 입증합니다. 본 연구는 수동 프롬프트 설계나 자연어 도메인 사양 없이도 실용적이고 확장 가능한 도메인 적응을 가능하게 하는 새로운 패러다임을 제시합니다.

Original Abstract

Large Language Models have demonstrated remarkable progress in general-purpose capabilities and can achieve strong performance in specific domains through fine-tuning on domain-specific data. However, acquiring high-quality data for target domains remains a significant challenge. Existing data synthesis approaches follow a deductive paradigm, heavily relying on explicit domain descriptions expressed in natural language and careful prompt engineering, limiting their applicability in real-world scenarios where domains are difficult to describe or formally articulate. In this work, we tackle the underexplored problem of domain-specific data synthesis through an inductive paradigm, where the target domain is defined only through a set of reference examples, particularly when domain characteristics are difficult to articulate in natural language. We propose a novel framework, DOMINO, that learns a minimal sufficient domain representation from reference samples and leverages it to guide the generation of domain-aligned synthetic data. DOMINO integrates prompt tuning with a contrastive disentanglement objective to separate domain-level patterns from sample-specific noise, mitigating overfitting while preserving core domain characteristics. Theoretically, we prove that DOMINO expands the support of the synthetic data distribution, ensuring greater diversity. Empirically, on challenging coding benchmarks where domain definitions are implicit, fine-tuning on data synthesized by DOMINO improves Pass@1 accuracy by up to 4.63\% over strong, instruction-tuned backbones, demonstrating its effectiveness and robustness. This work establishes a new paradigm for domain-specific data synthesis, enabling practical and scalable domain adaptation without manual prompt design or natural language domain specifications.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!