2601.03676v1 Jan 07, 2026 cs.CL

기술 분류 기반 데이터 생성 방법을 통한 LLM의 합성 일반화 향상 연구

Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis

Yifan Wei
Yifan Wei
Citations: 283
h-index: 10
Li Du
Li Du
Citations: 184
h-index: 3
Xiaoyan Yu
Xiaoyan Yu
Citations: 77
h-index: 5
Yang Feng
Yang Feng
Citations: 38
h-index: 2
Angsheng Li
Angsheng Li
Citations: 21
h-index: 2

대규모 언어 모델(LLM) 및 에이전트 기반 시스템은 종종 복잡한 기술 조합이 긴 꼬리 분포 및 파워 법칙을 따르는 데이터 병목 현상으로 인해 합성 일반화에 어려움을 겪습니다. 이는 지시 사항 준수 성능과 에이전트 중심 작업에서의 일반화를 모두 제한합니다. 이러한 문제를 해결하기 위해, 우리는 STEPS라는 기술 분류 기반 엔트로피 기반 추가 학습 데이터 생성 프레임워크를 제안합니다. STEPS는 구조 정보 이론을 사용하여 기술 간의 잠재적인 관계를 파악하고 이를 해석 가능하고 계층적인 기술 분류 체계로 구성함으로써 명시적으로 합성 일반화를 목표로 합니다. 이 분류 체계를 기반으로, 우리는 데이터 생성을 제약된 정보 최대화 문제로 정의하고, 계층 구조 내에서 주변 구조 정보를 최대화하면서 의미적 일관성을 유지하는 기술 조합을 선택합니다. 어려운 지시 사항 준수 벤치마크에 대한 실험 결과, STEPS는 기존 데이터 생성 방법보다 우수한 성능을 보였으며, 다운스트림 에이전트 기반 평가에서 향상된 합성 일반화를 달성했습니다.

Original Abstract

Large Language Models (LLMs) and agent-based systems often struggle with compositional generalization due to a data bottleneck in which complex skill combinations follow a long-tailed, power-law distribution, limiting both instruction-following performance and generalization in agent-centric tasks. To address this challenge, we propose STEPS, a Skill Taxonomy guided Entropy-based Post-training data Synthesis framework for generating compositionally challenging data. STEPS explicitly targets compositional generalization by uncovering latent relationships among skills and organizing them into an interpretable, hierarchical skill taxonomy using structural information theory. Building on this taxonomy, we formulate data synthesis as a constrained information maximization problem, selecting skill combinations that maximize marginal structural information within the hierarchy while preserving semantic coherence. Experiments on challenging instruction-following benchmarks show that STEPS outperforms existing data synthesis baselines, while also yielding improved compositional generalization in downstream agent-based evaluations.

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!