2607.28048v1 Jul 30, 2026 cs.AI

SKILL-KD: LLM 에이전트를 위한 대비 학습 기반 기술 증류

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Di Weng
Di Weng
Citations: 891
h-index: 16
Zhaolu Kang
Zhaolu Kang
Citations: 30
h-index: 2
Qiming Shi
Qiming Shi
Citations: 23
h-index: 2
Yulong Tao
Yulong Tao
Citations: 0
h-index: 0
Linbo Jin
Linbo Jin
Citations: 0
h-index: 0
Yibo Dou
Yibo Dou
Citations: 0
h-index: 0
J. Zhu
J. Zhu
Citations: 75
h-index: 3
Yunfan Zhou
Yunfan Zhou
Citations: 9
h-index: 2

기술 기반 프롬프팅은 대규모 언어 모델(LLM) 에이전트의 성능을 향상시키는 실용적인 방법으로 자리 잡았습니다. 그러나 기존의 기술 습득 방법들은 종종 기술을 경험 요약, 메모 항목 또는 성공적인 시연의 직접적인 요약으로 취급합니다. 이는 성능이 낮은 학생 에이전트에 문제를 야기하는데, 학생 에이전트가 과제 지식이나 운영 전략 부족으로 실패할 경우, 실패한 과정에서 누락된 행동을 추론할 충분한 증거를 포함하지 못하는 반면, 교사 에이전트의 과정은 재사용 가능한 지침으로 내재화하기에는 너무 암묵적일 수 있습니다. 본 논문에서는 SKILL-KD라는 대비 학습 기반 기술 증류 프레임워크를 제안합니다. 이 방법은 기술을 서로 다른 능력 수준의 에이전트 간의 명시적인 증류 매개체로 취급합니다. SKILL-KD는 동일한 과제에 대한 학생 에이전트의 실패와 교사 에이전트의 과정을 기반으로, 두 과정 간의 실행 가능한 차이를 텍스트 기술 패치로 추출하고, 학생 에이전트를 다시 실행하여 패치의 효과를 평가하며, 학생 에이전트가 여전히 실패할 경우 반복적으로 패치를 개선합니다. 또한, SKILL-KD는 반복적인 국소적 업데이트로 인한 기술 편향을 방지하기 위해 수정 내역을 추적하고, Drift-Aware Skill Consolidation(편향 인지 기술 통합)을 수행하여 각 패치가 새로운 규칙을 추가해야 하는지, 기존 규칙을 삭제하거나 수정해야 하는지, 또는 건너뛰어야 하는지를 결정합니다. 본 논문에서는 5개의 에이전트 벤치마크와 두 가지 학생 에이전트 설정에서 SKILL-KD가 고정된 모델 적응 기반 방법보다 학생 에이전트의 성능을 지속적으로 향상시키는 것을 확인했습니다.

Original Abstract

Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a mismatch for weaker student agents: when a student fails because it lacks task knowledge or operational strategy, its failed trajectory may not contain enough evidence to infer the missing behavior, while the teacher trajectory may be too implicit to be internalized as reusable guidance. We propose SKILL-KD, a contrastive skill distillation framework that treats skills as an explicit distillation medium between agents of different capabilities. Given a student failure and the teacher trajectory on the same task, SKILL-KD distills their actionable discrepancy into a textual skill patch, evaluates the patch by re-running the student, and iteratively refines the patch when the student still fails. To prevent repeated local updates from causing skill drift, SKILL-KD further maintains trace-linked edit histories and performs Drift-Aware Skill Consolidation, deciding whether each patch should add a new rule, delete or modify an existing rule, or be skipped. Across five agent benchmarks and two student settings, SKILL-KD consistently improves frozen student agents over fixed-model adaptation baselines.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!