2605.27115v1 May 26, 2026 cs.AI

도메인 보존과 함께 일반적인 능력 회복을 위한, 역작용 인지 다중 티처 온폴리시 증류

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

Tianle Chen
Tianle Chen
Citations: 112
h-index: 3
Ruiming Tang
Ruiming Tang
Citations: 103
h-index: 4
Jian Liang
Jian Liang
Citations: 48
h-index: 2
Jiao Ou
Jiao Ou
Citations: 78
h-index: 5
Ziyuan Liu
Ziyuan Liu
Citations: 37
h-index: 4
Han Li
Han Li
Citations: 8
h-index: 2

도메인 특화는 LLM의 수직 도메인 내 동작을 향상시킬 수 있지만, 종종 원래 모델로부터 상속된 일반적인 능력을 약화시키는 경향이 있습니다. 최근의 다중 티처 온폴리시 증류(MOPD) 파이프라인은 학생이 생성한 경로를 티처 피드백으로 감독하여 모델 능력을 회복하지만, 일반적으로 티처와 일치하는 프롬프트 커버리지를 가정하며, 이는 티처의 학습 분포와 일치하는 프롬프트가 필요합니다. 이 가정은 특히 일반적인 티처가 오픈 소스 모델이고 그 이후의 학습 데이터가 알려지지 않은 경우 충족하기 어렵습니다. 우리는 숨겨진 분포를 재구성하려고 시도하는 대신, 쉽게 구할 수 있는 프록시 일반 프롬프트를 사용하여 일반적인 능력 회복을 연구합니다. 불완전한 커버리지 상황에서 기존 MOPD의 두 가지 실패 요소를 파악했습니다: 상반되는 회복 및 보존 그래디언트가 혼합되어 발생하는 역작용으로 인한 회복-보존 균형 깨짐, 그리고 교정 요구량이 서로 다른 샘플을 균일하게 평균화하여 신호가 희석되는 현상입니다. 우리는 이러한 문제를 해결하기 위해 역작용 인지 다중 티처 온폴리시 증류(CaMOPD)를 제안합니다. CaMOPD는 분리된 교차 학습과 격차 기반 샘플 선택을 통해 일반적인 능력 회복에 특화된 업데이트를 제공하고, 주기적으로 도메인 프롬프트를 검토하여 보존을 유지하며, 평균 토큰 수준의 티처-학생 로그 확률 차이가 큰 샘플을 선택하여 교정 신호를 집중시킵니다. 역할극 대화 및 의료 추론 QA 시나리오에서 CaMOPD는 기존 방식보다 일반적인 능력 회복 측면에서 가장 우수한 성능을 보였으며, 동시에 도메인 특이적인 동작을 유지합니다. 그래디언트 일관성 분석은 CaMOPD가 더 일관된 교정 신호를 생성하는 데 의도한 효과를 뒷받침합니다.

Original Abstract

Domain specialization can improve LLM behavior in vertical domains, but often weakens the general capabilities inherited from the original model. Recent Multi-Teacher On-Policy Distillation (MOPD) pipelines recover model capabilities by supervising student-generated trajectories with teacher feedback, but typically assume teacher-aligned prompt coverage, requiring prompts to match the teachers' training distributions. This assumption is difficult to satisfy when the general teacher is an open-source model whose post-training data are unknown. Instead of attempting to reconstruct this hidden distribution, we study general capability recovery with readily available proxy general prompts. We identify two failure modes of vanilla MOPD in this incomplete-coverage situation: recovery-preservation counteraction from mixing conflicting recovery and preservation gradients, and weak-signal flattening from uniformly averaging samples with unequal correction demand. We propose Counteraction-Aware Multi-Teacher On-Policy Distillation (CaMOPD), which addresses these issues with decoupled alternating training and gap-based sample selection. CaMOPD gives general recovery dedicated updates, periodically reviews domain prompts for preservation, and selects samples with larger averaged token-level teacher-student log-probability gaps to concentrate correction signals. Across role-play dialogue and medical reasoning QA scenarios, CaMOPD performs best in general recovery over baselines while maintaining domain-specific behavior. Gradient coherence analyses further support the intended effect of CaMOPD in producing more coherent correction signals.

2 Citations
1 Influential
2.5 Altmetric
16.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!