2607.27248v1 Jul 28, 2026 cs.AI

다이버전스 디코딩: 학습이 필요 없는 능력 결합 방법

Divergence Decoding: Training-Free Capability Fusion

Dechen Zhang
Dechen Zhang
Citations: 9
h-index: 1
He Cao
He Cao
Citations: 58
h-index: 3
Hao Li
Hao Li
Citations: 187
h-index: 7
Zhiyuan Yan
Zhiyuan Yan
Citations: 473
h-index: 7
Li Yuan
Li Yuan
Citations: 47
h-index: 2
Yiming Wang
Yiming Wang
Citations: 151
h-index: 5
Shuo Yang
Shuo Yang
Citations: 150
h-index: 7
Ziang Wu
Ziang Wu
Citations: 111
h-index: 2
Fan Mo
Fan Mo
Citations: 49
h-index: 3

대규모 언어 모델은 추론 능력이 뛰어나지만, 종종 특정 과학 분야에 대한 지식이 부족합니다. 반면, 해당 분야에 특화된 모델(전문가)은 전문성으로 인해 논리적 사고 능력이 저하되고 견고성이 감소하는 단점이 있습니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 학습이 필요 없는 능력 결합 프레임워크인 '다이버전스 디코딩'을 제안합니다. 다이버전스 디코딩은 추론 과정에서 발생하는 모델 간의 불일치를 감지하여, 일반적인 추론 능력을 동적으로 통합하는 적응형 라우팅 메커니즘을 활용합니다. 핵심 아이디어는 젠슨-섀넌 발산(Jensen-Shannon divergence)을 사용하여 각 토큰 단계에서 두 모델 간의 분포 차이를 모니터링하는 것입니다. 전문 모델이 상당한 불일치를 보이는 경우, 본 방법은 이를 잠재적인 추론 위험으로 판단하고 즉시 제어 권한을 일반 모델로 전환합니다. 이를 통해 일반적인 추론 능력을 유지하면서도 해당 분야의 전문성을 보존하여, 추론 시점에 일반 모델과 전문 모델의 정책을 통합할 수 있습니다. 우리는 Qwen 및 Llama 시리즈와 같은 다양한 모델 계열에서 GPQA, ChemBench, ChemCoTBench 등 다양한 과학 벤치마크를 사용하여 다이버전스 디코딩의 성능을 평가했습니다. 실험 결과는 다이버전스 디코딩이 전문화된 모델과 범용 모델 모두보다 우수한 성능을 보이며, 대부분의 단일 모델 기준 성능을 능가한다는 것을 보여줍니다. 이는 다이버전스 디코딩이 다양한 LLM의 능력을 적응적인 추론 시간 협업을 통해 결합하는 일반적이고 학습이 필요 없는 패러다임을 제공함을 시사합니다.

Original Abstract

While large language models excel in reasoning, these generalists often lack knowledge for specialized scientific domains. Conversely, domain models~(specialists), while knowledgeable, suffer from specialization side-effects including diminished logic and reduced robustness.To address this dilemma, we introduce Divergence Decoding, a training-free framework for capability fusion. It reconstructs the "draft-and-verify" skeleton of speculative decoding into an adaptive routing mechanism. The core is using Jensen-Shannon divergence to monitor the distributional disagreement between the two models at each token. When the specialist exhibits significant divergence, our method identifies it as a potential reasoning risk and instantaneously routes control to the generalist. This allows the dynamic injection of general reasoning while preserving domain expertise, achieving inference-time policy composition of the generalist and the specialist.We evaluate Divergence Decoding across diverse model families (Qwen and Llama series) on challenging scientific benchmarks (GPQA, ChemBench, and ChemCoTBench). Experimental results demonstrate that Divergence Decoding outperforms both the domain-specialized and general-purpose models, effectively surpassing the performance of most single-model baseline. This suggests that Divergence Decoding provides a general, training-free paradigm for fusing diverse LLM capabilities through adaptive inference-time collaboration.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!