2602.16189v1 Feb 18, 2026 cs.CL

학습을 넘어: 모델 적응을 위한 훈련 불필요한 대안

Beyond Learning: A Training-Free Alternative to Model Adaptation

Namkyung Yoon
Namkyung Yoon
Citations: 62
h-index: 5
Kyeonghyun Yoo
Kyeonghyun Yoo
Citations: 1
h-index: 1
Wooyong Jung
Wooyong Jung
Citations: 48
h-index: 3
Sanghong Kim
Sanghong Kim
Citations: 1
h-index: 1
Hwangnam Kim
Hwangnam Kim
Citations: 91
h-index: 6

언어 모델에 대한 지속적인 연구와 발전에도 불구하고, 때로는 이전 버전보다 성능이 저하되는 경우가 있습니다. 이러한 문제를 해결하기 위한 기존 방법들은 자원 집약적인 경향이 있으며, 즉각적인 대응을 가능하게 하는 대안의 필요성이 강조됩니다. 본 연구는 각 언어 모델 내부에 특정 기능에 적합한 로컬 모듈이 존재한다는 가정하에 진행되었습니다. 먼저, 활성화 기반 분석을 통해 추론 작업 환경에서 일관되고 국소적인 활성화 변화를 보이는 모듈들을 식별합니다. 그 후, 특정 작업에 적절하게 활성화된 내부 모듈을 대상 모델에 이식하여, 추가적인 훈련이나 미세 조정 없이 즉각적이고 측정 가능한 기능적 변화를 유도합니다. 이식 기술의 효과를 실험적으로 입증하기 위해, 두 개의 언어 모델에 대해 다양한 조건에서 이식 강도와 성능 향상 간의 관계를 정량적으로 분석했습니다. 교차 생성 설정에서, 활성화 기반으로 선택된 모듈을 이식하면 성능이 저하된 모델을 크게 개선할 수 있으며, 최대 2배의 목표 기준 성능을 달성하고 100% 이상의 격차 기반 회복률을 보였습니다. 또한, 기본 모델과 instruction-tuned 모델 간의 이식 실험에서, 이식은 성능이 저하된 모델을 더 강력한 기준 모델에 더욱 가깝게 만들었으며, 최대 약 2.33배의 목표 기준 성능을 달성하고, 최상의 경우 100%의 격차 기반 회복률을 보였습니다. 이러한 결과는 언어 모델 내에 내재된 고도로 국소화된 모듈을 활용하여 의미 있는 용량 이전이 가능하다는 것을 보여줍니다. 전반적으로, 본 연구는 언어 모델의 작업별 모듈성을 뒷받침하는 경험적 증거를 제시하고, 모델 이식이라는 새로운 연구 분야를 제시합니다.

Original Abstract

Despite the continuous research and evolution of language models, they sometimes underperform previous versions. Existing approaches to overcome these challenges are resource-intensive, highlighting the need for alternatives that enable immediate action. We assume that each language model has a local module inside that is suitable for a specific function. First, this work identifies a set of modules showing consistent and local activation changes under an inference workload through activation-based analysis. Subsequently, we transplant an internal module that is properly activated for a specific task into the target model, leading to immediate and measurable functional changes without additional training or fine-tuning. To experimentally demonstrate the effectiveness of the transplant technique, we quantify the relationship between transplant strength and performance improvement under different conditions for two language models. In the cross-generation setting, we find that transplanting activation-selected modules can substantially improve the underperforming model, reaching up to twice the target baseline and achieving gap-based recovery above 100%. Moreover, in transplant experiments between a base model and its instruction-tuned counterpart, transplantation improves the underperforming model toward the stronger baseline, yielding up to about 2.33 times the target baseline with gap-based recovery reaching up to 100% in the best case. These results show that meaningful capacity transfer can be realized through the implantation of highly localized modules implied by language models. Overall, this work provides empirical evidence for task-localized modularity in language models and presents a new research area: model transplantation.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!