2607.13425v1 Jul 15, 2026 cs.LG

어텐션 헤드 재가중치를 이용한 데이터 효율적인 LLM 적응

Data-Efficient Adaptation of LLMs via Attention Head Reweighting

Zixiao Chen
Zixiao Chen
Citations: 6
h-index: 1
C. Singh
C. Singh
Citations: 438
h-index: 11
Tuomas P. Oikarinen
Tuomas P. Oikarinen
Citations: 900
h-index: 11
Charlotte Siska
Charlotte Siska
Citations: 85
h-index: 4
Tsui-Wei Weng
Tsui-Wei Weng
Citations: 4,698
h-index: 24
Jianfeng Gao
Jianfeng Gao
Citations: 1,144
h-index: 13

제한된 데이터만으로 효과적으로 학습하는 능력은, 레이블이 첨부된 예제가 부족한 보안 분야와 같은 영역에서 매우 중요합니다. 대규모 언어 모델(LLM)은 특히 파라미터 효율적인 적응 방법을 통해 데이터 효율적인 학습 능력을 일부 보여주었지만, 어려운 작업에 대한 몇 안 되는 샘플만 주어질 경우 여전히 어려움을 겪습니다. 이러한 과제를 해결하기 위해, 우리는 어텐션 헤드 재가중치(AHR)라는 데이터 효율적인 방법을 제안합니다. AHR은 각 어텐션 헤드당 하나의 스칼라 값만을 학습하여 LLM을 새로운 텍스트 분류 작업에 적응시킵니다. 이는 개별 어텐션 헤드의 기능적 전문성을 활용하여 학습해야 하는 파라미터 수를 크게 줄입니다. 다양한 오픈 소스 텍스트 분류 데이터 세트에 대한 실험 결과, AHR은 제한된 샘플로 학습할 때 LoRA와 같은 표준 모델보다 더 우수한 성능을 보이며, 학습 가능한 파라미터가 200~1000배 적음에도 불구하고 모델의 약 0.0001%에 해당하는 파라미터만 수정합니다. 또한, 학습된 가중치는 해석이 용이하며 LLM에서 문맥 내 학습 능력에 책임 있는 메커니즘과 어텐션 헤드를 더 잘 이해하기 위해 분석할 수 있습니다.

Original Abstract

Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs) have demonstrated some capabilities for data-efficient learning, especially through parameter-efficient adaptation methods, but continue to struggle when faced with few samples for difficult tasks. To meet this challenge, we propose Attention Head Reweighting (AHR), a data-efficient method that adapts LLMs to new text-classification tasks by learning only a single scalar per attention head. This drastically reduces the number of parameters that need to be learned by making use of the functional specialization of individual attention heads. Experiments on diverse open-source text classification datasets show that AHR can outperform standard baselines like LoRA when learning from limited samples, despite having 200-1000x fewer trainable parameters, as our AHR only modifies ~0.0001% of the model's parameters. In addition, our learned weights are easy to interpret and can be analyzed to better understand the mechanisms and attention heads responsible for in-context learning abilities in LLMs.

0 Citations
0 Influential
12 Altmetric
60.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!