2608.01624v1 Aug 03, 2026 cs.CL

차원이 아닌 규범: 언어 모델의 기울기 기반 가중치 변동에서 중요한 것은 무엇인가

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

Taehyeon Kim
Taehyeon Kim
KAIST
Citations: 686
h-index: 10
Unggi Lee
Unggi Lee
Citations: 484
h-index: 11
Taeyeon Kim
Taeyeon Kim
Citations: 32
h-index: 2
Ahhyun Kim
Ahhyun Kim
Citations: 5
h-index: 1

언어 모델을 특정 작업에 맞게 조정할 때 더 이상 모든 가중치를 학습시킬 필요가 없으며, 파라미터 효율적인 방법들이 학습해야 할 파라미터 수를 수십억 개에서 소수의 스칼라 값으로 줄이는 데 기여했습니다. 기울기 기반 변동 방법을 사용하지 않고 무작위로 가중치 변동을 샘플링하고 성능이 좋은 것을 선택하는 방식은 이러한 추세를 따르지 못하고 여전히 가중치 텐서의 모든 요소를 변동시킵니다. 전체 가중치 검색이 필수적인지는 불확실하며, 더 근본적으로 어떤 특성이 변동을 효과적으로 만드는지는 기존 방법들이 검색 공간, 변동 크기 및 집계 방식을 동시에 변경하기 때문에 명확하지 않습니다. 본 연구에서는 고정된 파이프라인 내에서 한 번에 하나의 요소를 변화시키면서 후보 점수 및 투표는 일정하게 유지하면서, 검색 차원, 변동을 전달하는 부분 공간, 그리고 그 규범을 변화시켜 이 문제를 해결합니다. 12~16개의 스칼라 값으로 고정된 프레임을 사용하는 것은 평균적으로 49개의 모델-벤치마크 조합에서 전체 가중치 검색보다 1.8점 정도의 정확도 손실을 보이며, 그 중 36개 조합에서는 성능이 더 낮습니다. 이러한 성능 차이는 차원이나 기저 선택에 의해 설명되지 않습니다. 특이값 분해(SVD) 프레임과의 그래스만 중첩도가 무작위 수준인 프레임은 단일 스케일 인자를 조정하면 동일한 성능을 보이며, 큰 스케일에서는 SVD 방향이 먼저 사라집니다. 생존하는 것은 변동 규범이며, 그 유효 범위는 7개의 모델에서 최대 다섯 배 이내로 좁혀지며 내부적으로 일정하게 유지됩니다. 따라서 변동 규범은 실패 요인을 가진 유일한 요소이며, 안전 영역은 스케일과 모델 계열에 걸쳐 전송될 수 있습니다. 설계 질문은 어떤 부분 공간을 변동시킬 것인가에서 얼마나 강하게 변동시킬 것인가로 좁혀집니다.

Original Abstract

Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a handful of scalars. Gradient-free adaptation, which samples random weight perturbations and keeps the ones that score well, has not followed that trajectory and still perturbs every entry of the weight tensor. It is unknown whether that full-weight search is necessary, and more fundamentally which property of a perturbation makes it work at all, because existing methods vary the search space, the perturbation scale, and the aggregation together. We resolve this by intervening on one factor at a time inside a fixed pipeline, holding candidate scoring and voting constant while we vary the search dimension, the subspace that carries the perturbation, and its norm. Perturbing a frozen frame of 12 to 16 scalars stays 1.8 accuracy points behind full-weight search on average across 49 model-benchmark cells, trailing it in 36 of them. Neither the dimension nor the choice of basis explains that performance. A random frame whose Grassmann overlap with the SVD frame is at chance level performs identically once a single scale factor is matched, and at large scales the SVD directions collapse first. What survives is the perturbation norm, whose usable range closes within a factor of five across seven models and stays flat inside. The perturbation norm is therefore the one factor with a failure mode, and its safe region transfers across scale and family. The design question narrows from which subspace to perturb to how hard to shake.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!