2602.15799v1 Feb 17, 2026 cs.LG

정렬 붕괴의 기하학: 미세 조정이 안전성을 저하시키는 시점

The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety

M. Springer
M. Springer
Citations: 3
h-index: 1
Chungwoo Lee
Chungwoo Lee
Citations: 21
h-index: 3
Blossom Metevier
Blossom Metevier
Citations: 197
h-index: 6
J. Castleman
J. Castleman
Citations: 14
h-index: 3
Bohdan Turbal
Bohdan Turbal
Citations: 7
h-index: 2
Hayoung Jung
Hayoung Jung
University of Washington
Citations: 65
h-index: 4
Zeyu Shen
Zeyu Shen
Citations: 82
h-index: 4
Aleksandra Korolova
Aleksandra Korolova
Citations: 88
h-index: 5

정렬된 언어 모델을 양성적인 작업에 대해 미세 조정하면 예측 불가능하게 안전 장치가 저하됩니다. 이는 학습 데이터에 유해한 내용이 포함되어 있지 않고 개발자가 적대적인 의도를 가지고 있지 않더라도 발생합니다. 현재의 주된 설명은 미세 조정 업데이트가 고차원 파라미터 공간에서 안전에 중요한 방향과 직교해야 한다는 것이지만, 이는 잘못된 안심을 제공합니다. 우리는 이러한 직교성이 구조적으로 불안정하며 경사 하강법의 역학에 의해 붕괴된다는 것을 보여줍니다. 우리는 새로운 기하학적 분석을 통해 이를 해결하고, 정렬이 급격한 곡률을 가진 저차원 부분 공간에 집중되어 있어, 1차 방법으로는 감지하거나 방어할 수 없는 취약한 구조를 형성한다는 것을 증명합니다. 초기 미세 조정 업데이트는 이러한 부분 공간을 피할 수 있을 수 있지만, 미세 조정 손실의 곡률은 2차 가속을 생성하여 체계적으로 경로를 정렬에 민감한 영역으로 유도합니다. 우리는 이 메커니즘을 '정렬 불안정 조건'을 통해 형식화하며, 이 세 가지 기하학적 속성이 동시에 충족되면 안전 저하가 발생합니다. 우리의 주요 결과는 4차 스케일링 법칙을 제시합니다. 즉, 정렬 손실은 학습 시간의 4승에 비례하며, 이는 정렬 기하학의 날카로움과 미세 조정 작업과 안전에 중요한 파라미터 간의 곡률 결합 강도에 의해 결정됩니다. 이러한 결과는 현재 안전 패러다임의 구조적 결함을 드러냅니다. 안전한 미세 조정에 대한 주요 접근 방식은 근본적으로 동적인 문제의 초기 스냅샷만을 다룹니다. 정렬의 취약성은 수정해야 할 버그가 아니라, 곡면 다양체에 대한 경사 하강법의 고유한 기하학적 속성입니다. 우리의 결과는 곡률을 고려하는 방법 개발을 촉진하며, 정렬 안전 분석이 반응적인 레드 팀 활동에서 예측적인 진단으로 전환될 수 있기를 바랍니다. 이는 특히 개방형 가중 모델 배포에 중요한 역할을 할 것입니다.

Original Abstract

Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial intent. We show that the prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show this orthogonality is structurally unstable and collapses under the dynamics of gradient descent. We then resolve this through a novel geometric analysis, proving that alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend. While initial fine-tuning updates may indeed avoid these subspaces, the curvature of the fine-tuning loss generates second-order acceleration that systematically steers trajectories into alignment-sensitive regions. We formalize this mechanism through the Alignment Instability Condition, three geometric properties that, when jointly satisfied, lead to safety degradation. Our main result establishes a quartic scaling law: alignment loss grows with the fourth power of training time, governed by the sharpness of alignment geometry and the strength of curvature coupling between the fine-tuning task and safety-critical parameters. These results expose a structural blind spot in the current safety paradigm. The dominant approaches to safe fine-tuning address only the initial snapshot of a fundamentally dynamic problem. Alignment fragility is not a bug to be patched; it is an intrinsic geometric property of gradient descent on curved manifolds. Our results motivate the development of curvature-aware methods, and we hope will further enable a shift in alignment safety analysis from reactive red-teaming to predictive diagnostics for open-weight model deployment.

4 Citations
0 Influential
3 Altmetric
19.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!