정렬 붕괴의 기하학: 미세 조정이 안전성을 저하시키는 시점
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
정렬된 언어 모델을 양성적인 작업에 대해 미세 조정하면 예측 불가능하게 안전 장치가 저하됩니다. 이는 학습 데이터에 유해한 내용이 포함되어 있지 않고 개발자가 적대적인 의도를 가지고 있지 않더라도 발생합니다. 현재의 주된 설명은 미세 조정 업데이트가 고차원 파라미터 공간에서 안전에 중요한 방향과 직교해야 한다는 것이지만, 이는 잘못된 안심을 제공합니다. 우리는 이러한 직교성이 구조적으로 불안정하며 경사 하강법의 역학에 의해 붕괴된다는 것을 보여줍니다. 우리는 새로운 기하학적 분석을 통해 이를 해결하고, 정렬이 급격한 곡률을 가진 저차원 부분 공간에 집중되어 있어, 1차 방법으로는 감지하거나 방어할 수 없는 취약한 구조를 형성한다는 것을 증명합니다. 초기 미세 조정 업데이트는 이러한 부분 공간을 피할 수 있을 수 있지만, 미세 조정 손실의 곡률은 2차 가속을 생성하여 체계적으로 경로를 정렬에 민감한 영역으로 유도합니다. 우리는 이 메커니즘을 '정렬 불안정 조건'을 통해 형식화하며, 이 세 가지 기하학적 속성이 동시에 충족되면 안전 저하가 발생합니다. 우리의 주요 결과는 4차 스케일링 법칙을 제시합니다. 즉, 정렬 손실은 학습 시간의 4승에 비례하며, 이는 정렬 기하학의 날카로움과 미세 조정 작업과 안전에 중요한 파라미터 간의 곡률 결합 강도에 의해 결정됩니다. 이러한 결과는 현재 안전 패러다임의 구조적 결함을 드러냅니다. 안전한 미세 조정에 대한 주요 접근 방식은 근본적으로 동적인 문제의 초기 스냅샷만을 다룹니다. 정렬의 취약성은 수정해야 할 버그가 아니라, 곡면 다양체에 대한 경사 하강법의 고유한 기하학적 속성입니다. 우리의 결과는 곡률을 고려하는 방법 개발을 촉진하며, 정렬 안전 분석이 반응적인 레드 팀 활동에서 예측적인 진단으로 전환될 수 있기를 바랍니다. 이는 특히 개방형 가중 모델 배포에 중요한 역할을 할 것입니다.
Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial intent. We show that the prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show this orthogonality is structurally unstable and collapses under the dynamics of gradient descent. We then resolve this through a novel geometric analysis, proving that alignment concentrates in low-dimensional subspaces with sharp curvature, creating a brittle structure that first-order methods cannot detect or defend. While initial fine-tuning updates may indeed avoid these subspaces, the curvature of the fine-tuning loss generates second-order acceleration that systematically steers trajectories into alignment-sensitive regions. We formalize this mechanism through the Alignment Instability Condition, three geometric properties that, when jointly satisfied, lead to safety degradation. Our main result establishes a quartic scaling law: alignment loss grows with the fourth power of training time, governed by the sharpness of alignment geometry and the strength of curvature coupling between the fine-tuning task and safety-critical parameters. These results expose a structural blind spot in the current safety paradigm. The dominant approaches to safe fine-tuning address only the initial snapshot of a fundamentally dynamic problem. Alignment fragility is not a bug to be patched; it is an intrinsic geometric property of gradient descent on curved manifolds. Our results motivate the development of curvature-aware methods, and we hope will further enable a shift in alignment safety analysis from reactive red-teaming to predictive diagnostics for open-weight model deployment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.