2608.01835v1 Aug 03, 2026 cs.AI

재작성인가, 재가중화인가? 언어 모델에서의 기하학적 분석

Rewriting or Reweighting? A Geometric Account in Language Models

Juntong Wang
Juntong Wang
Citations: 0
h-index: 0
Xiyuan Wang
Xiyuan Wang
Citations: 727
h-index: 12
Muhan Zhang
Muhan Zhang
Citations: 129
h-index: 7
Shengkun Yang
Shengkun Yang
Citations: 0
h-index: 0

사후 학습은 언어 모델의 동작을 크게 변화시킬 수 있지만, 전체적인 동작 비율만으로는 학습이 기존 메커니즘을 제거하는 것인지, 새로운 메커니즘을 생성하는 것인지, 아니면 상속된 메커니즘의 사용 방식을 변경하는 것인지를 알 수 없습니다. 우리는 디코딩 과정에서 발생하는 반복 현상과 선호도 관련 정렬 실패로 인한 아첨(sycophancy)이라는 두 가지 서로 다른 유형의 오류를 통해 이 질문을 연구합니다. 우리는 행동 특이적인 기하학적 구조를 분리하기 위해, 행동과 관련된 희소 좌표를 선택하고 이를 저차원 지역 지도로 변환하는 '행동 다기관 분석'을 소개합니다. 우리는 두 개의 상호 보완적인 공간에서 이러한 지도를 구성합니다. ACT는 런타임 활성화 상태를 포착하고, NOC는 모델이 공유되는 행동 관련 부분 공간을 통해 기능 정보 흐름을 얼마나 강하게 전달하는지를 정량화합니다. 여러 모델 패밀리에 걸쳐, 결과적으로 생성된 지도는 매우 압축되어 있으며 아키텍처 간에 부분적으로 일치합니다. 기여 공간(contribution space) 지도는 더 많은 아키텍처에 적용 가능한 공유 코어를 드러내는 반면, 활성화 공간(activation space) 지도는 더욱 강력한 패밀리별 구조를 유지합니다. 제어된 사후 학습을 통해 이러한 지도를 추적하면 일관된 비대칭성이 나타납니다. 지도 학습 미세 조정은 상속된 행동 기하학적 구조를 크게 변화시키는 반면, 보상 최적화는 행동을 변경하면서도 기본 지도의 대부분을 유지합니다. 이 기하학적 관점은 두 가지 목표의 메커니즘적 차이를 이해하기 위한 통합적인 프레임워크를 제공합니다. 지도 학습 미세 조정은 주로 행동 기하학적 구조를 재작성하는 경향이 있는 반면, 보상 최적화는 주로 이를 재가중화합니다. 코드는 https://github.com/ronglingze/Manifold-Analysis 에서 확인할 수 있습니다.

Original Abstract

Post-training can substantially alter language-model behavior, yet aggregate behavior rates do not reveal whether training removes an existing mechanism, creates a new one, or changes how an inherited mechanism is used. We study this question through two mechanistically distinct failures, repetition as a decoding-attractor pathology and sycophancy as a preference-related alignment failure. We introduce behavioral manifold analysis, which isolates behavior-specific geometry by selecting sparse behavior-associated coordinates and lifting them into low-dimensional local charts. We construct these charts in two complementary spaces. ACT captures runtime activation states, while NOC quantifies how strongly the model routes functional information flow through the shared behavior-associated subspace. Across multiple model families, the resulting charts are highly compressed and partially alignable across architectures. Contribution-space charts expose a more architecture-robust shared core, whereas activation-space charts retain stronger family-specific structure. Tracking these charts through controlled post-training reveals a consistent asymmetry. Supervised fine-tuning substantially alters the inherited behavioral geometry, whereas reward optimization changes behavior while largely preserving the underlying chart. This geometric perspective provides a unified framework for understanding the mechanistic distinction between the two objectives. SFT tends to rewrite behavioral geometry, whereas reward optimization primarily reweights it. Code is available at https://github.com/ronglingze/Manifold-Analysis

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!