2602.10134v1 Feb 07, 2026 cs.CR

언어 모델의 역공학적 모델 편집 연구

Reverse-Engineering Model Editing on Language Models

Zhiyu Sun
Zhiyu Sun
Citations: 50
h-index: 2
Minrui Luo
Minrui Luo
Citations: 4
h-index: 1
Yu Wang
Yu Wang
Citations: 271
h-index: 6
Zhili Chen
Zhili Chen
Citations: 6
h-index: 2
Tianxing He
Tianxing He
Citations: 38
h-index: 3

대규모 언어 모델(LLM)은 수조 개의 토큰을 포함하는 데이터셋으로 사전 훈련되므로, 필연적으로 민감한 정보를 암기하게 됩니다. 모델 편집의 주요 패러다임인 '찾아서 편집' 방식은 모델을 재훈련하지 않고도 모델 파라미터를 수정하여 유망한 해결책을 제시합니다. 그러나 본 연구에서는 이러한 패러다임의 중요한 취약점을 밝혀냅니다. 즉, 파라미터 업데이트가 의도치 않게 사이드 채널 역할을 하여 공격자가 편집된 데이터를 복구할 수 있게 된다는 것입니다. 우리는 이러한 업데이트의 저랭크 구조를 활용하는 두 단계의 역공학 공격인 extit{KSTER} ( extbf{K}ey extbf{S}paceRecons extbf{T}ruction-then- extbf{E}ntropy extbf{R}eduction)을 제안합니다. 첫째, 업데이트 행렬의 열 공간이 편집된 대상의 "지문"을 포함하며, 이를 스펙트럼 분석을 통해 정확하게 복구할 수 있음을 이론적으로 보여줍니다. 둘째, 편집의 의미적 맥락을 재구성하는 엔트로피 기반 프롬프트 복구 공격을 소개합니다. 여러 LLM에 대한 광범위한 실험 결과, 우리의 공격은 높은 성공률로 편집된 데이터를 복구할 수 있음을 보여줍니다. 또한, 업데이트 지문을 의미적 속임수로 가려 재구성 위험을 완화하는 방어 전략인 extit{subspace camouflage}를 제안합니다. 이 접근 방식은 편집 유용성을 손상시키지 않고 재구성 위험을 효과적으로 줄입니다. 저희의 코드는 https://github.com/reanatom/EditingAtk.git 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) are pretrained on corpora containing trillions of tokens and, therefore, inevitably memorize sensitive information. Locate-then-edit methods, as a mainstream paradigm of model editing, offer a promising solution by modifying model parameters without retraining. However, in this work, we reveal a critical vulnerability of this paradigm: the parameter updates inadvertently serve as a side channel, enabling attackers to recover the edited data. We propose a two-stage reverse-engineering attack named \textit{KSTER} (\textbf{K}ey\textbf{S}paceRecons\textbf{T}ruction-then-\textbf{E}ntropy\textbf{R}eduction) that leverages the low-rank structure of these updates. First, we theoretically show that the row space of the update matrix encodes a ``fingerprint" of the edited subjects, enabling accurate subject recovery via spectral analysis. Second, we introduce an entropy-based prompt recovery attack that reconstructs the semantic context of the edit. Extensive experiments on multiple LLMs demonstrate that our attacks can recover edited data with high success rates. Furthermore, we propose \textit{subspace camouflage}, a defense strategy that obfuscates the update fingerprint with semantic decoys. This approach effectively mitigates reconstruction risks without compromising editing utility. Our code is available at https://github.com/reanatom/EditingAtk.git.

1 Citations
0 Influential
23 Altmetric
6.9 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!