2608.03919v1 Aug 04, 2026 cs.CV

저차원 고영향 부분 공간 최적화: 신경망 양자화를 위한 전체 파라미터 결합 학습의 한계를 넘어

Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Peng Xia
Peng Xia
Citations: 185
h-index: 4
Junbiao Pang
Junbiao Pang
Citations: 18
h-index: 2

낮은 비트 수 양자화는 특히 소형 네트워크에서 심각한 정확도 저하를 초래하는데, 이는 모든 파라미터를 동시에 훈련하는 일반적인 방식이 파라미터 부분 공간의 이질성을 고려하지 않기 때문입니다. 이러한 제한된 특징 중복성은 양자화 오차를 흡수할 여지를 거의 남기지 않습니다. 기존 방법은 단일 최적화를 사용하는데, PTQ(Post-Training Quantization)는 개선 없이 고정된 사전 훈련 모델을 재구성하고, QAT(Quantization-Aware Training)는 백본 가중치와 보정 파라미터 간의 기울기 결합으로 인해 모든 파라미터를 함께 업데이트하여 성능 저하를 야기합니다. 본 논문에서는 정규화 아핀 파라미터가 양자화 강건성을 지배하는 저차원 고영향 부분 공간임을 확인하고, 이러한 부분 공간을 대상으로 하는 Normalization Affine Preconditioning (NAP) 기법을 제안합니다. PTQ의 경우, NAP는 백본 가중치를 고정하고 전체 정밀도 모델에서 목표로 하는 가짜 양자화 그래프 하에 아핀 파라미터만 미세 조정하여, 이후 재구성을 위한 양자화 적합성을 미리 향상시킵니다. QAT의 경우, 특징 학습과 수치 보정을 분리하는 교차 QAT-NAP 방식을 도입하여 포화된 공동 훈련의 성능 한계를 극복합니다. 이론적 분석 결과, BN(Batch Normalization) 아핀 파라미터는 채널별 아핀 성분을 완전히 상쇄하며, 비선형 라운딩 및 클리핑 잔여 오류가 최소 불가피한 오차 경계를 형성합니다. 지식 증류를 활용한 NAP는 방향성 평탄화 최적화를 수행하여, 교사-학생 로그릿 불일치를 제한된 부분 공간으로 투영합니다. ImageNet과 CIFAR-100 데이터셋에 대한 실험 결과, NAP는 심각하게 성능이 저하된 낮은 비트 수 양자화를 복구하고, 재구성 기반 PTQ의 성능을 꾸준히 향상시키며, 미미한 튜닝 비용으로 포화된 전체 파라미터 QAT보다 우수한 성능을 보였습니다. 본 연구는 효율적인 심층 학습을 위한 전체 파라미터 결합 학습의 한계를 넘어, 목표 지향적인 저차원 부분 공간 최적화의 원리를 제시합니다.

Original Abstract

Low-bit quantization suffers severe accuracy degradation on compact networks, rooted in the dominant full-parameter coupled training paradigm that ignores parameter subspace heterogeneity. Their limited feature redundancy leaves little room to absorb quantization errors. Conventional pipelines adopt monolithic optimization: PTQ reconstructs fixed pretrained models without improving inherent quantization friendliness; QAT updates all parameters jointly, suffering from gradient coupling between backbone weights and calibration parameters. In this paper, we identify normalization affine parameters as a low-dimensional high-leverage subspace dominating quantization robustness, and propose Normalization Affine Preconditioning (NAP) for targeted subspace optimization. For PTQ, NAP freezes backbone weights and fine-tunes only affine parameters under the target fake-quantization graph on full-precision models, proactively boosting quantization friendliness before downstream reconstruction. For QAT, we introduce an alternating QAT-NAP schema that decouples feature learning and numerical calibration, breaking the performance ceiling of saturated joint training. Theoretical analysis confirms BN affine parameters fully cancel the channel-wise affine component of quantization distortion, while nonlinear rounding and clipping residuals form the irreducible error boundary; distillation-guided NAP acts as directional flatness optimization, projecting teacher-student logit mismatch onto the restricted subspace. Experiments on ImageNet and CIFAR-100 show NAP recovers severely collapsed low-bit quantization, consistently boosts reconstruction-based PTQ, and outperforms saturated full-parameter QAT with negligible tuning cost. This work reveals the principle of targeted low-dimensional subspace optimization, offering a new perspective beyond full-parameter coupled training for efficient deep learning.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!