2606.05682v1 Jun 04, 2026 cs.AI

NVFP4 LLM 증류 과정에서 출력 매칭을 넘어: 내부 기하 구조 보존

Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillatio

C. Liu
C. Liu
Citations: 46
h-index: 1
Fang Tu
Fang Tu
Citations: 3
h-index: 1
Srinivasan Manoharan
Srinivasan Manoharan
Citations: 1
h-index: 1
Junhua Zhao
Junhua Zhao
Citations: 50
h-index: 3
Xin Chen
Xin Chen
Citations: 468
h-index: 4
Haifeng Wu
Haifeng Wu
Citations: 7
h-index: 2
Jianmin Wan
Jianmin Wan
Citations: 876
h-index: 17

대규모 언어 모델이 지연 시간 및 비용 제약 환경에서 점점 더 많이 사용됨에 따라, NVFP4 기반 접근 방식을 포함한 저정밀 추론에 대한 수요가 증가하고 있습니다. 양자화 인식 증류(QAD)는 낮은 비트 양자화로 인해 발생하는 정확도 손실을 복구하기 위해, 훈련된 양자화 모델(학생 모델)이 고정된 더 높은 정밀도의 모델(교사 모델)의 출력 분포를 KL-발산 손실을 통해 일치시키도록 합니다. 본 연구에서는 먼저 QAD의 표현 수준 진단을 수행했습니다. 출력 매칭만으로는 내부 성능 저하를 숨길 수 있습니다. 왜냐하면 많은 중간 활성화 기하 구조가 유사한 교사 모델 정규화 로그 값을 생성할 수 있기 때문입니다. CKA(Contrastive Language-Aware)를 사용하여 KL-전용 QAD가 BF16 교사 모델에 비해 계층별 표현 유사성을 감소시키고, 특히 강화 학습으로 추가 훈련된 모델에서 심각한 성능 저하를 유발한다는 것을 확인했습니다. 이러한 성능 저하는 추론 및 코딩 작업에서의 병목 현상과 관련이 있으며, 이는 저비트 복구에 있어서 출력 매칭뿐만 아니라 내부 기하 구조를 보존하는 것이 중요하다는 것을 시사합니다. 이러한 발견에 따라, NVFP4 QAD 및 저비트 LLM 정확도 복구를 위한 CKA 기반 표현 정렬 방법인 extbf{CKA-QAD}를 제안합니다. 이 방법은 CKA를 통해 계층별 Gram 행렬을 정렬시켜 증류 과정에서 내부 표현 기하 구조를 보존하는 가벼운 규제항을 추가합니다. Nemotron 3 Nano 및 Qwen3-4B-Thinking-2507 모델에 대해 CKA-QAD는 표현 정렬을 크게 개선하고, 적절한 훈련 오버헤드로 다운스트림 추론 및 코딩 정확도를 향상시켰습니다. 우리의 연구 결과는 양자화된 LLM 복구를 위한 출력 매칭의 실용적인 보완 방법으로 CKA 기반 표현 정렬의 중요성을 강조합니다.

Original Abstract

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production environments. Quantization-aware distillation (QAD) helps recover accuracy lost under low bit quantization by training a quantized student to match the output distribution of a frozen higher precision teacher via a KL-divergence loss. In this work, we first provide a representation level diagnosis of QAD: output matching alone can mask internal degradation, because many intermediate activation geometries can yield similar teacher-aligned logits. Using CKA, we show that KL-only QAD can reduce layerwise representational similarity relative to the BF16 teacher, with especially severe drift in RL-post-trained models. This drift correlates with downstream bottlenecks on reasoning and coding tasks, suggesting that low bit recovery requires preserving internal geometry rather than matching outputs alone. Motivated by this finding, we propose \textbf{CKA-QAD}, a CKA-guided representational alignment method for NVFP4 QAD and low bit LLM accuracy recovery. The method adds a lightweight regularizer that preserves internal representational geometry during distillation by aligning layerwise Gram matrices through CKA. Across Nemotron 3 Nano and Qwen3-4B-Thinking-2507, CKA-QAD substantially improves representational alignment and improves downstream reasoning and coding accuracy with modest training overhead. Our findings position CKA-guided representational alignment as a practical complement to output matching for quantized LLM recovery.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!