2606.13054v1 Jun 11, 2026 cs.LG

TWLA: 사후 양자화(Post-Training Quantization)을 통한 LLM의 삼진 가중치 및 저비트 활성화 구현

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

Xing Hu
Xing Hu
Citations: 290
h-index: 9
Zhe Jiang
Zhe Jiang
Citations: 126
h-index: 4
Zukang Xu
Zukang Xu
Citations: 189
h-index: 5
Zhixiong Zhao
Zhixiong Zhao
Citations: 21
h-index: 3
Zhixuan Chen
Zhixuan Chen
Citations: 70
h-index: 4
Dawei Yang
Dawei Yang
Citations: 30
h-index: 2

대규모 언어 모델(LLM)은 뛰어난 일반적인 언어 처리 능력을 보이지만, 메모리 및 계산 비용으로 인해 배포에 어려움이 있습니다. 삼진화는 유망한 압축 기술로, 모델 크기와 추론 복잡성을 크게 줄일 수 있습니다. 그러나 기존 방법은 꼬리가 긴 활성화 분포 문제를 해결하지 못하여 활성화를 높은 정밀도로 유지하며, 이는 전체적인 추론 속도 향상을 제한합니다. 이러한 한계를 극복하기 위해, 우리는 1.58비트의 가중치 압축과 4비트의 활성화 양자화를 달성하면서도 높은 정확도를 유지하는 사후 양자화 프레임워크인 TWLA를 제안합니다. TWLA는 다음 세 가지 구성 요소로 이루어져 있습니다: (1) Euclidean-to-Manifold Asymmetric Ternary Quantizer (E2M-ATQ)는 두 단계 최적화를 통해 유클리드 초기화에서 매니폴드로 이동하면서 가중치 삼진화 과정에서 레이어 출력 오류를 최소화합니다. (2) Kronecker Orthogonal Tri-Modal Shaping (KOTMS)는 크로네커 구조의 직교 회전을 적용하여 가중치를 삼진화에 적합한 삼모드 분포로 변환하고, 공유된 회전을 통해 활성화 값의 이상치를 통계적으로 억제합니다. (3) Inter-Layer Aware Activation Mixed Precision (ILA-AMP)는 비트 할당 시 인접 레이어 간의 이차 상호 작용 비용을 명시적으로 고려하고, 공유된 직교 변환으로 인해 발생하는 각 레이어별 활성화 양자화 이득 차이를 동시에 최적화하여 몇몇 약한 레이어가 유발할 수 있는 연쇄적인 문제를 방지합니다. 광범위한 실험 결과는 TWLA가 W1.58A4 환경에서 높은 정확도를 유지하면서도 상당한 추론 속도 향상을 제공한다는 것을 보여줍니다. 코드 및 관련 자료는 <https://github.com/Kishon-zzx/TWLA> 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment. Ternarization has emerged as a promising compression technique, offering significant reductions in model size and inference complexity. However, existing methods struggle with heavy-tailed activation distributions and therefore keep activations in high precision, fundamentally limiting end-to-end inference acceleration. To overcome this limitation, we propose TWLA, a post-training quantization (PTQ) framework that achieves 1.58-bit weight compression and 4-bit activation quantization while maintaining high accuracy. TWLA comprises three components: (1) Euclidean-to-Manifold Asymmetric Ternary Quantizer (E2M-ATQ) minimizes layer-output error under weight ternarization via a two-stage optimization from Euclidean initialization to manifold relocation; (2) Kronecker Orthogonal Tri-Modal Shaping (KOTMS) applies a Kronecker-structured orthogonal rotation to reshape weights into ternary-friendly tri-modal distributions, while the shared rotation statistically suppresses activation outliers; and (3) Inter-Layer Aware Activation Mixed Precision (ILA-AMP) explicitly introduces adjacent-layer second-order interaction costs in bit allocation and jointly optimizes for the layer-wise disparity of activation quantization gains induced by the shared orthogonal transform, preventing cascades triggered by a few weak layers. Extensive experiments demonstrate that TWLA maintains high accuracy under W1.58A4, while delivering significant inference acceleration. The code is available at <https://github.com/Kishon-zzx/TWLA>.

0 Citations
0 Influential
34.229550745277 Altmetric
0.0 Score
Original PDF
6

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!