QuantWAMs: 월드 액션 모델의 효율적인 배치를 위한 적절한 세분성으로의 양자화
QuantWAMs: Calibrating at the Right Granularity for World Action Models
월드 액션 모델(WAM)은 미래 관측값과 행동을 동시에 예측하지만, 반복적인 노이즈 제거 및 폐루프 실행 방식 때문에 효율적인 배치가 어렵습니다. 기존의 사후 훈련 양자화(PTQ) 방법은 WAM에 적합하지 않는데, 이는 개방형 루프 목표, 균일한 모델 가정, 그리고 실제 배치 환경을 반영하지 못하는 보정 분포에 의존하기 때문입니다. 본 논문에서는 모델 구조, 실행 분포, 그리고 작업 목표로 정의된 보정 컨텍스트와 양자화 결정을 일치시키는 PTQ 프레임워크인 QuantWAMs를 제안합니다. QuantWAMs는 세 가지 전략을 도입합니다: 좌표 호환 가능한 모듈 간 활성화 정보를 공유하는 아웃라이어 보정, 공동 훈련 목표에 따른 중요도를 계산하여 안정적인 보정 레이어 단위로 가중치 정밀도를 할당하는 경험적-피셔 점수 활용, 그리고 수정된 노이즈 제거 단계 보호 일정을 채택 가능한 폐루프 상태를 사용하여 정밀도 예산을 변경하지 않고 적용하는 고정 개입 실행 감사. RoboTwin 2.0, LIBERO, 그리고 AgiBot G2 로봇을 사용한 Fast-WAM 및 LingBot-VA 모델에 대한 실험 결과, W4A4 환경에서 QuantWAMs는 FP16과 비교하여 평균 성능이 0.2~0.7%p 차이를 보입니다. 실제 로봇 실험에서는 세 가지 조작 작업에서 배포 가능성을 입증했습니다. 또한, 목표로 하는 비디오 및 액션 블록에서 QuantWAMs는 가중치와 활성화 메모리 사용량을 FP16의 약 29% 수준으로 줄이고, 블록 단위 속도를 1.4~1.6배 향상시켰습니다.
World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-training quantization (PTQ) methods are poorly suited to WAMs because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment. We present QuantWAMs, a PTQ framework that aligns quantization decisions with the calibration context defined by model structure, rollout distribution, and task objective. QuantWAMs introduces three strategies: shared-basis outlier calibration, which pools activation evidence only across coordinate-compatible modules; co-training-objective saliency, which computes empirical-Fisher scores from the joint video--action gradient and assigns weight precision at a calibration-stable layer granularity; and fixed-intervention rollout auditing, which revises denoising-step protection schedules using reachable closed-loop states without changing the precision budget. We evaluate QuantWAMs on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and real-robot manipulation with an AgiBot G2. Under a W4A4-dominant setting, the reported simulation means differ from FP16 by 0.2--0.7 percentage points. Real-robot trials further establish deployment feasibility on three manipulation tasks. For the targeted video and action blocks, QuantWAMs reduces peak weight-and-activation memory to about 29\% of FP16 and provides 1.4--1.6$\times$ block-level speedups.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.