Neural Information Processing Systems (NeurIPS 2024 Workshop on Edge AI)
·
2024
Sub-Bit Precision Quantization and Weight-Activation Slicing for Edge Transformer Inference
Deploying large transformer models on embedded hardware is severely bottlenecked by memory bandwidth and thermal throttling. We introduce WAS-Quant (Weight-Activation Slicing Quantization), an asymmetric 3.2-bit quantization algorithm that preserves perplexity within 0.12 of FP16 baselines while reducing KV-cache footprint by 68%.