Joan Gregorio Pérez - Publicación Científica Título: Sub-Bit Precision Quantization and Weight-Activation Slicing for Edge Transformer Inference Año: 2024 Venue: Neural Information Processing Systems (NeurIPS 2024 Workshop on Edge AI) DOI: 10.5555/neurips.2024.edge.104 Abstract: Deploying large transformer models on embedded hardware is severely bottlenecked by memory bandwidth and thermal throttling. We introduce WAS-Quant (Weight-Activation Slicing Quantization), an asymmetric 3.2-bit quantization algorithm that preserves perplexity within 0.12 of FP16 baselines while reducing KV-cache footprint by 68%. BibTeX: @inproceedings{pérez2024subbit, author = Pérez, Joan Gregorio and Silva, Marco Aurelio, title = {Sub-Bit Precision Quantization and Weight-Activation Slicing for Edge Transformer Inference}, booktitle = {Neural Information Processing Systems (NeurIPS 2024 Workshop on Edge AI)}, year = 2024, doi = {10.5555/neurips.2024.edge.104}, eprint = {2409.08311}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, url = {https://joangregorioperez.com/papers/latency-optimized-quantization-edge-transformers}, abstract = {Deploying large transformer models on embedded hardware is severely bottlenecked by memory bandwidth and thermal throttling. We introduce WAS-Quant (Weight-Activation Slicing Quantization), an asymmetric 3.2-bit quantization algorithm that preserves perplexity within 0.12 of FP16 baselines while red...}, }