Sub-2-Bit and 1-Bit Quantization Reshape AI Compute
Extreme quantization techniques compressing models to under 2 bits per weight are rewriting the economics of AI infrastructure, enabling larger models to run on cheaper hardware without catastrophic accuracy loss.