What is quantization in AI?
Short answer
Quantization is a technique that makes AI models smaller and faster by storing their weights with fewer bits, for example 8-bit or 4-bit integers instead of 16- or 32-bit floating-point numbers. A quantized model needs less memory and compute, so it can run on cheaper hardware or a laptop, usually with only a small loss in quality.
Why it matters
A model with 8 billion parameters needs about 16 GB of memory at 16-bit precision. At 4-bit, it needs roughly 4–5 GB, small enough for a consumer GPU or a recent laptop. Less memory traffic also means faster responses and lower hosting costs.
The trade-off
Lower precision loses information. 8-bit quantization is usually close to indistinguishable from the original; 4-bit is good for most tasks; below that, quality drops noticeably, especially for reasoning, maths and code. Always test a quantized model on your own task rather than relying on general benchmarks.
Where you will see it
- Running models locally: tools like Ollama and llama.cpp use quantized formats such as GGUF, labelled Q4, Q5 or Q8.
- Self-hosting open models in production to cut GPU costs.
- On-device AI in phones and browsers.
Quantization changes how a model is stored, not what it knows; to change behaviour you would use fine-tuning or better prompts.
