Glossary · AI Engineering

What is quantization in AI?

Short answer

Quantization is a technique that makes AI models smaller and faster by storing their weights with fewer bits, for example 8-bit or 4-bit integers instead of 16- or 32-bit floating-point numbers. A quantized model needs less memory and compute, so it can run on cheaper hardware or a laptop, usually with only a small loss in quality.

Why it matters

A model with 8 billion parameters needs about 16 GB of memory at 16-bit precision. At 4-bit, it needs roughly 4–5 GB, small enough for a consumer GPU or a recent laptop. Less memory traffic also means faster responses and lower hosting costs.

The trade-off

Lower precision loses information. 8-bit quantization is usually close to indistinguishable from the original; 4-bit is good for most tasks; below that, quality drops noticeably, especially for reasoning, maths and code. Always test a quantized model on your own task rather than relying on general benchmarks.

Where you will see it

  • Running models locally: tools like Ollama and llama.cpp use quantized formats such as GGUF, labelled Q4, Q5 or Q8.
  • Self-hosting open models in production to cut GPU costs.
  • On-device AI in phones and browsers.

Quantization changes how a model is stored, not what it knows; to change behaviour you would use fine-tuning or better prompts.

Published · Updated · By · All terms

Go deeper