What is quantization?
Quantization is a technique that reduces the numerical precision of a language model's weights. Instead of using 32-bit (FP32) or 16-bit (BF16) floating point values, the weights are stored with lower precision — typically 8 bits, 4 bits or even 2 bits.
This drastically reduces the memory needed to load a model, speeds up inference and allows large models to run on consumer hardware.
The GGUF format
GGUF (GPT-Generated Unified Format) is the file format used by llama.cpp to distribute quantized models. It replaced the old GGML with improvements such as extensible metadata, memory alignment and forward compatibility.
A GGUF file contains the model weights already quantized at one of several precision levels. When loading a GGUF, llama.cpp reads the weights directly without any additional conversion.
Quantization levels
The most common levels, ordered from highest to lowest quality:
Q8_0: 8 bits, virtually no perceptible loss. The best option if you have enough VRAM.
Q6_K: 6 bits, minimal loss. Excellent quality/size ratio.
Q5_K_M: 5 bits, a good balance for large models.
Q4_K_M: 4 bits, the most popular. Reduces size ~2.5× compared to FP16.
IQ4_XS: 4 bits with adaptive importance. Better than Q4_K_M in many benchmarks.
IQ3_XXS: 3 bits. Aggressive but useful for very limited hardware.
Quantization process with llama.cpp
To quantize a model you need a compiled llama.cpp:
git clone https://github.com/ggerganov/llama.cpp cd llama.cpp make -j
Convert the original model to FP16 GGUF:
python3 convert_hf_to_gguf.py /path/to/model --outtype f16 --outfile model-f16.gguf
Quantize to Q4_K_M:
./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
Inference with GGUF
Once quantized, you can run inference directly:
./llama-cli -m model-Q4_K_M.gguf -p "Explain what a transformer is" -n 512
Or spin up a server with llama-server:
./llama-server -m model-Q4_K_M.gguf --port 8080
Practical results
In my tests, quantizing Qwen3.6-35B-A3B from BF16 (65 GB) to Q8_0 (38 GB) keeps metrics practically identical on reasoning benchmarks. Quantizing to Q4_K_M brings it down to 19 GB with a loss of only 2-3 percentage points on MMLU.
The GGUF format is currently the most efficient way to run language models locally. Combining aggressive quantization with consumer hardware makes it possible to deploy 30B+ parameter models on a single 24 GB GPU.