Back to articles

A practical guide to GGUF quantization for local inference

15 May 2026
GGUFllama.cppvLLMcuantización

What is quantization?

Quantization is a technique that reduces the numerical precision of a language model's weights. Instead of using 32-bit (FP32) or 16-bit (BF16) floating point values, the weights are stored with lower precision — typically 8 bits, 4 bits or even 2 bits.

This drastically reduces the memory needed to load a model, speeds up inference and allows large models to run on consumer hardware.

The GGUF format

GGUF (GPT-Generated Unified Format) is the file format used by llama.cpp to distribute quantized models. It replaced the old GGML with improvements such as extensible metadata, memory alignment and forward compatibility.

A GGUF file contains the model weights already quantized at one of several precision levels. When loading a GGUF, llama.cpp reads the weights directly without any additional conversion.

Quantization levels

The most common levels, ordered from highest to lowest quality:

  • Q8_0: 8 bits, virtually no perceptible loss. The best option if you have enough VRAM.

  • Q6_K: 6 bits, minimal loss. Excellent quality/size ratio.

  • Q5_K_M: 5 bits, a good balance for large models.

  • Q4_K_M: 4 bits, the most popular. Reduces size ~2.5× compared to FP16.

  • IQ4_XS: 4 bits with adaptive importance. Better than Q4_K_M in many benchmarks.

  • IQ3_XXS: 3 bits. Aggressive but useful for very limited hardware.

Quantization process with llama.cpp

To quantize a model you need a compiled llama.cpp:

git clone https://github.com/ggerganov/llama.cpp cd llama.cpp make -j

Convert the original model to FP16 GGUF:

python3 convert_hf_to_gguf.py /path/to/model --outtype f16 --outfile model-f16.gguf

Quantize to Q4_K_M:

./llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M

Inference with GGUF

Once quantized, you can run inference directly:

./llama-cli -m model-Q4_K_M.gguf -p "Explain what a transformer is" -n 512

Or spin up a server with llama-server:

./llama-server -m model-Q4_K_M.gguf --port 8080

Practical results

In my tests, quantizing Qwen3.6-35B-A3B from BF16 (65 GB) to Q8_0 (38 GB) keeps metrics practically identical on reasoning benchmarks. Quantizing to Q4_K_M brings it down to 19 GB with a loss of only 2-3 percentage points on MMLU.

The GGUF format is currently the most efficient way to run language models locally. Combining aggressive quantization with consumer hardware makes it possible to deploy 30B+ parameter models on a single 24 GB GPU.