Back to articles

vLLM vs llama.cpp: which server to use for local inference?

16 Aug 2026
vLLMllama.cppGGUFinferenciaGPULLM

Running language models locally has become increasingly accessible, but choosing the right inference engine is still an important decision. Two of the most interesting options are vLLM and llama.cpp. Although both allow running and serving LLMs, they are designed with quite different priorities.

After working with both, my conclusion is that there is no absolute winner: vLLM stands out as a GPU inference server, while llama.cpp offers flexibility that is hard to match for running quantized models on practically any hardware.

vLLM: performance and serving on GPU

vLLM is designed mainly to serve models efficiently on GPU. It can load models distributed via Hugging Face directly and offers an OpenAI-compatible API, which makes it easy to integrate with existing applications and agents.

One of its main advantages is efficient memory management through PagedAttention, together with batching techniques that make better use of the GPU when there are several simultaneous requests.

It also offers features that are especially interesting on dedicated AI machines:

  • Tensor Parallel to distribute models across several GPUs.
  • FP8 quantization.
  • KV Cache in lower-precision formats.
  • Continuous batching.
  • Speculative decoding.
  • OpenAI-compatible API.
  • Good performance with multiple concurrent requests.

On systems with several GPUs, vLLM can split a model using Tensor Parallel. This makes it possible to run models that would not fit on a single card and to make joint use of the available VRAM.

Its main drawback is that it is considerably more demanding in terms of model and hardware compatibility. New architectures may need specific support and certain quantized models do not always work immediately.

llama.cpp: flexibility and GGUF

llama.cpp starts from a different philosophy. Its great strength is being able to run models in GGUF format using CPU, GPU or a combination of both.

This allows using different quantization levels —Q4, Q5, Q6, Q8, among others— and adapting the model to the available memory.

For example, a model that would need around 60 GB in BF16 can considerably reduce its requirements through GGUF quantization, keeping a quality surprisingly close to the original when high quantizations such as Q8 are used.

Another important advantage is offloading. If the model does not fit entirely in VRAM, part of its layers can be kept in RAM and run on CPU. Performance decreases, but it allows working with models that would simply be impossible to load fully on GPU.

llama.cpp stands out especially for:

  • Broad support for the GGUF format.
  • A wide variety of quantizations.
  • Execution on CPU and GPU.
  • Partial GPU offloading.
  • Relatively low hardware requirements.
  • Compatibility with numerous architectures.
  • OpenAI-compatible API via llama-server.

For experimenting with local models it is especially practical: downloading a GGUF and starting to test it is usually much simpler than preparing a full GPU serving environment.

FP8 versus GGUF

An important difference between the two ecosystems lies in the type of quantization used.

vLLM is especially oriented to formats used directly by accelerators, such as BF16 and FP8. llama.cpp, on the other hand, mainly uses the quantizations defined within the GGUF ecosystem.

An FP8 model usually needs about half the memory for weights as its BF16 equivalent. GGUF allows reducing it even further through quantizations such as Q6, Q5 or Q4.

In simplified form:

FormatMemoryQualityUsual use
BF16Very highMaximumGPU with lots of VRAM
FP8HighVery highGPU serving
GGUF Q8HighVery highLocal inference
GGUF Q6/Q5MediumHighQuality/memory balance
GGUF Q4LowGoodLimited hardware

However, comparing FP8 and GGUF only by file size would be a mistake. The model architecture, the kernel implementation, the context size and the KV Cache also considerably influence the real memory consumption and performance.

Which one delivers more performance?

When the model fits entirely on GPU and the hardware is well supported, vLLM is usually the most suitable option for offering a high-performance inference service, especially when there are several concurrent requests.

llama.cpp can offer excellent speeds, but its main advantage is not necessarily achieving maximum throughput, but rather its flexibility.

The difference becomes especially clear when the server must serve several users simultaneously. The batching and memory management techniques used by vLLM make much better use of the GPU in this scenario.

For individual use, development or experimentation, that advantage can be much less important.

Multi-GPU

Another interesting point appears when we have several GPUs.

vLLM can use Tensor Parallel to distribute the model across different cards. NVLink is not essential: it can also work over PCIe, although communication between GPUs can become a limiting factor.

llama.cpp also allows splitting the model across several GPUs, although its approach is still more oriented to running the model with the available resources than to maximizing the throughput of a server with many requests.

In both cases, adding two 24 GB GPUs does not literally turn the system into a single 48 GB GPU. The way weights, KV Cache and operations are distributed across devices matters as much as the total amount of VRAM.

Which one to choose?

The choice depends mainly on the goal.

I would choose vLLM when:

  • The model fits entirely on GPU.
  • I want to use BF16 or FP8 models.
  • I need to serve multiple requests.
  • I want to take advantage of several GPUs.
  • I am deploying an API for other applications.
  • Performance and throughput are a priority.

I would choose llama.cpp when:

  • I want to use GGUF models.
  • Available VRAM is limited.
  • I need to combine RAM and VRAM.
  • I want to experiment quickly with different quantizations.
  • I am looking for maximum compatibility with hardware and models.
  • Simplicity matters more than maximum throughput.

Conclusion

vLLM and llama.cpp are not really direct competitors in every scenario.

vLLM is closer to a high-performance serving platform, designed to take advantage of accelerators and efficiently serve multiple requests. llama.cpp is an extraordinarily flexible inference tool, capable of running large models even on machines where the available resources are limited.

On a dedicated AI workstation, it makes sense to use both.

For models that work correctly in FP8 and fit entirely on the GPUs, vLLM can provide a fast and efficient serving environment. For testing new models, using GGUF, experimenting with different quantizations or running models that exceed the available VRAM, llama.cpp remains one of the most versatile tools in the local AI ecosystem.

The question, therefore, is not so much "vLLM or llama.cpp?", but rather "which tool fits this model and this inference scenario best?".