Back to articles

Setting up an inference server with Ollama and vLLM locally

10 May 2026
OllamavLLMNVIDIAinferencia

Motivation

Running large language models locally offers clear advantages: total data privacy, no rate limits, no per-token costs and full control over the infrastructure. With suitable hardware, setting up a local inference server is a weekend project.

Ollama: the fast route

Ollama is the simplest way to start with local inference. One-command install:

curl -fsSL https://ollama.com/install.sh | sh

Download and run a model:

ollama pull qwen3:8b ollama run qwen3:8b

Ollama exposes an OpenAI-compatible REST API at localhost:11434:

curl http://localhost:11434/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"qwen3:8b","messages":[{"role":"user","content":"Hello"}]}'

Advantages of Ollama

  • Automatic model management (download, cache, cleanup)

  • Support for multi-layer GGUF (split between GPU and RAM)

  • OpenAI-compatible REST API

  • Per-model preconfigured chat templates

Limitations

  • Does not support batched inference

  • Higher latency than vLLM under concurrent load

  • Less control over advanced sampling parameters

vLLM: high performance

vLLM is an inference engine designed for maximum throughput. It supports continuous batching, PagedAttention and native FP8 quantization.

pip install vllm vllm serve Qwen/Qwen3.6-35B-A3B --dtype float16 --max-model-len 32768

Advantages of vLLM

  • PagedAttention: efficient KV cache management

  • Continuous batching: groups requests automatically

  • Native FP8 support without significant quality loss

  • OpenAI-compatible API

  • Prefix caching: reuses computation between similar requests

NVIDIA GPU setup

For NVIDIA GPU inference you need:

  • NVIDIA drivers ≥ 535

  • CUDA toolkit ≥ 12.1

  • Environment variables: CUDA_VISIBLE_DEVICES=0,1

Verify that the GPUs are detected:

nvidia-smi python3 -c "import torch; print(torch.cuda.device_count())"

My setup

I usually keep two inference servers running simultaneously:

  • Ollama: 14 models for general tasks, chat and embeddings

  • vLLM: Qwen3.6-35B-A3B in FP8 for high-demand tasks

Both serve through a proxy layer that routes according to the request type. The combination offers flexibility for development and performance for production.