Motivation
Running large language models locally offers clear advantages: total data privacy, no rate limits, no per-token costs and full control over the infrastructure. With suitable hardware, setting up a local inference server is a weekend project.
Ollama: the fast route
Ollama is the simplest way to start with local inference. One-command install:
curl -fsSL https://ollama.com/install.sh | sh
Download and run a model:
ollama pull qwen3:8b ollama run qwen3:8b
Ollama exposes an OpenAI-compatible REST API at localhost:11434:
curl http://localhost:11434/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"qwen3:8b","messages":[{"role":"user","content":"Hello"}]}'
Advantages of Ollama
Automatic model management (download, cache, cleanup)
Support for multi-layer GGUF (split between GPU and RAM)
OpenAI-compatible REST API
Per-model preconfigured chat templates
Limitations
Does not support batched inference
Higher latency than vLLM under concurrent load
Less control over advanced sampling parameters
vLLM: high performance
vLLM is an inference engine designed for maximum throughput. It supports continuous batching, PagedAttention and native FP8 quantization.
pip install vllm vllm serve Qwen/Qwen3.6-35B-A3B --dtype float16 --max-model-len 32768
Advantages of vLLM
PagedAttention: efficient KV cache management
Continuous batching: groups requests automatically
Native FP8 support without significant quality loss
OpenAI-compatible API
Prefix caching: reuses computation between similar requests
NVIDIA GPU setup
For NVIDIA GPU inference you need:
NVIDIA drivers ≥ 535
CUDA toolkit ≥ 12.1
Environment variables:
CUDA_VISIBLE_DEVICES=0,1
Verify that the GPUs are detected:
nvidia-smi python3 -c "import torch; print(torch.cuda.device_count())"
My setup
I usually keep two inference servers running simultaneously:
Ollama: 14 models for general tasks, chat and embeddings
vLLM: Qwen3.6-35B-A3B in FP8 for high-demand tasks
Both serve through a proxy layer that routes according to the request type. The combination offers flexibility for development and performance for production.