Local models and LLMs
Running language models on your own hardware has gone from being an eccentricity to a reasonable decision: predictable cost, controlled latency and zero data leakage. My job is to make sure that decision is taken for performance, not for fashion.
What I do
- Inference server deployment — vLLM, Ollama and llama.cpp configured for your load: concurrency, context length and memory.
- Model selection and comparison — which model (Qwen, Llama, Mistral, DeepSeek…) performs best on your concrete task, with measurable tests.
- Quantization and GGUF formats — adjust precision and size to squeeze the available hardware without losing relevant quality.
- Inference optimization — KV cache, batching, long contexts and multi-GPU: getting real performance out of the GPU investment.
The criterion
First I measure, then I optimize. Often the problem is not the model but the prompt, the context or badly sized hardware. If a quantized 8B model on a single GPU covers 90% of the case, that is what I will propose to you.
Related
This area technically underpins applied AI and relies on suitable infrastructure.
Undecided between cloud, API or your own server? I'll put the numbers in front of you.