Skip to main content

4 posts tagged with "inference"

View All Tags

We Replaced Ollama With vLLM on CPU-Only Kubernetes — Here Is What Changed

· 9 min read
Software Engineer & Cloud Architect

Most write-ups about vLLM versus Ollama assume you have a GPU. The benchmarks show impressive VRAM utilisation. The architecture diagrams include CUDA drivers. The recommendation — vLLM wins — comes with an implicit asterisk: assuming you have the hardware for it.

We did not. Our inference cluster is four ThinkPad laptops, each with an i7-8565U (4 cores, 8 threads, up to 4.6 GHz turbo), 16 GB of RAM, and no GPU of any kind. We ran Ollama first. Then we replaced it with vLLM. This is the honest account of that migration: what broke, what improved, and what the tradeoffs actually look like when you run LLM inference on commodity x86 CPUs.

Why Enterprises Should Run vLLM Instead of Ollama for AI Inference

· 8 min read
Software Engineer & Cloud Architect

Ollama is how most teams first run a large language model locally. You install it in five minutes, run ollama pull mistral, and you have a working API. It feels like magic.

Then you try to serve ten users at once. Or a hundred. Or you need to audit every request for compliance. Or your legal team asks where the data goes. That is when you realise Ollama was built for something else entirely.

LLMs on Bare Metal: Quantization, SLMs, and Replacing phi4-mini with Qwen 2.5 7B

· 9 min read
Software Engineer & Cloud Architect

Running LLMs on bare-metal CPU hardware forces you to understand the numbers behind model files. This post documents how we reason about model size, quantization, and inference serving on minicloud — and the concrete change we made: replacing phi4-mini with Qwen 2.5 7B across all three Ollama instances.