Skip to main content

Inference Optimization β€” CPU Tuning & Horizontal Scaling

Phase complete: 2026-07-05
Issue closed: #41

Hardware baseline​

4Γ— Intel Core i7-10510U (set-hog, fast-skunk, fast-heron, star-kitten)
β”œβ”€β”€ 4 cores / 8 threads per node
β”œβ”€β”€ 1.8 GHz base β†’ 4.9 GHz turbo
β”œβ”€β”€ 15.6 GB RAM per node
β”œβ”€β”€ AVX2 βœ… (used automatically by llama.cpp inside Ollama)
└── No GPU β€” all optimizations are CPU-path only

Part 1 β€” Ollama env var tuning​

Applied to all three Ollama instances (primary on fast-heron, secondary on star-kitten, tertiary on fast-skunk):

# Primary + secondary (7B model tier β€” full context)
OLLAMA_NUM_PARALLEL: "6" # 6 concurrent requests per instance (raised from 4 on 2026-07-05)
OLLAMA_NUM_THREADS: "6" # leave 2 threads for k8s node agent + OS
OLLAMA_MAX_LOADED_MODELS: "2" # keep 2 most-used models hot in RAM at once
OLLAMA_FLASH_ATTENTION: "1" # reduced memory bandwidth on CPU attention path
OLLAMA_KV_CACHE_TYPE: "q8_0" # 8-bit KV cache: half the memory of fp16
OLLAMA_NUM_CTX: "8192" # raised from 4096 β€” gives deepseek-r1:7b think phase room (500–1500 tokens)

# Tertiary (light-model tier β€” fast-skunk, no 7B models)
OLLAMA_NUM_PARALLEL: "6" # 6 concurrent requests
OLLAMA_NUM_CTX: "4096" # lower context β€” light conversational use, more parallel headroom

Config is in minicloud-ansible/helm-values/ollama-values.yaml, ollama-secondary-values.yaml, and ollama-tertiary-values.yaml.

Part 2 β€” CPU governor: performance mode​

Intel i7-10510U throttles from 4.9 GHz turbo to 1.8 GHz base under sustained load in powersave mode. Setting performance keeps turbo boost active.

Verify current governor on all nodes​

for node in set-hog fast-skunk fast-heron star-kitten; do
ssh $node "cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor"
done
# All should return: performance

How it was set (persistent across reboots)​

A systemd unit cpu-performance.service was deployed on all 4 nodes:

[Unit]
Description=Set CPU governor to performance
After=multi-user.target

[Service]
Type=oneshot
ExecStart=/bin/sh -c 'echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor'
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target
# Verify service status on a node:
ssh fast-heron "systemctl status cpu-performance"

Re-apply if needed (after a new node joins)​

ssh <new-node> "cat > /etc/systemd/system/cpu-performance.service << 'EOF'
[Unit]
Description=Set CPU governor to performance
After=multi-user.target
[Service]
Type=oneshot
ExecStart=/bin/sh -c 'echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor'
RemainAfterExit=yes
[Install]
WantedBy=multi-user.target
EOF
systemctl daemon-reload && systemctl enable --now cpu-performance"

Part 3 β€” Horizontal scaling: three Ollama instances​

Three independent Ollama Deployments, each pinned to a dedicated node (2026-07-05):

InstanceNodeServiceModelsPVCNUM_CTX
ollamafast-heronollama.ai.svc:11434all 7 (incl. 7B + vision)local-path 10Gi8192
ollama-secondarystar-kittenollama-secondary.ai.svc:11434all 7 (incl. 7B + vision)local-path 10Gi8192
ollama-tertiaryfast-skunkollama-tertiary.ai.svc:11434light only (1B + 3.8B + embed)local-path 10Gi4096

LiteLLM routes across all three with routing_strategy: least-busy. The tertiary adds 6 concurrent slots for light workloads without competing with 7B inference.

Total concurrent capacity (2026-07-05)​

What changedBeforeAfter
NUM_PARALLEL per instance46
Ollama instances23
Light-model slots (phi4-mini + llama3.2:1b)8 (4Γ—2)18 (6Γ—3)
deepseek-r1:7b40+ s on CPU5–8 s via DeepSeek cloud API (fallback: local)
Estimated concurrent users~20~80–90

Models per instance​

Primary + Secondary (fast-heron, star-kitten):
llama3.2:1b (1.3 GB) β€” ultra-fast tier
phi4-mini (2.5 GB) β€” instruction following, Microsoft CPU-optimised
phi3-financial (2.2 GB) β€” financial domain specialist (local-only)
qwen3.5:4b (3.4 GB) β€” primary smart tier, native tool calling
deepseek-r1:7b (4.7 GB) β€” reasoning specialist (local fallback only)
moondream (1.7 GB) β€” vision: fast OCR + image description
llava-phi3 (2.9 GB) β€” vision: detailed image analysis + document OCR

Tertiary (fast-skunk β€” light workloads only):
llama3.2:1b (1.3 GB) β€” ultra-fast tier
phi4-mini (2.5 GB) β€” instruction following
bge-m3 (1.2 GB) β€” 1024-dim embedding for RAG (100+ languages)

deepseek-r1:7b cloud routing: The reasoning model routes to DeepSeek cloud API (5–8 s) as first choice. If the cloud is unavailable (balance, outage), LiteLLM automatically falls back to deepseek-r1:7b-local on the primary/secondary Ollama instances. The fallbacks block in router_settings handles this automatically β€” no client-side changes needed.

Verify all three instances​

kubectl exec -n ai deploy/ollama -- ollama list
kubectl exec -n ai deploy/ollama-secondary -- ollama list
kubectl exec -n ai deploy/ollama-tertiary -- ollama list

Add a model to all three instances​

Model pulls require a temporary egress NetworkPolicy (ai namespace has default-deny-egress).
Use podSelector: matchLabels: app.kubernetes.io/instance: <instance-name> to target specific pods:

# Apply per-instance temp egress
kubectl apply -f - <<'EOF'
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: temp-allow-model-pull
namespace: ai
spec:
podSelector: {}
egress:
- ports:
- port: 443
- port: 80
policyTypes:
- Egress
EOF

# Pull via REST (avoid interactive TTY)
kubectl port-forward -n ai deploy/ollama 11434:11434 &
kubectl port-forward -n ai deploy/ollama-secondary 11435:11434 &
kubectl port-forward -n ai deploy/ollama-tertiary 11436:11434 &

curl -s -X POST http://localhost:11434/api/pull -d '{"model":"<model-name>","stream":false}'
curl -s -X POST http://localhost:11435/api/pull -d '{"model":"<model-name>","stream":false}'
curl -s -X POST http://localhost:11436/api/pull -d '{"model":"<model-name>","stream":false}'

# Remove temp policy
kubectl delete networkpolicy temp-allow-model-pull -n ai

Part 4 β€” Department tier routing​

Virtual keys created in LiteLLM control which models each department can call:

TierDepartmentsLocal modelsCloud accessBudget/30dTPM
premiumIT, Data Analytics, Actuariat, Transformationall (incl. vision)all cloud models$100200k
standardCybersecurity, Audit, Finance, Reinsurance, Juridique, Souscription, Commercialphi3-financial, qwen3.5:4b, phi4-mini, llama3.2:1b, moondreamgroq + deepseek + mistral-small + gpt-4o-mini + gemini-2.0-flash + hf-*$30100k
basicSinistres, Operations, RH, Services GΓ©nΓ©rauxphi3-financial, phi4-mini, llama3.2:1bnone$550k

All 15 keys persist in the litellm PostgreSQL database. Keys are stored at ~/.litellm-department-keys (mode 600) on the controller.

Verify a key's model restriction​

kubectl port-forward -n ai svc/litellm 4000:4000 &

# Attempt a model the key doesn't have access to β€” should return 403
DEPT_KEY=$(grep direction-sinistres ~/.litellm-department-keys | awk '{print $2}')
curl -X POST http://localhost:4000/v1/chat/completions \
-H "Authorization: Bearer $DEPT_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet","messages":[{"role":"user","content":"test"}]}'
# Expected: 403 β€” model not in basic tier allowlist

Part 5 β€” Vision and multimodal capability​

Two vision-language models are deployed alongside the text models:

ModelSizeBest for
moondream1.7 GBFast image description, basic OCR, scene classification
llava-phi32.9 GBDetailed document analysis, structured form extraction, complex OCR

How to call a vision model​

Vision requests use the standard OpenAI image_url format through LiteLLM:

import openai, base64

client = openai.OpenAI(
base_url="http://localhost:4000", # LiteLLM gateway
api_key="sk-direction-it-..." # premium tier key β€” llava-phi3 access
)

# Base64-encode a local image
with open("invoice.png", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
model="llava-phi3",
messages=[{
"role": "user",
"content": [
{"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{img_b64}"}},
{"type": "text",
"text": "Extract all text from this document as structured JSON."}
]
}]
)
print(response.choices[0].message.content)

Vision model limitations on CPU​

  • Image encoding adds 2–5 seconds before token generation starts (no GPU encoder).
  • Maximum reliable image resolution: ~1024Γ—1024 β€” larger images slow encoding significantly.
  • Multi-image requests are not recommended in CPU mode (memory pressure).
  • Neither model supports real-time video or streaming video frames.
  • moondream handles basic printed text well; llava-phi3 handles handwritten text and complex layouts better.

Part 6 β€” RAG pipeline: pgvector + bge-m3 + Open WebUI​

Retrieval-Augmented Generation (RAG) allows Open WebUI to answer questions using uploaded documents. All components run on-cluster β€” no external embedding API needed. For full detail see the RAG Pipeline runbook.

Architecture​

User uploads document β†’ Open WebUI chunks text (500 chars, 100 overlap)
β†’ bge-m3 (1024-dim) via Ollama β†’ pgvector HNSW index
↓
User asks question β†’ bge-m3 embeds query β†’ pgvector cosine search (TOP_K=10)
+ Python BM25Retriever (French stemmer) β†’ RRF merge β†’ re-ranker β†’ top 3 β†’ LLM

Components​

ComponentRoleConfig
bge-m3 (1.2 GB)Embedding model β€” 1024-dim, 100+ languages, strong FrenchPulled on primary Ollama instance
postgresql-aipgvector 0.8.4 host β€” ragdb databaseExisting pod, new DB
document_chunk tablevector(1024) + HNSW index + GIN index (french dictionary)Migrated from 768-dim 2026-07-06
Open WebUIRAG orchestrator β€” chunking, embedding, retrieval, context injectionREVISION: latest

Key configuration (open-webui-values.yaml)​

- name: VECTOR_DB
value: "pgvector"
- name: PGVECTOR_DB_URL
value: "postgresql://aiplatform:<password>@postgresql-ai.ai.svc.cluster.local:5432/ragdb"
- name: VECTOR_LENGTH
value: "1024"
- name: PGVECTOR_INITIALIZE_MAX_VECTOR_LENGTH
value: "1024"
- name: PGVECTOR_INDEX_METHOD
value: "hnsw"
- name: PGVECTOR_HNSW_M
value: "16"
- name: PGVECTOR_HNSW_EF_CONSTRUCTION
value: "64"
- name: RAG_EMBEDDING_ENGINE
value: "ollama"
- name: RAG_EMBEDDING_MODEL
value: "bge-m3"
- name: RAG_OLLAMA_BASE_URL
value: "http://ollama.ai.svc.cluster.local:11434"
- name: CHUNK_SIZE
value: "500"
- name: CHUNK_OVERLAP
value: "100"
- name: RAG_TOP_K
value: "10"
- name: ENABLE_RAG_HYBRID_SEARCH
value: "true"

RAG_EMBEDDING_ENGINE: ollama bypasses LiteLLM for embeddings β€” avoids polluting cost/spend metrics with internal embedding calls.

Why HNSW instead of IVFFlat​

pgvector 0.8.4 was compiled with SIMD instructions not supported on the ThinkPad i7-10510U (SIGILL during IVFFlat K-means build). HNSW uses a different code path that builds incrementally and does not trigger the SIGILL. Set PGVECTOR_INDEX_METHOD=hnsw in the Open WebUI env.

How to use RAG in Open WebUI​

  1. Log in at https://chat.devandre.sbs
  2. Start a new chat
  3. Click the paperclip icon β†’ upload a PDF, text, or CSV
  4. Ask questions β€” Open WebUI retrieves the top 5 relevant chunks and injects them into the LLM context

Verify the embedding chain​

# Direct Ollama embedding (what Open WebUI uses internally)
kubectl port-forward -n ai svc/ollama 11434:11434 &
curl -s -X POST http://localhost:11434/api/embeddings \
-d '{"model":"bge-m3","prompt":"Kubernetes manages containers"}' | \
python3 -c 'import sys,json; e=json.load(sys.stdin)["embedding"]; print(f"dims={len(e)}, ok")'
# Expected: dims=1024, ok
kill %1

# Via LiteLLM (for external API users)
curl -s -X POST http://localhost:4000/v1/embeddings \
-H "Authorization: Bearer <direction-it-key>" \
-H "Content-Type: application/json" \
-d '{"model":"bge-m3","input":"test"}' | \
python3 -c 'import sys,json; r=json.load(sys.stdin); print(f"dims={len(r[\"data\"][0][\"embedding\"])}")'
# Expected: dims=768

# pgvector table status
kubectl exec -n ai postgresql-ai-0 -- psql postgresql://aiplatform:<pw>@localhost/ragdb \
-c "SELECT indexname, indexdef FROM pg_indexes WHERE tablename='document_chunk';"

What is NOT possible on CPU-only nodes​

Document these so engineers don't spend time investigating:

TechniqueReason not available
CUDA / ROCm GPU accelerationNo NVIDIA/AMD GPU β€” Intel UHD 620 integrated only
vLLM PagedAttentionGPU VRAM management β€” not applicable on CPU
TensorRT / ONNX Runtime GPUNVIDIA-only compilation target
Flash Attention v2 kernelCUDA kernel β€” the OLLAMA_FLASH_ATTENTION=1 env var uses a CPU-path approximation
AWQ / GPTQ inferenceRequire GPU for dequantization at inference time
Speculative decodingNot yet natively supported by Ollama

Performance observed​

After applying Parts 1–3:

ScenarioThroughput
Single request (llama3.2:1b)~25–35 t/s
Single request (phi4-mini, 3.8B)~12–18 t/s
Single request (qwen3.5:4b)~10–15 t/s
Single request (deepseek-r1:7b)~6–10 t/s (includes think phase)
Vision request (moondream)~5–8 t/s + image encode time
Vision request (llava-phi3)~4–7 t/s + image encode time
4 concurrent requests (before 2026-07-05)All 4 start immediately (no queuing until 5th)
6 concurrent requests per instance (after 2026-07-05)All 6 start immediately (no queuing until 7th)
deepseek-r1:7b via cloud API5–8 s (vs 40+ s on CPU)
18 concurrent light-model slots (3 instances Γ— 6 parallel)~80–90 simultaneous users
Cold start after governor setCPU stays at 3.5–4.9 GHz under sustained load