Skip to main content

Capacity Planning — Gap Analysis & Fixes

Capacity planning answers the question every operations team ignores until it is too late: when will we run out? Not "are we over threshold today" — that is monitoring. Capacity planning is about reading the trend and acting before the trend wins.

The cluster had good reactive alerts (NodeCPUWarning, NodeMemoryCritical, LonghornNodeDiskPressure) but zero forward-looking tooling. swift-mac was at 71% CPU and 71% memory on the day of this audit with no signal for when it would saturate.


Cluster state at audit (2026-07-25)​

NodeCPU (actual)CPU %Memory (actual)Memory %
swift-mac2840m / 4000m71%5645 Mi / 7844 Mi71%
star-kitten2048m / 8000m25%8851 Mi / 15648 Mi56%
set-hog (CP)1166m / 8000m14%10612 Mi / 15656 Mi67%
fast-skunk1003m / 8000m12%7502 Mi / 15648 Mi47%
fast-heron725m / 8000m9%8404 Mi / 15648 Mi53%

swift-mac (MacBook Pro 2012, 4 CPU / 8 GB) is the capacity ceiling for the cluster. It runs Longhorn storage (most write-heavy workloads prefer local replicas) and is the natural landing zone for stateful pods. Without forecasting, there is no warning before it becomes unschedulable.


Gap inventory​

#GapRoot causeImpact
C1No memory/CPU exhaustion forecastNo predict_linear() rulesSilent saturation of swift-mac
C2No controller disk forecastOnly reactive disk alert at 85/90%MinIO crashed in July with no advance warning
C3No scheduler pressure trackingActual usage ≠ committed requestsNode can be "underloaded" but unschedulable
C4No Longhorn PVC fill forecastOnly reactive disk pressure alertClickHouse (40 GB quota) grows ~100 MB/day
C5No quota saturation alerts14 ResourceQuotas, zero alertsNew pods silently blocked when quota hits 100%
C6VPA covers 5 of 35+ DeploymentsNo right-sizing feedback for most servicesWasted requests or OOMKill risk
C7No capacity planning dashboardNo cluster-wide resource viewCapacity decisions made by intuition

Gap C1 — No node memory and CPU exhaustion forecast​

Problem​

The existing NodeMemoryWarning (>85%) and NodeCPUWarning (>80%) fire when the problem already exists. For swift-mac at 71%, you get one warning period before the critical alert, which may be only hours.

Fix​

manifests/monitoring/27-capacity-forecasting-rules.yaml — predict_linear() alerts with a 7-day horizon:

# NodeMemoryExhaustionForecast
predict_linear(
node_memory_MemAvailable_bytes[24h],
7 * 24 * 3600 -- project 7 days forward
) < 0

predict_linear(v[d], t) fits a linear regression over the last d of data and returns the projected value at t seconds in the future. If the projected available memory is below zero, the alert fires — giving 7 days of lead time to act.

The [24h] lookback window averages out daily cycles (e.g., nightly Velero backup causes a memory spike that would distort a shorter window). The alert requires for: 1h so a single anomalous hour doesn't trigger a week-long response.

# NodeCPUExhaustionForecast
predict_linear(
(sum without(mode) (avg without(cpu) (rate(node_cpu_seconds_total{mode!~"idle|iowait"}[2m]))))[24h:5m],
7 * 24 * 3600
) * 100 > 90

The [24h:5m] is a subquery: compute the 2-minute CPU rate at 5-minute intervals over the last 24 hours, then apply predict_linear to that series. Without the subquery, predict_linear would see a single instantaneous value instead of a time series.

:::caution predict_linear requires historical data predict_linear(metric[24h], ...) returns no data until Prometheus has 24 hours of samples for that metric. The alert will be absent (not firing, not pending) on a fresh install until the lookback window fills. :::


Gap C2 — No controller disk forecast​

Problem​

The MinIO service on the controller crashed in July 2026 when the disk hit ~94% utilization. The ControllerDiskWarning alert at 85% existed but fired only hours before the actual outage. MinIO caches the disk-full state in memory, so even after disk was freed it required a manual docker restart minio.

Fix​

# ControllerDiskExhaustionForecast
predict_linear(
node_filesystem_free_bytes{
instance="10.0.0.1:9100",
mountpoint="/"
}[24h],
7 * 24 * 3600
) < 1073741824 -- < 1 GiB remaining projected in 7 days

The controller disk is scraped by the node-exporter Docker container running at 10.0.0.1:9100. The threshold is 1 GiB (not zero) because MinIO needs headroom to create temp files during uploads. Running out below 1 GiB is operationally equivalent to running out completely.


Gap C3 — Scheduler pressure vs actual usage​

Problem​

The distinction that most teams miss:

Node actual CPU usage: 10% ← seems fine
Node CPU requests total: 95% ← new pods CANNOT schedule here

Actual usage measures how many CPU cycles are burning right now. Requests are the committed reservation — the scheduler uses requests, not actual usage, to decide where to place pods. A node can be nearly idle but unable to accept new pods if all its allocatable CPU is claimed.

On this cluster, swift-mac with 71% actual CPU may have a much higher committed request fraction — especially from stateful pods that over-request CPU as a safety margin.

Fix​

# NodeCPUSchedulerPressure
(
kube_node_status_allocatable{resource="cpu"}
- on(node) group_left()
sum by (node) (kube_pod_container_resource_requests{resource="cpu", node!=""})
) < 0.5 -- less than 500m uncommitted

This subtracts total pod CPU requests from allocatable CPU. The result is the scheduling headroom: how much new work can land on this node before the scheduler rejects new pods with Insufficient cpu.

The on(node) group_left() is required because kube_node_status_allocatable carries a node label while kube_pod_container_resource_requests also labels by node — the vector matching must be explicit to avoid a cross-product.

Same pattern for memory with a 512Mi threshold.


Gap C4 — No Longhorn PVC fill forecast​

Problem​

The langfuse ClickHouse PVC has a 40 GB Longhorn allocation and grows at roughly 100 MB/day from LLM trace data. No alert would fire until it hit 85% (~34 GB used), at which point there would be approximately 60 days of runway — but that 60-day window was invisible.

Fix​

# LonghornVolumeFillForecast
predict_linear(
longhorn_volume_usage_bytes[24h],
7 * 24 * 3600
) > longhorn_volume_capacity_bytes

The right side (longhorn_volume_capacity_bytes) varies per volume — this comparison is element-wise (same volume/pvc/namespace labels on both sides), so each volume's forecast is compared to its own capacity. The alert fires for any volume individually approaching full within 7 days.


Gap C5 — No quota saturation alerts​

Problem​

All 14 namespaces with ResourceQuotas had no alerting when they approached their hard limits. When a namespace hits 100% of requests.memory, every new pod creation in that namespace is blocked with:

Error creating: pods "X" is forbidden:
exceeded quota: requests.memory

Deployments, DaemonSets, Argo Rollouts canary pods, and ArgoCD Jobs all fail silently from the end-user perspective — ArgoCD shows OutOfSync, the CI job times out, but nothing in the alert feed explains why.

Fix​

manifests/monitoring/28-quota-saturation-alerts.yaml — four rules, all using the kube_resourcequota metric from kube-state-metrics:

# Memory requests > 80% of hard limit
kube_resourcequota{resource="requests.memory", type="used"}
/
kube_resourcequota{resource="requests.memory", type="hard"}
> 0.80

The ratio query works because kube_resourcequota exposes both type="used" and type="hard" as separate time series with identical namespace and resourcequota label sets. Element-wise division gives the utilization fraction for every namespace that has a quota — no per-namespace hardcoding needed.

AlertThresholdSeverity
NamespaceMemoryRequestsQuotaNearing> 80% for 10mwarning
NamespaceCPURequestsQuotaNearing> 80% for 10mwarning
NamespacePodCountQuotaNearing> 85% for 10mwarning
NamespaceQuotaExhausted≥ 100% for 1mcritical

The pod count threshold is 85% (not 80%) because Argo Rollouts and ArgoCD Jobs can create burst of short-lived pods that briefly push the count up during a deploy but settle back down. The 10-minute for window filters those transient bursts.


Gap C6 — VPA covers only 5 of 35+ Deployments​

Problem​

VPA recommendations are the cheapest form of right-sizing evidence — passive, continuous, and based on actual usage. Without them, requests and limits are set once at deploy time and never revisited. Services gradually drift: some waste allocations (over-requested), some accumulate OOMKill risk (under-limited).

At audit: 5 VPA objects existed (litellm, backstage, langfuse-web, langfuse-clickhouse, prometheus). All other services ran with no right-sizing feedback.

Fix​

Added 5 Off-mode VPA objects to manifests/vpa/00-vpa-objects.yaml:

WorkloadNamespaceKindNote
open-webuiaiStatefulSetMemory grows with concurrent sessions
matrix-synapsechatStatefulSetRWO PVC — Off only (eviction unsafe)
vaultvaultStatefulSetAWS KMS unseal during eviction is slow
minicloud-planeminicloud-plane-devRolloutOff to collect baseline first
platform-demoplatform-demo-devRolloutOff to collect baseline first

VPA can target Argo Rollouts (kind: Rollout, group argoproj.io/v1alpha1) natively — the VPA admission webhook patches the pod template regardless of the controller kind.

Off vs Auto decision matrix:

ConditionMode
< 7 days of dataOff
StatefulSet with RWO PVCOff (forced eviction may cause Multi-Attach)
External auth dependency at startup (Vault, KMS)Off
Stateless Deployment with 7+ days dataAuto

Reading a recommendation:

kubectl get vpa platform-demo -n platform-demo-dev -o json \
| python3 -c "
import json, sys
vpa = json.load(sys.stdin)
for c in vpa['status']['recommendation']['containerRecommendations']:
print(c['containerName'])
print(' target:', c['target'])
print(' upperBound:', c['upperBound'])
print(' lowerBound:', c['lowerBound'])
"

target is the recommended setting. upperBound is the 95th-percentile maximum — setting limits to upperBound gives a safety margin without over- provisioning.


Gap C7 — No cluster capacity dashboard​

manifests/monitoring/29-capacity-planning-dashboard.yaml — Grafana ConfigMap, UID minicloud-capacity, title "Capacity Planning". 7-day default time range.

Panel layout​

PanelWhat it shows
Cluster CPU allocatable / requested / actualThree bargauge bars per node — distinguishes scheduling pressure from load
Node CPU scheduler pressureGauge: requested / allocatable — red at >90%
Node memory scheduler pressureSame as above for memory
Node memory 7d trend + forecastTimeseries with current (solid) + predict_linear (dashed); if dashed hits 0 → forecast alert fires
Namespace quota saturation — memoryHorizontal bargauge per namespace; yellow >80%, red >95%
Namespace quota saturation — CPUSame for CPU requests
Longhorn PVC usage — top 15Horizontal bargauge by % full
Longhorn PVC growth rate (7d)Table with deriv(usage[7d]) * 86400 = bytes/day fill rate + current %
VPA recommendations vs limitsTable of kube_verticalpodautoscaler_status_recommendation_containerrecommendations_target for Off-mode services
Active capacity alertsALERTS{capacity="true"} — any firing forecast or quota alert

The PVC growth rate table uses:

deriv(longhorn_volume_usage_bytes[7d]) * 86400

deriv() returns the per-second rate of change. Multiplied by 86400 it gives bytes added per day. A negative value means the volume is shrinking (data being deleted). The table is sorted descending by growth rate so the fastest-filling volumes appear first.


Sizing runway at audit​

Based on predict_linear with current 7-day trends (approximate — trend varies with workload):

ResourceNode/VolumeEstimated runway
Memoryswift-macStable (low growth rate)
CPUswift-macStable (Ollama inference spikes but avg low)
Diskcontroller /~60+ days at current MinIO + k3s snapshot rate
Longhornlangfuse-clickhouse~120 days at ~100 MB/day fill rate

No immediate exhaustion risk, but swift-mac has no buffer — a single new memory-hungry workload (e.g., a second large Ollama model) would push it past the reactive alert threshold within hours, not days.


Files changed (gitops PR #273)​

FileWhat
manifests/monitoring/27-capacity-forecasting-rules.yamlNew — 6 PrometheusRules: memory/CPU/disk forecast + scheduler pressure (CPU + memory) + Longhorn PVC forecast
manifests/monitoring/28-quota-saturation-alerts.yamlNew — 4 quota alerts covering all 14 namespaces
manifests/monitoring/29-capacity-planning-dashboard.yamlNew — Grafana ConfigMap UID minicloud-capacity, 10 panels
manifests/vpa/00-vpa-objects.yaml+5 Off-mode VPA objects; coverage 5 → 10 workloads

Verification​

# Capacity forecasting rules loaded
kubectl port-forward svc/kps-prometheus -n monitoring 9090:9090 &
curl -s 'http://localhost:9090/api/v1/rules' \
| python3 -c "
import json, sys
for g in json.load(sys.stdin)['data']['groups']:
for r in g['rules']:
if r.get('labels', {}).get('capacity') == 'true':
print(r['name'], r['health'])
"

# No quota alerts firing (cluster should be healthy)
curl -s 'http://localhost:9090/api/v1/query?query=ALERTS{capacity="true"}' \
| python3 -c "import json,sys; r=json.load(sys.stdin); print(len(r['data']['result']), 'active capacity alerts')"

# Scheduler pressure — check remaining headroom
curl -s 'http://localhost:9090/api/v1/query?query=kube_node_status_allocatable{resource="cpu"}-on(node)group_left()sum+by(node)(kube_pod_container_resource_requests{resource="cpu",node!=""})' \
| python3 -c "
import json, sys
for r in json.load(sys.stdin)['data']['result']:
node = r['metric'].get('node', '?')
remaining = float(r['value'][1])
print(f'{node}: {remaining:.2f} CPU cores uncommitted')
"

# VPA coverage
kubectl get vpa -A --no-headers | awk '{print $1, $2, $3}'

# Grafana dashboard visible
/usr/bin/curl --cacert ~/minicloud-ca.crt -sf \
https://grafana.10.0.0.200.nip.io/api/dashboards/uid/minicloud-capacity \
| python3 -c "import json,sys; d=json.load(sys.stdin); print(d['dashboard']['title'], '—', len(d['dashboard']['panels']), 'panels')"