Skip to main content

Security & Resource Boundaries — RBAC, Admission Control, Quotas & LimitRanges

Evaluated 2026-08-02. Two gaps found and closed in gitops PR #525.

CriterionBeforeAfter
RBAC & Admission Control✅ 11 constraints, 0 deny violations🔧 Gatekeeper exclusion list corrected; collab removed
Resource Quotas & Limits⚠️ 22/63 namespaces covered; 10 app namespaces unprotected🔧 10 new ResourceQuota + LimitRange files; 32/63 now covered

1 — RBAC & Admission Control​

Gatekeeper (OPA) — 11 constraints active​

ConstraintModeViolations
no-privileged-containersdeny0
require-non-rootdeny0
no-privilege-escalationdeny0
require-resource-limitsdeny0 (see exclusion list below)
block-latest-tagdeny0
allowed-registriesdeny0
no-host-pathdeny0
block-net-rawdeny0
block-capabilitiesdeny0
require-ingress-tlsdeny0
no-loadbalancer-in-devdeny0
require-seccompwarn158 (unchanged — graduated enforcement backlog)

RBAC bindings​

ClusterRoleBindings (cluster-wide):

SubjectRoleType
oidc:Direction ITcluster-adminGroup
oidc:kanmegneacluster-adminDirect user
oidc:Cybersécurité, oidc:AuditviewGroup

RoleBindings (per namespace):

GroupRoleNamespaces
oidc:Développeurseditplatform-demo-dev, minicloud-plane-dev, collab-dev, insurance-dev
oidc:QAviewplatform-demo-staging, minicloud-plane-staging, collab-staging, insurance-staging

No human RoleBinding on any *-prod namespace.

Gap found — require-resource-limits exclusion list​

Gatekeeper's require-resource-limits constraint had a namespace exclusion list that included active production workloads: chat, erp, mail, collab, productivity, sign. These namespaces were excluded because upstream Helm charts don't allow setting resource limits on all their sub-components (StatefulSet sidecars, init containers, bundled PostgreSQL/Redis/RabbitMQ).

Confirmed effect: several pods in excluded namespaces were running without limits:

NamespacePodStatus
chatmatrix-synapse-redis-masterNo limits on Redis container
erperpnext-scheduler, erpnext-conf-bench-*No limits
productivityplane-ce-pgdb, plane-ce-minio, plane-ce-rabbitmq, plane-ce-redisNo limits

Gap closed — collab removed from exclusion list (PR #525)​

After verifying that all 4 Jitsi containers (jicofo, jvb, prosody, web) have explicit resource limits in helm-values/minicloud-1/jitsi-values.yaml, collab was removed from the exclusion list. New Deployments/StatefulSets in collab will now be blocked by Gatekeeper if they lack limits.

Accepted gap — direct user ClusterRoleBinding​

oidc-cluster-admin-kanmegnea grants cluster-admin to the individual OIDC user kanmegnea rather than going through the group model. This is intentional: the group binding (oidc:Direction IT → cluster-admin) relies on Authentik's group token claim, and a direct binding serves as a break-glass fallback if the group claim is misconfigured.

Path to close: Add kanmegnea to the Direction IT Authentik group, then delete oidc-cluster-admin-kanmegnea. The group binding alone is sufficient.

Remaining exclusion list (known backlog)​

These namespaces remain excluded from require-resource-limits due to upstream chart limitations:

NamespaceReason
chatmatrix-synapse-redis StatefulSet has no limit support in chart values
erperpnext-scheduler Deployment has no resource limit in chart values
mailses-inbound sidecar container has resources: {} — chart doesn't expose per-sidecar limits
productivityPlane CE StatefulSets (pgdb, minio, rabbitmq, redis) don't expose resource limit config
signPending chart audit
argo-rolloutsUpstream chart doesn't set limits on the controller Deployment
system-upgradeSUC chart doesn't set limits on the controller
vpa-systemVPA admission controller has no limits in chart values
temporalTemporal server chart generates init/config containers without resource limits

LimitRange injection mitigates this: Each of these namespaces now has a LimitRange (added in PR #525). New pods without explicit limits will receive defaults from the LimitRange automatically. The Gatekeeper exclusion only affects admission of Deployments/StatefulSets/DaemonSets — the LimitRange is the primary defense for existing workloads.


2 — Resource Quotas & LimitRanges​

Before (gap)​

22 of 63 namespaces had a ResourceQuota. 10 application namespaces with active, resource-intensive workloads had nothing:

NamespaceWorkloadRisk
aiOllama (4 CPU/8 GiB per pod), LiteLLM, Open WebUI, docling, RAGHigh — scaling Ollama to 2 replicas consumes 16 GiB with no guard
langfuseClickHouse, Langfuse web/workerMedium — ClickHouse can spike memory
erpERPNext + MariaDB + workersMedium — 11 pods, no ceiling
chatMatrix Synapse, ElementLow-medium
mailStalwart, ses-inboundLow
nextcloudNextcloud, OnlyOfficeMedium — OnlyOffice can spike CPU
automationn8nLow
productivityPlane CE (11 pods)Medium — bundled databases
signDocuSealLow
collabJitsi (JVB, prosody, jicofo, web)Low-medium

Fix — 10 new quota files (PR #525)​

Each file in manifests/quotas/ contains both a ResourceQuota and a LimitRange. Deployed automatically by the existing apps/platform/quotas.yaml ArgoCD Application.

Sizing approach:

  • ResourceQuota ceiling = ~2× measured current usage to allow headroom without unbounded scaling
  • LimitRange defaults = sane per-container defaults for new pods (prevents pods from inheriting cluster maximums)
NamespaceQuota (requests.cpu / limits.memory / pods)LimitRange default
ai12 CPU / 80 GiB / 30 pods1 CPU / 2 GiB
langfuse1 CPU / 12 GiB / 15 pods500m / 1 GiB
chat500m CPU / 6 GiB / 15 pods200m / 512 MiB
mail250m CPU / 4 GiB / 10 pods200m / 512 MiB
erp1 CPU / 12 GiB / 25 pods200m / 256 MiB
nextcloud500m CPU / 8 GiB / 15 pods500m / 512 MiB
automation100m CPU / 2 GiB / 10 pods200m / 512 MiB
productivity1 CPU / 10 GiB / 20 pods200m / 256 MiB
sign100m CPU / 2 GiB / 10 pods200m / 512 MiB
collab500m CPU / 4 GiB / 15 pods300m / 256 MiB

LimitRange effect on pods without explicit limits​

LimitRanges inject defaults at pod admission time. Existing pods that were created before the LimitRange existed are not retroactively updated. They will receive the default limits on their next restart (rolling update, node drain, OOMKill restart, etc.).

To verify injection is working for a new pod:

# Create a test pod without resource spec
kubectl run test-lr --image=busybox:1.36 --restart=Never -n erp -- sleep 60

# Confirm LimitRange injected defaults
kubectl get pod test-lr -n erp -o jsonpath='{.spec.containers[0].resources}'
# Expected: {"limits":{"cpu":"200m","memory":"256Mi"},"requests":{"cpu":"50m","memory":"64Mi"}}

kubectl delete pod test-lr -n erp

Summary​

RBAC
cluster-admin → oidc:Direction IT (group) + oidc:kanmegnea (direct user, break-glass) ✅
view → oidc:Cybersécurité, oidc:Audit ✅
dev edit → oidc:Développeurs in *-dev namespaces (4 services) ✅
staging view → oidc:QA in *-staging namespaces (4 services) ✅
prod → no human binding, ArgoCD SA only ✅

Gatekeeper
11 constraints active: 10 deny (0 violations) + 1 warn (158, graduated backlog) ✅
collab removed from require-resource-limits exclusion list (PR #525) ✅
9 namespaces still excluded — upstream chart limits required before removal ⚠️

Resource Quotas & LimitRanges
32/63 namespaces now have ResourceQuota (was 22) ✅
31/63 namespaces have LimitRange (was 21) ✅
All 10 application workload namespaces now bounded ✅
Existing pods without limits get defaults on next restart from LimitRange ✅

Files changed (gitops PR #525)​

FileChange
manifests/quotas/ai.yamlNew — ResourceQuota (50 CPU/80 GiB/30 pods) + LimitRange
manifests/quotas/langfuse.yamlNew
manifests/quotas/chat.yamlNew
manifests/quotas/mail.yamlNew
manifests/quotas/erp.yamlNew
manifests/quotas/nextcloud.yamlNew
manifests/quotas/automation.yamlNew
manifests/quotas/productivity.yamlNew
manifests/quotas/sign.yamlNew
manifests/quotas/collab.yamlNew
manifests/gatekeeper-policies/12-constraint-require-resource-limits.yamlRemove collab exclusion; fix misplaced comment; add per-entry comments