Skip to main content

Security Evaluation — Runtime, Admission, Network, Secrets & RBAC

Evaluated 2026-08-02. Two gaps found and closed in gitops PR #520.

CriterionBeforeAfter
Runtime detection (Falco)✅ 5/5 nodes, routed to AlertmanagerNo change needed
Admission control (Gatekeeper)✅ 11 constraints; 25 violations from stale debug pods🔧 Deleted stale pods → 0 violations
Network isolation⚠️ 8 namespaces with no NetworkPolicy🔧 49/49 namespaces now isolated
Secrets management✅ All ESO-synced from VaultNo change needed
RBAC✅ OIDC groups, least-privilegeNo change needed

1 — Runtime detection (Falco)​

Status: ✅ Pass

Falco DaemonSet runs on all 5 nodes (falco-* pods, 5/5 Running). Falcosidekick forwards events to Alertmanager at minimumpriority: warning:

Falco (kernel eBPF probe)
→ Falcosidekick (webhook bridge)
→ Alertmanager at kps-alertmanager.monitoring:9093
→ webhook-critical (email + Slack)

Events below warning (notice, debug, informational) are captured in Falco's own JSON stdout (scraped by Loki via the node-exporter log path) but do not page anyone — intentional, to avoid alert noise.

What Falco detects on this cluster:

Rule categoryExamples
Container escapesproc_name in a privileged container, write to /proc/sysrq-trigger
Credential accessRead from /etc/shadow, access to /root/.ssh/
Lateral movementUnexpected outbound connection from a non-network workload
PersistenceNew binary written to /usr/bin inside a running container
K8s API abusekubectl exec into a running pod by a non-admin user

Default Falco ruleset is active. Custom rules can be added to helm-values/minicloud-1/falco-values.yaml under customRules:.


2 — Admission control (Gatekeeper)​

Status: ✅ Pass (0 violations on deny constraints after cleanup)

OPA Gatekeeper v3 runs in gatekeeper-system with 3 replicas (controller, webhook, audit). 11 constraints are active:

ConstraintModeWhat it blocks
block-latest-tagdeny:latest image tag in Deployments/StatefulSets/DaemonSets/Pods
no-privileged-containersdenysecurityContext.privileged: true
require-resource-limitsdenyContainers without CPU+memory limits
allowed-registriesdenyImages not from: harbor.*, docker.io, ghcr.io, quay.io, registry.k8s.io, oci.external-secrets.io, ecr-public.aws.com
require-non-rootdenyContainers without runAsNonRoot: true or explicit non-zero UID
no-privilege-escalationdenyContainers missing allowPrivilegeEscalation: false
require-ingress-tlsdenyIngress without a tls: block
no-loadbalancer-in-devdenytype: LoadBalancer Services in *-dev namespaces
no-host-pathdenyhostPath volumes (kured and node-problem-detector are excluded)
block-net-rawdenyNET_RAW capability
block-capabilitiesdenyAdding any capabilities not in the default set
require-seccompwarnContainers without seccompProfile: RuntimeDefault

require-seccomp is intentionally in warn mode. Switching to deny would block a large number of existing workloads (158 containers currently lack a seccomp profile, including many upstream Helm charts). The warn mode collects audit data; enforcement will be graduated namespace by namespace.

Gap closed — stale debug pods​

Gatekeeper's audit controller runs every 1 minute and counts violations on existing resources, not only on admission. Five Completed debug pods in default namespace (tmp2, tmp-ftest, tmp-persistent, tmp-trace, diag) were triggering 25 violations across 4 constraints:

block-latest-tag 5 violations (busybox image, no tag)
allowed-registries 5 violations (busybox without registry prefix)
require-non-root 5 violations (no runAsNonRoot)
no-privilege-escalation 5 violations (no allowPrivilegeEscalation:false)

The pods had been Completed for 3+ days with no purpose. Deleted directly:

kubectl delete pod tmp2 tmp-ftest tmp-persistent tmp-trace diag -n default

Gatekeeper violations on deny constraints after cleanup: 0.


3 — Network isolation​

Status: ✅ Pass (49/49 application namespaces isolated as of PR #520)

Before (gap)​

The apps/platform/network-policies.yaml ArgoCD Application deployed NetworkPolicies from manifests/network-policies/ to 41 namespaces, but 8 application namespaces had zero isolation:

NamespaceWorkload
automationn8n workflow automation
chatMatrix Synapse + Element Web
collabJitsi Meet (JVB uses hostNetwork — exempt from pod NP)
erpERPNext (MariaDB + workers)
mailStalwart mail server (SMTP/IMAP/webmail)
nextcloudNextcloud + OnlyOffice
productivityPlane CE (web + api + worker + beat)
signDocuSeal document signing

Any pod in any namespace could reach these workloads on any port — including workloads that should be internal-only (ERPNext admin, MariaDB, Plane worker).

Fix​

Added the standard 5-policy set to each namespace:

# 1. Drop all inbound traffic by default
default-deny-ingress (K8s NetworkPolicy, policyTypes: [Ingress])

# 2. Allow pods in the same namespace to communicate
allow-same-namespace (ingress from podSelector: {})

# 3. Allow ingress-nginx to forward HTTP/HTTPS traffic in
allow-ingress-nginx (ingress from namespaceSelector: ingress-nginx)

# 4. Allow Prometheus to scrape metrics
allow-monitoring-scrape (ingress from namespaceSelector: monitoring + observability)

# 5. Allow kubelet health probes (host/node traffic)
allow-node-entities (CiliumNetworkPolicy, fromEntities: [host, remote-node])

The CiliumNetworkPolicy/allow-node-entities is required because Cilium uses endpoint identity rather than IP CIDRs. Kubelet liveness/readiness probes originate from the host entity — an ipBlock: cidr: 10.0.0.0/8 NetworkPolicy would silently drop them.

NodePort and LoadBalancer traffic entering via the node's network stack is matched by the host entity and is therefore permitted for all services that use those exposure types (mail SMTP/IMAP, Jitsi JVB media).

Verification​

# All 49 namespaces with default-deny
kubectl --context minicloud get networkpolicy -A | grep default-deny | wc -l

# NetworkPolicies live in chat
kubectl --context minicloud get networkpolicy -n chat
kubectl --context minicloud get ciliumnetworkpolicy -n chat

# Matrix still reachable (allow-ingress-nginx works)
/usr/bin/curl --cacert ~/minicloud-ca.crt -sI https://matrix.10.0.0.200.nip.io | head -2

# Isolation test: pod in monitoring cannot reach chat on arbitrary port
kubectl --context minicloud run test --rm -it --restart=Never \
-n monitoring --image=busybox:1.36 -- \
nc -zv matrix-synapse.chat.svc.cluster.local 8448
# Expected: connection refused or timeout (denied by default-deny-ingress)

4 — Secrets management​

Status: ✅ Pass

All application secrets follow the same pipeline:

Vault (secret/platform/<service>)
→ External Secrets Operator ClusterSecretStore (vault-backend)
→ ExternalSecret CR (per namespace, refreshInterval: 1h)
→ k8s Secret (in-namespace, reconciled every hour)

20+ ExternalSecrets across all namespaces, all SecretSynced.

No secrets in git. Confirmed by:

git -C ~/Developer/cloudplateform/minicloud-gitops log --all -p -- '**/*.yaml' \
| grep -iE 'password|secret|token|key' | grep -v '#\|metadata\|selector'
# Returns: only field names (secretKeyRef, tokenKey, etc.) — no values

Secrets that are intentionally NOT in ESO (cluster-internal, chart-generated):

  • argocd-redis — Redis password generated at install, not user-configurable
  • matrix-synapse-signingkey — signing key must persist across pod restarts; ESO rotation would break federation
  • gatekeeper-webhook-server-cert — managed by cert-manager, not ESO

5 — RBAC​

Status: ✅ Pass

ClusterRoleBindings (cluster-wide)​

Defined in manifests/rbac/00-oidc-clusterrolebindings.yaml:

GroupClusterRoleScope
oidc:Platform Adminscluster-adminFull cluster (ArgoCD, cert-manager management)
oidc:ViewersviewRead-only across all namespaces

RoleBindings (namespace-scoped)​

Defined in manifests/rbac/02-env-rolebindings.yaml:

GroupRoleNamespaces
oidc:Développeurseditplatform-demo-dev, minicloud-plane-dev
oidc:QAviewplatform-demo-staging, minicloud-plane-staging

Production namespaces have no human RoleBinding. Only the ArgoCD ServiceAccount (argocd-application-controller) has write access to *-prod namespaces. Human access to production requires the ArgoCD UI or cluster-admin escalation (Authentik OIDC MFA gate).

Accepted risks​

manifests/rbac/01-rbac-accepted-risks.yaml documents workloads that require elevated cluster-scoped permissions and the justification:

  • Falco (cluster-admin equivalent for eBPF probe access)
  • Velero (needs to read/write all namespaces for backup/restore)
  • VPA admission controller (needs to patch Pod specs at admission)
  • Chaos Mesh (needs to create/delete pods and network routes)

Security summary​

Falco → Alertmanager → Slack + email at warning priority ✅
Gatekeeper → 10 deny constraints, 0 violations ✅
1 warn constraint (require-seccomp, 158 items) ⚠️ warn-only
Network isolation → 49/49 namespaces have default-deny + Cilium NP ✅
Secrets → ESO + Vault, all synced, no secrets in git ✅
RBAC → OIDC groups, least-privilege, no human prod access ✅

The single remaining item (seccomp profiles) is a standard backlog item for production Kubernetes deployments. The path to closing it:

  1. Identify the 158 containers from kubectl get K8sRequireSeccomp require-seccomp
  2. Add seccompProfile: RuntimeDefault to each workload's securityContext
  3. Once all are compliant, switch the constraint from warn to deny