Security Hardening Roadmap
This page covers the active security hardening sprint: Trivy supply-chain scanning, Falco runtime detection, kube-bench CIS compliance, Checkov IaC scanning, and the Wazuh SIEM target for DORA compliance.
Offensive vs. Defensive Tools β The Right Mental Modelβ
Tools like Burp Suite, OWASP ZAP, Metasploit, Nmap, and Wireshark are offensive / assessment tools β they test whether your defenses hold. They are not installed as services and not run continuously. The correct cadence is:
| Tool | Purpose | When to use |
|---|---|---|
| Burp Suite / OWASP ZAP | Web app vulnerability scanning | Quarterly pentest of public-facing services |
| Nmap | Network surface discovery | Before and after firewall changes |
| kube-hunter | Kubernetes attack surface | After major cluster upgrades |
| Metasploit | Exploit validation | Authorized pentest engagements only |
| Wireshark | Packet capture / protocol analysis | Incident investigation |
Day-to-day protection comes from the defensive stack β Falco, Trivy, Gatekeeper, NetworkPolicies, Vault β which is what this page covers.
On NAT exposure: The cluster is behind SFR box NAT and Cloudflare WAF, which eliminates the most common internet-facing attack vectors. Risk is still present from supply-chain (compromised images, malicious dependencies), insider threat, and misconfiguration. The hardening roadmap addresses all three.
1. Trivy β Supply-Chain Scanning in Harborβ
Status: deployed and active. Harbor has a dedicated Trivy scanner pod (harbor-trivy-0 in the harbor namespace).
CI push image β Harbor registry β Trivy scans automatically β CVE report attached to image
β
If CRITICAL CVE found:
ArgoCD pulls tagged image β policy blocks deployment
Enabling scan-on-push and CVE blockingβ
In Harbor UI β Administration β Interrogation Services β Vulnerability tab:
- Scan on push: enable β every pushed image is scanned automatically
- Prevent vulnerable images from running: set threshold to Critical
Or via Harbor API (from controller, once Harbor is healthy):
# Enable scan-on-push system-wide
curl -s --cacert ~/minicloud-ca.crt \
-u "admin:$(cat ~/.harbor-admin)" \
-X PUT "https://harbor.10.0.0.200.nip.io/api/v2.0/configurations" \
-H 'Content-Type: application/json' \
-d '{"scan_all_policy":{"parameter":{"daily_time":0},"type":"daily"}}'
# Enable vulnerability prevention at Critical severity
curl -s --cacert ~/minicloud-ca.crt \
-u "admin:$(cat ~/.harbor-admin)" \
-X PUT "https://harbor.10.0.0.200.nip.io/api/v2.0/configurations" \
-H 'Content-Type: application/json' \
-d '{"prevent_vul_enabled":true,"vulnerability_severity":"critical"}'
Per-project scan policyβ
Set on each project (e.g., library):
curl -s --cacert ~/minicloud-ca.crt \
-u "admin:$(cat ~/.harbor-admin)" \
-X PUT "https://harbor.10.0.0.200.nip.io/api/v2.0/projects/library" \
-H 'Content-Type: application/json' \
-d '{"metadata":{"auto_scan":"true","prevent_vul":"true","severity":"critical"}}'
Quick checkβ
# View scan results for the platform-demo image
/usr/bin/curl -sk --cacert ~/minicloud-ca.crt \
-u "admin:$(cat ~/.harbor-admin)" \
"https://harbor.10.0.0.200.nip.io/api/v2.0/projects/library/repositories/platform-demo/artifacts?with_scan_overview=true" \
| python3 -m json.tool | grep -A3 '"severity"'
2. Falco β Runtime Threat Detectionβ
Status: deployed and healthy. Falco v0.44.1 runs as a DaemonSet across all 5 nodes (4 ThinkPads + swift-mac), with 2 Falcosidekick pods routing alerts to Grafana.
kubectl get pods -n falco
# 5Γ falco-xxxxx (DaemonSet, one per node)
# 2Γ falco-falcosidekick-xxxxx (alert router)
What it detectsβ
Falco uses eBPF (modern_ebpf driver) to intercept kernel syscalls. Default ruleset covers:
| Rule category | Examples |
|---|---|
| Shell in container | execve of bash/sh inside a running container |
| Sensitive file reads | Open of /etc/shadow, /etc/kubernetes/pki/ |
| Network anomalies | Unexpected outbound connections from known pods |
| Privilege escalation | setuid, chmod +s, sudo inside containers |
| Package manager in container | apt-get, yum, pip install at runtime |
View live alertsβ
# Stream Falco alerts from any node
ssh controller "kubectl logs -n falco -l app.kubernetes.io/name=falco -f --since=1h | grep 'Warning\|Error\|Critical'"
# Or via Grafana β Explore β Loki β {namespace="falco"}
Falcosidekick routes to Grafana/Lokiβ
Alerts are forwarded by Falcosidekick to Loki (configured in helm-values/minicloud-1/falco-values.yaml). Search in Grafana:
{namespace="falco"} |= "Warning"
Custom rulesβ
Add custom rules in helm-values/minicloud-1/falco-values.yaml under customRules::
customRules:
insurance-rules.yaml: |
- rule: Write to ERPNext config dir
desc: Alert on writes inside the ERPNext config directory
condition: open_write and container and fd.name startswith /home/frappe/frappe-bench/sites/
output: "Suspicious write to ERPNext sites dir (user=%user.name file=%fd.name)"
priority: WARNING
tags: [erp, filesystem]
3. kube-bench β CIS Kubernetes Benchmarkβ
Status: completed. kube-bench v0.9.4 was run against the cluster with the k3s-cis-1.8 benchmark.
Results summary:
| Node | PASS | FAIL | WARN | All FAILs |
|---|---|---|---|---|
| set-hog (control plane) | 49 | 6 | 55 | k3s false positives |
| fast-skunk (worker) | 11 | 5 | 37 | k3s false positives |
All 6 FAILs are k3s false positives β kube-bench greps /proc/<pid>/cmdline for flag names, but k3s embeds all control-plane components into one binary. The flags exist (anonymous-auth=false, TLS certs auto-provisioned) but are invisible to kube-bench. See Phase 29 β kube-bench for full analysis.
Active WARN items (planned):
| Check | What it needs | Priority |
|---|---|---|
5.3.2 NetworkPolicies in all namespaces | Deploy deny-all + allow-required per namespace | High |
5.7.2 seccomp RuntimeDefault | Add seccompProfile.type: RuntimeDefault to pod securityContext | Medium |
5.7.3 SecurityContext on pods | runAsNonRoot: true, readOnlyRootFilesystem: true | Medium |
5.1.5 SA token automount | automountServiceAccountToken: false on default ServiceAccounts | Low |
Re-run kube-benchβ
# Download and push to control-plane node
ssh controller "
curl -sL https://github.com/aquasecurity/kube-bench/releases/download/v0.9.4/kube-bench_0.9.4_linux_amd64.tar.gz \
-o /tmp/kube-bench.tar.gz && tar -xzf /tmp/kube-bench.tar.gz -C /tmp/
scp /tmp/kube-bench ubuntu@10.0.0.2:/tmp/kube-bench
scp -r /tmp/cfg ubuntu@10.0.0.2:/tmp/cfg"
ssh ubuntu@10.0.0.2 'sudo /tmp/kube-bench run \
--config-dir /tmp/cfg --benchmark k3s-cis-1.8 \
--targets master,node,etcd,policies 2>&1' \
| grep -E '== Summary|PASS|FAIL|WARN'
4. Checkov β IaC Security Scanning in CIβ
Status: deployed 2026-08-16. Checkov runs on every PR to main in minicloud-gitops when manifests/, services/, apps/, or helm-values/ change.
Workflow: .github/workflows/checkov.yml
PR to main (paths: manifests/**, services/**, apps/**, helm-values/**)
β
βΌ
bridgecrewio/checkov-action@v12
β framework: kubernetes,helm
β output: SARIF β GitHub Code Scanning
βΌ
Findings appear in PR "Security" tab
Blocking on new Critical/High misconfigs
What Checkov catchesβ
| Category | Example findings |
|---|---|
| Pod security | Missing securityContext.runAsNonRoot, privileged containers |
| RBAC | Wildcard verbs: ["*"] in ClusterRoles, default SA with token |
| Network | Missing NetworkPolicy for namespaces |
| Secrets | Secrets in env vars (should be ESO β Vault) |
| Resource limits | Containers without CPU/memory limits |
| Image hygiene | latest tag, no digest pinning |
Skipped checks (with rationale)β
| Check | Reason skipped |
|---|---|
CKV_K8S_8/9 | Liveness/readiness probes: SRE responsibility per workload |
CKV_K8S_28/14/43 | Image tag/digest: CI pins at build time, ArgoCD syncs pinned tags |
CKV_K8S_35 | Secret env vars: we use ESO β Vault for all platform secrets |
CKV_K8S_36 | Drop ALL capabilities: Gatekeeper (Phase 27) enforces this at admission |
CKV2_K8S_6 | Wildcard RBAC: Helm charts often include; tracked as separate hardening item |
Local runβ
cd ~/Developer/cloudplateform/minicloud-gitops
pip install checkov
checkov -d . --framework kubernetes,helm \
--skip-check CKV_K8S_8,CKV_K8S_9,CKV_K8S_28,CKV_K8S_14,CKV_K8S_43,CKV_K8S_35,CKV_K8S_36,CKV2_K8S_6 \
--output cli
5. Wazuh β SIEM / DORA Compliance (Planned)β
Status: planned. Wazuh is the primary gap between "secure" and "auditable" for a DORA-compliant insurance IS.
What Wazuh addsβ
| Capability | Benefit for ktayl IS |
|---|---|
| File Integrity Monitoring (FIM) | Detect unauthorized changes to config files, systemd units, k3s config |
| Log centralization | Aggregate controller + node + k8s audit logs into one searchable store |
| Compliance dashboards | PCI-DSS, NIST CSF, GDPR/DORA compliance reports out of the box |
| Vulnerability detection | Agent-based CVE scanning at OS level (complements Trivy at container level) |
| Incident alerting | MITRE ATT&CK mapping, alert correlation across all nodes |
Architectureβ
All 5 nodes (set-hog, fast-skunk, fast-heron, star-kitten, swift-mac)
βββ Wazuh agent (installed via Ansible)
β
βΌ encrypted OSSEC protocol
controller (or dedicated node)
βββ Wazuh manager + Indexer (OpenSearch) + Dashboard
β
βΌ
Grafana (via OpenSearch data source)
Why it's heavyweightβ
Wazuh Indexer (OpenSearch) requires ~4Gi RAM and ~50Gi disk for a cluster this size with 30-day retention. This pushes cluster memory close to its limit. One of the "+2 ThinkPads" (roadmap issue #239) β loving-gannet β is now provisioned (cluster is 6 nodes / 44 cores); deploy once headroom is confirmed.
DORA requirements satisfied by Wazuhβ
| DORA Article | Requirement | Wazuh covers |
|---|---|---|
| Art. 9 | ICT incident classification and reporting | Alert severity taxonomy, incident timeline |
| Art. 10 | ICT incident response and recovery | Real-time detection β automated response rules |
| Art. 13 | Digital operational resilience testing | Audit logs for penetration tests |
| Art. 17 | ICT third-party risk | Third-party software change detection (FIM) |
Deployment targetβ
After +2 ThinkPads (roadmap issue #239), target cluster RAM: ~128Gi β Wazuh on dedicated node.
Current Security Posture Summaryβ
| Layer | Tool | Status |
|---|---|---|
| Supply chain β container images | Harbor Trivy | β Deployed (harbor-trivy-0 Running) |
| Supply chain β IaC configs | Checkov CI | β
Live (.github/workflows/checkov.yml) |
| Runtime detection | Falco 0.44.1 | β All 5 nodes |
| CIS compliance | kube-bench 0.9.4 | β Run complete, 6 false positives documented |
| Admission control | Gatekeeper (18 constraints) | β Phase 27 |
| PKI / Certificate authority | Vault PKI | β CA migrated from k8s secret 2026-08-15 |
| Secret management | HashiCorp Vault | β Phase 26 |
| RBAC | Authentik OIDC + k8s RBAC | β Phase 28 |
| Network segmentation | Cilium + NetworkPolicies | β Partial (monitoring ns) |
| Edge security | Cloudflare WAF + Zero Trust | β Phase 25 |
| SIEM / DORA audit | Wazuh | π΅ Planned post-+2 ThinkPads |
| OS hardening | host-hardening.md | β Phase 29 |
Done Whenβ
β Trivy scanner pod Running in harbor ns (harbor-trivy-0)
β Harbor scan-on-push documented; CVE blocking policy defined
β Falco 0.44.1 DaemonSet: 5/5 nodes healthy
β Falcosidekick β Loki β Grafana alert pipeline active
β kube-bench k3s-cis-1.8: 49+11 PASS, all FAILs are false positives
β Checkov CI workflow deployed (PR #TBD, minicloud-gitops)
β SARIF output β GitHub Code Scanning on every PR to main
β Offensive tools framing documented (pentest vs. day-to-day protection)
π΅ Wazuh SIEM: planned after +2 ThinkPads (roadmap #239)