Skip to main content

3 posts tagged with "argocd"

View All Tags

Anatomy of a Cascading Kubernetes Outage: When the First Symptom Is Three Layers From the Cause

· 7 min read
Software Engineer & Cloud Architect

A monitoring alert fired: the platform watchdog was DOWN. Within minutes the picture looked ugly — cluster DNS was refusing connections, and shortly after, distributed-storage volumes stopped rebuilding. A textbook cascade.

This is the full post-mortem: how the outage propagated, why the obvious cause turned out to be a symptom three layers from the root, the fix, and — more importantly — the prevention that shipped so it can't recur the same way. Every change referenced here is a public, reviewable pull request.

The platform in question — minicloud — is a production-grade Kubernetes environment running on refurbished ThinkPads, built and operated solo as a simulation of an enterprise information system. Bare metal, no managed control plane, no safety net. Which is exactly why it's a good teacher.

Moving Your Kubernetes CA Private Key Into Vault PKI — Without Changing a Single Certificate

· 14 min read
Software Engineer & Cloud Architect

Every Kubernetes cluster that uses cert-manager for TLS has the same quiet risk buried in it: the CA private key that signs all your internal certificates is sitting in a Kubernetes secret, stored in plaintext in your cluster's datastore.

On managed clusters with etcd encryption at rest, this is adequately mitigated. On k3s with kine and SQLite — which is how many bare-metal clusters run — the secrets table is plaintext. Anyone who can read state.db from the control plane node can extract your CA private key and forge certificates your entire cluster trusts.

This post covers how we migrated the minicloud root CA private key into HashiCorp Vault's PKI secrets engine, with the same CA cert so nothing else needed to change — no re-trust, no downtime, no changes to any of our 43 Certificate resources.

Automating ConfigMap Reloads: Why We Added Stakater Reloader

· 5 min read
Software Engineer & Cloud Architect

Every time I updated the Homer dashboard config, I had to run kubectl rollout restart deployment/homer -n homer after ArgoCD finished syncing. Same for LiteLLM when routing changed. Same for Backstage after any catalog or proxy update. The pattern was identical every time: push to git, wait for ArgoCD sync, then manually trigger a pod restart.

That is an operational smell. If git is the only write path, the restart should be automatic too.