Skip to main content

8 posts tagged with "gitops"

View All Tags

The Engineering Harness: Turning an AI Agent Into a Reliable Colleague on a Real Information System

· 14 min read
Software Engineer & Cloud Architect

The frontier conversation in AI has quietly shifted. For two years it was "which model is smartest?" Now the people actually shipping agentic systems are asking a different question: "what do you put around the model?"

That surrounding machinery has a name — the harness. The model is the engine. The harness is the chassis, the steering, the seatbelt and the brakes. A brilliant engine bolted to nothing kills you at the first corner; a modest engine in a well-built car gets you home every time. On the ktayl-solution information system — a six-node Kubernetes platform running a simulated insurer's entire IS — I've spent months building the harness that lets an AI agent do real engineering work against live infrastructure without me holding my breath.

This post is that harness, concept by concept. Not the theory — the actual rules I put in place, why each one exists, and the incident that usually forced it.

The Layer Below GitOps: How ~200 Lines of Ansible Keep a Bare-Metal Cluster Reproducible

· 8 min read
Software Engineer & Cloud Architect

ArgoCD reconciles everything inside my cluster: 95 applications, from Vault to the AI gateway, all declared in Git and continuously synced. But ArgoCD cannot format a disk, install open-iscsi, or fix a default route. There is a layer below GitOps — the operating system on each bare-metal node — and if that layer isn't codified, "reproducible infrastructure" is a half-truth.

For the ktayl-solution information system, that layer is owned by one small repo: minicloud-ansible. It's about 200 lines of task code across four roles, and it does exactly one job well: make the node OS prerequisites for a 6-node k3s cluster reproducible and auditable. This post is how I use it to operate the organisation's platform.

Anatomy of a Cascading Kubernetes Outage: When the First Symptom Is Three Layers From the Cause

· 7 min read
Software Engineer & Cloud Architect

A monitoring alert fired: the platform watchdog was DOWN. Within minutes the picture looked ugly — cluster DNS was refusing connections, and shortly after, distributed-storage volumes stopped rebuilding. A textbook cascade.

This is the full post-mortem: how the outage propagated, why the obvious cause turned out to be a symptom three layers from the root, the fix, and — more importantly — the prevention that shipped so it can't recur the same way. Every change referenced here is a public, reviewable pull request.

The platform in question — minicloud — is a production-grade Kubernetes environment running on refurbished ThinkPads, built and operated solo as a simulation of an enterprise information system. Bare metal, no managed control plane, no safety net. Which is exactly why it's a good teacher.

Moving Your Kubernetes CA Private Key Into Vault PKI — Without Changing a Single Certificate

· 14 min read
Software Engineer & Cloud Architect

Every Kubernetes cluster that uses cert-manager for TLS has the same quiet risk buried in it: the CA private key that signs all your internal certificates is sitting in a Kubernetes secret, stored in plaintext in your cluster's datastore.

On managed clusters with etcd encryption at rest, this is adequately mitigated. On k3s with kine and SQLite — which is how many bare-metal clusters run — the secrets table is plaintext. Anyone who can read state.db from the control plane node can extract your CA private key and forge certificates your entire cluster trusts.

This post covers how we migrated the minicloud root CA private key into HashiCorp Vault's PKI secrets engine, with the same CA cert so nothing else needed to change — no re-trust, no downtime, no changes to any of our 43 Certificate resources.

Self-Hosted Kubernetes: What I Built vs What OpenShift Ships

· 15 min read
Software Engineer & Cloud Architect

OpenShift Container Platform is an opinionated enterprise Kubernetes distribution. My minicloud cluster is a 5-node k3s stack assembled component by component from CNCF projects. After going through the full build — GitOps, observability, secrets, registry, OIDC, ingress, storage replication, chaos testing, security patching, upgrades — I can say with some precision what the difference actually is.

It is not that OpenShift does more. It is that OpenShift has already made every choice you would have to make yourself, packaged those choices as a versioned, tested, supported unit, and enforced them at the architecture level. Whether that is a benefit or a constraint depends entirely on what you are trying to do.

Kubernetes Upgrades: What Managed Providers Handle for You and What You Own Yourself

· 12 min read
Software Engineer & Cloud Architect

A Kubernetes upgrade is never just changing a version number. There is a node drain, a binary swap, a control plane migration, a pod eviction sequence, and — if you are running bare-metal — nobody to call when it goes wrong.

Managed Kubernetes providers handle most of that for you. Self-managed clusters make you own all of it. This post documents both sides concretely: what EKS, GKE, and AKS actually do during an upgrade, and what I built to automate the same process on my 5-node k3s cluster running on ThinkPad laptops.

Automating ConfigMap Reloads: Why We Added Stakater Reloader

· 5 min read
Software Engineer & Cloud Architect

Every time I updated the Homer dashboard config, I had to run kubectl rollout restart deployment/homer -n homer after ArgoCD finished syncing. Same for LiteLLM when routing changed. Same for Backstage after any catalog or proxy update. The pattern was identical every time: push to git, wait for ArgoCD sync, then manually trigger a pod restart.

That is an operational smell. If git is the only write path, the restart should be automatic too.

Platform Engineering on a Budget: Running Production Kubernetes on 5 ThinkPads

· 5 min read
Software Engineer & Cloud Architect

Most cloud platforms hide the infrastructure from you. MAAS provisioning, PXE boot sequences, NIC bonding, storage backends, certificate chains — all of it abstracted behind a few CLI flags or a dashboard. That abstraction is valuable in production, but it can also keep engineers at arm's length from the system they're supposed to understand deeply.

This project started from a simple question: what does it actually take to build a production-grade Kubernetes platform from scratch?