Skip to main content

14 posts tagged with "platform-engineering"

View All Tags

Modular-Monolith-First: Sizing Architecture to the Problem, Not the Fashion

· 7 min read
Software Engineer & Cloud Architect

There's a reflex in our industry, and I had it too: a new business domain shows up, and the hand reaches for "…so that's a new microservice." It feels modern. It feels scalable. It feels like what serious engineers do.

On the ktayl-solution information system — a six-node Kubernetes platform running a simulated insurer's entire IS — I'd built four domains exactly that way: policy, claims, underwriting, identity, each its own repo, CI pipeline, database, promotion track. Then I stopped and asked the question that matters more than any framework choice: is this the right architecture, or just the fashionable one?

This post is the answer I arrived at — modular-monolith-first — and, honestly, the more valuable half: the decision I deliberately didn't make.

The Engineering Harness: Turning an AI Agent Into a Reliable Colleague on a Real Information System

· 14 min read
Software Engineer & Cloud Architect

The frontier conversation in AI has quietly shifted. For two years it was "which model is smartest?" Now the people actually shipping agentic systems are asking a different question: "what do you put around the model?"

That surrounding machinery has a name — the harness. The model is the engine. The harness is the chassis, the steering, the seatbelt and the brakes. A brilliant engine bolted to nothing kills you at the first corner; a modest engine in a well-built car gets you home every time. On the ktayl-solution information system — a six-node Kubernetes platform running a simulated insurer's entire IS — I've spent months building the harness that lets an AI agent do real engineering work against live infrastructure without me holding my breath.

This post is that harness, concept by concept. Not the theory — the actual rules I put in place, why each one exists, and the incident that usually forced it.

3-2-1, Not 1: Why I Kept MinIO Instead of Going Cloud-Only for Backups

· 5 min read
Software Engineer & Cloud Architect

While hardening disaster recovery on the ktayl-solution information system, a fair question came up: we already mirror backups to Cloudflare R2 — so why keep a MinIO server running on the controller at all? Couldn't we retire it and go cloud-only?

R2's free tier is generous, cloud object storage is durable, and one less service to operate is always appealing. So it's a reasonable instinct. It's also wrong — and about an hour after I talked myself through why, I proved it the hard way by deleting the wrong bucket prefix. This post is that reasoning, and the honest incident that validated it.

The Layer Below GitOps: How ~200 Lines of Ansible Keep a Bare-Metal Cluster Reproducible

· 8 min read
Software Engineer & Cloud Architect

ArgoCD reconciles everything inside my cluster: 95 applications, from Vault to the AI gateway, all declared in Git and continuously synced. But ArgoCD cannot format a disk, install open-iscsi, or fix a default route. There is a layer below GitOps — the operating system on each bare-metal node — and if that layer isn't codified, "reproducible infrastructure" is a half-truth.

For the ktayl-solution information system, that layer is owned by one small repo: minicloud-ansible. It's about 200 lines of task code across four roles, and it does exactly one job well: make the node OS prerequisites for a 6-node k3s cluster reproducible and auditable. This post is how I use it to operate the organisation's platform.

Self-Hosted Kubernetes: What I Built vs What OpenShift Ships

· 15 min read
Software Engineer & Cloud Architect

OpenShift Container Platform is an opinionated enterprise Kubernetes distribution. My minicloud cluster is a 5-node k3s stack assembled component by component from CNCF projects. After going through the full build — GitOps, observability, secrets, registry, OIDC, ingress, storage replication, chaos testing, security patching, upgrades — I can say with some precision what the difference actually is.

It is not that OpenShift does more. It is that OpenShift has already made every choice you would have to make yourself, packaged those choices as a versioned, tested, supported unit, and enforced them at the architecture level. Whether that is a benefit or a constraint depends entirely on what you are trying to do.

Designing High Availability on Bare-Metal Kubernetes — Layer by Layer

· 13 min read
Software Engineer & Cloud Architect

"High availability" is one of those terms that everyone agrees matters and almost nobody defines precisely. On a managed Kubernetes provider, you tick a checkbox for multi-AZ and move on. On a self-managed cluster, you make six independent architectural decisions, each with its own failure mode, trade-off, and operational cost.

This post documents every HA layer on minicloud — a 5-node k3s cluster on ThinkPad laptops — explains the reasoning behind each trade-off, and maps each layer to how EKS, GKE, and AKS solve the same problem.

Goodbye Microsoft Teams: Running Your Own Video Conferencing with Jitsi Meet on Kubernetes

· 19 min read
Software Engineer & Cloud Architect

Microsoft Teams costs money. It sends your meeting data to servers you don't control. It requires accounts in a Microsoft tenant. And if the licensing changes, your video conferencing disappears overnight.

Jitsi Meet costs nothing, runs on hardware you own, keeps your data inside your network, and works with any browser — no app install required.

This is the story of how I deployed it on a bare-metal Kubernetes cluster, solved a tricky network problem with SFR 5G mobile users, and wired it up to a company-wide SSO system. By the end of this post, you'll understand how WebRTC video calls actually work, and have a clear map for deploying Jitsi yourself.

Your Kubernetes Cluster Doesn't Run etcd — Mine Doesn't Either

· 9 min read
Software Engineer & Cloud Architect

Every Kubernetes tutorial mentions etcd. The architecture diagrams show a three-node etcd cluster with Raft consensus, leader election, and peer replication. If you run EKS, GKE, or AKS, that cluster exists somewhere — you just can never see it.

If you run k3s, you don't have etcd at all.

This post explains what k3s actually uses, what it means to manage backups and recovery yourself, and where that leaves you compared to a managed provider.

Kubernetes Upgrades: What Managed Providers Handle for You and What You Own Yourself

· 12 min read
Software Engineer & Cloud Architect

A Kubernetes upgrade is never just changing a version number. There is a node drain, a binary swap, a control plane migration, a pod eviction sequence, and — if you are running bare-metal — nobody to call when it goes wrong.

Managed Kubernetes providers handle most of that for you. Self-managed clusters make you own all of it. This post documents both sides concretely: what EKS, GKE, and AKS actually do during an upgrade, and what I built to automate the same process on my 5-node k3s cluster running on ThinkPad laptops.

Security Patching a Self-Managed Kubernetes Cluster — Five Layers, Zero Magic

· 17 min read
Software Engineer & Cloud Architect

Managed Kubernetes providers make security patching look simple. You enable auto-upgrade on GKE, you click "update node group" on EKS, and the CVE goes away. What actually happens is that the provider patches the OS image, replaces the node, validates the binary, and restores your workloads — all in the time it takes to refresh the AWS console.

On a self-managed cluster, none of that is automatic. You own the OS. You own the runtime. You own the base images. You own the Helm chart versions. And when a CVE drops, you own the decision about which layer it lives in and which mechanism will close it.

This post documents all five patching layers on minicloud — a 5-node k3s cluster on bare-metal ThinkPads — what's automated, what requires a PR, and where the gaps were and how they were closed.