Skip to main content

14 posts tagged with "k3s"

View All Tags

The Layer Below GitOps: How ~200 Lines of Ansible Keep a Bare-Metal Cluster Reproducible

· 8 min read
Software Engineer & Cloud Architect

ArgoCD reconciles everything inside my cluster: 95 applications, from Vault to the AI gateway, all declared in Git and continuously synced. But ArgoCD cannot format a disk, install open-iscsi, or fix a default route. There is a layer below GitOps — the operating system on each bare-metal node — and if that layer isn't codified, "reproducible infrastructure" is a half-truth.

For the ktayl-solution information system, that layer is owned by one small repo: minicloud-ansible. It's about 200 lines of task code across four roles, and it does exactly one job well: make the node OS prerequisites for a 6-node k3s cluster reproducible and auditable. This post is how I use it to operate the organisation's platform.

Anatomy of a Cascading Kubernetes Outage: When the First Symptom Is Three Layers From the Cause

· 7 min read
Software Engineer & Cloud Architect

A monitoring alert fired: the platform watchdog was DOWN. Within minutes the picture looked ugly — cluster DNS was refusing connections, and shortly after, distributed-storage volumes stopped rebuilding. A textbook cascade.

This is the full post-mortem: how the outage propagated, why the obvious cause turned out to be a symptom three layers from the root, the fix, and — more importantly — the prevention that shipped so it can't recur the same way. Every change referenced here is a public, reviewable pull request.

The platform in question — minicloud — is a production-grade Kubernetes environment running on refurbished ThinkPads, built and operated solo as a simulation of an enterprise information system. Bare metal, no managed control plane, no safety net. Which is exactly why it's a good teacher.

One Registry to Rule Them All: Harbor as the Single Image Entry Point for a Bare-Metal k3s Cluster

· 9 min read
Software Engineer & Cloud Architect

If you run a Kubernetes cluster seriously, you eventually hit the Docker Hub wall. One morning your CI pipeline starts failing with toomanyrequests: You have reached your pull rate limit. Or worse — it fails during a cluster recovery at 2 AM, when your nodes are pulling images to restart critical workloads and Docker Hub decides you've exceeded your 100-pulls-per-6-hours anonymous quota.

The standard advice is to add authentication. But that only raises the limit — it doesn't eliminate the dependency on an external service during your most vulnerable moments. The production answer is a pull-through proxy cache, and on a self-hosted k3s cluster, Harbor + k3s mirrors makes every node behave as if Docker Hub were local.

Every Non-Human Identity on a Self-Hosted k8s Platform: A Complete Taxonomy

· 10 min read
Software Engineer & Cloud Architect

When you build a production Kubernetes platform from scratch, you accumulate non-human identities faster than you expect. By the time the minicloud platform reached operational maturity — a 5-node bare-metal k3s cluster running 70+ workloads — it had more than 35 distinct non-human identities spanning six different categories. Most of them are invisible during normal operations. You only notice them when one breaks.

This post maps every non-human identity type in use on the platform, explains how they differ, and documents the operational lessons learned from the ones that caused incidents.

Moving Your Kubernetes CA Private Key Into Vault PKI — Without Changing a Single Certificate

· 14 min read
Software Engineer & Cloud Architect

Every Kubernetes cluster that uses cert-manager for TLS has the same quiet risk buried in it: the CA private key that signs all your internal certificates is sitting in a Kubernetes secret, stored in plaintext in your cluster's datastore.

On managed clusters with etcd encryption at rest, this is adequately mitigated. On k3s with kine and SQLite — which is how many bare-metal clusters run — the secrets table is plaintext. Anyone who can read state.db from the control plane node can extract your CA private key and forge certificates your entire cluster trusts.

This post covers how we migrated the minicloud root CA private key into HashiCorp Vault's PKI secrets engine, with the same CA cert so nothing else needed to change — no re-trust, no downtime, no changes to any of our 43 Certificate resources.

Bare Metal First, Cloud at the Edge — Our Hybrid Architecture Decision

· 10 min read
Software Engineer & Cloud Architect

Most Kubernetes tutorials start with eksctl create cluster or gcloud container clusters create. A managed control plane, auto-scaling node groups, load balancers that appear with a single annotation. The cloud abstracts the hardware entirely.

We went the other direction. Five physical machines — four ThinkPad laptops and a 2012 MacBook Pro — running k3s, with every workload scheduled and operated by us. No managed control plane. No auto-scaling node group. No cloud load balancer. Just Linux, containerd, and Flannel on iron we can touch.

But we do use cloud services. AWS delivers our email. Cloudflare sits in front of every HTTP request. A Lightsail instance relays our video call UDP traffic. Tailscale connects us to the cluster from anywhere.

This post explains how we decided what goes where, and why the resulting architecture is not a compromise — it is a deliberate design.

Self-Hosted Kubernetes: What I Built vs What OpenShift Ships

· 15 min read
Software Engineer & Cloud Architect

OpenShift Container Platform is an opinionated enterprise Kubernetes distribution. My minicloud cluster is a 5-node k3s stack assembled component by component from CNCF projects. After going through the full build — GitOps, observability, secrets, registry, OIDC, ingress, storage replication, chaos testing, security patching, upgrades — I can say with some precision what the difference actually is.

It is not that OpenShift does more. It is that OpenShift has already made every choice you would have to make yourself, packaged those choices as a versioned, tested, supported unit, and enforced them at the architecture level. Whether that is a benefit or a constraint depends entirely on what you are trying to do.

Designing High Availability on Bare-Metal Kubernetes — Layer by Layer

· 13 min read
Software Engineer & Cloud Architect

"High availability" is one of those terms that everyone agrees matters and almost nobody defines precisely. On a managed Kubernetes provider, you tick a checkbox for multi-AZ and move on. On a self-managed cluster, you make six independent architectural decisions, each with its own failure mode, trade-off, and operational cost.

This post documents every HA layer on minicloud — a 5-node k3s cluster on ThinkPad laptops — explains the reasoning behind each trade-off, and maps each layer to how EKS, GKE, and AKS solve the same problem.

Your Kubernetes Cluster Doesn't Run etcd — Mine Doesn't Either

· 9 min read
Software Engineer & Cloud Architect

Every Kubernetes tutorial mentions etcd. The architecture diagrams show a three-node etcd cluster with Raft consensus, leader election, and peer replication. If you run EKS, GKE, or AKS, that cluster exists somewhere — you just can never see it.

If you run k3s, you don't have etcd at all.

This post explains what k3s actually uses, what it means to manage backups and recovery yourself, and where that leaves you compared to a managed provider.

Kubernetes Upgrades: What Managed Providers Handle for You and What You Own Yourself

· 12 min read
Software Engineer & Cloud Architect

A Kubernetes upgrade is never just changing a version number. There is a node drain, a binary swap, a control plane migration, a pod eviction sequence, and — if you are running bare-metal — nobody to call when it goes wrong.

Managed Kubernetes providers handle most of that for you. Self-managed clusters make you own all of it. This post documents both sides concretely: what EKS, GKE, and AKS actually do during an upgrade, and what I built to automate the same process on my 5-node k3s cluster running on ThinkPad laptops.