Aller au contenu principal

8 articles tagués avec « self-hosted »

Voir tous les tags

One Registry to Rule Them All: Harbor as the Single Image Entry Point for a Bare-Metal k3s Cluster

· 9 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

If you run a Kubernetes cluster seriously, you eventually hit the Docker Hub wall. One morning your CI pipeline starts failing with toomanyrequests: You have reached your pull rate limit. Or worse — it fails during a cluster recovery at 2 AM, when your nodes are pulling images to restart critical workloads and Docker Hub decides you've exceeded your 100-pulls-per-6-hours anonymous quota.

The standard advice is to add authentication. But that only raises the limit — it doesn't eliminate the dependency on an external service during your most vulnerable moments. The production answer is a pull-through proxy cache, and on a self-hosted k3s cluster, Harbor + k3s mirrors makes every node behave as if Docker Hub were local.

Every Non-Human Identity on a Self-Hosted k8s Platform: A Complete Taxonomy

· 10 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

When you build a production Kubernetes platform from scratch, you accumulate non-human identities faster than you expect. By the time the minicloud platform reached operational maturity — a 5-node bare-metal k3s cluster running 70+ workloads — it had more than 35 distinct non-human identities spanning six different categories. Most of them are invisible during normal operations. You only notice them when one breaks.

This post maps every non-human identity type in use on the platform, explains how they differ, and documents the operational lessons learned from the ones that caused incidents.

Bare Metal First, Cloud at the Edge — Our Hybrid Architecture Decision

· 10 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

Most Kubernetes tutorials start with eksctl create cluster or gcloud container clusters create. A managed control plane, auto-scaling node groups, load balancers that appear with a single annotation. The cloud abstracts the hardware entirely.

We went the other direction. Five physical machines — four ThinkPad laptops and a 2012 MacBook Pro — running k3s, with every workload scheduled and operated by us. No managed control plane. No auto-scaling node group. No cloud load balancer. Just Linux, containerd, and Flannel on iron we can touch.

But we do use cloud services. AWS delivers our email. Cloudflare sits in front of every HTTP request. A Lightsail instance relays our video call UDP traffic. Tailscale connects us to the cluster from anywhere.

This post explains how we decided what goes where, and why the resulting architecture is not a compromise — it is a deliberate design.

Why We Skipped LDAP, Active Directory, and Entra ID — And What We Built Instead

· 8 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

Most enterprise identity architectures did not start as what they are today. They started in 1999 with Active Directory, accumulated LDAP integrations over the following decade, and are now partway through a migration toward cloud-based identity via Microsoft Entra ID — carrying the weight of every layer that came before.

We built our platform from scratch. We never had an on-premises domain. We never configured LDAP. We skipped directly to the protocol stack that enterprises are spending years and significant money trying to reach. This post explains what we chose, why, and how the resulting identity architecture compares to the enterprise standard.

Building an Enterprise AI Gateway on Kubernetes: LiteLLM, Local Models, and Zero-Trust Guardrails

· 9 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

Most enterprise AI deployments make the same architectural mistake early on: they give every team a direct API key to OpenAI or Anthropic and call it done. The result is predictable — no cost visibility, no access control, no audit trail, and sensitive data being sent to cloud APIs without any guardrails.

A proper enterprise AI gateway changes the shape of the problem. Instead of many teams talking to many APIs, you have one endpoint that handles routing, rate limiting, PII scrubbing, caching, and observability. Teams consume it the same way regardless of whether the model is running on your own hardware or on a cloud provider's GPU fleet.

This post covers the full design of such a gateway, built on Kubernetes with LiteLLM as the proxy layer, Ollama and vLLM for local inference, and Presidio for PII protection — with real configuration that is running in production.

We Replaced Ollama With vLLM on CPU-Only Kubernetes — Here Is What Changed

· 9 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

Most write-ups about vLLM versus Ollama assume you have a GPU. The benchmarks show impressive VRAM utilisation. The architecture diagrams include CUDA drivers. The recommendation — vLLM wins — comes with an implicit asterisk: assuming you have the hardware for it.

We did not. Our inference cluster is four ThinkPad laptops, each with an i7-8565U (4 cores, 8 threads, up to 4.6 GHz turbo), 16 GB of RAM, and no GPU of any kind. We ran Ollama first. Then we replaced it with vLLM. This is the honest account of that migration: what broke, what improved, and what the tradeoffs actually look like when you run LLM inference on commodity x86 CPUs.

Why Enterprises Should Run vLLM Instead of Ollama for AI Inference

· 8 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

Ollama is how most teams first run a large language model locally. You install it in five minutes, run ollama pull mistral, and you have a working API. It feels like magic.

Then you try to serve ten users at once. Or a hundred. Or you need to audit every request for compliance. Or your legal team asks where the data goes. That is when you realise Ollama was built for something else entirely.

Goodbye Microsoft Teams: Running Your Own Video Conferencing with Jitsi Meet on Kubernetes

· 19 minutes de lecture
Ingénieur Logiciel & Architecte Cloud

Microsoft Teams costs money. It sends your meeting data to servers you don't control. It requires accounts in a Microsoft tenant. And if the licensing changes, your video conferencing disappears overnight.

Jitsi Meet costs nothing, runs on hardware you own, keeps your data inside your network, and works with any browser — no app install required.

This is the story of how I deployed it on a bare-metal Kubernetes cluster, solved a tricky network problem with SFR 5G mobile users, and wired it up to a company-wide SSO system. By the end of this post, you'll understand how WebRTC video calls actually work, and have a clear map for deploying Jitsi yourself.