CloudNativePG is the Postgres standard
:::note Verified, live Summary + pointer page. The decision of record (with the full Langfuse migration runbook + the remaining-debt list) is the ADR in the GitOps repo β this page is the discoverable map entry. Snapshot: 2026-10-06. :::
PostgreSQL is the house relational engine, historically deployed three inconsistent ways (CNPG operator Β· custom-base StatefulSet Β· plain/vendor StatefulSet β see the Data stores inventory). The standard is now explicit:
CloudNativePG (CNPG) is the standard for every Postgres we operate. New databases are born as a CNPG
Cluster; raw/vendor instances are migration candidates (need-first, not a big-bang sweep). Each cluster is a dedicated instance per workload (no cross-namespace shared DB), backs up to R2 with PITR (dailyScheduledBackup+ WAL archiving), and ships aPodMonitor+ stale-backup/unhealthy alerts.
Whyβ
- PITR + automated backup β CNPG archives WAL + base backups to R2; raw StatefulSets have none.
- No shared-instance coupling β a DB hosted inside another app's Postgres (across namespaces) drags a cross-ns dependency + noisy-neighbour + a shared SPOF. A dedicated cluster per workload removes all three.
- Consistent operations β one backup/restore/monitoring shape for every DB (restore drills, alerts).
First migration β Langfuse (2026-10-06)β
Langfuse's metadata DB (projects / API keys / managed prompts β the LLMOps auth anchor) was moved off the
shared postgresql-ai StatefulSet (ai ns, custom -noavx512 base, no PITR) onto a dedicated
langfuse-postgres CNPG cluster in the langfuse namespace. Additive-first (cluster proven healthy +
R2 backup landing before touching Langfuse), then a pg_dumpβrestore at exact parity (71 tables;
70 prompts / 166 models / 3 api_keys), then a one-line host: cutover (same Vault password β drop-in).
New pod logged "412 migrations, No pending migrations to apply" β zero schema drift.
Second migration β Synapse / Matrix (2026-10-06)β
Matrix's DB moved off the -noavx512 postgresql-synapse STS to a dedicated synapse-postgres CNPG cluster
(C collation, as Matrix requires). This verified nothing actually uses pgvector (RAG vectors are in Qdrant)
β the custom -noavx512 base is retireable with stock CNPG. The DB cutover was clean, but it surfaced a
stateful-RWO-PVC incident: Synapse's media PVC (separate from the DB) wedged its Longhorn engine during
the rollout, and manual scale/patch recovery attempts made it worse (selfHeal + the root app-of-apps revert
them). Fixed by holding replicas=0 via a git ignoreDifferences on /spec/replicas and recreating the
near-empty media PVC fresh. Lesson carried forward: quiesce a workload before a DB host cutover when it owns
its own stateful RWO volume. Full postmortem in the ADR.
Operate / verifyβ
kubectl -n langfuse get cluster langfuse-postgres -o wide # Cluster in healthy state
kubectl -n langfuse get scheduledbackup,backups.postgresql.cnpg.io # daily schedule + R2 backups
Full ADRβ
The decision, the step-by-step migration runbook (incl. the peer-auth / 18β17 restore gotchas), and the
remaining-debt list live in the GitOps ADR:
docs/cnpg-standard.md.
Backup/restore operations: Database backup.