Kicking the hornet's nest
Breaking my homelab on purpose
A Kubernetes cluster is a fantastically complex structure with countless moving parts, and therefore countless points of failure. Every guide will happily walk you through standing one up, but very few of them tell you what it looks like when one of those parts gives out, how the rest of the cluster responds, and what you're supposed to do about it. The best way I could think of to find out was to build one and then make it fall over, one layer at a time.
I'd wanted to try Talos Linux for a while, since I have an interest in projects that stretch Linux into unusual shapes, and an OS purpose-built to run Kubernetes, with no shell, no SSH and no package manager, is about as unusual as it gets. My excuse was giving my Forgejo instance some runners for CI and a pair of cron jobs that I was running on my desktop and was meaning to move to the homelab. I built the cluster on a single Talos VM on my Proxmox host, pairing with an AI coding assistant, and then set about knocking it down. Most of the failures below were deliberate. A few were my own doing and entirely unplanned, and I've kept those in, since the cluster had to respond to them all the same. Everything lives in a git repository, which you can find here.
The setup
The cluster is called sakuya, and everything on it is declared in git and applied by Flux. Nothing on it is reachable from the public internet. The rest of this post walks up the stack from the bottom, in the order shown here:
+- - - - - - - - - - - +
' sakuya: '
' '
' +------------------+ '
' | CI | '
' | Forgejo runners | '
' +------------------+ '
' ^ '
' | '
' | '
' +------------------+ '
' | Apps | '
' | CronJobs | '
' +------------------+ '
' ^ '
' | '
' | '
' +------------------+ ' +---------------+
' | Database | ' WAL | Cloudflare R2 |
' | CloudNativePG | ' -----> | |
' +------------------+ ' +---------------+
' ^ '
' | '
' | '
' +------------------+ '
' | Network | '
' | MetalLB, Traefik | '
' +------------------+ '
' ^ '
' | '
' | '
+---------+ ' +------------------+ '
| nitori | git ' | GitOps | '
| Forgejo | -----> ' | Flux + SOPS | '
+---------+ ' +------------------+ '
' ^ '
' | '
' | '
' +------------------+ '
' | Node | '
' | Talos on Proxmox | '
' +------------------+ '
' '
+- - - - - - - - - - - +
| Layer | Chose | Over | Why |
|---|---|---|---|
| Node | Talos | k3s on Ubuntu | Immutable and API-driven, with nothing to hand-patch |
| GitOps | Flux, SOPS + age | Argo CD, Sealed Secrets | Everything is a CRD, secrets live encrypted in git, and it's light on a 12 GB node |
| Network | MetalLB, Traefik, cert-manager | ingress-nginx | ingress-nginx is archived; Traefik speaks Gateway API |
| Database | CloudNativePG, backups to R2 | A plain StatefulSet, in-cluster MinIO | Offsite backups |
| Apps | CronJobs | My desktop's crontab | I needed an excuse, frankly |
| CI | Forgejo runners with Docker-in-Docker | The host's Docker socket | Talos has no Docker |
Leaking the keys to the cluster
The first failure was not really planned. The cluster was about an hour old, and I was encrypting the Talos secrets file, which holds every certificate authority private key for the cluster, before committing it. SOPS matches its encryption rules against the input filename, I'd written my rule for the output name, and so it found no rule, wrote an empty file, and carried on without complaint. My check that encryption had worked was a diff against the plaintext, which on an empty file prints the whole plaintext, so every CA key for the cluster ended up printed into a session the AI assistant was logging. Nothing reached git, but the keys had gone to Anthropic's servers.
Talos can rotate its CAs, but not all of them rotate cleanly, and the cluster had nothing on it yet, so I decided to treat the whole root of trust as compromised and rebuild from scratch. That's where I made my second mistake. I assumed a bare talosctl reset would wipe the cluster's state and keep the OS, and the node promptly rebooted into a "no bootable disk" loop, since without any flags it wipes the entire system disk. After reinstalling from the ISO, I found the command I should have used:
talosctl reset --system-labels-to-wipe STATE --system-labels-to-wipe EPHEMERAL
With new secrets in place, the Proxmox console filled up with the node refusing my old client certificate, which had been signed by the leaked CA. It was the most reassuring wall of errors I've ever seen.
Rebuilding from git
A rebuild from nothing is also the best possible test of the claim at the heart of GitOps, that the repository fully describes the cluster. I bootstrapped Flux against the same repository, and it had nothing to commit:
component manifests are up to date
sync manifests are up to date
Everything else reconciled from git on its own. The only things I had to create by hand were the two secrets that can't live in the repository, the age key that decrypts it and the deploy key Flux uses to read it.
The deliberate breakage at this layer was to test ordering. Kubernetes can't create an object whose type doesn't exist yet, so I put MetalLB's Helm release in the same Flux Kustomization as an address pool, a custom resource type that MetalLB's own install defines, to see what would happen:
IPAddressPool/metallb-system/lan dry-run failed: no matches for kind
"IPAddressPool" in version "metallb.io/v1beta1"
Flux validates the whole batch before applying any of it, so the pool's failure meant the release that would have created the pool's type never got applied either, causing a deadlock. Splitting the two into Kustomizations chained with dependsOn fixed it, and a second experiment showed why the wait: true flag on that chain matters. Without it, the dependency counts as ready the moment it's applied, rather than once it's actually running, and the whole thing only worked because a 30-second retry happened to take longer than the Helm install. A slow image pull would break it.
Fighting my own DNS
My home network uses split-horizon DNS, where the same names resolve to LAN addresses at home and to mesh VPN addresses everywhere else. It's probably interesting enough to deserve its own article at some point, so I'll be brief. It also turned out to be the most fragile part of the whole setup. Right after bootstrap, kubectl started timing out while trying to reach a domain parking page. My router only overrides A records for local names, so the AAAA query fell through to the public zone, where my registrar's default wildcard CNAME answered it, and since a CNAME applies to every record type, the resolver followed it for the A record as well. Deleting the wildcard fixed it.
Later on, cert-manager's DNS challenge for my wildcard certificate sat stuck waiting for a record that was already public. The router intercepts every DNS query on port 53, even ones addressed to the authoritative nameservers, and it had cached a "no such record" answer for half an hour, which it had picked up from my own check a few minutes before the record existed. The fix was to change cert-manager to resolve over DNS-over-HTTPS, which the router can't intercept.
Killing the database
For the database, CloudNativePG runs Postgres as two instances and points a -rw Service at whichever one is currently the primary. I left a pod inserting a row every second through that Service, and deleted the primary:
08:01:28 7|10.244.0.38 INSERT 0 1
08:01:29 psql: error: ... FATAL: the database system is shutting down
08:01:30 36|10.244.0.41 INSERT 0 1
One insert failed, the replica was promoted, and the old primary rejoined as a replica fifteen seconds later. The jump from id 7 to 36 looked like lost rows at first, but Postgres logs sequence values 32 at a time ahead of use, so a promoted replica simply starts after the batch.
A replica copies mistakes as faithfully as everything else, though, so I also dropped a table and restored it from backups, which are a daily base backup plus a continuously archived write-ahead log in a Cloudflare R2 bucket. The restore, a new cluster replayed from the bucket up to two seconds before the drop, brought back every row, including three I'd written after the last base backup. I was surprised when checking the bucket right after the drop, as the log segment holding those three rows hadn't been uploaded yet. Postgres only ships a segment when it fills or after five minutes, and a restore at that moment would have lost them.
Moving the cron jobs in
The meal planner, chefcal, turned out to be the harder of the two jobs to move, because it keeps state. It stores the weeks it has already generated in a file, and it plans a week ahead, so a fresh container with no file would have re-rolled the week already on my calendar, and possibly already shopped for. It got a persistent volume, seeded with the desktop's copy before its first run. The next problem was retries. Kubernetes retries a failed Job by running the whole thing again, and the planner's generate step always adds one more week, so a retry after a failed push would have pushed the plan a week further ahead, permanently. I split it into two CronJobs, and only the push, which is safe to repeat, gets retries.
The calendar job, einkcal, had a more insidious problem. With all three of its calendars returning 401, it exited 0 and would have happily uploaded a blank calendar, so in Kubernetes' eyes nothing had gone wrong. It now fails when every source fails.
Letting CI loose on the node
The CI runners use a privileged Docker-in-Docker sidecar, and I expected a job's CPU to be billed to that container and capped by the pod's limits. To check, I ran three busy loops for two minutes. The node's CPU peaked at 3.16 cores, and the Docker container's at 0.40. A privileged container sees the node's root cgroup, and the Docker-in-Docker startup script assumes that cgroup is its own, so job containers were landing at the root of the node, outside the pod's memory limit and outside anything Kubernetes was accounting for. Bind-mounting the container's own cgroup over /sys/fs/cgroup before starting Docker fixed it:
args:
- |
self=$(cut -d: -f3 /proc/self/cgroup)
mount --bind "/sys/fs/cgroup$self" /sys/fs/cgroup
exec dockerd-entrypoint.sh --mtu=1450 ...
On the rerun, the node and the container both peaked at 3.21 cores. The cleanup produced one more unplanned failure. I moved what I took for PID 1 out of a stray cgroup from a debug pod, but on Talos, PID 1 in a pod that shares the host's process namespace isn't the host's init, and I had moved one of Talos' own services instead. A "resource busy" error gave it away, and I put it back where it belonged.
Why bother
Not everything I found was in the cluster. While adding a watchdog alert for sakuya to my monitoring VM, I discovered that its own alerting had been silently broken for four days, which had effectively killed alerts for my entire homelab.
This was my first serious attempt to go all-in on GitOps, and I was surprised to find that it's a very natural way of deploying once you wrap your head around secret management. The fact that I could keep kicking the cluster down, and felt like I had enough of a safety net to do so comfortably, is remarkable, and it really pushes me towards doing more repeatable infrastructure as code. The repository is here if you want to check it out.