Configuring the Cluster: TLS, Secrets, and Learning to Not Lose a Node

Standing up cert-manager with a Cloudflare DNS-01 solver, the three-stage evolution from plaintext secrets to Sealed Secrets, and the day I stopped waiting five minutes for a failed node to give up its pods.

Homelab Kubernetes Series: 1. Intro · 2. Installation · 3. Configuration · 4. Networking · 5. Storage · 6. Workloads · 7. ArgoCD · 8. GitOps

Recap

Last time I had three MicroK8s nodes joined into a cluster. That’s a cluster you can kubectl get nodes against, but it’s not yet a cluster you’d trust with anything real. This post covers the configuration work that closed that gap: certificates, secrets, and a couple of lessons I learned the hard way about how the cluster behaves when a node actually dies.

TLS: real certificates for an internal-only cluster

Every host I run is a subdomain of joeyaxtell.com, but none of them are reachable from the public internet — they all resolve to addresses inside my LAN. That rules out the usual HTTP-01 ACME challenge, which needs Let’s Encrypt to reach your server directly. The fix is a DNS-01 challenge through Cloudflare, which proves domain ownership by writing a TXT record instead of serving an HTTP response:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-cloudflare
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: YOUR_EMAIL@EXAMPLE.COM
    solvers:
      - dns01:
          cloudflare:
            apiTokenSecretRef:
              name: cloudflare-api-token
              key: api-token

That’s the production ACME endpoint, not staging. I went straight for real certificates on the first ClusterIssuer I ever wrote, which was a little bold in hindsight — Let’s Encrypt’s production rate limits aren’t generous if you get the config wrong and end up retrying in a loop. It worked out, but staging first is the safer habit.

I proved it worked the same way I proved everything else in this series: a disposable nginx pod at test.joeyaxtell.com, watched until a real certificate showed up in its Secret, then deleted.

Secrets: three stages, and only the first one is embarrassing

My secrets story is basically a timeline of learning why each previous approach doesn’t scale.

Stage 1 — plaintext templates committed to git. The Cloudflare API token secret started life as a literal kind: Secret manifest with api-token: YOUR_CLOUDFLARE_API_TOKEN committed to the repo, meant to be hand-edited locally before applying. Never a real credential in git, but also not a pattern I’d want to repeat as the number of secrets grew.

Stage 2 — imperative, out-of-band kubectl create secret. For a while, secrets simply didn’t exist in the repo at all, just a comment documenting the kubectl create secret generic ... --from-literal=... command I’d run once by hand and never again. It works, but it means the repo doesn’t actually describe the cluster. There’s a whole category of state that only exists in my shell history.

Stage 3 — Sealed Secrets. This is where I landed, and where new secrets go today. The Bitnami Sealed Secrets controller lets you encrypt a Secret client-side with its public key, commit the encrypted SealedSecret to git, and only the controller running in-cluster can decrypt it back into a real Secret. The workflow, including the PowerShell-specific piping since I do this from Windows:

1
2
kubeseal --fetch-cert --controller-namespace default --controller-name sealed-secrets-controller > cert.pem
Get-Content secret.yaml | kubeseal --cert cert.pem -o yaml > sealedsecret.yaml

That finally makes both a secret’s existence and its encrypted value visible in git history, without the value ever being recoverable by anyone who doesn’t hold the cluster’s private key. It’s not applied everywhere yet — that migration is still going workload by workload — but it’s the pattern I reach for now.

Update tracking: an opt-in watcher

Rather than a blanket “check everything for updates” policy, I run Diun with Kubernetes provider support enabled, watching every six hours. Workloads opt in individually by adding one annotation to their pod template:

1
2
annotations:
  diun.enable: "true"

That opt-in model matters. I want to know when something has a new image available, but I don’t want that noise on things I’ve deliberately pinned. When I locked one workload to a specific version and set imagePullPolicy: IfNotPresent so it would stop drifting, the next thing I did was pull the diun.enable annotation back off it. No point getting paged about updates I’ve already decided not to take.

Configuring for node failure

The most useful configuration change I made didn’t happen until months in, after I actually lost a node and watched what happened: nothing, for a long time. Kubernetes’ default tolerance for an unreachable node is generous — pods on a node that goes NotReady or Unreachable aren’t rescheduled for five minutes by default. On a three-node homelab cluster, five minutes of a chunk of your services being down because one box hiccuped is a bad trade.

The fix, applied across every deployment in one pass once I understood the knob:

1
2
3
4
5
6
7
8
9
tolerations:
  - key: "node.kubernetes.io/not-ready"
    operator: "Exists"
    effect: "NoExecute"
    tolerationSeconds: 30
  - key: "node.kubernetes.io/unreachable"
    operator: "Exists"
    effect: "NoExecute"
    tolerationSeconds: 30

That cuts the eviction wait from five minutes to thirty seconds. I paired it with strategy: Recreate on the same deployments, since my persistent volumes are ReadWriteOnce and a rolling update trying to attach the same volume to a second pod before the first releases it just hangs. Recreate tears the old pod down before standing the new one up. It’s slower for a routine deploy, but it’s the only strategy that actually works with single-writer storage.

Up next

With TLS, secrets, and failure handling in place, the next post covers how traffic actually gets from my LAN to a pod — Traefik, MetalLB, and DNS.


◀ Previous: 2. Installation | Next ▶: 4. Networking

comments powered by Disqus
Built with Hugo
Theme Stack designed by Jimmy