Greg Ross

A homelab that recovers on its own

Created

My homelab is small: one Proxmox node for compute, one ZFS NAS for everything that has to last, a firewall between the networks, and the Mac where I work. Its services run as containers (guests) on the Proxmox node. This page is about one goal: when the node restarts, everything comes back without me touching anything, and I can see at a glance whether it is ok.

A note on scope first. What is proved today is a clean, planned reboot. A power cut and a hang are handled by mechanisms that are configured but not yet tested. I say which is which below.

Why I built it this way

The compute node had stopped a few times, and each time getting everything back needed me. That is the opposite of what a homelab is for. I wanted one where nothing stalls on me: if something stops, it comes back by itself, and if it can’t, I find out quickly and from evidence, not by noticing something is missing.

One rule shapes the whole design: all persistent state lives on the NAS; compute is disposable. A guest is replaceable and redeployed from git. That makes recovery a question of starting things in the right order with the right things ready, not of repairing state.

How it comes back after a reboot

The test: a clean, planned reboot against a written definition of done (every guest, every app, the assistant answering, scheduled tasks running). The node was back in about a minute and everything was green in about 11 minutes, inside a 15-minute target, with nobody touching anything.

The node is also configured to restart itself rather than stay stuck after a hang. That part is configured, not yet proved: a controlled hang test and a power-cut test are still to run.

How it watches itself

The monitoring, central logs, status page and a daily summary run on the NAS, not on the node they watch. An earlier stop went unnoticed for hours because the alerter ran on the machine that had stopped. A watcher has to live somewhere else.

It checks a fixed list of sources. Each one is judged fresh, stale or unknown, and a missing source counts as unknown, never as ok. Silence is not health.

A few real alerts, for things that are down or broken, reach me as a message. Everything softer, like a disk filling up slowly, goes only into the daily summary, so an alert still means something.

There is also a short list of safe fixes the system could apply itself, with locks, cooldowns and rate limits. Anything risky (the firewall, the assistant’s guest, storage) is never fixed automatically. Today the fixes run only by hand: they are proven, but held back until I trust them.

The shape of it

In words, the diagram is:

Mac (where I work) → works on → Proxmox node (guests) ← storage and state ← NAS (state, git, logs, monitoring) → alerts → me

The NAS sits in the middle on purpose: it holds the state, and it watches the node.

What I learned

  1. Put the watcher on a different box than the thing it watches. Otherwise the failure you most need to hear about is the one that silences the alarm.
  2. Refuse to start rather than start wrong. A service that is down is obvious; one that is up against the wrong storage is not.
  3. “Unknown” is not “ok”. A source that stops reporting has to show up as a problem.
  4. Every safety setting has a cost. An aggressive self-restart rule can turn small storage hiccups into reboots, so each one is weighed against what it would trigger on.
  5. Prove recovery with a real test against a written “done”, not a belief, and say which scenarios are still untested.
  6. Review every command an agent runs on a host before it runs, and keep checks free of destructive steps.
  7. Still open, said honestly: the cause of the stops is not found yet, and the harder recovery tests are still to run. This page will change when they are.