My Pi 4 runs a handful of always-on services: a small web dashboard, some automation glue, a media helper or two. (If you’re buying today, get a Raspberry Pi 5 instead — same idea, much more headroom.) Every few weeks it would go completely dark. No SSH, no services, and nothing short of pulling the plug would bring it back. I did that twice, which is another way of saying I fixed nothing at all.
It looked like a crash rather than a network problem. The hungriest thing on the box was a local AI web UI using around 800 MB of the 4 GB, and I had already found crash dumps from a couple of system services. Most likely the machine was running out of memory and wedging itself solid, with nobody around to kick it.
So I set it up to recover on its own. Two parts, neither of them took long.
The hardware watchdog
The chip in the Pi has a watchdog timer built in (bcm2835_wdt). The kernel is supposed to reset it every few seconds; if it stops doing that, the board reboots itself. No human required. systemd can drive it with two lines in /etc/systemd/system.conf:
[Manager]
RuntimeWatchdogSec=15
The watchdog module needs to load at boot. On Raspberry Pi OS and Ubuntu:
echo "bcm2835_wdt" | sudo tee /etc/modules-load.d/bcm2835_wdt.conf
Reboot, then check it actually took:
systemctl show -p RuntimeWatchdogUSec
# RuntimeWatchdogUSec=15s
If that comes back as 0 or infinity, it isn’t active. Check dmesg | grep -i watchdog for clues. On my Pi it worked first time. Now, fifteen seconds after a hard hang, the board resets itself instead of sitting there dead until I happen to notice.
There are limits. A watchdog can’t fix a dead power supply, a corrupted SD card, or hardware that has genuinely given up. It handles hangs, which is what I had.
earlyoom, so it never gets that far
A watchdog reboot is a blunt tool. Better if the box kills the runaway process before memory pressure wedges everything in the first place. That is what earlyoom does: it watches memory and kills the largest offender before the kernel’s own out-of-memory killer gets involved, which on a small machine tends to arrive too late.
sudo apt install earlyoom
sudo systemctl enable --now earlyoom
The defaults are sensible. It starts killing when free memory drops below 10%, picking the biggest process first. You can tune it in /etc/default/earlyoom, but I left it alone.
Reading the logs afterwards
The annoying part of the earlier hangs was the lack of evidence. Whatever happened was gone after the power cycle, because the useful bits lived in the system journal and my unprivileged service account couldn’t read it. One command fixed that:
sudo usermod -aG adm,systemd-journal myserviceuser
If it hangs again, I can read journalctl over SSH afterwards and see what the watchdog rebooted from, instead of guessing.
The whole thing as a script
I wrapped it up so it’s reproducible if I ever rebuild the SD card:
#!/bin/bash
set -euo pipefail
# Hardware watchdog via systemd (15s timeout)
grep -q '^RuntimeWatchdogSec' /etc/systemd/system.conf \
|| echo 'RuntimeWatchdogSec=15' >> /etc/systemd/system.conf
echo "bcm2835_wdt" > /etc/modules-load.d/bcm2835_wdt.conf
# earlyoom for memory pressure
apt-get install -y earlyoom
systemctl enable --now earlyoom
# Let the service user read system logs
usermod -aG adm,systemd-journal myserviceuser
echo "Done. Reboot, then verify with: systemctl show -p RuntimeWatchdogUSec"
About ten minutes’ work and one reboot. The Pi hasn’t needed a power cycle since. If it ever hangs hard again, the worst case is a fifteen-second self-inflicted reboot rather than a dead box waiting for me to walk over to it.
If I were setting up a new Pi today I’d do this on day one, not after the third hang. You can’t properly test a watchdog until something actually wedges the machine, and by then you’re already annoyed.
As an Amazon Associate, Headless Diaries earns from qualifying purchases made through links on this page.