Monitoring does not start with a tool but with the question of what you want to know before a customer notices. These guides set up measurement with what is already on the machine: /proc, systemd, df, docker stats, cron. You learn what a good alert is and why "CPU above 90 %" usually is not one.
Load 8 is quiet on a 16-core and a disaster on a 2-core. Load also counts processes waiting on a disk, not just CPU. And the 1-minute value is almost always noise. What the three numbers really mean, how to normalise them per core, how to separate CPU pressure from I/O pressure with /proc/pressure, and which threshold to put in an alert.
A service that crashes is back in seconds with Restart=on-failure. A service that hangs, one stuck in a restart loop, or one that never started after a reboot, systemd doesn't see by itself. These are the drop-in settings, the OnFailure handler that pushes to your phone, the WatchdogSec trick for hanging processes and the cron script that closes the gaps.
Everything you can see yourself with docker stats, docker events and a healthcheck in your compose file — plus the cron script that reports a container restarting every 30 seconds within five minutes. Including the logging setting that stops containers from filling your disk.
Four measurements that make 90% of incidents visible in advance, using only what ships with Ubuntu 24.04 or Debian 12: sysstat, journalctl, df and a thirty-line cron script that pushes to your phone. Plus why "CPU above 90%" is a bad alert.