The boring work that prevents 80 % of incidents. These guides cover patching without surprises, backups that actually exist, certificates that do not expire silently and a status page your customers can read themselves.
A status page on the same server as the service it describes is down when you need it. This is the setup that works: Uptime Kuma on a separate, cheap VPS, checks from outside on HTTP, TLS and TCP, a public page on status.yourdomain.com with its own certificate, incident updates you post from your phone, and the DNS and cache details that decide whether the page stays reachable when the rest is on fire.
A kernel update is the only routine update that requires a reboot and that can leave your server hanging. This is the checklist we use ourselves: pre-checks (/boot, DKMS, Secure Boot), snapshot, installation, a GRUB fallback that automatically picks the old kernel after one failed boot, and the verification after the reboot. In a fifteen-minute window.
Let's Encrypt renews itself — until the day the hook fails, DNS validation breaks or someone removed the cron. This script finds every certificate on every port of every host (internal ones too, mail too, the reverse proxy nobody remembers), computes the remaining days and pushes an alert at 14 and 7 days. No Nagios, no external service.
The default install of unattended-upgrades patches security updates but leaves services running on old libraries and the new kernel sitting on disk. This is the configuration that fixes that for Ubuntu 24.04 and Debian 12: origins, blacklist, needrestart, reboot policy, mail — plus the cron script that reports which hosts have been waiting days for a reboot.
A backup job that returns exit 0 proves the job ran — not that there's anything usable in the archive. Three checks you automate: is the newest backup fresh enough (mtime check), is the archive readable (restic/borg check) and can you actually get something back out (monthly restore test with a log line). Plus the pitfall of rsync preserving mtimes.