Monitoring and auto-restarting systemd services: Restart=, OnFailure=, watchdogs — and when that's not enough
A service that crashes is back in seconds with Restart=on-failure. A service that hangs, one stuck in a restart loop, or one that never started after a reboot, systemd doesn't see by itself. These are the drop-in settings, the OnFailure handler that pushes to your phone, the WatchdogSec trick for hanging processes and the cron script that closes the gaps.
Contents
- Step 1: know what isn't running right now
- Step 2: set Restart= properly (with a drop-in, not in the unit itself)
- Step 3: OnFailure= — an alert to your phone, built in
- Step 4: hanging services — the watchdog
- Step 5: the cron script that closes the gaps
- Pitfalls
- What you still don't have
- How monsys does it
- FAQ
Most "the service is down" tickets have a boring cause: the process crashed at 03:12 and nobody had set Restart=; or it was running but no longer answering; or it had simply been off since last week's reboot because nobody ran enable. systemd can catch all three — but not with the defaults, and not without you thinking about what "healthy" means for that service.
Step 1: know what isn't running right now
# Units in failed state (crash, exit ≠ 0, start limit reached)
systemctl --failed
# Services that are enabled but NOT running (e.g. didn't come up after a reboot)
systemctl list-unit-files --type=service --state=enabled --no-legend | awk '$1 !~ /@\.service$/ {print $1}' \
| while read -r u; do
# oneshots and units whose Condition isn't met (cloud-init on bare metal) are supposed to be inactive
[ "$(systemctl show "$u" -p Type --value)" = oneshot ] && continue
[ "$(systemctl show "$u" -p ConditionResult --value)" = no ] && continue
# finished cleanly (Result=success, e.g. cloud-init-main) is not a failure; a crash gives exit-code/signal
[ "$(systemctl show "$u" -p Result --value)" = success ] && ! systemctl is-active --quiet "$u" && continue
systemctl is-active --quiet "$u" || echo "enabled-but-inactive: $u"
done
# How often has a service restarted since boot?
systemctl show nginx -p NRestarts
# Why did it fail?
journalctl -u nginx -b -p warning --no-pager | tail -20
That second list is the surprising one: units that are "enabled" but inactive. Oneshots, template units and units whose Condition… isn't met (cloud-init on a physical machine) belong there; a web server doesn't — which is why the loop filters those three out.
Step 2: set Restart= properly (with a drop-in, not in the unit itself)
Never edit /lib/systemd/system/*.service — a package update overwrites it. A drop-in in /etc/systemd/system/<unit>.service.d/ survives everything.
sudo systemctl edit nginx
[Unit]
# Start limit: default 5 attempts in 10 s, then failed and NOTHING more. Too tight for
# a service waiting on a database. Wider window.
StartLimitIntervalSec=300
StartLimitBurst=10
[Service]
# on-failure: on crash or exit ≠ 0. always: also on clean exit 0 (for daemons that think they're "done").
Restart=on-failure
RestartSec=5s
Mind the section: StartLimitIntervalSec and StartLimitBurst belong in [Unit] (since systemd 230); older examples put them in [Service], which warns and is ignored on newer versions.
Then check what systemd actually uses:
sudo systemctl daemon-reload
systemctl show nginx -p Restart -p RestartUSec -p StartLimitBurst -p StartLimitIntervalUSec
For services with dependencies (an app that needs a database): After=postgresql.service in [Unit], and a connection retry in the app itself. Requires= sounds logical but makes your app stop along with the database when you restart it — usually not what you want.
Step 3: OnFailure= — an alert to your phone, built in
systemd can start another unit when a service fails. One generic notify unit for all services:
sudo tee /etc/systemd/system/notify-failed@.service >/dev/null <<'EOF'
[Unit]
Description=Push failure of %i to ntfy
[Service]
Type=oneshot
# %i = the name of the failed unit (via OnFailure=notify-failed@%n.service)
ExecStart=/bin/sh -c 'journalctl -u %i -n 15 --no-pager -o cat | curl -s -H "Title: FAILED %i on $(hostname -s)" -H "Priority: high" --data-binary @- https://ntfy.example.be/services >/dev/null'
EOF
# Attach to a service via drop-in
sudo systemctl edit nginx
[Unit]
OnFailure=notify-failed@%n.service
%n is the full unit name (nginx.service); it becomes the instance of the notify unit. Test with a throwaway unit:
sudo systemd-run --unit=crashtest -p OnFailure=notify-failed@crashtest.service /bin/false
# → push with the last 15 journal lines of crashtest
Note: OnFailure only fires when the unit really fails — i.e. after hitting the start limit, not on every individual crash that Restart= absorbs. That's exactly right: you don't want a push per restart, you do want one when restarting no longer helps.
Step 4: hanging services — the watchdog
Restart= reacts to a process that stops. A process stuck in a deadlock, all threads busy, doesn't stop. For services supporting systemd's sd_notify (nginx doesn't; many Go/Rust/Python daemons do, and anything talking via systemd-notify), there's WatchdogSec=:
[Service]
Type=notify
WatchdogSec=30s
# If the process doesn't send WATCHDOG=1 every 30 s, it gets SIGABRT and systemd restarts it
Restart=on-watchdog
If your service doesn't support that, build the watchdog outside the service: a timer that does a real request and restarts on failure.
sudo tee /etc/systemd/system/healthcheck-web.service >/dev/null <<'EOF'
[Unit]
Description=HTTP healthcheck for the web app; restart on failure
[Service]
Type=oneshot
ExecStart=/bin/sh -c 'curl -fsS -m 5 -o /dev/null http://127.0.0.1:8080/healthz || { echo "healthz failed, restarting"; systemctl restart webapp.service; }'
EOF
sudo tee /etc/systemd/system/healthcheck-web.timer >/dev/null <<'EOF'
[Unit]
Description=Run web healthcheck every minute
[Timer]
OnBootSec=2min
OnUnitActiveSec=1min
[Install]
WantedBy=timers.target
EOF
sudo systemctl daemon-reload && sudo systemctl enable --now healthcheck-web.timer
Two rules for such an external watchdog: the healthcheck must do something real (a query, not just "200 OK"), and the restart must be logged and counted — otherwise you mask a memory leak as "the occasional restart".
Step 5: the cron script that closes the gaps
What systemd itself doesn't report: enabled-but-inactive units, restart loops that stay just under the start limit, and services that didn't come back after a reboot.
sudo tee /usr/local/sbin/service-check.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
NTFY="https://ntfy.example.be/services"
HOST=$(hostname -s)
STATE=/var/lib/service-check.state
touch "$STATE"; MSG=()
# 1. Failed units
f=$(systemctl --failed --no-legend | awk '{print $1}' | paste -sd, -)
[ -n "$f" ] && MSG+=("failed: $f")
# 2. Enabled but not active (templates, oneshots and unmet-Condition units excluded)
while read -r u; do
[ "$(systemctl show "$u" -p Type --value)" = oneshot ] && continue
[ "$(systemctl show "$u" -p ConditionResult --value)" = no ] && continue
systemctl is-active --quiet "$u" && continue
# Finished cleanly (Result=success) = boot job like cloud-init-main, not a failure.
# A daemon that unexpectedly stops with exit 0 you catch with Restart=always (step 2).
[ "$(systemctl show "$u" -p Result --value)" = success ] && continue
MSG+=("enabled-but-inactive: $u ($(systemctl show "$u" -p Result --value))")
done < <(systemctl list-unit-files --type=service --state=enabled --no-legend | awk '$1 !~ /@\.service$/ {print $1}')
# 3. Restart loops: NRestarts up by ≥ 3 since the previous check (10 min)
while read -r u; do
n=$(systemctl show "$u" -p NRestarts --value); [ "${n:-0}" -eq 0 ] && continue
prev=$(grep "^$u " "$STATE" | awk '{print $2}'); prev=${prev:-0}
[ $((n - prev)) -ge 3 ] && MSG+=("restart-loop: $u ($((n - prev)) restarts in 10 min, total $n)")
sed -i "/^$u /d" "$STATE"; echo "$u $n" >> "$STATE"
done < <(systemctl list-units --type=service --state=running --no-legend | awk '{print $1}')
[ ${#MSG[@]} -gt 0 ] && printf '%s\n' "${MSG[@]}" | curl -s -H "Title: services@$HOST" -H "Priority: high" --data-binary @- "$NTFY" >/dev/null
EOF
sudo chmod 0755 /usr/local/sbin/service-check.sh
echo '*/10 * * * * root /usr/local/sbin/service-check.sh' | sudo tee /etc/cron.d/service-check
Pitfalls
Restart=alwayson a oneshot. A backup script withRestart=alwaysruns again forever.Restart=is for daemons; for jobs you use timers.- A failing
ExecStartPrecounts towards the start limit. Amkdirfailing because a mount is missing burns your ten attempts in seconds. Prefix the command with-(ExecStartPre=-/bin/mkdir …) if failure is acceptable. systemctl restartin a healthcheck without rate limiting. If the database is gone, you restart an app every minute that won't work anyway, and fill the logs. Count the restarts and stop after three: then it's an alert, not a fix.KillMode=and children. A service spawning worker processes (gunicorn, php-fpm): withRestart=, old workers sometimes linger ifKillMode=process. The default (control-group) is almost always right.- Forgetting
daemon-reload. A drop-in withoutsystemctl daemon-reloadis a text file.systemctl showtells you what's really active.
What you still don't have
- History.
NRestartsresets on every boot. "How often has this service restarted in the last 30 days" needs a log you keep yourself. - Fleet view. Thirty servers, thirty
--failedlists, thirty ntfy topics. - Context. The service that crashed at 03:12 and the
aptupgrade at 03:10 are in two logs; the correlation is manual.
How monsys does it
The monsys agent reads the systemd state of every host: failed units, enabled-but-inactive, NRestarts, restart events with timestamps, and the exit status. Restart loops and down services become deduplicated alerts with the journal lines and the host metrics of that moment next to them; the SLA engine computes an availability percentage per service over 5-minute buckets, so "99.7% last month" is a fact rather than an estimate. Restarting remains an action a human approves — via a signed Emergency Action Token, with the output of systemctl restart in the audit log.
FAQ
Does systemd restart a service automatically after a crash?
Only if Restart= is set (on-failure or always); the default is no. And even then systemd stops after StartLimitBurst attempts within StartLimitIntervalSec and marks the unit failed.
How do I see why a service failed?
systemctl status <unit> shows the last lines; journalctl -u <unit> -b -p warning the full context since boot. systemctl show <unit> -p Result -p ExecMainStatus gives the exit code and the reason (exit-code, signal, start-limit-hit, watchdog).
Does the watchdog work with every service?
Only with services that use Type=notify and periodically send WATCHDOG=1 via sd_notify. For other services you build an external healthcheck with a timer, as in step 4.
Written by the monsys team — sysadmins who do this every day.
Done it by hand? Let monsys keep it running.
Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.