On-call rotations with ntfy: push alerts with escalation, without PagerDuty
PagerDuty starts at 21 dollars per user per month and sends your alert data to the US. For a team of two to ten people it can be done with a self-hosted ntfy server, a rotation schedule in a text file and a fifty-line script: alerts go to whoever is on duty this week, escalate after ten minutes without acknowledgement to the backup, and break through the phone's silent mode. Including one-tap acknowledge.
Contents
All the guides on this site push alerts to ntfy. That works perfectly until there's more than one of you: then everyone gets everything, at 03:00, including whoever isn't on duty — and after two weeks everyone has muted the topic. What you need is small: a schedule (who's on duty when), routing (alert to that person), escalation (after X minutes without a response, to the next) and a way to say "I'm on it". PagerDuty does that for 21 dollars per person per month. This article does it with ntfy and cron.
Step 1: your own ntfy server
ntfy.sh is free to use, but for on-call you want your own server: no rate limits, no dependency on a third party, and your alerts (with hostnames and error messages) stay with you.
sudo install -d /opt/ntfy && cd /opt/ntfy
sudo tee docker-compose.yml >/dev/null <<'EOF'
services:
ntfy:
image: binwiederhier/ntfy:latest
command: serve
restart: unless-stopped
ports: ["127.0.0.1:8085:80"]
volumes:
- ./cache:/var/cache/ntfy
- ./etc:/etc/ntfy
environment:
NTFY_BASE_URL: https://ntfy.example.com
NTFY_CACHE_FILE: /var/cache/ntfy/cache.db
NTFY_AUTH_FILE: /var/cache/ntfy/user.db
NTFY_AUTH_DEFAULT_ACCESS: deny-all # nobody may read/write without an account
NTFY_BEHIND_PROXY: "true"
NTFY_ENABLE_LOGIN: "true"
NTFY_ATTACHMENT_CACHE_DIR: /var/cache/ntfy/attachments
NTFY_UPSTREAM_BASE_URL: https://ntfy.sh # needed for iOS push; Android works without
EOF
sudo docker compose up -d
# Reverse proxy (Caddy) — WebSocket support is default
printf 'ntfy.example.com {\n reverse_proxy 127.0.0.1:8085\n}\n' | sudo tee -a /etc/caddy/Caddyfile && sudo systemctl reload caddy
deny-all is deliberate: an open ntfy server is an open spam channel to your team's phones. Create accounts and permissions:
N="sudo docker compose exec -T ntfy ntfy"
$N user add --role=admin ops-admin
$N user add jeroen; $N user add marie; $N user add tom
$N user add --role=user monitoring # the servers that send alerts
$N access monitoring 'alerts-*' write-only
$N access jeroen 'alerts-*' read-write; $N access marie 'alerts-*' read-write; $N access tom 'alerts-*' read-write
$N token add monitoring # token for the scripts, instead of a password
On the phone: ntfy app (F-Droid, Play Store, App Store), server https://ntfy.example.com, log in with the personal account.
Step 2: the rotation schedule
One text file, on the host that routes the alerts. ISO weeks, or date ranges — whatever your team uses to make agreements.
sudo tee /etc/oncall/schedule.tsv >/dev/null <<'EOF'
# from to primary backup phone_primary (for the escalation text)
2026-09-14 2026-09-20 jeroen marie +32470000001
2026-09-21 2026-09-27 marie tom +32470000002
2026-09-28 2026-10-04 tom jeroen +32470000003
2026-10-05 2026-10-11 jeroen marie +32470000001
EOF
Who's on duty now:
oncall() { # oncall primary|backup|phone
awk -F'\t' -v d="$(date +%F)" -v f="$1" '!/^#/ && $1<=d && d<=$2 { print (f=="primary"?$3:(f=="backup"?$4:$5)); exit }' /etc/oncall/schedule.tsv
}
echo "primary: $(oncall primary), backup: $(oncall backup)"
One topic per person: alerts-jeroen, alerts-marie, alerts-tom. Everyone subscribes to their own topic and to alerts-all (informational, muted). The schedule determines which personal topic the critical alerts go to.
Step 3: the router — one entry point for all scripts
Instead of every monitoring script pushing to a topic itself, they send to one local script that routes, deduplicates and schedules escalation.
sudo tee /usr/local/bin/alert >/dev/null <<'EOF'
#!/usr/bin/env bash
# alert <severity: info|warn|crit> <title> [body via stdin]
# info → alerts-all (silent). warn → alerts-all + primary (normal prio). crit → primary (urgent, breaks silent mode) + escalation.
set -u
SEV=$1; TITLE=$2; BODY=$(cat); NTFY=https://ntfy.example.com; TOK=$(cat /etc/oncall/token)
S=/etc/oncall/schedule.tsv; D=$(date +%F)
PRIM=$(awk -F'\t' -v d="$D" '!/^#/ && $1<=d && d<=$2 {print $3; exit}' "$S")
BACK=$(awk -F'\t' -v d="$D" '!/^#/ && $1<=d && d<=$2 {print $4; exit}' "$S")
ID=$(printf '%s' "$TITLE" | sha1sum | cut -c1-12) # dedup key per title
ST=/var/lib/oncall; mkdir -p "$ST"
# Dedup: don't push the same title again within 30 min (except crit escalation)
if [ -f "$ST/$ID.sent" ] && [ "$(find "$ST/$ID.sent" -mmin -30)" ]; then exit 0; fi
touch "$ST/$ID.sent"
push() { # push <topic> <prio> <tags> [extra headers...]
local topic=$1 prio=$2 tags=$3; shift 3
curl -s -u ":$TOK" -H "Title: $TITLE" -H "Priority: $prio" -H "Tags: $tags" "$@" --data-binary "$BODY" "$NTFY/$topic" >/dev/null
}
case $SEV in
info) push alerts-all min information_source ;;
warn) push alerts-all default warning; push "alerts-$PRIM" default warning ;;
crit)
# Actions header: one tap = acknowledge (POST to ack topic) — see step 4
push "alerts-$PRIM" urgent rotating_light -H "Actions: http, ACK, $NTFY/ack, method=POST, body=$ID $PRIM, clear=true"
push alerts-all default rotating_light
# Schedule escalation: after 10 min without ack → backup; after 20 min → both + phone in the text
echo "$ID $(date +%s) $PRIM $BACK $TITLE" >> "$ST/pending.tsv" ;;
esac
EOF
sudo chmod 0755 /usr/local/bin/alert
sudo install -d -m 0750 /etc/oncall && echo "tk_xxxxxxxxxxxxxxxx" | sudo tee /etc/oncall/token >/dev/null && sudo chmod 0640 /etc/oncall/token
# Usage in every monitoring script:
echo "disk /var 91% on web-01" | alert crit "disk almost full web-01"
Priority: urgent is what makes the difference at 03:00: the ntfy app lets an urgent message through the "do not disturb" mode (Android: if the app has that permission; iOS: if "critical alerts" is allowed). min appears silently in the list.
Step 4: acknowledge and escalation
The Actions header puts an "ACK" button under the notification. One tap does a POST to the ack topic with the alert id. A second cron reads those acks and escalates whatever isn't acknowledged.
sudo tee /usr/local/sbin/oncall-escalate.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
# Every minute: read acks from the ack topic, escalate unacknowledged crit alerts.
NTFY=https://ntfy.example.com; TOK=$(cat /etc/oncall/token); ST=/var/lib/oncall
[ -s "$ST/pending.tsv" ] || exit 0
# Fetch acks from the last 24 h (JSON, poll)
curl -s -u ":$TOK" "$NTFY/ack/json?poll=1&since=24h" | jq -r '.message // empty' | awk '{print $1}' | sort -u > "$ST/acked.txt"
now=$(date +%s); : > "$ST/pending.new"
while IFS=$'\t' read -r id t0 prim back title; do
if grep -qx "$id" "$ST/acked.txt"; then continue; fi # acknowledged → done
age=$(( (now - t0) / 60 ))
if [ "$age" -ge 20 ] && [ ! -f "$ST/$id.esc2" ]; then
phone=$(awk -F'\t' -v d="$(date +%F)" '!/^#/ && $1<=d && d<=$2 {print $5; exit}' /etc/oncall/schedule.tsv)
for p in "$prim" "$back"; do
curl -s -u ":$TOK" -H "Title: ESCALATION 20m: $title" -H "Priority: urgent" -H "Tags: rotating_light,sos" \
-d "Not acknowledged after 20 min. Primary $prim ($phone) and backup $back: someone call." "$NTFY/alerts-$p" >/dev/null
done; touch "$ST/$id.esc2"
elif [ "$age" -ge 10 ] && [ ! -f "$ST/$id.esc1" ]; then
curl -s -u ":$TOK" -H "Title: ESCALATION 10m: $title" -H "Priority: urgent" -H "Tags: rotating_light" \
-H "Actions: http, ACK, $NTFY/ack, method=POST, body=$id $back, clear=true" \
-d "Primary $prim has not acknowledged. You are backup." "$NTFY/alerts-$back" >/dev/null
touch "$ST/$id.esc1"
fi
[ "$age" -lt 120 ] && printf '%s\t%s\t%s\t%s\t%s\n' "$id" "$t0" "$prim" "$back" "$title" >> "$ST/pending.new" # give up after 2 h
done < "$ST/pending.tsv"
mv "$ST/pending.new" "$ST/pending.tsv"
EOF
sudo chmod 0755 /usr/local/sbin/oncall-escalate.sh
echo '* * * * * root /usr/local/sbin/oncall-escalate.sh' | sudo tee /etc/cron.d/oncall-escalate
# The ack topic: everyone may write (the button), only the router reads
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access jeroen ack write-only
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access marie ack write-only
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access tom ack write-only
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access monitoring ack read-write
The chain: crit alert → primary (urgent, ACK button) → after 10 min without ack: backup (urgent, ACK button) → after 20 min: both, with phone number in the text so whoever reads it can call the other. Whoever taps ACK stops the escalation; the ack is in the ack topic with a name — that's your log of who responded when.
Step 5: test it, every week
# Monday 09:00: a test alert to the primary, with ACK. Not acknowledged after 10 min = the schedule or the app is wrong.
echo "0 9 * * 1 root echo 'Weekly on-call test — tap ACK' | /usr/local/bin/alert crit \"on-call test $(date +\%F)\"" | sudo tee /etc/cron.d/oncall-test
And once a quarter a real exercise: someone fills a disk on a test host at 22:00, and you measure how long it takes until the ack. That number (MTTA, mean time to acknowledge) is what a customer with an SLA wants to see.
Pitfalls
- iOS push without upstream. On iPhones ntfy pushes only arrive via Apple's push service, and for that your server must point
NTFY_UPSTREAM_BASE_URLto ntfy.sh. Only the message id goes to ntfy.sh, not the content — but it's a dependency, and you should know it. Android works fully standalone. - Battery optimisation. Android kills background connections; the ntfy app must be on the exception list (the app asks for it). Without that, urgent alerts arrive minutes late. That's one of the things the weekly test catches.
- A schedule that runs out. If
schedule.tsvhas no line for today,PRIMis empty and the alert goes toalerts-— to nobody. Add a fallback to the script (PRIM=${PRIM:-jeroen}) and a daily check that warns when the schedule ends within 14 days. - Dedup hiding a second incident. "disk almost full web-01" at 03:00 and again at 03:20 (after a cleanup that didn't help) is the same title within 30 minutes → suppressed. Put the measured value in the body, not the title, and shorten the dedup window for crit.
- One person in the schedule. A rotation with one name isn't a rotation but a burnout. At least two, preferably three; and this week's backup is next week's primary.
What you still don't have
- Overrides. "Marie is ill, Tom takes over until Thursday" is an edit in
schedule.tsvvia ssh — at 06:30, from a phone. A UI is missing. - Per-customer routing. An MSP wants alerts from customer A to go to customer A's team, and customer A to have a topic of their own. That's a second dimension in the schedule and the router.
- Reporting. How many alerts per week, how many at night, average time to ack, who was woken most often: scattered across the
acktopic andpending.tsv.
How monsys does it
monsys uses the same building block — a self-hosted ntfy server, no Twilio, no PagerDuty — but the on-call rotation lives per group or tenant in the dashboard: shifts with start, end and contact, overrides without ssh, and per alert the person on duty is looked up automatically and added to the push. Acknowledge and resolve happen in the dashboard or via the push action, with name and time in the audit log; MTTA and MTTR per group are in the operations metrics and the Trust Score. For MSPs: alerts from customer A go to the on-call of customer A's group, and the customer can follow along on their own ntfy topic.
FAQ
Is ntfy reliable enough for on-call?
For a team managing its own server and doing the weekly test: yes. The protocol is simple (HTTP + WebSocket), the app is stable and the urgent priority breaks through silent modes. The weak spot isn't ntfy but the phone: battery optimisation and forgotten permissions.
What about SMS or calling as a fallback?
Deliberately not in this setup: SMS gateways (Twilio and co) are US services with an SLA no better than ntfy's, and calling requires a telephony provider. The escalation after 20 minutes therefore puts the phone number in the text, so that a human calls. For an organisation under NIS2 that's also the more defensible choice.
Can I combine this with Alertmanager or Uptime Kuma?
Yes. Both can send to a webhook or directly to ntfy; have them send to a local endpoint that calls /usr/local/bin/alert, or configure them directly on alerts-all and use the router only for crit. The schedule and the escalation then stay in one place.
Written by the monsys team — sysadmins who do this every day.
Done it by hand? Let monsys keep it running.
Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.