MSP & multi-tenantintermediate5 min read

On-call rotations with ntfy: push alerts with escalation, without PagerDuty

PagerDuty starts at 21 dollars per user per month and sends your alert data to the US. For a team of two to ten people it can be done with a self-hosted ntfy server, a rotation schedule in a text file and a fifty-line script: alerts go to whoever is on duty this week, escalate after ten minutes without acknowledgement to the backup, and break through the phone's silent mode. Including one-tap acknowledge.

Contents
  1. Step 1: your own ntfy server
  2. Step 2: the rotation schedule
  3. Step 3: the router — one entry point for all scripts
  4. Step 4: acknowledge and escalation
  5. Step 5: test it, every week
  6. Pitfalls
  7. What you still don't have
  8. How monsys does it
  9. FAQ

All the guides on this site push alerts to ntfy. That works perfectly until there's more than one of you: then everyone gets everything, at 03:00, including whoever isn't on duty — and after two weeks everyone has muted the topic. What you need is small: a schedule (who's on duty when), routing (alert to that person), escalation (after X minutes without a response, to the next) and a way to say "I'm on it". PagerDuty does that for 21 dollars per person per month. This article does it with ntfy and cron.

Step 1: your own ntfy server

ntfy.sh is free to use, but for on-call you want your own server: no rate limits, no dependency on a third party, and your alerts (with hostnames and error messages) stay with you.

sudo install -d /opt/ntfy && cd /opt/ntfy
sudo tee docker-compose.yml >/dev/null <<'EOF'
services:
  ntfy:
    image: binwiederhier/ntfy:latest
    command: serve
    restart: unless-stopped
    ports: ["127.0.0.1:8085:80"]
    volumes:
      - ./cache:/var/cache/ntfy
      - ./etc:/etc/ntfy
    environment:
      NTFY_BASE_URL: https://ntfy.example.com
      NTFY_CACHE_FILE: /var/cache/ntfy/cache.db
      NTFY_AUTH_FILE: /var/cache/ntfy/user.db
      NTFY_AUTH_DEFAULT_ACCESS: deny-all      # nobody may read/write without an account
      NTFY_BEHIND_PROXY: "true"
      NTFY_ENABLE_LOGIN: "true"
      NTFY_ATTACHMENT_CACHE_DIR: /var/cache/ntfy/attachments
      NTFY_UPSTREAM_BASE_URL: https://ntfy.sh   # needed for iOS push; Android works without
EOF
sudo docker compose up -d
# Reverse proxy (Caddy) — WebSocket support is default
printf 'ntfy.example.com {\n    reverse_proxy 127.0.0.1:8085\n}\n' | sudo tee -a /etc/caddy/Caddyfile && sudo systemctl reload caddy

deny-all is deliberate: an open ntfy server is an open spam channel to your team's phones. Create accounts and permissions:

N="sudo docker compose exec -T ntfy ntfy"
$N user add --role=admin ops-admin
$N user add jeroen; $N user add marie; $N user add tom
$N user add --role=user monitoring                     # the servers that send alerts
$N access monitoring 'alerts-*' write-only
$N access jeroen 'alerts-*' read-write; $N access marie 'alerts-*' read-write; $N access tom 'alerts-*' read-write
$N token add monitoring                                # token for the scripts, instead of a password

On the phone: ntfy app (F-Droid, Play Store, App Store), server https://ntfy.example.com, log in with the personal account.

Step 2: the rotation schedule

One text file, on the host that routes the alerts. ISO weeks, or date ranges — whatever your team uses to make agreements.

sudo tee /etc/oncall/schedule.tsv >/dev/null <<'EOF'
# from	to	primary	backup	phone_primary (for the escalation text)
2026-09-14	2026-09-20	jeroen	marie	+32470000001
2026-09-21	2026-09-27	marie	tom	+32470000002
2026-09-28	2026-10-04	tom	jeroen	+32470000003
2026-10-05	2026-10-11	jeroen	marie	+32470000001
EOF

Who's on duty now:

oncall() {  # oncall primary|backup|phone
  awk -F'\t' -v d="$(date +%F)" -v f="$1" '!/^#/ && $1<=d && d<=$2 { print (f=="primary"?$3:(f=="backup"?$4:$5)); exit }' /etc/oncall/schedule.tsv
}
echo "primary: $(oncall primary), backup: $(oncall backup)"

One topic per person: alerts-jeroen, alerts-marie, alerts-tom. Everyone subscribes to their own topic and to alerts-all (informational, muted). The schedule determines which personal topic the critical alerts go to.

Step 3: the router — one entry point for all scripts

Instead of every monitoring script pushing to a topic itself, they send to one local script that routes, deduplicates and schedules escalation.

sudo tee /usr/local/bin/alert >/dev/null <<'EOF'
#!/usr/bin/env bash
# alert <severity: info|warn|crit> <title> [body via stdin]
# info → alerts-all (silent). warn → alerts-all + primary (normal prio). crit → primary (urgent, breaks silent mode) + escalation.
set -u
SEV=$1; TITLE=$2; BODY=$(cat); NTFY=https://ntfy.example.com; TOK=$(cat /etc/oncall/token)
S=/etc/oncall/schedule.tsv; D=$(date +%F)
PRIM=$(awk -F'\t' -v d="$D" '!/^#/ && $1<=d && d<=$2 {print $3; exit}' "$S")
BACK=$(awk -F'\t' -v d="$D" '!/^#/ && $1<=d && d<=$2 {print $4; exit}' "$S")
ID=$(printf '%s' "$TITLE" | sha1sum | cut -c1-12)          # dedup key per title
ST=/var/lib/oncall; mkdir -p "$ST"

# Dedup: don't push the same title again within 30 min (except crit escalation)
if [ -f "$ST/$ID.sent" ] && [ "$(find "$ST/$ID.sent" -mmin -30)" ]; then exit 0; fi
touch "$ST/$ID.sent"

push() { # push <topic> <prio> <tags> [extra headers...]
  local topic=$1 prio=$2 tags=$3; shift 3
  curl -s -u ":$TOK" -H "Title: $TITLE" -H "Priority: $prio" -H "Tags: $tags" "$@" --data-binary "$BODY" "$NTFY/$topic" >/dev/null
}
case $SEV in
  info) push alerts-all min information_source ;;
  warn) push alerts-all default warning; push "alerts-$PRIM" default warning ;;
  crit)
    # Actions header: one tap = acknowledge (POST to ack topic) — see step 4
    push "alerts-$PRIM" urgent rotating_light -H "Actions: http, ACK, $NTFY/ack, method=POST, body=$ID $PRIM, clear=true"
    push alerts-all default rotating_light
    # Schedule escalation: after 10 min without ack → backup; after 20 min → both + phone in the text
    echo "$ID	$(date +%s)	$PRIM	$BACK	$TITLE" >> "$ST/pending.tsv" ;;
esac
EOF
sudo chmod 0755 /usr/local/bin/alert
sudo install -d -m 0750 /etc/oncall && echo "tk_xxxxxxxxxxxxxxxx" | sudo tee /etc/oncall/token >/dev/null && sudo chmod 0640 /etc/oncall/token
# Usage in every monitoring script:
echo "disk /var 91% on web-01" | alert crit "disk almost full web-01"

Priority: urgent is what makes the difference at 03:00: the ntfy app lets an urgent message through the "do not disturb" mode (Android: if the app has that permission; iOS: if "critical alerts" is allowed). min appears silently in the list.

Step 4: acknowledge and escalation

The Actions header puts an "ACK" button under the notification. One tap does a POST to the ack topic with the alert id. A second cron reads those acks and escalates whatever isn't acknowledged.

sudo tee /usr/local/sbin/oncall-escalate.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
# Every minute: read acks from the ack topic, escalate unacknowledged crit alerts.
NTFY=https://ntfy.example.com; TOK=$(cat /etc/oncall/token); ST=/var/lib/oncall
[ -s "$ST/pending.tsv" ] || exit 0
# Fetch acks from the last 24 h (JSON, poll)
curl -s -u ":$TOK" "$NTFY/ack/json?poll=1&since=24h" | jq -r '.message // empty' | awk '{print $1}' | sort -u > "$ST/acked.txt"
now=$(date +%s); : > "$ST/pending.new"
while IFS=$'\t' read -r id t0 prim back title; do
  if grep -qx "$id" "$ST/acked.txt"; then continue; fi          # acknowledged → done
  age=$(( (now - t0) / 60 ))
  if [ "$age" -ge 20 ] && [ ! -f "$ST/$id.esc2" ]; then
    phone=$(awk -F'\t' -v d="$(date +%F)" '!/^#/ && $1<=d && d<=$2 {print $5; exit}' /etc/oncall/schedule.tsv)
    for p in "$prim" "$back"; do
      curl -s -u ":$TOK" -H "Title: ESCALATION 20m: $title" -H "Priority: urgent" -H "Tags: rotating_light,sos" \
        -d "Not acknowledged after 20 min. Primary $prim ($phone) and backup $back: someone call." "$NTFY/alerts-$p" >/dev/null
    done; touch "$ST/$id.esc2"
  elif [ "$age" -ge 10 ] && [ ! -f "$ST/$id.esc1" ]; then
    curl -s -u ":$TOK" -H "Title: ESCALATION 10m: $title" -H "Priority: urgent" -H "Tags: rotating_light" \
      -H "Actions: http, ACK, $NTFY/ack, method=POST, body=$id $back, clear=true" \
      -d "Primary $prim has not acknowledged. You are backup." "$NTFY/alerts-$back" >/dev/null
    touch "$ST/$id.esc1"
  fi
  [ "$age" -lt 120 ] && printf '%s\t%s\t%s\t%s\t%s\n' "$id" "$t0" "$prim" "$back" "$title" >> "$ST/pending.new"   # give up after 2 h
done < "$ST/pending.tsv"
mv "$ST/pending.new" "$ST/pending.tsv"
EOF
sudo chmod 0755 /usr/local/sbin/oncall-escalate.sh
echo '* * * * * root /usr/local/sbin/oncall-escalate.sh' | sudo tee /etc/cron.d/oncall-escalate
# The ack topic: everyone may write (the button), only the router reads
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access jeroen ack write-only
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access marie ack write-only
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access tom ack write-only
sudo docker compose -f /opt/ntfy/docker-compose.yml exec -T ntfy ntfy access monitoring ack read-write

The chain: crit alert → primary (urgent, ACK button) → after 10 min without ack: backup (urgent, ACK button) → after 20 min: both, with phone number in the text so whoever reads it can call the other. Whoever taps ACK stops the escalation; the ack is in the ack topic with a name — that's your log of who responded when.

Step 5: test it, every week

# Monday 09:00: a test alert to the primary, with ACK. Not acknowledged after 10 min = the schedule or the app is wrong.
echo "0 9 * * 1 root echo 'Weekly on-call test — tap ACK' | /usr/local/bin/alert crit \"on-call test $(date +\%F)\"" | sudo tee /etc/cron.d/oncall-test

And once a quarter a real exercise: someone fills a disk on a test host at 22:00, and you measure how long it takes until the ack. That number (MTTA, mean time to acknowledge) is what a customer with an SLA wants to see.

Pitfalls

  • iOS push without upstream. On iPhones ntfy pushes only arrive via Apple's push service, and for that your server must point NTFY_UPSTREAM_BASE_URL to ntfy.sh. Only the message id goes to ntfy.sh, not the content — but it's a dependency, and you should know it. Android works fully standalone.
  • Battery optimisation. Android kills background connections; the ntfy app must be on the exception list (the app asks for it). Without that, urgent alerts arrive minutes late. That's one of the things the weekly test catches.
  • A schedule that runs out. If schedule.tsv has no line for today, PRIM is empty and the alert goes to alerts- — to nobody. Add a fallback to the script (PRIM=${PRIM:-jeroen}) and a daily check that warns when the schedule ends within 14 days.
  • Dedup hiding a second incident. "disk almost full web-01" at 03:00 and again at 03:20 (after a cleanup that didn't help) is the same title within 30 minutes → suppressed. Put the measured value in the body, not the title, and shorten the dedup window for crit.
  • One person in the schedule. A rotation with one name isn't a rotation but a burnout. At least two, preferably three; and this week's backup is next week's primary.

What you still don't have

  • Overrides. "Marie is ill, Tom takes over until Thursday" is an edit in schedule.tsv via ssh — at 06:30, from a phone. A UI is missing.
  • Per-customer routing. An MSP wants alerts from customer A to go to customer A's team, and customer A to have a topic of their own. That's a second dimension in the schedule and the router.
  • Reporting. How many alerts per week, how many at night, average time to ack, who was woken most often: scattered across the ack topic and pending.tsv.

How monsys does it

monsys uses the same building block — a self-hosted ntfy server, no Twilio, no PagerDuty — but the on-call rotation lives per group or tenant in the dashboard: shifts with start, end and contact, overrides without ssh, and per alert the person on duty is looked up automatically and added to the push. Acknowledge and resolve happen in the dashboard or via the push action, with name and time in the audit log; MTTA and MTTR per group are in the operations metrics and the Trust Score. For MSPs: alerts from customer A go to the on-call of customer A's group, and the customer can follow along on their own ntfy topic.

FAQ

Is ntfy reliable enough for on-call?

For a team managing its own server and doing the weekly test: yes. The protocol is simple (HTTP + WebSocket), the app is stable and the urgent priority breaks through silent modes. The weak spot isn't ntfy but the phone: battery optimisation and forgotten permissions.

What about SMS or calling as a fallback?

Deliberately not in this setup: SMS gateways (Twilio and co) are US services with an SLA no better than ntfy's, and calling requires a telephony provider. The escalation after 20 minutes therefore puts the phone number in the text, so that a human calls. For an organisation under NIS2 that's also the more defensible choice.

Can I combine this with Alertmanager or Uptime Kuma?

Yes. Both can send to a webhook or directly to ntfy; have them send to a local endpoint that calls /usr/local/bin/alert, or configure them directly on alerts-all and use the router only for crit. The schedule and the escalation then stay in one place.

Written by the monsys team — sysadmins who do this every day.

Done it by hand? Let monsys keep it running.

Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.