Verifying backups that say "succeeded": freshness, integrity and a restore test that logs itself
A backup job that returns exit 0 proves the job ran — not that there's anything usable in the archive. Three checks you automate: is the newest backup fresh enough (mtime check), is the archive readable (restic/borg check) and can you actually get something back out (monthly restore test with a log line). Plus the pitfall of rsync preserving mtimes.
Contents
Every organisation that has lost data had backups. The job ran every night, the mail said "Backup completed successfully", and on the day it was needed the archive turned out to be empty, three weeks old, encrypted with a key nobody had, or simply unopenable. This article automates the three questions you should really be asking every day: is it fresh, is it intact, and can I get something out of it?
Step 1: freshness — did something really arrive last night?
The simplest check works for every backup tool, including the tar cron from ten years ago: look at the newest file at the destination.
# Newest file under the backup path, with age in hours
DEST=/srv/backup # or a mount of your NAS, or the restic/borg repo directory
newest=$(find "$DEST" -type f -printf '%T@ %p\n' 2>/dev/null | sort -n | tail -1)
ts=${newest%% *}; file=${newest#* }
age_h=$(( ( $(date +%s) - ${ts%.*} ) / 3600 ))
echo "newest: $file age: ${age_h}h"
[ "$age_h" -gt 26 ] && echo "STALE: no new backup in ${age_h} hours"
The 26-hour threshold is deliberate: a daily job plus two hours of slack. For a job running hourly, use 3.
With restic and borg there's a better source than mtime — the repository's own snapshot list:
# restic
restic -r "$DEST" snapshots --latest 1 --json | jq -r '.[0] | "\(.time) \(.hostname) \(.paths|join(","))"'
# borg
borg list "$DEST" --last 1 --format '{time} {hostname} {name}{NL}'
Compare the time with now, and also watch the hostname: a repo where a different host accidentally writes looks fresh while your host hasn't backed up in weeks.
Step 2: integrity — is the archive readable?
"There's a 40 GB file" says nothing about the contents. restic and borg can verify their own repository; do it weekly on a sample and monthly in full.
# restic: structure + actually read 5% of the data (weekly)
restic -r "$DEST" check --read-data-subset=5%
# restic: everything (monthly; can take hours on large repos)
restic -r "$DEST" check --read-data
# borg
borg check --verify-data "$DEST" # full
borg check "$DEST" # metadata only, fast
For tar/zip archives without their own check: tar -tzf archive.tgz >/dev/null reads the whole archive and fails on corruption. For database dumps: pg_restore --list dump.pgc >/dev/null (PostgreSQL custom format) or gzip -t dump.sql.gz (compression layer only).
Step 3: the restore test that logs itself
This is the step everyone skips, and the only one that counts. Once a month you put something real back in a different place and verify it's correct. Not the whole server — a targeted set that's representative: the config of your most important service and a database dump.
sudo tee /usr/local/sbin/restore-test.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
# Monthly restore test. Restores a known set into a temporary directory,
# compares with the original, and logs the result with a date.
set -u
DEST=/srv/backup
LOG=/srv/inventory/restore-tests.log
NTFY="https://ntfy.example.be/backup"
T=$(mktemp -d /tmp/restore-test.XXXX)
HOST=$(hostname -s)
result=OK; detail=""
# 1. Config files: restore and compare byte-for-byte with live
if restic -r "$DEST" restore latest --target "$T" --include /etc/nginx --include /etc/ssh/sshd_config.d >/dev/null 2>&1; then
if ! diff -rq /etc/nginx "$T/etc/nginx" >/dev/null 2>&1; then
result=FAIL; detail+="nginx config differs from live; "
fi
else
result=FAIL; detail+="restic restore failed; "
fi
# 2. Database dump: restore into a throwaway database and count
if [ -x /usr/bin/psql ]; then
dump=$(ls -t "$T"/srv/dumps/*.pgc 2>/dev/null | head -1)
if [ -n "$dump" ]; then
sudo -u postgres createdb restoretest 2>/dev/null
if sudo -u postgres pg_restore -d restoretest "$dump" >/dev/null 2>&1; then
tables=$(sudo -u postgres psql -tAc "select count(*) from information_schema.tables where table_schema='public'" restoretest)
[ "${tables:-0}" -gt 0 ] || { result=FAIL; detail+="pg_restore: 0 tables; "; }
else
result=FAIL; detail+="pg_restore failed; "
fi
sudo -u postgres dropdb restoretest 2>/dev/null
fi
fi
rm -rf "$T"
line="$(date -Is) restore-test $HOST $result ${detail:-config+db restored and compared}"
echo "$line" | sudo tee -a "$LOG" >/dev/null
[ "$result" = FAIL ] && echo "$line" | curl -s -H "Title: RESTORE-TEST FAILED $HOST" -H "Priority: urgent" --data-binary @- "$NTFY" >/dev/null
echo "$line"
EOF
sudo chmod 0755 /usr/local/sbin/restore-test.sh
echo '0 4 1 * * root /usr/local/sbin/restore-test.sh' | sudo tee /etc/cron.d/restore-test
The log line — 2026-10-01T04:00:12+02:00 restore-test web-01 OK config+db restored and compared — is exactly what an auditor means by "evidence of recovery tests". Twelve lines per year, per host.
Adapt the --include paths and the dump location to your situation; the principle is: restore something you can compare, and a database you can query.
Step 4: the daily check, with alert
Steps 1 and 2 in one script that runs every morning and only reports what's wrong:
sudo tee /usr/local/sbin/backup-check.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
DEST=/srv/backup; MAX_H=26
NTFY="https://ntfy.example.be/backup"
HOST=$(hostname -s); MSG=()
# Freshness via the repo itself (restic); falls back to mtime if restic is missing
if command -v restic >/dev/null; then
last=$(restic -r "$DEST" snapshots --latest 1 --json 2>/dev/null | jq -r '.[0].time // empty')
[ -n "$last" ] && age_h=$(( ( $(date +%s) - $(date -d "$last" +%s) ) / 3600 )) || age_h=9999
else
ts=$(find "$DEST" -type f -printf '%T@\n' 2>/dev/null | sort -n | tail -1); ts=${ts%.*}
age_h=$(( ( $(date +%s) - ${ts:-0} ) / 3600 ))
fi
[ "$age_h" -gt "$MAX_H" ] && MSG+=("STALE: last backup ${age_h}h old (max ${MAX_H}h)")
# Integrity, Sundays only (sample)
if [ "$(date +%u)" = 7 ] && command -v restic >/dev/null; then
restic -r "$DEST" check --read-data-subset=5% >/tmp/restic-check.log 2>&1 || MSG+=("CHECK FAILED: $(tail -1 /tmp/restic-check.log)")
fi
# Space at the destination
pct=$(df --output=pcent "$DEST" 2>/dev/null | tail -1 | tr -dc '0-9')
[ "${pct:-0}" -gt 85 ] && MSG+=("destination ${pct}% full")
if [ ${#MSG[@]} -gt 0 ]; then
printf '%s\n' "${MSG[@]}" | curl -s -H "Title: backup@$HOST" -H "Priority: high" --data-binary @- "$NTFY" >/dev/null
fi
EOF
sudo chmod 0755 /usr/local/sbin/backup-check.sh
echo '15 7 * * * root /usr/local/sbin/backup-check.sh' | sudo tee /etc/cron.d/backup-check
Pitfalls
- rsync preserves mtimes.
rsync -acopies files with their original modification date. An mtime check on an rsync destination therefore sees the date of the source file, not of the copy. For rsync backups use a marker:rsync ... && date -Is > "$DEST/.last-run"and check that file. - The job that "succeeds" with zero files. A wrongly mounted path, an empty Docker volume directory, an
--excludethat's too broad: exit 0, 12 KB archive. Besides age, also check the size of the last snapshot (restic stats latest) and alert on a drop of more than 50% versus the previous one. - Retention that cleans out everything.
restic forget --keep-daily 7 --prunewithout--keep-monthlymeans: after a week without new backups, the repo is empty. Always combine with a freshness check, or nobody notices. - The key on the same server. An encrypted restic repo whose password only lives in
/etc/restic/passwordon the backed-up server is worthless in a ransomware incident. Keep the key in a second place (password vault, paper in a safe). - Databases as files. A copy of
/var/lib/postgresqlfrom a running database is very likely inconsistent. Usepg_dump/pg_basebackup,mariadb-dump --single-transaction, or a filesystem-level snapshot with a consistent state. - Restore test on production. The script restores into
/tmpand a throwaway database. Neverrestore latest --target /. It sounds obvious until someone does it differently at 04:00.
What you still don't have
- Overview per host. Thirty hosts, thirty
restore-tests.log. "Which hosts had no successful restore test in Q3?" is a script over ssh. - Alerts that stop when it's fixed. The ntfy push comes back every day until someone fixes it, and then one more day because the check runs at 07:15.
- Evidence for ISO 27001 A.8.13. The log line is good; an auditor wants it in a signed, immutable form, across the whole fleet, and over twelve months.
How monsys does it
In monsys you register a backup watch per host: the path where the backup lands and the maximum age. The agent checks the path every 30 minutes (stat-only, it doesn't open the archives — so encrypted repos work without read access) and reports the newest mtime and size. The hub alerts on stale (warning) and on twice stale (critical), deduplicated to one open alert that closes by itself when the backup runs again. The control ISO 27001 A.8.13 automatically counts the watches within their max age and includes that in the monthly audit pack. See backup verification in the docs. The restore test remains human work — but you can attach its log line as evidence.
FAQ
How often should I do a restore test?
Monthly for a targeted set (config + database), and once a year a full recovery of a real service to a test environment, with a stopwatch. The latter is also the only way to know your RTO instead of estimating it.
Isn't the backup tool's own check enough?
restic check proves the repository is internally consistent. It doesn't prove the right paths were backed up, that the database dump is usable, or that you still have the key. That's what the restore test is for.
What if my backup is in the cloud (S3, Backblaze, Hetzner Storage Box)?
All commands work the same: restic and borg speak to those backends directly. The mtime check from step 1 doesn't work on object storage; use the tool's snapshot list there. And test the restore from the cloud location, not from a local cache — bandwidth is part of your RTO.
Written by the monsys team — sysadmins who do this every day.
Done it by hand? Let monsys keep it running.
Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.