Monitoring an Ubuntu server without agent sprawl: what you actually need to measure
Four measurements that make 90% of incidents visible in advance, using only what ships with Ubuntu 24.04 or Debian 12: sysstat, journalctl, df and a thirty-line cron script that pushes to your phone. Plus why "CPU above 90%" is a bad alert.
Contents
Most servers don't go down because of something exotic. They go down because a disk filled up, a process slowly ate all the memory, a service didn't come back after an update, or a journal wrote the root partition full. All four are visible days ahead — if someone is looking. This guide sets up that looking with what's already on the machine: no agent, no SaaS, no Grafana. Then you'll read what you're still missing.
Step 1: know what's happening right now (the five commands)
Before automating anything, you need to know the manual version. These are the five commands you run first on every "the server is slow" report.
# 1. Who is eating CPU and memory?
top -bn1 | head -20
# 2. Is it CPU, I/O or swap? (sysstat)
sudo apt install -y sysstat
vmstat 1 5 # columns r (run queue), si/so (swap in/out), wa (I/O wait)
iostat -xz 1 3 # %util per disk, await in ms
# 3. Disk: bytes AND inodes
df -h --output=target,pcent,avail | sort -k2 -rn | head
df -i | awk 'NR==1 || $5+0 > 80'
# 4. What failed?
systemctl --failed
journalctl -p err -b --no-pager | tail -30
# 5. Did the OOM killer shoot anything?
journalctl -k -b | grep -iE 'out of memory|oom-kill' | tail
Two things beginners misread here:
free -malmost always shows "free" as low. Look at the available column. Linux uses free memory as page cache and hands it back when needed.load averageis not a percentage. A load of 4 on an 8-core machine is quiet; on a 2-core machine it means processes are permanently waiting. Always compare againstnproc.
Step 2: turn on history (sysstat)
Without history you can never say "this has been going on since Tuesday". sysstat takes a snapshot every 10 minutes and keeps 28 days by default.
sudo apt install -y sysstat
# Ubuntu 24.04 and Debian 12 use systemd timers:
sudo systemctl enable --now sysstat-collect.timer sysstat-summary.timer
# On older Debian/Ubuntu: set ENABLED="true" in /etc/default/sysstat
# From tomorrow on you can look back:
sar -u # CPU today, per 10 min
sar -r # memory
sar -d -p # disk I/O per device
sar -n DEV # network per interface
sar -u -f /var/log/sysstat/sa12 # the 12th of this month
Want 90 days? Set HISTORY=90 in /etc/sysstat/sysstat. It costs a few MB per month.
Step 3: the four measurements that matter
You could measure a hundred things. These are the four whose absence costs you an incident:
| Measurement | Why | Threshold that works |
|---|---|---|
| Disk fill and its direction | Full disk = database stops, logs stop, updates fail | > 85% or "full within 7 days" based on growth |
| Failed units + reboot-required | A service that didn't come up after an update is noticed by the customer first | systemctl --failed ≠ 0, or /var/run/reboot-required exists > 3 days |
| OOM kills | Today's memory leak is next week's outage | ≥ 1 OOM kill in 24 h |
| Heartbeat | The server itself is gone, so it reports nothing | No ping in 5 min |
Note what is not on the list: "CPU > 90%". A backup, an apt upgrade, a nightly logrotate with compression, a cron job building a report — all 100% CPU, all normal. That alert teaches you to ignore it within a week, and then you miss the real one. Alert on symptoms (queues, I/O wait above 30% for 10 minutes, response time) or on deviation from the normal pattern for that hour, not on an absolute number.
Step 4: a cron script that talks to your phone
Thirty lines, no dependencies beyond curl. Push goes to ntfy — self-hostable, free, EU-hostable. Replace NTFY with your own topic.
sudo tee /usr/local/sbin/healthcheck.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
# Runs every 10 min via cron. Reports only what deviates.
NTFY="https://ntfy.example.be/servers"
HOST=$(hostname -s)
ALERTS=()
# Disk > 85% (bytes or inodes), tmpfs excluded
while read -r target pct; do
ALERTS+=("disk $target ${pct}")
done < <(df -x tmpfs -x devtmpfs --output=target,pcent | awk 'NR>1 && $2+0>85')
while read -r target pct; do
ALERTS+=("inodes $target ${pct}")
done < <(df -x tmpfs -x devtmpfs --output=target,ipcent | awk 'NR>1 && $2+0>85')
# Failed units
FAILED=$(systemctl --failed --no-legend | awk '{print $1}' | tr '\n' ' ')
[ -n "$FAILED" ] && ALERTS+=("failed: $FAILED")
# Reboot required for > 3 days
if [ -f /var/run/reboot-required ] && [ "$(find /var/run/reboot-required -mtime +3)" ]; then
ALERTS+=("reboot-required > 3d ($(tr '\n' ' ' </var/run/reboot-required.pkgs))")
fi
# OOM kills in the last 24 h
OOM=$(journalctl -k --since "24 hours ago" --no-pager | grep -ci 'out of memory')
[ "$OOM" -gt 0 ] && ALERTS+=("oom-kills: $OOM")
# Load > 2x cores, 15-min average (so no short spike)
CORES=$(nproc); LOAD15=$(awk '{print $3}' /proc/loadavg)
awk -v l="$LOAD15" -v c="$CORES" 'BEGIN{exit !(l > 2*c)}' && ALERTS+=("load15 $LOAD15 on $CORES cores")
if [ ${#ALERTS[@]} -gt 0 ]; then
printf '%s\n' "${ALERTS[@]}" | curl -s -H "Title: $HOST" -H "Priority: high" --data-binary @- "$NTFY" >/dev/null
fi
EOF
sudo chmod 0755 /usr/local/sbin/healthcheck.sh
echo '*/10 * * * * root /usr/local/sbin/healthcheck.sh' | sudo tee /etc/cron.d/healthcheck
Test with an artificial failure: sudo systemctl mask cron && sudo systemctl restart cron produces a failed unit. Don't forget to unmask.
Heartbeat: the server that stops talking
The script above can't report that the server itself is gone. For that you need a second machine (or your laptop) that expects something to arrive. Simplest form: the script does a curl to an ntfy topic heartbeat-$HOST every 10 minutes, and a cron on another host checks whether the last message is older than 15 minutes via curl "https://ntfy.example.be/heartbeat-$HOST/json?poll=1&since=15m". It's crude, and it works.
Step 5: predicting disk growth instead of discovering it
"85%" on a disk that grows 1% per month is not a problem. "60%" on a disk growing 5% per day is an outage in eight days. With sar -F (Ubuntu 24.04+) or your own little log you compute the slope:
# Log the fill level every day
echo "0 6 * * * root df -B1 --output=target,used / /var | tail -n +2 | sed \"s/^/\$(date +\\%F) /\" >> /var/log/diskgrowth.log" | sudo tee /etc/cron.d/diskgrowth
# After a week: bytes per day and days until full
awk -v total="$(df -B1 --output=size / | tail -1)" '
$2=="/" {n++; x[n]=n; y[n]=$3}
END {
for(i=1;i<=n;i++){sx+=x[i];sy+=y[i];sxy+=x[i]*y[i];sxx+=x[i]*x[i]}
slope=(n*sxy-sx*sy)/(n*sxx-sx*sx)
if (slope<=0) {print "shrinking or stable"; exit}
printf "growth %.1f MB/day, full in %.0f days\n", slope/1e6, (total-y[n])/slope
}' /var/log/diskgrowth.log
That's a linear regression in awk. Not elegant, but enough to plan a purchase or a cleanup two weeks ahead.
Pitfalls
- journald eats your disk.
journalctl --disk-usagesurprises people. SetSystemMaxUse=500Min/etc/systemd/journald.confandsystemctl restart systemd-journald. dfdoesn't see deleted-but-open files. A 20 GB log file that was deleted while nginx still has it open keeps occupying space.sudo lsof +L1shows them; reloading the process frees the space.- Snapshots and
/boot. LVM snapshots and old kernels fill/bootuntil anapt upgradefails. Include/bootin the disk check and runapt autoremove --purgeregularly. - Timezone in cron.
cronruns in the system timezone; so do your ntfy messages andsar. If the server is in UTC and you're looking from Brussels, "at 3 o'clock" is wrong.timedatectltells you. - Alerts without an owner. A script pushing to a topic nobody subscribes to anymore is not monitoring. Put it in the calendar: one test alert per quarter.
What you still don't have
This works excellently for one to five servers. After that you hit the same four walls:
- No history across hosts.
sarlooks per machine. "Which of my twenty servers is growing fastest?" is a for-loop over ssh, every single time. - No deduplication. A disk at 86% produces a push every 10 minutes until someone acts. After a day everyone ignores the topic.
- No baseline. The script doesn't know this host always has load 6 on Tuesday nights because of the backup. You do — your replacement doesn't.
- No evidence. A NIS2 or ISO 27001 auditor asks "how do you know your monitoring works?" and a cron file is a weak answer.
How monsys does it
The monsys agent (one static Rust binary, no dependencies) measures the same things — CPU, memory, disk per mount, network per NIC, failed units, OOM kills, reboot-required — every 15 seconds and ships aggregated signals to the hub. The hub builds an hour-of-day baseline per host so you can alert on deviation instead of an absolute number, deduplicates alerts to one open row per cause, and keeps 13 months of history for capacity questions. Heartbeat silence is a separate detection, distinguishing "briefly away" from "silent for days".
The first five servers are free, forever. Install is one line: curl -fsSL https://get.monsys.ai/install.sh | sudo bash.
FAQ
Is 90% CPU usage a problem?
Not by itself. A server using its CPU is doing its job. It becomes a problem when a queue builds at the same time (column r in vmstat structurally larger than the core count) or when your application's response time rises. Alert on those symptoms, not on the percentage.
Do I need Prometheus or Grafana for a few servers?
No. For fewer than five servers, sysstat plus a cron script with push notifications gives you 90% of the value for 5% of the maintenance. Prometheus becomes interesting when you want dashboards across multiple hosts or start scraping application metrics.
Does this also work on Debian 12?
Yes, all commands in this article were tested on Debian 12 and Ubuntu 24.04. The only difference is the name of some units; sysstat-collect.timer exists on both.
Written by the monsys team — sysadmins who do this every day.
Done it by hand? Let monsys keep it running.
Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.