Monitoring Docker containers on a single VPS: restart loops, memory leaks and image drift
Everything you can see yourself with docker stats, docker events and a healthcheck in your compose file — plus the cron script that reports a container restarting every 30 seconds within five minutes. Including the logging setting that stops containers from filling your disk.
Contents
A VPS with ten containers fails differently from a VPS with ten services. The container that crashes is restarted automatically (restart: unless-stopped), so nobody sees it — until the customer asks why the webshop disappears for three seconds every minute. The memory leak sits in a container with a limit, so the host stays healthy while the app gets killed by the OOM killer every few hours. And the latest tag you pulled in March is not today's latest. This article makes those three things visible using nothing but Docker itself.
Step 1: the five commands
# What's running, and for how long? "Up 2 minutes" on a service that should run for weeks = restart loop
docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'
# Live resource usage, one-shot (no TUI)
docker stats --no-stream --format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.MemPerc}}\t{{.NetIO}}\t{{.BlockIO}}'
# How often has each container restarted since it was created?
docker ps -q | xargs docker inspect --format '{{.Name}} restarts={{.RestartCount}} oom={{.State.OOMKilled}} exit={{.State.ExitCode}}'
# What happened in the last hour? (die, oom, restart, health_status)
docker events --since 1h --until now --filter 'type=container' \
--format '{{.Time}} {{.Actor.Attributes.name}} {{.Action}}'
# How much disk do images, containers, volumes and build cache use?
docker system df -v | head -40
Two interpretation mistakes everyone makes once:
docker statsshows CPU above 100% on multi-core hosts: 250% means 2.5 cores. Compare againstnproc.- MemUsage includes the container's page cache. A database container at "95%" of its limit is often just caching efficiently. Look at
OOMKilledand atdocker eventswithoombefore raising the limit.
Step 2: healthchecks that mean something
restart: unless-stopped restarts a process that stops. It does nothing for a process that hangs. For that you need a healthcheck, and it has to do what the customer does:
# docker-compose.yml
services:
web:
image: ghcr.io/example/shop:2026.09.1 # no :latest, see step 5
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-fsS", "http://localhost:8080/healthz"]
interval: 30s
timeout: 5s
retries: 3
start_period: 40s
deploy:
resources:
limits:
memory: 512m
logging:
driver: json-file
options:
max-size: "20m"
max-file: "5"
Three details:
- The
/healthzendpoint must touch the database. A healthcheck that just returns "200 OK" while every real request times out is worse than no healthcheck. - Docker does not restart an unhealthy container. The label goes to
unhealthyand that's it. You have to act yourself (step 4) or run a tool likeautoheal. - The
loggingblock is not optional. Withoutmax-size, a container in a crash loop writes hundreds of MB per hour to/var/lib/docker/containers//-json.log. Set it as the default in/etc/docker/daemon.jsontoo:
{
"log-driver": "json-file",
"log-opts": { "max-size": "20m", "max-file": "5" }
}
Existing containers only get the new default after docker compose up -d --force-recreate.
Step 3: catching memory leaks before the OOM killer does
You don't see a leak in one measurement but in the direction. Log memory usage per container every five minutes and look at what rises monotonically:
echo '*/5 * * * * root docker stats --no-stream --format "{{.Name}} {{.MemUsage}}" | sed "s/^/$(date +\%s) /" >> /var/log/docker-mem.log' \
| sudo tee /etc/cron.d/docker-mem
# After a day: per container the difference between first and last sample
awk '{
split($3,a,"MiB"); mb=a[1]; if ($3 ~ /GiB/) {split($3,g,"GiB"); mb=g[1]*1024}
if (!($2 in first)) first[$2]=mb; last[$2]=mb
} END { for (c in last) printf "%-30s %8.0f -> %8.0f MiB (%+.0f)\n", c, first[c], last[c], last[c]-first[c] }' /var/log/docker-mem.log | sort -k5 -rn
A container that goes from 180 to 470 MiB in 24 hours without traffic changing is leaking. Better to see that on Tuesday than to have the OOM killer tell you on Saturday night.
Step 4: the cron script
This script runs every five minutes and reports via ntfy only what deviates: restart loops, OOM kills, unhealthy containers and the usual disk check for /var/lib/docker.
sudo tee /usr/local/sbin/docker-check.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
NTFY="https://ntfy.example.be/docker"
HOST=$(hostname -s)
STATE=/var/lib/docker-check.state # previous RestartCount per container
ALERTS=()
touch "$STATE"
for id in $(docker ps -aq); do
read -r name restarts oom health status < <(docker inspect --format \
'{{.Name}} {{.RestartCount}} {{.State.OOMKilled}} {{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}} {{.State.Status}}' "$id")
name=${name#/}
prev=$(grep "^$name " "$STATE" | awk '{print $2}'); prev=${prev:-0}
# ≥ 3 restarts since the previous check (5 min) = loop
[ $((restarts - prev)) -ge 3 ] && ALERTS+=("$name: $((restarts - prev)) restarts in 5 min")
[ "$oom" = "true" ] && ALERTS+=("$name: OOM-killed")
[ "$health" = "unhealthy" ] && ALERTS+=("$name: unhealthy")
[ "$status" = "exited" ] && ALERTS+=("$name: exited")
sed -i "/^$name /d" "$STATE"; echo "$name $restarts" >> "$STATE"
done
# Docker data dir > 85%
pct=$(df --output=pcent /var/lib/docker | tail -1 | tr -dc '0-9')
[ "$pct" -gt 85 ] && ALERTS+=("/var/lib/docker at ${pct}%")
if [ ${#ALERTS[@]} -gt 0 ]; then
printf '%s\n' "${ALERTS[@]}" | curl -s -H "Title: docker@$HOST" -H "Priority: high" --data-binary @- "$NTFY" >/dev/null
fi
EOF
sudo chmod 0755 /usr/local/sbin/docker-check.sh
echo '*/5 * * * * root /usr/local/sbin/docker-check.sh' | sudo tee /etc/cron.d/docker-check
Test: docker run -d --name crashy --restart=always alpine sh -c 'sleep 5; exit 1' and wait five minutes. Clean up with docker rm -f crashy.
Step 5: image drift — are you running what you think you're running?
image: nginx:latest means: the version that was there when you last pulled. Two servers with the same compose file can be months apart. To check:
# Local digest vs. what the registry serves under that tag right now
for img in $(docker ps --format '{{.Image}}' | sort -u); do
local=$(docker image inspect --format '{{index .RepoDigests 0}}' "$img" 2>/dev/null | cut -d@ -f2)
remote=$(docker manifest inspect -v "$img" 2>/dev/null | jq -r '.[0].Descriptor.digest // .Descriptor.digest' 2>/dev/null)
[ "$local" != "$remote" ] && echo "DRIFT $img local=${local:0:19} remote=${remote:0:19}"
done
The structural fix isn't a script but a habit: pin to a version tag (nginx:1.27.2) or a digest (nginx@sha256:…), and update deliberately. Then "which version is running" is a grep in git instead of a question to the registry.
Pitfalls
docker compose up -dafter an image update only recreates containers whose image changed — but not if you haven't pulled the tag.docker compose pull && docker compose up -dis the pair.docker system prunein cron. It also removes stopped containers and dangling volumes. Do it deliberately, with--filter "until=168h", and never with--volumesautomated.- Healthcheck commands that aren't in the image.
curlis missing from many slim/alpine images. Usewget -qO- --spideror a built-in endpoint, and test withdocker inspect --format '{{json .State.Health}}'. - Memory limit without swap limit. With
memory: 512mbut nomemswap_limita container can still swap and slow down the whole host. On a VPS without swap that's no issue; with swap it is. - Timezone in containers. Logs in UTC while the host is in Europe/Brussels is the classic "the crash was at 02:14 — no, at 04:14" misunderstanding. Mount
/etc/localtime:roor setTZ=.
What you still don't have
- History per container across hosts. Today's
docker-mem.logdoesn't help with "since which deploy has this been leaking". - What's inside that image? Restart loops and memory are the operational side. The security side — which CVEs sit in the 340 packages of your
python:3.12-slim— you can't see with anydockercommand. For that, Trivy is the free reference. - Correlation. The container OOM-killed at 03:00 and the backup job starting at 02:55 live in two different logs.
How monsys does it
The monsys agent reads the Docker socket and reports per container resource usage, restart events, health status and the image digest. The hub compares digests against the registry (image-latest lookup), scans every image with Trivy and links CVEs to the container running them. Restart loops and OOM kills arrive as deduplicated alerts with the host metrics of the same moment next to them. See Applications in the docs.
FAQ
Does Docker restart an unhealthy container?
No. The restart policy reacts to a process that stops, not to a failing healthcheck. An unhealthy container keeps running with the label unhealthy. You need an external process (a cron script, autoheal, or an orchestrator) to intervene.
How much memory should I give a container?
First measure a week without a limit using docker stats and take the 95th percentile plus 30%. A limit that's too low causes OOM kills that look like crashes; no limit means one leaking container takes the whole host down.
Is :latest always wrong?
For a test environment where you deliberately always want the newest: no. For production: yes, because you can never say with certainty which version is running, and a docker compose up on a second server can yield a different version than on the first.
Written by the monsys team — sysadmins who do this every day.
Done it by hand? Let monsys keep it running.
Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.