Server monitoringintermediate5 min read

Monitoring Docker containers on a single VPS: restart loops, memory leaks and image drift

Everything you can see yourself with docker stats, docker events and a healthcheck in your compose file — plus the cron script that reports a container restarting every 30 seconds within five minutes. Including the logging setting that stops containers from filling your disk.

Contents
  1. Step 1: the five commands
  2. Step 2: healthchecks that mean something
  3. Step 3: catching memory leaks before the OOM killer does
  4. Step 4: the cron script
  5. Step 5: image drift — are you running what you think you're running?
  6. Pitfalls
  7. What you still don't have
  8. How monsys does it
  9. FAQ

A VPS with ten containers fails differently from a VPS with ten services. The container that crashes is restarted automatically (restart: unless-stopped), so nobody sees it — until the customer asks why the webshop disappears for three seconds every minute. The memory leak sits in a container with a limit, so the host stays healthy while the app gets killed by the OOM killer every few hours. And the latest tag you pulled in March is not today's latest. This article makes those three things visible using nothing but Docker itself.

Step 1: the five commands

# What's running, and for how long? "Up 2 minutes" on a service that should run for weeks = restart loop
docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'

# Live resource usage, one-shot (no TUI)
docker stats --no-stream --format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.MemPerc}}\t{{.NetIO}}\t{{.BlockIO}}'

# How often has each container restarted since it was created?
docker ps -q | xargs docker inspect --format '{{.Name}} restarts={{.RestartCount}} oom={{.State.OOMKilled}} exit={{.State.ExitCode}}'

# What happened in the last hour? (die, oom, restart, health_status)
docker events --since 1h --until now --filter 'type=container' \
  --format '{{.Time}} {{.Actor.Attributes.name}} {{.Action}}'

# How much disk do images, containers, volumes and build cache use?
docker system df -v | head -40

Two interpretation mistakes everyone makes once:

  • docker stats shows CPU above 100% on multi-core hosts: 250% means 2.5 cores. Compare against nproc.
  • MemUsage includes the container's page cache. A database container at "95%" of its limit is often just caching efficiently. Look at OOMKilled and at docker events with oom before raising the limit.

Step 2: healthchecks that mean something

restart: unless-stopped restarts a process that stops. It does nothing for a process that hangs. For that you need a healthcheck, and it has to do what the customer does:

# docker-compose.yml
services:
  web:
    image: ghcr.io/example/shop:2026.09.1     # no :latest, see step 5
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "curl", "-fsS", "http://localhost:8080/healthz"]
      interval: 30s
      timeout: 5s
      retries: 3
      start_period: 40s
    deploy:
      resources:
        limits:
          memory: 512m
    logging:
      driver: json-file
      options:
        max-size: "20m"
        max-file: "5"

Three details:

  • The /healthz endpoint must touch the database. A healthcheck that just returns "200 OK" while every real request times out is worse than no healthcheck.
  • Docker does not restart an unhealthy container. The label goes to unhealthy and that's it. You have to act yourself (step 4) or run a tool like autoheal.
  • The logging block is not optional. Without max-size, a container in a crash loop writes hundreds of MB per hour to /var/lib/docker/containers//-json.log. Set it as the default in /etc/docker/daemon.json too:
{
  "log-driver": "json-file",
  "log-opts": { "max-size": "20m", "max-file": "5" }
}

Existing containers only get the new default after docker compose up -d --force-recreate.

Step 3: catching memory leaks before the OOM killer does

You don't see a leak in one measurement but in the direction. Log memory usage per container every five minutes and look at what rises monotonically:

echo '*/5 * * * * root docker stats --no-stream --format "{{.Name}} {{.MemUsage}}" | sed "s/^/$(date +\%s) /" >> /var/log/docker-mem.log' \
  | sudo tee /etc/cron.d/docker-mem

# After a day: per container the difference between first and last sample
awk '{
  split($3,a,"MiB"); mb=a[1]; if ($3 ~ /GiB/) {split($3,g,"GiB"); mb=g[1]*1024}
  if (!($2 in first)) first[$2]=mb; last[$2]=mb
} END { for (c in last) printf "%-30s %8.0f -> %8.0f MiB (%+.0f)\n", c, first[c], last[c], last[c]-first[c] }' /var/log/docker-mem.log | sort -k5 -rn

A container that goes from 180 to 470 MiB in 24 hours without traffic changing is leaking. Better to see that on Tuesday than to have the OOM killer tell you on Saturday night.

Step 4: the cron script

This script runs every five minutes and reports via ntfy only what deviates: restart loops, OOM kills, unhealthy containers and the usual disk check for /var/lib/docker.

sudo tee /usr/local/sbin/docker-check.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
NTFY="https://ntfy.example.be/docker"
HOST=$(hostname -s)
STATE=/var/lib/docker-check.state   # previous RestartCount per container
ALERTS=()
touch "$STATE"

for id in $(docker ps -aq); do
  read -r name restarts oom health status < <(docker inspect --format \
    '{{.Name}} {{.RestartCount}} {{.State.OOMKilled}} {{if .State.Health}}{{.State.Health.Status}}{{else}}none{{end}} {{.State.Status}}' "$id")
  name=${name#/}
  prev=$(grep "^$name " "$STATE" | awk '{print $2}'); prev=${prev:-0}
  # ≥ 3 restarts since the previous check (5 min) = loop
  [ $((restarts - prev)) -ge 3 ] && ALERTS+=("$name: $((restarts - prev)) restarts in 5 min")
  [ "$oom" = "true" ] && ALERTS+=("$name: OOM-killed")
  [ "$health" = "unhealthy" ] && ALERTS+=("$name: unhealthy")
  [ "$status" = "exited" ] && ALERTS+=("$name: exited")
  sed -i "/^$name /d" "$STATE"; echo "$name $restarts" >> "$STATE"
done

# Docker data dir > 85%
pct=$(df --output=pcent /var/lib/docker | tail -1 | tr -dc '0-9')
[ "$pct" -gt 85 ] && ALERTS+=("/var/lib/docker at ${pct}%")

if [ ${#ALERTS[@]} -gt 0 ]; then
  printf '%s\n' "${ALERTS[@]}" | curl -s -H "Title: docker@$HOST" -H "Priority: high" --data-binary @- "$NTFY" >/dev/null
fi
EOF
sudo chmod 0755 /usr/local/sbin/docker-check.sh
echo '*/5 * * * * root /usr/local/sbin/docker-check.sh' | sudo tee /etc/cron.d/docker-check

Test: docker run -d --name crashy --restart=always alpine sh -c 'sleep 5; exit 1' and wait five minutes. Clean up with docker rm -f crashy.

Step 5: image drift — are you running what you think you're running?

image: nginx:latest means: the version that was there when you last pulled. Two servers with the same compose file can be months apart. To check:

# Local digest vs. what the registry serves under that tag right now
for img in $(docker ps --format '{{.Image}}' | sort -u); do
  local=$(docker image inspect --format '{{index .RepoDigests 0}}' "$img" 2>/dev/null | cut -d@ -f2)
  remote=$(docker manifest inspect -v "$img" 2>/dev/null | jq -r '.[0].Descriptor.digest // .Descriptor.digest' 2>/dev/null)
  [ "$local" != "$remote" ] && echo "DRIFT $img local=${local:0:19} remote=${remote:0:19}"
done

The structural fix isn't a script but a habit: pin to a version tag (nginx:1.27.2) or a digest (nginx@sha256:…), and update deliberately. Then "which version is running" is a grep in git instead of a question to the registry.

Pitfalls

  • docker compose up -d after an image update only recreates containers whose image changed — but not if you haven't pulled the tag. docker compose pull && docker compose up -d is the pair.
  • docker system prune in cron. It also removes stopped containers and dangling volumes. Do it deliberately, with --filter "until=168h", and never with --volumes automated.
  • Healthcheck commands that aren't in the image. curl is missing from many slim/alpine images. Use wget -qO- --spider or a built-in endpoint, and test with docker inspect --format '{{json .State.Health}}'.
  • Memory limit without swap limit. With memory: 512m but no memswap_limit a container can still swap and slow down the whole host. On a VPS without swap that's no issue; with swap it is.
  • Timezone in containers. Logs in UTC while the host is in Europe/Brussels is the classic "the crash was at 02:14 — no, at 04:14" misunderstanding. Mount /etc/localtime:ro or set TZ=.

What you still don't have

  • History per container across hosts. Today's docker-mem.log doesn't help with "since which deploy has this been leaking".
  • What's inside that image? Restart loops and memory are the operational side. The security side — which CVEs sit in the 340 packages of your python:3.12-slim — you can't see with any docker command. For that, Trivy is the free reference.
  • Correlation. The container OOM-killed at 03:00 and the backup job starting at 02:55 live in two different logs.

How monsys does it

The monsys agent reads the Docker socket and reports per container resource usage, restart events, health status and the image digest. The hub compares digests against the registry (image-latest lookup), scans every image with Trivy and links CVEs to the container running them. Restart loops and OOM kills arrive as deduplicated alerts with the host metrics of the same moment next to them. See Applications in the docs.

FAQ

Does Docker restart an unhealthy container?

No. The restart policy reacts to a process that stops, not to a failing healthcheck. An unhealthy container keeps running with the label unhealthy. You need an external process (a cron script, autoheal, or an orchestrator) to intervene.

How much memory should I give a container?

First measure a week without a limit using docker stats and take the 95th percentile plus 30%. A limit that's too low causes OOM kills that look like crashes; no limit means one leaking container takes the whole host down.

Is :latest always wrong?

For a test environment where you deliberately always want the newest: no. For production: yes, because you can never say with certainty which version is running, and a docker compose up on a second server can yield a different version than on the first.

Written by the monsys team — sysadmins who do this every day.

Done it by hand? Let monsys keep it running.

Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.