Patching, backup & uptimebeginner5 min read

TLS certificates that expire silently: inventory the whole fleet and alert 14 days ahead

Let's Encrypt renews itself — until the day the hook fails, DNS validation breaks or someone removed the cron. This script finds every certificate on every port of every host (internal ones too, mail too, the reverse proxy nobody remembers), computes the remaining days and pushes an alert at 14 and 7 days. No Nagios, no external service.

Contents
  1. Step 1: find all certificates on one host
  2. Step 2: the certificates that don't listen
  3. Step 3: the alert — 14 and 7 days ahead, and on expiry
  4. Step 4: the fleet version
  5. Step 5: check from the outside
  6. Pitfalls
  7. What you still don't have
  8. How monsys does it
  9. FAQ

An expired certificate is the most predictable outage there is: the date is literally written inside it. And yet it happens every year at organisations with six-figure monitoring budgets, because the certificate was on the internal API, or on the mail server's SMTP port, or on the load balancer of a customer migrated in 2022. The problem is never "when does this certificate expire", it's "which certificates do I actually have".

Step 1: find all certificates on one host

Don't ask the configuration which certificates should be there; ask every listening port which certificate it serves. That also catches the forgotten stunnel, the Postgres with ssl=on and the Docker container with its own nginx.

sudo tee /usr/local/sbin/cert-inventory.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
# Prints per listening TCP port: port, SNI name, days to expiry, subject, issuer.
# Tries plain TLS and STARTTLS (smtp, imap, pop3, postgres) on the well-known ports.
# SNI: servers like Caddy and strict nginx configs refuse an unknown servername,
# so per port we try every name from /etc/cert-inventory.names, the FQDN, and no SNI.
HOST=$(hostname -s)
now=$(date +%s)
NAMES=( $(cat /etc/cert-inventory.names 2>/dev/null) "$(hostname -f 2>/dev/null || hostname)" "" )
for port in $(ss -tlnH | awk '{sub(/.*:/,"",$4); print $4}' | sort -un); do
  case $port in
    25|587) st="-starttls smtp" ;;
    143)    st="-starttls imap" ;;
    110)    st="-starttls pop3" ;;
    5432)   st="-starttls postgres" ;;
    3306)   st="-starttls mysql" ;;
    *)      st="" ;;
  esac
  seen=""
  for name in "${NAMES[@]}"; do
    sni=(); [ -n "$name" ] && sni=(-servername "$name")
    cert=$(echo | timeout 4 openssl s_client -connect "127.0.0.1:$port" "${sni[@]}" $st 2>/dev/null \
           | openssl x509 -noout -enddate -subject -issuer -fingerprint -sha256 2>/dev/null) || continue
    fp=$(sed -n 's/^sha256 Fingerprint=//p' <<<"$cert")
    case " $seen " in *" $fp "*) continue ;; esac     # same cert under another name: show once
    seen="$seen $fp"
    end=$(sed -n 's/^notAfter=//p' <<<"$cert")
    days=$(( ( $(date -d "$end" +%s) - now ) / 86400 ))
    subj=$(sed -n 's/^subject=//p' <<<"$cert" | sed 's/.*CN *= *//')
    iss=$(sed -n 's/^issuer=//p' <<<"$cert" | sed 's/.*O *= *\([^,]*\).*/\1/')
    printf '%s\t%s\t%s\t%s\t%s\t%s\n' "$HOST" "$port" "${name:--}" "$days" "$subj" "$iss"
  done
done
EOF
sudo chmod 0755 /usr/local/sbin/cert-inventory.sh
# Names your server serves via SNI (one per line); without this file only the FQDN + no SNI
printf 'shop.example.be\napi.example.be\nmail.example.be\n' | sudo tee /etc/cert-inventory.names
sudo /usr/local/sbin/cert-inventory.sh | column -t -s $'\t'
# web-01  443   shop.example.be   61   shop.example.be   Let's Encrypt
# web-01  443   api.example.be    61   shop.example.be   Let's Encrypt    ← same cert (SAN), deduplicated
# web-01  8443  -                -12   internal-api      Example Internal CA     ← expired 12 days ago
# web-01  25    mail.example.be  203   mail.example.be   Let's Encrypt

Negative days are expired certificates. On a server you inventory for the first time, you almost always find one.

Step 2: the certificates that don't listen

Some certificates sit in files only read at use time: client certificates for mTLS, a SAML signing cert, your VPN's CA. You find those on disk:

sudo find /etc /opt /srv /var/lib -type f \( -name '*.crt' -o -name '*.pem' -o -name '*.cer' \) 2>/dev/null \
  | while read -r f; do
      end=$(openssl x509 -in "$f" -noout -enddate 2>/dev/null | sed 's/notAfter=//') || continue
      days=$(( ( $(date -d "$end" +%s) - $(date +%s) ) / 86400 ))
      printf '%s\t%s\t%s\n' "$days" "$f" "$(openssl x509 -in "$f" -noout -subject 2>/dev/null | sed 's/.*CN *= *//')"
    done | sort -n | head -20

Skip private keys (.key): they have no expiry date, and you don't want tooling touching them. Let's Encrypt files under /etc/letsencrypt/archive show all* historical versions; filter on /live/.

Step 3: the alert — 14 and 7 days ahead, and on expiry

sudo tee /usr/local/sbin/cert-check.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
NTFY="https://ntfy.example.be/certs"
WARN=14; CRIT=7
MSG=()
while IFS=$'\t' read -r host port name days subj iss; do
  if   [ "$days" -lt 0 ];        then MSG+=("EXPIRED ${days#-}d ago: $subj ($host:$port, $iss)")
  elif [ "$days" -le "$CRIT" ];  then MSG+=("CRITICAL ${days}d: $subj ($host:$port, $iss)")
  elif [ "$days" -le "$WARN" ];  then MSG+=("warning ${days}d: $subj ($host:$port, $iss)")
  fi
done < <(/usr/local/sbin/cert-inventory.sh)
if [ ${#MSG[@]} -gt 0 ]; then
  prio=default; printf '%s\n' "${MSG[@]}" | grep -qE '^(EXPIRED|CRITICAL)' && prio=urgent
  printf '%s\n' "${MSG[@]}" | curl -s -H "Title: certs@$(hostname -s)" -H "Priority: $prio" --data-binary @- "$NTFY" >/dev/null
fi
EOF
sudo chmod 0755 /usr/local/sbin/cert-check.sh
echo '0 8 * * * root /usr/local/sbin/cert-check.sh' | sudo tee /etc/cron.d/cert-check

Why 14 days: Let's Encrypt renews at 30 days before expiry. A certificate not yet renewed on day 14 has therefore had a failing renewal for sixteen days — that's the signal, not the expiry itself.

Step 4: the fleet version

Run the inventory across all your hosts and collect in one place. With just ssh:

# hosts.txt: one hostname per line
while read -r h; do ssh -o ConnectTimeout=5 "$h" sudo /usr/local/sbin/cert-inventory.sh; done < hosts.txt \
  | sort -t$'\t' -k4 -n > /srv/inventory/certs-$(date +%F).tsv
column -t -s $'\t' /srv/inventory/certs-$(date +%F).tsv | head -20

The result, sorted by remaining days, is a list you can maintain monthly — and it's exactly the piece of evidence NIS2 Art. 21(2)(h) and ISO 27001 A.8.24 mean by "management of cryptographic assets".

Step 5: check from the outside

The inventory looks from the inside. What a visitor sees can differ: a CDN or load balancer in front serves its own certificate. Check the public side separately, from a host outside your network:

for d in shop.example.be api.example.be mail.example.be; do
  end=$(echo | timeout 5 openssl s_client -connect "$d:443" -servername "$d" 2>/dev/null | openssl x509 -noout -enddate | sed 's/notAfter=//')
  printf '%-24s %4s days\n' "$d" "$(( ( $(date -d "$end" +%s) - $(date +%s) ) / 86400 ))"
done

And once a quarter look in the Certificate Transparency logs (https://crt.sh/?q=%25.example.be) at which certificates have been issued for your domain at all. A certificate you didn't request is an incident.

Pitfalls

  • Let's Encrypt renewal failing without mail. certbot renew logs to /var/log/letsencrypt/ and stops silently. systemctl status certbot.timer and certbot certificates belong in your monthly round. The 14-day alert catches it too, but then you have two weeks less.
  • Renewed but not reloaded. The file on disk is new, but nginx/postfix/haproxy still serves the old one from memory. That's why step 1 checks the port, not the file. Put a --deploy-hook "systemctl reload nginx postfix" on certbot.
  • SNI. A server with several virtual hosts returns the default certificate without -servername — and Caddy or a strict nginx simply refuses the handshake (tlsv1 alert internal error). That's why the script reads /etc/cert-inventory.names; put every name the host serves in there, or you won't see the certificate. It deduplicates on fingerprint, so a SAN certificate appears once.
  • Internal CAs with long lifetimes. A ten-year internal certificate is no reason to relax; the CA itself expires too, and you notice that on all clients at once. Include CA files in step 2.
  • Timezone and the last day. notAfter is in GMT. On the last day, "0 days" may already be expired for your customer in another timezone. Treat 0 as expired.

What you still don't have

  • Deduplication. The same wildcard certificate on twelve hosts gives twelve alerts. The script doesn't know it's one certificate.
  • History. When was this certificate last renewed, and how long did the renewal outage last time? That needs a database, not a tsv per day.
  • The outside view, continuously. Step 5 is a manual loop from one place. What you want is a check that looks from outside every few minutes, and that also sees the port is closed — because that's the outage that comes before the expiry.

How monsys does it

The monsys agent does steps 1 and 2 on every host (TLS port probe plus certificate files), the hub deduplicates on fingerprint and keeps the renewal history per certificate. The uptime checks do step 5 from the hub: every HTTPS check reads the certificate chain and raises a uptime.cert_expiring alert at 14 days, which closes by itself after renewal. Certificates getting too close to expiry appear as capacity-as-CVE in the same list as your real vulnerabilities — with the same prioritisation, because an expired certificate on an internet-facing host is, operationally, exactly that.

FAQ

Why alert at 14 days when Let's Encrypt renews at 30?

Because the alert then means "automatic renewal has been failing for two weeks", and you still have two weeks to fix it. Alerting at 30 days gives noise (the renewal may still be in progress that day); alerting at 3 days gives a weekend with no margin.

I use a CDN; do I still need to check my own certificates?

Yes. The CDN serves its certificate to visitors, but between the CDN and your origin there's usually TLS too, with your certificate. If that expires, the CDN returns a 526 or 502 error — just as dead for visitors as an expired public certificate.

Can I do this for Windows servers too?

The principle is the same; the command is Get-ChildItem Cert:\LocalMachine\My | Select Subject, NotAfter in PowerShell, and the port probe works with openssl from a Linux host against the Windows port. Watch out for IIS bindings still pointing to an old certificate after a renewal.

Written by the monsys team — sysadmins who do this every day.

Done it by hand? Let monsys keep it running.

Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.