Patching, backup & uptimeintermediate6 min read

Updating the kernel on a production server: sequence, rollback plan and when livepatch is worth it

A kernel update is the only routine update that requires a reboot and that can leave your server hanging. This is the checklist we use ourselves: pre-checks (/boot, DKMS, Secure Boot), snapshot, installation, a GRUB fallback that automatically picks the old kernel after one failed boot, and the verification after the reboot. In a fifteen-minute window.

Contents
  1. Step 0: know why you're rebooting
  2. Step 1: pre-checks (five minutes that save you a night)
  3. Step 2: snapshot or backup
  4. Step 3: install, but don't reboot yet
  5. Step 4: the fallback — one failed boot and GRUB picks the old kernel itself
  6. Step 5: reboot, in the window
  7. Step 6: verify, then make it permanent
  8. Rollback: if step 6 is NOT good
  9. When livepatch is worth it
  10. Pitfalls
  11. What you still don't have
  12. How monsys does it
  13. FAQ

Every other update you can roll back with apt install pkg=old-version. A kernel update you roll back by choosing the old kernel in GRUB — and that requires that you can reach GRUB, that the old kernel is still there, and that you do all of that under time pressure while the server is down. That's why preparation is 80% of the work. This article is the checklist, in the order we work through it ourselves.

Step 0: know why you're rebooting

uname -r                                          # what's running
dpkg -l 'linux-image-[0-9]*' | awk '/^ii/{print $2}' | sort -V   # what's installed
cat /var/run/reboot-required.pkgs 2>/dev/null     # why the system wants a reboot

If the new kernel is already there (installed by unattended-upgrades, not yet booted), skip step 3. If a CVE is involved: first check in kernel CVEs and backports whether it affects your kernel — half of the "urgent" kernel reboots don't.

Step 1: pre-checks (five minutes that save you a night)

# 1. Space on /boot — each kernel is 100–150 MB; too little = half an install
df -h /boot
sudo apt autoremove --purge -y            # old kernels out, keep at least two

# 2. Out-of-tree modules (DKMS): these must be rebuilt for the new kernel
dkms status 2>/dev/null                    # nvidia, zfs, wireguard (old), virtualbox, ...
# Empty = no worries. Otherwise: make sure linux-headers-<new version> gets installed too.

# 3. Secure Boot + DKMS = modules must be signed with a MOK key
mokutil --sb-state 2>/dev/null             # "SecureBoot enabled" + DKMS → mokutil --list-enrolled

# 4. Which kernel does GRUB boot by default, and is there a menu to fall back to?
grep -E '^GRUB_(DEFAULT|TIMEOUT|TIMEOUT_STYLE)' /etc/default/grub
# GRUB_DEFAULT=0 and GRUB_TIMEOUT=0 means: no chance to choose on a hanging boot.

# 5. Is there a console if SSH doesn't come back? (cloud: VNC/serial in the panel; bare metal: IPMI/iDRAC)
#    This isn't a command, it's: do you know where the button is BEFORE you hit reboot?

# 6. Is anything running that needs a clean stop?
systemctl list-units --type=service --state=running --no-legend | grep -Ei 'postgres|mysql|mariadb|redis|rabbit|docker'

Point 5 is the most important. A kernel that hangs on a host you have no console for is a ticket with the hoster and an hour of downtime.

Step 2: snapshot or backup

On a VPS: snapshot via the hoster's panel or API (Hetzner, OVH, Scaleway all have a snapshot create). On LVM:

sudo lvcreate -s -L 5G -n root-prekernel /dev/vg0/root
# Rollback later: sudo lvconvert --merge /dev/vg0/root-prekernel  (takes effect after reboot)

No LVM and no snapshot API? Then at least a restic backup /etc /boot and the certainty that your last full backup is from today. See verifying backups.

Step 3: install, but don't reboot yet

sudo apt update
sudo apt install -y linux-image-generic linux-headers-generic     # Ubuntu
# sudo apt install -y linux-image-amd64 linux-headers-amd64        # Debian
# Or, if unattended-upgrades already fetched it: nothing to do.

# Check: is the new kernel in /boot AND in GRUB?
ls -1 /boot/vmlinuz-*
grep -oE "menuentry '[^']+'" /boot/grub/grub.cfg | head -5
# Are DKMS modules built for the new kernel?
dkms status 2>/dev/null | grep -v "$(uname -r)"    # must say "installed" for the new version

If dkms status shows added or built without installed for the new kernel, do not reboot. sudo dkms autoinstall -k <new-version> and look at the error.

Step 4: the fallback — one failed boot and GRUB picks the old kernel itself

This is the step that makes the difference between "a bit tense" and "driving to the datacenter in the middle of the night". GRUB can remember whether a boot succeeded, and fall back otherwise.

# 1. Let GRUB remember the last-chosen entry and show a short menu on the console
sudo tee /etc/default/grub.d/90-fallback.cfg >/dev/null <<'EOF'
GRUB_DEFAULT=saved
GRUB_SAVEDEFAULT=false
GRUB_TIMEOUT=5
GRUB_TIMEOUT_STYLE=menu
GRUB_RECORDFAIL_TIMEOUT=30
EOF
sudo update-grub

# 2. Set the OLD kernel as default, and boot the new one ONCE
OLD=$(uname -r)
NEW=$(ls -1 /boot/vmlinuz-* | sed 's|/boot/vmlinuz-||' | sort -V | tail -1)
echo "old=$OLD new=$NEW"
sudo grub-set-default "Advanced options for Ubuntu>Ubuntu, with Linux $OLD"
sudo grub-reboot      "Advanced options for Ubuntu>Ubuntu, with Linux $NEW"
# Debian: replace "Ubuntu" with "Debian GNU/Linux"

What happens now: the next boot uses the new kernel (grub-reboot is one-shot). If the new kernel doesn't boot (kernel panic, hanging driver), the old kernel is the default again at the following power cycle — without you having to do anything. If it does boot, you make it permanent in step 6.

Get the exact menu names from grep -oE "menuentry '[^']+'" /boot/grub/grub.cfg; they differ per distribution and per kernel flavour (-generic, -cloud-amd64).

Step 5: reboot, in the window

# Clean stop of databases first — don't rely on the shutdown timeout
sudo systemctl stop postgresql mariadb 2>/dev/null
sudo needrestart -b | grep KSTA          # 3 = reboot needed, so we're doing the right thing
sudo reboot

On a second screen, start a ping and an ssh attempt in a loop. A VPS is usually back within 30–90 seconds; a bare-metal server with lots of RAM and a POST can take five minutes. Only after ten minutes without SSH do you go to the console.

Step 6: verify, then make it permanent

uname -r                                          # the NEW version?
systemctl --failed                                 # nothing
journalctl -b -p err --no-pager | head -30         # no driver errors
ip -br a; ip route | head -3                       # network as expected
lsmod | wc -l; dkms status 2>/dev/null             # modules loaded
df -h | grep -vE 'tmpfs|udev'                      # all mounts back (NFS! iSCSI!)
systemctl is-active postgresql mariadb docker 2>/dev/null
curl -fsS -o /dev/null -w '%{http_code}\n' http://localhost/healthz   # the application itself

# All good? Make the new kernel the default.
sudo grub-set-default "Advanced options for Ubuntu>Ubuntu, with Linux $(uname -r)"
# Or back to the plain first entry:
# sudo sed -i 's/^GRUB_DEFAULT=saved/GRUB_DEFAULT=0/' /etc/default/grub.d/90-fallback.cfg && sudo update-grub

Log the result with a date — that's the evidence for the patch lead time an auditor asks about:

echo "$(date -Is) $(hostname -s) kernel $OLD -> $(uname -r) OK, downtime $(( $(date +%s) - $(stat -c %Y /var/log/kernel-update.start) ))s" | sudo tee -a /var/log/kernel-updates.log

(Run sudo touch /var/log/kernel-update.start right before the reboot in step 5.)

Rollback: if step 6 is NOT good

The server came up, but something doesn't work (a network card, a storage driver, an application relying on a kernel feature):

# 1. Boot back to the old kernel
sudo grub-reboot "Advanced options for Ubuntu>Ubuntu, with Linux $OLD" && sudo reboot
# 2. After the reboot: freeze the kernel until you know what was wrong
sudo apt-mark hold linux-image-generic linux-headers-generic
# 3. Only remove the broken kernel once the cause is known
# sudo apt remove linux-image-$NEW linux-modules-$NEW

And the LVM snapshot from step 2 you merge only if files are broken (a half-done dpkg run), not for a kernel you simply don't like.

When livepatch is worth it

Livepatch (Ubuntu Pro, free up to five machines; kpatch on RHEL) patches a subset of critical kernel CVEs in memory, without a reboot. It's not a replacement for this procedure — needrestart keeps saying "reboot needed" — but it buys time: Monday's CVE is closed on Tuesday, and you plan the reboot in next week's window.

The trade-off: if you have a monthly maintenance window in which a reboot fits anyway, livepatch adds little. If you have a database or a customer SLA that allows only two reboots a year, it's worth it. Debian has no official equivalent; there, a planned window is the practice.

Pitfalls

  • GRUB_TIMEOUT=0 and no console. Then there's no way at all to intervene on a hanging kernel. The fallback from step 4 solves that, but only if you set it up before the update.
  • NFS and iSCSI mounts that don't come back. After a reboot with a new kernel a storage module may be missing or a mount may time out. df -h in step 6 is there precisely for that; an application starting with an empty data dir is worse than one that doesn't start.
  • Cloud kernels and the metapackage. On AWS/Azure/GCP images the metapackage is called linux-image-aws (etc.). Installing linux-image-generic on such an image gives you a kernel without the cloud drivers.
  • apt autoremove removing the running kernel. It can't — apt protects the running kernel. But it can remove the old kernel you wanted as fallback. Run autoremove after step 6, not before.
  • Two updates at once. Kernel + glibc + systemd in the same reboot: if something breaks, you don't know what. When in doubt, the rest first, then the kernel.
  • Reboots outside monitoring's knowledge. A reboot at 03:00 lasting two minutes generates "host down" alerts. Set a maintenance window in your monitoring, or accept the alert and document it.

What you still don't have

  • Order across the fleet. Rebooting fourteen servers in the right order (replicas before primaries, load balancers last, never two nodes of the same pair at once) is a runbook someone executes by hand.
  • Approval and separation of roles. Who decided this kernel went to production today, and who executed it? For a NIS2 or ISO audit those must demonstrably be two steps, with dates.
  • Correlation with what happens afterwards. The reboot at 03:00 and the slow database at 08:00 are related, but that link sits in two logs.

How monsys does it

In monsys you plan kernel updates as a batch over a group of hosts: the hub knows per host the running and installed kernel, reboot_required, the failover relations (never both nodes of a pair at once) and the group's maintenance window. A batch proposed by the AI assistant or an MCP client must be approved by a different person (segregation of duty, enforced in the database), and every step — proposal, approval, execution per host, result — is a signed line in the transparency log. Execution itself goes via Emergency Action Tokens with UpdateKernel and Reboot as actions; a host that doesn't come back after the reboot pauses the batch.

FAQ

How long does a kernel reboot take?

On a VPS 30–90 seconds between reboot and working SSH. On bare metal with lots of memory and a BIOS POST up to five minutes. Budget fifteen minutes in your window including verification; that you usually only need three is a bonus.

Can I roll back a kernel update?

Yes, as long as the old kernel is still in /boot: pick it in the GRUB menu or with grub-reboot, and set it as default again with grub-set-default. That's why you always keep at least two kernels and only run apt autoremove after successful verification.

Do I have to reboot for every kernel update?

To run the new kernel: yes. Whether it's urgent depends on the CVEs the update fixes and whether they apply to your configuration. Check that first; plan the rest in your regular maintenance window.

Written by the monsys team — sysadmins who do this every day.

Done it by hand? Let monsys keep it running.

Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.