Patching the Host Kernel Under Running MicroVMs
There is one upgrade in my building with no rollback, no canary that means anything on its own, and no version of "just restart it" a customer cannot detect. Not the API. Not the database. The Linux kernel on the hosts, underneath the hypervisor, underneath other people's live virtual machines. The host kernel is not a component on the machine. It is the machine.
I'm Ajay; I build PandaStack, a Firecracker microVM platform. Each Linux host runs a Go agent owning live guests: customer sandboxes, hosted apps, managed Postgres instances with real data in them. I can upgrade a guest kernel any time by re-baking a template. The host kernel is the one I cannot. Two sibling posts cover the neighbours — How to Drain a Host Without Dropping Everything on It is the drain procedure, Upgrading the Daemon That Owns Your Running VMs is replacing the agent binary without the init system murdering its children. The agent is a process you can teach to survive its own restart. The kernel is not.
The thing you are patching is the thing you sell
The entire security argument for a microVM platform is a statement about kernels. A container shares the host kernel, so a kernel privilege escalation reachable from inside one is an escape: the isolation was namespaces and a seccomp filter standing between untrusted code and several million lines of shared C. A microVM gives each guest its own kernel, and the host kernel is reached only through a narrow hypervisor interface guarded by the VMM's seccomp filter and jailer.
That difference is real, and it comes with maintenance attached. "Small, well-guarded surface" describes the interface, not the code behind it: KVM's ioctl surface and its instruction and MSR emulation, the virtio and vhost code parsing descriptors a guest wrote into a shared ring, the host network stack receiving whatever a guest emits through its tap, the block path under each guest's disk. All of it is host kernel code reachable, directly or at one remove, by code a stranger pasted into a sandbox.
So, plainly: a long host-kernel patch latency is a quiet, compounding form of exactly the risk the product claims to remove. It does not page anyone and it is not an error rate — it is a number that grows while everything looks fine. The failure mode is symmetric: treat every advisory as an emergency and the organisation learns that kernel work is a disaster and defers everything, and treat none as emergencies and you carry a six-month latency on the one that was reachable from a guest virtqueue.
Triage first: is this reachable from a guest?
Most host-kernel CVEs are not reachable from a guest on a well-configured hypervisor host. That is not complacency; it follows from what such a host does. Few filesystems, few drivers, few protocols, and no untrusted code in the host's own kernel context. So the first job is not to schedule a reboot. It is to work these in order and stop as soon as one says no.
- Does the vulnerable code exist in our kernel at all? Check the running config, not the distro default, then whether the module is loaded and whether anything could autoload it.
- What attacker position does exploitation need: local unprivileged code on the host, code inside a guest, a remote peer, or physical access? The first and last are usually not your threat model; the second is all of it.
- If it needs code in a guest, is there a path? Only a few channels exist: virtqueues the host parses, the KVM ioctls and emulation paths the VMM drives on the guest's behalf, packets emitted through the tap, vsock, and whatever the block device bottoms out in. A bug in a channel you do not expose is not reachable.
- Does the VMM's own seccomp filter block the syscall or ioctl needed to get there? That filter exists so "the guest made the VMM do something" has a short list of somethings.
- What is the impact: host code execution and escape, host denial of service that threatens the neighbours' availability, or memory disclosure? Three categories, three clocks.
- Is there a mitigation that is not a reboot — blacklisting a module, a sysctl, dropping a device model? One that closes the path turns an emergency into scheduled maintenance.
- Does the fix need microcode or firmware? Then livepatch and kexec are both wrong and you are cold-booting regardless.
- Is it a speculative-execution issue? Different shape: a kernel parameter plus microcode, a real performance cost, and a product decision as much as a security one.
Question one is mechanical, so make it a script. Question two carries an assumption worth writing on a wall: the moment anything untrusted executes in the host's own kernel context, every local privilege escalation becomes a guest escape. A build step running a customer's script on the host rather than inside a guest, a pipeline unpacking customer archives as root — each reclassifies a whole category of CVE from "not our threat model" to "ship it tonight".
The four options, graded honestly
Once triage says the patch matters there are exactly four moves. Three are techniques and one is a decision — a legitimate move rather than a failure to make one, provided it is written down.
| Strategy | Guest impact | Patch coverage | Time to apply | Residual risk | Operational cost |
|---|---|---|---|---|---|
| Drain and reboot | Every guest pauses and resumes elsewhere, or ends; connections drop | Everything: kernel, modules, microcode, firmware | The drain dominates, set by your slowest workload class | None from the patch; all of it lives in the evacuation | Spare capacity, a practised drain, a window per host |
| Live kernel patching | None, once the transition completes | Partial: function-body fixes only, no data structure or inlined code | Minutes per host, nothing to reschedule | Partial application, unconverted tasks, an unreleased kernel | A second version dimension, and a reboot you still owe |
| kexec fast reboot | Every guest still dies; shorter window, not no window | Kernel and modules; skips firmware, so no microcode | Cold boot minus firmware and bootloader time | A failed kexec means cold-booting anyway | Getting kernel, initrd and cmdline right every time |
| Deliberate deferral | None | None — the bug is still there | Immediate | The whole CVE while you defer, plus any triage error | A written risk acceptance, owner and review date |
Option 1: drain and reboot
The default, the only universally correct option, and the expensive one. Correct because its coverage is everything, including the microcode the firmware loads on cold boot. Expensive because the price is every running customer workload on the box, paid once per host.
What makes draining tolerable on a snapshot platform is that there is no warm pool of idle VMs to preserve. Every sandbox create here is a restore of a baked Firecracker snapshot, so snapshot-and-restore is not a path bolted on for drains — it is the ordinary boot path, exercised on every create, p50 179 ms and p99 203 ms. A running sandbox can be snapshotted at any point and its memory and disk become bytes in object storage, so "evacuate" means "hibernate and restore elsewhere", not "live-migrate". Restoring a machine's worth of captured state on a different host is the 1.2 to 3.5 second shape, since the far side must pull the artifacts first. The floor is higher for classes that do not restore: a template with no baked snapshot cold-boots in about 3 seconds, and a managed Postgres instance takes 30 to 90 seconds to reach ready.
A snapshot captures guest memory and guest disk. That is a lot of state and it is not all of it. Three things do not come along:
- In-flight TCP connections. The guest resumes with sockets it believes are established, whose peers gave up seconds ago. Its view of the connection is restored perfectly; the connection is not. Snapshot, Restore, and the Connections You Left Open is the full treatment — connection-oriented work must reconnect rather than resume.
- The clock. A resumed guest believes no time has passed: monotonic frozen through the pause, wall clock stuck at snapshot time. Harmless for most code and quietly catastrophic for TLS validity windows, token expiry and lease arithmetic. Resync from inside the guest on restore, or you get a VM that is alive and confidently wrong about the year.
- Anything holding a host-local resource — the tap device, the netns, the loopback-backed disk, the vsock path. Each sandbox here gets a netns from a pool of 16,384 pre-allocated /30 subnets, so a restore is also an allocation, and a drain that leaks them returns a host with a shrinking address pool and no error about it.
Why this rather than live migration: Snapshot-Restore vs Live Migration: Not the Same Problem argues it out. Live migration buys a cutover nobody perceives, and costs dirty-page tracking plus a network fabric assumption I have not taken on. What I have is a pause measured in seconds. For a kernel patch that is acceptable — it is just not zero.
Option 2: live kernel patching
The genuinely clever option, and worth describing as a mechanism rather than magic, because the mechanism is where the limits come from. The lineage is kpatch and kGraft, merged into mainline as the livepatch subsystem. A livepatch is a kernel module registering replacement functions, and the redirection is done with ftrace: the patched function's entry already has a hook site, and the livepatch installs a handler sending callers to the new implementation. Hence only function-body changes are patchable.
The hard half is consistency. At the instant you patch, tasks are inside the old versions of those functions on stacks with old assumptions. Mainline handles it with a per-task transition: every task is flagged pending, and each converts individually once the kernel can show, by walking its stack, that it is not executing a patched function. Tasks sleeping in user space convert at the kernel/user boundary. Until all have converted the kernel really is running both versions, per task.
Two limits fall out of that, and both are structural rather than tooling gaps. Data structure changes — layout, lifetime, the semantics of a field — cannot be expressed as a function replacement, because existing instances were allocated by the old code and nothing migrates them. Code inlined into its callers needs every caller patched. Functions effectively always on somebody's stack, like parts of the scheduler and idle paths, can never satisfy the consistency check. So a vendor's stream covers many high-severity fixes and never all, and the misses correlate with deep structural bugs.
The other limit is a task that refuses to convert — a kernel thread parked in a long loop, something blocked deep in a path containing a patched function. The patch then sits in transition indefinitely: not a crash, not a success, but a kernel where the fix applies to most of the system and not to that task. There is a mechanism to force it, and it is the loaded gun it sounds like, since forcing asserts the check you could not satisfy did not matter.
Then the governance reality. After a livepatch you are running a kernel corresponding to no released source tree, with an unchanged version string, so your inventory has two dimensions and most tooling knows one. A host that reboots for an unrelated reason comes back unpatched unless the module loads at boot. Where livepatch earns its place is as a scheduling instrument: it turns "empty every host this week" into "the exposure is closed today, we reboot on the normal cadence".
Option 3: kexec, for a shorter window rather than no window
kexec boots a new kernel directly from the running one. You load an image, initrd and command line with `kexec -l`, then `kexec -e` jumps into it, skipping firmware, self-tests and the bootloader. On a server whose firmware takes a geological interval to enumerate disks that is a real saving, and `systemctl kexec` wires it into the ordinary shutdown sequence.
Be absolutely clear: kexec destroys every guest on the host. The new kernel starts with its own view of memory, and a Firecracker process from the old kernel's world does not survive in any useful sense. kexec shrinks the reboot window; it does not avoid the reboot. Two further traps. Skipping firmware also skips the microcode the firmware would load, so for a fix that includes microcode kexec leaves you believing you patched something you did not. And wrong initrd, a lost command-line parameter, or a kernel that cannot find its root device puts you at the console cold-booting anyway, with every guest already gone.
One research direction, flagged as a direction and nothing more: there is ongoing upstream work on preserving memory regions across a kexec, handing specified memory from the old kernel to the new one. If it matures, hibernating guests into a region that survives the handover would turn a cross-host evacuation into an in-place pause. I do not plan around it; check the current state of that work yourself.
Option 4: do nothing yet, deliberately
Most fleets defer constantly without admitting it. The difference between a defensible deferral and an embarrassing one is whether it was written down with the reasoning attached. A deferral with a record is a risk acceptance; one without is patch latency wearing a lanyard. The record fits on half a page:
- The advisory identifier and the specific subsystem, not the headline. "A bug in the network stack" is not a record; the function and the path are.
- Our configuration state — compiled in, module loaded, autoloadable — and the exact command that established it, so it can be re-run against a kernel we have since changed.
- The attacker position exploitation needs, the architectural reason no such position exists here, and the thing that would invalidate that reason. "This stops being true if we ever run customer build steps outside a guest" is a sentence that has saved me twice.
- Mitigations in place now, and whether they are load-bearing for this decision.
- The decision, a named owner, and a review date — the part everyone skips and the only part that makes this different from forgetting.
- What would escalate it: a public proof-of-concept, reports of exploitation, a vendor reclassification, or discovering the triage was wrong about reachability.
Inspecting livepatch state on a host
If you livepatch, you need to answer "is this host patched, right now" from the host rather than from a vendor dashboard. The kernel exposes its own state under sysfs, and that is the authoritative answer.
#!/usr/bin/env bash
# livepatch-state.sh -- what the KERNEL thinks is patched, not what the
# vendor dashboard thinks. Run it on the host; read it with your own eyes.
set -uo pipefail
echo "== base kernel =="
# NOTE: a livepatch does NOT change this. The base image is unchanged, so
# uname is necessary and nowhere near sufficient.
uname -r
echo "== is livepatching even compiled in? =="
# No sysfs directory means either no patch is loaded or the kernel has no
# livepatch support at all. Those are very different answers.
grep -E 'CONFIG_LIVEPATCH|CONFIG_HAVE_RELIABLE_STACKTRACE' \\
"/boot/config-$(uname -r)" 2>/dev/null || echo "config not readable here"
test -d /sys/kernel/livepatch && echo "sysfs: present" || echo "sysfs: absent"
echo "== loaded patches and their state =="
# enabled = the patch is applied (or applying)
# transition = 1 means tasks are still being converted: the fix is live for
# some tasks and not others. A transition that never ends is a
# task that will not convert -- plan a reboot, do not force it.
for p in /sys/kernel/livepatch/*; do
test -d "$p" || continue
echo "patch: $(basename "$p")"
for f in enabled transition force; do
test -r "$p/$f" && echo " $f = $(cat "$p/$f")"
done
# Which functions this patch actually redirects -- useful when an advisory
# names a function and you want to know whether your patch covers it.
find "$p" -mindepth 3 -maxdepth 3 -type d -printf ' func: %f\\n' 2>/dev/null
done
echo "== the modules behind them =="
lsmod | grep -iE 'livepatch|kpatch' || echo "no livepatch modules loaded"
echo "== unconverted tasks, if the kernel will tell you =="
# Some kernels log the offending task on a stalled transition. Cheapest way
# to find out which thread is holding the patch open.
dmesg | grep -iE 'livepatch' | tail -20 || trueTreat those paths as a starting point, not gospel: distributions differ in where the config lives and whether it is readable, vendor products ship their own CLI with their own notion of state, and some hardened kernels restrict `dmesg`. Make the script say "I could not determine this" rather than printing nothing — a check that cannot tell "patched" from "unreadable" is worse than no check.
Fleet mechanics: one host at a time, and capacity is the binding constraint
Rolling a kernel across a fleet is not a deploy in the web sense, where a bad version affects requests in flight and then stops mattering. A bad kernel affects every guest on every host you have already done, for as long as those guests live. The pace is set by how long a subtle regression takes to surface: hours of production load, not minutes of smoke tests.
The real constraint is capacity. You cannot drain a host into a full fleet. Every guest you evacuate must be admitted somewhere else, and evacuation is a burst: a host's whole population of creates and wakes arrives at the remaining hosts within minutes. How Many Sandboxes Fit on One Machine? is the headroom arithmetic; Saying No Correctly: Admission Control for Sandbox Fleets is what should happen when a restore has nowhere to land — a retryable refusal and a queue, never a silent drop.
Cordoning deserves a note, because on this platform there is no cordon API and I suspect that is the common case. My scheduler reads the hosts' own self-reported capacity: each agent heartbeats every 10 seconds, agents older than 30 seconds are excluded from placement entirely, ownership is lease-based, and among candidates passing every hard capacity gate, placement is a deterministic score of `0.6 * free_cpu + 0.3 * free_mem_gb`. So "cordon" is implemented by making the host stop advertising capacity. What you must not do is cordon by stopping the heartbeat: silence is how a host says it is dead, the lease expires, and something fleet-wide starts deciding what to do about workloads it thinks are stranded. Make draining durable, too — the flag that bit me lived only in memory, so the agent came back advertising full capacity and the scheduler refilled the host I was about to reboot.
Then the ordering trap. The host is drained and empty, you are already in a window, the agent has a version bump waiting, and there is a config change you have been meaning to apply. Doing all three is obviously efficient, and it is also how you spend the next day unable to answer the only question that matters: which change broke the box? One change per reboot.
The runbook: cordon, evacuate, verify, patch, canary, return
This is the shape of what I run. The per-step gates are the point, because the two ways this goes badly are rebooting a host that still has guests on it and returning a host to service before it works.
#!/usr/bin/env bash
# patch-host-kernel.sh -- drain, patch, verify, return ONE host.
#
# Never run this in a loop over the fleet: the whole point of the cadence is
# that a regression gets time to show up before the next host goes. A
# non-zero exit means the host is cordoned and NOT back in service, which is
# the safe resting state.
set -euo pipefail
HOST="${1:?usage: patch-host-kernel.sh <agent-id>}"
AGENT="http://$HOST:7070"
API="${PANDASTACK_API:-https://api.pandastack.ai}"
abort() { echo "ABORT: $*" >&2; echo "host stays cordoned" >&2; exit 1; }
# ---------------------------------------------------------------- 1. cordon
# Durable flag in the control plane AND a direct call to the host, because
# the scheduler caches its capacity view and a database row alone is not a
# cordon until every cache in front of it agrees.
echo "== cordon =="
psql "$PGURL" -v ON_ERROR_STOP=1 -c \\
"UPDATE agents SET draining = true, drain_reason = 'kernel-patch' WHERE id = '$HOST';"
curl -fsS -X POST "$AGENT/v1/admin/cordon" \\
-H "X-Node-Token: $PANDASTACK_NODE_TOKEN" \\
-d '{"draining":true,"reason":"kernel-patch"}' \\
|| abort "host would not accept the cordon; fix that before draining"
sleep 35 # outlast the scheduler's capacity cache before counting anything
# -------------------------------------------------------------- 2. evacuate
# Hibernate what can pause (memory + disk go to object storage and restore
# elsewhere on wake), fail databases over, let short TTLs expire. This step
# is bounded by your slowest workload class, not by the kernel. Give it a
# deadline you are actually willing to reach.
echo "== evacuate =="
pandastack admin drain "$HOST" --deadline 30m --classes hibernate,failover,expire \\
|| echo "drain returned non-zero; the gate below decides, not this"
# ---------------------------------------------- 3. verify empty (the gate)
# Reconcile BOTH views. The control plane holds a claim; the host holds the
# fact. A process that outlived its row is the dangerous direction: nothing
# is watching it and it appears on no dashboard.
echo "== verify empty =="
ROWS=$(psql "$PGURL" -tAc \\
"SELECT count(*) FROM sandboxes WHERE agent_id='$HOST' AND status NOT IN ('deleted','stopped');")
PROCS=$(ssh "$HOST" 'pgrep -c firecracker || true')
echo "control plane rows: $ROWS firecracker processes: $PROCS"
[ "$ROWS" = "0" ] || abort "$ROWS guest rows still assigned to this host"
[ "$PROCS" = "0" ] || abort "$PROCS firecracker processes still running -- do NOT reboot"
# ----------------------------------------------------------------- 4. patch
# One change per reboot. Kernel now; the agent binary and any config change
# wait for the next pass, so a regression has exactly one candidate.
echo "== patch =="
OLD_KERNEL=$(ssh "$HOST" 'uname -r')
ssh "$HOST" 'sudo apt-get update -qq && sudo apt-get install -y --only-upgrade linux-image-generic'
ssh "$HOST" 'sudo systemctl reboot' || true
for i in $(seq 1 60); do
sleep 10
ssh -o ConnectTimeout=5 "$HOST" true 2>/dev/null && break
[ "$i" = "60" ] && abort "host did not come back after reboot"
done
# ---------------------------------------------------------------- 5. verify
echo "== verify host =="
NEW_KERNEL=$(ssh "$HOST" 'uname -r')
echo "kernel: $OLD_KERNEL -> $NEW_KERNEL"
[ "$NEW_KERNEL" != "$OLD_KERNEL" ] || abort "kernel unchanged -- the upgrade did not take"
ssh "$HOST" 'test -c /dev/kvm' || abort "/dev/kvm missing -- no virtualisation here"
ssh "$HOST" 'lsmod | grep -qE "^kvm_(intel|amd)"' || abort "kvm module did not load"
# Sysctls provisioning set are not automatically true again after a reboot.
# This one matters: hugepage-ness is a SNAPSHOT property, and such snapshots
# restore only through the UFFD backend -- a missing overcommit pool does not
# fail at boot, it fails later, as restores, on a host you called healthy.
ssh "$HOST" 'sysctl -n vm.nr_overcommit_hugepages'
ssh "$HOST" 'systemctl is-active pandastack-agent' | grep -qx active \\
|| abort "agent did not come back"
# ----------------------------------------------------- 6. canary, per host
# One real create on THIS host before it is allowed back into placement.
# Pin it, or the scheduler will helpfully place it somewhere healthy and
# hide the exact problem you are looking for.
echo "== canary =="
SBX=$(curl -fsS -X POST "$API/v1/sandboxes" \\
-H "Authorization: Bearer $PANDASTACK_API_KEY" \\
-H "Content-Type: application/json" \\
-H "X-Pandastack-Pin-Agent: $HOST" \\
-d '{"template":"base","ttl_seconds":300,"metadata":{"purpose":"kernel-canary"}}' \\
| python3 -c 'import json,sys; print(json.load(sys.stdin)["id"])') \\
|| abort "canary create failed on freshly patched host"
# Prove the guest runs code, not just that a row appeared. This prints the
# GUEST kernel -- a useful reminder that what you patched is not what the
# guest sees.
curl -fsS -X POST "$API/v1/sandboxes/$SBX/exec" \\
-H "Authorization: Bearer $PANDASTACK_API_KEY" \\
-H "Content-Type: application/json" \\
-d '{"cmd":"uname -r"}' \\
|| abort "canary exec failed -- guest booted but cannot run commands"
curl -fsS -X DELETE "$API/v1/sandboxes/$SBX" \\
-H "Authorization: Bearer $PANDASTACK_API_KEY" >/dev/null
# -------------------------------------------------------------- 7. uncordon
echo "== uncordon =="
psql "$PGURL" -v ON_ERROR_STOP=1 -c \\
"UPDATE agents SET draining = false, drain_reason = NULL WHERE id = '$HOST';"
curl -fsS -X POST "$AGENT/v1/admin/cordon" \\
-H "X-Node-Token: $PANDASTACK_NODE_TOKEN" -d '{"draining":false}'
echo "host $HOST in service on $NEW_KERNEL. Now WAIT. Do not start the next"
echo "host until this one has carried real load long enough for a regression"
echo "to have shown itself."Three details carry the weight. The verify-empty gate checks the control plane and the host's process list and aborts on either being non-zero, because a Firecracker process that outlived its row is the dangerous direction: nothing is watching it, and rebooting kills a guest no system you own believes exists. The canary is pinned, because an unpinned create goes wherever the score is best — precisely not the host you need to test.
Verification, and the host you leave behind on purpose
The checks are not interchangeable. `uname -r` proves the boot took and nothing else, and it proves nothing about livepatch, since a livepatch leaves the version string untouched. Confirming the KVM modules loaded and `/dev/kvm` exists sounds trivial, and it is exactly what a kernel upgrade breaks: a module not rebuilt, a lockdown policy that now refuses it, a device node whose permissions came from a rule that no longer matches. That failure is invisible unless you look, because the agent comes up, heartbeats happily, advertises capacity, and refuses every create.
Then the sneaky category: things your provisioning asserted that a reboot does not restore. Sysctls are the classic, and here the hugepage overcommit pool is the one that matters, because hugepage-ness is a property of a snapshot rather than of a running VM, and such snapshots restore only through the UFFD memory backend. A host that comes back without its pool does not fail at boot; it fails minutes later as restore errors, on a host you already called healthy.
And the last one, which costs a little capacity and buys the ability to reason: leave one host on the old kernel on purpose. A deliberate control, not an accident. Then when something gets slower next week you have a comparison rather than a correlation.
The honest limits
No option here makes a host-kernel patch free. Drain-and-reboot pauses every guest; livepatch covers a subset and leaves you a reboot you still owe; kexec shortens the window; deferral keeps the vulnerability. And the deeper limit cannot be engineered away: you are trusting a kernel. The honest microVM claim is not "there is no shared kernel surface" — it is "the shared surface is far smaller and far more scrutinised, and we patch it on a schedule we can show you". That second clause is operational work, and it is the part customers should ask about.
The summary
Triage before you plan: does the vulnerable code exist in your kernel, what attacker position does it need, and does a guest have a path to it through a virtqueue, a KVM ioctl, a tap device or vsock? Reachability sets the clock, not the severity score. Drain and reboot is the only strategy with complete coverage, and it is tolerable on a snapshot platform because evacuation is hibernate-and-restore over paths you exercise on every create. Livepatch buys a scheduling window and will never cover every fix. kexec shrinks the window, still kills the guests, and skips microcode. Deferral is legitimate with a written risk acceptance and a review date, and is otherwise just latency. Then the mechanics: one host at a time with real waiting between them, free capacity for the evacuation to land in, cordon as "stop advertising capacity" rather than "stop heartbeating", one change per reboot, a canary create before each host returns to placement, and one host left on the old kernel on purpose.
Frequently asked questions
Is a host-kernel CVE reachable from a guest?
Usually not, and establishing which case you are in is the whole job. A large share of kernel vulnerabilities require local unprivileged code on the host, and a properly built hypervisor host has no untrusted local user — the untrusted code sits on the far side of a VM boundary, inside its own kernel, with the VMM's seccomp filter between it and the host. Those you can schedule normally with a written note. The reachable ones live in a short list of channels: the virtqueues a guest writes and the host parses, including any in-kernel vhost backend you have enabled; KVM's ioctl surface and its instruction, MSR and page-table emulation paths; the host network stack receiving packets the guest emits through its tap device; and vsock. A bug in any of those is a guest-to-host path and gets emergency treatment. Work the questions in order: is the code compiled in, is the module loaded, what position does exploitation need, does a guest have a path, and is the impact escape, denial of service, or disclosure. One caveat: all of that depends on nothing untrusted ever executing on the host itself. A build step that runs a customer's script outside a guest reclassifies every local privilege escalation into a guest escape.
Can I livepatch my way out of ever rebooting?
No, and the reasons are structural rather than a tooling gap you can wait out. A livepatch replaces function entry points, installing an ftrace-based redirect so callers reach the new implementation, so the fixes it can express are the ones that fit inside a function body. A change to a data structure's layout or semantics cannot be applied this way, because existing instances were allocated by the old code and nothing migrates them. Code inlined into its callers may not be patchable at all. Functions effectively always on somebody's stack, like parts of the scheduler and idle paths, can never satisfy the consistency check, which converts each task individually only once the kernel can show by walking its stack that it is not inside a patched function. A task that will not sleep somewhere convertible leaves the patch stuck in transition, and the mechanism to force it through is a genuine loaded gun. There is a governance cost too: a kernel matching no released source tree, an unchanged version string most inventory tooling cannot see past, and an unplanned reboot that reverts you unless the module loads at boot.
How much spare capacity does a fleet-wide kernel roll need?
At minimum one host's worth of genuinely free capacity, and realistically more, because an evacuation is a burst rather than a trickle. When you drain a host its entire population has to be admitted elsewhere inside the drain window: every pausable guest hibernated and restored on another host, every stateless app redeployed, every database failed over from its archive. That arrives at the remaining hosts as a concentrated spike of creates and wakes. If the fleet is already near its admission limits, the first host drains fine and the second finds nowhere to put anything. Size the headroom against your largest host rather than the average, because the constraint is the worst case you might have to empty. And make sure the path when a restore cannot be placed is a retryable refusal with backpressure rather than a dropped workload.
Does kexec let me patch the kernel without the guests noticing?
No. kexec boots a new kernel directly from the running one, skipping firmware, self-tests and the bootloader, which on server hardware is a real saving — but every guest on the host dies in the transition. The new kernel starts with its own view of memory, and a hypervisor process from the old kernel's world does not survive it in any useful sense. kexec shrinks a drained host's downtime; it does not let you skip the drain. Two traps. Because it skips firmware it also skips the microcode the firmware would load, so for a fix that includes microcode it leaves you believing you patched something you did not. And a kexec is only as reliable as what you loaded: the wrong initrd, a missing command-line parameter, or a kernel that cannot find its root device, and you are cold-booting anyway with every guest already gone.
Should I patch the host kernel and the host agent in the same window?
No, even though it is obviously efficient and the temptation peaks exactly when you are already in the window with an empty host. The problem is diagnostic rather than technical: change the kernel and the daemon in one pass, and when something behaves badly afterwards — restores getting slower, a device model misbehaving, guests failing to boot on one host and not another — you cannot tell which change caused it without undoing both and re-testing. One change per reboot means a regression has exactly one candidate. There is a second reason: these are different hazards sharing a failure mechanism. The agent restart kills guests if the daemon's children inherit its cgroup and the init system signals the whole control group on stop, or if the shutdown path hangs past its stop timeout. The kernel reboot kills guests unconditionally. Both are handled by the same sequence: evacuate, verify empty against the actual process list, then touch it.
Keep reading
- How to drain a host without dropping everything on it — The drain procedure this runbook calls out to: cordon, classify by workload class, grace budgets, and verifying empty.
- Upgrading the daemon that owns your running VMs — The other half of the hazard class — why restarting the agent can take the guests with it, and the unit file that survives.
- What snapshot restore does to network connection state — The thing an evacuation does not preserve: sockets the guest still believes in, and peers that gave up seconds ago.
- Snapshot restore vs live migration — The capability that would make a kernel roll invisible, what it actually costs, and why this platform does not have it.
- MicroVM fleet capacity planning — The headroom arithmetic that decides whether a fleet-wide kernel roll is routine or a self-inflicted capacity incident.
Related posts
- Container Escape CVEs, by Class: What the Pattern Tells You
Container escapes aren't a list of unlucky bugs — they're four recurring shapes, plus one category that will never get a CVE because it's working as designed.
- Firecracker boot_args, argument by argument
Everyone copies the same magic `boot_args` string from the Firecracker docs and never reads it. It's a short, unusually honest description of what a microVM is — and what it has decided not to be.
- Guest kernel lockdown and module loading in Firecracker microVMs
Root in the guest is not a breach, it is Tuesday. But what root can then do to that kernel is the first half of every escape chain — and in a snapshot-restore world, the hardening has to be baked in, because nothing runs at boot ever again.
- The I/O Scheduler Inside a MicroVM Is Almost Always Wrong
Your guest's block layer is sorting requests by sector number to optimise seeks on a platter that is a file on somebody else's filesystem. Why `none` is the answer, why it still merges, and why readahead is the knob that actually moves the number.
- Conntrack Table Limits: The Shared Resource Under Your MicroVM Fleet
Your guests have separate kernels, separate namespaces and separate subnets. They share one fixed-size hash table, and one port scanner can fill it for everybody.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.