all posts

KVM Halt Polling Explained: Why an Idle microVM Can Burn a Core Doing Nothing

Ajay Kumar··10 min read

The host was doing nothing, loudly. A fleet agent packed with sleeping sandboxes — every one a `base` microVM with its 8 vCPUs and 4 GiB, every one parked waiting for a request that had not arrived in twenty minutes — and the host's load average was not zero. `top`, sorted by CPU, showed a long column of threads named `fc_vcpu 0`, `fc_vcpu 1`, `fc_vcpu 2`, each burning a steady, unmistakable slice of a core. None of them was executing a guest instruction anyone had asked for.

There was no runaway process. No busy-wait in a customer's code, no cryptominer, no metrics agent gone feral. I checked all three, in that order, because that is the order in which they are usually the answer. The guests genuinely were idle. The host CPU genuinely was being consumed. The explanation was a host kernel module parameter with a boring name that nobody on the team had ever typed: `/sys/module/kvm/parameters/halt_poll_ns`.

KVM halt polling is one of the best latency optimisations in the hypervisor, and one of the worst-behaved features you can run on a host packed with idle guests. It is also almost entirely undiscussed outside the kernel mailing list, which is a shame, because it is a knob that changes your cost per sandbox without changing a line of anybody's code. So: what the guest actually does when it has nothing to do, what KVM does about it, why the fix for latency is a tax on density, and how to decide which side of that trade your host is on.

The short version. When a guest vCPU runs out of work it executes `HLT`, which is a VM exit; the vCPU thread then blocks in the host kernel until an interrupt wakes it. Blocking and waking costs a scheduler round trip, and that cost lands on the next piece of work the guest wanted to do. Halt polling is KVM's bet that the interrupt will arrive soon: instead of blocking immediately, the vCPU thread busy-polls for a short window, and only blocks if the window expires. A poll that succeeds saves the whole round trip. A poll that fails burns host CPU for nothing. On a host running six fat VMs that is a great bet. On a host running a hundred sleepers it is a tax on density and a lie in your CPU accounting.

What a guest does when it has nothing to do

A Linux guest with no runnable tasks ends up in the idle loop, and the idle loop's job is to stop consuming the CPU. On x86 that means executing `HLT`, which halts the logical processor until an interrupt arrives. On bare metal this is a power-management primitive and the story ends there. In a virtual machine it is something else entirely: `HLT` is an instruction the hypervisor traps.

So the guest executes `HLT`, the CPU takes a VM exit with exit reason `HLT`, and control lands in KVM with a question to answer: this vCPU says it has nothing to do, so what should the host thread that is running it do now? The mechanics of an exit — the state save, the transition, the cost of the round trip itself — are in VM exits: the actual currency of virtualization overhead, and they matter here mainly as context. The expensive part is not the exit. The expensive part is what happens after it.

The obvious answer is: block. KVM puts the vCPU thread on a wait queue, calls into the host scheduler, and the thread sleeps. The host CPU is now free for something else, which is exactly what you want from an idle guest. When an interrupt eventually arrives — a virtio completion, a timer, an IPI from another vCPU — KVM wakes the thread, the scheduler picks it, and the guest resumes. Clean, correct, and genuinely cheap when the guest is going to be idle for a while.

It is also not free, and the part that is not free is the part you feel. A block-and-wake is two scheduler operations, a context switch out and a context switch in, a cache and TLB footprint that may not survive the gap, and a wake-up latency that depends on what else is runnable on that CPU at the moment the interrupt lands. Every nanosecond of it is charged to the next thing the guest wanted to do. Where a vCPU thread sits in the host's scheduling hierarchy, and what it competes with, is covered in How Firecracker Schedules vCPUs: The Threading Model.

For a workload that idles in long stretches, none of that matters: you pay one wake-up per minute of sleep. For a workload that idles in short gaps — a network server between packets, a Python interpreter between requests, a database between queries, essentially everything anyone actually runs in a sandbox — you pay it constantly, and the wake-up latency starts to dominate the thing you were trying to measure. That is the problem halt polling exists to solve.

Halt polling: host cycles spent to avoid a scheduler round trip

KVM's answer is to not block immediately. On a halt exit, before touching the wait queue, the vCPU thread enters a tight loop: check whether an interrupt or other wake-up event has become pending, `cpu_relax()`, check the clock, go round again. It keeps doing that for a bounded window. If an event shows up during the window, KVM returns straight into the guest without ever involving the scheduler — no block, no wake, no context switch, and the interrupt is delivered with roughly the latency of noticing it. If the window expires with nothing pending, KVM gives up and blocks exactly as it would have in the first place.

That is the whole mechanism, and the trade it makes is stark once you say it out loud. A successful poll converts a scheduler round trip into a short spin. A failed poll converts nothing into a short spin: you burned host CPU, you still blocked, and you added the poll window to the latency of the eventual wake. Halt polling is the kernel spending host cycles to buy guest latency, and like every such arrangement it works beautifully right up to the point where you multiply the number of accounts.

Two details are worth having precisely, because they change what the feature does to a loaded host. First, the poll loop bails out on `need_resched()` — if the scheduler wants the CPU back, the poll ends immediately. Second, and more interestingly, KVM only continues polling while the vCPU thread is the only runnable task on that CPU. The moment something else becomes runnable there, the poll is abandoned and the thread blocks. Verify both against `virt/kvm/kvm_main.c` on your own host kernel rather than trusting my summary, but the consequence of that pair is important and slightly counter-intuitive: halt polling does not starve other work. It is not a thread hogging a core that someone else needs.

Which means the harm is not contention. The harm is cycles burned, energy spent, and — the part that actually cost us an afternoon — accounting that now reports something untrue. And there is a nasty second-order effect hiding in that same mechanism: on a busy host, polls get cut short by the only-runnable-task check, so they fail more often, so you pay the poll cost and get the block anyway. Halt polling is least effective precisely when the host is most loaded.

The adaptive controller and its four knobs

KVM does not use a fixed poll window, because a fixed window would be wrong for every workload simultaneously. Each vCPU carries its own current window, and a small controller adjusts it based on what just happened: grow the window after a block short enough that polling would have helped, shrink it after a block long enough that polling was hopeless. The per-vCPU window is bounded above by a global ceiling. Four module parameters govern all of it, and all four live under `/sys/module/kvm/parameters/` and are writable at runtime.

  • `halt_poll_ns` — the ceiling, in nanoseconds, for any vCPU's poll window, and the master switch. Writing `0` disables halt polling entirely, host-wide, taking effect at the next halt exit on every vCPU. The x86 default has been 200000 (200 microseconds) for a long while, but defaults differ by architecture and have changed across releases, so read the file rather than quoting a number back at your future self.
  • `halt_poll_ns_grow` — the multiplier applied to a vCPU's window when a poll looks like it was worth growing. Default 2, so the window doubles. Setting it to `0` pins the window and disables growth without disabling polling.
  • `halt_poll_ns_grow_start` — the value the window jumps to when it grows from zero, default 10000 (10 microseconds). A vCPU does not start at the ceiling; it starts small, and climbs by doubling, which is why a freshly restored guest behaves differently from one that has been running for a minute.
  • `halt_poll_ns_shrink` — the divisor applied on a failed poll. This one has a trap in it, below.

The shrink default does not mean what it looks like

`halt_poll_ns_shrink` defaults to `0`, and the natural reading of that is "do not shrink". It is the opposite. The kernel special-cases zero to mean "collapse the window to nothing", so with the default configuration a single block long enough to count as a failure throws away all of a vCPU's accumulated growth and sends it back to square one, where it will climb again from `grow_start` by doubling. Set `halt_poll_ns_shrink=2` and you get a gentle halving instead; set it high and the window falls off a cliff on the first miss.

This is the knob I would reach for first if I wanted to keep polling for bursty guests while limiting the damage from sleepers, because it changes the shape of the controller rather than the size of the bet. It is also the knob with the least written about it, which is a reasonable proxy for how often people actually tune this feature at all.

The other half: polling inside the guest, before HLT

There is a mirror-image mechanism on the guest side, and it is the more elegant of the two. If the guest polls before executing `HLT`, and the interrupt arrives during the guest's own poll, there is no VM exit at all — not a cheap exit, not a trapped-and-returned exit, nothing. The guest never leaves the guest. Linux ships this as the `cpuidle_haltpoll` driver with a `haltpoll` governor, and the governor's own knobs (`guest_halt_poll_ns` and friends) live under `/sys/module/haltpoll/parameters/`, entirely separate from the host's.

The two mechanisms know about each other, which is the detail I find genuinely satisfying. When the guest-side driver comes up, the guest writes a KVM paravirt MSR to tell the host to stop polling on its behalf — because polling in both places means paying for the same bet twice. The driver also does not load on its own initiative: it wants a paravirtualisation hint from the hypervisor's CPUID leaf saying that guest-side polling is appropriate here, and absent that hint it needs a `force` parameter to come up anyway.

Whether any of this is available to you is a guest-kernel build question, not a configuration question. Our guest kernel is 5.10, and what a given 5.10 build exposes depends on whether `HALTPOLL_CPUIDLE` was enabled when it was compiled. Firecracker also normalises the CPUID the guest sees — the machinery is in Firecracker CPUID masking, explained — so the paravirt hint the driver looks for may simply not be presented. Read `/sys/devices/system/cpu/cpuidle/current_driver` in your own guest and believe that, not this post.

One more piece of the same family, for completeness: KVM has a capability a VMM can use to stop trapping `HLT` altogether, so the guest's idle loop halts the physical CPU and nothing exits. That is the configuration for a pinned, dedicated-core, latency-is-everything VM, and it means an idle guest pins a core at full occupancy forever. Firecracker does not do this, and on a density platform you would never want it to — but it is the far end of the same spectrum, and knowing it exists makes the middle of the spectrum easier to reason about.

# What a sandbox can and cannot see about halt polling.
#
# Short answer: a guest sees its own idle driver and its own HLT
# behaviour. It cannot see, and cannot change, the host's halt_poll_ns.
# That asymmetry is the entire operational point of this post -- on a
# managed platform the knob belongs to whoever owns the host.

from pandastack import Sandbox

PROBE = r'''
set -u
C=/sys/devices/system/cpu/cpuidle

echo "idle driver:   $(cat $C/current_driver 2>/dev/null || echo none)"
echo "idle governor: $(cat $C/current_governor 2>/dev/null || echo none)"

# Per-state counters. If the guest-side haltpoll driver is NOT loaded,
# expect a very short list -- the guest's idle loop goes more or less
# straight to HLT, and the host decides what that costs.
for s in /sys/devices/system/cpu/cpu0/cpuidle/state*; do
  [ -d "$s" ] || continue
  printf '  %-12s usage=%-10s time_us=%s\n' \
    "$(cat $s/name)" "$(cat $s/usage)" "$(cat $s/time)"
done

# Is MWAIT even exposed? If not -- and on a Firecracker guest you should
# expect not, because the VMM normalises CPUID -- the guest idle loop
# uses HLT, which is the case this post is about.
if grep -qw mwait /proc/cpuinfo; then
  echo "mwait: exposed"
else
  echo "mwait: not exposed; the guest idle loop will use HLT"
fi

# If the guest-side haltpoll driver IS built and loaded, its knobs live
# here, and they are completely separate from the host's kvm module
# parameters. Verify on your own guest kernel instead of assuming.
ls /sys/module/haltpoll/parameters/ 2>/dev/null \
  || echo "no haltpoll module parameters (driver absent)"

# Proof that the host knob is invisible from in here:
cat /sys/module/kvm/parameters/halt_poll_ns 2>/dev/null \
  || echo "halt_poll_ns: not visible from inside the guest (correct)"
'''

sbx = Sandbox.create(
    template="base",      # 4 GiB / 8 vCPU, baked into the snapshot.
                          # Firecracker cannot change vCPU or RAM at
                          # snapshot restore, so a cpu= or memory_mb= on
                          # create is overridden to the baked values.
    ttl_seconds=900,      # an IDLE timeout, not a wall-clock lifetime
    metadata={"demo": "halt-poll"},
)
try:
    # Write the probe to a file rather than quoting it through a shell:
    # the script contains both kinds of quote and a lot of $.
    sbx.filesystem.write("/tmp/idleprobe.sh", PROBE)

    # timeout(1) in-guest, because a one-shot exec() does not reliably
    # enforce a timeout argument -- bound anything that can hang on the
    # guest side instead of passing timeout_seconds.
    r = sbx.exec("timeout --kill-after=5s 30 sh /tmp/idleprobe.sh")
    print(r.stdout or r.stderr, "exit", r.exit_code)

    # Now park it. Eight idle vCPUs with nothing to do is exactly the
    # condition the host-side measurement above wants to see.
    sbx.exec("sync; pkill -f unattended-upgrade >/dev/null 2>&1 || true")
    print("sandbox parked -- measure its fc_vcpu threads on the host now")

    # And the dense-idle condition, which is where this stops being
    # trivia. Fan out sleepers and watch host CPU climb while every
    # guest does nothing. Mind your own capacity before uncommenting.
    # sleepers = [Sandbox.create(template="base") for _ in range(8)]
    # ... measure, then: for s in sleepers: s.kill()
finally:
    sbx.kill()            # teardown is kill(); there is no .delete()

Why this specifically breaks a dense microVM fleet

Halt polling was designed on and for hosts running a handful of large virtual machines. In that setting the arithmetic is lovely: a few vCPU threads, each with real work arriving frequently, each idle gap short, each successful poll saving a round trip on the critical path of a request somebody is waiting for. The host has spare capacity, the polls mostly hit, and the feature is pure win. It shipped on by default for exactly that reason.

Now change the workload to the one a sandbox platform actually runs. Put a large number of mostly-idle microVMs on one host. Every one of them has multiple vCPUs — our `base` template is 8 — and every one of those vCPUs hits the idle loop constantly, because "mostly idle" means "idling, in a great many short gaps". Each of those idle gaps now opens a poll window. Multiply.

And here is the sharp edge. Remember that KVM only keeps polling while the vCPU thread is the only runnable task on its CPU. On a host full of sleepers, that condition is overwhelmingly true — the pCPUs are mostly idle, which is the entire point of packing sleepers — so every single one of those polls runs to its full window before giving up. The configuration that makes halt polling cheapest in aggregate (a busy host, where polls get cut short) is precisely the configuration a density platform is trying not to be in. A host of sleeping guests gets the maximum possible polling cost in exchange for the minimum possible benefit, because the thing those guests are waiting for is usually not about to arrive.

Your CPU accounting is now reporting fiction

A blocked vCPU thread is in state `S` and consumes nothing. A polling vCPU thread is in state `R`, spinning in the kernel, and therefore: it counts toward the host's load average; it accrues system time in `/proc/<pid>/task/<tid>/stat`; it accrues `usage_usec` in its cgroup's `cpu.stat`; and it shows up in every tool that samples any of those. The guest is asleep. The accounting says it is working. Nothing in the stack is lying — the CPU time is real, it was really spent on behalf of that VM — but the number no longer means what everyone downstream assumes it means.

Three things downstream break on that, in ascending order of how much it will cost you:

  • Capacity dashboards. Load average and host CPU utilisation both read high on a host that is, functionally, empty. You look fuller than you are, and the operator's instinct — stop packing this host — is exactly the wrong call.
  • CPU-based autoscaling. If you scale a fleet on average host CPU, halt polling puts a floor under that metric proportional to how many idle guests you are holding. Scale-to-zero platforms are built on the premise that a sleeping tenant costs nothing; a polling vCPU thread puts a non-zero number under that premise. The placement side of this is in How a Sandbox Scheduler Decides Where Your VM Lands, and it only works if the capacity signal it reads is honest.
  • CPU-seconds billing. This is the one that makes it our problem and not just an interesting kernel fact. PandaStack bills CPU by active CPU-seconds actually burned, rather than by committed vCPU-hours — that is deliberate, it is the dimension that makes an idle sandbox nearly free, and it is documented on pricing. It is also exactly the metric halt polling pollutes. An idle tenant whose vCPU threads are politely spinning in `kvm_vcpu_halt` accrues CPU-seconds for work nobody requested. Charging for that would be indefensible, and the reason it is not a live problem for us is that we made a deliberate decision about the knob, not because the kernel sorted it out on our behalf.

This is the mirror image of CPU steal time, where a guest's own accounting under-reports because time passed that it was not scheduled for — see CPU Steal Time in microVMs: Your VM Isn't Slow, It's Waiting. Halt polling is the host's accounting over-reporting, for the same underlying reason: the thing being measured is a thread, and a thread is a poor proxy for "did useful work happen".

How to see it

The crude check takes thirty seconds and is almost always enough to confirm the diagnosis. Find a guest you are confident is doing nothing — no cron, no agent heartbeat, no metrics scrape — and look at its vCPU threads on the host. If a guest that is provably idle has vCPU threads showing non-trivial CPU, and you have ruled out device emulation and timer interrupts, you are looking at halt polling. The giveaway is that the time is almost entirely system time, which also explains why it is invisible from inside the guest: no guest-side profiler can see cycles that the guest never ran.

The precise check is the per-vCPU statistics. KVM keeps counters for attempted polls, successful polls, nanoseconds spent in successful polls, nanoseconds spent in failed polls, and nanoseconds spent genuinely blocked. Those five numbers tell you everything: the hit rate tells you whether the bet is paying off, and the ratio of failed to successful poll nanoseconds tells you precisely how much host CPU you are donating to the cause.

#!/usr/bin/env bash
# Run this on the HOST, the KVM side. Halt polling is a host-kernel
# feature: nothing inside a guest can see it, read it, or change it.
set -euo pipefail

P=/sys/module/kvm/parameters

# 1. The four knobs. All mode 0644, all writable at runtime, and all
#    GLOBAL to this host's kvm module. There is no per-process scope.
for k in halt_poll_ns halt_poll_ns_grow halt_poll_ns_grow_start \
         halt_poll_ns_shrink; do
  printf '%-24s %s\n' "$k" "$(cat "$P/$k")"
done
# x86 defaults at the time of writing: 200000 / 2 / 10000 / 0. Read the
# files rather than trusting that line -- the defaults differ by
# architecture and have changed across releases.

# 2. Per-vCPU statistics, interface (a): debugfs. One directory per VM
#    (named <pid>-<vmfd>), one subdirectory per vCPU inside it.
for d in /sys/kernel/debug/kvm/*/vcpu*; do
  [ -d "$d" ] || continue
  printf '%s\n' "$d"
  for s in halt_successful_poll halt_attempted_poll halt_poll_invalid \
           halt_wakeup halt_exits halt_poll_success_ns \
           halt_poll_fail_ns halt_wait_ns; do
    [ -r "$d/$s" ] && printf '  %-22s %s\n' "$s" "$(cat "$d/$s")"
  done
done

# 2b. Interface (b): the binary statistics interface (KVM_GET_STATS_FD,
#     Linux 5.18+), which is what a current kvm_stat prefers. If your
#     host is older, only the debugfs path above exists.
kvm_stat --once 2>/dev/null | grep -iE 'halt|exits' || true

# 3. The two ratios that answer "is polling paying for itself here":
#      hit rate      = halt_successful_poll / halt_attempted_poll
#      useful cycles = halt_poll_success_ns
#                      / (halt_poll_success_ns + halt_poll_fail_ns)
#    halt_poll_fail_ns is, by definition, host CPU you burned and got
#    nothing for. halt_wait_ns is the time actually spent blocked, which
#    is the cost polling was trying to avoid. Compare the three.

# 4. Watch the adaptive controller move each vCPU's window. The
#    tracepoint class is kvm_halt_poll_ns, with _grow and _shrink
#    instances; kvm_exit tells you the halt exits are happening at all.
perf list 2>/dev/null | grep -iE 'halt_poll|kvm_exit' || true
# perf record -a -e kvm:kvm_halt_poll_ns_grow \
#   -e kvm:kvm_halt_poll_ns_shrink -e kvm:kvm_exit -- sleep 10

# 5. Turn it off host-wide, right now. No reboot, no VM restart: the
#    next halt exit on every vCPU on this host reads the new value.
# echo 0 > "$P/halt_poll_ns"
#
# A sysfs write does not survive a reboot. Persist it with a
# /etc/modprobe.d/kvm.conf "options kvm halt_poll_ns=0" line, or with
# kvm.halt_poll_ns=0 on the host kernel command line if kvm is built in.

Measure it yourself, because I am not going to give you a number

I am deliberately not publishing a figure for how much CPU halt polling burns on an idle guest, and you should be suspicious of anyone who does. The answer depends on the host CPU, the guest's vCPU count, how often the guest idles, how long its idle gaps are, how loaded the host is, and the host kernel version. A single number from my fleet would be worse than no number, because you would plan with it.

What I will give you is the experiment, which is simple enough to be trustworthy: hold the guest constant, move only the host knob, and measure the vCPU threads' CPU time over a fixed wall-clock window. Run it with an idle guest and the default is the expensive arm. Then run it again with a guest that idles in short bursts — an HTTP server taking a request a second is ideal — and watch the sign of the trade flip, with polling now winning on request latency while still costing more host CPU. Both results are real. Which one describes your fleet is the actual question.

#!/usr/bin/env bash
# Measure the polling tax on YOUR hardware. Do not take a number from me
# or from anyone else: it depends on the host CPU, the guest's idle
# pattern, how many vCPUs the guest has, and the host kernel version.
#
# Shape of the experiment: one deliberately idle guest; measure the CPU
# time its vCPU threads accrue over a fixed wall-clock window, first
# with halt polling at its default and then with it disabled. Nothing
# inside the guest changes between arms. Only the host knob moves.
set -euo pipefail

P=/sys/module/kvm/parameters/halt_poll_ns
WINDOW=${WINDOW:-60}                 # seconds of wall clock per arm
ORIG=$(cat "$P")
trap 'echo "$ORIG" > "$P"' EXIT      # always put the host knob back

PID=${PID:?set PID to the firecracker process you want to measure}

# Firecracker names its vCPU threads fc_vcpu N, so the thread comm is a
# reliable selector. Everything else in the process (the API thread, the
# VMM thread, the virtio queue handlers) is not a vCPU and must not be
# counted, or you are measuring device emulation instead.
mapfile -t TIDS < <(
  for t in /proc/"$PID"/task/*; do
    grep -q '^fc_vcpu' "$t/comm" 2>/dev/null && basename "$t"
  done
)
[ "${#TIDS[@]}" -gt 0 ] || { echo "no fc_vcpu threads found"; exit 1; }
echo "vcpu threads: ${TIDS[*]}  (clock ticks/sec: $(getconf CLK_TCK))"

# utime+stime, summed across the vCPU threads. A polling vCPU spins
# inside the kernel, so the cost lands in SYSTEM time, not user time --
# which is also why it never shows up in a guest-side profile.
cpu_ticks() {
  local sum=0 t
  for t in "${TIDS[@]}"; do
    # Strip through the comm field (it can contain spaces and parens),
    # then utime/stime are fields 12 and 13 of what remains.
    read -r -a f < <(sed 's/.*) //' /proc/"$PID"/task/"$t"/stat)
    sum=$(( sum + f[11] + f[12] ))
  done
  echo "$sum"
}

arm() {                              # arm <halt_poll_ns value>
  echo "$1" > "$P"
  sleep 3                            # let the adaptive window settle
  local a b
  a=$(cpu_ticks); sleep "$WINDOW"; b=$(cpu_ticks)
  printf 'halt_poll_ns=%-9s %6s ticks over %ss\n' "$1" "$((b-a))" "$WINDOW"
}

# Before you start: make sure the guest really is idle. No cron, no
# agent heartbeat, no metrics scrape, no package manager. An idle guest
# still takes timer interrupts, so the floor is not zero -- what you are
# looking for is the DIFFERENCE between the arms, not the absolute.
arm "$ORIG"
arm 0
arm "$ORIG"   # repeat arm one. If the two default arms disagree much,
              # the guest was not idle and the measurement is worthless.

# Then repeat the whole thing with a guest that idles in SHORT bursts --
# an HTTP server taking one request a second is ideal -- and watch the
# sign of the trade flip: now the polling arm is the one that wins, on
# request latency, while still costing more host CPU.

One correction to a reflex I had myself, and watch for it in your own reasoning. The latency halt polling protects is not sandbox create latency. Every PandaStack create is a snapshot restore — p50 179 ms, p99 around 203 ms, with the `/snapshot/load` step itself in the 49-80 ms range — and that path is dominated by network setup, a reflink and loading the memory snapshot, not by a vCPU waking from a halt. The whole sequence is in The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms. Halt polling affects latency inside an already-warm sandbox, per idle gap. If you tune it hoping to improve cold start, you have tuned the wrong thing and you will get a bill for it.

The four configurations, and how to choose

The same trade, four ways. No row is universally correct — the right one depends on what the host is for.
ConfigurationWake latency after an idle gapHost CPU burned while idleAchievable densityAccounting and billing honesty
Poll window far too small (a few microseconds)Most polls expire before the interrupt lands, so you pay the full block-and-wake anyway, plus the pollSmall but real — you bought the cost and almost none of the benefitBarely affectedSlightly polluted. The worst value for money in the table: all tax, no service
Adaptive default (ceiling 200 microseconds on x86, grow 2, shrink 0)Good for guests with frequent short idle gaps: a successful poll skips the scheduler entirelyProportional to how often your guests idle and how often the poll missesFine on a handful of fat VMs — this is the case the feature was designed for and shipped on forIdle guests accrue system time on their vCPU threads; load average reads high on an empty host
Deliberately raised ceilingHighest hit rate; for a latency-sensitive guest the wake path all but disappearsHighest — every failed poll burns the whole, larger windowWorst. Sleepers consume cores they are not using, and on a mostly-idle host every poll runs to full lengthWorst. An idle tenant can be indistinguishable from a busy one in CPU terms
Disabled (halt_poll_ns=0)Every idle gap pays a full block and wake-up; the cost lands on the guest's next piece of workEffectively none from this mechanismBest. A sleeping guest's vCPU thread is genuinely blocked and the pCPU is genuinely freeHonest. vCPU CPU time tracks guest work, which is what every tool downstream assumes

The decision framework is two questions and one constraint. The questions: are my idle gaps short enough that a poll will usually hit, and is the latency on the far side of those gaps something a customer is waiting for? Two yeses means poll, and consider raising the ceiling. A no to either means turn it down, and if you are packing sleepers, turn it off.

The constraint is the one that makes this an unusual tuning problem: the module parameter is global to the host. Every VM on that host gets the same answer. You cannot give the latency-sensitive tenant a generous poll window and the sleeping tenant none, not from the module parameter and absolutely not from inside a guest. KVM does have a per-VM capability a VMM can use to set a per-VM ceiling, which in principle breaks the deadlock — but the VMM has to implement and expose it, and Firecracker offers no configuration for it today, so for anyone running Firecracker this is a host-wide policy decision in practice. Verify that against your own VMM's current documentation if it matters to you; it is the kind of thing that changes.

Which turns the question into a placement question, and that is the genuinely useful reframing. If you run a mixed fleet, the right move is not to find a compromise value — a compromise is the first row of that table, which is the row with no upside. The right move is to split the fleet: hosts configured for latency with polling on, hosts configured for density with polling off, and a scheduler that knows which is which. Memory Overcommit & Page Sharing: How MicroVMs Get Dense and MicroVM Density: The Economics of Per-Tenant Isolation cover the economics on the density side of that split.

Who owns this knob, honestly

On a managed platform, this is our decision and not the customer's, and I would rather say that plainly than let it look like a feature you can request. `halt_poll_ns` lives in the host kernel. A sandbox cannot read it, cannot write it, and cannot tell from the inside what it is set to — the probe in the Python block above demonstrates exactly that, which is the only reason it is in the post. If you run on us, you get whatever we chose, and what we chose follows from the business we are in: hosts that pack large numbers of mostly-idle sandboxes, with CPU billed by the second, which makes density and accounting honesty worth more to us than a saved scheduler round trip on a warm guest's idle gap.

If you run Firecracker yourself, the inverse is true and it is a gift. This is one of a very small number of knobs that changes your cost per sandbox without changing a line of your application code, without a redeploy, without a template rebuild, and without a VM restart: one `echo` into one sysfs file, effective at the next halt exit, reversible just as fast. That is a rare shape for an optimisation. Measure it on your own hardware with your own guests, using the harness above, and make it a documented fleet policy rather than a setting that drifts between hosts — ours is pinned in configuration management precisely because a host-wide knob that nobody owns will eventually differ between two hosts and produce an incident nobody can reproduce.

Where I would not reach for PandaStack at all, while I am being honest about limits: we have no GPUs and no GPU passthrough, so anything accelerated is off the table regardless of how the host is tuned. Our guest kernel is 5.10, so if you depend on a newer kernel feature — including, ironically, anything in the guest-side idle machinery that landed after it — check You Ship the Kernel: Firecracker Guest 5.10 vs 6.1 before you build on us. And if your workload is one long-lived latency-critical machine rather than thousands of disposable ones, a dedicated instance with pinned cores and halt polling turned up is a better purchase than a density platform, and you should buy that instead.

The bottom line

Halt polling is a good feature with a default tuned for a workload that is not a sandbox fleet. It buys guest latency with host cycles, and the exchange rate depends entirely on how often your guests are about to be woken — which on a host full of sleepers is "rarely", at the exact moment the mechanism is at its most expensive, because an idle host lets every poll run to its full window.

So go and read `/sys/module/kvm/parameters/halt_poll_ns` on your hosts, and read the per-vCPU hit rate next to it. If you are packing idle guests and billing or scaling on CPU, that file is reporting a number into your invoices that your guests did not earn. Measure the trade on your own hardware, pick a side deliberately, and write the choice down somewhere a future operator will find it — because the single most expensive version of this is a fleet where half the hosts poll, nobody remembers deciding, and the capacity graphs have been quietly wrong for a year.

Frequently asked questions

Is halt polling the same thing as CPU steal time, or as a busy-wait in my application?

No to both, and the distinctions matter because they point at completely different fixes. A busy-wait in your application is guest CPU time: the guest is running your instructions, it shows up as user time in the guest, and any guest-side profiler will find it. Halt polling is host CPU time spent while the guest is not running at all — the vCPU thread is spinning inside the host kernel, in `kvm_vcpu_halt`, after a halt exit. No guest-side tool can see it, because from the guest's point of view no time passed and no instructions were executed. The cost appears only in host-side accounting, as system time on a thread named `fc_vcpu N`. CPU steal time is the opposite direction again: that is the guest noticing that wall-clock time advanced during intervals when the host scheduler had not given it a CPU, and it tells you the host is oversubscribed. You can have all three at once, and they are diagnosed in different places — guest profiler for the first, host per-thread CPU time plus the KVM halt statistics for the second, the guest's own `st` figure for the third. If you are chasing "why is this host busy when every guest is idle", only the second one is a candidate.

Can I set halt_poll_ns per VM, or per tenant, or from inside a guest?

From inside a guest, no, not in any form. `halt_poll_ns` is a host kernel module parameter; a guest cannot read it, write it, or determine its value by observation. That is correct behaviour and not a limitation to work around — it is host policy about how host CPU is spent, and a tenant does not get a vote on that. From the host, the module parameter is global: every VM on that host gets the same ceiling, the same grow multiplier, the same shrink divisor. KVM does additionally provide a per-VM capability that a VMM can use to set a per-VM ceiling, which in principle lets a hypervisor give different VMs different answers, but the VMM has to implement and expose it. Firecracker currently offers no configuration for it, so if you run Firecracker the practical situation is one value per host. Verify that against Firecracker's current documentation rather than taking my word for it, because this is exactly the sort of thing that gets added. The useful consequence of the constraint is architectural: because you cannot differentiate within a host, you differentiate between hosts. Dedicate some hosts to latency-sensitive guests with polling enabled, others to density with polling off, and let the scheduler place workloads onto the right pool. That is a cleaner design than any single compromise value, because the compromise value is the configuration with the cost of polling and almost none of its benefit.

Should I enable the guest-side haltpoll driver instead of host-side polling?

If it is available to you and your guests are the latency-sensitive kind, it is the better-shaped mechanism, because a poll that succeeds inside the guest avoids the VM exit entirely rather than merely avoiding the block that follows it. The guest never leaves the guest, which is strictly cheaper than exiting and then deciding not to block. The two mechanisms are also designed not to double-charge you: when the guest-side driver comes up it writes a KVM paravirt MSR asking the host to stop polling on its behalf, so you are not paying for the same bet in two places. The catch is availability, and it is a real catch. The driver has to be built into your guest kernel, and it will not load unless the hypervisor presents a paravirtualisation hint in its CPUID leaf saying guest-side polling is appropriate, absent which it needs a force parameter. Our guest kernel is 5.10 and Firecracker normalises the CPUID a guest sees, so whether any given guest can use this is a build-and-platform question rather than something you can configure at runtime. Read `/sys/devices/system/cpu/cpuidle/current_driver` and `current_governor` inside your own guest and believe what they say. And note that moving the polling into the guest does not change the economics for a dense fleet of sleepers — it changes who is spinning, not whether cycles are being burned on a bet that is usually going to lose.

If I disable halt polling, how much latency am I giving up?

Measure it, because the number is workload-dependent and anyone quoting one from their own fleet is handing you a figure that will not transfer. What I can give you is where to look and what shape to expect. The cost you re-introduce is a block-and-wake round trip on every idle gap, so the magnitude of the regression scales with how often your guest idles and how latency-sensitive the work on the other side of each gap is. A batch job that idles once a minute will not notice at all. A request-response server handling steady traffic with short gaps between requests will notice on its tail latencies, because each request now starts with a wake-up. Two things make this less alarming than it sounds for a sandbox platform specifically. First, the per-vCPU statistics tell you in advance what you are giving up: `halt_wait_ns` is the time already being spent blocked, and the ratio of successful to failed polls tells you how often polling is currently rescuing anything. If the hit rate is low, you are giving up very little. Second, cold-start latency is unaffected — on a snapshot-restore platform a create is dominated by network setup, a reflink and loading the memory snapshot, not by a vCPU waking from a halt, so disabling it does not touch the 179 ms create path. Run the A/B harness in this post against a guest that mimics your real idle pattern rather than a fully idle one, and the answer will be yours.

How do I tell whether halt polling is actually paying for itself on a given host?

Read the per-vCPU KVM statistics and compute two ratios, which is a five-minute job and more informative than any amount of reasoning about it. The first is the hit rate: `halt_successful_poll` divided by `halt_attempted_poll`. That tells you how often the bet is won. The second, and the one I would lead with, is `halt_poll_success_ns` divided by the sum of `halt_poll_success_ns` and `halt_poll_fail_ns` — the fraction of polled nanoseconds that bought something. Failed-poll nanoseconds are, by construction, host CPU you burned for no benefit whatsoever, so that denominator is your tax bill in units you can reason about. Put `halt_wait_ns` next to them for context: that is time genuinely spent blocked, which is the cost polling was trying to avoid in the first place, and if it dwarfs the poll numbers then your guests are idling in long stretches and polling was never going to help them. Those counters are exposed either through KVM's debugfs directories, one per VM with a subdirectory per vCPU, or through the newer binary statistics interface that a current `kvm_stat` prefers; which one you get depends on your host kernel version. Collect them per host rather than fleet-wide, because the answer legitimately differs between a host running a few busy guests and a host holding a hundred sleepers — and if your fleet contains both, that difference is the argument for splitting the fleet into two pools rather than picking one global value and hoping.

Keep reading

Related posts

  • DAMON: Measuring What a Sandbox Actually Touches

    Charging a host for every guest's configured RAM is what caps your density, and it is wrong by a lot — a guest handed 4 GiB to run an idle Node server touches a fraction of it. DAMON is the kernel subsystem that can tell you which fraction, at a cost bounded by region count instead of memory size. Here is how it works, and the three different wrong answers you can get instead.

  • Tidal: Turning the Host Into a Cache, and Paying for the Seconds You Touch

    Committed capacity is the address space; the working set is what deserves host RAM. We rebuilt our compute layer around cache semantics — burst freely when the host is quiet, fair-share when it isn't, and a bill that meters the seconds you actually touch.

  • Firecracker vCPU Hotplug, and How CPU Scaling Actually Works

    Somebody always asks for "more CPU" on a running microVM. Firecracker's answer is no — vcpu_count is fixed before InstanceStart and frozen into every snapshot. The good news: the lever you actually wanted was never vCPU count in the first place.

  • Your benchmark ran at a different clock speed than production

    The same core does not run at the same speed twice. Governor, turbo bin, how many neighbours are busy, thermal headroom and ramp latency all move it — which is why the first sandbox on a quiet host looks fast and the fiftieth on a busy one gets blamed on the platform.

  • Bare metal vs cloud VMs for running Firecracker

    If you run microVMs, you run a hypervisor — and the layer underneath it is an architecture decision, not a procurement detail. Nested virt taxes every exit and every page fault; bare metal hands you the machine and the pager. Here is the ledger I actually use.

More in Firecracker & microVMs · See Firecracker microVM sandboxes

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.