all posts

The I/O Scheduler Inside a MicroVM Is Almost Always Wrong

Ajay Kumar··10 min read

There is a specific fiction running inside every virtual machine's block layer, and in a microVM it gets absurd enough to be worth staring at. The guest believes it is driving a disk. It has opinions about that disk: that seeking costs something, that sorting requests by sector makes them cheaper, that the controller holds only so many in flight, that reads and writes compete for a physical arm. Every one of those beliefs feeds a scheduling decision.

What is actually underneath is a virtio-blk device — a ring buffer in shared memory, serviced by a thread in the Firecracker process, which calls `pread` and `pwrite` on an ordinary file on the host's filesystem. There is no arm. The sectors the guest is so carefully sorting are offsets into a file whose extents were placed by a host filesystem the guest cannot see, on a device it cannot name. Sorting them is astrology with better tooling.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs, so I spend most of my time on the host side of that boundary. This post is the guest side. The conclusion up front: use `none`, leave merging alone, and spend the attention you were going to spend on the elevator on readahead and on warming up your benchmark instead. The rest is why, the honest exceptions, and the large class of sandbox workloads where none of this moves a number at all.

What one I/O actually crosses

Before picking a scheduler it helps to count the layers, because the guest's elevator is the first of several things that will reorder the same work.

  • Guest filesystem and block layer — ext4 or XFS builds bios, blk-mq turns them into requests, merges what it can, and an elevator may sort and batch before dispatch. This is the only layer you touch when you write to `/sys/block/vda/queue/scheduler`.
  • The virtio-blk ring — descriptors in memory shared with the VMM, a notification out, an interrupt back. One queue, with a depth fixed by the device's ring size; read Firecracker's block device source for the version you run. This is the layer that makes request COUNT expensive and request SIZE cheap.
  • Firecracker's block device thread — an event loop turning each descriptor into host file I/O. Synchronous `pread`/`pwrite` by default, with an opt-in io_uring engine available. Whether a guest flush becomes a host `fsync` is the drive's cache mode, fixed at bake time.
  • The host filesystem, page cache, and the host's own block layer — with its own scheduler (on NVMe, usually `none`), its own merging, its own readahead, and a page cache that absorbs an enormous fraction of everything above it.

Two consequences fall out. First, anything your guest elevator achieves by reordering is re-reordered, or ignored, twice more before it reaches hardware. Second, the genuinely expensive thing here is crossing the boundary: each request costs a descriptor, a notification, a handoff and a host syscall. Ten 4 KiB requests and one 40 KiB request move the same bytes at wildly different prices. Hold onto that, because it is the only reason the next section is not simply "turn it off and stop reading."

Why the answer is usually `none`

`none` is not the absence of a block layer. It is the absence of an elevator: requests enter the software queues and dispatch roughly in arrival order, with merging still fully active. Four arguments point at it in a microVM, and they stack.

The guest cannot know the real queue depth. An elevator's batching heuristics model how much the device will usefully swallow at once, and the guest's only view of that is the virtio ring — a software constant chosen by the VMM, not a property of the NVMe drive many layers down that is also serving a few dozen other sandboxes. Sorting by sector is meaningless for a related reason: the guest's LBAs are offsets into a file, and on PandaStack that file is a copy-on-write clone of a template image, so its extents are shared with the parent and new blocks land wherever the host filesystem put them. Two adjacent guest sectors may be nowhere near each other physically; two distant ones may sit in the same host page.

Idling is actively harmful. Several elevators hold the queue briefly hoping a nearby request arrives, because on a spinning disk that bet pays for itself many times over. Here it is dead latency on a burstable, shared vCPU, waiting for a locality that does not exist.

And the isolation unit is the VM, not the process. Fairness schedulers exist to stop one greedy process starving another on a shared device. In a sandbox platform the device is shared between virtual machines, and that contention is the host's problem — its block layer, cgroup I/O controllers, and in Firecracker's case a per-drive rate limiter. Running per-process fairness inside a guest, when the guest exists precisely because you did not want to share, solves the wrong problem one layer too high.

The detail everyone gets wrong: `none` still merges

This is the most common error in guest I/O tuning advice, and it matters because it makes people keep an elevator for a benefit the elevator is not providing.

Request merging is done by the block layer, not by the I/O scheduler. Switching to `none` turns off sorting, batching and deadline accounting. It does not turn off merging: adjacent bios are still coalesced into single requests through the plug list and the per-CPU software queues. To disable merging you have to say so explicitly via `nomerges` — and you almost certainly do not want to.

Internalise that, because merging is the part that actually pays in a microVM. Fewer, larger requests mean fewer ring descriptors, fewer notifications, fewer host syscalls, and better odds the host page cache serves the whole thing from one place. Same economics as batching packets, and `none` leaves it untouched.

There is exactly one mechanism by which an elevator can beat `none` on a virtual device, and it is a merging effect rather than a sorting one. With a scheduler attached, blk-mq allocates a separate pool of scheduler tags, so the block layer parks more requests before dispatching — giving merging more chances to find neighbours. You can see it in `nr_requests`, which often reads higher with `mq-deadline` attached than with `none`, where the bound is the driver's own tag depth. If an elevator wins on your workload, check whether you can get the same win by letting requests grow instead: `max_sectors_kb` (bounded by `max_hw_sectors_kb`), and fixing whatever is issuing a thousand small writes.

Two more knobs to know and leave alone. `rotational` is already 0 for virtio-blk — the driver sets the non-rotational flag itself — so writing to it confuses your own heuristics rather than tuning anything. `nr_requests` is bounded by what the hardware queue can back, and the kernel will clamp or refuse a value it cannot honour. `Documentation/block/queue-sysfs.rst` in the kernel tree is the authority for all of these, and it is short enough to read in full.

The four schedulers you may find in `/sys/block/vda/queue/scheduler`, judged as a microVM guest.
SchedulerWhat it doesWhen it helps in a microVMThe catch
noneNo elevator. Requests dispatch roughly in arrival order. The block layer still merges.The default answer. Lowest CPU cost per request, no idling, no sorting against meaningless sector numbers.No fairness inside the guest at all. One process running a giant `dd` can push read latency up for everything else in that VM, and nothing will stop it.
mq-deadlineSeparate read and write FIFOs with expiry deadlines, plus sector sorting and batched dispatch. Tunables under `queue/iosched/`.The only credible exception: a single guest running a latency-sensitive reader alongside heavy writeback, where reads must not starve.The sorting half is wasted work on a file-backed device, so you pay for the whole algorithm to get the fairness half — which cgroup I/O controllers may give you more directly.
kyberLatency-targeted throttling: you give it read and write latency targets and it throttles to try to hit them.Rarely. It was built for genuinely fast multi-queue devices with a stable latency signal.The latency it measures is the host's queue, page cache and noisy neighbours — not a device. Targeting a number that moves for reasons outside the guest produces throttling you did not ask for.
bfqProportional-share fairness with per-process and per-cgroup weights, heavy heuristics, and idling to protect interactivity.Almost never. It was designed for desktop interactivity on rotational disks.Highest per-request CPU cost of the four, and its idling is unambiguously wrong here. Often not compiled into a minimal guest kernel, so you may not have the option.

Readahead is the knob that actually moves numbers

If you change one thing in a guest's block layer, change `read_ahead_kb`. It has a far larger effect on realistic workloads than the scheduler choice, for a boring reason: it decides how many bytes the kernel fetches that nobody asked for.

The default is sized for a device where a seek is expensive and the marginal bytes after it are nearly free — a precise description of a spinning disk and a poor one of anything in this stack. For long sequential reads, raising it means fewer and larger requests crossing the virtio boundary, which is a real win. For the random small-file pattern that dominates actual sandbox work — unpacking a dependency tree, a compiler walking thousands of headers, a test suite stat-ing fixtures — it is pure waste: you fetch far more than you use, pay for it on the ring and in host syscalls, then evict something useful to hold it.

And it is applied twice. The guest reads ahead into the guest page cache; those reads become `pread` calls on the host file; the host filesystem then reads ahead for that file into the host page cache. A sequential guest read can pull substantially more off the physical device than it requested, at two independent layers neither of which knows about the other. The guest half is charged against RAM fixed at bake time — Firecracker cannot change guest memory at snapshot restore, so our `base` template's 4 GiB is 4 GiB forever — and page cache competes with your build's resident set inside that budget. On a 2 GiB template, aggressive readahead is a reclaim generator. Raise it for bulk sequential reads, lower it for metadata storms, and measure.

#!/usr/bin/env bash
# What the guest's block layer thinks it is driving, and what you can tell it.
# Run this INSIDE a sandbox. /dev/vda is a virtio-blk device whose "platter"
# is a file on the host -- most of these numbers were chosen for hardware the
# guest does not have.
set -euo pipefail

Q=/sys/block/vda/queue

echo "=== what the guest believes ==="
# The active scheduler is in brackets; the others are what this kernel built.
# On a device with ONE hardware queue -- which is what Firecracker's virtio-blk
# looks like -- blk-mq's default is an elevator, not 'none'. Read the file
# rather than trusting any blog post about what the default is.
cat "$Q/scheduler"          # e.g. [mq-deadline] kyber none   <- bracket = live
cat "$Q/rotational"         # 0: virtio_blk sets NONROT itself. Don't write here.
cat "$Q/read_ahead_kb"      # the number that actually moves benchmarks
cat "$Q/nr_requests"        # how many requests the block layer will park
cat "$Q/max_sectors_kb"     # largest request it will build...
cat "$Q/max_hw_sectors_kb"  # ...and the ceiling it cannot exceed
cat "$Q/nomerges"           # 0 = merge freely. Leave it at 0. See below.

echo "=== set it ==="
# 'none' does NOT mean "no merging". Merging is done by the block layer, not by
# the elevator, and it survives switching to none. What you are turning off is
# sorting, batching and read/write deadline accounting -- the parts that model a
# seek penalty that does not exist here.
echo none > "$Q/scheduler"

# Readahead: the default is sized for a spinning disk you do not have, and it is
# applied twice -- once here, and again by the HOST filesystem underneath the
# virtio-blk file. Raise it for sequential bulk reads; lower it for the random
# small-file storm that is `npm install` or `pip install`. Measure, don't guess:
# every page readahead fills is guest RAM, and your template's RAM was baked.
echo 256 > "$Q/read_ahead_kb"

# There is no `elevator=` boot parameter to reach for. It belonged to the legacy
# single-queue block layer, was deprecated, and went away with it; on a blk-mq
# kernel it does nothing except possibly log a complaint. Confirm against
# Documentation/admin-guide/kernel-parameters.txt for the kernel you run.

echo "=== make it stick: a udev rule in the TEMPLATE ==="
# Written into the template image, not into a running sandbox. The subtlety is
# in the next section: this rule fires on device add, and a snapshot restore has
# no device add.
install -d /etc/udev/rules.d
cat > /etc/udev/rules.d/60-pandastack-io.rules <<'RULE'
# virtio-blk: no elevator, readahead chosen on purpose.
ACTION=="add|change", SUBSYSTEM=="block", KERNEL=="vd[a-z]", \
  ATTR{queue/scheduler}="none", ATTR{queue/read_ahead_kb}="256"
RULE
udevadm control --reload
udevadm trigger --subsystem-match=block --action=change

# Verify, because a udev rule that silently did not match is the most common
# outcome here:
udevadm test /sys/block/vda 2>&1 | grep -iE 'scheduler|read_ahead' || true
cat "$Q/scheduler" "$Q/read_ahead_kb"

Why this only works if you bake it

Here is the microVM-specific twist that makes all of the above either trivially easy or completely futile, depending on where you put it. Those sysfs values live in kernel memory. A Firecracker snapshot is, among other things, kernel memory. So a value set in the guest before the snapshot is taken is frozen into it and comes back on every restore, for free, with no startup cost. That is the good half.

The bad half is that a udev rule does not fire on a restore. udev applies `ATTR{}` assignments in response to device events, and the only `add` event `/dev/vda` ever sees is during a cold boot. A restore resumes a kernel that already has the device enumerated: no new device, no event, rule never runs. On PandaStack that cold boot happens exactly once per template — the first spawn takes roughly 3 seconds, captures a snapshot, and every create after that is a restore of that snapshot at p50 179 ms. So the udev rule in your template image runs precisely once, at bake time, and that single run is what every subsequent sandbox inherits.

That is the mechanism, not a bug, and once you see it the design falls out: put the tuning in the template, let the bake's cold boot apply it, let the snapshot carry it. The alternative — an `exec` writing sysfs after each create — works, but it costs a round trip on a create whose selling point is finishing in under 200 ms, and it has to be remembered in every start script forever.

And no, you cannot do this on the kernel command line. The `elevator=` parameter belonged to the legacy single-queue block layer; it was deprecated and then removed along with it, and on a blk-mq kernel it does nothing but possibly log a complaint about being obsolete. Our guests boot with `console=ttyS0 reboot=k panic=1 pci=off` and there is no scheduler token to add. Confirm against `Documentation/admin-guide/kernel-parameters.txt` for the kernel you actually run, which is the only honest source for which parameters still exist.

# templates/base-tuned/Dockerfile
# The only place these settings survive a snapshot restore is the template bake.
FROM pandastack/base:latest

# A udev rule fires on a block device ADD or CHANGE event. The ONLY time that
# happens for /dev/vda is the single cold boot that precedes the snapshot --
# a restore resumes a kernel that already has the device, so udev never runs
# again. That is not a problem; it is the mechanism. The bake's cold boot
# applies the rule, and the snapshot freezes the applied sysfs state into
# kernel memory, which is what every later restore gets back.
COPY 60-pandastack-io.rules /etc/udev/rules.d/60-pandastack-io.rules

# If you would rather not rely on udev ordering during boot, a oneshot unit
# ordered before local-fs.target does the same job and is easier to read in
# `systemctl status`. Either way: bake it.
COPY tune-vda.service /etc/systemd/system/tune-vda.service
RUN systemctl enable tune-vda.service

# Bake it, with RAM chosen here because it cannot be chosen later:
#   pandastack template build -f Dockerfile -n base-tuned \
#     --size-mb 8192 --memory-mb 4096
#
# Firecracker cannot change vCPU count or guest RAM at snapshot restore, so
# --memory-mb is the one and only place it is decided. Passing memory_mb on a
# create is silently corrected to the baked value, and --cpu is deprecated and
# ignored -- every template gets 8 burstable vCPUs.

Measuring it without fooling yourself

The first write is a different operation from the second

This invalidates more microVM I/O benchmarks than every scheduler argument combined, and it has nothing to do with scheduling. A sandbox rootfs is a copy-on-write clone of a template image — an XFS reflink, or a dm-snapshot. CoW is why a create hands you a multi-gigabyte filesystem in single-digit milliseconds, and it is local-only by necessity: copy-on-write needs a real local block device, which is why we stream a snapshot's memory from object storage over userfaultfd but never stream the disk.

So a write that looks like an overwrite to the guest is a first-touch allocation on the host. With reflink, writing into a shared extent forces an unshare — allocate, copy, remap. With dm-snapshot, a write to an untouched chunk triggers a chunk-sized read-modify-write into the COW device. Either way the first write to a region does strictly more work than the second, and the gap is not subtle. A benchmark that creates a file and immediately measures writes to it is measuring host allocation and calling it I/O — and it will show a steady improvement across the run that you may credit to whatever you just tuned. Pre-allocate, run a full warm-up pass, throw those numbers away, then compare. fio's `--ramp_time` handles the short tail of a run; it does not handle allocation, which is a property of the file, not of the run.

If your two schedulers differ by more than your warm-up pass did, you have a finding. Otherwise you have a copy-on-write filesystem and a hypothesis.

What a guest `fsync` actually guarantees

A guest `fsync` pushes dirty pages out of the guest page cache, through the block layer, across the virtio ring, into Firecracker's device thread. What happens next is not up to the guest — it is the drive's cache mode on the host side. Firecracker's default for a drive does not honour guest flush requests: the `fsync` returns success while the data is still only in the host page cache. A clean shutdown flushes it; a hard power-off does not, and for a journalled filesystem that is worse than losing recent writes, because the journal's ordering guarantees break and the image can come back inconsistent rather than merely stale.

We keep the rootfs in that mode on purpose: it is copy-on-write and discarded with the sandbox, so a real host flush per guest write would slow every workload to protect data that is throwaway by definition. Durable volumes are the opposite — a volume is the durable store, so those drives run in a write-back mode where a guest flush becomes a real host flush. Cache mode is a snapshot property, like hugepage backing: it cannot be patched onto a running VM, so changing it means re-baking. The practical rule is that no scheduler, readahead value or `fsync` call makes a sandbox rootfs durable storage, because it was never trying to be. Anything that must survive goes on a durable volume, into object storage, or into a managed database.

"""Measure your own workload under `none` vs an elevator, in your own sandbox.

The point of this harness is not to produce a number I can put in a blog post.
It is that the only honest answer to "which scheduler" is "run your workload
twice", and the two things people get wrong when they do that are (a) no warm-up
pass over a copy-on-write rootfs, and (b) believing `--direct=1` in a guest
measures a device.
"""

import json
import re

from pandastack import Sandbox

TEMPLATE = "base"          # base is baked at 4 GiB RAM / 8 burstable vCPUs
WORKLOAD = "mixed"         # see FIO_JOBS below

# fio if the template has it, dd if not. dd measures almost nothing useful for a
# mixed workload -- it is here so the harness still runs on a bare template.
FIO_JOBS = {
    # Sequential bulk read: the case where readahead, not the elevator, decides.
    "seqread": "--rw=read --bs=1m --size=512m --numjobs=1",
    # The one shape where an elevator can plausibly help inside a single guest:
    # a reader competing with a writer for the same virtio queue.
    "mixed":   "--rw=randrw --rwmixread=70 --bs=16k --size=256m --numjobs=4",
    # Small-file metadata storm -- i.e. unpacking a dependency tree.
    "smallio": "--rw=randwrite --bs=4k --size=128m --numjobs=8",
}

PROBE = r"""
set -eu
command -v fio >/dev/null && echo fio || echo nofio
cat /sys/block/vda/queue/scheduler
"""

RUN = r"""
set -eu
Q=/sys/block/vda/queue
SCHED="$1"; PASS="$2"

echo "$SCHED" > "$Q/scheduler"
# Confirm it took. A kernel that did not build mq-deadline will not tell you
# twice, and a benchmark comparing none against none is a long way to learn it.
grep -q "\[$SCHED\]" "$Q/scheduler"

# Drop the GUEST page cache between passes. Note what this does NOT do: the host
# page cache sits underneath the virtio-blk file and you cannot reach it from in
# here. Neither can --direct=1 -- O_DIRECT in the guest bypasses the guest's
# cache and then lands in an ordinary pread/pwrite on the host. "Direct I/O" in
# a microVM is not direct to anything.
sync; echo 3 > /proc/sys/vm/drop_caches

cd /work
# Hard limit in the shell, where it belongs: neither exec endpoint enforces
# timeout_seconds server-side. That argument is a CLIENT deadline.
timeout 600 fio --name="$PASS" --filename=/work/testfile \
  --ioengine=libaio --iodepth=32 --group_reporting --output-format=json \
  --ramp_time=5s --runtime=30s --time_based \
  {jobargs}
"""

def main() -> None:
    with Sandbox.create(template=TEMPLATE, ttl_seconds=1800) as sbx:
        probe = sbx.exec(PROBE)
        if "nofio" in probe.stdout:
            raise SystemExit(
                "fio is not in this template. Add it to the Dockerfile and "
                "re-bake -- installing it here measures the install, not the disk."
            )
        # Only compare against elevators this guest kernel actually built. The
        # guest kernel is 5.10 and you do not get to swap it per template, so
        # bfq in particular is often simply absent.
        available = re.findall(r"[\w-]+", probe.stdout.splitlines()[-1])
        candidates = [s for s in ("none", "mq-deadline", "kyber", "bfq")
                      if s in available]
        print("schedulers this kernel built:", candidates)

        sbx.exec("mkdir -p /work && fallocate -l 1G /work/testfile", check=True)

        # ---- THE WARM-UP PASS, which is the whole reason this script exists.
        #
        # The rootfs is a copy-on-write clone of the template image (XFS reflink,
        # or dm-snapshot). To the guest, writing a block that already has
        # contents is an overwrite. To the host it is the FIRST write to a shared
        # extent: unshare it, allocate, copy, remap. First-write latency and
        # steady-state latency are structurally different numbers, and a
        # benchmark that skips this measures allocation and calls it I/O.
        #
        # fio's own --ramp_time handles the short tail; this pass handles the
        # allocation, which is a property of the file, not of the run.
        print("warm-up pass (allocating the CoW extents) ...")
        sbx.exec_stream(
            RUN.format(jobargs=FIO_JOBS[WORKLOAD]).replace('"$1"', "none")
               .replace('"$2"', "warmup") + "\n",
            on_stdout=lambda c: None,
            on_stderr=lambda c: print(c, end=""),
            timeout_seconds=900,
        )

        results = {}
        for sched in candidates:
            if sched in ("kyber", "bfq"):
                continue  # run them if you like; see the post for why I don't
            script = (RUN.format(jobargs=FIO_JOBS[WORKLOAD])
                      .replace('"$1"', f'"{sched}"')
                      .replace('"$2"', f'"{sched}"'))
            out = []
            rc = sbx.exec_stream(
                script,
                on_stdout=out.append,
                on_stderr=lambda c: print(c, end=""),
                timeout_seconds=900,
            )
            if rc != 0:
                print(f"{sched}: rc={rc}, skipping")
                continue
            blob = "".join(out)
            doc = json.loads(blob[blob.index("{"):])
            job = doc["jobs"][0]
            results[sched] = {
                "read_iops":  round(job["read"]["iops"]),
                "write_iops": round(job["write"]["iops"]),
                "read_p99_us":  job["read"]["clat_ns"]["percentile"].get(
                    "99.000000", 0) / 1000,
                "write_p99_us": job["write"]["clat_ns"]["percentile"].get(
                    "99.000000", 0) / 1000,
            }

        for sched, r in results.items():
            print(sched, r)

        # Read the result with the right amount of suspicion. Both passes ran on
        # ONE sandbox on ONE host, sharing a real NVMe with real neighbours. A
        # gap inside the run-to-run noise of that host is not a finding. Repeat
        # the pair a few times, interleaved, before you change a template.
        if results:
            print("\nIf none and mq-deadline are within noise of each other on "
                  "your workload -- which is the usual outcome -- the honest "
                  "conclusion is `none`, and the next thing to tune is "
                  "read_ahead_kb and your template's baked RAM.")


if __name__ == "__main__":
    main()

io_uring, and why size beats cleverness

io_uring shows up here twice, in two different places, and conflating them is common. In the guest it is an application-level submission interface: it cuts syscall overhead and gives you genuine asynchronous I/O with batched submission and completion. What it does not do is change anything below the block layer — requests still go through blk-mq, still get merged there, still cross the same single virtio ring. Note too that our guest kernel is 5.10, one kernel per host and not swappable per template, so io_uring's feature coverage in there is older than current liburing documentation assumes. Check what your kernel supports before building on an opcode.

On the host side, Firecracker has an optional io_uring-based block engine that replaces blocking per-request `pread`/`pwrite` in its device thread with submitted batches. That is the one that changes the shape of the pipe, and it is a host configuration decision rather than something a guest can request. Read Firecracker's own docs for the version you run; this area has moved across releases.

The thread connecting both: the win in a virtualised block path almost always comes from issuing fewer, larger I/Os rather than from reordering the ones you issue. A build that writes 4 KiB at a time through unbuffered I/O will be slow under every scheduler in the table above. Buffer it, let the page cache and the block layer merge, and the elevator question becomes uninteresting — which is the correct outcome.

When none of this matters, which is most of the time

I would rather say this than let the preceding two thousand words imply otherwise. For the overwhelmingly common sandbox workload — restore a snapshot, unpack a dependency tree, compile something, write a few artifacts, exit — the I/O scheduler choice is noise. Not a small effect: noise, below the run-to-run variance of a shared host. What actually decides whether that job is fast, roughly in order:

  • The template's baked RAM. It sets how much of the working set the guest page cache can hold, and it is fixed at build time via `--memory-mb` — `memory_mb` on a create is silently corrected to the baked value. A job that fits in cache does almost no block I/O at all; a job that does not does a great deal, and no scheduler recovers that.
  • Whether the host page cache absorbed your writes. With the rootfs cache mode we run, most writes land there and never see a device during the sandbox's life. That is the single largest performance fact in the whole stack, and it is a host property.
  • What you baked into the template versus what you install at runtime. Installing a toolchain per job is the most expensive I/O decision available, and it is a packaging decision, not a tuning one.
  • Whether you are doing many small I/Os where few large ones would do — buffering, archive handling, how your build writes object files.
  • Readahead, if your access pattern is genuinely sequential or genuinely random and the default is wrong for it.
  • And then, a long way down, the elevator.

The reason to know all this is not to tune every template. It is so that when someone reports that a sandbox has slow disk I/O, you can work out which of six layers they are actually describing — and so the first thing you reach for is the layer that moves the number rather than the one with the most interesting documentation. Set `none`, bake it, pick a readahead value you can justify, and go argue about something that matters.

Frequently asked questions

Should I just set the I/O scheduler to `none` in my microVM template?

For a virtio-blk device backed by a file on the host, yes, that is the right default, and you should set it at template build time so the snapshot carries it. The reasoning is that every job an elevator does is either meaningless or redundant here: sorting by sector number optimises seeks on a device the guest cannot see, whose extents were placed by a host filesystem; batching is tuned against a queue depth that is a software constant rather than hardware; and idling to wait for a nearby request is dead latency on a shared burstable vCPU. Meanwhile the host has its own scheduler, its own merging and a page cache that absorbs most of the traffic. The one genuine exception is a single guest running a latency-sensitive reader alongside heavy writeback, where `mq-deadline`'s read/write deadline accounting can help. Test that specific case on your own workload rather than adopting it defensively, and check `/sys/block/vda/queue/scheduler` to see which schedulers your guest kernel actually built — a minimal kernel often omits bfq entirely.

Does switching to `none` turn off request merging?

No, and this is the detail that most tuning advice gets backwards. Merging is performed by the block layer, not by the I/O scheduler. Adjacent bios are coalesced into single requests through the plug list and the per-CPU software queues regardless of which scheduler is attached, so `none` merges perfectly well. What `none` removes is sorting, batching and deadline accounting. If you genuinely want to disable merging you have to say so explicitly through the `nomerges` sysfs attribute, and that is a debugging tool rather than a tuning knob — merging is the single most valuable thing the guest block layer does in a microVM, because it directly reduces the number of ring descriptors, notifications and host syscalls per unit of data moved. There is one indirect merging effect worth knowing: attaching a scheduler gives blk-mq a separate pool of scheduler tags, so more requests are parked before dispatch and merging gets more chances to find neighbours. If an elevator beats `none` on your workload, that is usually why — and growing `max_sectors_kb` or fixing the application's write sizes gets you the same win more cheaply. Read `Documentation/block/queue-sysfs.rst` for the kernel you run.

Will raising `read_ahead_kb` make my sandbox faster?

It depends entirely on your access pattern, and it is a much bigger lever than the scheduler either way. Readahead decides how many bytes the kernel fetches that nobody asked for. The default is sized for a device where a seek is expensive and the bytes after it are nearly free, which describes a spinning disk and not a file-backed virtio device. For long sequential reads, raising it means fewer and larger requests crossing the virtio boundary, which is a real improvement. For the random small-file pattern that dominates real sandbox work — unpacking a dependency tree, a compiler walking headers, a test suite stat-ing fixtures — readahead is waste: you fetch far more than you use, pay for it on the ring, and evict something useful to hold it. Two further wrinkles specific to microVMs: readahead is applied again by the host filesystem underneath your device file, so a sequential guest read can pull much more off the physical disk than it requested at two independent layers; and the guest half is charged against RAM that was fixed when the template was baked, so on a small template aggressive readahead mostly generates reclaim pressure.

Does a guest `fsync` mean my data is safe in a sandbox?

Not on a sandbox rootfs, and that is a design decision rather than an oversight. A guest `fsync` pushes dirty pages through the guest block layer and across the virtio ring, but whether that becomes a durable write depends on the drive's cache mode on the host side. Firecracker's default mode does not honour guest flush requests: the `fsync` returns success while the data sits in the host page cache. A clean shutdown flushes it; a hard power-off does not, and for a journalled filesystem that is worse than losing recent writes, because the journal's ordering guarantees break and the image can come back inconsistent rather than just stale. We keep the sandbox rootfs in that mode deliberately — it is a copy-on-write clone discarded with the sandbox, so a real host flush per write would slow every workload to protect throwaway data. Durable volumes are the opposite: a volume is the durable store, so those drives run in a write-back mode where a guest flush becomes a real host flush. The cache mode is a snapshot property and cannot be patched onto a running VM, so changing it means re-baking the template.

Keep reading

Related posts

  • EEVDF vs CFS: What the New Linux Scheduler Means for MicroVM Density

    Linux 6.6 swapped the default CPU scheduler out from under everyone. If the runnable threads on your host are the vCPUs of a few hundred oversubscribed microVMs, the change is not academic — and half the tuning advice you'll find now edits sysctls that no longer exist.

  • Your benchmark ran at a different clock speed than production

    The same core does not run at the same speed twice. Governor, turbo bin, how many neighbours are busy, thermal headroom and ramp latency all move it — which is why the first sandbox on a quiet host looks fast and the fiftieth on a busy one gets blamed on the platform.

  • Firecracker boot_args, argument by argument

    Everyone copies the same magic `boot_args` string from the Firecracker docs and never reads it. It's a short, unusually honest description of what a microVM is — and what it has decided not to be.

  • Build, Boot, Test: Kernel CI on MicroVMs (and Where It Stops)

    "Does this patch boot" should be a per-commit check, not a nightly. A microVM makes the boot step sub-second and a git bisect a coffee break — right up to the device model, the nested-virt wall, and the fact that you cannot swap our guest kernel.

  • MGLRU and MicroVM Density: Reclaim When Every Page Is Someone's Guest

    The host kernel picks reclaim victims using access information already filtered through a second kernel. MGLRU gives it generations instead of two buckets — which changes the shape of your bad days, not how many guests fit on the box.

More in Internals · See PandaStack benchmarks

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.