all posts

Guest Writeback Throttling: Why dirty_ratio Is Wrong for a 2 GiB MicroVM

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and the most consistently misdiagnosed performance problem I see has nothing to do with Firecracker. It is a Linux kernel tunable last sanely defaulted for a machine with a spinning disk bolted to it, which now ships unexamined inside every 2 GiB sandbox anybody runs.

The symptom is always described the same way. A build, or a dataset write, or a test suite that produces artefacts runs beautifully for a few seconds and then falls off a cliff. Not gradually — it stops. `top` shows almost no CPU and a pile of processes in `D` state. Then it resumes, and stops again. Somebody says “the disk is slow,” somebody else adds a rate limiter to rule it out, and three people lose an afternoon to a problem whose cause is one sysctl and whose evidence was in `/proc/meminfo` the whole time.

The underlying fact: a Firecracker guest is a complete Linux system that believes it owns a block device. It has a page cache, flusher threads, a writeback bandwidth estimator and a dirty-page throttle, all calibrated for a physical machine with a real disk. Put that inside a small guest whose “disk” is a virtio-blk device backed by a file on a host filesystem with its own page cache and its own flusher threads, and the defaults produce a recognisable pathology. The funniest part: the guest's estimate of its own disk bandwidth — the number the throttle's maths is built on — is a work of fiction.

The short version: use `vm.dirty_bytes` instead of `vm.dirty_ratio` in any fleet with heterogeneous guest sizes, size it to the stall you are willing to take rather than to a fraction of RAM, and understand that several independent layers can throttle a write — guest writeback, the VMM's rate limiter, host cgroup I/O, host memory pressure — whose symptom is identically “the disk is slow.” Get per-layer counters before you touch a knob.

Where a guest write actually goes

Trace a single buffered `write(2)` inside a sandbox, because the number of places the data rests is the whole story.

  1. Your process calls `write()`. The kernel copies the bytes into a guest page-cache page and marks it dirty. The syscall returns. Nothing has been written anywhere. This is a memcpy with a receipt.
  2. Later — on a timer, or because there are now too many dirty pages — a guest flusher thread walks the dirty inodes and submits bios to the guest block layer.
  3. The guest block layer schedules them (or does not, if the scheduler is `none`) and virtio-blk puts descriptors in a virtqueue and kicks the device.
  4. Firecracker, on the host, picks up the descriptors and writes to the backing file with `pwrite` or, with the async engine, io_uring.
  5. The host kernel copies those bytes into the host page cache and marks those pages dirty. Firecracker's write returns. Still no disk.
  6. Later still, a host flusher thread writes them to the real NVMe, subject to the host's own `dirty_ratio`, scheduler and cgroup limits.

Six steps, two page caches, two independent writeback engines — and exactly one of them has any idea what the storage underneath can do. That one is scheduling against a device whose latency is dominated by a queue hand-off to a userspace process on a machine it cannot see. Stacked, these layers produce a system where “the write was fast” usually means “it landed in two caches and no disk.”

The writeback machinery, precisely

Two thresholds, two clocks

`vm.dirty_background_ratio` (default 10) is where the kernel wakes the flusher threads to write dirty pages out in the background. Nobody is penalised; your process runs at memory speed while a kworker drains behind it. This is the threshold you want to live at. `vm.dirty_ratio` (default 20) is the one that hurts: the point at which the kernel decides the writer is dirtying faster than the system can clean, and makes the writing task do the cleanup itself — the “throttle.” Hence the cliff. Below it your process is doing memcpy; above it, storage I/O.

A detail that matters more in a guest than on a server: these are not percentages of total RAM. They are percentages of dirtyable memory — free plus reclaimable pages — so the ceiling moves as the guest fills up. A sandbox fresh off a restore and one forty seconds into a Node build have different dirty thresholds, which is much of why the stall point feels non-deterministic. The kernel publishes the computed values as `nr_dirty_threshold` and `nr_dirty_background_threshold` in `/proc/vmstat`.

Then two clocks. `vm.dirty_expire_centisecs` (default 3000, thirty seconds) is how old a dirty page must be before the flusher writes it regardless of thresholds. `vm.dirty_writeback_centisecs` (default 500) is how often the flushers wake to look. Thirty seconds of unflushed data is a strange default for a sandbox whose entire natural lifetime might be ninety seconds.

balance_dirty_pages(), and the fiction it is built on

The function doing the pausing is `balance_dirty_pages()` in `mm/page-writeback.c`, reached from `balance_dirty_pages_ratelimited()` on the write path. Two things about it are routinely misunderstood.

First, the throttle is applied to the task dirtying pages, not to the system. There is no global brake. Each task carries its own counters — `nr_dirtied`, `nr_dirtied_pause`, `dirty_paused_when` — and enters the full balancing path only once it has dirtied its allowance. It looks global because when dirty pages are over the limit, every process writing to that filesystem exceeds it at the same moment. A process that only reads sails past the queue of stalled writers, which is why “the disk is slow” is so often filed by the one process that was fine.

Second, modern kernels do not implement a hard wall at `dirty_ratio`. Since the throttling rework in 3.2 it is a feedback control loop, with a freerun region — roughly below the midpoint of the background and hard limits — where no throttling happens. Above it, the kernel computes a target rate per dirtier and sleeps it in increments, holding the dirty count near a setpoint. The pause length derives from an estimate of the device's writeout bandwidth, kept as a moving average.

Elegant, and entirely dependent on that estimate being a property of a device. In a guest it is not. What the guest measures is the rate Firecracker drains a virtqueue, which is the rate the host page cache accepts memcpy — enormous, while the host has dirty headroom. So the estimator learns it owns a very fast disk and lets dirty pages climb. Then the host crosses its own background threshold, starts real writeback to real NVMe, and apparent bandwidth collapses. The guest finds out by overshooting the hard limit, where the pauses stop being gentle and your build stops being a build.

The per-BDI limit and strictlimit

The global limit is not the only one. Each backing device gets a share of the dirty budget, derived from its observed proportion of total writeback bandwidth and bounded by `min_ratio` and `max_ratio` in `/sys/class/bdi/`, so one slow device cannot hoard the global allowance. There is also a `strictlimit` mode that turns the per-BDI limit into a hard cap enforced independently of global state — FUSE uses it, because a userspace filesystem that can be arbitrarily slow should not be able to fill a machine's memory with dirty pages.

You will read advice to enable `strictlimit` for your virtio device. Check your kernel first. On the 5.10 guest kernel we ship, that directory exposes `read_ahead_kb`, `min_ratio`, `max_ratio` and `stable_pages_required` — and no `strict_limit`, which is a later addition. And the honest conclusion: in a guest with one virtio-blk device the per-BDI machinery does essentially nothing. A single BDI accounting for all observed writeback bandwidth gets, near enough, the whole global limit. Those knobs arbitrate between devices; you have one device.

#!/bin/sh
# Inside the guest. What does this sandbox currently believe about its disk?
set -eu

sysctl vm.dirty_ratio vm.dirty_background_ratio \
       vm.dirty_bytes vm.dirty_background_bytes \
       vm.dirty_expire_centisecs vm.dirty_writeback_centisecs

# Do not compute the threshold from the ratio yourself -- the kernel already
# did, against *dirtyable* memory (free + reclaimable), not total RAM. Both
# numbers move as the guest fills up, which is why the stall point drifts.
awk '/^nr_dirty_threshold|^nr_dirty_background_threshold/ \
     { printf "%-34s %9d pages %6d MiB\n", $1, $2, ($2*4096)/1048576 }' \
    /proc/vmstat

# What this device's BDI actually exposes. On the 5.10 guest kernel: no
# strict_limit and no *_bytes variants. Both arrived in later kernels.
dev=$(lsblk -ndo MAJ:MIN /dev/vda | tr -d ' ')
ls -1 "/sys/class/bdi/$dev/"

# ---- Now set it in bytes, not percent. -------------------------------------
# Writing dirty_bytes zeroes dirty_ratio and vice versa: they are two views
# of one limit, and the kernel keeps exactly one of them authoritative.
# 64 MiB hard / 16 MiB background, sized for a ~2 GiB sandbox guest.
sysctl -w vm.dirty_bytes=67108864
sysctl -w vm.dirty_background_bytes=16777216
sysctl -w vm.dirty_expire_centisecs=500    # overdue after 5s, not 30s
sysctl -w vm.dirty_writeback_centisecs=100 # flusher wakes every 1s

# Bake it into the template image, not into one boot. A sysctl applied by a
# post-boot script is a sysctl that is missing on the restore path.
cat > /etc/sysctl.d/60-sandbox-writeback.conf <<'CONF'
vm.dirty_bytes = 67108864
vm.dirty_background_bytes = 16777216
vm.dirty_expire_centisecs = 500
vm.dirty_writeback_centisecs = 100
CONF

Ratios are the wrong unit in a small guest

The argument is one line of arithmetic. Twenty percent of a 2 GiB guest is about 400 MiB of dirty pages queued against a virtio device. Twenty percent of a 16 GiB guest is 3.2 GiB — the same sysctl, unchanged, producing an eightfold difference in worst-case drain time. If your fleet runs one template at 2 GiB and another at 4 GiB — ours does; `agent` and `code-interpreter` are baked at 2 GiB, `base` and `browser` at 4 GiB — a ratio is not a policy. It is a policy generator, and nobody reviewed the policies it emitted.

What you care about is a duration: how long can one task be made to wait while the backlog it created drains? That is bytes divided by bandwidth, and bytes is the half you can set. So set it in bytes, and size it to the stall you are willing to take. Pick a worst-case pause you can defend — under a second for an interactive sandbox, several seconds for a throughput-measured batch job — multiply by write bandwidth you have actually measured through the virtio path, and set `dirty_background_bytes` to about a quarter of it.

Two caveats I will not hide. Measure that bandwidth on a busy host, not an idle one — an idle host will tell you the host page cache is infinitely fast, which is exactly the fiction we are trying to stop believing. And a smaller budget costs throughput on bulk sequential writes, because writeback happens in more, smaller batches. If your workload is one enormous `tar` and the only metric is wall time, the defaults may legitimately win.

Do not set `dirty_bytes` absurdly low to “keep it smooth.” The minimum the kernel accepts is two pages, and anything near that turns every write into synchronous I/O with the full writeback path's overhead attached. The useful range for a sandbox guest is tens of megabytes.

Catching the stall in the act

You do not have to infer any of this. Two numbers matter, and the difference between them is the diagnosis. `Dirty` in `/proc/meminfo` is data waiting for somebody to write it; `Writeback` is data in flight to the device. `Dirty` pinned against `nr_dirty_threshold` with small `Writeback` means the backlog is not draining and your writers are throttled while the device sits idle-ish. Large sustained `Writeback` while `Dirty` falls slowly means the backlog is draining and the device really is the constraint. Identical from userspace, completely different fixes.

#!/bin/sh
# Inside the guest. Deliberately overcommit dirty pages, then watch.
# Write more than the guest has RAM; the point is to cross the threshold.
set -eu

dd if=/dev/zero of=/work/fill bs=1M count=4096 2>/dev/null &
writer=$!

printf '%-10s %-10s %-10s %-5s %s\n' DIRTY_kB WB_kB THRESH_kB STAT WCHAN
while kill -0 "$writer" 2>/dev/null; do
  # Dirty = waiting for someone. Writeback = in flight to the device.
  d=$(awk '/^Dirty:/{print $2}' /proc/meminfo)
  w=$(awk '/^Writeback:/{print $2}' /proc/meminfo)
  # The computed ceiling, converted to kB so it compares with the two above.
  t=$(awk '/^nr_dirty_threshold/{print ($2*4)}' /proc/vmstat)
  # Is the writer asleep in the kernel, and where? If WCHAN says
  # balance_dirty_pages, the argument is over: this is guest writeback.
  s=$(awk '{print $3}' "/proc/$writer/stat" 2>/dev/null || echo -)
  c=$(cat "/proc/$writer/wchan" 2>/dev/null || echo -)
  printf '%-10s %-10s %-10s %-5s %s\n' "$d" "$w" "$t" "$s" "$c"
  sleep 1
done

# Fleet-level proof, and the one to alert on. "full" is the share of wall
# time in which EVERY runnable task was stalled on I/O -- not a sample of
# one process, a statement about the whole guest.
cat /proc/pressure/io
rm -f /work/fill

The `wchan` column is the thing I wish more people knew about. When a process is parked in `balance_dirty_pages`, the kernel tells you so by name — no inference, no flame graph, no afternoon. And `/proc/pressure/io` is worth exporting per sandbox: `full` is the proportion of time in which every runnable task was stalled on I/O, which is what your user experiences as “it froze.”

Two page caches, stacked

Now the structural problem, which tuning improves but does not remove. The guest dirties a page. Writeback pushes it through virtio-blk. Firecracker writes it to the backing file. The host kernel copies it into its page cache and marks it dirty. The page is now dirty in two places, in two kernels, each with its own thresholds, expiry clock and flusher threads, neither aware of the other's state.

The first is memory accounting. Host page cache dirtied on behalf of a VM is host memory, and it grows with that VM's write traffic. Where the host filesystem supports cgroup writeback it is charged to the cgroup whose threads dirtied it; where it does not, writeback is attributed to the root cgroup and your per-VM I/O accounting quietly stops being per-VM. Either way, a guest “only” using its baked 2 GiB can be responsible for considerably more than 2 GiB of host memory while it writes hard. If your admission control models a VM as a fixed number, this is the term it is missing.

The second is that latency becomes bimodal in a way neither side can explain alone. While the host has dirty headroom, guest writes complete at memcpy speed. When the host crosses its own background threshold, every VM on the box sees apparent bandwidth drop at once, and each guest independently concludes its disk has become slow and throttles its writers. Synchronised stalls across unrelated tenants, driven by a threshold none of them can see.

Whose flush is it, anyway?

The third consequence is durability, and the failure mode is silent. When your program calls `fsync()`, the guest kernel writes the file's dirty pages out and issues a flush to the block device — a virtio-blk `FLUSH` request, if the device advertises the feature. What the VMM does with it is a configuration choice, and on Firecracker that is the drive's `cache_type`. With `Writeback`, a guest flush becomes an `fsync()` on the backing file. With `Unsafe` — the default, and the rare option named after its consequences — it is not: depending on version the flush is either never advertised or acknowledged without being performed. Either way, `fsync()` returning zero inside the guest does not mean anything left the host's page cache. Your database thinks it has a durable commit. It has a very fast one.

There is no host-cache-bypass option either. Firecracker's drive object gives you `cache_type` and `io_engine`, not an `O_DIRECT` mode, so the host page cache is unconditionally in the path: double buffering is not something you can turn off here. Read the drive JSON, because this is exactly the kind of fact teams assume rather than check. For an ephemeral sandbox rootfs it genuinely does not matter, since the disk is deleted with the VM — What "persistent" actually means in a sandbox covers why durability is usually the wrong question for a throwaway machine. It matters enormously for a durable volume or a managed database.

import os
import time

PATH = "/work/durability-probe"
PAYLOAD = b"x" * (8 * 1024 * 1024)  # 8 MiB


def timed_write(path: str, do_fsync: bool) -> float:
    t0 = time.perf_counter()
    fd = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_TRUNC, 0o600)
    try:
        os.write(fd, PAYLOAD)
        if do_fsync:
            os.fsync(fd)  # guest page cache -> virtio-blk FLUSH -> ???
    finally:
        os.close(fd)
    return (time.perf_counter() - t0) * 1000.0


print("no fsync: %7.1f ms" % timed_write(PATH, False))
print("fsync   : %7.1f ms" % timed_write(PATH, True))
os.unlink(PATH)

# The honest caveat, and the whole point of the measurement:
#
# The fast number is a memcpy into the guest page cache. The slow number is
# that cache drained through a virtqueue into a userspace process on the
# host. NEITHER number tells you the bytes reached storage.
#
# os.fsync() emits a virtio-blk FLUSH. Whether that becomes an fsync() on
# the backing file is a property of the drive's cache_type, not of your
# program -- and Firecracker's default is literally named "Unsafe". Beneath
# that, the host filesystem has its own dirty pages and its own opinion
# about when to write them, which your guest cannot see or influence.
#
# So a suspiciously small gap between the two numbers is not good news. It
# is evidence that your flush is being acknowledged rather than honoured.

blk-wbt inside the guest: a second control loop

There is another writeback throttle further down, solving a different problem. Block-layer writeback throttling — blk-wbt, exposed as `wbt_lat_usec` in `/sys/block/<dev>/queue/` — watches read completion latency and, when reads start missing the target, reduces the queue depth available to buffered and background writes. `vm.dirty_*` decides how much dirt the system tolerates; blk-wbt decides how much of the queue a writeback flood may occupy while somebody is reading. The kernel enables it by default, and writing `0` turns it off.

Does it help in a guest? Honestly: it depends, and you should measure rather than reason. For leaving it on: the problem it solves is real inside a sandbox. A build writing hard while reading thousands of small files from the same virtual device is the exact read-starvation case, and a read stuck behind 128 queued writeback requests is slow for ordinary reasons. Against: blk-wbt is a latency-feedback controller, and latency through a virtio queue is not a property of storage — it is sometimes memcpy-fast and sometimes NVMe-slow. Stacked on the dirty throttle, that is two controllers acting on one signal neither should trust.

One thing to know before experimenting: blk-wbt is disabled automatically when you select a scheduler that does its own latency management, notably BFQ. So “I set the scheduler and also tuned `wbt_lat_usec`” is sometimes two changes and sometimes one. If you are picking a scheduler at all, The I/O Scheduler Inside a MicroVM Is Almost Always Wrong is the longer argument. My position: `vm.dirty_bytes` is where the leverage is, and without a read-latency win demonstrated A/B on a loaded host you have not earned the right to change this.

Four places a write is throttled, and one symptom

This is what turns a kernel-internals discussion into an operational one. A write can be slowed by several independent mechanisms, owned by different people, configured in different places, all presenting as “the disk is slow.” Guest writeback is the one above: invisible to the host, a property of the template you baked. Firecracker's token-bucket rate limiter sits in the VMM above the host page cache, capping bytes and operations per virtual drive — see Firecracker's Rate Limiter, Explained. Host cgroup I/O control acts at bio submission, and Per-Sandbox Disk I/O Throttling and the Neighbour You Cannot See covers what each of those is blind to.

On PandaStack, each VM gets its own cgroup, `vm-<id>`. The agent writes `cpu.weight` — the proportional share that makes the eight baked vCPUs burst capacity rather than a reservation — and `cpu.max`, a hard ceiling on a 100 ms period. It reads `cpu.stat`'s `usage_usec` for billing CPU-seconds burned, `memory.current` for working-set admission, and `io.stat` for per-VM I/O metrics. So I/O is measured per VM and not capped per VM; I would rather say that than imply a fence that is not there. And note what is absent: `memory.max` is never written. A hard host-side cap underneath a guest size baked into a snapshot does not produce backpressure — it produces a well-behaved VM killed mid-syscall.

`memory.high` is a different matter, and it is the fourth layer. A pressure ladder watches host `MemAvailable` and `/proc/pressure/memory` on a ten-second loop and, when the host is genuinely short, squeezes `memory.high` to roughly 70% of current residency on provably cold VMs only — never a database, never a hugepage-backed guest, never one that is CPU-active or was touched in the last sixty seconds — lifting the caps back as pressure falls. The rung is disarmed unless the host has swap plus zswap, which is what makes it a reclaim lever rather than a kill. It matters here because reclaim and writeback are one machinery seen from two ends: shrink a write-heavy guest's residency and you shrink the dirtyable memory its thresholds are computed from. The guest does not know why its ceiling moved. It just stalls, and reports that the disk is slow.

#!/bin/sh
# On the host. One cgroup per VM: vm-<sandbox-id>.
set -eu
SBX="$1"
cg=$(find /sys/fs/cgroup -maxdepth 4 -type d -name "vm-$SBX" | head -1)
[ -n "$cg" ] || { echo "no cgroup for $SBX"; exit 1; }

# What the agent writes: proportional share, and a hard ceiling.
echo "cpu.weight: $(cat "$cg/cpu.weight")"        # 1..10000
echo "cpu.max   : $(cat "$cg/cpu.max")"           # "<quota> <period_us>"

# What the agent reads: billing, admission, and per-VM I/O.
grep usage_usec "$cg/cpu.stat"
echo "memory.current: $(cat "$cg/memory.current") bytes"
cat "$cg/io.stat"   # per device: rbytes wbytes rios wios dbytes dios

# Read it twice and diff wbytes. That delta is the one fact that separates
# the layers:
#   wbytes climbing while the guest insists the disk is slow
#     -> the writes ARE leaving the VM; look lower (host device, neighbours).
#   wbytes flat while the guest stalls in D state
#     -> they never left. The stall is above the host block layer:
#        guest writeback, or the VMM's rate limiter.
Five throttles, one symptom. Firecracker's drive options come from its own documentation and move between releases -- verify against the version you run.
LayerWhat it limitsWhat it looks likeThe counter that proves itWho should own the setting
Guest `vm.dirty_*` writebackDirty pages the guest accepts before making writers do the writeback themselvesMemory-speed writes, then every writer parks in D state together; throughput sawtooths`Dirty` pinned near `nr_dirty_threshold`; writer `wchan` = `balance_dirty_pages`; `full` rising in `/proc/pressure/io`The template image — baked guest config, not a per-tenant dial
Guest blk-wbt (`wbt_lat_usec`)Queue depth for buffered and background writes, to protect read latency in the guestReads stay crisp under a write storm; bulk throughput below what the path could doNo dedicated counter: `wbt_lat_usec` non-zero plus a read-latency delta measured A/B under loadThe template, and only after measuring
Firecracker drive `rate_limiter`Bytes/s and IOPS per virtual drive per VM, in the VMM above the host page cacheSmooth, boring slowness no guest counter explains; guest queue depth looks healthyInvisible to the guest. Read the drive config; host `io.stat` `wbytes` holds a flat rateThe platform — the noisy-neighbour fence, never tenant-settable
Host cgroup `io.max` / `io.weight`Bytes/s and IOPS per cgroup per host block device, at bio submissionAs above, and it also bites the restore path, the seed pull and the snapshot uploadHost `io.stat` for the `vm-<id>` cgroup, plus `/proc/pressure/io` on the hostThe platform, per host, with the device's measured numbers
Host cgroup `memory.high`Residency of a cold guest's cgroup, forcing reclaim — which shrinks the dirtyable memory its thresholds derive fromA previously idle sandbox is slow on its first write burst after the host got busy`memory.high` below `memory.current` for that `vm-<id>`, plus host `/proc/pressure/memory`The platform, and only with swap plus zswap present and cold-VM gating

The copy-on-write wrinkle, and the benchmark it ruins

One more thing changes the arithmetic, specific to how sandbox rootfs images are made. A sandbox does not get a freshly allocated disk. It gets a copy-on-write clone of the template's rootfs — XFS reflink, or dm-snapshot — which is why create includes a reflink of roughly 4 ms rather than a multi-gigabyte copy. The whole 179 ms p50 create depends on not copying that disk.

So a first write to a shared region costs more than a steady-state overwrite, and not by a little. Under dm-snapshot, a write to a chunk never written before triggers a copy-out: the original chunk is read from the origin and written to the COW device before your write lands. Chunks are typically tens of kilobytes, so a 4 KiB write to a cold chunk becomes a read, plus a chunk-sized write, plus your write. Under XFS reflink, writing into a shared extent allocates new blocks and XFS pads that allocation with the `cowextsize` hint.

Two things follow. A tuning point: the real drain rate early in a sandbox's life is lower than its steady-state rate, so the guest's estimator is most optimistic precisely when the device is least capable. The first thirty seconds of a build are the worst moment to be holding 400 MiB of dirty pages, which is exactly when the default ratio is happiest to let you. And a benchmarking point I have watched smart people lose a week to: on CoW storage, writing a fresh file measures something different from overwriting an existing one, by a margin large enough to reverse a conclusion. Copy-on-Write Rootfs: dm-snapshot vs reflink for MicroVMs has the mechanics of both.

A recipe, and when to leave it alone

For a typical 2 GiB sandbox guest doing interactive or agent-driven work — builds, tests, code execution, a human or a model watching output appear — this is where I would start. A starting point you then measure, not a constant.

  • `vm.dirty_bytes = 67108864` (64 MiB). A hard limit sized to the stall you will accept, not to 20% of whatever RAM the template was baked with — so the 2 GiB and 4 GiB templates get the same policy.
  • `vm.dirty_background_bytes = 16777216` (16 MiB). A quarter of the hard limit. Keeping these two well apart is what makes the throttle a slope instead of a cliff.
  • `vm.dirty_expire_centisecs = 500` (5 s). Thirty seconds of unflushed data is a laptop default; a sandbox may not live thirty seconds.
  • `vm.dirty_writeback_centisecs = 100` (1 s). Wake the flushers more often so the backlog is worked continuously rather than in lumps.
  • Leave blk-wbt at its default. Change `wbt_lat_usec` only with an A/B read-latency measurement on a loaded host in hand.
  • Put it in `/etc/sysctl.d/` in the template image, then re-bake. A sysctl applied by a post-boot script is absent on every restore — and restore is the only boot path that matters after the first spawn.

And when to leave the defaults alone, because tuning is a cost too. If your sandbox's whole output is a few megabytes of JSON, it will never approach any threshold. If the workload is one large sequential write measured purely on throughput, a bigger budget is genuinely better and my recommendation is the wrong trade. And if you are copying these numbers because they appeared in a blog post, you have swapped a default nobody chose for one somebody else chose for a workload that is not yours. The reasoning transfers; the constants are mine.

A related trap: putting scratch space on tmpfs to “avoid writeback entirely.” It does, by charging the data to guest RAM instead. In a 2 GiB guest whose memory ceiling is fixed by the baked snapshot, a tmpfs that grows with your build is not an optimisation. It is an OOM kill with extra steps.

Which layer is stalling you

A diagnostic order, cheapest first. The discipline is doing these before touching a knob, because every layer here answers being tuned with a slightly different version of the same symptom.

  1. In the guest, is anything in `D` state? If not, this is not an I/O stall and you are in the wrong post.
  2. Read that process's `wchan`. If it says `balance_dirty_pages`, you are done: guest writeback, go and set `dirty_bytes`. One `cat`, and the highest-yield check here.
  3. Compare `Dirty` against `nr_dirty_threshold`. Pinned at the ceiling with low `Writeback` means the backlog is not draining; high sustained `Writeback` means it is, and the constraint is below you.
  4. Check `/proc/pressure/io` in the guest. A rising `full` figure is the number to alert on.
  5. On the host, diff `io.stat` `wbytes` for that VM's cgroup. Climbing means the writes are leaving the VM; flat while the guest stalls means the cause is guest writeback or the VMM limiter.
  6. Flat `wbytes` with a `wchan` that was not `balance_dirty_pages` sends you to the drive's `rate_limiter` config — a token bucket produces conspicuously smooth slowness.
  7. Compare `memory.high` with `memory.current`. A cap below residency means something is deliberately reclaiming this guest, and its dirty thresholds shrank with it.
  8. Still unexplained? A real host storage problem: host `/proc/pressure/io`, the other tenants, and whether somebody is pulling a multi-gigabyte seed onto the same device right now.

That last one is the failure that taught me to build the ladder. A host restoring snapshots, streaming memory chunks and pulling template artefacts is doing heavy I/O that belongs to no tenant's cgroup in any way a tenant can see. From inside a sandbox, a neighbour's seed pull is indistinguishable from your own disk being slow — and from a dashboard, both look like a guest sysctl being wrong.

The summary

A Firecracker guest is a complete Linux system with a complete writeback subsystem, and it believes it owns a disk: page cache, flusher threads, two dirty thresholds, two clocks, and a feedback controller in `balance_dirty_pages()` pacing each dirtying task against an estimate of its device's writeout bandwidth. The estimate is a fiction, because what the guest measures is the rate a userspace process on another machine accepts memcpy into a second page cache — enormous, right up until the host starts real writeback.

The consequences are small and specific. Use `vm.dirty_bytes`, not `vm.dirty_ratio`, because 20% of 2 GiB and 20% of 16 GiB are one policy producing an eightfold difference in worst-case stall. Size the limit to the pause you will accept, measured through the real stack on a busy host. Keep the background threshold well below the hard one, shorten the expiry, bake it into the template, and ignore the per-BDI knobs in a single-device guest.

And remember the stack is double-buffered, which is not optional: Firecracker gives you `cache_type` and `io_engine`, not a host-cache bypass, so your guest's `fsync()` is only as durable as the drive configuration beneath it — and the default is named `Unsafe` for a reason worth reading twice. Two kernels, two page caches, no shared understanding of what the storage can do. The guest is extremely confident about a disk that does not exist. The host is also confident, about different things. Which works fine, right up until the afternoon it does not.

Frequently asked questions

What does vm.dirty_ratio actually do, and why is the default wrong in a 2 GiB guest?

`vm.dirty_ratio` (default 20) is the point at which the kernel stops letting a writing task dirty pages freely and makes it participate in writeback itself — the throttle, implemented in `balance_dirty_pages()`. Below it, `write()` is a memcpy into the page cache; above it your process is doing storage I/O, which is why the slowdown presents as a cliff rather than a slope. The default is wrong in a small guest for two reasons. First, it is a percentage of dirtyable memory, so the same number means roughly 400 MiB of backlog in a 2 GiB guest and 3.2 GiB in a 16 GiB guest — one sysctl generating a different policy per template, none of them reviewed. Second, the pause length is computed from an estimate of the device's writeout bandwidth, and in a guest that estimate measures how fast a userspace process on the host accepts memcpy into a second page cache. It is enormous while the host has dirty headroom and collapses when the host starts real writeback, so the controller is most optimistic exactly when it has the least room to correct.

Should I use dirty_bytes or dirty_ratio for microVM guests?

`dirty_bytes` and `dirty_background_bytes`, in any fleet where guest sizes differ. They are two views of one limit — writing the byte form zeroes the ratio form and vice versa, and the kernel keeps exactly one authoritative — so this is a choice of unit, not an extra control. Bytes win because the thing you care about is a duration: how long can one task be stalled while the backlog it created drains? That is bytes divided by bandwidth, and bytes is the half you can set. The rule I use is to size it to the stall I am willing to take rather than to a fraction of RAM: pick a defensible worst-case pause — under a second for an interactive sandbox, several seconds for a throughput-measured batch job — multiply by write bandwidth measured through the whole stack on a busy host, and set the background threshold to about a quarter of the result. Measure on a loaded host, because an idle one will tell you the host page cache is infinitely fast. And accept that a smaller budget really does cost bulk sequential throughput: you are trading tail latency against it, not getting both.

Does fsync() in a microVM guest actually make data durable?

Only if the layers beneath were configured to let it. The guest kernel writes the file's dirty pages out and issues a virtio-blk FLUSH, and what the VMM does with that request is a configuration choice. On Firecracker it is the drive's `cache_type`. With `Writeback`, the device advertises the flush feature and a guest flush becomes an `fsync()` on the backing file. With `Unsafe` — the default, named after its consequences — it does not: depending on version the feature is either not advertised or the flush is acknowledged without being performed. Either way `fsync()` returning zero inside the guest does not mean anything left the host's page cache. There is also no host-cache-bypass option to reach for; Firecracker's drive object gives you `cache_type` and `io_engine`, not `O_DIRECT`, so the host page cache is unconditionally in the path and double buffering cannot be switched off. For an ephemeral sandbox rootfs none of this matters, because the disk is deleted with the VM. For a durable volume or a managed database it matters enormously, and it is exactly the kind of fact teams assume rather than read out of the drive configuration.

Is blk-wbt (wbt_lat_usec) useful inside a guest, or should I disable it?

It depends, and this is a measure-don't-reason situation. blk-wbt watches read completion latency and reduces the queue depth available to buffered and background writes when reads start missing a target. The problem it solves is real inside a sandbox: a build writing hard while reading thousands of small files from the same virtual device is precisely the read-starvation case, and a read stuck behind a hundred queued writeback requests is slow for completely ordinary reasons that do not care whether the device is virtual. The argument against is that blk-wbt is a latency-feedback controller whose target is calibrated against real device behaviour, and latency through a virtio queue is not a property of storage — it is sometimes memcpy-fast and sometimes NVMe-slow depending on the host's dirty state and on neighbours the guest cannot see. Stacked on the dirty throttle you then have two controllers acting on the same traffic from the same bad signal. One practical note: selecting a scheduler that does its own latency management, notably BFQ, disables blk-wbt automatically, so changing both is sometimes one change.

How do I tell whether the guest, the hypervisor or the host is throttling my writes?

In that order, cheapest first. In the guest, check whether anything is in `D` state; if not, it is not an I/O stall. Then read that process's `wchan` — if it says `balance_dirty_pages` you are finished: guest writeback, go and set `dirty_bytes`. That is one `cat` and the highest-yield check in the list. Next compare `Dirty` in `/proc/meminfo` against `nr_dirty_threshold` in `/proc/vmstat`: pinned at the ceiling with low `Writeback` means the backlog is not draining, while high sustained `Writeback` means it is and the constraint is lower down. Export `full` from `/proc/pressure/io` per sandbox — the share of time in which every runnable task was stalled on I/O, which is what a user calls “it froze.” On the host, diff `io.stat` `wbytes` for the VM's cgroup: climbing means the writes are leaving the VM, flat while the guest stalls means the cause is above the host block layer. Then check `memory.high` against `memory.current`, because a cold guest being deliberately reclaimed has smaller dirty thresholds than it did an hour ago.

Keep reading

Related posts

  • The I/O Scheduler Inside a MicroVM Is Almost Always Wrong

    Your guest's block layer is sorting requests by sector number to optimise seeks on a platter that is a file on somebody else's filesystem. Why `none` is the answer, why it still merges, and why readahead is the knob that actually moves the number.

  • Firecracker Block Device Cache Modes Explained

    A Firecracker drive has a cache_type field with two values, and the choice is a bet about a host crash. The default respects the guest's flushes so a durable disk survives a power loss; Unsafe lets those flushes ride the host page cache for speed. One is right for a throwaway CoW rootfs, the other for a database volume — and picking wrong is how you lose data you promised to keep.

  • How much memory does your Postgres actually need?

    The advice 'give Postgres as much RAM as your data' is expensive and usually wrong. What matters is the working set — and there's a query that tells you whether yours fits.

  • Firecracker vCPU Hotplug, and How CPU Scaling Actually Works

    Somebody always asks for "more CPU" on a running microVM. Firecracker's answer is no — vcpu_count is fixed before InstanceStart and frozen into every snapshot. The good news: the lever you actually wanted was never vCPU count in the first place.

  • Conntrack Table Limits: The Shared Resource Under Your MicroVM Fleet

    Your guests have separate kernels, separate namespaces and separate subnets. They share one fixed-size hash table, and one port scanner can fill it for everybody.

More in Internals · See PandaStack benchmarks

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.