Readahead Inside a microVM: The Kernel Guessing Wrong, 4 KiB at a Time
Someone will tell you, with total confidence, that the first thing you do to a fresh Linux box is raise `read_ahead_kb`. I did that to a microVM template once. The cold boot did not get faster, so I raised it again, because that is what engineers do when a knob disappoints them. It got worse. Not dramatically worse, just reliably and repeatably worse, in the direction that means you have misunderstood something rather than mismeasured it.
What I had misunderstood is that there is no disk. I build PandaStack, an open-source Firecracker microVM platform, and the thing I had been tuning was the guest's opinion about a block device that is, in reality, a file on a host that has its own page cache and its own opinions. I had taught one layer to speculate harder without knowing what the layer underneath it was already doing.
This post is about that stack: what readahead actually is, what it costs in a guest specifically, and the one place it genuinely moves cold start. The honest headline up front, because the rest of the internet will not give it to you: readahead tuning is a lever on the cold path, and the bigger lever is not taking the cold path at all.
What readahead actually is
When a process reads a file, the kernel does not fetch only the pages it was asked for. The page-cache readahead machinery sits in front of every buffered read and tries to be useful: it watches the access pattern, keeps a window, and speculatively fetches pages beyond the request when the pattern looks sequential.
The state is per open file, not per file and not per process. In 5.10, which is the guest kernel we ship, it is a small struct hanging off `struct file` with a window start, a window size, an async-size marker, and the per-file cap in `ra_pages`. The loop is simpler than its reputation:
- A read arrives. If its offset continues where the last one left off, the kernel treats the stream as sequential and issues the request plus a window beyond it.
- The window grows on continued hits -- roughly doubling, with a faster initial ramp -- and is clamped at `ra_pages`, which is where your `read_ahead_kb` setting lands.
- Inside each window the kernel flags one page as the readahead marker. When the reader touches that page, the next batch is issued asynchronously, so the device stays busy ahead of the reader instead of stalling at the window edge.
- If an offset arrives that continues nothing the kernel has recorded, the window collapses: it reads what was asked for and nothing more. There is a secondary heuristic that inspects what is already cached to rescue interleaved readers, because two sequential streams sharing one descriptor look exactly like random access.
The controls, from blunt to precise. `/sys/block/<dev>/queue/read_ahead_kb` is the per-device cap in kibibytes, and it is the default every newly opened file on that device inherits. `blockdev --getra /dev/vda` prints the same number in 512-byte sectors, so a stock 128 KiB device reports 256 and people conclude they have found two knobs. They have found one knob and two units.
Per file range, `posix_fadvise` is the precise instrument. `POSIX_FADV_SEQUENTIAL` doubles this descriptor's cap, `POSIX_FADV_RANDOM` sets it to zero so reads fetch exactly what was requested, `POSIX_FADV_NORMAL` puts it back, and `POSIX_FADV_WILLNEED` is the odd one out: it performs real asynchronous I/O for the range right now rather than setting a policy. `POSIX_FADV_DONTNEED` drops a range's clean pages, which makes it the civilised way to run a cold-cache benchmark on a shared box. For mappings the equivalents are `madvise` with `MADV_SEQUENTIAL`, `MADV_RANDOM` and `MADV_WILLNEED`.
Most of a cold start is mmap faults, not read() calls
This is the part that costs people an afternoon. A dynamic linker does not `read()` a shared object; it maps it and takes page faults. A Python interpreter loading an extension module does the same. So the readahead code that runs during a cold start is mostly the fault path, and the fault path has different heuristics from the read path.
In 5.10's `mm/filemap.c`, the mmap readahead is readaround rather than a forward window: on a miss it starts the window half of `ra_pages` BEFORE the faulting page, on the theory that code and data get touched in both directions. It checks the VMA flags first, so `MADV_RANDOM` on a mapping disables it outright, which `posix_fadvise` on the underlying descriptor will not do. And it keeps a miss counter per file: once a file accumulates more than a hundred readaround misses, the kernel gives up on that file and stops trying. The kernel has already conceded the argument you are about to have with it.
There is a neighbouring mechanism that is easy to confuse with readahead and is not: fault-around. On a minor fault the kernel maps already-resident neighbouring pages into the page table so you do not fault on them individually. Default window is 64 KiB, readable in debugfs. It issues no I/O at all, so it is free in a way readahead never is -- and it means some of the cold-start fault reduction you might credit to readahead was actually this.
Readahead is a bet, not a feature
A hit turns N small I/Os into one large one. That is a real and substantial win: fewer requests, fewer completions, better device-queue behaviour, and on a guest, fewer trips across the virtqueue. A miss reads bytes nobody wanted. Those bytes evict pages somebody did want, consume I/O bandwidth, and in a guest cost VM exits that the useful work then has to queue behind.
The kernel's default is a defensible compromise tuned for a world where a block device is a real device, seeks are expensive, and a wrong 128 KiB is cheap insurance against a right one you did not issue. That is a good default for a laptop. It is a questionable default for a microVM whose disk is a file, where the seek you were insuring against does not exist and the wrong bytes land in a page cache whose size was frozen at template bake time.
Readahead is not an optimisation. It is a wager the kernel places on your behalf, with your memory, at odds it inferred from the last few reads.
There is no disk: what a guest read becomes
Follow one 4 KiB guest read all the way down. The guest filesystem issues a bio. The guest block layer merges what it can -- and this is why readahead helps in a guest when it hits, because a 128 KiB readahead batch becomes one multi-segment request rather than thirty-two separate ones. The virtio-blk driver places a descriptor chain on a virtqueue and writes the notification register. That write traps to the VMM: an exit, the first cost that has no analogue on bare metal.
Firecracker then reads the backing file on the host, through the host's own filesystem and the host's own page cache, with its own readahead window over that file. On our hosts the store is XFS and the guest rootfs image is ext4, so you have two filesystems with two sets of metadata-readahead heuristics stacked on each other. The completion comes back, an interrupt is injected, the guest driver finishes the request. For a hot read where the host already has the page, the data motion is trivial and the exit is most of the cost.
Which produces the thing nobody warns you about: the same bytes can sit in the guest's page cache and the host's page cache at the same time. On a host packing many microVMs, the host copy is arguably the more valuable one, because it is shared across every guest reading the same template image -- our rootfs clones are XFS reflinks of one template file, so identical blocks are literally the same host pages. The guest copy is private to one guest and consumes RAM that guest cannot grow. Readahead in the guest manufactures the less valuable copy, at the cost of exits, and bills it to the budget you have less of.
The cold-start shape, and why it is near-worst-case
Here is the honest shape of a cold start. A first spawn of a template with no baked snapshot is a full cold boot, about 3 s, and a large fraction of the reads in that window are scattered, small and metadata-heavy. The dynamic linker opens shared objects one at a time. An interpreter walks `sys.path` and stats hundreds of directories that do not exist before finding the one that does. A single import pulls in a tree of modules spread across the filesystem. Package metadata, locale files, CA bundles, timezone data.
That pattern is close to the worst case for sequential readahead, and not because it is random in the textbook sense. Within any one file the reads often ARE sequential, so the window dutifully ramps up. Then the file ends, the next access is somewhere else entirely, and the window is thrown away. You pay the ramp over and over and collect the payoff almost never. Meanwhile a meaningful share of the I/O is not file data at all but inode and directory blocks, which is governed by the filesystem's own metadata readahead -- on ext4, `inode_readahead_blks`, default 32 -- and not by `read_ahead_kb` at all.
Now the contrast that makes this post worth writing. Our actual create path is not a cold boot. Every create restores a baked Firecracker snapshot: p50 179 ms, p99 around 203 ms, with `/snapshot/load` itself in the 49-80 ms range. On that path the pages the guest is about to touch are already in the memory image, because they were resident when the template was baked. The interesting fault stream is memory, not disk. The import storm that dominated the cold boot does not read the device at all, because the interpreter's bytes are in a page cache that was snapshotted along with everything else.
So: readahead tuning is a lever on the cold path. The cold path is the first spawn of a template, a pull of an artifact that is not local yet, a workload reading data that was never resident at bake time. For everything else, the lever that matters is the one that removes the cold path, and we wrote about that separately at How to Optimize MicroVM Cold Start and The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms. I would rather you believe that than believe I tuned a sysfs file into sub-second boots.
You can watch the difference yourself with the guest's own block counters. Run an import storm on a cache-dropped sandbox, then run it again on a sandbox restored from a snapshot taken after the storm had already warmed the cache:
# Which path does read_ahead_kb actually move? Run the identical workload on
# the two paths a microVM platform has -- cold, and restored -- and read the
# guest's own block counters instead of a stopwatch.
from pandastack import Sandbox
STORM = "\n".join([
"import email, json, logging, sqlite3, decimal",
"import xml.etree.ElementTree, urllib.request",
"import concurrent.futures, unittest, pydoc",
])
PROBE = "\n".join([
"#!/bin/sh",
"D=$(lsblk -no PKNAME \"$(findmnt -no SOURCE /)\")",
"S=/sys/block/$D",
"A=$(awk '{print $3}' \"$S/stat\")", # field 3 = sectors read
"timeout --kill-after=5s 60 python3 /tmp/storm.py",
"B=$(awk '{print $3}' \"$S/stat\")",
"echo ra=$(cat $S/queue/read_ahead_kb)KiB disk=$(( (B - A) / 2 ))KiB",
])
sbx = Sandbox.create(
template="base", # 4 GiB / 8 vCPU, baked. Firecracker cannot
# change vCPU or RAM at restore, so a cpu= or
# memory_mb= here is overridden to the baked
# values -- and that is exactly why a page
# cache you inflate with readahead is competing
# for a fixed budget.
ttl_seconds=900, # an IDLE timeout, not a wall clock
metadata={"demo": "readahead"},
)
try:
sbx.filesystem.write("/tmp/storm.py", STORM)
sbx.filesystem.write("/tmp/probe.sh", PROBE)
sbx.exec("chmod +x /tmp/probe.sh")
# Path A, the cold shape. This sandbox itself came up by snapshot restore
# (p50 179 ms), so its page cache holds whatever was resident at bake
# time. Drop it and you are looking at the import storm the way a true
# first-spawn cold boot sees it: scattered, small, metadata-heavy.
sbx.exec("sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'")
print("cold guest cache :", sbx.exec("/tmp/probe.sh").stdout.strip())
# Now sweep the window on the cold shape. Note that exec() one-shot does
# not reliably enforce a timeout_seconds argument, so the bound lives
# in-guest as timeout(1) inside probe.sh.
for ra in (0, 128, 1024):
# vda is the rootfs device in our first-party templates; discover
# it with lsblk if you have built your own.
sbx.exec(f"sh -c 'echo {ra} > /sys/block/vda/queue/read_ahead_kb'")
sbx.exec("sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'")
print(" sweep :", sbx.exec("/tmp/probe.sh").stdout.strip())
# Path B. Run the storm once so the bytes are resident, THEN snapshot.
# The page cache is guest RAM and guest RAM is what vm.mem is, so those
# pages travel inside the memory image.
sbx.exec("/tmp/probe.sh")
snap = sbx.snapshot() # synchronous; returns a snapshot id
finally:
sbx.kill() # teardown is kill(); no .delete()
warm = Sandbox.create(from_snapshot=snap, metadata={"demo": "readahead"})
try:
# Same workload, same read_ahead_kb, near-zero sectors off the device.
# There is nothing left for readahead to be right or wrong about, which
# is the entire point of the post.
print("restored, warm :", warm.exec("/tmp/probe.sh").stdout.strip())
finally:
warm.kill()
The second number should be close to zero, and that is the whole argument. There is a real cost, though, and it is the honest other side: those warm pages are guest RAM, guest RAM is what the memory image contains, so a deliberately warm cache makes `vm.mem` bigger -- which is exactly the problem Guest page cache: why it bloats microVM snapshots argues for solving by dropping the cache before you snapshot. Both posts are right. Warm what your workload touches on the first request; drop what it does not.
The prefetch trace: readahead that works, because it is not guessing
This next part is an analogy, and I want to label it as one before I draw a lesson from it. It is about memory, not disk, and the mechanisms are not interchangeable.
On a streamed restore we do not download the guest's memory image before booting it. Firecracker's memory is served by a userfaultfd handler: the guest touches a page, the kernel raises a fault, our handler maps that fault to an offset in the snapshot and fetches the containing 4 MiB chunk from object storage over an HTTP Range GET, then installs it. A bitmap baked alongside the image records which chunks are entirely zero, so absent chunks are zero-filled with no fetch at all. And a prefetch trace, recorded at bake time, lists the chunks the guest actually touched during startup; on restore it is replayed in the background so the faults the guest is about to take become cache hits instead of round trips.
That prefetch trace is readahead for memory. Same shape: fetch something before it is asked for, so the asking is cheap. It works where blind readahead does not for exactly one reason, and the reason is not that the chunks are bigger or the transport smarter. It is that the trace is a RECORDED access pattern rather than an inferred one. We are not predicting the pattern from the last few faults; we watched this exact workload do this exact thing at bake time and wrote it down.
The lesson generalises cleanly, and it is the one sentence from this post worth keeping: if you know the pattern, prefetch it explicitly -- `FADV_WILLNEED` on the range, a warm cache baked into your image, a trace you recorded. If you do not know the pattern, a bigger window is not more knowledge. It is a bigger bet.
The decision table
Workload shape decides this, not the storage medium and definitely not a blog post's recommended value. The column that matters is the last one.
| Workload shape | Window | Mechanism | Why |
|---|---|---|---|
| Sequential bulk read: `dd`, `tar -x`, restoring a database dump, streaming a dataset | Large -- hundreds of KiB up, measured not guessed | `read_ahead_kb` on the template image, or `FADV_SEQUENTIAL` per file | Every hit folds dozens of 4 KiB requests into one virtqueue round trip and one completion interrupt. This is the case readahead was built for. |
| Interpreter import storm or dynamic-linker walk on a first boot | Leave the default; the win is somewhere else | Fewer files (zipapp, vendored tree, precompiled bytecode), or a snapshot baked warm | Within-file reads ramp the window, then the next file throws it away. The mmap path also stops trying after enough misses, so a bigger cap mostly buys wasted bytes. |
| A store doing its own buffer management: Postgres, an LSM engine, a vector index | Small, or off for data files | `FADV_RANDOM` on data files; leave the sequential-scan and WAL paths alone | It knows its own access pattern better than the kernel's last-few-reads inference. Anything with its own cache should be told to stop competing for a page cache it is already double-buffering. |
| Random small reads over a file much larger than guest RAM | Off | `MADV_RANDOM` for mappings, `FADV_RANDOM` for `read()` -- they are different switches | Every speculative page is pure eviction pressure on a page cache that cannot grow, because the guest's RAM was baked into the template. |
| Read a known range once, immediately: a model file, a wasm module, a big config | Not a window problem | `FADV_WILLNEED` on the exact range, then read it | One explicit large request beats a window that has to ramp to the same size, and you already know the answer the kernel would spend four reads inferring. |
| Write-heavy build step: an install, a bundler, a compiler | Irrelevant | Tune writeback instead -- see Guest Writeback Throttling: Why dirty_ratio Is Wrong for a 2 GiB MicroVM | Readahead never touches the write path. The stall you are chasing is dirty-page writeback, not speculation. |
The mechanism column is not decoration. A mapping ignores `posix_fadvise`, and a `read()` loop ignores `madvise`, and the number of production tuning changes that pick the wrong one and then get credited with an improvement from somewhere else is not small.
# fadvise_demo.py -- how to tell the kernel to stop guessing, per file.
#
# POSIX_FADV_RANDOM -> this file's ra_pages = 0. Reads fetch exactly
# what was asked for. Nothing speculative.
# POSIX_FADV_SEQUENTIAL -> this file's ra_pages = 2x the device default.
# POSIX_FADV_WILLNEED -> starts a real asynchronous read of the range NOW.
# The only hint that performs I/O instead of
# setting a policy.
# POSIX_FADV_DONTNEED -> drops this range's clean pages. Your per-file
# drop_caches, and the honest way to run a
# cold-cache benchmark without nuking the box.
import mmap, os, random
PATH = "/var/tmp/index.bin"
RECORD = 4096
def point_lookups(hint, n=20_000):
# Random 4 KiB reads through read()/pread(): the fadvise path.
fd = os.open(PATH, os.O_RDONLY)
try:
size = os.fstat(fd).st_size
# offset=0, length=0 means "the whole file". The hint lives on this
# struct file, so it dies with this descriptor: a sibling process
# that opens the same path gets the device default back. There is no
# way to make FADV_RANDOM sticky for a path, which is why people keep
# reaching for read_ahead_kb and hitting everything on the device.
os.posix_fadvise(fd, 0, 0, hint)
rnd = random.Random(1234)
for _ in range(n):
off = rnd.randrange(0, size - RECORD) & ~(RECORD - 1)
os.pread(fd, RECORD, off)
finally:
os.close(fd)
def mapped_lookups(n=20_000):
# The same pattern through mmap: a DIFFERENT switch entirely.
fd = os.open(PATH, os.O_RDONLY)
try:
mm = mmap.mmap(fd, 0, access=mmap.ACCESS_READ)
finally:
os.close(fd) # the mapping holds its own ref
# MADV_RANDOM sets VM_RAND_READ on the VMA. In mm/filemap.c the fault
# handler checks that flag FIRST and skips readaround altogether.
# posix_fadvise would not help here: a page fault never consults the
# read() path's per-descriptor window. If your hot loop is mmap-based --
# and a dynamic linker, a vector index and most embedded key-value
# stores are -- madvise is the lever and fadvise is a no-op.
mm.madvise(mmap.MADV_RANDOM)
rnd = random.Random(1234)
size = len(mm)
for _ in range(n):
off = rnd.randrange(0, size - RECORD) & ~(RECORD - 1)
mm[off:off + 64] # touch one page
mm.close()
def warm_then_read(path):
# When you DO know the pattern: one explicit request beats a ramp.
fd = os.open(path, os.O_RDONLY)
try:
os.posix_fadvise(fd, 0, 0, os.POSIX_FADV_WILLNEED)
return os.read(fd, os.fstat(fd).st_size)
finally:
# Hand the pages back rather than letting them squat in a page cache
# whose size was baked into the template.
os.posix_fadvise(fd, 0, 0, os.POSIX_FADV_DONTNEED)
os.close(fd)
if __name__ == "__main__":
for name in ("NORMAL", "SEQUENTIAL", "RANDOM"):
point_lookups(getattr(os, "POSIX_FADV_" + name))
print("pread ", name, "done")
mapped_lookups()
print("mmap MADV_RANDOM done")
`O_DIRECT` is the sledgehammer, and it is usually wrong here. It bypasses the guest page cache entirely, which sounds like exactly the fix for double caching until you notice which cache it bypasses: the guest's, the private one, while the host keeps its shared copy and keeps doing its own readahead over the backing file. You have removed the layer you could tune and kept the layer you cannot see from inside. You also inherit strict alignment requirements, lose block-layer merging of your small reads, and turn every access into a virtqueue round trip with no caching to absorb repeats. It is the right call for a process with a large, well-managed buffer pool of its own that genuinely wants the kernel out of the way. It is the wrong call as a reflex.
Measuring it, so you can stop arguing about it
Everything above is mechanism. Here is how to see it on your own files, which is the only readahead number you should trust. Three rules before you run anything. Drop caches on BOTH sides, or you are measuring the host's memory and calling it the guest's readahead. Count bytes, not seconds, because bytes are attributable and seconds are not. And compute an amplification ratio -- bytes off the device divided by bytes you asked for -- because that single number is the whole story and it needs no baseline from me.
#!/usr/bin/env bash
# Run this INSIDE the guest. It finds the rootfs block device, prints the
# readahead cap three different ways, then measures how many bytes actually
# came off the virtqueue for a workload you can account for.
set -euo pipefail
SRC=$(findmnt -no SOURCE /) # /dev/vda1, /dev/mapper/..., etc
DEV=$(lsblk -no PKNAME "$SRC" 2>/dev/null || true)
DEV=${DEV:-$(basename "$SRC")} # vda
Q=/sys/block/$DEV/queue
# Three views of ONE number. read_ahead_kb is kibibytes. blockdev --getra is
# 512-byte sectors, so it reads exactly double. Both write bdi->ra_pages, so
# setting one changes the other -- a surprising number of runbooks set both
# and then congratulate themselves on a "belt and braces" fix.
cat "$Q/read_ahead_kb" # stock kernels: 128
blockdev --getra "/dev/$DEV" # stock kernels: 256 sectors
cat "$Q/rotational" # 0 -- virtio_blk sets NONROT
# ext4 keeps its own metadata readahead, and THIS is the one that moves during
# an import storm, because an import storm is mostly inode and dentry reads.
cat "/sys/fs/ext4/$(basename "$SRC")/inode_readahead_blks" \
2>/dev/null || true # 32 blocks
# Optional: the mmap fault-around window. Not readahead -- it issues no I/O,
# it only maps already-resident neighbours into the page table to save
# follow-up faults. Worth knowing it exists before you blame readahead.
cat /sys/kernel/debug/fault_around_bytes 2>/dev/null || true # 65536
amplification() { # amplification <label> <cmd...>
local label=$1; shift
sync; echo 3 > /proc/sys/vm/drop_caches # GUEST cache only. The host's
# page cache over the backing
# file is still warm. See below.
local a b sectors
a=$(awk '{print $3}' "/sys/block/$DEV/stat") # field 3 = sectors read
"$@" >/dev/null 2>&1 || true
b=$(awk '{print $3}' "/sys/block/$DEV/stat")
sectors=$(( b - a ))
printf '%-24s %9d sectors %8d KiB off the device\n' \
"$label" "$sectors" "$(( sectors / 2 ))"
}
STORM='import email, json, logging, sqlite3, decimal, pydoc, unittest'
dd if=/dev/urandom of=/var/tmp/blob.bin bs=1M count=256 conv=fsync \
status=none
for ra in 0 32 128 512 2048; do
echo "$ra" > "$Q/read_ahead_kb"
amplification "ra=${ra}KiB import" python3 -c "$STORM"
amplification "ra=${ra}KiB seqread" dd if=/var/tmp/blob.bin of=/dev/null bs=4k
done
# The arithmetic: KiB-off-the-device divided by KiB-you-asked-for.
#
# seqread asked for exactly 262144 KiB, so the ratio should sit near 1.0 at
# every window. That is readahead being free when it is right: the bytes
# were wanted, they just arrived in bigger batches.
#
# import asked for a few hundred KiB of .py, .so and inode bytes. Whatever
# number shows up above that is the kernel's guess, billed to your guest's
# RAM and your host's I/O queue. Watch it grow with ra and ask yourself
# which column got faster.
#
# Cross-check system-wide with /proc/vmstat pgpgin. Do NOT trust its unit
# from memory -- it has been sectors in some trees and pages in others.
# Calibrate it once against a dd of known size, then use deltas only.
grep -E '^(pgpgin|pgpgout|pgfault|pgmajfault)' /proc/vmstat
# Persist it for the template, because this is an image property. A udev rule
# survives reboots and re-attachments; an echo into sysfs does not.
cat >/etc/udev/rules.d/60-virtio-readahead.rules <<'RULE'
ACTION=="add|change", SUBSYSTEM=="block", KERNEL=="vd[a-z]", \
ATTR{queue/read_ahead_kb}="128"
RULE
udevadm control --reload
udevadm trigger --subsystem-match=block --action=change
The host side needs its own pass. Dropping the guest's cache leaves the host's copy of the backing file warm, so a guest-only measurement shows you amplification across the virtqueue but tells you nothing about real storage I/O. On a dedicated test host, drop the host page cache too; on a shared one, use `posix_fadvise` with `POSIX_FADV_DONTNEED` against the specific backing file rather than clearing everything. Then watch the host's read volume on that file -- `iostat -x` on the store device, or the per-process counters in `/proc/<firecracker pid>/io` -- while the guest runs the same workload at two different windows. If guest-side bytes go up and host-side bytes do not, the extra reads were being served from the host's cache and you were paying only in exits and guest RAM. That is the common case, and it is why the regression is usually modest, consistent and maddening rather than dramatic.
One trap in the tooling: `/proc/vmstat`'s `pgpgin` and `pgpgout` are useful cumulative counters whose unit has not been stable across kernel versions. Do not state a KiB figure derived from them without calibrating first -- read a file of known size, see what the delta is, and then use deltas only. `/sys/block/<dev>/stat` field three is sectors read in 512-byte sectors and is the more dependable primary, with `vmstat 1` and `iostat -x 1` inside the guest for the shape over time.
What we do, and where I would not reach for this
What we actually do about the guest block read path is mostly structural rather than tuned. Rootfs clones are copy-on-write -- XFS reflink, or dm-snapshot -- so many guests share host pages for identical blocks, which makes the host's cache the efficient place for template bytes to live. Memory on a streamed restore comes from object storage in 4 MiB chunks with a zero-chunk bitmap and a recorded prefetch trace. Templates ship with their block queue configured for a device that is a file, which matters more than the readahead number: see The I/O Scheduler Inside a MicroVM Is Almost Always Wrong.
Where I would not reach for readahead tuning. If your create path is a snapshot restore, this is a third-order lever and the honest advice is to go look at the memory fault stream instead. If you are chasing a build that stalls, you are probably looking at writeback, not reads. If your workload is GPU-bound, we are not the platform at all: PandaStack has no GPUs and no GPU passthrough, and no amount of block-layer tuning changes that. And if you want the newer readahead work in the kernel -- the folio-era rewrites and the improved large-folio behaviour -- our guest kernel is 5.10, so you are tuning the generation of this code I described above, not the current one. Verify the function names and the heuristics against the kernel you are actually running before you quote me in a review.
I would also not spend a sprint on this before checking the dumb things: whether the file is sparse and you are faulting in holes, whether the workload opens the same file thousands of times and gets a fresh window each time, whether a container-ish init is reading a locale archive you could have deleted. Readahead is a satisfying thing to tune because it has a number in a file. That is not the same as it being your bottleneck.
The bottom line
Readahead is the kernel inferring your future from your recent past, and in a guest it does that on top of a host doing the same inference with better information and cheaper mistakes. Raising the window is not a performance setting, it is a larger wager. The workloads where it pays are the ones where you already know the pattern is sequential, and in that case you could have said so explicitly and skipped the ramp.
The cold-start angle is the one that reframed it for me. The scattered, metadata-heavy read pattern of a first boot is close to the worst thing you can hand a sequential predictor, which means readahead tuning on the cold path is damage control on a path you should be trying to eliminate. We eliminated it by restoring a snapshot in 179 ms instead of booting in 3 s, and on that path the disk reads we spent years thinking about mostly do not happen. The best readahead is the read you never issued because the page was already there.
So: measure your amplification ratio, decide per workload shape rather than per device, use `fadvise` and `madvise` to say what you know, and treat a recorded pattern as categorically different from a guessed one. And if someone tells you the first thing to do with a new Linux box is raise `read_ahead_kb`, ask them what the disk is.
Frequently asked questions
Is raising read_ahead_kb inside a microVM ever the right call?
Yes, for one clearly identifiable workload shape: large sequential reads of files, where you can account for the bytes you asked for and the amplification ratio comes out near one. Restoring a database dump, extracting a large archive, streaming a dataset through a pipeline, a single-threaded reader walking a multi-gigabyte file front to back. In a guest the payoff is actually bigger than on bare metal, because the guest block layer merges a readahead batch into one multi-segment virtio request, so you save not just I/O operations but the notification exit and the completion interrupt that each separate request would have cost. The conditions matter, though. Only one or two such readers should be active, because several interleaved sequential streams defeat the sequentiality detection and the kernel's rescue heuristic is not magic. The bytes must fit in a page cache whose size was baked into the template and cannot grow. And the file should not already be resident in the host's page cache for free, which on a copy-on-write template clone it often is. Set it where it belongs, which is per file with POSIX_FADV_SEQUENTIAL in the code doing the bulk read, rather than per device where it also applies to every random reader on the same disk. If you must set the device cap, set it in the image as a udev rule so it survives reboots, and then go and measure the amplification ratio rather than trusting the change.
Why did making the readahead window bigger make my guest slower?
Almost always one of three things, and often all three at once. First, double caching: the bytes you speculatively fetched now exist in both the guest page cache and the host page cache, and the guest copy is the less useful one because it is private to one VM while the host's copy can be shared across every guest reading the same template image. You spent guest RAM to duplicate something you already had. Second, eviction: a guest's RAM is fixed at template bake time, because Firecracker cannot change vCPU or RAM at snapshot restore. A bigger readahead window does not get more memory, it takes memory from the hot working set that was already resident, so you converted cache hits into faults somewhere you were not measuring. Third, exits: every speculative batch crosses the virtqueue, which means a notification that traps to the VMM and a completion that injects an interrupt, and useful work queues behind that. On bare metal a wrong readahead costs bandwidth you had spare. In a guest it costs a round trip through the hypervisor. The reason this regression is maddening rather than obvious is that it is usually modest and consistent: the host's cache absorbs the extra reads, so the storage device never shows the problem, and all you see is a workload that got slightly worse for no visible reason.
Does readahead tuning help a snapshot-restore create path at all?
Barely, and understanding why is more useful than the tuning. A create on PandaStack restores a baked Firecracker snapshot, p50 179 ms and p99 around 203 ms, with the /snapshot/load step itself around 49-80 ms. The pages the guest is about to touch were resident when the template was baked, so they are inside the memory image rather than on the virtual disk. The interesting fault stream on that path is memory, not block I/O: for a local restore it is copy-on-write memory faults against a mapped image, and for a streamed restore it is userfaultfd faults served from object storage in 4 MiB chunks. Neither consults read_ahead_kb, because neither goes through the guest's block layer. Readahead tuning is a lever on the cold path, and the cold path is the first spawn of a template with no baked snapshot (about 3 s of real cold boot), an artifact that is not local on the host yet, or a workload reading data that was never resident at bake time. If your workload's disk reads happen on every single request rather than once at startup, then yes, this matters and you should go measure your amplification ratio. If they happen at startup and startup is a restore, the thing to tune is what your bake captures, not what your kernel guesses.
Should I use O_DIRECT in the guest to avoid the double-caching problem?
Usually not, because it bypasses the wrong cache. O_DIRECT removes the guest page cache from the path, which does eliminate the duplicate copy and does eliminate guest readahead. What it does not touch is the host's page cache, because the VMM is reading the backing file with ordinary buffered I/O on the host side, with the host's own readahead over that file. So you have removed the layer you could see and tune from inside the guest, and kept the layer you cannot. You also inherit the costs: strict alignment requirements on offset, length and buffer, no block-layer merging of your small reads into larger requests, and no caching at all to absorb repeated access, which means every read becomes a virtqueue round trip with an exit. For a process with its own large, well-managed buffer pool that genuinely wants the kernel out of the way, that trade can be correct, and it is why database engines offer the option. As a general fix for double caching it is a sledgehammer that misses. The cheaper tool for the same intent is POSIX_FADV_RANDOM to stop the speculation while keeping the cache, or POSIX_FADV_DONTNEED after a one-shot read to hand the pages back. And note separately that O_DIRECT in a guest says nothing about durability: whether a write has really reached stable storage depends on the VMM's cache mode on the host, not on your open flags.
How is a UFFD prefetch trace different from just making readahead more aggressive?
The difference is information, not size, and it is the point of the comparison. Kernel readahead is an inference: it looks at the last few accesses through one descriptor and extrapolates. It has no idea what your program is about to do, so a bigger window is simply a larger bet on the same guess, which is why it pays on a sequential scan and loses on an import storm. A prefetch trace is a recording. At bake time we observe which 4 MiB chunks of the memory image the guest actually touches during startup and write that list down; at restore we replay it in the background so the faults the guest is about to take find their chunks already local. There is no prediction involved, because the workload already told us the answer once. The accompanying zero-chunk bitmap is the same idea taken further: chunks that are entirely zero are recorded as such and zero-filled on fault with no fetch at all, which is not a faster guess but the elimination of a fetch that would have been pure waste. The transferable lesson for block I/O is to move your tuning from the first category to the second. If you know a file will be read in full, say so with POSIX_FADV_WILLNEED on the range instead of letting a window ramp up to the same size over four reads. If you know access is random, say so with MADV_RANDOM or POSIX_FADV_RANDOM instead of hoping the heuristic notices. A recorded or declared pattern beats an inferred one every time, and the inferred one is the only thing read_ahead_kb can give you.
Keep reading
- The I/O scheduler inside a microVM is almost always wrong — The other half of the guest block queue, and why `none` is the right scheduler when the disk is a file.
- Firecracker block device cache modes — What the host does with the backing file, which is the layer your guest readahead is guessing against.
- Guest page cache: why it bloats microVM snapshots — The cost of a warm cache at bake time, and the argument for dropping it before you snapshot.
- Snapshot memory prefetch and the working set — The recorded-pattern idea applied to memory, which is where it actually pays.
- The snapshot-restore boot path — What happens in those 179 ms, and why almost none of it is block I/O.
Related posts
- Guest Writeback Throttling: Why dirty_ratio Is Wrong for a 2 GiB MicroVM
The guest is extremely confident about a disk that does not exist. The host is also confident, about different things. Between those two confidences live your tail latencies.
- virtio-blk Discard and TRIM in Firecracker, Explained
A job downloads 6 GB, does its work, deletes it, and the host-side image is still 6 GB. Deleting a file is a metadata edit in the guest's own allocator; the host never hears about it unless somebody issues a discard. Every link in that chain fails silently.
- kvm-clock vs TSC: How a Firecracker Guest Tells the Time, and What a Snapshot Does to It
A guest clock is not one thing. It is a rated list of clocksources, a shared memory page the host writes into, and a userspace fast path that silently becomes a trap when you pick wrong. Then you snapshot it, and the whole stack starts lying in a very specific way: TLS first, loudly.
- virtiofs vs virtio-blk: How Files Actually Get Into a MicroVM
One gives the guest a disk it owns. The other gives it a window onto a directory the host owns. That single difference decides whether you can fork a machine in 400ms, who parses guest-controlled input, and what a multi-tenant escape looks like.
- Why the disk under your sandbox fleet decides your boot time
People pick a sandbox platform on features and then get bitten by storage hardware. A reflink clone is metadata-only and nearly free; every copy-on-write byte afterwards is a real read-modify-write against a real device. Under 50 concurrent restores, that device is either local NVMe or it is your bottleneck.
More in Snapshots & forking · See Thaw: sub-second cold restore
49ms p50 cold start. Fork, snapshot, and scale to zero.