all posts

EROFS vs SquashFS for a Read-Only microVM Rootfs: Compression You Pay For on Every Page Fault

Ajay Kumar··11 min read

I spent a week last year convinced I had found free money. Our `base` template rootfs is an Ubuntu 24.04 tree with four language runtimes pre-warmed in it, and essentially none of it is ever written. The guest reads `/usr/lib`, it reads `site-packages`, it reads a hundred megabytes of shared objects, and it writes to `/tmp` and `/home` and nowhere else that matters. A filesystem that is 95% read-only and is stored as plain uncompressed ext4 looks, at first glance, like a rounding error someone forgot to optimise.

So: make the read-only part a compressed read-only filesystem, put a thin writable layer on top, and ship a fraction of the bytes. SquashFS has done this since forever. EROFS was designed for it. Both are in the kernel. The arithmetic on image size is immediate and flattering, and the whole idea takes about an afternoon to prototype.

The reason we do not ship it is the subject of this post, and it is not "it was hard". It worked fine. The problem is that compression does not remove the cost of those bytes, it relocates it — off the network and the disk, and onto the CPU of whichever guest happens to fault the page in. If your bottleneck was the network, that is a magnificent trade. If your bottleneck was the cold read path, you have just made your cold read path worse and written a press release about your image size.

The short version. SquashFS compresses a fixed-size block of INPUT (commonly 128 KiB) into a variable-size output, so a random 4 KiB read decompresses the whole block — amplification of roughly block size over page size. EROFS inverts it: fixed-size compressed OUTPUT in physical clusters, indexed per logical cluster, so a small read decompresses roughly one unit. Both store compressed bytes on disk and DECOMPRESSED pages in the page cache, which is the part that matters for microVM density. We ship plain ext4 cloned with an XFS reflink, because our create path is a snapshot restore and not an image pull, and the bytes-on-the-wire problem these formats solve is one we solved somewhere else entirely.

The premise: your rootfs is already read-only, nobody told the filesystem

Take an honest inventory of a sandbox rootfs after it has done a day's work. The OS is untouched. The interpreters are untouched. `site-packages` and `node_modules` are read thousands of times and written once, at bake time. The things that change are a working directory, some logs, a package cache if the user installs something, and `/tmp`. On most of our templates the written set is a low single-digit percentage of the image, and the read set is smaller than people expect too — a sandbox that runs one Python script touches a surprisingly thin slice of the hundreds of megabytes it could.

That is the exact shape the compressed read-only filesystems were built for, and the standard construction is well worn: mount the compressed image read-only, stack a writable upper layer on it with overlayfs — a tmpfs, or a small second block device — and let the guest believe it has a normal writable root. overlayfs Inside the Guest: Read-Only Rootfs, Writable Upper Layer covers the assembly. You get three things for free: the image is smaller, there is no journal to replay because there is nothing to write, and the workload physically cannot corrupt its own base. If you want that last property enforced rather than merely true, Immutable microVM Rootfs with dm-verity is the next step up.

The cost is one sentence long and the rest of this post is its footnotes: the bytes still have to be decompressed, and that happens in the guest, on the fault path, every time a page is read that is not already in the page cache.

SquashFS: the veteran, and the block-size bargain at its centre

SquashFS has been in mainline since 2.6.29 and is in every initramfs, every live CD, and every router you have ever flashed. Its design is a compact, careful piece of 2000s engineering, and the shape of it is worth knowing precisely because the shape is the performance model.

File data is split into blocks of a chosen size — selectable from 4 KiB up to 1 MiB, with 128 KiB the usual default — and each block is compressed independently. Independently is the important word: it means a block can be located and decompressed without touching its neighbours, which is what makes random access possible at all. The compressor is pluggable, which in practice means gzip, lzo, lz4, xz, lzma or zstd, and the choice moves both the ratio and the per-block decode cost.

Small files are handled by fragment packing. A file's tail — or a whole file smaller than a block — is appended into a shared fragment block along with other tails, and the fragment block is compressed as a unit. This is why SquashFS is genuinely good at a tree full of tiny files on the size axis. It is also why reading one small file can mean decompressing a fragment block that is mostly other people's files.

Metadata is compressed too, in 8 KiB chunks: the inode table, the directory table, the fragment lookup table and the id table all live compressed, with compact variable-length inodes and an optional directory index for large directories. That is excellent for image size and it means a path walk can require decompression — `stat()` on a cold cache is not free in the way it is on ext4.

Now the consequence. You want 4 KiB from the middle of a file. SquashFS must locate the containing block, read all of its compressed bytes, decompress the entire thing, and hand you your page out of the result. With a 128 KiB block and a 4 KiB page, the worst case read amplification is about 32× in decompressed bytes, plus whatever the compressed read cost. The kernel softens this with its own small caches — a metadata cache, a fragment cache, a decompressed-block cache, each a handful of entries, tunable at build time and not at mount time — and the decompressed file pages do land in the page cache, so a second read of the neighbourhood is cheap. On a cold cache, the first read of each neighbourhood is not.

The block size is one knob controlling two things you want to move in opposite directions. A bigger block gives the compressor more context and a better ratio; it also makes every small random read decompress more. There is no setting that is good at both, which means there is no correct default — only a correct answer for a specific access pattern. If you do not know your access pattern, you cannot pick `-b`, and someone picked 128 KiB for you based on a 2008 live CD.

EROFS: fixed-size output, and metadata you can seek into

EROFS — the Enhanced Read-Only File System — came out of Huawei, was built for Android system images, reached mainline around 5.4, and has since become the on-disk format of choice for lazily-loaded container images. It is not "SquashFS but newer". It changed two specific design decisions, and both of them are aimed squarely at random small reads.

Uncompressed, fixed-size metadata

EROFS does not compress its metadata. Inodes are fixed-size — a compact 32-byte form and an extended 64-byte form — in a flat, directly-addressable table, so resolving an inode number is arithmetic rather than a search. Directories are stored as blocks of fixed-size entries with a name-offset array at the front and names sorted, so a lookup within a block is a binary search over an array you just read. No decompression anywhere on the path-walk hot path.

This costs image size, obviously, and EROFS pays it deliberately. The bet is that metadata is a small fraction of a system image and the overwhelming majority of filesystem operations touch metadata, so making lookups free is worth more than making the inode table 60% smaller. On a tree with a lot of directories — a `node_modules`, a `site-packages`, a Linux source tree — that bet pays. It is the single clearest difference you will feel.

Fixed-size output compression

For file data, EROFS inverts the SquashFS arrangement. Instead of compressing a fixed-size chunk of input into a variable-size output, it packs compressed data into physical clusters that fill whole 4 KiB blocks — a fixed-size output — each covering a variable amount of logical file data. Because the output unit is fixed and block-aligned, no space is wasted on padding, and because a compact per-logical-cluster index maps file offsets to the clusters that cover them, finding the right cluster for an arbitrary offset is a direct lookup rather than a scan.

The payoff: to serve one 4 KiB page you decompress roughly one physical cluster, not a 128 KiB block. Pair that with LZ4, whose decode speed is its entire reason for existing, and the per-fault cost gets genuinely small. There is also a decompress-in-place design in the kernel side of this: because the compressed input for a cluster is small and block-aligned, it can be read into the tail of the destination pages and expanded in place, which removes a round of allocation and copying from the fault path. That is the sort of detail that only matters when the fault path is your bottleneck, which is exactly the case we are discussing.

Codecs: LZ4 is the classic pairing and still the safe one. `-zlz4hc` spends more effort at build time for a better ratio while keeping the cheap LZ4 decode. DEFLATE and zstd arrived later and buy a real ratio improvement for a real CPU cost. There is also a big-pcluster mode (`-C`) that raises the cluster size for a better ratio — note carefully that this moves EROFS back toward SquashFS's trade-off, which is a fine thing to do on purpose and an embarrassing thing to do by accident.

Every one of those codecs, and big pclusters, and the fscache integration, landed in a LATER kernel than EROFS itself. The filesystem being present tells you nothing about which `-z` your guest can mount. Check the kernel you ship to guests, not the one you build images on — our guest kernel is 5.10, which realistically means LZ4, and you should verify the rest against your own config rather than against this paragraph. An image compressed with a codec the mounting kernel lacks fails at mount time, which is at least honest, but it fails in production rather than in your build.

fscache and on-demand loading, which is why container people care

The feature that made EROFS interesting beyond Android is on-demand loading through fscache: a mounted EROFS image can be backed by a local cache that is populated from a remote source as blocks are actually requested, instead of requiring the whole image to be present first. That is lazy container image pulling, and it is the mechanism behind Nydus — whose RAFS v6 format is EROFS — and a neighbour of composefs. If your pain is "a container starts by downloading gigabytes it will never read", this is the answer to it, and Snapshot Restore vs Container Image Pull: Two Ways to Start Fast is the comparison with how we attack the same pain.

It is also the feature most sensitive to kernel version and to how your platform is put together, so treat it as a thing to verify rather than a thing to assume. A design that depends on a cachefiles daemon and a specific kernel's on-demand API is a design with a deployment story, and you should write that story down before you fall in love with the benchmark.

The three axes that actually decide it for a microVM

Strip away the format trivia and there are three questions. Only the third one is interesting, and almost nobody asks it.

1. Bytes on the wire and in object storage

This is the axis the compressed formats win, and they win it decisively. A compressed system image is a fraction of the uncompressed one, and if your create path begins with "fetch an image this host has never seen", that fraction is latency and egress spend you simply stop paying. Measure it on your own tree — and measure allocated size, not apparent size, because a sparse ext4 image makes `ls -l` lie in a direction that flatters the compressed candidates for free.

Be careful not to compare the wrong things, though. The honest baseline is not "uncompressed ext4 image"; it is "uncompressed ext4 image, sparse, after hole-punching, shipped once per host and then cloned". We have written at length about what sparseness does to artifact size in Sparse Files and Hole Punching: Why Your Snapshot Lies About Its Size, and the gap between a compressed image and a properly sparse one is smaller than the naive comparison suggests.

2. Latency and CPU on the cold read path

Here is the workload that decides it, and it is not a `dd`. The first `import pandas` in a fresh sandbox is a few hundred `stat()` calls and a few hundred small reads, scattered across the image in an order determined by Python's import machinery and the order packages were installed — which is to say, an order nobody laid the image out for. That is the precise adversarial case for fixed-input-block compression: many small reads, each landing in a different block, each costing a full block decompression.

It is also the case EROFS was designed against, between the uncompressed metadata (the `stat()` storm costs plain block reads) and the small decompression unit (each small read costs roughly one cluster). If you only benchmark one thing, benchmark this. And note where the CPU goes: into the guest, on the fault path, synchronously, in front of the user's first request. If you bill CPU time, that is CPU time someone is paying for, and the compression ratio you chose at build time is now a line on their invoice.

3. Page-cache sharing and copy-on-write — the one that changed my mind

A compressed read-only filesystem stores compressed bytes on disk and decompressed pages in the page cache. Write that sentence down and look at it, because it has a consequence that does not appear in any size comparison.

The on-disk bytes are shared beautifully: every guest mounts the same read-only image, nobody needs copy-on-write at all, and the host holds one copy. So far so good — better than good, actually, since read-only means there is no divergence to manage. But the useful cached unit is the decompressed page, and the decompressed page lives in the guest's page cache, which is guest RAM. Two guests reading the same file each build their own decompressed copy of it, each pay their own decompression CPU for it, and neither can help the other.

Now compare with a plain ext4 image cloned per sandbox. The clone is an XFS reflink, which is O(metadata) to create and leaves every unwritten block as literally the same physical extent on the host — or a dm-snapshot over a shared origin device, where reads of unwritten blocks go to one device that one host page cache covers. Copy-on-Write Rootfs: dm-snapshot vs reflink for MicroVMs has the comparison, and Shared Pages & Copy-on-Write: Packing MicroVMs Densely has the density argument in full. The point is that in the uncompressed case, the bytes the host is caching are the bytes the guest will use, unchanged, with no CPU in between. In the compressed case the host can only cache compressed bytes that every guest must independently transform before they are useful.

If your density strategy is built on shared bytes, compression is in tension with it. That is not a reason never to compress; it is a reason to know which of the two you are optimising, because you cannot have both at once.

SquashFS, EROFS and a plain ext4 image cloned with a reflink, on the seven properties that decide a microVM rootfs.
PropertySquashFSEROFSext4 + reflink
Metadata layoutCompressed in 8 KiB chunks: compact inodes, directory table, optional directory index. A cold path walk can require decompression.Uncompressed and fixed-size: 32-byte compact or 64-byte extended inodes in a flat table, sorted dirents with a name-offset array. Lookup is arithmetic plus a binary search.Standard inodes, extent trees, htree directories. Uncompressed; the journal exists but a read-only mount never replays it.
Compression unitFixed INPUT. A 4 KiB to 1 MiB block of original data (commonly 128 KiB) compressed independently; small files and tails packed into shared fragment blocks.Fixed OUTPUT. Physical clusters filling whole 4 KiB blocks, with a compact per-logical-cluster index from file offset to cluster. `-C` enlarges the cluster and trades back toward the SquashFS behaviour.None.
Small random read costDecompress the whole containing block to produce one 4 KiB page: amplification on the order of block size over page size. Kernel block, fragment and metadata caches soften repeats, not the first touch.Decompress roughly the cluster covering the offset; LZ4 keeps the per-unit decode cheap, and decompress-in-place removes a copy from the fault path.One 4 KiB read. No decompression, no CPU, no amplification.
Image sizeVery strong, especially with xz or zstd at a large block. The ratio and the amplification are the same knob, so the best number here is the worst number above.Good, and closer to SquashFS with big pclusters plus a strong codec than with LZ4 alone. The uncompressed metadata costs real bytes. Measure on your own tree.Largest in principle — the bytes are the bytes. Sparseness and hole-punching close more of the gap than people expect.
Page-cache / CoW sharingCompressed on disk, DECOMPRESSED in the page cache. One shared on-disk image, no per-guest CoW needed or possible; every guest builds and pays for its own decompressed pages.Same model. fscache-backed on-demand loading adds a host-side shared cache of compressed blocks, which helps the fetch and not the decode.Guests share physical extents via reflink, or a dm-snapshot origin; writes divert per guest. The host caches exactly the bytes the guest uses.
Tooling maturity`mksquashfs` is everywhere, has been for fifteen years, and its flags are well understood and well documented. Nothing will surprise you.`mkfs.erofs` from erofs-utils: newer, moving quickly, excellent, and the codec you want may need a newer kernel than the one you ship.`mkfs.ext4`, `e2fsprogs`, `resize2fs`, `debugfs` — the most boring and best-understood storage tools in Linux.
Create cost per sandboxMount a shared read-only image, then assemble an overlayfs upper layer (tmpfs or a second device) for writes.Same, plus an fscache/cachefiles setup if you use on-demand loading.`cp --reflink` of the image: O(metadata), no data copied, then boot or restore.

Build both from the same directory and look at your own numbers

Everything above is structure. The ratios are content. Here is the build script; the numbers it prints are yours and I am deliberately not printing mine, because mine are a property of our template tree and would be worse than useless to you.

#!/usr/bin/env bash
# Build the SAME directory tree as squashfs and as erofs, several ways
# each, then compare. Needs squashfs-tools and erofs-utils. Read step 0
# before you get attached to a codec: an image compressed with something
# your GUEST kernel cannot decode is a very compact brick.
set -euo pipefail

SRC=${1:?usage: build-ro-images.sh /path/to/rootfs-tree [outdir]}
OUT=${2:-/var/tmp/roimg}
mkdir -p "$OUT"

# 0. What can the kernel that will MOUNT this actually read? erofs reached
#    mainline well before its newer codecs did -- LZ4 first, then LZMA,
#    then DEFLATE, then zstd, each in a later release, and big pclusters
#    and fscache later still. Check the kernel you SHIP to guests, not the
#    one you build on. Ours is 5.10, which in practice means LZ4 and you
#    should verify anything beyond that against your own config.
modinfo erofs 2>/dev/null | sed -n '1,3p' || echo "erofs: not modular here"
( zgrep -E 'EROFS|SQUASHFS' /proc/config.gz \
  || grep -E 'EROFS|SQUASHFS' "/boot/config-$(uname -r)" ) 2>/dev/null || true

# 1. SquashFS. -b is the fixed INPUT block and it is THE knob: a bigger
#    block compresses better and decompresses more bytes for every 4 KiB
#    page you actually wanted. Build the extremes so you can see the cost
#    of the ratio instead of arguing about it.
mksquashfs "$SRC" "$OUT/sq-zstd-128k.sqfs" -noappend -quiet \
  -comp zstd -Xcompression-level 15 -b 131072
mksquashfs "$SRC" "$OUT/sq-zstd-16k.sqfs"  -noappend -quiet \
  -comp zstd -Xcompression-level 15 -b 16384
mksquashfs "$SRC" "$OUT/sq-lz4-128k.sqfs"  -noappend -quiet \
  -comp lz4 -Xhc -b 131072
# -noI -noD -noF leave the inode, data and fragment tables UNcompressed,
#    which is the crude way to buy back metadata lookup speed:
# mksquashfs "$SRC" "$OUT/sq-noi.sqfs" -noappend -quiet -noI -b 131072

# 2. EROFS. Note the argument order: destination first, then source.
#    -zlz4hc is the classic pairing -- expensive LZ4 at build time, cheap
#    LZ4 decode on every fault. -C raises the pcluster size ("big
#    pcluster"), which buys ratio by handing back the small-read
#    advantage, so build both and measure the trade on your content.
mkfs.erofs -zlz4hc,12         "$OUT/er-lz4hc.erofs"     "$SRC"
mkfs.erofs -zlz4hc,12 -C65536 "$OUT/er-lz4hc-64k.erofs" "$SRC"
mkfs.erofs -zzstd,19          "$OUT/er-zstd.erofs"      "$SRC" \
  || echo "zstd unsupported by this erofs-utils -- itself a finding"
# No -z at all: uncompressed erofs. You keep the compact fixed-size
# inodes and the flat metadata and pay no decompression whatsoever. A
# genuinely underrated option that almost nobody builds. (In COMPRESSED
# mode, -Eztailpacking and -Efragments are the flags that shrink the
# small-file tail; they do nothing here, with no -z to pack into.)
mkfs.erofs                    "$OUT/er-plain.erofs"     "$SRC"

# 3. The baseline you are trying to beat: a plain ext4 image, sized to the
#    tree plus slack. This is what we actually ship to guests.
BYTES=$(du -sb "$SRC" | cut -f1)
truncate -s $(( BYTES + BYTES / 4 + 256 * 1024 * 1024 )) "$OUT/base.ext4"
mkfs.ext4 -F -q -d "$SRC" "$OUT/base.ext4"

# 4. Compare apparent size AND allocated size. The ext4 image is sparse,
#    so `ls -l` overstates it, which flatters the compressed formats for
#    free. du -b is apparent, du -B1 is allocated. Quote the second one.
printf '%13s %13s  %s\n' APPARENT ALLOCATED FILE
for f in "$OUT"/sq-*.sqfs "$OUT"/er-*.erofs "$OUT"/base.ext4; do
  [ -f "$f" ] || continue
  printf '%13s %13s  %s\n' "$(du -b "$f" | cut -f1)" \
    "$(du -B1 "$f" | cut -f1)" "$(basename "$f")"
done

The honest benchmark: cold cache, many small reads, count the sectors

Sizes are the easy half. The half that decides it is a cold-cache random-small-read measurement, and it has two failure modes worth naming. The first is forgetting that SquashFS keeps its own small caches independent of the page cache, so dropping caches under a live mount does not fully reset it — unmount first, then drop. The second is reporting only wall-clock time, when the interesting quantity is the amplification: how many sectors the backing device actually read to satisfy your reads. Read that straight off the loop device's `stat` and the argument stops being a matter of opinion.

#!/usr/bin/env bash
# The only benchmark that answers the question, because the answer is a
# property of YOUR file-size histogram and YOUR access order. Needs root
# (mount + drop_caches) on a machine you are allowed to abuse. Run each
# case three times, keep the median, and distrust the first run.
set -euo pipefail

IMG=${1:?usage: coldread.sh /var/tmp/roimg/er-lz4hc.erofs}
MNT=/mnt/ro-bench
mkdir -p "$MNT"

case "$IMG" in
  *.sqfs)  FSTYPE=squashfs ;;
  *.erofs) FSTYPE=erofs ;;
  *.ext4)  FSTYPE=ext4 ;;
  *) echo "unknown image type: $IMG" >&2; exit 2 ;;
esac

LOOP=$(losetup --find --show --read-only "$IMG")
STAT=/sys/block/$(basename "$LOOP")/stat
sectors() { awk '{print $3}' "$STAT"; }   # field 3 = sectors read

cold() {
  # Order matters. Unmount FIRST, so the filesystem's own caches go away
  # too -- squashfs keeps small metadata, fragment and decompressed-block
  # caches that are separate from the page cache and survive a drop while
  # the mount is live. Then drop. vmtouch -e "$IMG" is the surgical
  # version when you cannot drop caches globally on that box.
  umount "$MNT" 2>/dev/null || true
  sync
  echo 3 > /proc/sys/vm/drop_caches
}

run() { # run <label> <command...>
  local label=$1; shift
  cold
  mount -t "$FSTYPE" -o ro "$LOOP" "$MNT"
  local s0 s1
  s0=$(sectors)
  /usr/bin/time -f "$label  %e s wall  %P cpu" "$@" || true
  s1=$(sectors)
  echo "$label  sectors read from backing device: $((s1 - s0))"
}

# A -- the realistic worst case: hundreds of small reads scattered across
# the image in an order nobody laid the image out for. A Python import is
# exactly this shape: stat a dozen candidate paths, then read a few KiB
# each out of fifty .py and .so files. This is where large-block
# compression is punished and the sector counter proves the
# amplification.
run A-import timeout 120 chroot "$MNT" /usr/bin/python3 -c \
  'import json, ssl, email, logging.config, sqlite3, xml.etree.ElementTree'

# B -- metadata only, no file data decompressed at all. erofs's
# uncompressed fixed-size inode table shows up here; squashfs has to
# decompress metadata blocks to walk the tree.
run B-stat timeout 300 find "$MNT" -type f -printf ''

# C -- read every byte, sequentially. The flattering case for large-block
# compression, and the one your users will never actually run.
run C-readall timeout 600 bash -c \
  "find '$MNT' -type f -print0 | xargs -0 cat > /dev/null"

umount "$MNT" 2>/dev/null || true
losetup -d "$LOOP"

# Report A, B and C per image, with the sector counts next to the wall
# clock. Notice what I have NOT printed anywhere in this post: a number.
# Anyone who quotes you a single "format X is N percent faster" figure is
# quoting you their rootfs, not yours.

If you would rather not install three sets of filesystem tools on your laptop, build the images in a sandbox. Image building is pure userspace — neither `mksquashfs` nor `mkfs.erofs` mounts anything — so this part runs anywhere, and you can snapshot the result and come back to the same bake-off later instead of recreating it.

# Build the candidates inside a disposable sandbox, so you are not
# installing erofs-utils on your laptop and not arguing with a Mac about
# loop devices. Image BUILDING is pure userspace -- mksquashfs and
# mkfs.erofs never touch a mount -- so this half works anywhere. The
# mount-and-drop-caches half needs a Linux box with root.
from pandastack import Sandbox

TREE = "/work/tree"

sbx = Sandbox.create(
    template="base",                      # 4 GiB / 8 vCPU, baked
    ttl_seconds=900,                      # IDLE timeout, not a deadline:
    metadata={"job": "ro-fs-bakeoff"},    # a busy sandbox is not reaped
)
try:
    with open("build-ro-images.sh") as f:
        sbx.filesystem.write("/work/build.sh", f.read())

    r = sbx.exec(
        "apt-get -qq update && apt-get -qq install -y "
        "squashfs-tools erofs-utils e2fsprogs"
    )
    assert r.exit_code == 0, r.stderr

    # Something worth measuring: a real site-packages tree, which is the
    # long tail of thousands of small files the two formats disagree
    # about most. Bound long commands with timeout(1) IN THE GUEST --
    # exec() does not reliably enforce a client-side timeout, so the
    # guest owns the clock.
    sbx.exec(f"mkdir -p {TREE} && python3 -m venv {TREE}/venv")
    sbx.exec(
        f"timeout 600 {TREE}/venv/bin/pip -q install "
        "pandas pyarrow httpx jinja2 pydantic"
    )

    for ev in sbx.exec_stream(f"timeout 900 bash /work/build.sh {TREE}"):
        print(ev)          # stdout / stderr / exit events as they happen

    # Pull one artifact back out if you want it locally.
    img = sbx.filesystem.read("/var/tmp/roimg/er-lz4hc.erofs")
    print(f"er-lz4hc.erofs: {len(img)} bytes")

    # Snapshot now and the whole bake-off is restorable later instead of
    # rebuildable: Sandbox.create(from_snapshot=snap) brings the images
    # back with the page cache already warm.
    snap = sbx.snapshot()
    print("restore with Sandbox.create(from_snapshot=%r)" % snap)
finally:
    sbx.kill()             # teardown -- there is no sbx.delete()

Why we ship plain ext4 and a reflink, and when I would change my mind

Our rootfs is a plain ext4 filesystem inside a file. Creating a sandbox clones it with an XFS reflink — O(metadata), no data copied — or a dm-snapshot, and the host shares every unwritten block between every sandbox on the box. That is the whole mechanism, described properly in Copy-on-Write Rootfs: Why MicroVM Create Is O(metadata).

The reason compression does not help us is that we do not have the problem it solves. Our create path is not an image pull; it is a Firecracker snapshot restore, which is why creates land at a p50 of 179 ms and a p99 around 203 ms rather than at image-pull timescales. Only the very first spawn of a template with no baked snapshot does a full cold boot — about 3 seconds — and then it auto-bakes and nobody pays that again. The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms has the step-by-step.

And the bytes-on-the-wire problem is real for us; we just attacked it in a different place. What crosses the network on a cold host is guest MEMORY, not the rootfs, and we stream it on demand: HTTP Range GETs of 4 MiB chunks from object storage as the guest faults, a zero-chunk bitmap so pages that were never written are filled with zeroes and never fetched at all, and a prefetch trace recorded at bake time and replayed in the background so faults become cache hits. Memory Prefetch: The Working Set Is the Real Unit of a Fast Restore and UFFDIO_ZEROPAGE vs UFFDIO_COPY: Stop Paying RAM for Zeros cover the two halves. We compressed the artifact that actually travels, by not sending most of it.

There is a second-order reason too, and it is the one I find most convincing. Because we restore from a snapshot rather than booting, a sandbox starts with its page cache already populated from bake time. The decompression cost of a compressed rootfs would be paid once, during the bake, and then frozen into the memory image as ordinary decompressed pages — which sounds like a win until you notice that the thing we stream over the network is exactly that memory image. Compression would shrink the artifact that never leaves the host and is already shared by reflink, while leaving untouched the artifact we actually pay to move. That is precisely the wrong way round.

Where it would win, for us or for you, is specific and worth stating plainly:

  • A very large template with a long tail of rarely-read bytes. Three JDKs, a full TeX distribution, every Chrome locale, a vendored toolchain nobody invokes. Compression is free money on bytes you ship and never read.
  • A cold host that must fetch an image it has never seen. Fleet churn, autoscaling, a host replaced by the managed instance group, a brand-new region: the first create on that host pays the full fetch, and that is the moment a compressed image earns its keep.
  • A fleet where object-storage egress is the dominant line item. If you are paying per gigabyte to move images and your hosts are short-lived, the compression ratio is directly a cost ratio.
  • Many hosts and no ability to pre-place images. Our model leans hard on placing artifacts on hosts ahead of time; if you cannot do that, lazy loading over EROFS plus fscache is a genuinely good answer and you should look at Nydus before you build your own.
  • Immutability as a requirement rather than a nicety. If the base must be provably unmodifiable, a read-only format plus dm-verity is a cleaner story than "ext4 that we promise nobody writes to".

Which one to reach for

If you are building your own microVM platform, the decision is mostly determined by which of the three axes is your binding constraint:

  1. Reach for EROFS when you have many hosts, cold pulls are common, image size matters, and your read pattern is random small reads — which it is, if guests run interpreted languages. The uncompressed metadata and the small decompression unit are aimed at exactly that. Check your guest kernel's codec support before you choose `-z`, and leave big pclusters alone until you have measured with them off.
  2. Reach for SquashFS when the tooling maturity is worth more to you than the last 15% of random-read performance, when you want to pick the block size yourself and can defend the number, or when your access pattern is sequential — a model being loaded once, a dataset being streamed — in which case large-block compression is a pure win and the amplification never materialises.
  3. Reach for plain ext4 plus reflink when create latency is the metric you are judged on, when density comes from sharing bytes between many guests on one host, and when you can pre-place images so cold pulls are rare. This is our situation and it is why we are boring about it.
  4. Reach for uncompressed EROFS — no `-z` at all — if you want the compact fixed-size metadata and the no-journal immutability without any decompression on the fault path. Almost nobody builds this and it is a perfectly reasonable middle position.
  5. Do not reach for any of them before you have run the cold-read benchmark above on your real tree. The variance between content mixes is larger than the variance between formats, which means a published comparison of someone else's rootfs predicts yours poorly.

Where I would not use this, and what we do not do

I would not use a compressed read-only rootfs for a workload that writes a lot, which includes any sandbox where a user runs `pip install` or `npm ci`. Those writes land in the overlay upper layer, and if the upper is a tmpfs they land in guest RAM — so you traded disk bytes the host was sharing for free for memory bytes that are per-guest and, in our architecture, streamed over the network. A second writable block device avoids that and gives you two artifacts to version and assemble instead of one. Neither is obviously wrong; both are more moving parts than `cp --reflink`.

I would not use one for a database. Our managed Postgres runs on a durable volume rather than the ephemeral rootfs for exactly this reason, and compressing a filesystem whose entire job is to be written is an elaborate way to be slower.

And the things we genuinely do not do, since this is where these posts usually get vague: we have no GPUs and no GPU passthrough, so if your compressed-image problem is a multi-gigabyte CUDA stack, we are not the platform to solve it on. Our guest kernel is 5.10, which bounds what EROFS features a guest of ours could use even if we shipped one. Our `base` template is 4 GiB of RAM and 8 burst vCPU, and Firecracker cannot change vCPU or memory at snapshot restore, so the template's baked configuration governs the size and a `memory_mb` on create is overridden to it. And we have not shipped a compressed rootfs, so everything above about how one would behave on our platform is reasoning from the architecture plus an afternoon's prototype, not a production measurement. I would rather say that than imply a benchmark I do not have.

The general lesson is the one I keep relearning in this job. Compression is not a reduction in cost; it is a change in where the cost is denominated. Moving bytes from the network to the CPU is a brilliant trade when the network is your constraint and a self-inflicted wound when the fault path is. The formats are both excellent. The question was never which filesystem is better — it was which of my resources I am short of, and I had to be wrong about that for a week to find out.

Frequently asked questions

Is EROFS simply a better choice than SquashFS now?

For a microVM or container rootfs with random small reads, EROFS is usually the better default, and the two reasons are structural rather than a matter of tuning. Its metadata is uncompressed and fixed-size, so a cold path walk is plain block reads plus a binary search instead of decompressing metadata chunks — and filesystem operations are overwhelmingly metadata operations. And its data is stored in fixed-size-output physical clusters with a per-logical-cluster index, so serving one 4 KiB page decompresses roughly one small unit rather than a whole 128 KiB input block. That is a difference in kind, not degree, on exactly the access pattern an interpreted-language runtime generates. But "usually better" is not "always better", and three things push back. SquashFS at a large block with xz or zstd can still produce a smaller image, and if image size is your binding constraint that is the number that matters. SquashFS tooling is fifteen years more mature, documented and predictable, which is worth real money when something goes wrong at 3am. And EROFS's newer features — the stronger codecs, big pclusters, fscache-backed on-demand loading — each landed in a later kernel than the filesystem itself, so what you can actually use is bounded by the kernel you ship to guests rather than by the one in the release notes. Build both from the same tree, run the cold random-read benchmark, and let your own content decide.

Does a compressed rootfs reduce how much memory my guests use?

No, and this is the misconception worth killing first because it is so intuitive. A compressed read-only filesystem stores compressed bytes on disk and decompressed pages in the page cache. When a guest reads a file, the kernel decompresses the relevant unit and the resulting pages go into the page cache in their normal, expanded form — exactly as they would from an uncompressed filesystem. The memory footprint of "the files this guest has read" is therefore roughly the same either way. Compression actually adds a little memory pressure: transient buffers for the compressed input, plus the filesystem's own caches, which in the SquashFS case are a small metadata cache, a fragment cache and a decompressed-block cache that are separate from the page cache. The saving is on disk and on the network, full stop. Worse, for a platform like ours the asymmetry runs the wrong way: the artifact that travels over the network on a cold host is guest memory, not the rootfs, and compression does nothing for it. If your goal is to reduce guest memory, the levers are elsewhere — a smaller working set, dropping the page cache when you know you are done with it, ballooning, host-side deduplication. /blog/firecracker-guest-page-cache-drop-explained and /blog/microvm-memory-oversubscription-explained are the places to start.

Can I use a compressed read-only rootfs together with Firecracker snapshots?

Yes, and there is a pleasing interaction plus a trap. The pleasing part: a snapshot captures guest memory and device state, and the rootfs is a separate file, so if you bake a template by booting it, exercising it, and then snapshotting, the guest's page cache in that snapshot already holds the decompressed pages of everything the bake touched. The decompression cost is paid once, at bake time, and every subsequent restore gets those pages as ordinary memory. That genuinely removes most of the cold-read penalty for the hot set, which is the main argument against compression in the first place. The trap is twofold. First, restoring from a snapshot does not help pages outside the baked working set; the first guest to read something the bake never touched pays full decompression, and that is precisely the long tail you compressed most aggressively. Second, the memory image now contains those decompressed pages, so on any platform that moves memory images around — streaming them, replicating them, storing them per generation — you have moved bytes out of the compressed disk artifact and into the uncompressed memory artifact. Whether that is a win depends entirely on which artifact you pay to transport. Also check that your mount is assembled deterministically across restore: an overlayfs upper on tmpfs is in the snapshot, so every clone of one snapshot shares whatever was in it.

Isn't the compression codec the only thing that really matters — zstd versus LZ4?

The codec sets the ratio and the per-byte decode cost, but the format sets how many bytes you have to decode to answer a small read, and that multiplier usually dominates. Consider a 4 KiB read against a SquashFS image with a 128 KiB block: you decode 128 KiB regardless of codec. Switching that image from zstd to LZ4 makes each of those 128 KiB cheaper to decode, but you are still decoding 32 pages to deliver one. Against an EROFS image you decode roughly one small cluster for the same read. The format change reduces the work; the codec change reduces the price of work you should not be doing. That is why the most interesting experiment in the build script is not zstd versus LZ4 but 128 KiB versus 16 KiB SquashFS blocks — same codec, same content, and the amplification moved by 8x. The sector counter on the loop device shows it directly. There is also a second-order effect people miss: smaller compression units compress worse because the compressor has less context, so shrinking the block to help random reads costs you image size, and the two knobs fight. EROFS's fixed-output-size design is an attempt to escape that fight rather than to tune it. Once you have chosen the format for your access pattern, then pick the codec by what your guest kernel supports and how much CPU you are willing to burn on the fault path.

What about EROFS plus fscache for lazy image loading — should I build that for a microVM platform?

Look at it seriously if your dominant cost is cold hosts fetching images they have never seen, and look at Nydus before you build anything, since its RAFS v6 format is EROFS and the lazy-loading problem is already solved there in production. The mechanism is good: a mounted EROFS image backed by a local cache populated on demand, so a guest can start before the whole image is present and pulls only the blocks it touches. For a fleet with high churn and large images, that is the right shape. Two cautions. First, it depends on kernel support for on-demand loading plus a cachefiles daemon, and that dependency moved across several releases — so pin the kernel version in your design document, verify against your own config, and accept that a guest kernel older than the feature cannot use it at all. Second, be clear about what it fixes: it fixes the fetch, not the decode. Blocks still arrive compressed and are still decompressed per guest on the fault path, so the random-small-read cost discussed above is unchanged. On our platform we attack the same problem from the other side — the rootfs is pre-placed locally because copy-on-write needs a local block device, and the artifact we stream on demand is guest memory, in 4 MiB Range GETs with a zero-chunk bitmap and a prefetch trace. Different artifact, same instinct: do not transport what nobody will read.

Keep reading

Related posts

  • overlayfs Inside the Guest: Read-Only Rootfs, Writable Upper Layer

    Mount the template read-only, stack a writable layer on top, let the kernel merge them: overlayfs is three lines of mount options and a surprising number of sharp edges. Whole-file copy_up, whiteouts, metacopy, inode weirdness — and why PandaStack does its copy-on-write below the guest instead.

  • Should You Compress Firecracker Memory Snapshots?

    Compression looks like free money when your snapshot bucket is measured in terabytes. Then you discover that a compressed stream has no byte N — and your lazy 49ms restore turns into reading four gigabytes you were never going to touch.

  • Immutable microVM Rootfs with dm-verity

    "Read-only" is a promise the guest makes to itself. dm-verity is a Merkle tree the kernel checks on every block read — so tampering is detected rather than merely discouraged. Here is the wiring, the Firecracker boot args, and an honest account of what it does not cover.

  • Firecracker Boot Sources: initrd vs a Root Block Device

    A microVM has no BIOS, no bootloader, and no GRUB menu — just a kernel, a command line, and one decision: does userland arrive in RAM as a cpio archive, or on a virtio-blk device you mount? The answer changes your memory bill more than your boot time.

  • virtiofs vs virtio-blk: How Files Actually Get Into a MicroVM

    One gives the guest a disk it owns. The other gives it a window onto a directory the host owns. That single difference decides whether you can fork a machine in 400ms, who parses guest-controlled input, and what a multi-tenant escape looks like.

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.