all posts

Host Swap, zram and zswap: How Many MicroVMs Actually Fit on One Machine

Ajay Kumar··10 min read

I build PandaStack, an open-source Firecracker microVM platform, which means I spend a lot of my life staring at one number on a host: how many guests fit before something goes wrong. It is a more interesting question than it looks, because the honest answer depends entirely on which definition of "fit" you use, and the two obvious definitions differ by a factor of five.

Here is the shape of it. A `base` sandbox bakes 4 GiB of guest RAM. A customer spins one up, runs a dependency install, builds a Next.js app, and the guest touches maybe 400 MiB of real pages. The host is holding 400 MiB. The accounting says 4 GiB. Multiply by thirty sandboxes and you have a 32 GiB machine that reports itself completely full while roughly 90% of its physical memory is doing absolutely nothing. That gap — committed versus resident — is the entire density question, and every lever you have for closing it eventually routes through host memory overcommit, which eventually routes through swap.

Swap inside the guest is a different question with a different answer, and I wrote that one up separately in Swap and zram Inside a Firecracker MicroVM: What Actually Happens. This post is strictly about the host: the machine running dozens of Firecracker processes, and whether its swap configuration lets you safely sell more guest RAM than you own. The short version of the conclusion is that it does not, and that the useful thing swap gives you is not capacity at all.

If you read one paragraph: measure the working set, admit against the measurement with a reserve per in-flight create, keep a hard floor of genuinely free RAM, and treat swap as an airbag rather than an engine. zram and zswap are worth running as a compression play. Neither is capacity you can sell. And watch PSI in /proc/pressure/memory, not free memory, because free memory tells you nothing on a healthy host.

A microVM host is an overcommit problem by construction

A Firecracker guest's memory is not a special kind of memory. It is an anonymous mapping in an ordinary Linux process. The VMM creates a region the size of the configured guest RAM, hands the guest physical address space to KVM, and from the host kernel's point of view it is a large private mapping belonging to a process that happens to be named `firecracker`. Everything the host kernel knows how to do to anonymous memory — reclaim it, compress it, swap it, kill the process holding it — it can do to a customer's VM.

The problem is what the guest kernel does with that memory, which is: keep it. A Linux guest will page-cache every byte it reads and will not voluntarily give pages back to anyone, because from inside the VM there is no one to give them to. Free memory, to a guest, is wasted memory — that is correct behaviour and it is also exactly what makes density hard. So a guest that genuinely needed 400 MiB will, after an hour of builds and file reads, be sitting on most of its 4 GiB as page cache it will never be asked to drop. Resident grows monotonically and does not come back down.

This is why PandaStack has two admission modes rather than one, and why the choice is the single most consequential capacity decision in the whole platform. In committed mode, the scheduler charges every sandbox its full configured size at create time. It is safe, it is trivially correct, and it caps a 32 GiB host at roughly seven 4 GiB sandboxes. In working-set mode the scheduler admits against a measured usage figure the host reports instead, with each in-flight create charged a reserve — a quarter of its ask, with a floor — so a burst of simultaneous creates cannot all admit themselves against the same free pages. Working-set mode is the fleet default.

I will be blunt about how we learned the difference. We ran committed mode in production and a fleet started refusing creates with a capacity error while the large majority of its physical RAM was idle. Nothing was broken. The arithmetic was working perfectly, on promises instead of pages. The incident was not a bug in the scheduler, it was the scheduler believing the wrong number — and that is the failure mode this whole topic exists to produce.

One carve-out, and it is deliberate: managed Postgres databases are always charged committed and never overcommitted. A database is the one workload whose memory is genuinely hot, whose tail latency is the product, and whose pages are the last thing on the machine you want a host reclaim pass to consider. Everything in the rest of this post about trading latency for density is a bet, and you do not make that bet with someone's primary datastore.

MAP_PRIVATE is the reason any of this works at all

Overcommit is only viable because a restored guest does not start out resident. When PandaStack restores a snapshot, the guest memory image is mapped `MAP_PRIVATE`: the kernel fills each page from the file on first fault and copies it on write. A 4 GiB guest that restores and then touches 200 MiB has 200 MiB of resident private pages, not 4 GiB. Resident grows with what the guest touches, not with what it was sold. That single property is what turns "179 ms to create a sandbox" into "179 ms to create a sandbox without paying 4 GiB of RAM for the privilege."

There is a second, stronger version of the same idea: userfaultfd. With streaming enabled the guest's memory is not backed by a local file at all — a handler process registers the region, and when the guest faults, the handler fetches a 4 MiB chunk from object storage and installs it. Resident memory then tracks the chunks the guest has actually demanded, and the host never has to hold, or even download, the parts of the image nobody touched. If you want that mechanism in detail it is userfaultfd: Lazy Memory for Instant VM Restore. For this post the relevant fact is narrower and comes up again in a minute: a page the UFFD handler installed is, from then on, ordinary host anonymous memory. Which means the host can swap it.

The overcommit sysctls, and why strict mode is the wrong answer here

`/proc/sys/vm/overcommit_memory` takes three values and they are less symmetric than the numbering suggests. Mode 0, the default, is a heuristic: the kernel refuses allocations that are obviously absurd — a single mapping larger than total RAM plus swap — and waves everything else through. Mode 1 is "always overcommit": never refuse an allocation, let the OOM killer sort out the consequences later. Mode 2 is strict accounting: the kernel maintains a hard `CommitLimit` of swap plus `overcommit_ratio` percent of RAM, and refuses any mapping that would push `Committed_AS` past it.

Mode 2 sounds like the responsible choice and is almost always wrong for this workload, for a reason that is structural rather than a matter of taste. Firecracker's guest memory is one big private writable mapping, so it counts against `Committed_AS` in full the moment it is created — the whole 4 GiB, before the guest has touched a page. Strict accounting therefore refuses precisely the thing that makes microVM density work. You would have to raise `overcommit_ratio` well above 100 to start any meaningful number of guests, at which point you have a limit that admits everything and a false sense of having configured something. The useful enforcement point is one layer up, in a scheduler that knows the difference between a promise and a page, not in a kernel counter that only ever sees promises.

So: leave `overcommit_memory` at 0, or set it to 1 if you have a VMM or allocator that trips the heuristic, and do your admission control where the information is. Then go read the numbers that actually describe the machine.

# ---- 1. The number everyone reads first, and it is the wrong one. ---------
free -h
#                total        used        free      shared  buff/cache   available
# Mem:            125Gi        96Gi       1.1Gi       0.0Ki        28Gi         24Gi
# Swap:           8.0Gi       128Mi       7.9Gi
#
# "free" is near zero on every healthy Linux box and means nothing -- unused
# RAM is wasted RAM and the kernel knows it. "available" is the honest column:
# a (conservative) estimate of what a new workload can get without swapping.

# ---- 2. The accounting numbers. One of these is the business model. ------
grep -E '^(MemTotal|MemAvailable|Cached|AnonPages|Committed_AS|CommitLimit|SwapTotal|SwapFree)' \\
  /proc/meminfo
# MemTotal:       131072000 kB   ~125 GiB of physical memory
# MemAvailable:    25165824 kB   ~24 GiB you can hand out right now
# Cached:          29360128 kB   page cache: rootfs images, snapshot files
# AnonPages:      100663296 kB   ~96 GiB -- nearly all of it is guest memory
# Committed_AS:   184549376 kB   ~176 GiB of PROMISES
# CommitLimit:     71303168 kB   only enforced when overcommit_memory=2
# SwapTotal:        8388608 kB
# SwapFree:         8257536 kB
#
# Committed_AS (176 GiB) > MemTotal (125 GiB) is not an error, it is the
# product. The gap between Committed_AS and AnonPages is the density you are
# selling. The gap between AnonPages and MemTotal is your remaining runway.

cat /proc/sys/vm/overcommit_memory   # 0 heuristic | 1 always | 2 strict
cat /proc/sys/vm/overcommit_ratio    # 50 by default; only read in mode 2

# ---- 3. zswap: a compressed cache IN FRONT OF a real swap device. --------
grep -r . /sys/module/zswap/parameters/ 2>/dev/null
# .../enabled:Y
# .../compressor:zstd          lzo | lz4 | zstd -- zstd compresses hardest
# .../zpool:zsmalloc           zsmalloc | zbud (z3fold is on its way out)
# .../max_pool_percent:20      cap on RAM the compressed pool may occupy
# .../accept_threshold_percent:90
# .../shrinker_enabled:Y       lets reclaim push the pool out to real swap

# Flip it at runtime (survives nothing; set it on the kernel cmdline for real):
echo zstd     | sudo tee /sys/module/zswap/parameters/compressor
echo zsmalloc | sudo tee /sys/module/zswap/parameters/zpool
echo 20       | sudo tee /sys/module/zswap/parameters/max_pool_percent
echo Y        | sudo tee /sys/module/zswap/parameters/enabled
# Durable version, in your bootloader config:
#   zswap.enabled=1 zswap.compressor=zstd zswap.zpool=zsmalloc \\
#   zswap.max_pool_percent=20

# ---- 4. zram: a compressed block device in RAM, used AS swap. ------------
# No disk anywhere in this path by default. This is RAM-for-CPU, not capacity.
sudo modprobe zram num_devices=1
echo zstd | sudo tee /sys/block/zram0/comp_algorithm
echo 16G  | sudo tee /sys/block/zram0/disksize      # ADDRESSABLE size, not RAM
echo 4G   | sudo tee /sys/block/zram0/mem_limit     # the actual RAM ceiling
sudo mkswap /dev/zram0 && sudo swapon --priority 100 /dev/zram0
zramctl
# NAME       ALGORITHM DISKSIZE DATA COMPR TOTAL STREAMS MOUNTPOINT
# /dev/zram0 zstd           16G 2.9G  742M  771M      64 [SWAP]
# DATA 2.9G stored in TOTAL 771M -- that ratio is the whole value proposition.

# ---- 5. The signal that actually tells you you are over the line. --------
cat /proc/pressure/memory
# some avg10=0.42 avg60=0.31 avg300=0.19 total=81204991
# full avg10=0.00 avg60=0.00 avg300=0.00 total=1902233
#
# "some" = at least one task stalled on memory. "full" = EVERYTHING stalled.
# A nonzero and rising "full avg10" is the host telling you it is reclaiming
# instead of working. That is the alert. MemFree is not an alert.

Plain swap, zram and zswap are three different things

People conflate these constantly, including in production runbooks, so it is worth being pedantic. Plain swap is a partition or file on a real block device; cold anonymous pages are written there and read back on fault, and the cost is a disk round trip. zram is a compressed block device that lives in RAM and is then used as a swap device; pages "swapped" to it never leave physical memory, they are compressed in place, and you have traded CPU cycles and some RAM for the pool against the ability to hold more logical pages. zswap is the third thing and the one most often misdescribed: it is a compressed write-back cache sitting in front of a real swap device. Pages go into the compressed pool first; when the pool hits `max_pool_percent`, older entries are decompressed and written out to the actual swap device underneath. zswap without a swap device configured does approximately nothing.

The knobs are the same shape in both compressed paths. The allocator — `zsmalloc`, `zbud`, or the deprecated `z3fold` — decides how tightly compressed pages can be packed: `zbud` pairs two compressed pages per physical page and so caps you at 2:1 no matter how compressible the data is, `zsmalloc` packs variable-size objects and gets the real ratio. Use `zsmalloc` unless you have a specific reason. The compressor — `lzo`, `lz4`, `zstd` — is the usual speed-versus-ratio dial, and for guest memory on a machine that is CPU-rich and RAM-poor, `zstd` is usually the right end of it. Guest memory compresses well, because a lot of it is zeroes, page cache of text files, and runtime heaps with a great deal of structural repetition.

Four host swap arrangements for a machine running many microVMs
ArrangementWhere cold pages liveCPU costCapacity or only densityFailure mode under a real spikeFit for a microVM host
No swap at allNowhere — they stay residentNoneNeitherOOM killer, immediately and without warningHonest but brittle; needs a real free-RAM floor
Plain disk swapA block deviceNegligible CPU, real I/OLooks like capacity; is not sellable capacityMulti-millisecond faults inside vCPU threads; guest soft lockupsKeep it small, as an airbag only
zram (as swap)Compressed, still in RAMHigh — compress on evict, decompress on faultDensity only; the pool eats RAM tooPool fills, reclaim finds nothing left to compress, then OOMGood: microsecond "swap", no disk in the path
zswap + disk swapCompressed pool, spilling to diskHigh, with a disk tailDensity, plus a genuine overflowGraceful until the pool fills, then it degrades to disk swapBest general default; cap max_pool_percent

Note the fourth column, because it is the one that matters commercially. Only plain disk swap adds anything resembling capacity, and it adds the worst possible kind. The compressed paths do not add capacity at all — they increase the number of logical pages the machine can hold at a given moment, in exchange for CPU, which is a density play. If your business plan involves selling guest RAM you do not have, and the mechanism is swap, then you are not selling memory. You are selling latency, and you are quoting the price in gigabytes.

Swappiness is folklore in reverse, and PSI is the only honest alarm

`vm.swappiness` is not "how much the kernel swaps." It is the relative cost the reclaim path assigns to evicting anonymous pages versus reclaiming file-backed pages. The folk wisdom — lower is better, set it to 1, 10 if you are feeling generous — comes from a world of desktops and database servers with a large page cache full of re-readable file pages, where reclaiming file pages really is cheaper.

A microVM host is the opposite world. Look at the `/proc/meminfo` output above: `AnonPages` is nearly all of memory, because guest RAM is anonymous. There is barely any file cache to reclaim, and the file cache that does exist is your template rootfs images and snapshot memory files — the pages that make a 179 ms restore possible. So a low swappiness on this machine does not prevent swapping; it just makes reclaim thrash the small amount of page cache you have first, evicting exactly the snapshot pages your next twelve creates are about to read, and then swap anyway because there was nothing else left. The result is slower creates and the same amount of swapping. On a host whose memory is almost entirely anonymous, a higher swappiness is the correct setting — it lets reclaim choose the genuinely cold thing instead of the hot thing that happens to be file-backed. Also note the range: since the anon/file cost rework, swappiness goes up to 200, not 100, and values above 100 are meaningful.

For alarms, forget free memory entirely. Free memory on a healthy host is approximately zero and stays there. Watermarks tell you the kernel is doing its job, not that anything is wrong. Pressure Stall Information is the measurement you want: `/proc/pressure/memory` reports the share of wall time tasks spent stalled waiting on memory, as `some` (at least one task stalled) and `full` (everything stalled), averaged over 10, 60 and 300 seconds. A rising `full avg10` means the machine is spending its time reclaiming rather than running guests, which is the actual condition you care about and it shows up before anything dies. On cgroup v2 you also get `memory.pressure` per cgroup — put each Firecracker in its own cgroup and you can attribute pressure to a specific tenant instead of to "the host."

Asking a guest what it actually uses

Admitting against a measurement requires a measurement, and there are three honest sources. The first is the host's own view: `/proc/<pid>/smaps_rollup` gives you `Rss`, `Pss`, `Swap` and friends for an entire process in one cheap read, which beats walking `smaps` line by line on a process with a multi-gigabyte mapping. `Rss` is what is resident now; `Pss` divides shared pages by the number of sharers, which matters once reflinked rootfs pages or deduplicated guest pages enter the picture; `Swap` is what the host has already pushed out, and if that number is not near zero you have already made the bet described in the next section.

The second source is cooperation from the guest. `MADV_FREE` and `MADV_DONTNEED` are how a process tells the kernel it no longer needs pages — `DONTNEED` drops them immediately and the next touch faults in a zero page, `FREE` marks them reclaimable but lets an untouched page be reused without a fault. These are what the virtio-balloon device turns into a usable lever: when the host inflates a guest's balloon, the guest driver hands pages back and the VMM madvises them away, and the guest's own allocator has to live with less. Firecracker's balloon also reports statistics and can deflate on guest OOM, which makes it the one mechanism that can actually reclaim memory from a guest that has decided it owns everything. It is also cooperative, which means a hostile or wedged guest ignores it — the balloon is a capacity tool, never a security boundary.

The third is to measure what the guest touches over time rather than what it holds: DAMON is the kernel's answer there, and I covered it separately. For capacity admission, resident-plus-pressure is usually enough and a great deal cheaper.

#!/usr/bin/env bash
# What each Firecracker guest is REALLY using, against what you sold it.
# Run this on the agent host. The kernel prints kB; we show MiB.
set -euo pipefail

# The `base` template bakes 4 GiB / 8 vCPU. vCPU and RAM are frozen INTO the
# snapshot -- Firecracker cannot change them at restore -- so this is exactly
# the number the scheduler charges in committed mode. It is the "sold" column.
BAKED_MIB=4096

sold=0; rss=0; swap=0
printf '%-8s %-38s %9s %9s %9s\\n' PID SANDBOX RSS_MiB PSS_MiB SWAP_MiB

for pid in $(pgrep -x firecracker); do
  # Firecracker is launched with `--id <sandbox-uuid>`; cmdline is NUL-separated.
  sbx=$(tr '\\0' '\\n' < "/proc/$pid/cmdline" | awk '/^--id$/{getline; print}')

  # ONE read for the whole address space. Walking /proc/<pid>/smaps instead
  # means parsing thousands of VMAs per guest; smaps_rollup pre-sums them.
  #   Rss  = pages resident right now
  #   Pss  = Rss with shared pages divided by their number of sharers
  #   Swap = pages the HOST has already evicted to swap. Should be ~0.
  read -r r p s < <(awk '
    /^Rss:/  {rr=$2}
    /^Pss:/  {pp=$2}
    /^Swap:/ {ss=$2}        # matches Swap:, not SwapPss:
    END      {print rr+0, pp+0, ss+0}' "/proc/$pid/smaps_rollup")

  printf '%-8s %-38s %9d %9d %9d\\n' "$pid" "${sbx:-?}" $((r/1024)) $((p/1024)) $((s/1024))
  sold=$((sold + BAKED_MIB)); rss=$((rss + r/1024)); swap=$((swap + s/1024))
done

echo
echo "configured (what the committed scheduler charges): ${sold} MiB"
echo "resident   (what the host is really holding):      ${rss} MiB"
echo "host swap  (pages already evicted -- want ~0):     ${swap} MiB"
awk -v s="$sold" -v r="$rss" 'BEGIN{
  if (s > 0) printf "working set is %.1f%% of committed\\n", 100*r/s }'

# Two numbers, one decision. A ratio around 20-30% is the normal steady state
# for build and agent workloads, and it is the entire argument for admitting
# against measurement. A ratio near 100% means your guests really are using
# what they asked for, and you should not be overcommitting this host at all.
#
# Caveat worth stating out loud: this is a snapshot of a monotonic curve. A
# long-lived guest's page cache only grows, so the honest planning figure is
# the resident high-water mark over a sandbox's lifetime, not this instant.

Why swap under a VM is a different animal

This is the part worth the page. Host swap under a normal process is a well-understood, mildly unpleasant performance event. Host swap under a VM is something else, because of one fact: when the host swaps out a page of guest memory, the guest does not know. It cannot know. The guest kernel believes it owns physical RAM with the latency characteristics of RAM, and it is making scheduling and locking decisions on that belief.

So consider what happens when a vCPU touches a swapped-out page. The guest issues what it thinks is a memory read — nanoseconds, no scheduling decision required, possibly with a spinlock held. On the host, that read is an EPT violation that lands in the vCPU thread, which now blocks in the page fault handler waiting on a disk read. The guest's clock keeps running. The guest's other vCPUs keep running, and keep queueing behind whatever lock the stalled one is holding. The guest's own memory reclaim cannot help, because from inside the VM there is no pressure to respond to: `MemFree` in the guest looks fine, PSI in the guest looks fine, every signal the guest has says the machine is healthy while one of its CPUs has been inexplicably standing still for forty milliseconds.

The host and the guest fight about which pages are hot

It gets worse than slow, because both kernels are running their own LRU over the same physical pages with completely different information. The host sees access through EPT accessed bits and concludes a region is cold. The guest knows that region is its hottest page cache but has no way to say so. So the host evicts it, the guest immediately touches it, the host faults it back in, and the pair of them oscillate. The classic pathological version is double paging: the guest decides to swap a page out to its own swap device, which requires reading that page — which the host has already swapped out — so the host must page it in from host swap purely so the guest can write it to guest swap. You have performed two disk round trips to free one page, and the page is now stored twice.

What this looks like when you are paged at 03:00

The symptoms do not look like memory pressure, which is why this costs people entire nights:

  • vCPU stalls with no corresponding guest-visible load — the guest's own metrics show an idle CPU that is somehow not making progress.
  • `BUG: soft lockup - CPU#2 stuck for 23s!` in the guest console, and RCU stall warnings, from a guest that is doing nothing wrong. The guest watchdog is correct: a CPU really was stuck. It just was not stuck on anything inside the guest.
  • Timeouts that make no sense in application terms — a health check failing against a process that is running, a heartbeat missed by a service with no work queued.
  • Tail latency that moves with other tenants' behaviour, which is the part customers notice and cannot explain, because the cause is not in their VM.
  • And the worst case: the host OOM killer, scanning for the largest RSS, picking a `firecracker` process. From the customer's side a VM simply stopped existing, with nothing in the guest logs, because the guest was never consulted. One tenant's burst, another tenant's dead sandbox.
Set `oom_score_adj` deliberately on the host, or the kernel will make this choice for you and it will choose badly. The largest RSS on a microVM host is always a guest, which means the default policy's favourite victim is a paying customer's VM — and possibly also your agent process, whose death turns one dead sandbox into an entire host that has stopped answering. Decide the order yourself, in advance, in a config file.

Two specific landmines

First, userfaultfd. A page installed by a UFFD handler is ordinary anonymous memory afterwards, so the host can and will swap it. Now you have two demand-paging systems stacked on the same region, each unaware of the other: a guest fault may be served by the handler from object storage, and a later fault on that same page may be served by the host from swap, and if the handler's own working memory gets reclaimed you can be waiting on a fault in order to serve a fault. The failure mode is a VM that appears to be running and is in fact wedged in page faults forever, which is specific enough that we run a watchdog for it — streamed guests are polled, and one whose faults have stopped making progress gets killed and reported as memory-unbacked rather than left to sit there looking healthy. Relatedly, a hugepage-backed guest cannot be swapped at all, because hugetlbfs pages are not swappable. That is a feature if you want predictable latency and a hard constraint if you wanted overcommit, and you should know which one you chose.

Second, and this one is purely self-inflicted: do not put swap on the same device as your snapshot store. The moment the host starts swapping, swap-in competes for queue depth with exactly the reads that serve snapshot restores and streamed memory chunks — so your create path slows down precisely when memory pressure is already causing you to want more hosts, and the two problems amplify each other. A separate device, or no swap, but not shared.

What I would actually do

An opinionated list, in the order I would implement it.

  1. Measure the working set instead of guessing it. `smaps_rollup` per Firecracker process, reported in the host heartbeat, high-water mark retained. You cannot admit against a number you do not collect.
  2. Admit against the measurement, with a reserve per in-flight create. This is the step people skip and it is the one that breaks: without a reserve, twenty simultaneous creates all evaluate against the same free pages and all succeed. A fraction of each ask with a sensible floor is enough.
  3. Run zswap with `zstd` and `zsmalloc`, cap `max_pool_percent` somewhere around 20, and give it a small real swap device on its own disk to spill into. Treat the compression ratio as a density bonus you discover, not as inventory you sell.
  4. Keep a hard floor of genuinely free RAM and refuse creates below it. A scheduler that can say no is worth more than one that is clever. On PandaStack a host below its store or memory floor is skipped by the scheduler and the create goes elsewhere, which is the only correct answer: the right response to a full host is a different host.
  5. Alert on PSI, specifically a rising `full avg10` on `/proc/pressure/memory`, plus nonzero `Swap` in any guest's `smaps_rollup`. Do not alert on free memory; you will either page constantly or never.
  6. Set `oom_score_adj` so the kernel's last resort is your choice. Protect the control-plane agent above the guests.
  7. Keep databases, and anything else whose tail latency is the product, entirely out of the overcommit pool. Charge them committed.
  8. Never let swap be the plan for the common case. It is the airbag, not the engine. If swap is load-bearing in your capacity model on a normal Tuesday, you have already shipped the incident and are waiting for the trigger.

And the sentence I would want on the whiteboard: selling swap as RAM is a bet against your customers' simultaneity. The bet is usually fine, because tenants genuinely are uncorrelated most of the time — right up until the morning everyone's CI starts at once, every guest demands its working set in the same ninety seconds, and correlation arrives all at the same time, which is the only way correlation ever arrives.

Here is the experiment that makes the gap concrete. It does not prove swap is safe or unsafe; it shows you the two numbers that decide whether the question even comes up on your workload.

# The committed-vs-working-set gap, measured rather than argued about.
# Creates N sandboxes, has each touch a known amount of guest memory, and
# prints what you SOLD against what each guest actually reports holding.
# Then you run the host-side script from earlier and compare the totals.
import json, time
from pandastack import Sandbox

N = 6
TOUCH_MIB = 400     # a realistic "did some actual work" footprint per guest
BAKED_MIB = 4096    # `base` bakes 4 GiB / 8 vCPU. vCPU and RAM are frozen
                    # into the snapshot -- Firecracker cannot change them at
                    # restore -- so passing memory_mb here would simply be
                    # overridden to the baked value. Do not pretend otherwise.

# Allocate, then WRITE one byte per page. A read of untouched anonymous memory
# is served by the shared zero page and costs the host nothing; only a write
# makes a page private and resident. Benchmarks that only read measure nothing.
TOUCH_PY = (
    "import time\n"
    "n = {mib} * 1024 * 1024\n"
    "buf = bytearray(n)\n"
    "for off in range(0, n, 4096):\n"
    "    buf[off] = 1\n"
    "print('touched', n // (1024 * 1024), 'MiB', flush=True)\n"
    "time.sleep(600)\n"            # hold the pages so the host can be measured
)

boxes = []
try:
    for i in range(N):
        sbx = Sandbox.create(
            template="base",
            ttl_seconds=900,                  # IDLE clock, not a wall clock
            metadata={"exp": "host-overcommit", "n": str(i)},
        )
        sbx.filesystem.write("/tmp/touch.py", TOUCH_PY.format(mib=TOUCH_MIB))

        # Fence it IN-GUEST with timeout(1). One-shot exec() does not actually
        # enforce timeout_seconds and exec_stream only widens the HTTP timeout,
        # so an allocator loop you got slightly wrong is otherwise unbounded.
        sbx.exec(
            "nohup timeout --kill-after=10s 300 python3 /tmp/touch.py "
            "> /tmp/touch.log 2>&1 & echo started",
            check=True,
        )
        boxes.append(sbx)
        print(sbx.id, "created")   # snapshot-restored: p50 179 ms, p99 203 ms

    time.sleep(25)                 # let every guest finish writing its pages

    for sbx in boxes:
        # The guest's own view, for comparison with the host's view. These two
        # disagree, and the disagreement is the subject of this whole post.
        inside = sbx.exec(
            "cat /tmp/touch.log; "
            "awk '/^(MemTotal|MemAvailable|Cached)/{print $1, $2}' /proc/meminfo"
        ).stdout.strip().replace("\n", " | ")
        print(sbx.id, inside)

    print(json.dumps({
        "sandboxes": N,
        "sold_mib": N * BAKED_MIB,              # committed-mode arithmetic
        "touched_mib": N * TOUCH_MIB,           # the floor on resident memory
        "ratio": round(N * TOUCH_MIB / (N * BAKED_MIB), 3),
    }, indent=2))
    print("now run the smaps_rollup script on the agent host and compare")
finally:
    # Unconditional teardown. kill() is the call -- there is no sbx.delete().
    # Leaving six 4 GiB guests resident because an assertion blew up is how
    # a density experiment becomes a capacity incident.
    for sbx in boxes:
        sbx.kill()

The honest summary is that host swap does not solve the density problem, it relocates it from a capacity error you can see in a dashboard to a latency distribution you will find out about from a customer. The thing that actually buys density is knowing the difference between committed and resident, admitting against the measured one, and having somewhere else to send a create when a host is full. Compression — zram or zswap — is a real and worthwhile multiplier on top of that. Disk swap is an airbag. If you are running your own fleet and your current plan is "add swap," add the measurement first; you will usually find you had considerably more room than committed accounting was willing to admit, and that you did not need to make the bet at all.

And the case against using PandaStack for this: if your workload's memory really is hot — a long-running analytics engine, a big in-memory cache, anything whose steady-state resident set is close to its configured size — then overcommit buys you nothing and you should rent a machine with the RAM you need and no neighbours. Our model earns its keep when guests are numerous, short-lived and mostly idle, which describes agents, CI jobs and preview environments very well and describes a 200 GiB JVM not at all. If you want the economics in full, MicroVM Density: The Economics of Per-Tenant Isolation has the arithmetic, and pricing has the rates it is computed from.

Frequently asked questions

Should I enable zswap or zram on a host running many microVMs?

Both are reasonable, and the choice follows from whether you have a disk you are willing to swap to. zram is the simpler proposition: a compressed block device in RAM used as swap, so pages "swapped" to it never leave physical memory and a fault against them costs a decompression rather than a disk round trip. That is microseconds instead of milliseconds, which matters enormously when the faulting thread is a vCPU. The catch is that the compressed pool occupies RAM too, so zram does not create headroom so much as improve the exchange rate — once the pool is full there is nothing underneath it and reclaim has nowhere left to go. zswap is the better general default on a server: a compressed cache in front of a real swap device, so you get the same cheap-fault behaviour for the hot portion of the evicted set plus an actual overflow path when the pool fills. Configure it with the zstd compressor and the zsmalloc allocator — zbud caps you at two compressed pages per physical page regardless of how compressible the data is, and z3fold is on its way out of the kernel — and cap max_pool_percent so the pool itself cannot become the memory problem. Put the backing swap device on a different disk from your snapshot store. The thing neither of them does is give you capacity you can sell, and that distinction is where most of the trouble comes from.

Why is swapping a VM's memory worse than swapping a normal process?

Because the guest has no idea it happened and keeps making decisions as if it had RAM. When the host evicts a page of guest memory, the guest kernel's view of the world does not change: it still believes it owns physical memory with nanosecond access latency. So when a vCPU touches that page, what the guest thinks is a memory read becomes a disk read that blocks the vCPU thread on the host — possibly while the guest holds a spinlock, with its other vCPUs queueing behind it. The guest cannot respond to pressure it cannot observe; its own MemFree and PSI look perfectly healthy throughout. Worse, both kernels run their own LRU over the same pages with different information, so the host evicts what it judges cold while the guest knows it is the hottest thing it has, and the two oscillate. The pathological case is double paging: the guest decides to swap a page to its own swap device, which requires reading a page the host already swapped out, so the host faults it in purely so the guest can write it out again. Two disk round trips to free one page. What you observe is not graceful degradation — it is soft lockups, RCU stalls and health-check timeouts in guests doing nothing wrong, and tail latency that moves with other tenants' behaviour.

How many microVMs actually fit on one host?

It depends on which number you admit against, and the two answers differ by roughly a factor of five on real workloads. In committed accounting — charge every sandbox its full configured size at create time — a 32 GiB host holds about seven 4 GiB sandboxes and that is the end of the conversation. It is safe and it is also why a fleet of ours once started refusing creates with a capacity error while most of its physical RAM sat idle: the arithmetic was correct, it was just arithmetic over promises rather than pages. Admitting against measured working set is what changes the number, because build, agent and CI workloads typically sit at 20-30% of their configured size, so the same machine holds several times as many guests. The ceiling that is not memory is networking: we pre-allocate 16,384 /30 subnets per host, the hard cap on sandboxes per agent, and nothing has come close to it because memory and CPU bind first by a wide margin. The real answer for your fleet is a measurement, not a formula. Run smaps_rollup across your Firecracker processes, keep the resident high-water mark per sandbox lifetime rather than an instantaneous sample, and set your admission target below the point where PSI full pressure starts rising. Then keep a hard free-RAM floor underneath all of it, because the correct response to a full host is a different host.

Is vm.swappiness=1 a good idea on a hypervisor host?

Usually not, and the advice is inherited from a workload that looks nothing like this one. Swappiness is not a volume dial for swapping; it is the relative cost the reclaim path assigns to evicting anonymous pages versus reclaiming file-backed ones. The classic low-swappiness advice comes from machines with a large page cache full of cheaply re-readable file pages, where preferring file reclaim genuinely is cheaper. A microVM host inverts that: guest memory is anonymous, so anonymous pages are nearly all of memory, and the file cache that remains is your template rootfs images and snapshot memory files — precisely the pages that let a create complete in 179 ms. Set swappiness to 1 there and reclaim will dutifully shred that cache first, slowing every subsequent restore, and then swap anyway because there was nothing else to take. You paid the cost twice. On a host whose memory is almost entirely anonymous, a higher swappiness is correct, because it lets reclaim pick the genuinely cold thing rather than the hot thing that happens to be file-backed. One detail worth knowing: the valid range goes to 200, not 100, since the anon/file cost model rework, so values above 100 are meaningful rather than clamped. And whatever you set, swappiness changes which pages get evicted, not whether your capacity plan is sound — it is a tuning knob, not an admission policy.

What should I monitor to know I have overcommitted too far?

Not free memory, which is the instinct and is useless. MemFree on a healthy Linux host is close to zero and stays there by design, because unused memory is wasted memory; watching it tells you the kernel is working, never that you are in trouble. Watch three things instead. First, PSI: /proc/pressure/memory reports the fraction of wall time tasks spent stalled on memory, as some (at least one task stalled) and full (everything stalled), averaged over 10, 60 and 300 seconds. A rising full avg10 means the machine is spending its time reclaiming instead of running guests, and it moves before anything dies. On cgroup v2, memory.pressure per cgroup lets you attribute that pressure to one tenant rather than to the host in general, which is the difference between a useful page and a shrug. Second, the Swap field in each guest's /proc/<pid>/smaps_rollup. Any sustained nonzero value there means the host has started paging a customer's VM, and from the customer's side that is invisible — they just see unexplained stalls. Treat it as a capacity alarm, not an informational metric. Third, the gap between Committed_AS and AnonPages in /proc/meminfo, which tells you how much of your overcommit is still theoretical. And set oom_score_adj deliberately, because the default policy's preferred victim on this kind of host is whichever guest is largest, which is a customer, with no evidence of the cause visible anywhere inside their VM.

Keep reading

Related posts

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.