all posts

MGLRU and MicroVM Density: Reclaim When Every Page Is Someone's Guest

Ajay Kumar··11 min read

Reclaim is the part of the kernel that decides, when memory runs short, which pages to take away from whoever has them. On a laptop that is a question about your browser. On a host running dozens of Firecracker microVMs it is a question about somebody else's kernel's working set, answered with information already filtered through that other kernel — and the classic two-list page LRU is close to perfectly designed to answer it badly.

I'm Ajay, I build PandaStack, and memory on one of our hosts is an odd population: a few dozen guests whose RAM is either an `mmap`ed snapshot file or anonymous pages installed by a userfaultfd handler, most idle, a few suddenly very much not. Linux 6.1 merged a replacement for the aging half of reclaim — MGLRU, `CONFIG_LRU_GEN`, generations instead of two lists. Here is what the old scheme did, what MGLRU changes, the files you poke, and what it actually buys: the shape of your bad days, not how many guests fit.

One distinction for the whole post: this is the HOST kernel's reclaim, arbitrating between VMM processes. Every guest runs its own LRU over its own memory on its own kernel — ours is 5.10, which predates MGLRU entirely — so enabling MGLRU on the host changes nothing about how a guest ages its own pages. A modern host with an older guest kernel is a normal configuration, and conclusions do not transfer between layers.

What the host actually knows about guest memory

From the host's point of view a running microVM is one process with one enormous mapping, and its shape depends on how the guest started. Restore from a local snapshot file and the VMM `mmap`s that file `MAP_PRIVATE`: guest memory is file-backed page cache, clean until the guest writes, at which point written pages become anonymous copy-on-write pages. Stream the memory instead and the mapping is anonymous from the start, registered with userfaultfd and populated by a handler that resolves faults out of object storage. Same guest, same bytes, two completely different reclaim objects.

Now the hard part: how does the host learn a page is hot? Guest accesses are translated by the second-level page tables KVM maintains, so the accessed bits recording guest activity live in the secondary MMU and reclaim reaches them through the `mmu_notifier` interface rather than by reading the VMM's own page tables. Whether a given kernel's aging path consults the secondary MMU cheaply, expensively, or at all is genuinely version-dependent — this area has seen active work alongside MGLRU, so read the MGLRU documentation and the KVM `mmu_notifier` code for the kernel you run rather than trusting any summary.

Even when that machinery works perfectly, look at what the signal means. The guest read a gigabyte of files into its own page cache, because free RAM is waste and every kernel believes that. To the guest those pages are trivially reclaimable: clean cache, drop, re-read later. To the host they are indistinguishable from the guest's heap, from a Postgres buffer pool, from the guest kernel's own text. The second kernel's opinion — the one thing you actually want — is not in the signal at all.

The two-list LRU, and what it reads

The classic scheme keeps four lists per NUMA node and per memory cgroup: active and inactive, split anonymous and file. New pages land on inactive. Reclaim takes candidates from the inactive tail; a page found to have been accessed gets a second chance and is promoted to active, and pressure pushes pages back down. Eviction takes clean file pages for free, writes dirty ones back, and swaps anonymous ones if there is swap, with `vm.swappiness` biasing the split.

It is a cache of recency with exactly two buckets of resolution, and that is the first problem. On a host where almost every page has been touched by somebody at some point, 'recently-ish' versus 'not recently-ish' flattens under sustained pressure. Everything looks hot, so the kernel scans a great deal to free a little, and that work is charged to whoever is allocating.

The second problem is how it finds accessed bits. A scan starts from a page on a list and walks that page's reverse mappings to find every PTE pointing at it — reasonable when pages are mapped once or twice. When one process maps four gigabytes of contiguous guest memory, you are doing per-page reverse-map work across a region you could have walked sequentially, so aging cost scales with how much memory you are deciding about rather than how much you are freeing.

To its credit the old scheme is not blind to its own mistakes: an evicted page leaves a non-resident shadow entry, so a page that comes straight back is recognised as a refault — that is what the `workingset_*` counters expose, and it survives into MGLRU unchanged. It is also feedback after the fact: you learn the eviction was wrong by paying for it.

MGLRU: generations instead of two buckets

MGLRU replaces the two lists with a small, bounded set of generations numbered by sequence numbers running from `min_seq` to `max_seq`, maintained per memory cgroup, per node, and per type. Aging increments `max_seq`, creating a new youngest generation, and pages found to have been accessed are promoted into it. Eviction takes from the oldest generation and increments `min_seq`. The kernel defines a minimum and maximum generation count — read the docs for the number on your kernel rather than taking one from me.

Two things follow. The first is resolution: instead of a binary hot/cold you get an ordered sequence of 'how long since anybody touched this', a far better basis for choosing a victim when the honest answer about most pages is 'a while ago, but not forever'. Pages are further ordered within a generation by how often they have been referenced.

The second is how aging is performed, and for a VM host this is the real change. MGLRU ages by walking page tables in address order and harvesting accessed bits, rather than starting from pages and walking reverse maps, and accessed bits on non-leaf entries let the walker skip whole subtrees nothing has touched. For a process mapping gigabytes of guest memory in one region — which is to say, every VMM on the box — a sequential walk with subtree skipping is simply a better algorithm for the same question.

Per-memcg generations deserve their own sentence. Each cgroup ages on its own clock, so a guest that goes cold ages out without its recency data being flattened by a noisy neighbour. If you run one cgroup per VMM — and you should, if only so `memory.pressure` means something — this is what makes per-sandbox reclaim behaviour legible.

The files you poke

#!/usr/bin/env bash
# mglru-inspect.sh -- is MGLRU available, is it on, and is reclaim the problem?
# Run this on the HOST (the kernel doing the reclaiming), not inside a guest.
set -uo pipefail

uname -r

# ---------------------------------------------------------------- 1. compiled?
# MGLRU is CONFIG_LRU_GEN, merged in 6.1. CONFIG_LRU_GEN_ENABLED decides
# whether it is on at boot; without it the code is present and dormant.
if [ -r /proc/config.gz ]; then
  zgrep -E 'CONFIG_LRU_GEN(_ENABLED)?=' /proc/config.gz
else
  grep -E 'CONFIG_LRU_GEN(_ENABLED)?=' "/boot/config-$(uname -r)" 2>/dev/null \
    || echo "no readable kernel config -- use the sysfs probe below"
fi

# ------------------------------------------------------------------ 2. sysfs
# The directory only exists if the code is compiled in. Missing => your distro
# kernel does not have MGLRU and nothing below applies.
[ -d /sys/kernel/mm/lru_gen ] || { echo "!!! no MGLRU"; exit 0; }

cat /sys/kernel/mm/lru_gen/enabled       # a FEATURE BITMASK, not a boolean
cat /sys/kernel/mm/lru_gen/min_ttl_ms    # 0 = no youngest-generation protection

# Enabling: `y` turns on every feature this kernel supports, `n` reverts to the
# two-list active/inactive LRU. Both at runtime, no reboot:
#
#   echo y | sudo tee /sys/kernel/mm/lru_gen/enabled
#   echo n | sudo tee /sys/kernel/mm/lru_gen/enabled
#
# You will also find posts (this one included) that want to hand you a hex
# value for individual features. DO NOT copy a bitmask out of a blog. The bit
# assignments and the feature set are kernel-version dependent, and a wrong bit
# silently disables the thing you wanted. Read
# Documentation/admin-guide/mm/multigen_lru.rst in the tree of the kernel YOU
# run -- for this file and for min_ttl_ms, it is the only trustworthy source.

# ---------------------------------------------------------------- 3. debugfs
# The generations themselves: per memcg, per NUMA node, the sequence numbers
# (min_seq..max_seq) and what sits in each generation. This is how you watch
# aging happen instead of inferring it.
mountpoint -q /sys/kernel/debug || sudo mount -t debugfs none /sys/kernel/debug
sudo sed -n '1,60p' /sys/kernel/debug/lru_gen        # compact
# sudo sed -n '1,80p' /sys/kernel/debug/lru_gen_full # + per-tier detail
#
# This file is also WRITABLE: commands to force aging or eviction for a given
# memcg/node. That is a testing interface -- exactly what you want in a
# benchmark harness and exactly what you do not want in a cron job. Syntax is
# in the same .rst, again for your kernel and not for mine.

# ----------------------------------------------------- 4. is reclaim hurting?
# The part people skip, and the only part that answers a question. Every
# vmstat counter here is cumulative since boot: one reading tells you nothing,
# two readings a minute apart tell you everything.
cat /proc/pressure/memory
#   some  = at least one task stalled on memory
#   full  = EVERY non-idle task stalled -- the box is doing nothing but reclaim
#   totals are monotonic microseconds. Alert on the derivative of `total`, not
#   on avg10, if you want something that survives your scrape interval.

grep -E '^(pgscan|pgsteal|pgrefill|workingset|pgmajfault|pswp|allocstall)' \
     /proc/vmstat
#   pgscan_* vs pgsteal_*     pages examined per page freed. A big ratio is the
#                             signature of "everything looks hot".
#   pgscan_direct/allocstall  reclaim done BY the allocating task: latency
#                             charged straight to whoever wanted memory. On a
#                             microVM host, that is a guest.
#   workingset_refault_*      evicted it, it came right back. Thrash.
#   pswpin/pswpout            anonymous memory moving. On a streamed (UFFD)
#                             host, that is guest RAM going to swap.
# Counter NAMES move between kernels -- the workingset_refault, pgscan and
# pgsteal families were split by anon vs file at different points. Grep for the
# family; never hardcode one name.

# Per-cgroup attribution: host-wide numbers say the box is unhappy, cgroup
# numbers say which sandbox is paying for it. This is why MGLRU's per-memcg
# generations matter operationally.
for d in /sys/fs/cgroup/pandastack.slice/*; do
  [ -r "$d/memory.pressure" ] || continue
  printf '%s  %s\n' "$(basename "$d")" "$(head -1 "$d/memory.pressure")"
done

The instruction in there worth repeating: do not copy a feature bitmask out of a blog post. `Documentation/admin-guide/mm/multigen_lru.rst`, in the tree of the kernel you actually run, is the only source worth trusting for `enabled` and for `min_ttl_ms`.

Two-list LRU vs MGLRU, for someone packing microVMs
DimensionTwo-list active/inactive LRUMGLRU (CONFIG_LRU_GEN, 6.1+)
Recency resolutionTwo buckets per type plus a referenced bit for a second chance.A bounded set of numbered generations per memcg and per node; access promotes a page to the youngest.
How accessed bits are foundStart from a page on a list, walk its reverse mappings to every PTE.Walk page tables in address order; non-leaf accessed bits let the walker skip untouched subtrees.
Cost over a huge contiguous mappingPer-page rmap work scaling with the size of the region, not with the amount being freed.Sequential walk with skipping — a much better fit for a VMM mapping gigabytes of guest RAM.
Isolation between guestsLists are per-memcg, but aging pressure flattens recency broadly under load.Generations advance per memcg, so one busy sandbox does not destroy another's aging information.
Thrash protectionRefault detection via non-resident shadow entries; grow the inactive list when refaults say it was too small.The same refault accounting, plus min_ttl_ms to protect the youngest generation — see the warning below.
What it still cannot doSee inside a guest.See inside a guest. Guest page cache and guest heap remain indistinguishable.
Effect on density—None directly. Better victim choice and cheaper aging show up as tail latency, not capacity.
`min_ttl_ms` is not a gentle knob. Setting it non-zero asks the kernel to protect the youngest generation for at least that long, and the documented consequence of being unable to is that the OOM killer is invoked instead of the working set being evicted. On a host full of microVMs the OOM killer's victim is a VMM process — a customer's sandbox dies instantly rather than getting slow. Sometimes that is the trade you want: a fast attributable death beats a host that thrashes for twenty minutes while everyone's p99 quietly triples. Decide it on purpose, read the docs for your kernel, and watch for the kills.

Two kernels, one page: the double LRU

Every microVM on the box runs a second, complete memory-management subsystem over the same physical pages, and it does not know you exist. The guest was told it has a fixed amount of RAM at template build time — we bake `base` and `browser` at 4 GiB, `code-interpreter` and `agent` at 2 GiB, `postgres-16` at 1 GiB, all with eight burstable vCPUs — and it manages that RAM as though it owned hardware. That is the contract of a VM. It is also the problem.

So two LRUs hold overlapping jurisdiction with no channel between them. The guest fills its page cache because unused RAM is waste; the host sees touched pages and keeps them, because that is what its signal says. Nobody reclaims, and the host sits at high utilisation holding memory one of the two kernels considers free. Or the other failure: the host reclaims a page the guest had already written off as cold cache, the guest re-reads it, and accounts for that as a slow disk. Your customer opens a ticket about I/O latency. The problem is host memory, and nothing in their metrics will say so.

The asymmetry that makes policy matter

Here is the part specific to snapshot-restore platforms. A restored guest's memory is file-backed until it writes, and clean file-backed pages are the cheapest thing reclaim can take: no writeback, no swap, just drop them. That is a large part of why you can overcommit such a host at all.

Then it has to fault them back, and that cost is not symmetric with the eviction. On a local-file restore it is a read from a file probably still in page cache. On a streamed restore it can be a network round trip: our handler resolves a missing page by fetching a 4 MiB chunk from object storage over a Range GET, with a zero-chunk header so an all-zero chunk costs no fetch, and a prefetch trace that pulls hot chunks in the background. Cheap to evict, potentially expensive to recover — that asymmetry is the whole reason better victim choice is worth having here.

One precision the folklore gets wrong: the network round trip is the cost of the FIRST touch of a page, not of a reclaim refault. Once a page is installed into an anonymous userfaultfd-registered mapping, reclaiming it is a swap-out rather than a drop — and a swap entry in a PTE is not a missing page, so the kernel resolves the next fault from swap without consulting your handler. The uncomfortable corollary: on a streaming host with no swap configured, that guest memory is not reclaimable at all. It is anonymous, resident and immovable, and all the pressure lands instead on page cache and on whatever guests were restored the old way. The restore path you chose decided whether guest memory is a cheap reclaim target or a brick.

This is a memory story only: our rootfs copy-on-write is local-only (XFS reflink or dm-snapshot) because copy-on-write wants a local block device, so streaming replaces the memory download and never the disk. The disk side stays ordinary page cache over local files — pleasantly, the exact case the classic LRU was designed for.

Hugetlbfs pages are not on these lists at all

Plainly, because it is the most common confusion here: hugetlbfs pages are not on the LRU lists and are not reclaimable. They are a separate allocator with a separate pool. A 2 MiB-hugepage-backed guest is not a reclaim candidate at any pressure level — it is a reservation, and MGLRU has no opinion about it. Transparent hugepages are a different thing entirely and do live on the lists, managed as one large folio, splittable and reclaimable. Conflate THP with hugetlbfs and you will tune the wrong subsystem with great confidence.

We use hugetlbfs anyway, because the fault economics of a restore are overwhelming: one fault covers 2 MiB instead of 4 KiB. The costs are real. Hugepage-ness is a property of the snapshot rather than the restore, so a hugepage snapshot can only come back through the userfaultfd path — the plain memory-file path is rejected outright, which is why every snapshot from a hugepage guest carries a marker forcing the streaming path regardless of any flag. And because the pages are a reservation, a boot-time static pool pins RAM idle whether anyone needs it or not, so we use the overcommit pool sysctl instead. Backing guest memory with hugetlbfs opts that memory out of the entire reclaim discussion. That can be the right call; it is never a discovery you want to make mid-incident.

The balloon is the only honest signal

There is exactly one mechanism in this stack where the guest's own memory manager tells the host something true. Inflate a virtio-balloon and the guest kernel allocates pages through its own allocator — its own LRU, its own free lists, its own idea of what it can spare — and hands the page numbers back for the VMM to release. That is not the host guessing from accessed bits filtered through a foreign kernel; that is the kernel that actually knows saying 'I genuinely do not want these'. It is strictly better information than any amount of scanning, and the only strictly better information available.

Firecracker has a balloon device with deflate-on-OOM and a statistics interval. Whether a particular version implements free-page hinting or free-page reporting as distinct virtio features, and how the device behaves across snapshot and restore, are version questions with real caveats attached — read the Firecracker balloon and snapshot documentation for the version you run rather than taking anyone's summary as current.

Why not lean on it? Because it is a cooperative protocol with a guest running arbitrary customer code. A guest can be slow to inflate, or busy, or deliberately uncooperative; deflate-on-OOM hands the memory back at the worst possible moment, by design; and a guest baked at 4 GiB and ballooned down hard is a guest whose own reclaim is now thrashing — you moved the problem somewhere you cannot see or fix it. There is also a structural reason the balloon looms large: Firecracker cannot change vCPU count or guest RAM at snapshot restore. RAM is chosen once, at template build time, and a `memory_mb` passed on a create is silently corrected to the baked value. The balloon is the only runtime size knob that exists, which makes it conceptually important and a bad default at once.

Is reclaim actually hurting you?

The symptom profile is why this goes undiagnosed for weeks. A guest under host reclaim pressure does not OOM and sees no pressure of its own: it believes it has all its RAM, because it does — touching some of it just occasionally takes an improbably long time. Inside the guest that presents as inexplicable latency, a p99 moving with no change in guest CPU, memory or I/O, so the owner profiles their application for a week. The information is on the host, in two places.

First, PSI, in `/proc/pressure/memory`. The `some` line counts time when at least one task was stalled on memory; `full` counts time when every non-idle task was. On a host, `full` above zero means whole intervals spent doing nothing but reclaim, and it is the number that correlates with people reporting latency they cannot account for. Alert on the derivative of the monotonic microsecond total, which survives a scrape interval, rather than on `avg10`. Per-cgroup `memory.pressure` gives attribution, and attribution is what turns a graph into an action.

Second, `/proc/vmstat`. The ratio of `pgscan_*` to `pgsteal_*` is how many pages the kernel examined to free one — the direct measure of 'everything looks hot'. `pgscan_direct` and the `allocstall_*` family mean reclaim was performed by the task that wanted the memory, so the latency was charged straight to a guest instead of absorbed by kswapd. `workingset_refault*` is the thrash signal, and `pswpin`/`pswpout` show anonymous memory moving, which on a streaming host is guest RAM going to swap. Counter names have been split and renamed across versions, so grep for the family rather than hardcoding one.

Then measure instead of believing. Turning MGLRU on is one `echo`, which makes it an unusually easy thing to A/B — so do that, with the same guest count and the same workload, and compare deltas:

#!/usr/bin/env python3
"""mglru-harness.py -- measure, do not believe.

Run it twice with the same N and the same workload: once with MGLRU off, once
with it on, and compare the DELTAS. Absolute vmstat numbers are cumulative
since boot and mean nothing alone; absolute PSI averages mean nothing without
a baseline.

The SDK half runs anywhere. The sampling half must read /proc on the AGENT
HOST, because that is the kernel making the reclaim decisions. Single-node
install: leave SAMPLE_CMD empty. Fleet: point it at one agent AND make sure
the sandboxes land there, or you are averaging reclaim behaviour across
machines and measuring nothing.
"""
import json
import os
import shlex
import subprocess
import time

from pandastack import Sandbox

N = int(os.environ.get("N", "8"))
TEMPLATE = os.environ.get("TEMPLATE", "code-interpreter")   # 2 GiB baked
TOUCH_MB = int(os.environ.get("TOUCH_MB", "900"))
SECONDS = int(os.environ.get("SECONDS", "180"))
SAMPLE_CMD = os.environ.get("SAMPLE_CMD", "")               # e.g. "ssh agent-1"

# Guest-side toucher. Allocating memory is not the same as having it: an
# untouched page is not a page the host has to back. So write one byte per
# 4 KiB page, then keep re-touching -- that is what makes these pages
# genuinely hot rather than merely allocated.
TOUCH_PY = """
import os, time
mb = int(os.environ["TOUCH_MB"])
buf = bytearray(mb * 1024 * 1024)
for off in range(0, len(buf), 4096):
    buf[off] = 1
deadline = time.time() + float(os.environ["HOLD_SECONDS"])
while time.time() < deadline:
    for off in range(0, len(buf), 1 << 20):
        buf[off] = (buf[off] + 1) & 0xFF
    time.sleep(0.05)
"""


def host_read(path):
    out = subprocess.run(shlex.split(f"{SAMPLE_CMD} cat {path}".strip()),
                         capture_output=True, text=True)
    out.check_returncode()
    return out.stdout


def sample():
    vmstat = {}
    for line in host_read("/proc/vmstat").splitlines():
        k, _, v = line.partition(" ")
        vmstat[k] = int(v)
    psi = {}
    for line in host_read("/proc/pressure/memory").splitlines():
        f = line.split()
        psi[f[0]] = int(f[-1].split("=")[1])      # total=, microseconds
    return {"t": time.time(), "vmstat": vmstat, "psi": psi}


def report(first, now):
    v, f = now["vmstat"], first["vmstat"]
    d = lambda k: v.get(k, 0) - f.get(k, 0)
    scan = d("pgscan_kswapd") + d("pgscan_direct")
    steal = d("pgsteal_kswapd") + d("pgsteal_direct")
    # The refault family was split by anon/file at different points in
    # different kernels, so sum whatever is present instead of naming one
    # counter and getting a KeyError on half the fleet.
    refault = sum(val - f.get(k, 0) for k, val in v.items()
                  if k.startswith("workingset_refault"))
    print(f"  +{now['t'] - first['t']:5.0f}s scan={scan:<9} steal={steal:<9} "
          f"s/s={(scan / steal if steal else 0):5.1f} refault={refault:<8} "
          f"psi_some={(now['psi']['some'] - first['psi']['some']) / 1e6:6.2f}s "
          f"psi_full={(now['psi']['full'] - first['psi']['full']) / 1e6:6.2f}s")


def main():
    boxes = []
    for i in range(N):
        # No memory_mb= and no cpu=: RAM is fixed at template build time
        # (--memory-mb) and the agent SILENTLY corrects whatever you pass here
        # to the baked value, so passing it would make this script read like a
        # lie about what it is testing.
        boxes.append(Sandbox.create(template=TEMPLATE,
                                    metadata={"harness": "mglru"},
                                    ttl_seconds=SECONDS + 300))

    for sbx in boxes:
        sbx.filesystem.write("/tmp/touch.py", TOUCH_PY)
        # Two gotchas in four lines.
        #  1. export BEFORE setsid, or the detached process never sees these.
        #     (Docker ENV from the template Dockerfile does not reach a
        #     detached process either -- only /etc/environment does.)
        #  2. `timeout` in the SHELL is the real limit. Neither exec endpoint
        #     enforces timeout_seconds server-side -- it is a CLIENT deadline --
        #     so a runaway toucher would happily outlive this script.
        sbx.exec("set -eu\n"
                 f"export TOUCH_MB={TOUCH_MB} HOLD_SECONDS={SECONDS}\n"
                 f"setsid sh -c 'timeout --signal=KILL {SECONDS + 30} "
                 "python3 /tmp/touch.py' >/tmp/touch.log 2>&1 &\n"
                 "echo started", timeout_seconds=60, check=True)

    first = sample()
    series = [first]
    try:
        while time.time() - first["t"] < SECONDS:
            time.sleep(10)
            now = sample()
            series.append(now)
            report(first, now)
    except KeyboardInterrupt:
        pass
    finally:
        for sbx in boxes:
            try:
                sbx.kill()
            except Exception as exc:      # a leaked sandbox costs real money
                print(f"  WARN: kill failed for {sbx.id}: {exc}")

    with open("mglru-run.json", "w") as fh:
        json.dump(series, fh)

    # What to look at, in order:
    #   psi_full > 0      the host spent whole intervals doing only reclaim.
    #                     This is the number that correlates with people
    #                     reporting latency they cannot explain.
    #   s/s rising        more pages examined per page freed: everything looks
    #                     hot, which is the flattening MGLRU should reduce.
    #   refault climbing  you are evicting the working set and reading it
    #                     straight back. That is thrash, not reclaim.
    # psi_full and s/s down with refault flat is the win -- and it is a TAIL
    # LATENCY win. It will not let you fit another guest on the box.


if __name__ == "__main__":
    main()

What you want to see is `psi_full` and the scan/steal ratio going down while refaults stay flat. If that is what you get, MGLRU is helping — and what it gave you is a better tail, not another guest on the box.

What actually moves density (it is not the LRU)

MGLRU is not a density multiplier and I would be suspicious of anyone selling it as one. It is a better algorithm for a question you should be trying not to ask: it changes the shape of degradation — fewer pathological thrash episodes, better victim choice, much cheaper aging over the enormous mappings a VM host is full of. If you are out of memory you are still out of memory, and a better eviction policy has simply picked a more deserving victim.

The levers that actually move density, roughly in order of how much they move:

  • Do not promise RAM you cannot back — admission control is the lever. Ours used to admit creates against committed memory, the sum of every guest's baked size, which refuses work while most of the host's RAM sits idle. Moving admission to a working-set view — the agent reports what guests actually touch, the scheduler admits against that, and every in-flight create reserves a slice so a burst cannot all land on one host — is the single change that raised density. Control plane, not kernel knob.
  • Elide the zeros. Most of a freshly booted guest's RAM is zeros. A snapshot header recording which chunks are non-zero means an all-zero chunk costs no fetch and installs a zero page, so the memory is never allocated at all. Memory you never allocate needs no reclaim policy, which is the best reclaim policy going.
  • Choose the restore path deliberately. A clean file-backed page is the cheapest reclaim target in the kernel; anonymous memory on a host with no swap is the most expensive, because it is not a target at all.
  • Give memory back for real instead of negotiating for it. Our sleep path deletes the guest and keeps a published seed, and the wake restores it — a fresh snapshot-restore create is p50 179 ms, a same-host fork 400–750 ms. Scale-to-zero beats balloon-to-small because the memory is genuinely gone rather than the subject of an ongoing conversation with a guest that may not be listening.
  • Then, and only then, reclaim policy. Turn MGLRU on because the aging is cheaper and the victim choice is better, measure, and expect a tail-latency result.
A better eviction policy is a better answer to the question of who suffers. Density is the business of never having to ask it.

What I keep relearning here is that the host kernel will do an impressive job of allocating a shortage fairly, and that this is not the same as having enough memory. MGLRU is the best version of that fairness I have run, and worth enabling on a microVM host for the aging cost alone. Just do not file it under capacity. File it under 'the bad days are less bad' — a real and underrated category — and keep the capacity work where it belongs: in admission control, in the zeros you never allocate, and in not telling a guest it has four gigabytes you cannot produce.

Frequently asked questions

Should I just turn MGLRU on, on a microVM host?

Probably yes, if your kernel has it, and for a narrower reason than the hype suggests. MGLRU's aging walks page tables in address order and can skip untouched subtrees using accessed bits on non-leaf entries, instead of starting from pages and walking reverse mappings. A microVM host is the pathological case for the old approach — one process mapping gigabytes of contiguous guest memory, repeated dozens of times — so the aging cost improvement is structural rather than incidental. Better victim choice from having many generations instead of two buckets is a bonus on top. Do it as an A/B, not as a leap of faith: it is a single write to `/sys/kernel/mm/lru_gen/enabled`, reversible at runtime with `n`, so run the same guest count and the same workload both ways and compare deltas in `/proc/pressure/memory` and in the `pgscan`, `pgsteal` and `workingset_refault` families in `/proc/vmstat`. Expect a tail-latency and thrash-frequency result. Do not expect to fit more guests on the box, and leave `min_ttl_ms` alone until you have read its documentation for your kernel, because its failure mode is invoking the OOM killer.

Does enabling MGLRU on the host improve memory behaviour inside the guests?

No, and keeping the layers separate matters here more than almost anywhere else. MGLRU is a host-kernel reclaim change: it alters how the host ages and evicts the pages backing guest memory. Each guest runs a complete, independent memory-management subsystem over the RAM it believes it owns, with its own LRU, its own page cache and its own OOM killer. Our guests run a 5.10 kernel, which predates MGLRU entirely, so the guest side is firmly two-list regardless of what the host does — and a modern host kernel with an older guest kernel is a normal, correct configuration rather than a mismatch to fix. If you want better aging inside a guest, that is a guest-kernel question, and on a platform like ours it is not a per-template choice: there is one guest kernel per host and you cannot swap it per template. The practical consequence for debugging is to decide which layer you are looking at before forming a theory. Guest-visible memory pressure and guest OOM kills are the guest's own reclaim. Host reclaim shows up in a guest as time that went missing, with no local explanation.

Is min_ttl_ms a good idea on a dense host?

It is a sharp instrument and it deserves a decision rather than a default. Setting it non-zero asks the kernel to protect the youngest generation for at least that long, and the documented behaviour when the kernel cannot honour that is to invoke the OOM killer rather than evict the protected working set. On a host full of microVMs the OOM killer's natural victim is a VMM process, so the policy you have actually chosen is: when memory gets tight, kill a sandbox outright instead of making everyone slow. There are real situations where that is the better outcome — a fast, attributable death is operationally far easier than a host that thrashes for twenty minutes while every tenant's p99 triples and nothing in anyone's dashboard says why. But it converts a latency problem into a liveness problem, which changes what your control plane must handle: something has to notice the kill, attribute it, and either restart or reschedule the work. Read the `min_ttl_ms` section of `Documentation/admin-guide/mm/multigen_lru.rst` for your kernel, and if you set it, instrument the kills before you ship it.

My guest shows unexplained latency but no memory pressure inside it. How do I confirm host reclaim is the cause?

That profile is the classic signature, because a guest under host reclaim pressure sees no pressure of its own: it believes it has all its RAM, and it does — touching some of it just occasionally takes an improbably long time. Confirm it from the host, in three places. First, `/proc/pressure/memory`: the `full` line's monotonic total growing means the box spent whole intervals doing nothing but reclaim, and per-cgroup `memory.pressure` tells you which sandbox is paying. Second, `/proc/vmstat`: a rising `pgscan_*` to `pgsteal_*` ratio means the kernel is examining many pages to free one, `pgscan_direct` and the `allocstall_*` family mean the allocating task did the reclaim itself so the latency was charged directly to a guest, and `workingset_refault*` climbing means you are evicting the working set and reading it straight back. Third, correlate timestamps against the guest's latency spikes, because reclaim pressure is bursty and an averaged graph will hide it. Note that `pgmajfault` on a snapshot-restore host is ambiguous: a major fault may be a first touch of never-yet-restored memory rather than a refault of something evicted, so read it alongside the refault counters rather than on its own.

Keep reading

Related posts

  • UFFDIO_ZEROPAGE vs UFFDIO_COPY: Stop Paying RAM for Zeros

    A demand-paged microVM restore answers thousands of page faults. If your handler answers all of them with UFFDIO_COPY, you are allocating anonymous pages to store nothing — the single biggest hidden cost in a restore fleet.

  • Guest page cache: why it bloats microVM snapshots

    A microVM snapshot captures the guest's RAM verbatim — including megabytes of file data Linux cached from a disk that's sitting right there. Drop the cache before you bake, and the memory image shrinks to the live working set.

  • KSM Memory Deduplication for MicroVMs — And Why It's a Trap

    Eighty microVMs running the same kernel and the same libc, and the host is storing eighty copies. KSM will merge them — for a permanent CPU tax, a merge that evaporates on first write, and a timing oracle across tenant boundaries. Here's how ksmd actually works, and why sharing at the snapshot layer beats scanning for duplicates you should never have made.

  • EEVDF vs CFS: What the New Linux Scheduler Means for MicroVM Density

    Linux 6.6 swapped the default CPU scheduler out from under everyone. If the runnable threads on your host are the vCPUs of a few hundred oversubscribed microVMs, the change is not academic — and half the tuning advice you'll find now edits sysctls that no longer exist.

  • tmpfs in a microVM: The Filesystem That Eats Your RAM

    Writing a gigabyte into /tmp inside a microVM does not consume disk. It consumes the guest RAM your process was counting on — and if you snapshot that VM, the gigabyte gets cloned into every copy.

More in Internals · See PandaStack benchmarks

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.