DAMON: Measuring What a Sandbox Actually Touches
Every capacity decision a sandbox platform makes reduces to one number: how much memory does this guest actually need. Not how much it was promised — how much it touches. Too high and you refuse paying work while the host's RAM sits idle. Too low and you discover the OOM killer's taste in victims.
I'm Ajay, I build PandaStack, and our answer used to be the laziest one available: charge the host for every guest's configured RAM, sum it, refuse creates when the sum hits the budget. That is committed sizing. It can never overdraw, and it is wrong by an enormous margin, because a guest handed 4 GiB to run an idle Node server touches a fraction of it and holds the rest as an accounting fiction. Replacing that fiction with a measurement is what moves density.
Residency is not access, and that is the whole problem
The accounting most platforms reach for first is residency. Read `/proc/<pid>/smaps_rollup`, sum `Rss` plus the hugetlb terms, and you have the bytes of physical memory this process currently holds. One file read, no kernel feature flags, and it is exactly what we meter today.
It is not a working set. Residency answers 'has this page been faulted in and not yet reclaimed'; a working set answers 'is anybody still using it'. Those diverge immediately, because nothing takes a page away from a process that stopped caring about it. A guest boots, its kernel touches a gigabyte laying out structures and filling page cache, its allocator later recycles that memory internally — and every one of those pages is still resident on the host. The guest reusing a page does not hand it back.
So residency ratchets toward a high-water mark and, on an idle VM, parks there. That is written into our pressure controller as a problem statement: the comment above its proactive-reclaim cycle says that without it an idle VM keeps its peak residency forever, because guest page reuse never returns host pages.
The fix is to measure access instead. The classic way is idle page tracking: write a bitmap to clear pages' accessed state, wait, read back what was touched. The mechanism is sound; the cost model is the problem, because clearing and re-reading accessed state is work proportional to the address space you ask about rather than to the information you get back. A VM host is dozens of processes each mapping gigabytes of mostly-untouched, sparse address space, so that scheme charges you most for the guests you learn least about.
DAMON's trick: regions, not pages
DAMON — Data Access MONitor, in `mm/damon` — refuses to be a per-page census. It is a sampler, and the unit it samples is a region. Every sampling interval, for each region, it picks one page at random, checks whether that page's accessed state has been set since last time, and clears it. Every aggregation interval it reports each region with the count it accumulated — that region's estimated access frequency for the window — and resets.
One random page per region would be a useless estimator for a region that is half hot and half cold. The piece that makes it work is adaptive region adjustment, running alongside aggregation: regions with similar measured frequencies are merged, and regions are split so the total moves toward a configured target. The partition migrates toward regions that are internally homogeneous — and where every page behaves the same, one sampled page is representative.
Two bounds govern the cost: a minimum and a maximum number of regions. The maximum is why DAMON is interesting on a VM host at all — overhead is bounded by how many regions you allow, not by how much memory they cover, so watching a 4 GiB guest mapping and a 64 GiB one cost the same. That inverts the economics of idle page tracking exactly where you need it inverted. The minimum stops the merger collapsing everything into one region and declaring the guest uniformly medium-warm.
The three intervals, and what each one buys
Three time constants decide what question you are actually asking. Take the values from your kernel's documentation; what follows is what each does to the answer.
- Sampling interval — how often one page per region is checked. Shorter means less noise per window and more sampling work: the knob that trades CPU for confidence. It also sets your detection floor, since a pattern faster than a sampling interval can slip between samples.
- Aggregation interval — the window a reported count covers. Not a sampling knob: it is the definition of 'hot'. For a sandbox serving a request every few minutes, a short aggregation interval reports — correctly and uselessly — that the guest is cold.
- Region update interval — how often DAMON re-derives its target from the operations set, so address space that appeared or vanished is accounted for. Nearly a no-op for a VMM whose mapping is established at restore; where mappings churn, too long means measuring ranges that no longer exist.
There is also a choice of what DAMON walks, the monitoring operations set: a virtual-address mode for a target process, a fixed-range variant where you name the ranges, and a physical-address mode. Availability depends on your kernel's configuration, and the choice is not cosmetic for a VM — it decides whether the accessed state you sample has anything to do with guest activity.
DAMOS: doing something about it, carefully
Measurement is half the subsystem. The other half is DAMOS, DAMON-based Operation Schemes: rather than exporting numbers for a daemon to act on, you declare a target access pattern and an action, and the kernel applies it to matching regions. A scheme has four parts, in order of operational importance.
- The access pattern — a size range, an access-frequency range, and an age range. Age is the underrated term: how long a region has held its current access state, which lets you say 'cold, and has been for a while' rather than 'cold right now'. That is the difference between a policy and a twitch.
- The action — page it out, mark it cold or wanted-soon, deprioritise or prioritise it on the LRU lists, apply or suppress hugepage treatment, or a do-nothing action that only counts. Which actions exist, and which are valid for which operations set, is version-dependent. Filters narrow what an action may touch, by properties like cgroup membership or anonymity.
- The quotas — the part that decides whether you can run DAMOS in production. A quota bounds the work a scheme may do per reset interval: time spent applying the action, and bytes it may apply it to. When matching regions exceed the quota, DAMOS prioritises among them using configurable weights over size, frequency and age. Without quotas, a scheme that matches too much under a pressure spike is an unbounded reclaim loop with root privileges.
- The watermarks — a metric plus a band. The scheme runs only while the metric sits inside it: do not bother when memory is plentiful, do not pile on when things have already gone wrong.
My position: run DAMOS with the counting action long before any reclaim action. Counting turns DAMON into the thing most platforms are missing — a kernel-side, quota-bounded, continuously updated answer to 'how much of this mapping is in play' — with no behavioural side effects at all.
The interfaces, with the version caveat they deserve
DAMON has had two kernel-side interfaces. The current one is sysfs, under `/sys/kernel/mm/damon/admin/`: a tree of kdamonds with contexts, monitoring attributes, targets and schemes, driven by writing to a `state` file. The older debugfs interface is deprecated and on a removal path, so a guide built around debugfs files predates the interface you have. And there is `damo`, the userspace front-end, which is what you should use for exploration.
#!/usr/bin/env bash
# damon-session.sh -- what does one sandbox's guest memory actually touch?
#
# Run on the HOST (the kernel that owns the mapping), not inside a guest.
# Paths and field names are KERNEL-VERSION DEPENDENT -- see the warning above
# this block. Verify against Documentation/admin-guide/mm/damon/ for your tree.
set -uo pipefail
uname -r
[ -d /sys/kernel/mm/damon/admin ] || { echo "!!! no DAMON sysfs on this kernel"; exit 0; }
# The target: the Firecracker process backing one sandbox. Its guest RAM is
# ONE host-side mapping, so from here "the guest's working set" is a question
# about ranges of host virtual addresses inside that mapping -- NOT about any
# process inside the guest. That limit is structural; see below.
SBX="${1:?usage: damon-session.sh <sandbox-id>}"
PID=$(pgrep -f "firecracker.*${SBX}" | head -1)
[ -n "${PID:-}" ] || { echo "no firecracker pid for ${SBX}"; exit 1; }
# What we are replacing: residency. One read, cheap, and it is a high-water
# mark rather than a working set. Keep it on screen for comparison.
grep -E '^(Rss|Shared_Hugetlb|Private_Hugetlb):' "/proc/${PID}/smaps_rollup"
# --------------------------------------------------------- the easy path
# Explore with damo first. Flags move between damo releases; read --help.
# damo start --ops vaddr --target_pid "$PID" \
# --sample_interval 5ms --aggr_interval 200ms
# damo show # current regions + access frequencies
# damo report access # access pattern snapshot
# damo report wss # working-set-size distribution over a record
# damo stop
# `damo report wss` is the number this whole post is about: not "bytes
# resident" but "bytes that anyone touched", as a distribution rather than a
# single figure -- which is the honest shape of the answer.
# --------------------------------------------------- the raw sysfs path
# Worth doing once by hand, because it is the only way to see that the cost
# bound is a REGION COUNT and not a byte count.
D=/sys/kernel/mm/damon/admin
echo 1 | sudo tee "$D/kdamonds/nr_kdamonds" >/dev/null
K="$D/kdamonds/0"
echo 1 | sudo tee "$K/contexts/nr_contexts" >/dev/null
C="$K/contexts/0"
cat "$C/avail_operations" # which operations sets THIS kernel offers
echo vaddr | sudo tee "$C/operations" >/dev/null
# The three intervals, in microseconds. These are the question, not the
# plumbing: aggr_us is literally your definition of "hot".
I="$C/monitoring_attrs/intervals"
echo 5000 | sudo tee "$I/sample_us" >/dev/null # per-region sample cadence
echo 200000 | sudo tee "$I/aggr_us" >/dev/null # window a count covers
echo 60000000 | sudo tee "$I/update_us" >/dev/null # re-derive target ranges
# The cost bound. Overhead tracks THIS, not the size of the mapping.
R="$C/monitoring_attrs/nr_regions"
echo 10 | sudo tee "$R/min" >/dev/null # floor: don't merge into mush
echo 1000 | sudo tee "$R/max" >/dev/null # ceiling: the actual cost knob
echo 1 | sudo tee "$C/targets/nr_targets" >/dev/null
echo "$PID" | sudo tee "$C/targets/0/pid_target" >/dev/null
# A scheme whose ACTION IS NOTHING. `stat` matches regions and counts them and
# touches no memory -- all of the measurement, none of the second-guessing.
# Start here. Always start here.
echo 1 | sudo tee "$C/schemes/nr_schemes" >/dev/null
S="$C/schemes/0"
echo stat | sudo tee "$S/action" >/dev/null
# Wide-open access pattern: match everything, so tried_regions becomes a
# snapshot of the whole mapping rather than a filter.
P="$S/access_pattern"
echo 0 | sudo tee "$P/sz/min" >/dev/null; echo 17592186044416 | sudo tee "$P/sz/max" >/dev/null
echo 0 | sudo tee "$P/nr_accesses/min" >/dev/null; echo 4294967295 | sudo tee "$P/nr_accesses/max" >/dev/null
echo 0 | sudo tee "$P/age/min" >/dev/null; echo 4294967295 | sudo tee "$P/age/max" >/dev/null
# Quotas, even for `stat`. Build the habit here, where it costs nothing, so
# that the day you set action=pageout you are not learning quotas for the
# first time on a host full of other people's VMs.
Q="$S/quotas"
echo 1000 | sudo tee "$Q/reset_interval_ms" >/dev/null
echo 10 | sudo tee "$Q/ms" >/dev/null # time budget per reset
echo 0 | sudo tee "$Q/bytes" >/dev/null # 0 = unlimited bytes
echo on | sudo tee "$K/state" >/dev/null
cat "$K/pid" # the kdamond kernel thread doing the work
# ------------------------------------------- reading the access pattern
# Snapshot route: ask the kernel to refresh the regions the scheme tried, then
# read them. Field names here are the ones most likely to have drifted.
sleep 10
echo update_schemes_tried_regions | sudo tee "$K/state" >/dev/null
for r in "$S/tried_regions"/*/; do
[ -d "$r" ] || continue
printf '%s-%s nr_accesses=%s age=%s\n' \
"$(cat "$r/start")" "$(cat "$r/end")" \
"$(cat "$r/nr_accesses")" "$(cat "$r/age")"
done | sort -n | head -40
# Read it as: a partition of the mapping, each part carrying an ESTIMATED
# access frequency for the last aggregation window and how long it has held
# that state. Sum the sizes of the regions with nr_accesses > 0 and you have a
# working-set estimate for that window. Compare it to the Rss figure printed
# at the top. The gap between those two numbers is your density argument.
# Scheme bookkeeping: how much the scheme matched, how much it applied, and
# how often the quota cut it off. qt_exceeds climbing means your quota is the
# binding constraint -- which is the correct way for it to fail.
cat "$S/stats/nr_tried" "$S/stats/sz_tried" \
"$S/stats/nr_applied" "$S/stats/sz_applied" "$S/stats/qt_exceeds" 2>/dev/null
# Streaming route, for anything continuous: the damon:damon_aggregated
# tracepoint emits per-region records each aggregation interval, which is what
# `damo record` consumes. Prefer this over polling sysfs in a loop.
# perf record -e damon:damon_aggregated -a -- sleep 60
echo off | sudo tee "$K/state" >/dev/null
The instruction in there worth tattooing on a runbook: sum the sizes of the regions with a non-zero access count, and compare that to the residency figure printed at the top of the same script. Two numbers, same process, same second. The gap between them is the entire density argument.
Why this is a density question: a microVM is one mapping
There is no warm pool on our platform — every create restores a baked Firecracker snapshot, p50 179 ms, p99 203 ms — so a host's population is a few dozen VMMs each holding one large memory region, most doing nothing at any instant. Guest sizing is baked into the snapshot: Firecracker cannot change vCPU or RAM at restore, so a `memory_mb` on a create is silently corrected to the template's baked value. Our `base` template is 4 GiB. An idle Node server on `base` has 4 GiB whether it wants it or not.
Now the admission question. A create arrives; does this host have room? The committed answer sums every live guest's baked size against a budget. It cannot overdraw, and it is why a 32 GiB host admits roughly seven 4 GiB sandboxes and then refuses work with most of its RAM genuinely unused. Committed sizing is not a conservative version of the right answer; it answers a different question — what did we promise — and capacity is a question about what is in use.
So the fleet default moved to a working-set view: the agent measures what guests actually hold, and the admission gate budgets against that measurement, plus a reserve for in-flight creates not yet measured, plus headroom the pressure controller needs. Database-class VMs are charged their full committed size, because 'never overcommitted' is a promise we make about managed Postgres. Every other create is charged a fraction of its baked size with a floor — and that fraction is a guess doing a lot of work.
Three ways to get the number, three ways to be wrong
First, ask the guest. The guest kernel is the only entity that actually knows what it needs, and a balloon device is the honest channel: inflate it and the guest allocates through its own machinery and reports what it can genuinely spare. The problems are structural. It is a cooperative protocol with a guest running arbitrary customer code, so a slow or unhelpful guest does not answer. The honest answer is usually unhelpful anyway: free RAM is waste, so every kernel fills page cache until 'used' means 'all of it'.
Second, host residency — `Rss` plus hugetlb from `smaps_rollup`, which is what we meter. Cheap, immune to a hostile guest, and the only one of the three you can read in a hot admission path without thinking about it. Its wrongness is one-directional: it over-counts, because it is a ratchet, so your capacity view slowly converges back on committed sizing — the thing you were escaping.
Third, access-frequency monitoring: DAMON. The number you wanted, bounded cost, continuously updated, with age information nothing else gives you. Its wrongness comes in three flavours. It is a sample, so it is a distribution estimate and will be wrong about any individual region. It cannot attribute anything. And it is silent on the cost of being wrong: it will tell you a region is cold, with no opinion about whether refaulting it means a swap read, a local file read, or a Range GET to object storage.
What each instrument actually answers
Five instruments get treated as interchangeable by anyone who has not had to act on their output. They answer genuinely different questions, and the differences decide which one you may pack against.
| Instrument | What it measures | Overhead shape | What it cannot tell you | What you can safely act on |
|---|---|---|---|---|
| Committed sizing | What you promised: the sum of every guest's baked memory size. | A database query. Free. | Anything about actual use — it is an accounting identity, not an observation. | Hard guarantees: the right basis for a class you promised never to overcommit, like managed Postgres. |
| Host residency (smaps_rollup Rss + hugetlb) | Bytes of physical memory a VMM currently holds. | One file read per sandbox per tick. | Whether anyone still wants those bytes. It ratchets to a high-water mark and parks there. | Admission and billing today — cheap, immune to a hostile guest, systematically high. |
| PSI (memory.pressure) | Time spent stalled waiting on memory, per cgroup and host-wide. | Always on, negligible, already there. | How much memory anyone needs, or which pages. It is an alarm, not a gauge. | Alerting and control-loop trip points. Never capacity planning. |
| MGLRU (lru_gen) | Nothing, as a product: it consumes access information to pick reclaim victims. | Ages by page-table walks rather than reverse-map walks — a better fit for huge mappings. | A working set you can export, bill or schedule against. Its debugfs file is a testing interface. | Reclaim quality. Expect a tail-latency result, not capacity. |
| DAMON (+ DAMOS stat) | Estimated access frequency and age per region of an address space. | Bounded by the configured maximum region count, independent of mapping size. | Which guest process is hot, how expensive a refault would be, or per-page truth. | Capacity decisions and tiering — once you have validated that the accessed state reflects guest activity. |
The honest limits, including the one that can invalidate everything
Start with the limit that is structural. From the host, a guest's memory is one flat region. DAMON will partition it and tell you the range at some offset is hot — and that offset is a guest physical address, allocated by the guest's own page allocator, which reuses frames continuously and tells nobody. You cannot learn that the customer's Node process is hot and their Postgres client is cold; you learn that this much of the guest's physical memory is in play. Sufficient for a capacity decision, maddening for a debugging session.
Then the limit that can invalidate the exercise. Guest memory accesses are translated through the second-level page tables KVM maintains, so the accessed bits recording guest activity are set in the secondary MMU, not necessarily in the VMM's own host page table entries. Whether a given kernel's monitoring path consults the secondary MMU — and the young-bit helpers used by the physical-address operations set do have notifier-aware variants, unlike walking a process's page tables directly — is version-specific, with active work behind it. Validate it empirically rather than trusting any summary, mine included: run a guest that touches a known fixed amount of memory in a loop, and check that DAMON's non-zero regions sum to roughly that amount. If they sum to approximately nothing, you have learned something far more important than any tuning parameter.
And one more time, because it is the limit people act on anyway: DAMON prices nothing. A locally-restored VM's cold page is clean file-backed cache, free to drop. A streamed VM's cold page is anonymous memory from a userfaultfd handler — a swap-out with swap, and not a reclaim target at all without it.
What we actually run: residency, a ladder, and a reclaim cycle
We do not run DAMON on the fleet. We run residency measurement plus a cgroup-v2 pressure controller, and the gap between the two is the interesting part.
The controller is a ladder on a ten-second tick. It reads `MemAvailable`, `MemTotal` and `SwapTotal` from `/proc/meminfo` and the `some` and `full` averages from `/proc/pressure/memory`, and classifies the host OK, elevated or critical — with hysteresis, so a level rises immediately but falls only after two consecutive quiet ticks.
On elevated or critical it sorts VMs coldest-first — lowest CPU rate, then oldest activity — and squeezes a bounded number per tick by writing `memory.high` to a fraction of current residency, with a floor so a guest keeps room to breathe. Two guards are worth stealing. It squeezes against the largest residency ever observed for that VM, not the post-squeeze figure, so repeated squeezes cannot ratchet a guest to nothing one tick at a time. And the rung is gated on swap being present, because `memory.high` without a reclaim destination does not shrink a cgroup, it stalls it.
And on a calm host, the anti-ratchet: periodically, for one long-idle VM at a time, it asks the kernel to proactively reclaim a slice via cgroup-v2's `memory.reclaim`. Why not just leave a cap in place? Because a `memory.high` cap only reclaims through the cgroup's own allocation paths, and an idle VM never allocates — we watched a capped idle VM hold several gigabytes under a cap it never tripped. Proactive reclaim is a push; a cap is a tollbooth on a road nobody is driving.
#!/usr/bin/env bash
# sandbox-cgroup.sh -- read one sandbox exactly as the pressure ladder does.
#
# Run on the agent host. The ladder lives in the agent's own delegated cgroup
# subtree, which it discovers from /proc/self/cgroup at startup -- so DERIVE
# the path rather than hardcoding the systemd unit name.
set -uo pipefail
SBX="${1:?usage: sandbox-cgroup.sh <sandbox-id>}"
AGENT_PID=$(pgrep -x pandastack-agent | head -1)
SVC="/sys/fs/cgroup$(awk -F: '/^0::/{print $3}' "/proc/${AGENT_PID}/cgroup")"
G="${SVC}/vm-${SBX}"
[ -d "$G" ] || { echo "no cgroup for ${SBX} under ${SVC}"; exit 1; }
# --- residency: the number admission and the squeeze target are computed from
awk '{printf "memory.current %.0f MiB\n", $1/1048576}' "$G/memory.current"
# "max" => unsqueezed. Anything else => the ladder (or you) capped this VM.
printf 'memory.high %s\n' "$(cat "$G/memory.high")"
printf 'memory.max %s\n' "$(cat "$G/memory.max")"
# --- the breakdown. anon vs file is the restore path showing through: a
# locally-restored guest is file-backed page cache until it writes; a streamed
# (userfaultfd) guest is anonymous from the first page.
grep -E '^(anon|file|shmem|slab|active_anon|inactive_anon|active_file|inactive_file|pgmajfault|pgscan|pgsteal|workingset_refault)' \
"$G/memory.stat"
# --- did the squeeze actually cost anybody anything?
# some = at least one task in this cgroup stalled on memory
# full = EVERY non-idle task stalled. On a squeezed VM, full climbing means
# you took memory it wanted back.
cat "$G/memory.pressure"
# high/max counters here are the squeeze's receipt: `high` climbing means the
# cap is being hit, which is the cap working; `oom` climbing means it isn't.
cat "$G/memory.events"
# --- the billing basis. usage_usec is cumulative ACTIVE CPU microseconds for
# this VM's Firecracker process; differentiate it across your own sampling
# window. CPU is metered from measurement. Memory is still billed committed.
# That asymmetry is the subject of the last section of this post.
grep -E '^(usage_usec|user_usec|system_usec)' "$G/cpu.stat"
# --- what the ladder writes, for reference. The cap:
# echo $(( $(cat "$G/memory.current") * 70 / 100 )) | sudo tee "$G/memory.high"
# Release it:
# echo max | sudo tee "$G/memory.high"
# Proactive reclaim -- a PUSH, not a cap. An idle VM never allocates, so it
# never trips a cap; this is the only rung that moves a quiet guest. Needs a
# reclaim destination (swap/zswap) or it is theatre. The write BLOCKS until the
# kernel has reclaimed the amount or given up with partial progress:
# echo $((512*1024*1024)) | sudo tee "$G/memory.reclaim"
Notice what that ladder does not have: any idea which pages to take. It picks a victim VM by coldness — CPU rate and last activity, both proxies — and delegates the page-level decision entirely to host reclaim. That is the gap DAMON fills: the ladder knows which guest looks idle, where DAMON would know which parts of that guest's memory have not been touched in a measured number of aggregation intervals.
The prefetch trace is a bake-time working set
Here is the part I find genuinely interesting, and the reason I went looking at DAMON: we already ship a recorded working set. We just record it once, at the wrong time, for a different purpose.
With streaming enabled, a restored guest's `vm.mem` is not downloaded — it is demand-paged from object storage over HTTP Range GETs in 4 MiB chunks by a userfaultfd handler, with a header recording which chunks are non-zero so an all-zero chunk is never fetched. Alongside that we bake a prefetch trace: during the bake's own warm-up, resume plus readiness probe, the resolver records which chunks it actually had to fetch, and that set is replayed in the background on every later restore so the guest's faults become local cache hits.
Read that again with the measurement question in mind. The set of chunks a representative run had to fetch is a recorded hot set — a working-set measurement at 4 MiB granularity, produced by instrumenting faults rather than sampling accessed bits, and nearly free, because the fetch had to happen anyway. What it is not is current.
- The trace is one-shot and historical: one warm-up of one template at bake time, then frozen for the life of that seed generation. A guest an hour into real traffic has a working set the trace has never seen.
- It is fault-granular, not access-granular. It records first touches, so it cannot distinguish a chunk touched once at boot from one hammered continuously. Our implementation even sorts the chunk indices ascending for replay, so temporal order — the part an access monitor cares most about — is discarded.
- It is 4 MiB coarse, because that is the transfer unit, chosen for object-storage economics rather than measurement resolution.
- DAMON is the continuous runtime version of the same idea: regions instead of chunks, frequency and age instead of fetched-or-not, a self-tuning partition instead of a fixed grid.
Which suggests the obvious and so far unbuilt thing — a design, not a shipped feature: a DAMOS counting scheme over a streamed guest's mapping produces exactly the input a prefetch trace wants, from real traffic rather than a synthetic warm-up, and chunk indices are a trivial transform of region offsets. One extra objection on top of the ones above: a trace recorded from one guest's traffic describes that guest, not the template.
The billing consequence, which is where it gets uncomfortable
Our rate card is $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same across every class. Look at how the two halves are computed and you find the asymmetry this post is really about.
CPU is measured. The agent reads each sandbox's cgroup `cpu.stat` and takes `usage_usec` — cumulative active CPU microseconds actually burned by that VM's Firecracker process — as the metering counter. An idle guest with eight burstable vCPUs accrues almost nothing. You pay for the seconds you touch.
Memory is committed. It bills the template's baked GiB, because that is what Firecracker gave the guest, and an app cannot opt out, because guest sizing is a property of the snapshot. The measurement to fix that already exists: our residency sampler integrates resident gibibyte-seconds per sandbox, which is precisely what a measured-memory bill would be computed from. The flip is not an engineering decision.
Charging by measured working set is a fairness improvement and a forecasting nightmare. Under committed billing a customer can compute their memory bill before deploying: template size times hours. Under measured billing it is a function of their own access pattern, which most teams have never measured. And a sharper problem falls out of the pressure ladder: if the bill is computed from residency, every housekeeping action the platform takes is a pricing input. The controller squeezes a cold VM and its bill drops; the reclaim cycle fires and it drops again. My reclaim tuning becomes a silent discount, and a bug in it becomes a billing incident.
Which is why access monitoring matters more for billing than residency, not less. Residency is a high-water mark the vendor can move; measured access frequency is a property of the workload. That does not solve forecasting — nothing does except a floor, a cap, or a committed baseline with measured burst on top — but the number you are arguing about is the customer's.
Committed sizing is an answer to 'what did we promise'. Capacity is a question about what is in use. Confusing the two is how a host ends up refusing work while most of its RAM does nothing.
So: measure before you pack, measure for a long time before you act, and run the action that does nothing until the numbers stop surprising you. DAMON is the best-shaped instrument Linux has for the question a dense sandbox host asks, because it is the only one whose cost does not scale with the thing you have most of. Prove that the accessed state it samples reflects guest activity and you get the number your density limit is made of.
Frequently asked questions
Can I just use DAMON to measure a microVM's working set from the host?
You can point it at the VMM process, and the shape of the answer is right — a partition of the guest memory mapping with an estimated access frequency and age per region, at a cost bounded by your configured maximum region count rather than by the size of the mapping. But validate one thing before you believe any of it. Guest memory accesses are translated through the second-level page tables KVM maintains, so the hardware accessed bits that record guest activity are set in the secondary MMU rather than necessarily in the VMM's own host page table entries. Whether a given kernel's monitoring path consults the secondary MMU is version-dependent and has seen active work, and the young-bit helpers used by the physical-address operations set have notifier-aware variants that walking a process's page tables directly does not. The empirical test is cheap and conclusive: boot a guest that touches a known fixed amount of memory in a loop, run a counting scheme over the VMM, and check that the regions with a non-zero access count sum to roughly that amount. If they sum to nearly nothing, your monitoring path is blind to guest activity on this kernel — and you found out before it became a capacity policy.
Should I run DAMOS with pageout on guest memory to free up RAM?
Not as a first move, and not without quotas. The structural problem is that you are inserting a third policy into a stack that already has two. The guest kernel runs complete reclaim over the memory it believes it owns, on its own clock, with its own LRU. The host kernel runs reclaim over the pages backing that memory, using access information filtered through the guest. A scheme that pages out regions by measured coldness means three independent algorithms hold an opinion about the same physical page and two of them cannot see the third. The failure mode is double reclaim: your scheme evicts a region the guest considers warm, the guest touches it, the host faults it back in, and the guest's own reclaim drops something else to make room — a feedback loop built from two individually correct algorithms, presenting as unexplained I/O latency inside a guest whose own metrics look fine. If you do it anyway, the order is: a time and byte quota per reset interval so the scheme cannot run away, prioritisation weights so that when the quota binds it drops the least valuable work rather than an arbitrary slice, watermarks so the scheme is inactive both when memory is plentiful and when the host is already in trouble, and refault counters on a dashboard from day one.
How do PSI, MGLRU and DAMON relate? They all seem to be about memory pressure.
They answer three different questions and are not substitutes in any direction. PSI is an alarm: time spent stalled waiting on memory, host-wide and per cgroup, with a some line for at least one task stalled and a full line for every non-idle task stalled. It is the best trip signal in the kernel and our pressure controller classifies host state on it directly, but it says nothing about how much memory anybody needs or which pages are involved — there is a fire, not where the fuel is. MGLRU is a reclaim policy. It replaces the aging half of reclaim with numbered generations and ages by walking page tables in address order rather than walking reverse maps from pages, which fits a host full of processes each mapping gigabytes contiguously far better. It consumes access information to pick better victims; the generation sizes exposed through its debugfs file are a byproduct and a testing interface, not an exportable working set, and the result to expect is better tail latency rather than extra capacity. DAMON is the only one of the three that is a measurement service: a quota-bounded, continuously updated estimate of access frequency and age per region, designed to be read and acted on from outside reclaim. Alert with PSI, reclaim better with MGLRU, pack with DAMON.
If you can measure a sandbox's working set, why not bill for that instead of committed memory?
Because it is a fairness improvement and a forecasting nightmare at the same time. Our rate card is 0.054 dollars per vCPU-hour and 0.0162 dollars per GiB-hour, and the CPU half is already measured: the agent reads each sandbox's cgroup cpu.stat and meters usage_usec, cumulative active CPU microseconds actually burned, so an idle guest with eight burstable vCPUs accrues almost nothing. The memory half bills the template's baked GiB, which an app cannot opt out of, because Firecracker cannot change guest RAM at snapshot restore and the only sizes available are the template sizes — so an idle Node app on our 4 GiB base template pays for 4 GiB while touching much less. The measurement to fix that exists; our residency sampler already integrates resident gibibyte-seconds per sandbox. What stops me is twofold. A customer can compute a committed memory bill before deploying and cannot compute a measured one at all, because it depends on an access pattern they have never measured. And worse: if the bill is computed from residency, every housekeeping action we take becomes a pricing input — the pressure controller squeezes a cold VM and its bill drops, the reclaim cycle fires and it drops again, so my reclaim tuning becomes a discount and a bug in it becomes a billing incident. Measured access frequency is better than residency precisely because it is a property of the workload rather than a high-water mark the vendor can move.
Keep reading
- Memory oversubscription in microVM fleets, explained — the bet this measurement is meant to bound — lazy paging, copy-on-write, the balloon and swap
- PSI: the right instrument for memory oversubscription — the alarm our pressure ladder classifies on, and how to alert on it without lying to yourself
- MGLRU and microVM density: reclaim when every page is someone's guest — the reclaim policy that consumes access information — a better tail, not more capacity
- Memory prefetch: the working set is the real unit of a fast restore — the bake-time recorded hot set that DAMON would be the continuous version of
- userfaultfd: lazy memory for instant VM restore — why a streamed guest's memory is anonymous from the first page, and what that does to reclaim
Related posts
- Tidal: Turning the Host Into a Cache, and Paying for the Seconds You Touch
Committed capacity is the address space; the working set is what deserves host RAM. We rebuilt our compute layer around cache semantics — burst freely when the host is quiet, fair-share when it isn't, and a bill that meters the seconds you actually touch.
- KSM Memory Deduplication for MicroVMs — And Why It's a Trap
Eighty microVMs running the same kernel and the same libc, and the host is storing eighty copies. KSM will merge them — for a permanent CPU tax, a merge that evaporates on first write, and a timing oracle across tenant boundaries. Here's how ksmd actually works, and why sharing at the snapshot layer beats scanning for duplicates you should never have made.
- Shared Pages & Copy-on-Write: Packing MicroVMs Densely
The cheapest page of memory is the one you never had to copy. When a hundred microVMs restore from the same baked snapshot, their RAM starts as copy-on-write over a single shared backing — identical pages physically shared until a guest writes. Here's the mechanic, where it stops helping, and why KSM is the messier cousin.
- Guest page cache: why it bloats microVM snapshots
A microVM snapshot captures the guest's RAM verbatim — including megabytes of file data Linux cached from a disk that's sitting right there. Drop the cache before you bake, and the memory image shrinks to the live working set.
- The Firecracker virtio-balloon Device, Explained
The balloon is the guest politely handing RAM back, on the honor system. Here's how virtio-balloon lets a host reclaim idle guests' memory to pack more microVMs per box — and the honest caveats of a cooperative mechanism.
More in Internals · See PandaStack benchmarks
49ms p50 cold start. Fork, snapshot, and scale to zero.