memory.high vs memory.max: Throttle or Kill
Every explanation of cgroup v2 memory limits reaches for the same two words. `memory.max` is the hard limit; `memory.high` is the soft limit. I have written that sentence myself, on this blog, and it is close enough to be useful and wrong in exactly the place that matters. "Soft limit" makes `memory.high` sound advisory — a threshold the kernel honours when convenient and ignores when busy. It is not advisory. It is enforced hard, by a mechanism that is different in kind from the one behind `memory.max`, and conflating the two is how people end up setting both badly.
The accurate framing is this. `memory.max` is a failure boundary: cross it and the kernel reclaims aggressively, and if reclaim cannot free enough, the OOM killer fires inside that cgroup and something dies. `memory.high` is a throttle: cross it and the task doing the allocating performs its own reclaim and is then handed a deliberate sleep that scales with how far over it is. Nothing dies. The workload stays alive and gets slow, on purpose, as a design decision the kernel made for you.
One knob converts an overcommit mistake into a dead process. The other converts the same mistake into a latency problem. For a service you own, you may not care which. For a multi-tenant compute fleet that distinction is the whole design, because a dead tenant workload is a support ticket with a confused human attached and a slow one is a line on a graph you can look at on Monday.
I have already written the tour of cgroup v2 for sandboxing — the unified hierarchy, every controller, where cgroups stop being a security boundary. This is not that post. This is the narrow one: two files, the handful you need to tell which of them fired, and the decision between them for someone choosing a memory policy for a fleet that runs other people's code. It ends where people do not expect, which is that for one important class of workload — a hypervisor process — the right answer is neither.
The files, precisely
There are more than two. The memory controller exposes a protection pair, a throttle, a wall, a swap bound, and a set of files whose only job is to say which of those is happening. Most incidents here are someone having set one of the first four and none of the last.
`memory.min` and `memory.low`: protection, not limits
These two point the opposite way from the rest: they do not cap a cgroup, they defend it. Memory within `memory.min` is never reclaimed, full stop — if every candidate is protected by a `min`, the kernel invokes the OOM killer rather than taking pages. Strong promise, correspondingly strong failure mode: a `min` set carelessly across enough cgroups turns a reclaimable host into one that OOMs with gigabytes of page cache technically present.
`memory.low` is the best-effort version: reclaim avoids a cgroup under its `low` unless there is nothing unprotected left, and the protection decays as the overage grows. Neither file is a limit, and people who set `memory.low` expecting a cap are surprised in the least useful direction.
`memory.high`: the throttle
This is the interesting file. When a charge would push the cgroup above `memory.high`, two things happen to the task that made the allocation. First it performs direct reclaim itself — not kswapd in the background, the allocating task, synchronously, so the latency is charged to the thing responsible rather than smeared across the host. Second, on its way back to userspace, the kernel puts it to sleep for a penalty that grows with the overage, steeply and nonlinearly, up to a cap of a couple of seconds per return. Go further over and you do not get a bigger single sleep, you get that sleep more often.
The crucial property: breaching `memory.high` never invokes the OOM killer. Which means `memory.high` is not a guarantee of anything. A cgroup can sit above its `high` indefinitely, for as long as it likes, if reclaim cannot make progress. The kernel documentation says this plainly and it is still the thing that surprises people most: under extreme conditions the limit may simply be breached and stay breached. What you have bought is not a ceiling. You have bought a feedback signal with teeth.
`memory.max`: the wall
`memory.max` is the one that ends things. The kernel first tries hard: reclaim runs, repeatedly, and will drop page cache, swap anon pages if swap is available, and generally do unreasonable work to keep the charge under the line. Only when reclaim cannot deliver does the in-cgroup OOM killer run. Note the scoping, because it is the whole value of the file: the victim is chosen from inside that cgroup, not from across the host. One tenant's runaway kills one tenant's process.
Some kernel-side allocations that cannot be charged fail with `ENOMEM` rather than triggering a kill, which is why a cgroup at its limit sometimes produces failures in odd places instead of the clean `Killed` you expected. Unlike `high`, though, usage genuinely cannot exceed `max` for long — the kernel will kill until it does not.
There is a joke in here that is also the honest description: `memory.max` is the kernel agreeing to kill your process so you do not have to think about capacity planning. A good deal — worth knowing you took it.
`memory.swap.max`: the file that changes what the other two mean
Swap is not a side issue here; it decides whether either knob can work. Reclaim can take two kinds of page: clean file-backed pages, which it drops for free, and anonymous pages, which it must write somewhere first. With `memory.swap.max` at 0 there is no somewhere. An anon-heavy workload under `memory.high` therefore has almost nothing reclaimable, so the throttle engages and never relieves; the same workload under `memory.max` reaches the kill sooner than its numbers suggested.
We measured exactly this on our own hosts, and it is the most decisive comparison in the post. A guest capped at `memory.high=1G`, asked to burn a 2.5 GiB working set, on a host with no swap: pressure pegged and the job did not finish — not slow, not killed, just permanently throttled. Same cap, same burn, on a host with 16 GiB of swap and zswap in front of it: 45 seconds, 1.79 GiB compressed into a 32 MB pool, guest alive and instantly responsive afterwards. The knob did not change. The thing behind it did.
The files that tell you which knob fired
`memory.current` is residency including page cache and kernel allocations — slab, socket buffers, page tables — so it is not "your heap" and never was. `memory.peak` on recent kernels is the high-water mark: the number you wanted an hour ago, when the page fired and `memory.current` had already relaxed. `memory.stat` is the breakdown — `anon` versus `file`, the inactive and active lists, `slab`, `sock`, `shmem`, plus `pgscan`, `pgsteal` and `workingset_refault`, which say whether reclaim is working or thrashing.
`memory.events` is the one that settles arguments. Five counters: `low` and `high` for the protection and throttle paths, `max` for times usage would have exceeded the wall, `oom` for times an allocation could not proceed, `oom_kill` for processes actually killed. `high` climbing with `oom_kill` at zero is a latency bug and nothing died; `max` climbing with `oom_kill` at zero means reclaim is still winning at the wall and you are one bad burst away. It is hierarchical and includes descendants; `memory.events.local` is the cgroup alone.
And `memory.pressure` is PSI, per cgroup: `some` is the share of time at least one task was stalled on memory, `full` the share where every task was. `full` above zero means the cgroup is not absorbing pressure, it is drowning in it; `some` rising with `full` flat is a cgroup doing its job. Between those two files you can always answer "throttle or kill", which means you never have to write "we think it OOMed" in an incident review again.
#!/usr/bin/env bash
# Which knob actually fired? This is the only honest way to find out, and it
# takes four files. Guessing from a dashboard is how "we think it OOMed" ends
# up in an incident review next to "we think it was the disk".
set -euo pipefail
CG=${1:-/sys/fs/cgroup/jobs/run-1}
echo "--- what is even set ---"
# "max" means unset. A cgroup with memory.high=max and memory.max=max has no
# memory policy at all, however confident your config management looks.
paste <(echo high) <(cat "$CG/memory.high")
paste <(echo max) <(cat "$CG/memory.max")
paste <(echo swap) <(cat "$CG/memory.swap.max")
echo "--- the counters that distinguish a throttle from a kill ---"
cat "$CG/memory.events"
# low 0 protection was breached and we reclaimed anyway
# high 42 42 times a task was throttled for being over memory.high
# max 3 3 times usage would have exceeded memory.max
# oom 1 1 time the limit was hit and the allocation could not proceed
# oom_kill 1 1 process was actually killed
#
# Read it as a decision tree:
# high > 0, max = 0, oom_kill = 0 -> you have a LATENCY bug, not a memory
# bug. Nothing died. Something is slow
# and it is the kernel doing it to you.
# max > 0, oom_kill = 0 -> reclaim is still winning at the wall.
# You are one unlucky burst from a kill.
# oom_kill > 0 -> contained death. Check oom.group to
# know whether it took one process or
# the whole tree.
# memory.events is HIERARCHICAL -- it includes descendants. If the numbers do
# not match the processes you are looking at, read memory.events.local.
cat "$CG/memory.events.local" 2>/dev/null || true
echo "--- how much time was spent stalled (PSI) ---"
cat "$CG/memory.pressure"
# some avg10=18.44 -> at least one task stalled on memory 18% of the last 10s
# full avg10=4.10 -> ALL tasks stalled 4% of it: nothing ran but reclaim.
# "full" above zero on a cgroup you care about is the signal. "some" rising
# with "full" at zero is a cgroup absorbing pressure, which is the job.
echo "--- is it reclaimable, or is it anon with nowhere to go ---"
grep -E '^(anon|file|slab|sock|shmem|swap|inactive_file|active_anon) ' "$CG/memory.stat"
grep -E 'pgscan|pgsteal|workingset_refault' "$CG/memory.stat"
# A high file:anon ratio means memory.high can do its job cheaply -- clean page
# cache is dropped, no I/O, no stall. A high anon ratio with swap at 0 means
# reclaim has almost nothing to take, so the throttle degenerates into a stall.
# pgscan climbing much faster than pgsteal is "everything looks hot": the
# kernel is examining many pages to free one. That is thrash, not reclaim.
echo "--- and the host's own opinion ---"
journalctl -k --no-pager | grep -i 'memory cgroup out of memory' | tail -5
cat "$CG/memory.current" "$CG/memory.peak" 2>/dev/null
# memory.peak (recent kernels) is the high-water mark, which is the number you
# wanted three hours ago when the pager went off and memory.current had
# already relaxed back to something reasonable.
The five files, side by side
| File | What the kernel does on breach | Who feels it | Where you see it | Fit for a VMM process | Fit for an in-guest task |
|---|---|---|---|---|---|
| `memory.min` | Nothing below it is ever reclaimed; if every candidate is protected, the OOM killer runs instead | Everyone else on the host | Host PSI and global OOM logs; no counter of its own | Rarely — only for a guaranteed class, and only if host-wide OOM is acceptable | No; a disposable sandbox is not what you protect |
| `memory.low` | Reclaim avoids the cgroup unless nothing unprotected is left; decays with overage | The cgroups that are not protected | `memory.events` `low` counts reclaim that happened anyway | Reasonable as a soft floor under a latency-sensitive VM | Rarely useful |
| `memory.high` | The allocating task reclaims, then sleeps for a penalty scaled to the overage. Nothing dies; usage can stay above it indefinitely | The task that allocated, and nobody else | `memory.events` `high`, `memory.pressure`, `pgscan`/`pgsteal` | Only as a reversible squeeze on an idle VM with swap behind it | Yes — this is the knob you want |
| `memory.max` | Reclaim hard; if that fails, the in-cgroup OOM killer picks a victim inside it | One process inside the cgroup, fatally | `memory.events` `max`/`oom`/`oom_kill`, kernel log | No — turns a well-behaved guest into a VM that vanishes mid-syscall | Yes, as the blast-radius fence well above `high` |
| `memory.swap.max` | Swap charges past it fail, so anon pages become unreclaimable | The workload's own reclaimability | `memory.swap.events`, `swap` in `memory.stat` | Set deliberately: a squeeze with no swap behind it is a stall | `0` is honest with no guest swap — but then `high` cannot help anon-heavy code |
The two ways people set these badly
The first failure is setting `high` just under `max`. It looks prudent — a little grace before the wall — and it buys nothing. The throttle's whole value is the band: the window in which a workload is slowed enough to notice, with room left to finish, flush its state and exit on its own terms. A 32 MiB band gives a job one allocation's worth of warning. You pay the latency penalty and then die anyway: slow and dead, the worst cell in the matrix.
The second failure is the mirror image, more common and much harder to see. Set `high` far below `max` with no monitoring and you ship a workload that lives permanently in reclaim. It never dies, so nothing alerts. It is just mysteriously slow, forever, in a way that looks exactly like slow storage — direct reclaim and page-cache refaults produce real disk reads, so your disk metrics are genuinely elevated and genuinely not the cause. I have watched a week-long I/O investigation whose entire root cause was a `MemoryHigh=` someone set during an incident eighteen months earlier and never removed.
The pattern that works is boring and has three parts. `memory.high` is the working-set target you genuinely believe, from measurement rather than from a round number you like. `memory.max` is the blast-radius fence well above it, whose job is not accuracy but bounding the damage when your belief turns out wrong. And `memory.events` plus PSI tell you which of the two you got wrong — because you will get one wrong, and the only question is whether you hear it from a graph or from a customer.
Reclaim mechanics that decide whether the throttle works at all
Whether `memory.high` is a gentle throttle or a slow-motion hang is not a property of the knob. It is a property of what is inside the cgroup: two workloads with identical `memory.current` behave completely differently under the same cap, and the difference is the `anon` to `file` ratio in `memory.stat`.
A cgroup's page cache is charged to whichever cgroup first faulted each page in, and that charge outlives the process. Clean page cache is also the cheapest thing in the world to reclaim: the kernel drops the page and moves on, no I/O, no stall. So a file-heavy job can sit happily at its `memory.high` forever, shedding and re-reading cache, with PSI barely registering. The same cap on an anon-heavy job of the same size is a different universe: every page has to be written to swap first, and with no swap it cannot be taken at all.
This is also why "how much memory is it using" is a worse question than it looks: a process that reads a large file once and exits leaves its cgroup looking fat. The number that behaves like a working set is `memory.current` minus `inactive_file` — residency with the cold, trivially-droppable cache subtracted, the same arithmetic Kubernetes uses.
Dirty pages and cgroup-aware writeback
Here is a story you will still read in blog posts, including, I am sorry to say, in older corners of this one. A process does buffered writes, the dirty pages pile up in the page cache, the eventual device I/O is performed by a kernel writeback thread that belongs to no cgroup, and so the write escapes your limits entirely while kswapd and the host pay for it. Set a memory limit all you like; dirty pages go around it.
That was true, and it is pre-cgroup-writeback behaviour. Modern kernels track the owning cgroup of each dirty page through the memory controller and attribute the eventual writeback to that owner, so the I/O is charged to the cgroup that dirtied the page rather than to nobody. Dirty throttling is cgroup-aware too: `balance_dirty_pages` works against a per-cgroup writeback domain, so a cgroup gets a proportional share of the dirty budget instead of being able to dirty the whole host's allowance. Repeating the old story in 2026 is how you end up designing around a problem the kernel fixed.
The honest caveats. Cgroup writeback needs filesystem support — ext4 and btrfs have it; check anything more exotic. It needs the memory controller on the same v2 hierarchy as the I/O controller, which is a thing you can accidentally not have. And pages whose owner cannot be established are attributed to the root cgroup, which from your point of view is the same as unattributed, so attribution blurs through device-mapper stacks, loop devices and network filesystems. It works; it is not magic; verify it on your own storage stack rather than believing the old story or this paragraph.
The OOM path, and what `memory.max` actually buys
When reclaim loses, the in-cgroup OOM killer runs. It picks a victim from the processes inside that cgroup by badness — roughly, footprint adjusted by `oom_score_adj` — and kills it. By default it kills one process, which on a process tree is frequently the wrong one: the worker dies, the parent waits forever on a pipe that will never produce another byte, and the job reads as "running" until something times out hours later.
`memory.oom.group` fixes that: set it to 1 and the entire cgroup is killed as a unit. For any job you treat as atomic — a build, a test run, a model-generated script and its children — this is what you want, because a cleanly dead cgroup is a non-event to debug and a half-dead process tree is a Tuesday. There is no systemd property for it, so you write the file.
And then the operationally important part, which is the real reason to set `memory.max` at all. An OOM kill inside a tenant's cgroup is a contained failure: one workload dies, every neighbour continues, and your incident is one customer's job failing with a diagnosable cause. A host-level OOM is an incident. The global OOM killer chooses by badness across the whole machine, so it is biased toward whatever is biggest, which on a compute host is approximately never the thing that caused the problem. It kills a neighbour. Possibly several. Then you find out your monitoring agent was the biggest resident process on the box.
Setting `memory.max` does not prevent the failure. It decides whose failure it is. That is a smaller promise than people think and a much more useful one than they expect.
Why a hypervisor process is the wrong workload for either knob
Everything above assumes the thing in the cgroup is a normal process: it allocates, it can be slowed down, it can be killed and restarted. A VMM process breaks all three, and this is where a memory policy that is correct everywhere else becomes actively dangerous.
Start with the ceiling. A Firecracker guest's RAM is decided when the VM is configured, and on PandaStack that means baked into the template snapshot — `base` is 4 GiB, `code-interpreter` and `agent` are 2 GiB, `postgres-16` is 1 GiB — because Firecracker cannot change vCPU or RAM at snapshot restore. The guest already has a hard ceiling, enforced by the only entity that can enforce it coherently: the guest's own kernel, with its own OOM killer and its own dmesg. That machinery works. A program inside the guest that allocates too much gets killed there, inside a VM that was going to be deleted anyway, and the guest tells you about it.
Now put a host-side `memory.max` under the baked size. The guest touches a page it legitimately owns, the host charge crosses the line, reclaim runs, and when reclaim cannot deliver — which for guest anon memory with no swap is immediately — the host OOM killer kills the VMM. You do not get graceful degradation, and you do not get the guest's OOM killer making a sensible local decision. You get a machine that vanishes mid-instruction, nothing in the guest log, because the guest never ran again. From the tenant's side an entire computer disappeared while it was working correctly.
`memory.high` on a VMM is stranger, and in some ways worse, because it does not announce itself. The throttle is applied to the task that made the allocation — and inside a VMM, that task is a vCPU thread. Injecting a sleep into a vCPU thread's path back to userspace does not slow down "the workload". It stops that virtual CPU. From inside the guest, time jumps: soft lockup warnings, an RCU stall, a clocksource the kernel declares unstable, TCP retransmits, a disk the guest decides has failed. Every symptom is real and every symptom is a lie, because the guest cannot know it was frozen and attributes the gap to whatever it was touching when the music stopped.
So what do we actually write?
Here is the honest architecture, including the part more subtle than I would like. Each Firecracker process lives in a per-VM cgroup, `vm-<id>`, under the agent's own slice. The agent writes two files there: `cpu.weight`, the vCPU count times 100 clamped to the kernel's 1–10000 range, which gives proportional fairness under contention and lets a single VM use every idle core when nobody is competing; and `cpu.max`, a hard ceiling in physical cores resolved from the workspace tier, re-asserted every 15-second tick so a restart or tier change converges instead of leaving a VM uncapped. Those are the knobs where a VM behaves like a normal workload: CPU throttling is survivable, and a guest that runs slowly is still a guest that runs.
It reads `cpu.stat`'s `usage_usec` for active CPU time, `memory.current` and `inactive_file` from `memory.stat` for working-set residency, and `io.stat`. Those feed per-VM metrics, the billing basis, and admission.
It never writes `memory.max`. Not on any VM, not at any size, deliberately. The file is there — the agent delegates `+memory` into the subtree, so every `vm-<id>` cgroup has the complete memory interface from birth — and the policy declines to use that one file, for exactly the reason in the previous section. Memory safety comes from admission instead. Working-set admission is our fleet default: the scheduler admits against a measured `memory_mb_admittable` rather than the sum of baked sizes, and charges a reserve of `max(512, memory_mb × 0.25)` per non-database create so an in-flight burst cannot all be admitted against the same free memory. Managed databases are exempt from the overcommit entirely — charged committed, always, because a database is the one workload where "it got slow" and "it is broken" are the same ticket.
And `memory.high`? We do write it — this is the part the simple version of the story gets wrong, including an earlier draft of this post. But never as a limit. A 10-second controller watches host `MemAvailable` and `/proc/pressure/memory`, and above its trip line it squeezes the coldest VMs: `memory.high` at 70 percent of that VM's current residency, floor 256 MiB, at most two per tick, lifted again one per tick on the way back down. The gates are the interesting part, because each exists because of the previous section. Never a managed database. Never a VM that has used CPU recently or been touched in the last minute. Never a hugepage-backed VM — hugetlb pages are not accounted the way you would expect, so such a VM's `memory.current` reads a few megabytes and the knob would act on a number that means nothing. And the rung is disarmed unless the host has swap with zswap in front of it: a cap with nowhere to put the pages is not a throttle, it is a stall.
One more wrinkle, which is the single most useful thing here for anyone building something similar: `memory.high` cannot reclaim from an idle workload. Reclaim runs on the victim's own allocation paths, and a guest doing nothing does not allocate — so capping an idle VM achieves precisely nothing. We measured a capped idle VM holding 3.26 GB indefinitely, cap in place, kernel entirely unbothered. The fix is `memory.reclaim`, the proactive-reclaim file added in 5.19: write a byte count and the kernel does the work itself rather than waiting for a task that is never coming back. Same VM, same cap: 3258 MB down to 897 MB in 12.9 seconds, with zswap absorbing the difference.
#!/usr/bin/env bash
# The real per-VM cgroup on a PandaStack host. One Firecracker process per
# vm-<id>, siblings of an agent/ leaf that holds the agent's own pids (the
# no-internal-processes rule means the service root cannot hold both).
CG=/sys/fs/cgroup/system.slice/pandastack-agent.service/vm-$SBX
# ---- WRITTEN, every 15s reconcile tick, idempotently ----
cat "$CG/cpu.weight"
# vcpus * 100, clamped to the kernel's 1..10000 -- so 800 for an 8-vCPU
# guest. A proportional share: under contention it decides who wins, and on
# an idle host it lets a single VM use every core. That is the burst.
cat "$CG/cpu.max"
# "<cores * 100000> 100000" -- a hard ceiling in physical cores, resolved
# from the workspace tier. Re-asserted every pass so an agent restart or a
# tier change converges within one tick instead of leaving a VM uncapped.
# ---- READ, every tick: metrics, billing basis, admission ----
awk '/usage_usec/ {print $2}' "$CG/cpu.stat"
awk '/inactive_file/ {print $2}' "$CG/memory.stat"
cat "$CG/memory.current" # residency, INCLUDING host page cache for the
# VM's disk image -- which is why the working-set
# number is memory.current minus inactive_file,
# the same arithmetic Kubernetes uses.
cat "$CG/io.stat"
# ---- NEVER WRITTEN ----
cat "$CG/memory.max" # "max". Always. On every VM. Deliberately.
# The file exists -- the agent delegates +memory so these cgroups have the
# whole memory interface from birth. The policy simply declines to use this
# one, because the guest's real ceiling is the RAM baked into its snapshot
# and Firecracker cannot renegotiate that at restore. A memory.max under
# the baked size does not degrade the guest. It kills the VMM while the
# guest is mid-syscall, with nothing in the guest log, because the guest
# never got to run.
# ---- WRITTEN ONLY AS A TEMPORARY, REVERSIBLE SQUEEZE ----
cat "$CG/memory.high" # "max" normally; 70% of residency while squeezed
# A 10s controller watches MemAvailable and /proc/pressure/memory. Above the
# trip line it caps the coldest VMs -- idle, cpu-quiet, not a managed
# database, not hugepage-backed, floor 256 MiB, at most two per tick -- and
# lifts the cap one per tick on the way back down. It is a reclaim lever,
# not a limit, and it is only armed when the host has swap plus zswap behind
# it, because a cap with nowhere to put the pages is not a throttle. It is
# a stall.
echo $((3 * 1024 * 1024 * 1024)) > "$CG/memory.reclaim"
# And for an IDLE VM, the cap does nothing at all -- reclaim runs on the
# victim's own allocation paths, and an idle guest does not allocate. That
# is what memory.reclaim (5.19+) is for: ask the kernel directly, instead of
# waiting for a task that is never going to come back.
The summary: for a guest, admission control replaces the limit file, because the limit file's failure modes are unacceptable for a VM. `memory.high` survives as a reversible reclaim lever on VMs that are provably idle, with swap behind it. `memory.max` does not survive at all. If that sounds like a lot of machinery to avoid writing one number, consider what the number is: a standing promise to kill a customer's machine while it is behaving correctly.
Inside the guest, these knobs are exactly right
It would be easy to read that as "cgroup memory limits are a trap". They are a trap for one workload. One layer down, inside the guest, they are precisely the right tool, and we use them without hesitation.
Think about what has changed. The thing in the cgroup is now an ordinary process — a Python script a model wrote ninety seconds ago, a test run, a build. It allocates, so the throttle has something to act on. It can be slowed down meaningfully, because "took 40 seconds instead of 12" is an acceptable outcome. And when `memory.max` finally fires, the guest kernel's OOM killer takes one Python process inside a virtual machine with a TTL measured in minutes. That is not an incident; that is the system working. The job gets a non-zero exit code and a line in dmesg, which is a thousand times nicer to debug than a guest-wide OOM that takes the `pandastack-init` agent with it and leaves you a sandbox you cannot even ask what happened.
So inside a sandbox: set both, set them far apart, set `memory.oom.group=1`, and set `pids.max` while you are there, because a fork bomb needs PID space rather than memory and no memory limit will save you from one. Throttle-then-contained-kill is exactly the sequence you want for untrusted code — a band in which a well-written job notices and bails, then a fence that bounds the damage when it does not.
#!/usr/bin/env bash
# A memory policy for code you did not write, set INSIDE a sandbox guest.
# This is the place these knobs are exactly right: a throttle that buys the
# job time to finish, then a contained kill that takes the job and nothing
# else. The guest is deleted in a few minutes either way.
set -euo pipefail
# Controllers are handed DOWN. Enable them in each parent, for its children.
echo "+memory +pids" > /sys/fs/cgroup/cgroup.subtree_control
mkdir -p /sys/fs/cgroup/jobs
echo "+memory +pids" > /sys/fs/cgroup/jobs/cgroup.subtree_control
CG=/sys/fs/cgroup/jobs/run-1
mkdir -p "$CG"
# The working-set target you actually believe. Past this, the allocating task
# does its own reclaim and gets a sleep that scales with the overage.
echo 1536M > "$CG/memory.high"
# The blast-radius fence, well above it. Past this the cgroup OOM killer
# fires. 25% of headroom between the two is a degradation band, not a rounding
# error -- set them 32M apart and you pay the throttle and then die anyway.
echo 2G > "$CG/memory.max"
# No swap in this guest, so say so out loud. It also tells you something:
# with swap at 0, anon pages have nowhere to go, so memory.high can only
# reclaim page cache. An anon-heavy job will sit above high, throttled, until
# it reaches max. If you want the throttle to mean anything for such a job,
# give the guest zram.
echo 0 > "$CG/memory.swap.max"
# Kill the whole cgroup as a unit. Without this the in-cgroup OOM killer picks
# one victim by badness, which on a Python job is usually a worker -- leaving a
# parent that waits forever on a pipe that will never close. A clean group kill
# is a non-event to debug; a half-dead process tree is a Tuesday.
echo 1 > "$CG/memory.oom.group"
# The only real defence against a fork bomb. Costs one integer.
echo 256 > "$CG/pids.max"
echo $$ > "$CG/cgroup.procs"
exec python3 /work/main.py
# The systemd equivalent, which is what you should ship if systemd owns the
# tree (it does, on any modern distro -- hand-written cgroupfs gets reset on
# the next daemon-reload). Note there is no MemoryOOMGroup= property; if you
# want oom.group you still write that one file.
# systemd-run --scope -q -p MemoryHigh=1536M -p MemoryMax=2G \
# -p MemorySwapMax=0 -p TasksMax=256 -- python3 /work/main.py
One honest caveat, straight from the swap section: our guests have no swap, so an anon-heavy job under `memory.high` inside a guest hits the same wall a VM does — reclaim has only page cache to take. For a build or a test run, which touch a lot of files, the throttle behaves. For a script that allocates a 3 GiB array, `high` throttles it into mud rather than helping, and `max` is what resolves the situation. If you want a real degradation band for anon-heavy code, give the guest zram: compressed swap in RAM gives reclaim somewhere to put pages and turns the stall back into a throttle. Same lever we pull on the host, one layer down.
And if you want to see the boundary rather than read about it, the script below walks it: it grows 32 MiB at a time, touches every page so the allocation is real, and prints its own cgroup's counters beside each step's wall time. The step times sit flat, then jump, then climb, then the process stops printing. The band between the jump and the silence is what this whole post is about.
#!/usr/bin/env python3
"""throttle-then-kill.py -- walk the boundary between the two knobs.
Run it inside a cgroup that has both set, e.g. from the bash above, or:
systemd-run --scope -q -p MemoryHigh=512M -p MemoryMax=768M \
-p MemorySwapMax=0 -- python3 throttle-then-kill.py
It grows in 32 MiB steps, TOUCHING every page (allocating memory is not the
same as having it -- an untouched page is not a page anyone has to back), and
prints the wall time of each step next to its own cgroup's counters.
What you should see, and the whole point of the post:
step 13 416 MiB 12 ms high=0 max=0 <- free
step 16 512 MiB 14 ms high=0 max=0 <- at memory.high
step 17 544 MiB 290 ms high=1 max=0 <- the throttle engages
step 19 608 MiB 980 ms high=3 max=0 <- superlinear in overage
step 22 704 MiB 1890 ms high=6 max=0 <- ~2s/return-to-userspace
Killed <- memory.max, finally
The step times are the kernel deliberately sleeping the allocating task on its
way back to userspace, on top of the direct reclaim it just performed itself.
Nothing is broken. This is the feature. Note how many steps happen after the
throttle starts: that band is the time a well-written job has to notice, flush
and exit on its own terms, which is the entire argument for setting `high`
below `max` rather than setting one number and hoping.
"""
import os
import time
CHUNK = 32 << 20 # 32 MiB
PAGE = 4096
LIMIT_STEPS = 64
def own_cgroup() -> str:
"""cgroup v2 writes one line: '0::/the/path'."""
with open("/proc/self/cgroup") as fh:
for line in fh:
parts = line.strip().split(":", 2)
if parts[0] == "0":
return "/sys/fs/cgroup" + parts[2]
raise RuntimeError("no cgroup v2 membership -- v1 or hybrid hierarchy?")
def events(cg: str) -> dict:
out = {}
try:
with open(os.path.join(cg, "memory.events")) as fh:
for line in fh:
k, _, v = line.partition(" ")
out[k] = int(v)
except FileNotFoundError:
pass # memory controller not delegated to this cgroup
return out
def current_mib(cg: str) -> int:
try:
with open(os.path.join(cg, "memory.current")) as fh:
return int(fh.read().strip()) >> 20
except FileNotFoundError:
return -1
def main() -> None:
cg = own_cgroup()
print(f"cgroup: {cg}")
for name in ("memory.high", "memory.max", "memory.swap.max"):
try:
with open(os.path.join(cg, name)) as fh:
print(f" {name:16} {fh.read().strip()}")
except FileNotFoundError:
print(f" {name:16} (absent -- controller not enabled here)")
held = []
base = events(cg)
for step in range(1, LIMIT_STEPS + 1):
t0 = time.monotonic()
buf = bytearray(CHUNK)
for off in range(0, CHUNK, PAGE):
buf[off] = 1 # fault it in for real
held.append(buf)
dt_ms = (time.monotonic() - t0) * 1000
ev = events(cg)
print(f"step {step:3d} {current_mib(cg):5d} MiB {dt_ms:7.0f} ms "
f"high={ev.get('high', 0) - base.get('high', 0)} "
f"max={ev.get('max', 0) - base.get('max', 0)}",
flush=True) # flush, or the kill eats your log
# If you get here, memory.max was never reached: either the limits are
# wider than LIMIT_STEPS * CHUNK, or reclaim found enough page cache to
# drop that you never pressured anon at all. The second case is the one
# people misread as "the limit does not work".
print("finished without a kill -- widen the run or narrow the limits")
if __name__ == "__main__":
main()
What to set, in order
For a multi-tenant fleet, the order that has worked for me: collect `memory.events` and `memory.pressure` before setting anything, because every later decision depends on seeing which knob fired; set `memory.max` generously as a blast-radius fence, with `memory.oom.group` so the kill is clean; measure `memory.peak` across a representative week; only then set `memory.high`, leaving real headroom to `max`. Alert on the `high` counter and PSI `full`, not on `memory.current`, which spends its life near the limit doing nothing wrong.
And for any workload where the thing in the cgroup cannot be slowed down or cannot be killed — a hypervisor, a storage daemon holding a lease, anything whose death cascades — do not reach for a limit file at all. Its two outcomes are slow and dead, and if neither is acceptable, no value of the limit helps. That is an admission control problem: decide what you let onto the host before it arrives, because once it is there your only remaining tools are the two in the title.
The mental model, one more time, because it is the thing worth keeping. `memory.max` is a failure boundary and `memory.high` is a throttle. Not hard and soft — dead and slow. Pick the failure mode you can operate, build the observability that tells you which one you got, and be suspicious of any workload where the honest answer is "neither".
Frequently asked questions
What is the difference between memory.high and memory.max in cgroup v2?
They enforce by completely different mechanisms, which is why calling them hard and soft limits misleads people. memory.max is a failure boundary: cross it, the kernel reclaims aggressively, and if reclaim cannot free enough the in-cgroup OOM killer kills a process inside that cgroup. memory.high is a throttle: cross it, and the task that made the allocation performs direct reclaim itself and is then given a deliberate sleep scaled to the overage. Nothing is killed. Breaching memory.high never invokes the OOM killer, which also means it guarantees nothing — a cgroup can sit above its high indefinitely if reclaim cannot make progress. In short, memory.max turns an overcommit mistake into a dead process and memory.high turns the same mistake into a latency problem. Set high as the working-set target you believe, max as a blast-radius fence well above it, and monitor memory.events to learn which one you got wrong.
Does memory.high guarantee a cgroup stays under the limit?
No, and this is the most common misunderstanding about the file. memory.high is a rate limiter on allocation, not a bound on residency. When a charge would push the cgroup over, the allocating task does reclaim and then sleeps for a penalty proportional to the overage — but if reclaim cannot free pages, the cgroup simply stays above the line. The kernel documentation says so explicitly: going over high never invokes the OOM killer, and under extreme conditions the limit may be breached. The usual cause is anonymous memory with no swap. Reclaim drops clean file-backed pages for free, but anonymous pages must be written somewhere first, and with memory.swap.max at 0 there is nowhere. The result is a workload permanently throttled above its limit, alive and useless — a state nothing will alert on unless you watch memory.events and memory.pressure, because from inside the cgroup it is indistinguishable from slow storage.
How do I tell whether memory.high throttling or a memory.max OOM kill hit my workload?
Read memory.events, which has five counters that answer it directly. The high counter increments each time a task was throttled for being over memory.high. The max counter increments each time usage would have exceeded memory.max. The oom counter records times the limit was hit and an allocation could not proceed, and oom_kill counts processes actually killed. Read it as a decision tree: high above zero with max and oom_kill at zero means a latency problem and nothing died; max above zero with oom_kill at zero means reclaim is still winning at the wall and you are one burst from a kill; oom_kill above zero means contained death. memory.events is hierarchical and includes descendants, so use memory.events.local when the numbers do not match the processes in front of you. Pair it with memory.pressure, where full above zero means intervals in which nothing ran but reclaim.
Should I set memory.max on a Firecracker VMM process?
No. A guest's RAM is fixed when the VM is configured — on PandaStack it is baked into the template snapshot, because Firecracker cannot change vCPU or RAM at snapshot restore — so the guest already has a hard ceiling enforced by its own kernel and its own OOM killer. That machinery works and produces diagnosable failures inside a VM that is about to be deleted anyway. A host-side memory.max under the baked size replaces it with something worse: the guest touches a page it legitimately owns, the host charge crosses the line, reclaim cannot free guest anon memory, and the host OOM killer kills the VMM. The machine vanishes mid-instruction with nothing in the guest log. memory.high on a VMM is subtler and also bad, because the throttled task is a vCPU thread, so the guest sees soft lockups, RCU stalls and apparently failed hardware with no way to know it was frozen. Our agent writes cpu.weight and cpu.max per VM and reads memory.current; memory safety comes from admission control instead.
Do dirty page-cache pages escape a cgroup's memory limit?
Not on a modern kernel, and the claim that they do is pre-cgroup-writeback behaviour that should stop being repeated. The old story was that buffered writes dirty pages in the page cache, the device I/O is later performed by a kernel writeback thread belonging to no cgroup, and so the write escapes accounting while the host pays. Cgroup-aware writeback fixed this: the kernel tracks the owning cgroup of each dirty page through the memory controller and attributes the eventual writeback to that owner, and dirty throttling works against a per-cgroup writeback domain so one cgroup cannot dirty the whole host's budget. The real caveats are narrower. It needs filesystem support — ext4 and btrfs have it; check anything more exotic. It needs the memory controller on the same v2 hierarchy as the I/O controller. And pages whose owner cannot be established are attributed to the root cgroup, so attribution blurs through device-mapper stacks, loop devices and network filesystems.
Keep reading
- Guest writeback throttling and dirty_ratio — The same reclaim machinery seen from the write side: why squeezing a guest’s residency shrinks the dirtyable memory its dirty_ratio is computed from, and the guest then blames the disk.
- cgroups v2 explained for sandboxing — The wide tour this post narrows: the unified hierarchy, every controller, and why resource limits are accounting rather than a security boundary.
- PSI for memory oversubscription — The other half of the observability story — what some and full actually measure, and how to build a pressure signal you can act on.
- The OOM killer and guest memory — What happens when the kill lands inside the guest instead of on the host, and how the guest kernel chooses its victim.
- Admission control and backpressure — The mechanism that replaces the limit file for VMs: deciding what gets onto a host before it arrives.
Related posts
- MGLRU and MicroVM Density: Reclaim When Every Page Is Someone's Guest
The host kernel picks reclaim victims using access information already filtered through a second kernel. MGLRU gives it generations instead of two buckets — which changes the shape of your bad days, not how many guests fit on the box.
- DAMON: Measuring What a Sandbox Actually Touches
Charging a host for every guest's configured RAM is what caps your density, and it is wrong by a lot — a guest handed 4 GiB to run an idle Node server touches a fraction of it. DAMON is the kernel subsystem that can tell you which fraction, at a cost bounded by region count instead of memory size. Here is how it works, and the three different wrong answers you can get instead.
- Per-Sandbox Disk I/O Throttling and the Neighbour You Cannot See
CPU contention is visible and roughly fair. Memory contention has pressure signals. Disk contention shows up as latency in somebody else's unrelated workload, with no attribution and no ticket. Here are the three places you can actually throttle it, what each one really controls, and the two layers that quietly lie to you.
- EEVDF vs CFS: What the New Linux Scheduler Means for MicroVM Density
Linux 6.6 swapped the default CPU scheduler out from under everyone. If the runnable threads on your host are the vCPUs of a few hundred oversubscribed microVMs, the change is not academic — and half the tuning advice you'll find now edits sysctls that no longer exist.
- tmpfs in a microVM: The Filesystem That Eats Your RAM
Writing a gigabyte into /tmp inside a microVM does not consume disk. It consumes the guest RAM your process was counting on — and if you snapshot that VM, the gigabyte gets cloned into every copy.
More in Internals · See PandaStack benchmarks
49ms p50 cold start. Fork, snapshot, and scale to zero.