Per-Sandbox Disk I/O Throttling and the Neighbour You Cannot See
Of the four resources a multi-tenant host shares, three fail politely. CPU contention is visible in every dashboard and the scheduler divides it defensibly. Memory contention announces itself: a pressure signal, a reclaim path, eventually a very loud kill. Network contention at least has a victim who can say their downloads are slow, and a byte counter that names the culprit.
Disk fails rudely. One tenant starts an `npm ci` against a cold cache, a `git clone` of a decade-old monorepo, an `fsync`-heavy Postgres import, or the honest old-fashioned `dd` somebody left running in a loop — and p99 latency gets worse for every other guest on the host. Not their throughput, which they would notice. Their latency, in operations they never thought of as I/O: a `require()` that reads forty thousand files, a compiler stat-ing its include path, a health check that touches a log.
And the symptom is "my build is slow today", which nobody files correctly. They file a flake, or they retry and it works and quietly form an opinion about your platform. I build PandaStack, a Firecracker microVM sandbox platform, and this is the dimension where I have the least enforcement and the most respect for the problem. So: where you can clamp disk I/O, what each clamp controls, which two layers lie to you, and what we do and do not do today.
The dimension with no victim-side signal
CPU is divided by a scheduler whose whole job is fairness. Our agent gives each Firecracker process a child cgroup with `cpu.weight` proportional to its vCPU entitlement, so an idle host lets any sandbox burst to every physical core and a busy host divides cores in weight proportion. And `cpu.stat usage_usec` per VM cgroup is an exact count of what a tenant burned — we bill from it. Memory is divided worse but louder: a stall percentage in `/proc/pressure/memory`, residency in `memory.current`, and an OOM killer with a strong opinion if you ignore both.
Disk has neither shape. No fair-share scheduler divides a device between tenants the way CFS divides cores — the block layer's elevator reorders for device efficiency, not tenant equity. No victim-side counter says "I waited" unless you collect one. And the unit of damage is a queue: a tenant issuing a deep stream of small operations does not take a slice and leave the rest, it lengthens the line for everybody, and the cost lands on whoever was behind them.
Worse, the loudest offender is usually innocent. The tenant running `dd` is easy to find. The tenant doing real damage is installing packages that write a hundred thousand small files, or importing a database with an `fsync` per batch — from their side, that is just their job. There is no abuse to ban, only a shape of work that is expensive on shared hardware, so the answer has to be a mechanism rather than a policy conversation.
Three places you can throttle, and they are not interchangeable
A single guest write passes three layers that each have an opinion about rate: the virtual block device inside the VMM, the host block layer underneath the VMM process, and the filesystem holding the backing file. You can clamp at any of them. They do not clamp the same thing.
Layer one: Firecracker's per-drive rate limiter
Firecracker exposes a `rate_limiter` on each block device: two independent token buckets, one counting `bandwidth` in bytes and one counting `ops` in operations, each with a `size`, a `refill_time` in milliseconds, and an optional one-shot burst. The field names are not the rate, so get the semantics exact.
- `size` is bucket capacity — the maximum tokens held, and therefore the largest burst spendable in one gulp. Bandwidth tokens are bytes; ops tokens are single I/O operations.
- `refill_time` is how long a full refill takes, and it is the denominator of your rate. Sustained rate is `size / refill_time`, so a bigger size with a proportionally longer refill buys a larger spike at the same average. That trade is the design decision.
- The one-shot burst is an extra pool spent once and never refilled, for the workload that legitimately spikes at the start and then settles. Check its exact spelling against the API spec for your Firecracker version — the token-bucket shape is stable, names drift between releases, and we run v1.16 here.
- A request needs tokens in every configured bucket to proceed. Configure one and the other axis stays wide open.
You want both, and that is the whole argument for the ops bucket. A bandwidth-only cap stops a tenant moving a lot of data and does nothing against a flood of tiny random writes. Four-kibibyte scattered writes hurt a device most per byte transferred: they defeat merging, defeat readahead, become a long queue of independent requests, and register as a rounding error on any byte counter you are watching. Bandwidth-only catches `dd` and waves through the thing causing your complaints.
The drive object is configured through Firecracker's REST-over-Unix-socket API: a `PUT` to the drive path before boot, and a `PATCH` to the same path to change `rate_limiter` on a live VM. The digits below are illustrative arithmetic chosen so the division is legible — 67,108,864 bandwidth tokens refilling every 1,000 ms, a one-shot pool of four times that, 4,096 ops tokens on the same window. Not a recommendation; real numbers come from measuring your own device.
{
"drive_id": "rootfs",
"path_on_host": "/var/lib/pandastack/vms/7f3c9a21/clone.ext4",
"is_root_device": true,
"is_read_only": false,
"cache_type": "Unsafe",
"io_engine": "Async",
"rate_limiter": {
"bandwidth": {
"size": 67108864,
"one_time_burst": 268435456,
"refill_time": 1000
},
"ops": {
"size": 4096,
"refill_time": 1000
}
}
}One detail matters enormously on a snapshot-restore platform. Most drive configuration is frozen into the snapshot — the cache mode is, because a partial drive update carries only the drive id, the host path and the rate limiter. The limiter is the exception: of all the drive properties, it is precisely the one a partial update can still carry. For networking I would prefer host-side shaping over a baked device setting; for block I/O the limiter is patchable on a live guest, so the argument for leaving it in the VMM is much stronger.
Layer two: the host's cgroup v2 I/O controllers
Underneath the VMM, Firecracker is just a process reading and writing a file, and cgroup v2 has three controllers that want to govern it. They answer different questions and fail differently.
`io.max` is a hard cap, per cgroup per block device, with separate read and write limits in bytes per second and operations per second. Wonderfully predictable: a tenant gets what you wrote in the file and not a byte more. Also wasteful exactly the way a fixed CPU quota is — sandbox workloads are bursty and most sandboxes are idle most of the time, so a per-sandbox hard cap leaves the device idle while throttling the one tenant who wants it.
`io.latency` inverts the question. You set a latency target on the group you want protected; the controller leaves the device alone until that group misses its target, then throttles the others until it recovers. That is the shape for a clear priority ordering — a managed-database VM whose commit latency is a product promise, against a build sandbox whose only promise is to finish eventually. It costs nothing when nobody suffers, which `io.max` cannot offer. But it protects rather than allocates, so it has nothing to say when ten equal peers want the device at once.
`iocost` is the ambitious one: it builds a cost model of the device — the relative expense of sequential reads, sequential writes and random operations — then splits it proportionally by `io.weight`, the way `cpu.weight` splits cores. When it works you get the property I actually want: idle device means anybody bursts, contention means the split follows the weights. The price is the model. Wrong optimistically it fails to protect; wrong pessimistically it throttles a device with capacity to spare.
# Our agent already runs one delegated child cgroup per VM for cpu.weight,
# so the enforcement point for I/O is a file in a directory that exists.
CG=/sys/fs/cgroup/system.slice/pandastack-agent.service/vm-7f3c9a21
PARENT=$(dirname "$CG")
# cgroup v2's no-internal-process rule: a controller must be enabled in the
# PARENT's subtree_control before any child can use it.
echo '+io' > "$PARENT/cgroup.subtree_control"
# io.max is keyed on the BLOCK DEVICE, not the filesystem path. Resolve the
# major:minor of the disk the store actually lives on -- where the store is a
# loopback image on another filesystem this is not the disk you first thought
# of, and capping the wrong device caps nothing.
SRC=$(df --output=source /var/lib/pandastack | tail -1)
MAJMIN=$(lsblk -ndo MAJ:MIN "$SRC" | tr -d ' ')
# A hard cap. rbps/wbps are bytes per second; riops/wiops are operations per
# second. Fill these from your own measured device baseline -- a number
# borrowed from a blog post is a guess about somebody else's hardware.
echo "$MAJMIN rbps=$RBPS wbps=$WBPS riops=$RIOPS wiops=$WIOPS" > "$CG/io.max"
# Protection instead of a cap: throttle EVERYONE ELSE when this group's
# completion latency exceeds the target, expressed in milliseconds.
echo "$MAJMIN target=$TARGET_MS" > "$CG/io.latency"
# What this cgroup has actually done, per device.
cat "$CG/io.stat"
# 253:0 rbytes=... wbytes=... rios=... wios=... dbytes=... dios=...
# How long tasks in this cgroup spent stalled waiting on I/O. This is the
# victim-side signal, and it is the one number worth alerting on.
cat "$CG/io.pressure"
# some avg10=... avg60=... avg300=... total=...
# full avg10=... avg60=... avg300=... total=...Note which file is the interesting one. `io.stat` is volume: who is busy. `io.pressure` is stall time: who is suffering — and in a noisy-neighbour investigation that is the question with a victim attached. A tenant moving bytes on an idle device is not a problem; a tenant with a rising `full` stall percentage is, however few bytes they moved.
The blind spot: dirty pages, writeback, and who gets charged
The host's controllers throttle where a block request is submitted to the device. A buffered write does not submit one: it dirties a page in the host page cache and returns, and the device write happens later, from a writeback context, possibly long after the process that caused it has exited. So the obvious model — cap the VMM process, therefore cap its writes — is wrong in the only direction that matters. A guest can dirty host memory far faster than the cap allows, return success, and leave the bill for a flusher thread to pay later, in a burst, while somebody else wanted the device.
Modern kernels try to fix this. Cgroup-aware writeback tracks the owning cgroup of each dirty page through the memory controller and charges the eventual device write to that owner. But that support is per filesystem, it needs the memory controller enabled on the same hierarchy, and pages whose ownership cannot be established — shared mappings, anything reclaimed where the owner is ambiguous — land on the root cgroup, which is to say on nobody. Your cap is not ignored; it is enforced on a subset of the traffic you cannot enumerate.
On PandaStack this meets a decision we already made. Our rootfs drives run in the unsafe cache mode: guest flushes are not turned into real host `fsync` calls, because a sandbox rootfs is copy-on-write and discarded with the sandbox. Volumes are the opposite — a volume is the durable store, managed Postgres keeps its data directory on one. Which is inconvenient: the drives where tenants write heaviest are the ones whose writes linger longest in host memory, and the drives whose I/O is honest and attributable are the database volumes I least want to throttle.
Layer three: where a guest write is not one host write
A PandaStack sandbox does not get a freshly written disk image. It gets a copy-on-write clone of a template rootfs — an XFS reflink, or a dm-snapshot device. That is why a create is cheap: a reflink is an O(metadata) operation that shares every data extent with the template and copies nothing, part of why a create lands at a p50 of 179 ms rather than the three seconds an unbaked template's cold boot costs.
So write amplification is not constant. It is highest at the start and it decays. The first write to a shared extent cannot happen in place, because that extent belongs to the template every other sandbox is also using — the filesystem allocates, copies, then writes. The second write to the same region is ordinary. A freshly created sandbox doing scattered small writes across its rootfs generates more host I/O than its guest-side counters suggest, at the moment of creation, which is probably the moment several sandboxes were created at once.
The conclusion is counter-intuitive. Two sandboxes running identical workloads on the identical template load the device differently depending on how much rootfs they have already touched. A long-lived sandbox is cheaper per guest write than a new one, and a fleet of short-lived sandboxes — the entire point of the product — pays the premium continuously and never amortises it. The guest's metrics show none of it: from inside, it issued a write and the block device said yes.
The controls side by side
| Control | What it caps | Enforced where | Burst behaviour | Blind to page-cache writeback? | Operational cost |
|---|---|---|---|---|---|
| Firecracker drive rate_limiter | Bytes and operations per virtual drive, per VM | In the VMM, before the drive's file engine | Token bucket: size is the burst, plus an optional one-shot pool | No — it sits above the host page cache and sees requests as issued | Low: JSON on the drive object, patchable on a live VM |
| cgroup v2 io.max | Bytes and operations per cgroup, per block device | Host block layer, at bio submission | None — a flat ceiling, always | Partly: depends on cgroup writeback support and page ownership | Low to set, high to tune; wastes an idle device by design |
| cgroup v2 io.latency | Nothing directly — throttles others to protect one group | Host block layer, reactive to completion latency | Free rein until the protected group misses target | Same caveat as io.max | Medium: needs a priority ordering you can defend |
| cgroup v2 iocost with io.weight | Proportional share of a modelled device | Host block layer, with a device cost model | Idle device means full speed; contention splits by weight | Same caveat as io.max | High: a wrong model under- or over-protects |
| No throttle (shares and admission only) | Nothing | Nowhere | Unbounded | Entirely | Zero until the first incident, then a blind afternoon |
No row is a complete answer. The VMM limiter sees the guest's intent and nothing about the device or its neighbours; the host controllers see the device and may be lied to about who caused the traffic. They compose rather than compete — the effective limit is whichever binds first — so use the VMM limiter for the per-tenant ceiling and a host controller for host-level protection.
The restore path is not a steady-state guest
There is no warm pool of idle VMs here. Every create restores a baked Firecracker snapshot, so every create does a burst of I/O unrelated to the tenant's workload. On the simple path the memory file is mapped `MAP_PRIVATE` and paged in lazily as the guest touches it. On the streaming path — the alternative being to download a multi-gibibyte memory file before the VM can start — memory is demand-paged from object storage over HTTP range requests in four-mebibyte chunks, with a zero-chunk header so all-zero chunks are never fetched and a bake-time prefetch trace replayed in the background.
Read that again with a throttle in mind. Those chunks arrive over the network and land on the disk, in a shared on-disk chunk cache keyed per seed generation: the first restore of a seed on a host pays object-storage latency once, every later restore is local. A cold create is therefore a burst of writes into that cache followed by a burst of reads out of it, and neither is the tenant doing anything — it is the platform assembling the machine they asked for. A throttle tuned for a steady-state guest turns a 179 ms create into something you will have to explain.
What PandaStack actually clamps today
In the spirit of not describing a technique as a feature, here is the current state with the gaps in it. I read our agent rather than trusting my memory of it, which I recommend before writing anything of this kind.
- CPU: enforced. Each Firecracker process gets a delegated child cgroup with `cpu.weight` proportional to its vCPU entitlement, reconciled every fifteen seconds so create, fork, wake, restore and post-restart recovery all converge. Weights bind only under contention.
- CPU accounting: enforced and billed, from `cpu.stat usage_usec` per VM cgroup — active CPU seconds, not wall-clock reservation.
- Memory: a ladder, not a cap. A controller watches host `MemAvailable` and the memory PSI stall percentages and squeezes the coldest VMs via `memory.high`; under genuine pressure it snapshot-freezes the coldest idle sandbox, which wakes on touch. Database VMs are never touched.
- Disk I/O rate: not enforced. No `rate_limiter` on our drive objects, and no `io.max`, `io.latency` or `iocost` anywhere in the agent. The only mention of a rate limiter in the source tree is a comment about which fields a partial drive update carries. This is the gap.
- Disk I/O accounting: not collected per sandbox either. The metrics poller reads real cgroup values for CPU and memory and leaves the disk fields at zero rather than lying with a wrong number; the comment notes that filling them means reading `io.stat` per VM.
- Disk capacity: enforced, which is a different thing. A free-space floor on the store disk: below it creates are refused with a 507, the floor is published on the heartbeat so the scheduler skips the host, and every out-of-space error inside a create maps to the same refusal.
That last bullet is the distinction people collapse. Admission control protects capacity: not here, not now, and the control plane retries elsewhere. Throttling protects latency: yes, but slowly. The floor exists because a host once ran at full for ten days while still receiving creates. A busy disk is a different problem and the floor is blind to it: a host with free space and a saturated device passes every gate we have.
Why has the limiter not shipped? Ordering, not disagreement. What came first bounded unbounded losses: a tenant reaching another tenant's network, a tenant stealing a host credential, a host wedging out of space and silently failing creates. A noisy disk is a bad afternoon for other people's build times, which is genuinely bad and genuinely not the same category. When it ships the shape is clear from everything above — and per-sandbox `io.stat` and `io.pressure` come first, so the numbers are measured rather than invented.
Finding the offender
Before you can limit you have to attribute, and attribution is where this problem earns its reputation. The complaint arrives as "builds are slow" and your job is to get from there to a sandbox id. The obstacle: at the block layer every microVM looks identical — a Firecracker process reading and writing a file. No per-tenant process names, no helpful command lines, one `fio`-shaped blur.
# 1. Is the device the bottleneck at all? Queue depth and utilisation say
# whether it is saturated; the await columns say what one request costs.
# If utilisation is low and await is low, your slow build is not a disk
# problem and you are about to waste an hour.
iostat -x 2
# 2. Who is moving bytes? Diff io.stat across every VM cgroup. Volume alone
# convicts nobody, but it narrows the field fast.
AGENT=/sys/fs/cgroup/system.slice/pandastack-agent.service
for cg in "$AGENT"/vm-*; do
printf '%-24s %s\n' "$(basename "$cg")" "$(tr '\n' ' ' < "$cg/io.stat")"
done
# 3. Who is STALLING? This is the victim-side signal and the reason you were
# called. A group with a climbing 'full' average is waiting on I/O that it
# is not necessarily causing.
for cg in "$AGENT"/vm-*; do
printf '%-24s %s\n' "$(basename "$cg")" "$(awk '/^full/ {print $2}' "$cg/io.pressure")"
done
# 4. What does the block-layer latency distribution look like? A long tail
# with a healthy median is queueing; a uniformly shifted histogram is the
# device itself, or a neighbour below you that you cannot see.
biolatency-bpfcc -D 10 1
# 5. Individual slow requests with the issuing task and sector. Every row
# will say 'firecracker', which IS the attribution problem -- but the pid
# maps to a cgroup and the cgroup name carries the sandbox id.
biosnoop-bpfcc | awk -v ms="$SLOW_MS" '$NF+0 > ms'
# 6. Close the loop.
cat /proc/$PID/cgroupStep one is not optional and it is the step everyone skips. Disk is a fashionable thing to blame, and a good fraction of "the disk is slow" is a guest with the wrong readahead, a lock, a DNS timeout, or a workload that was always like this. If the device is not saturated and requests are not slow, stop.
The eBPF tools see the host and only the host: a VMM process doing I/O against a file, no guest filenames, no guest processes. The boundary that makes the platform safe makes this investigation hard — the whole subject of our post on what eBPF can and cannot see from outside a guest. Attributing to a sandbox still works: the pid maps to a cgroup and the cgroup name carries the sandbox id. Going deeper needs in-guest instrumentation.
If you are building this, in this order
- Collect `io.stat` and `io.pressure` per sandbox before throttling anything. You cannot pick a limit without a distribution, and the stall signal is the one that correlates with the complaints you receive.
- Decide what you are protecting. "Fairness" is three answers: a database's commit latency against build sandboxes wants `io.latency`; every tenant against one tenant wants weights; the host against an unbounded abuser wants a cap.
- Exempt or separately budget the platform's own I/O — restores, seed pulls, chunk-cache fills, reflinks, hygiene passes. Put it in the guest's budget and you will regress create latency and spend a week blaming object storage.
- Set the ops bucket, not just bandwidth. Bandwidth-only catches the obvious offender and waves through the small-random-write pattern behind your real complaints.
- Give the burst real room, make the numbers changeable per plan on a live host, and confirm the knob survives a restore. The first time a paying customer's workload hits your cap at 2am you want a command, not a template re-bake.
- Verify the limit binds, with a real workload, by watching the throttle counters move. If they never move you capped the wrong device, the wrong direction, or a path the page cache routes around.
The summary
Disk is the dimension where contention is real, attribution is hard, and the victim's symptom is too vague to report. CPU has a fair scheduler and an exact per-tenant counter. Memory has a pressure signal and a reclaim path. Disk has a queue, and a queue punishes whoever is behind the person who filled it.
Three enforcement points, not substitutes. Firecracker's per-drive limiter is a clean token-bucket pair sized by `size` and `refill_time`, and on a snapshot-restore platform it is patchable on a live VM when most drive properties are frozen. The cgroup v2 controllers offer a predictable-but-wasteful cap, a protection target that costs nothing until someone suffers, or a proportional model that is correct as far as its cost model is. And the filesystem changes the arithmetic underneath: on a reflinked rootfs the first write to a shared extent costs a copy, so amplification is highest on the newest sandboxes, and a platform built from short-lived sandboxes never amortises it.
Two things lie to you: the page cache, which lets a guest dirty host memory far faster than any cap and defers the bill to a context charged to nobody in particular, and a capacity gate that looks like a throttle. We clamp CPU with weights, manage memory with a pressure ladder, gate capacity with a store floor, and do not yet rate-limit block I/O per sandbox. The measurement comes first anyway — every limit chosen without a distribution is a guess with a changelog entry.
Frequently asked questions
Should I cap IOPS or bandwidth?
Both, and if you only get to set one, set operations. The two limits defend against different abuses and neither constrains the other. A bandwidth cap bounds how much data a tenant can move, which catches the obvious offender — the sequential bulk transfer, the runaway `dd`, the enormous file copy. It does nothing against a flood of small scattered operations, and small scattered operations are what actually degrade a shared device: they defeat request merging, they defeat readahead, they become a long queue of independent round trips, and they can make a device thoroughly miserable while barely registering on any byte counter you are watching. So: set the ops ceiling first because that is where the latency damage comes from, add bandwidth second, and give both a generous one-shot burst so the legitimate ninety-second dependency install never touches the limit while the tenant transferring continuously for six hours does. Firecracker requires tokens in every configured bucket, so configuring both is a conjunction rather than a choice, and whichever is tighter for a given workload binds.
Why did capping I/O make my memory pressure worse?
Because memory reclaim is itself I/O, and you just throttled it. When the kernel needs to free memory it writes dirty pages out, and if swap is involved it writes anonymous pages out too. Both are block device traffic submitted from a reclaim context, and if that context sits inside a cgroup whose `io.max` you tightened, reclaim gets slower. Slower reclaim under memory pressure does not degrade gracefully — it stalls, because the allocation that triggered the reclaim is waiting on it. So you tightened a disk cap and the symptom appeared as memory stalls, which is the most confusing failure mode in this area. We hit a related trap directly: our pressure ladder sets `memory.high` on a VM's cgroup to squeeze its cold pages out, and we had to gate that rung on the host actually having swap, because `memory.high` with nowhere to put the pages does not reclaim, it stalls — a stall percentage in the nineties and a workload that stopped making progress. Memory and I/O limits are coupled through the reclaim path: change an I/O cap, re-check your memory pressure signals, and make sure a memory controller's own reclaim traffic is not subject to the cap it is working within.
Does Firecracker's drive rate limiter survive a snapshot restore?
The baked configuration comes back with the snapshot, and — unusually among drive properties — the limiter is also one of the few things you can change afterwards on a live restored VM. That matters a great deal when every create is a restore rather than a boot. Most of a drive's configuration is effectively frozen at bake time: the cache mode is, because a partial drive update carries only the drive identifier, the host path and the rate limiter, so changing a cache mode means re-baking every template rather than shipping a config change. We have a capitalised comment in our own driver to that effect, written by somebody who learned it the hard way. The limiter being on that short list inverts the advice I would give for networking, where I prefer host-side shaping precisely because a baked device setting cannot be varied per customer or per incident. For block I/O you can patch the limiter on a running guest, which makes the in-VMM enforcement point practical for per-plan limits and for incident response. Confirm which fields are runtime-mutable against the API spec for your exact Firecracker version before building a control plane on it — we run v1.16, and this detail shifts between releases.
Why can I not see which process inside the sandbox is causing the I/O?
Because the isolation boundary that makes a microVM safe is the same boundary that makes this attribution impossible from outside. From the host's point of view a sandbox is one process — the VMM — doing reads and writes against a backing file. There is no guest process table, no guest file names, no guest syscalls crossing into host-visible territory. A host eBPF tool tracing block requests will faithfully report that the responsible task is `firecracker`, for every sandbox on the box, which is correct and useless. This is the inverse of the container case, where the host kernel is the guest kernel and a host tracer sees everything down to individual guest syscalls — containers are a panopticon, microVMs are opaque, and which you want depends on whether you are debugging your own workload or isolating somebody else's. What you can do from the host is attribute to the sandbox: the VMM's pid maps to its cgroup and the cgroup name carries the sandbox identifier. Getting from there to a process inside that tenant requires in-guest instrumentation, which is a feature you choose to build rather than a technique you apply mid-incident.
Is a disk free-space limit the same as disk I/O throttling?
No, and they protect against completely different failures. A free-space gate is admission control: it answers a yes-or-no question at create time, it protects capacity, and its failure mode is a host that fills up and starts failing everything routed to it. We run one — the agent checks free space on the store disk before a create, refuses with a 507 below the floor, publishes the floor on its heartbeat so the scheduler skips that host, and maps any out-of-space error mid-create to the same refusal so the control plane retries elsewhere instead of marking the app broken. It exists because a host once ran at full for ten days while still receiving creates. Throttling is a rate: it answers how fast, protects latency, and its failure mode is one tenant lengthening the device queue for everybody else. The two are orthogonal in the inconvenient direction. A disk can be ninety percent empty and completely saturated, and in that state a capacity gate sees nothing wrong and keeps accepting work — exactly the scenario where a noisy neighbour is ruining somebody's afternoon. If you have shipped a capacity gate and you describe it as protection against noisy neighbours, you have not.
Keep reading
- Rate-limiting the network per sandbox — The same argument on the other dimension, where the enforcement point is a veth pair you own both ends of.
- Firecracker's rate limiter, explained — The in-VMM token bucket in depth: size, refill_time, one-shot burst, and how it composes with cgroups.
- dm-snapshot vs reflink for CoW rootfs — The layer that makes a guest write cost more than one host write, and why the premium decays.
- What eBPF can and cannot see — Why host tracing attributes every slow block request to a process called firecracker.
- cgroups v2 explained for sandboxing — The unified hierarchy, delegation, and the I/O controller this post keeps writing files into.
Related posts
- What "persistent" actually means in a sandbox
You wrote the file. You ran cat and saw it. Neither of those facts says the bytes are on a disk. Here is every layer a write passes through inside a microVM, which of them a crash erases, and why the honest answer for an ephemeral rootfs is not "fsync harder" but "get the artifact out".
- EEVDF vs CFS: What the New Linux Scheduler Means for MicroVM Density
Linux 6.6 swapped the default CPU scheduler out from under everyone. If the runnable threads on your host are the vCPUs of a few hundred oversubscribed microVMs, the change is not academic — and half the tuning advice you'll find now edits sysctls that no longer exist.
- Firecracker vCPU Hotplug, and How CPU Scaling Actually Works
Somebody always asks for "more CPU" on a running microVM. Firecracker's answer is no — vcpu_count is fixed before InstanceStart and frozen into every snapshot. The good news: the lever you actually wanted was never vCPU count in the first place.
- Per-Customer Cron Jobs with microVM Isolation
One customer's runaway cron shouldn't be able to page your whole on-call. Here's how to run every tenant's scheduled job in its own throwaway microVM.
- virtio-blk Discard and TRIM in Firecracker, Explained
A job downloads 6 GB, does its work, deletes it, and the host-side image is still 6 GB. Deleting a file is a metadata edit in the guest's own allocator; the host never hears about it unless somebody issues a discard. Every link in that chain fails silently.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.