firecracker-containerd: Running OCI Images in MicroVMs, and What It Costs You
There is a question that comes up roughly every second week in my inbox, phrased about six different ways, and it always reduces to the same thing: “I already have OCI images, a registry, and a cluster. Can I just keep all of that and swap the thing at the bottom for a microVM?”
The answer is yes, there is a project for exactly that, it is called firecracker-containerd, and it is a more interesting piece of engineering than its README suggests. It is also, for a good half of the people who ask me about it, the wrong tool — not because it is bad, but because they are describing a sandbox API and it is a container runtime. Those are different products that happen to share a hypervisor.
I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. I have a horse in this race, so I am going to spend the first two-thirds of this post explaining firecracker-containerd's architecture as fairly as I can, including the places where it is straightforwardly better than what I built, and only then get to the decision framing.
What firecracker-containerd actually is
It is not a binary you run. It is a set of components that slot into stock containerd and collectively make a microVM look like a container runtime target. There are four pieces worth knowing by name, because when something breaks you will need to guess which one broke.
- The runtime shim — `containerd-shim-aws-firecracker`. containerd's shim API is how containerd delegates “actually run this” to something else; runc has a shim, and so does this. One shim process per microVM, and it is the thing that holds the Firecracker VMM socket and the vsock connection to the guest.
- The firecracker-control plugin — a containerd plugin that exposes a gRPC service for VM lifecycle (create a VM, stop a VM) independently of container lifecycle. This is the piece that makes the “one VM, several containers” model possible, and it is also why the project has an API surface that stock containerd does not.
- A snapshotter — in practice the device-mapper `devmapper` snapshotter, which materialises image layers as thin-provisioned block devices rather than as an overlay mount. A VM cannot mount your host's overlayfs; it needs something that looks like a disk.
- The in-guest agent — a small daemon inside the microVM that receives the OCI runtime spec over vsock and runs it with runc. Yes, runc. Inside the VM. There is a container in there, and it is a real one.
That last point is the one that surprises people, and it is the key to understanding the whole design. firecracker-containerd does not “turn a container into a VM”. It boots a VM and then runs a perfectly ordinary runc container inside it. The VM is the isolation boundary; the container is still the packaging and the process model.
Where the image comes from, and how it becomes a disk
The image comes from your registry, pulled by containerd exactly as it always was. Nothing novel happens there, which is precisely the selling point. What is novel is the next step, because an OCI image is a stack of tar layers and a JSON config, and a Firecracker guest can consume neither.
The snapshotter's job is that conversion. The devmapper snapshotter keeps a device-mapper thin-pool on the host; each image layer becomes a thin device, each container gets a writable thin snapshot of its parent, and the result is a block device with a path. The shim then attaches that block device to the microVM as a virtio-block drive, and inside the guest it shows up as `/dev/vdb` or similar for runc to use as the container's rootfs.
There is a wrinkle here that I find genuinely clever. Firecracker's drive configuration is largely a boot-time concern, so the runtime pre-attaches a number of placeholder drives — stub drives, backed by nothing interesting — when it boots the VM, and then patches a real backing file into one of them when a container needs a rootfs. That is how you get a container's disk into an already-running VM without depending on hot-plug. It also means the number of containers a VM can host has a ceiling set by configuration, which is the kind of detail that is invisible until it is the only thing that matters.
Why there is an agent inside the guest at all
People read “agent in the guest” as bloat. It isn't; it is load-bearing, and the reason is that a VM boundary is opaque in both directions.
Think about what `ctr run --tty` promises. You get the container's stdout on your terminal, your keystrokes go to its stdin, and when it exits you get its exit code. Now put a hypervisor between the two halves of that sentence. The host side has no process table entry for the container, because the container's processes live in another kernel. Something on the inside has to be the one that forks runc, wires up the fds, notices the exit, reads the status, and ships all of it back across a transport that is — in Firecracker's case — a vsock socket, not a pipe.
So the agent is the thing that: owns the guest's init responsibilities so there is a sane PID 1 reaping orphans; receives the OCI spec and bundle metadata; invokes runc with it; proxies stdio streams; and reports the exit code. If you removed it you would have to re-invent all five jobs, worse, in the serial console.
A nice consequence of doing it this way is spec fidelity. Because a real runc consumes a real OCI runtime spec inside the guest, image semantics behave the way the OCI spec says they should — `ENTRYPOINT`, `CMD`, `ENV`, `WORKDIR`, the user, the capability set. I will be blunt about the contrast: on PandaStack, a template's Docker `ENV` does not reach a detached process in the guest, because we flatten the image into a rootfs and the thing that sets environment for a detached process is `/etc/environment` via PAM. firecracker-containerd simply does not have that class of bug, and that is a genuine architectural advantage of keeping runc in the picture.
Running it, and the one-VM-many-containers model
# firecracker-containerd is four moving parts, not one binary. On the host:
#
# containerd -- stock containerd, plus two plugins
# containerd-shim-aws-firecracker -- the runtime shim: one process per microVM
# firecracker-control plugin -- a gRPC VM lifecycle service (CreateVM /
# StopVM), so a VM can outlive and host
# more than one container
# devmapper snapshotter -- turns image layers into a thin block device
#
# ...and INSIDE the guest: an `agent` that is what PID 1 answers to, running the
# OCI bundle with runc and relaying stdio and the exit code back over vsock.
# 1. The device-mapper thin-pool. This is real storage administration and it does
# not administer itself: you size the data and metadata volumes up front, and
# then you watch them. A full metadata device is an outage, not a warning.
sudo dmsetup create fc-dev-thinpool --table \
"0 ${DATA_SIZE_SECTORS} thin-pool ${META_DEV} ${DATA_DEV} 128 32768 1 skip_block_zeroing"
sudo dmsetup status fc-dev-thinpool # used_data / used_metadata live here
# 2. The runtime config. Illustrative only -- the path and several key names have
# moved between releases, so read the repo's docs for the tag you build.
cat /etc/containerd/firecracker-runtime.json
# {
# "firecracker_binary_path": "/usr/local/bin/firecracker",
# "kernel_image_path": "/var/lib/firecracker-containerd/runtime/vmlinux",
# "kernel_args": "console=ttyS0 noapic reboot=k panic=1 pci=off nomodules ro",
# "root_drive": "/var/lib/firecracker-containerd/runtime/rootfs.img",
# "cpu_count": 1,
# "mem_size_mib": 512,
# "default_network_interfaces": [],
# "debug": false
# }
#
# Read that file again and notice what is in it and NOT in your image: the guest
# kernel, and the VM's OWN root filesystem -- the one that carries the agent.
# `debian:bookworm` brings neither. Your image supplies the CONTAINER's rootfs.
# 3. Run a container in a microVM.
sudo ctr --address /run/firecracker-containerd/containerd.sock \
image pull --snapshotter devmapper docker.io/library/debian:bookworm
sudo ctr --address /run/firecracker-containerd/containerd.sock \
run --snapshotter devmapper \
--runtime aws.firecracker \
--rm --tty \
docker.io/library/debian:bookworm demo /bin/bash
# What that one line actually did, in order: resolve and pull the image; commit a
# thin snapshot and expose it as a block device; boot a Firecracker VM; attach
# the block device to it; ship the OCI runtime spec over vsock to the agent; the
# agent runs runc inside the guest. Every one of those steps costs something,
# and none of them is "restore a guest that already finished booting".One VM, possibly several containers
By default the shim gives each container its own microVM, which is the configuration most people want and the one that matches the isolation story you are presumably here for. But the control plugin lets you create a VM explicitly and then place several containers into it.
This is worth doing when a group of containers belongs to the same tenant and you want to amortise the VM's memory overhead and boot cost across them. It is worth not doing the moment those containers belong to different tenants, because co-tenanted containers inside one microVM share that microVM's kernel with each other. You have bought a VM boundary around the group, not between its members. A container is still, as ever, a polite suggestion to the kernel — it is just that now you have chosen which kernel it is being polite to.
The gotchas that cost people a weekend
In rough order of how often I see them.
- The guest kernel is yours to supply, and it is not the image's kernel — because images do not have kernels. `alpine` is a userspace. `ubuntu` is a userspace. The kernel comes from `kernel_image_path` in the runtime config, you build or fetch it, and one kernel serves every image you run. This confuses people badly, and the confusion usually surfaces as “why does my container behave differently” when the honest answer is “because you are running it on a kernel you did not choose carefully”.
- That kernel needs a config suited to Firecracker's machine: virtio over MMIO rather than a PCI bus to walk, and the 8250 serial console, which is your only window into a boot that fails before the agent starts. Start from Firecracker's own guest configs in `resources/guest_configs/` rather than from `defconfig`.
- The VM's root filesystem is a second artifact you own, separate from both the kernel and your images. It carries the agent. When you upgrade firecracker-containerd you are often rebuilding that image too, and if you treat it as a one-time setup step you will eventually run a new shim against an old agent and get an error message that explains nothing.
- Networking is a tap device per VM, and the integration path is CNI. The trick that makes this work — running a CNI plugin chain that ends by redirecting a CNI-managed interface onto a tap with `tc` rules — is a real and separate project, and your CNI config is now part of your runtime's config. Debugging it means debugging inside a network namespace you did not create by hand.
- Volumes and extra block devices are block-device plumbing, not bind mounts. A host path is not automatically visible in the guest; something has to be attached as a drive or shared over a filesystem transport, and which of those is available to you depends on the versions in play.
- Logs and stdio cross the vsock boundary. When something hangs, the question “is it the container, the agent, the vsock, the shim, or containerd” has five answers and you will need instrumentation at more than one layer to tell them apart.
- The thin-pool is a database you did not know you had deployed. It needs sizing, monitoring of both data and metadata usage, and a plan for what happens when it fills. Overlayfs fails by returning ENOSPC; a thin-pool with exhausted metadata fails in more creative ways.
Who should use it, and who should not
Here is the framing I actually give people, and it has nothing to do with which hypervisor is cooler.
Use it when containerd is already your world. If you have images in a registry, `ctr` or CRI in your tooling, a Kubernetes-shaped mental model, and a workload where the shared host kernel is the specific thing keeping you up at night, firecracker-containerd is the cheapest path to a VM boundary that does not force you to re-package anything. Your images stay images. Your registry stays your registry. You change a runtime handler. That is a remarkably small diff for a remarkably large change in blast radius.
Do not reach for it if what you actually want is a fast sandbox API. This is the distinction that matters and it is structural, not a tuning problem. Count the work on the critical path of a cold request: resolve and pull an image, commit a snapshot through the snapshotter, boot a VM, attach the device, hand a spec across vsock, start runc. Compare that to a fast path whose entire job is to map a memory image of a guest that already finished booting and resume it. On our path that is a p50 of 179 ms and a p99 of 203 ms, with the snapshot load itself around 49 to 80 ms; a cold boot of a template before any snapshot exists costs about 3 seconds, which is the price we pay once per template rather than once per request.
Those are not the same chain with different constants. They are different chains. And to be fair to the longer one: it buys OCI compatibility, which snapshot-restore does not give you for free, and “arbitrary customer-supplied image, right now, with no build step” is a product requirement that my architecture simply cannot satisfy.
Pick the runtime whose fast path is the thing you do most often. Everything else is a tuning exercise against the wrong chain.
| Runtime | Isolation boundary | Packaging | Fast path | Best for |
|---|---|---|---|---|
| runc | The host kernel, partitioned by namespaces, cgroups and seccomp. Shared kernel means shared attack surface. | OCI images, natively. It is the reference implementation. | Unpack layers, clone namespaces, exec. Milliseconds once the image is local. | Code you wrote, or code you trust as much as code you wrote. |
| gVisor / runsc | A user-space kernel. Syscalls are intercepted and re-implemented in a userspace sentry rather than passed to the host kernel. Not a VM, and there is no guest kernel. | OCI images, natively — it is an OCI runtime and drops into Docker and containerd. | Start the sentry, run. Fast to start; syscall-heavy workloads pay an ongoing interception cost. | Container platforms that want a much narrower host-kernel surface without taking on VM operations. |
| firecracker-containerd | The microVM. Containers co-located in one VM share that VM's kernel with each other. | OCI images, natively — your registry, your `ctr` and CRI tooling, a different runtime handler. | Pull, snapshotter, boot a VM, then start a container in it. | Teams already living in containerd who want VM-level isolation without re-packaging anything. |
| Kata Containers | The VM. The standards-track CRI and OCI VM runtime, with a broader VMM and device story. | OCI images via CRI; a RuntimeClass in Kubernetes. | The same shape — VM, then container — with more hypervisors, devices and knobs available, and correspondingly more surface to learn. | Kubernetes clusters that need VM-isolated pods and want the option to change hypervisor later. |
| Ignite / Flintlock | The VM, used as a machine rather than as a container — init system inside, long-lived. | OCI images used as a distribution format for the VM's rootfs and kernel, not as a container to execute. | Boot a VM the way you boot a VM. | Declarative VM fleets and Cluster API. Treat Ignite as a reference design and check project health before betting on either. |
| PandaStack | The microVM, one per sandbox, with its own network namespace, veth pair and tap device. | A template you build once from a Dockerfile, baked into a snapshot. | Restore the memory image of a guest that already booted: p50 179 ms, p99 203 ms. | A sandbox API where create latency is the product. No arbitrary OCI image at create time. |
One correction I make a lot: gVisor is not “Firecracker but lighter”. It is a different kind of thing entirely — there is no guest kernel and no hypervisor, just a process that pretends to be Linux convincingly enough. And Ignite and Flintlock are not competitors to firecracker-containerd at all; they manage Firecracker VMs as machines, with OCI images borrowed as a convenient way to ship a rootfs. Confusing the two leads to architecture diagrams that cannot be built.
An OCI image is a filesystem plus some metadata
Strip the ecosystem away and an OCI image is two things: a stack of tar layers that compose into a filesystem, and a JSON config describing how to start a process in it. That is the whole format. It is a remarkably good format, and its goodness comes from being boring.
A VM, however, does not consume filesystems. It consumes block devices, or a filesystem share over a virtio transport. So “running an OCI image in a VM” always, without exception, involves converting that layer stack into something a guest can mount. There is no design that avoids the conversion. There is only the question of when you pay for it.
- Pay per run — firecracker-containerd's answer. A snapshotter does the conversion on the create path, which means a thin-pool to operate and work on the critical path, and in exchange any image in your registry can run right now with no build step.
- Pay once, at build time — our answer. Flatten the image into an ext4 rootfs, cold-boot it once, snapshot the running guest, and then every create is a reflink of the disk plus a restore of the memory image. The conversion is amortised across every sandbox that template will ever create.
Baking has a second effect that is easy to miss and is actually the bigger one. Once you have snapshotted a guest that finished booting, you are no longer competing on boot time at all — you have moved userspace initialisation into the artifact too. The interpreter is already warm. The service already bound its port. That is not an optimisation of the container start path; it is the deletion of it.
# The PandaStack equivalent pays the image-to-block-device conversion ONCE, at
# build time, and never again on the create path.
pandastack template build -f Dockerfile -n my-template \
--size-mb 8192 \ # rootfs size. Measure `du -sh` of your tree, then
# double it; running out of disk in a baked guest is a
# re-bake, not a resize.
--memory-mb 2048 # guest RAM, BAKED INTO THE SNAPSHOT. Firecracker cannot
# change guest RAM or vCPU count at snapshot restore, so
# this is the only place it is ever chosen. --cpu is
# deprecated and ignored; every template gets 8
# burstable vCPUs.
# Roughly what happens: flatten the Dockerfile's filesystem into an ext4 rootfs,
# cold-boot it once (~3 s -- the only cold boot this template will ever pay), and
# capture a Firecracker snapshot: a memory image plus device state. Every create
# after that is a restore of that memory image, not a boot.
#
# The trade, stated plainly, because it is the whole point of the comparison:
# there is no --runtime flag here that lets a caller hand me an arbitrary OCI
# image at create time. You build a template first. firecracker-containerd does
# not make you do that, and for some teams that single sentence decides it.from pandastack import Sandbox
# Where the time goes on a create: nowhere interesting. There is no image pull,
# no snapshotter, no thin-pool. The rootfs is reflinked from the template --
# copy-on-write, and local, because CoW needs a local block device -- and the
# guest's memory comes back from a snapshot. p50 179 ms, p99 203 ms on our fast
# path; the /snapshot/load call itself is roughly 49-80 ms of that.
sbx = Sandbox.create(template="my-template", ttl_seconds=600)
# Note what is NOT passed: cpu= and memory_mb=. The agent silently corrects them
# to the baked snapshot's values, so passing them reads like a lie in your code.
# This guest did not boot. It woke up holding the state the template froze --
# which is also why /proc/uptime is a liar and the clock needs re-syncing.
print(sbx.exec("cat /proc/uptime").stdout)
# Credit where it is due: Docker ENV from the template's Dockerfile does NOT
# reach a detached process in our guests -- only /etc/environment, via PAM,
# does. firecracker-containerd is genuinely better on this point, because a real
# runc in the guest applies the image config, so ENV and ENTRYPOINT mean exactly
# what the image says they mean.
r = sbx.exec("printenv MY_VAR || echo 'not set -- export it yourself'")
print(r.stdout, r.exit_code)
# timeout_seconds is a CLIENT deadline. Neither exec endpoint enforces it
# server-side, so a hard limit belongs in the shell, next to ulimit.
sbx.exec("timeout 30 ./long-thing.sh", timeout_seconds=60)
sbx.kill()And the bill for that trade, itemised honestly: you build a template first, so there is a build step in your onboarding that firecracker-containerd does not have. RAM is chosen at bake time and cannot be changed at restore, because Firecracker cannot resize guest memory or vCPU count when loading a snapshot — which is why our `base` template is 4 GiB, `browser` is 4 GiB, `code-interpreter` and `agent` are 2 GiB, and `postgres-16` is 1 GiB, with 8 burstable vCPUs each. Copy-on-write for the rootfs is local-only, because reflink and dm-snapshot need a local block device; we can stream a snapshot's memory on demand from object storage over userfaultfd, but never the disk. And the guest kernel is 5.10, one per host, not swappable per template — so if your workload needs a newer kernel's netfilter, that is a wall and not a config change.
Where I would draw the line
If your packaging is the constraint — you have images, you have a registry, you have CRI, and re-packaging is politically or practically impossible — then the OCI-native runtimes are correct and you should be choosing between firecracker-containerd and Kata on the basis of how much you need Kubernetes to be a first-class citizen. Kata is the standards-track answer with the broader hypervisor and device story; firecracker-containerd is the smaller, more direct one if containerd rather than Kubernetes is your actual interface.
If your latency is the constraint — if your product is “give me an isolated environment in the time it takes to make an HTTP request” — then the chain “pull, convert, boot, start” has a floor you cannot optimise your way under, and you want something that starts from a memory image. That is the bet I made, and the price of the bet is that I cannot run your arbitrary image on demand.
Both of those are defensible engineering positions. The failure mode I see most often is picking one of them for a workload shaped like the other, then spending a quarter tuning the wrong chain. Count the steps on your own critical path before you pick a runtime, and be honest about which step you will be doing ten thousand times a day.
Frequently asked questions
Does firecracker-containerd give me Kubernetes pods in microVMs?
Not directly, no. firecracker-containerd operates at the containerd level and has its own VM lifecycle API through the firecracker-control plugin, which is a layer below where Kubernetes expects to plug in. Kubernetes talks to a container runtime over CRI and selects runtimes via RuntimeClass, and the project that is built for that path is Kata Containers — it is the standards-track CRI and OCI VM runtime, it exposes itself as a RuntimeClass, and it can use Firecracker as one of its supported hypervisors as well as others. So if the goal is literally VM-isolated pods in an existing cluster, start with Kata and read its current documentation on which hypervisors and which Kubernetes versions it supports, because that matrix changes. firecracker-containerd is the better fit when containerd itself is your interface — a build service, a job runner, a function platform you operate directly — rather than when the kubelet is the thing doing the asking.
Do I need a separate guest kernel for each container image?
No, and the question itself reveals the most common misconception about running images in VMs. Container images do not contain kernels. An `alpine` image is a userspace: a libc, a shell, some binaries. It has always borrowed the host's kernel, which is exactly why a container is a weaker boundary than a VM. When you run that image in a microVM, the kernel comes from the runtime's configuration — `kernel_image_path` in firecracker-containerd's runtime config — and one kernel serves every image you run through it. What you do need is a kernel configured for the machine Firecracker presents: virtio over MMIO rather than a PCI bus to enumerate, and the 8250 serial console, which is your only diagnostic channel if the boot fails before the in-guest agent comes up. Start from Firecracker's own guest configs rather than `defconfig`. On PandaStack the same constraint exists in a harder form: the guest kernel is 5.10, there is one per host, and you cannot swap it per template.
Is running a container in a microVM faster than running it in a container?
No, and anyone telling you otherwise is comparing different things. Against runc on a warm host with the image already local, firecracker-containerd is strictly slower: you have added a VM boot, a snapshotter conversion, a device attach and a vsock round-trip to a path that previously consisted of unpacking layers and cloning namespaces. The thing you bought is not speed, it is a hardware-virtualisation boundary between the workload and your host kernel, and that is worth paying for whenever the workload is something you would not run as a trusted process. Separately, do not confuse “microVM” with “fast” in general. Firecracker boots very quickly for a VM, but a boot is still a boot. The genuinely fast paths in this space are the ones that skip booting altogether by restoring a snapshot of a guest that already booted, and that is an architectural choice about your artifact, not a property of the hypervisor.
What is the operational burden people underestimate?
Four things, consistently. First, the device-mapper thin-pool: it needs capacity planning for both its data and its metadata volumes, active monitoring of both, and a documented response for when either fills, because exhaustion there does not fail as cleanly as a full overlay filesystem. Second, the artifact sprawl: you now own a guest kernel and a VM root filesystem carrying the agent, in addition to your application images, and those two artifacts must be upgraded in step with the shim. Third, networking: a tap device per VM integrated through a CNI chain, which means your runtime config and your CNI config are coupled and your debugging happens inside network namespaces you did not create. Fourth, observability across the vsock boundary — when a workload hangs, distinguishing the container from the agent from the shim from containerd requires instrumentation at more than one layer, and the serial console is a poor substitute for it.
Why can't PandaStack just run any OCI image I give it?
Because of the trade I made deliberately, and I would rather state it than let you discover it. An OCI image is a layer stack that has to be converted into something a guest can mount, and I pay that conversion once at template build time instead of once per create: the image is flattened into an ext4 rootfs, cold-booted once in about three seconds, and snapshotted. Every create afterwards is a reflink of the disk plus a restore of the memory image, which is how the create path lands at a p50 of 179 ms. Accepting an arbitrary image at create time would put the pull and the conversion back on that path and give up the thing the whole platform is built around. It also would not fix the related constraint: Firecracker cannot change guest RAM or vCPU count when loading a snapshot, so memory is a property of the baked template rather than of the request. If arbitrary images at request time is a hard requirement for you, firecracker-containerd or Kata is the honest recommendation.
Keep reading
- Firecracker vs Kata vs gVisor — The three-way comparison in full — what each boundary actually is, and which one your threat model needs.
- Snapshot restore vs container image pull — The latency argument from this post, taken apart step by step on both chains.
- Golden images vs snapshot baking — Why baking a booted guest beats baking a disk, and what the artifact costs you.
- Firecracker orchestration tools — The wider field: go-sdk, firecracker-containerd, Kata, Ignite, flintlock, and building it yourself.
- Firecracker vs runc — The shared-kernel baseline this whole post is trying to get away from.
Related posts
- Best Open-Source Sandboxes for Running Untrusted Code
The honest set of open-source, self-hostable options for isolating untrusted code execution — characterized by license and isolation model, not a leaderboard.
- Kata Containers vs gVisor: the two secure-container runtimes
Two runtimes both promise "container UX, stronger isolation" and reach it by opposite routes: Kata puts a real hardware VM under your container, gVisor re-implements Linux in Go so it can say 'no' to your syscalls more politely. Here's the honest head-to-head.
- virtiofs vs virtio-blk: How Files Actually Get Into a MicroVM
One gives the guest a disk it owns. The other gives it a window onto a directory the host owns. That single difference decides whether you can fork a machine in 400ms, who parses guest-controlled input, and what a multi-tenant escape looks like.
- Kata Containers vs Firecracker: Honest Head-to-Head
The framing is slightly off: Kata Containers is an OCI runtime that can run ON Firecracker. One gives you Kubernetes-shaped ergonomics, the other is the minimal VMM doing the isolating. Here's the honest comparison.
- Firecracker vs crun vs youki: a VMM vs OCI runtimes
"Firecracker vs crun vs youki" compares a hypervisor to two OCI container runtimes — a category mismatch worth explaining. crun (C) and youki (Rust) make faster, safer runc-equivalents, but they don't change the isolation boundary: your one shared host kernel.
More in Firecracker & microVMs · See Firecracker microVM sandboxes
49ms p50 cold start. Fork, snapshot, and scale to zero.