all posts

Firecracker vs FreeBSD Jails: Two Honest Answers to the Same Question

Ajay Kumar··10 min read

FreeBSD jails are better designed than Linux containers. I want to open with that, because I build a Firecracker microVM platform for a living and the usual version of this post is a strawman: somebody lines up a twenty-five-year-old BSD feature against a hypervisor, declares the hypervisor safer, and goes home. That comparison is both true and useless, and it skips the part where jails got a lot of things right that Linux is still assembling out of seven separate subsystems.

I'm Ajay. I built PandaStack on Firecracker, so my bias is obvious and I would rather name it than pretend. What follows is the comparison I would want if I were running a FreeBSD fleet and someone was trying to sell me a hypervisor: what a jail actually is, where it is genuinely the better tool, and the one axis on which a microVM is not negotiable.

What a jail actually is

`jail(2)` landed in FreeBSD 4.0 in 2000, roughly a decade before Linux containers were a thing people deployed. The crucial design difference is not age, it is coherence. A jail is a first-class kernel object — a `struct prison`, declared in `sys/sys/jail.h`, reference-counted, with a name, an ID, and a parent. It was designed once, as one feature, by people who were trying to solve one problem: give a hosting customer root without giving them the machine.

Linux arrived at a similar place by a completely different route. There is no `struct container` in the Linux kernel. There is a pile: seven-odd namespace types, cgroups v1 and v2, capabilities, seccomp-bpf, LSMs, and a userspace runtime whose job is to assemble all of it correctly on every single process spawn. It works, it is enormously flexible, and it is also a construction kit where forgetting one piece is a vulnerability rather than an error message. A jail, by contrast, is a thing you either are in or are not in, and the kernel knows which.

The model has real depth. Jails are hierarchical — a jail can create child jails, and a child can never be granted a permission its parent lacks, which gives you nesting with a monotonic security property rather than nesting with a prayer. Permissions are a set of `allow.*` parameters, every one of which defaults to off: `allow.mount`, `allow.raw_sockets`, `allow.sysvipc`, `allow.chflags`, `allow.socket_af`, and friends. Each jail has its own `securelevel`, which may be raised above the host's but never lowered below it. `VNET` gives a jail an entire independent network stack — its own interfaces, routing table, firewall state, not a filtered view of the host's — on a kernel built with `VIMAGE`, which is standard on current releases. `RCTL`, on top of `RACCT`, gives you per-jail limits on memory, process count, thread count, open files and CPU.

That is a complete, internally consistent compartmentalisation story that has had one coherent permission model for twenty-five years. Over the same period Linux has been finding and fixing user-namespace-reachable privilege escalations at a steady clip: CVE-2022-0492 via the cgroup v1 `release_agent`, CVE-2022-0185 via a heap overflow in filesystem-context parsing, CVE-2024-1086 via a use-after-free in nf_tables. Those are not exotic research curiosities; they are escapes from exactly the configuration that unprivileged containers rely on. Anyone who tells you jails are the sloppy option has not read `jail(8)`.

jail.conf, ZFS, and the tooling

The operational experience is also better than its reputation. A jail is declared in `/etc/jail.conf`, started with `service jail start <name>`, listed with `jls`, and entered with `jexec`. If you want something more opinionated there is `bastille`, `iocage`, `cbsd` or `pot`, all of which wrap the same kernel primitives with templates, networking helpers and ZFS integration.

And ZFS is the part Linux people should look at honestly, because FreeBSD has had the storage layer for this since before anyone was arguing about layer counts. A jail's whole world is a dataset. `zfs snapshot` plus `zfs clone` gives you a new jail root in metadata time, copy-on-write, deduplicated against the parent — which is to say that jails had a first-class CoW story years before Linux container people rebuilt a worse version of it out of overlayfs and union mounts. Our own disk layer is the same idea reached independently: XFS reflink or dm-snapshot instead of ZFS clones.

# /etc/jail.conf — shared defaults first, then one block per jail.
# Start: `service jail start web`.  List: `jls`.  Enter: `jexec web sh`.

exec.start    = "/bin/sh /etc/rc";          # run the jail's OWN rc(8)
exec.stop     = "/bin/sh /etc/rc.shutdown";
exec.clean;                                 # scrub the inherited environment
mount.devfs;                                # mount a FILTERED /dev inside
devfs_ruleset = 4;                          # ruleset 4 = the stock jail set
persist;

web {
    path          = "/jails/web";           # root of this jail's ZFS dataset
    host.hostname = "web.internal";

    # --- VNET: an independent network stack, not a filtered view of the
    # host's. Needs a kernel with VIMAGE (standard on current releases).
    vnet;
    vnet.interface = "epair0b";             # the jail-side end of an epair(4)
    exec.prestart  = "ifconfig epair0 create && ifconfig bridge0 addm epair0a up";
    exec.poststop  = "ifconfig epair0a destroy";

    # --- Permissions. Every allow.* bit defaults to OFF; you opt in one at a
    # time. Spelled out here only to make the model explicit.
    allow.nomount;          # no mount(2) at all — historically the best lever
    allow.noraw_sockets;    # no raw packet crafting
    allow.nosysvipc;        # no SysV IPC shared with the host or siblings
    allow.nochflags;        # cannot clear schg on host-immutable files
    securelevel = 3;        # immutable flags locked, even for jail root
}

# ------------------------------------------------------------------
# /etc/rctl.conf — limits live OUTSIDE jail.conf, keyed per jail.
# RACCT/RCTL ship compiled in but DISABLED: set kern.racct.enable=1 in
# /boot/loader.conf and reboot, or none of these rules do anything.

jail:web:memoryuse:deny=2g
jail:web:vmemoryuse:deny=4g
jail:web:maxproc:deny=200
jail:web:openfiles:deny=4096
jail:web:nthr:deny=400
jail:web:pcpu:log=80   # Not every resource supports every action. Check
                       # rctl(8) on YOUR release for which of deny/log/
                       # devctl/throttle apply to %CPU before relying on it.

# Live usage for one jail:  rctl -hu jail:web
Everything version-dependent here — VIMAGE in GENERIC, which rctl actions a resource accepts, devfs ruleset numbering — should be checked against the man pages for the release you actually run. FreeBSD moves these details, and a blog post is a worse source than `man jail.conf` on the box in front of you.

The one axis that matters

Here it is, once, plainly: a jail shares the kernel; a microVM does not.

Everything else in this post is downstream of that sentence. When a process inside a jail makes a syscall, the host's kernel executes it — the same kernel serving every other jail and the host itself. The jail boundary is a set of checks inside that kernel: is this PID in your prison, is this mount visible to you, does this `allow.*` bit permit the operation. The checks are good. They are also checks, and a check only runs in the code path someone remembered to put it in.

A Firecracker guest makes its syscalls to its own kernel, loaded from its own image, running under KVM with the CPU's virtualisation extensions enforcing the boundary. To reach the host, an attacker has to own the guest kernel first — which buys them root on a disposable machine that is about to be deleted — and then find a bug in the VMM.

A shared-kernel boundary is a promise that every syscall in a kernel of tens of millions of lines will behave correctly for a caller who is actively trying to make it misbehave. A container is a polite suggestion to the kernel. A jail is a much better-worded suggestion to a much better-organised kernel. It is still a suggestion.

Syscall surface vs device surface

The asymmetry is in what each boundary has to defend. A jail's attack surface is the full kernel ABI: every syscall, every `ioctl`, every `sysctl`, every filesystem the kernel can parse, reachable by an attacker who is already inside and already root in their own prison.

Firecracker's guest-facing surface is a short device list: virtio-net, virtio-block, virtio-vsock, virtio-balloon, virtio-rng, pmem, a serial port, an i8042 controller. That is the whole menu a hostile guest gets to order from. Two corrections to the stale version of this claim, because they get repeated everywhere: current Firecracker does have an opt-in virtio-PCI transport behind `--enable-pci` (with hot-plug in developer preview), and it does support ACPI. MMIO remains the default transport and our guests still boot `pci=off`. The device model is small, not nonexistent, and the right statement is that it is deliberately narrow rather than that it is magic.

Underneath that, Firecracker runs itself the way you would run anything you did not fully trust. The jailer puts the VMM process in a chroot, pivots root, drops capabilities, joins a pre-created network namespace, and applies cgroup limits. A seccomp-bpf allowlist then restricts the VMM to the small set of syscalls it genuinely needs, so a successful virtio exploit lands in a process that cannot do very much. There is a pleasing symmetry in the naming: FreeBSD's jail is the boundary, and Firecracker's jailer is the boundary around the thing that implements the boundary.

What each does on a bad day

On a bad day in a jail, a local privilege escalation in the host kernel is an escape affecting every jail on the box and the host, and your mitigation is patching quickly. Historically jails have had far fewer reported escapes than Linux containers, and some of that is genuinely the better design — but some of it is that vastly fewer people are looking. Absence of published escapes in a small-attention-surface OS is weaker evidence than it feels like.

On a bad day in a microVM, the guest kernel falls over and a VM that was going to be deleted anyway gets deleted sooner. To matter, the attacker needs a guest-kernel bug and a VMM bug and a way past seccomp and the jailer, and then they are in a chrooted, capability-stripped process in its own namespace. That is not invulnerable. It is several independent things going wrong instead of one.

Performance and density, honestly

This is where jails win and it is not close, so I am not going to dress it up. A jail has no hypervisor tax. There is no guest kernel to boot, no second page cache, no virtio round-trips, no vCPU threads. Starting a jail is forking a process into a new prison and running `rc`; the kernel-side cost is in the milliseconds. Memory is shared with the host at the page-cache level, so a hundred jails running the same binaries share those pages for free, and RCTL accounting measures what each one actually used rather than what it was promised.

A microVM pays for all of that. Each guest has its own kernel, its own page cache, its own idea of how much RAM exists. Our honest answer to the boot cost is not that booting is free — it very much is not, about three seconds for a template's first cold boot — but that you only have to do it once. We boot a template, snapshot the running machine, and every subsequent create restores that snapshot instead of booting: p50 179 ms, p99 203 ms, with the `/snapshot/load` step itself somewhere around 49 to 80 ms. That is a different trick from a jail's millisecond start, not a match for it.

On the memory side we do what we can rather than what we would like. Guest RAM is fixed at template build time, because Firecracker cannot change vCPU count or guest memory at snapshot restore — pass `memory_mb` on a create and it is silently corrected to the baked value. Our templates are sized accordingly: 4 GiB for `base`, 2 GiB for `code-interpreter` and `agent`, 4 GiB for `browser`, 1 GiB for `postgres-16`, each with eight burstable vCPUs. The mitigations are real but indirect: snapshot memory can be paged in on demand from object storage over userfaultfd in 4 MiB chunks, with a header recording which chunks are non-zero so all-zero pages cost no fetch at all, and a prefetch trace that warms the hot set. Clever, and still not as cheap as sharing the host's page cache, because nothing is.

If the code you are isolating is code you wrote, a microVM is overhead you are paying for a threat you do not have. Compartmentalising your own services on a FreeBSD box is what jails are for, and buying a hypervisor to do it is a worse answer. I would rather tell you that than sell you a VM you do not need.

Operations, side by side

FreeBSD jails vs Firecracker microVMs, by axis.
AxisFreeBSD jailFirecracker microVM
Isolation boundaryChecks inside the shared host kernelKVM plus a separate guest kernel per VM
Guest-facing attack surfaceThe full kernel ABI — every syscall and ioctlA short virtio device list, behind seccomp and the jailer
Permission modelOne coherent allow.* set, stable for 25 yearsA pile: seccomp, caps, cgroups, netns, chroot — assembled per process
Start timeMilliseconds; no kernel to bootp50 179 ms from a snapshot; ~3 s for a template's first cold boot
Memory costShares the host kernel and page cacheOwn kernel and page cache; RAM fixed at bake time
Resource limitsRCTL/RACCT rules in /etc/rctl.confcgroups on the VMM process plus the baked guest size
NetworkVNET — a real independent stack per jail (needs VIMAGE)netns plus veth plus tap per sandbox; 16,384 pre-allocated /30s per host
Storage copy-on-writeZFS snapshot plus cloneXFS reflink or dm-snapshot — local disk only, CoW needs a local device
Snapshot of a RUNNING instanceNo mainstream equivalent; ZFS captures the disk, not the processesYes — memory plus device state, which is what fork and branch are built on
Kernel choice per tenantNone; one host kernel, though older userlands run under compat ABIsA kernel image per guest in principle; on PandaStack one 5.10 build per host
Host OSFreeBSD onlyLinux with KVM only
Ecosystem fitports and pkg; Linux binaries need the compat layerLinux userspace unmodified — the rootfs comes out of a Dockerfile build

The row worth staring at is the running-instance snapshot. A ZFS clone of a jail gives you the filesystem at an instant; the clone then boots `rc` like a fresh machine. There is no resume-where-it-was, because the processes were never captured. A VM snapshot includes the vCPU registers, the device state and the guest's RAM page for page, so a restore continues mid-syscall. That is not a nicer version of a ZFS clone, it is a different capability, and it is the one that makes fork-and-branch workflows possible at all.

# --- FreeBSD: a jail's whole world is a dataset, so "clone it" is one op.
zfs snapshot zroot/jails/web@pre-upgrade
zfs clone    zroot/jails/web@pre-upgrade zroot/jails/web-test   # O(metadata)
# ...then point a new jail.conf block at /jails/web-test and start it.
# What you did NOT copy: the running processes. The clone boots from rc(8).

# --- Linux side: the disk half is the same idea, a different primitive.
cp --reflink=always rootfs.ext4 clone.ext4    # XFS/Btrfs reflink, O(metadata)

# The difference is the second and third files a VM snapshot also has:
#   vm.state  -> vCPU registers and virtio device state
#   vm.mem    -> the guest's RAM, page for page
# Restore all three and the machine continues mid-syscall. Same-host fork of a
# live sandbox: 400-750 ms. Cross-host, pulling the memory image: 1.2-3.5 s.
#
# The honest gotcha: a fork gets its parent's RNG state and clock EXACTLY.
# Two forks of one parent will generate the same "random" bytes and both think
# it is still the moment the snapshot was taken. Reseed, and re-sync the clock.

The elephant in the room

Jails are FreeBSD-only. That is usually the decisive factor, it has nothing to do with which design is better, and pretending otherwise wastes everybody's afternoon.

Your base images are Dockerfiles. Your binaries are linked against glibc. Your GPU stack, if you have one, is a Linux kernel driver. Your observability vendor ships a Linux agent, your security vendor ships a Linux agent, and your team's entire muscle memory — `journalctl`, `ip`, `nsenter`, `systemd` unit files, `strace` — is Linux muscle memory. Adopting jails means adopting FreeBSD, which means re-platforming everything above the kernel, in exchange for a boundary that is still a shared kernel. That is a very large bill for a non-answer to the untrusted-code problem.

FreeBSD's Linux ABI compatibility layer — the linuxulator, the `linux` and `linux64` modules plus a `/compat/linux` userland from the Linux-base ports — is better than people expect and worth knowing about. It is syscall translation, not emulation, so statically or dynamically linked Linux CLI tools and plenty of compiled server software genuinely run. What it does not give you is anything that depends on Linux kernel interfaces as interfaces: cgroups, namespaces, eBPF, io_uring, the finer corners of `/proc` and `/sys`. You are not running Docker under it. And from a security standpoint, note the direction of travel: running Linux binaries in a jail means the FreeBSD kernel is now implementing a second syscall ABI for an attacker to poke at. That is more surface in the shared kernel, not less.

If you are on FreeBSD and you want a VM-grade boundary, the answer is `bhyve`, not Firecracker. Firecracker requires Linux and KVM; it does not run on a FreeBSD host, and no amount of enthusiasm changes that.

Which one to pick

  • Compartmentalising your own software on a FreeBSD box — jail, clearly. It is the native, cheapest, best-designed answer and a hypervisor is dead weight.
  • A FreeBSD hosting business selling root to customers who are not actively hostile — jail. This is literally the problem jail(2) was written for, and the economics have always worked.
  • Many instances of trusted workloads where density is the binding constraint — jail. No hypervisor tax, shared page cache, millisecond starts. Nothing with a guest kernel competes.
  • Running code you did not write — a customer's, a model's, a stranger's pull request — microVM. The shared kernel IS the threat model, and no allow.* bit list fixes a privilege-escalation bug in the syscall you forgot existed.
  • Multi-tenant platform where tenants are mutually hostile — microVM, or bhyve if you are staying on FreeBSD. One kernel bug cannot be the whole blast radius.
  • You need to snapshot a running machine and fork it — microVM. A jail has no equivalent; a ZFS clone is the disk only.
  • You need the Linux ecosystem as-is (Docker-built rootfs, glibc binaries, vendor agents) — microVM, by elimination rather than by merit.

What that last category looks like on our side is unremarkable, which is the point. One sandbox per untrusted job, a TTL so the platform reaps it even if your orchestrator dies, hard limits in the shell rather than in a hopeful API parameter, and an explicit kill.

from pandastack import Sandbox

# One sandbox per untrusted job. The API is not what you are buying here --
# `jexec` is a perfectly nice API too. You are buying a different boundary.
sbx = Sandbox.create(
    template="code-interpreter",          # 2 GiB RAM, 8 burstable vCPUs (baked)
    ttl_seconds=600,                      # reaped even if this process dies
    metadata={"job": "pr-4812", "trust": "none"},
)

try:
    # timeout_seconds is a CLIENT deadline. Neither exec endpoint enforces it
    # server-side, so it stops YOU waiting -- it does not stop anything inside
    # the guest. Hard limits belong in the shell, exactly as they would in a
    # jail. The difference is what is on the other side of the limit.
    r = sbx.exec(
        "timeout -s KILL 120 sh -c '"
        "ulimit -v 1048576 -t 100 -u 256; "   # address space, CPU secs, procs
        "python /work/job.py'",
        timeout_seconds=150,
    )
    print(r.exit_code)
    print(r.stdout[-2000:], r.stderr[-2000:])
finally:
    sbx.kill()   # not "the jail is stopped" -- the machine is gone

If I were starting a FreeBSD shop today I would use jails for everything I wrote myself and I would not feel like I was settling. The design is cleaner than Linux's, the ZFS integration is a genuine advantage, and the permission model has been coherent for longer than some of my colleagues have been programming. I would also not put a stranger's code in one, and neither would the people who wrote `jail(8)` — because the question a jail answers brilliantly is 'how do I keep my own services out of each other's way', and the question a microVM answers is 'what happens when the thing inside is trying to get out'. Two honest answers. Two different questions.

Frequently asked questions

Are FreeBSD jails more secure than Docker containers?

Better designed, same class of boundary. A jail is a single first-class kernel object with a default-deny permission set, hierarchical nesting where a child can never exceed its parent, and a per-jail securelevel — versus Docker, which composes namespaces, cgroups, capabilities, seccomp and an LSM profile per container and depends on the runtime getting every piece right. Fewer moving parts means fewer configuration mistakes, and configuration mistakes are the most common real-world container compromise, so in practice jails are meaningfully harder to misconfigure into uselessness. But both execute untrusted syscalls in the host kernel, so both inherit every local privilege escalation in that kernel as a potential escape. Jails have had far fewer published escapes, and some of that is the design while some of it is that the FreeBSD kernel gets a tiny fraction of the offensive research attention Linux gets. Treat 'better engineered' and 'a different class of boundary' as separate claims, because only the first one is true.

Can I run Linux binaries inside a FreeBSD jail?

Often yes, with real caveats. FreeBSD's Linux ABI layer — the `linux` and `linux64` kernel modules plus a `/compat/linux` userland from the Linux-base ports — translates Linux syscalls to FreeBSD ones. It is translation, not emulation, so there is no interpreter tax, and a large amount of ordinary compiled software and CLI tooling simply works. What does not work is anything that treats Linux kernel interfaces as the product: cgroups, namespaces, eBPF, io_uring, and the deeper corners of `/proc` and `/sys`. You are not running Docker, Kubernetes components, or most container runtimes under it. Coverage is also genuinely version-dependent and has improved a lot over recent releases, so check the handbook and `linux(4)` for the release you run rather than any blog post. One security note people miss: enabling the layer means the FreeBSD kernel is now implementing a second syscall ABI that an attacker inside the jail can reach. That is additional surface in the kernel you are sharing.

Can I snapshot and clone a running jail the way I would a VM?

Not the running part. ZFS gives you an excellent snapshot and clone of the jail's dataset in metadata time, and that is a real operational superpower for upgrades and rollbacks — but it captures the filesystem, not the processes. Start a clone and it runs `rc` from scratch like a fresh machine; there is no resume mid-execution, because the memory and the kernel state belonging to those processes were never part of the snapshot and live in a kernel shared with everyone else. There is no mainstream production equivalent of CRIU for jails. A VM snapshot is categorically different: it contains vCPU registers, virtio device state and the guest's RAM page for page, so a restore continues exactly where it paused. That is what makes fork-and-branch workflows possible — on our platform a same-host fork of a live sandbox runs 400 to 750 ms, or 1.2 to 3.5 seconds cross-host when the memory image has to be pulled. If you want that capability on FreeBSD, you want bhyve, not a jail.

Isn't a microVM per job wasteful compared to a jail?

Yes, measurably, and the question is whether you are buying anything with the waste. A jail has no hypervisor tax, no guest kernel, no second page cache, and starts in milliseconds; a hundred jails running the same binaries share those pages with the host for free. A microVM pays for its own kernel and its own page cache, and on Firecracker the guest's RAM is fixed when the template is built — it cannot be changed at snapshot restore, so passing `memory_mb` on a create is silently corrected to the baked size. Our mitigations narrow the gap without closing it: restore from a snapshot rather than boot, so a create costs p50 179 ms instead of about three seconds, and page guest memory in on demand from object storage over userfaultfd in 4 MiB chunks, with a header marking which chunks are non-zero so all-zero pages cost no fetch. If your workload is trusted, that overhead buys nothing and a jail is the right call.

I'm on FreeBSD and need VM-grade isolation. Can I use Firecracker?

No. Firecracker is a Linux-and-KVM program — it needs `/dev/kvm` and a set of Linux-specific interfaces, and it does not run on a FreeBSD host. The native answer is `bhyve`, FreeBSD's own type-2 hypervisor, which gives you the property you actually care about: each guest runs its own kernel, and a guest-to-host escape requires a hypervisor bug rather than a kernel bug. bhyve is a more conventional VMM than Firecracker — a broader device model, a general-purpose feature set — so expect a different performance and density profile, and expect to build the snapshot and fast-create machinery yourself if you need it rather than finding it ready-made. A reasonable FreeBSD architecture is both: jails for your own services, where they are the best tool available on any operating system, and bhyve guests for anything running code you did not write. Check current bhyve documentation for snapshot support on your release, since that area has moved.

Keep reading

Related posts

More in Firecracker & microVMs · See Firecracker microVM sandboxes

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.