all posts

Best bare metal providers for Firecracker in 2026

Ajay Kumar··11 min read

Firecracker needs /dev/kvm. Not a Linux box with enough RAM, not a container runtime, not a generous CPU allowance -- a character device at /dev/kvm that your process can open. There is no software fallback, no emulation mode, no flag that makes it work anyway. That single requirement quietly disqualifies most of the cheap compute on the market, including nearly every instance in every "best value VPS" list you have open in another tab.

So this post is a shortlist, not a ranking. I have provisioned Firecracker hosts across more provider families than I would choose to again, and the pattern is always the same: the decision is made on price per core, and then two weeks are spent discovering that the chosen instance family cannot expose hardware virtualization, that the root filesystem cannot do reflink, or that the machine bills monthly when the workload is spiky. I am Ajay; I build PandaStack, a Firecracker microVM sandbox platform whose agents run on Linux KVM hosts, so these are my own prerequisites rather than a survey of someone else's.

Ground rules. No prices appear in this post, deliberately: they change faster than I can update a blog, and a stale price is worse than no price. Instance families, nested-virtualization support, available regions and the shape of local storage all move on a timescale of months, so every provider claim below carries the same instruction -- verify it against that vendor's current documentation before you commit. One of the providers in this post has already shut its bare-metal product down, which is exactly the point: a 2026 roundup that nobody re-checked is a liability.

The requirement that decides your shortlist

Firecracker is not a hypervisor in the way VMware is a hypervisor. It is a small userspace virtual machine monitor that asks KVM, in the host kernel, to do the actual virtualization. Hardware virtualization extensions, the kvm module, and the device file are therefore not optional dependencies -- they are the product. Firecracker runs on x86_64 and aarch64 Linux hosts with KVM, and nowhere else.

There are exactly three ways to have that device. You own the metal, so KVM talks straight to the CPU's virtualization extensions. Or you are on a cloud VM whose hypervisor deliberately exposes those extensions to the guest, which is nested virtualization. Or you are on a laptop hypervisor that does the same thing -- Apple's Virtualization.framework on recent Apple Silicon, which is how a Mac developer runs a real Firecracker host locally.

Three failure modes cover almost every "it does not work" ticket. The CPU flags are present but /dev/kvm is missing, which usually means the module is not loaded or the provider does not expose the feature at all. The device exists but your service user cannot open it, which is a udev and group problem that produces an error message pointing nowhere useful. Or everything works and nobody noticed the host is itself a guest, so the capacity model, the support story and the per-exit cost are all quietly wrong. Check for all three before you believe a host.

Verify KVM in 60 seconds

Run this on a trial instance before you build an image, and on every new host before it joins the fleet. It is cheap, it is boring, and it has saved me from two multi-week commitments.

#!/usr/bin/env bash
# kvm-check.sh -- run this on a candidate host BEFORE you install anything and
# certainly before you sign a twelve-month term. Six checks, about a minute,
# non-zero exit if the host cannot run Firecracker.
set -uo pipefail
fail=0
note() { printf '%-26s %s\n' "$1" "$2"; }

# 1. CPU flags. x86 only: Intel reports vmx, AMD reports svm. On arm64 there is
#    no flag to grep -- virtualization lives in EL2 -- so an empty result there
#    is not a failure, and the device check below is the real test.
lscpu | grep -iE '^(Architecture|Virtualization|Hypervisor vendor)' || true
if [ "$(uname -m)" = x86_64 ]; then
  grep -qE '^flags.*\b(vmx|svm)\b' /proc/cpuinfo \
    && note "cpu flags" "ok" \
    || { note "cpu flags" "MISSING vmx/svm"; fail=1; }
fi

# 2. The device. This is the actual requirement, and the permissions are part of
#    it: a /dev/kvm that root can open and your agent's service user cannot is
#    the same outage with a more confusing error message.
if [ -e /dev/kvm ]; then
  ls -l /dev/kvm
  { [ -r /dev/kvm ] && [ -w /dev/kvm ]; } \
    && note "/dev/kvm rw as $(id -un)" "ok" \
    || note "/dev/kvm" "present, not rw for $(id -un) -- fix the group + udev rule"
else
  # Worth one modprobe before declaring defeat: on a fresh image the module is
  # frequently just not loaded yet. If this fixes it, persist it in
  # /etc/modules-load.d or you will rediscover it after the next reboot.
  sudo modprobe kvm_intel 2>/dev/null || sudo modprobe kvm_amd 2>/dev/null || true
  [ -e /dev/kvm ] \
    && note "/dev/kvm" "appeared after modprobe -- persist the module" \
    || { note "/dev/kvm" "ABSENT"; fail=1; }
fi

# 3. kvm-ok (Debian/Ubuntu package cpu-checker) distinguishes "this CPU cannot"
#    from "this CPU can but it is disabled in firmware". On a freshly racked
#    dedicated server the second one is common, and it is a support ticket
#    rather than a design problem.
command -v kvm-ok >/dev/null || sudo apt-get install -y cpu-checker >/dev/null 2>&1 || true
command -v kvm-ok >/dev/null && kvm-ok

# 4. Are we nested? Not fatal. It must simply not be a surprise six weeks later,
#    because it changes your per-exit cost, your support story and who you call
#    when a guest wedges.
if systemd-detect-virt -q; then
  note "nesting" "this host is itself a guest ($(systemd-detect-virt)); your microVMs are L2"
else
  note "nesting" "bare metal"
fi

# 5. Firecracker itself, and then proof rather than inference. --version shows
#    the binary runs; pin that version while you are looking at it, because
#    snapshots are version-bound and a fleet running two builds cannot restore
#    each other's snapshots. Then ask KVM directly: open the device, read its API
#    version, and actually create a VM. If KVM_CREATE_VM hands back a file
#    descriptor, this host can run Firecracker and nothing else needs believing.
firecracker --version || { note firecracker "not installed"; fail=1; }
python3 - <<'KVMPROBE' || fail=1
import fcntl, os, sys
KVM_GET_API_VERSION = 0xAE00   # _IO(KVMIO, 0x00)
KVM_CREATE_VM       = 0xAE01   # _IO(KVMIO, 0x01)
try:
    fd = os.open("/dev/kvm", os.O_RDWR | os.O_CLOEXEC)
except OSError as e:
    print(f"open /dev/kvm failed: {e}"); sys.exit(1)
print("kvm api version", fcntl.ioctl(fd, KVM_GET_API_VERSION, 0), "(expect 12)")
vm = fcntl.ioctl(fd, KVM_CREATE_VM, 0)
print("KVM_CREATE_VM returned fd", vm, "-- this host can run Firecracker")
os.close(vm); os.close(fd)
KVMPROBE

# 6. 2 MiB hugepages, only if you intend to back guest memory with them. Two
#    different questions: can the sysctl be SET, and can the kernel actually
#    ASSEMBLE a 2 MiB page right now. The last line is the one that bites. In
#    /proc/buddyinfo the free-run counts start at field 5 (order 0), so order 9
#    -- a 2 MiB contiguous run on 4 KiB pages -- is field 14. That number
#    collapses toward zero on a host that has done real work, however much free
#    RAM /proc/meminfo cheerfully reports.
sudo sysctl -q -w vm.nr_overcommit_hugepages=64 \
  && note "overcommit sysctl" "settable" || note "overcommit sysctl" "REFUSED"
note "hugepages total/free" "$(awk '/HugePages_(Total|Free)/{printf "%s ", $2}' /proc/meminfo)"
note "free 2 MiB runs" "$(awk '$4=="Normal"{print $14}' /proc/buddyinfo | paste -sd, -)"

exit "$fail"

Read the results together rather than one at a time. Flags present and device absent means the module or the provider, in that order. Device present but not readable by your service account means you will see a permission error at VM start that reads like a Firecracker bug. A non-empty "Hypervisor vendor" line with a working device means nested virtualization is on, which is a perfectly legitimate place to develop and a questionable place to run a fleet. And on arm64 the flag grep finds nothing even on healthy metal, so trust the device.

Tier 1: nested virtualization on a cloud VM

Nesting stacks hypervisors. Your provider's hypervisor is L0, your VM is L1, and the microVM you boot inside it is L2. It works, and I want to be careful not to sneer at it, because PandaStack runs on nested virtualization in production and the create latencies we publish -- p50 179 ms, p99 around 203 ms for a snapshot-restore create -- are measured on nested hosts. What nesting costs you is paid on VM exits that L1's KVM cannot resolve alone and must trap down to L0, plus a second CPU scheduler between your guest's vCPU thread and a physical core. The mechanics are their own post; the shopping decision is below.

The other cost is support posture. On most providers nested virtualization is a documented but lightly trodden feature, and "it runs" is not the same as "it is supported for your workload". That distinction becomes interesting the day you open a ticket about a guest that wedged in a page fault.

Google Compute Engine

The best-documented nested path of the three. Nesting is a property of the instance or instance template -- an advanced-machine-features flag, with an enable-vmx license on the image in some configurations -- and it is restricted by machine family and CPU platform. Historically it needs an Intel Haswell-or-later platform, and the shared-core, memory-optimized, Arm and most AMD families are excluded, with exceptions added over time. The catch is that the exclusion list is effectively the API: pick a family off it and every host in your managed instance group comes up without /dev/kvm, having reported success. Verify the current list in their docs; it moves.

Amazon EC2

Ordinary EC2 instances do not expose hardware virtualization to the guest. There is no flag, because there is no feature. The AWS answer to "I want KVM" is a *.metal instance type, which belongs in tier two. There is a pleasing irony in AWS having written Firecracker, running it under Lambda and Fargate on their own metal, and leaving you no nested option -- but it does at least make the decision quick. If your platform must run on EC2, your instance type is decided for you.

Azure

Azure exposes nested virtualization on a subset of sizes, by way of Hyper-V nesting, and a Linux guest on a supported size gets working KVM. Support has historically been Intel-based v3-and-later general-purpose and compute families; burstable B-series sizes and several AMD-based sizes have not been on the list. Sizes join and leave that list, so check your exact size in your exact region. Of the three hyperscaler nested paths this one feels the most like a feature built for Hyper-V labs that you are borrowing, which is fine until it is not.

Apple Silicon, for local work only

On an M3 or later Mac, nested virtualization is a hardware feature, and Lima running on Virtualization.framework will give you an Ubuntu VM with a real /dev/kvm inside it. That is how our own local end-to-end script runs a genuine Firecracker agent on a laptop, and it is excellent: you get a correct host, the same code path as production, and no cloud bill. What you do not get is capacity, a support story, or architecture parity -- your guests and your snapshots are arm64, Firecracker snapshots do not cross architectures, and they do not cross Firecracker versions either. Develop here, ship elsewhere.

My honest read on tier one: it is the right place to develop, to run your control plane's integration tests, and to prove your provisioning code. It is also where every platform in this category starts, including this one. Just make the nesting a known property of the fleet rather than something a new engineer discovers in /proc/cpuinfo nine months later.

Tier 2: bare metal from the hyperscalers

Here you get real KVM and real cost. The trade is not subtle: you pay the highest per-core rate on this page and you keep the API, the images, the object storage, the IAM, the private networking and the ability to hand the machine back this afternoon. For a platform whose control plane already lives in one of these clouds, that last point is worth more than the price delta suggests.

AWS *.metal

Available across several families, with genuine bare-metal access, cloud-shaped billing granularity, and the rest of your AWS estate one security group away. Local NVMe depends entirely on the family: instance-store families give you real local disks, and an EBS-only metal instance means your copy-on-write story runs over the network, which is the wrong end of the trade for a dense snapshot fleet. The honest catches are that metal sizes are large, so your minimum fleet granularity is coarse and you buy a whole machine's worth of capacity per step, and that metal capacity in a specific availability zone is not always instantly available. Verify current families, local-storage configurations and billing increments in their docs.

Google Cloud bare metal machine types

GCE now offers bare-metal machine types in the C3 and C4 series -- large Intel machines, with local SSD capacity on the C4 metal configurations. Two things people confuse here and should not. Sole-tenant nodes are not bare metal: the hardware is dedicated to you, but Google's hypervisor is still underneath, so every nested-virtualization rule from tier one still applies. And Bare Metal Solution is a different, specialised product aimed at workloads like Oracle, not a general KVM host. The catches on the real metal types are that there are very few of them, they are enormous, and regional availability is narrower than the virtualized families. Verify both.

Azure, and the bare-metal gap

Azure has no general-purpose bare-metal compute of the shape this post is about. BareMetal Infrastructure is a specialised offering for workloads like SAP HANA, Oracle and Nutanix, sales-led and not something you provision from a CLI on a Tuesday. A Dedicated Host gives you dedicated physical capacity with Azure's hypervisor still in place, which is dedicated, not bare. So on Azure the realistic Firecracker path remains a nested-capable size -- tier one, with a larger invoice and the same nesting tax.

Tier 3: independent bare metal and dedicated servers

Cheaper per core, often dramatically so, and you own considerably more. The operational tax is the real subject of this tier, so let me name it in full rather than hiding it in a sentence: you own the host kernel and its upgrades, the disks and their layout, the NICs and their offloads, the BIOS and firmware revisions, the out-of-band management interface and its password, DNS and reverse DNS, and the abuse desk. That last one is not a joke. If you host other people's code, your provider's abuse team becomes part of your on-call rotation, and you will learn more about cryptocurrency mining pools than you ever wanted to.

Hetzner

Dedicated root servers, including a server-auction market of returned machines, and the cheapest route to a large amount of real RAM and real NVMe that I know of. KVM works; a freshly delivered box occasionally arrives with the extensions disabled in firmware, which is a ticket rather than a problem. Billing on the dedicated line has traditionally been monthly with a one-time setup fee -- waived on auction machines -- and their billing model has been moving toward hourly with a monthly cap, so verify the current terms before you model a bursty fleet. Provisioning is order-driven: minutes for instantly available models, longer for anything that needs a human. The catches: the ordering API is an ordering API rather than an autoscaler, regions are few, and the cheapest boxes use consumer-grade parts. The last point is survivable, provided your design already assumes a host can die, which it should.

OVHcloud

A very wide dedicated-server range, from budget lines through to high-core-count performance machines with NVMe. Mostly monthly commitments, with some hourly options on specific products. Delivery time varies by range from a few minutes to long enough that you should plan around it. The honest catch is the product matrix itself: the cheap ranges and the business ranges differ in SLA, delivery, bandwidth policy and hardware generation, and a conclusion you reach about one range does not transfer to another. Price the exact range and read its exact terms.

Vultr Bare Metal

Hourly-billed bare metal, provisioned from an API in minutes, with NVMe on the relevant SKUs -- the shape of product that lets a fleet actually have a shape. We costed this seriously for a staging footprint and the model is attractive precisely because hourly metal makes a campaign-shaped workload affordable without a twelve-month commitment. The catch is the SKU catalogue: fewer configurations than a hyperscaler, and the binding constraint is often per-region availability of the specific machine you want rather than its price. Verify what is actually in stock in the region you need.

Equinix Metal, and why it is still on this list

Equinix Metal was the template for this category: hourly bare metal, API-first, deeply automatable, in a lot of facilities. Equinix announced in November 2024 that it was sunsetting the product, with 30 June 2026 given as the final day of service. That date is now behind us. I have left the entry in because the lesson is more useful than the entry: your hosting provider is a dependency with a product roadmap, and a roundup that still recommends Metal today is a roundup nobody re-checked. Keep your host provisioning in code, keep your artifacts in object storage rather than on a host, and keep the exit cheap. Then a provider shutting down is a migration rather than an incident. Verify directly with Equinix if you hold a contract.

Latitude.sh

Bare metal with hourly-through-annual billing, a REST API, a Terraform provider and a CLI -- built for the use case this post describes, and a frequent destination for the Metal migration above. NVMe and recent AMD platforms are the norm. One catch specific to snapshot-based platforms, and it applies to any fleet you grow across SKUs rather than to Latitude in particular: a Firecracker snapshot restored on a host whose CPU exposes a different feature set than the host that took it can fail, or worse, hand the guest an instruction it had previously been told did not exist. If your fleet is not CPU-homogeneous, the mitigation is Firecracker's CPU templates plus a placement rule, and you want to know that before the second machine arrives, not after.

Scaleway

Two products under one brand, and the difference matters. Elastic Metal is hourly, API-provisioned bare metal; Dedibox is the classic monthly dedicated-server line. Both give you real KVM. The catch is that local storage varies by generation and SKU -- some configurations are NVMe, others are not what you assumed -- so read the specific machine's spec rather than the product page, and confirm which billing model the SKU you clicked actually uses.

Honourable mentions: Oracle metal, and your own hardware

The cheapest per core by a wide margin, and the most operational tax by a wider one. Oracle Cloud's bare-metal shapes also belong in a sentence here as an hourly hyperscaler-adjacent option worth pricing. But owning the hardware means owning procurement lead times, spares, remote hands, and a conversation about redundant power. It is the right answer at a scale where the hardware line item is large enough to justify a person, and the wrong answer roughly everywhere before that.

Firecracker hosting options, graded on the five things that actually decide it. Every cell describing another company is from their public documentation at the time of writing and changes without warning -- verify against their current docs before you commit, and note that one entry has already reached end of service.
OptionReal /dev/kvmBilling granularityLocal NVMeProvisioning speedWhat you still operate
GCE nested-capable VMYes, nested (flag + Intel family)Sub-hourlyOptional local SSDSeconds, API/MIGHost OS, kernel, fleet, artifacts
AWS EC2, non-metalNo -- not exposed to guestsn/aFamily-dependentSecondsNothing; it cannot run Firecracker
Azure nested-capable sizeYes, nested (Hyper-V, Intel sizes)Sub-hourlySome sizes, ephemeral diskMinutes, APIHost OS, kernel, fleet, artifacts
Apple Silicon M3+ via LimaYes, nested, arm64Your laptopYes, it is an SSDMinutes, onceLocal dev only; arm64 snapshots
AWS *.metalYes, nativeCloud-shaped, verify incrementsYes on instance-store familiesMinutes, API, AZ capacity variesHost OS, kernel, disks, fleet
GCE C3/C4 bare metalYes, nativeCloud-shaped, verifyYes on C4 metalMinutes, API, fewer regionsHost OS, kernel, disks, fleet
GCE sole-tenant nodeNested only -- it is not bare metalCloud-shapedOptional local SSDMinutesAs a normal VM; family rules still apply
Azure BareMetal / Dedicated HostSpecialised / still virtualizedContract / cloud-shapedVariesSales-led / minutesNot the shape for this
Hetzner dedicatedYes, nativeMonthly + setup fee; hourly terms moving, verifyYes on most linesMinutes to hours, order-drivenEverything, including BMC and firmware
OVHcloud dedicatedYes, nativeMostly monthly, some hourly, per rangeYes on performance rangesMinutes to a day, per rangeEverything
Vultr Bare MetalYes, nativeHourlyYes on NVMe SKUsMinutes, APIEverything above the host image
Equinix MetalYes, nativeWas hourlyYesWas minutes, APINothing -- announced end of service has passed
Latitude.shYes, nativeHourly through annualYesMinutes, API/TerraformEverything; mind CPU-feature homogeneity
Scaleway Elastic Metal / DediboxYes, nativeHourly (EM) / monthly (Dedibox)Per SKU, verifyMinutes (EM)Everything
Own hardware in coloYes, nativeRack, power, monthlyWhatever you boughtWeeksLiterally everything, plus remote hands

What a Firecracker fleet actually needs from a host

KVM gets you a single VM. A fleet needs a handful of other things, and most of them are settled by choices made in the first hour of a host's life -- which is why they belong in a provisioning script rather than in a runbook nobody reads.

KVM, loaded and openable

The module loaded at boot, the device present, and the device's group and mode set so the process that boots VMs does not need to be root to open it. Our host provisioning adds a kvm group, a udev rule setting GROUP=kvm and MODE=0660 on the device, and the module in modules-load.d. It is three lines, and skipping them produces a failure that looks like a Firecracker bug for about an hour.

An XFS data volume, if you want reflink copy-on-write

This is the prerequisite people skip and then feel. When a sandbox is created we clone the template rootfs; on XFS with reflink enabled that is a FICLONE, which costs metadata rather than bytes, and the data stays shared until something writes. ext4 has no reflink at all, so the same operation becomes a full copy of a multi-gigabyte file. Reflink is not strictly mandatory -- our bake path tries FICLONE and falls back to a byte copy where it is unsupported, and dm-snapshot is the other route to host-side copy-on-write -- but the fallback converts an O(metadata) clone into an O(size) one, and you will meet that difference the first time you fan out.

The wrinkle is that most providers hand you one large ext4 root filesystem. You have two decent options: carve a dedicated XFS filesystem on the NVMe device, or create an XFS loopback image and mount it at the data directory. The second is what our own provisioning does -- mkfs.xfs with reflink and crc on an image file, mounted at the agent's data directory -- precisely because it works identically on every provider regardless of what their installer did to the root disk. Worth being explicit about one point of confusion: the guest rootfs images are ext4 files, and reflink is a property of the host filesystem they sit on. Those are different layers.

#!/usr/bin/env bash
# storage-check.sh -- the half of this decision that nested virtualization gets
# all the credit for. Point it at the volume you will use as the data directory,
# not at /.
set -uo pipefail
DATA=${1:-/var/lib/pandastack}
sudo mkdir -p "$DATA"

# Is the device local, and is it flash? ROTA=0 is non-rotational; TRAN tells you
# nvme vs sata vs something arriving over a network while politely presenting
# itself as a disk.
lsblk -d -o NAME,ROTA,SIZE,MODEL,TRAN

# Filesystem type decides the copy-on-write story for every rootfs clone:
#   xfs with reflink=1 -> FICLONE works; a clone costs metadata, not bytes
#   btrfs              -> FICLONE works too
#   ext4               -> no reflink, ever. A "clone" is a full byte copy.
findmnt -no FSTYPE,SOURCE,OPTIONS --target "$DATA"
[ "$(findmnt -no FSTYPE --target "$DATA")" = xfs ] && xfs_info "$DATA" | grep -E 'reflink|crc'

# And now the only authority that counts: not the docs, the kernel.
sudo head -c 64M /dev/urandom | sudo tee "$DATA/.rl-src" >/dev/null
if sudo cp --reflink=always "$DATA/.rl-src" "$DATA/.rl-dst" 2>/dev/null; then
  echo "reflink: OK -- rootfs clones are O(metadata)"
else
  echo "reflink: UNSUPPORTED on $DATA -- every rootfs clone is a full copy"
fi
sudo rm -f "$DATA/.rl-src" "$DATA/.rl-dst"

# Headroom. Template images, snapshots and per-sandbox CoW deltas all land here.
# An ENOSPC in the middle of a snapshot restore is a far worse failure than a
# refused create, which is why our agent statfs's this volume and answers 507
# below a floor (20 GiB by default) instead of trying and dying halfway.
df -h --output=target,fstype,size,avail,pcent "$DATA"

Free store disk, treated as a hard gate

Template images, baked snapshots, memory files and per-sandbox copy-on-write deltas all accumulate on the data volume, and the failure mode of running out is spectacularly worse than the failure mode of refusing to start another VM. Our agent statfs's the store volume on every create and answers 507 below a configurable floor, 20 GiB by default, with extra headroom reserved when a create will first need to pull an artifact from object storage. Size the disk for the artifact set rather than for the number of guests, and put a floor in front of it.

Hugepage overcommit, if you back guests with 2 MiB pages

If you use 2 MiB hugetlbfs pages for guest memory -- the attraction being that one page fault covers 2 MiB instead of 4 KiB, which matters a great deal on a demand-paged restore -- there are two sysctls and they are not interchangeable. This is the one I would most like to save you from learning the hard way.

  • vm.nr_overcommit_hugepages is permission to assemble a hugepage on demand. Under memory fragmentation that assembly fails, and plenty of free RAM does not help: what matters is contiguous 2 MiB runs, which collapse toward zero once a host has done real work. MemAvailable cannot see this. The failure surfaces somewhere unhelpful, such as mid-write during a snapshot dump.
  • vm.nr_hugepages is a boot-time reservation of contiguous memory, and it is what actually works on a long-lived host. Apply it as early in boot as you can, while RAM is still unfragmented.
  • The reservation is not free while idle. Reserved hugepages are not returned to the page cache and are subtracted from MemAvailable, so an oversized pool reads as memory pressure that does not exist -- on one of our 31.3 GiB hosts, a 9.4 GiB pool sitting almost entirely idle removed the same amount from MemAvailable. Size it to demand you can name, not to a share of RAM.
  • Hugepage-ness is a property of the snapshot, not just of the host. A snapshot taken from a hugepage-backed guest can only be restored through the userfaultfd path, so this choice propagates into your restore code and your template bakes.

Two smaller prerequisites round it out. IP forwarding, because each sandbox lives behind NAT in its own network namespace. And a pinned Firecracker version across the fleet, because snapshots are version-bound: two Firecracker builds in one fleet cannot restore each other's snapshots, which turns a rolling upgrade into a re-bake.

#!/usr/bin/env bash
# survive-a-reboot.sh -- the boring half. Every check above is worthless if the
# configuration evaporates the first time your provider reboots the host for a
# firmware update, which they will, at 04:00, during your busiest week.
set -euo pipefail

# KVM: load the module at boot, and make the device group-accessible so the
# agent does not need to run as root just to open it.
cat > /etc/modules-load.d/firecracker.conf <<'MODS'
kvm_intel
dm_snapshot
MODS
# kvm_amd rather than kvm_intel on AMD hosts. This is exactly the line that gets
# copied between providers unexamined and then fails on precisely one machine, so
# derive it instead of trusting your memory of what you ordered.
grep -qm1 ' svm ' /proc/cpuinfo && sed -i s/kvm_intel/kvm_amd/ \
  /etc/modules-load.d/firecracker.conf
groupadd -f kvm
printf 'KERNEL=="kvm", GROUP="kvm", MODE="0660"\n' > /etc/udev/rules.d/99-kvm.rules
udevadm control --reload-rules && udevadm trigger --name-match=kvm

# An XFS data volume even when the provider handed you one big ext4 root. A
# dedicated partition on the NVMe is cleaner; a loopback image is what you do
# when reinstalling the host is not on the table. Both give you FICLONE.
if [ ! -f /opt/fcdata.img ]; then
  fallocate -l 300G /opt/fcdata.img
  mkfs.xfs -m reflink=1 -m crc=1 -q /opt/fcdata.img
fi
mkdir -p /var/lib/pandastack
grep -q fcdata.img /etc/fstab || \
  echo "/opt/fcdata.img /var/lib/pandastack xfs loop,noatime,nofail 0 0" >> /etc/fstab
mount -a

# Sysctls. ip_forward because each sandbox sits behind NAT in its own network
# namespace. The two hugepage knobs only if you run 2 MiB-backed guests, and
# they are not interchangeable: nr_hugepages is a boot-time RESERVATION of
# contiguous memory and is what actually works, while nr_overcommit_hugepages is
# permission to assemble pages on demand and fails under fragmentation. Size the
# reservation to demand you can name -- idle reserved hugepages are not returned
# to the page cache and are subtracted from MemAvailable, so an oversized pool
# reads as memory pressure that does not exist.
cat > /etc/sysctl.d/99-firecracker.conf <<'SYSCTL'
net.ipv4.ip_forward=1
vm.nr_hugepages=1152
vm.nr_overcommit_hugepages=16
SYSCTL
sysctl --system

# Then verify after the NEXT reboot, not now. "It worked when I ran the script"
# and "it comes back" are different claims, and only one of them pages you.

That script is, with the vendor-specific parts removed, what our cloud-init does to a fresh agent host: KVM device and group, the XFS reflink volume, the hugepage sysctls, the dm-snapshot module, forwarding. The list is short, every line is there because something broke without it, and none of it is Firecracker-specific. And the one-line version of the pitch: PandaStack runs this on managed Linux KVM hosts so you can call an API instead, which is worth exactly as much as your appetite for owning BIOS firmware revisions.

Billing granularity is an architecture decision

Of the five columns in that table, the one that changes your design rather than your invoice is billing granularity. Hourly metal means your fleet can have a shape: you can add capacity for a launch, a campaign or a Tuesday, and your scheduler's job is placement. Monthly metal means your capacity is a step function you already paid for, and your scheduler's job quietly becomes bin-packing into a fixed floor -- a different problem, with different failure modes, mostly involving a refusal to create at 80 percent utilisation because the remaining space is the wrong shape.

For a sandbox fleet the honest answer is usually both. The part of your load that exists regardless of the hour belongs on the cheapest metal you can get on a monthly or annual term, because that floor is where the per-core saving compounds. The part above it belongs on hourly capacity, metal or nested, that you can hand back. That split also keeps a provider's roadmap from being your problem: if the hourly tier shuts down, you resize the floor and re-point your provisioning.

Two line items to check while you are in the pricing docs, neither of which is a price I will quote: what egress costs, because it differs by orders of magnitude between hyperscalers and independents and is the bill that surprises people; and what a partial month costs on a monthly product when you cancel. The second determines whether you can experiment at all.

How I would choose, in five lines

  • Developing, or testing your control plane? A nested-capable cloud VM, or Lima on an M3-or-later Mac. Do not over-engineer this; you need correctness, not capacity.
  • Platform already lives in AWS or GCP and you want one vendor? A *.metal EC2 instance or a C3/C4 bare-metal machine type. You will pay the most per core and save it back in integration and the ability to hand the machine back.
  • Stuck on Azure? A nested-capable size, and plan for the nesting tax, because the bare-metal product is not aimed at this.
  • Building a fleet where per-core cost is the business model? An independent provider, with a monthly or annual floor on the cheapest good metal and hourly capacity above it. Hetzner for density per currency unit, Latitude.sh or Vultr for an hourly API, Scaleway or OVHcloud depending on region.
  • Whatever you pick: run the check scripts above on the first machine, put their contents in your provisioning repo, and make a host that fails them refuse to join the fleet.
Every hosting decision in this category is eventually re-made by someone else's product roadmap. Keep the provisioning in code and the artifacts in object storage, and that day is a migration rather than an incident.

The durable advice here is not a provider name, it is a shortlist procedure: confirm /dev/kvm before anything else, confirm reflink on the volume you will actually use, confirm the billing granularity matches the shape of your load, and confirm you know who owns the firmware. Everything else on this page will have moved by the time you read it, including the parts about my own platform -- so check it, including against my docs.

Frequently asked questions

Can I run Firecracker on a regular cloud VM or a cheap VPS?

Usually not, and the test takes ten seconds: ls -l /dev/kvm on the machine. Most general-purpose virtualized instances do not expose hardware virtualization to the guest, because the provider's hypervisor is using those extensions itself and nesting is an explicit, per-family feature they have to choose to offer. Ordinary EC2 instances do not expose it at all -- on AWS the answer is a *.metal instance type. Google Compute Engine does expose it, but only on specific machine families with a flag set at instance or template creation time, and the excluded list has historically included the shared-core, memory-optimized, Arm and most AMD families. Azure exposes it on a subset of sizes. Smaller VPS providers vary wildly and their marketing copy is not a reliable guide, so test the device rather than reading the feature list. One trap worth naming: some providers expose the CPU flags without exposing a working device, so grepping /proc/cpuinfo for vmx or svm is not sufficient evidence. The device has to exist, your service user has to be able to open it, and Firecracker has to be able to start against it. All three, verified against the provider's current documentation, because this changes.

Is nested virtualization good enough for production, or do I need real metal?

It is good enough more often than purists admit, and I can say that with a straight face because PandaStack runs on nested virtualization in production -- the snapshot-restore create latency we publish, p50 179 ms and p99 around 203 ms, is measured on nested hosts. What you pay is real but workload-dependent: every VM exit that the middle hypervisor cannot resolve on its own traps down another level, and your guest's vCPU thread is scheduled by two schedulers instead of one. Workloads dominated by guest-side computation notice this far less than workloads doing constant I/O and device work. I deliberately will not quote you an overhead percentage, mine or anyone else's, because the honest version of that number comes from measuring your own workload on both -- which, with the check script above and a hundred creates, is an afternoon's work rather than a research project. The non-performance arguments are often the deciding ones: on metal you get real local NVMe, which decides your copy-on-write and fan-out story, you get a support relationship where nobody can suggest your use of the feature is unusual, and you get a per-core cost that makes a dense fleet a business rather than a hobby.

Does local NVMe actually matter, or is network-attached SSD fine?

It matters, and it matters for a structural reason rather than a benchmark one. Copy-on-write cloning of a rootfs needs a local block device: a reflink is a filesystem operation on a real filesystem, and dm-snapshot wants a real device underneath it. You cannot reflink a network volume into existence. So if your host has only network-attached storage, every sandbox create either copies a multi-gigabyte image over the network or your platform needs an entirely different disk strategy. The second place it shows up is read latency under fan-out: restoring a snapshot touches a large memory file, and doing that for many guests at once on shared network storage behaves differently from doing it on a device in the chassis. There is a software answer to part of this -- we stream guest memory on demand from object storage over userfaultfd, with a local chunk cache, so the multi-gigabyte memory file does not have to be downloaded first -- but note what that does and does not solve. It streams memory. The rootfs is always a local file, because copy-on-write needs a local block device. When you are comparing SKUs, check the TRAN column in lsblk and whether the filesystem on that device supports FICLONE, not the marketing word "SSD".

Hourly or monthly bare metal for a sandbox fleet?

Both, split by load shape rather than by preference. The portion of your fleet that is busy regardless of the hour -- your floor -- belongs on the cheapest metal you can get, which in practice means a monthly or annual commitment with an independent provider, because that is where per-core savings compound into the unit economics of the product. Everything above the floor belongs on capacity you can return: hourly bare metal from an API-driven provider, or nested cloud VMs if the workload tolerates it. The reason this matters beyond the invoice is that billing granularity changes what your scheduler is for. With hourly capacity the scheduler places work and you add machines; with a fixed monthly floor the scheduler is bin-packing into capacity you already bought, and the interesting failures are refusals at high utilisation because the free space is the wrong shape rather than too small. Two things to check before signing anything monthly: what a cancelled partial month actually costs, which determines whether you can experiment, and what provisioning a replacement takes in wall-clock time, which determines whether a dead host is an incident or a shrug. A monthly commitment with a one-time setup fee and a multi-hour delivery window is a very different operational posture from an API call that returns a machine in minutes.

What breaks when my fleet spans different providers or CPU models?

Three things, all worth designing for before the second machine arrives. First, CPU features: a Firecracker snapshot captures a guest that was told it had a particular feature set, and restoring it on a host whose CPU differs can fail, or can hand the guest an instruction it had been told did not exist -- which is a worse outcome than failing. The mitigation is Firecracker's CPU templates to present a consistent, lowest-common-denominator guest CPU, plus a placement rule so restores land on compatible hardware. Second, the Firecracker version itself: snapshots are version-bound, so a fleet running two builds cannot restore each other's snapshots. Pin the version in provisioning, and treat a Firecracker upgrade as a re-bake of every template rather than a rolling restart. Third, host-side capability drift: a host whose data volume is ext4 rather than XFS will silently behave differently from its neighbours, because the clone that was a metadata operation everywhere else is now a full copy there. All three have the same remedy, which is unglamorous: make host provisioning a script in your repository, make every host prove it passes the checks before it accepts work, and have the scheduler know which hosts hold which artifacts. Heterogeneity is fine. Undetected heterogeneity is a pager at 3 a.m.

Keep reading

Related posts

  • Firecracker vs Proxmox: a VMM Is Not a Platform

    Proxmox VE manages VMs. Firecracker runs one. The comparison is a layer mismatch — but the question underneath it is real, and the answer usually comes down to whether your guests are pets or cattle.

  • KVM explained for developers: the hardware boundary under microVMs

    You keep seeing "KVM" and "hardware virtualization" in isolation discussions. Here's the real mental model: a kernel module, the CPU's VT-x/AMD-V extensions, and the VM-exit trap that is the security boundary.

  • Can You Run a Sandbox Fleet on Spot Instances?

    The discount is real and the question people ask about it is the wrong one. Spot is not a pricing decision, it is a statement about whether your workload can be interrupted and rebuilt — which is a different answer for a 30-second code execution than for a customer's database.

  • The best Kafka hosting platforms in 2026

    Kafka is a distributed log that has convinced a generation of teams they have a streaming problem. If you genuinely do, here is how the hosting market actually differs.

  • Top 5 AI Agent Hosting Platforms in 2026

    Hosting an agent is not hosting a web app. It's long-running, bursty, stateful, and it executes untrusted code. Here's a 2026 shortlist of where to actually run agents in production — judged by execution model, isolation, state, and idle economics.

More in Firecracker & microVMs · See Firecracker microVM sandboxes

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.