all posts

Build, Boot, Test: Kernel CI on MicroVMs (and Where It Stops)

Ajay Kumar··10 min read

The loop a kernel developer actually wants is embarrassingly simple to describe and historically miserable to own. Apply a patch. Build it. Boot it. See whether the box came up. Run something. Throw the box away. Repeat about two hundred times a day, in parallel, across a few configs.

The reason that loop has traditionally lived in a nightly job rather than in a pre-commit hook is the third step. Booting a machine is slow, booting a machine you cannot trust is risky, and booting a machine you have to physically own is an entire discipline with its own serial consoles and remote power switches. A microVM collapses the boot step to a fraction of a second, which changes what class of check "does this patch boot" belongs to: it stops being a nightly and becomes a per-commit assertion, the same way a unit test is.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This post is about using a microVM as a kernel CI harness, and it is going to spend an unusual amount of its length on where that stops working — including the place it stops working on my own platform. That is not false modesty. The ceiling is sharp, it is easy to walk into, and the people most likely to try this are precisely the people who will not thank me for burying it under the demo.

Read this first: the device model is the ceiling

Firecracker is a minimal virtual machine monitor, and "minimal" there is a statement about the hardware it presents, not only about its source tree. The machine it presents is a short list: virtio-net, virtio-block, virtio-vsock, virtio-balloon, virtio-rng, a pmem device, a serial console, and an i8042 keyboard controller whose only real job is to notice that the guest asked to reboot. Current versions also offer an opt-in virtio-PCI transport behind `--enable-pci`, with hot-plug of virtio-block, virtio-net and virtio-pmem marked developer preview, and ACPI is supported — so the flat claims "Firecracker has no PCI" and "Firecracker has no ACPI" are stale, even though MMIO is still the default and our own guests boot with `pci=off`.

Device-model facts here decay on a timescale of about two releases. Read `docs/`, `resources/guest_configs/` and `src/firecracker/swagger/firecracker.yaml` in the Firecracker repository for the version you actually run, rather than trusting this paragraph in six months.

So draw the line by subsystem rather than by vibe. A microVM is a superb harness for work in the core kernel, the scheduler, memory management, the networking stack, filesystems, the VFS, cgroups, namespaces, seccomp, BPF, and the virtio drivers themselves. All of that exercises code paths the machine genuinely has. It is the wrong harness for real hardware drivers, GPU work, anything that depends on PCI passthrough or VFIO, anything that needs a device quirk to reproduce, and anything whose bug report begins "on this specific controller."

QEMU remains the right tool for a wide device matrix, and I would rather say that plainly than pretend otherwise. If your patch touches a driver, you want the emulator that can pretend to be the hardware. If your patch touches the page allocator, you want the VMM that boots in a fraction of a second and that you can run two hundred of.

The build half is just a big compile

Half of kernel CI is not interesting, and that is good news. `make -j`, a toolchain, and enough disk. There is no clever trick here, only two numbers that quietly decide whether the job runs at all, and on a snapshot-based platform both of them are chosen before the job exists.

The first is rootfs size. Our template build takes `--size-mb`, the agent's default is 1024 MB, and the bundled Go CLI passes 2048 — both of which are comedy next to a kernel source tree plus an object tree. The second is RAM, set with `--memory-mb` at template build time and baked into the snapshot. Firecracker cannot change a guest's vCPU count or RAM at snapshot restore, so passing `memory_mb` on a create is not an error; it is silently corrected to the baked value, which is worse. A link step that wants more memory than the template was baked with needs a different template, not a different argument.

Three smaller things that cost people an afternoon each. Our `base` template installs `build-essential`, `git`, `curl`, `pkg-config` and `xz-utils` — and not `flex`, `bison`, `bc`, `libelf-dev`, `libssl-dev`, `cpio` or `rsync`, all of which a kernel build reaches for; `Documentation/process/changes.rst` is the authoritative list of minimum tool versions. Docker `ENV` is not injected into the Firecracker guest, so `CCACHE_DIR` set in your Dockerfile will not reach a detached build process — only `/etc/environment` does. And every template gets 8 burstable vCPUs, which means `make -j$(nproc)` is `make -j8` whether you wanted that or not.

On config generation: the folk advice is `make localmodconfig`, and it is the wrong tool here. `localmodconfig` trims a config down to the modules currently loaded on the machine you run it on, which in a microVM built with `CONFIG_MODULES=n` is an empty set. The microVM equivalent is to start from Firecracker's own guest config in `resources/guest_configs/` and run `make olddefconfig`. That gets you a kernel configured for the machine you are actually going to boot, which is a different machine from the one `defconfig` assumes.

from pandastack import Sandbox

# A kernel build is a disk-and-CPU hog and both of the numbers that decide
# whether it runs at all are fixed at TEMPLATE BUILD TIME, not here:
#
#   pandastack template build -f Dockerfile -n kernel-build \
#     --size-mb 40960 \      # rootfs. The agent's default is 1024 MB and the
#                            # bundled Go CLI passes 2048 -- a kernel tree plus
#                            # an object tree laughs at both. Measure du -sh on
#                            # your own config and then double it.
#     --memory-mb 8192       # baked into the snapshot. Firecracker cannot
#                            # change guest RAM at restore, so this is the only
#                            # place it is chosen. --cpu is deprecated and
#                            # ignored; every template gets 8 burstable vCPUs.
#
# The Dockerfile needs more than `base` ships. base installs build-essential,
# git, curl, pkg-config and xz-utils -- and not flex, bison, bc, libelf-dev,
# libssl-dev, cpio or rsync, every one of which a kernel build wants. Check
# Documentation/process/changes.rst for the current minimum versions.
sbx = Sandbox.create(template="kernel-build", ttl_seconds=3600)

# cpu= and memory_mb= are not passed on purpose: the agent SILENTLY corrects
# them to the baked snapshot's values, so passing them reads like a lie.

# Docker ENV does not reach a detached process in the guest -- only
# /etc/environment (PAM) does -- so export what the build needs inline rather
# than assuming the image's ENV survived the flattening.
BUILD = r"""
set -euo pipefail
export CCACHE_DIR=/cache/ccache CCACHE_MAXSIZE=20G
export KBUILD_BUILD_TIMESTAMP='@0'
cd /src/linux
cp /opt/fc-configs/microvm-kernel-ci-x86_64-5.10.config .config
make olddefconfig
make CC="ccache gcc" HOSTCC="ccache gcc" -j"$(nproc)" vmlinux
ccache -s | sed -n '1,12p'
xz -T0 -c vmlinux > /out/vmlinux.xz
sha256sum /out/vmlinux.xz
"""

# exec_stream, not exec. Neither exec endpoint enforces timeout_seconds
# server-side -- the agent decodes it and applies it on neither -- so the
# client gives up at 30 seconds by default, which a kernel build will reach
# while looking exactly like a hang. Passing it to exec_stream raises the
# CLIENT timeout and hands you the build log as it happens, which is the only
# way to tell a slow link from a wedged one. Hard limits go in the shell.
rc = sbx.exec_stream(
    BUILD,
    on_stdout=lambda chunk: print(chunk, end=""),
    on_stderr=lambda chunk: print(chunk, end=""),
    timeout_seconds=5400,
)
if rc != 0:
    raise SystemExit(f"kernel build failed: rc={rc}")

# filesystem.read() returns bytes and handles ONE file -- there is no
# directory copy, and upload() of a directory raises IsADirectoryError. For a
# compressed vmlinux that is fine. For a module tree, tar it in the guest
# first, or better: push it to your own object storage from inside the guest
# and keep the artifact out of the control path entirely.
open("vmlinux.xz", "wb").write(sbx.filesystem.read("/out/vmlinux.xz"))

# Note what has NOT happened: nothing booted this kernel. It is a file now,
# and the next machine that touches it needs /dev/kvm.
sbx.kill()

The boot half, and the sentinel that makes it a test

A boot test needs an unambiguous answer, and the nice property of a microVM is that "unambiguous" is cheap to arrange. Firecracker writes the guest's serial console to its own stdout, so the whole verdict mechanism is: start the process, capture stdout, watch for a string your probe init prints. No agent in the guest, no network, no ssh, nothing to poll.

Two kernel command line tokens do most of the work. `panic=1` makes a guest panic end the VM instead of leaving a kernel spinning in a reboot loop, and `reboot=k` routes reboot through the i8042 stub — which is also how your probe init exits cleanly with `reboot -f`. In a per-commit harness, a test that hangs is strictly worse than a test that fails, because the hang costs you the full timeout on every step of a bisect while the failure costs you nothing.

#!/usr/bin/env bash
# build-and-boot.sh -- build the kernel in $PWD and prove it boots under
# Firecracker. Designed to be handed straight to `git bisect run`:
#
#   exit 0    booted to userspace
#   exit 1    built, did not boot (panic, hang, init died)
#   exit 125  did not build -- UNTESTABLE, tell git to skip it
#
# Runs on a machine with /dev/kvm. See the note later about why this cannot
# run inside a Firecracker guest.
set -uo pipefail

KDIR=${KDIR:-$PWD}
OUT=${OUT:-/tmp/kbuild}                    # out-of-tree: keeps the source clean
FC_SRC=${FC_SRC:-/opt/firecracker}         # a checkout, for its guest configs
PROBE_ROOTFS=${PROBE_ROOTFS:-/var/lib/kci/probe.ext4}
SENTINEL=KCI_USERSPACE_REACHED
BOOT_TIMEOUT=${BOOT_TIMEOUT:-45}

# ---------------------------------------------------------------- 1. config
# Start from Firecracker's own guest config, not `defconfig`. defconfig builds
# a kernel for a PC; you are booting something that is not a PC. The configs
# live in resources/guest_configs/ in the Firecracker repo -- read the ones
# for the version you actually run, the filenames move between releases.
mkdir -p "$OUT"
cp "$FC_SRC/resources/guest_configs/microvm-kernel-ci-x86_64-5.10.config" \
   "$OUT/.config"

# olddefconfig is the load-bearing line in this whole script. During a bisect
# you jump across commits that add and remove Kconfig symbols; without it an
# old .config either answers new questions wrong or stops the build to ask a
# human, and an interactive prompt inside `git bisect run` is a hang.
make -C "$KDIR" O="$OUT" olddefconfig >/dev/null || exit 125

# ----------------------------------------------------------------- 2. build
# `vmlinux`, not `all`: on x86_64 Firecracker loads the uncompressed ELF, so
# bzImage and the module tree are work nobody is going to read. (On aarch64
# you want `make Image` -- arch/arm64/boot/Image, the PE-format kernel.)
#
# ccache is worth the setup here for a reason specific to bisect: between two
# adjacent commits almost every object file is identical, so steps 2..N of a
# bisect are mostly cache hits and a link.
if ! make -C "$KDIR" O="$OUT" \
          CC="ccache gcc" HOSTCC="ccache gcc" \
          KBUILD_BUILD_TIMESTAMP='@0' \
          -j"$(nproc)" vmlinux; then
  echo "BUILD FAILED -- reporting untestable (125)" >&2
  exit 125
fi

# ------------------------------------------------------------------ 3. boot
# Firecracker writes the guest's serial console to its own stdout, so the
# entire boot test is: start it, capture stdout, watch for a sentinel that the
# probe rootfs prints from its init.
#
# panic=1 and reboot=k matter more than they look. panic=1 turns a guest panic
# into an exiting VM instead of a kernel sitting in a loop, and reboot=k
# routes the reboot through the i8042 stub -- which is also how the probe
# init exits cleanly with `reboot -f`. A boot test that hangs is worse than
# one that fails, because it burns the timeout on every single bisect step.
SOCK=$(mktemp -u /tmp/fc-XXXXXX.sock)
LOG=$(mktemp /tmp/console-XXXXXX.log)
CLONE=$(mktemp /tmp/probe-XXXXXX.ext4)
cp --reflink=auto "$PROBE_ROOTFS" "$CLONE"   # never boot the pristine image

firecracker --api-sock "$SOCK" >"$LOG" 2>&1 &
FC=$!
trap 'kill -9 "$FC" 2>/dev/null; rm -f "$SOCK" "$CLONE"' EXIT
for _ in $(seq 50); do [ -S "$SOCK" ] && break; sleep 0.05; done

fcapi() { curl -sf --unix-socket "$SOCK" -X PUT "http://localhost$1" \
               -H 'Content-Type: application/json' -d "$2"; }

fcapi /boot-source "$(cat <<JSON
{ "kernel_image_path": "$OUT/vmlinux",
  "boot_args": "console=ttyS0 reboot=k panic=1 pci=off nomodule init=/sbin/kci-probe" }
JSON
)" || exit 1

fcapi /drives/rootfs "$(cat <<JSON
{ "drive_id": "rootfs", "path_on_host": "$CLONE",
  "is_root_device": true, "is_read_only": false }
JSON
)" || exit 1

fcapi /machine-config '{"vcpu_count": 2, "mem_size_mib": 512, "smt": false}' || exit 1
fcapi /actions        '{"action_type": "InstanceStart"}'                      || exit 1

# ------------------------------------------------------------------ 4. verdict
rc=1
deadline=$(( $(date +%s) + BOOT_TIMEOUT ))
while [ "$(date +%s)" -lt "$deadline" ]; do
  if grep -q "$SENTINEL" "$LOG";                                 then rc=0; break; fi
  if grep -qE 'Kernel panic|Attempted to kill init' "$LOG";       then rc=1; break; fi
  if ! kill -0 "$FC" 2>/dev/null; then
    grep -q "$SENTINEL" "$LOG" && rc=0 || rc=1; break
  fi
  sleep 0.2
done

if [ "$rc" -ne 0 ]; then
  echo "=== last 40 lines of guest console ===" >&2
  tail -40 "$LOG" >&2
fi
exit "$rc"

The probe rootfs is deliberately boring: a few megabytes, an init script that prints the sentinel, optionally runs one check, and calls `reboot -f`. Bake it once and treat it as a fixture. Everything that varies in this harness is the kernel; everything else should be a file you have not thought about in a month.

The payoff: git bisect run over a boot test

This is the part that justifies the whole exercise. A bisect is a binary search, so its cost is the number of steps times the cost of a step, and the number of steps is fixed by arithmetic — about ten for a thousand-commit range, about fourteen for sixteen thousand. You cannot make the search shorter. You can only make a step cheap, and a step whose boot phase is sub-second and whose build phase is mostly ccache hits is a fundamentally different thing from a step that provisions hardware.

The mechanics have one sharp edge, and it is the exit codes. `git bisect run` reads 0 as good, 1 through 124 as bad, 125 as "skip, untestable", and anything from 128 up as "abort now." In a range worth bisecting, some commits will not compile. Score those as bad and the search walks into the wrong half and returns a plausible, wrong commit — which you will then spend a day defending to the person who wrote it.

#!/usr/bin/env bash
# bisect-boot.sh -- bisect a boot regression where every step is a fresh guest.
#
# git's exit-code contract for `bisect run`, which is the part people get
# wrong and the reason bisects confidently blame innocent commits:
#
#   0            this commit is GOOD
#   1..124       this commit is BAD          (125 is excluded from this range)
#   125          SKIP -- untestable
#   128 or more  abort the whole bisect immediately
#
# The 125 case is not a nicety. In any range worth bisecting, some commits do
# not compile, and scoring those as "bad" walks the search into the wrong half
# and hands you a plausible, wrong answer that you will then spend a day
# defending. build-and-boot.sh returns 125 on a build failure for exactly
# this reason.
set -u

git bisect start
git bisect good v6.6          # last release known to boot
git bisect bad  v6.7          # first release known not to

# One guest per step. A 1,000-commit range is about ten steps, because
# log2(1000) is about ten -- which is the entire argument for making a single
# step cheap rather than making the search cleverer.
git bisect run ./build-and-boot.sh

# ALWAYS keep the log. It replays without rebuilding anything, which turns
# "I think it was that commit" into something you can hand to someone else.
git bisect log > bisect.log
git bisect reset
# later, or on another machine:
#   git bisect replay bisect.log

# ---------------------------------------------------------------------------
# Two refinements worth the extra lines once the easy version works.
#
# 1. A boot test that is really a flake test. If the regression is a race,
#    one boot proves nothing. Boot N times and only report GOOD if all N
#    pass; the whole point of a sub-second guest is that N=20 is still
#    cheaper than one QEMU boot:
#
#      for i in $(seq 20); do ./build-and-boot.sh || exit 1; done; exit 0
#
#    (Keep the 125 pass-through if you do this -- `|| exit 1` above would
#    turn a build failure into a BAD verdict, which is the bug this comment
#    block exists to prevent.)
#
# 2. Bisecting a *test* rather than a boot. Same contract, different inner
#    command: boot the guest, run one kselftest over vsock or ssh, and map
#    its exit status. Keep the mapping explicit -- a harness that cannot
#    distinguish "the test failed" from "the test did not run" is a harness
#    that produces confident nonsense.
Keep the `git bisect log`. It replays on another machine without rebuilding anything, which is the difference between "I'm fairly sure it was that commit" and a result someone else can check. It is also the only artifact that survives you closing the terminal.

The test harnesses, named accurately

There is an existing ecosystem here and pretending otherwise would be both rude and unhelpful. Each of these has moved recently enough that you should read its current documentation rather than my description of it.

  • `kselftest` — the in-tree selftests under `tools/testing/selftests`. Mostly userspace programs that exercise kernel interfaces, driven by `make kselftest` or the `run_tests` targets, and installable as a standalone tree (`make -C tools/testing/selftests install`) that produces a `run_kselftest.sh` you can drop into a rootfs. That installable form is what makes it a good fit for a boot-and-run harness: the tests ship with the kernel you just built.
  • KUnit — in-tree unit tests for kernel code, configured with `CONFIG_KUNIT` and reporting results in KTAP. Its driver, `tools/testing/kunit/kunit.py`, builds and runs the kernel as a User-Mode Linux binary by default and can target other architectures through QEMU. You can also build the tests into a real kernel and read the results out of the console on boot, which is the form that fits a microVM harness.
  • LTP, the Linux Test Project — a large userspace test suite that runs on a booted system rather than a built tree. It is the "and now actually exercise the kernel" half, and it is long-running enough that it belongs in a scheduled job rather than per commit.
  • `virtme-ng` — the closest existing thing to this post's idea, and it deserves the acknowledgement. It boots the kernel you just built in your working tree directly, against your host filesystem, with no rootfs to assemble; QEMU is underneath. If your loop is "edit, rebuild, boot, poke at it," start here rather than building any of the above.
  • KernelCI — the upstream project doing distributed build-and-boot testing of mainline across a federation of labs with real hardware. If you are wondering whether a tree builds and boots across a hundred platforms, that question already has an answer and it is not yours to rebuild.

Nested virtualisation: the limit on my own platform

Here is the honest constraint, and it is the reason this post is not an advert. Firecracker does not expose virtualisation extensions to its guests. There is no `/dev/kvm` inside a microVM, the CPU features it advertises do not include VMX or SVM, and no configuration changes that — it is a design property, not a missing flag. So a kernel CI job that itself wants to boot VMs — KVM unit tests, nested-virt tests, anything that runs QEMU with acceleration — cannot run inside a Firecracker guest. It needs bare metal, or a VMM that exposes nested virtualisation to its guests.

There is one real escape hatch and it is worth knowing: QEMU with TCG rather than KVM acceleration boots real guest kernels inside a microVM, because software emulation never wanted hardware virtualisation in the first place. It is slow in the way emulating a CPU in software is slow. For a boot check — a short workload whose answer is a single bit — that is frequently an acceptable trade, and it is one of the few cases where "slow" and "fine" overlap. Measure it against your own kernel and config before designing around it.

The guest kernel is not your kernel

Two flavours of this, and the second one disqualifies my own product for part of this job.

The first is a config story. Minimal microVM kernels are minimal on purpose, and the omissions surface as failures that look nothing like a missing Kconfig symbol. We measured this properly while evaluating whether to ship a Kubernetes template: k3s installs cleanly in a sandbox, the node reaches Ready, pods run, pod-IP networking serves traffic — and ClusterIP services never work at all, because kube-proxy emits `-m comment` on every rule and the guest kernel lacks `xt_comment`, so the entire `iptables-restore` batch fails atomically and zero service rules land. Five netfilter config symbols are missing from both of Firecracker's CI guest kernels, 5.10 and 6.1, so upgrading does not help; they are config options, not version features. We did not ship the template, because a product that boots and then fails the first `kubectl expose` is the worst kind of half-working.

The second is structural and specific to us. Our agent picks the guest kernel by globbing `DataDir/kernels/vmlinux-*` and taking the last match — one kernel per agent, for every template on it. The bake script already writes a `"kernel"` field into each template's `meta.json`, and nothing reads it at restore time, so dropping a second kernel on a host today would silently repoint every template and invalidate every baked snapshot. The practical consequence is blunt: you cannot boot a kernel you just built inside a PandaStack sandbox. We are the build farm and the place to run userspace tests against the platform's kernel. We are not, today, a kernel-boot-test harness.

Those two constraints compound. You cannot swap the sandbox's kernel, and you cannot run an accelerated VMM inside the sandbox to boot your own. The boot half of this loop wants a host with `/dev/kvm` — or the TCG escape hatch above, with its performance consequences accepted up front.

Four harnesses, compared on what actually trades off

Kernel build-and-boot harnesses. Boot times are qualitative on purpose: the only figure here I will put a number on is Firecracker's own advertised startup for a minimal guest, and you should measure your own kernel and config rather than inherit anybody's.
HarnessDevice coverageBoot of a fresh kernelNested virt inside itSetup cost
QEMU with KVMWidest: full PC chipset, PCI, a large emulated device catalogue, plus TCG for foreign machinesFast, and `-kernel` direct boot skips firmware entirelyYes, if the host below you exposes it — not your decisionLarge configuration surface, but every distro ships it and every kernel developer already knows it
virtme-ngQEMU's, minus what it configures away for speedFastest to iterate: no rootfs to assemble, boots your build tree against the host filesystemInherits QEMU's answerLowest: one command inside a kernel tree
FirecrackerSmallest: virtio-net/block/vsock/balloon/rng, pmem, serial console, i8042 reboot stub; opt-in virtio-PCI, ACPI supportedSub-second; upstream advertises roughly 125 ms for a minimal guestNo — no VMX/SVM exposed, no `/dev/kvm` in the guest, by designLow once a probe rootfs exists; the boot path is four API calls
Bare metalEverything, including the quirks nothing emulatesSlowest: POST, firmware, bootloader, then the kernelYes — it is the real thingHighest: provisioning, netboot, serial capture, remote power, and a human when all of that fails

These compose rather than compete, and the composition is the actual recommendation. `virtme-ng` for the inner loop on a developer's machine. Firecracker for the per-commit boot gate and the bisect, where step cost dominates. QEMU for the device matrix and anything driver-shaped. Bare metal, or a federation like KernelCI, for the platform coverage you cannot emulate. Picking one and defending it is a worse engineering position than running three.

Why our usual trick does not apply here

Almost every post on this blog eventually arrives at snapshot-restore, because it is the thing that makes our creates cost 179 ms at p50 and 203 ms at p99 with the `/snapshot/load` step itself around 49 ms. This is the one workload where that fast path is beside the point, and saying so is more useful than finding an angle.

A snapshot restores a guest that has already booted a particular kernel. Kernel CI deliberately boots a new kernel every time — the kernel is the variable under test, and restoring a memory image of the previous one is exactly the thing you must not do. You are cold-booting on purpose. The first cold boot of one of our templates, before its snapshot exists, is about 3 seconds for kernel plus a full Ubuntu userspace coming up; a purpose-built probe rootfs is a great deal less than that, because almost all of those 3 seconds are userspace.

What snapshots do buy you is the other half. The toolchain, the ccache, the kernel source tree, the installed selftests, the probe rootfs — all of that is identical across every step of a bisect and all of it is what you bake. A build sandbox that comes up in 179 ms with gcc, flex, bison and a warm compiler cache is a real win on a harness that starts two hundred of them a day. The honest framing is a one-liner: the rootfs template and the toolchain are what you bake; the kernel is what you vary.

What to actually build, in order

  1. A probe rootfs: a few megabytes, an init that prints a sentinel and calls `reboot -f`. Bake it once, version it, and stop thinking about it.
  2. A boot test that exits 0, 1 or 125 and prints the guest console on failure. The exit codes are the contract; everything else is implementation.
  3. A build template sized deliberately — `--size-mb` for the object tree, `--memory-mb` for the link step — with the packages `base` does not ship and a ccache directory that survives the sandbox.
  4. `git bisect run` over the boot test, with the log saved. This is the deliverable. Everything above it exists to make this step cheap.
  5. Installed kselftests in the probe rootfs, so the harness can answer "and does it still work" as well as "does it boot".
  6. An honest note in the README about which subsystems this harness covers, so the next person does not spend a week trying to reproduce a driver bug in a machine that has no such device.
A boot test is one bit of information. The entire engineering problem is making that bit cost less than the engineer's attention.

Which is, in the end, the whole argument. Nobody needed a cleverer bisect algorithm; binary search was already optimal. What was missing was a machine you could boot two hundred times before lunch and a harness honest enough to tell you when the answer it gave you does not apply.

Frequently asked questions

Can I boot a kernel I just built inside a PandaStack sandbox?

No, and this is the limitation to internalise before designing anything. Our agent selects the guest kernel by globbing its data directory for `kernels/vmlinux-*` and taking the last match: one kernel per host, for every template running on it. The template bake already records a `"kernel"` field in `meta.json`, but nothing reads it at restore time, so adding a second kernel to a host today would silently repoint every template and invalidate every baked snapshot. On top of that, Firecracker does not expose virtualisation extensions to its guests, so you cannot run an accelerated VMM inside a sandbox to boot your own kernel either. What a sandbox is genuinely good for in this workflow is the build half — a snapshot-restored toolchain in 179 ms, 8 burstable vCPUs, and a rootfs you sized for an object tree — plus userspace tests that run against the platform's kernel. For the boot half you want a host with `/dev/kvm`, or QEMU with TCG inside the sandbox if you can accept software-emulation speed for what is ultimately a one-bit answer.

Is Firecracker a reasonable harness for kernel testing at all, given how little hardware it presents?

For a large and important slice of kernel work, yes, and the slice is defined by subsystem rather than by preference. Core kernel, scheduler, memory management, the networking stack, filesystems and the VFS, cgroups, namespaces, seccomp, BPF, and the virtio drivers themselves all run against code paths the machine genuinely has, and they benefit enormously from a boot that costs a fraction of a second. For device drivers, GPU work, PCI passthrough or VFIO, or any bug that needs a hardware quirk to reproduce, it is the wrong harness and QEMU is the right one — QEMU's whole value proposition is being able to pretend to be hardware you do not have. Note that the old blanket statements about Firecracker have gone stale in one direction: current versions have an opt-in virtio-PCI transport behind `--enable-pci`, with hot-plug in developer preview, and ACPI is supported. That widens the device model slightly; it does not turn it into a PC. Read the docs for the version you run, because this area changes across releases.

My kernel boots under QEMU but not under Firecracker. What is usually wrong?

Almost always the config, and almost always because `defconfig` builds a kernel for a machine Firecracker is not. The usual suspects: virtio over MMIO rather than PCI (Firecracker's default transport declares devices on the kernel command line instead of presenting a bus to walk), the 8250 serial console, which is the only output path you have and therefore the only way to see why the boot failed, and the KVM guest support bits. The fix is not to debug `defconfig` but to start from Firecracker's own guest configuration in `resources/guest_configs/` and run `make olddefconfig`. Check the artifact format too: on x86_64 Firecracker loads an uncompressed ELF `vmlinux` (recommended) or a `bzImage`, and on aarch64 a PE-format `Image` from `arch/arm64/boot/` — the `vmlinux` target is also the faster thing to build, since you skip the compressed image and the module tree. And if the symptom is a hang rather than a message, add `panic=1` so a panic exits the VM instead of looping, which converts a timeout into an answer.

What about KVM unit tests and nested-virtualisation tests?

Those need a machine that can be a hypervisor, which a Firecracker guest cannot be. Firecracker does not advertise VMX or SVM to its guests and provides no `/dev/kvm` inside them; this is a deliberate design property, not a flag waiting to be flipped, and the error you get is the truth rather than a permissions puzzle. So kvm-unit-tests, nested-virt selftests, and any job that boots an accelerated VM belong on bare metal or on a VMM that explicitly exposes nested virtualisation to its guests. The one workaround worth evaluating is QEMU with TCG rather than KVM: software emulation never needed hardware virtualisation, so it works inside a microVM, and for a short boot check that is often an acceptable trade. It is genuinely slow, so measure it on your own kernel before you build a pipeline on it. Note also that nested setups can look healthy and be quietly slow for reasons that have nothing to do with your patch — if you benchmark inside one, count the layers before blaming the code.

Does snapshot-restore help kernel CI?

For the boot half, no, and a post that claimed otherwise would be selling something. A snapshot is a memory image of a guest that has already booted a specific kernel, and kernel CI boots a new kernel every time by definition — the kernel is the variable under test, so restoring a previous one is the exact thing you must not do. You are cold-booting deliberately, which is the one workload where our fast path does not apply. For the build half it helps a great deal, because everything that is not the kernel is identical on every step: the toolchain, the compiler cache, the source tree, the installed selftests, the probe rootfs. Bake those into a template and each build sandbox starts from a restored snapshot rather than a boot — p50 179 ms, p99 203 ms — which matters on a harness that creates hundreds of them a day. The rule of thumb is to bake what is constant across the bisect and vary only the thing under test.

Keep reading

Related posts

  • Guest kernel lockdown and module loading in Firecracker microVMs

    Root in the guest is not a breach, it is Tuesday. But what root can then do to that kernel is the first half of every escape chain — and in a snapshot-restore world, the hardening has to be baked in, because nothing runs at boot ever again.

  • Firecracker boot_args, argument by argument

    Everyone copies the same magic `boot_args` string from the Firecracker docs and never reads it. It's a short, unusually honest description of what a microVM is — and what it has decided not to be.

  • The PVH Boot Protocol: How Firecracker Skips Firmware

    A normal x86 VM spends its first second reminiscing about 1981: firmware, option ROMs, real mode, a bootloader. Firecracker declines the nostalgia. It loads the kernel directly and jumps straight into it — via the Linux 64-bit boot protocol or the PVH entry point. Here's exactly how, ELF note and all.

  • What Runs as PID 1 Inside a MicroVM (and Why It Matters)

    The kernel boots, mounts a rootfs, and executes exactly one program. That program's job description is short, strange, and easy to get wrong — which is why sandboxes hang, leak zombies, lose their last log line, or take thirty seconds to stop.

  • Why your visual regression tests are flaky

    A visual regression suite that cries wolf gets muted within a month. Almost every false positive traces to the rendering environment, not the code — and that's fixable.

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.