Debugging Inside a MicroVM: gdb, strace, and Core Dumps
A worker in one of my sandboxes died at 02:14. I found out at 09:30, from a dashboard, and what I had was an exit code: 139. That is 128 plus 11, so I knew it was `SIGSEGV`, and that was the complete inventory of what I knew. The sandbox was gone. The idle reaper had taken it around 02:45, the rootfs clone had been unlinked, and the stack that would have told me which of eleven thousand lines did it had never been written down anywhere.
I build PandaStack, an open-source Firecracker microVM platform, so I get to make this mistake at a volume most people do not. What I have slowly accepted is that it is not one mistake. Ephemeral infrastructure is hostile to debugging in four specific, separable ways, and each has a countermeasure you have to put in place before the failure rather than after.
The tooling below applies to any Linux box. The countermeasures are PandaStack-flavoured because that is the platform whose source I can read, and I read it for this post rather than repeating what I assumed. Two things I assumed turned out to be wrong, and they are both in here.
The four places your evidence goes to die
These are not four descriptions of one problem. They fail independently, they have different fixes, and in the incident above I managed three of the four at once.
1. The machine is gone before you get there
Something deleted it. The candidates, in rough order of how often they are the culprit: a TTL, an idle reaper, a wall-clock lifetime cap on your plan, an autoscaler, or -- most commonly, and most embarrassingly -- your own `finally` block.
It is worth being precise about the TTL, because the mental model is usually wrong. On PandaStack `ttl_seconds` is an IDLE TTL, not a countdown from create. The reaper compares the time since the sandbox's last activity against its TTL; a busy sandbox is never reaped by it. So `ttl_seconds=1800` does not promise you thirty minutes of life, it promises deletion thirty minutes after the last thing you did -- and a crashed workload with nobody poking it goes quiet immediately. The clock that kills your evidence starts running at the moment the thing you wanted to debug stopped working. There is a separate hard wall-clock lifetime cap on capped tiers, which is what it sounds like and is not negotiable from the API.
Worse, two details about what counts as activity will catch you specifically while you investigate. Read-only inspection GETs -- the status row, `/lifecycle`, `/metrics`, `/logs`, `/events`, `/ports`, and the filesystem METADATA routes `/fs/stat` and `/fs/dir` -- deliberately do not bump the idle clock. That gate exists for a good reason (a dashboard tab polling every thirty seconds used to reset the TTL faster than it could ever elapse, so nothing was ever reaped) but it means an hour spent reading a wedged sandbox's logs does nothing to keep it alive. Genuine use of the guest does bump: an `exec`, an actual byte read through `/fs`, a PTY or SSH session, a fork, a snapshot. And the bump happens once per request, at the start -- so a single forty-minute `exec` under a thirty-minute idle TTL is a job that gets reaped out from under itself at minute thirty-one.
The cheap countermeasure is better than the obvious one. The obvious one is to keep the machine: `set_ttl()` widened, and `set_persistent(True)` to make the reaper skip the sandbox entirely. That works, and it bills by the GiB-hour until somebody remembers. The better default is to make the evidence outlive the machine -- take a `snapshot()`, and then let the reaper have the VM. The idle reaper's delete path deliberately does NOT cascade-delete the sandbox's snapshots, because a snapshot is meant to outlive its sandbox; it leaves them as orphans that a separate sweep only reclaims after a grace period defaulting to seven days. So "snapshot, then let it be reaped" costs you nothing after the snapshot and still leaves you a full machine image to restore on Monday.
Keep the live VM only when you actually need to attach to the running process -- an interactive gdb session, a `gcore` of a process you do not want to kill. Note then that your own debugging keeps it alive by accident: every exec bumps the clock, so it is the gaps between your commands that are dangerous, not the commands. And one honest caveat on the persistence flag: `set_persistent(True)` is refused with a 403 on any tier that has a hard lifetime cap, because otherwise a capped sandbox could be created and immediately promoted out of its own cap. Catch it and fall back to the extended TTL.
2. The evidence is inside a filesystem nobody will mount again
A sandbox rootfs is a copy-on-write clone. When the sandbox goes, so does the clone, and there is no `/var/lib/docker/overlay2` equivalent to go spelunking in afterwards. Anything you have is something you pulled out while the guest was alive.
Which makes the discipline boring and non-negotiable: artifacts go to one known path, decided at design time, the same one every run. Not a temp directory your code chose at runtime, not the cwd of whichever process happened to launch -- both of those are things you then have to work out under pressure while the clock runs.
Getting bytes out is `sbx.filesystem.read(path)`, which underneath is a `cat` over the SSH bridge, buffered whole in the agent and again in your client. Two consequences worth knowing before an incident rather than during one. The Python client's default HTTP timeout is 30 seconds and `read()` does not override it, so a large pull dies on the clock rather than on the content -- raise it with `set_default_client(Client(timeout=300))`. And writes in the other direction are capped at 32 MiB per request, which matters the day you try to push a debug-symbol package in.
3. The logs you are looking at are the wrong logs
This one is a genuine and frequently-confused distinction, and I had it slightly wrong myself until I read the handler. There is no single log. There are four streams, they contain different things, and the one most people reach for is the one least likely to contain their program's output.
| Stream | What is actually in it | How you read it |
|---|---|---|
| Serial console capture | The VM's serial console: kernel messages, plus anything the guest writes to /dev/console. The guest boots with console=ttyS0. The capture file is truncated every time the VM starts or restores. | GET /v1/sandboxes/{id}/logs, or sbx.logs(follow=True). This is what the endpoint serves when a console capture exists -- which it does on every path. |
| firecracker.log | The VMM's own log: device errors, API-socket complaints, snapshot-load failures. The log of the hypervisor, not of anything inside the guest. | The same endpoint, as the fallback when there is no console capture. It is the log you want when the failure took the whole guest down. |
| App runtime log | An app's stdout and stderr, captured to the guest file /var/log/pandastack-app.log by the deploy pipeline's launcher. | GET /v1/apps/{id}/runtime-logs, with ?follow=1 to stream. Apps only. |
| Your own redirect | Whatever your process wrote to wherever you pointed it. For a plain sandbox this is the only place your program's output exists, because nothing captures it for you. | You read the file: exec a tail, or filesystem.read() it. Nobody else is going to. |
Two consequences of that truncation. First: every normal create here is a snapshot restore rather than a boot, so a fresh sandbox's console capture starts empty and contains no kernel boot messages at all -- those happened once, at template bake time, in a VM you never saw. If you are hunting a boot problem you are hunting it in the wrong artifact. Second: a restore truncates, so hibernate-and-wake cycles quietly discard the previous generation's console output.
One free trick falls out of the mechanism. Because the guest's kernel command line carries `console=ttyS0`, anything written to `/dev/console` inside the guest lands in that capture, and therefore in `sbx.logs()`. It is a perfectly good out-of-band marker channel for a process whose real output is going somewhere you cannot reach yet.
# Lands in sbx.logs() without touching your process's stdout at all.
echo "MARKER phase=post-migrate pid=$$ ts=$(date -Is)" > /dev/console4. The thing you want to attach to is a shell session that ended
An exec API is not a terminal. Each call is a fresh `/bin/sh -c` over the SSH bridge, running as root, and when it returns that shell is gone. Nothing carries between calls except the filesystem and whatever processes you deliberately detached. So `ulimit -c unlimited` in one call and the launch in the next is not a two-step setup; it is a no-op followed by a crash with no core.
Three countermeasures for three needs. For live output from something long-running, `exec_stream` with `on_stdout`/`on_stderr` callbacks -- server-sent events underneath, so you see output as it happens rather than at the end. For anything that must outlive one call, detach it: `setsid`, output redirected to a file you named, pid written somewhere you will look. And when you genuinely want a terminal -- an interactive gdb session, mostly -- there is a WebSocket PTY at `GET /v1/sandboxes/{id}/exec/pty`, xterm.js-compatible, which is what the dashboard terminal is built on.
One more trap in the same family: on the one-shot exec path, `timeout_seconds` is not the bound you think it is -- the real ceiling is your HTTP client's timeout. Put the bound inside the guest, enforced by a program whose entire job is enforcing it (`timeout --kill-after=5s 900 ...`, or `ulimit -t` for CPU time), and treat `ttl_seconds` as the platform backstop rather than the mechanism.
Core dumps, the long way round
A core dump is the highest-value artifact available from a crash, and the number of places the setup silently does nothing is genuinely impressive. Here is the whole chain, because skipping any link produces the same symptom: no file, no error, no explanation.
The limit, in the right shell
`RLIMIT_CORE` defaults to 0 nearly everywhere, which means the kernel is willing to write you a core of at most zero bytes. `ulimit -c unlimited` raises it for the current shell and anything it goes on to spawn, and for nothing else. Same call as the launch. Root can raise the hard limit, and exec in a PandaStack guest runs as root, so you will not hit that wall here -- you will hit it the moment you correctly drop privileges somewhere else.
The pattern, which decides where it goes
`/proc/sys/kernel/core_pattern` is a system-wide setting controlling the filename, and it has a second mode most people meet by accident: if the pattern begins with a pipe character, the kernel writes no file at all -- it runs the named program and streams the core to its stdin. That is how apport and systemd-coredump work, and it is a sensible design on a desktop.
In a minimal guest it is a trapdoor. The first-party templates here are built from the `ubuntu:24.04` image and install neither handler, and nothing in the platform writes `core_pattern`, so the value is the kernel's default: the bare relative name `core`. Relative means relative to the crashing process's working directory, and a fixed name means two crashes overwrite each other. Set an absolute pattern with expansions -- `%e` executable name, `%p` pid, `%t` unix timestamp. Read the file first, though, rather than trusting any of us about what it currently says.
The filter, which decides what goes in it
`/proc/<pid>/coredump_filter` is a bitmask over which VMA classes get dumped: anonymous private and shared, file-backed private and shared, ELF headers, huge pages, DAX. The default omits some of them to keep cores small, which is right until the thing you need is in a shared mapping or a huge-page region. `0x3f` includes the lot. Per-process, inherited across fork, so set it on the launcher and the children get it.
The part everyone forgets: the core is not enough
A core file is a memory image plus some notes, and on its own it is unreadable. To turn it into a stack trace you need the exact binary that produced it, its debug info, and the shared objects it had mapped -- matched by build-id, not by sharing a version number in the filename. In an ephemeral guest the binary was built in that guest or pulled into it, and when the guest goes so does the only copy with a matching build-id. So the artifact is never the core. It is the core plus the binary plus anything under `/usr/lib/debug` plus the libraries plus the mapping layout, tarred together.
The mappings come two ways: from `/proc/<pid>/maps` while the process is alive, or out of the core's own `NT_FILE` note afterwards, which is what gdb's `info proc mappings` reads and which needs no live process at all.
# Inside the guest. Every exec here is a FRESH `/bin/sh -c` over the SSH
# bridge, as root -- so `ulimit` and the launch must be in the SAME call.
mkdir -p /var/crash/artifacts
# 1. WHERE the kernel writes it. Read this before trusting anyone, me
# included -- a `|` prefix pipes the core to a handler process, and this
# guest has neither apport nor systemd-coredump installed to receive it.
cat /proc/sys/kernel/core_pattern
# So: an absolute path. %e = exe name, %p = pid, %t = unix time.
# A relative pattern (the kernel default is the bare name `core`) lands
# in the process's cwd, which is never where you looked.
echo '/var/crash/core.%e.%p.%t' > /proc/sys/kernel/core_pattern
# 2. WHAT goes in it. The default filter omits some shared and huge-page
# mappings; 0x3f includes them. Per-process, inherited across fork.
echo 0x3f > /proc/self/coredump_filter
# 3. Raise the limit and launch IN THE SAME SHELL, detached, with its own
# log file -- nothing on the platform captures a sandbox's stdout for you.
ulimit -c unlimited
cd /srv/app
setsid sh -c 'exec ./bin/worker --jobs 4' >/var/log/worker.log 2>&1 &
echo $! > /run/worker.pid
# ... later it dies with 139, which is 128 + SIGSEGV ...
# 4. The bundle. A core alone is a bag of bytes: gdb needs the EXACT binary,
# its debug info and the objects it had mapped, matched by build-id.
cd /var/crash
CORE=$(ls -t core.worker.* 2>/dev/null | head -1)
APP=/srv/app/bin/worker
cp "$CORE" "$APP" artifacts/
readelf -n "$APP" | grep -i 'build.id' > artifacts/binary.txt
mkdir -p artifacts/libs
for so in $(ldd "$APP" | awk '/=>/ && $3 ~ /^\// {print $3}'); do
cp --parents "$so" artifacts/libs/
done
[ -d /usr/lib/debug ] && cp -a /usr/lib/debug artifacts/ 2>/dev/null
# Mappings: from the live process if it is still there, otherwise out of
# the core's own NT_FILE note (see `info proc mappings` below).
PID=$(cat /run/worker.pid 2>/dev/null)
[ -d "/proc/$PID" ] && cp "/proc/$PID/maps" artifacts/
# 5. Extract the ANSWER before you think about extracting the core. This
# file is kilobytes; the core is gigabytes.
gdb --batch -q \
-ex 'set pagination off' \
-ex 'thread apply all bt full' \
-ex 'info registers' \
-ex 'info sharedlibrary' \
-ex 'info proc mappings' \
"$APP" "$CORE" > artifacts/backtrace.txt 2>&1
tar -C /var/crash -czf /var/crash/evidence.tar.gz artifacts
ls -lh /var/crash/evidence.tar.gz artifacts/backtrace.txtNow the honest part. A core from a multi-gigabyte process is a multi-gigabyte file, and the extraction path is a `cat` over an SSH bridge, buffered in memory twice, inside an HTTP request. That is not a transport for three gigabytes. The practical move is step 5: run gdb in the guest and pull the kilobyte-sized backtrace. If you genuinely need the core itself, compress it where it lives (cores are mostly zeroes and compress absurdly well), serve it over the sandbox's own preview URL, or push it to your own object store from inside the guest -- sandbox egress is open by default, so a `gsutil cp` to a bucket you control skips the control plane entirely.
gdb in the guest
First, the thing I wish someone had told me: no first-party template ships gdb or strace. `base` carries build-essential, which is a compiler toolchain, not a debugger. So the first line of any debugging session is an `apt-get install`, which costs you forty-odd seconds of an incident. If you do this more than occasionally, bake a template with the tools in it -- that is a Dockerfile and a `template build` call, and the baked snapshot restores at the same p50 179 ms as everything else.
Second: you have an exec API, not a terminal, so the interesting mode is batch. `gdb --batch` runs a list of `-ex` commands and exits, which turns a debugger into something a script can call and whose output you can capture:
PID=$(cat /run/worker.pid)
# Every thread's stack, from a live process, non-interactively.
gdb --batch -q -p "$PID" \
-ex 'set pagination off' \
-ex 'thread apply all bt' \
-ex 'info threads' \
> /var/crash/artifacts/live-bt.txt 2>&1
# A core WITHOUT killing the process. gcore ships with gdb: it attaches,
# dumps, detaches, and the process carries on. This is the single most
# useful command in this post -- it gets you the artifact while the
# workload keeps running, so you are not trading evidence for uptime.
gcore -o /var/crash/core-live "$PID"
# Read a core, offline, with the binary beside it.
gdb --batch -q -ex 'thread apply all bt full' ./bin/worker /var/crash/core-live.$PID`thread apply all bt` is the workhorse -- a deadlocked process tells you almost everything in one shot, because two stacks sitting in `futex_wait` on each other's locks is a readable picture. `bt full` adds locals and gets long fast. `info sharedlibrary` is how you confirm gdb actually found symbols rather than printing a column of `??` at you, which is the failure mode that wastes the most time because it looks like output.
For an interactive session, use the WebSocket PTY rather than fighting the exec API: gdb's command loop wants a terminal, and `GET /v1/sandboxes/{id}/exec/pty` gives it one.
gdbserver and the preview URL: the correction
I had this one wrong, so here is the corrected version. The appealing idea is to run `gdbserver` in the guest, expose its port on the sandbox's tokenless preview URL (`https://<port>-<sandbox-id>.<suffix>`, where the sandbox UUID is the credential), and attach a local gdb or an IDE. It is a lovely idea and it does not work, for a mechanical reason: the preview route terminates as an HTTP reverse proxy, and gdb's remote serial protocol is raw TCP that is not HTTP and does not upgrade from it. There is nothing for the proxy to forward. The only raw-TCP tunnel on the platform is deliberately allow-listed to managed-database sandboxes and two Postgres ports, specifically so it cannot become a generic guest-port primitive.
What the preview URL does carry is HTTP and HTTP upgrades, so the debug surfaces that work over it are the HTTP-shaped ones: a Go service's `net/http/pprof` endpoints, a `python3 -m http.server` over your evidence directory, a browser-based debugger front end that speaks HTTP and WebSocket rather than the raw protocol. Verify any of those on your own deployment before you need them, rather than discovering the shape of the proxy at 03:00 as I did.
# The preview route is an HTTP reverse proxy, so it carries HTTP (and HTTP
# Upgrade) but NOT the raw GDB remote serial protocol -- see above.
# A Go service with net/http/pprof on :6060. Plain HTTP, so it just works:
url = sbx.preview_url(6060)
print(f"{url}/debug/pprof/goroutine?debug=2") # every stack, as text
print(f"go tool pprof {url}/debug/pprof/profile?seconds=30")
# Or serve the evidence directory, instead of dragging a 3 GB core through
# the control plane:
sbx.exec("setsid python3 -m http.server 8000 -d /var/crash "
">/var/log/http.log 2>&1 & echo ok")
print(sbx.preview_url(8000))
# That URL's only credential is the sandbox UUID and it needs no login.
# A pprof endpoint is a read. A debugger front end is remote code execution.
# Decide which of those you just pasted into a Slack channel.And a real decision rather than a disclaimer: that URL needs no login and is live for the sandbox's lifetime. The proxy at least strips your platform `Authorization` header before forwarding, so you cannot accidentally hand your API token to a process in the guest -- but what you put on that port is your call, and "paste the debug URL into the incident channel" is a thing to do on purpose.
ptrace, strace, and the oldest joke in debugging
`strace` answers a different question from gdb. gdb tells you where the program is; strace tells you what it is asking the kernel for, which is usually the faster route to "it is blocked on a connect to a host that does not resolve" or "it is reading from a pipe nobody will ever write to".
# Launch under it. -f follows forks (essential: the interesting process is
# usually a child), -tt gives microsecond wall-clock stamps, -T records how
# long each call took, -s 256 stops truncating strings at 32 characters,
# and -o writes to a file so the trace is not competing with the output.
strace -f -tt -T -s 256 -o /var/crash/artifacts/trace.txt ./bin/worker --jobs 4
# Attach to something already running (-p). Keep it bounded: an
# investigation command that hangs is its own small incident.
timeout 30 strace -f -tt -T -p "$(cat /run/worker.pid)" -o /tmp/attach.txt
# Cut the volume with -e trace=, by class or by name. A full trace of a
# busy process is tens of megabytes a minute and mostly futex noise.
strace -f -e trace=%network,%file -o /tmp/net-fs.txt ./bin/worker
strace -f -e trace=openat,connect,execve -o /tmp/three.txt ./bin/worker
# Counts instead of lines -- the cheapest first look at a process that is
# busy rather than stuck.
timeout 20 strace -c -f -p "$(cat /run/worker.pid)"Then the two gotchas that actually bite.
The first is `ptrace_scope`. Attaching to a process that is not your own descendant is privileged, and on kernels carrying the Yama module it is governed by `/proc/sys/kernel/yama/ptrace_scope`: 0 is classic permissive, 1 restricts you to descendants, 2 requires admin capability, 3 disables it entirely with no way back short of a reboot. Read it rather than assuming -- and if the file does not exist, that is the answer too, because it means the module is not compiled into this kernel. In a PandaStack guest, exec runs as root, so `CAP_SYS_PTRACE` carries you past a scope of 1 anyway; the day it bites is the day you correctly dropped privileges inside the guest and then tried to attach as the unprivileged user.
The second is the joke. `strace` is a ptrace-based tool, so it stops the tracee at every syscall entry and exit and resumes it once the tracer has looked -- two context switches per syscall, on every syscall. A syscall-heavy process can run an order of magnitude slower under it, and the thing you were chasing was a race. So the bug does not reproduce, and you have learned something real but not what you wanted: that it is timing-dependent. That is the oldest joke in debugging and it is still funny to everyone except the person on call.
When the observer effect is the problem, change instrument. `ltrace` traces library calls instead of syscalls and pays the same ptrace tax. `perf trace` and eBPF tooling are the genuinely low-overhead answers, because they sample or attach in the kernel rather than stopping the process. The caveat is real and I will not wave it away: both depend on kernel configuration, the guest kernel here is 5.10, and whether `perf` and the tracepoints you want exist is a question to settle on the actual guest, on a quiet day, before you plan a workflow around it.
The wrapper: debug on failure, not on success
Everything above assembles into one pattern that I now put in front of anything whose failures are expensive to reproduce. Run the workload. On a zero exit, tear the sandbox down and move on. On a non-zero exit, stop deleting things and start collecting.
# debug_on_failure.py -- run a workload; keep the machine only if it fails.
import pathlib
from pandastack import Sandbox, Client, set_default_client
# filesystem.read() does NOT override the client's HTTP timeout, and the
# default is 30 seconds -- which an artifact pull will blow straight through.
# (snapshot() is the exception: it sets its own 180s timeout.)
set_default_client(Client(timeout=300))
SETUP = pathlib.Path("setup_and_launch.sh").read_text() # the bash above
ARTIFACT = "/var/crash/evidence.tar.gz"
LOCAL = pathlib.Path("./evidence.tar.gz")
# Keep the whole VM alive for interactive attach, or just keep the snapshot?
# The snapshot is free after it is written; the VM bills by the GiB-hour.
NEED_LIVE_PROCESS = False
sbx = Sandbox.create(
template="base",
ttl_seconds=1800, # IDLE ttl: 30 min with no activity
metadata={"job": "nightly-worker", "run": "4412"},
)
print("sandbox:", sbx.id)
def sh(cmd: str, label: str) -> int:
# Stream it. exec_stream widens the HTTP timeout to at least 900s;
# one-shot exec() is bounded by the client timeout set above. The only
# bound I actually trust is the in-guest `timeout` in the command itself.
return sbx.exec_stream(
cmd,
on_stdout=lambda c: print(f"[{label}] {c}", end=""),
on_stderr=lambda c: print(f"[{label}!] {c}", end=""),
)
# No first-party template ships gdb or strace. Install them, or bake a
# template that has them and stop paying for this on every run.
sh("apt-get update -qq && apt-get install -y -qq gdb strace", "setup")
sbx.filesystem.write("/srv/app/setup_and_launch.sh", SETUP)
code = sh("timeout --kill-after=5s 900 sh /srv/app/setup_and_launch.sh", "run")
if code == 0:
sbx.kill()
raise SystemExit(0)
# ---- FAILURE PATH. There is no kill() below this line. ----
print(f"exit {code} -- preserving sandbox {sbx.id}")
# FIRST, make the evidence outlive the machine. Full VM: guest RAM verbatim
# plus the disk, at the moment of the failure. The idle reaper's delete does
# NOT cascade to snapshots, so this survives the VM being reaped -- which an
# explicit kill() would not.
snap = sbx.snapshot()
print("snapshot:", snap)
# THEN widen the idle TTL, and only take the persistence flag if you need the
# live process to attach to. A persistent VM bills until someone deletes it.
sbx.set_ttl(86400)
if NEED_LIVE_PROCESS:
try:
sbx.set_persistent(True) # 403 on a tier with a lifetime cap
except Exception as e: # noqa: BLE001 -- keep going either way
print("persistent refused, running on the extended TTL only:", e)
# Get the small thing first.
bt = sbx.exec("tail -c 20000 /var/crash/artifacts/backtrace.txt")
print(bt.stdout or bt.stderr)
# Then the bundle, if it is a size a `cat` over SSH can reasonably carry.
size = int(sbx.exec(f"stat -c %s {ARTIFACT} 2>/dev/null || echo 0").stdout or 0)
if 0 < size < 200 * 1024 * 1024:
LOCAL.write_bytes(sbx.filesystem.read(ARTIFACT))
print(f"wrote {LOCAL} ({size} bytes)")
else:
print(f"{ARTIFACT} is {size} bytes -- analyse it in the guest instead")
print(f"restore: Sandbox.create(from_snapshot={snap!r})")
if NEED_LIVE_PROCESS:
print(f"attach: sandbox {sbx.id}, gdb -p $(cat /run/worker.pid)")
print("COST: this VM bills by the GiB-hour until someone deletes it --")
print(" and deleting it CASCADES to the snapshot above. Restore first.")
else:
print(f"the VM will be reaped when idle; snapshot {snap} outlives it")The shape is the point. There is exactly one `kill()` in the whole script and it is on the success path. The failure path snapshots first -- the artifact that outlives the machine -- then widens the TTL, pulls the small thing, and prints the restore command. It only pins the VM into persistence when something genuinely needs the live process, because a forensic VM nobody deletes is a line item that grows while you sleep, and deleting it later takes the snapshot with it.
What a snapshot gives you that a container does not
This is the part a microVM does that a container cannot, and it is worth being precise about which call does what, because getting it backwards has shipped in posts before -- including mine.
`sbx.snapshot()` captures the full virtual machine: guest RAM verbatim, plus the disk. Not a filesystem layer, not a checkpoint of one process tree -- the memory image of the machine, including the wedged process, its locks, its heap and its socket state. The SDK gives that call a 180-second timeout rather than the default 30, because writing a multi-gigabyte memory image takes real time. What you get back is a durable artifact you can restore days later into a different sandbox. That is a reproduction you did not have to construct, which is the single most valuable object in debugging.
Then the fork distinction, stated plainly:
- sbx.fork_tree(n) is the call that inherits memory, and so the one that gives you N copies of the same frozen process state -- N hypotheses tried against an identical wedge rather than against N slightly different ones. Hard-capped at 16 children per call: ask for more and you get an error, not a truncation. For a wider tree, grow it breadth-first, popping a node from the frontier and calling fork_tree(min(16, remaining)) on it.
- sbx.fork() clones the DISK ONLY, and the child cold-boots. It does not inherit the parent's memory, so it does not give you the frozen process -- it gives you a machine with the same filesystem and its own fresh entropy. Useful, and not the tool for this job. Reach for fork() expecting the stuck thread to still be stuck and you get a clean boot and a confusing hour.
- Latency: same-host fork 400-750 ms, cross-host 1.2-3.5 s. Both are the restore path, which is also why an ordinary create lands at p50 179 ms with no warm pool involved.
One artefact of restore will mislead you specifically in a trace: the guest clock is re-synced to host time on restore, resume and wake. That is deliberate -- it was added after a frozen guest clock produced TLS failures that looked like a network problem -- but it means wall-clock timestamps jump across a restore boundary. In an `strace -tt` log that reads as a multi-minute stall inside a syscall that actually returned instantly. It is not a stall. Reconcile against monotonic time, or just know where your restore boundaries are.
If the failure is intermittent rather than fatal, the reproduction half of this story is its own post: Snapshot the Failure, Not the Log Line. And if the process is merely stuck rather than dead, triage before reaching for any of this machinery -- How to Debug a Hung Process in a Sandbox is the cheapest-first order of operations, and ninety seconds spent on process state beats an hour of core-dump plumbing.
The honest limits
- A forensic VM bills until someone deletes it. CPU is metered on CPU-seconds actually burned, so an idle crashed sandbox is cheap there, but memory bills as committed GiB-hours at $0.0162 -- so a 4 GiB base sandbox you pinned persistent and forgot is about $1.56 a day, forever. A workflow that keeps machines on failure and has no reaper of its own is a budget incident with a delay fuse. Prefer the snapshot to the pinned VM, tag what you keep in metadata, delete it on a deadline.
- Observing a sandbox does not keep it alive. The read-only inspection GETs do not bump the idle clock -- a deliberate gate, and a correct one, but it means the reaper can delete the thing whose logs you are reading while you read them. Snapshot before you investigate, not after you have a theory.
- strace changes the timing of the bug you are chasing. Not a caveat -- a property of ptrace, and for a whole class of races it means the tool cannot see the bug. Accept it, note that non-reproduction under strace is itself a finding, and switch instruments.
- perf and eBPF depend on guest kernel configuration you must verify rather than assume. The guest kernel here is 5.10, and I am not going to tell you which tracepoints or BPF features are compiled in, because the honest answer is that you check on the guest you actually run, on a day when it does not matter.
- A big core dump is genuinely awkward to extract. The fs path is a cat over SSH buffered twice; it is not a bulk transport. The alternative -- analysing in-guest and extracting only a backtrace -- is better most of the time anyway, and the only reason I can say that with a straight face is that it is also cheaper.
- A snapshot is guest RAM verbatim, which makes a forensic snapshot a secret-bearing artifact. Every token, key and decrypted buffer the process held at the moment of failure is in that file -- exactly why it is useful, and exactly why it should not sit in a shared bucket with a cheerful name. The Security Gotchas of Firecracker Snapshots (Secrets Frozen in RAM) is the long version.
- Your own tidy-up is the most destructive thing in the system. An explicit delete cascades to that sandbox's snapshots; the idle reaper does not. I have deleted my own evidence with a finally block more than once.
- None of this helps if the failure took the whole guest down rather than one process. A kernel panic, an OOM that ate init, a device error -- no process left to attach to, no core to collect, and you are reading firecracker.log, which is, with some irony, the one log that endpoint will definitely give you.
The summary
Ephemeral infrastructure is hostile to debugging in four ways, and they are separable. The machine is gone before you arrive, so decide in advance that a failure is a reason to keep the evidence: snapshot first, because the snapshot outlives the reaper, and pin the VM itself only when you need the live process. The evidence is in a filesystem nobody will mount again, so write everything to one known path and pull it while the guest is alive. The logs you are reading are probably the wrong logs -- learn which of the four streams you are looking at, and remember that for a plain sandbox nothing captures your process's stdout except you. And the thing you want to attach to is a shell session that ended, so detach what must survive and stream what you want to watch.
Then the mechanics, in the order that works. `ulimit -c unlimited` in the same shell as the launch. An absolute `core_pattern` with `%e.%p.%t`, because a pipe pattern in a minimal guest pipes into nothing. `coredump_filter` when what you need is in a shared or huge-page mapping. The binary, its debug info and the libraries collected alongside the core, because a core without them is a bag of bytes. `gcore` to get an artifact without killing the process. `gdb --batch -ex 'thread apply all bt'` to turn a debugger into something a script can call. And `strace -f -tt -T -s 256 -o`, knowing you may have just made the race disappear.
The microVM payoff is the one at the end: a snapshot is the whole machine, memory included, so you can freeze a wedged process and come back to it tomorrow, and `fork_tree` gives you up to sixteen identical copies of that wedge to interrogate in parallel. A container cannot do that, and it is the first thing I have used that makes a heisenbug feel tractable rather than merely annoying.
Frequently asked questions
Why doesn't sbx.logs() show my program's output?
Because it is not that log. The endpoint serves the VM's serial console capture -- kernel messages and anything the guest writes to /dev/console -- falling back to the hypervisor's own firecracker.log when there is no console capture. Neither is your process's stdout. For a plain sandbox, nothing on the platform captures your process's output: if you did not redirect it to a file, it went to the stdout of an exec session that has since exited, and it is gone. Redirect it deliberately to a path you chose -- setsid sh -c 'exec ./worker' >/var/log/worker.log 2>&1 & -- then read that file with an exec tail or filesystem.read(). Apps are the exception: the deploy pipeline's launcher captures app stdout and stderr to the guest file /var/log/pandastack-app.log and serves it at GET /v1/apps/{id}/runtime-logs, which is a different endpoint from the sandbox one for exactly this reason. Two further details about the console stream, since they surprise people. The capture file is truncated every time the VM starts or restores, and every ordinary create is a restore rather than a boot -- so a fresh sandbox's console log starts empty and holds no kernel boot messages at all. The useful corollary: a line written to /dev/console inside the guest does land in that stream, which makes it a decent out-of-band marker channel while your real output is somewhere you cannot reach yet.
Can I attach a local gdb or an IDE debugger to a process inside a sandbox?
Not over the sandbox's preview URL, and I want to be exact about why because the idea is appealing enough that I believed it myself. The preview route terminates as an HTTP reverse proxy. gdb's remote serial protocol -- what gdbserver speaks -- is raw TCP that is neither HTTP nor an HTTP upgrade, so there is nothing for the proxy to forward. The only raw-TCP tunnel on the platform is deliberately allow-listed to managed-database sandboxes and two Postgres ports, precisely so it cannot become a generic guest-port primitive. What does work: run the debugger in the guest and drive it with gdb --batch -ex, which is the right shape for an exec API anyway, or use the WebSocket PTY at GET /v1/sandboxes/{id}/exec/pty when you want a real terminal. And for HTTP-shaped debug surfaces the preview URL is genuinely good -- a Go service's net/http/pprof endpoints, or a plain python3 -m http.server over your evidence directory. Remember what that URL is, though: its only credential is the sandbox UUID, it needs no login, and it is live for the sandbox's lifetime. A pprof endpoint is a read. A debugger front end is remote code execution for anyone holding the link.
How do I get a multi-gigabyte core dump out of a sandbox?
Preferably you do not. The filesystem read path is a cat over the SSH bridge, buffered whole in the agent and again in the client, inside an HTTP request whose default timeout in the Python SDK is 30 seconds and which filesystem.read() does not override. You can raise that with set_default_client(Client(timeout=300)), and for a compressed hundred-megabyte bundle that is fine. For three gigabytes it is the wrong transport and no timeout fixes that. Three better options, in the order I reach for them. First: analyse in the guest and extract only the answer -- gdb --batch -ex 'thread apply all bt full' against the core produces a file measured in kilobytes, and nine times out of ten the backtrace is the whole investigation. Second: compress where it lives, since cores are mostly zero pages and compress absurdly well. Third: skip the control plane -- sandbox egress is open by default, so a gsutil cp or aws s3 cp from inside the guest to a bucket you own moves the bytes directly, at the cost of metered transmit egress and a credential you should think carefully about putting in a sandbox. And consider not producing the giant core at all: gcore dumps a live process without killing it, so you can take one from a smaller, earlier state rather than from something that has grown for six hours.
Does strace actually change the bug I am chasing?
Yes, and often decisively. strace works through ptrace, which means the kernel stops the traced process at every syscall entry and again at every exit, hands control to the tracer, and resumes it afterwards -- two context switches per syscall, on every single syscall. For a syscall-heavy workload that can be an order of magnitude slowdown, and it does not slow things uniformly: it specifically inflates the time in and around syscalls, which is exactly the window a lot of races live in. So a timing-dependent bug very often will not reproduce under strace. That is frustrating but it is not nothing -- non-reproduction under tracing is itself a finding, and it narrows your hypothesis space to timing. When you need to observe without perturbing, change instrument rather than squinting harder. strace -c gives aggregate counts at lower cost than a full trace; -e trace= narrows to the calls you care about and cuts both volume and overhead; and perf or eBPF tooling samples or attaches in the kernel rather than stopping the process, which is the genuinely low-overhead answer. The honest caveat on that last one: both depend on kernel configuration, the guest kernel here is 5.10, and you should verify what is available on your actual guest on a quiet day rather than discovering it during an incident.
If I snapshot a crashed sandbox and then delete it, do I keep the snapshot?
No, and this is the sharpest edge in the whole workflow. An explicit delete -- sbx.kill(), or DELETE on the sandbox -- cascade-purges every snapshot taken from that sandbox: the local bytes, the object-store copy, and the database row. The idle reaper's delete deliberately does not, because a snapshot is meant to outlive its sandbox; it leaves them as orphans, which a separate sweep reclaims only after a grace period that defaults to seven days. The practical consequence is counter-intuitive and genuinely useful once you internalise it: your tidy finally: sbx.kill() is strictly more destructive than the reaper you were trying to beat, and "snapshot the failure, then let the machine be reaped" is therefore a perfectly good forensics pattern. It costs nothing after the snapshot is written, and you still have a full machine image to restore on Monday. What you must not do is take the snapshot and then delete the sandbox to stop the billing, because that is the one action that destroys both. If you need the VM gone and the snapshot kept, restore the snapshot into a new sandbox first, or just let the reaper do it. Treat the seven days as a default rather than a promise -- it is configurable, so confirm it on the deployment you actually run. And remember what that snapshot contains: guest RAM verbatim, which is every token and key the process was holding at the moment it died. It is a secret-bearing artifact and deserves the same handling as a database dump.
Keep reading
- Snapshot the failure, not the log line — The reproduction half of this story: turning an intermittent failure into an artifact you can restore.
- How to debug a hung process in a sandbox — Cheapest-first triage. Start here when the process is stuck rather than dead.
- The security gotchas of Firecracker snapshots — Why a forensic snapshot is a secret-bearing artifact: RAM verbatim, keys included.
- How to stream command output from a sandbox — The exec_stream and SSE mechanics behind watching a long job instead of waiting for it.
- Firecracker guest kernel config — What is and is not compiled into the guest kernel -- the thing to check before planning around perf or eBPF.
Related posts
- How to profile code running in a sandbox
Attaching a profiler to a pod is muscle memory. Inside a microVM, half of that muscle memory works and the other half quietly wastes your afternoon.
- How to tail logs and debug a running app
"Check the logs" is unhelpful advice when there are three kinds. Here's how to tell which one holds the answer, before you start guessing.
- Replaying and Debugging Webhooks in Disposable MicroVMs
You captured a stream of real Stripe/GitHub/Shopify webhook payloads and now you want to replay them against a candidate build to reproduce a bug — without a poison payload wiping your host on the way. Fork a fresh microVM per replay and let the blast radius die with it.
- Your Firecracker Snapshot Restore Failed: A Field Guide
A restore either works in tens of milliseconds or fails with an error that tells you almost nothing. Here are the seven classes of Firecracker snapshot restore failure, what each looks like, and how to bisect your way to the real cause.
- Debugging a Firecracker microVM That Won't Boot
A microVM that boots to a silent black hole is Firecracker's way of saying you assumed the rootfs, didn't you. Turn on the serial console and let the kernel tell you what actually went wrong.
49ms p50 cold start. Fork, snapshot, and scale to zero.