Flamegraphs Inside a microVM: When perf Has No Counters to Count
The flamegraph was 1,400 pixels wide and told me nothing. One box spanned ninety-six percent of it, labelled `[unknown]`. Underneath, a tidy little staircase of hex addresses. Above, nothing. The customer's note said "profiled it like you said, here's the graph, where's the hot spot" and the honest answer was that there was no hot spot in that image because there was no information in that image. It was a picture of a profiler failing, rendered in the house colours.
I build PandaStack, an open-source Firecracker microVM platform, and this is the single most common way profiling goes wrong on it. Not because `perf` is unavailable — it runs fine — but because the recipe everyone has memorised depends on three things a guest quietly does not provide, and `perf` is not in the habit of telling you which one bit you. It records something, you render it, and the something is garbage shaped exactly like data.
The general question of how to profile a workload in a sandbox is covered in How to profile code running in a sandbox, and the short version of that post is "use the in-language profiler first", which remains correct and which I will not re-argue here. This post is the other case: you genuinely want `perf`, you want a flamegraph, and you want to understand why the standard pipeline produces an empty one inside a microVM. Three failures, in the order you will hit them, and the workflow that avoids all three.
The usual recipe, and why it became the usual recipe
Four commands, and they have been four commands for over a decade:
- `perf record -F 99 -g -- ./cmd` — sample the program 99 times a second, capture a call stack with each sample, write it to `perf.data`.
- `perf script` — expand that binary file into one text block per sample: a timestamp, a process, and a stack, innermost frame first.
- `stackcollapse-perf.pl` — fold identical stacks into one line each, semicolon-separated, with a sample count on the end.
- `flamegraph.pl` — turn those folded lines into an SVG where box width is sample count and vertical position is stack depth.
It became the standard because every stage is dumb and inspectable. The folded format is grep-able text. The SVG is self-contained. There is no agent, no daemon, no language runtime cooperation, and no instrumentation in your binary — `perf` samples from outside the process, so the thing you measure is the thing you ship. You can diff two folded files. You can `wc -l` them. When it breaks you can read the intermediate files and see exactly where the information disappeared, which is more than most observability stacks will grant you.
That inspectability is the reason this post can be short and specific instead of a shrug. Every one of the three failures below shows up as a recognisable pattern in `perf script` output, before you ever draw a graph. The mistake is drawing the graph first.
Failure one: there is no PMU to count with
`perf record` with no `-e` does not sample a timer. It samples `cycles` — the hardware CPU-cycle counter — because on bare metal that is the best available signal: cheap, precise, and attributable to an instruction pointer with hardware help. Those counters live in the Performance Monitoring Unit, a block of host silicon with its own model-specific registers, and the PMU is not something a guest gets for free. The hypervisor has to deliberately virtualise it, trapping and multiplexing those MSRs on the guest's behalf, and a VMM optimised for boot time and a small attack surface tends not to.
Firecracker's CPUID normalisation hides the architectural performance-monitoring leaf from the guest, so the guest kernel does not register a hardware PMU and `perf` finds no hardware events to arm. Whether any PMU is exposed at all varies by architecture and by release — KVM itself can do a virtual PMU on x86, so this is a VMM and host policy decision rather than a law of physics. Check your own version with `perf list` and `perf stat` rather than trusting my summary or anybody else's.
What you see depends on your `perf` version, and neither outcome is friendly. Older builds fail outright with a line about the `cycles` event not being supported, which is at least honest. Newer builds have a fallback path: when arming `cycles` returns "not supported", `perf` retries with the `cpu-clock` software event and prints a warning to stderr. If you ran the command inside a script, or a CI job, or an `exec()` whose stderr you did not read, you now have a `perf.data` recorded with a different event than you asked for and no indication that a substitution happened.
cpu-clock and task-clock: timers, not silicon
The fix is one flag: ask for a software event explicitly. `perf record -e cpu-clock -F 99 -g -- ./cmd` works in a guest with no PMU because a software event is implemented by the kernel, not the CPU. The kernel arms a high-resolution timer, and every time it fires it takes a sample: current instruction pointer, current stack, attribute it to whatever was running. No model-specific registers are involved.
`task-clock` is the per-task sibling: it advances with the CPU time of the task you attached to, so for `perf record -- ./cmd` the two land in much the same place. The distinction matters when you record system-wide with `-a`, where `cpu-clock` is a timer per CPU and `task-clock` is not what you want. For profiling one command, either is fine; pick one and write it down so your next recording is comparable.
Run `perf list hw sw` in your guest and read the output. That command is the entire diagnostic. If the hardware section is thin or missing while the software section lists `cpu-clock`, `task-clock`, `page-faults` and `context-switches`, you now know precisely what you are allowed to ask about. Then run `perf stat -e cycles,instructions,task-clock -- sleep 1` and watch the hardware rows come back `<not supported>` beside two real software numbers. Five seconds, and you never have to wonder again.
The questions that become unanswerable
Here is the part worth being precise about, because the flag swap makes the problem look solved when actually a category of question has been removed from the menu. `cpu-clock` is a timer. A timer-driven sample tells you where on-CPU time went: which function the instruction pointer was in, how often, and what called it. That is enough to answer the overwhelming majority of real performance questions, and it is why the software-event fallback is not a consolation prize.
But a timer cannot count events inside the CPU, because it is not inside the CPU. So this whole family of questions has no answer available inside a guest without a PMU:
- Instructions per cycle. You cannot compute a ratio when one of the two terms does not exist.
- Cache behaviour: L1, L2 and LLC miss rates, miss latency, whether your restructured loop actually improved locality or just moved the time somewhere else.
- Branch mispredictions, and therefore any question about whether a branch is predictable enough to leave in a hot loop.
- TLB misses and page-walk cost — relevant precisely when you are reasoning about hugepages, which is a thing we do.
- Memory bandwidth saturation and anything derived from uncore counters.
- Precise event attribution. Without PEBS-style hardware assistance there is no `:pp` precision, so a sample's instruction pointer carries skid and attributing cost to one instruction rather than its neighbourhood is not sound.
If the question you walked in with was "is this loop memory-bound or compute-bound", a microVM is the wrong instrument and no amount of flag archaeology will change that. Rent a machine where you own the hardware, measure there, and bring the conclusion back. If the question was "which function is eating the wall clock", a timer answers it completely and you should stop reading about PMUs.
Failure two: the paranoia knob — and the best argument in this post
The second failure is permissions, and it produces the most confusing symptom of the three: `perf` runs, writes a file, reports samples, and the flamegraph is half-blind. User frames resolve. Kernel frames are hex. Or there are no kernel frames at all and your graph stops politely at the syscall boundary as though nothing happens on the other side of it.
Two sysctls control this, both in `/proc/sys/kernel/`, and they do different jobs that people routinely conflate.
What each perf_event_paranoid value permits
`kernel.perf_event_paranoid` gates what an unprivileged process may ask the perf subsystem for. It is a ladder, and each rung removes a capability:
- `-1` — effectively no restrictions. Raw tracepoint access is allowed and the `perf_event_mlock_kb` limit is ignored without `CAP_IPC_LOCK`. This is what you want inside a sandbox you own.
- `0` — raw and ftrace function tracepoint access is disallowed. CPU-wide measurement and kernel samples are still permitted.
- `1` — additionally, no CPU-wide measurement: `perf record -a` and `perf top` across all CPUs are refused. You may still profile your own processes, kernel frames included.
- `2` — additionally, no kernel samples at all. This is the common distro default and it is the one that produces the half-blind flamegraph, because your profile now contains only userspace and you were not told.
- `3` and above — Debian and Ubuntu kernels carry an out-of-tree patch adding a stricter rung meaning "unprivileged use of perf events is off entirely". Upstream stops at `2`. Read your own value rather than assuming a default; the number differs across distributions and across releases of the same distribution.
There is also `CAP_PERFMON`, added in Linux 5.8 and therefore present on the 5.10 guest kernel we ship. It lets you grant perf access to a process without handing it the whole of `CAP_SYS_ADMIN`, which is the right tool if you are building a profiling capability into a product rather than poking at one sandbox by hand.
kptr_restrict, or why your kernel frames are hex
`kernel.kptr_restrict` is a different knob for a different problem: it governs whether kernel pointers are exposed in a form that can be turned back into names. At `1`, kernel addresses are hidden from processes without `CAP_SYSLOG`; at `2`, from everyone. `/proc/kallsyms` is the casualty — it dutifully lists every kernel symbol with an address of `0000000000000000`, so `perf report` has a sorted table of zeros to look names up in and resolves nothing.
The resulting flamegraph is the specific flavour of useless that sends people down the wrong path for an afternoon: userspace reads fine, and above `entry_SYSCALL_64` there is a stack of raw hex that looks like corruption. It is not corruption. It is redaction, and it is undone with `sysctl -w kernel.kptr_restrict=0`. Set both knobs before you record, not after — a recording made under `perf_event_paranoid=2` contains no kernel samples to resolve later, so relaxing the sysctl afterwards buys you nothing but a second recording.
The part that only works because the kernel is yours
This is the strongest argument in the post, so let me make it plainly. On a shared-kernel container platform, `kernel.perf_event_paranoid` is a host sysctl. It is not namespaced. There is exactly one copy of it per machine, `/proc/sys` is typically mounted read-only in your container anyway, and asking your platform to set it to `-1` is asking them to let every tenant on that host sample instruction pointers belonging to every other tenant. The correct answer to that request is no, and the platforms that say no are being responsible rather than obstructive. So the usual outcome is that kernel-level profiling is simply not on offer, and you learn to live without it.
In a microVM the same knob is inside your own kernel, writing to your own `/proc`, governing sampling of your own vCPUs. Setting it to `-1` grants you nothing you did not already have, because the only thing within reach of that perf subsystem is you. There is no host-wide blast radius to weigh, no other tenant whose stacks become visible, and no platform decision to escalate — which means there is nothing to ask permission for. You are root in a kernel nobody else is using.
The container answer to "please loosen perf_event_paranoid" is a conversation about other people's security. The microVM answer is a sysctl.
That is the whole isolation argument in miniature, and it generalises past profiling. `kptr_restrict`, `dmesg_restrict`, `ftrace`, `/proc/kcore`, loading a kernel module, running `bpftrace`, unmasking `/sys/kernel/debug`: every one of them is a host-wide negotiation in a shared-kernel sandbox and a local decision in a microVM. The security property you paid for when you accepted a separate kernel turns out to also be an ergonomics property, which is not how these trades usually go. The sibling case for debuggers, `ptrace` scope and core dumps is in Debugging Inside a MicroVM: gdb, strace, and Core Dumps.
Failure three: the stacks are lies
Now the PMU is out of the way and the sysctls are open, and the flamegraph is still wrong — but wrong in a new way. Boxes one frame deep. Towers that terminate in the middle of nothing. A graph whose total width is right and whose structure is fiction. This failure has nothing to do with virtualisation; it is the same on bare metal, and it just happens to arrive third.
Frame pointers, DWARF, and the one option you do not have
`-g` asks `perf` to walk the frame-pointer chain: read the base pointer register, follow the saved-pointer links up the stack, collect return addresses. It is fast, it costs almost nothing per sample, and it only works if every frame on the stack actually maintained a frame pointer. At `-O1` and above, GCC and Clang take `-fomit-frame-pointer` as the default, because that register is more usefully spent on your variables than on bookkeeping the debugger might want. So the unwinder follows the chain until it reaches a frame that did not keep one, and then it stops — and `perf` does not label the result as truncated. You get a one-frame stack, or a three-frame stack that is missing the interesting middle, presented with total confidence.
Worth knowing: Ubuntu 24.04 — our rootfs — restored frame pointers for distribution-built packages on 64-bit architectures, specifically so that profiling and crash analysis would work. Fedora made the same call. That changes the odds for distro-shipped libraries in your stack, and it changes nothing at all about the code your own build produced at `-O2`. Check your binary rather than inheriting a claim about the archive.
Three ways out, two of which are available to you:
- `--call-graph dwarf` — the unwinder uses the CFI in `.eh_frame` and `.debug_frame` instead of a register chain, so it works without frame pointers. The cost is real: `perf` copies a slice of the stack with every sample (8 KiB by default, tunable as `--call-graph dwarf,<bytes>`) and the unwinding happens at report time. `perf.data` gets large quickly, deeply recursive code can exceed the copied slice and get truncated anyway, and report-time unwinding is slow. It is still usually the right answer in a sandbox, where the recording lasts minutes and lives on a disk you are about to throw away.
- `--call-graph lbr` — the unwinder reads the processor's Last Branch Record stack, which is nearly free and needs no frame pointers and no stack copying. It is also a PMU-adjacent hardware feature on specific Intel parts, which means it is exactly the thing a guest without a virtualised PMU does not have. Mentioned here so you recognise it in the manual page and do not spend twenty minutes trying to make it work.
- `-fno-omit-frame-pointer` — rebuild the code you care about. The right permanent answer for a service you own and profile often, and a non-answer for the dependency whose binary you were handed. A sandbox is a pleasant place to do this, since you can build the instrumented variant, profile it, and delete the machine.
Interpreted and JIT runtimes: you are profiling the runtime
The most demoralising version of a broken stack is the one where the stack is perfectly correct. Point `perf` at a Python process and it will unwind flawlessly and hand you a tower of `_PyEval_EvalFrameDefault`, repeated once per Python frame, with occasional `PyObject_Vectorcall` for texture. That is not a bug. `perf` samples machine code, and the machine code executing your Python is CPython's bytecode evaluation loop. The profile is an accurate description of the interpreter and contains no trace of your program, because your program is data that the interpreter is reading.
JIT runtimes fail differently and worse. The JVM, V8 and .NET generate machine code at runtime into anonymous memory that belongs to no file on disk, so there is no object for `perf` to look symbols up in. The samples are real and land at real addresses; `perf report` has nothing to map them to, and you get the single enormous `[unknown]` box that started this post.
Pick the fix before you record, because every one of these changes the recording and not the rendering:
- A `perf-<pid>.map` in `/tmp` — a plain text file of address, size and symbol that the runtime writes and `perf` reads at report time. Node emits one with `--perf-basic-prof`; the JVM has `perf-map-agent`; `perf inject --jit` handles the richer jitdump format. This is the mechanism that makes JIT frames resolvable at all.
- CPython 3.12 and later can emit perf trampolines so Python frames appear in a `perf` profile: `python3 -X perf script.py`, or `PYTHONPERFSUPPORT=1`, or `sys.activate_stack_trampoline("perf")` from inside the process. Ubuntu 24.04 ships 3.12, so this is plausibly available to you — verify it on your interpreter build rather than assuming, and note the documentation's own advice that an interpreter built with frame pointers unwinds considerably better.
- A language-native sampler, which is usually the shortest path: py-spy reads a running Python process's memory from outside and reconstructs real Python stacks; async-profiler does the JVM properly and, relevantly for us, has an `itimer` mode precisely for environments where perf events or the PMU are unavailable; Go's runtime profiler already knows its own frames and needs none of this.
- For Go, Rust and C++ the plain `perf` path is the correct one and the only question is frame pointers, which is the previous section.
The decision has to be made before the recording because none of it is recoverable afterwards. A `perf.data` captured without a JIT map is a file full of addresses that no longer correspond to anything; the process that could have explained them has exited and its anonymous mappings are gone. There is no post-processing step that recovers that, and in an ephemeral sandbox the machine is gone too.
Symbols must resolve at report time, not record time
One more trap, and it is the one that specifically punishes disposable infrastructure. `perf.data` does not contain symbols. It contains addresses plus build-ids identifying the objects those addresses came from. Resolution happens later, when `perf report` or `perf script` locates each object and reads its symbol table. Strip the binary, or build it in one container layer and run it in another that dropped the debuginfo, or record in a sandbox and report on a laptop that never had the binary, and you get hex — a correct recording rendered unreadable by a missing input.
`perf archive perf.data` is the fix and almost nobody uses it. It walks the build-ids in the recording, collects the matching objects from the local build-id cache, and produces a tarball you unpack into `~/.debug` on the machine where you will do the reporting. Run it in the guest, copy the tarball out alongside `perf.data`, and the recording stays resolvable for as long as you keep two files. `perf buildid-list -i perf.data` tells you what it thinks it needs, which is the thing to read when resolution fails anyway.
Symptom, cause, fix
| Symptom | Actual cause | Fix |
|---|---|---|
| `perf record` fails outright, or stderr mentions the `cycles` event not being supported | No virtualised PMU, so the default hardware event cannot be armed. Newer perf silently falls back to `cpu-clock` and warns on stderr you did not read. | Ask for a software event explicitly: `-e cpu-clock` or `-e task-clock`. Confirm with `perf list hw sw` and `perf stat -e cycles,task-clock -- sleep 1`. |
| Every frame is a hex address, user and kernel alike | Symbols cannot resolve at report time: stripped binary, no matching build-id, debuginfo left behind in a build layer or on another machine. | `perf archive perf.data` in the guest, copy the tarball out with the recording, unpack into `~/.debug`. Keep an unstripped copy of what you profile. |
| User frames have names, kernel frames are hex — or there are no kernel frames at all | `kptr_restrict` is hiding kallsyms addresses, and/or `perf_event_paranoid` is at `2` so no kernel samples were ever recorded. | `sysctl -w kernel.kptr_restrict=0` and `kernel.perf_event_paranoid=-1` in your own guest, BEFORE recording. A paranoid-2 recording cannot be fixed afterwards. |
| Stacks are one or two frames deep; the flamegraph is a flat row of wide boxes | Frame-pointer omission. `-g` walks a chain that stops at the first frame compiled with `-fomit-frame-pointer`, which is the default at `-O1` and above. | `--call-graph dwarf,16384`, or rebuild the code you own with `-fno-omit-frame-pointer`. Do not reach for `lbr`: it needs hardware the guest does not have. |
| One enormous `[unknown]` box eating most of the width | JIT-generated code in anonymous memory with no symbol map, so the addresses belong to no object on disk. | Record with a JIT map: Node `--perf-basic-prof`, JVM `perf-map-agent` or `perf inject --jit`. Must be decided before recording; unrecoverable after. |
| A tower of `_PyEval_EvalFrameDefault` repeated down the graph | The profile is accurate — it is CPython's evaluation loop. `perf` samples machine code, and your Python is data to that machine code. | py-spy for real Python stacks, or CPython 3.12+ perf trampolines (`python3 -X perf`). async-profiler with `itimer` mode is the JVM equivalent. |
| Sampled time is far smaller than the wall clock, and nothing in the graph explains the gap | The time is off-CPU, or it is host-side work the guest cannot observe: VM exits, the VMM's device thread, page faults on a streamed restore. | In-guest: `perf record -e sched:sched_switch -g` for off-CPU. For host-side time, stop profiling the guest — no guest-side tool can see it. |
Record in the guest, render outside it
The workflow that falls out of all three failures is the same one the ephemerality of a sandbox would have suggested anyway: the guest's job is to produce `perf.data`, and every step after that belongs on your workstation.
There are three independent reasons, and they stack. Resolution needs the binaries and the debuginfo, which your machine has and which you would otherwise be installing into a guest you are about to destroy. Flamegraph generation is a Perl pipeline plus a repo checkout, and baking Perl scripts into a VM image to use them once is how images get fat. And the output is an SVG you want to open in a browser, diff against last week's, and attach to a ticket — none of which is a thing the guest can do for you.
So: probe, loosen, record, archive, copy out, destroy. Render at home. The `perf.data` plus its `perf archive` tarball are the deliverable, they are small enough to commit next to the fix, and they stay readable long after the machine that produced them stopped existing.
#!/usr/bin/env bash
# Run this INSIDE the guest. The order matters -- probe, loosen, record --
# and most "perf is broken in a VM" reports are step 1 skipped.
set -euo pipefail
# 0. Is there even a usable perf binary? On an Ubuntu 24.04 rootfs running
# a 5.10 Firecracker kernel there is no linux-tools-$(uname -r) package,
# because nobody built a perf for your kernel. linux-tools-generic gives
# you one whose version differs from the running kernel. It mostly works
# and it tells you when it does not.
apt-get install -y linux-tools-generic >/dev/null
perf --version
# 1. What does this guest actually expose? Read it; do not assume it.
# Hardware events -- cycles, instructions, cache-misses, branch-misses --
# come from the PMU, which is host silicon the VMM usually does not hand
# through. Software events are timer-driven and need no PMU at all.
perf list hw sw 2>/dev/null | sed -n '1,40p'
# 2. Prove it rather than inferring it. On a guest with no virtualised PMU
# the hardware rows come back <not supported> while the software rows are
# real numbers. Five seconds, zero ambiguity.
perf stat -e cycles,instructions,task-clock,context-switches -- sleep 1
# 3. The sysctls. These are GUEST sysctls, which is the whole argument of
# this post: -1 permits everything including raw tracepoints, and
# kptr_restrict=0 makes /proc/kallsyms readable so kernel frames get
# names instead of hex.
sysctl -w kernel.perf_event_paranoid=-1
sysctl -w kernel.kptr_restrict=0
# Sampling is rate-limited too. Read the ceiling before perf quietly
# throttles you and mutters "Lowering default frequency rate".
sysctl kernel.perf_event_max_sample_rate kernel.perf_event_mlock_kb
# 4. The invocation that works. -e cpu-clock is a timer-driven SOFTWARE
# event, so no PMU is required. -F 99 is the convention: 99 and not 100
# so samples do not fall into lockstep with anything ticking at 1 Hz or
# 10 Hz and quietly over-represent it.
perf record -e cpu-clock -F 99 -g -o /root/fp.data -- ./workload
# Careful: -g is shorthand for --call-graph fp. Passing -g AND
# --call-graph dwarf is not "both" -- the last one wins. Pick one.
# DWARF unwinding copies a slice of the stack with every sample, so
# perf.data grows fast; 16k is a deliberate, bounded choice.
perf record -e cpu-clock -F 99 --call-graph dwarf,16384 \
-o /root/dwarf.data -- ./workload
# 5. On-CPU sampling cannot see blocked time. For off-CPU work on a 5.10
# guest the sched tracepoint recipe still works. perf's own --off-cpu
# mode is newer than this kernel, so check before you plan around it.
perf record -e sched:sched_switch -g -o /root/offcpu.data -- ./workload
# 6. The step everyone forgets. perf archive bundles the objects that
# perf.data's build-ids point at, so the recording is still resolvable
# after the sandbox is gone. Copy BOTH files out.
perf archive /root/dwarf.data || perf buildid-list -i /root/dwarf.data
# On your workstation, where the symbols, the toolchain and a browser are.
set -euo pipefail
# 1. perf.data came out of the guest through the filesystem API (see the
# Python below). Unpack the build-id objects into the local cache so
# report-time resolution can find them.
mkdir -p ~/.debug && tar xf dwarf.data.tar.bz2 -C ~/.debug
# 2. Sanity-check resolution BEFORE drawing anything. A column of hex here
# is a symbol problem, not a profiling result, and a flamegraph will
# render it just as beautifully as it renders real data.
perf report -i dwarf.data --stdio --sort symbol | head -30
# 3. The pipeline. stackcollapse-perf.pl folds perf script's output into one
# line per unique stack plus a count; flamegraph.pl draws the SVG. Two
# Perl scripts from Brendan Gregg's FlameGraph repo -- and two Perl
# scripts you do not want baked into a guest image.
perf script -i dwarf.data > out.perf
./FlameGraph/stackcollapse-perf.pl out.perf > out.folded
./FlameGraph/flamegraph.pl out.folded > profile.svg
# 4. Two recordings of the same warm state? Diff them instead of squinting
# at two SVGs in two browser tabs.
./FlameGraph/difffolded.pl before.folded after.folded \
| ./FlameGraph/flamegraph.pl > diff.svg
The thing a microVM gives you that nothing else does
Everything above is damage control — three ways the standard recipe breaks and how to route around each. Here is the part that is an actual advantage, and it is the reason I will take a guest without a PMU over a container with one for most of the work I do.
A confusing profile is normally a story. You recorded something odd on Tuesday, you cannot reproduce it, and all you have left is an SVG and a memory. In a microVM you can snapshot the machine you were profiling. Not a container image of the filesystem — the running state: memory, registers, device state, the page cache as it was, the process table, the warm JIT, the half-filled connection pool. Restore it and you are back in the machine that produced the confusing profile, not an approximation of it. Restore is 179 ms at p50, so "let me go and look again" costs less than reloading the SVG.
Then there is the fan-out, which is the bit I did not expect to care about and now use constantly. `fork_tree()` gives you up to 16 children that inherit the parent's running memory, so you can profile the same warm state several different ways — once with `cpu-clock`, once with DWARF unwinding, once with `sched_switch` for the off-CPU view, once with py-spy for a second opinion — without re-reaching that state four times and hoping it is the same state each time. A same-host fork lands in 400 to 750 ms.
Note the asymmetry, because it matters here and has been got wrong in print before: `fork()` clones only the disk and the child cold-boots, with its own entropy, its own PIDs and its own clock. That is the right primitive when you want a clean machine with the same files. Only `fork_tree()` inherits running memory, which is what "profile the same warm state" requires. The full comparison is in Fork a microVM for Tree-of-Thought Agents, and the reproducibility angle more generally in Snapshot the Failure, Not the Log Line.
from pandastack import Sandbox
# 1. ttl_seconds is an IDLE timeout, not a wall-clock lifetime. A sandbox
# busy recording will not be reaped out from under you; one you are
# staring at while reading a report might be.
sbx = Sandbox.create(
template="base",
ttl_seconds=1800,
metadata={"purpose": "perf", "ticket": "PERF-412"},
)
sbx.exec("apt-get update -qq && apt-get install -y -qq linux-tools-generic")
sbx.filesystem.write("/root/workload.py", open("workload.py").read())
sbx.exec("pip install -q -r /root/requirements.txt")
# 2. Loosen the knobs in OUR guest kernel. Nothing on the host moves.
sbx.exec("sysctl -w kernel.perf_event_paranoid=-1 kernel.kptr_restrict=0")
# 3. Freeze the ready state. This is the profiling equivalent of a core
# dump you can run: setup done, caches warm, nothing measured yet.
snap = sbx.snapshot()
# 4. Record. exec() does not reliably enforce a timeout argument, so bound
# the command in-guest with timeout(1). A wedged profiler should fail
# the step, not hold the sandbox. -X perf asks CPython 3.12+ to emit
# trampolines so Python frames appear instead of the interpreter loop
# -- verify it on your own interpreter build before relying on it.
r = sbx.exec(
"cd /root && timeout --kill-after=10s 300 "
"perf record -e cpu-clock -F 99 --call-graph dwarf,16384 "
"-o /root/perf.data -- python3 -X perf workload.py"
)
print(r.exit_code, r.stderr[-2000:])
# 5. Bundle the symbol objects, then pull both files out. The artifact is
# the deliverable; the sandbox is not.
sbx.exec("cd /root && perf archive /root/perf.data || true")
for name in ("perf.data", "perf.data.tar.bz2"):
open(name, "wb").write(sbx.filesystem.read(f"/root/{name}"))
# 6. Profile the SAME warm state several ways. fork_tree inherits the
# parent's RUNNING memory (capped at 16 children), so each child starts
# from the state you just measured instead of re-reaching it. A plain
# fork() would clone only the disk and cold-boot -- useful, different.
for i, child in enumerate(sbx.fork_tree(3)):
child.exec(
f"cd /root && timeout 300 perf record -e task-clock -F 99 -g "
f"-o /root/v{i}.data -- python3 workload.py --variant {i}"
)
open(f"v{i}.data", "wb").write(child.filesystem.read(f"/root/v{i}.data"))
child.kill()
sbx.kill()
# The confusing profile stays reproducible: restore the snapshot and you
# are back on the machine that produced it. Restore is 179 ms at p50.
again = Sandbox.create(from_snapshot=snap)
What a guest-side profile cannot tell you
A guest-side profiler samples guest instruction pointers. That is its whole epistemology, and it draws a hard boundary: work the host does on your behalf is not merely hard to see, it is categorically absent from the recording. There is no flag that fixes this.
- VM exits. Every trap out to the VMM and back is time your vCPU is not executing guest code. The mechanics are in VM exits: the actual currency of virtualization overhead. Your guest profile cannot show them; `perf kvm stat` on the host can, and as a tenant you do not have the host.
- The VMM's own threads. Your virtio block and network I/O is handled by Firecracker's device threads in a host process. From inside, that work appears as a syscall that took a while.
- Host-side page faults on a streamed restore. When guest memory is paged in on demand over the network — on our platform, 4 MiB chunks fetched from object storage with HTTP Range GETs — a first touch can park the vCPU in the host for the duration of a fetch. The guest's clock keeps advancing; no sample is attributable to the wait; your sampled time comes up short against the wall clock and nothing in the graph accounts for the difference.
- CPU steal. Another tenant on the physical host is runnable and you are not. CPU Steal Time in microVMs: Your VM Isn't Slow, It's Waiting covers how to see it; the point is that the guest-side flamegraph will not.
The practical rule: when a profile shows your process idle while the wall clock advances, stop reasoning about your code. Compare the summed sample time against elapsed time, look at off-CPU state with `sched_switch`, and if the gap persists accept that it is below your floor of visibility. Chasing it with a guest-side profiler is futile, and the measurement you actually want is on the host.
Two more honest caveats. First, sampling perturbs a guest more than it perturbs bare metal: arming and taking a high-frequency timer interrupt involves the hypervisor's timer path, so the overhead of the measurement scales worse in a VM than in the numbers you remember from a physical box. I am not going to invent a percentage for you — measure your own workload with and without the profiler attached. Stay near the conventional `-F 99` rather than reaching for a kilohertz, and treat a very high sampling rate as something you justify rather than something you default to. Also watch for `kernel.perf_event_max_sample_rate`, which will silently throttle you and say so in a line you were not reading.
Second, the guest kernel is 5.10. Check feature availability rather than assuming your `perf` manual page describes your kernel. BPF-based off-CPU profiling (`perf record --off-cpu`), newer `perf lock contention` modes and various `perf ftrace` subcommands all postdate it, which is why the off-CPU recipe above uses `sched:sched_switch` tracepoints — the technique that has worked since forever and still works. The version trade-offs are in You Ship the Kernel: Firecracker Guest 5.10 vs 6.1.
The bottom line
`perf` works inside a microVM. The recipe does not, and the three things it depends on fail in a fixed order: the default event needs hardware you do not have, the sysctls are at distro defaults that hide half the stack, and the frame pointers were optimised away years ago. Swap in `-e cpu-clock`, open `perf_event_paranoid` and `kptr_restrict` in your own kernel, unwind with DWARF, `perf archive` the symbols, and render at home. That is the entire fix, and none of it is exotic once you know which of the three bit you.
Where I would not use this. If your question needs real silicon counters — IPC, cache-miss rates, branch prediction, memory bandwidth — a guest without a virtualised PMU cannot answer it and you should go and measure on hardware you own. We have no GPUs and no GPU passthrough, so GPU kernel profiling is not a conversation. If you need a guest kernel newer than 5.10 for a specific `perf` feature, check the comparison post before you build on us. And if the answer was going to come from a language profiler in ten minutes, use the language profiler — How to profile code running in a sandbox makes that case properly and it is the right default.
What the microVM gives you in exchange is two things worth more than the counters you lost. The kernel is yours, so every tracing knob is a local decision instead of a request to a platform team that will reasonably say no. And the machine is snapshottable, so a profile you do not understand is a thing you can go back to rather than a story you tell. I would rather have a reproducible timer-based profile of the exact machine than a once-off cycle-accurate profile of a machine that no longer exists — and having had both, that is not a close call.
Frequently asked questions
Why does perf record fail or produce no useful samples inside a Firecracker microVM?
Because `perf record` with no `-e` flag does not sample a timer — it samples `cycles`, a hardware event read from the CPU's Performance Monitoring Unit. The PMU is host silicon, and a hypervisor has to deliberately virtualise it for a guest to see one. Firecracker's CPUID normalisation hides the architectural performance-monitoring leaf, so the guest kernel registers no hardware PMU and there is nothing for `perf` to arm. What happens next depends on your `perf` build: older versions fail with a message about the `cycles` event not being supported, and newer versions retry with the `cpu-clock` software event and print a warning to stderr — which, if you ran the command from a script or through an API call whose stderr you did not read, means you now hold a recording made with an event you did not request. The fix is one flag: `-e cpu-clock` or `-e task-clock`, both software events implemented by a kernel timer rather than by hardware counters. Confirm the situation in your own guest before planning around it, because PMU exposure varies by architecture and by Firecracker release and KVM itself is capable of virtualising one. `perf list hw sw` shows what is registered, and `perf stat -e cycles,instructions,task-clock -- sleep 1` is the decisive five-second experiment: the hardware rows come back `<not supported>` beside two real software numbers, or they do not, and either way you now know.
Can I get hardware counters — cache misses, IPC, branch mispredictions — inside a sandbox?
Not on a guest with no virtualised PMU, and it is worth understanding that this is a removed capability rather than a harder path to the same answer. A software event like `cpu-clock` is driven by a kernel timer: it fires, it reads the instruction pointer and walks the stack, and it attributes a sample to whatever was running. That tells you where on-CPU time went, which answers most real performance questions completely. What it cannot do is count events happening inside the CPU, because it is not inside the CPU. So instructions per cycle is unavailable — you cannot form a ratio when one term does not exist. Cache miss rates, miss latency, branch mispredictions, TLB misses and page-walk cost, memory bandwidth and anything from uncore counters are all equally out of reach. You also lose precise event attribution: without PEBS-style hardware assistance there is no `:pp` precision, so a sample's instruction pointer carries skid and blaming one specific instruction rather than its neighbourhood is not sound. If your question is genuinely "is this loop memory-bound or compute-bound", a microVM is the wrong instrument and no flag archaeology changes that — measure on hardware you control and bring the conclusion back. If your question is "which function is eating the wall clock", a timer answers it and you can stop worrying about PMUs entirely.
Is it safe to set kernel.perf_event_paranoid=-1 inside a sandbox?
Inside a microVM, yes, and this is the most useful structural difference between a microVM and a shared-kernel container. `kernel.perf_event_paranoid` is not a namespaced sysctl. On a container platform there is exactly one copy of it for the whole host, `/proc/sys` is usually mounted read-only in your container anyway, and asking the platform to set it to `-1` is asking them to let every tenant on that machine sample instruction pointers belonging to every other tenant. The right answer to that request is no, which is why kernel-level profiling is often simply unavailable on container-based sandboxes — not obstruction, just arithmetic about blast radius. In a microVM the knob lives in your own kernel, writing to your own `/proc`, governing sampling of your own vCPUs. Setting it to `-1` grants you nothing you did not already have, because the only thing reachable by that perf subsystem is you. There is no other tenant whose stacks become visible and no host-wide setting being weakened, so there is nothing to escalate. The same reasoning covers `kptr_restrict`, `dmesg_restrict`, `ftrace`, `/proc/kcore`, loading a module and running `bpftrace`. Two practical notes: set the sysctls before you record, since a recording made at `perf_event_paranoid=2` contains no kernel samples to resolve later; and if you are building this into a product rather than debugging by hand, `CAP_PERFMON` (Linux 5.8 and later, so present on a 5.10 guest) grants perf access without handing over all of `CAP_SYS_ADMIN`.
Why is my flamegraph a tower of _PyEval_EvalFrameDefault, and what should I use instead?
Because the profile is correct and you asked the wrong tool. `perf` samples machine code. The machine code executing your Python program is CPython's bytecode evaluation loop, so an accurate stack of a Python process is that loop repeated once per Python frame, with some `PyObject_Vectorcall` between the layers. Your program is data being read by the thing `perf` is profiling, and no unwinding option will surface it. There are three fixes and they differ in effort. The shortest is a language-native sampler: py-spy reads a running interpreter's memory from outside and reconstructs real Python stacks with no changes to the process, which is almost always the right call. Second, CPython 3.12 and later can emit perf trampolines so Python frames appear in a `perf` recording — `python3 -X perf script.py`, or `PYTHONPERFSUPPORT=1`, or `sys.activate_stack_trampoline("perf")` from inside the process. Ubuntu 24.04 ships Python 3.12, so this is plausibly available, but verify it on your own interpreter build and note the documentation's point that an interpreter compiled with frame pointers unwinds much better. Third, for JIT runtimes the equivalent mechanism is a symbol map the runtime writes: Node's `--perf-basic-prof`, the JVM's `perf-map-agent`, or the richer jitdump format consumed by `perf inject --jit`; async-profiler is the better JVM answer and has an `itimer` mode made for environments without usable perf events. All of these change the recording, not the rendering, so the decision has to be made before you press record — a `perf.data` full of unmapped JIT addresses cannot be repaired after the process exits.
Should I generate the flamegraph inside the sandbox or copy perf.data out first?
Copy it out, and for three independent reasons that all point the same way. First, symbol resolution happens at report time, not record time: `perf.data` holds addresses plus build-ids, and `perf report` or `perf script` has to locate the matching objects and read their symbol tables. Your workstation has the binaries and the debuginfo; the guest may not, and installing debuginfo into a machine you are about to destroy is work you pay for twice. Use `perf archive perf.data` in the guest to bundle the objects the recording's build-ids point at, copy that tarball out beside `perf.data`, and unpack it into `~/.debug` where you report — that is the step almost nobody takes and it is the difference between a readable profile and a column of hex. Second, flamegraph generation is a Perl pipeline plus a repo checkout, and baking `stackcollapse-perf.pl` and `flamegraph.pl` into a guest image to use them once is how images get fat for no benefit. Third, the output is an SVG you want to open in a browser, diff against last week's run and attach to a ticket, and the sandbox is going away. So the division of labour is: the guest produces the artifact, your machine reads it. On our platform there is a bonus — snapshot the sandbox before you tear it down and the exact machine that produced a confusing profile is restorable in 179 ms at p50, which turns "I cannot reproduce that" into "let me go and look again".
Keep reading
- How to profile code running in a sandbox — The general how-to, and the right place to start: in-language profilers first, strace when it is not CPU, and how the artifact gets out.
- Debugging inside a microVM: gdb, strace and core dumps — The same ownership argument applied to ptrace scope, core patterns and the knobs a shared kernel will not let you touch.
- KVM VM exits explained — The time a guest-side profiler cannot see, and why it is invisible rather than merely hard to find.
- Snapshot-based bug reproduction — Turning a confusing one-off recording into a machine you can go back to, which is the whole payoff of the workflow above.
- Firecracker guest kernel: 5.10 vs 6.1 — Which perf and tracing features postdate the kernel you are actually running, before you plan a recording around them.
Related posts
- Firecracker vCPU Hotplug, and How CPU Scaling Actually Works
Somebody always asks for "more CPU" on a running microVM. Firecracker's answer is no — vcpu_count is fixed before InstanceStart and frozen into every snapshot. The good news: the lever you actually wanted was never vCPU count in the first place.
- Readahead Inside a microVM: The Kernel Guessing Wrong, 4 KiB at a Time
There is no disk. A guest read is a descriptor on a virtqueue, an exit to the VMM, a file read on the host, and a second page cache doing its own speculation underneath yours. Both layers are guessing about the same access pattern, and only one of them is paying for the exits.
- CPU Steal Time in microVMs: Your VM Isn't Slow, It's Waiting
The CPU wasn't slow, it was in a meeting. Steal time is the one number that tells you whether a tenant's slowness is their code, their storage, or your overcommit — here's the mechanism, the misreads, and the fix order.
- kvm-clock vs TSC: How a Firecracker Guest Tells the Time, and What a Snapshot Does to It
A guest clock is not one thing. It is a rated list of clocksources, a shared memory page the host writes into, and a userspace fast path that silently becomes a trap when you pick wrong. Then you snapshot it, and the whole stack starts lying in a very specific way: TLS first, loudly.
- Your benchmark ran at a different clock speed than production
The same core does not run at the same speed twice. Governor, turbo bin, how many neighbours are busy, thermal headroom and ramp latency all move it — which is why the first sandbox on a quiet host looks fast and the fiftieth on a busy one gets blamed on the platform.
More in Internals · See PandaStack benchmarks
49ms p50 cold start. Fork, snapshot, and scale to zero.