Self-Hosting a Compiler Explorer: Running Strangers' Compilers Safely
A compiler explorer is a wonderful thing to run. Paste a function, pick a compiler, watch the assembly appear, argue with a colleague about whether the optimiser really did that. Godbolt's public instance made a generation of engineers better at reading the output of their own tools, and the natural next thought — for a language vendor, a chip company, a docs team, an educator — is to run one for your own toolchain.
The natural thought after that is where this post starts. "It's safer than a normal code playground, because we're not actually running their code — we're only compiling it." I have heard that sentence from very good engineers. It is wrong in a specific, enumerable way, and the enumeration is more interesting than the conclusion.
I'm Ajay; I built PandaStack, which runs untrusted code in Firecracker microVMs. Compilation is the workload where the gap between what people think they are exposing and what they are actually exposing is widest, so it is worth walking the surface properly before talking about any boundary at all.
The compiler is the program, and they get to program it
Start with the preprocessor, which is the part everyone forgets is a language. It is not a text filter with a few conveniences; it is an interpreter with filesystem access, and it runs before any notion of "your program" exists. Four primitives, all of them ordinary documented behaviour, all of them confirmed against clang 21 while writing this post rather than recalled from memory:
- `#include` takes a path. An absolute one works. The compile fails, of course — `/etc/passwd` is not valid C — but the diagnostics quote the offending lines verbatim, which is to say they quote the file. Your service renders those diagnostics in a browser. That is the exfiltration path, and it runs through the feature you most wanted to build.
- `__has_include` is a filesystem oracle with no error and no non-zero exit. Pair it with `#pragma message` and you get a successful compile whose warnings are a directory listing. Loop it over a wordlist and you have reimplemented `ls` through a preprocessor conditional.
- `#embed`, standardised in C23 and shipping in current compilers, is not a probe and not a leak through diagnostics. It turns a whole file into an array initialiser in the artefact you hand back. If you offer any path by which the submitter receives compiled output, that is a complete arbitrary-file-read.
- `.incbin` in inline assembly does the same thing one layer down, because the assembler is also a parser with file access and it predates every sandbox you have ever configured. `strings` on the resulting object file is the whole exploit.
#!/usr/bin/env bash
# Four file-read primitives in a compiler that never runs your program.
# Every one of these was confirmed against clang 21 before this post shipped;
# behaviour varies by compiler and version, so re-run them on the exact
# toolchains you intend to host.
set -u
printf 'MY_FAKE_TOKEN_abcdef123\nsecond_line_here\n' > secret.txt
# 1. INCLUDE AN ABSOLUTE PATH. The compile fails -- and the diagnostics quote
# the offending source lines verbatim, which is to say they quote the file.
# Your service then renders those diagnostics in a browser.
printf '#include "%s/secret.txt"\nint main(void){return 0;}\n' "$PWD" > t1.c
clang -c t1.c -o /dev/null
# secret.txt:1:1: error: unknown type name 'MY_FAKE_TOKEN_abcdef123'
# 1 | MY_FAKE_TOKEN_abcdef123
# 2. __has_include AS A FILESYSTEM ORACLE. One bit per probe, no error, no
# non-zero exit -- a clean compile with a warning that happens to be a
# directory listing. Loop it over a wordlist and you have `ls`.
cat > t2.c <<'EOF'
#if __has_include("/opt/secrets/api.key")
#pragma message("PRESENT")
#else
#pragma message("ABSENT")
#endif
int main(void){return 0;}
EOF
clang -c t2.c -o /dev/null # warning: PRESENT -- exit code 0
# 3. #embed (C23). Not a probe and not a leak through diagnostics: the whole
# file becomes an initialiser in the artefact you hand back.
cat > t3.c <<'EOF'
const unsigned char leak[] = {
#embed "secret.txt"
};
#include <stdio.h>
int main(void){ fwrite(leak,1,sizeof leak,stdout); return 0; }
EOF
clang -std=c23 t3.c -o t3 && ./t3 # prints the file
# 4. THE ASSEMBLER, WHICH NOBODY REMEMBERS IS ALSO A PARSER WITH FILE ACCESS.
# .incbin predates every sandbox you have ever configured.
cat > t4.c <<'EOF'
__asm__(".incbin \"secret.txt\"");
int main(void){return 0;}
EOF
clang -c t4.c -o t4.o && strings t4.o | grep FAKE
# 5. AND THE ONE THAT IS NOT EVEN SUBTLE: a flag that asks the compiler to
# dlopen() a path of the submitter's choosing.
clang -fplugin=/tmp/theirs.so -c t1.c
# error: unable to load plugin '/tmp/theirs.so': 'dlopen(...)'
# -- note WHICH part failed. Not the permission check. The file lookup.
None of that is a vulnerability in anything. It is the preprocessor working as documented, being asked by a stranger to work on their behalf. There is no patch to wait for and no version to pin to.
Compile-time execution: which half is real
This is where I want to correct a claim I see made loosely, including in my own earlier notes, because getting it wrong makes the real risk harder to take seriously.
C++ `constexpr` and `consteval` are frequently described as "arbitrary code execution at compile time". They are not, and the language is careful about it: a constant expression may not call a non-`constexpr` function, which rules out every syscall wrapper in the standard library. Ask `constexpr` to `fopen` and clang refuses with `non-constexpr function 'fopen' cannot be used in a constant expression`. Rust's const evaluator is the same shape — a dedicated interpreter with no I/O, deliberately so. Compile-time evaluation in both languages is a sandbox, and it was built as one.
What it is emphatically not is a *bounded* sandbox, and that is the part that matters. Template instantiation and constant evaluation are both Turing-complete-adjacent enough to allocate until something dies. A 295-byte C++ file with no includes — a recursive template whose two branches instantiate distinct types, so nothing memoises — kept clang pinned for the entire two minutes I let it run before I killed it. That is not a cleverly crafted exploit; it is the first thing anyone tries.
The genuinely arbitrary-code-at-build-time surfaces are narrower and much worse. Rust's `build.rs` and procedural macros are native code that the build system and the compiler respectively execute, by design, with the user's privileges — the dedicated post on sandboxing Cargo builds covers that properly and I won't re-derive it. And `-fplugin=<path>` asks the compiler to `dlopen` a shared object the submitter names. When I pointed clang at a nonexistent plugin it told me exactly which step failed: the file lookup. Not a permission check. There wasn't one.
The flags are user input, which breaks the usual advice
Every compiler has knobs for the exhaustion problem. Clang takes `-ftemplate-depth=`, `-fconstexpr-steps=`, `-fconstexpr-depth=` and `-fbracket-depth=`; GCC has `-ftemplate-depth=` and `-fconstexpr-ops-limit=` (check `--help=common` on the version you ship, these move). `-ftime-report` tells you where the time went. Good flags. Set them.
Then notice what makes a compiler explorer different from a build service: the command line is the feature. The entire product is that somebody gets to try `-O3 -march=native` and see what changes. So every limit you express as a flag is a limit the submitter can raise by passing the same flag with a bigger number, and every flag you ban is a line in a denylist that the next compiler release will quietly route around.
Which collapses the design question in a useful way. There are exactly two flags you must refuse on principle — the plugin loaders, because they are direct code execution, and the response-file forms like `@file` that make the compiler read a path you did not choose. Everything else goes through, the per-compiler limits become hints to the cooperative majority, and the actual enforcement lives somewhere the submitter cannot address. Not a smaller list of flags. A different kind of boundary.
Exhaustion is the common case, not exfiltration
If you run this service, the thing that will actually take it down is not a cleverly exfiltrated secret. It is Tuesday, and somebody's template is eating the host.
What makes this harder than it sounds is that compile jobs are CPU and memory pigs by nature. A malicious submission and an enthusiastic one produce the same telemetry. "This process has been at 100% of eight cores for forty seconds and has six gigabytes resident" describes both an attack and a perfectly reasonable C++ translation unit with too many headers in it. There is no signal to alert on, because the signal is also the product.
So the bound has to be structural rather than diagnostic, and it wants to exist at two levels: inside the guest, where a limit produces a clean error message you can show the submitter, and at the VM edge, where a limit produces a dead VM. You want both, because the first one is good UX and the second one is the one that holds.
One finding from our own source belongs here, because it changes how you write the in-guest half. One-shot `exec()` takes a `timeout_seconds` and the agent does not apply it: the handler decodes the field and then runs the command with no deadline and no cancellation path. The client's 30-second default gives up on the HTTP response, and the compiler carries on burning CPU inside the guest until it finishes on its own. For most workloads that is a footnote. For a service whose whole failure mode is unbounded compiles, it is the difference between a bounded request and a VM you are paying for while nobody is watching. `exec_stream` does propagate cancellation to the guest. Better still, bound it in the shell, which is where the next code block does it.
Multi-toolchain hosting is a template-catalogue problem
The actual product requirement is never one compiler. It is "GCC 9 through 15, Clang 11 through 21, three Rust channels, and somebody will ask for an old MSVC you cannot legally ship". That is dozens of compilers times dozens of versions, and how you package it decides your disk bill and your latency.
One fat image containing everything is the obvious first move and the wrong one. Every request pays the disk footprint of every toolchain it did not want, the image becomes a thing nobody dares rebuild, and a single apt conflict between two versions blocks the whole catalogue. The shape that works is one template per toolchain *family* — all your GCC versions together, because they share a libc, a binutils and a sysroot and a request that wants GCC plausibly wants a choice of GCC — and separate templates for Clang, for Rust, for Go.
Two build flags deserve more attention than usual. `--size-mb` is the rootfs size and it defaults to 1024 MB on the build API (the bundled Go CLI sends 2048), which will not hold three GCC versions with headers and binutils; the failure is loud at bake time rather than subtle later, but it catches everyone exactly once. And `--memory-mb` is the only place guest RAM is ever chosen, because Firecracker cannot change a guest's RAM at snapshot restore — the agent silently corrects whatever `cpu` or `memory_mb` a create request asks for to the baked values. For most workloads that is the most annoying constraint in the platform. Here it is the single most useful property available: the memory ceiling is not a parameter the submitter can raise, because it is not a parameter at all.
#!/usr/bin/env bash
# One template per TOOLCHAIN FAMILY, several versions inside it. The unit of
# baking is "the set of compilers a single request might reasonably need",
# because a request picks one compiler and a template boots as a whole.
set -euo pipefail
cat > Dockerfile.gcc <<'DOCKERFILE'
FROM pandastack/base:latest
# Several GCC versions in one image. They share a libc, a binutils and a
# sysroot, which is exactly why they belong in one template -- and why Clang
# and Rust do not, they would only make this image bigger for every request
# that wanted GCC.
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc-12 g++-12 gcc-13 g++-13 gcc-14 g++-14 \
binutils libc6-dev && rm -rf /var/lib/apt/lists/*
# A compile account with no business anywhere. Resource limits in Linux are
# largely per-UID, so a shared uid means one submitter's fork bomb is
# everybody's fork bomb for the life of this (very short-lived) guest.
RUN useradd -m -s /bin/bash -d /home/compile compile \
&& install -d -o compile -g compile -m 0700 /work
# Docker ENV does NOT reach a detached process in the Firecracker guest --
# only /etc/environment (PAM) does. This line is the difference between a
# working PATH and a baffling "gcc-14: not found" at request time.
RUN printf '%s\n' \
'PATH=/usr/lib/gcc-14/bin:/usr/local/bin:/usr/bin:/bin' \
'TMPDIR=/work/tmp' \
'SOURCE_DATE_EPOCH=1700000000' >> /etc/environment
# Nothing in here needs to resolve a name at request time. Make that explicit
# rather than leaving a resolver configured and hoping nobody calls it.
RUN : > /etc/resolv.conf
DOCKERFILE
# --size-mb DEFAULTS TO 1024. Three GCC versions plus headers plus binutils
# is not going to fit in a gigabyte, and the failure is a build error at
# bake time rather than anything subtle -- but it catches everyone once.
# Measure the image and size this deliberately.
# --memory-mb is the ONLY place guest RAM is chosen. Firecracker cannot change
# a guest's RAM at snapshot restore, so the agent silently corrects whatever
# cpu/memory_mb a create request asks for to the baked values. For a compile
# service that normally-annoying constraint is the single most useful
# property in this file: the memory ceiling is not a parameter the submitter
# can raise, because it is not a parameter at all.
# --cpu is deprecated and ignored. Every template gets 8 burstable vCPUs.
pandastack template build \
--name toolchain-gcc \
-f Dockerfile.gcc \
--context . \
--size-mb 8192 \
--memory-mb 4096
# Then the same shape for the families you actually offer. Each one is its own
# snapshot, its own disk, and its own line item.
# pandastack template build --name toolchain-clang -f Dockerfile.clang ...
# pandastack template build --name toolchain-rust -f Dockerfile.rust ...
# pandastack template build --name toolchain-go -f Dockerfile.go ...
Note the `/etc/environment` line, which is the gotcha that costs an afternoon. Docker `ENV` is not injected into the Firecracker guest; only `/etc/environment`, read through PAM, reaches a detached guest process. Set `PATH` in the Dockerfile the usual way and your request-time compile fails with `gcc-14: not found` in an image that demonstrably contains `gcc-14`.
Why a cold VM per request is affordable here
The architecture a compiler explorer wants is one VM per compile request, destroyed afterwards, no reuse. Nothing persists between submitters, there is no state to scrub, and a submission that wedged its toolchain wedged a machine with a sixty-second life expectancy. The reason nobody builds it that way is that it has historically been too slow: a cold container start plus however long the toolchain takes to become usable is a latency budget a human notices, and this is a feature people drive interactively.
Snapshot restore is what moves it. There is no warm pool of idle VMs here; every create restores the template's baked Firecracker snapshot on demand, at p50 179 ms and p99 203 ms, of which the `/snapshot/load` step itself is around 49 ms. The first create of a template before its snapshot exists is a real cold boot around 3 seconds, after which the snapshot is captured and every create lands on the fast path. The toolchain is not installed at request time, or warmed at request time; it is a block device that was baked once and is reflinked now.
Which means the per-request cost of a fresh VM sits below the cost of the compile it is there to contain. That is the inequality the whole design depends on, and it is the one that was not true five years ago.
import hashlib, json, shlex
from pandastack import Sandbox
# Dozens of compilers x dozens of versions is a ROUTING table over templates,
# not one fat image. The family decides which snapshot gets restored; the
# binary name decides what runs inside it.
FAMILIES = {
"gcc": "toolchain-gcc",
"clang": "toolchain-clang",
"rust": "toolchain-rust",
}
# Flags are user input. On a compiler explorer they are THE PRODUCT -- the
# whole point is that people get to try -O3 -march=native -fno-exceptions --
# so an allowlist of exact flags is a worse product and a denylist of
# dangerous ones is a game you lose to the next release. Reject the handful
# that are not "compile this", keep the VM, and stop pretending the list is
# the boundary.
BANNED_PREFIXES = ("-fplugin", "-load", "-Xclang", "-specs=", "-B", "-wrapper",
"@") # @file makes the compiler read a response file
def sanitize(flags: list[str]) -> list[str]:
for f in flags:
if f.startswith(BANNED_PREFIXES) or "\n" in f:
raise ValueError(f"flag not accepted: {f!r}")
return flags
# Identical source + identical flags + identical compiler build = identical
# output, so the cache in front of this is where almost all of your latency
# goes. Include the compiler's own version string in the key: a toolchain
# re-bake must invalidate every entry it produced, and "we shipped GCC 14.2
# and the cache kept serving 14.1's assembly" is a bug report you will not
# enjoy reading.
def cache_key(family: str, binary: str, toolchain_build: str,
flags: list[str], source: str) -> str:
h = hashlib.sha256()
for part in (family, binary, toolchain_build, "\x00".join(flags), source):
h.update(part.encode()); h.update(b"\x1e")
return h.hexdigest()
# The compile runs under three nested bounds, in the shell, inside the guest.
# This is not belt-and-braces paranoia -- it is the only layer that reports a
# clean error. The VM's bounds report a dead VM.
#
# timeout -s KILL wall clock. KILL, not TERM: a compiler wedged inside
# template instantiation is not reliably interruptible.
# ulimit -t RLIMIT_CPU. Catches a spin that a wall clock misses
# because the process was descheduled, and it is the
# bound that costs you money.
# ulimit -f file size. A compiler asked for -S can emit assembly
# far larger than its input; ~400 KB of dull array
# initialiser expanded ~22x in a test for this post, and
# that was the laziest possible input.
# ulimit -u processes, per-UID -- hence the dedicated account.
# ulimit -c 0 no core dumps, which are both large and interesting.
#
# ulimit -v (RLIMIT_AS) is the obvious missing one, and deliberately: a
# sanitizer build reserves enormous address-space ranges and will simply
# refuse to start under a tight -v. If you offer -fsanitize=address -- and a
# compiler explorer should -- let the guest's baked RAM be the memory bound
# instead. It is a ceiling nobody can argue with.
COMPILE_SH = r"""
set -u
cd /work
# filesystem.write lands as the guest's SSH user, so hand the inputs and the
# output files to the compile account before dropping to it. Pre-creating
# out.s matters: the compiler opens that path itself, as `compile`, and
# /work is 0700. diag.txt is opened by THIS shell before the exec, so the
# redirect below works on an inherited fd either way.
install -d -o compile -g compile -m 0700 /work/tmp
chown compile:compile /work/in.{ext} && chmod 0400 /work/in.{ext}
: > /work/out.s && chown compile:compile /work/out.s
ulimit -t 20 -f 262144 -u 64 -c 0
umask 077
exec timeout -s KILL 15 \
setpriv --reuid=compile --regid=compile --clear-groups \
{binary} {flags} -S -o /work/out.s /work/in.{ext} 2>/work/diag.txt
"""
def compile_once(family: str, binary: str, flags: list[str], source: str,
ext: str = "cpp") -> dict:
sbx = Sandbox.create(
template=FAMILIES[family],
ttl_seconds=120, # IDLE timeout, not a walltime cap
metadata={"purpose": "compile-explorer"},
)
try:
# filesystem.write takes ONE file. That is a fine fit here: the
# submission is one translation unit, and anything that wants a
# project tree is a different product with a different threat model.
sbx.filesystem.write(f"/work/in.{ext}", source)
cmd = COMPILE_SH.format(
binary=shlex.quote(binary),
flags=" ".join(shlex.quote(f) for f in sanitize(flags)),
ext=ext,
)
# exec_stream, not exec. One-shot exec() accepts timeout_seconds and
# the agent never applies it: the handler reads the field, then runs
# the command with no deadline and no cancellation path, so the
# client's 30 s default abandons the HTTP response while the compiler
# keeps burning CPU in the guest until it finishes. exec_stream's
# context cancellation does reach the guest (SIGTERM), and the shell
# bounds above hold regardless. Belt, braces, and a trapdoor.
chunks: list[str] = []
rc = sbx.exec_stream(cmd, on_stdout=chunks.append,
on_stderr=chunks.append, timeout_seconds=30)
# Read back bounded. The output is attacker-influenced text: string
# literals survive into .asciz byte-for-byte, escape sequences
# included, and #pragma message puts an arbitrary string on stderr.
# Treat all of it as data on the way out -- text/plain, nosniff, no
# interpolation into markup -- and cap it before it reaches a browser.
diag = sbx.exec("head -c 65536 /work/diag.txt", timeout_seconds=10)
asm = sbx.exec("head -c 1048576 /work/out.s", timeout_seconds=10)
return {
"exit_code": rc,
"stderr": diag.stdout,
"asm": asm.stdout,
"truncated": len(asm.stdout) >= 1048576,
}
finally:
# Unconditional. A compile that wedged the guest is exactly the case
# where you most want the teardown to be somebody else's problem.
sbx.kill()
if __name__ == "__main__":
out = compile_once("gcc", "g++-14", ["-O2", "-std=c++20"],
"int square(int x){ return x*x; }\n")
print(json.dumps(out, indent=2)[:2000])
The output path is attacker-influenced too
You are returning disassembly, IR and diagnostics to a browser. All three are derived from submitted input, and the derivation preserves more than people expect.
- String literals survive into the assembly byte-for-byte. A literal containing terminal escape sequences comes out the other side as an `.asciz` directive containing terminal escape sequences — fine in a `<pre>`, considerably less fine in anyone's terminal, and exactly the sort of thing that gets piped into one during debugging.
- `#pragma message` places an arbitrary string on stderr with no transformation. I confirmed that a pragma message containing `</script>` and an `onerror` handler arrives on stderr intact. If your frontend interpolates diagnostics into markup rather than rendering them as text, you have an XSS bug whose payload is delivered by your own compiler.
- Identifiers reach the output through mangled and demangled symbol names, so the attacker has partial control of tokens in a region of the response that looks structural.
- Volume is a denial-of-service vector on its own. Roughly 400 KB of dull array initialiser expanded to about 9 MB of assembly in a test for this post — a 22x amplification from the laziest input imaginable, and a motivated one does far better. Cap the bytes you read out of the guest, cap what you send to the client, and say "truncated" rather than streaming until something falls over.
The rule is boring and absolute: compiler output is untrusted data on the way out, exactly as source was on the way in. Serve it as `text/plain` with `nosniff` if you serve it raw, render it as text and never as markup, and cap it. The symmetry is the thing worth internalising — most people harden the input side thoroughly and treat the output side as their own data, because it came out of their own binary.
Cache in front, and be honest about what the VM is for
Identical source plus identical flags plus an identical compiler build produces identical output. That makes a content-addressed cache in front of the whole apparatus not an optimisation but the primary serving path, and it is where essentially all of your latency improvement lives. A link shared on social media is one cache entry served ten thousand times; a demo in a conference talk is one entry served by a room full of people who all clicked the same thing.
Two things to get right. The key must include the toolchain's own version or build identifier, not just the family and the flags, so that re-baking a template invalidates everything it produced — "we shipped the new GCC and the cache kept serving the old assembly" is a confusing bug to receive. And the cached value is still attacker-influenced text, so it gets the same treatment on the way out as a fresh result; a cache is not a laundering step.
Which leads to the honest framing of this entire architecture: the sandbox exists for the cache misses. On a popular instance that is a small fraction of traffic, and it is the fraction containing every submission nobody has ever compiled before — which is also, precisely, the fraction containing anything interesting. The cache serves the traffic. The VM serves the risk.
A compiler service wants no egress at all
Ask what a compile of a single translation unit legitimately needs from the network. Nothing. Not a package index, not a name server, not a clock. A compiler explorer is the rare workload where "no egress" is not a hardening posture you grudgingly adopt but a plain statement of the requirement, and treating it that way removes most of the abuse economy at a stroke: no crypto mining with a useful destination, no proxy relay, no beacon for whatever the preprocessor managed to read.
Mechanically, every sandbox gets its own Linux network namespace with its own interfaces and its own rules — the agent pre-allocates 16,384 /30 subnets per host so that namespace is already built when a create arrives, rather than costing around 100 ms of `ip` and `iptables` calls on the request path. Two rules in that scheme are worth knowing because they are the ones you would otherwise have to write yourself: traffic from the sandbox pool to the sandbox pool is dropped, so one guest cannot reach another, and traffic to 169.254.0.0/16 is dropped, which is the cloud metadata endpoint that turns a file-read primitive into a credential.
What each boundary actually stops, for this workload specifically
| The thing that happens | Container + seccomp | gVisor | Per-request microVM |
|---|---|---|---|
| Compiler reads a path it can name (`#include`, `#embed`, `.incbin`) | Only what the mount namespace excludes — and a toolchain needs a large, permissive mount | Same: this is a filesystem question, not a syscall question | Same question, but the filesystem is a private block device you baked, containing nothing else |
| `-fplugin=` dlopens a path the submitter chose | Permitted; it is an ordinary open plus mmap | Permitted | Permitted, and confined to a kernel nothing else is using |
| Template instantiation allocates until something dies | cgroup memory limit; the OOM killer then picks a victim, not always the right one | Same cgroup story, plus the sentry's own footprint | Guest RAM is baked into the snapshot and cannot be changed at restore — a ceiling with no parameter attached |
| A CPU spin that is indistinguishable from a real build | cgroup `cpu.max`; noisy-neighbour effects still land on co-tenants | Same, plus interception overhead on every syscall | 8 burstable vCPU per guest, and CPU billed on the seconds actually burned |
| A kernel bug reached from the compiler process | One shared kernel, one blast radius, every tenant in it | A reimplemented syscall surface in userspace: smaller, and its own code to trust | A separate guest kernel behind KVM; the host-facing surface is the VMM |
| Stopping a wedged compile instantly, without cooperation | SIGKILL the process tree, assuming you can still enumerate it | Kill the sandbox process | Kill the VM; there is nothing to enumerate and nothing to negotiate with |
| Hosting forty compiler versions | Image layers, and a genuinely useful shared page cache across containers | As containers | One baked template per family; disk is the cost and snapshot restore is the payoff |
The row that decides it is not the kernel-bug row, which is the one these comparisons usually turn on. It is the memory row. In every shared-kernel option, the compile's memory ceiling is a policy you configure and can misconfigure, enforced by a killer with its own opinions about which process deserves to die. In the microVM, it is a number that was written into a snapshot at bake time and that nothing in the request path can address. For a workload whose defining characteristic is an unbounded appetite for RAM, a ceiling nobody can argue with is worth more than a smaller syscall surface.
What this does not fix
- The catalogue is a maintenance commitment, not a build step. Forty toolchain-version combinations is forty templates to re-bake when a base image gets a CVE, and the honest cost of a compiler explorer is mostly this, not the isolation.
- You still have to decide about plugins and sanitizers, and both are things users genuinely want. `-fsanitize=address` is a reasonable feature request and it fights the memory limits in the same breath. `-fplugin=` is a reasonable feature request from the exact users you built this for, and it is remote code execution. There is no configuration that makes both of those comfortable.
- Compiling is not running, and if you also offer an Execute button you have a second, larger threat model bolted onto the first. Deal with it as a separate thing with separate limits, not as a checkbox on the compile path.
- Cold-boot amortisation has a floor. Snapshot restore puts a fresh VM at p50 179 ms, which is excellent and is not zero. Interactive type-and-see-the-asm needs the cache to be doing its job; if your traffic is overwhelmingly unique submissions, you are paying the restore on nearly every keystroke-triggered compile and should debounce hard.
- `fork()` on a sandbox is disk-only — children cold-boot from a clone of the parent's rootfs with no parent memory and no running processes. If you reach for forking to amortise anything here, know which primitive you have: `fork_tree(count)` does carry memory and disk but caps at 16 children pinned to the parent's host, and wider fan-out means `snapshot()` then creating from the returned id.
- Guest clocks are frozen at bake time and re-synced best-effort on restore, so a toolchain that validates a TLS certificate in a build step can fail with "certificate not yet valid" rather than "expired" — a diagnostic that sends people looking in entirely the wrong direction. Mostly irrelevant for a no-egress compile, which is one more argument for no egress.
- None of this helps with the licence question. If somebody asks for an old proprietary compiler, your problem is a licence agreement, and no amount of isolation engineering is the answer to it.
The shape of the thing
A content-addressed cache serving almost all the traffic. Behind it, one microVM per cache miss, restored from a snapshot of a template baked per toolchain family, with the compile bounded in-shell under a hard-kill wall clock and a CPU limit, running as a dedicated uid, with no egress and nothing on the disk worth reading. Output capped, treated as text, never interpolated into markup. The VM destroyed afterwards because it is cheaper to destroy it than to reason about what it became.
That is not an elaborate architecture. It is roughly the architecture you would have drawn anyway for running untrusted *programs* — which is the actual lesson, and the reason the opening instinct is so expensive.
A program does what its author wrote. A compiler does what its input tells it to. You did not take delivery of a safer workload by refusing to run the binary; you took delivery of a larger one, and then pointed it at your filesystem.
Frequently asked questions
Is compiling untrusted code really as dangerous as running it?
Treat it as at least as dangerous, and for some toolchains worse. The preprocessor is an interpreter with filesystem access that runs before any notion of "the program" exists: `#include` of an absolute path gets the file's contents echoed into diagnostics you are about to render in a browser, `__has_include` is a silent filesystem oracle that exits zero, `#embed` puts a whole file into the artefact, and `.incbin` in inline assembly does the same through the assembler. Add `-fplugin=`, which asks the compiler to dlopen a path the submitter named, and Rust's `build.rs` and procedural macros, which are native code the build system and the compiler respectively execute by design. The one place the intuition holds is compile-time evaluation: C++ `constexpr` and Rust's const evaluator genuinely cannot do I/O — ask `constexpr` to call `fopen` and the compiler refuses — but they are unbounded in resources, which is a denial-of-service surface rather than an exfiltration one. The useful summary is that a program does what its author wrote while a compiler does what its input tells it to, and on a compiler explorer the input includes the command line.
Can't I just put resource limits on the compiler instead of using a VM?
You should do both, and you should understand why the flags are not the boundary. Clang accepts `-ftemplate-depth=`, `-fconstexpr-steps=`, `-fconstexpr-depth=` and `-fbracket-depth=`; GCC has `-ftemplate-depth=` and `-fconstexpr-ops-limit=`. Those are excellent against accidental blowups from cooperative users. But a compiler explorer's entire product is that the submitter chooses the flags, so any limit you express as a flag is one they can raise by passing it again with a bigger number, and any flag you ban is a denylist entry the next release routes around. The limits that actually hold are the ones outside the process they control: a hard-kill wall-clock `timeout`, `ulimit -t` for CPU seconds, `ulimit -f` so a request for assembly cannot fill the disk, a dedicated uid so `ulimit -u` means something, and above all a memory ceiling the submitter cannot address. Deliberately skip `ulimit -v` if you offer sanitizers, since a sanitizer build reserves enormous address-space ranges and will not start under a tight RLIMIT_AS — let the guest's baked RAM be the memory bound instead.
How should I package forty compiler versions — one image or many?
Many, grouped by toolchain family rather than by individual version. All your GCC versions belong in one template because they share a libc, a binutils and a sysroot, and because a request that wants GCC plausibly wants a choice of GCC; Clang, Rust and Go each get their own. One fat image containing everything makes every request pay the footprint of every toolchain it did not want, turns the image into something nobody dares rebuild, and lets a single apt conflict between two versions block the whole catalogue. Size the rootfs deliberately when you build: `--size-mb` defaults to 1024 MB, which three GCC versions with headers and binutils will not fit into. And pick guest RAM at build time with `--memory-mb`, because Firecracker cannot change a guest's RAM at snapshot restore and the agent silently corrects whatever a create request asks for to the baked values. Budget for the catalogue being the real ongoing cost: forty combinations is forty templates to re-bake when a base image picks up a CVE.
Is a cold VM per compile request actually fast enough for an interactive tool?
With snapshot restore, yes, and the honest answer has two parts. A create restores the template's baked Firecracker snapshot on demand at p50 179 ms and p99 203 ms, with the `/snapshot/load` step itself around 49 ms — there is no warm pool, every create takes that path, and the toolchain is a block device baked once rather than anything installed or warmed at request time. The first create of a template before its snapshot exists is a real cold boot around 3 seconds, after which the snapshot is captured and subsequent creates are on the fast path. That puts the per-request VM cost below the cost of the compile it contains, which is the inequality the design needs. The second part: a content-addressed cache in front does far more for your p50 than any VM optimisation, because identical source plus identical flags plus an identical compiler build gives identical output, and a shared link is one entry served many times. The sandbox exists for the cache misses — a small fraction of traffic that happens to contain every submission nobody has compiled before.
What do I need to do about the disassembly I send back to the browser?
Treat it as untrusted data, exactly as you treated the source on the way in, because it is derived from the source and the derivation preserves more than you would guess. String literals survive into the assembly byte-for-byte including terminal escape sequences, so output that is harmless in a `<pre>` is not harmless in the terminal somebody pipes it into while debugging. `#pragma message` places an arbitrary string on stderr with no transformation — a pragma message containing `</script>` and an `onerror` handler arrives intact, so a frontend that interpolates diagnostics into markup has an XSS bug delivered by its own compiler. Identifiers reach the response through mangled and demangled symbol names, giving partial control of tokens in a region that looks structural. And volume is its own denial of service: around 400 KB of dull array initialiser expanded to roughly 9 MB of assembly in a test for this post, a 22x amplification from the laziest possible input. Render as text and never as markup, serve raw output as `text/plain` with `nosniff`, cap the bytes you read out of the guest, cap what you send, and label truncation rather than streaming until something falls over.
Keep reading
- Building untrusted Rust code in a microVM — The closest neighbour to this post's thesis: `build.rs` and proc macros are arbitrary code the compiler runs, so `cargo build` on a hostile crate is RCE before your program exists.
- Runnable code playgrounds in your docs — The sibling problem for executing snippets rather than compiling them — the anonymous-visitor threat model, WASM's ceiling, and the latency budget.
- Compiling user-submitted LaTeX safely — Another compiler whose language has a documented shell escape. Same lesson, a different macro expander.
- Cross-compilation build farms on microVMs — If the product is producing binaries for many targets rather than showing assembly for one, the template-per-target-triple version of the catalogue problem.
- PandaStack templates — The first-party catalogue and how a custom template build like the one above works, including the flags that decide rootfs size and guest RAM.
- PandaStack sandboxes — Snapshot restore, the per-sandbox network namespace, and the lifecycle calls this architecture leans on.
Related posts
- Running Customer Trading Strategies in Isolated microVMs
Your customers write strategies. Your infrastructure runs them, next to each other, holding credentials that place orders. This is the untrusted-code problem with a P&L attached.
- Running User-Generated Game Mods in Isolated microVMs
A mod is arbitrary code from a stranger that you execute on your infrastructure. "We removed the io library" is not a security boundary — a guest kernel is.
- Testing Against Ten Toolchains Without Ten Broken Runners
If you ship a library you own a matrix. On a shared runner the legs quietly contaminate each other. A microVM per leg makes the matrix mean what it says.
- CircleCI Self-Hosted Runners on MicroVMs
The moment you move a CircleCI job onto your own machine runner, you quietly trade a fresh VM per job for a box that remembers every build that ever ran on it. That trade is the whole security story, and you do not have to make it.
- Running Customer Git Hooks in Isolated microVMs
A git hook is a shell script your user wrote that your server agreed to run. The exploit is not the scary part — the scary part is a policy hook that greps a 4 GB monorepo on every push and takes the whole push queue with it.
More in Code execution · See Code interpreter sandboxes on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.