all posts

Compiling a Stranger's Shader: Isolate the Toolchain, Because You Cannot Isolate a GPU You Do Not Have

Ajay Kumar··10 min read

Somewhere in your product there is a text box, and a stranger types a fragment shader into it. Maybe it is a Shadertoy-style gallery. Maybe it is a WebGPU playground, or a materials marketplace where artists upload node graphs that lower to WGSL, or a creative-coding course where two thousand students submit homework on a Sunday night. Increasingly it is an AI agent that writes a shader, reads the compiler's error output, and tries again — in a loop, at machine speed, forever.

Then your backend takes that text and hands it to `glslangValidator`. That is a C++ preprocessor, a C++ parser, a C++ type checker and a C++ code generator, consuming a string your attacker wrote, in a process that probably also holds your database credentials. After that the SPIR-V goes to `spirv-opt`, which is a different pile of C++ running twenty-odd passes over a binary format, and then to `spirv-cross`, which is a third pile of C++ that parses the same binary again to emit MSL and HLSL. Three independent parsers, three independent fuzzing backlogs, one untrusted input, all running on the same kernel as the rest of your build fleet.

I build PandaStack, an open-source Firecracker microVM platform, and I want to start this post with the thing a vendor blog usually buries: PandaStack has no GPU. There is no PCI or VFIO GPU passthrough. You cannot run a shader on a GPU in a PandaStack microVM, you cannot test a vendor driver's shader compiler, and you cannot measure a frame time. If that is what you need, stop reading and go buy GPU hosts.

The reason to keep reading is that almost none of the pipeline is GPU work. Translating GLSL to SPIR-V, validating a module, optimising it, cross-compiling it to five platform shading languages, and even rasterising a golden image through a software Vulkan implementation are all CPU work — and the CPU work is where the untrusted input lands. The GPU driver is somebody else's attack surface. The compiler is yours.

The short version: you cannot isolate a GPU you do not have, and you do not need to. What you need to isolate is a chain of shader toolchain binaries — glslang, spirv-val, spirv-opt, spirv-cross, Naga, DXC, Slang — each of which is a parser plus an optimiser plus a code generator fed entirely by a stranger. Give each compile job its own kernel and its own hard memory ceiling, fence every binary with `timeout` and `ulimit` in-guest, and treat the compiler's stderr as the product, because it is.

What actually executes when you compile a stranger's shader

Worth being specific, because "shader compiler" sounds like one program and it is a chain of four or five, each with a different shape of exposure.

The preprocessor, which is a macro language with no budget

GLSL has a C-style preprocessor. It has `#define`, it has conditionals, and its macro expansion is the oldest trick in the denial-of-service book: a handful of mutually-referencing object-like macros that expand combinatorially, so forty bytes of source becomes gigabytes of token stream before the parser has seen a single declaration. Implementations have recursion guards; implementations have had bugs in those guards. If your pipeline also supports `#include` through an include resolver you wrote, you now own a path-traversal surface as well, and "the resolver only reads from the shader's own directory" is a sentence that holds right up until someone writes `../`.

The parser and the code generator

This is the ordinary memory-safety story and I will not pretend it is dramatic: a large, mature, heavily-fuzzed C++ front end, where the historical bug class is deeply nested expressions and deeply nested struct or array types blowing the recursive-descent stack, plus the occasional integer overflow in an array-size computation. Heavily fuzzed is genuinely good news. Heavily fuzzed also means there is a public corpus of inputs that used to crash it, and the gap between "fixed upstream" and "deployed in your container image" is however long it has been since you rebuilt.

SPIR-V, if you accept modules instead of source

This is the part people underestimate. The moment your API accepts a pre-compiled `.spv` upload — and plenty of asset pipelines do, because the artist's tool already emitted one — you have stopped accepting a text format and started accepting an untrusted binary format with offsets, ids, type graphs and a control-flow graph. A malformed SPIR-V module is a hostile input to `spirv-opt`, which was written to transform valid modules.

`spirv-val` exists precisely because the downstream consumers are not robust against invalid modules. That is the design intent, stated plainly in the SPIRV-Tools documentation: validate first, because the optimiser and the drivers assume validity. So validation is not an optional quality gate you run in CI, it is the first thing that should touch the bytes. It is also not a safety proof — passing the validator means the module is structurally valid, not that it is cheap, and the resource-exhaustion inputs in the next section are all perfectly valid SPIR-V.

The optimiser's own pass pipeline

`spirv-opt -O` is a sequence of passes, and several of them are superlinear in quantities the attacker controls directly: basic-block count, loop trip count, function-call depth, struct nesting depth. Loop unrolling is the obvious one — a loop whose trip count is a compile-time constant of two million is a legal shader and an instruction to the optimiser to materialise two million copies of the body. Inlining plus dead-code elimination plus store-load forwarding over the result of that is where the memory goes. The same applies to the cross-compiler afterwards: `spirv-cross` has to reconstruct structured control flow and name every temporary, over however many instructions the optimiser produced.

The failure that actually happens is resource exhaustion

I have watched a lot of teams design shader-upload features, and the threat model they bring is always remote code execution through a parser bug. That is a real risk and it is not the one that pages you. The one that pages you is that a shader compiles for nine minutes and allocates 30 GB, and this is not even a security bug. It is a Tuesday.

Four shapes of it, all of which you will see from honest users before you see them from attackers:

  • A loop with a huge constant trip count that the author wrote because they were computing a lookup table at compile time, and which the unroller takes literally.
  • Nested macro expansion in a shader that got generated by another program — a node-graph exporter, a template engine, an LLM — where nobody ever looked at the emitted source.
  • A permutation matrix that is combinatorial by construction. N materials times M quality tiers times K platforms times the `#ifdef` soup that shipped last sprint. Nobody attacked you; multiplication did.
  • A pathological control-flow graph: thousands of tiny blocks, deep switch nesting, a function called from everywhere. Valid, cheap to write, expensive to optimise.

On a shared CI runner, any one of those takes the runner with it. The job that dies is not only the hostile one — it is the three sibling jobs that were sharing the box, plus whatever the kernel's OOM killer decided looked tastiest, which is frequently the thing with the largest resident set and the best reason to exist. If you want the longer version of this argument in a different domain, the zip-bomb post makes the same point about archive extraction: Extracting Untrusted Archives: Zip Bombs, Zip Slip, Symlink Escape. The input is different, the arithmetic is identical.

Why the usual boundaries are partial answers here

The reflex answer is a container, and a container is the wrong boundary for this specific job for two reasons that compound. First, a shader toolchain is several C++ binaries run in a chain, which means the boundary has to hold for the weakest of them, not the strongest. Second, a container's memory limit is a cgroup, which the job can argue with: it is enforced by the host kernel's OOM path, the kill is a kernel decision made under pressure, and the pressure is shared with everything else on that kernel. A container is a polite suggestion to a kernel you also run your build fleet on.

Isolation choices for a single untrusted shader compile job
BoundaryWhat it actually isHonest drawback for a shader compile
The same build container as everything elseNamespaces plus a cgroup on your CI host kernelOne 30 GB optimiser allocation and the runner's other jobs die with it
Separate process plus seccompA syscall filter around glslang in your own serviceFilters syscalls; does nothing about CPU, RAM, or a heap overflow inside the parser
gVisorA userspace kernel intercepting the sandbox's syscallsA real boundary with a real syscall-cost profile, and you now trust the sentry instead
glslang compiled to WASM, in the browserThe compiler runs on the visitor's machine, not yoursPerfect for a playground; produces no server-side artefact you can trust or cache
One microVM per compile jobIts own guest kernel and its own RAM ceiling, deleted afterPer-job boot and teardown cost, no GPU inside, and egress is open unless you fence it

The WASM row deserves more than one line, because for a playground it is genuinely the right answer and I will not pretend otherwise. If the only thing you want is for a visitor to see their own compile errors, compile glslang to WebAssembly and ship it — the untrusted code runs on the untrusted person's machine, which is the cheapest isolation in existence. It stops working the moment you need the artefact: a cached SPIR-V blob, a signed asset, a cross-compiled MSL variant you will serve to other users, or a model's feedback loop that you are paying for. At that point the compile has to happen on your hardware and you are back to this table. The general comparison is at WebAssembly vs MicroVMs for Sandboxed Code.

Everything I say about other people's projects here is qualitative on purpose — substrate, fit, one honest drawback, no numbers I did not measure myself. Shader toolchains in particular move fast: default pass pipelines change, validator rules are added, and the flags I name below have changed spelling before. Verify every specific against the current upstream docs before you build on it, including mine.

The microVM shape: one kernel per compile job

A microVM gives you the two things the compile chain needs and a cgroup cannot provide. The job gets its own guest kernel, so a parser bug that escalates lands in a kernel that exists for this one shader and gets deleted afterwards. And the memory ceiling is the VM's, which is not negotiable from inside: the guest sees 4 GiB because the virtual machine has 4 GiB, and when the optimiser asks for the thirty-first gigabyte it gets a failed allocation inside its own guest, not a pressure event on your build fleet's host.

On PandaStack the per-job arithmetic is a snapshot restore rather than a boot: every create restores a baked template snapshot, p50 179 ms and p99 203 ms, with no warm pool of idle VMs behind it. The first spawn of a template that has no snapshot yet is a real cold boot plus an auto-bake, around 3 seconds, paid once. That matters here because a per-job VM only makes sense if the VM is cheaper than the job, and a 179 ms create against a compile that takes seconds is noise.

Inside the guest, fence everything anyway. Defence in depth is not a slogan in this case, it is the difference between a failed job and a stuck one, because a compiler blocked on something is not burning CPU and will outlive a CPU-time rlimit indefinitely.

#!/bin/sh
# /work/compile.sh -- runs INSIDE the guest. /work/untrusted.frag arrived from
# a stranger, or from a model, which is the same threat model with better
# manners. Every stage below is a separate C++ binary reading attacker bytes.
set -u
cd /work

# --- Cheap fences, applied once to every child of this shell. ---------------
# These cost nothing and they stop the boring 90%: the shader that unrolls a
# loop two million times and asks the optimiser to hold all of it at once.
ulimit -v 4194304   # 4 GiB of address space, in KiB. glslang and spirv-opt are
                    # malloc-happy; runaway macro expansion grows here first.
ulimit -t 180       # 180s of CPU TIME, then SIGKILL. Not the same fence as a
                    # wall clock -- see the timeout(1) calls below.
ulimit -f 2097152   # 2 GiB max file size. A degenerate module can emit a
                    # SPIR-V blob orders of magnitude larger than its source.
ulimit -c 0         # No core dumps. A crashing compiler should not write 4 GiB
                    # of itself into the rootfs on the way out.

# 1. GLSL -> SPIR-V. Preprocessor + parser + type checker + codegen, all fed by
#    the stranger. timeout(1) is the WALL-CLOCK fence: ulimit -t only counts
#    CPU, so a compiler wedged in a futex or a read() sleeps there forever.
#    --kill-after sends SIGKILL if it ignored the SIGTERM, which C++ that is
#    deep in a recursive descent sometimes effectively does.
timeout --kill-after=10s 60 \
  glslangValidator -V --target-env vulkan1.2 \
    -o out.spv untrusted.frag 2> compile.err \
  || { echo "stage=compile rc=$?" >&2; exit 10; }

# 2. VALIDATE BEFORE YOU OPTIMISE. spirv-val exists because the consumers of
#    SPIR-V -- optimisers, cross-compilers, drivers -- assume valid input. If
#    you accept pre-compiled .spv uploads at all, this is the first thing that
#    touches the bytes, and it is not optional.
timeout --kill-after=10s 30 \
  spirv-val --target-env vulkan1.2 out.spv 2> val.err \
  || { echo "stage=validate rc=$?" >&2; exit 11; }

# 3. Optimise. Several passes are superlinear in things the attacker picks:
#    block count, loop trip count, nesting depth. This is the stage that turns
#    forty lines of GLSL into nine minutes of CPU, so it gets the longest
#    leash and the hardest stop.
timeout --kill-after=10s 180 \
  spirv-opt -O --strip-debug out.spv -o out.opt.spv 2> opt.err \
  || { echo "stage=optimize rc=$?" >&2; exit 12; }

# 4. Cross-compile to the platform shading languages. A separate codebase
#    parsing the same binary format again -- separate bugs, same bytes.
for t in msl hlsl; do
  timeout --kill-after=10s 60 \
    spirv-cross --$t out.opt.spv --output out.$t 2>> cross.err \
    || { echo "stage=cross:$t rc=$?" >&2; exit 13; }
done

# 5. stderr IS the product. The user reads compile.err to fix their shader and
#    the agent reads it to fix its next attempt. Nobody reads out.spv. Cap it,
#    because an error message can quote arbitrary attacker source back at you.
head -c 65536 compile.err val.err opt.err cross.err > diagnostics.txt
echo ok

Note which fence does which job, because people pick one and are surprised. `ulimit -v` bounds the allocation; `ulimit -t` bounds CPU burned; `timeout` bounds wall clock. A compile that is spinning dies to the second. A compile that is blocked dies to the third. A compile that is allocating dies to the first, inside its own guest, where the failure is a `std::bad_alloc` and not an OOM kill on a box with other tenants.

Driving it from outside, and treating stderr as the deliverable

The orchestration side is small, and the one thing it must get right is the timeout story. On PandaStack, one-shot `exec(timeout_seconds=)` does not actually enforce the timeout, and `exec_stream` only widens the HTTP read timeout — neither of them kills the process in the guest. So the in-guest `timeout` in the script above is doing all of the real enforcement, and the outer bound exists only so your own worker does not hang.

# One untrusted shader, one microVM, and the compiler's stderr handed back as
# the actual product of the job.
import pathlib
from pandastack import Sandbox

COMPILE_SH = pathlib.Path("compile.sh").read_text()

def compile_shader(source: str, upload_id: str) -> dict:
    sbx = Sandbox.create(
        template="base",        # 4 GiB / 8 vCPU -- BAKED into the snapshot.
                                # Firecracker cannot change vCPU or RAM at
                                # restore, so a cpu=/memory_mb= you pass here
                                # is overridden to match the baked values.
                                # "give the optimiser more RAM" is a template
                                # you bake, not an argument you pass. Current
                                # sizes are on /templates.
        ttl_seconds=600,        # An IDLE clock, not a wall clock: exec traffic
                                # resets it, a status GET deliberately does
                                # not. When it fires, the reaper DELETES this
                                # VM -- which is exactly what we want, since
                                # everything we care about is read out below.
        metadata={"job": "shader-compile", "upload": upload_id},
    )
    try:
        sbx.filesystem.write("/work/untrusted.frag", source)
        sbx.filesystem.write("/work/compile.sh", COMPILE_SH)

        diags: list[str] = []
        # The in-guest `timeout` calls in compile.sh do the real enforcement.
        # exec_stream's timeout_seconds only widens the HTTP read timeout and
        # one-shot exec() does not enforce it at all -- neither reaches into
        # the guest and kills anything. Belt (ulimit) and braces (timeout),
        # both in-guest; this outer bound just protects the worker.
        rc = sbx.exec_stream(
            "timeout --kill-after=10s 420 sh /work/compile.sh",
            on_stdout=lambda chunk: None,
            on_stderr=diags.append,       # <-- this is the deliverable
            timeout_seconds=480,
        )
        if rc != 0:
            # A failed compile is a NORMAL outcome, not an incident. Hand the
            # diagnostics to the user, or feed them straight back to the model
            # as the next turn of its loop. This is the whole reason the job
            # exists; the SPIR-V is a bonus.
            return {"ok": False, "rc": rc, "diagnostics": "".join(diags)}

        return {
            "ok": True,
            "spirv": sbx.filesystem.read("/work/out.opt.spv"),          # bytes
            "msl": sbx.filesystem.read("/work/out.msl").decode("utf-8"),
            "diagnostics": "".join(diags),
        }
    finally:
        # Teardown is kill(). There is no sbx.delete(). Do it in finally, or
        # you will learn about idle billing from an invoice.
        sbx.kill()

One design note that is easy to miss: the compiler's stderr is the only thing in this system that an attacker gets to influence and you then show to someone. It can contain arbitrary source text quoted back from their shader. Cap its size, do not interpolate it into HTML without escaping, and if you are feeding it to a model, remember that the stranger's shader comments are now in your prompt.

The permutation matrix: one warm toolchain, many children

The gallery case is one shader at a time. The game-engine case is a permutation matrix: every material against every quality tier against every platform, which is thousands of compiles that all want the same toolchain, the same headers and the same warm page cache. Installing glslang and SPIRV-Tools a thousand times is the slow, expensive way to do this, and it is what most pipelines do.

Build one warm VM, snapshot it, and fan out from the snapshot. The call you want is `fork_tree`, and the distinction here is the single most-misused pair in the whole API: `fork_tree(n)` snapshots the parent once and boots n children that inherit its memory, while `fork()` clones the disk only and the child cold-boots. Disk-only is not what you want when the valuable state is a hot page cache and a loaded set of binaries. The longer treatment is at Snapshots and Forks: Copy-on-Write for Running Machines.

# A permutation matrix is combinatorial by construction: materials x quality
# tiers x platforms x whatever #ifdef soup shipped last sprint. The toolchain
# install is identical for every cell, so pay for it exactly once.
from pandastack import Sandbox

# 1. ONE warm root: toolchain installed, headers unpacked, page cache hot.
root = Sandbox.create(template="base", ttl_seconds=1800,
                      metadata={"role": "shader-matrix-root"})
root.exec("apt-get update -qq && apt-get install -y glslang-tools spirv-tools",
          check=True)
root.filesystem.write("/work/compile.sh", COMPILE_SH)
root.exec("glslangValidator --version && spirv-val --version", check=True)

# 2. Freeze it. snapshot() is synchronous and writes the FULL guest memory
#    image, so budget real seconds on a 4 GiB guest -- fine once per toolchain
#    version, unacceptable per cell. Pleasant side effect: a snapshot pins the
#    exact glslang / SPIRV-Tools / spirv-cross builds, which is what makes the
#    compiler reproducible. That matters a lot if you cache compiled artefacts
#    by source hash, because the hash has to cover the compiler too.
snap = root.snapshot()
print("toolchain snapshot:", snap)
# Tomorrow's matrix run skips step 1 entirely:
#   root = Sandbox.create(from_snapshot=snap)   # restore, p50 179 ms

# 3. Fan out. fork_tree(n) snapshots the parent and boots n children that
#    INHERIT ITS MEMORY -- warm page cache, binaries already resident,
#    copy-on-write until each child writes. The parent is untouched
#    (pause -> snapshot -> resume).
#
#    fork() is the OTHER call and it is not this one: disk-only clone, and the
#    child COLD-BOOTS with its own entropy and its own PIDs. You would keep
#    the installed toolchain and throw away every byte of warm state.
#
#    Hard cap: 16 children per fork_tree() call. Asking for 17 is an error,
#    not a clamp, so a wider fan-out is a breadth-first loop.
def fan_out(parent, total):
    frontier, kids = [parent], []
    while total > 0 and frontier:
        node = frontier.pop(0)
        batch = node.fork_tree(min(16, total))
        kids += batch
        frontier += batch
        total -= len(batch)
    return kids

PERMUTATIONS = [
    {"defines": "-DSHADOWS=1 -DFOG=0", "target": "msl"},
    {"defines": "-DSHADOWS=0 -DFOG=1", "target": "hlsl"},
    # ... the other 38
]

workers = fan_out(root, len(PERMUTATIONS))   # 40 cells -> three rounds
try:
    for w, perm in zip(workers, PERMUTATIONS):
        # Still fenced in-guest. A permutation that unrolls badly kills its own
        # VM and nothing else; the other 39 cells never notice.
        w.exec_stream(
            "timeout --kill-after=10s 300 sh /work/compile.sh "
            + perm["defines"] + " --target " + perm["target"],
            on_stderr=print,
        )
        open(f"out/{perm['target']}.txt", "wb").write(
            w.filesystem.read("/work/diagnostics.txt"))
finally:
    for w in workers:
        w.kill()
    root.kill()

Software rasterization for golden images, and what it costs you

You can go one step past "it compiled" without a GPU. Mesa's `llvmpipe` gives you an OpenGL implementation on the CPU, and `lavapipe` gives you a Vulkan one, both of which run fine in a microVM because they are just LLVM-generated code running on cores. That is enough to render a shader to an offscreen buffer at a fixed resolution with a fixed set of uniforms and compare the result to a stored reference — a golden-image test, which catches the class of regression where a shader still compiles and no longer draws the right thing.

It is slow. Software rasterization is one of those things where the honest number is "orders of magnitude," and I will not invent a multiplier for your shader. Plan for a golden-image pass to be a batch job measured in seconds per frame at modest resolutions, not an interactive preview, and keep the resolution small on purpose — 256 by 256 catches almost every structural regression and costs a fraction of 1080p.

Two honest caveats. First, `lavapipe` is not a vendor driver: a shader that validates and renders correctly under it can still fail on a real Adreno or RDNA driver, because the real driver has its own compiler with its own bugs and its own limits. Software rasterization proves your shader is self-consistent, not that it is portable. Second, floating-point results differ between implementations — precision, rounding of transcendentals, interpolation — so an exact-pixel comparison between a CPU reference and real hardware will fail for reasons that are nobody's fault. Golden images are only stable against the same rasteriser, which in practice means your golden reference has to be a lavapipe render too.

Caching, and the economics of a VM per compile

The thing that makes this affordable is that shader compilation is beautifully cacheable. The input is a string and a set of flags; the output is deterministic for a fixed toolchain. Key your cache on a hash of the preprocessed source plus the flags plus the exact toolchain versions, and the steady state of a gallery or a marketplace is overwhelmingly cache hits. You only spin a VM for a cache miss — which is to say, for a shader nobody has compiled before, which is exactly the shader you least want running on a shared kernel.

That toolchain-version component is not pedantry. If you cache by source hash alone and then upgrade SPIRV-Tools, you are serving artefacts built by a compiler that no longer exists, and the day a pass changes behaviour you get a bug report that is impossible to reproduce. This is the quiet argument for the snapshot approach above: a snapshot of a VM with pinned toolchain builds is a reproducible compiler, and its snapshot id is a perfectly good cache-key component.

On cost, the shape to understand is that CPU bills on CPU-seconds actually burned while memory bills committed GiB-hours for as long as the sandbox exists. A compile job is a CPU burst followed by nothing, which is the good side of that split — but a VM you forgot to `kill()` is paying for 4 GiB of committed RAM while doing nothing at all. At the published $0.0162 per GiB-hour and $0.054 per vCPU-hour, one forgotten sandbox is pocket change and two hundred of them is a line item somebody is going to ask you about. The `finally: sbx.kill()` in the code above is not a style preference. Current rates are on pricing.

The honest limits

  • No GPU, and no path to one. There is no PCI or VFIO passthrough on PandaStack, so anything GPU-accelerated simply does not run. That rules out three things people will ask you for: testing a vendor driver's own shader compiler, validating against real hardware limits and extension support, and measuring performance. If any of those is your requirement, you need real GPU hosts and this post cannot help you.
  • Software rasterization is a different thing from a GPU, not a slower one. llvmpipe and lavapipe have their own precision behaviour and their own feature support, so a golden image rendered on the CPU is only comparable to other CPU renders.
  • The guest kernel is 5.10 and the guest userland is Ubuntu 24.04. Most shader toolchains do not care at all; newer Mesa builds and anything that wants a recent kernel interface might. Find out before you design around it rather than after — You Ship the Kernel: Firecracker Guest 5.10 vs 6.1 has the detail.
  • Egress is open by default. Sibling sandboxes cannot reach each other's subnets and the cloud metadata range is dropped at the host, but your own VPC, your internal registries and your databases are not fenced, and that work is yours. A compile job genuinely has no reason to reach the network, so this is one of the easier cases to lock down: Controlling Network Egress for Untrusted Code.
  • vCPU and RAM are baked into the template snapshot and Firecracker cannot change them at restore, so a per-create `memory_mb` is overridden to the baked value. Sizing a compile VM is a template-baking exercise, not a parameter.
  • The compiler's diagnostics are attacker-influenced output. Size-cap them, escape them before display, and think about what it means that they are now in your model's context window.
  • `spirv-val` is a validity check, not a cost check. Every resource-exhaustion input in this post is valid SPIR-V. Validation and fencing are two separate jobs and you need both.
  • A per-job VM is not free at extreme scale. A 179 ms create against a multi-second compile is noise; a 179 ms create against a 20 ms cached trivial shader is the dominant cost. Batch the cheap cells per VM and reserve one-VM-per-job for cache misses and for anything a stranger just uploaded.

The summary

If your product compiles shaders that other people wrote, you are running three or four independent C++ parsers over hostile input, and the thing that will actually hurt you first is not a parser bug, it is an honest user whose generated shader unrolls a loop two million times and takes the shared runner down with it.

The GPU is not the part you need to isolate, which is convenient, because PandaStack cannot isolate it — there is no passthrough and no GPU in the guest. What you need to isolate is the toolchain: glslangValidator, spirv-val, spirv-opt, spirv-cross, and whatever else your pipeline chains behind them. That is all CPU work, it all runs fine in a microVM, and a microVM is the boundary that gives each job its own kernel and a memory ceiling it cannot argue with. Validate before you optimise, fence every stage with `timeout` and `ulimit` in-guest because the SDK's `timeout_seconds` does not reach into the guest, hand the stderr back as the product, snapshot the warm toolchain once and `fork_tree` the permutation matrix off it, and cache aggressively with the toolchain version in the key.

And do not use PandaStack for this if what you actually need is a GPU. If you are validating against real driver compilers, checking hardware extension support, or measuring frame times, nothing in this post applies and you should buy GPU machines. If you are running a browser-only playground where the visitor is the only person who ever sees the result, compile glslang to WASM and let their laptop do it. The microVM earns its place in the middle case, which is also the most common one: you need the artefact, you need the diagnostics, the input came from a stranger or from a model, and the compile has to happen on hardware you are responsible for.

Frequently asked questions

Can I actually run the shader on a GPU in a PandaStack microVM?

No. PandaStack has no GPU and no GPU passthrough — no PCI passthrough, no VFIO, nothing. A Firecracker guest here sees virtio devices and that is it, so anything that needs a real graphics or compute device does not run at all. I would rather say that in the second paragraph of a post than have you discover it in week three. What you can do is everything that is CPU work, which turns out to be most of a shader pipeline: GLSL or WGSL to SPIR-V with glslang, glslc or Naga, validation with spirv-val, optimisation with spirv-opt, cross-compilation to MSL, HLSL and GLSL ES with spirv-cross, and software rasterisation through Mesa's llvmpipe for OpenGL or lavapipe for Vulkan, which is slow but real and good enough for golden-image tests. If your requirement is one of the three genuinely GPU-bound jobs — testing a vendor driver's own shader compiler, validating against real hardware limits and extensions, or measuring actual performance — then you need GPU hosts from a cloud that sells them, and the right architecture is usually a small pool of real GPU machines for those jobs plus microVMs for the compile and validate stages, which are the ones that run untrusted input and therefore the ones that need the isolation.

Is spirv-val enough to make an uploaded SPIR-V module safe to optimise?

It is necessary and it is nowhere near sufficient, and conflating those two is the most common mistake I see in asset pipelines that accept pre-compiled modules. What spirv-val gives you is a structural guarantee: ids are defined before use, types are consistent, control flow is structured the way the spec requires, decorations are legal for their targets. That matters enormously, because the SPIRV-Tools documentation is explicit that the downstream consumers — the optimiser, the cross-compilers, the drivers — are written against valid input and are not robust against malformed modules. Running spirv-opt on an unvalidated module is handing attacker-controlled offsets to code that trusts them. So validate first, always, as the first thing that touches the bytes. What validation does not tell you is anything about cost. A module with a two-million-iteration loop, ten thousand basic blocks, or struct nesting forty deep is perfectly valid SPIR-V and will still take the optimiser into superlinear territory and your host into swap. Validity and cost are orthogonal properties, and you need two separate mechanisms: spirv-val for the first, and a hard resource ceiling — a VM's memory limit, ulimit, timeout — for the second.

How do I stop a shader that compiles for nine minutes and allocates 30 GB?

With three fences, because they catch three different failures and people routinely install only one. A wall-clock bound: wrap each toolchain invocation in `timeout --kill-after=10s N`, which handles the compiler that is blocked rather than busy — stuck on a read, spinning in a futex, waiting on something that will never arrive. A CPU-time bound: `ulimit -t`, which handles the compiler that is genuinely burning cores and would otherwise sit just under your wall clock forever. And an address-space bound: `ulimit -v`, which turns the thirty-first gigabyte into a failed allocation inside the guest instead of memory pressure on your host. Then put all three inside a microVM whose RAM ceiling is the virtual machine's, so even if every in-guest fence is misconfigured the worst outcome is one dead VM. The reason the in-guest fences carry the real weight on PandaStack specifically is that the SDK's `timeout_seconds` does not reach into the guest: one-shot `exec()` does not enforce it, and `exec_stream` only widens the HTTP read timeout. Neither kills the process. The enforcement has to be `timeout` and `ulimit`, in the guest, in your script.

For a permutation matrix, should I use fork() or fork_tree()?

`fork_tree()`, and getting this wrong silently costs you the entire benefit, which is why it is worth the paragraph. `fork()` clones the disk only and the child cold-boots: you keep the installed toolchain and the unpacked headers on a copy-on-write disk, and then the child starts from nothing — cold page cache, nothing resident, its own fresh entropy and PIDs. That is not this case. `fork_tree(n)` snapshots the parent once and boots n children from that snapshot, so each child inherits the parent's memory: the warm page cache, the binaries already resident, the whole working set, shared copy-on-write and diverging only as each child writes. For a matrix of a thousand cells that all want the same toolchain, that is the difference between paying for the environment once and paying for it a thousand times. The constraint to design around is the hard cap of 16 children per `fork_tree()` call — asking for 17 is an error, not a clamp — so a wider fan-out is a breadth-first loop that pops a node off the frontier and forks up to 16 more from it. And kill every worker in a `finally`: teardown is `kill()`, there is no `delete()`, and committed memory bills for as long as the sandbox exists.

Can I use llvmpipe or lavapipe to golden-image test shaders without a GPU?

Yes, and it is one of the more useful things you can do in a CPU-only sandbox, with two caveats that decide how you design the test. The mechanism works: Mesa's llvmpipe is a software OpenGL implementation and lavapipe is a software Vulkan one, both of which are LLVM-generated code running on ordinary cores, so they run in a microVM with no special device access. Render your shader offscreen at a fixed resolution with fixed uniforms, hash or diff the result against a stored reference, and you catch the regression class where a shader still compiles and no longer draws the right thing — which compilation alone will never catch. Caveat one: it is slow enough that you should treat it as a batch job rather than an interactive preview, and keep the resolution deliberately small, since 256 by 256 catches almost every structural regression at a fraction of the cost. Caveat two, and this is the one that bites: software rasterisation is a different implementation, not a slower GPU. Floating-point precision, transcendental rounding and interpolation all differ, so an exact-pixel comparison between a CPU render and real hardware will fail for reasons that are nobody's bug. Your golden reference has to be a lavapipe render too, and the test proves self-consistency rather than portability.

Keep reading

  • Untrusted archive extraction and zip bombs — The same resource-exhaustion arithmetic in a different parser. Forty bytes of input, gigabytes of output, one dead runner.
  • WASM vs microVM — Why glslang-in-the-browser is the right answer for a playground and the wrong one the moment you need the artefact.
  • Snapshot and fork, explained — The full treatment of the distinction this post leans on — fork clones disk and cold-boots, fork_tree inherits memory.
  • PandaStack vs gVisor — The other serious boundary for a C++ toolchain on hostile input, and what you are actually trusting in each case.
  • Per-tenant image and thumbnail generation — The sibling case: another chain of decades-old C++ media libraries pointed at files a stranger uploaded.

Related posts

  • A Font Is a Program: Rendering User-Uploaded Typefaces Without Trusting Them

    The moment you accept a .ttf from a user, you have accepted a program: TrueType carries a stack machine, CFF carries another one, and woff2 puts a Brotli decoder in front of both. Here is why that parser should not run in the process holding your database credentials, and the microVM shape that fixes it.

  • Compiling User-Submitted LaTeX Without Handing Over a Shell

    You added LaTeX because the typography is beautiful. You also added a macro language with a documented primitive for running shell commands, and pointed it at strangers.

  • Pickle Is a Virtual Machine You Have Been Downloading

    pickle is a stack machine with an eval loop, and you have been downloading it from the internet as model weights. Here is the disassembly, why the allowlist promises less than it looks, and the conversion gateway that fixes it once instead of forever.

  • Server-Side Rendering Someone Else's React Component

    If your product server-renders components your customers wrote, you are running their code in your process, with your env vars and your database socket. node:vm is a sandbox for values, not for a module graph that can require("fs"). Here is the boundary that actually holds, and what it costs.

  • Self-Hosting a Compiler Explorer: Running Strangers' Compilers Safely

    "I'm not running their code, I'm only compiling it" is the most expensive sentence in this genre. A compiler is a programmable machine with a filesystem, a plugin loader and an unbounded appetite for RAM — and on a compiler explorer, the flags are user input too.

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.