Letting Customers Run Their Own Container Image
The feature request arrives in the same shape every time. Your product has a runtime — a workflow engine, an ETL platform, an ML training service, a CI product, an agent marketplace — and a customer does not want it. They have an image. It holds their CUDA version, or their internal wheel index, or a Fortran binary compiled in 2011 that nobody can rebuild. They would like to run that on your infrastructure, and they are large enough that “no” is expensive.
So you ship “bring your own image”, and what you have shipped is a remote-code-execution endpoint with a product manager attached. That is not a reason not to ship it. It is a reason to be clear about what it is, because the design that follows from “this is a container feature” and the design that follows from “this is an RCE endpoint we have chosen to operate” are different designs, and only one survives the first tenant who is not who they said they were.
I'm Ajay; I build PandaStack, which runs customer code in Firecracker microVMs. Here is the pipeline I would build on any substrate, the things that break long before anyone reaches for a CVE, and the two constraints that surprise everyone porting a container product onto VMs: your image needs an init, and its memory is decided at bake time, not at request time.
What “bring your own image” actually agrees to
Strip the tooling away and an OCI image is two things: a stack of tar files that become a filesystem, and metadata saying which program to start. That is the whole artifact. It contains no isolation and cannot express any — isolation is entirely a property of how you choose to run it.
So the interesting question is never “is this image safe”, it is “what am I running it with”. On a shared kernel that means a process in some namespaces with some capability set and some seccomp filter, sharing one Linux kernel with every other tenant on the box. The customer did not supply a container. They supplied a filesystem and a command; your host kernel supplied the container. A container is a configuration of the host kernel, so “customer-supplied container” and “customer-supplied syscalls against my kernel” are the same sentence written twice.
The escape surface therefore stops being Docker's and becomes the kernel's: every syscall your seccomp profile permits, reaching every filesystem driver, netlink interface, eBPF verifier edge case and ioctl on every device node you left in place. That surface is enormous, written in C, and maintained by people not thinking about your multi-tenancy.
The things that bite well before a CVE does
This part gets underweighted, because kernel escapes are the interesting risk and absorb the whole conversation. In practice almost everything that goes wrong in year one is boring, and all of it is cheap to catch at admission.
- The image whose ENTRYPOINT is a miner. Not hidden, not clever — often it is the entire point of the image, under a stolen base with a plausible name.
- The 40 GiB image. Usually not malice: a model checkpoint, a .git directory and the apt cache all COPY'd into the final stage. It exhausts the disk of whichever host pulls it first, and if that host also serves traffic it takes the other tenants with it.
- The image that runs as root and expects to own /. No USER, an entrypoint writing to /etc, a package install at startup. On a read-only runtime it fails instantly, with an error filed as your bug.
- No USER, no healthcheck, and a CMD that forks forever. Nothing to probe, so the only signal is the CPU graph.
- The image that pulls a dependency at runtime from a host it cannot reach — the classic “works locally” failure. Their laptop has a warm cache and their office network's egress; your runner has neither.
- The image with a registry credential in an ENV line. Layers are readable by anyone who can pull. You will find these; it is an awkward conversation worth having early.
The miner, and the unglamorous way you find it
Mining is the most common abusive workload that turns up when you let strangers execute images, and people build elaborate defences against the wrong thing. You do not need behavioural analysis, binary scanning or a threat-intel feed. A miner has one property that does not hide: it burns essentially 100% of every core it is given, continuously, for as long as you allow, while producing almost no I/O. So the detector is average vCPU burn per tenant over a window. Legitimate workloads are spiky — builds compile and exit, ETL jobs read and wait, notebooks idle between cells, training runs punctuate compute with checkpoint writes. A sustained high average with flat egress and flat disk growth is the signature, and the cheapest possible query against metrics you already collect for billing.
The blunt lever that pairs with it: drop the well-known Stratum pool ports outbound, in the host FORWARD chain where the guest's root cannot reach the rule. We ship that list by default, ahead of the egress ACCEPTs and tunable by environment variable. No normal application dials those ports, so it costs legitimate workloads nothing and kills the pool handshake for most miners. It also matters beyond your costs: sustained mining trips your cloud provider's abuse detection, which can flag the whole project rather than the one tenant.
The pipeline: inspect, reject, flatten, bake, restore
The shape that works moves all the dangerous and all the slow work off the request path, and does it somewhere disposable.
- Resolve and pin. Turn the customer's reference into an immutable digest once. Everything downstream speaks the digest.
- Inspect metadata only. The manifest and config blobs are small; reject on platform, size and layer count before a gigabyte moves.
- Pull and flatten in a throwaway builder — the first point at which you hold the customer's bytes, so hold them somewhere you are prepared to lose. The builder needs registry read credentials and nothing else: no push rights, no control-plane access, no metadata service.
- Bake a snapshot. Boot the flattened rootfs once, let it settle, snapshot memory and disk. The ~3s cold boot plus the conversion is paid here once, asynchronously — the tenant watches an “image is being prepared” state, not a request.
- Restore per run. Every subsequent job is a restore, not a pull and not a boot.
The fifth stage is what makes this affordable. An unbounded registry pull on the hot path is the worst latency property a platform can have: minutes for a large image, a cold-cache cliff whenever scheduling lands somewhere new, and variance the customer controls rather than you. On PandaStack a create is a snapshot restore — p50 179ms, p99 around 203ms, with /snapshot/load inside that in the 49–80ms range — and there is no warm pool anywhere; every create restores. The only cold boot is a template's first spawn, about 3 seconds.
#!/usr/bin/env bash
# accept-image.sh -- the gate a customer-supplied image passes BEFORE anything
# on your platform executes it.
#
# Run this in a throwaway builder: a per-build microVM, a fresh Cloud Build
# worker, an instance you destroy afterwards. Not on a host that serves
# traffic, and not anywhere with push credentials to your registry or a
# reachable metadata service. Every line below either handles
# attacker-controlled bytes or runs attacker-controlled code, and the ones that
# run code are marked.
set -euo pipefail
REF="${1:?usage: accept-image.sh <registry/repo:tag-or-digest> <tenant>}"
TENANT="${2:?tenant id}"
MAX_COMPRESSED=$(( 8 * 1024 * 1024 * 1024 )) # 8 GiB of layers. Pick a number.
MAX_LAYERS=120
WANT_PLATFORM="linux/amd64"
# ------------------------------------------------------------- 1. pin the digest
# A tag is mutable. If you resolve acme/etl:latest at admission and resolve it
# again at run time, you have built a supply-chain hole with a nice UI: whoever
# controls that registry account can swap the bytes between the scan that
# passed and the container that ran, and your audit log will say "latest" for
# both. Resolve ONCE, store the digest, and never speak the tag again.
DIGEST="$(crane digest "$REF")" # sha256:...
CANON="${REF%%:*}@${DIGEST}" # the only reference you persist
echo "pinned: $CANON"
# -------------------------------------------------------- 2. read metadata only
# crane manifest fetches the manifest, not the layers. Most bad submissions die
# here, before a single gigabyte moves. (docker buildx imagetools inspect --raw
# "$CANON" prints the same JSON if you would rather not add another binary.)
MANIFEST="$(crane manifest "$CANON")"
# A multi-arch index has no layers of its own, only child manifests. Resolve the
# ONE platform you can actually boot and fail loudly if it is absent. Do not let
# the registry choose for you: "it works on my laptop" is frequently an arm64
# image, your hosts are amd64, and the failure arrives as an exec-format error
# inside the guest hours later instead of as a 400 at submit time.
MEDIA="$(jq -r '.mediaType // ""' <<<"$MANIFEST")"
case "$MEDIA" in
*".index.v1+json"|*"manifest.list.v2+json")
SUB="$(jq -r --arg p "$WANT_PLATFORM" '
.manifests[]
| select((.platform.os + "/" + .platform.architecture) == $p)
| .digest' <<<"$MANIFEST" | head -1)"
[ -n "$SUB" ] || { echo "no $WANT_PLATFORM manifest in index" >&2; exit 2; }
CANON="${REF%%:*}@${SUB}"
MANIFEST="$(crane manifest "$CANON")"
;;
esac
# ------------------------------------------------- 3. size and layer-count gates
# These are compressed sizes from the manifest. An extracted rootfs is routinely
# 2-3x larger, so budget on the multiple, not the number you just read. A 40 GiB
# image is not a security finding, it is a disk-exhaustion finding, and the
# tenant who submitted it is usually not malicious -- they COPY'd a model
# checkpoint, a .git directory and the apt cache into the final stage.
TOTAL=$(jq '[.layers[].size] | add // 0' <<<"$MANIFEST")
LAYERS=$(jq '.layers | length' <<<"$MANIFEST")
echo "layers=$LAYERS compressed_bytes=$TOTAL"
[ "$TOTAL" -le "$MAX_COMPRESSED" ] || { echo "image too large" >&2; exit 2; }
[ "$LAYERS" -le "$MAX_LAYERS" ] || { echo "too many layers" >&2; exit 2; }
# ------------------------------------------- 4. read the config, believe none of it
# The config blob carries Entrypoint, Cmd, User, Env and WorkingDir. Record all
# of them against the job -- this is the "what did we actually run" row you will
# want during an incident -- and treat Env as tainted. Customers put credentials
# in ENV lines and are then surprised that an image layer is a world-readable
# tar on a registry.
CFG="$(crane config "$CANON")"
jq -r '
"entrypoint: \(.config.Entrypoint // [] | join(" "))",
"cmd: \(.config.Cmd // [] | join(" "))",
"user: \(.config.User // "<root>")",
"env keys: \(.config.Env // [] | map(split("=")[0]) | join(","))"
' <<<"$CFG"
# No USER means uid 0. Inside a microVM that is root in a machine you are going
# to delete, which is a tolerable kind of root. On a shared kernel it is root
# holding a syscall table you also use. Decide which world you are in before you
# decide how much to care.
[ -n "$(jq -r '.config.User // ""' <<<"$CFG")" ] || echo "WARN: image runs as root"
# -------------------------------------- 5. flatten the layers into one ext4 disk
# Firecracker boots a block device, not an OCI image. There is no overlay stack
# in the guest, so the layers have to become one filesystem.
#
# docker create + docker export is the useful trick: `create` materialises the
# merged filesystem WITHOUT running the entrypoint, and `export` streams it as a
# tar. You are handling the bytes; you are not executing them. (`/bin/true` as
# the command is cosmetic -- nothing in the image runs.)
IMG="/var/tmp/rootfs-${TENANT}.ext4"
SIZE_MB=$(( (TOTAL / 1024 / 1024) * 3 + 2048 )) # 3x compressed + headroom
truncate -s "${SIZE_MB}M" "$IMG"
mkfs.ext4 -q -F -O ^has_journal -m 0 "$IMG" # disposable: skip the journal
mkdir -p /mnt/rootfs && mount -o loop "$IMG" /mnt/rootfs
CID="$(docker create --platform "$WANT_PLATFORM" "$CANON" /bin/true)"
# --numeric-owner keeps uids numeric, so a /etc/passwd difference between the
# builder and the guest cannot silently reassign ownership of the whole rootfs.
docker export "$CID" | tar -x --numeric-owner -C /mnt/rootfs
docker rm -f "$CID" >/dev/null
# ------------------------------------------------------------ 6. the guest contract
# This is the step people do not see coming. A Firecracker kernel booted with
# init=/sbin/init needs /sbin/init to exist. Almost no application image has an
# init -- node:22-slim does not, python:3.12-slim does not -- so the kernel
# falls through to /bin/sh and you get a VM that boots perfectly into nothing.
# You also need some channel into the guest to run the entrypoint and copy
# results out; ours is sshd on :22, which doubles as the readiness probe.
#
# Installing that needs a working chroot with DNS, and policy-rc.d returning 101
# so dpkg does not try to START systemd in a chroot that has no init to start it
# with. This is also where the honest constraint lives: apt. A musl or
# distroless image has nothing for this step to use.
cp /etc/resolv.conf /mnt/rootfs/etc/resolv.conf
for d in dev proc sys; do mount --bind "/$d" "/mnt/rootfs/$d"; done
chroot /mnt/rootfs /bin/sh -c '
set -e
export DEBIAN_FRONTEND=noninteractive
printf "#!/bin/sh\nexit 101\n" > /usr/sbin/policy-rc.d && chmod 0755 /usr/sbin/policy-rc.d
[ -s /etc/machine-id ] || echo deadbeefdeadbeefdeadbeefdeadbeef > /etc/machine-id
apt-get update -qq
apt-get install -y --no-install-recommends systemd systemd-sysv openssh-server sudo
apt-get clean && rm -rf /var/lib/apt/lists/* /usr/sbin/policy-rc.d'
for d in sys proc dev; do umount -lf "/mnt/rootfs/$d"; done
umount /mnt/rootfs
echo "rootfs ready: $IMG (${SIZE_MB} MiB)"
# ------------------------------------------------------------------ on PandaStack
# All of the above is one command, and it runs on a builder that is not a host
# serving traffic. The wrapper Dockerfile is one line, because the image bytes
# come from the registry rather than from your 50 MiB build context:
#
# printf 'FROM %s\n' "$CANON" > Dockerfile.tenant
# pandastack template build -f Dockerfile.tenant -n "tenant-${TENANT}" \
# --size-mb "$SIZE_MB" --memory-mb 4096 --replace
#
# The bake installs the init+sshd contract itself if the image lacks it, probes
# /etc/os-release on the BUILT image to confirm it is Debian-derived, builds
# --platform linux/amd64, associates the 5.10 guest kernel, and snapshots the
# result. --memory-mb is the number that matters and the one you cannot change
# later without re-baking. The default kernel association is vmlinux-5.10.
The journal-less ext4 is deliberate: the filesystem is disposable, and a journal buys nothing when recovery means “delete it and restore the snapshot again”. And the 3x size multiplier is a floor, not a rule — running out of space mid-extraction produces a rootfs that looks fine and is subtly truncated.
What you must pin: digest, platform, entrypoint
Digest, never tag. A tag is a mutable pointer in somebody else's database. If your pipeline scans acme/job:latest and your runner resolves acme/job:latest, those are two fetches of a name, and nothing guarantees the same bytes came back — a supply-chain hole with a UI on it, and it only takes a CI job of theirs that pushes to the same tag. Resolve to repo@sha256:... once, persist it, and make the digest what your audit log records. Asked what ran last Tuesday, you should be able to answer with a digest.
Platform and architecture, explicitly. A multi-arch index lets the registry pick based on who is asking, which is convenient for laptops and wrong for a platform. Our builds are --platform linux/amd64, so an arm64-only image should be rejected at submit time with a message a human can act on, not accepted and then failing as an exec-format error inside a guest. The subtler version is the kernel: ours is 5.10 in the guest, so an image built against something much newer will build, pass every gate, bake, and then fail at run time reaching for a syscall or io_uring opcode that is not there. Inspecting the image cannot catch that — only running it can, which argues for a smoke-run stage in the bake and for publishing your guest kernel version where integrators will read it.
The entrypoint you will actually execute. Resolve ENTRYPOINT and CMD at admission, store the resolved argv, and run that rather than re-deriving it from an image you fetch again. It makes the job row self-describing and gives you one place to apply policy. There is also an ugly practical reason: a probe that runs the image's default command hangs forever on an image whose CMD is a server. Override the entrypoint for anything you run as a check — our os-release probe overrides it to cat for exactly this reason.
The bit nobody expects: the image needs an init
This constraint surprises every team porting a container feature onto microVMs, and it has nothing to do with security. A container runtime starts your process as PID 1 in a fresh namespace set; the init problem is somebody else's. A VM kernel boots and executes init, so if init= points at a path that does not exist you get a kernel that came up perfectly, found nothing to run, and sits there. Almost no application image ships an init: node:22-slim does not, python:3.12-slim does not, and a distroless image has no shell to notice with.
You also need a channel into the guest. On a shared kernel the control plane can look at the process; across a VM boundary it cannot, so something inside has to accept commands and move files. Ours is sshd on port 22, which doubles as the readiness probe — the create path polls TCP :22, and that is how it knows the restore landed.
So the platform supplies both. PandaStack's bake checks the extracted rootfs for /sbin/init and /usr/sbin/sshd and, if they are missing, chroots in and installs systemd-sysv, openssh-server and sudo before baking. That is what lets a tenant hand over a plain application image rather than a specially prepared one. It is idempotent: our first-party templates install them in their Dockerfile, so the step does nothing.
Resource reality: RAM is a bake-time decision
The other surprise, and the one that reliably generates an awkward conversation with whoever wrote the pricing page. Firecracker cannot change vCPU count or guest RAM when it restores a snapshot — the memory size is part of the snapshot's identity, and there is no resize-on-restore. So “give this tenant 64 GiB” is not a create-time parameter. It is a template decision, made with --memory-mb at bake time, and changing it means baking another template.
Concretely, a memory_mb passed on a create is accepted and then silently corrected to the baked value. Silent correction is the wrong default, and we document it loudly precisely because it is: a request that asks for 16 GiB and gets 4 GiB without complaining produces a bug report about the customer's code being slow, three weeks later, from a different team. If you are building this yourself, reject the mismatch instead.
The product consequence is a small set of memory tiers rather than a slider, because each tier is a baked artifact somebody maintains. Less flexible than a cgroup limit; in exchange the ceiling cannot be exceeded and cannot be mis-enforced. A tenant process allocating without bound is OOM-killed by its own guest kernel inside its own RAM, and no other tenant notices — where on a shared host the same storm is everyone's latency problem. CPU is looser: every template runs 8 burstable vCPUs shared by cgroup weight, so cpu= on a create is deprecated and ignored.
Registries, egress, and the snapshot that is a memory dump
Private registries are where credential handling goes wrong, because the obvious implementation is the broken one. It passes the customer's registry credential into the environment that pulls the image, and that environment eventually becomes — or sits next to — the environment that runs it. Then the customer's own entrypoint can read their own credential, which sounds harmless until you remember the entrypoint came from whoever controls the registry account, who is not necessarily the person who gave you the credential. So keep the pull and the run separated: the builder authenticates, the guest never does.
Which brings me to the rule that is easiest to break by accident: a snapshot is a memory dump. Whatever is in RAM at bake time is in the artifact, and the artifact is stored, replicated and potentially streamed to other hosts. A registry token in an environment variable, a key a warm-up step decrypted, a session cookie in a browser process — all of it is bytes in vm.mem. So bake the generic template and inject per-run credentials into the restored sandbox, where they live only in that VM's copy-on-write memory and die with it. On our streaming path that memory image is Range-GET from object storage by a userfaultfd handler at restore, which makes the point viscerally: it is a file, and files go places.
Egress is the last piece. VM isolation means a tenant's code cannot reach your kernel or another tenant's data; it does nothing about that code opening a socket. An image you accepted, holding data you mounted, with unrestricted outbound, is a data-transfer service with your name on the invoice.
Be precise about what a platform gives you versus what you add. PandaStack enforces three targeted DROPs in the host FORWARD chain, ahead of the egress ACCEPTs and out of reach of guest root: sandbox-to-sandbox, so one guest can never reach another even though the pool is one routed /16 with every guest's ports DNAT'd; the whole 169.254.0.0/16 link-local range, because on a cloud host that address hands out the host's service-account token; and the Stratum ports above. What it does not give you is default-deny — outbound internet is open, and there is no per-sandbox allowlist knob.
Four ways to run a stranger's image
The comparison that matters for this job. Every cell is qualitative except our own measured numbers, because the others move with version and configuration.
| Approach | Escape surface | Cold start | Enforceable RAM ceiling | Arbitrary OCI image unchanged? | Operational cost |
|---|---|---|---|---|---|
| Shared-kernel runc | The whole host kernel syscall surface, filtered by your seccomp profile and capability set | Process-fast once the image is local — but the pull is on the hot path, and unbounded | cgroup limit: real, but enforced by code you configure and can misconfigure | Yes — native format, no conversion | Lowest. Everyone already knows how to run it |
| gVisor (runsc) | A user-space kernel reimplementation; host syscalls narrowed to what the sentry makes. Smaller than runc, not a hardware boundary | Slower than runc, by a version- and platform-dependent margin — check upstream | cgroup, as runc | Mostly — the gap is syscalls and /proc details, so some workloads fail unpredictably | Moderate. A compatibility failure needs people who understand the sentry |
| Kata Containers | Hardware virtualisation: a separate guest kernel per pod, VMM as the boundary | A VM boot, unless you add snapshot or template support | Guest RAM — hard ceiling; resizing is a VM-level operation | Yes, via an OCI shim — designed as a drop-in runtime | Moderate to high. You operate a hypervisor and a Kubernetes integration |
| Per-run microVM from a baked snapshot | Hardware virtualisation: separate 5.10 guest kernel, Firecracker's device model as the boundary | p50 179ms, p99 ~203ms restore; the ~3s cold boot and the conversion are paid once, async | Baked guest RAM — cannot be mis-enforced, and cannot be changed at restore | Nearly — must be flattened to a rootfs and needs an init plus an exec channel, which the bake installs. Ours also wants a Debian-derived base | Highest at build time, lowest at run time. You own a pipeline; you stop owning cross-tenant escape risk |
The honest summary is that the per-run microVM approach front-loads all of its cost. You build and operate a conversion pipeline, accept a base-image constraint, give up create-time memory sizing, and explain all three to integrators. In return the run path is a fixed 179ms restore with a kernel boundary between tenants, and “what happens if this image is hostile” has a short answer instead of a threat model.
Running one job
With a template baked from the pinned digest, the per-job path is short. The interesting parts are in the comments, because they are where a reasonable assumption turns out to be wrong.
# run_one_job.py -- one customer image, one job, one microVM, then nothing.
import os
import uuid
from pandastack import Sandbox # PANDASTACK_API_KEY in the environment
TENANT = "acme"
TEMPLATE = f"tenant-{TENANT}" # baked from the pinned digest, above
DIGEST = os.environ["IMAGE_DIGEST"] # carried through so the job row is honest
JOB = str(uuid.uuid4())
# What the customer asked for in the ticket: "we need 64 GiB and 16 cores".
# What you can hand them at create time: neither. Firecracker cannot change
# vCPU count or guest RAM when it restores a snapshot, so both are properties of
# the TEMPLATE, fixed by --memory-mb at bake time. Passing memory_mb here is
# accepted and then silently corrected to the baked value -- which is worse than
# an error, because nothing tells you. If a tenant genuinely needs 64 GiB, that
# is a re-bake into a second template and a line on your pricing page, not a
# field in a request body. cpu= exists and is deprecated: every template gets 8
# burstable vCPUs, shared by cgroup weight under contention.
sbx = Sandbox.create(
template=TEMPLATE,
ttl_seconds=1800, # platform-side reaper, so a dead orchestrator cannot
# leak a VM. A backstop, not your timeout.
metadata={
"tenant": TENANT,
"job": JOB,
# Pin the digest into the job record. "which image ran" should never be
# answerable only by re-resolving a tag that has since moved.
"image_digest": DIGEST,
},
# Never put a registry token, a customer secret or anything else sensitive in
# metadata. It is labelling, not a vault.
)
# The entrypoint, under a guest-side hard limit. Two things to be explicit about:
#
# 1. timeout_seconds on exec is a CLIENT deadline. The server does not enforce
# it. If your client process dies, the command inside the guest keeps
# running happily until the TTL reaps the VM. The real limit is the
# `timeout` binary in the shell, which is why it is in the script and not
# only in the keyword argument.
# 2. The memory ceiling enforces itself, and this is the best property of the
# whole arrangement. A customer process that allocates forever is OOM-killed
# by the guest's own kernel inside its own baked 4 GiB, and your host never
# learns about it. "Enforceable RAM ceiling" is not a policy you wrote and
# have to keep correct; it is the size of the DIMM you gave the VM.
script = r"""
set -eu
mkdir -p /work && cd /work
# ulimit here is NOT about memory -- the VM is the memory limit. It is about
# making the two other classic shapes legible instead of mysterious:
ulimit -u 512 # fork bombs hit a wall instead of chewing the guest
ulimit -f 8388608 # 8 GiB max single file: no filling the rootfs with zeros
ulimit -n 4096
# --kill-after is the part people forget. A process that ignores SIGTERM is not
# a hypothetical; for customer-supplied entrypoints it is most of them.
exec timeout --kill-after=30s --signal=TERM 900 \
/usr/local/bin/etl --input /work/in.csv --output /work/out.json
"""
try:
sbx.exec("mkdir -p /work", timeout_seconds=30, check=True)
sbx.filesystem.upload("in.csv", "/work/in.csv")
# Stream rather than wait: for a job whose runtime is the customer's problem,
# a live log is the difference between a support ticket and a self-service
# answer. exec_stream returns the exit code.
code = sbx.exec_stream(
script,
on_stdout=lambda s: print(s, end=""),
on_stderr=lambda s: print(s, end=""),
timeout_seconds=960,
)
if code == 0:
# filesystem.read returns bytes and handles ONE file -- there is no
# directory copy. For a tree of outputs, tar it in the guest first.
with open(f"out-{JOB}.json", "wb") as fh:
fh.write(sbx.filesystem.read("/work/out.json"))
elif code == 124:
print("entrypoint hit the 900s wall") # timeout(1)'s exit code
else:
print(f"entrypoint exited {code}")
finally:
# Delete it on the success path too. A customer-supplied image you left
# running is a bill, a live preview URL and somebody else's shell.
sbx.kill()
One thing is deliberately absent: a preview URL. Ours are tokenless — a port is reachable at https://<port>-<sandbox-id>.<suffix> for the sandbox's lifetime, and the UUID is the credential. For a customer-supplied image that is a sharper edge than usual, because the software behind the URL is something you neither wrote nor audited. If you expose it, keep TTLs short and treat the UUID as a bearer secret in your own logs.
The ticket titled “my image works locally”
You will get this ticket. It is always the same ticket, and it is never lying: the image genuinely does work locally. Their laptop has a warm package cache, a Docker daemon running everything as root with a capability set nobody trimmed, an arm64 CPU, a kernel from this year, and sixty-four gigabytes of RAM no policy is metering. Your platform has none of those things, and each absence is a choice you made on purpose.
So the fix is not technical. It is to make each of those differences produce a specific, early, actionable error instead of a late mysterious one. Wrong architecture: a 400 at submit time naming the architectures you accept. Unsupported base: a rejection naming the base and the constraint. Blocked outbound host: a log line naming the host and how to request it, not a timeout the customer reads as your platform being slow. Oversized image: the limit and the measured size. That is most of the real work in a BYO-image feature, and the alternative is that your support queue becomes the error-handling layer for other people's Dockerfiles.
And keep the thing that made the feature tractable in view: the customer's image is not running next to your control plane, next to another tenant, or on a kernel you share with either. It is in a VM restored in 179 milliseconds from a snapshot you baked, with its own kernel, its own RAM ceiling, and a delete at the end of the job. The image can be as hostile as it likes. It is still going to be deleted in a minute.
Frequently asked questions
Can I run a customer's container image unchanged in a microVM?
Almost, but not literally. Two conversions are unavoidable. Format: Firecracker boots a block device, not an OCI image, so the layers must be flattened into one filesystem. docker create followed by docker export does that without ever running the entrypoint, which is also what you want for safety. And the guest contract: a VM kernel executes init, so if init= points at a path that does not exist you get a machine that booted correctly into nothing. Almost no application image ships an init, and you also need a channel into the guest to start the entrypoint and copy results out, because the control plane can no longer just look at the process. PandaStack's bake handles both — it checks the extracted rootfs for /sbin/init and /usr/sbin/sshd and chroots in to install systemd-sysv, openssh-server and sudo if they are missing. Using apt for that imposes one constraint: the base must be Debian-derived. Alpine, distroless, Wolfi, busybox and the RPM family are rejected, by a substring check on the final stage's FROM and authoritatively by reading /etc/os-release inside the built image. Only the final stage counts, so building in an Alpine stage and shipping a Debian one is fine.
Why can't the customer choose RAM at run time if they are paying for it?
Because Firecracker cannot change vCPU count or guest RAM when it restores a snapshot — the memory size is part of the snapshot's identity, and there is no resize-on-restore. Memory is therefore a property of the template, set with --memory-mb at bake time, and changing it is a re-bake rather than a field in a request. On PandaStack a memory_mb passed on a create is silently corrected to the baked value. That is the wrong default, and we document it loudly for that reason: a request asking for 16 GiB that quietly gets 4 GiB does not fail, it produces a performance bug report from another team three weeks later. If you are building this yourself, reject the mismatch instead. What follows is a small set of memory tiers rather than a slider, because each tier is a baked artifact somebody maintains. In exchange the ceiling cannot be mis-enforced: a tenant allocating without bound is OOM-killed by its own guest kernel inside its own RAM, where on a shared host the same storm is everyone's latency problem. CPU is looser — every template gets 8 burstable vCPUs shared by cgroup weight, so cpu= is deprecated and ignored.
How do I detect a crypto miner in a customer-supplied image?
Not by looking at the image. Scanning binaries, matching hashes and buying a threat feed are all more work than the signal justifies, and a recompile defeats all three. Mining has one property that does not hide: it burns essentially all of every core it is given, continuously, for as long as you allow, while producing almost no disk or network traffic. So the detector is average vCPU burn per tenant over a window, computed from metrics you already collect for billing. Legitimate workloads are spiky in a way mining is not — builds compile and exit, ETL jobs read and wait, notebooks idle between cells, training runs punctuate compute with checkpoints. A sustained near-100% average with flat egress and flat disk growth is the signature, and it is a query rather than a product. Pair it with a blunter lever: drop the well-known Stratum pool ports outbound in the host FORWARD chain, out of reach of guest root. We ship that list by default — though note it is a targeted port block, not default-deny egress. Two operational notes: measure average burn rather than counting executions, since one long-running sandbox is the whole attack; and make suspending an account and reaping its workloads one action, because a ban does nothing about VMs already burning cores.
When is baking a snapshot per customer image NOT worth it?
Two cases, both worth checking before you build the pipeline. Genuinely one-shot images, where you pay the conversion and the bake and never restore again. And tenants who push a new digest for every run — if your median job uses a brand-new digest, a per-run pull with a well-placed layer cache may be the better trade, and you should know which world you live in rather than assuming. Everywhere else, baking wins on variance rather than on the mean. A registry pull on the hot path is unbounded work whose size the customer chooses, with a cold-cache cliff every time scheduling lands somewhere without the layers; you cannot make a latency promise on top of that. Baking moves it to a one-off asynchronous step and turns every subsequent start into a snapshot restore — p50 179ms, p99 around 203ms, with /snapshot/load itself in the 49–80ms range, and no warm pool behind the number. The first spawn of a template is a cold boot of about 3 seconds, which belongs in an “image is being prepared” state the tenant can watch.
What stops a customer's image from exfiltrating data, given the VM boundary?
Nothing in the VM boundary does, and conflating the two is the most common mistake here. Hardware virtualisation means the tenant's code cannot reach your host kernel or another tenant's memory; it says nothing about that code opening a socket and sending whatever it can read somewhere of its choosing. An image holding data you mounted, with unrestricted outbound, is a data-transfer service with your name on the invoice. Be clear about what your substrate enforces. PandaStack drops three things in the host FORWARD chain ahead of the egress ACCEPTs — sandbox-to-sandbox traffic, the 169.254.0.0/16 link-local range that would otherwise hand a tenant the host's cloud service-account token, and the well-known Stratum ports. Beyond those, outbound internet is open: there is no per-sandbox allowlist knob, so default-deny is a rule you add yourself in host iptables or security groups, outside the guest where the customer's root cannot edit it. Two adjacent rules matter as much. Registry credentials belong to the builder, never the guest. And nothing secret may be in RAM at bake time, because a snapshot is a memory dump — a token in an environment variable is bytes in vm.mem, and that file gets stored, replicated and streamed.
Keep reading
- Building images from untrusted Dockerfiles — The other half of this problem: when the customer hands you a Dockerfile instead of an image, the build itself is the untrusted workload.
- Why Docker is not a sandbox — The argument this post leans on: a container is a configuration of the host kernel, so the escape surface is the kernel's.
- Firecracker vs runc — The boundary comparison in detail, without a customer-supplied image on top of it.
- Snapshot restore vs container image pull — Why moving the pull off the hot path changes the latency story rather than just improving it.
- Controlling egress for untrusted code — The default-deny half of the design — the part the VM boundary does nothing for, and that you add yourself.
- BuildKit vs Kaniko vs microVM builds — If the throwaway builder in this pipeline is where your design questions are, start here.
Related posts
- firecracker-containerd: Running OCI Images in MicroVMs, and What It Costs You
It lets you keep your images, your registry and your containerd tooling, and swap the thing at the bottom for a microVM. That is a genuinely small diff for a large change in blast radius — and it is still the wrong tool if what you wanted was a fast sandbox API.
- The Best Multi-Tenant Isolation Platforms in 2026
You are not shopping for a sandbox. You are shopping for an isolation boundary between one paying customer and the next — and the right one depends entirely on what your tenants are allowed to do.
- Running Jobs With Customer-Supplied Cloud Credentials
Every SaaS eventually asks customers to hand over cloud credentials. Those credentials then sit in a worker process shared with everyone else's job. One dependency with a postinstall script and you've breached all of them at once.
- A Model File Is a Program: Per-Tenant Serving Isolation
You let customers upload their own models and you serve inference for them. Congratulations: you built a remote code execution endpoint with a friendly onboarding flow. Here's the real threat model and the isolation shape that survives it.
- Firecracker vs Bottlerocket: One Is a Hypervisor, One Is a Host OS
These two aren't rivals; they're stacked. Bottlerocket hardens the host you share. Firecracker removes the sharing. If you're running untrusted code, that distinction is the whole ballgame.
More in Firecracker & microVMs · See Firecracker microVM sandboxes
49ms p50 cold start. Fork, snapshot, and scale to zero.