all posts

Every In-Process Python Sandbox Leaks, and the Reason Is Structural

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and there is a conversation I now recognise from its first sentence. Someone needs to run Python they did not write — a customer's automation script, a notebook cell, a formula in a product that grew a scripting feature, increasingly the output of a language model. They have built something, they are pleased with it, and the thing they have built is a dictionary.

Not a boundary. A dictionary passed as the globals argument to `exec`, `__builtins__` set to an empty mapping, maybe an AST walk in front rejecting `import` and anything with a dunder in the name. A reasonable first move, and structurally rather than incidentally not a security boundary — for reasons more interesting than a list of bugs.

This post is about the in-process layer specifically: the thing people build before they discover that nsjail, bubblewrap, gVisor and microVMs exist. The full ladder with trade-offs is How to Sandbox Untrusted & AI-Generated Code. What I want to do here is explain why the in-process version cannot be completed, be fair to the projects that do this work honestly, and say where the boundary has to go instead.

The short version: CPython has no security boundary inside a process, because its object graph is fully connected and introspection is a design feature rather than an oversight. Removing a name from a namespace removes a name, not reachability. An AST allowlist is a static answer to a language that redefines attribute access at run time. And none of it survives `import ctypes`, a documented foreign function interface — you cannot allowlist your way out of a feature whose purpose is calling arbitrary native code. The boundary has to be a process at minimum, a kernel-enforced filter to be interesting, and a hypervisor to be the thing you bet the company on.

The four stages, in the order everyone does them

The progression is consistent enough that I could set it to music. Nobody arrives at stage four first, and everybody believes stage four is different.

  1. A string blocklist: reject the submission if it contains `import os`, `subprocess`, `open(`, `eval`. This survives reality for about a day, because `getattr(x, "__imp" + "ort__")` is not on your list and neither are the next hundred spellings. Everyone abandons it quickly, for the right reasons.
  2. Strip the builtins: `exec(code, {"__builtins__": {}})`. No `open`, no `__import__`, no `eval`. This feels categorically better than stage one because it is structural rather than a pattern match — you removed the capability rather than guessing at its spelling. That feeling is the error this post is about.
  3. Walk the AST first. Reject `Import`, `Attribute` nodes whose `attr` starts with an underscore, calls to names not on your allowlist. Now you have an allowlist rather than a denylist, which is what security people always tell you to build, so you are presumably done.
  4. Reach for RestrictedPython, where someone has already done stage three properly, with guard hooks and two decades of maintenance. You install it, feel the relief of standing on real work, and its own documentation tells you plainly that it is not a defence against a determined attacker. People install it anyway and point it at a hostile internet.

Reason one: the object graph has no partitions

Python's introspection is not a leak in the design. It is the design. A type is an object, a function carries its defining module's globals, a frame is a first-class value, every object can tell you its type. That is what makes the debugger, the ORM, the serialiser and the profiler possible, and the language would be worse without it.

It also means the graph is connected. From any object you reach its type. From a type you reach `__base__` and climb to `object`. From `object` you call `__subclasses__()`, which returns every subclass the interpreter has loaded — in any process with dependencies, a long list of things that know how to open files, load modules and call into C. So setting `__builtins__` to an empty dict changes which names resolve in that namespace. It does not change what is reachable from the values the code can construct, and the code can construct a tuple by typing two parentheses.

# The naive sandbox. It is a dictionary, and a dictionary is not a boundary.
#
# This is a specimen. Do not ship it.

def run(code: str, fetch_row):
    env = {"__builtins__": {}, "fetch_row": fetch_row}   # "no builtins"
    exec(code, env)                                       # noqa: S102
    return env


# What the restriction actually did: it removed NAMES. It did not remove
# REACHABILITY. Both of these run with __builtins__ == {}:

#   1. The object graph has no partitions. Walk type -> object -> every loaded
#      subclass of object, using only syntax and dunder attributes. No name
#      lookup is required at all -- note `.__len__()` standing in for `len`,
#      because `len` is a builtin and `__len__` is an attribute.
#
#        ().__class__.__base__.__subclasses__().__len__()
#
#      On CPython 3.14 in a plain interpreter that returns 234. Somewhere in
#      those 234 are the import machinery's loaders, file-backed objects, and
#      whatever your dependencies registered on the way in. The famous
#      `__subclasses__` walk is not a trick; it is the graph being connected.

#   2. Anything you hand in is a bridge back out. You passed `fetch_row`
#      because a sandbox with no API is useless -- and a function object
#      carries the module globals it was defined in:
#
#        fetch_row.__globals__["__builtins__"]      # the real builtins module
#
#      I checked this one rather than assuming it: the leaked function's
#      __globals__ resolves to the real builtins, while a generator created
#      INSIDE the restricted exec inherits the restricted globals. The route
#      opens the moment you expose a callable, which is always.

I ran that walk rather than quoting it from memory. On CPython 3.14 a plain interpreter reports 234 direct subclasses of `object`, reachable from inside an `exec` with `__builtins__` set to `{}`, using nothing but punctuation and dunder attribute access — no name lookup at all. `len` is gone, so you write `.__len__()`. That substitution is the lesson in miniature: you removed a name, and the capability was never attached to the name.

The routes nobody puts on the list

The `__subclasses__` walk is famous, which makes it the one people defend against. The point is not that chain; it is that the graph has no partitions, so there is always another. A representative sample, none of it exotic:

  • Any callable you hand in. A function object carries `__globals__` — the module dict it was defined in, which contains the real `__builtins__`. I checked this one specifically: a helper exposed into a zero-builtins `exec` does resolve the genuine builtins that way, while a generator created inside the restricted `exec` inherits the restricted globals. The route opens the instant you expose an API, which is always, because a sandbox that computes nothing useful is not a feature.
  • Frames, reachable from more things than you think. A frame exposes `f_globals`, `f_builtins`, `f_locals` and `f_back`; a generator has `gi_frame`; a caught exception carries `__traceback__`, which carries `tb_frame`. If your sandbox reports errors back — and reporting errors is the point of a sandbox — you have handed over a namespace walk.
  • Bound methods and descriptors. A bound method knows `__func__`, which knows `__globals__`. A classmethod, a property's `fget`, a `functools.partial`'s `func`: all one attribute from a module dict.
  • Class definition, which invokes the metaclass machinery: `__init_subclass__`, `__set_name__` and `__mro_entries__` run submission code at points your AST visitor was not thinking about.

You can patch each of these. You will be patching them for the rest of your life, against a language whose next release adds objects you have not enumerated. That is not a sandbox; it is a subscription.

Reason two: deleting __builtins__ is a namespace trick

The gap between what `exec(code, {"__builtins__": {}})` does and what people believe it does is where the whole misunderstanding lives, so be precise. CPython inserts a reference to the builtins module into a globals dict if one is absent; if you supply the key yourself, it respects your value. That is the entire mechanism. You have participated in name resolution.

You have expressed no policy about what the code may do, because CPython has nowhere to express one: no capability table, no object-level permission bit, no reference monitor between an expression and the object it evaluates to. The interpreter's job is to evaluate. The question "should this code be allowed to do that" has no representative in the data model.

Compare a kernel, where that question has an implementation: a syscall traps and something that is not your code decides. In a process there is no trap and nothing that is not your code. You are the bouncer, and the attacker is inside the building wearing your jacket.

The one in-process mechanism that is genuinely one-way is the audit hook, PEP 578: `sys.addaudithook` installs a callback that sees audit events for imports, `exec`, `open` and socket operations, and CPython deliberately provides no API to remove a hook once added. A real asymmetry and a genuinely useful veto and telemetry surface. Still not a boundary: a Python-level callback in the same process, blind to native code calling the C API or issuing syscalls directly, with no opinion whatsoever about memory corruption. Use it for tripwires. Do not file it under isolation.

Reason three: a static allowlist against a dynamic language

An AST allowlist is a better idea than a blocklist, and it fails for a reason specific to it: you are making a decision at parse time about a language that resolves almost everything at run time.

An AST allowlist is a bouncer who has memorised a list of names and is standing next to an open window.

Consider what you have to forbid, noticing that each item is a separate maintenance commitment rather than one rule:

  • Attribute access, because `getattr` is a name you can block and `__getattribute__` is a protocol you cannot. A class the submission defines can override attribute access for its own instances, so your static reasoning about `x.y` is now reasoning about a method the submission wrote.
  • Dunder names — the standard move, which breaks a surprising amount of legitimate code while being incomplete, since dunders are not the only interesting attributes and string construction defeats name-based rules wherever any dynamic lookup survives your filter.
  • Decorators, which apply an arbitrary callable at definition time in a position people forget to visit; and the class-creation hooks `__init_subclass__`, `__set_name__`, `__mro_entries__` and `__prepare__`.
  • Comprehension scopes, whose semantics moved in 3.12 when comprehension inlining landed; and walrus assignment, which binds in positions a visitor written against "assignments appear in statements" does not check.
  • f-strings, which PEP 701 formalised in 3.12 so the expression part is parsed by the full grammar, with nested same-quote characters and arbitrary expressions. A filter that treated f-strings as mostly-string is now looking at mostly-code.
  • `match` statements, with capture patterns, class patterns that call `__match_args__`, and guards; plus the 3.12 type-parameter syntax and its lazily evaluated annotation scopes — frames that evaluate later, exactly the shape a static pass is bad at.

Every CPython release is a new surface. That is not a criticism of CPython: a living language adds syntax, and "we cannot add `match` statements because it would invalidate third-party AST sandboxes" is not a constraint anyone should accept. It does mean an AST allowlist is a denylist wearing an allowlist's clothes — you enumerate the node types you understand, and the language enumerates faster.

In defence of RestrictedPython

RestrictedPython deserves better than being the punchline. It is a real, carefully maintained project with a coherent design: it compiles the submission with a restricted transformer and routes attribute access, subscription and writes through guard hooks — `_getattr_`, `_getitem_`, `_write_` and friends — that the embedding application supplies. That is a genuine policy mechanism, considerably more thoughtful than anything you will write on a Thursday.

Its use case is real and it is the one it was built for: through-the-web scripting in Zope and Plone, where the script's author is a site administrator or a content editor — somebody with an account, a contract and a manager. In that setting a restriction layer is the right tool, and the project says plainly that it is not a defence against a determined attacker. The failure is not RestrictedPython's. It is people reaching for it against the open internet, or against a language model — adversarial by construction rather than by intent, and trained on every sandbox-escape write-up ever published, including the ones about the design you just invented.

And then somebody types import ctypes

Everything above concerns pure-Python reachability, and you could imagine, with enough years, a perfect answer to it. Suppose you had one. It would be irrelevant, because the environment has `ctypes` in it.

# ctypes making the point. Three lines, no cleverness, nothing undocumented.

import ctypes

libc = ctypes.CDLL(None)      # dlopen(NULL): the symbols already in THIS process
libc.syscall                  # <_FuncPtr>. The raw syscall gate, callable.

# That is the whole demonstration. `ctypes` is a documented foreign function
# interface whose stated purpose is calling arbitrary native code. You cannot
# allowlist your way out of a feature that exists to do the thing you are
# trying to prevent.
#
# From here the Python-level restriction is simply not in the control path:
#   - libc.syscall(n, ...) issues syscalls your AST checker never sees.
#   - libc.mprotect plus ctypes.memmove make a page writable and overwrite
#     bytes in the interpreter that is enforcing your policy.
#   - CDLL("libanything.so") loads code that was never parsed by Python.
#
# And `cffi` is the same capability with a nicer API. So is any wheel with a
# .so in it: a native extension is C running in your address space, and your
# AST allowlist inspected the Python file that imported it.

# Verified on the machine I wrote this on: CDLL(None) resolves `syscall` and
# `getpid`, and `libc.getpid()` returns the real pid. Nothing exotic.

`ctypes` is in the standard library, it is documented, and its stated purpose is calling functions in shared libraries. `CDLL(None)` is a `dlopen` of the process image, handing you the symbols already loaded — libc included, raw syscall gate included. From there your Python-level restriction is not weakened; it is simply not in the control path. Syscalls do not consult your AST visitor. `mprotect` and `memmove` do not consult your guard hooks. You can make the page holding your policy enforcement writable and change it.

The standard response is "so I will remove `ctypes`," which is where the argument gets genuinely hard rather than merely losing.

The useful sandbox is the porous one

The reason to offer a Python sandbox is usually the libraries. Someone wants analysis, so they want `numpy`. Dataframes, so `pandas`. A model, so `torch`. Those are not Python libraries in the sense your AST visitor cares about. They are native extensions — compiled C, C++, Fortran and CUDA running in your address space, which your parser inspected exactly as much as it inspected the kernel. Three consequences, in increasing order of how much they ruin your week:

  1. The FFI is often reachable through the libraries themselves. `numpy.ctypeslib` is a documented, public part of NumPy's API whose entire purpose is interoperating with `ctypes`, and `ndarray.ctypes` is a documented attribute exposing the buffer's address. You did not import the FFI; the library you wanted imported it for you.
  2. Native code is outside your checker by construction. A C extension with a bounds bug is memory corruption in the process enforcing your policy, with no Python-level escape required — the submission merely passed a hostile array shape to a parser written in C in 2009. The relevant CVE history is not Python's; it is libjpeg's and every image codec you have ever shipped.
  3. Several of these libraries are compilers. Numba emits machine code through LLVM at run time, so a sandbox that allows Numba has allowed a native code generator past the only layer you were inspecting.

I can be concrete because it is our own product surface. PandaStack's `code-interpreter` template exists to give people exactly this environment, and it ships NumPy 2.3.5, pandas, SciPy, scikit-learn, Numba, OpenCV, Pillow and a Playwright install on Python 3.13 — dozens of shared objects. Any claim I made about Python-level containment inside that guest would be theatre. The containment is the guest.

So the trap closes. A restriction tight enough to be defensible excludes the entire native ecosystem that made anyone want the feature; a sandbox that includes the ecosystem has an FFI in it. There is no interior point on that curve. The sandbox that is useful is the sandbox that is porous.

The quiet one: nothing above is needed to take you down

All of that is about escape. The more common incident needs no escape at all, which is why it gets less attention: no exploit, no CVE, no write-up, no name to put in the postmortem.

  • `while True: pass` — two keywords, no escape, one pinned core.
  • A list literal multiplied by a trillion: a memory bug with no exploit attached. The same outage arrives from a typo in an exponent, summing a range with one zero too many.
  • `open("/dev/zero").read()`, which is a plausible line in a notebook; or ten thousand ten-megabyte `bytearray` allocations in a loop that looks perfectly reasonable in review.

Your AST allowlist permits every one of these, because every one is valid, idiomatic, allowlisted Python doing nothing forbidden. There is no name to ban, no node type to reject, nothing for a transformer to transform.

The in-process mitigations are worse than they look. A signal-based timeout needs the interpreter to reach a bytecode boundary to run the handler, and a C extension holding the GIL inside a long `numpy` reduction will not reach one — your alarm fires and queues behind code that is not checking. A watchdog thread has the same problem plus no way to stop native code. `sys.settrace` can count instructions at a large constant factor, and the submission can remove your trace function if it can reach `sys` — which you were going to allow, because `sys.stdout` is how it prints.

A separate process fixes this cleanly, and is the ladder's first rung, because `RLIMIT_AS`, `RLIMIT_CPU` and `SIGKILL` are enforced by something that is not the code being limited. That is the whole argument of this post in one sentence.

And while I am naming layers that do not do what people assume: a platform-level sandbox TTL does not cover this either. Ours is an idle TTL — the reaper compares time since last activity against it — so it reclaims a guest somebody walked away from and will never touch a guest pinning a core, because a spinning loop is busy rather than idle. The fence that stops a spin is `ulimit -t` inside the guest. Most incidents come from assuming a layer has a job it does not.

What the boundary actually has to be

A separate process

A distinct address space, so corrupting memory corrupts the submission's memory. `rlimit`, enforced by the kernel. A kill you can actually perform, because `SIGKILL` is not deliverable-if-convenient. A different uid with no credentials in its environment. A genuine and large improvement over everything above, and cheap: fork, exec, `python3 -I`, read a pipe.

What you do not get is a kernel boundary. The submission still calls the host kernel directly, with the full syscall surface, as some uid. If your threat model is "a colleague's script should not take down the web server," stop here. If it is "a stranger's code should not become root," a process is a speed bump.

seccomp, namespaces, Landlock

The next rung is real and good: filter the syscalls, unshare the namespaces, restrict filesystem access with Landlock, and you have the tooling behind nsjail, bubblewrap and the hardened sandboxes inside browsers. I will not re-explain it; seccomp explained for developers: filtering syscalls to shrink the kernel attack surface covers the filtering model, and Firecracker vs nsjail: which for running untrusted code? and Firecracker vs Bubblewrap: sandboxing untrusted code cover where those tools land against a VM.

The honest statement about this rung: it is a filter in front of a kernel that is still shared. The syscall surface is hundreds of entry points, each with its own argument space, and a container is a polite suggestion to the kernel about who things belong to. A well-written seccomp policy removes most of that surface and is worth doing. It does not change who the kernel is.

A virtual machine

A VM moves the boundary to the hypervisor. The submission gets its own kernel, so a kernel-reachable bug in its code is a bug in a kernel nobody else shares, and it talks to a device model rather than your kernel's syscall table. With Firecracker that device model is a handful of virtio devices — block, net, vsock, a serial console, the metadata service — rather than four decades of emulated hardware, and the VMM runs behind a seccomp filter, optionally under a jailer.

The honest version of this claim is not "secure." KVM has had CVEs and will have more, as has every hypervisor ever shipped. The claim is that the reachable interface went from hundreds of syscalls into a kernel shared by everything you own, to a small device model in front of a kernel that exists for this one job. That is a quantitative argument, and quantitative arguments are the only honest kind here.

The in-process layer compared with the rungs below it. Architectural properties, not vendor claims.
DimensionIn-process restrictionSeparate processseccomp + namespacesMicroVM
What enforces itYour own code, at the privilege of the code it restricts.The kernel: rlimit, uid, signals.The kernel, via a syscall filter it applies to itself.The hypervisor, plus a kernel the submission does not share.
Stops `import ctypes`No. An FFI is a feature, not a name you can ban.Moot: ctypes runs, in a process you can kill.In effect: the syscalls it would make are filtered.In effect: the syscalls go to a throwaway kernel.
Stops `while True: pass`Not reliably. A C extension holding the GIL ignores your alarm.Yes. RLIMIT_CPU and SIGKILL are not negotiable.Yes, plus cgroup accounting.Yes, and it burns a guest's CPU rather than yours.
Blast radius of a native memory bugYour application process, with everything it holds.One child process.Shared kernel: an escape is a host problem.A guest kernel deleted when the job returns.
Cost per executionMicroseconds. The honest reason people stop here.Milliseconds to tens of ms.Tens of ms, plus policy work.Snapshot restore every create: p50 179 ms, p99 203 ms.

Where the pure-Python approaches are genuinely right

I would be making a weaker argument if I pretended this layer had no use. It has a real one, and the test is a single question: what happens to the author if they escape? If the answer is "a conversation with their manager," a restriction layer is good engineering. If it is "nothing, we do not know who they are," it is decoration. Three settings where it is the right tool:

  • An internal scripting hook: a rules engine, a workflow step, a report expression, a formula a colleague writes in a cell. The author has an account and a name, and what you want is for mistakes to be legible and bounded. A restricted evaluator that refuses to touch the filesystem gives exactly that.
  • A correctness guardrail rather than a security boundary. Forbidding `open` in a notebook cell meant to be a pure transformation catches the reviewer's mistake before it becomes a reproducibility bug. That is linting with teeth, valuable for reasons unrelated to attackers.
  • Narrowing the language on purpose. The strongest version is not a restricted Python at all: it is a small expression language you designed, with no attribute access, no imports, bounded iteration and a total evaluator. If your requirement is user-authored formulas, that beats a restricted CPython, because you can reason about its whole semantics in an afternoon.

Defence in depth is real too. A restricted evaluator inside a microVM is two controls, and the ordering matters: the microVM makes the restriction's inevitable failure survivable. Having only the restriction is the wrong one to have picked.

WASM is a genuinely different answer, with a genuinely different bill

One option deserves separating from the ladder, because it is not a rung on it. Running Python inside WebAssembly — Pyodide, or CPython's own wasm targets — gives you a real memory-safe boundary at the module level. Linear memory is bounded by construction, there is no raw pointer arithmetic that escapes the sandbox, and capabilities are deny-by-default: a wasm module can do nothing the host did not explicitly import into it, which is precisely the property every in-process Python restriction is trying and failing to simulate.

The bill is the native ecosystem. A wheel with a `.so` in it does not run under wasm unless somebody compiled that package for wasm. Pyodide maintains a substantial set of such builds, NumPy, pandas and SciPy among them — and that set is not PyPI. Arbitrary `pip install` is off the menu and numeric performance is a real tax. So: an excellent answer when you control the dependency set, and the wrong answer when the entire point was that users bring their own packages. WebAssembly vs Firecracker for Untrusted Code has the longer comparison.

Moving the boundary, concretely

The shape I would actually ship, stated as properties rather than a diagram, because properties survive your framework choice.

  1. The submission crosses as bytes, written to a file in a guest — not as a string handed to a function in your process. The trusted side never dereferences anything the submission names.
  2. The guest holds no credential. Not a narrowly scoped one — none. No cloud role, no database password, no internal service token, no API key in the environment. The useful question is not "what can this code do" but "what is in reach," and the right answer is a file and a temp directory.
  3. The CPU ceiling and the memory ceiling are enforced inside the guest by the guest kernel, with `timeout` and `ulimit`. This is the part people get wrong, and I got it wrong in a draft of this post: a sandbox TTL is an idle TTL, so it reclaims an abandoned guest and has no opinion about a busy one.
  4. A validated value crosses back, not an object. JSON conforming to a schema you own, length-capped, parsed by you. Never a pickle: unpickling is arbitrary code execution by design, and accepting one from the sandbox undoes the entire exercise in one line on the trusted side.
  5. One submission, one VM, unconditional teardown. No reuse, because the second submission in a recycled guest inherits the first one's consequences and the point was that there are none.
"""Move the boundary. The code still runs; it just stops running in your process.

Honest framing up front: this does not make the code safe. It makes the code
someone else's kernel's problem -- specifically a kernel that exists for this
one execution and will not exist in a moment.
"""
import json
from pandastack import Sandbox


# The fence is INSIDE the guest, because the in-guest bound is the real one:
# `timeout` sends a signal the kernel delivers whether or not Python is holding
# the GIL in a C extension, and `ulimit` is enforced by the kernel rather than
# by an interpreter the code can reach. `python3 -I` is isolated mode: no
# site-packages path injection from the CWD, no PYTHONPATH, no user site.
FENCE = (
    "cd /work && umask 077 && "
    "ulimit -v 1572864; "          # ~1.5 GiB address space: [0]*10**12 dies here
    "ulimit -t 20; "               # 20s of CPU: `while True: pass` dies here
    "ulimit -u 64; ulimit -f 65536; ulimit -c 0; "
    "exec timeout --signal=TERM --kill-after=5s 20 "
    "python3 -I job.py /work/out.json"
)


def run_untrusted(code: str, job_id: str) -> dict:
    """One submission, one microVM, one audit record. No reuse."""
    sbx = Sandbox.create(
        template="code-interpreter",     # 2 GiB / 8 vCPU, baked into the snapshot
        # NOTE: ttl_seconds is an IDLE ttl. The reaper deletes a sandbox whose
        # last activity is older than this -- so it reclaims an ABANDONED guest
        # and does nothing about a busy one. `while True: pass` is not idle. The
        # thing that bounds a spin is `ulimit -t` in the fence above.
        ttl_seconds=120,
        metadata={"job": "untrusted-python", "id": job_id},
    )
    try:
        # Bytes in. The guest never dereferences anything the submission names,
        # and there is no credential in the environment to find.
        sbx.filesystem.write("/work/job.py", code.encode())

        r = sbx.exec(FENCE, timeout_seconds=30)   # keep this <= 30; see note below
        if r.exit_code != 0:
            # 124 = timeout fired. 137 = killed. 1 = it raised. All of them are
            # a deleted VM and a log line rather than an incident.
            raise JobFailed(job_id, r.exit_code, r.stderr[-2000:])

        # A VALUE crosses back, validated against a schema you own. Not a
        # pickle, not a live object, not a repr you will later eval.
        out = sbx.filesystem.read("/work/out.json")
        if len(out) > 1 << 20:
            raise JobFailed(job_id, 0, "output implausibly large")
        return validate(json.loads(out))
    finally:
        sbx.kill()    # unconditional. A leaked sandbox bills by the GiB-hour,
                      # patiently, for as long as you fail to notice it.

# Note on the 30: on a one-shot exec() the client-side HTTP timeout is the
# binding one, so do not write exec(timeout_seconds=900) and believe it. For
# genuinely long work use exec_stream(), which does honour it -- and in either
# case let the in-guest `timeout`/`ulimit` be the fence you actually rely on.

The economics are not the interesting part, but people ask: $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same rate for every class, CPU billed on CPU-seconds actually burned and memory on committed GiB-hours. No per-request price. A submission occupying a 2 GiB guest for three seconds is a small fraction of a cent. What you buy is that a submission which tries something interesting becomes a deleted VM and a log line with a job id, rather than an incident with a timeline.

The honest limits

  • A microVM does not stop resource exhaustion; it bounds who pays for it. `while True: pass` inside a guest still burns CPU-seconds, and our billing charges CPU-seconds actually burned. Worth stating plainly about our own product: a sandbox TTL is an idle TTL — our reaper compares time since last activity against it — so it reclaims an abandoned guest and will not touch a spinning one, which is not idle. The only thing that bounds a spin is the in-guest fence. The VM converts "my host is pinned" into "a guest I am about to delete is pinned, and I am paying for it."
  • Egress is open by default, with no default-deny policy. The host's root FORWARD chain does drop pool-to-pool traffic, so no sandbox reaches another sandbox's subnet, and drops all of 169.254.0.0/16, so the cloud metadata endpoint is genuinely unreachable. Your own VPC, databases and internal APIs are not fenced. If your threat model includes the submission reaching an internal service, those deny rules are yours to add, as code, for every sandbox.
  • A VM boundary has had CVEs and will have more. KVM is software. The claim is a smaller and better-audited interface, not an absolute one, and anyone selling you the absolute version is selling you something else.
  • 179 ms is not free on a hot path. Every create is a snapshot restore with no warm pool, at p50 179 ms and p99 203 ms end to end, and the first spawn of a template is a cold boot of roughly 3 seconds before the snapshot is baked. An in-process evaluator is microseconds. That gap is precisely why people build the thing this post argues against.
  • If your requirement is "untrusted code that returns an answer in 5 ms," none of this ladder is your answer. Narrow the language instead: a small total expression evaluator with no attribute access, no imports and bounded iteration, and accept that users cannot bring arbitrary Python. That is a product decision rather than a security one, and it is correct more often than people like.
  • Per-execution VMs mean capacity can run out. A create refused for capacity is a real 503 you have to handle, with a retry policy and a queue, and at 3am a capacity incident and an attack look remarkably similar on a dashboard.

The summary

CPython has no security boundary inside a process, and the reason is structural rather than a backlog. The object graph is fully connected, because introspection is a design feature that makes the whole ecosystem possible, so removing a name from a namespace removes a name and not reachability. An AST allowlist is a static decision about a language that resolves attribute access at run time and adds syntax every release — a denylist wearing an allowlist's clothes. Then the layer nobody plans for ends the argument: `ctypes` is a documented FFI, and `numpy`, `pandas` and `torch` are native extensions with their own reachable internals. The libraries that made anyone want a Python sandbox are exactly the ones that put that feature in reach.

RestrictedPython and its relatives are good tools pointed at the wrong threat. Keep them where the author has a name and a manager, where you want a correctness guardrail and legible accidents. Then put the execution somewhere disposable: a process at minimum, a syscall filter to be interesting, a separate kernel to be the thing you bet on. One submission, one guest, no credential in reach, an in-guest fence you own, a validated value as the only thing crossing back, an unconditional `kill()`.

None of that makes the code safe. It makes the code someone else's kernel — and that kernel is deleted when the job returns, which after twenty-five years of people trying to build a wall inside the interpreter is the honest target.

Frequently asked questions

Is RestrictedPython safe for running untrusted code?

No, and the project itself says so — which is the part most people skip. RestrictedPython is real, carefully maintained software with a coherent design: it compiles submissions with a restricted transformer and routes attribute access, subscription and writes through guard hooks the embedding application supplies. That is a genuine policy mechanism. But its design target is through-the-web scripting in Zope and Plone, where the script's author is a site administrator or a content editor — somebody with an account, a contract and a manager — and its documentation is explicit that it is not a defence against a determined attacker. The structural reason it cannot be one is that it runs at the same privilege, in the same address space, on the same object graph as the code it restricts, and that graph has no partitions: there is always another route from an ordinary value to a module dict, and every CPython release adds syntax a static transformer did not anticipate. The decisive problem is one level below the language anyway. Even a perfect pure-Python restriction is irrelevant if the environment can reach ctypes or cffi, because an FFI exists to call arbitrary native code and syscalls do not consult your guard hooks. Keep RestrictedPython for semi-trusted authors. For a hostile author, use a separate process at minimum, and a separate kernel if the consequences matter.

Can I just remove ctypes from the sandbox environment?

You can delete the module, and it buys less than you would hope. First, removal has to be total rather than nominal: the C extension is compiled into the interpreter on most builds, so deleting a Python wrapper is not the same as removing the capability, and the import system has more than one route to a built-in module. Second and more important, the libraries you are offering are themselves native extensions, and several re-export the FFI as documented public API. numpy.ctypeslib exists specifically to interoperate with ctypes, and ndarray.ctypes is a documented attribute exposing a buffer address. You did not import the FFI; the library you wanted did it on your behalf. Beyond re-export, native code is outside your checker by construction: a C extension with a bounds bug is memory corruption in the process enforcing your policy, and the submission only had to pass a hostile array shape. Numba goes further, generating native code at run time. This is the trap: a restriction tight enough to be defensible excludes the entire scientific ecosystem that made anyone want the feature. Build the boundary outside the interpreter instead, where the libraries can be whatever people need.

Why doesn't a signal-based timeout reliably stop runaway Python?

Because a Python-level signal handler runs at a bytecode boundary, and long-running native code does not reach one. When you install a SIGALRM handler, CPython records that a signal arrived and runs your handler the next time the eval loop checks — fine for a pure-Python loop, useless inside a single large numpy reduction, a regex backtracking in C, or an image decode, all of which can hold the GIL for a long time without returning to the interpreter. A watchdog thread is worse: it cannot interrupt native code at all, and it shares the process it is policing. sys.settrace can count instructions, but it costs a large constant factor, and any submission that can reach the sys module can remove your trace function — and you were going to let it reach sys, because sys.stdout is how it prints. The reliable mechanisms live outside the interpreter: RLIMIT_CPU, which the kernel enforces with SIGXCPU and then SIGKILL; RLIMIT_AS, which makes a large allocation fail rather than swap your host to death; the timeout utility, whose signal the kernel delivers regardless of what the GIL is doing; and cgroup limits one level up. That is why the ladder's first rung is a separate process rather than a cleverer in-process timer — not because a process is more secure in the abstract, but because the thing doing the limiting is finally not the thing being limited.

Do subinterpreters give me an isolation boundary?

They give you isolation of module state, which is real and useful and not a security boundary. A subinterpreter has its own interpreter state, its own module dicts and — since the per-interpreter GIL work landed — its own GIL, which is why they are interesting for parallelism. What they do not have is their own address space. They share the process, the C heap, the loaded shared objects and the operating system's view of the world, so a native extension that corrupts memory in one subinterpreter corrupts it for all of them, and anything reachable through the FFI is reachable from any of them. The useful mental model: subinterpreters are a better module system, not a sandbox. They answer "can these two workloads have independent copies of this library" rather than "can this workload harm that one." If you need the second question answered, it is still a process boundary at minimum, and a kernel boundary if the author is someone you have never met.

Keep reading

Related posts

  • Pickle Is a Virtual Machine You Have Been Downloading

    pickle is a stack machine with an eval loop, and you have been downloading it from the internet as model weights. Here is the disassembly, why the allowlist promises less than it looks, and the conversion gateway that fixes it once instead of forever.

  • A Regex Is Untrusted Code That Does Not Look Like Code

    Nobody reviews a regex in a form field as code. But a nested quantifier on a backtracking engine is a denial-of-service primitive with a one-line source file — and the attack payload is usually a badly-typed email address, not a pattern.

  • Running User-Generated Game Mods in Isolated microVMs

    A mod is arbitrary code from a stranger that you execute on your infrastructure. "We removed the io library" is not a security boundary — a guest kernel is.

  • Sandboxing User-Written Webhook Transformations

    Somebody added a textarea labeled 'Transform (optional)' and shipped it on a Thursday. Congratulations: you are a code-execution company now, and nobody told your threat model.

  • The Best WebAssembly Runtimes in 2026 (From a MicroVM Vendor)

    I sell microVMs, which makes me the wrong person to write this and therefore obliged to write it fairly. A buyer's guide to the WASM runtimes themselves: what the artifact is, what the guest can reach, what your language actually compiles to, and what stops a while(1).

More in Code execution · See Code interpreter sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.