all posts

The Best Sandbox APIs for PHP and Laravel AI Agents in 2026

Ajay Kumar··10 min read

Somewhere in your PHP application there is a line that takes a string a model wrote and does something with it. If you are lucky it is `shell_exec`. If you are unlucky it is `eval`, and it is executing inside the same process that holds your `.env`, your PDO handle, your session store and whatever your framework decrypted into memory at boot. PHP makes this mistake unusually easy to reach, partly because `eval` is right there in the language and partly because PHP has shipped things called sandboxes, repeatedly, for twenty years, and none of them were one.

The agent version of the problem is worse than the classic web-app version, because an agent does not want to run one snippet. A model writing PHP wants a loop: write a file, run it, read the real error output, fix it, run it again. It wants a working `composer`, because almost nothing in modern PHP is a single file. It wants a specific PHP version, because 8.1, 8.3 and 8.4 are genuinely different languages to user code. And it usually wants a database and an HTTP server, because PHP code is web code. A platform that can run `php -r` and nothing else will not get you through a single Laravel ticket.

So the question is not really "which sandbox runs PHP". They all run PHP. The question is which one gives a model a project directory, a chosen interpreter, a package manager that is allowed to execute code, a port you can look at, and a blast radius that ends somewhere you can describe to a security reviewer.

I am Ajay. I build PandaStack, an open-source Firecracker microVM platform, and it is one of the entries below — so read this as a vendor's roundup and discount the vendor accordingly. What I can do honestly is be specific about my own system and careful about everyone else's.

Ground rules. PandaStack is mine, so it is the only platform I give numbers for — latency, pricing, limits — because it is the only one where I can point at the code that produces them. Every other option here is described qualitatively: what it is built for, what substrate it sits on, one honest drawback. PHP support specifically is the fastest-moving fact in this comparison, so I will not assert what any vendor's base image carries today. Check their docs and write down the date you checked.

First, the bad news: PHP cannot sandbox PHP

PHP people reach for a PHP answer first, and PHP has an unusually rich catalogue of wrong answers, because several of them were once official.

eval() is not a boundary and was never advertised as one

`eval()` compiles a string in the current scope. That is the whole feature. It has no allowlist, no capability model, no separate heap and no way to say "this code may not touch that object". Code inside `eval()` can read `$_ENV`, open your config file, instantiate your ORM, redefine an autoloader, and call any function the process can call. Nothing about it is a confinement mechanism, and unlike some languages PHP has never pretended otherwise. If you find `eval()` wrapped in a function called `safe_eval`, the function name is the only safety feature present.

The version that actually gets shipped is worse, because it looks careful: a regular expression that rejects `system`, `exec` and `passthru` before the `eval`. PHP has variable functions, `call_user_func`, `$$name` indirection, string concatenation and a reflection API. A pattern match over source text loses to runtime name construction every time, and it loses silently, which is the part that matters.

disable_functions and open_basedir are speed bumps

These two directives are what most teams mean when they say their PHP is sandboxed. They are worth setting. They are not a boundary, and the reason is structural rather than a bug anyone can fix: `disable_functions` is a denylist of function names, maintained by hand, over a standard library of thousands of functions plus every extension you load. Multiple distinct functions reach the same kernel capability. Spawning a process, writing a file, opening a socket and loading native code each have several front doors, some of them in extensions that exist for entirely unrelated reasons.

So the list is only correct until the next extension is enabled, the next release adds a function, or someone turns FFI on because a library needed it. "The `disable_functions` bypass" is not an event; it is a genre of blog post, with new entries whenever a popular image ships a new extension. `open_basedir` has the same shape: it constrains PHP's own stream layer, not extension code that opens files through C, and it has lost to symlink and realpath-cache races often enough that treating it as containment is a choice rather than an oversight.

We have been here before, and PHP itself gave up on it

The strongest argument against in-language PHP sandboxing is that PHP tried hardest and then stopped. `safe_mode` was the official answer for years — uid checks on file access, restricted environment manipulation — and it was deprecated in 5.3 and removed in 5.4 on the grounds that trying to solve the problem at the PHP level was architecturally wrong. That was the right call and it is still the right call. The `runkit` and Suhosin era that followed produced genuinely clever work and then ran out of maintainers somewhere around the PHP 7 transition; building new isolation on top of it today means adopting an unmaintained extension as a security control.

And `php -d` flags are not in this conversation at all. They are configuration passed on a command line. Configuration that the code under test cannot change is a useful property; it is not the same property as a boundary.

What that leaves

Isolation outside the language, in exactly four flavours: a separate process with seccomp and a dropped user, a container, a WebAssembly build of PHP, or a virtual machine. Everything in the field section below is one of those four with an API in front of it. The honest framing of the whole decision is: how far down the stack does your boundary go, and who operates it.

# A php.ini that looks like a sandbox in code review and is not one. Every
# directive here is worth setting. Not one of them is a boundary.
cat > /etc/php/8.3/cli/conf.d/99-looks-hardened.ini <<'INI'
; --- The denylist ---------------------------------------------------------
; disable_functions is PHP_INI_SYSTEM, so guest code cannot ini_set() it back.
; That is the one genuinely good property this file has.
disable_functions = exec,passthru,shell_exec,system,proc_open,popen,pcntl_exec,pcntl_fork,dl,putenv,symlink,link
; Why it is a speed bump: it is a list of NAMES, enumerated by hand, over a
; standard library of thousands of functions plus every extension you load.
; Several distinct functions reach the same kernel capability, so the list is
; only correct until the next extension is enabled or the next release adds a
; function. Keeping it correct is a permanent job, and "the disable_functions
; bypass" is a genre of blog post rather than an incident.

; --- Path confinement ----------------------------------------------------
open_basedir = /srv/app:/tmp
; Why it is a speed bump: it constrains PHP's own stream layer. Extension code
; that opens files through C, and anything that hands a path to a child
; process, is a different code path. It has also been defeated often enough by
; symlink and realpath-cache races that nobody treats it as a boundary.

; --- The obvious remote shapes -------------------------------------------
allow_url_include = Off
allow_url_fopen   = Off
; Worth having. Says nothing about the sockets PHP can still open directly.

; --- Native-code escape hatches ------------------------------------------
ffi.enable = false
; FFI calls into libc. If this is on, the denylist above is decoration.
enable_dl  = Off
; Same category: loading a compiled extension at runtime ends the argument.
zend.assertions = -1
; assert() used to evaluate strings. Off, and then stop thinking about eval.

; --- Resource caps --------------------------------------------------------
memory_limit = 256M
max_execution_time = 30
; These are about your bill, not containment. max_execution_time counts time
; spent executing the SCRIPT; time inside blocking syscalls is not counted, so
; a script asleep in a socket read outlives its limit comfortably.
INI

# Prove it even loaded. Half of all php.ini hardening is in a file PHP is not
# reading, because CLI, FPM and the web SAPI each have their own ini scan path.
php --ini
php -i | grep -E '^(disable_functions|open_basedir|ffi.enable)'

# What this file does not cover, and structurally cannot:
#   * network egress -- PHP can still reach anything routable from the box
#   * composer install, which runs before your application's php.ini matters
#   * every other interpreter on the machine: python3, node, perl, sh
#   * the kernel, which is shared with everything else on the host
# A denylist inside the interpreter is configuration. A boundary is a kernel
# you do not mind losing.

The PHP-specific costs nobody budgets for

Everything above is the security argument, and it is the same argument in every language. What is genuinely PHP-shaped is the ecosystem, and this is where a generic sandbox either fits your agent loop or quietly ruins it.

composer install is arbitrary code execution, by design

This is the single most common way untrusted PHP executes on your machine, and it happens before any line of the agent's code runs. Composer supports install scripts, `post-install-cmd` and friends, and a plugin API whose plugins are PHP classes loaded and executed inside Composer's own process. None of that is a vulnerability — it is the documented feature set, and real packages depend on it. It does mean that "I will review the code before running it" is not a plan, because dependency resolution ran first.

Practical consequences for an agent loop: pass `--no-scripts --no-plugins` when you can and accept that some packages then do not work; never let a model's `composer.json` run on your laptop or your CI runner's shared kernel; and remember that `composer install` needs network access to Packagist, which is the same network access it needs to reach anywhere else. I wrote up the specifics separately at How to Sandbox an Untrusted composer install.

pecl install is a compiler run

If the task touches an extension that is not in the image, you are running a build: a toolchain, headers, make, and a few minutes of CPU. That is fine in a VM you are going to throw away and a disaster in a per-call sandbox created fresh for every tool invocation, because you will pay it on every turn. The fix is the same as in every compiled-dependency ecosystem: bake it into the template, or snapshot the sandbox after the first successful install and start from there.

artisan and bin/console boot your entire application

An agent asked to "check the user count" will reach for `php artisan tinker`, and that command boots the framework kernel: service providers, config cache, `.env`, database connections, queue credentials, mail credentials, your API keys. If that happens anywhere near your real environment then the agent has your production secrets, and whether it meant to is not a security control. Console commands in PHP frameworks are full application entry points, so they belong on the far side of the boundary with a `.env` you wrote for the occasion.

Opcache and JIT change what cold start means

PHP's execution model has a per-request compile step that opcache exists to amortise, and a freshly created sandbox has an empty opcode cache. The first request pays compilation for every file the framework touches, which on a large application is not a rounding error; preloading moves that cost to startup instead of removing it. JIT changes the shape again for CPU-bound code. The practical point for benchmarking: if your agent times one request in a brand-new sandbox and calls it the app's latency, it has measured compilation, and it will report the wrong conclusion confidently.

Long-running PHP breaks the disposable-process assumption

The comfortable old property of PHP was that a process handled one request and died, so leaked state and half-corrupted globals cleaned themselves up. FrankenPHP, RoadRunner and Swoole deliberately remove that. A worker now serves thousands of requests, which means untrusted code that modified a static property, poisoned a connection pool or left a file handle open has done so for every request that follows. If your agent's job is to work on a worker-mode application, the per-request reset you were implicitly relying on is not there, and the only clean reset left is destroying the whole machine — which is an argument for cheap machines.

The criteria a PHP team should actually use

Ordered roughly by how much time each will cost you if you get it wrong. The first four are PHP-specific; the rest apply to any agent sandbox and are still where the surprises live.

  1. Can you choose the PHP version, exactly, and is that choice an image/template property rather than something you install on every call? 8.1 to 8.4 spans enums, readonly, fibers, property hooks and a pile of deprecations. "You can install it" is not a yes.
  2. Can it run composer install at all — meaning: is there a persistent writable project directory, an outbound network path to Packagist-style registries, and enough CPU and time to build native extensions?
  3. Can it serve HTTP and can you reach that port from outside? PHP code is web code. If the only way to see the result is stdout, the agent cannot check its own work on anything with a template in it.
  4. Can you attach a real database cheaply and disposably? A Laravel or Symfony task without a database is a syntax check.
  5. Is the isolation boundary a kernel you can lose, or a kernel you share? This is the question a security reviewer will ask, and the answer determines whether model-generated code is an acceptable input class at all.
  6. Is there a filesystem API, or are you base64-ing files through a shell command? You will be moving a lot of small files.
  7. Does exec stream? A composer install behind a blocking request with no output is indistinguishable from a hang, and your agent will time out and retry it.
  8. Are timeouts enforced server-side, in the guest? Many client libraries' timeout parameter only widens the HTTP wait. Verify it, then bound hostile commands in-guest anyway.
  9. What is the egress default? Open egress is convenient and is also how a sandbox becomes someone's proxy. Mine is open by default, which I am telling you because it is your problem to fence: see Controlling Network Egress for Untrusted Code.
  10. What happens to the machine when the agent crashes mid-loop? Idle reaping, TTLs and explicit teardown are the difference between a bill and a line item.

The field, through a PHP lens

Grouped by the job each is genuinely built for, not ranked, because ranking requires pretending they compete for the same job and they do not.

PandaStack (mine — read accordingly)

Firecracker microVMs behind a small REST API, with Python and TypeScript SDKs. Every create is a snapshot restore rather than a boot, which is p50 179 ms and p99 203 ms with no warm pool behind it; the first spawn of a template that has no snapshot yet is a cold boot at around 3 seconds and then bakes one. Each sandbox gets its own Linux 5.10 guest kernel under KVM, its own network namespace and tap device, and a tokenless HTTPS preview URL per guest port — so a `php -S` or an nginx in the guest is something a human can actually open. Pricing is one rate card: $0.054 per vCPU-hour and $0.0162 per GiB-hour, with CPU billed on seconds actually burned.

The PHP story, stated plainly: there is no PHP SDK, and the `base` template does not pre-warm PHP. `base` is Ubuntu 24.04 at 4 GiB and 8 vCPU with mise managing runtimes, and mise pre-warms Node, Python, Go and Bun. PHP is an install step on top of that, which is fine for an experiment and wrong for a hot path — for production you bake a template with your PHP, your extensions and your vendor directory already in it. The interesting primitive for an agent loop is `fork_tree(n)`: get one sandbox to the warm state with dependencies installed and the database migrated, then boot up to 16 children that inherit the parent's memory, and try several candidate fixes at once. Note `fork()` is not that — it clones the disk and cold-boots the child.

Where it is not the right fit: no PHP SDK means you drive it from Python, TypeScript or REST. There is no GPU and no GPU passthrough. The guest kernel is 5.10, which is deliberate and occasionally inconvenient. Egress is open by default. And vCPU and RAM are baked into the snapshot, so a template's 4 GiB is the sandbox's 4 GiB regardless of what you asked for on the create call.

The agent-sandbox platforms: E2B, Daytona, Runloop, Modal

E2B is the most focused entry in the category and focus is a feature: sandboxes for AI agents, microVM-backed, with custom templates you define from a Dockerfile — which is the right shape for PHP, because it means your PHP version and extensions are an image concern rather than a per-call install. Its SDKs are Python and JavaScript, so a PHP shop is writing glue either way.

Daytona comes from the development-environment direction rather than the ephemeral-invocation one, which suits a long agent session working on a repo. The thing to check carefully is which runtime class you are actually getting, because container-by-default and VM-class are very different answers to the "is this a kernel I can lose" question.

Runloop is aimed specifically at coding agents — devboxes, snapshots, the lifecycle an agent loop wants — which means the primitives line up well with a PHP repo task. It is hosted, and it is opinionated about the agent workflow, so confirm its image story covers the PHP version matrix you need.

Modal's centre of gravity is serverless Python and ML compute, and its images are defined in Python. You can absolutely run PHP there, and you will be declaring your PHP toolchain inside someone else's language, on a platform whose documentation, examples and community are about something else. That is a real tax on a PHP team even when every feature is technically present.

TypeScript-native and edge-adjacent: Vercel Sandbox, Cloudflare

Vercel Sandbox is microVM-backed and TypeScript-first by design, built for ephemeral work inside the Vercel ecosystem. If your agent already lives in a Next.js app it is a short hop. If your agent is a PHP service, you are adopting an ecosystem to get an API. Cloudflare's container story is driven from a Worker, which is elegant if you are already on Workers and a whole programming model to adopt if you are not; what you get in exchange is your own container image, so any PHP build you like, with an HTTP front door that already exists.

Both are products whose main value is being inside an ecosystem. Judge them on whether you are inside it.

Machines and roll-your-own: Fly.io, Docker, gVisor

Fly.io Machines are Firecracker-based VMs you start and stop over an API with your own image, and for PHP that is genuinely attractive: full control of the build, real HTTP, no language opinions. What it is not is an agent sandbox. The exec endpoint, the filesystem API, the per-call lifecycle, the idle reaper and the quota accounting that an agent loop needs are yours to build on top. That is a few weeks of work you should price honestly before you call it the cheap option.

Plain Docker driven from PHP is the default thing teams build, and it is a real improvement over a child process: namespaces, cgroups, a seccomp profile. Be precise about what you bought, though — one kernel, shared between your host and every container on it, and a container escape is a kernel bug away rather than a hypervisor bug away. A container is a polite suggestion to the kernel. gVisor is a genuine step up you can operate today, intercepting syscalls in a user-space kernel, at a syscall-heavy performance cost that a `composer install` will find for you. Raw Firecracker is the seductive one: the VMM is small, well audited, and a proof of concept boots in an afternoon. The platform around it — networking, snapshots, scheduling, reaping, billing — is the part that takes a year, which I know because I am still doing it. Start at What is a microVM? Firecracker, isolation, and why agents need it if you want the substrate explained before you commit.

Judges and public evaluators: Judge0, Piston, 3v4l-style sites

Judge0 and Piston are legitimate answers to a narrower question. You submit source, they run it under a confined one-shot runner, you get stdout, stderr and an exit code back, across dozens of language versions, and you can self-host them. For "evaluate this standalone snippet" they are simpler and more honest than anything else on this list. They stop being the answer the moment your agent needs a project directory that persists between calls, a `composer install`, a database, or a port — which for PHP is roughly immediately.

Public paste-and-run evaluators in the 3v4l mould are a wonderful tool for checking how one expression behaves across three hundred PHP builds, and are not somewhere you send a customer's code or a model's output. Treat them as documentation, not infrastructure.

php-wasm and PHP on WebAssembly

PHP compiled to WebAssembly is a real and lovely thing — it is how WordPress runs in a browser tab — and it has the strongest story on this page for the narrow case of evaluating self-contained PHP with no ambient capabilities, because the sandbox is the Wasm runtime and the host grants capabilities explicitly rather than denying them by name. That is the opposite of the `disable_functions` model and it is the right way round.

The limits are equally real: process spawning and raw sockets are absent or emulated, the extension set is whatever was compiled in, and a lot of ordinary PHP tooling assumes a POSIX machine underneath it. So `composer install` is at best partial and often a non-starter, and anything that shells out does not work. Great for a language playground, an in-browser demo or a plugin host; not a place to run a Laravel test suite. WebAssembly vs MicroVMs for Sandboxed Code has the longer version of this tradeoff.

And the baseline you should not ship

A hardened `php.ini` and nothing else. I am listing it as an option because it is, empirically, the most widely deployed option, usually without anyone deciding on it. It is configuration that a code reviewer mistakes for a boundary, which makes it worse than nothing in one specific way: it ends the conversation.

At a glance

PHP execution options for AI agents, compared qualitatively on substrate, version control, package-manager reality and HTTP. Competitor cells describe positioning, not benchmarks.
OptionIsolation boundaryPHP version + composer installServes HTTP you can reachHonest drawback
PandaStackFirecracker microVM per sandbox, own Linux 5.10 guest kernel under KVMYou install or bake it: base is Ubuntu 24.04 with mise and does not pre-warm PHP; composer install then works normallyYes — tokenless HTTPS preview URL per guest portMine, so discount accordingly. No PHP SDK, no GPU, guest kernel 5.10, egress open by default
E2BmicroVM per sandboxCustom templates from a Dockerfile, so PHP is an image concernVerify against current docsNo PHP SDK; you drive it from Python, JS or REST
DaytonaContainers by default, with VM classes availableImage-defined, so PHP is a Dockerfile lineVerify against current docsCheck which runtime class you are actually getting; the default is a shared kernel
RunloopManaged devboxes aimed at coding agents, with snapshotsImage or blueprint definedVerify against current docsHosted only, and opinionated about the agent workflow
ModalgVisor by default, with VM-backed sandboxes availableImages are declared in Python, so your PHP toolchain lives in someone else's languageVerify against current docsThe platform, docs and community are about Python ML compute; PHP is a guest
Vercel SandboxmicroVM-backed ephemeral sandboxTypeScript-first API; verify what the base image carriesVerify against current docsBuilt for short-lived work inside the Vercel ecosystem, not for parking a Laravel app
Cloudflare Containers / Sandbox SDKContainer alongside Workers, driven from a WorkerYour own container image, so any PHP build you likeYes, through the Worker in front of itYou adopt the Workers programming model and its lifecycle to get there
Fly.io MachinesFirecracker-based VMs started and stopped over an APIYour own image, full control, no language opinionsYesInfrastructure, not an agent sandbox: exec, filesystem, reaping and quotas are yours to build
Judge0 / PistonConfined one-shot runner, self-hostableFixed set of language builds; no persistent project directory, so composer install is out of scopeNoBuilt for submitting snippets, not for agents that want a repo
php-wasm / PHP on WebAssemblyWebAssembly runtime; host grants capabilities explicitlyWhatever PHP build was compiled; limited extension setOnly what the host emulatesNo real process spawning or raw sockets, so most PHP tooling does not run
Docker or gVisor you run yourselfNamespaces plus seccomp, or a user-space kernel intercepting syscallsEntirely yoursYesYou are now the platform: quotas, reaping, multi-tenancy, API, on-call
Hardened php.ini onlyNone. A denylist inside the interpreterNot applicableNot applicableNot a boundary. Configuration a reviewer mistakes for one
Every competitor cell above is qualitative on purpose and is stale the moment it is published. Base images change, isolation backends occasionally get swapped, SDKs and language support get added, licences shift, and pricing moves faster than this post's shelf life. Use this to build a shortlist and to understand which questions matter, then pull every specific claim live from each vendor's own current documentation — especially anything about PHP versions, extensions or whether composer is in the image. Then spend one afternoon spiking your top two with your real composer.json. An hour of measurement settles more than a week of roundups, this one very much included.

What this looks like in code

The shape below is a single agent turn on PandaStack: pin a PHP, write the agent's files, install dependencies behind a fence, run the script, read both streams, tear the machine down in a `finally`. Two details in it are not decoration. The first is the mise environment, which has to be re-exported in every command because exec runs non-login shells. The second is `timeout --kill-after=10s` in the guest: the one-shot `exec()` timeout parameter widens the HTTP wait rather than enforcing a limit on the command, and `composer install` is the exact case where you need a real one, because install scripts are arbitrary code with no obligation to finish.

import json
from pandastack import Sandbox

# The base template carries mise at /opt/mise, and exec runs non-login shells,
# so every command re-exports the runtime environment itself. This is the one
# bit of boilerplate you will forget exactly once.
MISE = "export MISE_DATA_DIR=/opt/mise MISE_CONFIG_DIR=/opt/mise PATH=/opt/mise/shims:$PATH"

def fenced(cmd, seconds):
    """Bound the command IN-GUEST.

    One-shot exec() does not actually enforce timeout_seconds -- the parameter
    widens the HTTP wait and nothing more. For composer, which runs arbitrary
    install scripts, and for model-written PHP, which loops, the only timeout
    that fires is the one inside the VM.
    """
    return "timeout --kill-after=10s {0} sh -c '{1}; {2}'".format(seconds, MISE, cmd)

AGENT_PHP = """<?php
require __DIR__ . '/vendor/autoload.php';
$code = file_get_contents('php://stdin') ?: '<?php echo 1;';
$parser = (new PhpParser\\ParserFactory())->createForNewestSupportedVersion();
var_dump(count($parser->parse($code)));
"""

sbx = Sandbox.create(
    template="base",          # Ubuntu 24.04, 4 GiB / 8 vCPU, baked into the snapshot
    ttl_seconds=1800,         # IDLE clock, not a wall clock. Status GETs do not reset it.
    metadata={"job": "php-agent-turn"},
)
try:
    # 1. Pin the PHP the user's code expects. base pre-warms Node, Python, Go
    #    and Bun -- PHP is NOT pre-warmed, so this is a real install step and
    #    on some backends a source build. In production, bake a template.
    boot = sbx.exec(fenced(
        "mkdir -p /work && mise use -g php@8.3 && mise reshim && php -v", 900))
    if boot.exit_code != 0:
        raise RuntimeError(boot.stderr[-2000:])

    # 2. Write the manifest and the agent's file.
    sbx.filesystem.write("/work/composer.json", json.dumps({
        "require": {"nikic/php-parser": "^5.0"},
        "config": {"allow-plugins": False},
    }, indent=2))
    sbx.filesystem.write("/work/analyse.php", AGENT_PHP)

    # 3. composer install IS code execution: install scripts, post-install-cmd,
    #    and plugins that run as PHP inside composer's own process. The flags
    #    remove the easy paths; the microVM covers the rest. Note that egress
    #    is open by default, so composer reaches packagist -- and anywhere else.
    dep = sbx.exec(fenced(
        "cd /work && composer install --no-interaction --no-progress "
        "--no-scripts --no-plugins --prefer-dist && mise reshim", 600))
    print(dep.exit_code, dep.stdout[-4000:], dep.stderr[-4000:])

    # 4. Run it behind a much shorter fence. Separate budgets: a slow install
    #    is normal, a slow script is the thing you are defending against.
    run = sbx.exec(fenced(
        "cd /work && echo '<?php $x = 1 + 1;' | php -d memory_limit=512M analyse.php", 120))
    print("exit", run.exit_code)
    print(run.stdout[-8000:])
    print(run.stderr[-8000:])  # warnings and fatals land here with display_errors off
finally:
    sbx.kill()   # teardown. There is no sbx.delete().

Two things worth copying out of that even if you use a different platform. Separate time budgets for the install and the run, because a slow install is normal and a slow script is the thing you are defending against. And `mise reshim` after the install, because Composer drops executables into the project's `vendor/bin` and a shim-based runtime manager cannot see them until it is told to look again — which presents as "phpunit: command not found" ten minutes into an agent run, every time.

Then there is the part that makes PHP agents feel different from Python agents: the model can look at the page it generated.

# Serve what the agent built, then hand a human (or the agent) a real URL.
start = sbx.exec(fenced(
    "cd /work && setsid nohup php -S 0.0.0.0:8000 -t public "
    ">/var/log/php-dev.log 2>&1 < /dev/null & "
    "sleep 2; curl -sS -o /dev/null -w '%{http_code}' http://127.0.0.1:8000/", 60))
print(start.stdout)          # "200", or the reason it is not 200

# Tokenless: the port is reachable over HTTPS for the sandbox's lifetime and
# the sandbox UUID is the credential. Treat the URL like a secret, because it
# is one, and raise the TTL or set persistent if a human will look later.
print(sbx.preview_url(8000))

# Two honest caveats about php -S:
#   * it is single-threaded, so one slow request blocks the next one, and an
#     agent benchmarking against it is measuring the dev server
#   * it rewrites nothing, so a framework needs its own router script:
#         php -S 0.0.0.0:8000 -t public public/index.php
# For anything a reviewer will judge, run the real stack in the guest --
# php-fpm behind nginx, or FrankenPHP as a single binary.

On the database: for a throwaway rehearsal, run Postgres or MySQL inside the same sandbox, because it is instant, disposable and dies with the VM. Use a managed database only when the agent needs state that outlives the machine — on my platform that is a separate Firecracker VM with a durable volume and a 30 to 90 second create, which is the right call for a persistent environment and the wrong one for a throwaway test run.

The bottom line

There is no best sandbox API for PHP agents, and there is barely a PHP-specific one — nobody ships a PHP SDK, so you are writing glue whichever way you go. What there is, is a clear ordering of boundaries, and a PHP-shaped set of questions that will decide whether the thing is usable: can you pin the version, can `composer install` run, can you reach a port, and can you attach a disposable database.

My honest recommendation. If you only need to evaluate self-contained snippets, self-host Judge0 or Piston and stop reading roundups; it is less machinery and a better fit. If the code is semi-trusted — your own team's generated migrations, not the public internet — a container with seccomp on a box you already own is a defensible answer for a long time. If you want PHP in a browser with no server at all, php-wasm is a genuinely good answer to that question and a bad one to any other. Reach for microVMs, mine or someone else's, when the input is model-generated or customer-supplied, when `composer install` on an unreviewed manifest is part of the job, or when you need to describe the boundary to an auditor without using the word "mostly".

And do not pick PandaStack if you need a GPU, a guest kernel newer than 5.10, a first-party PHP SDK, or a vendor with a support contract and a decade of compliance paperwork. Those are real requirements and I do not meet them. If that trade still looks right, the primitives are documented on AI agent sandboxes, the one rate card is on pricing, and the PHP-specific follow-ups are below.

Frequently asked questions

Are disable_functions and open_basedir enough to sandbox untrusted PHP?

No, and the reason is structural rather than a bug waiting to be fixed. `disable_functions` is a denylist of function names, written by hand, over a standard library of thousands of functions plus every extension loaded into the process. Spawning a process, writing a file, opening a socket and loading native code each have several reachable front doors, and some of them live in extensions that exist for unrelated reasons, so the list is only correct until the next extension is enabled or the next release adds a function. That is why "the disable_functions bypass" reads as a genre of blog post rather than a series of incidents. `open_basedir` constrains PHP's own stream layer and not extension code that opens files through C, and it has lost to symlink and realpath-cache races often enough that it is not treated as containment by anyone who has looked. PHP itself reached this conclusion: `safe_mode` was the official in-language answer and was removed in 5.4 because solving the problem at the PHP level was architecturally wrong. Set both directives anyway — they are cheap and they remove the laziest paths. Then put a real boundary around the process: a container with seccomp at minimum, and a microVM with its own guest kernel if the code came from a model.

Can I run composer install on an untrusted composer.json?

Only inside something you are willing to lose, because `composer install` executes code by design. Composer supports install scripts and lifecycle hooks like `post-install-cmd`, and it has a plugin API whose plugins are PHP classes loaded and run inside Composer's own process. None of that is a vulnerability; it is the documented feature set that real packages rely on. The consequence for an agent pipeline is that "a human will review the code before it runs" is not a plan, because dependency resolution already ran. Mitigations in order of effectiveness: run it in a microVM or VM you destroy afterwards; pass `--no-scripts --no-plugins` and accept that some packages then misbehave; pin with a committed lockfile so resolution is not deciding anything at install time; and fence the command in-guest with something like `timeout --kill-after=10s 600 sh -c '...'`, because many SDK timeout parameters only widen the client's HTTP wait. Also remember the install needs network egress to a package registry, and that same egress reaches everywhere else, so if you care about exfiltration you need an allowlist rather than an open path.

Does the PHP version in the sandbox actually matter for an AI agent?

More than in most languages, because the 8.x line moved quickly and user code notices. Enums, fibers and readonly properties arrived in 8.1; 8.3 and 8.4 brought typed class constants, property hooks, asymmetric visibility and a tranche of deprecations that turn previously quiet code noisy. A model trained on a mix of all of it will cheerfully write 8.4 syntax into an 8.1 project and then be confused by the parse error, or write code that triggers a deprecation your CI treats as a failure. So treat the version as part of the task specification, not an environment detail: tell the model which version it is targeting, and make the sandbox match the project rather than matching the platform's default image. The operational version of this requirement is that version choice should be a template or image property, baked once, not an install step you pay for on every agent turn. If a platform's answer is "you can install any version you like", ask how long that install takes and then multiply it by the number of turns in a typical session.

How do I let an agent check its own work on a PHP web app?

Give it a port it can fetch, because PHP code is usually web code and stdout proves almost nothing about a template, a route or a middleware. The minimal version is a dev server in the guest plus a loopback curl, which catches fatals and 500s. The better version is a URL, so that both the agent and a human reviewer can look at the rendered page — on PandaStack that is `sbx.preview_url(8000)`, which maps a guest port to an HTTPS host for the sandbox's lifetime with no token, meaning the sandbox UUID is the credential and the URL should be treated as a secret. Two caveats on the quick path: `php -S` is single-threaded, so one slow request blocks the next and any benchmark you run against it is measuring the dev server, and it does no rewriting, so a framework needs its own router script passed on the command line. For anything a reviewer will judge, run the real stack in the guest — php-fpm behind nginx, or FrankenPHP as a single binary — because "works under the dev server" has been a false positive for twenty years.

Is a container enough isolation for model-generated PHP?

It depends entirely on where the code came from, and that is the question to answer before the technology question. A container gives you namespaces, cgroups and a seccomp profile, which is a genuine improvement over a child process: the code gets its own filesystem view, its own PID space and a capped share of CPU and memory. For semi-trusted input — your own team's generated migrations, a refactor suggested by a model you then read — that is a defensible boundary and plenty of serious companies stop there. What you have bought is one kernel, shared between the host and every container on it, with a syscall surface measured in hundreds of entry points, so an escape is a kernel bug away rather than a hypervisor bug away. For code that arrives from the public internet, from a customer, or from a model acting on a customer's instructions, I would not accept that, which is why I build on Firecracker: each sandbox gets its own guest kernel, and the thing an escape has to get through is a small, well-audited VMM rather than the whole Linux syscall table. Between those two sits gVisor, a real option with a real syscall-heavy performance cost that a composer install will find for you.

Keep reading

Related posts

  • The best sandbox APIs for .NET agents in 2026

    Roslyn scripting is not a sandbox, AssemblyLoadContext is not a boundary, and every sandbox vendor ships a Python SDK first. Here is what a C# agent actually does.

  • The best PHP hosting platforms in 2026

    PHP hosting advice is thirty years deep and most of it is stale. The useful question in 2026 is which of three very different runtime models you are buying.

  • Best Code-Execution Sandbox APIs for Go AI Agents in 2026

    You're building an agent in Go, and every sandbox vendor's docs open with Python and TypeScript. This is the buyer's checklist for the person who's going to be reading the REST reference instead — plus how to wrap any sandbox API in a Go client that doesn't embarrass you.

  • The Best Sandbox APIs for Java AI Agents in 2026

    You're building an agent in Java or Kotlin, and every sandbox vendor's quickstart opens with Python. The honest first finding: almost nobody ships a first-party Java SDK, so the real question is how good the REST API is and how well it generates a client.

  • How to deploy a Laravel app without Docker

    Your repo already declares the PHP version, the packages and the boot command. A Dockerfile that restates them is ceremony. Here is the Laravel deploy without one, and the eight decisions that actually decide whether it comes up healthy.

More in AI agent sandboxes · See AI agent sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.