Sandboxes for AI agents

The AI agent sandbox that boots
before your model finishes a token.

Every sandbox is a Firecracker microVM with its own Linux kernel, restored from a baked snapshot, with no warm pool behind it. Run model-written code, tool calls and full dev environments behind hardware virtualization, then throw the VM away.

49ms
p50 create → ready
1 → N
fork a running VM
1 kernel
per sandbox, never shared
sandbox.create() — snapshot restore

What is an AI agent sandbox?

An AI agent sandbox is an isolated, disposable computer where an AI agent runs the code, shell commands and tool calls a model generated, without access to your host, its credentials or other users' data. On PandaStack, each sandbox is a Firecracker microVM with its own Linux kernel, created from a snapshot in 49ms p50.

Isolated

Every sandbox is its own Firecracker microVM on KVM, with a dedicated Linux guest kernel, its own network namespace and a private copy-on-write disk. Root inside the guest is the normal, supported state.

Disposable

A create restores a baked snapshot; a delete tears the VM down. Pass ttl_seconds on create as a backstop, so a sandbox your agent forgets reaps itself.

Programmable

One-shot and streaming exec, an interactive PTY, file upload and download, a preview URL for any port, snapshots and forks. Drive it from the Python SDK (pip install pandastack), the TypeScript SDK (@pandastack/sdk), the CLI or MCP.

Most agents need one of three templates: code-interpreter (Python and Node with pandas, numpy and Jupyter), browser (headless Chromium, Playwright and Xvfb) or base (Ubuntu 24.04 with Node 24, Python 3.12, Go and Bun). The agent template adds the claude, codex, opencode, amp, grok, gemini and copilot CLIs. All templates.

Why AI agents need a sandbox

A model emits actions that nobody reviewed before they run. Executing them on a laptop, a CI runner or a production node hands that output your shell, your files and your credentials.

Untrusted by construction

Every tool call is code from an author you can't vet: the model, plus whatever ended up in its context. Treat it like a pull request from a stranger that merges itself.

Injection becomes execution

A README, a web page or an issue comment can steer an agent into running an attacker's command. The sandbox decides what that command can reach. The prompt-injection-to-RCE chain.

Ambient credentials leak

On a developer machine or CI runner the agent inherits SSH keys, cloud tokens and often the cloud metadata endpoint. A PandaStack sandbox starts with none of them and can't reach the metadata service.

Blast radius

In a shared-kernel container, a kernel exploit is a host escape. In a microVM it gets the attacker a guest kernel; reaching the host still takes a break through KVM and Firecracker. Why Docker isn't a sandbox.

A secure AI agent sandbox starts with its own kernel.

Containers share the host kernel, so every syscall untrusted code makes lands in the kernel every other tenant uses. gVisor narrows that with a user-space kernel. A microVM gives each workload its own guest kernel behind hardware virtualization. PandaStack runs every sandbox in Firecracker, the open-source VMM AWS developed for Lambda and Fargate, unmodified.

Isolation models for running untrusted agent code: Firecracker microVM, gVisor, and shared-kernel containers
Firecracker microVM (PandaStack)gVisorShared-kernel container
Kernel serving the codeIts own Linux guest kernel, on KVMgVisor's user-space application kernel; the host kernel sits underneathThe host kernel, shared with every container on the node
Host-facing surfaceKVM plus a small virtio device model (block, net, vsock)A syscall layer reimplemented in user spaceThe full host syscall interface, narrowed by seccomp and LSM profiles
What an escape takesA guest-to-host break through KVM or FirecrackerA break out of gVisor's kernel, then the hostOne kernel bug or runtime misconfiguration
CompatibilityLinux syscall semantics from a real guest kernel; no PCI passthroughNot every syscall, /proc or /sys file is implementedFull, since it is the host kernel
Cost of the boundaryBoot time, removed here by snapshot restore (49ms p50 create)Higher per-syscall overheadLow, but for untrusted code it is a packaging boundary, not a security boundary

What a microVM doesn't do: the host kernel and KVM still have to be correct, the network and disk plumbing around each VM is host-side code, and CPU side channels are mitigated on the host rather than eliminated by virtualization. The full model is in the isolation model docs; for a longer comparison, see gVisor vs Firecracker for agents.

Security controls in every sandbox

The VM is the boundary. These are the host-side controls around it, plus the one thing that is deliberately left open.

Own guest kernel on KVM

Each sandbox boots its own Linux kernel under hardware virtualization. Firecracker's device model is minimal (virtio block, net and vsock) with no PCI passthrough, no BIOS and no legacy device emulation.

Network namespace per VM

Every VM gets a dedicated network namespace with one TAP device and a unique /30. Outbound traffic is source-NATed per VM, so each connection is attributable to one sandbox and one workspace.

Sandbox-to-sandbox traffic dropped

An explicit DROP at the top of the host's FORWARD chain blocks VM-to-VM traffic on every port, including SSH and Postgres. Unit tests pin that it is inserted first in the chain.

No cloud metadata, no ambient identity

The whole link-local range is dropped, so a guest can't query the cloud metadata endpoint for host credentials. Sandboxes carry no cloud service account and no instance credentials.

Private disk and memory

Each sandbox writes to its own copy-on-write clone of the template disk, and snapshot memory is mapped MAP_PRIVATE. A sandbox's writes never reach the template, its parent or a sibling fork.

Mining-port egress blocked

Outbound TCP to the standard Stratum mining-pool ports is dropped by default. Self-hosted deployments can change the list with PANDASTACK_BLOCKED_EGRESS_PORTS.

Outbound internet is on by default. Agents install packages, call APIs and clone repos, so beyond the blocks above there is no outbound allowlist today. Keep long-lived secrets out of the sandbox, hand the agent short-lived scoped tokens per task, and keep the orchestrator that holds your model API keys outside the VM that runs the code.

More in controlling egress for untrusted code and the egress controls docs. Compliance status, including the certifications PandaStack does not hold yet, is on the security page.

Snapshot restore on every create.

There is no pool of idle VMs waiting for you. A create allocates a pre-built network slot, clones the root filesystem copy-on-write and restores the template's memory snapshot, so the guest wakes with its runtime already warm.

Why restore beats warm pools

Warm pools bill someone for idle capacity and still run dry under load. A snapshot restore pages memory in on demand, so there is no pool to exhaust. On the live API, server-measured create → ready is 49ms p50 across 50 consecutive creates of the base template (June 17, 2026).

Concurrent creates take longer per request, and your client adds its own network round-trip. The method and raw numbers are on the benchmarks page.

alloc net slotclone rootfs (CoW)restore + resume
python — pip install pandastack
from pandastack import Sandbox

# restored from the template's snapshot;
# ttl_seconds reaps it if your agent forgets
sb = Sandbox.create(template="code-interpreter", ttl_seconds=600)

out = sb.exec("python3 -c 'print(2**64)'")
print(out.exit_code, out.stdout)

snap_id = sb.snapshot()   # memory + disk, returns the snapshot id
sb.kill()                 # delete the VM

Branch the whole computer, keep the winner.

fork_tree() snapshots a running sandbox once and restores N children from that snapshot. Each child gets the parent's memory and disk, so running processes survive, plus a fresh network identity and copy-on-write storage. Give an agent N attempts, promote() the one that worked, and the siblings are reaped.

fork_tree & fork

fork_tree carries memory and disk, so a warm interpreter or a running dev server is already there in every child, each restored in a few hundred milliseconds after the one-time parent snapshot. A plain fork copies only the disk and cold-boots the child, which suits fan-out from a prepared environment.

Snapshots

Capture a running sandbox's memory and disk at any instant, then create new sandboxes from that point with from_snapshot.

Hibernate & wake

A hibernated sandbox writes its memory and disk to a snapshot and stops billing compute. The next request wakes it with files, processes and memory where they were.

python — best of N with fork_tree
from pandastack import Sandbox

parent = Sandbox.create(template="code-interpreter", ttl_seconds=900)
parent.filesystem.upload("solve.py", "/tmp/solve.py")

kids = parent.fork_tree(count=3)   # memory + disk, fresh network per child
runs = [k.exec(f"python3 /tmp/solve.py --seed {i}") for i, k in enumerate(kids)]

winner = next((k for k, r in zip(kids, runs) if r.exit_code == 0), None)
if winner:
    winner.promote(cleanup_siblings=True)   # keep one, reap its siblings
else:
    for k in kids:
        k.kill()
parent.kill()

Every way an agent wants to talk to a machine.

One-shot REST exec for tools and CI, an interactive PTY for terminals and human handoff, and a multiplexed WebSocket for high-frequency agent loops that can't pay a handshake on every command.

REST · PTY · WS exec

POST a command, open a resizable shell, or multiplex tagged commands over one socket.

Files in & out

Stream source trees, artifacts and prompt bundles into a running microVM, and pull results back out.

Preview URLs

Any port is reachable at a stable per-sandbox URL for demos and webhooks.

Durable volumes

Attach named volumes for caches, repos, model shards and user state.

Works with the agent stack you already run

Every sandbox exposes an MCP endpoint with shell, read_file, write_file and list_dir tools. On Pro, Team and Enterprise, a hosted workspace MCP server at api.pandastack.ai/mcp can also create sandboxes, run commands and manage databases and apps. Framework guides:

AI agent sandbox use cases

The same primitive (a disposable microVM with exec, files, ports, snapshots and forks) covers most of what agents do with a computer.

Coding agents

Clone a repo, install dependencies, run the tests and open a preview URL for the result. The agent template ships the claude, codex, opencode, gemini and copilot CLIs.

Code interpreter

create_code_context() keeps a Jupyter-style kernel alive across calls, and charts and tables come back as PNG, HTML or JSON results instead of scraped stdout.

Browser automation and computer use

The browser template runs headless Chromium, Playwright and Xvfb inside the VM, so the pages an agent visits never touch your network.

Evals and parallel attempts

Prepare an environment once, fork_tree it N ways, score each branch and promote the winner. Every attempt starts from identical memory and disk.

Agent CI and PR checks

A fresh VM per run for code nobody has reviewed yet, thrown away when the job finishes.

Long-running, stateful agents

Keep one persistent sandbox per agent and call hibernate() between turns. It bills no compute while hibernated and wakes on the next request with its state intact.

How PandaStack differs

The design choices behind this sandbox, in short.

  • Forks carry memory.fork_tree restores the parent's memory and disk into every child, so a loaded dataset or a running server is already there in each branch.
  • Isolation you can read. The firewall rules, network namespaces and snapshot paths described on this page are in the Apache-2.0 repository.
  • Managed or self-hosted. Use the hosted API, or run the same stack on any Linux host with /dev/kvm.
  • One substrate. Managed Postgres (each database in its own microVM) and git-driven app hosting run on the same platform, so an agent can provision the database and deploy the app it just built.
  • Billed on use. CPU is metered on active vCPU-seconds and memory on the working set, per second; a hibernated sandbox bills no compute.

Weighing hosted options? See code execution sandboxes compared. For the rest of the agent stack (state between turns, Postgres for memory, audit tags), see AI agent infrastructure.

Open-source AI agent sandbox you can self-host

PandaStack is Apache-2.0. Read the code that isolates your agents, run it in your own VPC or on-prem, and patch it.

What a self-hosted deployment is

A control plane (API + Postgres), an agent on each KVM-capable Linux host, and the Firecracker microVMs they manage. Hosts need /dev/kvm: bare metal, or a cloud VM with nested virtualization.

Guides cover GCP (Terraform on Compute Engine with nested virtualization), AWS bare metal (c5n.metal), on-prem hardware, a single Ubuntu or Debian box, and Apple Silicon via Lima. Sandbox workloads, filesystems and snapshots stay in your environment.

bash — one Linux host with /dev/kvm
git clone https://github.com/pandastack-io/pandastack-ai
cd pandastack-ai
bash scripts/linux-local-e2e.sh    # Postgres, API, dashboard and agent

# the same CLI and SDKs, pointed at your API
npm install -g @pandastack/sdk     # provides the pandastack command
export PANDASTACK_API=https://sandbox-api.internal.example
export PANDASTACK_API_KEY=pds_...
pandastack sandbox create --template base --ttl 3600

The internals are public.

The whole substrate is open source: read exactly how boot, fork and isolation work.

AI agent sandbox FAQ

What is an AI agent sandbox? +

An AI agent sandbox is an isolated, disposable execution environment where an AI agent runs the code, shell commands and tool calls a model generated, without access to your host, its credentials or other users' data. On PandaStack each sandbox is a Firecracker microVM with its own Linux kernel. It is created from a snapshot in 49ms p50 (server-measured) and deleted when the task is done.

Is Docker enough to sandbox an AI agent? +

For trusted first-party code, often yes. For code a model wrote, a container is a packaging boundary rather than a security boundary: it shares the host kernel, so a kernel exploit or a runtime misconfiguration becomes a container escape. A microVM runs its own guest kernel behind hardware virtualization, so untrusted code would have to escape through KVM and Firecracker's small device model instead.

MicroVM or gVisor: which is better for AI agents? +

gVisor intercepts system calls in a user-space application kernel, which is a real improvement over plain containers. Its own documentation notes reduced application compatibility and higher per-syscall overhead, and the host kernel is still underneath. A Firecracker microVM gives each sandbox a full Linux guest kernel on KVM with a minimal device model. The usual cost of a VM is boot time; PandaStack restores every sandbox from a snapshot, so create is 49ms p50 (server-measured).

Can a sandbox reach the internet, and how is egress controlled? +

Yes. Outbound access is on by default so agents can install packages, call APIs and clone repositories, and there is no outbound allowlist today. Each VM egresses through its own network namespace and NAT, so traffic is attributable to one sandbox and workspace. VM-to-VM traffic, the link-local range that serves cloud metadata, and the standard Stratum mining-pool ports are dropped at the host firewall. Keep long-lived secrets out of the sandbox and pass short-lived, scoped tokens instead.

Does state persist between tool calls? +

Within a sandbox, yes: files, installed packages and running processes stay until you delete it or its TTL expires. For longer gaps, snapshot() captures memory and disk, and hibernate() writes state to disk and stops compute billing; the next request wakes the sandbox. On the Free plan a sandbox is deleted one hour after it is created. On paid plans the TTL is an idle timeout, up to 7 days on Pro and 30 days on Team.

How fast are create and fork? +

On the live API, server-measured create to ready is 49ms p50 across 50 consecutive creates of the base template (June 17, 2026). Concurrent creates take longer per request, and your client adds its own network round-trip; the method and raw numbers are at pandastack.ai/benchmarks. fork_tree() snapshots the parent once, then restores every child with the parent's memory and disk in a few hundred milliseconds. A plain fork() copies only the disk and cold-boots the child.

Can I self-host the AI agent sandbox? +

Yes. PandaStack is open source under Apache-2.0 at github.com/pandastack-io/pandastack-ai. A self-hosted deployment is a control plane (API + Postgres), an agent on each Linux host with /dev/kvm, and the Firecracker microVMs they manage. There are guides for GCP, AWS bare metal, on-prem hardware, a single Linux box and Apple Silicon via Lima, and the same Python and TypeScript SDKs point at your own API endpoint.

Does it work with MCP, LangChain, the OpenAI Agents SDK and Claude? +

Yes. Every sandbox exposes an MCP endpoint with shell, read_file, write_file and list_dir tools, and a hosted workspace MCP server (Pro, Team and Enterprise) can create sandboxes, run commands and manage databases and apps. The docs have guides for LangChain, LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK, the Vercel AI SDK, LlamaIndex, Pydantic AI, smolagents, Mastra and Claude Managed Agents.

How is an AI agent sandbox billed? +

Per second, at $0.054 per active vCPU-hour and $0.0162 per working-set GiB-hour. CPU counts the seconds your code actually burns, not the 8 vCPUs a sandbox can burst to, and memory counts what stays resident. A hibernated sandbox bills no compute. The Free plan includes $5.40 of usage credit a month, 5 concurrent sandboxes and 60 creates an hour.

Ship on the millisecond cloud.

Free tier with $5.40/mo usage credit. No card. Apache-2.0.