The AI agent sandbox that boots
before your model finishes a token.
Every sandbox is a Firecracker microVM with its own Linux kernel, restored from a baked snapshot, with no warm pool behind it. Run model-written code, tool calls and full dev environments behind hardware virtualization, then throw the VM away.
What is an AI agent sandbox?
An AI agent sandbox is an isolated, disposable computer where an AI agent runs the code, shell commands and tool calls a model generated, without access to your host, its credentials or other users' data. On PandaStack, each sandbox is a Firecracker microVM with its own Linux kernel, created from a snapshot in 49ms p50.
Isolated
Every sandbox is its own Firecracker microVM on KVM, with a dedicated Linux guest kernel, its own network namespace and a private copy-on-write disk. Root inside the guest is the normal, supported state.
Disposable
A create restores a baked snapshot; a delete tears the VM down. Pass ttl_seconds on create as a backstop, so a sandbox your agent forgets reaps itself.
Programmable
One-shot and streaming exec, an interactive PTY, file upload and download, a preview URL for any port, snapshots and forks. Drive it from the Python SDK (pip install pandastack), the TypeScript SDK (@pandastack/sdk), the CLI or MCP.
Most agents need one of three templates: code-interpreter (Python and Node with pandas, numpy and Jupyter), browser (headless Chromium, Playwright and Xvfb) or base (Ubuntu 24.04 with Node 24, Python 3.12, Go and Bun). The agent template adds the claude, codex, opencode, amp, grok, gemini and copilot CLIs. All templates.
Why AI agents need a sandbox
A model emits actions that nobody reviewed before they run. Executing them on a laptop, a CI runner or a production node hands that output your shell, your files and your credentials.
Untrusted by construction
Every tool call is code from an author you can't vet: the model, plus whatever ended up in its context. Treat it like a pull request from a stranger that merges itself.
Injection becomes execution
A README, a web page or an issue comment can steer an agent into running an attacker's command. The sandbox decides what that command can reach. The prompt-injection-to-RCE chain.
Ambient credentials leak
On a developer machine or CI runner the agent inherits SSH keys, cloud tokens and often the cloud metadata endpoint. A PandaStack sandbox starts with none of them and can't reach the metadata service.
Blast radius
In a shared-kernel container, a kernel exploit is a host escape. In a microVM it gets the attacker a guest kernel; reaching the host still takes a break through KVM and Firecracker. Why Docker isn't a sandbox.
A secure AI agent sandbox starts with its own kernel.
Containers share the host kernel, so every syscall untrusted code makes lands in the kernel every other tenant uses. gVisor narrows that with a user-space kernel. A microVM gives each workload its own guest kernel behind hardware virtualization. PandaStack runs every sandbox in Firecracker, the open-source VMM AWS developed for Lambda and Fargate, unmodified.
| Firecracker microVM (PandaStack) | gVisor | Shared-kernel container | |
|---|---|---|---|
| Kernel serving the code | Its own Linux guest kernel, on KVM | gVisor's user-space application kernel; the host kernel sits underneath | The host kernel, shared with every container on the node |
| Host-facing surface | KVM plus a small virtio device model (block, net, vsock) | A syscall layer reimplemented in user space | The full host syscall interface, narrowed by seccomp and LSM profiles |
| What an escape takes | A guest-to-host break through KVM or Firecracker | A break out of gVisor's kernel, then the host | One kernel bug or runtime misconfiguration |
| Compatibility | Linux syscall semantics from a real guest kernel; no PCI passthrough | Not every syscall, /proc or /sys file is implemented | Full, since it is the host kernel |
| Cost of the boundary | Boot time, removed here by snapshot restore (49ms p50 create) | Higher per-syscall overhead | Low, but for untrusted code it is a packaging boundary, not a security boundary |
What a microVM doesn't do: the host kernel and KVM still have to be correct, the network and disk plumbing around each VM is host-side code, and CPU side channels are mitigated on the host rather than eliminated by virtualization. The full model is in the isolation model docs; for a longer comparison, see gVisor vs Firecracker for agents.
Security controls in every sandbox
The VM is the boundary. These are the host-side controls around it, plus the one thing that is deliberately left open.
Own guest kernel on KVM
Each sandbox boots its own Linux kernel under hardware virtualization. Firecracker's device model is minimal (virtio block, net and vsock) with no PCI passthrough, no BIOS and no legacy device emulation.
Network namespace per VM
Every VM gets a dedicated network namespace with one TAP device and a unique /30. Outbound traffic is source-NATed per VM, so each connection is attributable to one sandbox and one workspace.
Sandbox-to-sandbox traffic dropped
An explicit DROP at the top of the host's FORWARD chain blocks VM-to-VM traffic on every port, including SSH and Postgres. Unit tests pin that it is inserted first in the chain.
No cloud metadata, no ambient identity
The whole link-local range is dropped, so a guest can't query the cloud metadata endpoint for host credentials. Sandboxes carry no cloud service account and no instance credentials.
Private disk and memory
Each sandbox writes to its own copy-on-write clone of the template disk, and snapshot memory is mapped MAP_PRIVATE. A sandbox's writes never reach the template, its parent or a sibling fork.
Mining-port egress blocked
Outbound TCP to the standard Stratum mining-pool ports is dropped by default. Self-hosted deployments can change the list with PANDASTACK_BLOCKED_EGRESS_PORTS.
Outbound internet is on by default. Agents install packages, call APIs and clone repos, so beyond the blocks above there is no outbound allowlist today. Keep long-lived secrets out of the sandbox, hand the agent short-lived scoped tokens per task, and keep the orchestrator that holds your model API keys outside the VM that runs the code.
More in controlling egress for untrusted code and the egress controls docs. Compliance status, including the certifications PandaStack does not hold yet, is on the security page.
Snapshot restore on every create.
There is no pool of idle VMs waiting for you. A create allocates a pre-built network slot, clones the root filesystem copy-on-write and restores the template's memory snapshot, so the guest wakes with its runtime already warm.
Why restore beats warm pools
Warm pools bill someone for idle capacity and still run dry under load. A snapshot restore pages memory in on demand, so there is no pool to exhaust. On the live API, server-measured create → ready is 49ms p50 across 50 consecutive creates of the base template (June 17, 2026).
Concurrent creates take longer per request, and your client adds its own network round-trip. The method and raw numbers are on the benchmarks page.
from pandastack import Sandbox
# restored from the template's snapshot;
# ttl_seconds reaps it if your agent forgets
sb = Sandbox.create(template="code-interpreter", ttl_seconds=600)
out = sb.exec("python3 -c 'print(2**64)'")
print(out.exit_code, out.stdout)
snap_id = sb.snapshot() # memory + disk, returns the snapshot id
sb.kill() # delete the VMBranch the whole computer, keep the winner.
fork_tree() snapshots a running sandbox once and restores N children from that snapshot. Each child gets the parent's memory and disk, so running processes survive, plus a fresh network identity and copy-on-write storage. Give an agent N attempts, promote() the one that worked, and the siblings are reaped.
fork_tree & fork
fork_tree carries memory and disk, so a warm interpreter or a running dev server is already there in every child, each restored in a few hundred milliseconds after the one-time parent snapshot. A plain fork copies only the disk and cold-boots the child, which suits fan-out from a prepared environment.
Snapshots
Capture a running sandbox's memory and disk at any instant, then create new sandboxes from that point with from_snapshot.
Hibernate & wake
A hibernated sandbox writes its memory and disk to a snapshot and stops billing compute. The next request wakes it with files, processes and memory where they were.
from pandastack import Sandbox
parent = Sandbox.create(template="code-interpreter", ttl_seconds=900)
parent.filesystem.upload("solve.py", "/tmp/solve.py")
kids = parent.fork_tree(count=3) # memory + disk, fresh network per child
runs = [k.exec(f"python3 /tmp/solve.py --seed {i}") for i, k in enumerate(kids)]
winner = next((k for k, r in zip(kids, runs) if r.exit_code == 0), None)
if winner:
winner.promote(cleanup_siblings=True) # keep one, reap its siblings
else:
for k in kids:
k.kill()
parent.kill()Every way an agent wants to talk to a machine.
One-shot REST exec for tools and CI, an interactive PTY for terminals and human handoff, and a multiplexed WebSocket for high-frequency agent loops that can't pay a handshake on every command.
REST · PTY · WS exec
POST a command, open a resizable shell, or multiplex tagged commands over one socket.
Files in & out
Stream source trees, artifacts and prompt bundles into a running microVM, and pull results back out.
Preview URLs
Any port is reachable at a stable per-sandbox URL for demos and webhooks.
Durable volumes
Attach named volumes for caches, repos, model shards and user state.
Works with the agent stack you already run
Every sandbox exposes an MCP endpoint with shell, read_file, write_file and list_dir tools. On Pro, Team and Enterprise, a hosted workspace MCP server at api.pandastack.ai/mcp can also create sandboxes, run commands and manage databases and apps. Framework guides:
AI agent sandbox use cases
The same primitive (a disposable microVM with exec, files, ports, snapshots and forks) covers most of what agents do with a computer.
Coding agents
Clone a repo, install dependencies, run the tests and open a preview URL for the result. The agent template ships the claude, codex, opencode, gemini and copilot CLIs.
Code interpreter
create_code_context() keeps a Jupyter-style kernel alive across calls, and charts and tables come back as PNG, HTML or JSON results instead of scraped stdout.
Browser automation and computer use
The browser template runs headless Chromium, Playwright and Xvfb inside the VM, so the pages an agent visits never touch your network.
Evals and parallel attempts
Prepare an environment once, fork_tree it N ways, score each branch and promote the winner. Every attempt starts from identical memory and disk.
Agent CI and PR checks
A fresh VM per run for code nobody has reviewed yet, thrown away when the job finishes.
Long-running, stateful agents
Keep one persistent sandbox per agent and call hibernate() between turns. It bills no compute while hibernated and wakes on the next request with its state intact.
How PandaStack differs
The design choices behind this sandbox, in short.
- Forks carry memory.fork_tree restores the parent's memory and disk into every child, so a loaded dataset or a running server is already there in each branch.
- Isolation you can read. The firewall rules, network namespaces and snapshot paths described on this page are in the Apache-2.0 repository.
- Managed or self-hosted. Use the hosted API, or run the same stack on any Linux host with /dev/kvm.
- One substrate. Managed Postgres (each database in its own microVM) and git-driven app hosting run on the same platform, so an agent can provision the database and deploy the app it just built.
- Billed on use. CPU is metered on active vCPU-seconds and memory on the working set, per second; a hibernated sandbox bills no compute.
Weighing hosted options? See code execution sandboxes compared. For the rest of the agent stack (state between turns, Postgres for memory, audit tags), see AI agent infrastructure.
Open-source AI agent sandbox you can self-host
PandaStack is Apache-2.0. Read the code that isolates your agents, run it in your own VPC or on-prem, and patch it.
What a self-hosted deployment is
A control plane (API + Postgres), an agent on each KVM-capable Linux host, and the Firecracker microVMs they manage. Hosts need /dev/kvm: bare metal, or a cloud VM with nested virtualization.
Guides cover GCP (Terraform on Compute Engine with nested virtualization), AWS bare metal (c5n.metal), on-prem hardware, a single Ubuntu or Debian box, and Apple Silicon via Lima. Sandbox workloads, filesystems and snapshots stay in your environment.
git clone https://github.com/pandastack-io/pandastack-ai
cd pandastack-ai
bash scripts/linux-local-e2e.sh # Postgres, API, dashboard and agent
# the same CLI and SDKs, pointed at your API
npm install -g @pandastack/sdk # provides the pandastack command
export PANDASTACK_API=https://sandbox-api.internal.example
export PANDASTACK_API_KEY=pds_...
pandastack sandbox create --template base --ttl 3600The internals are public.
The whole substrate is open source: read exactly how boot, fork and isolation work.
AI agent sandbox FAQ
What is an AI agent sandbox? +
An AI agent sandbox is an isolated, disposable execution environment where an AI agent runs the code, shell commands and tool calls a model generated, without access to your host, its credentials or other users' data. On PandaStack each sandbox is a Firecracker microVM with its own Linux kernel. It is created from a snapshot in 49ms p50 (server-measured) and deleted when the task is done.
Is Docker enough to sandbox an AI agent? +
For trusted first-party code, often yes. For code a model wrote, a container is a packaging boundary rather than a security boundary: it shares the host kernel, so a kernel exploit or a runtime misconfiguration becomes a container escape. A microVM runs its own guest kernel behind hardware virtualization, so untrusted code would have to escape through KVM and Firecracker's small device model instead.
MicroVM or gVisor: which is better for AI agents? +
gVisor intercepts system calls in a user-space application kernel, which is a real improvement over plain containers. Its own documentation notes reduced application compatibility and higher per-syscall overhead, and the host kernel is still underneath. A Firecracker microVM gives each sandbox a full Linux guest kernel on KVM with a minimal device model. The usual cost of a VM is boot time; PandaStack restores every sandbox from a snapshot, so create is 49ms p50 (server-measured).
Can a sandbox reach the internet, and how is egress controlled? +
Yes. Outbound access is on by default so agents can install packages, call APIs and clone repositories, and there is no outbound allowlist today. Each VM egresses through its own network namespace and NAT, so traffic is attributable to one sandbox and workspace. VM-to-VM traffic, the link-local range that serves cloud metadata, and the standard Stratum mining-pool ports are dropped at the host firewall. Keep long-lived secrets out of the sandbox and pass short-lived, scoped tokens instead.
Does state persist between tool calls? +
Within a sandbox, yes: files, installed packages and running processes stay until you delete it or its TTL expires. For longer gaps, snapshot() captures memory and disk, and hibernate() writes state to disk and stops compute billing; the next request wakes the sandbox. On the Free plan a sandbox is deleted one hour after it is created. On paid plans the TTL is an idle timeout, up to 7 days on Pro and 30 days on Team.
How fast are create and fork? +
On the live API, server-measured create to ready is 49ms p50 across 50 consecutive creates of the base template (June 17, 2026). Concurrent creates take longer per request, and your client adds its own network round-trip; the method and raw numbers are at pandastack.ai/benchmarks. fork_tree() snapshots the parent once, then restores every child with the parent's memory and disk in a few hundred milliseconds. A plain fork() copies only the disk and cold-boots the child.
Can I self-host the AI agent sandbox? +
Yes. PandaStack is open source under Apache-2.0 at github.com/pandastack-io/pandastack-ai. A self-hosted deployment is a control plane (API + Postgres), an agent on each Linux host with /dev/kvm, and the Firecracker microVMs they manage. There are guides for GCP, AWS bare metal, on-prem hardware, a single Linux box and Apple Silicon via Lima, and the same Python and TypeScript SDKs point at your own API endpoint.
Does it work with MCP, LangChain, the OpenAI Agents SDK and Claude? +
Yes. Every sandbox exposes an MCP endpoint with shell, read_file, write_file and list_dir tools, and a hosted workspace MCP server (Pro, Team and Enterprise) can create sandboxes, run commands and manage databases and apps. The docs have guides for LangChain, LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK, the Vercel AI SDK, LlamaIndex, Pydantic AI, smolagents, Mastra and Claude Managed Agents.
How is an AI agent sandbox billed? +
Per second, at $0.054 per active vCPU-hour and $0.0162 per working-set GiB-hour. CPU counts the seconds your code actually burns, not the 8 vCPUs a sandbox can burst to, and memory counts what stays resident. A hibernated sandbox bills no compute. The Free plan includes $5.40 of usage credit a month, 5 concurrent sandboxes and 60 creates an hour.
Ship on the millisecond cloud.
Free tier with $5.40/mo usage credit. No card. Apache-2.0.