solutions

AI agent infrastructure:
give your agent a real computer.

An agent that can only eval a string isn't doing real work. Each session gets a full Linux microVM — filesystem, shell, packages, long-lived state — with its own kernel behind hardware virtualization, so model-written commands can't reach your host, other tenants, or the cloud metadata service.

49ms
p50 sandbox create
N-way
fork for tree search
$0
compute while hibernated
sandbox.create() — snapshot restore

What is AI agent infrastructure?

It is everything under the agent loop that isn't the model or the framework: the computer where tool calls execute, the state that survives between turns, and the boundary that keeps model-written code away from everything else. An agent that only writes text needs none of it. An agent that installs packages, edits a repository, runs tests, or queries a database needs a machine to do that on, and because the model decides what runs, that machine has to be isolated from yours.

The four layers of an agent loop

The layers of an AI agent loop and where PandaStack fits
layerits jobon PandaStack
ModelDecides the next action: which tool to call, with which arguments.Any provider. PandaStack runs the agent's tools, not the model; your code calls the model API as it does today.
FrameworkRuns the loop: sends messages, parses tool calls, retries, hands results back to the model.Unchanged. You register one PandaStack-backed tool in LangGraph, CrewAI, the OpenAI Agents SDK, the Vercel AI SDK, Mastra and others.
RuntimeThe machine where tool calls execute: shell, filesystem, packages, processes, network.A Firecracker microVM per session, the AI agent sandbox: 49ms p50 create, fork N ways, hibernate between turns.
Tools and stateWhat outlives a single tool call: databases, files, a snapshot of a working session.Managed Postgres for agent memory, including ephemeral databases for agents that you branch or clone per run and delete afterwards. Snapshots capture a whole session.

Frameworks and runtimes get confused because both are sold as “the agent platform”. The AI agent runtime guide covers the difference in depth, what to check when you choose a runtime, and how to self-host one.

A computer for AI agents, not a code-eval endpoint.

Real agent plans mutate filesystems, install dependencies, run test suites, and keep state across tool calls. That takes a machine, not a stateless function, and it takes isolation you can hand untrusted model output without flinching.

A shell, not a string eval

Run any command, stream stdout/stderr back as the tool result, and let the model iterate: pip install, git, pytest, whatever the plan needs.

A filesystem that persists

Upload a workspace, let the agent edit files across many tool calls in one session, and download the result when the turn ends.

Hardware isolation by default

Every sandbox is its own Firecracker microVM with its own kernel. Egress is NAT'd in a private network namespace; inter-sandbox traffic and the cloud metadata service are blocked.

Fork the machine N ways, keep the winner.

Agents are tree search. When yours hits a fork in the road (which fix, which library, which algorithm), explore() forks the sandbox once per candidate, runs each approach in parallel real execution, scores the outcomes, promotes the winner, and reaps the losers.

Best-of-N over real execution

Branches come from one snapshot of the running parent, so each child starts with the parent's memory and disk: the loaded interpreter, the data, the installed packages. Forks are copy-on-write, ~400ms on the same host, and share the parent's unwritten pages until a branch writes.

Each branch reports exit code, output tails, files changed, and duration: a structured outcome the model can reason over. Up to 16 branches per experiment. Losers are reaped when the winner is promoted, with a TTL backstop if your process dies mid-run.

fork per candidatescore outcomespromote winner
python — the agent loop
from pandastack import Sandbox

sb = Sandbox.create(template="code-interpreter",
                    ttl_seconds=3600,
                    metadata={"agent": "solver"})

# every tool call runs behind hardware isolation
out = sb.run_code(tool_call["code"], language="python")

# the agent hits a fork in the road — try all three
sb.filesystem.write("/work/solve.py", solver_code)
result = sb.explore(
    ["python3 /work/solve.py --strategy greedy",
     "python3 /work/solve.py --strategy dp",
     "python3 /work/solve.py --strategy beam"],
    score_fn=lambda b, o: 1.0 if o.ok
        and "PASS" in o.stdout_tail else 0.0,
)
winner = result.winner   # promoted; losers reaped

Idle agents should cost nothing.

Most of an agent's wall-clock is the model thinking, or the user away. A running sandbox bills active CPU-seconds and resident memory, so waiting is already cheap; hibernate it between turns and compute billing stops, then it wakes with its state intact when the next tool call lands.

Hibernate & wake

Park a session while the model reasons or the human sleeps. Memory and disk are written to a snapshot; files, processes, and memory come back where they left off.

Postgres for agent memory

Attach a managed PostgreSQL database for durable agent memory and state that outlives any single sandbox or session. An idle database suspends and wakes on the next connection.

Tagged for audit

Tag every agent sandbox with metadata at create time, so cleanup, audit, and per-agent accounting stay one query away.

Plugs into the agent framework you already use.

For most frameworks the integration is one tool: the framework keeps the loop, and the tool's body runs the model's code in a sandbox and returns the result. The recipes differ only in how that tool is registered. Claude Managed Agents and MCP clients connect differently, and each has its own guide below.

One tool, any framework

A persistent code context keeps variables and imports across tool calls, so the agent can load a DataFrame in one step and plot it in the next. The tool returns stdout, a traceback the model can debug and retry against, or a chart as a base64 PNG.

MCP clients can skip the wrapper. Every sandbox exposes an MCP endpoint with shell, read_file, write_file, and list_dir tools, and the hosted MCP server (Pro, Team, and Enterprise) lets a model create sandboxes, provision databases, and trigger app deploys from one URL.

python — the tool body
from pandastack import Sandbox

# one sandbox + one persistent kernel per agent session
sandbox = Sandbox.create(template="code-interpreter",
                         ttl_seconds=3600)
ctx = sandbox.create_code_context()

# wrap with @tool (LangGraph) or @function_tool
# (OpenAI Agents SDK); the body is the same
def run_python(code: str) -> str:
    ex = ctx.run_code(code)
    if ex.error:
        return ex.error        # a traceback the model can fix
    return ex.text or ex.stdout

Two primitives, one agent runtime.

This solution is composed from the platform's building blocks: isolated compute for the agent's hands, a managed database for its memory.

AI agent infrastructure: common questions

What is AI agent infrastructure? +

It is the layer under the agent loop that is neither the model nor the framework: the machine where tool calls execute, the state that survives between turns, and the isolation boundary around model-written code. On PandaStack that is a Firecracker microVM per agent session, plus managed Postgres for anything that has to outlive the session.

What is the difference between an AI agent runtime and an agent framework? +

A framework such as LangGraph, CrewAI or the OpenAI Agents SDK runs the loop: it calls the model, parses tool calls and feeds results back. The runtime is where those tool calls execute. You need both. PandaStack is the runtime; you register it in your framework as a tool whose body runs code in a sandbox and returns the output to the model.

Why run agent tool calls in a microVM instead of a container? +

A container shares the host kernel with every other container on the machine, so model-written code talks to that kernel's full syscall interface. A Firecracker microVM gives each session its own Linux kernel behind KVM hardware virtualization, with a minimal virtio device model as the host-facing surface. PandaStack restores each VM from a snapshot instead of booting it, so a sandbox is ready in 49ms p50 on the live API.

Can the agent's sandbox reach the internet? +

Yes. Outbound access is open by default because agents install packages, call APIs and clone repositories, and there is no outbound allowlist today. Traffic to other VMs, to the cloud metadata service and to common crypto-mining pool ports is dropped at the host firewall, and every connection is attributable to one VM. Don't put a secret in a sandbox unless you are fine with the model being able to send it somewhere.

How long can an agent session run? +

On Free, a sandbox is deleted one hour after it is created. Pro allows TTLs up to 7 days and Team up to 30 days, and on paid plans the TTL is an idle timeout, so a sandbox that is doing work keeps running. Set ttl_seconds at create time so a crashed agent loop can't leak a VM.

What does an idle agent cost? +

Sandboxes bill $0.054 per active vCPU-hour and $0.0162 per working-set GiB-hour, metered per second. A sandbox waiting on the model burns almost no CPU, so it bills mostly for the memory it keeps resident. Hibernate it and compute billing stops; the next request to the sandbox wakes it with its files and memory intact.

Can I self-host the agent runtime? +

Yes. The core is Apache-2.0. A self-hosted deployment is the control-plane API with Postgres plus one or more agents on KVM-capable Linux hosts, and the same SDKs and CLI point at your API endpoint through PANDASTACK_API.

Ship on the millisecond cloud.

Free tier with $5.40/mo usage credit. No card. Apache-2.0.