all posts

AgentCore's Code Interpreter, or a Sandbox You Hold?

Ajay Kumar··10 min read

I build PandaStack, an open-source Firecracker microVM platform, so I get a version of this question a few times a month now. We are on AWS. Bedrock AgentCore ships a Code Interpreter tool. Why would we run a sandbox ourselves?

The usual answer to that question is a latency table, and the usual answer is bad. Nobody picks their agent's code-execution tool on boot time, and anyone who does will regret the basis for the decision about four months in, when the thing they actually needed turns out to be a session that outlives one request.

The real question is how much of your agent stack you want to buy as one unit. AgentCore is a very large unit that happens to contain a code interpreter. A standalone sandbox is one primitive that happens to be the only piece some teams need. Those are different purchases, and which one is right has almost nothing to do with milliseconds.

Disclosure and method. I sell one of the two options, so discount accordingly. I keep it honest the only way that works: every claim about AgentCore below comes from reading the current AWS developer guide on 2026-10-09, and is deliberately qualitative — no latency, throughput or cost figures for AWS, because I have not measured them and would not publish a competitor's numbers if I had. AgentCore is evolving quickly; its own component list grew while I was reading it. Verify anything load-bearing against the AWS documentation yourself, because some of this will have changed by the time you read it. The only numbers in this post are PandaStack's.

What a code-execution tool actually has to do

Strip the marketing off both options and the job has four parts. Accept code the model wrote. Run it. Survive it. Return something structured enough that the model can reason about what happened.

Three of those are easy. The third is the entire product.

The code arriving at your tool was written by a text predictor that has no idea what else is on the machine. It will shell out. It will `pip install` something at version `*`. It will allocate forty gigabytes to sort a CSV. It will fork a child, detach it, and leave it running after your timeout fires on the parent. And if the task you handed it was "clean up the workspace", it will eventually write the command you are afraid of. None of that requires an adversary. Incompetence and malice emit identical syscalls, and your isolation boundary cannot tell them apart either.

So surviving it means three concrete things: a boundary that holds when the code is wrong, a teardown that happens even when the code hangs, and a blast radius of exactly one disposable machine.

On that much, AWS and I agree, and it is worth saying plainly before the comparison starts. The AgentCore documentation describes per-session isolation as a dedicated microVM with isolated CPU, memory and filesystem, fully terminated with its memory sanitised when the session completes. That is the same architectural answer I give, for the same reasons. The isolation argument between a hyperscaler's agent sandbox and a Firecracker-based one is largely settled — see Firecracker vs AWS Lambda: Same Engine, Different Boundary and PandaStack vs AWS Lambda for Running Untrusted / LLM-Generated Code for why. Which is precisely why the interesting differences are somewhere else.

What AgentCore actually is, from its own docs

AgentCore is not a sandbox product with some extras bolted on. AWS describes it as an agentic platform for building, deploying and operating agents using any framework and foundation model, made of modular services you can use together or independently. The services documented today include a managed agent loop (Harness), a serverless agent host (Runtime), Memory, Gateway, Identity, Observability, Evaluations, Optimization, Policy, a Registry, Payments, and two built-in tools — Browser and Code Interpreter.

That list is the point. Code execution is one tile in a mosaic, and the mosaic is the thing being sold. Framework and model agnosticism are documented explicitly: Runtime is described as working with CrewAI, LangGraph, LlamaIndex, Strands and custom loops, with models inside or outside Bedrock, over MCP or A2A.

The Code Interpreter tile itself, as documented on the day I read it:

  • You create a Code Interpreter resource in your account — or use the system one — choosing a network mode and an IAM execution role that defines which AWS resources the running code can reach.
  • Network modes are Sandbox (limited external network access), Public network (public internet), or VPC, which is required if you want to mount your own storage.
  • Execution is session-based. You start a session with a configurable timeout, documented as 900 seconds by default and extendable up to eight hours, after which it terminates automatically.
  • Within a session, state is maintained across executions, and many sessions can be active at once against one Code Interpreter, each with its own state and environment.
  • Supported languages are documented as Python, JavaScript and TypeScript, with common libraries pre-installed.
  • Files created in a session live for the session and are cleaned up when it ends. Inline upload is documented up to 100 MB; uploading to S3 through terminal commands up to 5 GB.
  • For durable storage you bring your own: S3 Files or EFS access points mounted under a `/mnt/<name>` path, which requires VPC network mode, mount targets in matching availability zones, NFS on port 2049 and the right IAM. The docs state plainly that, unlike Runtime, Code Interpreter offers no managed session-storage option.
  • Everything is driven through IAM verbs — `CreateCodeInterpreter`, `StartCodeInterpreterSession`, `InvokeCodeInterpreter`, `StopCodeInterpreterSession` — with CloudTrail logging and CloudWatch metrics.

Read that list as an architecture rather than a feature inventory and its shape is clear. You are not handed a machine. You are handed a resource in your account, from which you start sessions, whose reach is defined by an IAM role and a network mode, and whose contents evaporate on termination. That is a very AWS-shaped object, and if your organisation is AWS-shaped, that is a feature.

The decision is bundling, not benchmarking

Here is where I actually send people, and it is not a vendor's hedge.

Buy the bundle when you are already deep in AWS and the gravity is real. Your data is in S3 and EFS, and the bundle mounts them rather than making you invent a transfer layer. Your compliance boundary is an AWS account, so a tool that executes code inside that account with an IAM role attached to it is a paragraph in a review rather than a project. Your auth is Cognito or Entra via AWS, your traces go to CloudWatch, your procurement is done, and when something breaks at 3am you want one phone number. Every one of those is a genuine, quantifiable saving, and a sandbox vendor who tells you otherwise is selling you something.

Buy the primitive when code execution is the only piece you actually need. Your agent loop already exists — in a Go service, a Next.js route handler, a Temporal worker — and it does not want to be redeployed into somebody's runtime to get a `run_code` tool. You are routing across model providers this quarter and would rather the execution substrate not move when the model does. You are multi-cloud, or deliberately not on a hyperscaler. Or, most concretely, you need primitives a managed tool does not expose, which is the next section.

There is a third answer that people miss, and the AWS docs support it: use both. The services are documented as usable independently, so running the loop on AgentCore Runtime with Memory while pointing the code-execution tool at a sandbox you control is a supported shape, not a hack. The AI Agent Runtime Infrastructure: What It Is, How to Choose, How to Self-Host post draws that line in more detail — where the loop runs and where the tool calls run are separate purchases.

The lock-in question, stated in both directions

The lazy version of this argument is that bundles are traps. The lazy version is wrong, and if you want the full taxonomy I wrote it up at Vendor Lock-In in Code Execution Infrastructure. The short version here is that integration is real value and you should price it as such. Identity that already knows your IdP, a tool gateway that already speaks MCP, traces that already land where your dashboards are, a memory store you did not have to design — that is six months of platform work you are buying, and the alternative is not freedom, it is a backlog.

The honest cost is not an API padlock. AWS documents these services as independently usable and the Code Interpreter is callable from third-party frameworks, so there is no export button being withheld from you. The cost is coupling of a different kind: your tool-execution layer inherits your model-hosting layer's release cycle, its region map, and its IAM model. When the platform ships a change to session semantics, that is now a change to where your untrusted code runs. When a feature you want lands in one region first, your sandbox's region list is somebody else's roadmap. When your security review asks what the code could reach, the answer is an execution role, which is a good answer right up until the day you want that answer to be "nothing at all".

That is a trade, not a trap. It is just a trade worth making consciously, because the people who regret it are the ones who never noticed they made it.

What a first-class sandbox handle buys you

The thing a standalone sandbox gives you that a managed tool generally does not is a handle. Not a session you invoke, but an object with an id, a lifetime you set, a filesystem you address by path, and operations you can perform on it that have nothing to do with running code.

On PandaStack every create is a snapshot restore — there is no warm pool of idle VMs waiting around — which lands at p50 179 ms and p99 203 ms end to end. The first spawn of a template is a cold boot that bakes the snapshot, roughly 3 seconds, once. What that latency buys is the licence to treat a VM as a cheap object rather than a scarce resource, which is what makes the rest of this list usable:

  • `sbx.snapshot()` returns an id you can create from later. The session your agent built — packages installed, data loaded, a half-finished analysis — becomes a thing you can put down and pick up tomorrow. Snapshots outlive their sandbox on purpose: an idle auto-reap deliberately does not cascade-delete them.
  • `sbx.fork()` clones the disk and the child cold-boots from it: 400–750 ms on the same host, 1.2–3.5 s cross-host. It does not inherit the parent's memory, so the child gets its own entropy, which matters more than people expect.
  • `sbx.fork_tree(n)` is the one that inherits memory from the parent's snapshot, capped at 16 children per call. This is branch-and-explore: park the agent at a decision point, fan out eight candidate plans into eight identical children, score the outcomes, keep the winner, reap the rest. See Snapshots and Forks: Copy-on-Write for Running Machines for the mechanics.
  • `sbx.filesystem.read`, `.write`, `.listdir`, `.walk` and `.exists` address the guest's own rootfs by path. No mount, no access point, no availability-zone alignment, no port 2049.
  • `ttl_seconds` on create is an idle clock, not a wall clock. The reaper deletes a sandbox once it has gone that long without activity, so a continuously busy one is never cut off mid-task — it is the backstop for the VM your code forgot to kill, not a cap on useful work. `persistent=True` makes the reaper skip the sandbox altogether.
  • `sbx.preview_url(port)` gives a tokenless `https://<port>-<id>.<suffix>` address, valid for the sandbox's lifetime, where the UUID is the credential. When the model writes a Streamlit app or a chart server, it is reachable without you building an ingress.
  • `sbx.create_code_context()` is a persistent Jupyter-style kernel: variables and imports survive between calls, and rich outputs come back as objects rather than files you have to go and find.

A note on what is not on that list, because accuracy matters more than the pitch. Nothing here is unique in kind — the AgentCore Code Interpreter has in-session file operations too, and its BYO mounts solve a durability problem that my `filesystem` API does not solve at all. The difference is the shape of the object: snapshot and fork of a live session are primitives I expose and the documentation I read describes no equivalent for a Code Interpreter session, where a terminated session is cleaned up and gone. If you are reading this later, check — that is exactly the kind of thing that changes.

Defining the tool

Here is the whole thing: a function the model calls, which creates a VM, writes the code, runs it, reads a result file back, and tears the VM down unconditionally. Register it with whatever framework you use — the sandbox does not know or care which one that is.

# tools.py - one code-execution tool, no framework assumed.
from pandastack import Sandbox
from pandastack.exceptions import APIConnectionError, SandboxTimeout

PREAMBLE = (
    "# Write your answer as JSON to /workspace/result.json.\n"
    "# Anything on stdout is shown to you; the VM is destroyed afterwards.\n"
)

def run_python(code: str) -> dict:
    """Run model-written Python in a throwaway microVM and return its result."""
    sbx = Sandbox.create(
        template="code-interpreter",
        ttl_seconds=900,                      # IDLE backstop: reaped if abandoned
        metadata={"tool": "run_python"},
    )
    try:
        sbx.filesystem.write("/workspace/cell.py", PREAMBLE + code)

        # The REAL bound is in-guest: SIGTERM at 20s, SIGKILL 5s later.
        # 124 (or 137 if it had to be killed) means the clock fired --
        # that is a result to report, not an exception to raise.
        r = sbx.exec(
            "cd /workspace && timeout --kill-after=5s 20 python3 cell.py",
            timeout_seconds=30,
        )

        result = None
        if sbx.filesystem.exists("/workspace/result.json"):
            result = sbx.filesystem.read("/workspace/result.json").decode("utf-8")

        return {
            "ok": r.exit_code == 0,
            "timed_out": r.exit_code in (124, 137),
            "stdout": r.stdout[-4000:],        # keep the context window honest
            "stderr": r.stderr[-2000:],
            "result": result,
        }
    except (SandboxTimeout, APIConnectionError) as e:
        return {"ok": False, "error": str(e)}
    finally:
        sbx.kill()                             # teardown is not optional

Three details in there are load-bearing. The in-guest `timeout` is the real wall clock, not the client's timeout argument — put the bound inside the guest and the hung process dies in the guest, where it cannot take your request handler with it. `ttl_seconds` is an idle clock rather than a wall clock, which is exactly the right shape for a backstop: on the day your process is OOM-killed between `create` and `finally`, the orphan stops being touched, goes idle and gets reaped — and meanwhile a sandbox that is genuinely busy is never cut off for running long. And truncating stdout before it reaches the model is not cosmetic: one `print(df)` on a wide frame will eat your context window and the model will then reason about the half it can see.

For a multi-turn analysis you want state to survive between tool calls, which is the other shape. One sandbox per conversation, one kernel inside it, torn down when the conversation ends:

from pandastack import Sandbox

# One sandbox + one kernel per conversation. Store sbx.id in your session row.
sbx = Sandbox.create(template="code-interpreter", ttl_seconds=3600)
ctx = sbx.create_code_context()          # Jupyter-style kernel

try:
    sbx.filesystem.upload("./sales.csv", "/workspace/sales.csv")

    # Turn 1: load. State persists across run_code calls.
    ctx.run_code("import pandas as pd; df = pd.read_csv('/workspace/sales.csv')")

    # Turn 2: `df` is still there. The model never re-sends the setup.
    ex = ctx.run_code("df.groupby('region').revenue.sum()")
    print(ex.text)                       # plain-text repr of the last expression
    print(ex.logs["stderr"])             # captured stderr, if any

    # Turn 3: a chart comes back as an object, not a file you hunt for.
    ex = ctx.run_code(
        "import matplotlib.pyplot as plt\n"
        "df.groupby('region').revenue.sum().plot(kind='bar')\n"
        "plt.show()"
    )
    png_b64 = ex.results[0].png          # hand straight to a vision-capable model
finally:
    ctx.close()
    sbx.kill()

That is the piece that changes how an agent feels. Without it, every turn re-sends the imports and re-reads the CSV, and the model spends its tokens rebuilding a world it already built. With it, `df` is just there.

Side by side

AgentCore Code Interpreter compared with a standalone microVM sandbox. AWS cells are qualitative, from the developer guide as of 2026-10-09.
AgentCore Code InterpreterStandalone microVM sandbox (PandaStack)
What you buyA managed tool inside a larger agent platform: runtime, memory, gateway, identity, policy, observability and more sold alongside itOne primitive. A sandbox you create, address and destroy. Nothing else is implied or included.
Isolation boundaryDocumented as a dedicated microVM per session with isolated CPU, memory and filesystem, terminated and memory-sanitised at session endFirecracker microVM, own guest kernel, own network namespace; destroyed on kill(), or once the idle TTL expires
Session lifetimeSession-based with a configurable timeout; documented default 900 seconds, extendable up to eight hours, then automatic terminationttl_seconds is an idle clock, not a wall clock — a busy sandbox is never reaped by it; persistent=True skips the reaper entirely; free tier adds a hard max lifetime
Snapshot / forkNo snapshot or fork primitive for a Code Interpreter session in the docs I read; a terminated session's data is cleaned upsnapshot() returns an id; fork() clones disk and cold-boots the child; fork_tree(n) inherits memory, capped at 16 per call
Filesystem accessFile upload and download in-session, plus bring-your-own S3 Files or EFS access points mounted under /mnt, which require VPC network modefilesystem.read / write / listdir / walk on the guest's own rootfs by path, with no mount setup; durable volumes are separate
Framework couplingDocumented as usable independently and from third-party frameworks; the real coupling is the account, the execution role and the region mapHTTP API plus Python and TypeScript SDKs; the sandbox has no opinion about what called it
Where it runsAWS-managed in supported AWS Regions, inside your account's IAM and network boundaryHosted multi-tenant, or self-hosted on your own KVM hosts — the agent is Apache-2.0

Egress, and the metadata endpoint specifically

This is the part of any sandbox comparison where vendors get vague, so here is mine stated precisely enough to argue with. I will not characterise AgentCore's network posture beyond what I quoted above — it documents Sandbox, Public and VPC modes, and you should read that page yourself rather than take my summary of it.

On PandaStack, egress is open by default. There is no default-deny policy. A sandbox can reach the internet, which is what makes `pip install` work and is also exactly the exfiltration channel you would expect it to be. What the root FORWARD chain does drop, before any egress rule is evaluated, is pool-to-pool traffic — no sandbox can reach another sandbox's subnet — and the whole of `169.254.0.0/16`.

That second rule is the one worth understanding. The link-local range is where cloud metadata lives, and on GCP that endpoint hands out the host VM's service-account token. A tenant that could `curl` it would hold a credential reading every other tenant's snapshots. So it is dropped at the host, for the entire range, as the first FORWARD rule, and there is a test asserting the rule's exact form. Cloud metadata is not reachable from a PandaStack sandbox.

What is not fenced is your own network. Your VPC, your databases, your internal APIs that trust anything originating from inside the perimeter — if a sandbox can route to them, it can reach them. That is the customer's remaining work, and the honest contrast with the bundle is that AWS's execution-role model gives you a native vocabulary for that question while mine gives you firewall rules to write. Controlling Network Egress for Untrusted Code is the long version.

What moving a code-execution tool off a bundle actually involves

Less than people fear, and in a different place than they expect. The work is not the exec call. It is the five things the bundle was quietly doing around it.

  1. Inventory what you use, not what you bought. Write down every bundle feature your tool path actually touches. Teams are routinely surprised to find the answer is "start a session, run code, read a file" — in which case the migration is an afternoon. If the answer includes memory, a gateway and identity, you are not migrating a tool, you are replatforming, and you should decide that deliberately.
  2. The tool contract is the seam, so make it one. A function with a code string in and a structured dict out is portable by construction; the same function body can call a hosted tool or your own sandbox behind one `if`. If your framework registration, your retry policy and your exec call are the same forty lines, fix that first and the migration becomes a config change.
  3. Re-plumb data in and results out. A managed tool gives you inline upload with a size cap and, if you set up VPC mode, mounts. A sandbox gives you path writes and uploads. Neither is harder; they are just different, and the code that bridges them is where the real diff lands. Tar directories — single-file upload is single-file.
  4. Map session identity into your own state. The bundle's session id was probably doing double duty as your conversation key. Now you store a sandbox id on your session row yourself, and you own the question of what happens when a user returns after the TTL expired. Answer it explicitly: re-create from a snapshot, or start clean and say so.
  5. Pick up what the bundle was carrying. Idle teardown, per-tenant quotas, egress policy, secret injection, OTel wiring, and the alerting for when a sandbox leaks. None of it is hard. All of it is yours now. Budget a sprint, not a quarter, and do not pretend it is free.

The thing that makes this tractable is that you do not have to migrate the whole stack to move the execution tile. Keep Memory. Keep the Gateway. Keep the loop wherever it is happy. Move the one tile where you need a primitive the managed version does not expose, and leave the rest alone.

Honest limits

PandaStack is not an agent platform and I am not going to pretend the comparison is symmetric. There is no managed memory service, no identity broker, no tool gateway, no policy engine, no evaluation harness, no managed observability bundle. If you want those, you are assembling them from other vendors and your own code, and the integration bill is real. AgentCore's pitch is that it eliminates exactly that integration work, and for a lot of teams that pitch is simply correct.

Two more that cost me something specific. First: if your organisation's security review is an AWS-account-shaped hole, a non-AWS vendor is procurement work — a DPA, a questionnaire, an architecture diagram that now has an arrow leaving your cloud. I cannot make that cheaper by writing a blog post. Second, and I would rather you hear it here: PandaStack does not encrypt sandbox rootfs images, snapshots, guest memory files or durable volumes at the application layer. There is no dm-crypt, LUKS or fscrypt in the platform. What exists is provider-level encryption at rest — GCS and GCP Persistent Disk encrypt by default with Google-managed keys, and the AWS Terraform module turns on S3 server-side encryption — with no customer-managed KMS key wired up. Application secrets for hosted apps are AES-256-GCM encrypted with the owner and key name bound in as authenticated data, but that is a different subsystem from your sandbox's disk. If your threat model or your auditor requires CMEK over the execution environment itself, that is a real gap today, and "the cloud provider encrypts the disk" is the whole of my answer.

And one operational truth: a sandbox size comes from the template's baked snapshot, because Firecracker cannot change vCPU or RAM at restore. The `code-interpreter` template is 2 GiB and 8 burst vCPU; passing `cpu` or `memory_mb` on create is overridden to match the snapshot. Managed services that provision dynamically do not make you think about that. I do.

The bottom line

If your data, your identity, your compliance boundary and your on-call rotation all already live in one AWS account, buy the bundle. The integration you are paying for is real, the isolation story is sound, and the alternative is a platform team you have not hired. Do it with your eyes open about the coupling, and do not let the execution tile be the thing you never examined.

If code execution is the only piece you need, or your loop lives somewhere else, or you want one substrate underneath whichever model wins this quarter, or you need a session you can snapshot, fork sixteen ways and pick back up tomorrow — then you want a sandbox you hold a handle to, and the bundle is a lot of surface area to adopt for one tile. On pricing that is $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same rate for every workload class, CPU billed on seconds actually burned, and no per-request charge, now or ever.

Either way: go read the AWS docs before you commit, not my summary of them. That is not false modesty. It is the one claim in this post I am certain will still be true next quarter.

Frequently asked questions

Is AgentCore Code Interpreter's isolation weaker than a dedicated microVM sandbox?

No, and anyone selling you a sandbox on that basis is overreaching. The AgentCore documentation describes each tool session as running in a dedicated microVM with isolated CPU, memory and filesystem resources, fully terminated with its memory sanitised when the session completes — which is the same architectural answer a standalone microVM platform gives, for the same reason. Hardware-virtualisation boundaries around untrusted code is a settled argument and the hyperscalers were early to it, not late. The differences that remain are about what is exposed rather than how strong the wall is: whether you can snapshot or fork a live session, how long a session may live, how you get data in and results out, whether the thing has an identity in your cloud's IAM, and who owns the egress policy. Pick on those. If you want the isolation reasoning in full, the Firecracker-versus-Lambda comparison on this blog walks through why both camps converged on the same primitive. And verify the session-isolation language against the current AWS documentation yourself — it is the sort of claim that gets refined as a product matures.

Can I run my agent loop on AgentCore Runtime but execute code in a non-AWS sandbox?

Yes, and the AWS documentation supports it rather than merely tolerating it: the AgentCore services are described as modular and usable together or independently, and Runtime is documented as framework-agnostic with support for custom loops. A code-execution tool is just a function your agent calls, so nothing stops that function from talking to a sandbox API over HTTPS instead of calling InvokeCodeInterpreter. In practice you want two things from that arrangement. First, keep the tool contract narrow — a code string in, a structured dict out — so the choice of backend is one branch rather than a shape that leaks into your prompts. Second, think about where the credential for the external sandbox lives and how it is rotated, because you have added an outbound dependency to a runtime whose other dependencies are all in-account. This split is also the common case in reverse: plenty of teams keep the loop in their own service and use nothing from any bundle except the execution tile.

How long can a code-execution session actually live, and what happens when it ends?

For AgentCore Code Interpreter, the documentation states a session timeout of 900 seconds by default, configurable up to eight hours, with automatic termination after the timeout and the session's files and data cleaned up at that point. Durable storage is explicitly bring-your-own — S3 Files or EFS access points mounted into the session, which require VPC network mode. On PandaStack the clock works differently in a way worth being precise about: ttl_seconds is an idle TTL, not a wall clock. The reaper deletes a sandbox once it has gone that long without activity, so a continuously busy sandbox is never cut off by it, and an abandoned one is cleaned up without you having to track it. Set persistent=True and the reaper skips the sandbox entirely; free-tier accounts additionally carry a hard maximum lifetime that does not care about activity. The more useful difference, though, is what you can do at the end rather than how long the clock runs: snapshot() hands back an id you can create a fresh sandbox from later, and an idle auto-reap deliberately does not cascade-delete those snapshots, so a long-running analysis does not have to be a long-running VM. The pattern I recommend for agents is a short TTL plus a snapshot on the way out, which gets you resumable sessions without paying for idle memory between turns. Check the AWS timeout figures against the current docs before you design around them.

What am I actually giving up by not buying an agent bundle?

Concretely: a managed memory store with short-term and long-term semantics you did not have to design; a gateway that turns your existing APIs and Lambda functions into MCP tools; an identity service that already speaks to Cognito, Okta and Entra for both inbound user auth and outbound third-party credentials; tracing built for agent workflows rather than HTTP requests; and increasingly an evaluation and policy layer alongside them. That is a genuine platform team's worth of work, and PandaStack supplies none of it — it supplies sandboxes, managed Postgres, app hosting and functions, and expects you to bring the agent framework. The counter-question is whether you want those from the same vendor as your code execution. Memory and identity are long-lived, organisation-shaped concerns; a code-execution sandbox is a disposable, high-churn one. Coupling them means the disposable thing inherits the long-lived thing's release cycle and region map. Some teams want that single throat to choke. Others want the churny tile to be independently replaceable. Both are defensible; only one of them is usually examined.

Does a sandbox stop model-written code from reaching my internal database?

Not by itself, and this is the question I most often see answered dishonestly. On PandaStack, egress is open by default — there is no default-deny policy — because a code interpreter that cannot pip install is not a code interpreter. Two high-value targets are closed at the host, in the root namespace's FORWARD chain, before any egress rule is evaluated: pool-to-pool traffic is dropped so no sandbox can reach another sandbox's subnet, and the whole of 169.254.0.0/16 is dropped, so cloud metadata — the classic path from arbitrary code execution to a host service-account token — is unreachable. What is not closed is your own network. If the sandbox can route to your VPC, your Postgres or your internal API that trusts anything inside the perimeter, model-written code can reach it. So add filter rules for your own ranges and peerings, give the sandbox no credentials it does not need for this one task, and prefer handing it data over handing it a connection string. The bundle's answer to the same question is an IAM execution role, which is a genuinely nicer vocabulary for in-cloud resources and tells you nothing about the public internet.

Keep reading

Related posts

More in AI agent sandboxes · See AI agent sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.