Top 13 Ephemeral Development Environment Platforms in 2026: Who Is Allowed to Call Create
Every product in this category was designed around the same mental image: a person sits down, opens a laptop, clicks a button, and gets a machine with their repository already on it. Almost every feature grid still assumes that person is the customer. Browser IDE. Dotfiles. Port forwarding wired into the editor. An onboarding flow for the new hire who starts Monday.
That person still exists and still matters. They are no longer the fastest-growing caller. The thing creating environments in 2026 is increasingly a pull-request webhook, a CI matrix job, a nightly fan-out, or an agent that has decided it needs four checkouts of the same repository to compare three refactors. None of them have hands. None of them can click "rebuild prebuild." None of them will notice a warm-path miss, shrug, and go get coffee — they will time out, retry, and page someone.
So I want to grade this category on one narrow, slightly mean question: who is allowed to call create, and what happens when the caller is a program rather than a person? It is a better question than it sounds, because it decomposes into four things you can answer in an afternoon with curl and a shell loop, and because it sorts the platforms differently than any feature grid does. Two products with near-identical feature lists land in opposite cohorts the moment you ask whether their create endpoint is a product or a remote control for a UI.
I build PandaStack, an open-source Firecracker microVM platform, so I have an obvious interest in the answer. I have put my own product in the list as entry thirteen and said plainly what it does not do. There are at least three platforms below I would recommend over mine for specific, common cases, and I have named which cases.
The caller changed and most of the category did not
Nearly every platform here has an API. That is not the distinction. The distinction is whether anything in the design anticipated that you would call it fifty times in a row from a process with no human attached, and that shows up in small, consistent tells once you know to look for them.
The credential is the first tell. If the only way to authenticate is a token tied to a named user's identity, with scopes defined in terms of what that person may see, then the platform's model of a caller is a person, and your CI job is going to end up holding somebody's personal token in a secret store — which is both a security finding and an on-call incident waiting for that person to leave.
The second tell is what create returns and when. A create call that returns the moment a database row exists, with a status field you poll until the environment is maybe reachable, is a different contract from one that returns a thing you can immediately exec into. Both are legitimate designs. Only one of them lets you write a synchronous script without inventing your own readiness loop, and the readiness loop is where everybody's retry logic goes wrong.
The third tell is concurrency. Ask a vendor what happens at fifty concurrent creates. There are exactly three acceptable answers — it works, it queues, it refuses with a documented error — and the common fourth answer, which is a shrug followed by an undocumented limit you discover in production, tells you the fan-out case was never on anybody's roadmap.
An API nobody expected you to call in a loop is a user interface with a URL.
And then there is the part that changed hardest. When the caller is a program, the code inside the environment frequently was not written by anyone you employ. It came from a fork's pull request, from a dependency's post-install hook, or from a model that generated it ninety seconds ago. A model-generated `rm -rf` is not malice, it is a plausible token sequence — and on a shared host kernel the difference does not matter much. That moves the isolation boundary from a compliance checkbox to a load-bearing property of your fan-out design.
The four questions, and the script that answers them
- Is create a first-class API, or a remote control for a UI? The test: can a program holding a machine credential go from nothing to a running, exec-able environment in one HTTP call, without a browser ever being involved? Most platforms pass a weaker version of this (there is an endpoint) and fail the strong version (the endpoint is how the UI drives the product, and the docs are written for the UI).
- What is the isolation boundary when the author of the code is a model or a fork? Container on a shared kernel, pod on a shared node, WASM inside a browser origin, or a VM with its own kernel. This is the one axis you cannot change later, and when the caller is a program it decides whether "run the tests from this pull request" is a routine operation or a trust decision you make fifty times an hour.
- What happens when a program asks for fifty at once? Where does the ceiling live — the vendor's scheduler, your cloud quota, a Kubernetes node pool, or your own laptop? Fan-out failures are rarely a hard error. They are usually the twelfth create taking minutes while the first eleven took seconds, and a feature grid has no column for that.
- Is there a cold path at all, or only a warm path you can miss? If the fast case depends on a prebuild, an image cache, or a warm pool, then the platform has two latencies and the automated caller experiences both. A dependency bump that invalidates every cache you had is a Monday, not a hypothetical, and it is the day you find out which latency your pipeline's timeout was written against.
Here is the script. It is deliberately boring, it takes about ten minutes to adapt per vendor, and it has settled more evaluations for me than any amount of reading docs. Run it against each shortlisted platform with a non-human credential and a fresh account.
#!/usr/bin/env bash
# The whole evaluation, in one script. Run the equivalent against every
# vendor on your shortlist. What this measures is not speed -- it is whether
# a PROGRAM, holding a machine credential, can get to a running environment
# with no human, no browser tab and no prebuild warming anything first.
set -euo pipefail
API="${PANDASTACK_API_URL:?point this at your platform's API base}"
KEY="${PANDASTACK_API_KEY:?use a MACHINE credential, not your own login}"
N=25
ids=()
# Teardown FIRST. Write the trap before you write the create loop: the ids
# array fills as we go, so a Ctrl-C at environment 11 still reaps the 10 that
# already exist. Every forgotten environment I have ever been shown was born
# because this block was written last, or written inside an `except`.
cleanup() {
for id in "${ids[@]:-}"; do
curl -sS -X DELETE "$API/v1/sandboxes/$id" \
-H "Authorization: Bearer $KEY" \
-o /dev/null -w "delete $id -> %{http_code}\n" || true
done
}
trap cleanup EXIT
start=$(date +%s)
for i in $(seq 1 "$N"); do
# One POST. No OAuth dance tied to a human, no devcontainer image build,
# no "your workspace is starting" page to poll. If a platform cannot do
# this line, nothing further in your evaluation matters.
#
# ttl_seconds is an IDLE clock, not a wall clock. It is not a lifetime cap;
# it is the backstop for the case where this script is SIGKILLed and the
# trap above never runs at all.
id=$(curl -sS -X POST "$API/v1/sandboxes" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d "{\"template\":\"base\",\"ttl_seconds\":900,\"metadata\":{\"run\":\"caller-eval\",\"n\":\"$i\"}}" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["id"])')
ids+=("$id")
# The interesting number is not any single one of these. It is the SPREAD
# across 25 -- that is where a warm path you missed shows up.
echo "created $id at +$(( $(date +%s) - start ))s"
done
# Did you get usable environments, or rows in someone's database? Exec is the
# proof, and it is the step that fails on platforms whose create call returns
# before the machine is actually reachable.
#
# Bound the command IN-GUEST with timeout(1): an HTTP timeout parameter kills
# your request, not the process on the far side of it.
for id in "${ids[@]}"; do
curl -sS -X POST "$API/v1/sandboxes/$id/exec" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"cmd":"timeout --kill-after=5s 20 sh -c \"uname -r; nproc\""}' \
| python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["exit_code"], " ".join(d["stdout"].split()))'
done
# Nothing here cleans up explicitly. The trap does it, including on failure.
Three things this script surfaces that nothing in a sales conversation will. First, whether a machine credential exists at all, or whether you are about to put a human's token in CI. Second, whether create returning 2xx means the environment is usable — the exec loop at the bottom fails loudly on platforms where it does not. Third, the spread. Twenty-five creates whose durations cluster tightly tell you there is one path. Twenty-five that split into a fast group and a slow group tell you there is a warm path, and that you will be on the wrong side of it eventually.
Cohort one: a person is the intended caller, and the API knows it
These four are workspace products. They are good at being workspace products, and I want to be clear that "the API is a remote control for a UI" is a description, not an insult — if your caller is a human, that is the correct architecture and the extra abstraction a program-shaped API would impose is pure cost.
1. GitHub Codespaces
Underneath: a devcontainer-defined Linux container on managed hosts, sharing the host kernel. There is a genuine REST API and a CLI, and automation does use them. But the API's unit of work is "a codespace for a user," the fast path is a prebuild baked by a GitHub Actions workflow, and the gravitational centre of the product is the editor. The real value is that it lives where your code already lives, which is not a technical property at all.
Honest drawback for a program-shaped caller: the cold path is an image build and the warm path is a prebuild you can miss, which is the worst possible shape for an automated caller because your pipeline's timeout gets written against the good mode. Concurrency is governed by org policy and quota rather than by a documented per-second limit. And a stopped codespace is not a deleted one — compute stops, the disk keeps existing and keeps billing until a retention policy notices. State survival is genuinely good here, which is exactly why the abandoned-workspace bill is the thing to watch.
2. Gitpod / Ona
Underneath: containerised workspaces from a repo-committed definition. The important fact is the identity shift — the newer generation, now carrying the Ona name, runs runners inside your own cloud account rather than on the vendor's hosted platform, with an increasingly agent-shaped product on top. That direction is the thesis of this post made visible: a workspace company followed the caller.
Honest drawback: moving the runners into your account moves the concurrency ceiling there too, so "fifty at once" becomes a question about your node pool and your quota rather than theirs. And the positioning has moved twice in eighteen months, which means tutorials, blog posts and your own evaluation notes all go stale at different rates. Check which generation a document is describing before you follow it.
3. Replit
Underneath: container environments with a declarative package layer, inside a product that has spent two years becoming an AI app-builder with a dev environment in it rather than a dev environment with AI features. The caller it assumes is increasingly a model — but the model is theirs, and it runs inside their product. The browser-first path from nothing to running code remains among the best in this list, and handing a non-engineer a working environment with no local setup is a genuinely strong case.
Honest drawback: it is the cleanest example in the category of an API that is a remote control for its own UI. Excellent for a human or for Replit's agent; awkward as a substrate for your agent. The isolation boundary is a shared kernel, so if your requirement is running code you did not write under a boundary you can describe to a security reviewer, this is not the shelf — and that is not a criticism of the product.
4. Google Cloud Workstations
Underneath: a managed service that runs your container image on a Compute Engine VM, one VM per workstation, with a persistent home disk and a configurable idle timeout. It is the most programmable entry in this cohort almost by accident: it is a GCP service, so it has a real API, a Terraform provider, IAM, org policies and audit logs, because everything in GCP does. The isolation boundary is also the strongest in the cohort, since a workstation is a VM.
Honest drawback: programmable is not the same as fan-out-friendly. Create is VM-shaped, cost is VM-shaped and linear, and the ceiling at fifty is regional CPU quota — a hard limit that lives in a different console from the product and is nobody's job to raise until it blocks you. If your pattern is many short-lived environments, this is the cohort member that punishes it hardest. If your pattern is twenty long-lived workstations inside a VPC your auditors already approved, it is a serious answer.
Cohort two: you are the API
These three hand you a control plane and leave the substrate to you. The honest framing is that you are not buying environments, you are buying a scheduler and a UX — and therefore the quality of your answer to all four questions is the quality of a template somebody wrote in a sprint eighteen months ago.
5. Coder
Underneath: a self-hosted control plane that provisions workspaces through Terraform, so the substrate is whatever you wrote — Kubernetes pods, cloud VMs, your own hardware. The API and CLI are first-class and the product is genuinely designed to be driven programmatically; enterprises adopt it precisely because a workspace can be a VM in a subnet their auditors already signed off.
Honest drawback: the isolation boundary and the concurrency ceiling are both properties of your template rather than of the product, so neither question has an answer the vendor can give you. "Fifty at once" means fifty Terraform applies at once, which is a sentence that has ruined afternoons. And if your template provisions pods on a shared node, then model-authored code runs on a shared kernel and the control plane cannot help with that.
6. DevPod
Underneath: a client-side tool, not a service. It reads the devcontainer spec and starts the environment through a provider — local Docker, SSH to a box, Kubernetes, a cloud VM — with no vendor control plane in the middle. Open source, and the architectural choice is the interesting part: there is no central thing to run, so there is no central thing to pay for and nothing to rate-limit you.
Honest drawback for a program: the caller has to be somewhere. A CI job can absolutely shell out to the CLI and it works fine, but you have just made your CI runner the control plane — environment state lives on whichever machine ran the command, and fanning out means fanning out the runners too. The API is a subprocess, which is a real answer but a thin one. Isolation is provider-dependent, and the Docker provider is a shared kernel.
7. Okteto
Underneath: Kubernetes namespaces and pods per developer or per branch, driven by a manifest that describes the environment. The programmatic caller is a first-class citizen here in a way it is not in cohort one, because environment-per-pull-request is the product's native use case rather than an automation bolt-on. If your application is already a dozen services on Kubernetes, this cohort member understands your problem better than anyone else in this post.
Honest drawback: it is Kubernetes all the way down, which decides every one of my four questions for you. Concurrency is node capacity and scheduler behaviour. The cold path is schedule plus image pull, which is not fast and is not consistent. Teardown is Kubernetes teardown, so namespaces finalize, persistent volume reclaim policy decides what survives, and a stray load balancer can outlive the thing that created it. And the boundary is a shared node kernel.
Cohort three: create is not really an API call
These two are in the list because they are genuinely good at what they do and because they usefully break the axis. In one of them there is no server to call. In the other, create is a session a person starts.
8. StackBlitz (WebContainers)
Underneath: a WebAssembly runtime executing Node inside the browser's own sandbox. Nothing is provisioned, so nothing is metered, and the isolation boundary is the browser origin — which is, for the record, one of the most aggressively attacked and therefore most hardened boundaries in computing. The programmable surface is a JavaScript API inside a page, so the caller is page code and there is no server-side create at all.
Honest drawback: it is not a substrate for your CI or your agent fleet unless your agent happens to live in a browser tab. Native binaries, arbitrary syscalls, a real Postgres and anything that wants a kernel are out, by design. But for a documentation page that runs, a bug reproduction a stranger can open, or a teaching environment, zero provisioning and structurally zero idle cost are unbeatable, and the limits announce themselves immediately rather than lurking in a bill.
9. CloudShell and the Cloud9 successor story
Underneath: a managed shell host attached to your cloud account, with a small persistent home directory and a session-scoped lifetime. AWS retired Cloud9 for new customers and the path forward is the IDE-plus-CloudShell story rather than a hosted IDE product, so this is the entry in the list most likely to have moved since I wrote this sentence. Every major cloud has a version of it and they are all roughly the same shape.
Honest drawback for a program-shaped caller: create is a console session, not an endpoint you drive fifty times. For three commands against one cloud account with credentials already attached, it is unbeatable and you should not overthink it. As a fan-out substrate it simply is not one, and pretending otherwise would waste your afternoon.
Cohort four: a program is the intended caller
These four were designed after the caller changed, and you can tell within thirty seconds because their documentation opens with a code sample instead of a screenshot. This is also the cohort where you most need to verify everything, because it is moving fastest and two of the four have changed what they are since 2024.
10. Daytona
Underneath: an API-driven sandbox platform aimed at AI workloads, which is a repositioning from an earlier identity as a self-hosted development environment manager. I include it partly because that repositioning is this post's argument in corporate form: a dev-environment company watched who was calling create and followed them. The current product is built around programmatic creation with snapshot-style reuse.
Honest drawback: the identity shift means documents from different eras describe genuinely different products, so anything you read second-hand — including this paragraph — may be describing the wrong one. Pin the substrate and the isolation boundary down in writing before you commit, because that is the claim most likely to be stale in the blog post you found on page one.
11. E2B
Underneath: a sandbox API for AI-generated code, Firecracker-based, with SDKs as the primary and intended interface. It is the clearest case in the list of a product where a model is the assumed caller — the ergonomics of "run the code the model just wrote and give me back stdout" are among the best anywhere, and it is open source, which matters if you ever need to read the thing rather than trust it.
Honest drawback: it is a code-execution sandbox, not a development environment, and the difference is real. There is no repo-shaped onboarding, no editor, and the lifetimes are short by design. If your actual requirement is "a human reviews a branch in a browser with a terminal attached," this is the wrong shelf. If your requirement is the interpreter loop inside an agent, it is one of the two or three products I would shortlist.
12. Modal sandboxes
Underneath: sandboxes inside a serverless Python platform, defined in Python, with container-shaped images. It is the most "the environment is a function of my code" entry here, and the developer experience of defining an image in the same file as the thing that runs in it is genuinely pleasant. The GPU story is real, which none of the microVM entries in this post — mine very much included — can match at all.
Honest drawback: you are adopting a platform, not a primitive. Your environment definition lives inside Modal's object model, which is a cost you pay on the way in and again on any way out, and the fit is best when the rest of your workload already lives there. If your threat model specifically requires a separate kernel per environment, verify the current isolation claims against their docs rather than against anyone's summary, including mine.
13. PandaStack
Mine, stated plainly so you can discount it accordingly. Underneath: one Firecracker microVM per environment with its own kernel — Firecracker v1.16.0, guest kernel 5.10, Ubuntu 24.04 userland — and no UI-first anything. The interfaces are a Python SDK, a TypeScript SDK, a CLI and a REST API, with tokenless preview URLs for whatever ports the guest exposes. Every create is a snapshot restore at p50 179 ms and p99 203 ms, which matters less as a speed claim than as a shape claim: there is no warm pool, so there is no warm path to miss. The only slow case is the very first spawn of a template that has no snapshot yet, which cold-boots in about three seconds and bakes one. Fan-out is in the API rather than in a guide: `fork()` clones the disk and the child cold-boots, while `fork_tree(n)` snapshots the parent once and boots n children that inherit its running memory, capped at 16 children per call.
Honest drawback: there is no browser IDE and no plan for one, so for "onboard a new hire with an editor already open" cohort one wins and I would say so on a call. Memory bills as committed GiB-hours for as long as a sandbox exists, so idle is cheap but not free. Per-environment VMs mean capacity can genuinely refuse you, and a 503 on create is a real thing your orchestration code has to handle. There is no GPU and no PCI passthrough, so anything GPU-accelerated does not run. And egress to the internet is open by default — sibling sandboxes cannot reach each other's subnets and the cloud metadata range is dropped at the host, but your own VPC and databases are not fenced for you.
All thirteen, against the four questions
| Platform | Who create was designed for | Isolation boundary | Fifty at once | Cold path |
|---|---|---|---|---|
| GitHub Codespaces | A person with a GitHub identity | Container, shared kernel | Org policy and quota decide | Image build on prebuild miss |
| Gitpod / Ona | A person, increasingly an agent | Container, in your cloud account | Your quota, not theirs | Prebuild-dependent |
| Replit | A person, or Replit's own agent | Container, shared kernel | Product-shaped, not fleet-shaped | Image-cached, fast |
| Google Cloud Workstations | A GCP admin, then a person | VM per workstation | Regional CPU quota | VM-shaped boot |
| Coder | A platform team, then a person | Whatever your Terraform says | Fifty Terraform applies | Whatever you provisioned |
| DevPod | A CLI on somebody's machine | Provider-dependent | Fan out the runners too | Devcontainer build, locally cached |
| Okteto | A pull request | Kubernetes pod, shared node kernel | Node capacity and the scheduler | Schedule plus image pull |
| StackBlitz (WebContainers) | Page code in a browser tab | Browser origin sandbox (WASM) | Fifty tabs, on your laptop | Near-instant; nothing is provisioned |
| CloudShell-class shells | A person in a cloud console | Managed shell host | Not a fan-out substrate | Session start |
| Daytona | A program; verify current docs | Vendor-stated — pin it in writing | Program-driven by design | API-fast |
| E2B | A model writing code | Firecracker microVM, own kernel | Program-driven by design | Sandbox start |
| Modal sandboxes | Your Python, as a function | Platform sandbox; check their docs | Platform-managed scaling | Image-dependent |
| PandaStack | A program, only ever | Firecracker microVM, own kernel | Honest 503 when memory runs out | Flat 179 ms restore; no warm path |
The column that reorders this category is the second one. Nothing else in the table is as predictive, because the intended caller is upstream of all the rest: a platform built for a person optimises for the first environment being delightful, and a platform built for a program optimises for the fiftieth being identical to the first.
Fifty at once, and what it costs when it works
Here is the fan-out written the way it survives contact with reality, which is to say with three defences in a specific order: a TTL set at create time so the platform cleans up after your process dies, an unconditional teardown in a `finally` so it normally never comes to that, and an in-guest timeout so one wedged command cannot hold an environment open all night.
import concurrent.futures as cf
from pandastack import Sandbox
from pandastack.exceptions import PandastackError
REPO = "https://github.com/acme/widget.git"
BRANCHES = [f"pr/{n}" for n in range(50)]
def review(branch: str) -> tuple[str, int, str]:
# ttl_seconds is an IDLE clock, not a wall clock. It does not cap how long
# this environment may live -- it caps how long it may sit UNTOUCHED, so a
# busy sandbox is never reaped mid-build. The activity bump is gated on
# whether a request actually uses the guest: exec, byte reads and PTY
# traffic reset it, read-only status/metrics/log GETs deliberately do not.
# Without that gate a dashboard polling every 30s keeps the fleet alive
# forever, which is a bug we shipped once and then fixed.
sbx = Sandbox.create(
template="code-interpreter", # 2 GiB baked -- half the memory bill of base
ttl_seconds=600, # backstop for the case the finally never runs
metadata={"branch": branch, "caller": "pr-bot"},
)
# Note what is NOT here: cpu= and memory_mb=. Firecracker cannot change
# vCPU or RAM at snapshot restore, so for any template with a baked
# snapshot those arguments are overridden to the snapshot's values before
# anything persists. Passing them is a wish, not a setting.
try:
rc = sbx.exec_stream(
f"git clone --depth 1 --branch {branch} {REPO} /src",
timeout_seconds=120,
)
if rc != 0:
return branch, rc, "clone failed"
# The real bound is IN-GUEST. One-shot exec() does not actually enforce
# timeout_seconds, and exec_stream only widens the client's HTTP
# timeout -- so timeout(1) is the only thing in this function that can
# kill a wedged `npm ci` that decided to resolve the internet.
logs: list[str] = []
rc = sbx.exec_stream(
"cd /src && timeout --kill-after=10s 600 sh -c 'npm ci && npm test'",
on_stdout=logs.append,
on_stderr=logs.append,
timeout_seconds=660,
)
return branch, rc, "".join(logs[-20:])
finally:
# kill() is the teardown call. There is no sbx.delete(). Unconditional,
# in a finally -- not inside an except, and not after the return.
# `with Sandbox.create(...) as sbx:` does the same thing if you prefer
# the context manager; __exit__ calls kill().
sbx.kill()
# Fifty at once is a real question to a scheduler, not a formality: each one is
# a Firecracker microVM with its own kernel, so 50 x 2 GiB is 100 GiB of
# committed guest memory that has to exist somewhere in the fleet right now.
with cf.ThreadPoolExecutor(max_workers=50) as pool:
futures = {pool.submit(review, b): b for b in BRANCHES}
for fut in cf.as_completed(futures):
branch = futures[fut]
try:
_, rc, tail = fut.result()
print(f"{branch}: {'ok' if rc == 0 else f'FAILED rc={rc}'}")
except PandastackError as err:
# A create refused for capacity is a real 503, and it is the honest
# answer -- per-environment VMs can run out of memory in a way a
# shared-kernel pod cannot. Back off and run a narrower fan-out.
# Do not retry in a tight loop against a fleet that is already full.
print(f"{branch}: platform refused: {err}")
Let me price my own homework, because a roundup that will not do arithmetic on its own product is an advertisement. Fifty `code-interpreter` sandboxes bake at 2 GiB each, so that is 100 GiB of committed guest memory at $0.0162 per GiB-hour — about $1.62 an hour. A fan-out that finishes in ten minutes therefore costs roughly 27 cents in memory, plus CPU at $0.054 per vCPU-hour billed on seconds actually burned, which means the stretch where fifty environments sit waiting on `npm ci` to resolve the internet is not charged as eight idle vCPU apiece. One rate card, all classes.
The refusal branch is not decoration. One microVM per environment means memory is a real, finite resource and a create can genuinely be denied — there are 16,384 pre-allocated subnet slots per agent host, but memory is the binding constraint long before you get anywhere near that number. A 503 is the honest answer and the correct response is backoff plus a narrower fan-out, not a retry loop hammering a fleet that is already full. A shared-kernel pod platform will often say yes here and then quietly give you contention instead, which feels better and is not.
If what you need fanned out is warm in-memory state rather than fifty cold checkouts — fifty agents all starting from the same loaded index, the same imported libraries, the same running server — then the call is `fork_tree`, not `fork`. That distinction is the single most-misused pair in my own API: `fork()` clones the disk and the child cold-boots with its own entropy and its own PIDs, so you keep the files and lose the warmth. `fork_tree(n)` snapshots the parent once and boots n children from that snapshot, capped at 16 per call, so wider fan-outs grow the tree breadth-first. Same-host forks land in 400–750 ms; cross-host is 1.2–3.5 s because memory has to move. The mechanics are in Snapshot vs restore vs fork vs clone, explained.
One more detail that bites program-shaped callers specifically, and that I got to learn in production: a snapshot restores the guest exactly as it was, including its clock. Restored guests used to come up believing it was whenever the template was baked, which meant TLS handshakes failed with certificate-not-yet-valid errors in a way that looked like a network problem for an embarrassingly long time. The clock is force-synced on restore, resume and wake now. If you are evaluating any snapshot-based platform, that is a question worth asking out loud.
The honest limits on my side of the table
- No editor, no onboarding flow, no browser IDE. This is a compute primitive with dev-environment uses, driven by an SDK, a CLI and a REST API. Cohort one sells a product to a human opening a laptop; I sell one to a program. For the human case they win, and the substrate argument for why a kernel per environment changes what you can safely run is in Cloud Dev Environments on microVMs.
- A reaped sandbox loses uncommitted work. The idle reaper deletes the VM, and your unpushed diff goes with it unless you created with `persistent=True`, raised the TTL, or snapshotted first. One deliberate nuance: an idle auto-reap does not cascade into the sandbox's snapshots, because a snapshot is durable and is supposed to outlive the VM that made it. The pattern that genuinely survives abandonment is snapshot-and-let-it-die.
- The idle clock is an idle clock, and that cuts both ways. `ttl_seconds` measures time since the guest was last actually used, so a busy environment is never killed mid-build — and a continuously busy one is never killed at all, which is why read-only status, metrics and log GETs deliberately do not reset it. Full mechanics in How to control sandbox lifetime: TTL, idle, and cleanup.
- vCPU and RAM are baked into the template snapshot. Firecracker cannot change them at restore, so the `cpu` and `memory_mb` you pass to create are overridden to the snapshot's values before anything persists. If you need different memory, you need a different template, not a different argument. RAM is the only template knob: `base` is 4 GiB, `code-interpreter` and `agent` are 2 GiB, `browser` is 4 GiB, `postgres-16` is 1 GiB, all at 8 vCPU.
- No GPU, and no path to one. No PCI or VFIO passthrough, so anything GPU-accelerated does not run here. Software rasterizers work, slowly. If your environments need a GPU, Modal is the entry in this list to look at and I will not pretend otherwise.
- Egress is open by default. Sibling sandboxes cannot reach each other's subnets and the cloud metadata IP range is dropped at the host, but there is no default-deny policy, and your own VPC, internal APIs and databases are not fenced for you. If your fan-out runs code from pull requests nobody read, those deny rules are still your work — the options are in Controlling Network Egress for Untrusted Code.
- The guest kernel is 5.10. Most toolchains do not care and some do, so anything that wants a very recent kernel interface, an exotic filesystem, or specific container-in-container trickery is worth testing before you plan a migration around it.
Picking one
Sort by who is calling create. Let the isolation boundary veto, and let the cold-path shape break the tie.
- A human opening a laptop, repo on GitHub, code you trust: Codespaces. Spend the afternoon you saved on the prebuild workflow and the retention policy, because those two settings are the entire bill.
- A human opening a laptop, hard network or compliance perimeter: Coder if you have a platform team and Terraform competence, Google Cloud Workstations if you are GCP-committed and want someone else to operate it, Ona if you want the polished product with runners in your own account.
- An application that is already a dozen services on Kubernetes: Okteto. Budget for teardown semantics being Kubernetes teardown semantics, and for the cold path being schedule plus pull.
- A reproduction, a doc that runs, a teaching environment, a bug report a stranger can open: StackBlitz. Zero provisioning, structurally zero idle cost, honest limits.
- Three commands against one cloud account, right now: CloudShell. Do not overthink it, and do not try to build on it.
- The interpreter loop inside an agent — short-lived, code the model just wrote, stdout back: E2B or Daytona, or mine. All three assume a program is calling, which is most of the battle.
- Environments defined in Python, next to the rest of a workload that already lives there, or anything needing a GPU: Modal.
- A program creating environments at volume — a pull-request webhook, a CI matrix, an agent fanning out branches — especially when the code came from outside your org and a shared kernel is the part your security reviewer will not sign: PandaStack. Preview-environment specifics are at preview environments and the rate card is at pricing.
- A portable, local-first setup with no vendor at all: DevPod, with your eyes open about where the control plane actually ended up.
The recommendation I would actually give, including against my own product: if the caller is a person, buy a workspace product and stop reading comparison posts — the money you save by picking the clever option will be spent twice over on the onboarding story you now have to build. If the caller is a program, the question is only whether you need a kernel boundary. If you do not, the shared-kernel sandbox APIs are faster to adopt and perfectly honest about what they are. If you do, a microVM per environment is the trade, and mine is open source so you can read the boot path before you trust it. What I would not do is pick a platform designed for a person and then drive it from a webhook fifty times an hour, because that is how you end up maintaining a readiness loop, a quota ticket, and someone's personal access token in a secret store.
Frequently asked questions
How can I tell whether a platform's API is real or just a remote control for its UI?
Four checks, in increasing order of how much they annoy a sales engineer. First, can you get a machine credential — a token that belongs to a service rather than to a named human? If the only option is a personal access token scoped to what a person may see, then the platform's model of a caller is a person, and your CI is about to hold somebody's identity in a secret store. Second, does create return something you can immediately use, or a row you have to poll until maybe-ready? Both are legitimate, but only one lets you write a synchronous script without inventing a readiness loop, and readiness loops are where retry logic goes to die. Third, ask what happens at fifty concurrent creates. There are three good answers — it works, it queues, it refuses with a documented error — and a shrug is a fourth answer that tells you nobody designed for your case. Fourth, read the error catalogue. A product built for programs documents its failures, because a program cannot read a friendly red toast and decide what to do next. If the docs show screenshots where the error codes should be, you have your answer.
Why does the isolation boundary matter more now than it did two years ago?
Because the author of the code changed. Two years ago the realistic worst case inside an ephemeral environment was a colleague's buggy migration. Now a routine environment builds a fork's pull request, runs a dependency's post-install hook, or executes something a model generated ninety seconds ago. A model-generated `rm -rf` is not malice; it is a plausible token sequence that got sampled. On a shared host kernel that distinction does not buy you much — namespaces and cgroups are features of a large, actively-researched attack surface, and a CI-shaped environment is usually the most credential-rich thing an organisation owns, which makes it the best target rather than the least interesting one. A microVM moves the boundary to the hypervisor: a separate kernel per environment, with a device model of a handful of emulated devices instead of a whole kernel's syscall table. That said, the boundary is not a security posture on its own. On PandaStack, egress to the internet is open by default — siblings cannot reach each other's subnets and the cloud metadata range is dropped at the host, but your own VPC and databases are not fenced for you. If you run code nobody read, writing those deny rules is still your job, and no substrate choice does it for you.
What actually happens when I ask for fifty environments at once?
It depends on where the ceiling lives, and the ceiling is almost never where the marketing page implies. On the metered workspace products it is org policy and account quota, which is a number in a console that somebody set once and nobody owns. On Kubernetes-based platforms it is node capacity and the scheduler, so the honest answer is that the first fifteen schedule instantly and the rest wait for the autoscaler to add nodes, which takes however long your cloud takes. On bring-your-own-substrate platforms it is your own cloud quota and, if your templates are Terraform, fifty simultaneous applies — a thing that works right up until it does not. On PandaStack it is memory. Each environment is a microVM with committed guest RAM, so fifty 2 GiB sandboxes is 100 GiB that has to exist in the fleet right now; there are 16,384 pre-allocated subnet slots per host, but memory binds long before that. A create that cannot be placed returns a real 503, and the correct response is backoff and a narrower fan-out. I consider an honest refusal better than a yes that silently becomes contention, but it does mean your orchestration code has a branch most people forget to write.
Does PandaStack have a cold path, and what is the catch?
There is one, and it is small and specific. Every normal create is a restore of the template's baked Firecracker snapshot — p50 179 ms, p99 203 ms — and crucially there is no warm pool of idle VMs behind that number. That is the part that matters for a program-shaped caller: there is no warm path to miss, so the fiftieth create of the morning takes the same time as the first, and your pipeline's timeout only ever has to be written against one latency. The cold path appears exactly once per template: the very first spawn of a template that has no snapshot yet cold-boots in about three seconds and bakes the snapshot on the way, after which everything restores. Re-baking a template invalidates the old snapshot, so a template change moves you back to that one-time cold boot. The catch worth knowing is sizing. Firecracker cannot change vCPU or RAM at snapshot restore, so for any template with a baked snapshot the `cpu` and `memory_mb` you pass to create are overridden to the snapshot's values before anything persists — the API response, the database row and the billing event all show the corrected numbers rather than your request. Different memory means a different template, not a different argument.
Can an AI agent fan out branches on PandaStack, and is that fork or fork_tree?
Yes, and it is almost always `fork_tree` — this is the single most-misused pair in my own API, so it is worth being blunt. `fork()` clones the disk only and the child cold-boots: you keep the checkout, the installed toolchain and the files, and you lose every bit of warmth, because the child comes up with its own entropy, its own PIDs and nothing in memory. `fork_tree(n)` snapshots the parent once and boots n children from that snapshot, so each child inherits the parent's running memory — the loaded index, the imported libraries, the server that is already listening. If your agent wants fifty explorations from one warmed-up state, `fork_tree` is the only call that does what you meant. Two practical constraints. There is a hard cap of 16 children per `fork_tree` call, and exceeding it is an error rather than a clamp, so wider fan-outs grow the tree breadth-first: fork 16, then fork from those. And placement matters for latency — same-host forks land in 400 to 750 ms, while cross-host is 1.2 to 3.5 seconds because the memory image has to move across the network first. Teardown is `kill()` on each child, in a `finally`, unconditionally. There is no `delete()` method.
Keep reading
- Top 12 ephemeral development environment platforms — The same category graded on a different axis: what the second environment costs and what the fourteen you forgot cost.
- Top 10 disposable dev environment platforms — Graded on teardown correctness — whose leftovers survive the delete call, and which ones keep billing.
- Best sandbox APIs for LLM agents — Cohort four in much more detail, from the perspective of an agent as the caller rather than a dev team.
- Snapshot fork and tree-of-thought — What fork_tree is actually for: fanning out warm in-memory state instead of fifty cold checkouts.
- Snapshot restore vs warm pools — Why having no warm path is the point, and what a warm pool costs you when a program is the caller.
Related posts
- Top 11 Ephemeral Development Environment Platforms (2026)
Every roundup in this category sorts by substrate or by price list. Both are downstream of a question nobody asks on the evaluation call: who, or what, is going to call create? Eleven platforms sorted by their creator — human, CI job, pull-request webhook, agent API — because the creator decides how many environments exist at once and who pays for the idle ones.
- Top 8 Ephemeral Development Environment Platforms (2026)
Feature grids do not decide this. Two numbers do: how fast a fresh environment can exist, and what happens to it when nobody is looking. Eight platforms graded on both, with the substrate each one actually isolates with and the catch I would want before the purchase order.
- PandaStack vs Coder
These two products are shopped against each other constantly and compete almost never. One gives a person a machine for the week; the other gives a program a machine for four seconds. Naming that split is most of the decision.
- Top 7 Ephemeral Development Environment Platforms in 2026
An environment is ephemeral when it is created from a definition, nobody is sad when it dies, and the 400th costs the same as the 4th. Most "cloud dev environments" fail at least one of those tests.
- The Best Replit Alternatives in 2026
"Replit alternative" means three unrelated things: a browser IDE, a cloud dev environment, or the sandbox API that runs untrusted code under the hood. Split the intent first and the shortlist collapses from ten options to two.
More in AI agent sandboxes · See AI agent sandboxes on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.