Top 11 Ephemeral Development Environment Platforms (2026)
Every comparison in this category sorts the field by substrate or by price list. Both are downstream of a question that almost never gets asked on the evaluation call: who, or what, is going to call create?
That sounds like a detail and it is the whole decision, because the creator determines the only two properties that matter: how many environments exist at once, and who pays for the idle ones. Everything else — the IDE integration, the config format, the SSO story — is a preference. Those two are arithmetic, and arithmetic does not negotiate.
A platform that is excellent when a human opens one environment each morning is the wrong platform when an agent opens four hundred in an afternoon. The failure, when it comes, is never in the marketing copy. It is create latency multiplied by fan-out, and the multiplication is unforgiving: a create time nobody notices at a concurrency of one becomes the dominant term in your wall clock at four hundred, and an idle policy that was a rounding error on twelve seats becomes the largest line on the invoice.
I'm Ajay. I build PandaStack, which runs Firecracker microVMs behind an API, and it is entry eleven. I have sorted the other ten by creator rather than by merit, because the ranking inverts depending on which creator you have, and the most expensive mistake in this market is buying a platform built for a creator you do not have.
The create call is the whole product
An ephemeral environment platform is, functionally, two API calls with a billing system attached. Create, and destroy. Everything in between — the editor, the terminal, the port forwarding, the dotfiles — is interface. The create call deserves this much attention because its caller fixes the shape of your bill, and the caller is not a setting. It is a property of your organisation.
Three variables fall out of the caller. Concurrency: how many environments exist simultaneously, a different question from how many you create per day. Lifetime: seconds for an agent step, days for a pull request, months for a workstation somebody never deletes. And the active fraction: the share of that lifetime during which anything is actually executing — the quiet killer, because it is the number every cost model omits and the one that decides whether "keeps running when idle" is a reasonable product decision or a transfer of your budget. Identify your caller honestly, including the callers you are about to acquire, and the shortlist writes itself.
The four creators
A human clicking open
Concurrency equals headcount, bounded above and below by the same number. Lifetime is a working day at minimum and in practice much longer, because the environment accumulates uncommitted work and nobody deletes a machine that contains their afternoon. Active fraction is maybe a quarter: lunch, meetings, the sprint review, two hours reading someone else's pull request. Keep-warm is the correct decision here, and a seat price is an honest unit when concurrency is pinned to headcount.
A CI job
Concurrency equals jobs in flight, which equals push rate multiplied by matrix width. Zero at three in the morning, forty at 11:15, and the distribution is a sawtooth driven by when your team stands up. Lifetime is the job, minutes at most, and the active fraction is close to one: a CI environment that is idle is a CI environment that is broken. Create latency goes straight into the critical path of every build, multiplied by the matrix — a thirty-second create on a forty-way matrix is twenty minutes per push if you run them serially and a capacity cliff if you do not.
A webhook from a pull request
Concurrency equals open pull requests, a quantity no engineering organisation has ever successfully bounded. Lifetime is days to weeks, set by review latency rather than by anything you control. And the active fraction is the smallest in this list by a wide margin: a PR environment is woken by four CI runs and one reviewer at 16:00 on Thursday, and spends the remaining ninety-nine-point-something percent of its existence being billed for having a disk. This is the creator for which the idle verb is the entire decision, and the one where "suspends" sounds thrifty and costs the most over a quarter, because a stopped disk keeps billing silently for months after the branch was merged.
An agent calling an API
Concurrency equals whatever the loop decides, which is a function of the model's branching factor and not of your headcount. This is the creator that breaks every assumption the other three were built on, because the number of concurrent environments is finally decoupled from the number of people employed. Lifetime is seconds to minutes. Active fraction is high but spiky — an agent thinks between calls, and thinking is a period during which you are renting a machine to hold a filesystem. Create latency lands inside the agent's reasoning budget: a human waiting four seconds is mildly annoyed, and an agent waiting four seconds per step on a thousand-step run has turned a one-hour job into two.
Most platforms here serve exactly one of those four well. The other three are a page on the website.
The concurrency inversion
Here is the structural fact underneath the category. Once you see it you cannot unsee it in a pricing page.
Human-oriented platforms optimise for a long-lived environment per person, and they price per seat. That is coherent: when concurrency is pinned to headcount, a seat is a reasonable proxy for consumption, and keeping an environment warm is cheap insurance against a developer deciding the tool is slow. The engineering effort goes into resume latency and the editor, because those are what a person perceives.
Machine-oriented platforms optimise for create and teardown, and they charge for consumption — metered compute for as long as the environment existed, rather than a seat for as long as the contract runs. Also coherent, for the opposite reason: when concurrency is unbounded above and genuinely zero between bursts, a seat price is either a tax on an idle org or a bargain the vendor cannot survive. The effort goes into the cold path, the only path a machine ever takes.
The inversion is that these are not points on a spectrum. They are opposite bets about the shape of your demand curve, and each is approximately the worst possible fit for the other's workload. Pay per seat for a CI fleet and you are buying 730 hours a month of machines that work for forty minutes a day. Pay metered compute for sixty engineers who each want an editor open for eight hours and you have bought a workstation product with no workstation in it, at full rate, with the editor left as an exercise for the reader.
And most teams discover they bought the wrong one at the precise moment an agent loop starts. Not when they adopt agents as a policy — when one engineer wires a coding agent into a script that opens an environment per candidate patch. The next invoice is the discovery mechanism, and the conversation after it is about procurement rather than engineering.
The eleven, grouped by creator
Grouped, not ranked, with the strongest case stated before the catch. Within a group, the entry strongest for that creator comes first.
1. GitHub Codespaces — creator: a human
A container described by a `devcontainer.json` in the repository, running on a cloud VM allocated per codespace, driven from browser VS Code or your desktop editor. Strongest case: the distance between a repository and a working environment for somebody who has never seen the project, which is close to zero and was a genuinely hard problem. The devcontainer format it popularised is also the most portable asset in this list — the one thing you take with you when you leave. The catch is that the creator is firmly a human: creating four hundred of these because four hundred patches want testing is not what the pricing, the quotas or the UX are shaped around, and the idle timeout — the most consequential setting in the product — lives in a personal preference pane rather than the repository where a platform team could own it. Verify current idle, retention and storage billing behaviour in GitHub's docs.
2. Coder — creator: a human, on your hardware
A self-hosted control plane that provisions workspaces from Terraform templates, so a workspace is whatever your providers can create — a Kubernetes pod, a cloud VM, a bare-metal box. Strongest case: the regulated enterprise, without irony. If the requirement is that environments live in your VPC, under your IAM, with your audit trail and a security team that wants to read the definition before approving it, this entry treats that as the primary requirement rather than an enterprise-tier afterthought. The catch is that Terraform-shaped environments tend toward long-lived ones, because the underlying resource has a disk and recreating it is a full provision cycle — so ephemerality becomes an autostop policy you enforce rather than a property you receive. And because the substrate is whatever your template provisions, nothing stops a team believing it has a VM boundary while running a shared kernel.
3. Gitpod / Ona — creator: a human, with prebuilds
The project that argued for strong ephemerality before the rest of the category caught up: environments started per task from a declarative config, prebuilt ahead of time so the first command is not a dependency install. Its product direction and its naming have both moved substantially — treat any description of it, including this one, as a snapshot. Strongest case: prebuilds, which remain the most useful idea this category produced. Build the environment when the commit lands rather than when a developer asks, and much of the cold-start problem evaporates without anyone engineering a hypervisor. The catch is that prebuilds answer create latency for a human creator and answer it weakly for a machine one: a prebuild is warmed per commit, not per fan-out, so the four-hundredth environment from that commit still pays the create path in full.
4. DevPod — creator: a human, with no control plane at all
A client-side tool rather than a service. It reads the devcontainer spec and creates that environment on a provider you pick — your laptop, a cloud VM, a cluster — with nothing in the middle. Strongest case: it is the un-lock-in option, and that is a real requirement rather than a philosophical one. If you are about to standardise an engineering organisation's entire inner loop on one vendor's hosting, a client that runs the open spec anywhere answers the fear directly. The catch is symmetric: no control plane means no control plane. No registry of what exists, no policy enforcement point, no cost attribution, no reaper. Those become your problem — fine at ten engineers, a platform project at two hundred.
Those four share a creator and therefore a failure mode: they are session products, and a session implies a person. Now the group created by a push.
5. Vercel preview deployments — creator: a PR webhook
Every push produces a build and an immutable preview URL. The unit is a deployment, not a machine, and that distinction is the whole entry. Strongest case: zero-configuration per-pull-request URLs, the product that changed reviewer behaviour across the industry — largely because of how little effort it demanded from the team adopting it. On the creator axis this is close to ideal for a webhook: the create is a build, the idle cost of a finished immutable deployment is storage and routing rather than a machine, and nothing needs reaping. The catch is that it is a deploy target and not an environment. No persistent process you own, no shell, no daemon to attach a profiler to, and your database is not part of the preview unless you build that yourself. Note also that their separate programmatic code-execution sandbox is a different product with a shared brand.
6. Northflank — creator: a PR webhook or a CI job, optionally in your cloud
Build pipelines, services, managed databases and preview environments in one product, with the option to run the same control plane inside your own cloud account. Strongest case: the intersection of per-PR previews, managed data services and bring-your-own-cloud, which is genuinely uncommon — most products here give you two of those three and put the third on a roadmap. For a team whose previews are only useful with a real Postgres attached, and whose compliance story requires the nodes to be theirs, this is a shortlist of one or two. The catch is conceptual surface area: more objects in the model than in a two-field deploy product, and you will learn all of them whether you planned to or not. And BYOC changes the compliance conversation without changing the kernel-sharing story — if the real requirement is untrusted code, platform polish is not the variable that matters.
7. Okteto — creator: a PR webhook, on your cluster
Kubernetes, in your cluster, with a namespace as the unit and a manifest describing the environment and its dependencies. Strongest case: the honest answer for a team whose production already is a dozen services on Kubernetes, because the preview is the same shape as production rather than a simplified imitation that passes tests production would fail. Same manifests, same service discovery, same ingress. The catch is the cluster. You size it, you pay for its headroom, you carry its upgrades, and a per-PR environment that pulls up fifteen pods has a scheduling queue behind it — which turns the webhook creator's unbounded concurrency into a capacity problem on hardware you provisioned last quarter. The boundary is a namespace: a shared kernel with RBAC on top.
And now the group created by a program. This cohort has changed the most in two years, and it is where the substrate question stops being academic — the code being run was frequently written by a model thirty seconds ago, with no intent either way.
8. E2B — creator: an agent API
An open-source sandbox API built for AI code execution: an SDK call in, an isolated microVM out, with filesystem and process surfaces designed for a program to drive rather than a person to type into. Strongest case: it is the reference implementation of the creator axis taken seriously. The API reads as though designed by somebody who had actually tried to make a language model use one, which is a lower bar than it sounds and one most of this list does not clear. The boundary is a microVM with its own kernel — not the obvious call when they made it, clearly the right one for this creator. The catch is scope discipline, which is a virtue rather than a flaw: deliberately not a human IDE and not an application host, so if the task ends with "and now serve this under production traffic", that is a second product.
9. Daytona — creator: an agent API, after a pivot
Worth including precisely because it has changed category. It started as a developer-workspace manager — a human-creator product — and its current public positioning is explicitly about fast sandboxes for AI-generated code with an SDK-first surface. That pivot is the best single piece of evidence for this post's thesis: a company with the infrastructure to serve both creators looked at the demand curve and picked the machine. Strongest case: the agent case it moved into, with a declarative environment definition, so a program can be handed a repository and get a reproducible environment without a human first curating a Dockerfile. The catch is for the buyer who arrived looking for the old product: before planning a seat-based IDE-attached rollout, confirm from current docs that the product is still shaped like that. Isolation model and stopped-state billing are the two things most likely to have moved.
10. Modal — creator: a CI job or an agent, function-shaped
Python-first serverless compute: decorate a function, declare its image in code, run it remotely including on accelerators, plus a sandbox surface for arbitrary commands. Strongest case on the creator axis is its idle behaviour — the serverless shape means a resting workload costs close to nothing, which matters enormously for an agent that spends most of its wall clock thinking rather than executing. The catch is the mental model: the primitive is "run this function in this image", so reproducing an expensive state tends to mean re-running whatever produced it rather than branching a machine already in it — a materially worse fit for an agent exploring a tree where reaching the interesting state took four minutes of installation. The Python-centricity is a feature right up to the moment the thing you need to run is not Python, at which point you are wrapping shell in a decorator and asking yourself some questions. Read their current security documentation for the runtime boundary.
11. PandaStack — creator: any program, which is also the limitation
Mine, so weigh it accordingly. Firecracker v1.16.0 microVMs — Ubuntu 24.04 on a 5.10 guest kernel — where every create is a snapshot restore rather than a boot. There is no warm pool of idle VMs, and the rest follows from that: the create path allocates a pre-built network slot, patches the tap MAC, reflinks the rootfs, forks a Firecracker process, loads the baked snapshot and resumes. End to end, p50 179 ms and p99 203 ms; inside it, `/snapshot/load` is about 80 ms, the resume 6 ms, the TCP probe on port 22 about 40 ms. The first-ever spawn of a template cold-boots in roughly 3 seconds and bakes the snapshot every later create restores. Each sandbox gets its own netns, veth pair and tap device from 16,384 pre-allocated /30 subnets per host.
Why that matters for a machine creator: the create number is flat because there is nothing to resolve at create time — no registry lookup, no `apt`, no install step — so the multiplication by fan-out stays linear instead of compounding. And the idle verb is genuinely delete rather than suspend: for hosted apps, sleep bakes a seed to object storage and destroys the virtual machine, and waking restores from that seed in about 1.2 seconds. The rate card is $0.054 per vCPU-hour and $0.0162 per GiB-hour for every class, CPU billed on the CPU-seconds actually burned, with no per-request charge.
Now the catches, and they are real. There is no IDE — no browser editor, no `devcontainer.json` ingestion, no button on the pull request. You get an API, a CLI, Python and TypeScript SDKs, exec, a filesystem interface, SSH and tokenless preview URLs; if what you want is a tab with a code editor in it, entries one through four exist for that reason. Guest RAM is fixed when the template is baked, because Firecracker cannot change vCPU count or RAM at snapshot restore, so a `memory_mb` passed on a create against a snapshotted template is silently corrected rather than rejected — a 200 and a different machine than you asked for, which is worse than an error. Baked sizes run from `postgres-16` at 1 GiB to `base` at 4 GiB, each with 8 burstable vCPUs. Egress is open by default: a few targeted DROP rules for known-abuse protocols, not a default-deny network.
| Platform | Who calls create | Substrate | At idle | The honest catch |
|---|---|---|---|---|
| GitHub Codespaces | Human | Container on a per-codespace VM | Stops, then deleted after retention | Idle timeout is a personal preference, not repo policy |
| Coder | Human, self-hosted | Whatever your Terraform provisions | Autostop policy you enforce | You operate the control plane; boundary varies by template |
| Gitpod / Ona | Human | Container, prebuilt per commit | Verify current policy | Prebuilds warm per commit, not per fan-out |
| DevPod | Human, client-side | Devcontainer on a provider you pick | Whatever your provider does | No control plane means no reaper and no cost attribution |
| Vercel preview deployments | PR webhook | Managed function runtime, immutable deploy | Storage and routing, no machine | A deploy target, not an environment — no shell, no database |
| Northflank | PR webhook or CI | Containers, Kubernetes-shaped, BYOC option | Scales with open PRs | Large object model; BYOC does not change kernel sharing |
| Okteto | PR webhook, your cluster | Kubernetes namespace | Sleeps idle namespaces | You pay for the cluster headroom the webhook demands |
| E2B | Agent API | MicroVM, own kernel | You pay while it exists | Deliberately not an IDE and not an application host |
| Daytona | Agent API, after a pivot | Verify current model | Verify stopped-state billing | Changed category; confirm the product before a seat rollout |
| Modal | CI job or agent | Their runtime around containerised workloads | Near-free when resting | Functions and images, so no branching a live machine |
| PandaStack | Any program; API only | Firecracker microVM, snapshot-restored | Deleted (app wake ~1.2 s from a seed) | No IDE; guest RAM baked at template build time |
A per-pull-request environment, created from an API
The webhook creator is the one most teams get wrong, because the code looks trivial. A push arrives, you create an environment, you post a link. What makes it non-trivial is that the caller is a machine: nobody is watching a spinner, so latency failures are silent timeouts rather than complaints, and nobody notices a leak, so a duplicate create survives until the invoice.
#!/usr/bin/env python3
"""One environment per open pull request, created by a webhook handler.
The creator is a program, so the three defences that matter are idempotency,
a platform-side deadline, and an index of record that is YOURS. None of them
are about latency, and all of them are the difference between a preview system
and a slow leak with a nice URL.
"""
from pandastack import Sandbox
REPO = "acme/api"
def open_environment(pr: str, ref: str, db) -> str:
# 1. Idempotency. A webhook fires more than once -- the provider retries,
# somebody force-pushes, a reviewer clicks "re-run". Without this you
# get two environments per PR and find out from accounting.
#
# Your own table is the index of record, not a list call: you must know
# which sandbox belongs to which PR even when the API is having a day.
previous = db.get_sandbox_id(REPO, pr)
if previous:
try:
Sandbox.get(previous).kill()
except Exception:
pass # already gone, or the TTL reaper beat you to it
sbx = Sandbox.create(
template="base", # Ubuntu 24.04; Node 24 LTS + Python 3.12 via mise
ttl_seconds=72 * 3600, # 2. the deadline, enforced platform-side
metadata={"repo": REPO, "pr": pr, "ref": ref, "owner": "pr-bot"},
)
db.put_sandbox_id(REPO, pr, sbx.id)
# Why ttl_seconds and not a finally block: the reaper runs on the platform,
# not in this process. If this handler is SIGKILLed, if the pod is evicted
# mid-request, if the whole webhook service is rolled during a deploy --
# the sandbox still dies on schedule. A `finally:` cannot promise that,
# because the failure you are defending against is "this interpreter stops
# existing". The 72 hours is a backstop for your orchestrator's death, not
# the normal teardown; the PR-closed hook below is the normal teardown.
try:
# The branch's code is the one thing that cannot live in the template,
# because it changes per environment. Everything else -- toolchain,
# system packages, warm caches -- belongs in the baked snapshot.
sbx.exec(
f"git clone --depth 1 --branch {ref} "
f"https://github.com/{REPO} /work",
timeout_seconds=120,
check=True,
)
# timeout_seconds is a CLIENT deadline; the server does not enforce it.
# If you need a hard wall, put it in the guest where the kernel honours
# it -- `timeout 900 ...` -- and give the client a little more slack.
sbx.exec_stream("cd /work && npm ci", on_stdout=print)
sbx.exec("cd /work && timeout 900 npm run build", timeout_seconds=960,
check=True)
sbx.exec("cd /work && setsid nohup npm start >/var/log/app.log 2>&1 &")
except Exception:
sbx.kill()
raise
# 3. The reviewer's link. Tokenless: the sandbox UUID is the credential, so
# this string is a password with a protocol on the front. Short TTLs.
return sbx.preview_url(3000)
def close_environment(pr: str, db) -> None:
"""The PR-closed hook. This is the teardown that should actually run."""
sid = db.get_sandbox_id(REPO, pr)
if not sid:
return
try:
Sandbox.get(sid).kill()
finally:
db.clear_sandbox_id(REPO, pr)
That handler has the same shape on every platform in the table; what differs is what the second defence can even mean. Where the idle verb is "stops", a 72-hour deadline buys a stopped disk instead of a running machine — a smaller bill, not a zero one. Where the idle verb is delete, the deadline is the bill. That single difference is most of the cost gap between the rows.
An agent fanning out N environments
The agent creator inverts the problem. You are not creating one environment and keeping it a while; you are getting one environment into an expensive condition — repository cloned, dependencies installed, dev server listening — and then needing that exact condition in N places at once, so N candidate patches can be tried without contaminating each other.
There are two branching primitives on PandaStack and conflating them is the mistake I see most often, including in my own earlier writing. `sbx.fork()` is a disk clone plus a cold boot: it reflinks the rootfs with copy-on-write and boots a fresh guest, so the child inherits the filesystem — installed dependencies, build caches, files you wrote — and does not inherit the parent's memory. A bare `fork()` therefore has the roughly three-second cold-boot shape, and the child draws its own entropy. `sbx.fork_tree(count)` is the memory-inheriting path: it restores the children from the parent's snapshot, at 400 to 750 ms same-host and 1.2 to 3.5 seconds cross-host. Said plainly, because those numbers get paired with the wrong call in every comparison table including two of mine: 400 to 750 ms is the `fork_tree` restore path; a bare `fork()` has the ~3 s cold-boot shape.
#!/usr/bin/env python3
"""Fan out N isolated attempts from one expensive parent state.
The whole point: pay for `npm ci` once, then branch the machine that already
finished it. The alternative -- N independent creates, each re-running install
-- multiplies the slowest step in your pipeline by your branching factor.
"""
import concurrent.futures as cf
from pandastack import Sandbox
PATCHES = ["patch-a.diff", "patch-b.diff", "patch-c.diff", "patch-d.diff"]
parent = Sandbox.create(
template="agent", # 2 GiB / 8 burst vCPU, baked
ttl_seconds=3600,
metadata={"run": "agent-7781", "role": "parent"},
)
try:
# Get the parent into the state that is expensive to reach. Do it ONCE.
parent.exec("git clone --depth 1 https://github.com/acme/api /work",
timeout_seconds=120, check=True)
parent.exec("cd /work && npm ci", timeout_seconds=900, check=True)
parent.exec("cd /work && setsid nohup npm run dev >/var/log/dev.log 2>&1 &")
# fork_tree is the MEMORY-INHERITING path: the children are restored from
# the parent's snapshot, so that dev server is still running inside each
# one. 400-750 ms per child same-host, 1.2-3.5 s cross-host.
#
# `parent.fork()` is a different animal: a reflink of the rootfs plus a
# COLD BOOT. The child gets the filesystem and a fresh kernel -- no
# inherited memory, no running dev server, its own entropy -- and it has
# the ~3 s cold-boot shape. Reach for fork() when you want a clean boot on
# a warm disk; reach for fork_tree() when you want the parent's RAM.
children = parent.fork_tree(len(PATCHES))
def attempt(pair):
child, patch = pair
try:
# Children inherited the parent's memory EXACTLY, which includes
# the RNG state. Anything that seeds crypto or generates ids right
# after a fork_tree must reseed, or every sibling agrees on the
# "random" value and you lose an afternoon to it.
child.exec("head -c 32 /dev/urandom > /dev/null")
child.filesystem.write(f"/work/{patch}", open(patch, "rb").read())
child.exec(f"cd /work && git apply {patch}", check=True)
r = child.exec("cd /work && timeout 600 npm test",
timeout_seconds=660)
return patch, r.exit_code, r.stdout[-2000:]
finally:
child.kill() # the happy path; the TTL is only the backstop
with cf.ThreadPoolExecutor(max_workers=len(children)) as pool:
for patch, code, tail in pool.map(attempt,
zip(children, PATCHES)):
print(f"{patch}: exit={code}")
finally:
parent.kill()
Note which cost that structure eliminates. Not the create — the create was already a few hundred milliseconds. The `npm ci`, paid once instead of N times. That is the agent creator's economics in one sentence: the expensive thing is reaching the state, not acquiring the machine, so the primitive that matters is the one that branches a state you already paid for.
The arithmetic nobody does before buying
Create latency is the number vendors publish and the number buyers under-weight, because at a concurrency of one it genuinely is a detail. It stops being one the moment your creator changes. Multiply it by your fan-out and look at the wall clock.
#!/usr/bin/env python3
"""What create latency costs, as a function of who is calling create.
The create_s values below are placeholders. Measure your own on a trial
account, on your repository, stopping the clock on YOUR readiness signal
rather than on a status field that says "running" -- "running" means a machine
exists, not that your dev server is listening.
"""
def wall_clock(n, create_s, parallelism, step_s=0.0):
"""Seconds of wall clock to have n environments doing work."""
waves = -(-n // parallelism) # ceil division
return waves * (create_s + step_s)
SCENARIOS = {
"human, one per morning": (1, 1),
"CI, a 40-way matrix": (40, 8),
"PR webhook, 60 open PRs": (60, 10),
"agent loop, 400 candidates": (400, 16),
}
for label, (n, par) in SCENARIOS.items():
print(f"\n{label} (n={n}, parallelism={par})")
for create_s in (0.2, 5.0, 45.0, 300.0):
secs = wall_clock(n, create_s, par)
print(f" create={create_s:6.1f}s -> {secs/60:8.1f} min of create")
# Read the first row and the last row together. At n=1 the gap between a
# 200 ms create and a 5-minute create is the difference between "instant" and
# "go get a coffee" -- annoying, survivable, and the reason keep-warm exists.
# At n=400 it is the difference between about four seconds and roughly two
# hours, which is not a developer-experience question any more. It is whether
# the agent loop is a product or a demo.
#
# Then add the term this model deliberately omits: the setup your repo
# contributes. If each environment runs its own `npm ci`, multiply that by n
# too -- which is the argument for branching one warm state instead of
# creating n cold ones.
The second half of the arithmetic needs no script. Count the environments that exist at 09:00 on a Monday and divide by the number of people awake. Where the idle verb is "keeps running" that ratio is usually north of three, and three times your headcount in machine-hours is worth knowing before renewal rather than after.
Picking one
Identify your dominant creator — including the one you are acquiring — and read only that line:
- A human, trusted code, company lives in GitHub: Codespaces. Turn prebuilds on before judging the create time, and put somebody's name on deleting stopped ones, because nobody volunteers for that.
- A human, inside a compliance boundary, on your own hardware: Coder, with autostop policies written before the rollout rather than after the first invoice. Gitpod for prebuilds; DevPod if lock-in is the fear and you accept being the control plane.
- A PR webhook, and the app is a web front end: Vercel preview deployments, then stop evaluating — the per-PR URL is solved; spend the time on your database story.
- A PR webhook, and the preview is useless without a real database: Northflank, or Render's equivalent if the object model is more than you want.
- A PR webhook, and production is a dozen services on Kubernetes the preview must genuinely exercise: Okteto, with the cluster headroom budgeted honestly.
- An agent API, and the task is code interpretation in a well-scoped sandbox: E2B, or Modal if your orchestrator is already Python and idle cost dominates.
- An agent API, the code is untrusted or model-generated, and the work is "reach an expensive state once, then branch it N ways": a microVM sandbox API with a real branching primitive — mine, or another in that cohort.
- Both a human cohort and a machine cohort, which is most companies by now: buy two products. A workstation for the people, a sandbox API for the fleet. Making either do both is where the quarters go.
The bottom line
You are not choosing a development environment. You are choosing what your concurrency is allowed to be a function of.
Platforms here get chosen on feature grids and regretted on the creator axis, in one of two shapes. Either create was slow enough that the per-task model quietly died and everybody went back to one long-lived box they are afraid to delete — a pet with a devcontainer file in it — or the idle verb was "keeps running" and a quarter later somebody in finance asked what the compute line was and nobody in the room could answer.
Both failures are predictable from the creator, before any trial account exists. Write down who calls create, what the concurrency is a function of, and what the environment costs while nobody is looking at it. If those three answers do not match the product's own design centre, no amount of integration polish fixes it, and you will spend the difference in engineer-hours instead of dollars — the same money at a worse rate.
Frequently asked questions
How do I work out which creator my team actually has?
Look at what already calls your CI, not at what people say in planning. Three measurements settle it in an afternoon. First, count the environments that exist at 09:00 on a Monday and divide by the number of people awake: a ratio near one means a human creator, a ratio well above one means something automated is already creating them and nobody owns the cleanup. Second, plot creates per hour over a week. A human creator gives a flat line with a morning bump; a CI creator gives a sawtooth keyed to your standup; a webhook creator gives a slow accumulation that never returns to zero; an agent creator gives a spike with no relationship to working hours at all. Third, and most usefully, ask whether anyone is planning to wire a coding agent into anything this quarter. If the answer is yes, your dominant creator is about to change and your concurrency will stop being a function of headcount. Buy for the creator you will have in six months.
Why does create latency matter so much more for an agent than for a human?
Because of where the latency lands and what it gets multiplied by. For a human, create latency lands once per morning in a place a person can absorb: you wait, you get coffee, you are mildly annoyed, and keeping the environment warm afterwards makes the problem disappear. That is why keep-warm is the correct engineering decision for a human creator, and why vendors who serve humans invest in resume latency rather than in the cold path. For an agent, the same latency lands inside every step of a loop, multiplied by the branching factor, and there is no warm environment to come back to because the agent will never use that specific environment again. A five-second create at a fan-out of four hundred is over half an hour of pure create time before any work happens. The second thing that changes is who notices. A human reports a slow create as a complaint on day one. An agent reports it as a slightly worse success rate three weeks later, spread across a thousand runs — approximately undetectable without a histogram you did not think to build.
Can one platform serve both humans and agents well?
Not without a compromise you should pick deliberately rather than discover. The two creators want opposite engineering investments. A human workspace product spends its effort on the editor, the resume path, the dotfiles, the port forwarding and the session — and it prices per seat, which is honest when concurrency is pinned to headcount. A sandbox API spends its effort on the cold create path, the teardown guarantee, the branching primitive and the metering, and it ships no editor because no program has ever wanted one. You can automate a workspace product, and you can SSH into a sandbox API with your own editor, and both of those work until the moment they are load-bearing. My honest recommendation, including against my own interest, is two products: a workstation for the people, a sandbox API for the fleet. That sounds like vendor sprawl and it is cheaper than the alternative, because the alternative is six months of building a shim and then paying seat prices for machines that no person ever opened.
What is the difference between fork() and fork_tree() on PandaStack?
They are different primitives and the distinction is load-bearing. `sbx.fork()` is a disk clone plus a cold boot: it reflinks the parent's rootfs using copy-on-write and then boots a fresh guest. The child inherits the filesystem — installed dependencies, build caches, files you wrote, local database files — and does not inherit the parent's memory. Nothing the parent had running is running in the child, the child draws its own entropy, and the latency has the roughly three-second cold-boot shape. `sbx.fork_tree(count)` is the memory-inheriting path: it restores the children from the parent's snapshot, so a dev server the parent had listening is still listening inside each child, at 400 to 750 ms per child on the same host and 1.2 to 3.5 seconds across hosts. The attribution matters because those numbers get paired with the wrong call constantly: 400 to 750 ms is the `fork_tree` restore path, not `fork()`. One consequence of exact memory inheritance is that the children inherit the random number generator's state too, so anything that seeds crypto or generates identifiers immediately after a `fork_tree` must reseed explicitly, or every sibling agrees on a value it believes is random.
Keep reading
- Top 9 ephemeral dev environment platforms — The same field split by product category instead — human workspaces, per-branch previews, sandbox APIs — with Render and Fly Machines covered.
- Top 6 ephemeral dev environments for AI agents — The agent-creator group from this post, graded in depth against a six-question rubric for a machine driver.
- Always-on vs scale-to-zero agent infra — The idle half of the concurrency inversion, taken much further, with the arithmetic of a warm floor.
- PandaStack pricing — One rate card for every class, hourly units, no per-request charge — the numbers behind the eleventh row.
Related posts
- Top 8 Ephemeral Development Environment Platforms (2026)
Feature grids do not decide this. Two numbers do: how fast a fresh environment can exist, and what happens to it when nobody is looking. Eight platforms graded on both, with the substrate each one actually isolates with and the catch I would want before the purchase order.
- Top 7 Ephemeral Development Environment Platforms in 2026
An environment is ephemeral when it is created from a definition, nobody is sad when it dies, and the 400th costs the same as the 4th. Most "cloud dev environments" fail at least one of those tests.
- Top 10 Disposable Development Environment Platforms in 2026
Creating environments is the easy half. This is a roundup graded on the hard half: what actually gets destroyed when you destroy one, what survives that you did not plan for, and who else feels it.
- Top 7 Disposable Postgres Platforms for Developers (2026)
Three of the seven ways to get a disposable Postgres are not products you buy. Here are all seven shapes graded on time-to-DSN, isolation boundary, whether they branch real data, and the failure mode each one actually has.
- PandaStack vs Coder
These two products are shopped against each other constantly and compete almost never. One gives a person a machine for the week; the other gives a program a machine for four seconds. Naming that split is most of the decision.
More in AI agent sandboxes · See AI agent sandboxes on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.