Top 12 Ephemeral Development Environment Platforms in 2026: The Cost of the Second One
I build PandaStack, an open-source Firecracker microVM platform, which means I spend an unreasonable amount of time watching people evaluate ephemeral development environments. The evaluation is always the same. Someone opens one environment. It works. It is fast, the editor is attached, the repo is cloned, there is a terminal. Everybody nods. The purchase order moves.
The first environment proves nothing. Every product in this category can produce one good environment on demand — that is the demo, and the demo has been solved since about 2020. The interesting question is the second one, and then the fifteenth, and then the fourteen you opened on Tuesday and have not thought about since.
Because the actual business all twelve of these companies are in is selling you a computer you are going to abandon. That is the product. Not the editor, not the terminal, not the devcontainer spec — the abandonment. And almost all of the engineering that separates these platforms from each other lives in what happens after you stop paying attention: who notices, how fast, what it costs in the meantime, and what of yours is destroyed when something finally does notice.
Why the second environment is the only one worth evaluating
Here is the failure mode I see most often, and it has nothing to do with performance. A team adopts a cloud dev environment platform. It is good. Engineers like it. Nine months later the finance review surfaces a line item nobody can explain, and the investigation finds forty-one workspaces belonging to eleven people, six of whom have left, several holding a disk that has been billing since March for a branch that merged in April.
Nothing broke. Every component behaved exactly as designed. The platform created environments on demand, as advertised, and then applied its retention policy, as advertised, and the retention policy happened to be a number in a settings page that one admin set once in a hurry. That is the product working. The bill is not a bug report; it is a design document you did not read.
So the question to put to a vendor is not "how fast does one start." It is: I want fifteen of these at once, this afternoon, from a script. What does the fifteenth cost me, how do the first fourteen stop costing me, and when they stop, what of mine goes with them? A platform that answers all three crisply is a platform whose abandonment semantics someone designed on purpose. Most of the answers are a timer.
The four axes, chosen before the list
- Isolation boundary. A container on a shared kernel, a Kubernetes pod on a shared node, a WASM runtime in a browser tab, or a VM with its own kernel. This is the one axis you cannot change later, and it is the one that decides whether "run the test suite from this pull request" is a routine operation or a trust decision. A container is a polite suggestion to a kernel you share with strangers.
- Cold-start shape. Not a number — a shape. "Warm if a prebuild hit, otherwise several minutes of image build" is a different product from "the same couple of hundred milliseconds every time, because there is no warm path to miss." Averages hide the shape, and the shape is what your engineers feel on the Monday after a dependency bump invalidated every cache you had.
- Idle cost model. What an existing-but-unused environment costs per hour, and who is holding that cost: the vendor's meter, your cloud account, or a cluster node your platform team provisioned in advance. Compute and storage are usually metered separately and only one of them stops when the environment stops.
- State survival. When the thing is reclaimed, what happens to the uncommitted diff, the scratch file, the database you seeded by hand? Platforms split sharply here, and the split does not correlate with price. Some keep a disk forever and bill you for it. Some destroy everything and are honest about it. The dangerous ones are in between, where state survives some kinds of death and not others.
Teardown semantics sit across all four and deserve naming on their own: deletion is not a transaction anywhere in this category. Compute stops, a disk lingers, a DNS record outlives the thing it pointed at, a cloud load balancer survives the namespace that created it. The platforms differ mostly in how many of those leftovers they take responsibility for.
Cohort one: the metered workspace, where a timer owns the abandonment
These four own the computer, meter it, and manage your forgetting with an idle timeout plus a retention window. The abandonment is a billing event, which is at least legible. What makes this cohort tricky is that compute and storage are different meters with different stop conditions.
1. GitHub Codespaces
Underneath: a devcontainer-defined Linux container running on managed cloud hosts, with a shared kernel. The definition lives in your repo, prebuilds are GitHub Actions jobs that bake the image ahead of time, and the whole thing is wired into the place your code already is — which is most of the value and is not a technical property at all.
Who it fits: teams already standardised on GitHub who want onboarding to be "open this repo" and are willing to maintain prebuild workflows. The cold-start shape is bimodal and the prebuild is what moves you to the good mode.
Honest drawback: a stopped codespace is not a free codespace. Compute stops on the idle timer; the disk keeps existing, and keeps billing, until a retention policy deletes it — and the retention policy is the one setting that determines whether forty abandoned workspaces are a rounding error or a line item. State survival is genuinely good here, which is exactly why the idle bill is the thing to watch.
2. Gitpod / Ona
Underneath: containerised workspaces from a repo-committed definition. The important thing is the identity shift — Gitpod's newer generation, now carrying the Ona name, runs runners inside your own cloud account rather than on the vendor's hosted platform, with an increasingly agent-shaped product on top. Check which generation a tutorial is describing before you follow it; the older managed service and the newer self-hosted-runner model are different purchases with different bills.
Who it fits: teams that want a polished workspace product but need the compute inside their own perimeter for data or compliance reasons.
Honest drawback: moving the runners into your account moves the idle cost there too. You are no longer buying a workspace-hour; you are buying a workspace product and separately paying your cloud provider for whatever it leaves running. That is often the right trade, but it is a trade, and the churn in positioning means your evaluation notes go stale fast.
3. Google Cloud Workstations
Underneath: a managed service that runs your container image on a Compute Engine VM, one VM per workstation, with a persistent home disk and a configurable idle timeout. It is the most honest product in cohort one about what it is, because the pricing makes no attempt to hide that you are renting a VM.
Who it fits: GCP-committed organisations that want workstations inside their VPC, under their IAM, with their org policies applying — the compliance story is the reason to pick it.
Honest drawback: it is priced like a VM because it is a VM, so the marginal cost of the fifteenth is linear and unforgiving, and start time is VM-shaped unless you pay to keep capacity ready. If your usage pattern is "many short-lived environments," this is the cohort member that punishes it hardest.
4. Replit
Underneath: container-based environments with a declarative package layer, in a product that has spent the last two years becoming an AI app-building platform with a dev environment inside it rather than a dev environment with AI features. That reframing is the most important fact about evaluating it.
Who it fits: prototyping, teaching, demos, and the genuinely strong case of handing a non-engineer a working environment with no local setup at all. The browser-first path from nothing to running code remains among the best in this list.
Honest drawback: the isolation boundary is a shared kernel, and the product's centre of gravity has moved away from "the place my team's repo lives." If your requirement is running code you did not write, under a boundary you can describe to a security reviewer, this is not the cohort and that is not a criticism of the product.
Cohort two: you own the abandonment
These four give you the control plane and leave the substrate — and therefore the idle bill — to you. The honest framing is that you are not buying environments, you are buying a scheduler and a UX, and the computer is still yours to pay for. That is frequently correct and almost never cheaper by accident.
5. Coder
Underneath: a self-hosted control plane that provisions workspaces through Terraform, which means the substrate is whatever you write in the template — Kubernetes pods, cloud VMs, your own hardware. It is the most flexible entry in this list and the flexibility is the entire point: enterprises adopt it precisely because the workspace can be a VM in a subnet their auditors already approved.
Who it fits: organisations with a platform team, hard network requirements, and existing Terraform competence.
Honest drawback: you are the SRE for your own dev environments, and the quality of your answer to "what does the fifteenth cost" is the quality of the template someone wrote in a sprint eighteen months ago. Idle semantics are yours to implement, which means they are yours to get wrong.
6. DevPod
Underneath: a client-side tool, not a service. It reads the devcontainer spec and starts the environment through a provider — local Docker, SSH to a box, Kubernetes, a cloud VM — with no vendor control plane in the middle. Open source, and the architectural choice is genuinely interesting: there is no central thing to run, so there is no central thing to pay for.
Who it fits: individuals and small teams who want devcontainer portability without a platform, and anyone whose real requirement is "never be locked into a workspace vendor again."
Honest drawback: no control plane means no central quota, no central policy, and no central view of what exists. The abandonment problem does not disappear, it federates — fifteen forgotten environments scattered across eleven laptops and three cloud accounts, which is strictly harder to audit than fifteen rows in one vendor's dashboard.
7. Okteto
Underneath: Kubernetes, honestly and deliberately. An environment is a namespace on your cluster, containers run as pods on shared nodes, and the development loop is file sync into a running pod so your inner loop happens where your dependencies already are. For a team whose application genuinely only makes sense as a dozen interacting services, this is the cohort that does not pretend otherwise. Garden sits in adjacent territory with a different emphasis on the build-and-test graph; if you are shopping here, look at both.
Who it fits: teams with a real cluster, a real platform team, and an application whose shape is already Kubernetes.
Honest drawback: you inherit every Kubernetes semantic, including the teardown ones. Namespace deletion is asynchronous and can wedge on a finalizer; persistent volume claims and their reclaim policies decide whether your state survives in a way nobody remembers configuring; and the nodes bill whether or not a single environment exists on them, so "idle cost" is a step function in node-sized increments rather than anything per-environment.
8. Daytona
Underneath: re-check this one before you quote it to anyone. Daytona began as a self-hosted manager for human development environments and has moved decisively toward being runtime infrastructure for AI agents, with an API-first sandbox shape. The engineering is good and the repositioning is real, so a 2024 blog post about it is describing a product that no longer exists in that form.
Who it fits: teams whose creator is a program rather than a person — agent fan-out, automated code execution — and who want a sandbox API rather than a workspace UI.
Honest drawback: if you came here for a human workspace manager, the thing you read about may not be the thing you can buy. That is not dishonesty, it is a company following its market, but it makes evaluation notes expire unusually fast. I keep it in cohort two because self-hosting remains a first-class path.
Cohort three: abandonment is cheap by construction
The last four are grouped by a structural property rather than a pricing one: there is either nothing to abandon, or abandonment is the normal state and existence is the exception. This is the cohort where the marginal cost of the fifteenth environment stops being linear in anything you care about — and where the trade-offs get sharper, not softer.
9. StackBlitz (WebContainers)
Underneath: not a server. WebContainers runs a Node-compatible runtime compiled to WebAssembly inside the browser tab, with a virtual filesystem and a service-worker-backed network shim. The isolation boundary is the browser's origin sandbox, which is a boundary a very large number of people have attacked professionally for twenty years, and your code never leaves the laptop.
Who it fits: documentation that runs, bug reproductions, teaching, and framework playgrounds. The cold start is the best in this list by a wide margin for the simple reason that nothing is provisioned. Idle cost is genuinely, structurally zero. There is no meter because there is no computer.
Honest drawback: it is not Linux. No native binaries, no Docker, no arbitrary processes, no system package manager, and a dependency with a compiled component is where the demo stops. "Nothing to abandon" and "nothing to keep" are the same sentence: close the tab and the environment is gone, which is a feature for a reproduction and a disaster for a long-running branch.
10. CodeSandbox
Underneath: two products under one name. There is a browser-based sandbox lineage, and there is a cloud development environment built on microVMs that snapshots memory and resumes — so reopening a branch restores a machine that was already warm rather than booting one. Architecturally this is the most interesting entry in the list after the one I wrote, because it is the only other member that treats resume-from-memory as the default path rather than an optimisation.
Who it fits: frontend and full-stack teams who want a branch-per-environment review flow with an editor attached, and who value "pick up exactly where I left off" over substrate control.
Honest drawback: the resume model trades compute for storage — every hibernated environment is a memory image someone is storing, and a team that forks liberally is a team accumulating snapshots. It is also an opinionated product rather than a primitive: you get their editor and their flow, which is the point, and is a ceiling if you wanted an API you could build something else on.
11. CloudShell and the Cloud9 successors
Underneath: a managed shell session on a vendor-run host, with a small persistent home directory and a session that evaporates when you stop using it. AWS CloudShell is the archetype and the other large clouds ship an equivalent. The category matters here mostly for a historical reason: AWS Cloud9 was the browser IDE a lot of teams bet on, and AWS stopped taking new customers for it — the replacement story is a shell plus VS Code attached over a remote-access agent, not a hosted IDE.
Who it fits: operational work. A credentialed shell in the right account and region, one click from the console, with nothing to provision and nothing to clean up. For "I need to run three commands against this account right now," nothing in this list is better.
Honest drawback: it is a shell, not a development environment. Resource limits are modest, the home directory is small, long-running processes are not the contract, and the session is designed to die. The useful lesson from Cloud9's retirement is a procurement one: a vendor-owned IDE is a product decision you are renting, and the environment definition living in your repo is what makes a platform swap a Tuesday instead of a quarter.
12. PandaStack
Mine, so here are the numbers instead of adjectives. Underneath: one Firecracker microVM per environment, with its own kernel — not a namespace on a kernel you share. Every create is a snapshot restore rather than a boot, and there is no warm pool anywhere in the design, so the cold-start shape is flat: p50 179 ms and p99 203 ms end to end, measured, for the first environment and the fifteenth alike. The first spawn of a brand-new template does a real cold boot and bakes the snapshot, which takes about 3 seconds; after that, every create takes the restore path.
Each environment gets a pre-allocated Linux network namespace with a veth pair and its own tap device, from a pool of 16,384 slots per host — the hard per-host ceiling, though in practice you run out of memory long before you run out of subnets. Allocation from the warm pool is about 9 ms; when the warm depth drains, the agent builds a slot from scratch in roughly half a second rather than refusing the create.
Pricing is one rate card for every class: $0.054 per vCPU-hour and $0.0162 per GiB-hour. CPU bills on CPU-seconds actually burned, not on vCPUs provisioned, so an idle environment's compute line really does go to approximately nothing. Memory bills as committed GiB-hours for as long as the environment exists, which is the honest half of the sentence. There is no per-request pricing and there never will be.
Who it fits: when the creator is a program. Per-pull-request environments created by a webhook, CI jobs that need a real kernel, agents fanning out fifteen branches at once, anything running code whose author you do not know. The isolation boundary is the reason to pick it.
Honest drawback, and it is a big one: there is no browser IDE. Codespaces, Gitpod and CodeSandbox have an editor story and a human-onboarding story that I do not have. This is a compute primitive with dev-environment uses, driven by an SDK and a REST API, with tokenless preview URLs for whatever ports your environment exposes. If what you want is "my team opens their laptop and clicks into a workspace," cohort one is the right answer and I will say so on a call.
The twelve on four axes
| Platform | Isolation boundary | Cold-start shape | Idle cost model | State survival |
|---|---|---|---|---|
| GitHub Codespaces | Container, shared kernel | Fast on prebuild hit, slow on miss | Compute stops on timer; disk keeps billing | Disk survives stop, dies at retention |
| Gitpod / Ona | Container, your cloud account | Prebuild-dependent | Your cloud bill, not a workspace-hour | Depends on your runner config |
| Google Cloud Workstations | VM per workstation | VM-shaped unless you pay to stay warm | Priced like a VM; idle timeout stops it | Persistent home disk |
| Replit | Container, shared kernel | Fast, image-cached | Bundled in plan; sleeps when unused | Workspace files persist |
| Coder | Whatever your Terraform says | Whatever you provisioned | Your infrastructure, your timers | Your template decides |
| DevPod | Provider-dependent | Devcontainer build, locally cached | Whoever's account it ran in | Provider-dependent |
| Okteto | k8s pod, shared node kernel | Schedule plus image pull | Nodes bill whether or not it exists | PVC reclaim policy decides |
| Daytona | API-driven sandbox; re-check current docs | API-fast | Per-sandbox, program-driven | Snapshot-dependent |
| StackBlitz | Browser origin sandbox (WASM) | Near-instant; nothing is provisioned | Structurally zero — no server exists | None; close the tab and it is gone |
| CodeSandbox | microVM with memory resume | Resume from a warm memory image | Hibernates; you store the snapshot | Memory and disk both resume |
| CloudShell-class | Managed shell host | Session start | Free tier, session-bounded | Small home directory only |
| PandaStack | Firecracker microVM, own kernel | Flat: 179 ms p50 restore, no warm pool | Committed GiB-hours while it exists; CPU on seconds burned | Destroyed at reap unless persistent |
The arithmetic of fifteen at once, and fourteen forgotten
Let me do my own homework out loud, because a roundup that will not price its own product is an advertisement. The `base` template bakes at 4 GiB of RAM and 8 vCPU. Fifteen of those existing simultaneously and doing nothing costs fifteen times 4 GiB times $0.0162, which is about $0.97 per hour in memory, plus close to nothing in CPU because CPU bills on seconds actually burned. Leave all fifteen running and ignored for an eight-hour working day and that is roughly $7.80 of pure forgetting.
So my idle cost is not zero, and I would rather say that plainly than let you discover it. What makes it cheap is not the rate, it is the duration: the same fifteen environments with a 15-minute idle TTL cost about 24 cents in total, because the reaper destroys them a quarter of an hour after anyone last touched them. The lever is not the price per GiB-hour. The lever is that the environment stops existing.
Which makes the exact semantics of that timer the most load-bearing detail in this whole post, and there are three things about it worth knowing.
First, `ttl_seconds` is an idle clock, not a wall clock. The reaper compares now against the sandbox's last activity, not against its creation time, so a 15-minute TTL does not kill a sandbox you are using at minute 16. There is a separate hard wall-clock lifetime cap applied on the free tier, because an idle clock cannot bound a continuously busy sandbox — which we learned when an account ran a self-built template for ten and a half hours without ever going idle.
Second, and this is the detail I am most pleased with: observing a sandbox does not keep it alive. Both the activity bump and the auto-wake are gated on the same predicate — "does this request actually use the guest?" A read-only inspection GET for status, lifecycle, metrics, logs, events or a directory listing does neither. Without that gate the idle TTL silently never fires, because a dashboard polling every thirty seconds with a tab open resets the clock faster than the clock can ever elapse. We shipped that bug, measured `idle_seconds` stuck at zero on sandboxes that were never going to be reaped, and fixed it. Genuine use — exec, byte reads from the filesystem, a PTY, a fork, a snapshot — still bumps and still wakes.
from pandastack import Sandbox
# Auditing the fourteen you forgot about. Reading this does NOT reset
# the idle clock and does not wake a hibernated sandbox — that gate is
# the whole reason the TTL can ever fire.
for sbx in Sandbox.list():
info = sbx.lifecycle()
print(
sbx.id,
f"idle={info['idle_seconds']}s",
f"ttl={info['ttl_seconds']}s",
f"persistent={info['persistent']}",
sbx.metadata.get("branch", "-"),
)
# Shorten the leash on anything still breathing from last week.
if not info["persistent"] and info["idle_seconds"] > 3600:
sbx.set_ttl(900)
Third, an idle auto-reap deliberately does not cascade. Deleting a sandbox by hand removes its snapshots; the reaper's path does not, because a snapshot is durable and is supposed to outlive the sandbox that produced it. So the pattern that actually saves uncommitted work across an abandonment is to snapshot before you walk away and let the VM die — the snapshot is still there tomorrow, and restoring from it is the same restore path as any other create.
One sizing caveat, since this is the question that follows: Firecracker cannot change vCPU or RAM at snapshot restore, so for any template with a baked snapshot the `cpu` and `memory_mb` you pass to create are overridden to match the snapshot before anything persists — the API response, the database row and the billing event all show the corrected values rather than your wish. Disk is the exception: a per-request size can grow past the template default, never shrink. If you need a different memory size, you need a different template, not a different argument.
Fifteen environments, and the fourteen you forget
The shape that survives contact with reality has three defences, in this order: a TTL set at create time so the platform cleans up after your process dies, an explicit teardown in a `finally` so it normally never gets that far, and an in-guest timeout so a wedged command cannot hold the environment open indefinitely.
import concurrent.futures as cf
from pandastack import Sandbox
from pandastack.exceptions import PandastackError
REPO = "https://github.com/acme/widget.git"
BRANCHES = [f"feat/{n}" for n in range(15)]
def check(branch: str) -> tuple[str, int]:
# ttl_seconds is the backstop, not the plan: an IDLE clock, so a busy
# sandbox is not killed mid-test. If this process is OOM-killed before
# the finally runs, the reaper collects the VM 15 minutes later.
sbx = Sandbox.create(
template="base",
ttl_seconds=900,
metadata={"branch": branch, "owner": "pr-bot"},
)
try:
rc = sbx.exec_stream(
f"git clone --depth 1 --branch {branch} {REPO} /src",
timeout_seconds=120,
)
if rc != 0:
return branch, rc
# The real bound is in-guest: timeout(1) kills the build itself.
# exec_stream only widens the client's HTTP timeout -- neither path bounds the command, so stream long work and fence it in-guest.
logs: list[str] = []
rc = sbx.exec_stream(
"cd /src && timeout --kill-after=5s 600 make test",
on_stdout=logs.append,
on_stderr=logs.append,
timeout_seconds=660,
)
return branch, rc
finally:
# Unconditional. Not in an except, not after the return.
sbx.kill()
with cf.ThreadPoolExecutor(max_workers=15) as pool:
futures = {pool.submit(check, b): b for b in BRANCHES}
for fut in cf.as_completed(futures):
branch = futures[fut]
try:
_, rc = fut.result()
print(f"{branch}: {'ok' if rc == 0 else f'FAILED rc={rc}'}")
except PandastackError as err:
# A create refused for capacity is a real 503. Retry with
# backoff or run a smaller fan-out; do not swallow it.
print(f"{branch}: platform error: {err}")
Three things in there are the post's argument in code. The `finally` is unconditional, because the single most common cause of a forgotten environment is an exception path that skipped the cleanup someone wrote inside an `except`. The TTL is a backstop for the case the `finally` never runs at all, which is what a `SIGKILL` on your CI runner looks like from the platform's side. And the real timeout is `timeout(1)` inside the guest, because the bound you want is on the work, not on the HTTP call that started it.
The capacity branch is not decoration. Per-environment VMs mean a create can genuinely be refused, and the honest answer to a 503 is backoff and a smaller fan-out, not a retry loop that hammers a full fleet. Memory admission on the fleet works against admittable memory rather than committed memory, which buys real density, and it is still finite.
The honest limits
- There is no browser IDE and no plan for one. This is a compute primitive with dev-environment uses, driven by an SDK, a CLI and a REST API. Codespaces and Gitpod sell a product to a human opening a laptop; I sell one to a program. For "onboard a new hire in ten minutes with an editor already open," cohort one wins, and pretending otherwise would waste a call we could both use better.
- A reaped sandbox loses uncommitted work. The idle reaper destroys the VM, and your unpushed diff goes with it unless you created with `persistent=True`, raised the TTL, or snapshotted first. I think that default is correct for ephemeral compute and it will still bite someone in their first week. The mitigation is cultural as much as technical: push, or snapshot, or accept the loss.
- The guest kernel is 5.10, on Ubuntu 24.04. Most toolchains do not care and some do — anything wanting a very recent kernel interface, exotic filesystems or specific container-in-container trickery is worth testing before you plan a migration around it. I would rather you hit that on a Tuesday afternoon than during a cutover.
- Per-environment VMs mean capacity can run out, which is a class of page you do not get from a browser-based environment at all. There are 16,384 pre-allocated subnet slots per host, but memory is the binding constraint long before that, and a 503 on create is a real thing you now have to handle in your orchestration code.
- Idle is cheap, not free. Memory bills as committed GiB-hours for as long as the environment exists, which is why the TTL is doing most of the work in my own arithmetic above. If your pattern is "fifty environments that each live for days," measure it honestly against a platform that stops compute and keeps only a disk — their model may genuinely beat mine for that shape.
- Egress is open by default. Sibling sandboxes cannot reach each other's subnets and the cloud metadata range is dropped in the host's forward chain, but there is no default-deny policy, and your own VPC and databases are not fenced for you. If your environments run code from pull requests you have not read, those deny rules are work you still own.
Picking one
Sort by who is calling create and what the isolation boundary has to be, then let the idle model break the tie.
- A human opening a laptop, repo on GitHub, trusted code: Codespaces, and spend the afternoon you saved on the prebuild workflow and the retention policy. The retention policy is the bill.
- A human opening a laptop, hard network or compliance perimeter: Coder if you have a platform team and Terraform, Google Cloud Workstations if you are GCP-committed and want someone else to run it, Ona if you want the polished product with runners in your own account.
- An application that is already a dozen services on Kubernetes: Okteto, or Garden if your pain is the build-and-test graph more than the inner loop. Budget for teardown semantics being Kubernetes teardown semantics.
- A reproduction, a doc that runs, a teaching environment: StackBlitz. Zero provisioning, zero idle cost, and the limits are honest and immediate rather than lurking.
- A branch-per-environment review flow where resuming warm matters more than substrate control: CodeSandbox.
- Three commands against one cloud account, right now: CloudShell. Do not overthink it.
- A program calling create — a PR webhook, a CI matrix, an agent fanning out branches — especially when the code came from outside: PandaStack or the agent-shaped cohort. Mine if you want a kernel boundary per environment, a flat restore-on-create cost and one rate card; the others if you want a managed sandbox API and can live on a shared kernel.
- A portable local-first setup with no vendor at all: DevPod, with your eyes open about federated abandonment.
The bottom line
Every platform in this category is in the abandonment business, and the engineering that distinguishes them is almost entirely post-abandonment: whose meter keeps running, which timer notices, and what of yours is destroyed when it does. That is why the first environment tells you nothing and the second one tells you everything.
So run the real test. Create fifteen from a script, walk away for a day, and read the bill. Then read the invoice line by line and find out which of the fifteen are still alive, what the dead ones left behind, and whether anything you cared about survived. That experiment costs a few dollars on any platform here and answers the question a feature grid cannot. Whatever surprises you will be a timer or a disk, and it will be documented.
My own answer is the narrow one: one microVM per environment, its own kernel, snapshot-restored in 179 ms whether it is the first or the fifteenth, an idle clock that cannot be reset by watching it, and an unconditional `kill()` that you should write anyway. No editor, no onboarding flow, no browser IDE. If you need those, I have named the platforms that have them, and I meant it.
Frequently asked questions
What actually makes an environment "ephemeral" rather than just remote?
Three properties, and most products marketed as ephemeral fail at least one. First, it is created from a definition rather than configured by hand — if reproducing it requires a person remembering what they installed, it is a pet with a cloud IP address. Second, nobody is sad when it dies, which is a statement about state: if an environment holds the only copy of something, its destruction is an incident and you will therefore stop destroying it. Third, the marginal cost of the four-hundredth is close to the marginal cost of the fourth, in both money and latency. That third one is where remote-but-not-ephemeral platforms break down: a VM-per-workstation model is linear and unforgiving, so a team that wanted an environment per pull request quietly becomes a team with four shared staging environments and a booking spreadsheet. The test I would apply on a trial is deliberately unkind. Create fifteen at once from a script, with no console clicking, then walk away for a day without cleaning up. If the script cannot create them, the platform is not programmable. If the bill the next morning is a surprise, you have found the idle model. If anything you cared about is gone, you have found the state model. All three answers arrive in twenty-four hours for a few dollars.
Is idle cost on PandaStack really zero?
No, and I would rather be precise than flattering. CPU is billed on CPU-seconds actually burned rather than on vCPUs provisioned, so an environment doing nothing really does stop generating compute charges — that part is close to zero. Memory is billed as committed GiB-hours for as long as the sandbox exists, at $0.0162 per GiB-hour, and the `base` template bakes at 4 GiB. So fifteen idle base environments cost roughly 97 cents an hour in memory, which over an ignored eight-hour day is about $7.80 of pure forgetting. What makes the model cheap is not the rate, it is the duration: with a 15-minute idle TTL those same fifteen environments cost about 24 cents total, because the reaper destroys them a quarter of an hour after anything last touched the guest. The lever is that the environment stops existing, not that existing is free. This is also why the TTL semantics matter so much. It is an idle clock rather than a wall clock, so a busy sandbox is not killed mid-build, and critically the activity bump is gated on whether a request actually uses the guest — read-only status, metrics and log requests do not reset it. We learned that one the hard way: before the gate existed, a dashboard tab polling every thirty seconds kept sandboxes alive forever.
Does my uncommitted work survive when an ephemeral environment is reclaimed?
It depends entirely on the platform, and this is the axis where the answers diverge most sharply without correlating with price. In the metered-workspace cohort, stopping is usually not deleting: the compute halts on an idle timer while the disk persists — and keeps billing — until a retention policy deletes it, so your diff survives a stop and does not survive retention. On Kubernetes-based platforms the honest answer is "whatever your persistent volume claim's reclaim policy says," which is a sentence nobody wants to discover during an incident. In a browser-based runtime there is nothing to survive at all; closing the tab is the teardown. On PandaStack the default is destructive: the idle reaper deletes the VM and unpushed work goes with it unless you created the sandbox with `persistent=True`, raised the TTL, or took a snapshot first. One deliberate nuance there is worth knowing: an idle auto-reap does not cascade-delete the sandbox's snapshots, because a snapshot is durable and is meant to outlive the VM that produced it. So the pattern that genuinely survives an abandonment is to snapshot and let the VM die, rather than to keep a VM alive for the sake of a diff.
Can I use PandaStack as a replacement for Codespaces or Gitpod?
Only if you do not need the editor, and I would push back on anyone trying. The honest division is about who calls create. Codespaces, Gitpod and the rest of the metered-workspace cohort sell a product to a human who opens a laptop: a browser IDE, an onboarding flow, a dotfiles story, port forwarding wired into the editor. I do not have any of that and I am not building it. What I have is a compute primitive — one Firecracker microVM per environment with its own kernel, snapshot-restored on every create at p50 179 ms with no warm pool, driven by a Python or TypeScript SDK and a REST API, with tokenless preview URLs for whatever ports the guest exposes. That is the right shape when the creator is a program: a pull-request webhook, a CI matrix, an agent fanning out fifteen branches, or anything running code whose author you cannot vouch for, where a shared kernel is the part your security reviewer will not sign. Plenty of teams run both, and that is not a hedge — a workspace product for people and a sandbox API for programs are genuinely different purchases that happen to share a category page.
Why does the isolation boundary matter if the code is ours anyway?
Because "the code is ours" has a shorter half-life than people expect. The moment an ephemeral environment builds a fork's pull request, runs a dependency's post-install script, or executes something a model generated, the author of the bytes is not on your payroll. A container on a shared kernel is a polite suggestion to that kernel, enforced by namespaces and cgroups that are features of a large, actively-researched attack surface — and a CI-shaped environment is usually the most credential-rich thing an organisation owns, which makes it the best target rather than the worst. A microVM moves the boundary to the hypervisor: a separate kernel per environment, a device model measured in a handful of emulated devices instead of a whole kernel's syscall table. That said, the boundary is not a whole security posture. Egress on PandaStack is open by default — sibling sandboxes cannot reach each other's subnets and the cloud metadata range is dropped at the host, but your own VPC, databases and internal APIs are not fenced for you. If your environments run code you have not read, writing those deny rules is still your work, and no substrate choice does it for you.
Keep reading
- Cloud dev environments on microVMs — The substrate argument in full: why a kernel per environment changes what you can safely run in one.
- Controlling sandbox lifetime with TTL and idle timeouts — The idle clock in detail, including what counts as activity and what deliberately does not.
- Top 10 disposable dev environment platforms — The same category graded on teardown correctness rather than marginal cost.
- PandaStack vs GitHub Codespaces — The head-to-head, including the editor story I do not have.
- The carbon cost of idle compute — What the fourteen forgotten environments cost in something other than dollars.
Related posts
- Top 7 Ephemeral Development Environment Platforms in 2026
An environment is ephemeral when it is created from a definition, nobody is sad when it dies, and the 400th costs the same as the 4th. Most "cloud dev environments" fail at least one of those tests.
- Top 8 Ephemeral Development Environment Platforms (2026)
Feature grids do not decide this. Two numbers do: how fast a fresh environment can exist, and what happens to it when nobody is looking. Eight platforms graded on both, with the substrate each one actually isolates with and the catch I would want before the purchase order.
- Top 11 Ephemeral Development Environment Platforms (2026)
Every roundup in this category sorts by substrate or by price list. Both are downstream of a question nobody asks on the evaluation call: who, or what, is going to call create? Eleven platforms sorted by their creator — human, CI job, pull-request webhook, agent API — because the creator decides how many environments exist at once and who pays for the idle ones.
- The best Gitpod alternatives in 2026
Gitpod taught everyone to expect a dev environment from a URL. Picking a replacement means deciding which half of that promise you actually needed: the ephemerality, or the machine.
- A Disposable Dev Environment per Git Branch
Every branch gets its own fully-isolated microVM dev environment — checkout done, deps installed, services running — that forks in sub-second and disappears when the branch is gone. Ephemerality is what finally kills environment drift.
More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.