all posts

Top 7 Ephemeral Development Environment Platforms in 2026

Ajay Kumar··11 min read

Every engineering organisation past about thirty people has a machine with a name. Not an ID — a name. It is called something like dev-box-03, it has been running since a reorg two cycles ago, and it holds ninety days of uncommitted work belonging to someone who left in March. Once a quarter it kernel-panics in a way nobody can reproduce, and the fix is documented as 'ask Priya, she knows the trick.' There is a line item for it in the cloud bill that finance has stopped asking about because the last three people who asked got an explanation involving the word 'legacy.' It is not a dev environment. It is a pet, with a monthly stipend and a medical history.

The promise of the cloud-dev-environment category was that this would stop happening. Mostly it did not: a lot of teams moved the pet into a managed product and kept feeding it, because the platform happily let them. So this is a roundup of the category graded on the one property that actually prevents the pet — ephemerality — rather than on who has the nicest web editor.

A note on scope, because this site has a lot of adjacent pages and I would rather not waste your time. This is the category-wide comparison. If you have already decided to leave one specific vendor, the migration-shaped posts are better reading: there are dedicated pages on Codespaces alternatives, Gitpod alternatives, Daytona alternatives, and the CodeSandbox and StackBlitz cases, all linked at the bottom. If you want the mechanics of running this on microVMs rather than the vendor landscape, there is a post for that too. This one is the grading rubric and the field.

Disclosure, because it changes how you should read the rest: I'm Ajay, I build PandaStack, and PandaStack is entry seven. The rules I hold myself to in a roundup: specific numbers only for my own product, every other platform described qualitatively from its public documentation with no invented pricing, limits or internals, and a standing instruction to verify anything I say about someone else against their current docs — this category reshuffles faster than blog posts get updated. The section where I name what PandaStack is bad at is not a rhetorical device; it is the section I would read first if I were you.

What makes an environment ephemeral rather than merely remote

'Remote' and 'ephemeral' get used interchangeably and they are completely different properties. Remote means the CPU is somewhere else. Ephemeral means the environment has no history worth preserving. You can have a remote environment that is the most pet-like object in your entire infrastructure, and most teams do. Three tests, and an environment has to pass all three:

  1. It is created from a definition, not restored from a backup. The question to ask is: if this environment vanished right now, what would we read to rebuild it? If the answer is a file in the repo — a devcontainer spec, a Dockerfile, a .tool-versions, a Nix flake — it is ephemeral. If the answer is 'a snapshot of the disk as it was' or, worse, 'Priya's memory', it is a pet wearing a hostname. The definition is the environment; the running machine is a cache of it.
  2. Nobody is sad when it dies. This is a social test more than a technical one, and it is the one that actually predicts behaviour. If a developer would say 'wait, don't kill that, I had something in it', the environment has become stateful in a way nobody designed. The fix is not better backups. The fix is making creation so cheap that keeping things never becomes the default, so that the only state anyone trusts is the state in git.
  3. The 400th one costs the same as the 4th. Linear cost per environment is the whole point: it is what lets you give every branch, every pull request, every agent task its own machine instead of queueing them onto a shared one. The moment the marginal environment is expensive — in dollars, in a capacity ceiling, or in the thirty seconds a developer spends waiting — people start reusing, and reuse is how drift gets in.

Test three is where architecture leaks into product. A platform that keeps idle machines warm so it can hand you one quickly has to pay for those machines whether or not anyone is using them, and that cost comes back to you as a seat price, a concurrency cap, or an aggressive idle timeout that kills your environment mid-thought. A platform that can build an environment from cold in a time measured in milliseconds does not need a warm pool at all. Same user-visible feature, completely different economics, and the difference shows up in your bill rather than in the docs.

The seven criteria, before the list

Deciding the criteria after you have seen the products is how you end up buying the thing with the best landing page. So: here is what I grade on, in the order that usually decides the outcome.

  1. Cold-start-to-usable. Not to 'provisioning', not to 'container running' — to the moment you can type a command and get an answer. This number sets the lifetime model you can afford. A few hundred milliseconds and you can afford to destroy environments per branch, per task, per lunch break. Thirty seconds and you will keep them warm forever, and you will have chosen reuse for a latency reason rather than a design one.
  2. Isolation boundary. Shared kernel or your own? Most of the category runs your environment as a container on a multi-tenant host, which is fine right up until the moment you remember what a dev environment does: it runs npm install, which runs arbitrary postinstall scripts written by strangers, on a machine that also holds your cloud credentials. A shared kernel is a shared fate, and the thing standing between a stranger's postinstall script and your credentials is a configuration — configurations have typos.
  3. Statefulness. Three sub-questions that get conflated: can you suspend and come back with your processes still running? Does the filesystem survive a stop? And does any of it survive the platform garbage-collecting the host? Teams routinely assume all three and get only the second.
  4. How the environment is defined. Devcontainer JSON, a plain Dockerfile, a Nix expression, a bespoke YAML dialect, or a snapshot. This is your lock-in surface and your drift surface at the same time. Prefer a definition format you could run somewhere else, and prefer one your repo already contains.
  5. Idle behaviour. What the platform does when nobody has typed for twenty minutes, and what it charges you while it does it. Options observed in the wild: nothing (you pay), suspend-to-disk (you pay less, usually storage), stop (you pay storage, you lose processes), delete (you pay nothing, you lose everything not committed). The honest ones document which, with a timeout you can configure.
  6. Nested workloads. Can you run Docker, a local Kubernetes, a database container, a VM? This is not a nice-to-have for anyone whose test suite uses testcontainers. On a shared-kernel platform this requires either a privileged container or a daemon-in-a-daemon workaround; with your own kernel it is ordinary nesting.
  7. Who operates it, and where it runs. Fully managed SaaS, self-hosted in your cloud account, or a client-side tool that drives infrastructure you already own. This decides your compliance conversation and your worst-case bill, and it is the criterion people decide first and should decide last.

The seven

Numbered for the headline's benefit. This is not a ranking — the right answer genuinely inverts depending on which of the seven criteria is non-negotiable for you, and I have ordered these by shape rather than by merit. Everything about a platform that isn't mine is qualitative and should be checked against their current documentation before you commit.

1. GitHub Codespaces

What it is: a managed development environment attached directly to a GitHub repository, defined by the devcontainer specification, reachable from a browser VS Code or from your desktop editor over a remote connection. It is the default that most teams evaluate first because it is one click away from where the code already lives.

Genuinely best at: the gap between 'I have a repo' and 'I have a working environment' being nearly zero for someone who has never seen the project. Onboarding, drive-by contributions, reviewing a pull request in a real environment rather than a diff viewer. The devcontainer format it popularised is also the most portable environment definition in this list, which matters more than people credit — your devcontainer.json is the one asset here you can take with you.

Where it bites: it is the easiest platform in the category to accidentally turn into a pet. The environment persists, it retains your uncommitted work, and that convenience is precisely the mechanism by which drift accumulates. Teams end up with long-lived Codespaces holding state nobody has committed, which is the problem they adopted it to solve, and the isolation model is a container rather than a dedicated kernel. Verify current idle and retention behaviour, cost model and machine-size options against GitHub's docs before you plan around them.

2. Gitpod

What it is: the project that made the strong-ephemerality argument before the rest of the category caught up — environments as something you start per task from a declarative config, prebuilt ahead of time so the first command isn't a dependency install. Gitpod's product direction has moved significantly in recent years toward running environments on infrastructure you control, so treat any description of it, mine included, as a snapshot.

Genuinely best at: the prebuild idea, which is the single most useful concept this category produced. Build the environment when the commit lands, not when a developer asks for it, and the cold-start problem partly evaporates. If you take one thing from Gitpod into whatever you end up using, take that.

Where it bites: it is the entry whose shape has changed the most, so anything you read about it that is more than a few months old is probably describing a product that no longer exists in that form. Check the current deployment model, the supported config format, and what happens to environments on idle directly from their documentation — not from a roundup, including this one.

3. Coder

What it is: self-hosted development environments where the environment shape is expressed in Terraform. That one design decision explains the whole product: a workspace can be a Kubernetes pod, a cloud VM, a bare-metal box, whatever your Terraform provider can create, and the platform's job is lifecycle and access rather than owning the compute.

Genuinely best at: the enterprise case, honestly and without irony. If your environments must live inside your own VPC, on your own subnets, under your own IAM, with your own audit trail and a security team that wants to read the config, Coder is the entry that treats that as the primary requirement instead of an enterprise-tier afterthought. Terraform as the definition language is also a real advantage for a platform team that already speaks it.

Where it bites: a Terraform-defined workspace tends toward a long-lived workspace, because the underlying resource is usually a VM or a pod with a disk and the cost of recreating it is a full provision cycle. It is the most powerful entry here and also the one where you most easily end up with per-developer pets, just pets you provisioned declaratively. You are also running the control plane. Verify current workspace lifecycle, autostop behaviour and licensing tiers against their docs.

4. DevPod

What it is: a client-side tool, not a service. It reads the devcontainer spec and creates that environment on a provider you choose — your laptop, a cloud VM, a Kubernetes cluster — with no central control plane in the middle. Think of it as the devcontainer runtime decoupled from any particular vendor's hosting.

Genuinely best at: being the un-lock-in option. If your concern is that you are about to standardise your entire engineering org's inner loop on one vendor's proprietary environment format, a client that runs the open devcontainer spec anywhere is a strong answer to that specific fear. It is also the cheapest way to get the devcontainer workflow onto hardware you already pay for.

Where it bites: no control plane means no control plane. There is no shared place where environments are registered, policy is applied, costs are attributed, or an admin can see what exists — those become your problem, which is fine for a ten-person team and becomes a platform project at two hundred. Confirm the current provider list and the project's maintenance status before you build process around it.

5. Daytona

What it is: originally a developer-workspace manager, now positioned largely around fast sandboxes for AI agents — which tells you something true about where this entire category is heading. The interesting thing about Daytona for a human-dev-environment buyer in 2026 is that it has mostly stopped being that product, and the reason is instructive rather than embarrassing.

Genuinely best at: the agent case it pivoted into. An agent needs an environment created and destroyed programmatically, in large numbers, with no editor attached, and it does not care about tab completion. That is a different product from a human workspace even though the infrastructure underneath is nearly identical, and Daytona chose the side of that fork with more growth in it.

Where it bites: if you arrived looking for a seat-based IDE-attached workspace for sixty engineers, check carefully that the current product is still shaped like that before you plan a rollout. Its isolation model and lifecycle semantics should both be read from current docs rather than from anybody's comparison table — this entry in particular has moved.

6. Browser-native: CodeSandbox and StackBlitz

What they are: two different technical bets that land in the same product category. StackBlitz pioneered running a Node-compatible runtime inside the browser tab itself via WebAssembly, so for suitable projects there is no server in the loop at all. CodeSandbox runs real VMs with the ability to resume a previously running environment, so the editor is in the browser but the compute is not. Both have evolved well past 'frontend playground', and both now have programmable APIs alongside the editor — verify which product you would actually be buying.

Genuinely best at: the cold-start experience nothing else in this list can match, because for the browser-runtime case there is nothing to start. A reproduction link that opens into a running app in a tab, with no account and no provisioning, is a genuinely superior artifact for bug reports, documentation, teaching and code review. If your job is to get a stranger to a running version of your code in one click, this is the answer and the rest of this list is not close.

Where it bites: the in-browser runtime is a compatibility boundary, not a Linux box. Native modules, arbitrary binaries, Docker, a Postgres you can apt-get — these are the things that make a dev environment a dev environment for a backend team, and a browser tab is not where they live. Check which of the two execution models the product you're evaluating actually uses for your project, because the answer determines whether half your toolchain works.

7. PandaStack

What it is, with the disclosure already made: an open-source platform that gives each environment a Firecracker microVM with its own guest kernel, created by restoring a baked snapshot rather than booting. Create latency is 179 ms at p50 and 203 ms at p99, of which the snapshot load itself is about 49 ms. The first spawn of a template before its snapshot exists cold-boots in roughly 3 s and bakes the snapshot on the way through; everything after that takes the fast path. Networking is per-environment — its own Linux network namespace, veth pair and tap device, drawn from 16,384 pre-allocated /30 subnets per host — so per-environment egress policy is a property of the machine, not a rule someone has to remember to write.

Genuinely best at: three things that follow from snapshot-restore. First, there is no warm pool, so there is no idle fleet whose cost finds its way into your seat price — recreating is cheap enough that deletion is a reasonable idle strategy rather than a loss. Second, forking: a copy-on-write clone of the environment's disk in 400–750 ms on the same host, 1.2–3.5 s across hosts. Be precise about what that does and does not give you, because the word 'fork' oversells it — the clone inherits the filesystem, so your working tree, your installed dependencies, your build caches and your local database files all come along, but it boots fresh, so the parent's running processes do not. You restart your dev server in the clone. That is still a thing no container-based dev-environment product offers: a sub-second branch of a twenty-gigabyte working environment, which turns 'try this destructive migration against a copy of my real data' from an afternoon into one line. Third, the kernel boundary, which matters more for dev environments than for almost anything else, because a dev environment's whole job is to execute dependency postinstall scripts written by people you have never met, next to credentials.

Where it bites, plainly, and these are real: Firecracker cannot change vCPU or RAM when restoring a snapshot, so RAM is a property of the template, not of the environment. The cpu and memory_mb arguments on a create are silently corrected to the baked values whenever the template has a snapshot — not rejected, corrected, so you get a 200 and a different machine than you asked for rather than an error telling you so. Base is 4 GiB with 8 burst vCPU, the agent and code-interpreter templates 2 GiB, browser 4 GiB, postgres-16 1 GiB. If one team's monorepo needs 16 GiB to typecheck, the answer is a separate template baked at that size, not a parameter. Second: it is not an IDE and ships no browser editor. You connect your own editor over SSH, or you use it as an API; if what you want is a tab with a code editor in it, entries one and six are built for that and this is not. Third: no GPU, so model training and anything CUDA-shaped is out. Fourth: the guest kernel is trimmed, and trimmed kernels have sharp edges — k3s, for instance, wants an iptables module it does not have. Check your stack against a real environment before you plan a migration around it.

The seven against the criteria

Short cells, qualitative everywhere except the row I am allowed to be specific about. 'Verify' is not a hedge for its own sake — it means the honest answer depends on a product detail that changes faster than this table does.

Ephemeral development environment platforms, graded on the seven criteria. Non-PandaStack rows are qualitative; verify against current vendor docs.
PlatformIsolationDefined byOn idleNested DockerOperated by
GitHub CodespacesContainer, shared kerneldevcontainer.jsonStop, state retainedSupported patternVendor
GitpodContainer, shared kernelDeclarative configVerify current docsVerify current docsVendor or your infra
CoderWhatever Terraform makesTerraformAutostop, configurableDepends on providerYou
DevPodDepends on providerdevcontainer.jsonUp to your providerDepends on providerYou, no control plane
DaytonaVerify current docsImage or declarativeVerify current docsVerify current docsVendor
CodeSandbox / StackBlitzBrowser runtime or VMRepo plus configTab close or resumeNo for browser runtimeVendor
PandaStackOwn guest kernel, KVMDockerfile baked to snapshotIdle reaper, default 5 minOrdinary nesting, no daemon shippedVendor or self-host
The nested-Docker cell deserves a caveat even on my own row. A microVM has its own kernel, so running a container runtime inside one is ordinary nesting rather than a privileged-container exception — but the first-party templates do not ship a Docker daemon, and the guest kernel is trimmed. Install the runtime you want in a template and verify it works before you plan a testcontainers-based suite around it, the same way you should verify every other cell in that table.

The thing roundups skip: the economics of idle

Developers are not typing most of the time. They are in a meeting, at lunch, reading a design doc, arguing in a thread, or asleep. Across a working day, keystroke-adjacent activity is a small fraction of the wall-clock hours an environment exists, and across a working week it is smaller still, because weekends exist and environments do not know that. An environment that costs money during all of that is a fundamentally different product from one that costs nothing, and the difference is architectural rather than a pricing choice somebody could reverse on request.

There are two architectures underneath this whole category. The first keeps machines warm. Either your environment is a long-lived VM or pod that simply stays up, or the platform maintains a pool of pre-started instances so it can hand you one quickly. Either way, somebody is paying for silicon that is doing nothing — and because that somebody is the vendor, it comes back to you as a per-seat price, a concurrency cap, or an idle timeout tuned to protect their margins rather than your flow state. This is also why warm-pool platforms tend to be stingy with environment count: each additional one has a real standing cost.

The second architecture does not keep anything warm, because creation is fast enough that it does not need to. Bake the fully-configured environment into a snapshot once, then restore it on demand. PandaStack's restore path is 179 ms at p50 and 203 ms at p99, with the snapshot load itself around 49 ms, which is the number that makes the whole model work: at that latency, deleting an environment is not a loss, it is the idle strategy. You do not suspend at lunch. You delete, and when you come back you get a fresh one from the same snapshot in less time than it takes your editor to reconnect.

What that does to the per-seat maths, using only PandaStack's published rate card — $0.054 per vCPU-hour and $0.0162 per GiB-hour, one card for every workload class. CPU is billed on active CPU-seconds actually burned, so an environment nobody is typing into contributes essentially nothing on the CPU line; the honest cost of existing is memory, billed on committed GiB-hours. The base template commits 4 GiB, so arithmetic on the published card puts an environment that exists and does nothing at about 6.5 cents per hour. Leave it up around the clock, every day, and that is a real number multiplied by your headcount multiplied by the number of branches people have open. Delete it when the developer walks away and it is zero, because the thing you were paying to preserve is in a snapshot, and the snapshot is cheap.

The practical consequence is that the per-seat question dissolves into a per-active-hour question, and the capacity question changes shape with it. When the marginal environment costs nothing while idle, there is no reason to ration them: every branch gets one, every pull request gets one, every agent task gets one, and nobody queues. That is test three from the top of this post — the 400th costs the same as the 4th — and it is downstream of the restore latency, not of anyone's generosity.

One correction to a thing people assume: ttl_seconds on a PandaStack environment is an idle timeout, not a walltime budget. The reaper measures time since the last activity, and the default is five minutes — aggressive on purpose, because the restore is cheap. Worth knowing exactly what counts as activity: running a command, reading or writing file bytes, a PTY or SSH session, a fork or snapshot, and traffic through a preview URL all bump the clock. Read-only inspection does not — polling an environment's status or metrics deliberately does not keep it alive, or no environment would ever be reaped while a dashboard tab was open. So a teammate hitting your dev server keeps the environment up; a monitoring page watching it does not. If you want one to survive a long silence anyway, raise the TTL or mark it persistent.

How to actually migrate

The migration is less dramatic than it sounds, because most of the work is moving the environment definition out of a machine and into the repository — which you should do regardless of who hosts it. Start there. Step one is that every version your project needs is declared in a file a stranger can read.

# 1. The environment is a definition in the repo, not a machine someone built.
#    The base template runs mise, so these are read from the repo at setup time.
$ cat .tool-versions
node 22.14.0
python 3.12.8
go 1.24.0

# 2. Anything the repo cannot declare — system packages, CLIs, a seeded
#    database — goes in a template, baked once into a snapshot.
$ cat Dockerfile.dev
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
      git curl build-essential postgresql-client redis-tools ripgrep \
 && rm -rf /var/lib/apt/lists/*

$ pandastack template build -f Dockerfile.dev -n acme-dev

# The FIRST create on a fresh template cold-boots (~3s) and bakes a snapshot
# on the way through. Every create after that restores it: ~179ms at p50.
# Note what you cannot set here: RAM. Firecracker cannot change memory at
# snapshot restore, so the template's baked size IS the environment's size.
# Need 16 GiB for a monorepo typecheck? That is a separate template.

Step two is the per-branch environment. The shape below is deliberately boring: clone, install, build, serve, hand back a URL. The only unusual lines are the ones about sizing and timeouts, and both are there because getting them wrong is the most common way a first integration disappoints.

import os
from pandastack import Sandbox

BRANCH = os.environ["BRANCH"]
REPO = "https://github.com/acme/monorepo.git"

# Build steps run as non-login shells, so export mise's env explicitly.
MISE = (
    "export MISE_DATA_DIR=/opt/mise MISE_CONFIG_DIR=/opt/mise "
    "PATH=/opt/mise/shims:$PATH;"
)

# No cpu= or memory_mb= here: with a baked snapshot the agent overrides them
# to the template's values. ttl_seconds is an IDLE timeout, not a walltime cap.
sbx = Sandbox.create(
    template="acme-dev",
    ttl_seconds=1800,
    metadata={"branch": BRANCH, "owner": "dana"},
)

sbx.exec(f"git clone --depth 1 --branch {BRANCH} {REPO} /work", check=True)
sbx.exec(f"{MISE} cd /work && mise install && mise reshim", check=True)

# Long work goes through exec_stream, which honours timeout_seconds. Do NOT
# pass timeout_seconds > 30 to exec() — one-shot exec has no server deadline.
code = sbx.exec_stream(
    f"{MISE} cd /work && pnpm install --frozen-lockfile && pnpm build",
    on_stdout=lambda c: print(c, end=""),
    on_stderr=lambda c: print(c, end=""),
    timeout_seconds=1200,
)
if code != 0:
    raise SystemExit(f"build failed on {BRANCH} (exit {code})")

sbx.exec(f"{MISE} cd /work && setsid nohup pnpm dev >/var/log/dev.log 2>&1 &")
print(f"{BRANCH}: {sbx.preview_url(3000)}")

Step three is where the ephemerality actually pays for itself: warm one environment, snapshot it, and start every branch from that image instead of reinstalling the world. This is the pattern that turns a two-minute setup into a sub-second one, and it is also the pattern that makes deleting at lunch a non-event.

from pandastack import Sandbox

REPO = "https://github.com/acme/monorepo.git"
MISE = (
    "export MISE_DATA_DIR=/opt/mise MISE_CONFIG_DIR=/opt/mise "
    "PATH=/opt/mise/shims:$PATH;"
)

# Warm ONE environment: clone, install, prime the caches. Then freeze it.
warm = Sandbox.create(template="acme-dev", ttl_seconds=3600)
warm.exec(f"git clone {REPO} /work", check=True)
warm.exec(f"{MISE} cd /work && mise install && mise reshim", check=True)
warm.exec_stream(f"{MISE} cd /work && pnpm install", timeout_seconds=1800)

snap_id = warm.snapshot()   # a snapshot ID STRING, not an object
warm.kill()

# Every branch now starts from the warmed image instead of from apt-get.
envs = {}
for branch in ("feat/pricing", "feat/search", "fix/401-on-refresh"):
    sbx = Sandbox.create(from_snapshot=snap_id, metadata={"branch": branch})
    sbx.exec(f"cd /work && git fetch origin {branch} "
             f"&& git checkout {branch}", check=True)
    sbx.exec_stream(f"{MISE} cd /work && pnpm install", timeout_seconds=600)
    sbx.exec(f"{MISE} cd /work && setsid nohup pnpm dev "
             f">/var/log/dev.log 2>&1 &")
    envs[branch] = sbx

for branch, sbx in envs.items():
    print(f"{branch:20} {sbx.preview_url(3000)}")

# Branch the DISK of an environment to try something destructive against a copy
# of your real working tree + deps + local DB files (400-750ms, same host).
# fork() is disk-only: the child boots fresh, so the parent's processes do NOT
# come with it — restart the server in the clone. For running processes you
# want fork_tree() (<=16 children, same host) or snapshot-then-restore.
scratch = envs["feat/search"].fork(metadata={"why": "risky-migration"})
scratch.exec(f"{MISE} cd /work && ./scripts/migrate.sh --no-backup", check=False)
scratch.exec(f"{MISE} cd /work && setsid nohup pnpm dev >/var/log/dev.log 2>&1 &")
print("scratch:", scratch.preview_url(3000))
scratch.kill()

# Lunch. Nobody is sad: the snapshot is the environment.
for sbx in envs.values():
    sbx.kill()

Step four is the organisational half, and it is the half that decides whether the migration sticks: delete something on purpose, early, and see who complains. If a developer objects because they had uncommitted work in it, you have found a pet, and you have found it while it is still cheap to find. If nobody notices, the environments are ephemeral and you can stop thinking about them, which was the entire point.

Picking one

  • Your binding constraint is onboarding time for new or occasional contributors — Codespaces, or the browser-native entries for anything a stranger should be able to open in one click without an account.
  • Your binding constraint is that environments must run inside your own network under your own IAM — Coder, and budget for operating a control plane.
  • Your binding constraint is not getting locked into one vendor's environment format — DevPod on top of a devcontainer definition you own, and accept that governance is now your project.
  • Your binding constraint is that the environment runs untrusted dependency code next to credentials — anything with a per-environment kernel. That is PandaStack in this list; it is also why the entry exists.
  • Your binding constraint is idle cost at headcount scale — snapshot-restore rather than warm pools, and check the vendor's actual idle behaviour rather than their marketing word for it.
  • Your binding constraint is that your agents, not your humans, need thousands of environments a day — Daytona or PandaStack, and read the fork and concurrency semantics carefully before you commit.
  • Your binding constraint is GPUs, or a browser-based editor your team will not give up — none of my product's strengths help you. Choose from the others.

The bottom line

The category's real failure mode was never slow starts. It was that 'cloud dev environment' turned out to be compatible with the pet, so teams paid for a migration and kept the drift. The three tests at the top of this post are the whole check, and they are worth running against your current setup before you shop: is the environment created from a definition, would anyone be sad if it died, and does the 400th cost what the 4th did. Any platform here can pass them with discipline, and every platform here can be made to fail them with convenience. The reason I built PandaStack around snapshot-restore is that at 179 ms the discipline stops requiring any — deleting is simply easier than keeping, so the pet never forms. Grade the field on the criteria that bind you, spike your top two with your real repo, and delete something on purpose in the first week.

Frequently asked questions

What makes a development environment ephemeral rather than just remote?

Three tests, and it has to pass all three. First, it is created from a definition rather than restored from a backup: if the environment vanished, you would rebuild it by reading a file in the repo — a devcontainer spec, a Dockerfile, a .tool-versions, a Nix flake — not by restoring a disk image or asking the one person who remembers how it was set up. Second, nobody is sad when it dies. If a developer would say 'don't kill that, I had something in it', the environment has quietly become stateful and is now a pet. Third, the 400th one costs what the 4th did, so there is no reason to ration them and therefore no pressure to reuse one across branches. Remote only means the CPU is somewhere else, which is orthogonal: a remote environment can be the most pet-like object in your infrastructure, and in a lot of organisations it is. The test that predicts behaviour best is the second one, because it is social rather than technical.

Do I need a browser-based IDE for a cloud development environment?

Usually not, and assuming you do narrows your options for no benefit. Most engineers who try a browser editor go back to their local one within a week, because their keybindings, extensions, fonts, dotfiles and muscle memory all live there — and every serious platform in this category supports connecting a desktop editor over a remote connection or plain SSH. A browser IDE earns its place in two specific cases. The first is the stranger case: somebody who has never seen your project, has no account, and should get to a running version of it in one click — bug reproductions, documentation, teaching, code review. The second is a locked-down device where you genuinely cannot install a toolchain. Outside those, treat the browser editor as a feature you may never use rather than a requirement, and grade the platform on cold start, isolation and idle cost instead. PandaStack, for the record, ships no browser IDE at all — you bring your own editor over SSH, or you use it as an API — and that is a real reason to pick a different entry if the editor is what you want.

How do I size an ephemeral environment on PandaStack?

You size the template, not the environment, and this is a genuine constraint rather than a roadmap item. Firecracker cannot change vCPU count or RAM when it restores a snapshot, so whenever a template has a baked snapshot — which is the normal case, because that is what makes creates fast — the cpu and memory_mb arguments on a create are overridden by the agent to the template's baked values — silently, without an error, so a request for more memory succeeds and simply gives you the baked size. The first-party templates are baked at: base 4 GiB, agent and code-interpreter 2 GiB, browser 4 GiB, postgres-16 1 GiB, each with 8 burst vCPU shared fairly under contention by cgroup weight. If a team's monorepo needs 16 GiB to typecheck, the answer is to bake a template at 16 GiB and point that team at it, not to pass a larger number at create time and hope. The upside is that an environment's shape becomes a reviewable artifact rather than a parameter scattered across call sites — but if per-environment sizing is a hard requirement, know this before you start.

Can I run Docker or a local Kubernetes inside an ephemeral dev environment?

It depends entirely on the isolation model, and this is the criterion most likely to break a testcontainers-based test suite. On a shared-kernel platform, where your environment is a container on a multi-tenant host, running a container runtime inside it requires either a privileged container — which the platform may reasonably refuse — or a rootless or daemon-in-daemon arrangement with its own sharp edges. With a per-environment kernel the nesting is ordinary, because the kernel is yours to configure. The caveat on PandaStack specifically: the first-party templates do not ship a Docker daemon, and the guest kernel is trimmed, which means some things that work on a general-purpose distro kernel do not — k3s, for example, wants an iptables module the guest kernel does not include, so ClusterIP networking does not work out of the box. Install the runtime you need in a template, run your actual test suite in a real environment, and verify before you plan around it rather than after.

What does an idle development environment actually cost?

That depends on which of two architectures the platform uses, and the difference is structural rather than a pricing decision somebody could reverse for you. Warm-pool and always-on designs have machines running whether or not anyone is typing, and because the vendor pays for that silicon it comes back to you as a per-seat price, a concurrency cap, or an idle timeout tuned to their margins. Snapshot-restore designs keep nothing warm, because creating is fast enough not to need it. On PandaStack's published rate card — $0.054 per vCPU-hour and $0.0162 per GiB-hour, one card for all workload classes — CPU is billed on active CPU-seconds actually burned, so an environment nobody is using contributes almost nothing on the CPU line. The honest cost of merely existing is memory, billed on committed GiB-hours: the base template commits 4 GiB, which is roughly 6.5 cents an hour by arithmetic on that card. Deleting it instead costs nothing, and since the restore is 179 ms at p50, deletion is the idle strategy rather than a loss. Egress is metered separately.

Keep reading

Related posts

  • PandaStack vs Coder

    These two products are shopped against each other constantly and compete almost never. One gives a person a machine for the week; the other gives a program a machine for four seconds. Naming that split is most of the decision.

  • Best Preview Environment Platforms (2026)

    Every pull request can now get a live, clickable URL — reviewers want it, QA wants it, and increasingly AI coding agents need one to check their own work. A qualitative, hedged look at where Vercel, Netlify, Railway, Render, Northflank, GitHub Codespaces, and PandaStack each sit on isolation model, database handling, and boot latency.

  • The best CodeSandbox alternatives in 2026

    Half the people looking for a CodeSandbox alternative want an embeddable editor for their docs. The other half want an API that gives an agent a machine. Those are different products, and mixing them up is why the shortlists never help.

  • The Best Replit Alternatives in 2026

    "Replit alternative" means three unrelated things: a browser IDE, a cloud dev environment, or the sandbox API that runs untrusted code under the hood. Split the intent first and the shortlist collapses from ten options to two.

  • The Best Batch Compute Platforms for Bursty Jobs in 2026

    Nine ways to run a few thousand independent jobs, grouped by what they actually are rather than ranked. Most of your decision is made by two facts: whether you need GPUs, and whether you are willing to pay for a cluster that is idle most of the week.

More in Comparisons · See PandaStack pricing

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.