all posts

Top 8 Ephemeral Development Environment Platforms (2026)

Ajay Kumar··11 min read

Almost every comparison of ephemeral development environments is a feature grid: which IDEs, which VCS integrations, which SSO. Those are real questions and none of them is the question. The two numbers that decide whether a platform works for you are how fast a fresh environment can exist, and what happens to that environment when nobody is looking at it.

The first number decides the developer experience, and it does it indirectly. If a fresh environment takes thirty seconds, people create one per branch, per PR, per experiment, and the whole model works. If it takes six minutes, people keep one environment alive and reuse it forever, which is a long-lived cloud desktop with extra billing complexity. Nobody decides this in a meeting. They decide it by waiting, once, and then never waiting again.

The second number decides the bill, and it is the one vendors are quietest about. An environment that "stays warm so your next session is instant" is a machine you are renting while nobody types into it. The honest version of that sentence is on the invoice, under a SKU that sounds like infrastructure.

I'm Ajay. I build PandaStack, which runs Firecracker microVMs behind an API, and it is entry eight here. I have tried to write the other seven the way I would want a competitor to write mine: what substrate it actually isolates with, what "ephemeral" means inside its specific model, who it genuinely fits, and the catch I would rather know before the purchase order than after.

Ground rules, because they change how to read everything below. Specific latency numbers appear for exactly one platform -- mine -- because those are the only ones I can measure on hardware I control. Every other platform is described qualitatively from its public documentation, with no invented pricing, quotas or internals. And this category reshuffles faster than blog posts get updated: product names, isolation details, idle policies and price lists all move on a timescale of months. Verify anything I say about anyone else against that vendor's current docs before you commit budget, and verify that the smaller projects here are still actively maintained.

The two questions, stated precisely

How fast can a fresh environment exist

Not "how fast does the API return". Time-to-usable is the interval from the create call to the moment your code answers a request, your test command starts executing, or your editor can save a file and see the effect. It decomposes into three costs, and platforms differ wildly in which one they have attacked.

  • Machine acquisition: finding or starting the compute. A scheduler picking a node, a VM booting, a container being placed. Measured in seconds unless somebody has pre-allocated something.
  • Filesystem materialisation: getting the image or snapshot onto that machine. A registry pull if the layers are cold, a disk clone if they are not. This is where multi-gigabyte images go to be expensive.
  • Setup: the part your repo contributes. The postCreateCommand, the npm install, the migration, the "warming the cache" step somebody added in 2024 and nobody has timed since. On a lot of platforms this single term is larger than the other two combined, which means the vendor's impressive cold-start engineering is being eaten by your own install script.

When a platform quotes a start time, it is almost always a warm resume of a prebuilt image, which is a legitimate number measuring a completely different event. The number you care about is the cold one, on your repository, with your dependency tree. There is a harness for that at the end of this post.

What happens when nobody is looking

There are only four answers, and the answer is usually a product decision rather than a setting you can change.

  1. It keeps running. Your environment is a machine with an uptime, and you are buying all 730 hours in the month whether or not anybody opened a terminal.
  2. It suspends. Compute stops, state survives on a disk, and you pay a smaller but permanent storage charge plus a resume latency every time someone comes back. This is the most common answer and the easiest to under-budget, because the stopped disk keeps billing silently long after the branch was merged.
  3. It is deleted. Nothing survives except an artifact in object storage and whatever you deliberately persisted. The bill genuinely goes to approximately zero, and the entire burden shifts onto question one -- recreate has to be fast and it has to work every time.
  4. It never existed as remote compute at all. Browser-native environments run on the developer's own laptop inside the tab; the idle cost is a closed tab, and the trade is that you are no longer running Linux.

Notice that the two questions are not independent. Delete-on-idle is only rational if create is cheap; keep-warm is only rational if create is expensive. A platform that is slow to create AND deletes aggressively is the worst of both, and a platform that is fast to create AND keeps everything warm is leaving your money on the table by default.

Substrate sets the ceiling on both numbers

You cannot reason about either number without knowing what the thing is made of, because the substrate caps both the floor on create time and the strength of the boundary around whatever runs inside.

  • Container on a per-user or shared VM. Fastest conventional start, smallest disk footprint, and a boundary made of kernel namespaces and seccomp -- which is a well-engineered agreement rather than a wall. Perfectly appropriate for code your own team wrote. Less appropriate for code a model wrote thirty seconds ago.
  • Full virtual machine. The strongest boundary and the slowest acquisition; you are waiting on a hypervisor and a real kernel boot. This is what a per-developer cloud workstation is usually made of.
  • MicroVM. A hardware-virtualised guest with a minimal device model, which gets you the VM boundary at a start time competitive with containers -- especially when the platform restores a snapshot of an already-booted guest instead of booting one.
  • Browser WASM. A Node-compatible runtime compiled to WebAssembly, executing in the tab. Zero server, zero idle, zero cold start worth measuring, and zero Linux syscalls: no native modules, no Docker, no Postgres.
  • A routed slice of a shared cluster. Not a whole environment at all -- one or two replaced workloads inside a running baseline, with requests steered to your version by header or trace context. Dramatically cheaper than duplicating the stack, and it shares a blast radius with everyone else.

The eight

1. GitHub Codespaces

Substrate: a container, defined by devcontainer.json, running on a VM allocated per codespace. Ephemeral means time-bounded rather than task-bounded -- a codespace stops after an inactivity window and is deleted after a retention period, and between those two events it is a stopped disk that still costs money. Prebuilds, built by Actions, are the mechanism that makes the create tolerable; without them you are paying your own setup script on every create, which is exactly the third cost above. It fits teams already living inside GitHub who want a working environment for a human in one click, and the integration is genuinely excellent: the PR, the environment and the review are one surface. The catch is that the model is per-developer, not per-branch-at-scale. Spinning up forty of these because forty PRs are open is not what the pricing or the UX is shaped for, and the container boundary means you should not be running untrusted code in one. Verify current idle, retention and billing behaviour in GitHub's docs; this specific area has changed more than once.

2. Google Cloud Workstations

Substrate: a container image running on managed Compute Engine VMs inside your own VPC, with the machine type, image and idle policy described by a workstation configuration. Ephemeral here means auto-shutdown: the workstation stops when you stop using it and comes back with its persistent home directory intact. The reason to pick it is almost never developer experience -- it is that the environment sits inside your network perimeter, can reach your private services and databases directly, and inherits your existing IAM, VPC Service Controls and audit logging. For a regulated org trying to get source code off laptops, that is the whole argument. The catch is that this is a workstation product with an idle timer, not a per-PR fleet primitive: you are billed for a VM and a persistent disk per developer, the create path is a real VM start unless you are using a pre-started pool, and there is no notion of "create two hundred of these for the duration of a CI matrix". Check current configuration options and pool behaviour in Google's docs before designing around it.

3. Okteto

Substrate: Kubernetes, in your cluster, with a namespace as the unit of isolation and a manifest describing the development environment and its dependencies. This is the most honest answer for teams whose production is already a dozen services on Kubernetes, because the ephemeral environment is the same shape as production rather than a simplified imitation of it -- the same manifests, the same service discovery, the same ingress. Preview environments per pull request are a first-class feature, and idle namespaces can be put to sleep rather than left running. It fits a platform team that owns a cluster and wants per-PR environments that actually exercise the real topology. The catch is the cluster: you operate it, you size it, you pay for its headroom, and a per-PR environment that pulls up fifteen pods is a per-PR environment with a scheduling queue. The boundary is a namespace, so it is a shared kernel with RBAC and network policy on top, and a manifest is still a recipe executed at create time rather than a built artifact. Verify the current open-source versus commercial split before planning a self-hosted deployment.

4. Uffizzi

Substrate: Kubernetes again, but with virtual clusters as the per-environment unit, and an open-source core you can self-host. The model is tightly scoped and better for it: an environment per pull request, defined by a compose-style specification, built from images your CI already produces, torn down when the PR closes. Because it consumes images rather than building source from scratch, the create path skips most of the setup cost that dominates devcontainer-style platforms -- the artifact is already there. It fits a team that wants per-PR preview environments, wants the code to be inspectable, and does not want to buy a per-developer workstation product to get them. The catches are both about scale of project rather than scale of cluster: it is a much smaller ecosystem than the commercial entries here, so confirm the maintenance cadence before you build your release process on it, and virtual clusters are one more layer between a failing pod and an engineer trying to understand why. The isolation boundary is still a shared kernel underneath.

5. Signadot

Substrate: a shared baseline cluster plus request routing. This is the entry that answers the two questions by refusing the premise. Instead of creating an environment, you create a sandbox that replaces one or two workloads in an always-running baseline, and requests carrying the right header or trace context are steered to your version while everything else uses the shared copy. Time-to-exist is therefore the time to start one pod rather than a whole stack, and the idle cost of a sandbox approaches nothing because the expensive part -- the baseline -- is shared across everybody and amortised. For a microservice estate where a full per-PR environment would mean standing up forty services to test a change in one, this is the economically correct answer and the others are not close. The catch is written on the tin: it is not isolation. You share the baseline's databases and queues with every other sandbox, so a destructive migration or a poisoned message is everyone's problem, and the routing only works if your services propagate context correctly -- which means instrumenting them, including the ones nobody owns. Verify current routing and data-isolation options against their docs.

6. StackBlitz (WebContainers)

Substrate: a Node-compatible runtime compiled to WebAssembly, executing inside the browser tab. No remote machine exists, which makes both of our numbers degenerate in the most interesting way in this list: time-to-exist is roughly a page load, and the idle verb is "you closed the tab". There is no scheduler, no image pull, no cold start on a cloud provider's capacity, and no bill for an environment nobody is using, because the compute was always the laptop you were already holding. For reproductions, documentation that runs, interactive tutorials, teaching, and the first ninety seconds of evaluating a library, nothing else is in the same category. The catch is the substrate's ceiling, and it is a hard one: there is no Linux kernel in there. Native node modules, anything that shells out to a system binary, Docker, a local Postgres, a Go or Rust toolchain -- these are not slow, they are absent. It is also not where your team's actual monorepo with a native image-processing dependency is going to live. Judged on our two questions it wins outright; judged on "can it run my stack" it may not qualify at all.

7. Replit

Substrate: containers, behind a browser IDE that has largely become an agent surface -- you describe an app, the agent writes and runs it, and the environment is where that happens. Included here because it is one of the most widely used "cloud development environment" products in existence and buyers do put it on the shortlist, but it is worth being precise about the model: persistence is the product. A workspace is meant to still be there tomorrow with your files and your installed packages, which is the opposite of ephemeral by design. Where disposability does show up is at the edges -- deployments that scale down when traffic stops, and agent sessions that are inherently short-lived. It fits individuals, education, prototypes and anyone who wants the distance between an idea and a running URL to be as close to zero as possible. The catches, for the use case this post is about: it is not a fleet primitive with an API you would wire into CI to open an environment per pull request, the container boundary is a shared kernel, and the per-workspace persistence model means idle state accumulates rather than evaporates. Verify current plans and deployment behaviour directly; this product changes shape frequently.

8. PandaStack

Mine, so weigh it accordingly. Substrate: Firecracker microVMs -- a real guest kernel, 5.10, on Ubuntu 24.04, with a per-sandbox network namespace and tap device. There is no warm pool. Every create restores a template's baked snapshot on demand: allocate a pre-built network slot, patch the tap MAC, reflink the rootfs, fork a Firecracker process, PUT /snapshot/load, resume, probe port 22. That path measures p50 179ms and p99 around 203ms, with the snapshot load itself at roughly 49 to 80 milliseconds. The very first spawn of a brand-new template is the slow one, about 3 seconds, because it cold-boots and bakes the snapshot that every later create restores. Ephemeral is the default rather than a policy: a TTL is enforced by a platform-side reaper, so a sandbox outlives neither its deadline nor your orchestrator's unexpected death. For app hosting, sleep is implemented by deleting the virtual machine after baking a seed to object storage, and waking restores from that seed in about 1.2 seconds -- the idle verb is genuinely "deleted", not "suspended and quietly billed".

Who it fits: anyone running code that should not share a kernel with anything else -- model-generated code, customer-submitted code, a CI matrix on untrusted forks -- and anyone who needs environments created programmatically in bulk, since a fork of an existing sandbox hands the child the parent's filesystem -- dependencies already installed, files already written -- and a fork_tree snapshots the parent once and restores up to sixteen children from that snapshot, which on the same host is a 400 to 750 millisecond restore. Each agent host pre-allocates 16,384 network slots as its hard ceiling. Now the catches, and there are several. It is not an IDE in a browser, and this post should not pretend otherwise: there is no editor, no devcontainer.json ingestion, no "open this PR in a workspace" button. You get an API, a CLI and Python and TypeScript SDKs, with exec, a filesystem interface, SSH and tokenless preview URLs. Those preview URLs are tokenless in a specific sense -- the sandbox UUID is the bearer credential, so treat the string like a password and keep TTLs short. Guest RAM is fixed when the template is baked, because Firecracker cannot change memory or vCPU count at snapshot restore, so a memory_mb on a create is silently corrected rather than rejected. And the two fork calls are different animals, which is worth knowing before you design around either: fork gives the child the parent's disk and then cold-boots it, so what you get is a warm filesystem and a fresh kernel -- not a warm-process clone -- and it costs what a cold boot costs, seconds rather than milliseconds. The warm clone, where the parent's memory comes too and a process you left running is still running in the child, is fork_tree, and it is capped at sixteen children per call. Because those children inherit memory exactly, they also inherit the random number generator's state, so anything that seeds crypto or generates identifiers immediately after a fork_tree needs to reseed explicitly. If what you want is a comfortable cloud workstation for a human, one of the first two entries is a better answer than mine.

Eight platforms on the two questions. Substrate and idle behaviour for everything except PandaStack are from public documentation and move frequently -- verify before you buy.
PlatformSubstrateWhat "ephemeral" means hereIdle verbThe honest catch
GitHub CodespacesContainer on a per-codespace VMInactivity stop, then retention deleteSuspends, then deleted laterPer-developer model; stopped disks keep billing
Google Cloud WorkstationsContainer on a managed GCE VM in your VPCAuto-shutdown on idle, home dir survivesSuspendsA workstation with a timer, not a per-PR fleet
OktetoKubernetes namespace in your clusterEnvironment per PR, defined by a manifestSleeps idle namespacesYou operate and pay for the cluster's headroom
UffizziVirtual cluster on Kubernetes, OSS coreEnvironment per PR, from CI-built imagesDeleted on PR closeSmall ecosystem; check maintenance before adopting
SignadotRouted workloads in a shared baseline clusterA sandbox is a slice, not a whole stackNear-zero; baseline is sharedNot isolation -- shared data plane and blast radius
StackBlitzNode compiled to WASM, in the browser tabThe tab closes and it is goneNever existed as remote computeNo Linux: no native modules, Docker or Postgres
ReplitContainer behind a browser IDE and agentMostly persistent by design; edges scale to zeroKeeps state; deployments scale downPersistence is the product; not a CI fleet API
PandaStackFirecracker microVM, snapshot-restoredTTL-reaped by default; sleep deletes the VMDeleted (app wake ~1.2s from a seed)No IDE; RAM baked at template build time

What makes an ephemeral environment actually reproducible

Here is the failure that eventually happens on every platform in that table. Two developers create an environment from the same commit, twenty minutes apart, and get different behaviour. Not different code -- different behaviour. One has a transitively newer minor version of a library whose maintainer shipped on a Tuesday; the other picked up a base image whose tag moved. The test that fails is the test that was always slightly wrong, and now it is also intermittent, which converts a one-hour bug into a three-day one.

The cause is structural: the environment's definition was a program, and that program was executed separately for each environment. A setup script is not a specification. It is a set of instructions with a network dependency and a clock, and every execution is a separate opportunity to produce a different result. The contract has to be the artifact the script produced, not the script.

The litmus test: could you create this environment twice, a month apart, with your package registries unreachable? If the create path has to resolve a tag, hit a package index or run a compiler, the answer is no -- and every environment you create is a slightly different environment that happens to usually work.

In practice that means moving everything that resolves anything out of the create path and into a build that runs once, in CI, and produces something content-addressed: an image digest, a snapshot generation, a lockfile-pinned layer. The create path's job is to reference an immutable identifier and nothing else.

#!/usr/bin/env bash
# The question this script answers: "can I create the same environment twice,
# a month apart, with the package registries unreachable?" If the answer is no,
# the environment is not reproducible -- it is merely repeatable on a good day.
set -euo pipefail

# ---------------------------------------------------------------- BAD: a tag
# A tag is a mutable pointer. `node:22-bookworm` was a different filesystem
# last Tuesday and will be a different one next Tuesday, and nothing in your
# repo records which one your green CI run actually used.
#   FROM node:22-bookworm
#   RUN apt-get update && apt-get install -y build-essential   # <- and a clock

# -------------------------------------------------- GOOD: a content address
# Resolve the tag ONCE, in CI, and commit the digest. `docker buildx imagetools`
# reads the registry manifest without pulling the image, so this is cheap enough
# to run in a weekly "bump the base" job that opens a PR you can review.
DIGEST=$(docker buildx imagetools inspect node:22-bookworm \
           --format '{{json .Manifest.Digest}}' | tr -d '"')
echo "FROM node:22-bookworm@${DIGEST}" > base.pin
# Now the Dockerfile references base.pin's digest, and `apt-get install` happens
# at BUILD time -- once, in one place, with the result stored -- instead of on
# every environment create, 400 times a day, against a mirror that may have
# moved on.

# ------------------------------------------- the same idea, one layer deeper
# On PandaStack the create path restores a Firecracker snapshot, so the "image"
# is a snapshot of a machine that has already finished booting and already has
# the toolchain warm. That artifact is built once:
pandastack template build \
  -f templates/ci-node22/Dockerfile \
  -n ci-node22 \
  --memory-mb 4096          # BAKED. Not a default -- a constant.

# --memory-mb is load-bearing and surprising. Firecracker cannot change guest
# RAM or vCPU count when it restores a snapshot, so `memory_mb` passed on a
# create is silently corrected to the baked value. That is worse than an error,
# because a 512 MiB request against a 4 GiB template looks like it worked. Pick
# the number at bake time; every template gets 8 burstable vCPUs regardless.

# The first-ever spawn of a new template cold-boots and bakes the snapshot
# (~3s). Every create after that restores it: p50 179ms, p99 ~203ms, of which
# the /snapshot/load call itself is ~49-80ms. The reason those numbers are flat
# is that there is nothing to resolve at create time -- no registry lookup, no
# apt, no npm install. The artifact IS the contract.

Three things genuinely cannot be baked, and pretending otherwise is the other half of the problem. Your branch's code changes per environment, so it arrives as one shallow clone of one ref. Secrets must not be in an artifact that gets cached, replicated and eventually leaked, so they arrive as injected environment variables. And data -- the database contents your tests need -- is the hardest of the three, which is why database branching exists as a separate product category. Everything else belongs in the image or the snapshot.

The cost shape of idle

A platform that keeps your environment warm is billing you for a machine nobody is typing into. That is not an accusation of bad faith; keeping it warm is the only way to make a slow create tolerable, and vendors with slow creates make exactly the right engineering decision when they keep things warm. It is simply the trade, and it has a shape worth writing down before you sign.

The shape is: active hours at the machine rate, plus idle-but-running hours at the same machine rate, plus stopped hours at a disk rate, plus the cost of recreating. The first term is the only one anyone estimates. The second is the one that doubles the bill, because a developer's day is bursty and an idle timer fires several times a day rather than once. The third is the one that persists for months after the branch was merged, because nobody has ever enjoyed deleting other people's stopped environments.

#!/usr/bin/env python3
"""The idle model. Deliberately has no prices in it -- paste your own rate.

The output is machine-hours, because that is the unit every one of these
platforms ultimately bills in, whatever it calls the SKU. Run it with your
real numbers before you believe any vendor's cost comparison, including mine.
"""

HOURS_PER_MONTH = 730          # 24 x 365 / 12
WORKDAYS = 21

def machine_hours(
    envs,                      # concurrent environments (devs, or open PRs)
    active_h_per_day,          # hours a human is genuinely typing in one
    idle_verb,                 # "running" | "suspend" | "delete"
    idle_timeout_min=30,       # how long "idle" lasts before the verb fires
    creates_per_day=1,         # how often the environment is recreated
    create_minutes=0.0,        # time-to-exist, in minutes, from your harness
):
    # The subtlety: real work is bursty. Eight hours of "a session is open" is
    # not eight hours of typing, and an idle timeout fires several times a day,
    # not once. Four gaps per day is a conservative guess for one developer.
    idle_gaps_per_day = 4
    billed_idle_h = (idle_gaps_per_day * idle_timeout_min) / 60

    if idle_verb == "running":
        # Nothing stops it. You are renting the machine, not the work.
        per_env = HOURS_PER_MONTH
    elif idle_verb == "suspend":
        # You stop paying compute, start paying for a disk that still exists,
        # and pay the resume latency every time someone comes back.
        per_env = (active_h_per_day + billed_idle_h) * WORKDAYS
    elif idle_verb == "delete":
        # Nothing survives but an artifact in object storage. The create cost
        # is real, so it is charged here rather than hidden in a footnote.
        per_env = (active_h_per_day * WORKDAYS
                   + (create_minutes / 60) * creates_per_day * WORKDAYS)
    else:
        raise ValueError(idle_verb)
    return envs * per_env

for verb in ("running", "suspend", "delete"):
    h = machine_hours(envs=25, active_h_per_day=4.5, idle_verb=verb,
                      create_minutes=1.5, creates_per_day=3)
    print(f"{verb:>8}: {h:,.0f} machine-hours/month")

# The gap between "running" and "delete" is roughly an order of magnitude, and
# it is bigger for per-PR environments than for per-developer ones, because a
# PR environment's active fraction is dominated by one review that happens at
# 16:00 and a CI run that takes four minutes.
#
# But read the third line honestly: "delete" only wins if create_minutes is
# small AND the create works every time. Plug in create_minutes=6 with a 5%
# failure rate and your developers will start keeping environments warm by
# hand, which is the expensive verb with extra steps.

Run that with your own concurrency and your own measured create time. The gap between the keep-running line and the delete line is usually close to an order of magnitude, and it widens for per-PR environments, where the active fraction is one review at 16:00 plus a four-minute CI run. But read the delete line honestly: it is only the cheap option while create stays fast and reliable. The moment create takes six minutes and fails one time in twenty, your developers will start keeping environments warm manually, and manual keep-warm is the expensive verb plus a Slack thread.

This is why the two questions collapse into one decision. Fast, reproducible create is what buys you the right to delete aggressively; aggressive deletion is what makes a fast create worth engineering. Below is what the delete-first version looks like as code, including the part most teardown logic gets wrong -- that a teardown owned by your process cannot survive your process being killed.

from pandastack import Sandbox

# One environment per task, created at the moment the task starts and deleted
# when it ends. The interesting part of this script is not the create -- it is
# that there are three independent reasons the environment stops existing, and
# you want all three.
sbx = Sandbox.create(
    template="ci-node22",          # the baked artifact from the previous block
    ttl_seconds=1800,              # reason 1: the platform-side reaper
    metadata={                     # reason 2: so a human can audit the fleet
        "pr": "4821",
        "repo": "acme/api",
        "owner": "ci",
    },
)

# Reason 1, in detail: ttl_seconds is enforced by a reaper on the platform, not
# by this process. That distinction is the whole value. If this script is
# SIGKILLed, if the runner's spot instance is reclaimed mid-job, if someone
# force-pushes and GitHub cancels the workflow -- the sandbox still dies on
# schedule. A `finally: kill()` block cannot make that promise, because the
# failure mode you are defending against is "this interpreter stops existing".

# The code is the one thing that cannot live in the image: it changes per task.
# So the fast path is "immutable artifact + one shallow clone of one ref".
sbx.exec(
    "git clone --depth 1 --branch pr-4821 https://github.com/acme/api /work",
    timeout_seconds=120,
    check=True,
)

# Note on timeouts, because it bites everyone once: timeout_seconds is a CLIENT
# deadline. The server does not enforce it. If you need a hard wall, put it in
# the shell inside the guest, where the kernel will honour it:
r = sbx.exec("cd /work && timeout 900 npm test", timeout_seconds=960)
print(r.exit_code, r.stdout[-2000:])

# Reason 2: a tokenless preview URL for the reviewer -- https://<port>-<id>.<suffix>.
# There is no token endpoint because the sandbox UUID *is* the bearer credential.
# Treat the string like a password and keep the TTL short, because a leaked URL
# is a leaked environment for exactly as long as the sandbox lives.
print("preview:", sbx.preview_url(3000))

# Reason 3: delete it yourself, on the happy path, immediately. The TTL is the
# backstop for the day your orchestrator dies, not the normal teardown.
sbx.kill()

# `Sandbox` is also a context manager, which is the version you actually want
# in CI -- it kills on the way out of the block, including on an exception:
#
#   with Sandbox.create(template="ci-node22", ttl_seconds=1800) as sbx:
#       sbx.exec("cd /work && npm test", check=True)

Measure your own two numbers before you choose

Every number in every vendor comparison, including the table above, is a starting hypothesis. The two that decide your outcome are measurable in an afternoon on a trial account, and they are both measurements of your repository rather than of a demo.

#!/usr/bin/env bash
# time-to-usable.sh -- measure the number vendors do not publish.
#
# Two rules, both learned the hard way:
#   1. Stop the clock on YOUR readiness signal, not the platform's status field.
#      "running" means a machine exists. It does not mean the dev server is
#      listening, the migrations have run, or node_modules is on disk.
#   2. Measure the COLD path on YOUR repo. Every quoted start time in this
#      market is a warm resume of a prebuilt image, which is a different
#      number measuring a different thing.
set -euo pipefail

RUNS=${RUNS:-20}
READY_URL=${READY_URL:?set READY_URL to the health endpoint you actually wait on}
results=()

for i in $(seq "$RUNS"); do
  start=$(python3 -c 'import time; print(time.time())')

  # Whatever "create a fresh environment" is on the platform under test. Make
  # it genuinely fresh each time -- a new branch name, a cache-busting ref --
  # or you are benchmarking the prebuild cache and calling it a cold start.
  ENV_ID=$(your-platform create --branch "bench-$i-$RANDOM" --quiet)

  # Poll, do not sleep. A fixed sleep is simultaneously too long for a snapshot
  # restore and too short for a cold container build, so it hides both signals.
  until curl -sf -o /dev/null --max-time 2 "$READY_URL"; do
    sleep 0.05
    # Fail loudly rather than measuring an environment that never arrived; a
    # mean that silently drops the 8% of creates that hang is a lie.
    [ "$(python3 -c "import time;print(time.time() - $start > 600)")" = "True" ] \
      && { echo "run $i never became ready" >&2; break; }
  done

  end=$(python3 -c 'import time; print(time.time())')
  results+=( "$(python3 -c "print(f'{$end - $start:.3f}')")" )
  your-platform delete "$ENV_ID" >/dev/null
done

# p50 and p95, because the mean is the one statistic that cannot tell you
# whether people will trust the thing. A 20s p50 with a 4-minute p95 gets
# abandoned; developers remember the 4 minutes.
printf '%s\n' "${results[@]}" | sort -n | python3 -c '
import sys
v=[float(x) for x in sys.stdin if x.strip()]
q=lambda p: v[min(len(v)-1, int(round(p*(len(v)-1))))]
print(f"n={len(v)} p50={q(.5):.2f}s p95={q(.95):.2f}s max={v[-1]:.2f}s")'

# The second measurement needs no script. At 09:00 on a Monday, count the
# environments that exist and divide by the number of people who are awake.
# That ratio is your idle bill, and on a platform whose idle verb is "keeps
# running" it is usually somewhere north of three.

One more thing to look at while you have the trial open: what the platform does when it fails to create. A 503 with a retry hint is a platform you can build a queue in front of. A create that returns success and hands you an environment where the dev server never came up is a platform that will generate flaky CI for a year, and you will blame your own tests for most of it.

Picking one

Map your actual situation to the substrate, not to the feature grid:

  • A human needs a comfortable machine, the code is trusted, and the company lives in GitHub: Codespaces. Turn prebuilds on before you judge the create time, and put a calendar reminder on deleting stopped ones.
  • A human needs a machine inside the network perimeter, and compliance is the driver: Google Cloud Workstations, or the equivalent from whichever cloud already holds your IAM and audit trail.
  • Production is a dozen services on Kubernetes and a preview must exercise the real topology: Okteto if you want a supported platform on your cluster, Uffizzi if you want an open-source per-PR unit driven by images CI already builds.
  • Production is forty services and duplicating them per PR is absurd: Signadot, with eyes open about the shared data plane.
  • The environment is a reproduction, a tutorial or the first ninety seconds of evaluating a library: StackBlitz, and stop reading comparison posts.
  • The code is untrusted or model-generated, you need environments created programmatically in bulk, and nobody is going to sit in an editor inside them: a microVM sandbox API -- mine, or one of the others in that cohort.
The environment you can recreate in a second is the only one you will ever be willing to throw away.

If you remember one thing from this, make it the pairing. Platforms get chosen on feature grids and regretted on those two numbers, and the regret is predictable: either the create was slow enough that the per-branch model quietly died and everyone went back to one long-lived box, or the idle verb was "keeps running" and a quarter later somebody in finance asked what the compute line was. Measure both, on your repo, before you commit.

Frequently asked questions

What actually makes a development environment "ephemeral" rather than just remote?

Two properties, and both have to hold. First, the environment is created per unit of work -- a branch, a pull request, an agent task, a test run -- rather than per person. Second, destroying it is a non-event: nothing of value lives only inside it, so losing it costs nothing but a recreate. A remote machine that one developer keeps for eight months and occasionally rebuilds is not ephemeral however many times it has been rebuilt, because it accumulates state nobody can reproduce. The practical test is to delete one at random during working hours and see whether anybody notices. If somebody loses an hour of work, or a service becomes unreachable, or a half-configured dependency has to be rediscovered, the environment was long-lived and the label was marketing. The second-order test is whether recreating it is fast and deterministic, because an environment you are theoretically willing to destroy but practically afraid to destroy behaves exactly like a pet. That fear is rational when create is slow or flaky, which is why create latency and disposability are the same engineering problem rather than two.

Is a container enough for ephemeral environments, or do I need a VM boundary?

It depends entirely on who wrote the code. For your own team's source, running on infrastructure you control, a container on a per-user VM is a sensible boundary and the performance and density are better. Namespaces, cgroups and seccomp are genuinely well-engineered, and the threat model -- your developers' code accidentally doing something dumb -- is well matched to them. The calculus changes when the code is untrusted: a customer's submission, a fork's CI run, or output a language model generated ten seconds ago with no intent either way. There, the container boundary is a large shared kernel attack surface protected by a syscall filter, and container escapes are a recurring genre rather than a historical curiosity. A microVM gives you a separate guest kernel and a hardware virtualisation boundary, so the escape has to get through a hypervisor with a deliberately minimal device model. Historically the reason to accept a container's weaker boundary was speed, and that trade has mostly evaporated: snapshot-restoring a microVM puts a create in the few-hundred-millisecond range. If untrusted code runs in your ephemeral environments, the substrate decision is already made for you.

Is suspending or deleting an idle environment cheaper?

Deleting, almost always, and by more than people expect -- but only if recreating is fast and reliable. Suspending stops the compute charge and leaves a disk behind, and that disk is billed continuously, is rarely cleaned up, and survives the branch it belonged to by months. Multiply a modest per-environment disk by every PR your team opened last quarter and you have a line item nobody budgeted for. Deleting takes that to roughly the cost of an artifact in object storage, which is shared across all environments built from the same image or snapshot rather than paid per environment. The catch is entirely on the create side. If a cold create takes six minutes, deleting idle environments does not save money -- it moves the cost onto developers, who respond by keeping environments warm by hand, which costs more than suspension and annoys everybody. So the sequence matters: get create fast and deterministic first, then delete aggressively. A platform whose sleep path literally destroys the machine and restores it from an artifact in roughly a second can afford the aggressive policy; one whose cold path is minutes genuinely cannot.

How do I stop per-branch environments from drifting when the setup script keeps changing?

Stop treating the setup script as the definition. A script is a program with a network dependency and a clock, and running it once per environment means every environment is a separate roll of the dice: a moved tag, a transitively newer minor version, a mirror that was being reindexed. The fix is to execute that program once, in CI, and to promote its output to the contract -- a container image referenced by digest rather than tag, or a machine snapshot referenced by generation. The create path then resolves nothing. The useful litmus test is whether you could create the same environment twice, a month apart, with your package registries unreachable; if not, you have a recipe rather than a specification. Three things cannot be baked and should be handled explicitly: your branch's code, which arrives as one shallow clone of one ref; secrets, which arrive as injected environment variables because an artifact gets cached and replicated and eventually leaked; and test data, which is genuinely hard and is why database branching is its own product category. Everything else belongs inside the artifact, and the pleasant side effect is that create gets much faster, because installing dependencies at create time was most of the latency.

Can PandaStack replace GitHub Codespaces?

For the use case most people mean by Codespaces, no, and I would rather say so here than in a sales call. Codespaces is a workspace for a human: an editor in the browser or attached over SSH, a devcontainer.json your repo already has, and a button on the pull request. PandaStack has none of that. It is a sandbox API -- Firecracker microVMs created programmatically, with exec, a filesystem interface, SSH, tokenless preview URLs, snapshots and forks -- and the thing it is built to be good at is creating a lot of isolated environments very quickly and destroying them without ceremony. Where they overlap is the per-task case: an environment per CI job, per agent run, per untrusted fork, or a preview URL per branch, where nobody intends to sit inside the environment and type. There, the microVM boundary and a few-hundred-millisecond create matter and an IDE integration does not. Several teams run both for exactly that reason, and that is the honest recommendation: a workstation product for the humans, a sandbox API for the fleet. If you want one tool to do both, you will be compromising on one of them, and it is worth deciding in advance which.

Keep reading

Related posts

  • Cloud Dev Environments on microVMs

    Ephemeral dev environments want three things at once — real isolation, fast start, and the ability to freeze and resume. A microVM gives you all three, which is why it's a better substrate than a shared-kernel container.

  • A Disposable Dev Environment per Git Branch

    Every branch gets its own fully-isolated microVM dev environment — checkout done, deps installed, services running — that forks in sub-second and disappears when the branch is gone. Ephemerality is what finally kills environment drift.

  • PandaStack vs Coder

    These two products are shopped against each other constantly and compete almost never. One gives a person a machine for the week; the other gives a program a machine for four seconds. Naming that split is most of the decision.

  • Preview Environments on microVMs: a Live URL per PR

    Every pull request gets its own live URL backed by a real backend and database — on a Firecracker microVM, so even untrusted forked-PR code is isolated by hardware, not by a shared kernel.

  • Best Preview Environment Platforms (2026)

    Every pull request can now get a live, clickable URL — reviewers want it, QA wants it, and increasingly AI coding agents need one to check their own work. A qualitative, hedged look at where Vercel, Netlify, Railway, Render, Northflank, GitHub Codespaces, and PandaStack each sit on isolation model, database handling, and boot latency.

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.