Top 14 Ephemeral Development Environment Platforms in 2026: The Resume Path Nobody Benchmarks
The most honest benchmark in this category is a pull-request environment that has been idle since lunch, and nobody publishes it. A reviewer opens the link at three in the afternoon. The inactivity timer fired at 13:07. The page spins. Thirty seconds later there is a shell, and five minutes after that there is a dev server, because `node_modules` came back but the thing that built it did not.
Every platform here will tell you what create costs. Create is the demo: warm cache, prebuild hit, image already on the node, a fresh slot waiting. The number a human hits ten times a day and an agent hits ten times an hour is the other one -- the one after the platform decided nobody was looking. I have never once seen it on a pricing page.
So this is the same fourteen-platform category graded on the unglamorous half of the word ephemeral: what happens to an environment when it is unobserved, what of it survives, what it costs you at 3am, and how long the way back takes. I build one of the fourteen, which I will mark clearly and then grade harder than the rest.
The four verbs hiding behind the word idle
Read any twelve of these products' docs and you will find twelve nouns -- stopped, suspended, hibernated, archived, asleep, scaled to zero, cold, paused, terminated -- describing four actual behaviours. The nouns are marketing. The behaviours are these:
- Stop. The processes are killed and RAM is thrown away; the disk volume is kept. Resume means boot again and re-run everything that was not in the image. You pay for storage, and on most platforms for a volume that is pinned to one zone.
- Suspend. RAM is checkpointed to storage and the disk stays attached. Process state survives: your dev server is still listening, your language server still has its index, your shell still has its history in memory. You pay to store a disk AND a memory image, and usually you are still holding a slot.
- Hibernate. A suspend where the memory image is moved somewhere colder, or the host is released entirely. Cheaper at rest, slower to come back, and it may come back on a different machine -- which is the detail that decides whether your resume is milliseconds or seconds.
- Destroy. Nothing survives except what you deliberately exported: a commit, a volume, an artifact, a snapshot. You pay for that artifact and nothing else. Resume is not a resume. It is a create.
That list is the whole analytical frame of this post, because the second question follows from it automatically. Stop, suspend and hibernate all keep a thing somewhere -- a slot, a volume, a checkpoint -- and bringing it back means finding that thing again. Destroy keeps a file. Finding a file is a solved problem.
Ephemeral is a property of the environment. Durable is a property of the artifact. A platform that cannot name its artifact is not offering you ephemeral; it is offering you deletion with extra steps.
Why resume is structurally slower than create
It is tempting to assume resume should beat create -- the work was already done once, surely coming back is cheaper. It usually is not, and the reasons are architectural rather than sloppy:
- Re-placement. The scheduler has to find a host again. Your old host has moved on, drained, been patched, or filled up with somebody else's work. Create benchmarks are run against a fleet with room; resume happens whenever the user happens to come back.
- Re-attachment. A volume that was detached has to be attached, and a volume lives in one zone. That constraint quietly turns a placement decision into a placement requirement, which is slower and occasionally impossible.
- Re-hydration. Image layers get evicted from a node's cache. A create that hit a warm layer cache and a resume that pulls three hundred megabytes are the same API call with a 100x spread, and nothing in the response tells you which one you got.
- Re-binding. The URL, route, DNS record or editor port-forward that pointed at the old instance has to point at the new one. Propagation of that binding is in nobody's create benchmark and is frequently the single largest term in the resume time a user perceives.
- Re-running your setup. The half of a dev environment that is not in the image: installed dependencies, a dev server, a warm JIT, a language server index, an ssh-agent, a database with its cache populated.
Point five is why "the disk survived" is a weaker promise than it reads as. Your `pnpm install` survived. The dev server it started did not. The language server's index did not. The compiler's incremental cache survived as bytes but will be revalidated against a clock that jumped. Ephemeral, in a stop-based platform, is usually a euphemism for we deleted your running processes and kept your `node_modules` as a souvenir.
Suspend is the verb that fixes point five, and it is rare, because checkpointing memory correctly is genuinely hard: you have to deal with open sockets that are now stale, a monotonic clock that did not advance, a realtime clock that is now wrong, and any in-memory state that assumed time was continuous. That is not a theoretical list -- it is the list of bugs you get, in order, and the TLS failures come first.
Measuring the second visit
You cannot get this from documentation, because the thing you want to measure is the gap between the documented behaviour and the implemented one. You need two scripts: one that makes the platform admit its own taxonomy, and one that times the round trip.
First, ask the API what states it really has
#!/usr/bin/env bash
# Step zero: make the platform tell you its own taxonomy. Every one of these
# has a status enum, and the enum is the honest version of the lifecycle
# diagram on the marketing page.
set -euo pipefail
API="${ENV_API_URL:?platform API base}"
KEY="${ENV_API_KEY:?machine credential, not your browser session}"
H=(-H "Authorization: Bearer $KEY" -H "Content-Type: application/json")
curl -sS "${H[@]}" "$API/v1/sandboxes" | python3 -c '
import sys, json
d = json.load(sys.stdin)
rows = d if isinstance(d, list) else d.get("sandboxes", [])
print(sorted({r.get("status", "?") for r in rows}))'
# Then look for the verbs. A platform with no suspend/resume pair is one
# where "idle" means stopped or deleted, whatever the docs call it.
curl -sS "$API/openapi.json" \
| grep -ioE '(stop|start|suspend|resume|pause|hibernate|archive|destroy)' \
| sort -uTwo things fall out of that immediately. If the status enum has more values than the docs' lifecycle diagram has boxes, the extra ones are where your incidents will come from. And if there is no suspend/resume pair anywhere in the surface, then whatever the product page calls idle, the implementation is stop or destroy, and you should stop reading their latency claims as if they covered the resume case.
Then time the round trip, and check what survived
#!/usr/bin/env bash
# Resume-path probe. Point it at any platform that can create an environment,
# idle it, and bring it back. The number you want is not CREATE. It is the
# SECOND visit, after the platform decided nobody was looking.
set -euo pipefail
API="${ENV_API_URL:?platform API base}"
KEY="${ENV_API_KEY:?machine credential}"
H=(-H "Authorization: Bearer $KEY" -H "Content-Type: application/json")
TRIALS="${TRIALS:-10}"
IDLE_WAIT="${IDLE_WAIT:-0}" # 0 = idle it explicitly; >0 = sit out their timer
ms() { python3 -c 'import time; print(int(time.time()*1000))'; }
id_of() { python3 -c 'import sys,json; print(json.load(sys.stdin)["id"])'; }
body_of() {
python3 -c 'import json,sys; print(json.dumps({"cmd": sys.argv[1]}))' "$1"
}
run() { # run <id> <command>
curl -sfS "${H[@]}" -X POST "$API/v1/sandboxes/$1/exec" \
-d "$(body_of "$2")"
}
probe() {
local id t0 t1 t2 create resume
t0=$(ms)
id=$(curl -sS "${H[@]}" -X POST "$API/v1/sandboxes" \
-d '{"template":"base"}' | id_of)
until run "$id" true >/dev/null 2>&1; do sleep 0.02; done
t1=$(ms); create=$((t1 - t0))
# Put state in it that the IMAGE does not have. This is the half of a dev
# environment that "the volume persisted" quietly excludes.
run "$id" 'mkdir -p /work && head -c 48 /dev/urandom | base64 > /work/mark' \
>/dev/null
# A background counter tells stop from suspend with no docs at all: if the
# pid is still alive after resume, RAM survived. If only the file survived,
# it stopped, whatever the status string claims.
run "$id" 'nohup sh -c "while :; do date +%s > /work/tick; sleep 1; done" \
>/dev/null 2>&1 & echo $! > /work/pid' >/dev/null
# Idle it the way the platform wants to be idled. Do this BOTH ways at
# least once: an explicit stop and a timed-out stop are frequently
# different code paths with different resume costs, and usually only one
# of the two is documented.
if [ "$IDLE_WAIT" -gt 0 ]; then
sleep "$IDLE_WAIT"
else
curl -sS "${H[@]}" -X POST "$API/v1/sandboxes/$id/pause" >/dev/null
fi
t2=$(ms)
curl -sS "${H[@]}" -X POST "$API/v1/sandboxes/$id/resume" >/dev/null || true
until run "$id" 'cat /work/mark' >/dev/null 2>&1; do sleep 0.02; done
resume=$(( $(ms) - t2 ))
disk=$(run "$id" 'test -s /work/mark && echo disk-yes || echo disk-no')
ram=$(run "$id" 'kill -0 "$(cat /work/pid)" 2>/dev/null \
&& echo ram-yes || echo ram-no')
printf 'create=%sms resume=%sms %s %s\n' "$create" "$resume" "$disk" "$ram"
curl -sS "${H[@]}" -X DELETE "$API/v1/sandboxes/$id" >/dev/null
}
for _ in $(seq "$TRIALS"); do probe; doneThe two assertions at the bottom are the whole point. A background counter whose pid is still alive after resume proves RAM survived -- that is a real suspend. A file that survived while the pid did not proves it was a stop, no matter what the status string said. Run it with `IDLE_WAIT` set as well as unset, because an explicit stop call and a timed-out stop are often two different code paths, and typically only one of them is documented. There is a longer treatment of the measurement methodology in How to Benchmark Sandbox Cold Start Honestly.
Cohort one: it stops, and the disk is the product
These are the platforms built around a person and a repository. The design centre is that your work-in-progress must not be lost, so the volume is sacred and the compute is disposable. They are good at that. The cost is that resume is a boot plus your setup, every time.
1. GitHub Codespaces
The canonical shape of this cohort: an inactivity timer stops the container, the disk is retained, and a longer retention window eventually deletes the whole thing. Resume is a container start followed by whatever your `postStartCommand` does, and prebuilds only help the create case, not the resume case. Idle cost is storage, billed per unit time, which means a forgotten environment is a small permanent line item rather than a large temporary one -- the failure mode is forty of them, not one. Verify the current default timeout, the storage rate and the retention window against their docs; all three have moved and all three change the arithmetic.
2. Gitpod / Ona
Historically the most thoughtful implementation of the stop-and-restore model in this cohort, with explicit handling of workspace content across a stop rather than a naive volume detach. The newer runner model, where workspaces execute inside your own cloud account, moves the idle question off their bill and onto yours -- which is better for control and worse for predictability, since a stopped instance's idle rate is now your cloud provider's business. Resume is prebuild-dependent in a way worth measuring directly: a prebuild that is stale resumes very differently from one that is fresh. Verify the current state names and retention behaviour against their docs.
3. Coder
The honest outlier, because Coder does not really have an opinion about idle -- your Terraform template does. Auto-stop is a template setting, what persists is whatever you marked persistent, and the idle rate is your provider's stopped-instance rate. That makes it the most controllable platform here and also the slowest to come back, because resume is a provisioning run rather than a container start. If your template creates a VM, a disk, a DNS record and a load balancer entry, your resume path includes all four, and the only way to know what that costs is to run the probe against your own template rather than against Coder.
4. Replit
Product-shaped rather than fleet-shaped: an idle workspace or app sleeps, the filesystem persists, and waking it is fast enough that a human does not file a ticket. The cohort-one caveat applies in full though -- the filesystem came back and your processes did not, so anything you had running needs restarting, and the platform's own agent restarting it for you is a convenience rather than a guarantee. Fine for a person. Check the plan-level storage and wake behaviour before you build a fleet of them, which is not what this one is for.
Cohort two: there is no server, so there is no idle
Two platforms sidestep the entire question, in opposite directions, and both deserve credit for it. An idle cost of exactly zero is unbeatable, and they get there by having nothing to idle.
5. StackBlitz / WebContainers
The environment is a WASM runtime in your browser tab. Idle means you switched tabs; closed means it is gone. Nothing survives except what the project store or a git remote has, and resume is a cold boot of the whole runtime -- which is remarkably quick, costs nobody anything, and loses every byte of process state unapologetically. For teaching, reproductions, docs that run, and "click here to try it", this is the right answer and the rest of this table is overkill. For anything that needs a real kernel, a database, a native toolchain or more than one process tree that outlives a tab, it is not a candidate.
6. DevPod
A client that drives somebody else's compute, so its idle behaviour is a passthrough: DevPod will stop a machine on inactivity, and what that costs and what survives is entirely a property of the provider you pointed it at. Resume is a machine start plus a devcontainer re-up, and the re-up is the part that surprises people, because a devcontainer that has to rebuild is doing create-shaped work inside a resume-shaped wait. Excellent choice when you want the devcontainer standard without a vendor's control plane. Measure per provider, not once.
Cohort three: it suspends, and that is the rarest feature here
These are the three where RAM can genuinely survive, which means they are the only ones in the table where resume can be meaningfully cheaper than create rather than more expensive. They are also the three where you most need to read the fine print, because suspend and stop live behind adjacent API calls with very different bills.
7. Fly.io Machines
The most explicit lifecycle model in the roundup: a machine can be stopped, and it can be suspended with its memory checkpointed, and those are separate verbs with separate costs and separate resume shapes. That distinction is exactly the one this post is about, and having it in the API rather than in a blog post is worth a lot. A stopped machine costs you its rootfs and volumes; a suspended one additionally costs you the memory image. Resume from suspend reads memory back, which is a different and generally faster path than a start. Verify current availability, limits and rates against their docs -- this surface has been moving.
8. Northflank
A full application platform where scale-to-zero is a service-level setting rather than an environment concept: the container goes away, volumes and build artifacts stay, and the first request after idle pays a container start. The useful property for this axis is that preview environments and the services inside them are modelled separately, so you can let the compute go to zero while the database and the object storage behind it stay. Idle cost is storage plus whatever the control plane charges for the definition. Verify both against their current docs.
9. Railway
Similar shape, more opinionated defaults: idle services can be put to sleep and woken by the first inbound request, volumes persist, and you are billed for storage rather than compute while nothing is happening. Resume is a container start behind a request that is now waiting, which is the right trade for a staging app and the wrong one for an interactive shell. As with the rest of this cohort, confirm the current sleep semantics and which plan tiers they apply to before you design around them.
Cohort four: idle is a deletion, and the state lives somewhere else
The last cohort treats idle as a non-state. There is no stopped environment to pay for because there is no environment. What this buys you is an idle cost that rounds to the price of an object in a bucket. What it demands is that you have an answer to the question these four were designed around: what is the artifact?
10. Daytona
Of the sandbox-shaped platforms, the one with the most explicit ladder of idle states rather than a single idle flag -- a running sandbox, a stopped one, and a colder archived tier, each with its own cost and its own come-back time. That is the right model for this axis and I wish more of the category copied it. It also means the probe has to be run once per rung, because "resume time" is three different numbers on this platform and only the fastest one shows up in demos. Verify the current state names, archive thresholds and per-state rates against their docs.
11. E2B
Built for agent code execution, where the native lifetime of an environment is one task, so the default disposition is a timeout that ends it. There is pause-and-resume functionality in the surface -- verify the current semantics, retention and limits against their docs -- and the important question to settle before relying on it is whether a resumed sandbox gets its memory back or only its filesystem, because that is the difference between continuing an agent's session and restarting it. Strong default choice if your unit of work is a task rather than a workspace.
12. Modal
The scale-to-zero purist of the group: containers come and go, state belongs in volumes or object storage, and idle compute is simply not a concept you are billed for. Modal has also done real engineering on making the cold path fast rather than hiding it behind a warm pool, which is philosophically the same bet I made. It is also the platform on this list I would recommend over my own without hesitating for anything involving a GPU, because it has them and I do not. Verify snapshot and memory-snapshot behaviour against their current docs; this is an area they have been actively shipping.
13. Vercel Sandbox
Deliberately short-lived, with a maximum duration rather than an idle timer, which makes it the cleanest entry in the table: there is no idle state, no suspend, nothing to pay for at rest, and no resume path at all. You re-create. For build steps, untrusted transforms and agent tool calls inside a Vercel-shaped application, that constraint is a feature and the lack of a resume path costs you nothing. For an environment a human comes back to, the constraint is the whole answer. Verify the current maximum duration and resource limits against their docs.
14. PandaStack (mine)
My platform is in cohort four on purpose, and the design decision that puts it there is the one I would defend hardest: there is no warm pool. Every create restores a baked Firecracker snapshot through a pre-allocated network slot, which measures p50 179 ms and p99 around 203 ms, with the `/snapshot/load` step itself in the 49-80 ms range. The first ever spawn of a template with no baked snapshot is a real cold boot of roughly three seconds, and it bakes the snapshot on its way out so the next one is on the fast path.
Because create is that cheap, idle does not have to be a state. For app hosting, scale-to-zero means the sandbox is genuinely deleted -- netns released, CoW rootfs gone, no slot held -- and the thing that survives is a baked snapshot in object storage. Idle cost is therefore storage, not compute, and a wake is not a special slower code path: it is a create. The only extra variable is artifact locality. If the host that wins the placement already holds the snapshot bytes, you are on the published restore path. If it does not, it has to fetch them first, and the honest published shape of that is our cross-host fork number -- 1.2 to 3.5 seconds against 400 to 750 milliseconds on the same host. That is the same physics: local bytes versus remote bytes.
Two mechanisms make the remote case less bad than a download. Guest memory can be paged on demand from object storage over HTTP range requests in 4 MiB chunks, with a zero-chunk bitmap so empty memory is never fetched and a prefetch trace recorded at bake time so the hot set arrives before the guest faults on it. And memory restore is `MAP_PRIVATE`, so pages are copy-on-write rather than copied -- which is also why a fork can be hundreds of milliseconds instead of a memcpy of four gigabytes.
Fourteen platforms, four questions
| Platform | Idle behaviour | What survives | Cost while idle | Resume path |
|---|---|---|---|---|
| GitHub Codespaces | Stops on an inactivity timer | Disk; the container is killed | Storage, until retention deletes it | Container start plus your own setup |
| Gitpod / Ona | Stops; newer runners sit in your cloud | Workspace content, per their model | Theirs or your cloud's -- verify | Workspace start, prebuild-dependent |
| Coder | Template-defined auto-stop | Whatever Terraform marked persistent | Your provider's stopped-instance rate | A provisioning run, not a container start |
| Replit | Sleeps when nobody is connected | Filesystem; processes do not | Plan-level storage -- verify | Wake plus restarting your processes |
| StackBlitz / WebContainers | The tab closes; nothing is running | Only the project store or a git remote | Nothing -- there is no server | A fresh in-tab boot; fast and state-free |
| DevPod | Passthrough to the provider's stop | The provider's disk | Your own infrastructure's idle rate | Machine start plus a devcontainer re-up |
| Daytona | A ladder: running, stopped, archived | Per rung; archive is colder | Lower on the colder rung -- verify | Different per rung; probe all three |
| E2B | Timeout ends it; pause exists -- verify | What pause retains, else what you copied | Near zero when gone; storage if paused | Resume from pause, or a fresh create |
| Modal | Scales to zero by construction | Volumes and your own artifacts | Storage; no idle compute | Cold container start, snapshot-assisted |
| Vercel Sandbox | No idle state; a max duration instead | Nothing inside the sandbox | Nothing | Re-create; there is no resume |
| Fly.io Machines | Stop, or suspend with a memory image | Volumes always; RAM only on suspend | Rootfs and volumes, plus the image | A start, or a suspend read-back |
| Northflank | Scale-to-zero at the service level | Volumes and build artifacts | Storage plus control plane -- verify | Container start behind the first request |
| Railway | Sleeps, woken by a request -- verify | Volumes | Storage, not compute | Container start behind the first request |
| PandaStack | The sandbox is deleted; a snapshot remains | Disk and RAM, inside the snapshot | Object storage only; no compute | A create: restore at 179 ms p50 |
Read the last column as the real product comparison. Four of the fourteen have a resume path that is the same code path as create, and all four get there the same way -- by not keeping anything running. Three have a genuine suspend. The rest have a boot with your setup script bolted onto it, which is fine, as long as you measured it.
The round trip, on my own platform
Here is the same probe expressed in the PandaStack SDK, which has the advantage that destroy-and-restore is the normal path rather than an exotic one. Note what is being timed: not the snapshot, not the teardown, just the wait a developer or an agent actually experiences.
"""Resume-path measurement: create, dirty the disk, snapshot, destroy,
restore -- and time only the part a developer actually waits on."""
import statistics
import time
from pandastack import Sandbox
TRIALS = 20
def timed(fn):
t0 = time.perf_counter()
out = fn()
return out, (time.perf_counter() - t0) * 1000.0
# 1. Create. On a template whose snapshot is already baked this is the
# snapshot-restore path: p50 179 ms, p99 around 203 ms on our fleet.
# The FIRST ever spawn of an unbaked template is a real cold boot of
# roughly 3 s, and it bakes the snapshot on the way out.
# ttl_seconds is an IDLE clock, not a wall clock -- a busy sandbox is
# never reaped out from under a running build.
sbx, create_ms = timed(
lambda: Sandbox.create(template="base", ttl_seconds=900)
)
# 2. Put state in it that the image does not have: a repo, a dirty file, a
# timestamp. This is the half of an environment that "the volume
# persisted" never covers.
sbx.exec("mkdir -p /work && git -C /work init -q && date -Is > /work/when")
sbx.filesystem.write("/work/notes.md", "half-finished thought\n")
# 3. Snapshot is synchronous and hands back an id. It is the only thing
# that makes the next line survivable.
snap, snap_ms = timed(sbx.snapshot)
# 4. Destroy it. Not stop, not pause -- gone. Teardown is kill(); there is
# no sbx.delete(), and the netns, tap device and CoW rootfs go with it.
sbx.kill()
# 5. Resume IS create here. Same restore path, one extra variable: whether
# the host that wins the placement already holds the snapshot bytes.
resumes = []
for _ in range(TRIALS):
restored, ms = timed(lambda: Sandbox.create(from_snapshot=snap))
note = restored.exec("cat /work/notes.md")
assert note.exit_code == 0 and "half-finished" in note.stdout
resumes.append(ms)
restored.kill()
print(f"create {create_ms:8.0f} ms")
print(f"snapshot {snap_ms:8.0f} ms")
print(f"resume p50 {statistics.median(resumes):8.0f} ms")
print(f"resume max {max(resumes):8.0f} ms")
# Note on the OTHER two resume paths, which are not the same thing:
# sbx.fork() -> clones the DISK; the child COLD BOOTS with its own
# entropy, pids and clock. 400-750 ms same host,
# 1.2-3.5 s when the bytes have to cross hosts.
# sbx.fork_tree(4) -> inherits the parent's RUNNING memory, and therefore
# its RNG state too. Capped at 16 children.One API note that has bitten people, including me: a one-shot `exec()` does not reliably enforce a `timeout_seconds` argument, so bound long commands in the guest with `timeout(1)` instead of trusting the client. And if you want output as it happens rather than at the end, `exec_stream` yields stdout, stderr and exit events as they arrive. Idle timeouts and what does and does not reset them are covered properly in How to control sandbox lifetime: TTL, idle, and cleanup.
Where I would not use my own platform
If what you want is a managed browser IDE with dotfiles, an onboarding flow and a new hire productive on Monday morning, buy GitHub Codespaces or Gitpod/Ona and do not think about it again. That is a real product category with a decade of polish in it, and I have none of that polish: no browser IDE, no dotfiles story, no "open in editor" button that your team will actually use. My create call is faster than theirs and that is not the question those teams are asking.
If you need a GPU, use Modal, or a GPU cloud. PandaStack has no GPUs and no GPU passthrough, and nothing about Firecracker's device model is going to change that for me soon. The guest kernel is 5.10 and Ubuntu 24.04 on top of it, which is stable and boring and also means the newest kernel surfaces are not available to you -- if your workload needs a recent eBPF feature or a 6.x-era syscall, it will not find it. And if your model is one long-lived stateful service that should suspend with its RAM intact and come back on the same volume, Fly's explicit suspend verb is a better fit than my delete-and-restore model, because I am optimised for the case where the environment is disposable and the snapshot is the durable thing.
There is also a structural honesty point about cohort four generally. An idle cost of near zero is only a saving if the artifact is small and the restore is reliable. If you snapshot a 4 GiB guest and keep two hundred of them, you have replaced a compute bill with a storage bill, and you should go read your own object storage invoice before you congratulate yourself. Storage is cheaper. It is not free.
Picking one
Pick by the verb you need, not by the create latency on the landing page:
- A person comes back to it daily and must not lose work: cohort one. Codespaces, Gitpod/Ona, or Coder if you want the template under your own control. Accept that resume is a boot plus setup and budget for it.
- A long-lived service that should keep its memory across idle: cohort three, and specifically the platform whose API has an explicit suspend verb rather than a blog post about one.
- A program or an agent creates and discards these constantly: cohort four. The question becomes what the artifact is and how fast it restores, and the answer should be a published number, not a feature name.
- Teaching, demos and reproductions: StackBlitz, and stop evaluating.
- Anything with a GPU: not me.
And whichever one you pick, run the probe before you commit. The resume number is the one you will live with ten times a day, it is not on anyone's pricing page, and it takes about twenty minutes to measure. There is a deeper breakdown of where a wake's milliseconds actually go in Where Scale-to-Zero Wake Time Actually Goes, and a companion roundup graded on the cost of the second environment rather than the idle one in Top 12 Ephemeral Development Environment Platforms in 2026: The Cost of the Second One.
The bottom line
Ephemeral was never the hard part. Anyone can delete a container. The hard part is being able to delete it and still have the developer's afternoon back in under a second, and that requires having decided, explicitly, what the durable artifact is and how fast it restores. Platforms that answer that question tend to answer it in their API, with a verb. Platforms that have not answered it tend to answer it in their docs, with an adjective. The probe loop tells you which kind you are buying in about ten minutes, and it is the only part of this post I would ask you to actually run.
Frequently asked questions
What is the difference between stopping, suspending and destroying an ephemeral environment?
They differ in exactly one dimension that matters: whether RAM survives. A stop kills the processes and keeps the disk, so the volume comes back and nothing that was running does -- your dependencies are installed and your dev server, language server index and warm JIT are gone. A suspend checkpoints memory to storage and keeps the disk attached, so process state survives and a resume can genuinely be faster than a create; the price is storing a memory image and usually continuing to hold a slot. A hibernate is a suspend whose memory image has been moved somewhere colder, or whose host has been released, so it is cheaper at rest and slower to return, and it may return on a different machine. A destroy keeps nothing except what you deliberately exported -- a commit, a volume, a snapshot -- and its resume path is not a resume at all, it is a create from that artifact. The practical consequence: stop, suspend and hibernate all require finding something again at wake time, which is why their resume is a different and usually slower code path than create. Destroy has only one path, so its wake time is the number the vendor already publishes.
Why is resuming an idle development environment often slower than creating a brand new one?
Because resume has to reassemble a relationship that create gets to establish fresh, and five separate things can go wrong with that. Re-placement: the scheduler must find a host again, and your old host has drained, been patched, or filled with someone else's work, while create benchmarks are run against a fleet with room. Re-attachment: a detached volume has to be reattached and a volume lives in one zone, which turns a placement preference into a placement requirement. Re-hydration: image layers get evicted from node caches, so one resume hits a warm cache and the next pulls hundreds of megabytes, with nothing in the API response telling you which you got. Re-binding: the URL, DNS record or editor port-forward pointing at the old instance has to point at the new one, and that propagation is in nobody's create benchmark while frequently being the largest term a user actually perceives. And finally re-running your setup: everything that was not in the image, which on a stop-based platform is all of your running processes. Create pays none of these because it never had a previous incarnation to reconcile with.
How do I benchmark the resume path fairly across platforms?
Write a probe rather than reading docs, because the gap you are measuring is between documented and implemented behaviour. Create an environment and time to first successful command. Then write state into it that the image does not contain -- a file with random content, plus a background process writing a timestamp every second and recording its own pid. Idle the environment twice over separate trials: once with the platform's explicit stop or pause call, once by sitting out its inactivity timer, because those are frequently different code paths with different costs and usually only one is documented. Then resume and time to first successful command again. Finally check two things: does the file exist, and is the background pid still alive? A live pid proves RAM survived and you have a real suspend. A surviving file with a dead pid proves it was a stop, whatever the status string claimed. Run at least ten trials, report the median and the maximum rather than the best, and run it against your own template or image rather than the vendor's hello-world, since your setup script is part of the resume path whether the vendor counts it or not.
Does PandaStack keep anything running while a sandbox is idle, and what does idle actually cost?
No, and that is deliberate. There is no warm pool of idle VMs anywhere in the design. For app hosting, scale-to-zero means the sandbox is genuinely deleted: the network namespace is released, the copy-on-write rootfs is gone, no host slot is held. What survives is a baked Firecracker snapshot in object storage, so the idle cost is object storage, not compute. A wake is not a special slower path -- it is a create, and create via snapshot restore measures p50 179 ms and p99 around 203 ms, with the snapshot load step itself in the 49-80 ms range. The one extra variable at wake time is artifact locality: if the host that wins placement already holds the snapshot bytes you are on that published path, and if it does not it must fetch them first. The honest published shape of a remote fetch is our cross-host fork number, 1.2 to 3.5 seconds, against 400 to 750 milliseconds on the same host. On-demand memory paging over HTTP range requests, a zero-chunk bitmap and a bake-time prefetch trace exist specifically to make the remote case behave less like a download. One caveat on arithmetic: snapshots of a 4 GiB guest are not free, so a large sleeping fleet trades a compute bill for a storage bill rather than eliminating it.
If an environment is destroyed, how do I get my uncommitted work back?
Only through an artifact you created on purpose, which is why cohort-four platforms live or die on how good their snapshot story is. On PandaStack the primitive is a synchronous snapshot that returns an id, and a later create from that id restores both disk and memory state -- so the uncommitted file, the dirty git index and the process state are all inside it. The two branching verbs are not interchangeable and this has been shipped wrong before: fork() clones the disk and the child cold boots with its own entropy, pids and clock, taking 400 to 750 milliseconds on the same host and 1.2 to 3.5 seconds across hosts; only fork_tree() inherits the parent's running memory, and therefore also inherits its RNG state, which is a correctness concern rather than a performance one, and it is capped at 16 children. Teardown is kill(), not delete(). The habit worth building is treating the snapshot the way you treat a commit: something you take at a decision point, not something you hope the platform took for you. A platform that cannot name the artifact it keeps for you is not offering ephemeral environments, it is offering deletion with a longer timeout.
Keep reading
- Scale-to-zero wake latency anatomy — Where a wake's milliseconds actually go, step by step, on a platform that deletes instead of stopping.
- Controlling sandbox lifetime: TTL and idle timeouts — Why ttl_seconds is an idle clock rather than a wall clock, and which requests reset it.
- Snapshot restore vs warm pools — The architectural bet behind having no idle state at all, and what a warm pool costs you instead.
- Storage tiering for sleeping workloads — What a large fleet of sleeping snapshots costs once you read the object storage invoice.
- Top 13 ephemeral development environment platforms — The same category graded on the caller instead: whether create is a real API and what fifty at once does.
Related posts
- Top 12 Ephemeral Development Environment Platforms in 2026: The Cost of the Second One
The first ephemeral environment is a demo and every platform wins it. The second one is the product. Twelve platforms graded on idle cost, teardown semantics, whether uncommitted work survives a reap, and whether the isolation boundary is a kernel or a polite suggestion.
- Top 8 Ephemeral Development Environment Platforms (2026)
Feature grids do not decide this. Two numbers do: how fast a fresh environment can exist, and what happens to it when nobody is looking. Eight platforms graded on both, with the substrate each one actually isolates with and the catch I would want before the purchase order.
- Top 7 Ephemeral Development Environment Platforms in 2026
An environment is ephemeral when it is created from a definition, nobody is sad when it dies, and the 400th costs the same as the 4th. Most "cloud dev environments" fail at least one of those tests.
- Top 10 Disposable Development Environment Platforms in 2026
Creating environments is the easy half. This is a roundup graded on the hard half: what actually gets destroyed when you destroy one, what survives that you did not plan for, and who else feels it.
- Top 11 Ephemeral Development Environment Platforms (2026)
Every roundup in this category sorts by substrate or by price list. Both are downstream of a question nobody asks on the evaluation call: who, or what, is going to call create? Eleven platforms sorted by their creator — human, CI job, pull-request webhook, agent API — because the creator decides how many environments exist at once and who pays for the idle ones.
More in Snapshots & forking · See Thaw: sub-second cold restore
49ms p50 cold start. Fork, snapshot, and scale to zero.