all posts

Top 9 Ephemeral Development Environment Platforms (2026)

Ajay Kumar··11 min read

Three different products are sold under the phrase "ephemeral development environment," and almost every comparison I can find quietly mixes all three into one ranked list. That is why those articles are useless. It is a roundup of "vehicles" that grades a cargo bike, a minivan and a container ship on the same axis and declares a winner on cup holders.

The split that matters: category (a) is a cloud IDE or workspace for a human who will sit in front of it for hours — GitHub Codespaces, Coder, the Gitpod-shaped products, DevPod. Category (b) is a preview or deploy environment per branch or pull request — Vercel and Netlify previews, Railway and Render, Northflank, Okteto. Category (c) is a programmable sandbox API for machines — E2B, Modal, Daytona, Fly Machines, Vercel Sandbox, my own PandaStack. A platform excellent at one is usually mediocre at the other two, and that is design rather than failure. The expensive mistake is buying from the wrong category and then filing bug reports about a product being bad at a job it never applied for.

I'm Ajay. I build PandaStack, which is a Firecracker microVM sandbox platform and sits squarely in category (c). It is entry nine. I have tried to write the other eight the way I would want a competitor to write mine: what it is, where the isolation boundary sits, who it is for, what it is genuinely best at, and the drawback I would want before the purchase order rather than after.

Ground rules, because they change how you should read the rest. Specific numbers appear for exactly one platform — mine — because those are the only ones I can measure. Every other platform is described qualitatively from its public documentation, with no invented pricing, limits or internals. And this category reshuffles faster than blog posts get updated: pricing, quotas, isolation details and even product names move on a timescale of months, so verify anything I say about anyone else against that vendor's current docs before you commit budget.

The three categories, properly defined

(a) A workspace for a human

Optimised for session length. It lives for hours or days, holds your shell history and half-finished branch, and comes back the way you left it. A six-second start is fine for someone about to spend four hours in there. Idle is the enemy: a human closes a laptop and the environment keeps breathing unless something stops it. The boundary is usually a container on a per-user VM, which is reasonable because the code in it is yours.

(b) A URL per branch

Optimised for reviewability. The artifact is not a machine, it is an immutable deployment with an address you can paste into a pull request, and start time is usually dominated by your build rather than the platform. The hard part here is not compute, it is state: a preview of a frontend is easy, a preview that includes a database with representative data is where products in this category distinguish themselves or quietly punt.

(c) A sandbox as an API call

Optimised for create-and-destroy throughput at low per-unit cost. Nobody looks at these. A program calls an API, gets an isolated machine, runs something possibly hostile or merely stupid, reads the output, throws the machine away. Cold start is the headline number because something is literally blocked on it. Untrusted code is the default assumption rather than an edge case, which is why most serious entrants ended up at a VM boundary rather than a container one. A container is a polite suggestion to a shared kernel, and model-generated code does not reliably take hints.

The nine

1. GitHub Codespaces — category (a)

A workspace defined by a `devcontainer.json` in your repository, provisioned on managed cloud hardware and driven from VS Code in the browser or your desktop editor. Boundary: a container on a per-user cloud VM — root inside the container, the host is not yours. Fine for your own code, not what I would pick for code a model wrote. Best at the shortest distance between a repository and a human with a working editor; prebuilds take the sting out of the slow first create. Honest drawback: the most consequential setting in the product is the idle timeout, and it lives in a personal preference pane rather than in the repo. Billing has compute and storage dimensions, and storage keeps accruing on a stopped workspace. It is also not an API-first primitive — you can automate it, but creating thousands programmatically is not the shape of the product.

2. Coder — category (a), self-hosted

An open-source control plane that provisions workspaces from Terraform templates, so a workspace can be a Kubernetes pod, a cloud VM or a Docker container — whatever your providers can create. Boundary: therefore whatever your template provisions, which is both the best and most dangerous thing about it, because nothing stops you believing you have a VM boundary while running a shared kernel. Best at self-hosting with real policy attached: your cloud account, your VPC, your audit log, your existing modules. Honest drawback: you now operate a control plane, and ephemerality becomes a rule you enforce with autostop policies rather than a property you receive — the failure mode is a fleet of workspaces that are ephemeral in the same sense that a lease is temporary. Check their current licence tiers against their docs; the open-source and enterprise line has moved.

3. Vercel Preview Deployments — category (b)

Every push produces a build and an immutable preview URL; the unit is a deployment, not a machine. Boundary: your code runs in Vercel's managed function runtime and you do not get a shell you own — design, not omission. (They also ship a separate programmatic code-execution sandbox, which belongs in category (c); do not confuse the two because they share a brand.) Best at zero-configuration per-PR URLs; this is the product that changed reviewer behaviour across the industry, largely because of how little effort it took. Honest drawback: it is a deploy target, not an environment. No persistent process to SSH into, no daemon to poke, and your database is not part of the preview unless you wire branching up yourself. If your review needs a stateful backend with seeded data, you are assembling that from parts.

4. Render — category (b)

A container platform-as-a-service with pull-request preview environments declared in `render.yaml`, including copies of managed Postgres and Redis alongside your services. Boundary: containers on shared hosts, datastores as managed services. Best at previews that come with a database — the part most of category (b) skips, and the part that decides whether a preview is a demo or a real test. Honest drawback: a preview is a copy of your services, so cost scales with open pull requests, and open pull requests are a metric no engineering organisation has ever successfully bounded. Preview expiry is a config value set once during adoption and never read again; verify the current defaults for preview lifetime, and what stays billable after expiry, in their docs.

5. Northflank — category (b), with bring-your-own-cloud

Build pipelines, services, managed databases and preview environments in one product, with the option to run the same control plane inside your own cloud account. Boundary: containers, Kubernetes-shaped underneath; in BYOC mode the nodes are yours, which changes the compliance conversation substantially without changing the kernel-sharing story. Best at the combination of previews, managed data services and BYOC, a genuinely uncommon intersection. Honest drawback: more concepts than a two-field deploy product, and you will learn the object model whether you planned to or not. And the Kubernetes shape means a container boundary — if your real requirement is untrusted code, no amount of platform polish changes the boundary you are relying on.

6. E2B — category (c)

An open-source sandbox API built for AI code execution: an SDK call in, an isolated microVM out, with filesystem and process APIs designed for a model to drive rather than a person to type into. Boundary: microVM, own kernel, which is the right call for the job. Best at SDK ergonomics for the code-interpreter shape — the API reads like it was designed by someone who had actually tried to make a language model use it — and being open source is a real answer to the "what happens when you get acquired" question every infrastructure buyer should ask. Honest drawback: deliberately not a human IDE and not an application host. If you want a long-lived URL serving production traffic from a persistent process, that is a different category. Verify current limits, regions and the self-host story against their docs.

Python-first serverless compute: decorate a function, declare its image in code, run it remotely including on GPUs, plus a sandbox API for arbitrary commands. Boundary: their own hardened runtime around containerised workloads — the specifics have changed over the product's life, so read their current security documentation rather than any blog post, including this one. Best at Python developer experience and the container-start engineering underneath it, which is impressive work. Honest drawback: the mental model is functions and images, so an environment is something you define in code rather than something you shell into and explore. The Python-centricity is a feature right up until the thing you need to run is not Python, at which point you are wrapping shell commands in a Python decorator and wondering who you have become.

8. Fly Machines — category (c), the primitive

A REST API over Firecracker microVMs you create, start, stop and destroy individually, with anycast networking and fine-grained billing. Boundary: microVM, own kernel. It is infrastructure, offered as infrastructure. Best at stop-to-zero machines with a real public address in many regions — the closest thing on the market to a VM as an API call with a URL attached, where stop and start are first-class rather than bolted on. Honest drawback: you own the orchestration. Placement, lifecycle, reaping, snapshot semantics, what happens when a host dies — that is your code now. Exactly the right trade if you are building a product on top, and exactly the wrong one if you thought you were buying one.

9. PandaStack — category (c), mine

Firecracker microVM sandboxes where every create is a snapshot restore. There is no warm pool of idle VMs: the create path allocates a pre-built network namespace, reflinks the rootfs, forks Firecracker and loads a baked snapshot. Create is p50 179 ms, p99 203 ms, of which `/snapshot/load` is roughly 49 to 80 ms; the first cold boot of a template, before a snapshot exists, is about 3 seconds. Each sandbox gets its own network namespace, veth pair and tap device from 16,384 pre-allocated /30 subnets per host, and preview URLs are tokenless — a port is reachable at a hostname containing the sandbox UUID, and the UUID is the credential. Boundary: microVM, own kernel. `fork()` is a disk-plus-memory snapshot restore of a running parent, 400 to 750 ms same-host and 1.2 to 3.5 seconds cross-host, which is the feature I would actually buy this for: get a VM into an expensive state once, then branch it a hundred times. Managed Postgres databases are each a dedicated microVM with a durable volume, and creating one takes 30 to 90 seconds because it blocks until Postgres is genuinely accepting connections.

Now the drawbacks, which are sharper than the pitch. Firecracker cannot change a guest's vCPU count or RAM at snapshot restore, so RAM is chosen at template build time with `--memory-mb` and baked in; passing `memory_mb` on a create is not an error, it is silently corrected to the baked value, which is worse than an error. Our first-party templates bake at `base` 4 GiB, `code-interpreter` and `agent` 2 GiB, `browser` 4 GiB, `postgres-16` 1 GiB, each with 8 burstable vCPUs. The guest kernel is 5.10 under an Ubuntu 24.04 userspace, there is one kernel per host, and you cannot swap it per template. There is no `/dev/kvm` inside a guest, so no nested virtualisation. `fork()` restores the parent's RNG state and clock identically, so forked guests diverge in ways that genuinely surprise people — seed your RNG and re-sync the clock, or enjoy a long debugging session. Neither exec endpoint enforces `timeout_seconds` server-side; it is a client deadline. Rootfs copy-on-write is local-only because copy-on-write needs a local block device — we can stream snapshot memory on demand from object storage over userfaultfd, but the disk is always a local file. And there is no human IDE: if you want a person in an editor for six hours, buy from category (a) and do not let me talk you out of it.

Nine platforms, three categories. Isolation boundaries for other vendors are from their public documentation and may have changed — verify before you buy.
PlatformCategoryIsolation boundaryBest for
GitHub Codespaces(a) human workspaceContainer on a per-user cloud VMRepo-to-editor in one click for GitHub-native teams
Coder(a) human workspace, self-hostedWhatever your Terraform template provisionsWorkspaces inside your own compliance boundary
Vercel Preview Deployments(b) URL per branchManaged function runtime, no shellZero-config visual review of frontend changes
Render(b) URL per branchContainers plus managed datastoresPer-PR stacks that include a real database
Northflank(b) URL per branch, BYOCContainers, Kubernetes-shapedPreviews plus addons plus bring-your-own-cloud
E2B(c) sandbox APIMicroVM, own kernelAgent code execution with model-friendly SDKs
Modal(c) sandbox API, function-shapedHardened container runtime (check their docs)Python and GPU workloads without infrastructure work
Fly Machines(c) sandbox API, primitiveMicroVM, own kernelA stoppable VM with a public address, as an API call
PandaStack(c) sandbox APIMicroVM, own kernel and netnsSnapshot-restore creates and forking a running state

The questions that actually decide it

Nobody picks one of these on a feature matrix. They pick on about eight questions, and answering these honestly collapses the field to two or three candidates before you open a pricing page.

  • Who creates the environment — a human clicking a button, or an API call from a program? This one question eliminates two of the three categories.
  • Does untrusted code run in it? If a language model or a stranger's pull request generates the commands, you need a kernel boundary, and no amount of platform polish substitutes for one.
  • How long does one live — minutes, hours, or "until someone notices"? Minutes means cold start dominates. Hours means idle cost dominates. "Until someone notices" means you are buying a pet.
  • What does idle cost, and what exactly stops? Compute, the root disk, the attached volume, the managed database, the load balancer and the reserved address are six separate line items, and products stop different subsets.
  • Does it need a database, and does that database branch with the code? A preview without representative data tests the CSS and nothing else.
  • Does it need a public URL that a stranger, a webhook or a reviewer can reach? Some products give you this in one line; with others you are building ingress.
  • What is the cold-start budget of whatever is waiting? A human tolerates seconds. A PR check tolerates seconds. An agent in a tool-use loop, charged per token while it waits, does not.
  • Can you self-host it, and — separately — do you actually want to? Different questions, and the second has a higher bar than people admit at procurement time.

The cost trap: the ephemeral environment that isn't

Every platform here describes its environments as ephemeral. Ephemeral is a claim about intent; the invoice is a claim about fact, and the two part company with impressive regularity. Everything is ephemeral until the invoice arrives. The failure is always the same shape: someone adopts per-branch or per-task environments, adoption grows, and nothing reaps. Six months later there are two hundred environments, eleven of them in use, and the bill looks like a small production estate because that is precisely what it is. They did not stop being ephemeral by design; they stopped being ephemeral in practice, which is the only kind that bills.

Line item one: what idle actually costs

"Stopped" is not a single state. Ask specifically: when stopped, are you billed for compute? The root disk? The attached volume? The managed database the preview created? The load balancer and the reserved IP? I have seen products where stopping the compute leaves four of those five running, which from a finance perspective is not stopping anything, it is turning off the one part that was cheap.

Structurally, idle should cost roughly what the state costs to store and nothing for the machine. Hibernating a PandaStack sandbox snapshots it and releases the VM, so the resting cost is a snapshot on a disk rather than reserved RAM. That is the whole argument for snapshot-based platforms, and it is worth saying plainly that it is an argument about billing shape first and speed second.

from pandastack import Sandbox

# Lifetime as a declared property, not a hope. The TTL is enforced by the
# platform reaper, so a crashed orchestrator cannot leak this sandbox.
with Sandbox.create(template="base", ttl_seconds=900,
                    metadata={"branch": "feat/checkout", "owner": "ci"}) as sbx:
    sbx.exec("git clone --depth 1 https://github.com/acme/app /srv/app", check=True)
    sbx.exec("cd /srv/app && npm ci && npm run build", check=True)
    print(sbx.preview_url(3000))
# Context-manager exit kills it. The TTL is the backstop for the case where
# your own process is the thing that died.

Note the two independent mechanisms. The context manager handles the happy path; the TTL handles the path where your orchestrator gets OOM-killed halfway through, which is the path that actually generates the bill. If a platform offers only one of those, the one it offers is probably the happy path.

Line item two: TTL semantics

A time-to-live sounds like a single number and is actually a cluster of decisions, each of which can quietly convert an ephemeral environment into a permanent one.

  1. Does the clock run from creation, or from last activity? Absolute TTLs reap predictably; idle TTLs reap only if "idle" means what you assume.
  2. What counts as activity? This is where it goes wrong. We shipped a version where a read-only status GET bumped the idle timer, so a dashboard polling the sandbox list kept every one of them alive forever — the monitoring was the workload. We fixed it by gating the activity marker, but I tell the story because every platform with an idle timer has some version of this bug, and the way to find it is to ask which requests touch the timer.
  3. Does expiry delete or merely stop? A stopped environment still holds a disk, and a disk is a line item.
  4. Does the TTL survive a control-plane restart? If the reaper's state lives only in memory, a deploy is an amnesty.
  5. Can a user extend it, and is there a ceiling on the extension? "Extend indefinitely" is a button that manufactures pets.

None of this is exotic. It is just the part of the evaluation nobody does during the trial, because during the trial you have three environments and all three are in use.

If untrusted code runs in it, the category is decided for you

One asymmetry overrides most other preferences. If the code was written by a language model, submitted by a stranger, or pulled from a dependency tree you have not audited, then the boundary between that code and your other tenants is the product and everything else is ergonomics. A container shares the host kernel, so every local privilege escalation in a kernel you did not choose is a tenant escape in yours. A microVM gives the workload its own kernel and a device model small enough to reason about. (Firecracker's device list is short, but current versions do offer an opt-in virtio-PCI transport and support ACPI, so the old flat claims that it has neither are stale; our guests still boot MMIO with `pci=off`.)

And while drawing that boundary: a client-side timeout is not a limit. If an SDK accepts a timeout parameter, find out whether the server enforces it or merely stops waiting. On ours it is a client deadline, so real ceilings go in the guest where a client that hangs up cannot bypass them.

# Hard limits belong in the shell, inside the guest, where the client cannot
# opt out of them. A client-side timeout stops you waiting; it does not stop
# the process.
timeout --signal=KILL 30s \
  bash -c 'ulimit -v 1048576 -u 256 -f 524288; exec python3 /tmp/model_wrote_this.py'

Picking one

  • A human needs an editor and onboarding is broken — category (a). GitHub-native: Codespaces, and set the idle timeout on day one. Compliance boundary matters more than convenience and you have a platform team: Coder.
  • A reviewer needs a URL — category (b). Frontend-dominant: Vercel-shaped previews are hard to beat on effort. If the preview needs a database to mean anything: Render, or Northflank when the cloud account has to stay yours.
  • A program creates them — category (c). Agent code execution with good SDKs: E2B. Python and GPUs: Modal. Building a platform and want the raw primitive: Fly Machines. Snapshot-restore creates, forking a running state, and idle that costs storage rather than RAM: PandaStack, with the baked-RAM and 5.10-kernel constraints understood before you start rather than after.
The most expensive environment in your estate is the one that was described as temporary in a design document and has since acquired a DNS record.

If you take one thing from this: decide the category before you shortlist, and read the idle-cost and TTL pages before the feature pages. All three categories move quickly enough that the only durable advice is about which questions to ask — so ask them against each vendor's current documentation, including mine.

Frequently asked questions

What actually counts as an "ephemeral" environment?

Three tests, and it has to pass all of them. First, it is created from a definition rather than by hand — a devcontainer file, a Terraform template, a Dockerfile, an API call — so recreating it is mechanical. Second, nobody is upset when it dies: all state worth keeping lives in git, in object storage or in a real database, not in that environment's home directory. Third, the four-hundredth one costs about the same per unit as the fourth, which rules out anything whose cost structure depends on a human remembering to clean up. Notice that none of those tests mention the word "cloud" or how fast it starts. A remote environment can be the most pet-like thing you own, and plenty of them are: long-lived, hand-modified, holding uncommitted work, and known by a nickname. Ephemeral is a property of the lifecycle, not of the hosting location.

Container or microVM for untrusted code?

MicroVM, if the code is genuinely untrusted. A container shares the host kernel with every other tenant, so your isolation story reduces to "no local privilege escalation will ever be found in the kernel we happen to run," which is not a story, it is a hope with a CVE feed attached. A microVM gives each workload its own kernel and a deliberately small device model, so an escape has to get through a virtio driver and a VMM rather than through a syscall surface measured in hundreds. The cost used to be start time, which is why people accepted the container boundary for so long; snapshot restore removed most of that cost — on our platform a create is p50 179 ms, which is well inside what anything waiting on it will tolerate. There are legitimate reasons to stay on containers: your code, your team, your audited dependency tree, and a density target a VM boundary cannot meet. "A model generated this command" is not one of them.

Which category is cheapest when environments sit idle?

Whichever one stops the most line items, and that is a question about the specific product rather than the category. The thing to measure is not the hourly rate, it is what remains billable in the stopped state. A stopped environment can still be charging for its root disk, an attached volume, a managed database the preview created, a load balancer and a reserved IP address. I have seen setups where stopping the compute turned off the cheapest of the five. Architecturally, the lowest idle cost comes from platforms that treat the resting state as stored bytes rather than reserved capacity — snapshot the memory and disk, release the machine, pay for the snapshot. That is how hibernation works on PandaStack, and the honest framing is that it is a billing-shape argument first and a speed argument second. Ask every vendor the same question and make them itemise: with the environment stopped, what am I still paying for?

Can I self-host any of these?

Several, with different meanings of the word. Coder is open source and designed to be self-hosted; the workspaces are provisioned by your own Terraform into your own cloud. E2B is open source, so the sandbox runtime can be run yourself. Northflank offers bring-your-own-cloud, where their control plane operates against nodes in your account — not the same as self-hosting, but it answers most of the questions that drive the request. Codespaces and Vercel previews are managed services and that is the deal. Verify all of this against current docs, because licensing in this category changes. The harder question is whether you want to. A workspace control plane or a microVM fleet is a system with a pager: hosts to patch, snapshots to garbage-collect, network namespaces to reclaim, capacity to schedule. We run that, and it is most of where our engineering time goes. If the person who would operate it is already on four rotas, the hosted option is cheaper than it looks on the invoice.

Do I really need one environment per pull request?

Per pull request is usually the wrong granularity to start with — start per merge queue or per nightly and let demand pull you toward per-PR. The reason is that per-PR environments have a cost curve shaped like your open-PR count, which no engineering organisation has ever bounded, and a value curve that depends entirely on whether the environment contains representative data. A per-PR preview of a frontend with a mocked API tells a reviewer that the CSS compiles. A per-PR environment with a branched database containing real-shaped data tells them whether the change works, and that is worth paying for. So the honest sequencing is: get the environment definition into the repo first, make one correct environment with real data, then decide how many you want. If you do go per-PR, set an absolute TTL rather than an idle one, and confirm that expiry deletes the disk rather than merely stopping the compute.

Keep reading

Related posts

  • Top 10 Disposable Development Environment Platforms in 2026

    Creating environments is the easy half. This is a roundup graded on the hard half: what actually gets destroyed when you destroy one, what survives that you did not plan for, and who else feels it.

  • The Best Replit Alternatives in 2026

    "Replit alternative" means three unrelated things: a browser IDE, a cloud dev environment, or the sandbox API that runs untrusted code under the hood. Split the intent first and the shortlist collapses from ten options to two.

  • PandaStack vs Coder

    These two products are shopped against each other constantly and compete almost never. One gives a person a machine for the week; the other gives a program a machine for four seconds. Naming that split is most of the decision.

  • Cloud Dev Environments on microVMs

    Ephemeral dev environments want three things at once — real isolation, fast start, and the ability to freeze and resume. A microVM gives you all three, which is why it's a better substrate than a shared-kernel container.

  • A Disposable Dev Environment per Git Branch

    Every branch gets its own fully-isolated microVM dev environment — checkout done, deps installed, services running — that forks in sub-second and disappears when the branch is gone. Ephemerality is what finally kills environment drift.

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.