all posts

Top 10 Disposable Development Environment Platforms in 2026

Ajay Kumar··11 min read

There is an environment in your cloud account that was disposed of six weeks ago. The ticket is closed. The thread where somebody wrote 'torn down, we're good' has scrolled out of history. The Kubernetes namespace really is gone, the DNS record really is gone, and the GitHub deployment really does show as inactive. There is also a 40 GiB provisioned volume in the 'available' state, an idle load balancer, a database final-snapshot nobody will ever restore, and a log group with a retention policy of 'never'. Nothing here involves anyone lying. Everyone ran the teardown script. The teardown script ended every single line with || true.

Creating environments is the solved half of this category. Every platform below can create one; the demos are all good; the create path is where product teams spend their engineering effort because it is what you see in the first five minutes. Disposal is the unglamorous half, it is where the money and the incidents live, and almost nothing grades on it. So that is what this roundup grades on.

A note on scope. Yesterday I published this post's companion, a roundup of the same field graded on ephemerality as a property — created from a definition, nobody sad when it dies, the 400th costing what the 4th did. That one answers 'is it ephemeral at all'. This one assumes you said yes and asks the operational question underneath: what breaks when you throw environments away three hundred times a day. Different failure modes, different winners in a couple of places, and it is linked first at the bottom.

Disclosure, because it should change how you read the rest: I'm Ajay, I build PandaStack, and PandaStack is entry ten. The rules I hold myself to in a roundup — specific numbers only for my own product, everything else described qualitatively from public documentation with no invented pricing, limits or internals, and a standing instruction to verify anything I say about someone else against their current docs, because this category reshuffles faster than blog posts get updated. Entry ten also carries the longest 'what leaks' paragraph in the post, and that is not false modesty; it is the only row I can inspect the source code of.

Disposal is a transaction, and nothing gives you one

An environment is not a thing. It is a set of resources spread across systems that have never heard of each other: a namespace in a cluster, a record in a DNS zone, a database instance, a bucket, a secret in a vault, a row in a metering table, a load balancer, maybe a licence seat somewhere that bills per active workspace. Creating it is a sequence of calls. Destroying it is also a sequence of calls, and there is no two-phase commit spanning Kubernetes, Route 53, your secret store and your observability vendor. Every disposal is a best-effort fan-out that can partially succeed, and partial success is indistinguishable from success unless somebody designed for the difference.

Three distinct failure modes, which get conflated into 'cleanup is hard' and are worth separating because they have different fixes:

  1. Partial destroy. Some resources are gone, some are not, and the system of record says 'destroyed'. This is the expensive one, because the survivors are now invisible: no environment claims them, no dashboard groups them, and the only thing that will ever find them is a reconciliation job you have not written. The fix is not a better script — it is a destroy that reports per-resource outcomes plus something independent that re-checks later.
  2. State in the wrong place. Either something you needed survived inside the disposable thing — four hours of manual QA data, the auth session, the build cache that took eleven minutes to warm — or something you wanted gone survived outside it, like a cloned secret or a copy of production data in a bucket that outlives its creator. Same design error both ways: the boundary between 'disposable' and 'durable' was never drawn explicitly, so it got drawn by accident at 3am.
  3. A destroy with a blast radius. The environment shared something and tearing it down was felt by a neighbour: a node, a Docker daemon, a Postgres server with a schema per branch, a conntrack table, an ingress controller. The destroy succeeded and somebody else's test suite went red, which is how teams discover their disposable environments were never independent.

There is a fourth thing that is less a failure mode than a discipline: destroy order is not create order reversed, because some facts have to be read out of a resource before the resource containing them disappears. Two examples from our own delete path, both of which are comments in the code because both were bugs. The final egress meter has to read the host-side veth counter before the veth pair is destroyed — release the network first and every byte the environment transmitted since the last metering tick is gone, a revenue leak shaped exactly like a short burst followed by a delete. And a sleeper's tiered memory image lives in object storage, with a pointer file inside the environment's directory; read the pointer after removing the directory and nothing records that the object exists, so it bills until somebody audits the bucket. That is what 'ordering matters' means in practice.

The teardown script that cannot fail

Here is the artifact at the centre of this problem. You have one. I have written one. Look at the right-hand margin rather than the commands.

#!/usr/bin/env bash
# teardown-preview.sh — the script almost every team has.
set -euo pipefail          # disarmed, line by line, by everything below

ENV="$1"

kubectl delete namespace "pr-${ENV}" --wait=false                    || true
aws rds delete-db-instance --db-instance-identifier "pr-${ENV}" \
        --skip-final-snapshot                                        || true
aws s3 rb "s3://pr-${ENV}-assets" --force                            || true
aws route53 change-resource-record-sets --hosted-zone-id "$ZONE" \
        --change-batch "file://delete-${ENV}.json"                   || true
vault kv delete "secret/preview/${ENV}"                              || true
gh api -X DELETE "repos/acme/web/deployments/${DEPLOY_ID}"           || true

echo "environment ${ENV} destroyed"   # printed whether or not that is true
exit 0                                # green in CI. forever.

Each || true was added by a reasonable person for a reasonable reason: the line failed once, at 2am, on an environment that had already been partly removed, and it blocked a deploy. Making the destroy idempotent is the correct instinct — deleting a thing that is already gone should not be an error. The mistake is implementing idempotency by discarding the exit code, because that also discards 'the API rejected this', 'you lack the permission', 'the resource has a dependent that must go first' and 'the delete is still in progress'. The script stops distinguishing between absent and unreachable, which are the two outcomes you most need to tell apart.

The honest version is barely longer. Treat 'already gone' as success explicitly, treat everything else as failure, accumulate instead of aborting, and exit non-zero if any resource survived — then let a reconciler, not the script, have the last word.

#!/usr/bin/env bash
# A destroy that reports. Same resources, different contract.
set -uo pipefail
ENV="$1"; FAILED=()

# try <label> <cmd...> : success, or success-because-already-absent, or record it.
try() {
  local label="$1"; shift
  local out rc
  out=$("$@" 2>&1); rc=$?
  if [[ $rc -eq 0 ]]; then return 0; fi
  # The ONLY tolerated failure: the resource is genuinely not there.
  if grep -qiE 'notfound|does not exist|no such|NoSuchBucket' <<<"$out"; then
    echo "ok    ${label} (already absent)"; return 0
  fi
  echo "FAIL  ${label}: ${out%%$'\n'*}"; FAILED+=("$label"); return 0
}

# Order matters: read metering + logs BEFORE the thing that produced them dies.
try "usage-flush"  ./bin/flush-usage --env "$ENV"
try "logs-archive" ./bin/archive-logs --env "$ENV"

# Then dependents before owners, and anything billable before anything cosmetic.
try "dns"       aws route53 change-resource-record-sets --hosted-zone-id "$ZONE" \
                  --change-batch "file://delete-${ENV}.json"
try "rds"       aws rds delete-db-instance --db-instance-identifier "pr-${ENV}" \
                  --skip-final-snapshot
try "bucket"    aws s3 rb "s3://pr-${ENV}-assets" --force
try "namespace" kubectl delete namespace "pr-${ENV}" --wait=true --timeout=120s
try "secret"    vault kv delete "secret/preview/${ENV}"

if (( ${#FAILED[@]} )); then
  # Leave the environment marked as draining, not destroyed. The reconciler
  # retries; the dashboard keeps showing it; nobody closes the ticket.
  ./bin/mark-env --id "$ENV" --state draining --unresolved "${FAILED[*]}"
  echo "partial teardown: ${FAILED[*]}" >&2
  exit 1
fi
./bin/mark-env --id "$ENV" --state destroyed

Six axes, chosen before the list

Picking criteria after you have seen the products is how you end up buying the one with the best landing page. These are the six that actually decide whether you can throw environments away at volume, in roughly the order that settles the argument.

  1. The definition of record, and whether it is the only one. Not 'is there a config file' — every platform has one — but: if the environment vanished right now, is the file enough? The practical test is whether anyone has ever fixed an environment by shelling into it. The moment the running machine knows something the definition does not, disposal destroys the only copy of that knowledge. Drift is not an aesthetic problem; it is what makes a disposable environment undisposable.
  2. Teardown correctness. What a destroy actually destroys, what it leaves, and whether it tells you the difference. Count the systems one environment touches, then ask who owns the destroy of each, whether that destroy is idempotent for the right reason, and whether anything reconciles afterwards. A platform that owns every resource it created can give you a correct destroy; a platform orchestrating resources in your cloud account mostly cannot, and should say so.
  3. Where the state that must survive lives. Every real environment has some: the database somebody spent four hours seeding, the authenticated session that took an OAuth dance to get, the caches that are the difference between a 40-second rebuild and an 11-minute one. Each either lives inside the disposable thing — in which case disposal costs something real and people stop disposing — or outside it, in which case the outside thing is now the long-lived one and you have moved the problem rather than solved it. The good answers make the surviving state explicit, named, and separately destroyable.
  4. Blast radius of a destroy. Who else notices. Namespaces on a shared cluster: a destroy is API-server work plus a pile of terminating pods on nodes other people are using. A shared Docker daemon: a pruning mistake takes someone else's layers. A shared Postgres: a DROP SCHEMA holds locks. One machine per environment with its own kernel and subnet: the destroy is a kill and a free-list push, and the only shared resource is capacity.
  5. Reconstitution time from zero. Not resume-from-suspend — from nothing, with the definition as the only input. This number sets your disposal policy whether you choose one or not. Under a second and deletion is the obvious idle strategy. Thirty seconds and people suspend. Eight minutes and they keep environments for weeks, and the long-lived machine is back, now named pr-4417.
  6. Idle cost at rest — including the cost of an environment you believe you already disposed of. Two numbers: what a live-but-untouched environment bills, and what the residue of a destroyed one bills. Nobody measures the second, which is why it surfaces in a quarterly review as an unexplained line.

The ten

Numbered for the headline. This is not a ranking — four of these are approaches rather than vendors, and the right answer genuinely inverts depending on which axis binds you. Everything about a platform that is not mine is qualitative, drawn from public documentation, and should be checked against their current docs before you commit to anything.

1. The hand-rolled destroy (Terraform or Pulumi plus a bash wrapper)

What it is: the baseline, and statistically the thing you are actually running. An IaC module per environment, a CI job that applies it on branch creation and destroys it on merge, and a wrapper script around the edges doing the things the IaC module does not model — the DNS record somebody added by hand, the secret, the SaaS webhook, the seat.

Disposes well because: the definition of record is genuinely good. A Terraform module is reviewable, versioned and plannable, and terraform destroy against a healthy state file is the most correct destroy in this entire list. State and dependency ordering are the problem it was built to solve, and it solves them.

What leaks: everything outside the state file, which is always more than you think. Resources created by the application at runtime rather than by Terraform. Resources created by a previous version of the module and then removed from it. Anything a human touched during an incident. And the state file itself, which is the environment's soul and lives in a bucket — lose it and the resources become permanently invisible, because the destroy no longer knows they exist. Reconstitution from zero is a full provision cycle, usually minutes, which is why these environments get reused, which is how drift gets in. Honest grade: best definition, worst latency, and a teardown whose correctness is exactly as good as your discipline about never touching anything by hand.

2. GitHub Codespaces

What it is: a managed environment attached to a repository, defined by the devcontainer specification, opened from a browser or a desktop editor. The default most teams evaluate first, because it is one click from where the code lives.

Disposes well because: the platform owns the compute and the disk, so deleting a codespace is one operation against one system — no fan-out, no reconciliation problem. That is the single biggest advantage a vendor-owned environment has over an orchestrated one. The devcontainer file is also the most portable definition format in this list.

What leaks: not resources — habits. The environment retains uncommitted work by design, which is a feature right up until it becomes the mechanism by which a codespace turns into a thing people are afraid to delete. The resource leak is upstream: anything the environment creates in your cloud account during a test run is not the platform's to destroy, and the devcontainer spec has no opinion about it. Verify current retention, idle and billing behaviour against GitHub's docs rather than anyone's table.

3. Gitpod

What it is: the project that argued for strong ephemerality before the category caught up — an environment started per task from a declarative config, prebuilt when the commit lands rather than when a developer asks. Its product shape has moved significantly in recent years toward running on infrastructure you control, so treat any description, mine included, as a snapshot.

Disposes well because: prebuilds are the best answer anybody in this category has produced to the reconstitution-time axis, and reconstitution time is what determines whether your team is willing to dispose at all. Build the environment when the commit lands, and the cost of throwing one away collapses. If you take one idea from Gitpod into whatever you end up using, take that one.

What leaks: in the bring-your-own-infrastructure shape, teardown correctness moves to your side of the line along with the compute — a real trade, not a criticism. This is the entry whose details go stale fastest; read the current deployment model, config format and idle behaviour from their docs, not from a roundup, including this one.

4. Coder

What it is: self-hosted environments whose shape is expressed in Terraform. One design decision explains the product: a workspace can be a Kubernetes pod, a cloud VM or a bare-metal box — whatever a provider can create — and the platform's job is lifecycle and access rather than owning the compute.

Disposes well because: it is the only entry that puts a control plane in front of the hand-rolled approach without taking the Terraform definition away. You get a registry of what exists, attribution, autostop policy and an admin who can see the fleet — exactly the machinery entry one lacks.

What leaks: Terraform-shaped workspaces inherit Terraform-shaped teardown — the resources are in your account and the destroy is a provision cycle in reverse. A workspace backed by a VM with a disk trends long-lived because recreating it is expensive, so the common leak here is not an orphaned resource but an orphaned workspace: declaratively provisioned, functionally permanent. You also operate the control plane. Verify current lifecycle, autostop and licensing against their docs.

5. DevPod

What it is: a client-side tool rather than a service. It reads the devcontainer spec and creates that environment on a provider you choose — laptop, cloud VM, Kubernetes — with no central control plane in the middle.

Disposes well because: on a developer's own machine, disposal is genuinely trivial and genuinely local, and the open definition format means you are not betting your inner loop on one vendor's proprietary environment description.

What leaks: no control plane means no registry, and no registry means no reconciliation — the single most important mechanism in teardown correctness. Nothing holds the list of environments that should exist, so nothing can tell you about the one that still does. On a cloud provider that becomes real money distributed across engineers' personal workflows, where no cost dashboard groups it. Fine at ten people; a platform project at two hundred. Confirm the provider list and the project's maintenance status before building process on it.

6. Namespace-per-branch on Kubernetes (with or without vcluster)

What it is: not a product, but the most widely deployed disposable-environment architecture in the industry. A namespace per branch or per pull request, a Helm release or Kustomize overlay to fill it, an ingress wildcard for URLs, and an operator or CI job that destroys the namespace on merge. The virtual-cluster variants give each environment its own API server view on top of a shared data plane.

Disposes well because: the namespace is a real boundary with a real cascading delete. Everything with an owner reference inside it goes away in one API call, which is a stronger destroy primitive than most of this list has. Plain namespaces also reconstitute reasonably fast when images are cached on the nodes, and the definition is in the repo as charts.

What leaks: everything a namespace does not contain, which is a long list once you start writing it down — cluster-scoped objects, PersistentVolumes whose reclaim policy is Retain, external DNS records, cloud load balancers whose controller failed to finalise, webhook configurations, and anything created by an operator in its own namespace on your environment's behalf. Add namespaces stuck in Terminating behind a finalizer that no longer has a controller, which look destroyed in every dashboard and are not. Blast radius is the real weakness: environments share nodes, the API server, the ingress controller, the CNI and the conntrack table, so a destroy is work other tenants feel and a noisy environment degrades its neighbours. This entry is the clearest example of the gap between 'the delete returned 200' and 'the environment is gone'.

7. Docker Compose and devcontainers on a shared build host

What it is: the pragmatic middle. A compose file per environment, a beefy shared host or a per-developer VM, and docker compose down at the end. Often the fastest thing a small team gets working, and often what a bigger team is quietly still using in one corner.

Disposes well because: the definition is in the repo, reconstitution is fast when layers are cached, and compose down with the right flags really does remove containers and networks. Hard to beat on effort-to-value for a single-developer loop.

What leaks: named volumes, which compose down does not remove unless you ask, and which accumulate silently until a disk fills at the least convenient moment. Dangling images and build cache, which grow until somebody prunes and takes a colleague's layers with them. And the daemon is the shared dependency: one process, one image store, one network address space, one set of host ports — so two environments that both want 5432 are not independent. Blast radius is the whole host.

8. Nix, devenv and flake-defined environments

What it is: the environment as a pure function of a pinned input set, materialised into a store rather than into a machine. A flake and a lock file define the toolchain exactly; entering the environment is a shell, not a provision.

Disposes well because: it wins the first axis outright and it is not close. The definition is not merely authoritative, it is total — there is nothing to drift into, because the thing you enter is derived from the lock file every time. And disposal is close to a non-event: there is no environment to destroy, only a shell to exit, and the residue is a content-addressed store you garbage-collect on a schedule. If the question is 'can I guarantee two engineers and CI have the same toolchain', this is the answer and the rest of the list is a compromise.

What leaks: not resources — scope. A Nix environment defines your toolchain, not your running system. The Postgres with test data in it, the queue, the object store, the third-party credential: none of that is in the flake, so you still need one of the other nine entries underneath for the parts of an environment that are services rather than binaries. This is the best definition layer in the list and not a complete answer, and pairing it with something that owns the compute beats choosing between them. Budget for the learning curve, which no amount of enthusiasm from whoever introduced it will shorten.

9. Programmatic sandbox APIs (the agent-shaped cohort)

What it is: environments as an API call rather than a workspace — create, run, destroy, with no editor attached and no expectation that a human ever looks at one. E2B, Daytona's current direction, and the sandbox products attached to larger platforms all sit here. The category grew out of AI agents needing thousands of short-lived environments a day, which is the most disposal-intensive workload anyone has yet built.

Disposes well because: disposal is in the API from the first page of the documentation, and the usage pattern forces honesty. When a workload creates environments faster than any human could, a teardown bug is not a slow leak — it is an incident within the hour. Products shaped by that pressure tend to have real idle timeouts, a real per-environment destroy, and a registry you can list.

What leaks: the durable artifacts these APIs offer precisely because fast disposal makes them necessary — saved snapshots or images, persistent filesystems, attached volumes. Each has a lifecycle separate from the environment that created it, separate billing, and usually a separate delete call, which is exactly where the residue of a 'disposed' environment hides. Read the isolation model too: the cohort ranges from a shared-kernel container to a dedicated virtual machine and the word 'sandbox' is used for both. Verify against current docs; this is the fastest-moving part of the field.

10. PandaStack

What it is, with the disclosure already made: an open-source platform that gives each environment a Firecracker microVM with its own guest kernel, created by restoring a baked snapshot rather than booting. Create is 179 ms at p50 and 203 ms at p99, of which the snapshot load itself is about 49 ms; the first spawn of a template before its snapshot exists cold-boots in roughly 3 s and bakes the snapshot on the way through. Each environment gets its own Linux network namespace, veth pair and tap device from a pool of 16,384 pre-allocated /30 subnets per host. The template is a Dockerfile, baked once.

Disposes well because: the platform owns every resource an environment consists of, so the destroy is one call to one system with no cross-system fan-out — the structural advantage the vendor-owned entries have over the orchestrated ones, taken as far as it goes. Three specifics follow from that. Reconstitution from zero is 179 ms at p50, which makes deletion the default idle strategy rather than a loss, which in turn means environments do not accumulate in the first place. Blast radius is a kill and a free-list push: the subnet is released synchronously on the delete path so the slot is immediately reusable, and because the environment had its own kernel, its own netns and its own CoW disk, there is nothing for a neighbour to notice except capacity. And preview URLs cost nothing to clean up, because there is no per-environment DNS record to leak — the URL is a wildcard host label of the form port-id.suffix, routed by the proxy, and it stops resolving to anything the moment the environment is gone.

What leaks, plainly, and this is the longest such paragraph in the post on purpose. Snapshots are durable by design and that cuts both ways. The idle reaper deliberately does not cascade-delete an environment's snapshots, because a snapshot is the thing you explicitly saved and an automatic reap must not destroy it — so the environment you let expire can leave a full memory-plus-disk image behind, billing as storage until the orphan reaper expires it, which is 7 days after the source environment disappears by default. Meanwhile an explicit kill does cascade, which is the opposite footgun: a tidy try/finally destroy can delete the snapshot you took two lines earlier to capture a failure. Volumes are worse, because they are the correct answer to 'state that must survive' and so people use them: a named volume outlives every environment it was ever attached to, bills on provisioned size rather than bytes written at $0.15 per GiB-month, and needs its own delete call, which is refused while anything is attached. That is the 'disposed of six weeks ago and still in the bill' shape, on my own product, by design. Two more. Attaching a volume forfeits the fast path entirely — the snapshot-restore create is gated on there being no volumes, because Firecracker can only patch drive IDs that already exist in the snapshot's device topology, so a volume-attached create cold-boots in about 3 s and a restore with volumes is refused outright rather than silently dropping the mount. And custom templates are durable artifacts with their own lifecycle: deleting every environment built from a template does not delete the template. Managed databases fail in the other direction — their delete purges the backup archive and is genuinely irreversible, which is the correct behaviour and also the one that will hurt somebody eventually.

The ten against the axes

Short cells, qualitative everywhere except the row I am allowed to be specific about. 'Verify' is not a hedge for its own sake — it means the honest answer depends on a detail that changes faster than this table does.

Disposable development environment platforms, graded on disposal rather than creation. Non-PandaStack rows are qualitative; verify against current vendor docs.
Platform or approachDefinition of recordWhat survives a destroyBlast radiusRebuild from zero
Hand-rolled IaC + bashTerraform module (strong)Anything outside the state fileYour cloud accountFull provision cycle
GitHub Codespacesdevcontainer.jsonUncommitted work, by designVendor-side, isolatedVerify current docs
GitpodDeclarative configVerify current docsDepends on deploymentPrebuild, then fast
CoderTerraformWhatever Terraform missedYour cluster or cloudFull provision cycle
DevPoddevcontainer.jsonUnknown — no registryYour chosen providerDepends on provider
K8s namespace-per-branchCharts in the repoCluster-scoped + retained PVsShared cluster, feltFast with warm images
Compose on a shared hostcompose.yamlNamed volumes, build cacheThe whole hostFast with warm layers
Nix / devenv / flakesFlake + lock (total)The store, GC'd on scheduleNone (no services)Seconds from a warm store
Programmatic sandbox APIsImage or templateSaved snapshots, volumesVerify isolation modelVerify current docs
PandaStackDockerfile baked to snapshotSnapshots (7d orphan grace), volumes, templatesOne microVM, own kernel179 ms p50 restore
The 'what survives' column is the one to read twice, including on my own row. Every entry in this table has a durable-artifact story, because fast disposal makes durable artifacts necessary — you cannot throw an environment away cheaply unless something cheaper holds the parts you needed. The failure mode is never that the artifacts exist. It is that their lifecycle is separate from the environment's, nobody owns it, and no dashboard groups them. Whatever you pick, the first thing to build is a list of your durable artifacts with an owner and an expiry for each.

The audit that tells you the truth

You cannot grade your own teardown from the destroy script's exit code, because that is the thing under test. You grade it by reconciling two lists built independently: what your control plane says exists, and what the underlying systems say exists. The gap is your leak, and the first run is always the embarrassing one. Two halves — your cloud account, and then your environment platform's own durable artifacts.

#!/usr/bin/env bash
# leak-audit.sh — reconcile reality against the registry. Run it weekly,
# in a cron job somebody gets paged about, not by hand when you remember.
set -euo pipefail

# 1. The claim: environment ids your control plane still admits to.
#    (Whatever your source of truth is. Here: the platform's own list.)
pandastack sandbox list | jq -r '.[] | .metadata["env-id"]? // empty' \
  | sort -u > /tmp/claimed.txt

# 2. Reality: every cloud resource tagged with an env-id, whatever it is.
#    This is why tagging on create is non-negotiable — an untagged resource
#    cannot appear in this list and therefore cannot ever be reclaimed.
aws resourcegroupstaggingapi get-resources \
     --tag-filters 'Key=env-id' \
     --query 'ResourceTagMappingList[].[ResourceARN,Tags[?Key==`env-id`].Value|[0]]' \
     --output text | sort -k2 > /tmp/actual.txt

# 3. The gap: resources whose environment no longer exists.
join -v1 -1 2 -2 1 <(sort -k2 /tmp/actual.txt) /tmp/claimed.txt \
  | tee /tmp/orphans.txt | wc -l | xargs echo "orphaned resources:"

# 4. The classic untagged leaks, checked by shape instead of by tag.
aws ec2 describe-volumes --filters Name=status,Values=available \
     --query 'Volumes[].[VolumeId,Size,CreateTime]' --output text
aws elbv2 describe-load-balancers \
     --query 'LoadBalancers[].[LoadBalancerArn,CreatedTime]' --output text
kubectl get ns -o json | jq -r \
  '.items[] | select(.status.phase=="Terminating") | .metadata.name' \
  | sed 's/^/stuck terminating: /'

# Exit non-zero so this is a signal, not a report nobody opens.
[[ ! -s /tmp/orphans.txt ]]

The second half is the platform's own durable artifacts, which is where the residue lives once your cloud account is clean. On PandaStack that is three lists — environments, snapshots, volumes — and the interesting rows are the snapshots whose source environment is gone and the volumes nothing has attached recently. Note the asymmetry the loop below is written around: an explicit destroy cascades to that environment's snapshots, while an idle reap deliberately does not, so the orphans you find are almost always from environments that expired rather than ones you killed.

import os
from pandastack import Client

client = Client(api_key=os.environ["PANDASTACK_API_KEY"])

live = {s.id for s in client.sandboxes.list()}

# Snapshots whose source environment no longer exists. These are NOT a bug:
# an idle reap deliberately does not cascade-delete snapshots, because a
# snapshot is the artifact you explicitly saved. They are expired by the
# orphan reaper after a grace period (7 days by default) — but until then
# they are storage you are paying for, so decide deliberately.
for snap in client.snapshots.list():
    if snap.get("sandbox_id") and snap["sandbox_id"] not in live:
        gib = (snap.get("size_bytes") or 0) / 2**30
        print(f"orphan snapshot {snap['id']}  {gib:5.1f} GiB  {snap['created_at']}")
        # client.snapshots.delete(snap["id"])   # irreversible; removes the GCS copy

# Volumes outlive every environment they were ever attached to, and bill on
# PROVISIONED size, not bytes written. This is the line item that survives a
# teardown nobody questioned.
for vol in client.volumes.list():
    print(f"volume {vol['name']:28} {vol.get('size_mb', 0) / 1024:5.1f} GiB provisioned")
    # client.volumes.delete(vol["name"])  # refused while attached to a live sandbox

Designing the disposal into the create call

The best teardown script is the one that does not have to run, because the platform will dispose of the environment on its own if your code never gets the chance. That means two things at create time: an idle timeout as a backstop, and a deliberate decision about which parts of this environment are allowed to survive it. Both go in the create call, not in a runbook.

import os
from pandastack import Client

client = Client(api_key=os.environ["PANDASTACK_API_KEY"])
PR = os.environ["PR_NUMBER"]

# ttl_seconds is an IDLE timeout, not a walltime budget: the reaper measures
# time since last activity (default 5 minutes). It is the backstop for the
# case where your destroy never runs — cancelled pipeline, crashed runner,
# laptop lid closed on an interactive debug session. Size it to the longest
# silence you consider plausible, NOT to the length of the task.
#
# No cpu= / memory_mb= here: when the template has a baked snapshot the agent
# silently corrects both to the baked values. Guest RAM is chosen once, at
# template build time (`pandastack template build --memory-mb`).
sbx = client.sandboxes.create(
    template="acme-dev",
    ttl_seconds=1800,
    metadata={"env-id": PR, "owner": "ci", "disposable": "true"},
)
print("preview:", sbx.preview_url(3000))   # no DNS record created, none to leak

rc = 1   # anything that raises below leaves this non-zero, i.e. "do not kill"
try:
    sbx.exec(f"git clone --depth 1 --branch pr-{PR} git@github.com:acme/web /work",
             check=True)
    # Long work goes through exec_stream. timeout_seconds is CLIENT-side only
    # -- the agent decodes it on both exec endpoints and applies it on neither,
    # so one-shot exec() dies at 30s. Hard deadlines belong in the shell.
    rc = sbx.exec_stream(
        "cd /work && pnpm install --frozen-lockfile && pnpm test",
        on_stdout=lambda c: print(c, end=""),
        timeout_seconds=1500,
    )
    if rc != 0:
        # Capture evidence BEFORE the destroy, and know what the destroy does
        # to it: an explicit kill() cascade-deletes this environment's
        # snapshots. Saving the failure and then killing in a finally block
        # destroys exactly the thing you saved.
        snap_id = sbx.snapshot()          # a snapshot ID string, memory + disk
        print(f"failed: restore with Sandbox.create(from_snapshot='{snap_id}')")
        raise SystemExit(1)
finally:
    # The explicit destroy. Synchronous, single system, no fan-out: the VM is
    # killed, the /30 subnet goes back on the free list, the final egress
    # meter is read before the veth is torn down, and the row is gone before
    # this call returns. On the failure path above we do NOT reach here —
    # we exited first, and the idle reaper will take the environment in 30
    # minutes WITHOUT touching the snapshot we just took. That asymmetry is
    # deliberate and it is the only reason the evidence survives.
    if rc == 0:
        sbx.kill()
The asymmetry in that last block is worth stating on its own, because it inverts the usual advice. A try/finally destroy is normally the responsible pattern. Here, an explicit delete cascades to the environment's snapshots while an idle reap does not — so if the thing you want to survive the environment is a snapshot you just captured, the responsible pattern is to let the environment expire rather than to kill it. Durable volumes and custom templates are not cascaded by either path; they have their own delete calls and their own bills.

The three things that always need to survive

Every team that tries to make environments genuinely disposable hits the same three pieces of state, in the same order, and the quality of a platform is largely how good its answers are.

  • The database. Somebody spent four hours getting it into an interesting shape and that work is not in git. The wrong answers are 'keep the environment alive', which makes it long-lived, and 'share one database across environments', which makes them non-independent. The right answer is a database recreatable from a definition — a seed script, a sanitised dump, a branch of a reference database — so the interesting state is reproducible rather than precious. If a fresh seeded database is a sub-minute operation the problem dissolves; if it is eight minutes, people keep environments and you cannot talk them out of it.
  • The authenticated session. The OAuth dance against a third-party sandbox, the device-flow token, the cookie that took a human to obtain. This is the state most likely to turn a disposable environment into a permanent one, and the only clean answer is to make the credential an input rather than something acquired inside: minted on create from a secret store, scoped to that environment, revoked by the destroy. A credential the destroy does not revoke is the most dangerous residue there is, because it still works.
  • The build cache. The eleven minutes of dependency install and compilation paid on every create. This is where a memory snapshot beats a disk image categorically, because it restores the warmed process and not just the warmed directory: the cache is in page cache, the language server is indexed, the JIT is warm. Short of that, in descending order: a prebuilt image, a shared cache volume (now you own a durable artifact with a lifecycle), a remote cache service (now you own a service).

Picking one

  • The definition must be total and identical for every engineer and for CI — Nix or devenv for the toolchain layer, paired with something that owns the services. Do not make it a choice between them.
  • Environments must live in your VPC under your IAM with an audit trail — Coder, and accept that teardown correctness is now a reconciliation job you own and should staff.
  • You already have a cluster and a platform team — namespace-per-branch, and spend the first week on the orphan reconciler for cluster-scoped objects, retained PersistentVolumes and namespaces stuck in Terminating. Not the second week.
  • A stranger must reach a running environment in one click — Codespaces, and plan separately for whatever the environment creates in your cloud account, because the platform will not destroy that for you.
  • Something programmatic creates and destroys thousands a day — the sandbox-API cohort, and read the durable-artifact lifecycle before the create API, because that is where the bill accumulates.
  • Idle cost at rest plus a destroy with no blast radius — snapshot-restore on a per-environment kernel, which is PandaStack here and why the entry exists. Then re-read what my volumes and snapshots do after a destroy, because that is the part that will surprise you.
  • A browser IDE, a GPU, or Windows — none of my product's strengths help you, and two of the three are permanent. Choose from the others.

The bottom line

Disposability is not a feature you buy. It is a property you keep, and it decays — every time somebody fixes an environment by hand instead of fixing the definition, every time a destroy fails quietly, every time a durable artifact is created with no owner and no expiry. A platform's job is to make that decay slow and visible: own as many of an environment's resources as possible so the destroy is one call rather than a fan-out, make reconstitution fast enough that deleting is easier than keeping, and name the durable artifacts out loud so somebody can be responsible for them. I built PandaStack around the second of those — at 179 ms nobody has to be disciplined, because recreating is less effort than remembering. The longest 'what leaks' paragraph in this post sitting on my own row is the third. Pick on the axis that binds you, then do the thing no vendor can do for you: run the reconcile, read what it finds, and delete something on purpose in the first week.

Frequently asked questions

What does "teardown correctness" mean, and how do I test my own?

Teardown correctness is the gap between what your destroy claims and what is actually gone. You cannot test it from the destroy script's exit code, because that is the thing under test — and in most organisations that exit code is a constant, because every line of the script ends in a construct that discards failures. You test it by reconciling two independently built lists: the environment ids your control plane still admits to, and every resource in the underlying systems tagged with an environment id. Resources in the second list whose environment is absent from the first are your leak. Then check the classics that are usually untagged and therefore invisible to that join: unattached block volumes, idle load balancers, retained PersistentVolumes, namespaces stuck in Terminating behind a dead finalizer, orphaned database snapshots, log groups with no retention policy, and credentials that still authenticate. Run it on a schedule, as a job that pages someone, and make it exit non-zero — a report nobody opens is the same as no report. The first run is always embarrassing; the value is entirely in the second one being better.

What usually leaks when a preview or dev environment is destroyed?

In rough order of how often it bites: storage volumes, because the destroy removed the compute and the volume had a Retain or a default-keep policy; DNS records, because the record was created by a different system than the one that destroyed the environment; load balancers, because a cloud controller failed to finalise and nothing retried; secrets and tokens, because revocation was never wired into the destroy path, which makes them the most dangerous residue since they still work; cluster-scoped Kubernetes objects, which a namespace delete does not cascade to; database snapshots, often created by the delete itself as a safety default; metering and log data, which is less a leak than an ordering bug — the usage has to be read out before the resource that produced it disappears. Underneath all of them is one structural cause: the environment was a set of resources across several systems, the destroy was a best-effort fan-out, and nothing reconciled afterwards. The platforms that leak least are the ones that own all of an environment's resources themselves, so the destroy is a single call against a single system.

Where should a disposable environment's database live?

Inside the disposable thing, but recreatable from a definition — and the second half is what makes the first half work. Sharing one database across environments makes the environments non-independent: a migration in one breaks the others, and the destroy of one holds locks that everybody feels. Keeping the environment alive to preserve its database is how a disposable environment becomes a pet, and it is the single most common way teams lose this property. So: own database per environment, created from a seed script, a sanitised dump, or a branch of a reference database, and treat the time-to-seeded-database as the number that decides whether anyone is willing to dispose. If it is under a minute, people throw environments away without thinking about it. If it is eight minutes, they will keep them, and no policy document will change that. On PandaStack specifically, a managed Postgres create is 30-90 s wall-clock to usable — the API returns 202 immediately and you poll rather than block — and a clone from an existing database's archive gives you a point-in-time copy into a new id without touching the source, which is the shape you want for per-branch data.

Does deleting a PandaStack sandbox delete its snapshots and volumes?

Snapshots: it depends on who did the deleting, and the asymmetry is deliberate. An explicit delete — your kill() call, or a feature deleting its own sandbox — cascade-deletes that sandbox's snapshots, including the copies in object storage. An idle or TTL reap deliberately does not, because a snapshot is a durable artifact meant to outlive its sandbox; that is the entire point of having taken one, and an automatic reap must not destroy something you explicitly saved. Snapshots whose source sandbox is gone are expired later by an orphan reaper after a grace period, 7 days by default. The practical consequence is counter-intuitive: if you have just captured a snapshot as evidence of a failure, letting the sandbox expire preserves it and a tidy try/finally kill() destroys it. Volumes: never. A named volume outlives every sandbox it was ever attached to, bills on provisioned size rather than bytes written, and has its own delete call, which is refused while the volume is attached to a running sandbox. Custom templates are the same — deleting every sandbox built from a template does not delete the template. If you take one thing away: list your snapshots and volumes on a schedule, because nothing else will.

How fast does rebuilding have to be before deleting beats pausing?

The threshold is roughly a developer's patience for a context switch, which empirically sits somewhere under a couple of seconds — and it matters because the choice between delete and pause is not really a cost decision, it is a behavioural one. If rebuilding from zero takes minutes, people will pause, and a paused environment keeps its disk, keeps its drift, keeps its uncommitted work and keeps a line in your bill, so you have chosen reuse for a latency reason rather than a design one. If rebuilding is fast enough to be unnoticeable, deletion becomes the obvious idle strategy and the whole class of problems in this post stops accumulating, because there is nothing left to leak. That is the number PandaStack is built around: 179 ms at p50 and 203 ms at p99 to restore a baked snapshot, with the snapshot load itself about 49 ms. The first create of a template before its snapshot exists cold-boots in roughly 3 s and bakes the snapshot on the way through, so only the first one is slow. Attaching a durable volume forfeits that path entirely and cold-boots in about 3 s, because Firecracker can only patch drive IDs that already exist in the snapshot's device topology — which is a genuine trade between speed and surviving state, and worth knowing before you design around either.

Keep reading

Related posts

  • Cloud Dev Environments on microVMs

    Ephemeral dev environments want three things at once — real isolation, fast start, and the ability to freeze and resume. A microVM gives you all three, which is why it's a better substrate than a shared-kernel container.

  • PandaStack vs Coder

    These two products are shopped against each other constantly and compete almost never. One gives a person a machine for the week; the other gives a program a machine for four seconds. Naming that split is most of the decision.

  • The Best GitHub Codespaces Alternatives in 2026

    Cloud dev environments split into two jobs now: a place for humans to write code, and a place for agents to run it. Codespaces is good at the first and awkward at the second.

  • The best Gitpod alternatives in 2026

    Gitpod taught everyone to expect a dev environment from a URL. Picking a replacement means deciding which half of that promise you actually needed: the ephemerality, or the machine.

  • Preview Environments on microVMs: a Live URL per PR

    Every pull request gets its own live URL backed by a real backend and database — on a Firecracker microVM, so even untrusted forked-PR code is isolated by hardware, not by a shared kernel.

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.