all posts

PandaStack vs Gitpod / Ona: A Workspace Is Not a Sandbox

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and the comparison I get asked for most often lately is against Gitpod. It arrives in the same shape every time: a team already pays for cloud development environments, they have started shipping something agent-shaped, and someone has asked the reasonable question of whether the platform they already own can be where the agents run. Then the evaluation stalls, because nobody can work out whether they are comparing two implementations of one thing or two different things that share a vocabulary.

They are two different things. Both create Linux environments on demand, both talk about ephemerality, both can clone your repo and run your tests, and both will happily take your money. But they are built around different nouns, and the noun is the whole argument. Gitpod's noun is a workspace: an environment a person opens, works inside, and closes, defined declaratively in the repo and prebuilt so opening it is fast. PandaStack's noun is a sandbox: a microVM an API call brings into existence, with no editor, no session, no dotfiles and nobody attached to it.

Those overlap in the middle, and the middle is where buyers make expensive mistakes in both directions. So this is the honest version, including the direction where my answer is "do not buy PandaStack for that" — a real direction, and I will mark it clearly.

First, the name, because it moves

Gitpod is the product most readers know by that name. The company has since moved to the Ona name, with a positioning aimed much more squarely at software engineering agents than at humans opening a browser IDE. If you are reading a two-year-old comparison post, it is describing a product boundary that no longer exists in that shape; if you are reading this one in a year, treat it the same way.

Everything I say about Gitpod / Ona here is deliberately qualitative — capabilities and shapes, not versions, prices, limits, quotas or SLAs. That side of the comparison is from public documentation, and in this category the docs move faster than blog posts do. Read their current docs before you buy anything, including before you believe my table.

I am also not going to treat a rebrand as evidence of anything. Repositioning toward agents is a correct read of where the demand went. The interesting question is whether a workspace is the right container for an agent.

The strongest case for Gitpod / Ona, and I mean it

Let me put their argument properly, because the weak version of it is useless to you. Gitpod made the case that development environments should be disposable, reproducible and defined in the repository, and they made it while most of the industry was still arguing about whether a laptop setup script counted as infrastructure. They were early, they were right, and expectations changed permanently because of it. If you have ever opened a URL and found a cloned repo with dependencies installed and services running, you are consuming an idea they popularised, whoever you pay for it.

Concretely, the things they do well are things PandaStack does not do at all:

  • A declarative environment definition checked into the repo, versioned with the code it describes, reviewed in the same pull request. The environment becomes a reviewable artifact rather than tribal knowledge.
  • Prebuilds: do the slow, boring part of setup ahead of time, per commit, so that opening an environment does not mean watching a dependency install. This is the single feature I most often hear teams say they cannot give up.
  • Deep editor integration. Browser IDE, desktop editor attachment, the terminal where you expect it, port forwarding that works, and the sense of a real machine rather than a web toy.
  • Deep Git forge integration: the environment knows the repo, the branch, the pull request, the review you are in the middle of.
  • A credible self-hosted / bring-your-own-infrastructure story, which matters enormously the first time a security review says source code may not leave accounts you control.
  • Human ergonomics as a first-class concern — dotfiles, personalisation, shell history: the small things that decide whether a team adopts a cloud environment or quietly goes back to laptops.

If your problem is "onboarding takes three days, everyone's laptop is subtly different, and the tenth engineer this quarter lost a day to a broken native extension build", that is their problem and they solve it well. Nothing below should be read as a suggestion that you solve it with an API that returns a microVM and no editor. You would be building their product badly, and your engineers would hate you for it.

The divergence is the unit, not the quality

A workspace is something that is opened, and that single verb drags a lot of design with it. If a thing gets opened, something closes it, so it has a session; it has a person (or now an agent occupying a person-shaped slot); its lifecycle is a working day or a pull request; it accumulates state the occupant would be annoyed to lose; and it is worth real engineering effort to make the open fast, because a human is sitting there waiting. Every good decision in a workspace platform follows from someone waiting at the other end.

A sandbox is something that is created. Nobody opens it. It has no session, no occupant, no personalisation and nothing it would be sad to lose, because whatever created it put in exactly what the task needed and will read the result out before deleting it. Its lifecycle is seconds to minutes, and its interesting property is not how pleasant it is to be inside — nothing in there has feelings — but how many can exist at once, how fast they appear, and how completely they vanish.

On PandaStack a sandbox is a Firecracker microVM, created by a POST and addressed by a UUID. No IDE, no workspace list to browse, no concept of "my environment". You get a machine with a kernel, a network namespace, an exec channel and a filesystem, and you get it back in about the time it takes to blink.

The honest disqualifier, stated plainly: if what you want is a browser IDE, prebuilds, a dotfiles story and a per-developer workspace your engineers log into, PandaStack does not have any of that and is not going to pretend otherwise. Do not buy it for that. I would rather lose the deal than spend six months as the wrong answer in your stack.

The mirror-image mistake is more expensive and less often named: picking a workspace platform because it can obviously run your agent — it can, it has a terminal and a repo and credentials — and then discovering six months later that you are provisioning environments shaped like seats for a workload shaped like a queue. The architecture holds right up until an agent run wants four hundred concurrent environments for ninety seconds each: a sentence that makes perfect sense in a sandbox product and none at all in a workspace product.

Side by side

Gitpod / Ona column is from public documentation and is deliberately qualitative -- no versions, prices or limits. Their product boundaries and plan structure move fast; verify against their current docs before you buy. The PandaStack column is from my own system.
DimensionGitpod / OnaPandaStack
Product unitA workspace, opened and closedA sandbox, created and deleted
Who or what is attachedA developer, or an agent in a developer-shaped slotNobody; the caller holds a UUID
Isolation substrateGenerally container-based on shared-kernel infrastructureFirecracker microVM, own guest kernel 5.10, tiny virtio device model
Environment definitionDeclarative file checked into the repo, reviewed with the codeBaked template snapshot plus mise runtime detection at deploy time
How "fast" is achievedPrebuilds: front-load setup per commit so the open is quickSnapshot restore on every create: p50 179 ms, p99 203 ms
Branching live stateBranch the repo; the workspace follows the branch`fork_tree(n)` inherits guest memory, 400-750 ms same-host
EditorBrowser IDE and desktop editor attachment, first classNone. There is no UI inside a sandbox
Credentials inside by defaultGit and often cloud credentials, because a human needs themNothing. You inject exactly what the task needs
Natural concurrencyRoughly one per developer per taskHundreds per run; the ceiling is host RAM and CPU
Pricing shapeSeat and workspace-hour thinking$0.054 per vCPU-hour, $0.0162 per GiB-hour, every class
Self-hostingBring-your-own-infrastructure options existOpen source; run the control plane and agents on your own KVM hosts
Best fitHumans developing software, and agents working alongside themMachine-created, untrusted, high-fan-out, short-lived execution

Isolation substrate, and why it only sometimes matters

Workspace platforms in this category are generally container-based, running on shared-kernel infrastructure — namespaces, cgroups and seccomp over a host kernel that many workspaces use at once. Check their current docs for the runtime specifics, because that is exactly the detail that changes between generations of a product. Every PandaStack sandbox, by contrast, is a Firecracker microVM (v1.16.0) with its own guest kernel (5.10, Ubuntu 24.04 userland) and a deliberately tiny virtio device model, exec'd into its own pre-allocated network namespace with its own veth pair and tap device.

Now the part most vendor comparisons skip: this difference is irrelevant most of the time. For your own team's code, written and reviewed by people you employ, a container is completely fine. The kernel is not your threat model. Nobody in your company is escalating privileges out of their own dev environment to read their own repo, and if they were, you have an HR problem rather than a hypervisor problem.

The substrate becomes load-bearing at one specific moment: when the code being executed was not written by anyone you can call. Two things produce that — a customer uploaded it, or a model generated it — and both arrived in force in the last two years. That is the shift that happened to this entire category while the vocabulary stayed the same.

A container is a polite suggestion to the kernel. A hypervisor boundary is an argument the guest is not party to.

Containers share a kernel by design — that sharing is the efficiency. So the isolation is restrictions applied by the kernel to processes still talking directly to that kernel, across a syscall surface that is very large and has a long history of interesting afternoons. A microVM's guest talks to its own kernel, and what it can reach outside is a handful of virtio devices. Both are real boundaries. One is much smaller, and small is the only property of an attack surface that reliably helps you.

The operational consequence matters more to me than the theoretical one. On a shared kernel, a workload whose signature is "lots of CPU, lots of syscalls, unusual pattern" looks identical whether it is a legitimate build, an infinite loop a model wrote, or something probing your kernel. At 3am you cannot tell a capacity page from a security page. One microVM per unit of work does not prevent abuse, but it makes every signal attributable to one tenant and one task, and the remedy a delete call.

Prebuilds and snapshot restores: two answers to different halves of "fast"

This is the axis where comparisons get most confused, because both sides have a good story and the stories are not about the same thing. Prebuilds attack setup cost. The expensive part of an environment is almost never the machine; it is `npm ci`, the Gradle daemon warming up, the Python wheel that compiles from source, the Docker layer pull. Prebuilds move that off the critical path, per commit, so when a human opens an environment the slow part has already happened. That is why raw provisioning speed is a vanity metric in the workspace world: a platform that provisions in three seconds and then installs dependencies for four minutes is slower than one that takes thirty seconds from a post-install image.

PandaStack attacks machine-creation cost, because at the fan-out we care about that is what dominates. Every create restores a baked Firecracker snapshot — there is no warm pool of idle VMs, because a warm pool is something you pay for while it does nothing. End to end that is p50 179 ms, p99 203 ms: inside it, roughly 25 ms to fork and exec `firecracker`, 4 ms to reflink the rootfs, about 80 ms for the snapshot load, 6 ms to resume, and a 40 ms TCP probe to confirm the guest is answering. The first-ever spawn of a template with no snapshot baked cold-boots in about 3 s and the host agent then bakes it, so you pay that once per template per host rather than once per create.

The structural difference is what the one-time work amortises over. A prebuild amortises over a person's session: one commit's setup spread across however long a developer stays in that workspace, which is hours. A bake amortises over every create that template will ever serve on that host, which is thousands. Both are legitimate engineering, optimal for different ratios of setup-work to environment-count — and your workload picks the ratio, not your taste.

Fan-out, and the attribution people get wrong

Once creates are cheap, a different pattern opens up: take one prepared sandbox — repo cloned, dependencies installed, services up — and multiply it. On PandaStack `fork_tree(count)` is the memory-inheriting path: the children restore from the parent's snapshot, so they come up with the parent's process state, warmed caches and loaded interpreter already in place. That is 400-750 ms same-host, or 1.2-3.5 s cross-host, where the cross-host number includes pulling the memory image from object storage.

A bare `fork()` is a different operation and I want the attribution right, because people pair the fast number with the wrong call. `fork()` reflinks the rootfs — a copy-on-write disk clone — and then cold-boots a fresh guest. It does not inherit the parent's memory, and the child draws its own entropy. Its shape is the roughly 3 s cold-boot shape, not the 400-750 ms restore shape. So: 400-750 ms is `fork_tree` / restore; a bare `fork()` boots. Eight variants of a half-finished agent state is `fork_tree`; eight independent machines starting from the same disk is `fork()`.

Concurrency and pricing shape

Pricing is not a footnote here; it is the clearest expression of what each product thinks it is. Seat and workspace-hour thinking is correct for workspaces, because a workspace maps to a person and a person maps to a seat, so environment count is bounded by headcount. Capacity planning becomes a hiring question.

PandaStack meters compute per sandbox: $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same rate for every class, with CPU billed on CPU-seconds actually burned and memory on committed GiB-hours. Egress is metered. There is no per-request pricing, and the rate card is in hourly units — a 179 ms sandbox that burns almost no CPU costs almost nothing, but it accrues against hourly rates rather than being priced as an event. The practical effect: a run that creates four hundred sandboxes each living ninety seconds and burning ten CPU-seconds is a line item you squint at, not a budget conversation.

Baked guest sizes matter here. Firecracker cannot change vCPU or RAM at snapshot restore, so a template's size is fixed at bake time: `base` is 4 GiB / 8 vCPU, `agent` and `code-interpreter` 2 GiB / 8 vCPU, `browser` 4 GiB / 8 vCPU, `postgres-16` 1 GiB / 8 vCPU. The 8 vCPU are burst capacity, shared fairly under contention via `cpu.weight` in a per-VM cgroup, which is why billing CPU on actual CPU-seconds is the only honest way to bill burst. RAM is the one per-template size knob, so a 12 GiB environment is a template decision rather than a request parameter — a real limitation next to a platform where a human picks a size from a dropdown.

So the decisive commercial question is a counting question: how many environments exist at the same time? If the answer is "roughly one per developer, during working hours", seat pricing is not merely acceptable, it is simpler and you should take it. If the answer is "four hundred, for ninety seconds, triggered by a queue at 2am with nobody watching", you need metered compute and an API, and no amount of goodwill will make a seat model fit.

The agent angle, fairly

Ona's repositioning toward agents is sensible and the reasoning behind it is strong. A workspace is a genuinely defensible place to put an agent, because an agent needs almost exactly what a developer needs: the repo at a known commit, the environment definition already satisfied, build tooling, credentials to push a branch, and a terminal. A workspace platform has spent years getting all of that right. Dropping an agent into it is reuse of the correct kind, and the agent works in the same environment your humans review in, which kills a whole category of "works for the agent, not for me" bugs.

My counter-argument is not that this is wrong. It is that an agent does not need an editor, and it does need to be unable to matter when it misbehaves — two facts that pull against everything a workspace is optimised for. An agent will never appreciate your port-forwarding UX, and it will eventually do something stupid at a speed no human would. In a sandbox, a model-generated `rm -rf /` is a 179 ms replacement and a log line. In a long-lived workspace holding a day of accumulated state, it is a bad afternoon; on a shared host kernel with neighbours, it is a question about whether it stayed in its lane, asked under time pressure. The design goal is not "prevent the agent from being dumb" — you cannot, that is what sampling from a distribution means. It is "make the agent's worst action cheap to absorb".

The second difference is credential blast radius, and it is the most underrated point in this comparison. A workspace holding a developer's Git credentials, cloud tokens, registry auth and SSH agent is a different object from a sandbox created with nothing in it. For a human that credential set is the point. For an agent it is an enormous amount of standing authority handed to a process whose behaviour you can shape but not bound. A PandaStack sandbox starts with no credentials; whatever the task needs, the orchestrator injects for that task, and the blast radius of a compromise or a confused tool call is whatever you put in that one machine. That is a design you can reason about on a whiteboard, which is not something I can say about "the agent has what the developer had".

One honest note, because egress deserves it: PandaStack sandboxes have open egress by default. There is no default-deny. There are a few targeted DROP rules — crypto-mining Stratum traffic among them — but if your threat model needs an allowlist, you add it. Better you read that here than discover it after a comparison table.

Environment definition: declarative file vs baked template

A declarative, repo-checked-in environment definition is a genuinely good idea and I am not going to argue with it. It lives next to the code, it is reviewed in the same pull request, it diffs, and it makes the answer to "how do I run this?" a file rather than a Slack search. Here is a sketch of the shape — illustrative of the pattern, not quoted from anyone's documentation, written as JSON purely to make the structure legible:

{
  "image": "ghcr.io/acme/dev-base:2026-09",
  "tasks": [
    { "name": "deps", "init": "pnpm install --frozen-lockfile" },
    { "name": "db", "init": "docker compose up -d postgres", "command": "pnpm db:migrate" },
    { "name": "web", "command": "pnpm dev" }
  ],
  "ports": [
    { "port": 3000, "visibility": "private", "onOpen": "notify" },
    { "port": 5432, "visibility": "private" }
  ],
  "prebuild": {
    "branches": ["main", "release/*"],
    "runTasksOn": ["deps", "db"]
  },
  "editor": {
    "extensions": ["dbaeumer.vscode-eslint", "golang.go"]
  }
}

Read what that file is expressing. An `init` phase distinct from a `command` phase, because the slow deterministic part should be prebuilt and the long-running part should start when a human arrives. Ports with visibility, because a person will click a link. Editor extensions, because there is an editor. It is a description of a place to work, and every field is load-bearing for that.

PandaStack's equivalent is a different trade: a baked template snapshot plus runtime detection at deploy time via mise. The `base` template is Ubuntu 24.04 with mise and pre-warmed Node 24 LTS, Python 3.12, Go and Bun, baked at 4 GiB / 8 vCPU. A repo declares its runtime through the idiomatic files it probably already has — `.nvmrc`, `.python-version`, `.tool-versions`, `mise.toml` — and the deploy pipeline runs `mise install` on demand for whatever the bake did not cover. Beyond `base` the catalogue is `code-interpreter`, `agent`, `browser` and `postgres-16`, and you can bake your own.

That is less per-repo expressiveness in exchange for a much faster create and a simpler mental model. No `init` versus `command` distinction, because nothing arrives later; whatever the task needs, the caller does, in the call. No port visibility semantics, because a preview URL is tokenless at `https://<port>-<sandbox-id>.<suffix>` for the sandbox's lifetime, with the UUID as the credential. No editor extensions, because no editor. If your environment genuinely needs a bespoke eighteen-step setup per repository, the declarative file is the better tool and you will feel the difference immediately.

What the sandbox side actually looks like

Two blocks, because the shape of the code is the argument. The first is the per-task unit: create, do one thing, read the result, delete. The TTL and the `finally` are both load-bearing.

from pandastack import Sandbox, CommandFailed

# One task, one machine. ttl_seconds is the backstop: if this process dies
# holding the only reference, the platform reaps the sandbox anyway.
sbx = Sandbox.create(
    template="base",
    ttl_seconds=900,
    metadata={"job": "pr-4817-build", "repo": "acme/web"},
)

try:
    sbx.exec("git clone --depth 1 -b pr-4817 https://github.com/acme/web /work",
             timeout_seconds=120, check=True)

    # Stream the slow part so the caller sees progress instead of a timeout.
    code = sbx.exec_stream("cd /work && pnpm install --frozen-lockfile",
                           on_stdout=print)
    if code != 0:
        raise SystemExit(f"install failed: {code}")

    try:
        build = sbx.exec("cd /work && pnpm build", timeout_seconds=600, check=True)
    except CommandFailed as e:
        # A failed build is data, not an incident. Keep the logs, drop the VM.
        print("build failed\n", e.stderr)
        raise

    artifact = sbx.filesystem.read("/work/dist/bundle.js")
    print(f"built {len(artifact)} bytes in {build.stdout.splitlines()[-1]}")

finally:
    # There is no sbx.delete(). Teardown is kill(), and it belongs here --
    # a leaked sandbox bills by the hour and tells nobody.
    sbx.kill()

The second is the fan-out a workspace model cannot express, because it would mean four hundred workspaces. Prepare once, inherit memory, run the variants in parallel.

from concurrent.futures import ThreadPoolExecutor
from pandastack import Sandbox

# Pay the expensive setup exactly once: clone, install, warm the toolchain.
parent = Sandbox.create(template="base", ttl_seconds=3600,
                        metadata={"role": "fanout-parent"})
children = []
try:
    parent.exec("git clone --depth 1 https://github.com/acme/web /work",
                timeout_seconds=120, check=True)
    parent.exec("cd /work && pnpm install --frozen-lockfile",
                timeout_seconds=900, check=True)
    parent.exec("cd /work && pnpm tsc --noEmit", timeout_seconds=600)

    # fork_tree inherits the parent's MEMORY: warm caches, loaded interpreter,
    # resident node_modules. 400-750 ms same-host, 1.2-3.5 s cross-host.
    # (A bare parent.fork() is a disk clone plus a cold boot -- ~3 s, own entropy.)
    children = parent.fork_tree(8, metadata={"role": "candidate"})

    def attempt(pair):
        child, patch = pair
        try:
            child.filesystem.write("/work/candidate.patch", patch.encode())
            child.exec("cd /work && git apply candidate.patch",
                       timeout_seconds=60, check=True)
            r = child.exec("cd /work && pnpm test --silent", timeout_seconds=600)
            return child.id, r.exit_code, r.stdout[-2000:]
        except Exception as e:
            return child.id, 1, str(e)

    patches = model_generated_patches()  # 8 candidate diffs, from your agent
    with ThreadPoolExecutor(max_workers=8) as pool:
        results = list(pool.map(attempt, zip(children, patches)))

    winners = [r for r in results if r[1] == 0]
    print(f"{len(winners)}/8 candidates green")

finally:
    for c in children:
        c.kill()
    parent.kill()

Eight is the demo number; the real one is whatever your queue produces. The ceiling is host RAM and CPU, not the API.

What each one locks you into

Both sides have lock-in and anyone who tells you otherwise is selling.

A workspace platform locks in the environment definition and the habits around it. The definition file is specific to the platform that reads it, the prebuild semantics are specific, and over a year or two your repositories accumulate assumptions about what the environment provides — a tool that is just there, a service already running, a port already forwarded. Migrating is archaeology across every repo. The second-order lock-in is stronger than the file format: once developers can no longer run the stack locally, the platform is load-bearing for shipping anything. The mitigation is real, though — bring-your-own-infrastructure options mean source and data need not leave your accounts, which is the lock-in that ends up in security reviews.

PandaStack locks you into an API shape and a baked-template workflow. Your orchestrator calls `Sandbox.create`, `exec`, `fork_tree`, `kill`; port that to another sandbox vendor and the verbs mostly map, but the interesting parts do not. Memory-inheriting fork is not a commodity feature, so anything built around branching live guest state is the hardest thing to move. And the ops lock-in on self-hosting is honest: it is open source and you can run the control plane and per-host agents yourself, but "yourself" means Postgres, ClickHouse, KVM hosts and a Go agent per host. That is a platform team's worth of work, not a Helm chart and an afternoon.

The asymmetry worth noticing is where the lock-in lands. Theirs lands in your repositories, which is where change is slow and political. Mine lands in your orchestrator, which is usually one service owned by one team. Neither is free; one is easier to rip out on a Tuesday.

Pick this one if

Here is the split I would give a friend on a call.

Pick Gitpod / Ona if the environments are for people, or for agents working the way people work. Specifically: onboarding is your pain; laptop drift costs you real days; you want the environment defined in the repo and reviewed with the code; prebuilds would remove the worst part of your current loop; your engineers want a browser IDE or an editor attached to a remote machine; your agents do developer-shaped work on a branch and open a pull request a human reviews; and environment count is bounded by headcount, not by a queue. If most of that is you, buy their thing, and do not let an isolation argument about a threat model you do not have talk you out of it.

Pick PandaStack if the environments are for machines and the code inside them is not fully trusted. Specifically: something other than a person creates the environment; the code was generated by a model or uploaded by a customer; you need tens or hundreds concurrent, for seconds to minutes; you want the environment gone the instant the task ends, with no idle fleet accruing cost; you want to branch live state and run variants in parallel; a hardware isolation boundary is something you must explain to a customer or an auditor; and you are at peace with there being no IDE, ever.

Pick both — the most common correct answer for a team shipping an agent product — if your humans develop in workspaces and your agents execute in sandboxes. Different jobs, different budgets, different failure modes; running both is not indecision. The pattern that works: the workspace is where a change is authored and reviewed, the sandbox is where untrusted or generated code runs, and the boundary between them is an artifact — a patch, a build output — not a shared machine.

Pick neither if what you are describing is continuous integration. If the trigger is a push, the unit is a repo at a commit, the code is your own team's, and the output is a pass or a fail plus some artifacts, you want a CI runner. GitHub Actions, GitLab CI, Buildkite, a self-hosted runner fleet — that category has a decade of ecosystem, caching someone else maintains, and a UI your whole company already reads. Both products in this post can run your tests. Neither should be what you page someone about when the test matrix breaks at 4pm on a Friday, because you would be rebuilding caching, matrices and required-check integration from scratch.

The summary

Gitpod — now Ona — was early and correct about disposable development environments, and the workspace is a well-designed object: declarative, prebuilt, editor-integrated, built for someone waiting at the other end. PandaStack is not a better workspace; it is a different noun. A sandbox has no occupant, no session and nothing worth keeping, which is what makes it reasonable to create one per task, inherit memory into eight of them, and delete all nine without a thought.

The substrate difference matters less often than vendors imply and more than teams expect. The pricing difference is the same fact seen from accounting: seats are correct when environment count tracks headcount, metered vCPU-hours and GiB-hours when it tracks a queue. And on agents, the workspace is a defensible home — it just hands a process an entire developer's standing authority, where a sandbox starts empty and a model-generated `rm -rf /` is a 179 ms replacement instead of an incident report.

Frequently asked questions

Is Gitpod the same thing as Ona now, and does my existing setup still work?

Gitpod is the product most engineers know by that name, and the company has moved to the Ona name with a positioning aimed much more at software engineering agents than at humans opening a browser IDE. I am deliberately not going to tell you what that means for your specific plan, repository configuration or self-hosted deployment, because product boundaries and plan structure in this category change faster than blog posts get updated, and a confidently wrong answer is worse than no answer. Read their current documentation, and if you have a contract, ask your account contact about migration and continuity. What I will say is that the shape of the underlying good idea has not changed: a declarative environment definition in the repository, prebuilds that front-load the slow part of setup, and deep editor and Git forge integration. Depend on those and you are depending on the durable part of the product rather than on a name.

Can I just run my AI agents in a cloud development workspace instead of a sandbox platform?

Yes, and for a lot of teams that is the right first move, because a workspace already has everything an agent needs: the repo at a known commit, a satisfied environment definition, build tooling, credentials to push a branch, and a terminal. Reusing that is sensible rather than lazy, and the agent works in the same environment your humans review in, which kills a whole class of "it worked for the agent" bugs. Two things eventually push teams off it. The first is concurrency: a workspace model is sized around roughly one environment per developer per task, and an agent run wanting four hundred concurrent environments for ninety seconds each makes no sense in a seat-priced product. The second is blast radius: a workspace holds a developer's Git credentials, cloud tokens and registry auth, which is exactly the standing authority you do not want to hand to a process whose behaviour you can shape but not bound. If neither is biting yet, do not pre-solve them.

How do prebuilds compare with PandaStack's snapshot restore? Aren't they the same idea?

They are the same instinct applied to different costs, and the distinction decides which one fits your workload. A prebuild attacks setup cost: it runs the dependency install, the codegen and the container pull ahead of time, per commit, so that when a human opens an environment the slow deterministic work has already happened. A snapshot restore attacks machine-creation cost: on PandaStack every create restores a baked Firecracker snapshot rather than booting, which is p50 179 ms and p99 203 ms end to end, with no warm pool of idle VMs to pay for. The first-ever spawn of a template cold-boots in about 3 s and the host agent then bakes the snapshot, so that cost is paid once per template per host. The structural difference is amortisation: a prebuild amortises over one person's session, which is hours, while a bake amortises over every create that template will ever serve on that host, which is thousands. Setup-heavy with few environments favours prebuilds; many short environments favours restore.

Does hardware-level isolation actually matter for a development environment?

Usually not, and I would rather say that plainly than sell you a boundary you do not need. For code written and reviewed by people you employ, in an environment they opened themselves, a container on a shared kernel is completely appropriate — the kernel is not your threat model, and nobody is escalating out of their own dev environment to read their own repository. Buying a hypervisor boundary for an internal workspace fleet is paying for defence against an adversary you do not have. It becomes load-bearing at one moment: when the code being executed was not written by anyone you can call, because a model generated it or a customer uploaded it. Containers share a kernel by design, so the isolation is restrictions imposed by that same kernel on processes still talking to it across a very large syscall surface. A Firecracker microVM gives each sandbox its own guest kernel and a handful of virtio devices. Both are real boundaries; one is much smaller, and the operational payoff is being able to tell a capacity page from a security page at 3am.

What does PandaStack simply not do that a workspace platform does?

No IDE, in the browser or otherwise, and no plan to add one. No workspace lifecycle: nothing is opened or closed, there is no session, and there is no list of "my environments" to browse. No dotfiles or personalisation story, because nothing with preferences lives inside a sandbox. No prebuild system tied to commits and branches; the equivalent is a baked template snapshot plus mise-based runtime detection, which is less per-repo expressive in exchange for a much faster create. No per-request machine sizing either — Firecracker cannot change vCPU or RAM at snapshot restore, so a template's size is fixed at bake time, with `base` at 4 GiB / 8 vCPU and `postgres-16` at 1 GiB / 8 vCPU. And egress is open by default, with a few targeted DROP rules rather than a default-deny network, so an allowlist is yours to add. If the first three items are what you were shopping for, buy a workspace platform.

Keep reading

Related posts

  • The best CodeSandbox alternatives in 2026

    Half the people looking for a CodeSandbox alternative want an embeddable editor for their docs. The other half want an API that gives an agent a machine. Those are different products, and mixing them up is why the shortlists never help.

  • Cloud Dev Environments on microVMs

    Ephemeral dev environments want three things at once — real isolation, fast start, and the ability to freeze and resume. A microVM gives you all three, which is why it's a better substrate than a shared-kernel container.

  • Top 6 Ephemeral Dev Environment Platforms for AI Agents

    Every roundup of ephemeral development environments grades them for a human opening an IDE. This one grades them for an autonomous agent with no hands, no browser and no patience — six platforms, one rubric, including where mine loses.

  • The Best GitHub Codespaces Alternatives in 2026

    Cloud dev environments split into two jobs now: a place for humans to write code, and a place for agents to run it. Codespaces is good at the first and awkward at the second.

  • Top 11 Ephemeral Development Environment Platforms (2026)

    Every roundup in this category sorts by substrate or by price list. Both are downstream of a question nobody asks on the evaluation call: who, or what, is going to call create? Eleven platforms sorted by their creator — human, CI job, pull-request webhook, agent API — because the creator decides how many environments exist at once and who pays for the idle ones.

More in AI agent sandboxes · See AI agent sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.