PandaStack vs Namespace Cloud: Two Kinds of "Ephemeral"
Before anything else: there is a good chance you and I are not talking about the same thing, because the word in the title means three different things to three kinds of reader, and I would rather sort that out in the first screen than waste your click.
Meaning one is the Linux kernel's namespaces — the features (PID, mount, network, UTS, IPC, user, cgroup, time) that a container is assembled out of. Meaning two is a Kubernetes namespace, which is a label on API objects and isolates approximately nothing on its own. Meaning three is Namespace Cloud, the company at namespace.so, which sells fast managed CI runners and cached ephemeral compute for development and testing.
This post is about meaning three, compared against PandaStack, the microVM sandbox API I build. But I am spending the first section on all three, because the confusion is real and the second meaning in particular has cost people money. If you came here for the kernel mechanism rather than the vendor, the post you want is How Network Namespaces Isolate Each Firecracker MicroVM.
And the decision up front, because most traffic to a page titled "X vs Y" wants an answer rather than an essay. If your problem is "our CI is slow and our runners are a mess", a managed-runner-and-cache product is the right category and you should evaluate Namespace against Blacksmith and Depot — not against a sandbox API. Those three compete for the same job; I do not. You can leave now with my blessing, and The Best Ephemeral CI Runner Platforms in 2026 is a better next page than the rest of this one.
Three things called "namespace", and only one is a company
The kernel's namespaces: the mechanism, not the boundary
A Linux namespace is a scoped view of exactly one kernel subsystem. A PID namespace means a process sees a process table that starts at 1 and holds only its own cohort. A mount namespace means its own filesystem tree. A network namespace means its own interfaces, routing table, firewall rules and socket space. A user namespace means its root is not the host's root — usually.
The thing to hold onto is what has not changed: there is still one kernel. Every process in every namespace makes syscalls into the same code. The namespace changes what a process can name and reach; it does not change what it is talking to. That is why a container is best described as a polite suggestion to the kernel about what a process ought to be allowed to see. Suggestions are usually honoured. A local privilege escalation in a syscall interface is the kernel declining one.
PandaStack uses kernel namespaces heavily, and this is the counterintuitive part: we use them around the virtual machine, not instead of one. Each sandbox gets its own network namespace with a veth pair and a tap device, from 16,384 pre-allocated /30 subnets per host agent inside 10.200.0.0/16. That isolates one microVM's host-side plumbing from its neighbours'. Isolating the guest's code is the hypervisor's job.
Kubernetes namespaces: an authorization scope wearing a boundary's clothes
A Kubernetes namespace shares a word with the kernel feature and almost nothing else. It is a grouping for API objects: a string in metadata that scopes names so two teams can both own a Service called "api", gives RBAC something to bind roles against, gives ResourceQuota something to attach to, and gives cluster DNS a segment.
What it does not do, by itself, is isolate. Pods in two namespaces land on the same nodes and share those nodes' kernels. Traffic between them flows freely unless somebody wrote a NetworkPolicy and the CNI in use actually enforces one. Node-local paths, the container runtime's socket, the instance metadata endpoint: none of them grow a wall because an object got a different metadata string. The actual boundary is whatever your container runtime, your network policy and your node pools provide — three separate things you configure, each of which can be absent while the namespace sits there looking organisational and reassuring.
I have watched a team discover this live. They had a namespace per customer, a diagram with lovely boxes, and a sincere belief that the boxes meant something. What they had was consistent labelling. The namespace was not wrong — it was doing filing, and had been promoted to security by a naming coincidence.
Namespace Cloud: the company this post is about
Namespace Cloud is a developer-infrastructure vendor. From their public documentation at the time of writing, the product centres on managed CI runners you point existing workflows at, ephemeral compute and environments for development and testing, caching that persists across runs, and acceleration for container image builds. If you have read my Blacksmith post, that shape will be familiar: this is the managed-runner-and-cache category, and the category bet is that a fast, cached remote machine is worth more to you than tuning your own runners ever will be.
That bet is a good one. Most CI pipelines are slow for two unglamorous reasons — cold dependency installs and cold image-layer caches on a shared machine with unpredictable storage — and fixing both without changing how you express your builds is an excellent trade. Where Namespace's specifics differ from its siblings (how its caching integrates with the stock cache actions, what exactly an ephemeral environment is in their vocabulary, what their build acceleration covers), I will tell you to check their docs rather than characterise it for you. I know the category well. I do not know their current feature matrix well enough to be quoted on it.
What PandaStack is, so the comparison has two sides
I'm Ajay. PandaStack is an open-source Firecracker microVM sandbox platform: a Go control plane with Postgres and ClickHouse, a Go agent on each Linux KVM host, guests running Ubuntu 24.04 on a 5.10 kernel under Firecracker v1.16.
The entire interface is a function call. `Sandbox.create()` from Python or TypeScript, or a POST to /v1/sandboxes with a bearer token, and a microVM exists: its own kernel, its own network namespace, a guest agent for exec and filesystem operations. Then your code runs commands in it, streams their output, writes files in, reads files out, snapshots it, forks it, hibernates it, and deletes it. There is no workflow file, no job, no step, no matrix, no `runs-on`. Nothing above you decides when any of that happens. Your program is the scheduler.
The create path is the one engineering claim I would defend in a room. There is no warm pool of idle VMs — not an optimisation we have not got to, but a deliberate refusal, because a warm pool is a bill for machines nobody is using. Every create restores a baked Firecracker snapshot: allocate a pre-built network namespace slot, patch the tap MAC to match the baked guest identity, reflink the rootfs, fork and exec Firecracker, POST /snapshot/load, resume, probe TCP port 22. That is p50 179 ms and p99 203 ms end to end, with the snapshot load itself 49 to 80 ms. The only slow create is the first-ever spawn of a template, which cold-boots in about 3 seconds and bakes the snapshot so nobody after it pays that again. One constraint falls out of that design and matters later: guest vCPU and RAM are baked into the snapshot, because Firecracker cannot change either at restore, so per-request cpu and memory_mb are silently corrected to the baked values.
The boundary, drawn precisely
The actual difference is not speed or hardware. A CI-and-ephemeral-environment product accepts the CI-shaped interface wholesale. Something above you owns the job graph: it parses the workflow, expands the matrix, evaluates conditions, dispatches each unit of work, streams logs into a UI, retries flakes, and decides which secrets a given job may see. The vendor supplies the machine the job lands on, and makes it fast. Your tooling and your branch protection rules keep working, because none of them are being replaced.
Which is why changing a runner label is the shortest infrastructure migration in the industry: one line of YAML, reviewed in thirty seconds, reverted with a `git revert` if the week of measurement disappoints. I cannot compete with that reversibility and I do not try. Any comparison post that waves it away is selling something.
A sandbox API has no job graph at all. There is nothing to put a label in, because there is no file a label would go into. That absence is simultaneously the feature and the bill. The feature: your program decides a machine should exist, in response to an HTTP request, a queue message, or a model concluding it needs a shell — none of which a CI platform can represent. The bill: you own log streaming and retention, retries, concurrency limits, secret injection and its audit trail, artifact handling, and reporting the result back to whatever cares.
The asymmetry is instructive. The CI shape makes easy anything triggered by a git event, expressed declaratively, with a shared log UI and an on-call engineer who knows where to look. It makes impossible a machine that appears inside a user-facing request, a job that pauses at minute four and resumes tomorrow from exactly that state, and eight copies of one prepared environment. The sandbox shape makes all three easy, and makes impossible the thing that matters most on a Monday: handing your team a product they already know how to use, with a one-line diff.
Two code blocks make this more vivid than another paragraph will. First the runner shape, as the category documents it.
# .github/workflows/test.yml
#
# The whole managed-runner migration, and the reason it wins most arguments
# with a sandbox API: it is one line. GitHub still owns the control plane --
# it parses this file, expands the matrix, evaluates `if:`, dispatches jobs,
# streams logs into its own UI and gates which secrets a job can read. The
# vendor supplies the machine the job lands on, with warm caches on it.
name: test
on: [push, pull_request]
jobs:
test:
# BEFORE -- GitHub's own hosted runner:
# runs-on: ubuntu-latest
#
# AFTER -- the vendor's runner label. DO NOT COPY THE PLACEHOLDER BELOW.
# Runner label naming is the single most likely thing to have changed
# since this post was written; get the exact string from the vendor's
# current docs, not from a blog.
runs-on: <vendor-runner-label>
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
# Caching is where runner vendors differentiate and where you should
# read most carefully. Some ship drop-in replacements for the stock
# cache actions backed by storage local to the machine, which turns
# a cache hit from "download a tarball across the network" into
# something nearer "read a local disk". Whether that is what YOUR
# vendor does, and whether it covers image layers as well as
# package directories, is a docs question. Check the docs.
cache: npm
- run: npm ci
- run: npm test
# Rollback: one revert commit. Approval for this change is a Slack message.
# That is not a small advantage, it is most of the advantage.
Now the sandbox shape for the same intent — run an untrusted pull request's tests on a clean machine. It is longer, because everything the CI service was doing for you is now yours to do.
import os
import sys
from pandastack import Sandbox
# No runs-on. No job. Nothing dispatched this -- my program decided a machine
# should exist, so one does. Auth is a bearer token in PANDASTACK_API_KEY.
assert os.environ["PANDASTACK_API_KEY"], "set PANDASTACK_API_KEY"
PR_DIFF = open("/tmp/pr-4412.patch").read() # a stranger's code. act accordingly.
# create() is a snapshot restore, not a boot and not a pool checkout -- there is
# no pool. p50 179 ms, p99 203 ms, of which /snapshot/load is 49-80 ms. The
# first-ever spawn of a template cold-boots in ~3 s and bakes the snapshot so
# nobody after it pays that again.
sbx = Sandbox.create(
template="base", # Ubuntu 24.04 + mise: Node 24, Python 3.12, Go, Bun pre-warmed
ttl_seconds=1800, # enforced by a platform-side reaper, NOT by this process.
# a crashed orchestrator cannot leak a running VM at your cost.
metadata={"pr": "4412", "repo": "acme/api", "trust": "fork"},
# cpu= and memory_mb= are accepted and then corrected to the template's
# baked values. Firecracker cannot resize guest RAM or vCPU at snapshot
# restore, so sizing is a template-build decision, not a create argument.
)
try:
sbx.exec("git clone --depth 1 https://github.com/acme/api /work", check=True)
# Write the untrusted diff in as DATA rather than interpolating it into a
# shell command. One of those two habits survives a filename with a quote
# in it, and it is this one.
sbx.filesystem.write("/work/pr.patch", PR_DIFF)
sbx.exec("cd /work && git apply --3way pr.patch", check=True)
# exec_stream pushes stdout/stderr to callbacks as they arrive, which is how
# you get a live log without a log UI. You own retention, redaction and
# whatever your reviewers expect to read six weeks from now.
#
# `timeout 900` is NOT decoration. timeout_seconds on exec is a CLIENT
# deadline -- the SDK stops waiting, the command keeps running, the VM keeps
# billing. The hard wall has to live inside the guest.
sbx.exec_stream(
"timeout 900 sh -c 'cd /work && npm ci && npx jest --ci "
"--reporters=default --reporters=jest-junit'",
on_stdout=lambda line: print(line, end=""),
on_stderr=lambda line: print(line, end="", file=sys.stderr),
)
# Pull the machine-readable result out as bytes and report it yourself. On a
# CI platform this is "upload-artifact" plus a built-in test summary; here it
# is a filesystem read and some code you write.
junit = sbx.filesystem.read("/work/junit.xml")
with open("junit-4412.xml", "wb") as f:
f.write(junit)
print(f"wrote {len(junit)} bytes of junit xml")
finally:
# The only lifecycle discipline that survives contact with production. The
# TTL is the backstop, not the plan: delete on the success path AND on the
# exception path, every time.
sbx.kill()
Count the comments in the second block that are really "here is something the CI platform used to do for you". That list is the migration cost. If your workload is CI, paying it buys you nothing you did not already have — which is the whole reason the decision at the top of this post is what it is.
Overlap one: untrusted code, where the choice is real
Two places where these categories genuinely collide, and I want to be fair about both rather than manufacture a tie. The first is untrusted code. A fork pull request runs a stranger's code on a machine you pay for; model-generated commands are the same problem in newer clothes. Both need an isolation boundary, and "own kernel per job" versus "a container on a shared kernel" is a decision rather than a preference.
Be precise about the threat model first, because in CI the runner vendor is usually not the weak link. The two things that actually get people compromised by fork pull requests are `pull_request_target`, which runs workflow code in a privileged context with your secrets while checking out fork-authored code, and long-lived self-hosted runners that keep state between jobs so one job can poison the next. Both are configuration in your repository, not properties of who supplies the machine. Fix those before you shop: a vendor documenting per-job machine isolation on its own hardware is a reasonable posture on top of a correctly configured repo, and I will not pretend otherwise. Read their current isolation docs rather than my summary of a category.
Where I would push further is when the requirement changes from "reasonably isolated" to "demonstrably isolated to a named boundary". That is the microVM-per-job argument, and it is about provability rather than any vendor being careless. A Firecracker guest has its own kernel: a guest-kernel escalation gets the attacker root in a machine about to be deleted, with the host behind a hypervisor with a deliberately small device surface. That is a sentence you can put in front of an auditor. "Namespaces and seccomp, correctly configured" is also defensible, but every clause in it can quietly stop being true.
One honesty note so nobody over-reads the boundary: PandaStack's guest egress is open by default, with a few targeted DROP rules for known-abuse protocols. It is not a default-deny network. If your untrusted-code story requires that the sandbox cannot reach the internet, that is egress policy you configure, not a property you inherit from the word "microVM".
Overlap two: state between runs, which is the interesting one
This is the most technically interesting axis in the comparison and it deserves its own section rather than a table row. Both categories exist because the dominant cost in modern compute is not running the work — it is arriving at a state where the work can start. They attack it from opposite ends.
A cache product makes the second run fast by restoring bytes into a fresh machine. You hand it a key, it hands back a directory — a package tree, a compiler cache, an image layer store — and the machine receiving those bytes was coming anyway, because a job was dispatched and something had to run it. The build then does whatever turns those files into a running state: resolve, link, compile, start a process, warm a JIT. The cache saved you the download and possibly the compile. It did not save you the startup.
Snapshot-restore makes the second run fast by restoring a machine that was already warm. Guest RAM comes back mapped MAP_PRIVATE from the snapshot file and the VM resumes with its processes running, its page cache populated, its sockets bound. Nothing is re-done because nothing was undone. That is how a create can be 179 ms at p50 while doing strictly more than a cache hit: it is resuming the state, not rebuilding it.
A cache is a bet that you can rebuild the state faster than you can keep it. A snapshot is a bet that you can keep it cheaply enough never to rebuild it.
Both bets are right in context, and the cache bet has two structural advantages snapshot people routinely undersell. The first is composition: a cache key is content-addressed, so an entry produced by one job is usable by a completely different job. The lockfile hash that filled the package cache during a unit-test run serves the lint run, the release build and a developer's laptop. A snapshot does not compose — it is a specific machine in a specific state, and a workload wanting a slightly different state wants a different snapshot.
The second is that a cache survives arbitrary machine shapes. Restore it onto a bigger machine, a smaller one, next year's image; the bytes do not care. A snapshot cares enormously, because guest vCPU and RAM are fixed in it. If a job needs more memory than the template was baked with, the answer is a different template, not a different create call — and the create call will accept your `memory_mb` and quietly ignore it, which is worse than erroring because it is silent. That is a real limitation and I am not burying it in a footnote.
What a snapshot captures that a cache structurally cannot is warmth. There is no cache key for "a JVM that has already JIT-compiled its hot paths", or "a Postgres with populated shared buffers", or "a dev server past its first compile and listening", or "a model resident in memory". Those states are not files. They are the product of a process having run for a while, and the only way to restore them is to restore the process, which means restoring its memory. A cache hands you the ingredients for all four. It cannot hand you the finished thing.
So the test is one question: is the expensive part obtaining files, or reaching a running state? If it is files, a cache is simpler, composes better, and the platform was going to give you a machine anyway — buy the cache product.
One mechanism a cache has no equivalent for: idle snapshots cost storage, not reserved capacity. Hibernate snapshots the sandbox and releases the host entirely, and wake demand-pages the memory image back from object storage. A sleeping environment is a few objects in a bucket.
Overlap three, sort of: branching, which the runner category does not do at all
I list this as an overlap out of politeness, because it is not one: nothing in the managed-runner-and-cache category offers it, and if you need it the comparison is over. You have an environment that was expensive to reach — repo cloned, dependencies installed, project built, database migrated and seeded — and you want to try six mutually exclusive fixes against that exact state and keep whichever works. Six dependency bumps; six patches a model generated for one failing test. On a cache-based platform you run that six times, and each run restores the ingredients and cooks again. The cache makes each cook cheaper; it does not make the second one unnecessary, because each run is a fresh machine re-reaching the prepared state from files.
PandaStack has verbs for this, and the distinction between them is the most commonly botched fact about the platform. `fork()` clones the sandbox's disk — reflink or dm-snapshot, O(metadata), data shared until written — and then cold-boots the clone. The child inherits the installed dependencies and built artifacts on disk. It does not inherit guest memory: it gets its own kernel boot, its own PIDs, its own entropy pool. Only `fork_tree(count)` branches from live memory, snapshotting the parent once and restoring N children from that snapshot so they inherit RAM as well as disk. Get the latencies the right way round too. The published 400 to 750 ms same-host and 1.2 to 3.5 s cross-host figures measure the snapshot-restore path, which is what `fork_tree` children ride; the cross-host case is slower because the snapshot has to travel first. A bare `fork()` pays a cold boot, so it is the roughly 3 s shape instead. Both calls cap at 16 children.
What that unlocks is a different shape of experiment rather than a faster version of the same one: prepare once, branch six ways, keep the winner, delete the rest. The variance that normally pollutes this kind of comparison — different machine, different cache hit rate, different ordering — is mostly gone, because all six children came from the same parent at the same instant.
I will not oversell it. It is one verb, not a product, and it has sharp edges. Children of a `fork_tree` call share the parent's entropy pool and any PRNG seeded before the snapshot was taken, so anything generating IDs or tokens in them will cheerfully agree with its siblings; mix in bytes from outside the clone. A plain `fork()` avoids that by cold-booting, at the cost of not being warm. And six copies of a prepared environment are six copies you are billed for while they live, at $0.054 per vCPU-hour and $0.0162 per GiB-hour — cheap per minute, not cheap if you forget the cleanup. But no runner label gives you this, and no cache key either.
What PandaStack straightforwardly loses on
I asked you to verify everything I said about Namespace, so here is the list I would want handed to me about my own product. These are the reasons the recommendation at the top of this post is what it is.
- No first-class CI integration story. A managed-runner vendor's answer to "how do I use this with GitHub Actions" is a runner label. Mine is "register the microVM as an ephemeral self-hosted runner and write the lifecycle" — registration, deregistration, the runner that died before taking work, the job cancelled mid-run, capacity decisions. Work somebody else has already finished.
- No managed cache product. There is no PandaStack equivalent of a drop-in cache action with content-addressed keys shared across an organisation. Snapshots are a different, non-composable mechanism, and not a substitute for a shared dependency cache across heterogeneous jobs.
- No workflow UI. No job graph, no matrix expansion, no log viewer, no required-checks integration, no per-job secret scoping. `exec_stream` hands you bytes; everything you would do with them is yours.
- No build acceleration for container images. Remote layer caching and distributed image builds are a whole product category I am not in. If Docker build time dominates your pipeline, that is a Namespace-shaped or Depot-shaped problem.
- Guest sizing is fixed at template-bake time, so a create's cpu and memory_mb are silently overridden. A bigger machine for one job means building a template, not passing an argument.
- Guest kernel 5.10, no nested virtualisation, no GPU in the guest. If your build needs KVM inside the job, no amount of snapshot speed fixes that.
- `timeout_seconds` on exec is a client deadline, not a server-enforced kill. A hung command keeps running and keeps billing; the hard wall belongs inside the guest with `timeout`, and `ttl_seconds` on every create is the backstop.
The decision table
| What you are actually asking | Namespace Cloud (managed runners + cached ephemeral compute) | PandaStack (microVM sandbox API) |
|---|---|---|
| Our CI is slow and the runners are a mess | This is the product. Point your workflows at their runners, keep your YAML, measure for a week. Evaluate against Blacksmith and Depot | Wrong category. I would have to rebuild a CI control plane first, and I have not |
| Our CI is slow and the profile says one test file is most of it | A faster machine with warm caches makes a slow suite somewhat less slow, and does not change the shape of the problem | Neither of us. Your CI is fine and the slow part is your test suite. Go read the profile and fix the test |
| An AI agent runs model-generated code hundreds of times an hour | No workflow file, no commit, no job to dispatch — nothing to put a runner label in | This is the product. `Sandbox.create()` in a request handler, p50 179 ms, delete when the turn ends |
| We must run untrusted fork-PR code and prove the boundary | Per-job machine isolation on their own hardware is a reasonable posture — read their isolation docs, and fix `pull_request_target` first | Own guest kernel per sandbox, own netns from 16,384 pre-allocated /30 subnets per host. Provable to an auditor; you still write the CI glue |
| The second run has to be fast | Persistent caching restores bytes into a fresh machine, composes across unrelated jobs, and survives any machine shape | Snapshot restore resumes a machine that was already warm — processes running, page cache populated. Tied to the baked machine; vCPU and RAM fixed in it |
| Try six fixes against one prepared state | Six runs, six cache restores, six rebuilds of the prepared state | Prepare once, then `fork_tree(6)` to branch live memory via snapshot restore -- 400-750 ms per child same host, 1.2-3.5 s cross-host, 16 max. `fork()` is a disk clone that cold-boots (~3 s) |
| What does an idle environment cost | Entirely a function of their model — a pricing-page question I will not guess at | Storage, not reserved capacity. Hibernate releases the host; wake demand-pages memory back from object storage |
| What does the migration cost | One line per workflow, reverted in a commit. The cheapest infrastructure rollback that exists | You write the orchestration: lifecycle, retries, logs, status reporting, capacity. Days, not minutes |
| Honest ceiling | It is a CI-and-ephemeral-environment product. If your caller is not a CI system, no label fixes that | No CI control plane, no managed cache, no image-build acceleration, guest kernel 5.10, no GPU, RAM fixed at bake time |
How I would decide, in one question
Ask what dispatches the work.
If the answer is "a push, a pull request, a merge queue, a cron in a workflow file" — anything a CI service already knows how to trigger — the work is CI and you want a runner-and-cache product. Namespace is a credible one, as are the two siblings I keep naming. Change the label, measure for a week, revert if it disappoints: nearly free experiments should be run rather than deliberated.
If the answer is "an HTTP request, a message on a queue, a model deciding it needs a shell, or a line in my own program", there is no workflow file for a label to live in. You want an API that creates machines, and what matters is how fast one appears, how little an idle one costs, whether you can branch a warm one, and where the isolation boundary sits.
If the answer is both, buy both. CI spend scales with how often your engineers push; sandbox spend scales with how much your customers use the feature that runs their code. Those grow for unrelated reasons, and conflating them produces a consolidation project that makes both worse.
And if the answer is "a Kubernetes namespace, which we believed was a security boundary" — that is not a vendor decision at all. That is node pools, a CNI that enforces policy, and an honest conversation about which workloads may share a kernel. Have it before you compare any two products on this page.
Frequently asked questions
Is a Kubernetes namespace a security boundary?
No. It is an authorization and naming scope, and the difference matters more than the vocabulary suggests. A namespace gives you four concrete things: names that only have to be unique within it, a scope for RBAC role bindings, an attachment point for ResourceQuota and LimitRange, and a segment in cluster DNS. Those are useful, and they are all about the API server. What a namespace does not do is isolate workloads from each other at runtime. Pods in two namespaces are scheduled onto the same nodes and execute against the same host kernel. Traffic between them flows freely unless somebody authored a NetworkPolicy and the CNI in use actually enforces one — plenty of clusters have policies that are silently decorative. Node-local resources are shared regardless of namespace: hostPath mounts, the container runtime socket if anything may mount it, the instance metadata endpoint, the node's kernel keyring. A container escape in namespace A lands you on a node that is also running namespace B. The real boundary is whatever your container runtime, your network policy and your node pools provide — three separate things you configure, each of which can be missing while the namespace sits there looking tidy on a diagram. For a hard boundary around untrusted code the honest options are a sandboxed runtime, dedicated node pools per tenant, or a hypervisor, and only the last gives each workload its own kernel.
Should I evaluate PandaStack against Namespace Cloud at all?
Only if you are unsure which category you are in, which is more common than it sounds. If you already know the work is CI — triggered by git events, expressed in workflow files, read by engineers in a log UI — do not put a sandbox API on the shortlist. Compare Namespace against Blacksmith and Depot, because those three solve the same problem and differ on things you can actually evaluate: how their caching integrates with your existing cache steps, what their image-build acceleration covers, how pricing behaves under your real concurrency, and how the migration feels on your messiest workflow. A sandbox API will lose that evaluation and should. The time to look at my category is when the work is not CI-shaped: a machine that has to appear inside a user-facing request, a job that needs to resume from exactly its mid-run state, or N copies of one prepared environment. Those are not slow CI; they are things a CI interface cannot express, and faster hardware underneath a job graph does not produce them. The common end state is both products in one company, owned by different teams, with nobody trying to consolidate them.
Can I run my CI on PandaStack instead of buying a managed runner?
You can, and for most teams you should not. The pattern works: create a microVM, register it with your CI provider as an ephemeral self-hosted runner, let it accept exactly one job, then delete the VM. Because every create is a snapshot restore at p50 179 ms rather than a boot, a VM-per-job runner provisions roughly as quickly as a container-per-job one, and it gives you a boundary a container cannot — separate kernel, separate network namespace, nothing inherited from the previous job because the previous job's machine no longer exists. What you are signing up for is everything around that: registration and deregistration including the ugly cases (a runner that registered then died before taking work, a job cancelled mid-run, a rate limit during a burst), log retention, artifact handling, and whatever your branch protection expects to see. None of it is research and all of it is work a managed-runner vendor has already finished and tested against thousands of repositories. Do this when you need a provable isolation property no runner vendor documents — regulated workloads, build secrets that must never share a kernel with third-party code, an auditor with opinions. Do not do it to make CI faster.
If PandaStack uses Linux namespaces too, why does it need a virtual machine?
Because the two do different jobs, and we use them for different things. Each PandaStack sandbox gets its own network namespace with a veth pair and a tap device, drawn from 16,384 pre-allocated /30 subnets per host agent — pre-allocated because building a namespace, a veth pair and the firewall rules from cold costs enough milliseconds to be visible in a sub-200 ms create path. That namespace isolates the host-side plumbing of one microVM from its neighbours': routing, rules, interface names, address space. It is infrastructure hygiene on our side of the boundary. What it is explicitly not doing is isolating the guest's code, because the guest's code is not in that namespace — it is inside a Firecracker VM, making syscalls into its own kernel. Namespaces partition the view that processes on a shared kernel have of that kernel's subsystems: the right tool when you trust the processes and want them tidy, the wrong strength of guarantee when the code was written by a stranger or generated by a model. A namespace is a polite suggestion to the kernel. A hypervisor is a different kernel. We use the suggestion for our own plumbing and the different kernel for your untrusted workload.
Keep reading
- Network namespace isolation in Firecracker, explained — If you came here for the kernel meaning of "namespace", this is the post you actually wanted.
- PandaStack vs Blacksmith: mostly, buy Blacksmith — The same different-categories argument against the other managed-runner vendor, with more depth on cache versus snapshot.
- PandaStack vs Depot: build acceleration is not a sandbox — The third vendor in the category Namespace competes in, and the one most focused on container image builds.
- The best ephemeral CI runner platforms in 2026 — If the decision at the top of this post sent you to the runner category, start your shortlist here.
- Fork-PR CI without getting pwned — Why pull_request_target is the actual hazard, and what a per-job kernel buys you on top of fixing it.
Related posts
- Top 6 Ephemeral Dev Environment Platforms for AI Agents
Every roundup of ephemeral development environments grades them for a human opening an IDE. This one grades them for an autonomous agent with no hands, no browser and no patience — six platforms, one rubric, including where mine loses.
- The Best Self-Hosted CI Runners in 2026
You've decided to run your own build infrastructure. Now you have to pick the runner software and, whether you meant to or not, an isolation model — because a CI runner is a machine that executes whatever is in a YAML file someone can open a pull request against.
- Running Nix Builds Inside a Disposable VM
Nix's sandbox is a hermeticity fence, not a hypervisor — a hostile derivation still runs on your kernel. Here's how to put Nix inside a disposable VM without a cold /nix/store making every build miserable.
- Tekton Steps, MicroVM Bodies: Isolating Untrusted Tasks
A TaskRun is a Pod. That is Tekton's best feature and its sharpest constraint — because everything inside a Pod is inside one trust boundary, including the step running a stranger's build script.
- The Best Replit Alternatives in 2026
"Replit alternative" means three unrelated things: a browser IDE, a cloud dev environment, or the sandbox API that runs untrusted code under the hood. Split the intent first and the shortlist collapses from ten options to two.
More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.