all posts

PandaStack vs Blacksmith: Mostly, Buy Blacksmith

Ajay Kumar··11 min read

Let me put the answer at the top, because most of the traffic reaching a page titled "X vs Y" wants a decision rather than an essay. If your problem is that your GitHub Actions jobs are slow and expensive, buy Blacksmith. A drop-in managed runner is the correct shape of tool for that problem and a sandbox API is not. You can close the tab; I will not be offended.

I'm Ajay. I build PandaStack, which runs code in Firecracker microVMs behind an API. Blacksmith is a managed GitHub Actions runner product: you point your workflows at their machines and your CI gets faster. The two get compared because both show up in searches about ephemeral compute and both talk about bare metal and warm caches, and that is roughly where the resemblance ends. One replaces a line in a YAML file. The other is infrastructure you program.

The post still earns ten minutes for two reasons. The boundary between "CI runner" and "sandbox API" is genuinely blurry at one edge — untrusted fork-PR code, matrix fan-out, anything where a job must not share a kernel with a stranger — and knowing where it blurs saves you from buying the wrong category. And both products are bets on the same underlying thesis, which is the most interesting thing about either of them: that the cost of producing a fresh, warm environment is the product.

Ground rules. Every number here about PandaStack is one I measured. Every statement about Blacksmith is qualitative, from their public documentation as it read at the time of writing, with no latency, price, machine spec or limit — I will not state another vendor's numbers as fact. CI companies ship weekly: verify anything I say about Blacksmith against their current docs and pricing before it influences a purchase order. Where the docs and I disagree, the docs are right and I am stale.

What each one actually is

Blacksmith: a line in your workflow file

Blacksmith is a managed replacement for GitHub's hosted Actions runners. As documented at the time of writing, the migration is to change the runner label in your workflow YAML; your jobs then execute on their own bare-metal machines rather than shared cloud instances, with persistent and accelerated caching aimed at the two things that dominate CI wall-clock time — dependency installs and container image builds. Verify the specifics, including how their caching integrates with the stock cache actions, against their docs.

Architecturally the important property is not the hardware. It is that Blacksmith accepts the CI-shaped interface wholesale. GitHub's Actions service still owns the control plane: it parses your workflow, resolves the job graph, evaluates `if:` conditions, expands your matrix, dispatches each job, streams logs into its own UI and decides which secrets a job may see. Blacksmith supplies the machine the job lands on. Everything your team already knows keeps working, because none of it is being replaced. That is an enormous amount of value for a one-line diff, and the reason for the recommendation at the top of this post: a product that makes an existing system faster beats one that asks you to rebuild it, for any problem the first one covers.

PandaStack: a line in your program

PandaStack has no `runs-on`, no jobs, no steps and no YAML. It has `Sandbox.create()`. You call it from Python, TypeScript or raw HTTP, and a Firecracker microVM exists — Ubuntu 24.04, guest kernel 5.10, its own kernel and network namespace, a guest agent for exec and filesystem operations. Then your code runs commands in it, reads files out of it, snapshots it, forks it, hibernates it, exposes a port on a preview URL, and eventually deletes it. Nothing above you decides when any of that happens. Your program is the scheduler.

The shape follows from who the caller is, and the caller is usually not a CI system. It is an AI agent executing model-generated code, a multi-tenant backend running a customer's script, or a product feature that needs a shell. Those callers have no workflow file to put a label in, and the work they need — "create a machine now, in response to this request, and destroy it in ninety seconds" — is not expressible as a job in a CI graph.

You can build a CI runner on top of it, and people do: register the microVM as an ephemeral self-hosted Actions runner, let it take exactly one job, destroy it. But notice what happened — you wrote the registration loop, the lifecycle, the failure handling and the capacity logic yourself. Blacksmith hands you a finished runner; PandaStack hands you the primitives a runner is built out of, which is the right trade only if you need those primitives for something a runner could never do.

The interface difference, in two code blocks

This is more convincing to look at than to read about. First, the entire Blacksmith-shaped change, as documented at the time of writing.

# .github/workflows/test.yml
#
# The whole managed-runner migration. GitHub still owns the control plane: it
# parses this file, expands the matrix, evaluates `if:`, dispatches each job,
# streams the logs and decides which secrets the job can see. The vendor
# supplies the machine the job lands on.
name: test
on: [push, pull_request]

jobs:
  test:
    # BEFORE -- GitHub's own hosted runner:
    #   runs-on: ubuntu-latest
    #
    # AFTER -- the vendor's label. Get the EXACT string from their current
    # docs; runner label naming is the single most likely thing to have moved
    # since this post was written, so do not copy the placeholder below.
    runs-on: <vendor-runner-label>

    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          # Caching is where runner vendors differentiate, and it is worth
          # reading carefully: some ship drop-in replacements for the stock
          # cache actions, backed by storage local to the machine rather than
          # a remote blob store. That changes the arithmetic of a cache hit
          # from "download a tarball over the network" to something closer to
          # "read a local disk". Verify what your vendor actually does.
          cache: npm
      - run: npm ci
      - run: npm test

  # Rollback is a one-line revert. That matters more than it sounds: the
  # reversibility of this change is a large part of why it is the right first
  # move for a slow CI pipeline.

And the PandaStack-shaped equivalent of "run my test suite on a clean machine". It is longer, because everything the Actions service was doing for you is now yours.

import os
from pandastack import Sandbox

# No runs-on, no job graph, nothing dispatching this. Your program decided a
# machine should exist, so one does. Auth is PANDASTACK_API_KEY.
assert os.environ["PANDASTACK_API_KEY"], "set PANDASTACK_API_KEY"

# create() does not boot a VM and does not pull one out of a warm pool -- there
# is no warm pool. It restores a baked Firecracker snapshot: allocate a
# pre-built network namespace slot, patch the tap MAC, reflink the rootfs,
# fork/exec firecracker, POST /snapshot/load, resume, probe TCP :22.
# p50 179 ms, p99 ~203 ms, of which the snapshot load itself is 49-80 ms.
# The only slow create is the first-ever spawn of a template, which cold-boots
# and bakes the snapshot (~3 s) so that nobody after it pays that cost.
with Sandbox.create(
    template="base",        # Ubuntu 24.04 + mise, with Node 22 / Python 3.12 / Go / Bun pre-warmed
    ttl_seconds=900,        # a platform-side reaper enforces this, so a crashed
                            # orchestrator cannot leak a running VM at your expense
    metadata={"job": "pr-4412", "repo": "acme/api"},
    # memory_mb is accepted here and then silently corrected to the template's
    # baked value. Firecracker cannot change guest RAM or vCPU count at snapshot
    # restore, so RAM is a template-build decision (`--memory-mb`), not a create
    # argument. `base` is baked at 4 GiB with 8 burstable vCPUs.
) as sbx:
    sbx.exec(
        "git clone --depth 1 https://github.com/acme/api /work",
        timeout_seconds=120,
        check=True,
    )
    # timeout_seconds is a CLIENT deadline -- the server does not enforce it.
    # If you need a hard wall on a runaway build, put it in the guest shell:
    #   timeout 600 sh -c 'cd /work && npm ci && npm test'
    r = sbx.exec(
        "timeout 600 sh -c 'cd /work && npm ci && npm test'",
        timeout_seconds=660,
    )
    print("exit", r.exit_code)
    print(r.stdout[-4000:])

    # Things you now own that Actions was doing for you, and should write down
    # before you commit to this path:
    #   - log streaming and retention (exec_stream gives you the bytes)
    #   - retries, concurrency limits, and fan-out ordering
    #   - surfacing a pass/fail back to the commit status API
    #   - secret injection, and the audit trail for it
    #   - artifact upload, and whatever your team's required checks expect

# The context manager calls kill(). The VM stops existing on the success path
# and the exception path, which is the only lifecycle discipline that survives
# contact with production.

Count the comments in the second block that are really "here is a thing the CI platform used to do for you." That list is the migration cost, and if your workload is CI, paying it buys you nothing you did not already have.

Where they genuinely overlap

Three places, and I want to be fair about all three — this is where most "X vs Y" posts start cheating.

Untrusted fork-PR code. A pull request from a fork runs a stranger's code on a machine you pay for. The honest version of this threat model is that the dangerous part is usually not the runner vendor: it is `pull_request_target`, which hands fork-authored code a privileged context with your secrets in it, and the self-hosted runner that quietly keeps state between jobs. A managed-runner product that puts each job on a dedicated virtual machine on its own hardware is a real isolation story, and I will not pretend otherwise — better than a shared container on a shared kernel, and much better than a long-lived self-hosted runner with a Docker socket mounted into it. Read their current isolation docs rather than my summary.

What differs is not whether there is a boundary but who controls the lifecycle around it, and whether you can do anything with a job's state other than throw it away. On a managed runner the lifecycle is the job: it starts when GitHub dispatches and ends when the job ends — exactly right for CI, and all you get. On a sandbox API the lifecycle is a set of verbs you call — create, snapshot, fork, hibernate, wake, kill — so a "job" can be paused mid-way, duplicated eight times, or restored tomorrow from the state it was in at minute four.

Matrix fan-out. Both products care about running many similar things at once, and both get there by making a fresh environment cheap: a runner vendor with lots of fast machines carrying warm caches, a sandbox API by restoring the same baked snapshot N times at p50 179 ms apiece. The difference shows up when the N things all want the same built state — a snapshot fork hands you that state directly, where a cache hands you the ingredients to rebuild it.

Anything that must not share a kernel. This is the overlap that drove both products to bare metal. A container is, in the end, a polite suggestion to the shared kernel about what a process should be allowed to see, and "polite suggestion" is the wrong strength of guarantee for code you did not write. Both answers here are VM-shaped. They part company on what the VM is for: a CI job, or whatever your program needs a machine to do.

The decision table

Blacksmith rows describe a managed GitHub Actions runner product qualitatively, from public documentation at the time of writing, with no numbers — verify against their current docs and pricing. PandaStack rows are measured.
DimensionBlacksmith (managed CI runners)PandaStack (microVM sandbox API)
InterfaceA runner label in workflow YAML. Your existing jobs, steps and marketplace actions are unchanged`Sandbox.create()` from Python, TypeScript or REST. No YAML exists anywhere in the product
Who schedules the workGitHub's Actions service: it parses the workflow, expands the matrix, dispatches jobs, streams logs, gates secretsYour program. Nothing dispatches for you, and nothing retries for you either
Unit of workA job, composed of stepsA sandbox, and the commands you choose to run in it
Isolation unitPer-job machine on the vendor's own hardware, as documented — read their current isolation page rather than this cellFirecracker microVM: own kernel, own network namespace, own veth/tap, from 16,384 pre-allocated /30 subnets per host agent
What "ephemeral" meansThe job ends, the runner goes away. The lifecycle is the job's lifecycleYou call kill(), or a platform-side TTL reaper does it for you. The lifecycle is whatever your code says it is
Warm-start mechanismFast machines plus persistent and accelerated caching of dependencies and image layers (qualitative — verify)Snapshot restore of a baked memory + disk image on every create: p50 179 ms, p99 ~203 ms; ~3 s once, for the first spawn of a template
Snapshot and forkThe analogue is cache restore, which is a genuinely different mechanism: it rebuilds state from files rather than resuming itfork() clones the disk then cold-boots; fork_tree(N) snapshots once and restores N children with the parent's memory too — 400-750 ms same host, 1.2-3.5 s across hosts, 16 children max
Can a non-CI workload use itThe workload has to be expressible as a GitHub Actions workflow. That is the productYes — AI agents running model-generated code, per-tenant jobs, preview environments, scheduled functions, managed Postgres branches
Migration costOne line per workflow, revertible in a commit. Rollback is a `git revert`You write the orchestration: lifecycle, retries, log handling, status reporting. Days, not minutes
Honest ceilingIt is a CI runner. If your problem is not CI, there is no label you can set that fixes thatNo GPU in the guest, no nested virtualisation, guest kernel 5.10, 8 burstable vCPUs, and RAM fixed at template-bake time

Why the cache row and the snapshot row are not the same row

A cache restores files. You hand it a key, it hands back a directory — `node_modules`, a Go build cache, a layer store — and your build then re-does whatever turns those files into a running state: resolve, link, compile, start a process, warm a JIT, fill a page cache. The cache saved you the download and possibly the compile. It did not save you the startup.

A snapshot restores a machine. Guest RAM is mapped `MAP_PRIVATE` from the snapshot file, so the kernel copies pages on write and the VM resumes with its processes already running, its page cache already warm, its listening sockets already bound. Nothing is re-done because nothing was undone. That is how a create can be 179 ms at p50 while doing strictly more than a cache hit: it is not rebuilding the state, it is resuming it.

A cache is a bet that you can rebuild the state faster than you can keep it. A snapshot is a bet that you can keep it cheaply enough never to rebuild it.

Both bets are correct in context. A cache wins when the state is big, boring and reconstructible, and the machine you restore it into was coming anyway — which describes CI exactly. A snapshot wins when the expensive part is reaching the running state: a loaded model, a seeded database, a built app with a dev server already listening. And a snapshot has one property a cache structurally cannot: you can fork it.

Four scenarios, and the honest hybrid

Blacksmith: you want faster Actions tomorrow

Your test suite takes 22 minutes, your engineers context-switch while they wait, and someone has started asking about the hosted-runner bill. You have no appetite for a platform project. The correct move is a one-line change to your runner label, a week of measuring, and a `git revert` if it disappoints. A sandbox API cannot compete for this job, because winning it would require me to first rebuild the Actions control plane, and I have not.

Blacksmith: your build system is Actions YAML, and should stay that way

Some organisations have years of accumulated truth in their workflow files: reusable workflows, composite actions, environment approvals, required checks wired into branch protection, OIDC trust with three cloud accounts. That estate is institutional memory, not technical debt. Re-expressing it against a sandbox API means owning all of it, including the parts whose purpose nobody remembers — and the risk is not that it is hard, it is that you will faithfully reproduce nine of the ten behaviours and find the tenth during an incident. Keep the estate; change the machines under it.

PandaStack: an agent runs model-generated code hundreds of times an hour

Now a workload with no workflow file in sight. A language model emits Python, your product executes it, the result goes back into the conversation — on user request, hundreds of times an hour, unpredictably, and each execution must not see another user's data or the inside of your production network. The blast radius of a model-generated `rm -rf` should be a machine that was going to be deleted in ninety seconds anyway.

There is no label for this: no job to dispatch, no commit to key off, no YAML to put a runner in — just a function call in a request handler and a hard requirement that the environment appear inside a user-facing latency budget. Snapshot restore at p50 179 ms is what makes "a whole virtual machine per execution" a sane sentence. The tokenless preview URL — `https://<port>-<sandbox-id>.<suffix>` — then shows the user what their code started, and that UUID is a password: there is no token endpoint, the sandbox ID is the credential, and short TTLs are your mitigation for a URL pasted into Slack.

PandaStack: fork a warm, built environment into N variants

The second PandaStack-shaped job is the one that made me build the fork-tree verb. You have an environment that was expensive to reach — repo cloned, dependencies installed, project built, database migrated and seeded — and you want to try eight mutually exclusive things against that exact state and keep whichever works. Eight dependency-bump candidates; eight patches an agent generated for one failing test. A cache cannot do this: it restores the ingredients eight times and you cook eight times. A snapshot fork restores the finished state eight times.

import secrets
from pandastack import Sandbox

CANDIDATES = [
    "react@18.3.1", "react@19.0.0", "react@19.1.0", "react@19.2.0",
    "react@19.2.4", "react@canary", "react@experimental", "react@beta",
]

# ---- pay for the expensive state exactly once -----------------------------
parent = Sandbox.create(template="base", ttl_seconds=3600, persistent=True)
parent.exec("git clone --depth 1 https://github.com/acme/web /work", check=True)
parent.exec(
    "timeout 900 sh -c 'cd /work && npm ci && npm run build'",
    timeout_seconds=960,
    check=True,
)

# ---- branch the built machine, not the build ------------------------------
# Two different verbs, and for this job the difference is the whole point:
#
#   fork()      clones the parent's DISK only (reflink / dm-snapshot:
#               O(metadata), data shared until written) and the child then
#               COLD-BOOTS -- its own kernel, its own memory, its own entropy
#               pool. You inherit the installed deps, not the running build.
#
#   fork_tree() snapshots the parent ONCE (pause -> snap -> resume, so the
#               parent is not modified) and restores N children from that one
#               snapshot, so children inherit memory AND disk. This is the warm
#               one, and the warm one is what we want here.
#
# Restoring each child is the snapshot-restore path: 400-750 ms same host,
# 1.2-3.5 s cross-host because the snapshot has to move first. Both verbs are
# capped at 16 children per call.
children = parent.fork_tree(count=len(CANDIDATES), metadata={"exp": "react-bump"})

results = []
for candidate, kid in zip(CANDIDATES, children):
    # THE FOOTGUN, and it belongs to fork_tree specifically: these children were
    # all restored from ONE snapshot, so they share the parent's entropy pool
    # and any PRNG a process seeded before the snapshot was taken. Eight
    # children that agree on "random" values produce colliding IDs and duplicate
    # tokens. Mix in bytes generated OUT HERE, where they are actually distinct,
    # and restart anything that seeded a PRNG at boot. The guest clock is fine:
    # the platform re-syncs it on restore/resume/wake, which covers the
    # TLS-expiry class of failure -- it does not reseed a warm PRNG for you.
    nonce = secrets.token_hex(32)
    kid.exec(f"printf %s {nonce} > /dev/urandom", check=True)

    r = kid.exec(
        f"timeout 600 sh -c 'cd /work && npm install {candidate} && npm test'",
        timeout_seconds=660,
    )
    results.append((candidate, r.exit_code))

for candidate, code in results:
    print(f"{'PASS' if code == 0 else 'FAIL'}  {candidate}")

# ---- keep the winner, delete everything else ------------------------------
# `persistent=True` on the parent exempted it from the idle reaper, which also
# means the idle reaper will not clean up after you. Delete it yourself.
for kid in children:
    kid.kill()
parent.kill()
Know which verb you called, because they are not the same mechanism. `fork()` clones the parent's disk and then cold-boots the child, which gets a fresh kernel, fresh memory and a fresh entropy pool — you inherit the installed dependencies, not the running build. `fork_tree(count=N)` snapshots the parent once and restores N children from that snapshot, so they inherit memory as well as disk, and therefore share an entropy pool and any PRNG seeded before the snapshot. Anything generating IDs or tokens in those children will agree with its siblings: mix in bytes from outside the clone and restart processes that cached a seed. Both verbs cap at 16 children per call.

The hybrid, which is what most people who need both should do

Use a managed runner for CI and a sandbox API for the product workload. That is not a diplomatic fudge, it is the configuration I would recommend to a company with both problems, because the two are not competing for the same budget line. CI spend is a function of how often your engineers push. Sandbox spend is a function of how many customers use the feature that runs their code. Those grow for unrelated reasons and are owned by different teams, and conflating them produces a migration project that makes both worse.

The exception: if your isolation requirements have outgrown what any runner vendor documents — regulated workloads, build secrets that must never coexist with third-party code on one kernel, an auditor who wants to see the boundary — then you are buying a provable property rather than faster CI, and microVM-per-job is worth the build.

The shared thesis: the cost of a fresh environment is the product

Here is the part I find genuinely interesting, and the reason these two keep landing in the same search results despite having almost nothing in common at the interface. Both are bets that the dominant cost in modern compute is not running the work. It is arriving at a state where the work can start.

Decompose that cost and there are three pieces. The machine: something has to allocate hardware, boot a kernel, configure networking. The state: dependencies, caches, a built artifact, a migrated database, a warm process. And the proof of cleanliness: convincing yourself nothing from the last tenant is still in there. Every ephemeral-compute product is a set of answers to those three, and the answers are where the engineering lives.

A managed runner on bare metal attacks the first two. Dedicated hardware removes the noisy-neighbour variance and slow shared storage that make cloud CI unpredictable; caching keeps the expensive files close to where the job will run, so the dependency install stops being a download. Described qualitatively, because that is all I will say about someone else's internals: the bet is that CI's real enemy is cold caches on slow shared machines, and that fixing both without changing the user's interface is a product.

PandaStack attacks the same three costs from a different angle, which I can describe with numbers because they are mine. There is no warm pool of idle VMs — not an optimisation we skipped, but a design choice, because a warm pool is a bill for machines nobody is using. Every create restores a baked snapshot instead: allocate one of 16,384 pre-allocated per-host network namespace slots, patch the tap MAC, reflink the rootfs, fork/exec Firecracker, load the snapshot, resume, probe TCP :22. That lands at p50 179 ms and p99 around 203 ms, with the snapshot load itself 49-80 ms of it. The first-ever spawn of a template pays a cold boot of about 3 seconds and bakes the snapshot so nobody after it does.

The state cost gets the same treatment, and this is where the theses visibly diverge: rather than caching the files that produce the state, bake the state itself into the snapshot — processes running, page cache warm, dev server listening — then restore it, N times if you want. Restoring a child of a warm snapshot is 400-750 ms on the same host and 1.2-3.5 s across hosts, where the snapshot has to travel first. Idle cost goes to zero by deleting the VM entirely and keeping only the snapshot: app-hosting scale-to-zero sleeps by deleting the microVM after baking a seed to object storage, and wakes by restoring it in about 1.2 seconds. Cleanliness is free in this model — a fresh VM with its own kernel has no previous tenant.

Two bets, one premise: cold shared-tenant containers on slow shared storage are the thing to beat. The products diverge on what "warm state" means — files you restore versus a machine you resume — and on who the caller is. That is the whole comparison.

What PandaStack cannot do, so the table is not a sales document

I asked you to verify everything I said about Blacksmith, so here is the list I would want handed to me about my own product before a purchase order.

  • There is no CI control plane. No job graph, no matrix expansion, no log UI, no required-checks integration, no secret scoping. If you want Actions, you build the ephemeral-runner registration loop yourself, and a finished runner product is less work.
  • RAM is fixed at template-bake time. Firecracker cannot change guest RAM or vCPU count at snapshot restore, so `memory_mb` on a create is silently corrected to the baked value — worse than an error, because it is quiet.
  • Every template gets 8 burstable vCPUs, shared under contention by cgroup weight. You cannot ask for 64 cores for one job.
  • `fork()` and `fork_tree()` are both capped at 16 children per call, so a 32-way fan-out is two calls or a tree.
  • No GPU in the guest: Firecracker's device list has no graphics device, so there is nothing to pass through.
  • No nested virtualisation, and the guest kernel is 5.10 — not one you pick. If your build needs KVM inside the job, wrong substrate.
  • `timeout_seconds` on exec is a client deadline, not a server-enforced limit. A runaway command needs `timeout` or `ulimit` inside the guest.
  • Preview URLs are tokenless — the sandbox UUID is the bearer credential and there is no token endpoint. Short TTLs and not logging the UUID are load-bearing habits.

How I would decide, in one question

Ask what dispatches the work.

If the answer is "a push, a pull request, a cron in a workflow file, or anything else GitHub Actions already knows how to trigger", the work is CI and you want a runner. Change the label, measure for a week, revert if it disappoints. The one-line reversibility makes the experiment nearly free.

If the answer is "an HTTP request from a user, a message on a queue, a language model deciding it needs a shell, or a line in my own program", there is no workflow file to put a label in. You need an API that creates machines, and what matters is how fast one appears, how little an idle one costs, whether you can fork a warm one, and where the isolation boundary sits.

If the answer is both, buy both. They are not fighting over the same budget, and a comparison post that pretended otherwise would be the kind of thing you should stop reading at the second paragraph.

Frequently asked questions

Can I replace my GitHub Actions runners with PandaStack, and should I?

You can, and for most teams you should not. The pattern works: create a microVM, register it with GitHub as an ephemeral self-hosted runner, let it accept exactly one job, then delete the VM. Because every create is a snapshot restore at p50 179 ms rather than a boot, a VM-per-job runner is not obviously slower to provision than a container-per-job one, and it gives you a boundary a container cannot — separate kernel, separate network namespace, nothing shared with the previous job because the previous job's machine no longer exists. What you are signing up for is everything around that. You own the registration and deregistration lifecycle, including the ugly cases: a runner that registered and then died before taking a job, a job cancelled mid-run, a rate limit during a burst. You own capacity, log handling, artifact upload, and whatever your branch protection rules expect to see. None of that is research, but all of it is work, and a managed runner product has already done it. My honest recommendation: do this when you need a provable isolation property that no runner vendor documents — regulated workloads, build secrets that must never share a kernel with third-party code, an auditor with opinions. Do not do it to make CI faster. There is a cheaper way to do that and it is a one-line diff.

Is a managed runner on dedicated hardware safe enough for untrusted fork pull requests?

Often, yes — and it is worth being precise about why, because the runner is usually not the weak link. The two things that actually get people compromised by fork PRs are pull_request_target, which runs workflow code in a privileged context with access to your secrets while checking out fork-authored code, and long-lived self-hosted runners that keep state between jobs so one job can poison the next. Both of those are configuration problems in your repository, not properties of who supplies the machine. Fix them first: use pull_request rather than pull_request_target for anything that touches fork code, do not expose secrets to fork-triggered jobs, and never point a persistent self-hosted runner at a public repository. Once that is done, a vendor documenting per-job virtual-machine isolation on its own hardware is a genuinely reasonable posture and I would not try to scare you off it — check their current isolation docs rather than my summary, since this is exactly the detail that gets improved between blog posts. Where I would push further is when the requirement is not "reasonably isolated" but "demonstrably isolated to a named boundary": one kernel per job, and an architecture diagram an auditor will accept. That is a microVM-per-job argument, and it is about provability rather than any vendor being careless.

What is the real difference between a warm build cache and a snapshot restore?

A cache restores files; a snapshot restores a machine. When a cache hits you get a directory back — node_modules, a compiler cache, a layer store — and your build still does the work that turns those files into a running state: resolve, link, compile, start processes, warm a JIT. You saved the download and maybe the compile, not the startup. When a snapshot restores, guest memory is mapped MAP_PRIVATE from the snapshot file so the kernel copies pages on write, and the VM resumes with processes already running and sockets already bound. On PandaStack that is p50 179 ms, p99 around 203 ms. The practical test: is the expensive thing obtaining files, or reaching a running state? If it is files — a big dependency tree, a layer cache — a cache is simpler and better, especially when the platform was going to hand you a machine regardless. If it is the running state — a loaded model, a seeded database, a built app with a dev server listening — a snapshot skips the step a cache cannot. And a snapshot can be forked: fork_tree snapshots a warm environment once and restores each child in 400 to 750 ms on the same host, giving you N copies of a finished state rather than N rebuilds. Mind the verb, though: plain fork() clones the disk and cold-boots the child, and both verbs cap at 16 children per call.

If I use both, am I paying twice for the same compute?

No, because they are metered against different things that grow for different reasons. CI spend scales with how often your engineers push code. Sandbox spend scales with how much your users use the feature that runs their code: more customers, more executions, more VM-seconds. One is an internal engineering-productivity cost; the other is cost of goods sold for a product feature. Treating them as one budget line is how you end up with a migration project justified by a number that was never a duplicate. The legitimate version of the worry is a CI-shaped workload you moved onto a sandbox API while the Actions path still runs alongside it: real waste, and the fix is to pick one path per workload rather than to consolidate vendors. On the PandaStack side, the costs worth modelling are VM time while a sandbox exists — which is why TTLs and calling kill() on the success path matter — and egress, which is metered and is the dimension people forget. We do not bill per request, deliberately: a per-invocation charge makes developers avoid the platform for exactly the high-frequency, short-duration workloads it is best at. Idle cost on a sleeping app is approximately zero because the VM is deleted rather than paused, with a snapshot kept in object storage and a wake of about 1.2 seconds.

What is the single biggest thing people get wrong when they move a CI job to a sandbox API?

They assume the platform enforces the limits they set in their client code, and specifically that timeout_seconds on exec is a server-side kill. It is not. It is a client deadline: your SDK call gives up waiting, but the command inside the guest keeps running, and so does the VM it is running in. A build that hangs on a prompt or a network read sits there burning VM time while your orchestrator has already recorded a failure and moved on. There are two mitigations and you want both. Put a hard wall inside the guest — timeout 600 sh -c '...' is enough, plus ulimit for memory-shaped runaways — so the command dies on schedule. And set ttl_seconds on every create, because the TTL is enforced by a platform-side reaper rather than by your process, so a crashed orchestrator cannot leak a running sandbox at your expense. The second matters most for anyone coming from CI, where this bug class does not exist: a CI platform owns the job's lifecycle and kills it for you, so you have never had to think about it. A sandbox API hands you the lifecycle, including the responsibility for ending it. The related trap is persistent=True, which exempts a sandbox from the idle reaper: the right flag for a long-lived parent you intend to fork from, and a guarantee that nothing cleans up after you if you forget.

Keep reading

Related posts

  • Running Nix Builds Inside a Disposable VM

    Nix's sandbox is a hermeticity fence, not a hypervisor — a hostile derivation still runs on your kernel. Here's how to put Nix inside a disposable VM without a cold /nix/store making every build miserable.

  • Wiring preview environments into pull requests

    Reading a diff tells you whether the code is reasonable. A URL tells you whether the feature works. Getting from one to the other is about a hundred lines of CI and one decision about databases.

  • Bazel Remote Execution Workers on Firecracker microVMs

    A remote cache poisoned by a leaky worker is a supply-chain compromise with excellent uptime — every developer pulls it, nobody rebuilds it, and the digests all check out.

  • CircleCI Self-Hosted Runners on MicroVMs

    The moment you move a CircleCI job onto your own machine runner, you quietly trade a fresh VM per job for a box that remembers every build that ever ran on it. That trade is the whole security story, and you do not have to make it.

  • Isolating CI Build Caches Per-Job with MicroVMs

    Your build cache is a shared mutable global variable that ships to production. Here's how per-job microVMs give you the cache hit-rate without the cross-job poisoning.

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.