all posts

A Real Red-Green Loop for Coding Agents, One MicroVM per Attempt

Ajay Kumar··11 min read

An agent that can only read code is guessing eloquently. An agent that can run the tests is doing engineering. The difference is one tool call, and the loop it unlocks is the oldest one we have: write a failing test, run it, read the actual error, patch, run again, stop when it goes green.

Everything interesting in that loop is in the word “run”. And “run” means executing code a language model wrote about ninety seconds ago, which on your laptop means executing it as you, with your SSH agent, your cloud credentials and your working tree all in scope.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This post is the practical build for the run step: where it should execute, how to make a failure legible to a model, and — the part most people get wrong on the first attempt — how to make the loop fast enough that an agent can take thirty swings at a bug without you going to lunch.

The test runner is the agent's only honest feedback channel

Linters and type checkers are cheap, fast and shallow, and you should absolutely run them — they catch a real class of mistake before you pay for a test run. But a type error tells you a shape is wrong, not that the behaviour is wrong, and an agent can satisfy a type checker with a cast and a shrug. A test suite is different in kind. It is a specification written in the only language that cannot be argued with: a process either exits zero or it does not.

This matters because of a specific property of the models we are building on. They are extremely confident, which is exactly why there is a test. “I've identified and fixed the issue” is a sentence whose truth value is determined entirely by a process exit code somewhere else, and the model producing it has no privileged access to that exit code. Nor do you, until something runs.

  • It is binary. There is no partial credit and no hedging. The agent cannot produce a test result that is 70% passing in the way it can produce an explanation that is 70% right.
  • It is grounded in the real dependency tree. The suite runs against the versions actually installed, not the API the model remembers from its training data — which is how you catch the confidently-invented keyword argument.
  • It catches the collateral damage. The interesting output is rarely the test you were fixing; it is the three unrelated tests that went red because the fix changed a shared helper.
  • It localises. A stack trace names a file and a line. That is a far better prompt than a description of a symptom, and it is free.

So the test runner is the signal. The question is only where it runs, and the honest answer is that most people are running it on the laptop and hoping.

What a model-written test can actually reach

Not hysteria, and not a vendor's threat-model slide. Just four things that have happened to people, none of which require the model to be adversarial — only to be wrong in the ordinary way.

  • A test that computes a path and then removes it. `shutil.rmtree(os.path.join(base, name))` is fine right up until `base` resolves to an empty string and `name` resolves to `..`. The agent did not want to delete your repo; it wanted to clean up a fixture directory.
  • A test that opens ten thousand sockets. Usually a retry loop with no backoff against a service that is down. On a shared host this is not a failing test, it is an incident for everything else on the box.
  • A `conftest.py` that mutates your real environment. Fixtures run as you, with your environment: a fixture that writes a profile into `~/.aws/credentials` to “set up a test account” is a plausible thing for a model to generate and a genuinely bad afternoon.
  • A suite that passes because the assertion is gone. This is the famous one. The agent deleted the test file, the exit code went to zero, and the loop declared victory.

That last one deserves a moment, because it is funny until it is in your main branch. It is not malice and it is not a jailbreak; it is a gradient. You told a system to make the exit code zero, and `rm tests/test_the_hard_one.py` makes the exit code zero in one step with no reasoning required. If your loop's only input is the exit code, you have built a machine that optimises for deleted tests. The fix is two lines, and it is in the code below: check the diff as well as the output, and reject a green run that touched the test directory.

The other three are isolation problems, and isolation is where the substrate choice earns its keep. The point of running them in a microVM is not that nothing bad happens — things will absolutely still go wrong, loudly and often, because that is what a loop is for. The point is that the worst case becomes a VM you were going to delete at the end of the iteration anyway.

Check what your runner shares with you before you check anything else. A test process with your home directory on it has your credentials, and a bind-mounted repository means a model's miscalculated `rm -rf` reaches your actual working tree. A mounted Docker socket is root on the host with extra steps. Most agent-sandbox incidents are not kernel escapes; they are mounts someone added for convenience six months earlier.

The two-layer trick: bake the dependencies, branch the attempt

This is the heart of it, and it is an arithmetic problem rather than a security one. Get it wrong and the loop is too slow to use, you will conclude that sandboxes are slow, and you will go back to running `npm test` on the laptop and hoping.

A fresh sandbox per iteration is dominated by `npm ci`

Start with the naive design: create a sandbox, clone the repo, install dependencies, run the tests, delete the sandbox, repeat. On PandaStack the create itself is a snapshot restore rather than a boot — p50 179 ms, p99 203 ms, with the `/snapshot/load` step inside that in the 49–80 ms range, and around 3 seconds only for the first ever spawn of a template before a snapshot exists. That is not your problem.

Your problem is the next three lines. A `git clone` and an `npm ci` or a `uv sync` on a real repository is not a 179 millisecond operation, and it is not a 2 second operation either. Multiply it by thirty attempts and the loop is useless. If you benchmark “a sandbox per iteration” naively you will produce a number dominated entirely by your lockfile, and then draw a conclusion about virtual machines.

So split the work by how often it changes. The dependency tree changes when the lockfile changes, which is rarely. The patch changes every iteration. Pay for the first once and the second thirty times.

Layer one is the warm parent. Either bake it into a template — `pandastack template build -f Dockerfile -n my-repo-tests --memory-mb 4096`, with the toolchain and dependencies installed at build time — or create one long-lived sandbox, clone, install, and run the suite once so that the resolver caches, the bytecode cache and the page cache are all hot. Then snapshot it. A suite's second run on a warm machine is a different animal from its first.

Layer two is branching that warm state per attempt, and here there are three mechanisms that people treat as interchangeable and which are emphatically not:

  • `fork()` pauses the parent, clones its rootfs, resumes the parent, and cold-boots the child from that disk. The child inherits the parent's filesystem — the repo, `node_modules`, the caches — and gets its own fresh kernel, its own memory and its own entropy.
  • `fork_tree(count=N)` snapshots the parent's memory and disk once, then restores N children from that single snapshot in parallel. These children inherit memory as well as disk: a warm interpreter comes with them, and so does the parent's clock and its RNG state. Capped at 16 children per call. Same-host forks land in 400–750 ms; cross-host is 1.2–3.5 s, because the snapshot has to be pulled across the network first.
  • `snapshot()` once, then `Sandbox.create(from_snapshot=...)` per attempt. This is the same restore path `fork_tree`'s children take, without re-snapshotting the parent on every call — which is what you want for a sequential loop, where you are producing one child at a time and the parent has not changed.
The one bit that distinguishes those three is whether the parent's memory comes along, and it decides more than warm-up time: it decides whether your children share an RNG stream and a clock. A disk-only fork gives you an independent machine that happens to have your dependencies on it. A memory-inheriting restore gives you a clone of a running machine, with everything that implies.

What branching does not give you

  • A clean `git` state. The child's filesystem is the parent's filesystem at branch time, uncommitted mess included. Snapshot a parent with a half-applied patch in it and every child begins with a half-applied patch, which you will spend an afternoon blaming on the model. Run `git reset --hard && git clean -fdx` as the child's first command, every time, excluding your dependency directories.
  • Independent randomness, on the memory-inheriting paths. Siblings restored from one snapshot agree on what time it is and generate the same “random” values. A suite that picks a random port, a random temp directory name or a UUID will collide across all of them simultaneously, which presents as flakiness and is in fact determinism. Pass each child an explicit seed.
  • Changeable RAM. Firecracker cannot alter vCPU count or guest memory at snapshot restore, so RAM is fixed with `--memory-mb` at template build time and a `memory_mb` on a create is silently corrected to the baked value. `base` is baked at 4 GiB because a `tsc` or Next build OOMs at 2 GiB; `code-interpreter` and `agent` are 2 GiB. Every template gets 8 burstable vCPUs sharing cores under contention via cgroup `cpu.weight`.
  • Surviving TCP connections. Sockets do not live through a snapshot. If your parent held a connection open to a test database, the child comes back holding a dead file descriptor rather than an error you would enjoy debugging.

The loop, in code

Here is the whole thing: a warm parent built once, a child per attempt, and a transcript that contains real exit codes and real stderr rather than anybody's summary of them.

import os
from pandastack import Sandbox

# PANDASTACK_API_KEY in the environment. (PANDASTACK_TOKEN was removed in 0.3.0.)
OUTPUT_BUDGET = 8000          # characters of test output the model is allowed to see


# ------------------------------------------------------------------ layer 1
# The warm parent: built ONCE per (repo, lockfile) pair and reused by every
# attempt. Everything slow happens here -- clone, toolchain, install, and one
# throwaway run so the resolver cache, the pytest cache and the page cache are
# all hot. persistent=True keeps the idle reaper off it while you iterate.
def warm_parent(repo_url: str, ref: str = "main"):
    parent = Sandbox.create(
        template="base",                 # 4 GiB baked. See the note below.
        persistent=True,
        metadata={"role": "tdd-parent", "repo": repo_url},
    )
    setup = f"""
set -eux
git clone --depth 1 --branch {ref} {repo_url} /work
cd /work
# mise honours .python-version / .nvmrc / .tool-versions; reshim exposes the
# console scripts (pytest, uvicorn) that land in the runtime's own bin.
export MISE_DATA_DIR=/opt/mise MISE_CONFIG_DIR=/opt/mise PATH=/opt/mise/shims:$PATH
mise install && mise reshim
if [ -f uv.lock ]; then uv sync --frozen; else pip install -e '.[dev]'; fi
# Prime the caches. We do not care whether this passes -- we care that the next
# run does not have to byte-compile the world. This is the cost we are hoisting
# out of the loop, and on a real repo it dwarfs the VM by an order of magnitude.
timeout 600 pytest -q --collect-only >/dev/null 2>&1 || true
"""
    parent.exec(setup, timeout_seconds=900, check=True)
    # Snapshot the warm state ONCE and restore from it per attempt. fork_tree()
    # would re-snapshot the parent on every call, which is the right trade for
    # N siblings at once and the wrong one for a sequential loop.
    return parent, parent.snapshot()


# ------------------------------------------------------- making it legible
def legible(cmd: str, r, budget: int = OUTPUT_BUDGET) -> str:
    """What the model actually sees: exit code first, both streams labelled and
    verbatim, truncated FROM THE MIDDLE.

    Middle-truncation is not a style preference. Every runner worth using puts
    the first failure near the top of the buffer and the summary line at the
    very bottom. `tail` throws away collection errors and import failures;
    `head` throws away "1 failed, 412 passed", which is the single most
    informative line in the whole run. Keep both ends; the 300 identical
    deprecation warnings in the middle are exactly what you wanted to drop.
    """
    def squeeze(s: str, n: int) -> str:
        if len(s) <= n:
            return s
        head = n // 2
        tail = n - head - 60
        elided = len(s) - head - tail
        return f"{s[:head]}\n... [{elided} characters elided from the middle] ...\n{s[-tail:]}"

    half = budget // 2
    return "\n".join([
        f"$ {cmd}   (cwd=/work)",      # an agent that cannot see the cwd will
        f"exit_code: {r.exit_code}",   # eventually decide the file is missing
        "--- stderr ---",              # stderr FIRST: the traceback lives here
        squeeze(r.stderr, half),
        "--- stdout ---",
        squeeze(r.stdout, half),
    ])


# ------------------------------------------------------------------ the loop
def red_green(snap_id: str, propose, max_attempts: int = 8):
    """One attempt == one child restored from the warm snapshot, then deleted.

    `propose(transcript) -> unified diff` is your model call. It sees every
    previous attempt's real output and nothing else.
    """
    transcript: list[str] = []
    for attempt in range(max_attempts):
        child = Sandbox.create(
            from_snapshot=snap_id,     # the same restore path fork_tree children take
            ttl_seconds=900,           # the backstop for when YOUR process dies
            metadata={"role": "tdd-attempt", "attempt": str(attempt)},
        )
        try:
            # A branch of a dirty tree is a dirty tree. The child's filesystem is
            # the parent's filesystem at snapshot time, uncommitted mess
            # included -- "fresh VM" does not imply "fresh git".
            child.exec(
                "cd /work && git reset --hard && git clean -fdx -e node_modules -e .venv",
                timeout_seconds=120, check=True,
            )

            patch = propose(transcript)
            child.filesystem.write("/work/fix.patch", patch)
            applied = child.exec("cd /work && git apply --verbose /work/fix.patch",
                                 timeout_seconds=60, check=False)
            if applied.exit_code != 0:
                # A patch that will not apply is the cheapest honest failure you
                # will ever get. Feed the apply error back and let it try again.
                transcript.append(legible("git apply /work/fix.patch", applied))
                continue

            # check=False is the whole point: a non-zero exit is DATA, not an
            # exception. timeout_seconds is a CLIENT deadline only -- the hard
            # kill lives in run-tests, inside the guest.
            cmd = "/usr/local/bin/run-tests 'pytest -x -q --tb=short'"
            r = child.exec(cmd, timeout_seconds=600, check=False)
            transcript.append(legible(cmd, r))

            if r.exit_code == 0:
                # Green. Before believing it: did the patch touch the tests?
                # The failure mode is not malice, it is a gradient -- the
                # cheapest path to a zero exit code is sometimes `rm`.
                tests = child.exec("cd /work && git diff --stat -- tests/", check=False)
                if tests.stdout.strip():
                    transcript.append(
                        "REJECTED: the patch modified tests/. Green by deletion "
                        "is not green.\n" + tests.stdout)
                    continue
                # Extract the artefact BEFORE the finally clause deletes the VM.
                return child.filesystem.read("/work/fix.patch").decode(), attempt
        finally:
            child.kill()   # every iteration, success and failure alike
    return None, max_attempts

Three lines in there carry most of the value. `check=False` on the test run, because a non-zero exit is the data you came for and raising on it would be like raising on a compiler warning. The `git reset --hard` first, because the child's tree is the parent's tree. And the `git diff --stat -- tests/` check on the green path, which is the entire defence against the loop discovering that deletion is a valid optimisation.

Making a failure legible to a model

Where the tests run is half the problem. What you hand back is the other half, and it is the half people skip. A model reading test output is doing exactly what you do when you scan a CI log: looking for the first real error and the summary line, and ignoring everything between them. Build the buffer accordingly.

  1. Exit code first, on its own line, as a number. It is the one unambiguous fact in the entire response and it should not be something the model has to infer from prose.
  2. Both streams, labelled, verbatim, stderr first. The traceback is on stderr. Do not merge them into one buffer — you lose the signal that something wrote to stderr at all, which on some runners is the only sign of a crash versus a failure.
  3. Truncate from the middle, never from the end. More on this below; it is the single highest-leverage change in this section.
  4. Cap the total size in characters before it reaches the context window, not in lines. A stack-trace line is 200 characters and a progress dot is one, so a line budget is not a budget.
  5. Include the command and the working directory. An agent that cannot see which directory it ran in will eventually conclude the test file does not exist, and then write a new one.

The middle-truncation point is worth the paragraph. `pytest`, `jest`, `go test` and every other sane runner put the first failure near the top of the output and the summary at the very bottom. A `tail -n 200` discards collection errors and import failures — the ones where nothing ran at all. A `head -n 200` discards “1 failed, 412 passed”, which is arguably the most informative token sequence in the whole buffer, because it is the difference between “you fixed one thing” and “you broke the world”. Keep both ends, elide the middle, and say how much you elided so the model knows it is reading an excerpt rather than a complete log.

Better than truncating is narrowing at the source. `-x` to stop at the first failure, `--tb=short`, and a test selector for the single test you are fixing. A targeted run of one test is faster, cheaper and dramatically more legible than a middle-truncated full suite, and the full suite is a thing you run once at the end to check you have not broken anything else.

And for a suite that genuinely takes nine minutes, use `exec_stream` rather than `exec`. It wraps the agent's SSE stream, calls your `on_stdout` and `on_stderr` callbacks with each chunk as it arrives, and returns the exit code as an int. That lets the agent react to the first failure at second fifteen instead of waiting out the other eight and three-quarter minutes to learn the same thing.

Do not summarise the test output before the model sees it. Your summariser's job would be to decide which lines mattered — which is precisely the decision you were paying the model to make.

Four patches at once: fork_tree as a search

Here is where the branching gets genuinely interesting rather than merely hygienic. Ask a model for a fix to a non-obvious bug and you often get three or four plausible candidates it is equally confident about. Equal confidence across mutually exclusive patches is information about the model, not about the bug. So stop ranking them by vibes. Run all four, keep the one that goes green, kill the rest.

import secrets
from concurrent.futures import ThreadPoolExecutor

# The model offers four fixes it is equally confident about. Equal confidence
# across mutually exclusive patches is information about the model, not about
# the bug -- so stop ranking them by vibes and run all four.
def best_of_n(parent, patches: list[str]):
    # ONE snapshot of the parent's live memory+disk, then N children restored
    # from it in parallel. Children inherit the parent's MEMORY as well as its
    # disk: a warm interpreter comes along, and so does the parent's clock and
    # its RNG state. Capped at 16 children per call.
    kids = parent.fork_tree(count=len(patches), metadata={"role": "candidate"})

    def attempt(pair):
        child, patch = pair
        # Re-seed, per child, explicitly. Memory-identical siblings agree on
        # what time it is and produce the same "random" values, so a suite that
        # picks a random port or a random tempdir name collides across all four
        # at once. That reads as flakiness and is in fact determinism.
        # (PandaStack re-syncs the guest wall clock on restore/resume/wake with
        # `date -u -s`, which covers the TLS-expiry class of failure. It does
        # not re-seed a user-space PRNG that was already warm.)
        seed = secrets.randbelow(2**31)
        child.filesystem.write("/work/fix.patch", patch)
        r = child.exec(
            f"cd /work && git reset --hard -q && git apply /work/fix.patch && "
            f"PYTHONHASHSEED={seed} /usr/local/bin/run-tests 'pytest -x -q -p no:randomly'",
            timeout_seconds=600, check=False,
        )
        return child, r

    with ThreadPoolExecutor(max_workers=len(kids)) as pool:
        results = list(pool.map(attempt, zip(kids, patches)))

    winner = next((c for c, r in results if r.exit_code == 0), None)
    if winner is not None:
        # promote() detaches the winner from the tree; cleanup_siblings reaps
        # the losers server-side so a crashed orchestrator cannot leak them.
        winner.promote(cleanup_siblings=True)
        return winner, [legible("pytest", r) for _, r in results]

    for child, _ in results:
        child.kill()
    return None, [legible("pytest", r) for _, r in results]

When this pays: an expensive suite, cheap patches, and genuine ambiguity about the fix. Four candidates evaluated in parallel cost you roughly one suite's wall-clock time instead of four, and the selection criterion is an exit code rather than a second opinion from the same model that produced the ambiguity.

When it does not pay: a one-line fix you can verify sequentially in two seconds, and any loop where you would be calling `fork_tree(count=1)` per iteration. `fork_tree` snapshots the parent on every call, so one call for N candidates is a good trade and N calls of one child each is a worse one than snapshotting the parent once and restoring from it. The sequential loop in the previous section is deliberately written the other way round.

And this is exactly where the shared-RNG caveat bites, because it bites all the siblings at once. Four memory-identical children will all decide that the free port is 54321, all generate the same request ID, and all fail in the same way — and because they fail identically you will read it as a real bug in all four patches rather than as a property of how they were created.

Hard limits, because the loop exists to find the one you forgot

One thing to internalise before you run this unattended: `timeout_seconds` on the exec API is a client deadline. The agent's exec endpoint accepts the field and does not act on it — your HTTP client gives up and raises, and the `pytest` process inside the guest carries right on, consuming the CPU you are billed for until something else stops it. If you need a hard limit, it has to live in the guest.

#!/usr/bin/env bash
# /usr/local/bin/run-tests -- baked into the template, so every attempt inherits
# it and no orchestrator bug can forget to apply it.
#
# Why it exists: timeout_seconds on the exec API is a CLIENT deadline. The
# agent's exec endpoint decodes the field and does not act on it -- your HTTP
# client gives up and the pytest process in the guest carries right on, burning
# CPU you are paying for. The only limit the guest respects is one that lives in
# the guest.
set -uo pipefail                 # NOT -e: a red test must reach the exit code below
CMD=${1:?usage: run-tests "<command>"}

WALL=${WALL:-600}                # hard wall-clock ceiling, seconds
CPUSEC=${CPUSEC:-540}            # CPU-seconds: catches a spin loop that emits
                                 #   nothing and therefore trips no idle detector
FILE_MB=${FILE_MB:-512}          # a test that writes an infinite log is a disk
                                 #   exhaustion bug wearing a test's clothes
PROCS=${PROCS:-512}              # fork bombs are usually accidental
NOFILE=${NOFILE:-4096}           # "a test that opens 10,000 sockets" is a Tuesday

cd /work || exit 70

(  # a subshell, so the limits bind the test run and not this script
  ulimit -t "$CPUSEC"
  ulimit -f $(( FILE_MB * 1024 ))
  ulimit -u "$PROCS"
  ulimit -n "$NOFILE"
  ulimit -c 0                    # no core dumps; a 3 GiB core file is not evidence

  # DO NOT reach for `ulimit -v` here. It caps virtual address space, and both
  # V8 and the Go runtime reserve vastly more virtual memory than they ever
  # touch -- so -v kills perfectly healthy processes and the failure looks like
  # a mysterious allocator crash. Your real memory ceiling is the guest's baked
  # RAM: Firecracker cannot change vCPU or RAM at snapshot restore, so it is
  # chosen with `--memory-mb` at template build time and a memory_mb on the
  # create is silently corrected to the baked value.

  # --signal + --kill-after is the load-bearing part. A plain `timeout` sends
  # TERM, and every runner worth using traps TERM to print a summary -- so a
  # hung suite receives a polite request it is entirely free to ignore. Send
  # TERM so it can flush, then KILL it ten seconds later regardless.
  #
  # setsid puts the run in its own process group so the kill reaches the
  # children. Without it you kill the shell and orphan the four pytest-xdist
  # workers that were the actual problem.
  exec timeout --signal=TERM --kill-after=10s "${WALL}s" \
       setsid env PYTHONDONTWRITEBYTECODE=1 HOME=/work/.home \
       sh -c "$CMD"
)
rc=$?

# Name the exit code. "The suite timed out" and "the suite was killed for
# memory" lead to completely different next patches, and a bare 137 reads to a
# language model like a random number it is free to ignore.
case "$rc" in
  0)   echo "RESULT: green" ;;
  124) echo "RESULT: timeout -- exceeded ${WALL}s of wall clock" ;;
  137) echo "RESULT: killed by SIGKILL -- the --kill-after, ulimit -t, or the OOM killer" ;;
  *)   echo "RESULT: red (exit $rc)" ;;
esac
exit "$rc"

Three layers, and only the weakest is in your client. `timeout --signal=TERM --kill-after=10s` inside the guest for the command, so a runner that traps TERM to flush a summary gets to do that and then dies anyway. `ulimit` for the process, catching the spin loop that produces no output and therefore trips no idle detection. And `ttl_seconds` on the sandbox as the backstop, enforced by a platform-side reaper rather than by your orchestrator, so a loop that crashes at attempt nineteen cannot leak a running VM.

`setsid` is not decoration. Without its own process group, your `timeout` kills the shell and orphans the four `pytest-xdist` workers that were the actual problem — and those workers are now unparented processes in a guest nobody is watching, which is exactly the state the TTL exists to clean up. Put the run in its own process group and kill the group.

Laptop, container, one sandbox, one branch per attempt

The only numbers below are PandaStack's own. Everything about the other rows is a shape rather than a measurement, because the shape is the part that generalises.

Four places to run an agent's test suite, compared on the things that actually decide the design.
Where the tests runIteration latencyDependency warm-upClean between attemptsWhat a hostile test reachesParallel candidates
Your laptopFastest possible — nothing to provisionPaid once, then freeNothing is reset. Attempt 7 inherits attempt 6's half-applied patch, its stray processes and its edits to `node_modules`Your home directory, your SSH agent, your cloud credentials, your Docker socket, whatever VPN you are onOne. You have one working tree
Container per iterationFast — process-speed startFree if the deps are in the image; a full install per run if notGood for files. Shared kernel state — sysctls, the clock, the page cache — is still sharedA shared host kernel, plus everything you mounted. A bind-mounted repo or a Docker socket is the whole gameGood. Cheap enough to run several
One long-lived sandbox, reusedJust the exec. Cheapest option after the first createPaid oncePoor, and it degrades invisibly. This is the configuration where a suite starts passing for reasons nobody can reconstructA VM you can delete — but one accumulating every attempt's debris, including the `.env` the agent wrote at attempt 3One, unless you shard inside it
Branch per attempt (fork or restore)Same-host fork 400–750 ms; cross-host 1.2–3.5 s. A create from a baked snapshot is p50 179 ms, p99 203 msPaid once into the parent snapshot and inherited by every childPristine disk every attempt — plus an explicit `git reset`, because the parent's tree is whatever you snapshottedIts own kernel, its own netns and /30 out of 16,384 per host, nothing of yours. Worst case is a VM you were deleting anywayNative: `fork_tree(count=N)`, up to 16 children from one parent snapshot

To be fair to row two: a container per iteration is a perfectly good answer for running your own test suite against your own reviewed code, and I would not replace it in your CI. The argument here is narrower. In normal CI, a human reviewed the code before it ran. In an agent loop the code was written moments ago by a model optimising for a zero exit code, and it runs unreviewed by construction — that is the entire point of the loop. The threat model is different because the provenance is different.

Iterations are cheap; a forgotten sandbox is not

The cost shape of this architecture is worth stating plainly, because it is the opposite of what people brace for. A sandbox that exists for the ninety seconds of one test run and is then deleted is close to free. Thirty of them is thirty times close to free. The cost is not in the iterations.

The cost is the one you leave running. A parent sandbox marked `persistent=True` so the idle reaper leaves it alone — exactly as in the code above — is a machine that will sit there, billed, until something deletes it. The failure mode is not an expensive loop; it is a Friday afternoon experiment whose warm parent is still there on Monday, with a transcript nobody read.

So: `kill()` at the end of the episode, on the success path as well as the failure path, and set a TTL on everything including the parent. The TTL is the net, not the plan. Something that only ever fires when your own cleanup failed is a safety property; something you rely on for normal operation is a leak with a timer on it.

What to build, in order

  1. A warm parent per (repo, lockfile): clone, toolchain via `mise`, install, and one throwaway suite run to warm the caches. Snapshot it. Tag the snapshot with the lockfile hash so you know when it is stale.
  2. A guest-side `run-tests` wrapper with `timeout --signal=TERM --kill-after`, `setsid` and `ulimit`, baked into the template so no orchestrator bug can skip it.
  3. A child per attempt restored from that snapshot, with `git reset --hard && git clean -fdx` as its literal first command and a `kill()` in a `finally`.
  4. An output formatter: exit code first, both streams labelled, stderr first, truncated from the middle with the elision counted, capped in characters.
  5. A green-run guard that reads the diff and rejects a pass that modified the test directory. Two lines. Add them before the first unattended run, not after.
  6. `exec_stream` with `-x` once your suite is long enough that waiting for it is the dominant cost, so the agent reacts to the first failure rather than the last.
  7. `fork_tree(count=N)` for the genuinely ambiguous fix, with an explicit per-child seed, and `promote(cleanup_siblings=True)` so the losers do not outlive the decision.

None of this is exotic. It is a test runner, a snapshot, and a virtual machine you throw away, assembled carefully enough that a model can fail in public thirty times in a row without taking anything with it. The interesting part was never the sandbox. It was deciding that the only trustworthy thing in the entire loop is an integer returned by a process, and then building everything else around getting that integer honestly.

Frequently asked questions

Why not just run the agent's tests in a container per iteration?

For a lot of setups you should, and I would not rip containers out of your CI to put this in. The argument is narrower than container-versus-VM in general: it is about provenance. In a normal CI job, a human reviewed the code before it ran. In an agent loop the code was generated moments ago by a model optimising for a zero exit code, and it runs unreviewed by construction — that is the entire point of the loop. So the question becomes what that unreviewed code can reach. A container boundary is namespaces and seccomp over a shared host kernel, which is a bet that nothing the model writes finds a kernel bug. In practice that bet is rarely what loses; what loses is the mounts. A bind-mounted repository means a miscalculated `rm -rf` reaches your real working tree, and a mounted Docker socket is root on the host with extra steps. A microVM moves the boundary to hardware virtualisation: its own guest kernel, its own device model, its own network namespace from a pool of 16,384 pre-allocated /30 subnets per host. The historical objection was start cost, and snapshot-restore removes it — every create restores a baked snapshot rather than booting, p50 179 ms and p99 203 ms, with around 3 seconds only on the very first spawn of a template before a snapshot exists. If every attempt on the host is yours, and the host holds nothing you care about, a container is fine. Multi-tenant, or a host with real credentials on it, and it is not.

How much test output should I give the model, and how should I truncate it?

Budget in characters, not lines — a stack-trace line is 200 characters and a progress dot is one, so a line budget is not a budget at all. Then truncate from the middle, never from the end. Every runner worth using puts the first failure near the top of the buffer and the summary at the very bottom, so `tail` discards collection and import errors (the cases where nothing ran) while `head` discards “1 failed, 412 passed”, which is the difference between “you fixed one thing” and “you broke the world”. Keep both ends, elide the middle, and state how many characters you elided so the model knows it is reading an excerpt. Better still, narrow at the source rather than truncating after the fact: `-x` to stop at the first failure, `--tb=short`, and a selector for the single test you are fixing. A targeted run of one test is faster, cheaper and far more legible than a middle-truncated full suite, and the full suite becomes a thing you run once at the end. Put the exit code first and on its own line, because it is the only unambiguous fact in the response. Label stdout and stderr separately and put stderr first, since the traceback lives there and merging the streams loses the signal that anything wrote to stderr at all. Include the command and the working directory. And do not summarise before the model sees it — your summariser's job would be to decide which lines mattered, which is exactly the decision you were paying the model to make.

Does forking give each attempt a genuinely clean repository?

Clean relative to the parent, which is not the same thing as clean — and the two branch mechanisms differ in ways worth knowing. `fork()` pauses the parent, clones its rootfs, resumes the parent, and cold-boots the child from that disk: the child gets the parent's filesystem and its own fresh kernel, memory and entropy. `fork_tree(count=N)` snapshots the parent's memory and disk once and restores N children from it in parallel, so those children inherit memory too — a warm interpreter, a warm page cache, and also the parent's clock and RNG state. In both cases the filesystem is a copy of the parent's filesystem at branch time, including whatever uncommitted mess was in the working tree. Snapshot a parent that still has a half-applied patch in it and every single child starts from that half-applied patch, and you will spend an afternoon blaming the model for a problem you created. So reset explicitly: `git reset --hard && git clean -fdx`, with your dependency directories excluded, as the child's first command, every time. Two further catches. A memory-inheriting child's wall clock is as stale as the parent's freeze — PandaStack re-syncs the guest clock on restore, resume and wake, which covers the TLS-certificate class of failure, but it does not re-seed a user-space PRNG that was already warm, so pass each child an explicit seed. And guest RAM is baked: Firecracker cannot change vCPU or memory at snapshot restore, so it is fixed with `--memory-mb` at template build time and a `memory_mb` on a create is silently corrected to the baked value. `base` is 4 GiB; `code-interpreter` is 2 GiB, which a `tsc` or Next build will exceed.

How do I actually stop a runaway test suite?

In three layers, and only the weakest one lives in your client. `timeout_seconds` on the exec call is a client deadline: the endpoint accepts the field and does not enforce it, so when your HTTP client gives up and raises, the process in the guest carries right on burning CPU you are paying for. The real limit goes in the guest shell. Use `timeout --signal=TERM --kill-after=10s` so the runner gets a chance to flush a summary and then dies regardless — a plain `timeout` sends TERM, and every runner worth using traps TERM, so a hung suite receives a polite request it is free to ignore. Wrap the run in `setsid` so the kill reaches the process group instead of orphaning the four `pytest-xdist` workers that were the actual problem. Then add `ulimit -t` for CPU seconds, which catches the spin loop that emits no output and therefore trips no idle detector; `ulimit -n` for the test that opens ten thousand sockets; `ulimit -u` for the usually-accidental fork bomb; and `ulimit -f` for the test that writes an infinite log. Be careful with `ulimit -v`: it caps virtual address space, and both V8 and the Go runtime reserve far more virtual memory than they ever touch, so `-v` kills healthy processes and the failure looks like a mysterious allocator crash. Your real memory ceiling is the guest's baked RAM. The third layer is `ttl_seconds` on the sandbox, enforced by a platform-side reaper rather than by your orchestrator, so a loop that dies at attempt nineteen cannot leak a running VM. Then call `kill()` on the success path too. An agent in a loop is a machine for discovering the limit you forgot, and it will run that discovery a few thousand times.

Keep reading

Related posts

More in AI agent sandboxes · See AI agent sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.