all posts

Your Dependency Bot Runs Strangers' Install Scripts

Ajay Kumar··11 min read

State the setup plainly, because stated plainly it is slightly alarming. You have a robot. It holds a credential that can create branches and open pull requests in your repositories. It runs on a schedule, unattended, usually at an hour chosen because nobody is around. And its entire function — the reason it exists, the thing you are paying it to do — is to fetch code that was published to a public registry by someone you have never met, possibly minutes ago, and run it.

Not "review" it. Run it. An upgrade bot that only edits a version string in a manifest is nearly useless, because the whole value of the thing is the sentence "and the tests still pass." To produce that sentence it has to resolve a new dependency tree, install it, and execute your test suite against it. Installation executes code. Your test suite imports the new package, which also executes code. There is no configuration in which a useful upgrade bot does not run third-party code.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This post is not another round of "npm install is arbitrary code execution" — that argument is made, and the ecosystem has started acting on it. It is about the specific machine we have all quietly exempted from the conclusion: the bot's own execution environment.

The robot you gave push access to

Line up the properties of an upgrade bot against the properties of a thing you would deliberately expose to untrusted code, and they are close to inverted.

  • Its credentials are usually better than a contributor's. A human opening a dependency PR needs a fork. The bot pushes a branch to the real repository, which means a token with write access to the thing you are protecting. On a self-hosted runner it often also inherits whatever that runner has: a registry token, a cloud role, a cache credential.
  • It runs unattended. No human is watching the log while it happens, and the window between "a malicious version was published" and "your bot installed it" is a cron interval, not a code review.
  • It is long-lived and repeatedly used. Bots are the classic long-running service on a box someone set up once. Every candidate it processes runs on the same kernel as the last one, unless you did something about it.
  • It has to reach the network by definition. You cannot airgap the thing whose job is fetching from a registry.
  • And its output is the single most reflexively approved artefact in this industry. A PR titled `chore(deps): bump left-pad from 1.3.0 to 1.3.1` gets an approving emoji in four seconds from someone who has not opened the Files tab.
The one pull request nobody reads is the one whose entire content came from the internet.

That is the joke and it is also the threat model. The bot is a privileged, automated, scheduled path from a public registry into your main branch, with arbitrary code execution in the middle of it, terminating in a review step that social convention has agreed to skip.

What actually executes, per ecosystem

Check these against your own toolchain's current docs rather than taking anyone's word for it, mine included. Several of these defaults have moved recently and one of them moved this year.

npm moved the default, and that changes the advice

The old statement — npm runs `preinstall`, `install` and `postinstall` from your dependencies by default — was true for a very long time. It is no longer the whole truth. The `ignore-scripts` config is documented with a default of `false` and the effect "If true, npm does not run scripts specified in package.json files", which is the old world. But npm 12 changed the shipped default: install scripts from dependencies are off unless the package appears in an `allowScripts` field in your `package.json`, and git and remote-URL dependencies became opt-in too. The docs for `npm approve-scripts` describe the new behaviour precisely: "Install commands silently skip lifecycle scripts for any dependency that does not have a matching entry in `allowScripts`, and end with a list of the packages whose scripts were skipped." npm 12 is the current `latest` tag as of this writing.

This is a genuinely good change and it does not rescue an upgrade bot, for two reasons that are worth separating.

  • `allowScripts` entries are PINNED by default. `npm approve-scripts` writes `pkg@1.2.3` unless you pass `--no-allow-scripts-pin`. An upgrade is, definitionally, a version change — so the new version falls out of the allowlist and its scripts are skipped during the bot's install. Which means the install the bot tested is not the install a developer gets after they approve the bump. Your green checkmark attests to a tree that nobody will ever actually run.
  • Script-blocking does nothing for the test phase. The `ignore-scripts` docs note that commands designed to run a particular script — `npm test` among them — still run it. And `npm test` runs your test suite, which imports the upgraded package, which executes its module body. That is not a lifecycle script; it is the package working as intended. Blocking install hooks is the easy half of the problem and the test suite is the hard half.
The pinning behaviour creates a specific, quiet failure: on npm 12 your bot can produce a passing test run for a dependency tree in which the new version's install script never executed, and a developer then approves that version and runs the script on their laptop. The bot proved the code works. It proved nothing about the script. If you want the test to mean something, reproduce the approved shape deliberately inside the disposable guest.

Python: the wheel is the control, not the flag

pip's situation is cleaner than most people assume, and the clean part is the wheel. A wheel is installed by unpacking it; there is no install hook, so nothing in it runs at install time. An sdist is the opposite: to get so much as the package's metadata, pip has to invoke the project's build backend, which for a setuptools project means executing `setup.py`. The real lever is therefore `--only-binary`, documented as "Do not use source packages. Can be supplied multiple times... Accepts either `:all:` to disable all source packages, `:none:` to empty the set, or one or more package names with commas between them." `--only-binary :all:` is the pip equivalent of refusing to run install-time code, and it fails loudly rather than silently when no wheel exists.

One thing to be precise about, because the name invites the mistake: PEP 517 build isolation is not a security boundary. `--no-build-isolation` is documented as "Disable isolation when building a modern source distribution. Build dependencies specified by PEP 518 must be already installed if this option is used." That is dependency isolation — a clean environment so the build's requirements do not collide with yours. The build backend still runs as your user, with your filesystem and your network. Nothing about it is a sandbox.

Everyone else, briefly

  • Ruby: a gem with a native extension ships an `extconf.rb`, which RubyGems executes at install time to generate a Makefile before invoking `make`. It is not a config file; it is an arbitrary Ruby program that runs unreviewed. `bundle install` inherits this.
  • Rust: `cargo build` compiles and runs a crate's `build.rs` before building the crate. Install-time code execution with a different name and a nicer reputation.
  • Go: the notable exception. `go build` has no dependency-authored install hook — the module's code is compiled, not executed, until you run something that calls it. Which still covers your test suite, but it means the install step itself is genuinely boring.
  • Everything else with native extensions: the pattern generalises. Wherever a package has to compile something on your machine, somebody's script runs on your machine.

Three jobs, three trust levels

Here is the structural point, and it is the one I think most upgrade-bot deployments get wrong. The bot does three things that look like one thing. They need the network differently, they need credentials differently, and they trust the dependency differently.

  1. RESOLVE. Produce a new lockfile. This needs the registry's metadata and nothing else — no repository secrets, no cloud role, no write token. And it can be done without executing any third-party code at all: `npm install --package-lock-only` installs nothing, so no lifecycle script has anything to hook. This is the cheap, safe step, and it is the only one of the three that is genuinely low-risk.
  2. INSTALL. Materialise the tree. This is where third-party code executes — deliberately, because an install you neutered is an install you did not test. It needs the registry and a writable filesystem. It needs no credentials whatsoever.
  3. TEST. Run your suite against the new tree. This is the step that wants things: a database, fixtures, maybe an API key for an integration test, your repository's own code. It is also the step that executes the dependency's code by design, because that is what importing a library does.

Now notice what happens when you run all three in one process on one machine, which is the default shape of every bot deployment I have seen. The resolve step's zero-credential requirement is irrelevant, because the machine is holding the test step's secrets and the bot's push token the entire time. The blast radius of a hostile install script is not "the install" — it is the union of everything all three jobs needed, plus whatever the runner happened to have lying around. You took three tasks with three different trust levels and gave all of them the maximum of the three.

The fix follows directly and does not require any new technology to state: put the steps that execute third-party code somewhere that holds nothing, and keep the credential in a process that executes nothing.

The split: a guest with nothing to steal, a bot with nothing to run

Concretely: the install and the test happen in a disposable microVM that has no git token, no registry token and no cloud identity. The bot — which holds the token — never runs a line of dependency code. It starts the guest, feeds it a repository archive, waits, reads the result out, and opens the pull request. The token and the install script are never in the same execution context, which is the only property in this entire post that actually matters.

The usual objection is that the bot still trusts data produced inside a compromised guest. Correct, and worth being exact about how you handle it. A result "signed" inside the guest proves nothing — a compromised guest can sign whatever it likes with whatever key you put there. The trust anchors are elsewhere: the exit code comes from the host-side agent reporting on the process, not from the guest asserting its own success; and the guest's output is rendered into the PR body as inert text, never evaluated, never used as a command, never trusted to name a branch. Treat it exactly as you would a form submission from the internet, because that is what it is.

import json
from pathlib import Path

from pandastack import Sandbox, ExecInterrupted, SandboxTimeout

CANDIDATE = {"name": "sharp", "to": "0.34.2"}
REPO_TAR = Path("/tmp/repo.tar")          # `git archive` of the target ref


def run(script: str, sbx: Sandbox, budget: int):
    """Stream a long step and return (exit_code, log). exit_code None == UNKNOWN."""
    log: list[str] = []
    sbx.filesystem.write("/work/step.sh", script)
    try:
        # timeout_seconds is CLIENT-side only: the agent decodes it on both exec
        # endpoints and applies it on neither. On exec_stream it raises the HTTP
        # timeout; one-shot exec() dies at 30s, which is not a budget you can
        # run a test suite inside. The enforceable deadline is in the shell.
        rc = sbx.exec_stream(
            "sh /work/step.sh",
            on_stdout=log.append,
            on_stderr=log.append,
            timeout_seconds=budget,
        )
    except SandboxTimeout:
        return 124, log
    except ExecInterrupted as e:
        # The stream ended with no exit event: host gone, agent restarted,
        # sandbox reaped. The outcome is genuinely unknown -- and for a bot
        # with auto-merge switched on, unknown MUST NOT round to green.
        log.append(e.partial_stdout or "")
        return None, log
    return rc, log


# ---------------------------------------------------------------------------
# JOB 1 -- RESOLVE. Needs the registry's metadata and nothing else. No install,
# so no lifecycle script runs; `--ignore-scripts` is belt and braces.
# ---------------------------------------------------------------------------
resolver = Sandbox.create(template="base", ttl_seconds=900,
                          metadata={"job": "resolve", "pkg": CANDIDATE["name"]})
resolver.filesystem.write("/work/repo.tar", REPO_TAR.read_bytes())  # single files only
rc, _ = run(f"""set -eux
cd /work && mkdir -p repo && tar -xf repo.tar -C repo && cd repo
cp package-lock.json /work/before.json
npm install --package-lock-only --ignore-scripts \\
    {CANDIDATE['name']}@{CANDIDATE['to']}
""", resolver, budget=600)
assert rc == 0, "resolution failed; nothing to test"

new_lock = resolver.filesystem.read("/work/repo/package-lock.json")
pkg_json = resolver.filesystem.read("/work/repo/package.json")
resolver.kill()

# ---------------------------------------------------------------------------
# JOB 2+3 -- INSTALL AND TEST. This guest executes third-party code on purpose.
# It holds no git token, no registry token, no cloud credential, and it is
# never reused: one candidate, one kernel, one lifetime.
# ---------------------------------------------------------------------------
worker = Sandbox.create(template="base", ttl_seconds=2400,
                        metadata={"job": "install-test", "pkg": CANDIDATE["name"]})
worker.filesystem.write("/work/repo.tar", REPO_TAR.read_bytes())
worker.filesystem.write("/work/package-lock.json", new_lock)
worker.filesystem.write("/work/package.json", pkg_json)
worker.filesystem.write("/work/upgrade-surface.py", Path("upgrade-surface.py").read_text())

rc, log = run("""set -u
cd /work && mkdir -p repo out && tar -xf repo.tar -C repo
cp /work/package-lock.json /work/package.json repo/
cd repo

# Reproduce the install a developer WILL get once they approve the scripts, not
# the neutered one npm 12 hands a bot by default. Check the flag surface with
# `npm help approve-scripts` against the npm you pin -- this area landed
# recently and is still settling.
npm approve-scripts --all --no-allow-scripts-pin || true

npm ci > /work/out/install.log 2>&1; install_rc=$?
tail -40 /work/out/install.log

python3 /work/upgrade-surface.py /work/before.json package-lock.json \\
    package.json > /work/out/surface.md; surface_rc=$?

npm test > /work/out/test.log 2>&1; test_rc=$?
tail -200 /work/out/test.log

printf 'install_rc=%s\\nsurface_rc=%s\\ntest_rc=%s\\n' \\
    "$install_rc" "$surface_rc" "$test_rc" > /work/out/status

# This script's exit status is what the host-side agent reports back, so make
# it mean something rather than echoing the exit code of `tail`. surface_rc is
# advisory -- non-zero means new install scripts appeared, which gates
# auto-merge, not the PR.
[ "$install_rc" = 0 ] && [ "$test_rc" = 0 ]
""", worker, budget=1800)

# Read the evidence out BEFORE killing: an explicit kill() cascade-deletes this
# sandbox's snapshots, so a tidy `finally: kill()` can bin what you just made.
surface = worker.filesystem.read("/work/out/surface.md").decode("utf-8", "replace")
status = worker.filesystem.read("/work/out/status").decode("utf-8", "replace")
worker.kill()

# ---------------------------------------------------------------------------
# THE BOT. Holds the push token. Ran no third-party code -- not one line, not
# in a subshell. Everything below treats the guest's output as inert text.
# ---------------------------------------------------------------------------
# rc is None when the outcome was never reported, so this rejects unknown as
# well as failed. Only an explicit 0 from the host-side agent opens a PR.
if rc != 0:
    tail = "".join(log)[-2000:]
    raise SystemExit(f"candidate rejected (exit {rc!r})\n{status}\n{tail}")

body = "\n".join([
    f"Bumps `{CANDIDATE['name']}` to {CANDIDATE['to']}.",
    "",
    "Installed and tested in a disposable microVM holding no credentials.",
    "",
    surface.replace("```", r"\`\`\`"),   # defang fences in guest output
    "",
    f"<details><summary>raw status</summary>\n\n```\n{status}\n```\n</details>",
])
# Your SCM client. This is the ONLY place the push token is touched, in a
# process that has executed no third-party code.
open_pull_request(branch=f"deps/{CANDIDATE['name']}-{CANDIDATE['to']}", body=body)
The `ExecInterrupted` branch is not defensive padding. `exec_stream` raises it when the SSE stream ends with no terminal exit event — the host went away, the agent restarted, the sandbox got reaped mid-command. The SDK deliberately refuses to return 0 in that case, because the outcome is unknown. For a bot with auto-merge enabled, "unknown" rounding to "green" is the whole bug.

Egress is the control that matters here

The realistic attack on an upgrade bot is not destruction. Nobody is going to `rm -rf` your runner; that gets noticed in minutes and achieves nothing. The attack is reading `~/.npmrc`, `.git/config`, the process environment, and the cloud metadata endpoint, and sending what it finds somewhere. It is a quiet, fast, read-and-post operation that completes during an install you were not watching, and the compromised machine goes on working perfectly afterwards.

Which makes isolation and egress two separate controls answering two separate questions. Isolation decides what the code can reach locally. Egress decides whether it can tell anyone. You want both, and you should be honest about which ones you actually have:

  • Per-VM network isolation is structural, not a feature you enable. Each agent pre-allocates 16,384 /30 subnets, so every sandbox gets its own network namespace, its own tap device and its own NAT identity. There is no shared bridge for one candidate's install to sniff another's traffic on.
  • The cloud metadata hazard is closed by construction. The entire link-local range is dropped at the top of the host firewall chain, above any accept rule, so a guest cannot ask the hypervisor's metadata service for instance credentials. This is the single highest-value thing on this list and it is also the one people most often leave to an in-guest rule that the in-guest attacker can flush.
  • Cross-tenant VM-to-VM traffic is dropped the same way, with the rule ordering pinned by tests.
  • What you do NOT get from us is an outbound allowlist. Egress is open by default — that is a deliberate product decision, because agents install packages and clone repositories — paired with per-VM attribution and abuse monitoring. If you need a hard registry-only allowlist, build it outside the guest: point the guest at a registry proxy you run and give it no other route. An `iptables` rule inside the guest is a tripwire, not a boundary, because it is enforced by the kernel the install script is also running on.
  • Treat DNS as an exfil channel. An IP allowlist that still permits arbitrary resolution leaves the trick of encoding a stolen token into a subdomain lookup wide open. If you care about this, the resolver has to be yours too.

And one pleasant side effect of the guest being a microVM rather than a container: its environment is small enough to audit completely. Docker `ENV` from the template build is not injected into the Firecracker guest at all. `/etc/environment`, read by PAM for every session, is what actually reaches a detached process. So "what can this guest see?" has a one-file answer, which is a question containers rarely answer so cleanly.

#!/bin/sh
# Runs INSIDE the install-and-test guest, as its first action, before npm.
# Two jobs: prove the guest is empty of credentials, and make exfiltration
# loud. Nothing here is the security boundary -- the microVM is. These are a
# tripwire and a sanity check, and they are worth running anyway.
set -u

# 1. Prove the credential surface is empty. A bot that has ever shared a runner
#    with a build has probably inherited one of these by accident.
fail=0
for f in "$HOME/.npmrc" /etc/npmrc "$HOME/.git-credentials" "$HOME/.netrc" \
         "$HOME/.config/gh/hosts.yml" "$HOME/.docker/config.json" \
         "$HOME/.aws/credentials" /var/run/secrets; do
  [ -e "$f" ] && { echo "FAIL: credential path present in guest: $f"; fail=1; }
done
git config --get remote.origin.url | grep -q '@' \
  && { echo "FAIL: remote URL carries a credential"; fail=1; }
env | grep -Eiq '(^|_)(TOKEN|SECRET|PASSWORD|API_KEY)=' \
  && { echo "FAIL: secret-shaped environment variable present"; fail=1; }
[ "$fail" = 0 ] && echo "ok: guest holds nothing worth stealing"

# The guest's env is auditable precisely because it is small: Docker ENV from
# the template build is NOT injected into the Firecracker guest. /etc/environment
# (read by PAM for every session, login or not) is what actually reaches a
# detached process, so that one file is the whole inherited-env story.
echo "--- /etc/environment ---"; cat /etc/environment

# 2. Default-deny egress with the registry allowlisted, as a TRIPWIRE. Log
#    first, then drop, so the denied packets are evidence rather than silence.
#    Caveat worth being honest about: these rules are enforced by the same
#    guest kernel the install script is running on. Root in the guest can flush
#    them. If you need a hard allowlist, put it in a registry proxy outside the
#    guest and give the guest no other route.
REG=registry.npmjs.org
iptables -N EXFIL 2>/dev/null
iptables -A OUTPUT -o lo -j ACCEPT
for ip in $(getent ahostsv4 "$REG" | awk '{print $1}' | sort -u); do
  iptables -A OUTPUT -d "$ip" -p tcp --dport 443 -j ACCEPT
done
iptables -A OUTPUT -p udp --dport 53 -j ACCEPT   # DNS is also an exfil channel:
                                                 # see the note below on why an
                                                 # IP allowlist alone is leaky
iptables -A OUTPUT -j LOG --log-prefix "EXFIL-ATTEMPT "
iptables -A OUTPUT -j REJECT

# Link-local (169.254.0.0/16, cloud metadata) and cross-tenant VM traffic are
# already dropped at the host firewall on PandaStack, above any accept rule --
# a guest cannot ask the hypervisor's metadata service for instance credentials.
# You do not need a rule for it. Add one anyway if you run this elsewhere.

The lockfile diff is the actual review artefact

Everything above reduces the damage. This part reduces the probability, and it is the highest-leverage thing an upgrade bot can do that most upgrade bots do not do. A dependency bump's `package-lock.json` diff is where the transitive surprise shows up. You asked for one patch release of one package; the lockfile will tell you that it also pulled in four packages you have never heard of, one of which runs a `postinstall`.

The lockfile already carries the flag you need. In `lockfileVersion` 2 and 3, each entry under `packages` may have `hasInstallScript`, documented as "A flag to indicate that the package has a `preinstall`, `install`, or `postinstall` script." So the bot can compute, mechanically and before any human looks, the set of newly-arrived packages that execute code at install time. That is a sentence worth putting at the top of the PR body in bold, and it is worth gating auto-merge on.

#!/usr/bin/env python3
"""
upgrade-surface.py -- turn a lockfile diff into the two facts a reviewer
actually needs: which packages are NEW to the tree, and which of those
execute code at install time.

Run it in the disposable guest, write the output into the PR body.

  git show "HEAD:package-lock.json" > /tmp/before.json
  npm install --package-lock-only --ignore-scripts          # resolve only
  ./upgrade-surface.py /tmp/before.json package-lock.json package.json
"""
import json
import sys
from pathlib import Path

REGISTRY = "https://registry.npmjs.org/"   # your mirror, if you run one


def packages(lock_path: str) -> dict:
    lock = json.loads(Path(lock_path).read_text())
    if lock.get("lockfileVersion", 0) < 2:
        sys.exit("lockfileVersion 1 has no `packages` map; regenerate with npm >= 7")
    out = {}
    for path, meta in (lock.get("packages") or {}).items():
        if not path or meta.get("link"):
            continue            # "" is the root project; `link` entries are workspaces
        name = path.rsplit("node_modules/", 1)[-1]
        out[name] = {
            "version": meta.get("version", "?"),
            "resolved": meta.get("resolved") or "",
            # Documented meaning: "the package has a preinstall, install or
            # postinstall script". NOTE it does NOT cover `prepare`, which also
            # runs on install for git dependencies.
            "install_script": bool(meta.get("hasInstallScript")),
            "dev": bool(meta.get("dev")),
        }
    return out


def covered(allow: list, name: str, version: str) -> bool:
    # `npm approve-scripts` writes PINNED entries (pkg@1.2.3) unless you pass
    # --no-allow-scripts-pin. So a version bump falls OUT of the allowlist.
    return name in allow or f"{name}@{version}" in allow


before = packages(sys.argv[1])
after = packages(sys.argv[2])
allow = json.loads(Path(sys.argv[3]).read_text()).get("allowScripts", []) or []

added = sorted(set(after) - set(before))
bumped = sorted(n for n in set(after) & set(before)
                if after[n]["version"] != before[n]["version"])

print(f"### Upgrade surface\n")
print(f"- {len(added)} package(s) new to the tree")
print(f"- {len(bumped)} package(s) changed version")
print(f"- {len(after) - len(before):+d} net packages\n")

executes = [n for n in added + bumped if after[n]["install_script"]]
if executes:
    print("#### Runs code at install time\n")
    for n in executes:
        v = after[n]["version"]
        tag = "NEW" if n in added else f"{before[n]['version']} -> {v}"
        state = "pre-approved" if covered(allow, n, v) else "NOT in allowScripts"
        print(f"- `{n}@{v}` ({tag}) -- {state}")
    print()
else:
    print("#### No install-time scripts in the changed set\n")

offreg = [n for n in added + bumped
          if after[n]["resolved"] and not after[n]["resolved"].startswith(REGISTRY)]
if offreg:
    print("#### Resolved from outside the registry (review individually)\n")
    for n in offreg:
        print(f"- `{n}` <- {after[n]['resolved']}")
    print()

# Exit non-zero so the surrounding job can gate auto-merge on "nothing new
# executes". A clean bump of a pure-JS package is a different risk class from
# one that drags in a package with a postinstall, and the bot should say so.
sys.exit(1 if [n for n in executes if not covered(allow, n, after[n]["version"])] else 0)

Two deliberate details in there. `hasInstallScript` does not cover `prepare`, which also runs on install for git dependencies — so the off-registry `resolved` check is not redundant with it. And the allowlist check compares against both the bare name and the pinned `name@version` form, because `npm approve-scripts` pins by default; a bump that silently leaves the allowlist is exactly the case you want surfaced rather than smoothed over.

A bot that posts "this upgrade adds 4 transitive packages, 1 of which has a postinstall, and here it is" has done more for your security posture than a bot that bumps a version and goes quiet. It converts the reflexive approval from a liability into a reasonable one, because now the four-second glance at the PR actually contains the information.

Throughput: a monorepo does not generate one candidate

A repository with 400 direct and transitive dependencies generates a steady stream of upgrade candidates, and the whole design above insists that each one gets a clean install-and-test in a guest that is never reused. If a clean environment costs a minute to produce, that arithmetic kills the design before it starts and you will end up reusing the runner, which is where we came in.

Snapshot restore is what makes one-guest-per-candidate affordable rather than aspirational. Creating a sandbox from a baked template is a snapshot restore, not a boot: p50 179 ms, p99 203 ms, of which the `/snapshot/load` step itself is around 49 ms. The first create of a template before its snapshot exists is a real cold boot at roughly 3 seconds, after which the snapshot is captured and every subsequent create is on the fast path.

The honest note, which I would rather say than let you discover: the VM is not your bottleneck and never was. Your `npm ci` is tens of seconds at best and your test suite is minutes. Two hundred milliseconds of guest creation disappears entirely into the noise of the work the guest is doing. The point of the number is not that it makes your pipeline fast — it is that it removes the only reason you had for reusing a machine. There is no longer an efficiency argument for letting candidate 41 run on the kernel candidate 40 just finished with.

You can push this further by baking a template that already contains your toolchain and a populated module cache, so the common part of every install is paid once at bake time rather than 400 times. Two caveats: guest RAM is chosen at template-build time via `--memory-mb` and cannot be changed at restore, because Firecracker cannot change a guest's vCPU or RAM at snapshot restore — the agent silently corrects whatever `cpu` or `memory_mb` you pass on create to the baked values. And the rootfs flag `--size-mb` defaults to 1024 MB on the build API (the bundled Go CLI sends 2048), which a real toolchain plus a warm `node_modules` cache will exhaust without drama. Size both deliberately.

What each control actually stops

Graded against one concrete threat: a newly-published transitive dependency whose install script reads credentials and posts them out. Verify the third-party rows against the current docs of whatever you deploy; this is a shape, not a benchmark.

Controls for an upgrade bot, graded against a credential-stealing install script in a new transitive dependency.
ControlStops the install script runningStops the dep's code during `npm test`Stops the bot's token being stolenStops a kernel-level escape
`--ignore-scripts`, or the npm 12 defaultYes, for preinstall/install/postinstall — but pinned `allowScripts` means the tested tree is not the approved treeNo. `npm test` runs by design and importing the package runs its module bodyNo. Nothing here is a network or credential controlNot applicable — no boundary involved
Container per job on a shared runnerNoNoPartly, and only if the secrets genuinely are not in the container. Mounted sockets, shared caches and inherited env undo it routinelyNo. One kernel, shared by every candidate you process
gVisor or a user-space kernelNoNoPartly. A much better local boundary; the token is still in scope if the job holds itMostly. Syscalls are interposed in user space, at the cost of compatibility gaps on syscall-heavy installs
microVM per candidate, no credentials in the guestNo, and it does not try toNo, and it does not try toYes. The token was never in that execution context at allYes. Separate kernel, own network namespace, destroyed after one candidate

The column that matters is the third one, and the thing to notice is that the bottom row fails the first two columns on purpose. Isolation does not prevent the dependency's code from executing. It makes the execution not matter, by ensuring there is nothing in reach worth taking and nowhere useful to send it. Those are different goals and conflating them is why "we use containers" feels like an answer.

What this does not fix

A post that stopped at the happy path would be an advert. Several of these are structural rather than roadmap items.

  • The merge is still the hole. Isolating the test run does not stop a compromised package's code from being merged into your main branch and then running in your real CI, your real deploy and your users' browsers. This architecture protects the bot. It does not protect the repository from a bad upgrade — only review, provenance and a delay before adopting brand-new versions do that.
  • A green test run is weak evidence of safety. A supply-chain payload that cares about staying undetected will not break your tests; it will behave perfectly and read your environment. "Tests passed" means the upgrade is compatible, not that it is benign, and an upgrade bot that implies otherwise is doing harm.
  • Integration tests want secrets, and that is a genuine conflict with everything above. If your suite needs a real database and a real API key, you have put a credential in the guest that executes dependency code. Mitigate it with short-lived, narrowly-scoped, per-run credentials and separate the suites — the tests that need secrets and the tests that prove the upgrade compiles are usually not the same tests — but do not pretend the conflict is gone.
  • Egress is open by default on our platform, so the registry-only allowlist in this post is something you build, not something you switch on. The parts you get for free are the per-VM namespace, the link-local drop and cross-tenant separation.
  • Guest RAM and vCPU are fixed by the baked snapshot. A test suite that needs 16 GiB needs a template baked at 16 GiB; passing `memory_mb` on create is not an error, it is silently corrected, which is worse.
  • An explicit `kill()` cascade-deletes that sandbox's snapshots, while the idle reaper does not. A tidy `finally: sbx.kill()` after an interesting failure can bin the evidence you were about to investigate. Read the artefacts out first.

When this is overkill

If your bot resolves lockfiles and never installs or tests anything, it does not execute third-party code and this post does not apply to you. You have also not got an upgrade bot worth much, but that is a trade you are allowed to make, and for a small repository with a human who actually reads diffs it is a defensible one.

If you are on a hosted bot running entirely on someone else's infrastructure with no token of yours on the box, your exposure is their problem and their design. Worth reading how they do it; not worth rebuilding. The case for this architecture is the self-hosted bot on your own runner with your own tokens — which is what you end up with the moment you have private registries, internal packages or a monorepo big enough that the hosted option stops coping.

And if you take one thing: the dependency bot is not a special case that deserves an exemption from the rule you already believe. It is the strongest instance of the rule. Anything that runs code it downloaded should not also hold a credential that can write to your main branch.

You cannot stop the bot executing strangers' code — that is the job. You can make sure the machine executing it has nothing to give away.

Frequently asked questions

npm 12 turned install scripts off by default. Do I still need to sandbox my upgrade bot?

Yes, and the reason is specific. The npm 12 change is real and good: install scripts from dependencies now only run for packages listed in an `allowScripts` field, and the docs state that install commands "silently skip lifecycle scripts for any dependency that does not have a matching entry in `allowScripts`". But two things survive it. First, `npm approve-scripts` writes PINNED entries like `pkg@1.2.3` unless you pass `--no-allow-scripts-pin`, so a version bump falls out of the allowlist — meaning the bot's install skips the new version's script, a developer later approves the bump, and the script runs on their machine having never been tested. The bot's green check attests to a tree nobody will run. Second, and more fundamentally, blocking install hooks does nothing for the test phase. The `ignore-scripts` documentation notes that commands designed to run a particular script, `npm test` among them, still run it — and your test suite imports the upgraded package, which executes its module body. That is the package working as intended, not a lifecycle hook, and no script-blocking flag touches it. Install hooks were always the easy half.

Can I just run the bot with --ignore-scripts and skip all this?

You can, and for the resolve step you should — `npm install --package-lock-only --ignore-scripts` installs nothing and is genuinely low-risk. The problem is what you have left after that. Native modules do not build, so packages like `sharp` or `node-gyp`-based dependencies produce either a failed install or a silently incomplete one. More importantly you have not tested the tree a developer will actually get, which was the entire reason you wanted a bot. You end up with a machine that reliably tells you a lockfile resolves, which you could have determined without running anything. The honest version of the `--ignore-scripts` strategy is to split the work: use it for resolution, and then do the real install-and-test somewhere disposable where executing the script is fine because there is nothing in the guest to steal and nowhere useful to send it.

What exactly should the install-and-test guest be allowed to reach?

The registry, your source host if it needs to clone, and whatever services your test suite genuinely requires — nothing else, and ideally enforced outside the guest. On PandaStack some of this comes for free and some does not, so be precise about which is which. Each sandbox gets its own network namespace out of 16,384 pre-allocated /30 subnets, the entire link-local range including the cloud metadata endpoint is dropped at the top of the host firewall chain above any accept rule, and cross-tenant VM-to-VM traffic is dropped the same way. What is not provided is an outbound allowlist: egress is open by default, deliberately, because sandboxes exist to install packages and clone repositories. So if you want registry-only egress you build it — the robust way is a registry proxy you operate, with the guest given no other route, because `iptables` rules inside the guest are enforced by the same kernel the install script is running on and root in the guest can flush them. Also remember DNS: an IP allowlist that still permits arbitrary resolution leaves a stolen token encodable as a subdomain lookup.

How does the bot trust a test result produced inside an untrusted guest?

It does not trust the guest, and that is the point of the split. A result signed inside the guest is worthless — a compromised guest signs whatever it likes with whatever key you handed it. The trust anchors sit outside. The exit code comes from the host-side agent reporting on the process, not from the guest asserting its own success, and `exec_stream` is deliberate about the edge case: if the stream ends with no terminal exit event, it raises `ExecInterrupted` rather than returning 0, because the outcome is genuinely unknown. For a bot with auto-merge enabled, unknown rounding to green is the exact bug you are trying to avoid. Everything the guest produces — logs, the lockfile-diff report, test output — is treated as untrusted text: rendered into the PR body, never evaluated, never used to construct a shell command or a branch name. Treat it as a form submission from the internet, because that is what it is.

Does one microVM per upgrade candidate actually scale for a large monorepo?

The VM is not the cost. Creating a sandbox from a baked template is a snapshot restore rather than a boot: p50 179 ms, p99 203 ms, with the `/snapshot/load` step itself around 49 ms. The first create of a template before its snapshot exists is a real cold boot at roughly 3 seconds, and then the snapshot is captured and every create after that takes the fast path. Against an `npm ci` that takes tens of seconds and a test suite that takes minutes, two hundred milliseconds is noise. The honest framing is not that this makes your pipeline fast — it does not, your install and your tests dominate the wall clock and always will. It is that it removes the efficiency argument for reusing a machine, which was the only real reason anyone ran candidate 41 on the kernel candidate 40 just finished with. If you want to attack the install time itself, bake the toolchain and a warm module cache into the template so the common part is paid once at bake time; just size `--memory-mb` and `--size-mb` deliberately, because RAM is fixed by the snapshot and the rootfs default of 1024 MB is too small for a real toolchain.

Keep reading

Related posts

  • Sandbox Your AI Agent's Dependency Audit

    To audit a dependency tree you have to resolve and install it, and install scripts are arbitrary code execution you invited in. An AI agent running npm install on a stranger's lockfile is a supply-chain incident on a cron. Give each audit its own disposable Firecracker microVM.

  • Running npm install on untrusted code in a microVM

    `npm install` is arbitrary code execution with a friendly progress bar. A malicious postinstall can read your SSH keys and tokens the second you install. Here's how to make that safe.

  • Private npm and PyPI Mirrors in Front of a Sandbox Fleet

    A container fleet warms up. An ephemeral sandbox fleet cannot, because the whole point is that guest number four thousand is byte-identical to guest number one. That property is worth having and it means you will download left-pad four thousand times unless you do something about it.

  • How to Sandbox an Untrusted composer install

    `composer install` is a build step the way a stranger's USB stick is a file transfer. Scripts and plugins run arbitrary PHP before your app loads a single class — here's how to contain that.

  • Running a Plugin Marketplace Without Getting Owned

    Every plugin marketplace is a supply chain you don't control. Review passed v1.0; nobody re-reads v1.1. The fix isn't better review — it's running each plugin invocation in a Firecracker microVM that a malicious update can't escape.

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.