all posts

Giving an AI Agent a Language Server: The Index Is Code Execution

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and the most common complaint about AI coding agents is a navigation complaint wearing a competence costume. The agent renamed a function and missed four call sites. The agent invented a method signature that looked exactly right and does not exist. The agent grepped for a symbol, found it in a vendored copy, and edited that.

None of those are reasoning failures. They are information failures. A human engineer does not grep for a symbol and hope — they press one key, get every reference, and read the type. That information is real, it is already computed, and it comes out of a language server speaking the Language Server Protocol over stdio. Giving the agent the same thing is one of the highest-leverage changes you can make to a coding agent, and it costs you no model quality at all.

Then you read what a language server actually does on startup, and the shape of the problem changes.

rust-analyzer compiles and runs the repository's `build.rs`. gopls runs the go command, which reads the `toolchain` line in go.mod and will download and execute a different Go toolchain because the repository asked for one. tsserver reads tsconfig.json, which can extend a config out of node_modules and can name language-service plugins. clangd reads compile_commands.json, a file your build system produced by running your build, and with `--query-driver` it will execute the compiler binary that file names. A Python language server is gentler, right up until you notice that resolving imports requires a populated virtualenv, which means somebody ran `pip install`, which runs build backends.

So "index the repository so the agent can navigate it" turns out to be the same sentence as "execute the repository's build configuration." Indexing untrusted code is code execution with a reassuring name. The model-generated `rm -rf` everyone worries about is the least of it: the index already ran the repo's build scripts before the model saw a single token.

The short version: an agent with a language server is strictly better than an agent with grep, so give it one. But the indexer is a program the repository configures, so run it in a microVM that holds the checkout and nothing of yours — and because an index is expensive to build and worthless to throw away, index once, snapshot the warm server, and restore it per agent session instead of re-indexing.

What a language server gives an agent that grep cannot

Worth being concrete, because "semantic understanding" is the kind of phrase that survives review without meaning anything. Four specific capabilities change specific failure modes.

References before a rename. `textDocument/references` returns every use of the symbol under the cursor, resolved by the compiler front end rather than by string matching. That is the difference between renaming a method and renaming a method plus a same-named method on an unrelated type plus a string in a log line plus a comment. Grep cannot distinguish a call from a coincidence; this is the single most valuable request in the protocol for an agent that edits.

Diagnostics as a feedback loop. After an edit, the server pushes `textDocument/publishDiagnostics` — type errors, unresolved imports, borrow-check failures — in hundreds of milliseconds rather than the tens of seconds a full build takes. Give the agent that stream and its edit loop tightens from write-build-read-fix to write-read-fix. Agents are good at fixing a specific compiler error and bad at noticing they introduced one.

Workspace symbols instead of fuzzy file search. `workspace/symbol` answers "where is the thing named X" over the indexed symbol table. An agent asking that question with ripgrep gets a ranked list of files and burns context reading three wrong ones.

Hover types to stop the invention. `textDocument/hover` returns the resolved signature and doc comment. The failure mode where a model writes a plausible call to a function whose third parameter does not exist is a failure to look it up, and hover is the lookup. Add `textDocument/definition` to follow a type into a dependency and the agent reads the real source instead of recalling a version of the library from training data.

There is a nice symmetry here, which I have written about elsewhere: LSP is to editors roughly what the Model Context Protocol is to tools — see What is the Model Context Protocol (MCP)?. One wire format, N editors, M languages, instead of N times M integrations. Which means the agent-facing version of this is largely plumbing that already exists.

What a language server actually executes, per language

Here is the part that decides your architecture. I am going to be specific, because the defaults matter and because the advice "just turn the dangerous bits off" turns out to cost you the thing you were buying.

rust-analyzer: build scripts and proc macros, both on by default

This is the clearest case, and rust-analyzer is admirably honest about it. `rust-analyzer.cargo.buildScripts.enable` defaults to true, described in the manual as "Run build scripts (build.rs) for more precise code analysis." `rust-analyzer.procMacro.enable` also defaults to true, and its description states that it implies `cargo.buildScripts.enable`.

A `build.rs` is arbitrary Rust: compiled and executed on your machine, at index time, before any analysis happens. It is not an edge case, it is how Cargo discovers generated code — bindgen output, protobuf stubs, vendored C libraries, feature probes. Procedural macros are worse in kind: a proc macro is a compiler plugin, shipped as a dylib built from the dependency's source, loaded into a process and run on your code. Expanding `#[derive(Serialize)]` means executing serde_derive.

You can set both to false. Then derive-heavy code stops resolving — go-to-definition into a derived impl simply fails, hover on a macro-generated method returns nothing, and the diagnostics you wanted as a feedback loop fill with false positives about items that do exist. You have traded most of the index's value for a boundary. Hold that thought too, because it is the argument for the VM.

gopls: the module graph, and a toolchain the repository chose

gopls does not execute your package code to index it. It runs the go command — `go list` and friends — to resolve the module graph, and that is a network operation against a module proxy by design.

The sharper detail is toolchain selection. Standard Go distributions set `GOTOOLCHAIN=auto` in `$GOROOT/go.env`. With auto, the go command checks the main module's go.mod at startup: if there is a `toolchain` line naming a version newer than the bundled one, or a `go` line that is newer, it looks for that toolchain in PATH and otherwise downloads it — as the module `golang.org/toolchain` through GOPROXY — and runs it. Startup selection applies to every go command, `go list` included.

Read that as an indexing story. A repository you cloned contains two lines of text that cause your indexer to fetch a binary over the network and execute it. The download is checksum-verified and comes from a module proxy, so this is not a vulnerability — it is a documented feature working correctly. It is also not the behaviour most people picture when they say "index the repo." You can pin `GOTOOLCHAIN=local` and then fail to build repos that legitimately need a newer toolchain.

Separately: gopls exposes `go generate` as a code lens, which is user-triggered rather than index-time. Do not let an agent click it on a stranger's repository outside a sandbox, because `go generate` runs whatever the `//go:generate` comments say, and they say anything.

tsserver: a config that extends into node_modules, and the workspace server itself

tsconfig.json is configuration for a compiler, and it has reach. `extends` resolves a base config, which can be a package in node_modules — so the effective compiler configuration for the repository is partly code the repository downloaded. `compilerOptions.plugins` names language-service plugins: Node modules loaded into the tsserver process to change how the language service answers. tsserver's own CLI gates loading plugins out of a local node_modules behind `--allowLocalPluginLoads`, which defaults to false, and editors differ in whether and when they pass it — so check your client rather than assuming either answer.

The bigger surface is more mundane and almost universally accepted: editors offer to run the TypeScript the repository shipped. "Use workspace version" points your language service at `node_modules/typescript/lib/tsserver.js`, a JavaScript program that arrived with the checkout, and then runs it with your user's privileges. Editors have added workspace-trust prompts around this for exactly the reason you are now thinking of. An agent harness that silently prefers the workspace version has made that decision on your behalf and has no prompt.

And then type-checking a TypeScript monorepo is a build. Project references, composite projects, `.d.ts` emit into build info files — the index is downstream of real work.

clangd: compile_commands.json is build output, and query-driver is execution

clangd needs a compilation database, and the honest description of compile_commands.json is "a transcript of your build." You get it from CMake's `CMAKE_EXPORT_COMPILE_COMMANDS`, or from Bazel tooling, or from `bear` intercepting a real `make`. For a Makefile project the usual answer is: run the build, record what it ran. So before clangd indexes anything, the repository's build system has executed.

Then there is `--query-driver`, which clangd's man page describes as a comma-separated list of globs for white-listing gcc-compatible drivers that are safe to execute, used to extract system includes. The phrase "safe to execute" is doing real work in that sentence. clangd runs the compiler driver named in the database to ask it where its system headers live, which is the correct way to support a cross-compiler — and it is off by default and allowlist-shaped precisely because the database names the binary. A hostile repository supplies both the database and, if you let it, the driver path.

Python: the server is static, the environment is not

pyright and its relatives are the well-behaved end of this. They are static analysers: they do not import your modules to find out what is in them. But a Python language server is only useful if it can resolve imports, and resolving imports means pointing it at an interpreter and a populated site-packages. Which means somebody ran `pip install -r requirements.txt`, and installing a source distribution executes its build backend — `setup.py`, or a PEP 517 hook — as the install step, by design.

So Python moves the execution one step earlier and makes it explicit. That is genuinely better, and it is not a boundary; it is a different place to stand when the code runs. I have written the longer version of this at Running npm install in a MicroVM: Dependency Installation Is Arbitrary Code Execution. python-lsp-server adds a second surface: its plugin ecosystem loads from the same environment, so a pylsp plugin pinned in the repo's requirements is code in your language server's process.

What each language server executes at index time
ServerRuns repo code at index timeOpt-outWhat the opt-out costs
rust-analyzerYes — build.rs, and proc macros as compiled dylibsbuildScripts.enable / procMacro.enable, both default trueDerive and macro-generated items stop resolving; diagnostics fill with false positives
goplsNot package code — but the go command may download and run a toolchain go.mod namesGOTOOLCHAIN=localRepos that legitimately require a newer toolchain fail to resolve
tsserverLoads tsconfig extends from node_modules; plugins and workspace tsserver.js are Node codePin the editor's own TypeScript; --allowLocalPluginLoads defaults falseVersion skew against the repo's compiler; plugin-provided diagnostics disappear
clangdIndirectly — the database is build output; --query-driver executes the named driverOmit --query-driver (the default)Wrong or missing system include paths for cross-compilers
Python (pyright, pylsp)Server is static; populating the venv runs build backendsInstall from wheels only, no build isolation escapePackages with no wheel for your platform are unresolvable

The architecture: one microVM per repository

The arrangement that follows is simple enough to describe in one sentence, and the sentence is the whole design. One Firecracker microVM per repository, or per agent session: the checkout lives inside it, the language server runs inside it, the toolchain installs inside it, and the agent lives outside and receives only structured answers.

The important word is outside. The agent process is the thing holding your model API key, your git credentials, the user's session, and the orchestration logic that decides what to do next. It should not share an address space, a filesystem, or a kernel with a program whose configuration file was written by a stranger. So LSP crosses a bounded channel — a request goes in as a method name and parameters, a response comes back as JSON — and nothing else does.

Which is where the usual mistake happens, so let me be blunt about it. The boundary is not `rootUri`. The `initialize` request's `rootUri` and `workspaceFolders` tell the server where the project is; they are not a sandbox, they are a hint about scope. A language server asked to resolve a path outside the workspace will generally resolve it, because that is how you get go-to-definition into the standard library. The boundary is the kernel, the guest's own, which exists for this one repository and will stop existing shortly.

On PandaStack that means a per-sandbox network namespace with its own veth pair and tap device, from a pool of 16,384 pre-built slots; a guest kernel the indexer's syscalls go to instead of yours; a device model of a handful of virtio devices rather than four decades of emulated hardware; and a teardown that is a deletion rather than a cleanup routine that can go wrong. The general containers-versus-microVMs argument is covered at Agent runtimes: containers vs microVMs. The indexing-specific version is one line: a container's namespaces are enforced by the same kernel every other tenant calls into, and the program you are about to run is a compiler plugin the repository compiled.

One real consequence worth stating early: a microVM-per-repo design means the agent's filesystem tools and its LSP tools must address the same filesystem. If the agent edits a file through a host-side tool and the language server is looking at a guest-side copy, every answer is stale and the agent will fight the index. Put the edit path and the query path on the same side of the boundary — both in the guest — and let the agent hold neither.

Why the session is long-lived, and what a snapshot does about it

An LSP index is the most awkward kind of state: expensive to build, cheap to query, worthless the moment you discard it. A per-request VM is the right answer for parsing a document or compiling a snippet — I have argued exactly that at Self-Hosting a Compiler Explorer: Running Strangers' Compilers Safely — and it is the wrong answer here, because the cost you are trying to avoid is the setup, not the query.

Cold-index a real Rust workspace or a TypeScript monorepo and you are waiting minutes: dependency resolution, build scripts, macro expansion, type-checking enough of the graph to answer questions about it. Do that per agent request and your agent is a latency generator. Do it once per repository and keep the result, and every request after the first is tens of milliseconds.

This is where a Firecracker snapshot stops being a performance trick and becomes the architecture. A warm, indexed language server is exactly the state a full-memory snapshot was built to preserve. Index once; `snapshot()` captures guest memory and disk; every later agent session is created from that snapshot, which on PandaStack is a restore on every create with no warm pool at all: p50 179 ms, p99 203 ms end to end. The process table comes back with gopls in it, already indexed, because the memory came back. A cold boot of the same template is around 3 seconds and an index on top of that is minutes, so the snapshot is not a 10% improvement, it is a different interaction model.

Which changes what "long-lived" has to mean. The naive version is `persistent=True`, which takes the sandbox out of the reaper's scan entirely — and that is a capacity plan with no second half, because then nothing ever puts it down. The better arrangement uses the reaper deliberately, and it depends on a detail worth getting right: `ttl_seconds` is an idle TTL, not a wall clock. The reaper compares it against time since last activity, so a sandbox an agent is actively querying does not expire, and one nobody has touched does.

What counts as activity is narrow on purpose. Genuine use of the guest — `exec`, filesystem byte reads and writes, PTY and SSH sessions, fork, snapshot, and LSP traffic — bumps the idle clock and will transparently wake a hibernated sandbox. Read-only inspection calls do neither: status, lifecycle, metrics, logs, events, `fs/stat` and `fs/dir` are explicitly excluded, because observing a sandbox must not keep it alive. That exclusion is a bug fix rather than a nicety. A dashboard polling sandbox status every thirty seconds reset the idle clock faster than any TTL could elapse, so nothing was ever reaped and idle seconds sat at zero forever.

So the lifecycle that actually costs you nothing at rest is three states, not two. Index the repository, `snapshot()` the warm server, and then let the idle TTL delete the VM. The snapshot survives, because the idle reaper deletes through a path that deliberately does not cascade into the sandbox's snapshots — a snapshot is a durable artifact meant to outlive its sandbox, which is the whole point of being able to restore one. The next agent session creates from that snapshot in under 200 ms with the index already warm. You keep the index and stop paying for an idle 4 GiB guest, and an agent that is mid-task keeps its own session alive simply by using it.

Two mechanical facts make it cheap rather than merely fast. Memory restore is `MAP_PRIVATE` on the snapshot file, so the kernel copies pages on write and sibling VMs share the clean ones; disk is XFS reflink, so a clone is a metadata operation. The long version is at Copy-on-Write Memory: Why Forking a VM's RAM Is Cheap. And the guest clock is re-synced on restore, resume and wake — added after a frozen-clock TLS incident — so a restored indexer is not sitting at a stale timestamp wondering why certificates look wrong.

Building the warm index, with the Python SDK

The flow below tars a checkout, writes it in, installs the toolchain, warms the server, and snapshots. Two SDK details that are easy to get wrong and are both in here: `filesystem.upload()` takes a single file, not a directory, so tar first; and one-shot `exec()` does not enforce `timeout_seconds`, so the real bound on a long install is `timeout(1)` inside the guest -- `exec_stream` only widens the client's HTTP timeout so a long install is not cut off at 30 seconds, it does not bound the command either.

# Build a warm, indexed language-server VM once, then snapshot it.
# Go + gopls here because the `base` template already carries Go via mise;
# the rust-analyzer version is the same shape with a heavier install step.

import io, pathlib, tarfile
from pandastack import Sandbox
from pandastack.exceptions import CommandFailed

REPO = "/srv/checkouts/acme-api"   # a clean checkout at a known commit

# 1. Tar the checkout on the client. filesystem.upload() is single files only
#    -- it is write(remote, Path(local).read_bytes()) -- so handing it a
#    directory raises IsADirectoryError. Tar first, every time.
buf = io.BytesIO()
with tarfile.open(fileobj=buf, mode="w:gz") as tf:
    tf.add(REPO, arcname="repo")

sbx = Sandbox.create(
    template="base",          # 4 GiB / 8 vCPU, baked -- see the sizing section
    ttl_seconds=30 * 60,      # IDLE ttl, not a wall clock: exec and LSP traffic
                              # reset it, status GETs do not. Deliberately NOT
                              # persistent=True -- letting the reaper delete
                              # this VM is the plan, because the snapshot it
                              # takes with it is the thing we actually want.
    metadata={"repo": "acme/api", "commit": "9f31c0a", "role": "lsp-index"},
)

sbx.filesystem.write("/tmp/repo.tgz", buf.getvalue())
sbx.exec("mkdir -p /work && tar -xzf /tmp/repo.tgz -C /work --strip-components=1",
         check=True)

# 2. Toolchain + first index. THIS is the step that executes the repository's
#    build configuration: `mise install` honours .tool-versions, and the go
#    command honours the `toolchain` line in go.mod -- which means it may
#    download and run a Go toolchain this repo chose. That is the whole reason
#    we are inside a VM. Bound it in-guest with timeout(1): one-shot exec()
#    does not enforce timeout_seconds, and exec_stream only widens the HTTP timeout, so stream it and fence it in-guest.
setup = (
    "cd /work && timeout --kill-after=10s 900 sh -c '"
    "export MISE_DATA_DIR=/opt/mise MISE_CONFIG_DIR=/opt/mise "
    "PATH=/opt/mise/shims:$PATH; "
    "mise install && mise reshim && "
    "go install golang.org/x/tools/gopls@latest && "
    "go list ./... > /work/.pkgs'"
)
if sbx.exec_stream(setup, on_stdout=print, on_stderr=print) != 0:
    raise CommandFailed("toolchain setup failed; see stream above")

# 3. Warm the index. A snapshot of a language server that has not finished
#    indexing is a snapshot of a progress bar. ask.py (next block) sends a
#    real request and blocks until the server answers it.
sbx.exec("mkdir -p /work/.lsp", check=True)
sbx.filesystem.write("/work/.lsp/ask.py", pathlib.Path("ask.py").read_text())
sbx.exec_stream(
    "cd /work && timeout --kill-after=10s 900 python3 .lsp/ask.py --warm",
    on_stdout=print,
)

# 4. Freeze the warm index. snapshot() is synchronous and writes the full guest
#    memory image, so budget 30-60s on a multi-GB guest. You pay this once per
#    commit you care about, not once per agent.
snap = sbx.snapshot()
print("warm index:", snap)

# 5. Now let the idle TTL reap THIS VM. The reaper's delete path does not
#    cascade into its snapshots, so the warm index outlives the VM that built
#    it -- that is what makes "long-lived session" affordable.

# 6. Every later agent session restores that snapshot instead of re-indexing:
#    p50 179 ms, p99 203 ms, with gopls already running and already indexed.
session = Sandbox.create(from_snapshot=snap, metadata={"repo": "acme/api"})
print(session.exec("pgrep -a gopls").stdout)
session.kill()            # teardown is kill(); there is no sbx.delete()

Speaking LSP inside the guest, returning a value

The second piece is the client, and the point of running it inside the guest is that the agent never holds a connection to the language server — it holds a JSON array of file-and-line pairs. The framing is the part people get wrong: LSP's base protocol is HTTP-ish headers, `Content-Length` is a byte count of the UTF-8 body, and the body starts after a blank line.

#!/usr/bin/env python3
# /work/.lsp/ask.py -- a one-shot LSP client. Runs INSIDE the microVM, speaks
# JSON-RPC over gopls's stdin/stdout, prints one JSON value and exits. The
# agent outside receives a value, never a live socket to the indexer.
import json, os, pathlib, subprocess, sys, urllib.parse

ROOT = "/work"
TARGET = "internal/billing/meter.go"

def frame(obj):
    # LSP base protocol: "Content-Length: <n>\r\n\r\n<body>". n is the length
    # of the UTF-8 encoded body in BYTES. Count bytes, not characters -- one
    # non-ASCII identifier and a len(str) client desyncs for the rest of the
    # session, which is a fun afternoon.
    body = json.dumps(obj).encode("utf-8")
    return b"Content-Length: %d\r\n\r\n%s" % (len(body), body)

def read_message(out):
    headers = {}
    while True:
        line = out.readline()
        if not line:
            raise EOFError("language server closed the pipe")
        line = line.strip()
        if not line:                       # blank line ends the header block
            break
        k, _, v = line.decode("ascii").partition(":")
        headers[k.strip().lower()] = v.strip()
    return json.loads(out.read(int(headers["content-length"])))

# Plain `gopls` serves LSP over stdio -- the same way your editor launches it.
srv = subprocess.Popen(["gopls"], cwd=ROOT,
                       stdin=subprocess.PIPE, stdout=subprocess.PIPE)

def send(obj):
    srv.stdin.write(frame(obj)); srv.stdin.flush()

def await_id(want):
    # Everything that is not our response is progress and diagnostics noise:
    # window/logMessage, $/progress, textDocument/publishDiagnostics. A real
    # agent client KEEPS the diagnostics -- they are the post-edit feedback
    # loop. This one-shot drops them.
    while True:
        msg = read_message(srv.stdout)
        if "method" not in msg and msg.get("id") == want:
            return msg

root_uri = pathlib.Path(ROOT).as_uri()
send({"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {
    "processId": os.getpid(),
    "rootUri": root_uri,
    "workspaceFolders": [{"uri": root_uri, "name": "repo"}],
    "capabilities": {"textDocument": {"references": {"dynamicRegistration": False}},
                     "workspace": {"workspaceFolders": True}},
}})
await_id(1)
send({"jsonrpc": "2.0", "method": "initialized", "params": {}})

path = pathlib.Path(ROOT, TARGET)
uri = path.as_uri()
send({"jsonrpc": "2.0", "method": "textDocument/didOpen", "params": {
    "textDocument": {"uri": uri, "languageId": "go", "version": 1,
                     "text": path.read_text()},
}})

if "--warm" in sys.argv:
    # The warm pass exists only to make the server do its indexing work before
    # we snapshot it. gopls answers this once the workspace is loaded.
    send({"jsonrpc": "2.0", "id": 2, "method": "workspace/symbol",
          "params": {"query": "Meter"}})
    print(len(await_id(2).get("result") or []), "workspace symbols")
    srv.stdin.close(); srv.wait(timeout=10); sys.exit(0)

# position.line and .character are ZERO-BASED, and .character is a UTF-16
# code-unit offset unless you negotiated positionEncoding. Both bite once.
send({"jsonrpc": "2.0", "id": 3, "method": "textDocument/references", "params": {
    "textDocument": {"uri": uri},
    "position": {"line": 73, "character": 17},
    "context": {"includeDeclaration": False},
}})

refs = await_id(3).get("result") or []
out = []
for r in refs:
    p = urllib.parse.unquote(urllib.parse.urlparse(r["uri"]).path)
    out.append({"file": os.path.relpath(p, ROOT),
                "line": r["range"]["start"]["line"] + 1})
print(json.dumps(out, indent=2))          # the only thing that leaves the VM

srv.stdin.close(); srv.wait(timeout=10)

Three details in there earn their comments. Positions are zero-based, and `character` is a UTF-16 code-unit offset unless you negotiate `positionEncoding` — the two bugs every LSP client author writes exactly once. Most servers want a `textDocument/didOpen` before a position-based request, so send one even when the file is already on disk. And the response stream is full of `$/progress` and `publishDiagnostics` messages interleaved with your answer: this one-shot discards them, but a real agent client keeps the diagnostics, because that is the feedback loop from the first section.

Fanning out: one warm index, many agents

If you are running several agents against the same repository — a reviewer, a test-writer, and an implementer, or a best-of-N over the same task — you want one index and N VMs. This is where the distinction between the two fork calls matters more than anything else in this post, and where I have seen people lose an afternoon.

`fork_tree(n)` snapshots the parent once and boots n children from that snapshot, so each child inherits the parent's memory: the running language server, its loaded index, its caches. `fork()` clones the disk only and the child cold-boots — you keep the checkout and the installed toolchain, and you index again from nothing. For a warm language server, `fork_tree` is the call. The cap is 16 children per call, and above it you get an error rather than a clamp, so a wider fan-out is a breadth-first loop. The fuller treatment is at Snapshot-Restore vs Fork: When to Use Which.

# Give N agents the same warm index, without N cold indexes.
#
# fork_tree() snapshots the parent ONCE and boots the children from that
# snapshot, so each child inherits the parent's MEMORY: the running gopls
# process, its loaded index, its caches. The parent is untouched
# (pause -> snapshot -> resume).
#
# fork() is the OTHER call and it is not this one. fork() clones the DISK only
# and the child COLD-BOOTS. You get the checkout and the installed toolchain,
# and then you get to index again from nothing. For a warm language server,
# fork_tree is the only call that does what you want.
#
# Hard cap: 16 children per fork_tree() call (maxTreeChildren -- more is an
# error, not a clamp). For a wider fan-out, grow the tree breadth-first.

def fan_out(parent, total):
    frontier, kids = [parent], []
    while total > 0 and frontier:
        node = frontier.pop(0)
        batch = node.fork_tree(min(16, total))
        kids += batch
        frontier += batch
        total -= len(batch)
    return kids

# One question per worker: each is a symbol the agent wants references for.
QUESTIONS = ["Meter.Flush", "Meter.Record", "billing.Tier", ...]

workers = fan_out(sbx, 40)        # 40 agents, one warm index, three rounds

try:
    for w, question in zip(workers, QUESTIONS):
        # Each worker has its own kernel, its own page tables and its own copy
        # of the index -- copy-on-write at first, diverging as it edits.
        print(w.id, w.exec(f"cd /work && python3 .lsp/ask.py {question}").stdout)
finally:
    for w in workers:
        w.kill()

Resource discipline: rust-analyzer will eat what you gave it

Language servers are memory-shaped workloads. rust-analyzer on a large workspace holds the crate graph, the expanded macro output, and its salsa query caches, and on a big monorepo it is routinely the largest resident thing in the machine. Plan for it to meet your ceiling rather than sit comfortably below it.

Which brings up the one piece of PandaStack sizing you cannot work around with an argument: Firecracker cannot change vCPU or RAM at snapshot restore. A sandbox's size comes from the template's baked `meta.json`, and the agent overrides whatever `cpu=` and `memory_mb=` you passed to `create()` to match the snapshot — it logs the override rather than silently lying, because the alternative is a wrong number flowing into the API response, the database row and the billing event. `base` is baked at 4 GiB and 8 vCPU. So "give the indexer more memory" is not an argument you pass, it is a template you bake. Check the current set on templates before you design around a size.

A related trap specific to TypeScript: the `base` image sets `NODE_OPTIONS=--max-old-space-size=3072` in its environment, so tsserver's own V8 heap is capped below the VM's 4 GiB. On a large monorepo you will hit a Node heap failure before you hit a guest OOM, and the fix is a Node flag, not a bigger VM.

The 8 vCPU are burst capacity, shared under contention by cgroup `cpu.weight`. That suits indexing well — a burst of parallel type-checking, then near-idle. It also suits the billing model: CPU is metered on CPU-seconds actually burned while memory is committed GiB-hours, so an index VM left running overnight is paying for its RAM and almost nothing for its cores. At the published rate of $0.0162 per GiB-hour, a 4 GiB VM's memory is about 6.5 cents an hour whether or not anyone queries it. That is cheap for one and an actual line item for fifty, which is the arithmetic behind the lifecycle above: snapshot the warm index and let the idle TTL delete the guest, and the overnight cost becomes stored bytes instead of committed GiB-hours. Current rates are on pricing.

Now the honest part about memory pressure, because there is a controller involved and you should know its rules. A pressure ladder runs on each host. Under elevated host pressure it writes cgroup v2 `memory.high` on the coldest squeezable VMs, at 70% of their current residency with a 256 MiB floor, so the cold tail gets compressed rather than the VM getting killed. It never writes `memory.max`. It never touches database-class VMs. And it only squeezes a VM that is demonstrably not working: under half a core of CPU and untouched for at least a minute.

So an actively indexing VM is not squeezed — that gate exists for exactly this case. A warm index that nobody has queried for an hour is a legitimate squeeze target, and the consequence is honest and slightly annoying: the first query after a quiet stretch faults pages back in and is slower than the hundredth. Under sustained critical pressure the bottom rung freezes the coldest idle plain sandbox entirely, via the hibernate path, and it wakes on touch. If your product promises a fast first answer after idle, send a cheap keepalive — and note that it has to be real guest use, an `exec` or an actual LSP request, because a status GET bumps nothing by design.

Four arrangements, compared on what actually bites

Agent code-navigation arrangements compared
ArrangementAgent accuracyWho runs the repo's build configBlast radius of a hostile repo
Agent with grep onlyPoor — misses references, invents signaturesNobodySmall, and you paid for it in correctness
Language server in the agent's own processGoodYour agent process, with your credentialsYour machine, your tokens, your git remote
Language server in a containerGoodA container sharing your kernelOne kernel escape or one cgroup OOM from the neighbours
Language server in a microVM per repoGood — leave proc macros onA guest kernel that exists for this repoOne VM, deleted on teardown; egress is still yours to fence

The row that matters is the last cell of the fourth row. A microVM is what lets you leave build scripts and proc-macro expansion enabled, which is what gives you the accurate index, which was the entire point of the exercise. Disabling them for safety in a shared process is the trade you make when you have no boundary. Having a boundary means you do not have to make it.

The honest limits

  • The indexing cost does not disappear, it moves. A first cold index of a large monorepo is minutes, and no snapshot helps the first time — somebody pays it. If you have thousands of repositories and a long tail that is queried once a month, you are paying cold-index latency on most requests and the snapshot machinery buys you nothing on those.
  • Taking the snapshot is not free either. `snapshot()` is synchronous and writes the full guest memory image to disk; budget 30 to 60 seconds on a multi-GB guest. That is fine once per commit you care about and unacceptable per request.
  • A warm-index snapshot is durable, not permanent. Once its source sandbox is gone — which, in the lifecycle above, is the plan — it is an orphan, and orphaned snapshots are swept after a grace period that defaults to seven days. For a repository an agent touches weekly that is invisible. For a long tail of repositories queried once a month it means the snapshot is never there when you want it and you are paying the cold index every time, so measure your access pattern before you build around the warm path.
  • You need a staleness policy, and it is yours to write. A warm-index snapshot is a photograph of one commit. Source-only changes are usually absorbed by the running server in seconds. A change to Cargo.lock, package-lock.json, go.mod or requirements.txt invalidates the resolution the index was built on, and the re-resolve can be minutes. The rule I would start with: re-bake on any manifest or lockfile change, let the live server absorb source edits, and put the commit SHA in the sandbox metadata so you can tell what you restored.
  • A language server is not a security boundary, and neither is the agent's prompt. `rootUri` is scope advice. Telling the model not to read outside the workspace is a preference, not a control. The only thing in this design doing enforcement is the guest kernel, which is why it has to be the thing you rely on.
  • Egress is open by default. The host's FORWARD chain drops pool-to-pool traffic, so no sandbox reaches another sandbox's subnet, and drops all of 169.254.0.0/16, so cloud metadata is unreachable — but your own VPC, databases and internal APIs are not fenced, and that remaining work is yours. And you cannot simply deny the network to an indexer: a build.rs that fetches at build time is a legitimate, common pattern, and gopls downloading a toolchain is the documented default. Deny egress as a class and you break real repositories. Allowlist a module proxy and a package registry, and accept that an allowlisted proxy is an exfiltration channel with a cache in front of it.
  • The guest kernel is 5.10. Most toolchains do not care. Some newer tooling does, and you should find out which before you build a product on it rather than after.
  • Turning proc-macro expansion and build scripts off gives you a measurably worse index, and I want to say that plainly because the opposite advice is everywhere. In a derive-heavy Rust codebase you lose go-to-definition into generated impls, hover on generated methods, and the accuracy of the diagnostics stream. An agent working from that index will confidently tell you a method does not exist. If you cannot run them, run a worse index and know it is worse; do not pretend the trade is free.
  • Per-repo VMs mean capacity can run out. Memory admission on the fleet is working-set based rather than committed, which helps a lot with density, but a create that is refused for capacity is a real 503 your orchestrator has to handle — and at 3am a capacity incident and an attack look remarkably similar on a dashboard.

The summary

An agent that navigates code by grepping and guessing produces exactly the failures people complain about: missed references, invented signatures, edits in a vendored copy. A language server fixes that, and it is not a research problem — it is a protocol with mature implementations for every language you care about, and the agent-facing side is mostly plumbing.

The catch is that a language server is a program the repository configures. rust-analyzer compiles and runs `build.rs` and loads proc-macro dylibs, both enabled by default and for good reason. The go command will download and execute the toolchain go.mod names. tsconfig.json extends into node_modules and names plugins, and your editor offers to run the repo's own tsserver. clangd's compilation database is build output, and `--query-driver` executes the driver that database names. Python keeps the server static and moves the execution into `pip install`.

Which means the useful index and the dangerous index are the same index. So put it somewhere that can afford to run it: one microVM per repository, the checkout and the server inside, the agent outside holding structured answers, the boundary being a guest kernel rather than a `rootUri`. Index once, snapshot the warm server, restore at p50 179 ms per session, `fork_tree` when several agents need the same warm state. Then write down your staleness policy, fence your own VPC, and leave the proc macros on — because having a boundary is what lets you.

Frequently asked questions

Is giving an AI agent a language server actually better than letting it use ripgrep?

Yes, and the improvement is specific rather than vague. Four requests change four concrete failure modes. `textDocument/references` returns every real use of a symbol resolved by the compiler front end, which is the difference between renaming a method correctly and renaming it plus an unrelated same-named method plus a string in a log line — grep cannot tell a call from a coincidence. `textDocument/publishDiagnostics` gives the agent type errors and unresolved imports in hundreds of milliseconds instead of a full build, which tightens the edit loop from write-build-read-fix to write-read-fix; agents are good at fixing a named compiler error and bad at noticing they caused one. `workspace/symbol` answers "where is the thing called X" from an indexed symbol table instead of making the agent read three wrong files and burn context. And `textDocument/hover` plus `textDocument/definition` return the resolved signature and the real source, which is the direct fix for a model inventing a plausible third parameter. None of this requires a better model. It is the same model with the information a human engineer gets from one keystroke, and the main reason it is not standard yet is that running the indexer safely is a genuine infrastructure problem rather than a prompt problem.

Does running a language server really execute code from the repository?

For several of them, yes, by design and by default. rust-analyzer is the clearest: `rust-analyzer.cargo.buildScripts.enable` defaults to true and its own documentation describes it as running `build.rs` for more precise analysis, while `rust-analyzer.procMacro.enable` also defaults to true and implies build scripts. A `build.rs` is arbitrary Rust compiled and run at index time, and a proc macro is a compiler plugin built from a dependency's source, loaded as a dylib and run on your code. gopls does not execute your package code, but the go command it drives reads the `toolchain` line in go.mod and, with the default `GOTOOLCHAIN=auto`, will download that toolchain through GOPROXY and run it — startup selection applies to `go list` like everything else. clangd's compilation database is literally the output of running your build, and `--query-driver` executes the compiler driver that database names, which is why it is opt-in and allowlist-shaped. tsserver loads tsconfig `extends` out of node_modules, can load language-service plugins as Node modules, and most editors offer to run the repository's own `tsserver.js`. Python is the exception that proves the rule: pyright is a static analyser that does not import your modules, but it is only useful against a populated virtualenv, and populating one runs source-distribution build backends.

Should the sandbox be ephemeral per request, or long-lived?

Long-lived, and this is one of the clearest cases in the whole sandbox design space. The cost you are managing is the index, not the query. A cold index of a substantial Rust workspace or TypeScript monorepo is minutes of dependency resolution, build scripts, macro expansion and type-checking; a reference lookup against a warm index is tens of milliseconds. A VM per request pays the expensive part every time and throws away the result, which is the worst possible arrangement. But the answer is not a VM that lives forever either. Use a snapshot and you get both. Index once, call `snapshot()` to capture guest memory and disk, and create every later agent session from that snapshot — a restore on every create, p50 179 ms and p99 203 ms end to end, with the language server process already running and already indexed because the memory came back with it. Then let the idle TTL delete the indexing VM itself: `ttl_seconds` on PandaStack is an idle TTL rather than a wall clock, so an agent actively querying the server keeps its own session alive while an untouched one expires, and the reaper's delete path deliberately does not cascade into the sandbox's snapshots. So you keep the warm index and stop paying for an idle 4 GiB guest. That is the arrangement I would build: per-session isolation, warm-index latency, near-zero cost at rest. The operational cost you accept in exchange is a staleness policy, because the snapshot is a photograph of one commit and a lockfile change invalidates the resolution it was built on.

What is the difference between fork() and fork_tree() for this, and why does it matter?

It matters because only one of them gives you a warm language server, and picking the wrong one silently costs you the entire benefit. `fork()` clones the disk only and the child cold-boots. You get the checkout and the installed toolchain on a copy-on-write disk, and then the child starts from nothing: no running server, no index, so it indexes again. That is useful when disk state is what you wanted, and it is not what you wanted here. `fork_tree(n)` snapshots the parent once and boots n children from that snapshot, so every child inherits the parent's memory — the running language server, the loaded index, the caches — with copy-on-write page sharing, diverging only as each child writes. The parent is not modified; the sequence is pause, snapshot, resume. The cap is 16 children per call, enforced as a hard error rather than a clamp, so a wider fan-out is a breadth-first loop that pops a node off the frontier and forks up to 16 more from it. The practical shape for a best-of-N agent run is one warm index VM as the root, a fan-out to however many workers you need, and a `kill()` on each worker at the end — teardown is `kill()`, there is no `delete()`.

Can I just disable build scripts and proc macros instead of running a microVM?

You can, and you should know exactly what it costs, because the advice to do it is everywhere and the cost is rarely stated. In rust-analyzer, setting `cargo.buildScripts.enable` and `procMacro.enable` to false means macro-generated code stops resolving. Go-to-definition into a derived impl fails. Hover on a method a macro generated returns nothing. Code behind `cfg` flags a build script would have set is invisible. Worst for an agent, the diagnostics stream fills with false positives about items that do exist, and an agent reading that stream will confidently report that a method is missing and then write a workaround for a problem you do not have. In a serde-heavy or async-trait-heavy codebase that is most of the interesting surface. The equivalent trades in other ecosystems are smaller but real: `GOTOOLCHAIN=local` breaks repositories that legitimately need a newer toolchain, and omitting clangd's `--query-driver` gives you wrong system include paths for cross-compilation. So the honest framing is that disabling these is what you do when you have no boundary and have to choose between accuracy and safety. A microVM removes the choice: the indexer runs the repository's build configuration, because that is how you get a correct index, and it does so in a kernel that exists for that one repository and gets deleted afterwards.

Keep reading

Related posts

  • Running Customer Git Hooks in Isolated microVMs

    A git hook is a shell script your user wrote that your server agreed to run. The exploit is not the scary part — the scary part is a policy hook that greps a 4 GB monorepo on every push and takes the whole push queue with it.

  • Per-Tenant Search Indexing in Isolated microVMs

    Multi-tenant search is a noisy-neighbor and data-leak minefield — one tenant's reindex starves everyone's queries, and a shared index is a leak waiting to happen. Give each tenant its own microVM.

  • Giving AI Agents Persistent Memory & State via microVM Snapshots

    When people say they want their agent to 'remember,' they usually mean two different things: facts (a vector DB's job) and the machine it was working on — its filesystem, installed tools, scratch files, warm caches. A microVM snapshot captures that machine, exactly, and hands it back later.

  • What Is an AI Agent Sandbox?

    An LLM emits actions — shell commands, code, tool calls — that no human vetted before they ran. An agent sandbox is the isolated, disposable place those actions run so a mistake or an attack can't reach your host, your data, or another tenant.

  • Sandboxing IDE Agent Extensions

    Your Approve button is doing more security work than any control you actually designed. Here is how to move the agent's shell off every developer's laptop without giving up the part that made it good.

More in AI agent sandboxes · See AI agent sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.