A Linux Desktop per Agent: Xvfb, VNC, and One MicroVM per Session
Most “give the agent a browser” infrastructure is a headless browser and a CDP socket, and for a large class of work that is exactly right. Then a task arrives that involves a file picker, or a desktop application, or an OS-level print dialog, or a legacy internal tool whose vendor shipped a Java client in 2009 and an invoice in 2024, and the headless browser has nowhere to put any of it. What that task needs is not a browser. It is a computer: a display, a window manager, a mouse pointer that exists, and a screen you can photograph.
I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This post is the practical build for a disposable graphical Linux session — Xvfb, a window manager, x11vnc, noVNC — and an argument for why that session belongs in its own virtual machine rather than its own container. It also spends real length on the parts that do not work, because there is no GPU in here and that has consequences better learned before you design around them.
Why a headless browser is not a desktop
The distinction people blur is between a browser with no window and a machine with no screen. A headless browser is the first thing: one process, rendering offscreen, driven over a debugging protocol. A virtual display is the second: a real X server that happens to draw into ordinary memory, on which fully headed applications open fully real windows that nobody is looking at. The second one is strictly more general, and the difference shows up in five places.
- Native applications. A desktop email client, a CAD viewer, an installer, a vendor's thick client. There is no headless mode to ask for; the program wants an X connection and exits without one.
- OS-mediated dialogs. The file picker, the print dialog, the certificate prompt, the camera-permission chrome. The toolkit renders these on a display; a headless browser either suppresses them or routes them through a protocol call with no visual form.
- Human-in-the-loop RPA. The moment a workflow needs a person to look at a half-finished session and finish it — a two-factor prompt, a judgement call — you need a screen a human can attach to, not a log of CDP messages.
- Agents that look at pixels. A model handed a screenshot and emitting mouse coordinates operates on the rendered desktop, not the DOM. That is the whole point: it works on applications that have no DOM.
- Bugs that only exist in a window. Focus behaviour, drag-and-drop, anything involving window geometry. A headless suite cannot see these, which is why they reach production.
The cost of generality is that you now own a desktop: fonts, audio, a clipboard, a download directory, a window manager, and a display server whose resolution you have to decide on. All tractable, none free, and most of it fails in ways that look like the agent being stupid rather than the environment being wrong.
The bring-up script, with the gotchas inline
Here is the whole stack as one script. Xvfb provides the display, openbox provides focus, x11vnc exports the framebuffer over RFB on loopback only, and websockify wraps that in a WebSocket so a browser can connect. The ordering matters and so does the fact that it is a supervision loop rather than four backgrounded commands — the state where x11vnc has died and Xvfb has not is genuinely miserable to debug, because everything inside the guest reports healthy and the client simply hangs.
#!/usr/bin/env bash
# /usr/local/bin/desktop-up -- bring up a disposable graphical session inside
# the guest. Every line with a comment above it is a line someone (me) got
# wrong first. Run it from the template's init, not from an interactive shell.
set -euo pipefail
export DISPLAY=${DISPLAY:-:99}
GEOM=${GEOM:-1280x800x24} # ONE geometry. See the comment below.
VNC_PORT=5900 # loopback only
NOVNC_PORT=6080 # the only port we ever expose
# ------------------------------------------------------------------ 1. the X server
# Xvfb is a real X server that renders into ordinary memory instead of a card.
# Clients cannot tell the difference; there is simply nobody looking at it.
#
# GEOMETRY IS FIXED AT STARTUP. Xvfb creates its screens when it starts, and
# runtime resizing through RANDR is limited and version-dependent. If your
# agent computes click coordinates from a screenshot, a resize mid-session is
# a correctness bug, not a cosmetic one -- so pick one geometry, bake it into
# the template, and make the agent assert it on every connect.
#
# -nolisten tcp: no X over the network. The agent talks to the display through
# the unix socket in the same guest. -ac then only loosens access control for
# local clients, which inside a single-tenant microVM is the whole population.
Xvfb "$DISPLAY" -screen 0 "$GEOM" -dpi 96 -nolisten tcp -ac &
XVFB_PID=$!
# Do not sleep. Poll. A fixed `sleep 2` is both too long on a restored
# snapshot and too short on a cold boot.
for _ in $(seq 100); do
xdpyinfo -display "$DISPLAY" >/dev/null 2>&1 && break
sleep 0.05
done
xdpyinfo -display "$DISPLAY" >/dev/null 2>&1 || { echo "X never came up" >&2; exit 1; }
# DPI and scale factor: keep everything at 1.0 so that one screenshot pixel is
# one X coordinate. The moment GDK_SCALE or --force-device-scale-factor is not
# 1, the agent's "click at (x, y) from the screenshot" arithmetic is off by a
# constant it does not know about, and the failures look like a flaky model.
printf 'Xft.dpi: 96\nXft.antialias: 1\nXft.hinting: 1\n' | xrdb -merge
export GDK_SCALE=1 GDK_DPI_SCALE=1 QT_AUTO_SCREEN_SCALE_FACTOR=0
# ------------------------------------------------------------- 2. a window manager
# Not optional, and the reason is focus, not decoration. With no WM nothing
# owns the input focus, so synthetic keystrokes land nowhere, modal dialogs
# stack in the top-left corner, and menus close as soon as they open. Any
# lightweight reparenting WM does; openbox and fluxbox start in milliseconds.
openbox &
WM_PID=$!
# Fonts. A desktop with no fonts installed renders every string as tofu boxes,
# and an agent reading that screenshot is functionally blind. Liberation +
# DejaVu + Noto core covers Latin; emoji and CJK are separate packages and CJK
# is large enough that you should decide deliberately whether to ship it.
fc-cache -f >/dev/null 2>&1 || true
# ------------------------------------------------------------------- 3. x11vnc
# -localhost is load-bearing: the RFB port never leaves the guest's loopback,
# and the only thing listening on a routable address is the websocket proxy.
# -forever: without it x11vnc EXITS when the first client disconnects, which
# on a long-lived agent session means the display pipeline dies the first
# time a human closes a browser tab.
# -shared: a human and the agent's own viewer can both be attached.
# -repeat: x11vnc disables X key auto-repeat while a client is attached (it is
# protecting you from a laggy link). Held-key behaviour needs it back.
# -xkb: modifier and keymap handling that survives a non-US layout.
# Do NOT enable -ncache. Client-side caching parks pixels in an area below the
# visible screen; clients that show it hand your agent a screenshot twice as
# tall as the desktop, full of stale rectangles.
# -rfbauth: defence in depth. The preview URL is already unguessable, but a
# password costs nothing and means a leaked URL is not a leaked desktop.
x11vnc -display "$DISPLAY" \
-rfbport "$VNC_PORT" -localhost \
-forever -shared -repeat -xkb \
-rfbauth /etc/pandastack/vnc.passwd \
-o /var/log/x11vnc.log &
VNC_PID=$!
# --------------------------------------------------------- 4. noVNC + websockify
# websockify wraps the RFB TCP stream in a WebSocket so a browser can speak it.
# The novnc_proxy wrapper does both halves; the script's name and path move
# between distro packages (it was utils/launch.sh for years), so resolve it
# rather than hardcoding it.
NOVNC=$(command -v novnc_proxy || echo /usr/share/novnc/utils/novnc_proxy)
"$NOVNC" --vnc "localhost:${VNC_PORT}" --listen "0.0.0.0:${NOVNC_PORT}" \
>/var/log/novnc.log 2>&1 &
NOVNC_PID=$!
# ------------------------------------------------------------------- 5. audio
# Only if it matters -- but it matters more often than people expect. Chrome
# with no audio sink at all can fail WebAudio initialisation and take autoplay
# and media-capable pages down with it. A null sink is a device that exists and
# discards everything, which is exactly what a headless desktop wants.
# On distros that have moved to PipeWire, pactl still works through
# pipewire-pulse -- check which daemon your image actually ships.
pulseaudio --start --exit-idle-time=-1 >/dev/null 2>&1 || true
pactl load-module module-null-sink sink_name=agentsink >/dev/null 2>&1 || true
# --------------------------------------------------------------- 6. the clipboard
# X has two selections that both get called "the clipboard": PRIMARY (what
# middle-click pastes) and CLIPBOARD (what Ctrl-V pastes). Toolkits disagree
# about which one they own, so a paste that works in a browser silently does
# nothing in a Java app. autocutsel keeps them in sync; run one per selection.
autocutsel -selection CLIPBOARD -fork
autocutsel -selection PRIMARY -fork
echo "desktop up on ${DISPLAY} (${GEOM}); noVNC on :${NOVNC_PORT}"
# Hold the session open and take the whole thing down together. A desktop where
# x11vnc died but Xvfb is still running is the worst state to debug, because
# everything in the guest looks healthy and the client just hangs.
wait -n "$XVFB_PID" "$WM_PID" "$VNC_PID" "$NOVNC_PID"
echo "a desktop component exited -- tearing down" >&2
kill "$XVFB_PID" "$WM_PID" "$VNC_PID" "$NOVNC_PID" 2>/dev/null || true
Two choices deserve argument rather than a comment. Xvfb over Xorg with the `dummy` driver: the dummy driver is a real Xorg server with a stub video driver, so you inherit Xorg's module ecosystem and, depending on driver version, better RANDR support — read that driver's own documentation for the version your distro packages rather than any blog's summary. Xvfb's advantage is one binary and no configuration file, which in a baked template beats flexibility you will not use. And binding x11vnc to loopback while exposing only the websocket: RFB is not a protocol I want reachable from anything, and noVNC gives you a port whose only client is a browser speaking HTTP upgrade.
There is no GPU, and that caps the class of app
This is the honest ceiling. A Firecracker guest has a short device list — virtio-net, virtio-block, virtio-vsock, virtio-balloon, virtio-rng, pmem, a serial console, an i8042 stub — and recent versions add an opt-in virtio-PCI transport behind `--enable-pci`, with hot-plug in developer preview, plus ACPI support. What the list does not contain is a graphics device. There is no `/dev/dri`, no render node, nothing to pass through. Everything here is rendered by the CPU.
In practice that means Mesa's llvmpipe for OpenGL and lavapipe for Vulkan, both software rasterisers. Read the relevant environment variables (`LIBGL_ALWAYS_SOFTWARE`, `GALLIUM_DRIVER`, and the Vulkan driver-selection variable, whose name changed when `VK_ICD_FILENAMES` was deprecated) from your Mesa version's own documentation rather than copying them. Chrome additionally blocklists software GL by default and the opt-in switch has been renamed across majors, so check `chrome://gpu` on the exact build you ship.
The consequence is a line through the application space, and a usefully predictable one. Good side: browsers doing ordinary DOM and CSS work, forms, SaaS dashboards, spreadsheets, PDF viewers, administrative thick clients, anything 2D. Bad side: WebGL and WebGPU pages, canvas-heavy visualisation, GPU-rendered map tiles, and video at any appreciable resolution — there is no hardware decode either, so 1080p is software H.264 or VP9 on the same CPUs running the browser, and it drops frames. Games and 3D tools are off the table.
The three paths that always break: clipboard, downloads, screenshots
One piece of advice to anyone about to build this: solve the clipboard and the download directory on day one. They are where a desktop stops being a self-contained picture and starts exchanging data with the outside world, and both are historically awful.
The clipboard is awful because X has two selections both colloquially called “the clipboard” — PRIMARY, which middle-click pastes, and CLIPBOARD, which Ctrl-V pastes — and toolkits disagree about which they own. `autocutsel` run once per selection fixes the within-desktop case. Across the wire is worse: the base RFB cut-text message is Latin-1, and UTF-8 needs the extended-clipboard pseudo-encoding that both ends must implement, so a paste containing an em dash is a coin flip on versions.
My opinion, having been bitten: do not make the VNC clipboard a data path at all. The agent's data plane should be `exec` and the filesystem API — set the selection with `xclip` over an exec call, read files back with `filesystem.read` — and the RFB stream should carry pixels for human eyes and nothing else.
Downloads break for a different reason: a real browser downloads through its UI, so what you configure is a profile rather than a flag, and what you wait for is a rename rather than a file.
#!/usr/bin/env bash
# The download path. This is the single most reliably broken thing about
# driving a real browser on a real display, so set it up before the agent's
# first click rather than after its first bug report.
set -euo pipefail
PROFILE=/home/agent/.config/chromium
DL=/work/downloads
mkdir -p "$DL" "$PROFILE/Default"
# Chrome's download directory lives in the profile's Preferences JSON. Two
# traps: it is JSON written by Chrome, so a half-baked file makes Chrome reset
# the profile; and Chrome REWRITES it on exit, so editing it while Chrome is
# running silently loses your change. Write it with Chrome stopped.
python3 - "$PROFILE/Default/Preferences" "$DL" <<'PY'
import json, sys, os
path, dl = sys.argv[1], sys.argv[2]
prefs = json.load(open(path)) if os.path.exists(path) else {}
prefs.setdefault("download", {}).update({
"default_directory": dl,
"prompt_for_download": False,
})
# Stops the "unlock your keyring" modal that otherwise blocks the whole UI on a
# machine with no desktop keyring -- a dialog an agent cannot reason its way out
# of because it is not about the page at all.
prefs.setdefault("profile", {})["password_manager_leak_detection"] = False
json.dump(prefs, open(path, "w"))
PY
# --password-store=basic is the other half of the keyring fix.
# --disable-dev-shm-usage if /dev/shm in your guest is small; Chrome's renderers
# use it for shared memory and run out in ways that present as tab crashes.
# Window size is passed explicitly so it matches the Xvfb geometry exactly --
# a maximised window and a known window are not the same thing to an agent
# doing coordinate arithmetic.
DISPLAY=:99 chromium \
--user-data-dir="$PROFILE" \
--password-store=basic --disable-dev-shm-usage \
--window-position=0,0 --window-size=1280,800 \
--force-device-scale-factor=1 --no-first-run --no-default-browser-check &
# Waiting for a download to FINISH is not "does the file exist". Chrome writes
# to <name>.crdownload and renames on completion, so the file appears instantly
# and is truncated garbage for as long as the transfer takes. Wait for the
# rename, and for the partial file to be gone.
wait_for_download() {
local deadline=$(( $(date +%s) + ${2:-120} ))
while [ "$(date +%s)" -lt "$deadline" ]; do
if [ -e "$DL/$1" ] && ! compgen -G "$DL/*.crdownload" >/dev/null; then
echo "ok: $DL/$1"; return 0
fi
sleep 0.5
done
echo "download of $1 never completed" >&2; return 1
}
# The clipboard, from the control plane rather than through RFB. The base RFB
# cut-text message is Latin-1; UTF-8 needs the extended-clipboard pseudo-
# encoding, which BOTH ends have to implement -- so anything with an em dash or
# a non-Latin character in it is a coin flip across client versions. Read and
# write the selection with xclip over exec instead, and keep the VNC stream for
# human eyes only.
# read : DISPLAY=:99 xclip -selection clipboard -o
# write: printf %s "$text" | DISPLAY=:99 xclip -selection clipboard -i
Take the screenshot inside the guest
Third path, same theme. You can capture the X root window inside the guest with `import`, `scrot`, `xwd` or `ffmpeg -f x11grab`, or scrape a frame out of the VNC stream client-side. People reach for the second because the stream is already open.
Capture it in the guest. The in-guest capture is a lossless PNG at exactly the framebuffer resolution, with no cursor composited in, from a consistent framebuffer state. The VNC stream is the opposite of all four: encodings like Tight are lossy with chroma subsampling, the frame you grab may be mid-update and torn across rectangles, the client may be scaling, and the cursor is drawn in. Worse, the stream only exists while something is connected — so a pipeline that screenshots from it has a hidden dependency on a viewer being attached, which breaks the first time the agent runs unwatched.
Keep `ffmpeg -f x11grab` for the human-facing artefact — a session recording is genuinely useful when an agent does something inexplicable — but treat it as a separate output, not the agent's perception channel.
One microVM per session, and a URL that is a password
Now the isolation argument, which is why I am writing this instead of a Dockerfile. Consider what this session does: it drives a mouse over websites chosen by a language model, downloads files chosen by those websites, renders attacker-controlled content through a software rasteriser and a font stack, and runs shell commands the model generated at runtime. The blast radius is not a page. It is a whole desktop with a filesystem, a network, and a browser profile that may hold a real session cookie.
A container is the wrong boundary for that, and not because namespaces are badly built. It is because the thing you are isolating is a general-purpose desktop, the surface you are isolating it from is the shared host kernel, and a container is ultimately a polite suggestion to that kernel about what this process should be allowed to see. One kernel privilege-escalation bug — or one unlucky interaction between a model-generated command and a filesystem driver — and every other session on the host is in scope. A microVM moves the boundary to hardware virtualisation: separate kernel, separate device model, with the VMM a much smaller thing to attack than a kernel's full syscall surface.
The objection has always been cost, and snapshot-restore is what changes that arithmetic. Every create on PandaStack restores a baked snapshot rather than booting: p50 179 ms, p99 203 ms, with the `/snapshot/load` step itself in the 49–80 ms range. Only the first spawn of a template, before any snapshot exists, pays a cold boot of around 3 seconds. So “a VM per session” stops meaning “wait for a boot” and starts meaning something closer to starting a process. Each sandbox also gets its own network namespace, veth pair and tap device, from a pool of 16,384 pre-allocated /30 subnets per agent — the real ceiling, and not one you reach by accident.
from pandastack import Sandbox
# The template is built once, and the two numbers that decide whether a desktop
# is usable are both chosen HERE, not at create time:
#
# pandastack template build -f templates/agent-desktop/Dockerfile \
# -n agent-desktop --size-mb 8192 --memory-mb 4096
#
# --memory-mb is baked into the snapshot. Firecracker cannot change guest RAM
# or vCPU count at snapshot restore, so passing memory_mb on a create is not
# an error -- it is silently corrected to the baked value, which is worse than
# an error. A desktop running a modern browser wants 4 GiB; our `browser`
# template is baked at exactly that, and every template gets 8 burstable vCPUs.
# --cpu is deprecated and ignored.
sbx = Sandbox.create(
template="agent-desktop",
ttl_seconds=900, # the backstop. See the TTL note below.
metadata={"session": "demo-1", "owner": "agent-runner"},
)
# Wait for the display, not for the sandbox. `create` returns when the VM is
# reachable; X, the WM, x11vnc and websockify come up after that. On a baked
# snapshot they are already running and this returns immediately -- which is
# the entire point of baking the desktop instead of booting it.
#
# NOTE on timeouts: neither exec endpoint enforces timeout_seconds on the
# server. It is a CLIENT deadline. The real limit is the `timeout` in the
# shell, which is why it is there and not only in the keyword argument.
probe = r"""
timeout 60 sh -c '
until xdpyinfo -display :99 >/dev/null 2>&1; do sleep 0.2; done
until curl -sf -o /dev/null http://127.0.0.1:6080/vnc.html; do sleep 0.2; done
# Assert the geometry the agent was told to expect. If this ever disagrees,
# every coordinate the agent computes from a screenshot is wrong.
xdpyinfo -display :99 | awk "/dimensions:/ {print \$2}"
'
"""
r = sbx.exec(probe, timeout_seconds=90, check=True)
print("display:", r.stdout.strip()) # -> 1280x800
# Hand the agent (or a human) the desktop. Preview URLs on PandaStack are
# tokenless: https://<port>-<sandbox-id>.<suffix>. There is no token endpoint,
# because the sandbox UUID *is* the credential -- treat the string the way you
# would treat a password, and keep the TTL short so a leaked one expires.
#
# resize=off matters. noVNC's local scaling and remote-resize modes both change
# the mapping between the pixels the agent sees and the coordinates X accepts.
base = sbx.preview_url(6080)
desktop_url = f"{base}/vnc.html?autoconnect=1&resize=off&reconnect=1"
print("desktop:", desktop_url)
# ---- the screenshot the agent should actually look at ---------------------
# Take it INSIDE the guest. Capturing the X root window gives you a lossless
# PNG at exactly the framebuffer resolution, with no cursor, no JPEG chroma
# subsampling, and no partially-applied framebuffer update. Scraping frames out
# of the VNC stream gives you a lossy, torn, scaled approximation of that --
# and it only works while a client happens to be connected.
#
# ImageMagick 7 renamed the tools: `magick import` on 7, bare `import` on 6.
# xwd is in every X install and needs no ImageMagick at all, so it is the
# portable fallback.
shot = r"""
set -e
export DISPLAY=:99
if command -v magick >/dev/null; then magick import -window root /tmp/shot.png
elif command -v import >/dev/null; then import -window root /tmp/shot.png
else xwd -root -silent | magick xwd:- /tmp/shot.png; fi
"""
sbx.exec(shot, timeout_seconds=30, check=True)
# filesystem.read() returns bytes and handles ONE file. There is no directory
# copy, and upload() of a directory raises IsADirectoryError -- for a sequence
# of frames, tar it in the guest first or push it out from inside.
open("shot.png", "wb").write(sbx.filesystem.read("/tmp/shot.png"))
# Short sessions, deleted sessions. The TTL above is a backstop for the case
# where your orchestrator dies; the normal path is that you kill it yourself
# the moment the task is done, because an idle desktop you forgot about is both
# a bill and a live credential.
sbx.kill()
The preview URL deserves its own paragraph, because it is the part people get casually wrong. Preview URLs on PandaStack are tokenless: a port is reachable at `https://<port>-<sandbox-id>.<suffix>` for the sandbox's lifetime, and there is no token endpoint because the sandbox UUID is the credential. That is a deliberate trade — handing a URL to a model or a teammate is trivial — and it makes the string a bearer secret. A UUID in a log line, a Slack paste, a model's own transcript or a screenshot of your terminal is a live desktop someone else can now drive.
The mitigations are unglamorous and they work. Short TTLs, so a leaked URL is dead in minutes rather than days. `kill()` on the success path as well as the failure path, so a finished session stops existing. A VNC password on top, so a leaked URL is not by itself a leaked desktop. And treating the UUID like a password in your own observability — which mostly means not logging it at info level, the way you already do not log API keys.
Bake the desktop; don't boot it
The best property of this architecture is one you only get from snapshots. Everything in the bring-up script is identical on every session: the same X server at the same geometry, the same window manager, the same fonts in the same fontconfig cache, the same browser with the same warmed profile, the same null audio sink. None of it is session-specific, so bake it. Build the template with the desktop already running, snapshot that, and a create restores a desktop that is already up rather than starting four daemons and waiting on each.
That is also where the browser profile wants to live. A cold Chrome on first launch builds caches, writes a profile, probes for a keyring, and generally takes longer than you would like before it is ready to be clicked on. Do it once at bake time.
Three caveats, in increasing order of surprise. Memory is baked: Firecracker cannot change guest RAM or vCPU count at snapshot restore, so the desktop's memory is a template-build decision and a `memory_mb` on the create is silently corrected. Our `browser` template is baked at 4 GiB with 8 burstable vCPUs, which is the right shape for this. TCP connections do not survive a snapshot — a session restored with a viewer previously attached comes back with a dead socket, which is why `reconnect=1` is in the noVNC query string above.
Third, and this is the one that produces genuinely baffling bugs: forking. A `fork()` on PandaStack is a disk-and-memory snapshot restore of the parent, which means the child's RNG state and clock come back identical to the parent's. Fork eight desktop sessions from one warm parent — same-host forks land in 400–750 ms, cross-host in 1.2–3.5 seconds — and you have eight guests that agree on what time it is and will generate the same “random” values. On a desktop that means X magic cookies, browser session identifiers and TLS client randoms that collide across siblings, and a clock that is wrong by however long the parent sat frozen, which presents as certificate validation failures that look like a network problem. Re-seed the RNG and re-sync the clock on wake. It is two lines and it is not optional.
A desktop is the most general sandbox escape target you can hand a model, and also the most useful tool you can give it. Those two facts are not in tension; they are the reason the boundary has to be a machine.
Five ways to run an agent's desktop
Qualitative on purpose for everything that is not ours, and the vendor row is deliberately vague because these products move fast enough that any specific claim here is stale before you read it. Verify against their current documentation, not this table.
| Approach | Real desktop apps | Isolation boundary | Idle cost | Session start | Where it hurts |
|---|---|---|---|---|---|
| Headless browser only | No — a browser and nothing else | The process, plus whatever you wrapped it in | Low if you tear it down | Fast | A native app, an OS file picker or a human-in-the-loop handoff has nowhere to go |
| Container with Xvfb | Yes, and it is a genuinely good dev-loop answer | Namespaces and seccomp over a shared host kernel | Low | Fast | One kernel bug, reached by a model-generated command or a malicious download, is every other session on the box |
| Full cloud VM per session | Yes | Hardware virtualisation | High — you pay for a booted machine whether or not it is doing anything | Tens of seconds to minutes | Too slow and too expensive to hand one to every short agent run, so you end up pooling and sharing them |
| MicroVM per session | Yes | Hardware virtualisation, own kernel, own netns | Zero if you delete it | Snapshot restore: p50 179 ms, p99 203 ms. Around 3 s for the first cold boot before a snapshot exists | No GPU at all, no nested virtualisation, a guest kernel you do not choose, and RAM fixed at template-build time |
| Commercial browser-infra / RBI vendors | Varies — mostly browser-shaped; some offer a fuller desktop | The vendor's, and usually not documented in detail | The vendor's pricing model | Usually fast, often warm-pooled | You inherit their isolation model, their session semantics and their egress policy. Genuinely good at stealth, proxying and scale — verify the current docs |
To be fair to the middle row: a container with Xvfb is the right answer in plenty of situations, and I would not replace it for running your own screenshot tests in CI against your own code. The argument for a VM is about this threat model specifically — arbitrary websites, arbitrary downloads, model-generated shell commands — and about multi-tenancy. If every session on the host is yours and runs code you wrote, the shared kernel is a much smaller worry.
What to build, in order
- A template with the desktop baked in: Xvfb at one fixed geometry, a window manager, x11vnc on loopback, noVNC, fonts, a null audio sink, a warm browser profile. Build it with `--memory-mb 4096` and stop thinking about memory.
- A readiness probe the agent runs on every connect that asserts the geometry, not just that X answered. This catches the whole class of coordinate bugs at the door.
- An in-guest screenshot path, hashed. Feed the model the PNG from inside the guest; keep video as a separate human artefact.
- A download directory configured in the browser profile with Chrome stopped, plus a wait that looks for the `.crdownload` rename rather than file existence.
- A clipboard bridge over `exec` and `xclip`, and a written decision that the RFB clipboard is not a data path.
- A TTL on every session and a `kill()` on both the success and failure paths, with the sandbox UUID treated as a secret in your logging configuration.
- A re-seed and clock-sync step on wake, before you fork warm desktops — not after a day of debugging TLS.
None of this is exotic. It is a 1990s X stack, a websocket, and a virtual machine you throw away, assembled with enough care that an agent clicking through someone else's website cannot take anything with it when it goes. The interesting part was never the desktop. It was deciding that the blast radius of a mouse pointer driven by a language model is a whole machine, and then making a whole machine cheap enough to burn.
Frequently asked questions
Do I need a GPU to run an agent desktop, and what breaks without one?
You do not need one for a large class of work, but the line is sharp enough to design around. With no graphics device in the guest there is no /dev/dri and no render node, so everything goes through a software rasteriser — Mesa's llvmpipe for OpenGL, lavapipe for Vulkan — on the same CPUs running your application. What works well: ordinary DOM and CSS rendering, forms, SaaS dashboards, spreadsheets, PDF viewers, administrative thick clients, anything fundamentally 2D. What does not: WebGL and WebGPU pages, canvas-heavy visualisation, GPU-rendered map tiles, and video at any real resolution, because there is no hardware decode either and 1080p becomes software H.264 or VP9 competing with the browser for CPU. Games and 3D authoring tools are out. Chrome also blocklists software GL by default and the opt-in switch has been renamed across major versions, so check chrome://gpu on the exact build you ship rather than copying a flag. If your workload is genuinely GPU-bound, a microVM with no GPU is the wrong substrate and no amount of configuration fixes that.
Xvfb or Xorg with the dummy driver — which should I use?
Xvfb if you are baking a template, Xorg with the dummy driver if you need runtime resolution changes. Xvfb is a single binary with no configuration file: you give it a geometry and a display number and it draws into memory. That simplicity is worth a lot in an image you build once and restore thousands of times, and its main limitation — screens are created at startup, so runtime resizing through RANDR is limited and version-dependent — does not matter if you fixed the geometry anyway, which for a coordinate-clicking agent you should. The dummy driver is a real Xorg server with a stub video driver, so you inherit Xorg's full module ecosystem and, depending on the driver version, better RANDR behaviour. The honest answer on that last point is that dummy's RANDR support has changed over time and across distro packaging, so read the driver documentation for the version your image actually installs rather than any blog's summary. If you are on Wayland instead, the equivalent shape is a headless wlroots compositor with wayvnc, which is a cleaner architecture and a smaller application compatibility set.
Is a tokenless preview URL safe enough to expose a VNC desktop on?
It is safe in the sense that the URL is unguessable and unsafe in the sense that it is a bearer credential with no second factor. On PandaStack a port is reachable at https://<port>-<sandbox-id>.<suffix> for the sandbox's lifetime, and there is no token endpoint because the sandbox UUID is the credential. For a VNC desktop that is a higher-stakes string than for a dev-server preview: whoever holds it has a mouse. Treat it accordingly. Keep TTLs short so a leak expires in minutes, call kill() on the success path as well as the failure path, put a VNC password on x11vnc so the URL alone is not sufficient, and audit where the UUID ends up — logs, Slack, a model's own transcript, a screenshot of your terminal. The failure mode is not an attacker brute-forcing a UUID; it is a UUID being pasted somewhere durable. If your threat model does not tolerate a bearer URL at all, do not expose the port and tunnel to it from your own control plane instead.
Can I snapshot a desktop session with the browser already logged in, and what are the catches?
Yes, and it is the single biggest win available here — a session that starts from a restored desktop with the window manager up, fonts cached and a warm browser profile is a different experience from one that runs four daemons and waits. Create restores a baked snapshot at p50 179 ms. Three catches. Memory is baked: Firecracker cannot change guest RAM or vCPU count at restore, so the desktop's RAM is chosen with --memory-mb at template build time and a memory_mb on the create is silently corrected to the baked value. TCP connections do not survive, so a restored session comes back with any previously attached viewer's socket dead — have the client reconnect. And forking is where people get hurt: a fork is a disk-and-memory restore of the parent, so the child's RNG state and clock come back identical. Fork several desktops from one warm parent and they agree on the time and generate the same random values, which surfaces as colliding session identifiers and TLS failures that look like a network problem. Re-seed the RNG and re-sync the clock on wake.
Keep reading
- A sandbox for AI agent computer use — The threat model this post assumes: prompt injection, model-generated shell commands, and where the damage should land.
- Browser isolation on microVMs — The same boundary argument from the browser side, without the desktop stack on top.
- What is a headless browser — The distinction this post depends on: a browser with no window versus a machine with no screen.
- Form-filling and RPA in a microVM — The human-in-the-loop workflow that needs a screen somebody can actually attach to.
- Snapshotting browser session state — Baking a warm, logged-in profile into a template — and what the fork caveat does to it.
Related posts
- Best Remote Browser Isolation Platforms (2026)
RBI used to be about protecting a human from a drive-by exploit. In 2026 it's also about protecting you from your own agent — which will click things no human would, on a page that is actively trying to talk to it.
- Sandboxing AI-Agent 3D Rendering and Asset Pipelines in MicroVMs
Your agent writes a Blender script and something has to run it. Meanwhile the scene file it was handed can execute Python the moment it opens. Give that job a machine you can throw away.
- How many browsers can one machine actually run?
Everyone discovers the browser concurrency limit the same way: by exceeding it. The number is set by memory, it's lower than you'd guess, and the failure past it looks like flaky tests.
- Isolating AI Shopping Agents in MicroVMs
A jailbroken shopping agent with your saved credit card and a shared cookie jar is exactly the machine you don't want to build. Give every session its own throwaway microVM.
- Testing Browser Extensions with AI Agents in MicroVMs
An extension gets content-script access to every page you visit, your cookie jar, and a background worker that can phone home. Installing one to 'just test it' on your laptop is a trust decision. Do it in a microVM you throw away.
More in AI agent sandboxes · See AI agent sandboxes on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.