Document Conversion Is Remote Code Execution With a Progress Bar
Every product eventually grows a convert endpoint. The e-signature product needs a customer's docx as a PDF so it can stamp a signature field on page four. The helpdesk wants a preview of the xlsx somebody attached to a ticket. The document management system wants full-text search, which means text, which means opening the file. The request is two sentences long; the implementation is twenty lines. Write the upload to a temp directory, shell out to soffice, read the output back, delete the directory, return 200.
What those twenty lines do is accept an arbitrary program from an anonymous stranger and execute it on your infrastructure, inside your VPC, with your service account's credentials one HTTP request away. That is not a metaphor I am stretching for an opening: the office formats contain features that are executable by design and features that reach the network by design, implemented by a large body of C++ with a long CVE history. The only thing structurally separating your converter from an eval endpoint is the progress bar.
I build PandaStack, an open-source Firecracker microVM sandbox platform, and conversion services are a workload people arrive with already knowing something is wrong. They rarely phrase it as security. They phrase it as operations: the converter wedges, the pool fills with zombie soffice processes, and nobody can say whether the profile directory on worker three holds somebody else's temp files. Those complaints and the security problem share a root cause and a fix. What makes a conversion service distinct from an agent reading a PDF is that it is a daemon rather than a function.
The endpoint is an interpreter with a file picker
Start with what a modern office document is, because the usual mental model is wrong in a way that matters. OOXML — docx, xlsx, pptx — is a zip container holding a tree of XML parts, binary blobs, and a graph of relationships between them; ODF is the same idea with different schema names. A docx is not a document the way a PNG is an image. It is closer to a small application bundle: markup, a resource tree, a manifest of external references, optional embedded objects, optional fonts, optional code.
So opening one is not decoding a format. It is traversing an attacker-authored graph with a different parser at nearly every node: the zip reader, an XML parser per part, the relationship resolver that decides whether a target is external and whether to fetch it, an image decoder per image, the font engine, the formula evaluator, and the legacy binary readers, because somebody will upload a doc from 2003 and expect it to work. Each is a trust boundary and your twenty-line handler crossed all of them in one subprocess call. This is also why validation cannot be the answer: deciding a docx is safe requires parsing it.
The office-format features that dial out, by name
XXE: the document that reads your filesystem
XML external entities keep working because XML parsers are configured by whoever embedded them, and plenty of embedders left the defaults alone. A document part declares an entity whose body is a SYSTEM reference to a local path, then references it somewhere the contents land in rendered output. The parser resolves it; the file appears in the PDF. You have accidentally built an endpoint that reads files off your converter host and returns them, paginated, to the caller. The blind variant points the entity at a URL and reads the answer from timing, or from the attacker's server getting a request at all.
SSRF: external relationships, remote images, remote templates
This is the one I would worry about first in 2026, because it needs no memory-safety bug, no macro, and no interaction beyond the conversion itself. The relationship graph inside an OOXML package can mark a target as external, with a URL. Images can be referenced rather than embedded. Documents can inherit from a template fetched at open time. Stylesheets, OLE sources and linked spreadsheet ranges all point outward. A converter that resolves those references — correct behaviour for a desktop word processor — makes an HTTP request chosen by the uploader.
# Three payload shapes, none of which require a parser bug.
# (1) An external relationship inside word/_rels/document.xml.rels.
# The target is a URL and the parser is being helpful.
<Relationship Id="rId7"
Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image"
Target="http://169.254.169.254/latest/meta-data/iam/security-credentials/"
TargetMode="External"/>
# (2) XXE in any XML part the converter reads. The entity body is a path
# on YOUR host, and the reference puts it in the rendered output.
<!DOCTYPE doc [
<!ENTITY leak SYSTEM "file:///var/run/secrets/token">
]>
<w:t>&leak;</w:t>
# (3) A DDE field. Not a bug, not a macro -- a documented field type
# that names a program and its arguments.
DDEAUTO c:\\windows\\system32\\cmd.exe "/c calc.exe"
# The converter runs as your service. The request leaves from your network
# position. The output is a PDF you hand back to the uploader.Read payload (1) again and notice what it produces. A stranger uploads an invoice. Your converter fetches the cloud metadata endpoint, because the invoice said there was a logo there. The response — temporary credentials for the role your worker runs as — goes into the document as image content or an error string, is rendered, and comes back to the uploader as a nicely typeset PDF. There is a version of this where the credentials arrive in a table with alternating row shading, and I think about it more than is healthy.
Macros, OLE objects and DDE
VBA macros are the famous one and, headless, the least interesting: every serious suite disables macro execution at its highest security level. The caveat is that macro security is a setting, settings live in a per-user profile, and profiles on a shared converter are not reliably the profile you configured. OLE embedded objects are more interesting: an object is a blob plus a class identifier, so handling it means dispatching to whatever handler claims that class — code the document chose. Even declining to activate it means parsing enough to render a placeholder, through a legacy compound-file container with its own history. DDE fields are stranger still: a field names an external program and its arguments and asks the application to run it. Modern suites prompt or refuse, which headless means a hang, or a log line nobody reads.
Fonts: the parser inside the parser
Embedded fonts are the easiest item to forget. A document can carry its own font file, the renderer must parse it to lay out glyphs, and font formats are binary, tabular and — in their hinting instructions — a bytecode the font engine interprets. A font-shaped input reaching a font engine is one of the better-worn paths to memory corruption in all of client software, and uploading a document with a subsetted font in it is enough to reach yours.
The converter runs with your identity and your network position
All of that is only dangerous because of where the conversion happens. The process has your credentials, because somebody attached an instance role so the worker could write results to object storage. It has your network position, because it runs on your subnet, which routes to your internal services, your metadata endpoint and your database. It also has your shared temp volume. That is the problem in one sentence: an untrusted document is parsed inside your trust boundary, with your privileges, and the document gets a vote on what happens there. The fix is not a better parser, which you do not maintain anyway. It is somewhere with no credentials, no useful network position, and nothing of yours on disk.
A headless office suite is a desktop app in a daemon costume
Here the security story and the pager story merge, because even if every document is benign, running a headless office suite as a service is structurally miserable. It keeps a user profile: configuration, extension registrations, dictionaries and caches, created lazily on first run and written during normal operation. Two instances sharing one profile fight over it, which is why every production guide tells you to pass a per-instance profile path. It writes lock files, so a suite that dies uncleanly leaves a lock that makes the next job fail in a way resembling nothing like the cause. It often wants a session bus, so somebody ends up running dbus in a container whose purpose is converting spreadsheets.
And it hangs. A desktop application's answer to ambiguous input is to ask the user; a headless suite has no user, so the question goes into the void and the process waits, patiently and indefinitely, for a click that will never arrive. Password-protected files, corrupt-but-recoverable files, a repair prompt, a reference to an address that blackholes packets rather than refusing them: each becomes a process that is alive, holding a slot, never finishing. It also leaks descriptors and memory over a long run, which is why the oldest folk wisdom here is the one that works: restart the converter every N jobs.
Process per request, a pool, and why 'restart every N jobs' is the real answer
There are two honest architectures and the industry has mostly picked both, badly. Process per request is clean — fresh profile, no carryover, a hang bounded by killing one process — and pays the suite's full startup cost on every request, the slowest part of the job. A warm pool is fast, because the suite is up and you submit work over a socket, and it carries everything from the previous job: profile, caches, descriptors, memory, and any state a hostile document established.
So real deployments run a pool and bolt a lifecycle onto it. Recycle after N conversions. Health-check the listener, restart on probe timeout. Kill anything past a wall-clock limit. Give each listener its own profile directory and delete it on restart. If you have operated one of these you recognise the list, and that most of its complexity exists to make a stateful desktop application behave like a stateless function. That is the tell: a runbook that is mostly steps for erasing state is emulating disposability on a substrate that does not offer it.
Why a container is the wrong boundary for this job
To be fair, because the usual version of this argument is lazy: a container per conversion is dramatically better than a shared pool. Read-only image, non-root user, dropped capabilities, a seccomp profile, noexec mounts, a network policy — real controls, and you should want all of them regardless. The problem is what they are made of. Namespaces, cgroups, seccomp and LSM policy are host-kernel features, and the untrusted font parser, XML reader and image decoder are issuing syscalls to that same kernel. The boundary and the thing you are defending against are the same code. A container is a polite suggestion to the kernel, enforced by the kernel, about code that is talking to the kernel.
They also leave the operational half untouched. Namespaces isolate your view of the kernel, not its accounting, so gigabytes of dirty page cache are still flushed under the host's thresholds and every other writer notices. A hung soffice is still a process on your host. You moved the mess, not deleted it.
What a microVM buys: deletion instead of cleanup
A Firecracker microVM gives you four things here, and only three are the ones people expect.
- Its own guest kernel. Every sandbox boots a 5.10 guest kernel under Firecracker v1.16 in an Ubuntu 24.04 userland, so the font parser's syscalls go to a kernel that exists for this one conversion — and a guest kernel bug costs the attacker a machine already scheduled for destruction.
- A tiny device model: a deliberately small set of virtio devices and no legacy emulation zoo. That interface is small enough to reason about, which is a very different sentence from the one you can write about the Linux syscall table.
- Its own network namespace, pre-allocated. Each host agent carves 16,384 /30 subnets out of 10.200.0.0/16 and keeps namespaces, veth pairs and taps built in advance, so per-job network isolation is a slot allocation. What policy goes in that namespace is still your decision — see below.
- And the operationally decisive one: you delete the machine instead of cleaning it.
That last point collapses the whole runbook above. No profile directory to reset, because the profile was part of a disk image that no longer exists. No lock file to reap. No question of whether this soffice belongs to the previous tenant, because there is no previous tenant. No recycle-after-N to tune, because N is one and always was. A hung converter becomes a log line rather than a debugging exercise. You destroy a VM.
Snapshot the warm suite: this is the whole trick
All of this would be academic if a VM per conversion meant paying for a boot and an office-suite startup on every request. The conversion itself is usually fast. The expensive part of a conversion worker is becoming one: starting the suite, loading its filters and configuration registry, validating the profile, building the font cache. That is the cost a warm pool amortises, and the only reason anyone tolerates its state carryover.
So move the cost into the template. PandaStack keeps no warm pool of idle VMs; every create restores a baked Firecracker snapshot. A template's first spawn cold-boots, around three seconds, and the agent captures a snapshot of the running machine. Every create after that is a restore: p50 179 ms, p99 203 ms on our fleet. Guest memory returns through a private mapping, so pages fault in lazily and nothing is copied until written; the disk is a reflink clone. The consequence is the design: whatever was running when the snapshot was taken is running when it is restored. Bake so the snapshot is captured after soffice is up, its profile exists and its font cache is built, and every create hands you a warm suite in a fifth of a second.
# A converter template. The point of the bake is that the snapshot gets
# taken AFTER the office suite is warm and its profile exists -- so every
# create restores a running converter instead of starting one.
cat > soffice-warm.service <<'UNIT'
[Unit]
Description=Warm headless LibreOffice, captured in the template snapshot
[Service]
Type=simple
Environment=HOME=/root
# A per-machine profile path. On a shared pool this is the flag that stops
# two instances corrupting each other. Here it is just tidiness: the
# machine IS the instance.
ExecStart=/usr/bin/soffice --headless --norestore --nologo -env:UserInstallation=file:///srv/profile --accept=socket,host=127.0.0.1,port=2002;urp;
Restart=always
[Install]
WantedBy=multi-user.target
UNIT
cat > Dockerfile <<'DOCKER'
FROM ubuntu:24.04
ENV DEBIAN_FRONTEND=noninteractive HOME=/root
# The suite, the filters, the fonts, and the other usual suspects. Fonts
# matter more than people expect: a missing font is a silently wrong PDF,
# and building the cache is slow, so build it at bake time.
RUN apt-get update && apt-get install -y --no-install-recommends libreoffice-core libreoffice-writer libreoffice-calc libreoffice-impress fonts-dejavu fonts-liberation fonts-noto-core pandoc ghostscript fontconfig ca-certificates && rm -rf /var/lib/apt/lists/*
# Macro execution off at the strictest level the suite offers. Defence in
# depth, NOT the boundary. The boundary is the VM.
RUN mkdir -p /etc/libreoffice && echo 'MacroSecurityLevel=3' >> /etc/libreoffice/sofficerc
# Materialise the profile and the font cache at BUILD time, in the exact
# path the runtime will use. soffice creates the profile lazily on first
# run, and 'first run' is otherwise your first paying customer.
RUN mkdir -p /work/in /work/out /srv/profile && echo warm > /tmp/warm.txt && soffice --headless -env:UserInstallation=file:///srv/profile --convert-to pdf --outdir /tmp /tmp/warm.txt && fc-cache -f && test -s /tmp/warm.pdf
# Start the listener at boot so it is RUNNING when the snapshot is taken.
COPY soffice-warm.service /etc/systemd/system/soffice-warm.service
RUN systemctl enable soffice-warm.service
DOCKER
# Build the rootfs and register the template.
pandastack template build -f Dockerfile -n doc-convert
# The first create cold-boots (~3 s) and the agent bakes a snapshot of the
# running machine. Every create after that is a restore -- p50 179 ms --
# with a warm suite already inside it.
curl -s -X POST https://api.pandastack.ai/v1/sandboxes -H "Authorization: Bearer $PANDASTACK_API_KEY" -H 'Content-Type: application/json' -d '{"template":"doc-convert","ttl_seconds":300}'One constraint to design around: guest sizing is baked into the snapshot. Firecracker cannot change a restored machine's vCPU count or memory size, so the cpu and memory_mb you pass on a create are overridden to match the template. Mostly fine, occasionally annoying — a 300 MB spreadsheet with ten thousand charts wants a different machine than a one-page letter. The answer is tiers, not dynamic sizing: bake a small converter template and a large one, and route by input size. That decision is one comparison against a byte count you already have.
The per-request flow, end to end
Create a machine from the warm template, write one input file into it, convert under a timeout, check the output size before reading it, read it back, and destroy the machine in a finally block so a crash in your handler cannot leak a VM.
from pandastack import Sandbox
MAX_OUTPUT_BYTES = 50 * 1024 * 1024
HARD_SECONDS = 90
class ConversionRejected(Exception):
"""We reject the CONVERSION, never 'the document is malicious'."""
def convert_to_pdf(blob: bytes, upload_name: str) -> bytes:
# The filename and the Content-Type came from the uploader, so they are
# data, not metadata. Sniff the format ourselves and choose our own
# filename -- the uploader never gets to influence a path.
ext = sniff_format(blob) # your detection, from magic bytes
if ext not in {"docx", "doc", "odt", "xlsx", "pptx", "rtf", "md"}:
raise ConversionRejected("unsupported input format")
sbx = Sandbox.create(
template="doc-convert",
ttl_seconds=300, # idle backstop, NOT a deadline
metadata={"job": "convert", "fmt": ext},
)
try:
sbx.filesystem.write(f"/work/in/src.{ext}", blob)
# Two clocks. `timeout` in the guest is the one that reliably fires
# on a wedged suite; timeout_seconds is the caller-side backstop.
# -s KILL because a hung soffice waiting on a dialog that does not
# exist will not be talked out of it by SIGTERM.
cmd = (
f"timeout -s KILL {HARD_SECONDS} soffice --headless --norestore "
f"-env:UserInstallation=file:///srv/profile "
f"--convert-to pdf --outdir /work/out /work/in/src.{ext}"
)
r = sbx.exec(cmd, timeout_seconds=HARD_SECONDS + 30, check=False)
if r.exit_code in (124, 137):
raise ConversionRejected("conversion exceeded its time budget")
# Size-check BEFORE reading. A conversion can produce a far larger
# artefact than its input, and the read pulls it over the wire.
probe = sbx.exec("stat -c %s /work/out/src.pdf 2>/dev/null || echo 0")
size = int(probe.stdout.strip() or 0)
if size == 0:
raise ConversionRejected(f"no output produced: {r.stderr[:400]}")
if size > MAX_OUTPUT_BYTES:
raise ConversionRejected(f"output {size} B over cap")
return sbx.filesystem.read("/work/out/src.pdf")
finally:
# Not cleanup. Deletion. There is no profile to reset, no lock file
# to reap, and no question about whose soffice that was.
sbx.kill()The two clocks are deliberate. ttl_seconds is an idle timeout, not a deadline — the reaper measures time since last activity, so a converter chewing through a malicious spreadsheet never accumulates idle time and is never reaped. It is least effective against exactly the case you are worried about. Bound the wall clock yourself, and prefer exec_stream for long work, which honours long timeouts more reliably than one-shot exec does today. The size probe runs before the read because amplification goes in the direction you do not expect: a small presentation can render to an enormous PDF.
The filename is ours, sniffed from the bytes, because an uploader-supplied filename reaching a shell is a separate category of bad day and a Content-Type header is a claim typed by a client. And the error is about the conversion, not the document. You almost never know an input was malicious — you know your converter did not finish, produced nothing, or produced too much. 'We could not convert this file' is true and actionable. 'This document appears to be an attack' is a claim you cannot support, and the first time you say it to someone whose 2007-era template has an odd embedded object, you will wish you had not.
Network posture: the part you have to arrange yourself
I need to be straight about this, because it is where a post like this usually slides into implying a feature. PandaStack sandbox egress is open by default. A few targeted DROP rules exist for known-abuse protocols, and that is not a default-deny network, and I am not going to pretend otherwise. If your threat model says a converted document must not reach the internet or your metadata service — and for a conversion service it should — that is a posture you configure, not one you inherit. What the platform gives you is the place to put it: a netns, veth pair and tap per sandbox, so per-job policy is expressible per job rather than as a pool-wide compromise.
- Decide whether the converter needs egress at all. For document conversion the honest answer is almost always no. Every legitimate external reference — a logo, a web font, a linked image — can be resolved outside the sandbox by a fetcher that validates the URL, caps the size, refuses private ranges, and writes the bytes in as a file. Offline is also faster and more deterministic.
- Treat the link-local metadata range as a named target, not a side effect of a general rule. It is the highest-value destination for an SSRF in a renderer: on most clouds it hands the host role's credentials to anything that asks politely with a GET.
- Verify from inside the guest, not from your config. A preflight that curls the metadata endpoint and a public URL and fails the job if either answers is five lines, and turns an assumption into an assertion.
- Give the converter no credentials: no service account token, no instance role, no database URL. Its only job is turning bytes into other bytes.
- The output is still untrusted. A PDF that came out of a hostile docx is a hostile PDF, and if you rasterise or index it you are back at the top of this post with a different parser.
Shared pool, container per request, microVM per request
| Property | Shared soffice pool | Container per request | microVM per request |
|---|---|---|---|
| Blast radius of a parser bug | Every tenant's documents, the pool's credentials, the host | One request's namespace, then the shared host kernel under it | One guest kernel and one disk image, both discarded with the job |
| Hung-converter recovery | Find the PID, work out whose job it is, kill it, hope the profile lived | Kill the container; what it leaked into host kernel structures stays | Destroy the machine; there is nothing left to reset or inspect |
| Startup cost per request | None after the first job — the entire reason pools exist | Container start plus the suite's full startup, every request | Snapshot restore of an already-warm suite: p50 179 ms, p99 203 ms |
| State carried between jobs | Profile, caches, lock files, temp dirs, descriptors, environment | None by design, if the image is read-only and nothing is mounted in | None: memory and disk are copy-on-write copies of the snapshot |
| Network containment | Pool-wide, in your service's own network position | Per-container namespace, one shared policy, enforced by the host kernel | Per-sandbox netns from 16,384 pre-allocated /30s; policy is yours to set |
| Kernel shared with the untrusted parser | Yes, and with your API process too | Yes: the host kernel is both boundary and attack surface | No: each conversion gets its own 5.10 guest kernel |
The operational rules that survive contact with real uploads
- Convert out of band, never in the request path: the upload handler writes bytes to object storage and enqueues a job. A conversion inside an HTTP request is a hang with a direct line to your connection pool.
- One machine per conversion, destroyed in a finally block — not per tenant, per batch or per worker lifetime. The unit of isolation should be the unit of untrusted input.
- Bound the wall clock twice and expect the inner one to fire: timeout with SIGKILL in the guest, plus a caller-side deadline for when the guest is too wedged to honour its own.
- Cap the output before you read it, with a stat inside the guest, so amplification cannot become memory pressure on your side.
- Never trust the extension or the Content-Type. Sniff from the bytes, map to an allowlist, and pick your own filename; the uploader's string should never reach a path or a shell.
- Reject the conversion, not the document: report that you could not convert the file, because you usually cannot tell and sometimes the customer is right.
- Assert the posture from inside the guest on every bake rather than trusting the Dockerfile: macro level, metadata unreachable, public egress closed, shell-escape off in any LaTeX toolchain.
When this is overkill
If every document reaching your converter was produced by your own templating code, this is a lot of machinery for a problem you do not have. Keep the pool, keep a recycle limit — the descriptor leak is real whoever authored the input — and spend the effort elsewhere.
The line is whether a stranger can choose the bytes. The moment the answer is yes — a customer upload, an inbound email attachment, a file fetched from a user-supplied URL — your conversion service is an interpreter for an attacker-controlled program, and the only question is how much of your infrastructure that program holds while it runs. A microVM per conversion is not a clever answer; it is the boring one. Own kernel, own namespace, no credentials, no survivors. What makes it practical is that becoming a converter happens once, at bake time. A fifth of a second for a boundary you throw away is a good price, and the alternative is a runbook whose first step is finding out whose soffice that is.
Frequently asked questions
Can't I just strip the macros and external references before converting?
You can, and you should not rely on it, for a reason that is annoying but structural: a sanitiser is a parser. To remove macros from a docx you must open the zip container, walk the parts, parse the XML, work out which relationship targets are external, and rewrite the package. Every one of those steps is the dangerous operation you were trying to avoid, on the same hostile bytes, inside the same trust boundary, usually by a library far less battle-tested than the suite you were worried about. You have not removed a parser from your attack surface; you have added one. There is a second problem that bites in production rather than in a pentest: sanitiser and converter disagree. Your stripper walks the parts it knows about; the converter reads parts your stripper has never heard of — legacy binary streams, vendor extensions, relationships declared somewhere your rewriter did not look. A reference that survives because the sanitiser missed its container is one the converter will happily resolve. Sanitising is fine as defence in depth. It is not a boundary. If you do it, do it inside the sandbox, where being wrong is cheap.
Isn't LibreOffice in a hardened container with seccomp and no network good enough?
For a lot of threat models, genuinely yes, and I would rather you shipped that than nothing. A read-only image, a non-root user, dropped capabilities, a seccomp profile, noexec mounts and a default-deny network policy stops most opportunistic attacks and all of the lazy ones. Do it whether or not you also use VMs. The limitation is what the controls are made of. Namespaces, cgroups, seccomp and LSM policy are host-kernel features, and the untrusted font parser, XML reader and image decoder all talk to that same kernel, so the boundary and the attack surface are the same code. A seccomp profile shrinks the surface; it does not change what the surface is. Containers also leave untouched the operational half, which is the half that pages you: a hung soffice in a container is still a process on your host that something must decide to kill, its dirty page cache is still flushed under the host's writeback thresholds, and the shared-pool state problems reappear the moment you reuse a container to avoid paying startup twice. The microVM argument is not that containers are useless. It is that here the recovery action you want is 'destroy the machine', and only one boundary offers that honestly.
Does a VM per request mean I pay the office suite's startup time on every conversion?
No, and this is the single thing that makes the pattern viable rather than academic. PandaStack keeps no warm pool of idle VMs; every create restores a baked Firecracker snapshot. A template's first spawn cold-boots, around three seconds, and the agent snapshots the running machine at that point. Every create after that is a restore, measured at p50 179 ms and p99 203 ms on our fleet. Guest memory comes back through a private mapping so pages fault in lazily, and the disk is a reflink clone. That lets you move the expensive part of a conversion worker into the template: bake the image so soffice starts at boot, the profile is materialised and the font cache is built, and the snapshot captures a machine where all of that already happened. Restoring it gives you a warm, listening suite in a fifth of a second — the pool's startup economics without its state carryover, normally the two ends of a trade-off. The constraint is that guest sizing is baked in: Firecracker cannot change vCPU or RAM at restore, so per-request cpu and memory are overridden to match the template. Handle that with two or three size tiers routed by input size.
What actually stops a hung soffice, if the sandbox TTL does not?
Be clear on what ttl_seconds is, because the name misleads in a dangerous direction: it is an idle timeout. The reaper compares the current time against the sandbox's last activity and deletes it when the gap exceeds the TTL. A continuously busy workload never accumulates idle time and is never reaped, and a converter grinding through a hostile spreadsheet is the canonical continuously-busy workload — so the idle TTL is least effective against exactly the case you built the sandbox for. Bound the wall clock yourself in two places. Inside the guest, wrap the conversion in timeout with SIGKILL, because a suite waiting on a modal dialog no human will ever see will not be reasoned with by SIGTERM; timeout exits 124 when it fires and the shell reports 137 for a killed child, so your caller can tell a deadline from a failed conversion. On your side, set a deadline that deletes the sandbox outright, because a sufficiently wedged guest may not honour its own limits. For long commands prefer exec_stream over one-shot exec, which is more reliable with long timeouts today. All of this is less stressful than it sounds, because the recovery action is unconditional: you are deleting a machine, so there is no state to assess.
What about pandoc and Ghostscript — same problem, same answer?
Same answer, different reasons. Pandoc's risk is not a legacy C++ parser; it is that pandoc is extensible on purpose. Lua and external filters execute code by design, raw passthrough blocks let input carry constructs straight into the output format, and a LaTeX output path hands you a typesetting engine whose shell-escape feature is command execution as a documented capability. At the time of writing pandoc ships a sandbox mode that restricts filesystem access; read its current documentation rather than my paragraph, because that scope changes between releases. Ghostscript is the more classic case: a PostScript interpreter is an interpreter, PostScript is a programming language with file and pipe operators, and the history of its safety mode is a history of bypasses found in it. Both inherit the output problem too. A chain that goes docx to PDF to raster images runs three parsers over bytes an attacker influenced at every stage. So one template per toolchain, created and destroyed per job, is cheaper to reason about than one machine with the whole conversion zoo installed.
Keep reading
- Sandboxing untrusted uploads: ImageMagick, ffmpeg, LibreOffice — The decoder half of the same upload path — memory-safety bugs in media parsers rather than office-format features.
- Per-tenant PDF and invoice generation in microVMs — The outbound case: you own the template engine but the customer writes the template.
- Agents reading PDFs they found on the internet — What changes when the thing opening the document is a model with tools rather than a conversion endpoint.
- Controlling network egress for untrusted code — How to actually build the default-deny posture this post says you have to arrange yourself.
- PandaStack sandboxes — Snapshot-restore microVMs with their own guest kernel and netns — the disposable converter this post keeps describing.
Related posts
- A Spreadsheet Is a Programming Language Your Users Don't Call One
Your users write programs in your product every day. They spell them =SUMPRODUCT(...) instead of def main(), and that is the only difference. Then someone asks for PDF export and you shell out to an office suite.
- A Checkpoint Is a Program: Isolating Untrusted Model Weights
You downloaded a checkpoint from a model hub and called torch.load. Congratulations: you ran a stranger's Python as yourself. Here's the threat model, and the disposable-microVM ingest pipeline that turns hostile weights into boring data.
- How to Build a Remote Code Execution API Safely
An endpoint that runs arbitrary user code is a remote code execution vulnerability you shipped as a feature. The only sane boundary is a fresh microVM per request.
- Build a Headless Browser Screenshot Service on MicroVMs
A headless browser is a loaded gun that renders JavaScript. Run each screenshot/PDF/OG-render job in a throwaway Firecracker microVM with locked-down egress, and let it die after the shot.
- AI Contract Review Agents: Isolating Privileged Documents in MicroVMs
An LLM that hallucinates a clause is embarrassing. An LLM whose sandbox leaks one client's term sheet into another client's session is a malpractice claim with your firm's name on it.
49ms p50 cold start. Fork, snapshot, and scale to zero.