Parsing Untrusted Binary Protocols: A Length Field You Believed
I build PandaStack, an open-source Firecracker microVM platform, and the binary-protocol conversation goes better than the XML one. Someone has an endpoint that speaks a protocol with a number in it: MQTT from a device fleet, Modbus/TCP from a plant floor, DICOM C-STORE from a hospital PACS, gRPC from a partner, or — most often — an in-house framing protocol designed in 2017 with a two-byte length field and a comment reading `TODO: validate`.
I ask what the trust boundary around the parse is, and the answer is usually good: they fuzzed it, with a corpus in the repo and a week of AFL++ behind it. Then I ask which parser was fuzzed, and it is theirs — the framing layer they wrote — not the vendor SDK underneath, which arrives as a static library with a header and a PDF.
Text formats fail by doing too much: XML ships an interpreter in the specification, so an XXE is the parser conforming for the wrong author, and that argument is in Parsing Untrusted XML: The Format That Ships an Interpreter. Binary formats have no interpreter and no features to abuse. They fail by arithmetic.
The memory-safety CVEs in protocol parsers rhyme to the point of monotony. The wire said how long something was, and the parser allocated or copied based on that number. A 16-bit length in a 12-byte buffer. A negative count that became an enormous `size_t`. An offset pointing back into its own header. A nested element that outgrew its parent, with nobody checking containment. The attacker does not need a clever payload — they need one integer you did not range-check.
The wire format's length field is a suggestion from someone who wants your process.
The bug is always arithmetic
Six shapes cover nearly everything, and none are exotic. Each is a comparison that is either absent or written in a form that does not hold.
Unchecked length-prefix allocation
The original. Read a length, copy that many bytes, never compare it against how many bytes actually arrived. It is a write overflow when the destination is a fixed buffer and a read overflow when the source is the frame — and the read version is worse than it sounds, because the data goes back out to the sender.
# The entire exploit, as received. There is no shellcode in it.
# A 2-byte type, a 2-byte big-endian length, then the body.
00 0a ff f0 48 49
\___/ \___/ \___/
| | the two body bytes you actually received
| declared length = 0xfff0 = 65520
type = 10
# What the parser did, and it is four lines because it always is:
#
# uint16_t n = (p[2] << 8) | p[3]; /* 65520 */
# char *buf = malloc(n); /* succeeds -- 64 KiB is nothing */
# memcpy(buf, p + 4, n); /* reads 65518 bytes past the frame */
# emit(buf, n); /* and hands them back to the sender */
#
# Nothing in that is a parsing mistake. It is arithmetic performed on a
# number chosen by someone else. The 65518 bytes after the frame are
# whatever the allocator handed out most recently: another connection's
# buffer, a session key, the plaintext of the frame before this one. This
# is Heartbleed's exact shape, and the missing lines are two comparisons:
#
# if (avail < 4) return ERR_SHORT; /* header first, or the */
# if (n > avail - 4) return ERR_TRUNCATED; /* subtraction underflows */
#
# Now the comparison people DO write, which is wrong in a different way:
#
# if (off + n > total) return ERR; /* off+n wraps; the check passes */
# if (n > total - off) return ERR; /* subtract on the side you trust */
#
# Same bug, one layer up. The first form is a sum of two attacker-adjacent
# numbers inside a fixed-width type. The second never adds anything.
The exploit contains no payload: it is a frame that is honest about its type and lies about its size, and every byte of value comes from your address space rather than the attacker's. And `off + n > total` is the check a careful engineer writes, which is still wrong — it adds two numbers inside a fixed-width type before comparing. Subtraction on the side you control cannot wrap.
Signed and unsigned, the same bug wearing a type
A length field is unsigned on the wire and lands in a signed variable somewhere. A 32-bit field holding `0xFFFFFFF0` read into an `int` is `-16`. It passes `if (n > MAX) reject` with room to spare, then reaches `memcpy`, `read` or `malloc`, where the parameter is `size_t` and `-16` sign-extends to about eighteen quintillion.
No language saves you. Go's `int(binary.BigEndian.Uint32(b))` is negative on a 32-bit build above two billion, and Java has no unsigned 32-bit integer at all, so every `int` holding a four-byte length is one `Integer.toUnsignedLong` away from correct. Widen on the way in, into a type that cannot represent a negative or wrap at any value the field can hold, and never narrow until every bound is checked.
Offsets that point into their own header
Formats with internal pointers have a second class entirely. An offset table whose third entry points at byte four of the header is not malformed in any way the format can express — it is a legal number in a legal field. The parser follows it, reinterprets its own header as a record, and either loops until the budget is gone or builds an object out of structure it wrote itself. DICOM, TIFF, PE and half the device-telemetry formats I have been shown carry offsets like these.
TLV containment, the one teams miss
Nested type-length-value is the most common structure in the business and has the subtlest invariant. A child's declared length must not exceed the bytes remaining in its parent — not merely the bytes remaining in the buffer. Checking against the buffer is the natural mistake, it passes every test you wrote, and it lets a child reach past its parent's end to consume its siblings, the trailer, or the next message in the stream.
You almost certainly parse one already. ASN.1 BER and DER are TLV, which makes X.509 certificates, SNMP, LDAP, Kerberos and most 3GPP signalling TLV — and BER's indefinite-length form replaces the length with a terminator, a length field that says it will tell you later. Unbounded nesting plus deferred lengths has its own long CVE history. It is why the container in the Go snippet below is decoded from its parent's own slice: containment by construction means there is no comparison to forget.
Amplification: frame size and decompression
The sender picks the ratio in two places. A declared frame size lets a tiny packet reserve a large buffer, and a hundred connections each declaring the maximum legal frame is a memory bill paid before any authentication. Compression is worse, because the work happens before the limit can mean anything: you decompress, and only then know what you decompressed. An absolute output cap and a ratio cap are both required, enforced by the decompressor as it streams rather than asserted about the result. The archive version of this argument is Extracting Untrusted Archives: Zip Bombs, Zip Slip, Symlink Escape.
Recursion depth
Nesting is a stack-depth question, and stack exhaustion is not a catchable error in most runtimes: a crash in C, a crash in Go at the goroutine stack limit, a `StackOverflowError` in Java that leaves the heap in a state nobody designed. Several protobuf implementations ship a default recursion limit around a hundred for exactly this reason — check yours rather than assuming, because a thirty-byte message can be thirty levels of wrapper. The fix is a depth counter checked before descending, plus one global budget spanning every depth, because per-element caps are individually satisfiable: input that respects all of them and repeats ten million times is still a denial of service.
The protocols with the most exposure have the oldest parsers
The arithmetic above is universal. The tragedy is that the protocols where it matters most have been in the field longest, in the places least able to patch.
Modbus, designed when the threat model was the cable
Modbus is from 1979 — a serial protocol for talking to programmable logic controllers over RS-485, which got a TCP header stapled on in the late nineties: the MBAP header, carrying a transaction identifier, a protocol identifier, a two-byte length and a unit identifier. That is the entire modernisation. There is no authentication — not weak authentication, none, because the design predates the question. "Is this master allowed to write this register" is not something the protocol can express, so the only available answer is network position: whoever reaches port 502 is the master. The threat model when it was specified was someone tripping over the cable.
The parser specifics are exact. The MBAP length counts the unit identifier plus the protocol data unit — a number from the wire describing bytes you may not have received yet over a stream socket, so sizing a buffer from it rather than bounding it to what arrived is the first bug. Function codes carry their own caps, read holding registers topping out at 125 registers and write multiple at 123, and a response's byte count can disagree with its register count: trust one, allocate from the other, and the mismatch is reachable. These are tiny frames, so the parsers are correspondingly small, written in C, compiled into firmware, and frequently older than the engineer maintaining them.
MQTT, parsing varints before CONNECT is validated
MQTT's Remaining Length is a variable-length integer: up to four bytes, seven payload bits each, a continuation bit in the top bit, maximum 268,435,455. A broker decodes it before it knows what packet it holds — which is before CONNECT has been validated, which is before authentication. Every broker on the internet decodes attacker-supplied varints from unauthenticated sockets as its first act.
The specification caps the encoding at four bytes and requires a decoder to reject a longer one, which tells you somewhere there is a decoder that does not — and a varint loop with no cap is an unbounded shift. MQTT 5 then adds a properties block length-prefixed with its own variable-byte integer, with individual properties carrying lengths of their own: nested varints inside a varint-delimited region, which is TLV containment in a protocol people call lightweight. Lightweight is a statement about the wire, not about the parser.
DICOM, which is a parser and a network stack and a codec
DICOM packs three attack surfaces into one specification: a file format, a network protocol, and transfer syntaxes that embed JPEG, JPEG-LS, JPEG 2000 and RLE. A gateway accepting C-STORE from a modality therefore runs a TLV parser, an association state machine and an image codec stack — and association negotiation happens before anything resembling authentication, since the called application entity title is a string you compare, not a credential you verify.
The TLV details come with a twist. A data element is tag, value representation, length — and the length is two bytes or four depending on the representation, which is itself read from the wire, so the wire chooses how wide the next length field is. The undefined-length value `0xFFFFFFFF` is legal for sequences and encapsulated pixel data, terminated by a delimitation item instead of a count: a length field that declines to say how long the thing is, in a format whose consumers include a JPEG 2000 decoder. That is how a hospital ends up with a wavelet codec on the network. The compliance framing lives in Running Code Over PHI Without Expanding Your HIPAA Blast Radius; the point here is purely arithmetic.
Protobuf, memory-safe in two languages and C++ in the rest
Protobuf removes most of this class and not all of it. The Go and pure-Java implementations are memory-safe: a bad length is an error return, not a corruption. The reference C++ library is C++, and the fast Python and Ruby paths are bindings over it, so "we use protobuf" says nothing about which implementation is chewing the bytes in production.
Two things survive even a memory-safe runtime. Nesting is still a stack-depth question. And allocation is still driven by sender-chosen counts — a repeated field's element count comes off the wire, so a small message can ask a decoder to build a very large object graph, and a few hundred in flight is a memory curve rather than a bug. Memory safety converts a corruption into an exhaustion: an enormous improvement, and not the same thing as containment.
The in-house protocol, which nobody fuzzes
The one I see most often is in no specification at all. A device fleet speaks a framing protocol somebody designed in an afternoon: magic bytes, a version, a two-byte length, a type, a body, a CRC. It has worked for years and has never been fuzzed, because fuzzing is something you do to formats with names.
Why the usual answers are partial
None of these are wrong. Each is incomplete in a specific way worth naming, because each gets presented in review as though it closed the class.
- Fuzzing finds the bugs you have, not the ones you link. Coverage-guided fuzzing is the highest-value thing you can do to a parser and you should run it continuously. But it instruments what you built it against, and the vendor SDK compiled from source you were not given is not in the corpus — nor is the codec it dispatches to for transfer syntax seven. A clean run is a statement about your framing layer.
- A memory-safe rewrite is correct and will take years. Rewriting in Rust or Go is the right long-term answer and I would start today. It does not retire the C library speaking the vendor's proprietary extensions, the firmware you do not build, or the codec you call through FFI. In the meantime the old parser is still in the path, and in the meantime is where you live.
- A seccomp filter does not help when the bug is a heap overwrite. Narrowing the syscall set is genuine defence — see seccomp explained for developers: filtering syscalls to shrink the kernel attack surface — but think about what it stops. A corruption that rewrites a function pointer and then calls `write` on a descriptor the process legitimately holds is entirely inside the allowed set. Seccomp bounds what a process can ask the kernel for, not what it does with handles it already had.
What they share: each makes the parser less likely to be wrong. None changes what happens when it is. That is a different control, and it is architectural.
First, a parser that actually range-checks
Do this before anything architectural, because it is cheap and removes most real exposure. The snippet below is the one to copy: a declared-length ceiling that is a property of your format rather than of your buffer, containment enforced by re-slicing instead of by a comparison, a recursion cap checked before descent, and a single byte budget spanning the whole parse.
// Package wire decodes nested length-prefixed frames from a hostile peer.
// Every rule in here exists because some CVE did not have it.
package wire
import (
"encoding/binary"
"errors"
"fmt"
)
var (
ErrTruncated = errors.New("wire: declared length exceeds available bytes")
ErrTooLarge = errors.New("wire: declared length exceeds format ceiling")
ErrTooDeep = errors.New("wire: nesting deeper than the cap")
ErrTooMany = errors.New("wire: element count exceeds the cap")
ErrBudget = errors.New("wire: decode budget exhausted")
)
const (
hdrLen = 6 // 2 bytes type + 4 bytes length
maxValueBytes = 1 << 20 // ceiling: no single value may be bigger
maxDepth = 8 // the deepest nesting this format has
maxElements = 4096 // per container
totalBudget = 8 << 20 // across the WHOLE parse, all depths
)
type TLV struct {
Type uint16
Value []byte // a subslice of the input; never a fresh allocation
Children []TLV
}
// decoder carries the budget across the whole parse so a peer cannot stay
// under every individual cap and still make us do unbounded work.
type decoder struct{ budget uint64 }
// Parse decodes one buffer. buf is the ONLY authority on how many bytes
// exist. No number read off the wire is ever permitted to exceed it.
func Parse(buf []byte) ([]TLV, error) {
d := &decoder{budget: totalBudget}
return d.container(buf, 0)
}
func (d *decoder) container(buf []byte, depth int) ([]TLV, error) {
// 1. Depth, first and unconditionally. A recursive descent over an
// attacker-chosen nesting is a stack-exhaustion primitive, and a
// stack overflow is not a recoverable error in most runtimes.
if depth > maxDepth {
return nil, ErrTooDeep
}
var out []TLV
for off := 0; off < len(buf); {
// 2. The header must fit BEFORE you read any of it. A 6-byte read
// at len(buf)-1 is the panic you ship on a Friday afternoon.
if len(buf)-off < hdrLen {
return nil, fmt.Errorf("%w: %d trailing bytes", ErrTruncated, len(buf)-off)
}
typ := binary.BigEndian.Uint16(buf[off : off+2])
// 3. Widen, never narrow. int(binary.BigEndian.Uint32(...)) is
// NEGATIVE for values above 2^31 on a 32-bit build -- and a
// negative length sails through "if n > max" and then becomes a
// very large size_t in whatever C library you hand it to. uint64
// cannot wrap at any value a 32-bit field can hold.
n := uint64(binary.BigEndian.Uint32(buf[off+2 : off+hdrLen]))
off += hdrLen
avail := uint64(len(buf) - off)
// 4. Compare against REMAINING bytes, by subtraction on the trusted
// side. This is the check whose absence is most of the CVE list.
if n > avail {
return nil, fmt.Errorf("%w: declared %d, have %d", ErrTruncated, n, avail)
}
// 5. A ceiling that is a property of your format, not of the buffer.
// "it fits" and "it is plausible" are different statements, and
// only the second one bounds what a 100 MiB upload can ask for.
if n > maxValueBytes {
return nil, fmt.Errorf("%w: %d > %d", ErrTooLarge, n, maxValueBytes)
}
// 6. The global budget, decremented by bytes that really exist.
if n > d.budget {
return nil, ErrBudget
}
d.budget -= n
val := buf[off : off+int(n)]
off += int(n)
if len(out) >= maxElements {
return nil, ErrTooMany
}
node := TLV{Type: typ}
if isContainer(typ) {
// 7. Containment by construction. Children are decoded from the
// parent's OWN slice, so a child physically cannot name a byte
// outside its parent -- there is no comparison to forget. The
// classic nested-TLV bug is a child whose declared length
// exceeds its parent's; here that is ErrTruncated, for free.
kids, err := d.container(val, depth+1)
if err != nil {
return nil, fmt.Errorf("type %d: %w", typ, err)
}
node.Children = kids
} else {
node.Value = val
}
out = append(out, node)
}
return out, nil
}
// isContainer is a closed set, declared by you. Never "any type whose first
// byte looks like a tag" -- that is the wire choosing your control flow.
func isContainer(typ uint16) bool {
switch typ {
case 1, 2, 7:
return true
}
return false
}
Four things matter more than they look. Widening to `uint64` means no arithmetic in the function can wrap at any value a 32-bit field can hold. Comparing `n > avail` rather than `off + n > len(buf)` means nothing is ever added. Decoding a container from `val` — the parent's own slice — makes containment structural, because a child cannot address bytes it was not handed. And the budget lives on the decoder rather than in the loop, so an element count satisfying every local cap still terminates.
Then notice what the file does not protect. It protects the frames this code parses. The moment `isContainer` returns false and you hand `node.Value` to the vendor decoder, you are trusting somebody else's arithmetic — and that handoff is what the rest of this post is about.
The architecture: bytes in, a validated value out
The shape I would build, as properties rather than a diagram, because properties survive your framework choice.
- The trusted frontend owns both sockets: the one from the peer and the one to the historian, broker, database or PLC. The parser gets neither. This is the most important property here and the one people skip, because the obvious design is to move the whole connection handler into the sandbox.
- The frontend does exactly one piece of parsing: framing. Read the length, range-check it with the comparisons above, cap it, accumulate that many bytes, hand them over. That is twenty lines you can fuzz exhaustively, and it is the irreducible minimum — you cannot know where a frame ends without reading a number.
- Bytes cross as bytes, written to a file in the guest. Never a URL, never a shared buffer, never a handle. If the trusted side passes a location rather than contents, you have handed the parser a fetch primitive and the boundary is decorative.
- The parser holds no credential. Not a scoped one. None: no device certificate, no historian password, no broker token, no cloud role, nothing in the environment. The question about a parse process is not what it may do but what is in reach, and the right answer is one frame and a temp directory.
- Only a normalised, schema-validated value crosses back: JSON conforming to a schema you wrote, from fields you explicitly extracted, numbers range-checked and strings capped on the trusted side. Handing back a live parse tree defeats the exercise — a node that still knows how to resolve an offset is not a value, it is deferred work that will run in your trusted process.
- The wall clock is yours and it lives inside the guest. A `timeout` and a `ulimit -t` there are the real bound, and be exact about why: `ttl_seconds` on PandaStack is an IDLE TTL, so it reaps an abandoned sandbox and never a busy one — and a parser spinning in an amplification loop is extremely busy. The in-guest fence is load-bearing; the platform TTL is the backstop for a connection handler that died holding a sandbox.
What that buys: a guest kernel nobody else shares and which will not exist in a moment, a device model of a handful of virtio devices rather than four decades of emulated hardware, and a teardown that is a delete rather than a reset. And — the one I rank first operationally — a per-peer audit trail: a parse that does something interesting dies with its own kernel and leaves a record attributable to one peer, instead of a segfault in a shared worker that served four hundred devices this minute.
One VM per connection, per message, or per batch
This is the design decision, and the traffic sets it rather than taste. On PandaStack every create is a snapshot restore with no warm pool, at p50 179 ms and p99 203 ms end to end, so the question is how often you pay it and how long the guest then lives.
Per connection suits long-lived sessions: an MQTT client that stays connected, a Modbus master polling for weeks, a DICOM association. One create per connection is free against a session measured in hours, and you get what a per-message design cannot give you — protocol state, which a stateful stream parser needs across frames. The cost is memory held for the duration: a `code-interpreter` guest is 2 GiB at $0.0162 per GiB-hour, about 3.2 cents an hour per connection, plus CPU on the seconds actually burned. A few hundred industrial peers is a few dollars an hour. A hundred thousand MQTT clients is not something you will do this way.
Per message suits discrete, independent, low-rate work: batch DICOM ingest, a filing that arrives twice a day, a pcap replayed for forensics, a device check-in. The deciding arithmetic: a guest alive half a second per message means a hundred messages a second needs about fifty concurrent guests, and since Firecracker cannot change vCPU or RAM at snapshot restore each is the 2 GiB the template baked — roughly a hundred gigabytes of RAM. Per-message VMs are an ingest pattern, not a bus pattern, and anyone saying otherwise has not multiplied.
Per batch is the compromise, and the only honest way to draw a batch is by trust domain rather than by convenience: one guest per sending institution, per device, per tenant, not one per thousand messages from everybody. The second message in a reused guest inherits the first one's consequences, and the whole point was that there were none. Five hundred studies from one hospital is a defensible unit. Five hundred from fifty hospitals is one compromised study with access to forty-nine others.
| Dimension | In-process parse | VM per message | VM per connection | VM per batch |
|---|---|---|---|---|
| Blast radius of a parser bug | Your service, with every credential and peer buffer in it. | One message, in a kernel deleted when it returns. | One peer's session, for that session's life. | One trust domain's batch. |
| Protocol state across frames | Natural. | None. Stateless messages only. | Natural — the main reason to pick it. | Natural within the batch. |
| Added latency | None. The honest reason most people stop here. | Snapshot restore per create: p50 179 ms, p99 203 ms. | One create per connection, amortised away. | One create per batch. |
| Cost shape | Nothing beyond the service you run already. | A create plus a short life; CPU on seconds burned, no per-request price. | 2 GiB held: ~3.2 cents per connection-hour. | 2 GiB for the batch window. |
| Where it breaks | The first memory-safety bug in a linked SDK. | Message rate: ~100/s needs ~50 guests, ~100 GiB. | Connection count: thousands of peers is a RAM bill. | A batch drawn across trust domains, which is most batches. |
Per-connection, in code
The frontend keeps both sockets. The guest gets a file and returns a JSON object. Note what the in-guest fence does: `ulimit -t` bounds CPU seconds, which is what an amplification attack actually spends, and the guest's own `timeout` is the real wall clock, because the SDK's one-shot `exec` is not where a long bound belongs.
import json
from pandastack import Sandbox
from pandastack.exceptions import CommandFailed
# The parser, as a file the guest will run. It links the vendor SDK, it is
# half C, and it holds no credential because there is none in reach.
PARSE_PY = r"""
import json, sys
import vendor_sdk # the static library whose source you have a PDF of
frame = open(sys.argv[1], "rb").read()
msg = vendor_sdk.decode(frame) # the part you do not trust
out = { # the part you do: explicit fields only
"unit": int(msg.unit_id) & 0xFF,
"function": int(msg.function) & 0xFF,
"registers": [int(v) & 0xFFFF for v in msg.registers][:125],
}
with open(sys.argv[2], "w") as fh:
json.dump(out, fh) # a value, not a live parse tree
"""
# The in-guest fence is the REAL bound, and not as a style preference:
# ttl_seconds is an IDLE ttl, so it reaps an ABANDONED sandbox and never a
# busy one -- and a parser stuck in an amplification loop is very busy.
# ulimit -v is address space, not RSS, so give CPython room; ulimit -t is CPU
# seconds, which is what the amplification actually spends; timeout owns the
# wall clock.
FENCE = (
"cd /work && umask 077 && "
"ulimit -v 524288; ulimit -t 5; ulimit -f 4096; ulimit -c 0; "
"exec timeout --signal=TERM --kill-after=2s 10 "
"python3 -I parse.py {inp} {out}"
)
class HostileFrame(Exception):
"""What you page on. The peer gets a closed socket and nothing else."""
class ParserVM:
"""One microVM per CONNECTION. The socket never leaves the frontend; only
bytes cross in and only a normalised JSON object crosses back."""
def __init__(self, peer: str, proto: str):
self.seq = 0
self.sbx = Sandbox.create(
template="code-interpreter", # 2 GiB / 8 vCPU, fixed by the snapshot
ttl_seconds=900, # IDLE backstop for an abandoned VM
metadata={"job": "wire-parse", "proto": proto, "peer": peer},
)
self.sbx.filesystem.write("/work/parse.py", PARSE_PY.encode())
def parse(self, frame: bytes) -> dict:
"""frame has already passed the frontend's length check: it is exactly
as many bytes as its header claimed, and under the ceiling."""
self.seq += 1
inp, out = f"/work/f{self.seq:06d}.bin", f"/work/f{self.seq:06d}.json"
self.sbx.filesystem.write(inp, frame) # bytes in. never a URL.
# Keep the HTTP-side bound under 30s and let the guest's own timeout
# fire first; for anything genuinely long, use sbx.exec_stream().
r = self.sbx.exec(FENCE.format(inp=inp, out=out), timeout_seconds=25)
if r.exit_code != 0:
# 124 timeout, 137 OOM-killed inside a guest that is nobody
# else's problem, anything else the SDK rejecting the frame.
raise HostileFrame(f"seq={self.seq} exit={r.exit_code} {r.stderr[-800:]}")
raw = self.sbx.filesystem.read(out)
if len(raw) > 65536:
raise HostileFrame(f"seq={self.seq} normalised output implausible")
return validate(json.loads(raw)) # YOUR schema. Still hostile data.
def close(self):
self.sbx.kill() # unconditional. A leaked sandbox bills by the
# GiB-hour, patiently, until someone notices.
def serve(conn, peer):
vm = ParserVM(peer=peer, proto="modbus-tcp")
try:
for frame in read_framed(conn): # the frontend's ONE piece of parsing
value = vm.parse(frame)
historian.write(value) # the plant-side socket lives HERE,
# in the frontend, never in the guest
finally:
vm.close()
The economics are not the interesting part, but people ask: $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same rate for every class, CPU billed on CPU-seconds actually burned and memory on committed GiB-hours, with no per-request price and none coming. What you buy is that a frame which corrupts the vendor SDK's heap becomes a deleted VM and a log line with a peer address on it.
What the network boundary actually is, and what it is not
Be precise here, because this is where sandbox marketing goes vague. Each sandbox gets its own Linux network namespace with a veth pair and a tap device, from a pool of 16,384 pre-built slots in `10.200.0.0/16`. Egress is open by default — there is no default-deny policy, and I would rather say that plainly than let you discover it.
What exists instead is a small denylist in the root namespace's FORWARD chain, inserted ahead of the egress accepts. Pool-to-pool traffic is dropped, so no sandbox reaches another sandbox's subnet and cross-tenant scanning is closed at the host. The whole of `169.254.0.0/16` is dropped, so the cloud metadata endpoint is not reachable — on GCP it would hand out the host VM's full-scope service-account token, which makes it the highest-value target for a compromised parser and is why the rule exists. And there is a drop per well-known Stratum mining port, because free-tier abuse happens.
Your own SCADA network, historian, VPC and databases: not fenced. Say it loudly, because in an operational-technology context that is not an oversight, it is the premise. The entire reason a protocol gateway exists is that it can reach the plant network. Any vendor claiming their isolation solves your OT segmentation problem is selling you something.
Which is exactly why the architecture above keeps both sockets in the frontend. The parser VM needs no route to the PLC because it never speaks to the PLC — the frontend does, with the frontend's credentials, after validating a normalised value against a schema it owns. The sandbox's open egress then matters for one reason: it is an exfiltration channel for the bytes that parser saw. If that is in your threat model, and in OT it should be, the remaining work is yours — give the parse guest no network at all where you can, or add deny rules for your own ranges and peerings as code that runs for every sandbox. The general version is Controlling Network Egress for Untrusted Code.
The honest limits
- It does not make the parser correct. A microVM around an unchecked parser is one control where you should have two, and the wrong one to have picked first. Range-check the lengths, fuzz continuously, widen on the way in. The microVM makes the residual failure survivable, and residual is doing real work in that sentence.
- You still pay for the amplification. A guest that dies at its memory limit died alone, which is the point, but the CPU-seconds it burned getting there are billed, and nothing here stops forty thousand crafted frames arriving at once. Rate limiting and admission control are still yours, upstream.
- Per-message VMs do not survive a message rate. Around a hundred messages a second needs about fifty concurrent 2 GiB guests. If your bus does thousands, per-message is the wrong unit and no number of 179 ms creates fixes it.
- The normalised value is still attacker-influenced data. You extracted a register value from a hostile frame; it is still a number a stranger chose. If it flows into a historian, a dashboard, a control decision or an SQL statement, each of those needs its own validation. The sandbox has no opinion about your downstream.
- The latency is disqualifying for a closed loop. Roughly 200 ms to create, plus parse, plus teardown is fine for ingest, batch, north-bound telemetry, configuration import, conformance testing and forensics on captured traffic. It is not acceptable for anything in a control loop with a PLC, anywhere near a safety instrumented function, or on any path with a sub-100 ms deadline. In that class, harden the in-process parser and move the adjacent work — the historian feed, the backfill, the conformance suite — into VMs instead. Do not put a scheduler in a control loop.
- You cannot eliminate the frontend's own parsing. The framing reader in your trusted process is small and auditable, and it is still code reading a length field from a stranger. You reduced the surface by a large factor; you did not delete it.
- The guest is 2 GiB whether it needs it or not, because Firecracker cannot change vCPU or RAM at snapshot restore and a `code-interpreter` sandbox is whatever its snapshot baked. For a 12-byte Modbus frame that ratio is embarrassing — bake a leaner template, batch within a trust domain, or accept it.
The summary
Binary protocol parsers do not fail by doing too much. They fail because a number arrived on the wire and the code did arithmetic with it. Unchecked length-prefix allocation, a signed variable holding an unsigned field, an offset into its own header, a child element that outgrew its parent, an amplification ratio the sender chose, a nesting depth that is really a stack-depth question. Six shapes, and most of the CVE list.
Fix them where you can: widen on the way in, compare by subtraction against remaining bytes, enforce containment by re-slicing rather than by a comparison someone must remember, cap depth before descending, carry one budget across the whole parse, and fuzz continuously. Then be honest that the vendor SDK underneath has its own arithmetic, that its source is a PDF, and that the firmware speaking to you was compiled when the threat model was someone tripping over the cable.
So put the parse somewhere disposable. Both sockets in the trusted frontend, framing the only parsing it does, bytes handed to a guest with no credential and no route to the plant, a normalised schema-validated value as the only thing crossing back, and an unconditional `kill()`. Pick the unit from the traffic: per connection for long-lived sessions, per message for low-rate discrete ingest, per batch drawn along a trust boundary and never a convenience one. The frame that corrupts the decoder's heap then becomes a deleted VM and a log line with a peer address on it. That is not a correct parser. It is a survivable one, which for a 1979 protocol with a TCP header stapled on is the honest target.
Frequently asked questions
Is one microVM per message realistic at real protocol message rates?
Only at low rates, and the arithmetic is worth doing before you commit. On PandaStack every create is a snapshot restore with no warm pool, at p50 179 ms and p99 203 ms end to end. If a guest lives about half a second per message, sustaining a hundred messages a second needs roughly fifty concurrent guests, and because Firecracker cannot change vCPU or RAM at snapshot restore each of those is the 2 GiB the template baked — about a hundred gigabytes of RAM to parse a hundred messages a second. That is the honest ceiling, and memory is the binding constraint long before the 16,384 subnets per host become relevant. So per-message VMs are an ingest pattern: batch DICOM studies, a regulatory filing that arrives twice a day, a device check-in, a pcap replayed for forensics, anything discrete and stateless where the rate is tens per second rather than thousands. For a message bus you want one VM per connection, which pays one create per session and amortises it to nothing over a session measured in hours, or one VM per batch with the batch drawn along a trust boundary. For a genuinely high-rate stream, run the hardened in-process parser and move the adjacent work into VMs instead.
Does the microVM stop a compromised parser from reaching my PLCs or SCADA network?
Not by itself, and the division matters. Each sandbox gets its own network namespace with a veth pair and tap device, and egress is open by default — there is no default-deny policy. What exists is a denylist in the root namespace's FORWARD chain: pool-to-pool traffic is dropped, so no sandbox reaches another sandbox's subnet, and the whole of 169.254.0.0/16 is dropped, so the cloud metadata endpoint is unreachable, which closes the highest-value credential-theft path. Your own plant network, historian and VPC are not fenced. In an operational-technology context that is the premise rather than an oversight: the entire reason a protocol gateway exists is that it can reach the plant. The defence that actually works is architectural. Keep both sockets in your trusted frontend — the one to the peer and the one to the PLC or historian — so the parser never holds the plant-side connection or the credentials for it. The parser gets a file of bytes and returns a normalised JSON value. Then the sandbox's open egress matters for exactly one thing: it is an exfiltration channel for the bytes that parser saw. Close that by giving the parse guest no network at all where you can, or by adding deny rules for your own ranges and peerings as code that runs for every sandbox.
Isn't protobuf memory-safe? Why would I sandbox a protobuf decode?
It depends which implementation is actually running, and the answer is usually less comfortable than people expect. The Go implementation and the pure-Java one are memory-safe: a length that does not fit the input is an error return rather than a corruption. The reference C++ library is C++, and the fast paths in several other languages are bindings over it, so "we use protobuf" does not tell you what is chewing the bytes in production. Even in a fully memory-safe runtime, two things survive. Nesting is a stack-depth question, which is why several implementations ship a default recursion limit around a hundred — check yours rather than assuming, because a thirty-byte message can be thirty levels of wrapper. And allocation is still driven by counts the sender chose: a repeated field's element count comes off the wire, so a small message can ask a decoder to construct a very large object graph, and a few hundred of those in flight is a memory curve rather than a bug you can patch. Memory safety converts a corruption into an exhaustion. That is a genuine and large improvement, and it is not the same thing as containment, because an exhaustion in a shared process still takes down every other peer that process was serving.
Which traffic can I actually route through this pattern, and which can I not?
Split it by deadline. The pattern suits anything where a few hundred milliseconds of setup is noise against the work: batch ingest of DICOM studies or regulatory filings, north-bound telemetry heading for a historian or a data lake, historian backfill, gateway and device configuration import, protocol conformance testing against a vendor's firmware, and forensics on captured traffic where you are deliberately replaying something that already hurt you once. For all of those, a create at p50 179 ms plus the parse plus teardown is rounding error, and a per-connection or per-batch guest amortises even that away. It does not suit anything in a closed loop with a programmable logic controller, anything in or adjacent to a safety instrumented function, or any path with a sub-100 millisecond deadline. Putting a scheduler in a control loop is a bad idea on its own merits, before you even reach the question of what happens when a create is refused for capacity and returns a 503. If your traffic is in that class, harden the in-process parser properly — widen on the way in, compare against remaining bytes by subtraction, cap depth and total bytes — and move the adjacent, non-latency-sensitive work into disposable VMs instead. That split is usually the whole answer: the control path stays fast and audited, and everything downstream of it stops being a shared blast radius.
Keep reading
- Parsing untrusted XML: the format that ships an interpreter — The opposite failure mode — a text format that fails by doing too much, not by arithmetic.
- Extracting untrusted archives: zip bombs and Zip Slip — The amplification half of this post, where the attacker picks the ratio.
- Isolating a fuzzing harness in microVMs — Where the parser bugs come from, and why a crashing target wants its own kernel.
- Running IoT and embedded firmware emulation in disposable microVMs — The other end of the wire: the device whose parser you cannot patch.
- Controlling network egress for untrusted code — The deny-by-default rules you add yourself, since sandbox egress is open by default.
Related posts
- Document Conversion Is Remote Code Execution With a Progress Bar
A convert endpoint accepts an arbitrary program from a stranger and runs it with your service account and your network position. The office formats are not data. They are a feature set, and some of those features dial out.
- Processing PHI in Per-Job microVMs: Isolation for Regulated Healthcare Data
The senders are hospitals with 1998-era interface engines, the formats are containers full of compressed pixel data and XML, and the parsers are C. Give every PHI job a machine you can afford to lose — and know exactly which parts of compliance that does and does not buy you.
- A Spreadsheet Is a Programming Language Your Users Don't Call One
Your users write programs in your product every day. They spell them =SUMPRODUCT(...) instead of def main(), and that is the only difference. Then someone asks for PDF export and you shell out to an office suite.
- Per-Tenant Isolation for Vector Embedding Jobs
Your embedding worker pool parses customer PDFs, ships vectors to a model endpoint, and holds every tenant's documents in one address space. A PDF parser is a Turing-complete attack surface wearing a business-document costume. Give each tenant job its own microVM.
- Pickle Is a Virtual Machine You Have Been Downloading
pickle is a stack machine with an eval loop, and you have been downloading it from the internet as model weights. Here is the disassembly, why the allowlist promises less than it looks, and the conversion gateway that fixes it once instead of forever.
49ms p50 cold start. Fork, snapshot, and scale to zero.