Running a DICOM Pipeline in a MicroVM: Untrusted Medical Images, One Study at a Time
The ticket read: de-identification worker restarted eleven times overnight, backlog one thousand four hundred studies. It was not a memory leak. One object in one study — an ultrasound loop from a scanner model nobody at the hospital could still name — was taking the decoder down with it, and because the worker pulled studies from a shared queue into a shared scratch directory, each restart left half-written frames from three other patients' studies on disk for the next run to trip over.
That is the ordinary shape of a DICOM pipeline failure, and it is worth sitting with how many separate problems are stacked inside it. A native-code image decoder crashed on input it did not choose. The crash took down a process that was also holding three other patients' pixels. The scratch directory outlived the process that owned it. And every byte involved was protected health information, which turns an availability incident into a disclosure question: what exactly was in the core dump, and who can read it?
I build PandaStack, an open-source Firecracker microVM platform, so you can see where this is going. The argument stands without the product though, and the specifics are the useful part: what inside a DICOM file is actually executing, why "malformed" is the normal case rather than the attack case in a real archive, why the de-identifier is itself in the blast radius, and what one-microVM-per-study actually costs to run.
A DICOM file names its own decoder
The layout is a product of its era, and knowing it is what makes the rest of this concrete. A file begins with a 128-byte preamble that is allowed to contain anything at all, followed by the four characters `DICM`. Then comes the File Meta Information group — every element with group number `0002` — which is always encoded in Explicit VR Little Endian so that a reader can bootstrap. Inside that group sits one element that decides everything downstream: `(0002,0010)` TransferSyntaxUID.
That UID is not a version number. It is a dispatch table. `1.2.840.10008.1.2` means implicit VR little endian, so the rest of the dataset carries no type information at all. `1.2.840.10008.1.2.4.90` means the pixel data is JPEG 2000 and a wavelet codec is about to run. `1.2.840.10008.1.2.1.99` means deflate, so there is a decompressor in front of the parser. The file is telling your process which several thousand lines of C to execute next, and it is a file you received from outside.
The tag parser, before any codec runs
Under implicit VR there is no value-representation field on each element, so the parser must look the tag up in a data dictionary to learn whether those four bytes are a signed short, a UID string or a nested sequence. Tags with odd group numbers are private — vendor extensions, not in the dictionary — so the parser falls back to `UN`, unknown, and an `UN` element with undefined length is genuinely ambiguous about where it ends.
The standard also requires every value to be even-length, padded with a null or a space. Real files violate this constantly, so every production parser has an odd-length recovery path, and recovery paths are where the interesting bugs live. Sequences make it worse: an `SQ` element may declare undefined length and be terminated by delimiter items — `(FFFE,E000)` for an item, `(FFFE,E00D)` and `(FFFE,E0DD)` to close them. Get one of those wrong and the parser does not fail, it desynchronises, and every element after that point is read at the wrong offset with the wrong type.
Then there is pixel data itself. Element `(7FE0,0010)` can be a flat blob, or it can be encapsulated: a Basic Offset Table followed by fragments, where one frame may span several fragments and the offset table is a set of numbers the file supplies about itself. A multi-frame object whose offset table disagrees with its fragment boundaries is a classic, and the disagreement is attacker-controllable.
Then the codec, which is where the C lives
Image codecs are the most reliably exploited class of code in the history of software, for structural reasons rather than anecdotal ones: they are complex binary parsers, written in unsafe languages for speed, invoked automatically on content nobody vetted, and historically embedded in privileged processes. DICOM hands you most of that family at once, selected by a string in the file.
| Transfer syntax | What decodes it | Typical implementation | Failure mode to plan for |
|---|---|---|---|
| Implicit VR Little Endian | The tag parser itself; no codec | Pure Python in pydicom, C++ in dcmtk and GDCM | No VR on the wire, so types come from a dictionary and private tags are guessed. A desync here mis-parses everything after it. |
| Explicit VR Little Endian | Tag parser; pixel data is raw | Same | The safest case, and the one you should transcode to. Still carries odd-length values and undefined-length sequences. |
| JPEG Baseline and Extended | A libjpeg-lineage decoder | C | Decades-old C walking attacker-shaped entropy-coded data. The canonical heap-overflow surface in imaging. |
| JPEG Lossless and Lossless SV1 | A rarely-exercised lossless JPEG path | C | Low-traffic code inside a high-traffic library, which usually means less fuzzing coverage than the baseline path next to it. |
| JPEG-LS, lossless and near-lossless | CharLS | C++ | Golomb-Rice coding with run modes, plus a quantiser in the near-lossless variant. Small, fast, and not where most fuzzing effort has gone. |
| JPEG 2000 | OpenJPEG | C | Wavelets, tiles, precincts, code-blocks and markers: the largest attack surface in the set, with a long and public advisory history. Read it before deciding this is fine. |
| RLE Lossless | A segmented run-length decoder | C, or pure Python via pylibjpeg-rle | A format simple enough that people hand-roll it. Segment offsets come from the file, so one unchecked bound is an overflow in forty lines of code. |
| Deflated Explicit VR LE | zlib, then the tag parser | C | A decompressor in front of the parser. Your size caps have to apply after inflation, which is not where people put them. |
| MPEG-2, MPEG-4, HEVC | A video decoder | C and C++ | Yes, a DICOM object can contain H.265. Refuse these by UID unless you consciously decided to run a video decoder on hospital data. |
In a real archive, malformed is the normal case
If you have only ever parsed DICOM produced by one modern toolchain, the format feels strict. A PACS archive is the opposite experience. It holds a decade or more of objects written by many vendors' encoders, including ones whose authors are no longer in business, each with its own interpretation of the ambiguous parts and its own private tags carrying duplicates of the patient's name. There are retired big-endian objects, lengths that disagree with reality, sequences closed the wrong way, character sets declared in one place and ignored in another.
The operational consequence is the thing people miss. You cannot ship a strict parser, because a strict parser rejects a meaningful fraction of the archive you were hired to process. So you ship a lenient one — recovery paths on, best-effort length handling, dictionary fallbacks — and you cannot use parse failure as an attack signal either, because parse failures are your Tuesday. A deliberately malformed file arrives wearing exactly the same clothes as a 2009 ultrasound.
Every production DICOM parser is a lenient parser, and leniency is the attack surface. You do not get to choose strictness; you get to choose what the parser is standing inside when it is wrong.
Why "just run it in a container" is the shaky default
The default architecture is a worker pod with `pydicom`, `gdcm` and `dcmtk` installed, pulling study paths off a queue. Done well it has a read-only root filesystem, a non-root user, dropped capabilities, a seccomp profile and a memory limit. All of that is worth having and none of it changes the central fact: the decoder is native code processing hostile input, and it is doing so against the same kernel as every other container on that node. The syscall filter narrows which kernel code is reachable; it does not change which kernel is reachable. A container escape is a kernel bug away, and image-codec bugs are exactly the primitive you would use to go looking for one. The fuller ranking of these boundaries is in The Code Isolation Hierarchy.
A microVM changes the proposition rather than tightening it. The guest gets its own kernel, so the codec bug now buys an attacker root on a machine that holds one study and expires in a minute. Getting further means breaking out of that kernel and then through the VMM — and Firecracker's device model is deliberately tiny, a few virtio devices and a serial port, so there is strikingly little emulated surface to aim at. Two boundaries of different kinds, which is the only reason anyone bothers.
Be fair about when the container is enough, though. If you control the encoder, your inputs come from one scanner you own, volumes are modest and the data is already de-identified, a hardened container is a reasonable place to stop. The microVM argument gets strong when the inputs are a hospital's history, when the codec set is wide, and when the pixels are PHI.
The de-identification paradox: scrubbing is parsing
Here is the part that surprises people who arrive thinking de-identification is the safety step. To remove PHI from a DICOM object you must parse the DICOM object. Your scrubber sits inside the blast radius it was built to shrink, and it is the first code to touch the file. Any pipeline that treats de-identification as a trusted preprocessing stage has the ordering exactly backwards.
The standard's own guidance, PS3.15 Annex E, is a denylist: a long table of attributes with actions to remove, blank, replace or regenerate. Denylists are the wrong shape for this problem, because the identifiers you are chasing also live in private tags, in free-text fields like StudyDescription and SeriesDescription that clinicians type patient details into, in structured-report content, and in vendor duplicates of PatientName with tags no dictionary knows. Every new scanner model is a new chance for the denylist to be incomplete.
Where you can, invert it. If your downstream genuinely needs twelve attributes, build a fresh object containing those twelve rather than deleting several hundred from the original. Reconstruction fails closed; scrubbing fails open and quietly. There is more on that general pattern in Isolating PII Redaction Pipelines in MicroVMs.
Then there is the pixel problem, which has no clean answer. Ultrasound, secondary capture and anything screen-captured off a workstation routinely carry the patient's name burned into the image itself. The `(0028,0301)` BurnedInAnnotation attribute is supposed to tell you, and in practice is absent or wrong often enough to be unusable as a gate. Overlay planes in the `60xx` groups can carry the same text as a separate bitmap. So to find PHI in the pixels you have to read the pixels — which means OCR, which means Tesseract and Leptonica and another stack of image decoders, which is strictly more untrusted-input code, not less. Put that stage in its own VM too; the same shape works, and MicroVM Isolation for AI Invoice OCR and Extraction covers the OCR flavour of it.
One more trapdoor: DICOM has SOP classes for encapsulated PDF, CDA and STL. A perfectly standards-compliant object can contain a PDF, which means your imaging pipeline contains a PDF parser whether that was ever on your architecture diagram or not. Refuse those UIDs explicitly, or handle them with the same care as Document Conversion Is Remote Code Execution With a Progress Bar describes.
One microVM per study, and nothing to clean up
Scope the VM to the unit your blast radius should be, which is almost always one study, sometimes one series. A poisoned object can then reach the other objects belonging to the same patient — which it could anyway, since they are the thing being processed — and nothing else. There is no previous patient's pixels in the page cache, because there was no previous patient on this machine. There has never been anyone else on this machine.
That reframes cleanup as a non-problem, which matters more than it sounds. In a long-lived worker, every artefact of a crash is a disclosure question you have to answer: the core dump, the files in `/tmp`, decoded frames still resident in page cache, pages that went to host swap, the half-written output of the study that was interrupted. In a disposable VM, teardown deletes the whole machine. Set `core_pattern` empty and `ulimit -c 0` anyway, because defence in depth is cheap, but the structural answer is that the disk and the memory cease to exist. For the snapshots and volumes that do persist, Encryption at Rest for Sandbox Disks and Snapshots is the relevant piece.
The posture to aim for is a strange-sounding sentence that is exactly right: the parsing VM is inside your PHI boundary and outside your network. It has to see the pixels, because it is the thing parsing them. It must not be able to tell anyone. No egress, no credentials, no service-account token, no DNS. The only thing that leaves is the derived artefact, read back by the orchestrator over the control path, on paths the orchestrator chose.
The pipeline, concretely
- Ingest at the edge: terminate DIMSE or DICOMweb, write the bytes to object storage, hash them, and do not parse them. Checking for the `DICM` magic at offset 128 is not parsing; it is reading four bytes.
- One microVM per study, created from a baked template with your imaging stack already in it.
- Take the network away before the parse starts: default-deny in the sandbox's namespace, and delete the default route inside the guest as a second layer.
- Validate structurally and refuse rather than repair. Transfer syntax against an allowlist, geometry against your caps, encapsulated-document SOP classes rejected outright.
- De-identify by reconstruction, then decode pixels last, after every metadata gate has passed.
- Transcode to one boring syntax — explicit VR little endian, or straight out to PNG and a JSON record — so that nothing downstream ever runs a codec again.
- Emit only the derived artefacts, on paths the orchestrator named, and then kill the VM.
The host side of that is two short pieces: the network posture, and the fence that bounds the parse. The fence is the part people leave out, so it is written out in full here.
#!/usr/bin/env bash
# Host-side, around one study's sandbox. Two jobs: take the network away
# from the parsing VM, and build the fence that bounds the parse itself.
set -euo pipefail
STUDY=$1 # a directory of .dcm objects for ONE study
NS=$2 # this sandbox's netns: ns-<id>, veth + tap0
# 1. Bytes-only triage, on the host, with no parser involved. A DICOM file
# is a 128-byte preamble then the four characters DICM. Reading four
# bytes at a fixed offset is not parsing. Everything after this line is,
# which is exactly why everything after this line happens in the VM.
for f in "$STUDY"/*.dcm; do
[ "$(dd if="$f" bs=1 skip=128 count=4 2>/dev/null)" = DICM ] \
|| { echo "reject(no-magic): $(basename "$f")" >&2; exit 64; }
done
# 2. Take the network away. Default-deny in the sandbox's own namespace,
# then exactly one hole: the control path, so the agent's SSH bridge can
# still reach the guest. The codec has nothing legitimate to say to the
# internet, so it is given no way to say it. Note the log rule -- the
# line it writes is the evidence that something tried.
ip netns exec "$NS" nft -f - <<'RULES'
table inet dicomjail {
chain forward {
type filter hook forward priority 0; policy drop;
iifname "vg-*" oifname "tap0" tcp dport 22 ct state new accept
ct state established,related accept
iifname "tap0" ct state new log prefix "dicom-egress-denied " drop
}
}
RULES
# 3. The fence that actually bounds a runaway codec, written out here
# because it is the load-bearing part. This is the exact shell string
# the Python below hands to sbx.exec(). One-shot exec() does not
# reliably enforce a timeout_seconds argument, so the bound is in-guest
# or it does not exist at all.
#
# 124 = timeout(1) fired. 137 = the process ignored SIGTERM and needed
# the KILL from --kill-after. Either one means quarantine the study.
read -r -d '' GUEST_CMD <<'FENCE'
ip route del default 2>/dev/null || true # layer two: no path off-link.
echo > /proc/sys/kernel/core_pattern # a core file here is PHI on disk.
ulimit -v 2000000 # 2 GiB of VA inside a 4 GiB VM
ulimit -u 64 # no fork bombs out of a decoder
ulimit -c 0
cd /work
export PATH=/opt/mise/shims:$PATH
exec timeout --kill-after=10s 120 python3 deident.py
FENCE
# 4. Prove the posture in the run the auditor will ask about, and keep the
# proof. "We configured default-deny" is a claim; a denied connect from
# inside the guest, timestamped, is an artefact.
echo "$GUEST_CMD" > "/var/log/dicom/$NS.fence"
ip netns exec "$NS" nft list table inet dicomjail > "/var/log/dicom/$NS.nft"
Then the orchestration. Note what this process does and does not do: it holds credentials and a pseudonym map, it writes bytes in, it reads named artefacts out, and it never parses a DICOM object.
# One study, one microVM. The VM is the unit of cleanup, because there is
# no cleanup -- there is no afterwards.
import hashlib, json, pathlib, secrets
from concurrent.futures import ThreadPoolExecutor
from pandastack import Sandbox
GUEST = pathlib.Path("guest/deident.py").read_text()
# The fence from the bash above, as one string. Everything hostile happens
# on the far side of it.
FENCE = (
"ip route del default 2>/dev/null; echo > /proc/sys/kernel/core_pattern; "
"ulimit -v 2000000; ulimit -u 64; ulimit -c 0; cd /work; "
"export PATH=/opt/mise/shims:$PATH; "
"exec timeout --kill-after=10s 120 python3 deident.py"
)
def process_study(study: pathlib.Path) -> dict:
files = sorted(study.glob("*.dcm"))
if not 1 <= len(files) <= 4096:
raise ValueError(f"{study.name}: {len(files)} objects, refusing")
sbx = Sandbox.create(
template="base", # 8 vCPU / 4 GiB, baked into the snapshot.
# Firecracker cannot change vCPU or RAM at
# restore, so a cpu= or memory_mb= on create
# is overridden to the baked values. Size the
# template, not the request.
ttl_seconds=600, # An IDLE clock, not a wall clock. This is the
# backstop for an orchestrator that died
# holding a VM full of PHI. It is NOT a fence
# around the parse -- that is timeout(1).
metadata={"job": "dicom-deident",
"pseudonym": secrets.token_hex(8)},
) # metadata is host-side and searchable, so it
# carries a pseudonym and never an MRN. The
# pseudonym-to-record map lives in the
# orchestrator's store, which this VM has no
# route to and no credential for.
try:
sbx.exec("mkdir -p /work/in /work/out")
sbx.filesystem.write("/work/deident.py", GUEST)
sbx.filesystem.write("/work/job.json", json.dumps(
{"window": "p1-p99", "max_px": 4096 * 4096, "emit": "png"}))
# Our code first, then the untrusted bytes, byte-for-byte untouched.
inputs = {}
for i, f in enumerate(files):
raw = f.read_bytes()
if len(raw) < 132 or raw[128:132] != b"DICM":
raise ValueError(f"{f.name}: no DICM magic at offset 128")
name = f"{i:05d}.dcm"
sbx.filesystem.write(f"/work/in/{name}", raw)
inputs[name] = hashlib.sha256(raw).hexdigest()
res = sbx.exec(FENCE)
if res.exit_code in (124, 137):
raise TimeoutError(
f"{study.name}: decoder ran away, study quarantined")
if res.exit_code != 0:
# The guest's stderr is deliberately filenames and tag NAMES
# only. A traceback carrying a PatientName is a disclosure
# event wearing a log line.
raise ValueError(
f"{study.name}: refused -- {res.stderr.strip()[:300]}")
# Read back EXACTLY the paths we expect and nothing else. No
# directory walk, no glob: the derived-artefact list is ours, not
# the guest's, so a compromised guest cannot nominate a file for
# exfiltration by writing it somewhere we were going to look.
out = json.loads(sbx.filesystem.read("/work/out/manifest.json"))
for frame in out["frames"]:
store_derived(study.name, frame, sbx.filesystem.read(
f"/work/out/{frame}"))
return {"study": study.name, "frames": len(out["frames"]),
"inputs": inputs, "meta": out["meta"]}
finally:
sbx.kill() # Teardown is kill(). There is no sbx.delete().
def run_batch(studies: list[pathlib.Path], workers: int = 64) -> list[dict]:
# A burst of studies is a burst of VMs. They share a host and a tap
# driver and nothing else -- no /tmp, no page cache, no kernel.
with ThreadPoolExecutor(max_workers=workers) as pool:
return list(pool.map(process_study, studies))
And the guest script, which is the only code here allowed to be wrong in an interesting way. Every gate below runs before `pixel_array` is touched, because that single attribute access is where the C starts executing.
#!/usr/bin/env python3
# /work/deident.py -- runs INSIDE the microVM, with no default route and no
# credentials. Fails closed: anything surprising is a non-zero exit, never a
# best-effort render. The caller treats a refusal as a study to quarantine.
import json, pathlib, sys
import numpy as np
import pydicom
from PIL import Image
from pydicom.uid import (ExplicitVRLittleEndian, ImplicitVRLittleEndian,
RLELossless, JPEGLSLossless, JPEG2000Lossless)
# Transfer syntaxes we are willing to hand to a decoder, by UID. Everything
# else is refused unprobed: the retired big-endian syntaxes, the JPEG 2000
# multi-component variants, and -- yes, this is legal DICOM -- MPEG-2 and
# HEVC. If you did not plan to run a video decoder on hospital data, say so
# in a set literal now rather than in a post-mortem later.
ALLOWED = {ExplicitVRLittleEndian, ImplicitVRLittleEndian, RLELossless,
JPEGLSLossless, JPEG2000Lossless}
# An ALLOWLIST, not a denylist. PS3.15 Annex E enumerates hundreds of
# attributes to remove, and private tags plus vendor duplicates of
# PatientName mean a denylist is a bet you re-place on every new scanner
# model. Downstream needs these twelve, so these twelve are what leaves.
KEEP = ("Modality", "SOPClassUID", "Rows", "Columns", "BitsAllocated",
"BitsStored", "PhotometricInterpretation", "SamplesPerPixel",
"PixelSpacing", "SliceThickness", "WindowCenter", "WindowWidth")
ENCAPSULATED = {"1.2.840.10008.5.1.4.1.1.104.1", # Encapsulated PDF
"1.2.840.10008.5.1.4.1.1.104.2", # Encapsulated CDA
"1.2.840.10008.5.1.4.1.1.104.3"} # Encapsulated STL
job = json.load(open("/work/job.json"))
MAX_BYTES = int(job["max_px"]) * 2
def die(msg: str) -> None:
# Filenames and tag NAMES only -- never a tag VALUE. stderr lands in the
# orchestrator's logs, and your logging stack is not inside the PHI
# boundary no matter what the architecture diagram says.
print(msg, file=sys.stderr)
sys.exit(2)
frames, meta = [], {}
for src in sorted(pathlib.Path("/work/in").glob("*.dcm")):
raw = src.read_bytes()
if len(raw) < 132 or raw[128:132] != b"DICM":
die(f"{src.name}: no DICM magic")
# force=False: if the file meta group is unreadable we refuse, rather
# than let the library guess a transfer syntax. Guessing is how a
# lenient parser ends up feeding wavelet data to a JPEG decoder.
ds = pydicom.dcmread(src, force=False)
ts = ds.file_meta.TransferSyntaxUID
if ts not in ALLOWED:
die(f"{src.name}: refused transfer syntax {ts}")
if str(getattr(ds, "SOPClassUID", "")) in ENCAPSULATED:
die(f"{src.name}: encapsulated document -- a PDF, not a scan")
if int(getattr(ds, "NumberOfFrames", 1) or 1) != 1 \
or int(ds.SamplesPerPixel) != 1:
die(f"{src.name}: multi-frame or colour -- other stage, same jail")
if int(getattr(ds, "OverlayBitsAllocated", 0) or 0):
die(f"{src.name}: overlay plane present, may carry burned-in text")
if str(getattr(ds, "BurnedInAnnotation", "")).upper() == "YES":
die(f"{src.name}: declares burned-in annotation, send to OCR stage")
# The allocation bomb. Rows, Columns and BitsAllocated are
# attacker-controlled integers and numpy will cheerfully try to honour
# whatever product they describe, before a single pixel is decoded.
want = int(ds.Rows) * int(ds.Columns) * ((int(ds.BitsAllocated) + 7) // 8)
if not 0 < want <= MAX_BYTES:
die(f"{src.name}: declares {want} bytes of pixel data")
# Decode LAST, after every metadata gate, in the smallest possible
# scope. This one attribute access is where OpenJPEG, CharLS or a
# libjpeg descendant runs on bytes a stranger chose.
arr = ds.pixel_array.astype(np.float32)
lo, hi = np.percentile(arr, (1, 99))
img = np.clip((arr - lo) / max(hi - lo, 1e-6) * 255, 0, 255)
Image.fromarray(img.astype(np.uint8)).save(f"/work/out/{src.stem}.png")
frames.append(f"{src.stem}.png")
# Reconstruction, not scrubbing: a fresh dict of permitted scalars. The
# original dataset is never serialised back out, so nothing we forgot
# to delete can ride along inside it.
meta[src.stem] = {k: str(getattr(ds, k, "")) for k in KEEP}
pathlib.Path("/work/out/manifest.json").write_text(
json.dumps({"frames": frames, "meta": meta, "profile": "allowlist-v3"}))
Throughput: hundreds of short VMs beat a pool you have to trust
The objection to per-study VMs is always cost, and it rests on an assumption about boot time that does not hold on a snapshot-restore platform. On PandaStack a create is a restore of a baked snapshot, not a boot: p50 179 ms, p99 around 203 ms, of which the `/snapshot/load` step is roughly 49 to 80 ms. The first spawn of a template with no baked snapshot is a genuine cold boot at about three seconds, and then the snapshot is baked automatically — so you pay that once per template, not once per study.
Put that next to the work. A 300-slice CT study is one VM, not three hundred, and the decode of those slices is seconds to minutes. Two hundred milliseconds of create against that is rounding error, and it buys you a property no worker pool can offer: the machine is provably clean, because it is new. The alternative is a pool of long-lived workers that you trust to stay clean, where the trust is doing the work and the evidence is doing none.
For fan-out inside one study there is a warmer option. `fork_tree()` snapshots a parent that already has Python up, pydicom imported and your validated job configuration on disk, and boots children from that snapshot so they inherit the parent's running memory. It is capped at sixteen children per call, so wider fan-outs grow the tree breadth-first. Keep one rule absolute: one study per tree. Never fan a single parent out across patients, because the whole point of the parent was that it belongs to exactly one of them. `fork()` is the other call and it is not this one — it clones the disk and the child cold-boots with its own entropy and PIDs, which is the right choice when you want independence rather than warmth.
On timeouts, be precise, because this has bitten people. `ttl_seconds` is an idle clock, not a wall-clock lifetime; a busy sandbox is not reaped. It is the backstop for an orchestrator that died holding a VM full of PHI. It is not a fence around a wedged codec, and one-shot `exec()` does not reliably enforce a timeout argument either. The bound on a runaway decode is `timeout(1)` inside the guest, or it does not exist.
Isolation is a safeguard story, not a certification
This deserves to be said plainly rather than implied. Per-study kernel isolation gives you good answers to real questions an assessor asks — how one patient's data is prevented from reaching another patient's processing run, how you bound what a compromised parser can read, what happens to intermediate artefacts, how you evidence that a run had no network path out. Those are technical-safeguard arguments and they are stronger than the ones a shared worker pool can make.
What it is not: a certification, an attestation, or a business associate agreement. PandaStack does not sign a BAA today, and no architectural property substitutes for one. If you are processing real PHI in production you need a BAA with whoever holds the bytes, plus the administrative and physical safeguards that have nothing to do with hypervisors. Build the technical story on this shape if you like it; do not let it stand in for the paperwork. Running Code Over PHI Without Expanding Your HIPAA Blast Radius goes through that boundary in more detail.
Where I would not use this
We have no GPUs and no GPU passthrough, and medical imaging is the field where that bites hardest. Volumetric segmentation, a 3D reconstruction, or any of the deep-learning inference people actually want to run on a study wants a GPU, and GPU-accelerated JPEG 2000 decode exists for a reason. The honest split is to do parsing, validation and de-identification here — the part that is dangerous rather than compute-heavy — and hand the derived, de-identified volume to a GPU host that never sees an original object. If your workload is GPU-first, this is the wrong platform for that stage and I would rather say so than sell you a workaround.
The guest kernel is 5.10, so tooling that needs something newer will not be happy. The baked template size governs RAM — `base` is 4 GiB, and Firecracker cannot change vCPU or RAM at snapshot restore — so a reconstruction that wants 64 GiB resident is not a thing you squeeze out of this by passing a bigger number at create time. A long-lived DIMSE C-STORE listener is a server rather than a burst job: you can run one in a persistent sandbox, but it does not get the per-study-disposable property, so terminate DIMSE at the edge and fan out behind it. And if your inputs are one vendor's files from one scanner you administer, a hardened container is a defensible place to stop.
The failure you will actually meet, incidentally, is denial of service rather than remote code execution: a decoder that loops, a pixel geometry that asks for 40 GiB, a decompression bomb. Design for that first. The isolation is there for the rarer, worse day, and the nice thing about building for the rare case is that the cheap case comes free — a wedged codec in a disposable VM is one timed-out study and one deleted machine, not a restart loop that drops other patients' work on the floor.
Frequently asked questions
Is pydicom safe because it is written in Python?
Partly, and the part that is not safe is the part that matters. pydicom's dataset and tag parsing is pure Python, so the classic memory-safety bug classes — heap overflow, use-after-free, out-of-bounds read — are largely off the table in the tag parser itself. That is a real advantage over a C++ parser and worth having. The pixel handlers are a different story: pydicom does not decode JPEG, JPEG-LS or JPEG 2000 itself. It delegates to whatever plugin you installed, which is typically pylibjpeg with its openjpeg, libjpeg and rle backends, or GDCM, or Pillow. Those are C and C++ libraries, and the moment you touch `pixel_array` you are executing them on attacker-supplied bytes. Check which handler your environment actually resolves to rather than assuming; it depends on what is installed, and it can differ between your laptop and your production image. Even where the code is pure Python you still have availability and memory problems: unbounded allocation driven by Rows, Columns, NumberOfFrames and BitsAllocated, numpy reshapes sized from those same attacker-controlled integers, deflate bombs in the deflated transfer syntax, and pathological sequence nesting. None of those are remote code execution. All of them take out a shared worker, and in a shared worker that means taking out other patients' jobs too.
Can a library de-identify DICOM reliably, or do I need a human in the loop?
For a research release to third parties, you need a review step, and anyone who tells you otherwise has not audited the output of their own pipeline. For an automated internal pipeline you can get to a defensible place, but only by changing shape. PS3.15 Annex E is a denylist of attributes with actions, and denylists leak: identifiers hide in private tags that no dictionary describes, in free-text fields like StudyDescription and SeriesDescription where clinicians type names and dates, in structured-report content, in UIDs that encode site information, and in vendor-specific duplicates of PatientName. Every new scanner model is a fresh opportunity for your list to be incomplete. The stronger pattern is reconstruction: decide which attributes your downstream genuinely needs, build a new object containing only those, and discard the original rather than editing it. That fails closed. Then there is the pixel problem, which no attribute-level tool solves at all. Burned-in annotation is routine in ultrasound, secondary capture and workstation screen captures, the BurnedInAnnotation attribute is unreliable as a gate, and overlay planes in the 60xx groups can carry the same text as a separate bitmap. Finding it means OCR, which is more untrusted-input C code, and OCR on medical images has a false-negative rate. Practically: allowlist what you emit, refuse modalities where burned-in text is likely unless they have been through an OCR stage, sample the output for human review, and keep the review rate honest rather than aspirational.
One microVM per study or one per file?
Per study is the default, and the reason is that the blast radius should match an existing trust boundary rather than an arbitrary one. Objects within a study belong to the same patient, so a poisoned object reaching its siblings tells you nothing new — they were going to be processed together anyway. Objects from different patients have no business sharing a kernel, a page cache or a scratch directory. Per study also keeps the arithmetic sane: a 300-slice CT is one create at a p50 of 179 ms, against a decode measured in seconds to minutes, so the isolation is close to free. Per file makes sense in two situations. The first is triage of a mixed archive where you have no reliable study grouping yet, because the grouping metadata is the thing you do not trust — then each object genuinely is unrelated to its neighbours. The second is when a specific object is already suspicious: it failed a gate, it came from an unknown source, or it is the one that crashed the decoder last night. Give that one its own machine and keep the artefacts. For fan-out within a study, `fork_tree()` gives you children that inherit a warm parent's memory, capped at sixteen per call. The rule there is one study per tree, always: the parent's whole value is that it belongs to exactly one patient, and fanning it across patients throws that away for a few hundred milliseconds.
Does the microVM boundary actually help here, or is it security theatre on top of a container?
It helps, and it helps for a specific reason rather than a general one: it changes what a single bug buys. In a container, a memory-safety bug in OpenJPEG gives an attacker code execution in a process that shares a kernel with every other workload on the node, so the next step is a kernel bug and the prize is the node. In a microVM the same bug gives code execution inside a guest kernel that owns one study and will be deleted shortly, and the next step is a hypervisor escape through a device model that is deliberately minimal — a handful of virtio devices and a serial port, with no HPET, no sound card, no emulated legacy hardware. That is a much thinner target, and the VMM is Rust rather than decades-accumulated C. It is not magic, and I would rather be specific about the limits. The VMM is still software and still has a published advisory history; read it rather than assuming. Microarchitectural side channels cross the VM boundary — /blog/spectre-meltdown-microvm-isolation covers what that means in practice — so a microVM is not a confidentiality guarantee against a co-tenant with patience. Nothing here protects you from your own orchestrator leaking PHI into logs, which is the far more common incident. And none of it is a substitute for the boring controls: least-privilege credentials, no network egress from the parsing stage, and an emitted-artefact list that your code chose rather than the guest.
Keep reading
- HIPAA, PHI and code execution isolation — Where isolation genuinely helps the safeguard story, and the line past which you need a BAA instead of an architecture.
- Processing PHI in disposable microVMs — The same shape applied to healthcare data generally, rather than to one hostile binary format.
- PII redaction and anonymisation in a microVM — Why the redactor belongs inside the blast radius it is shrinking, and the allowlist-over-denylist argument in full.
- OCR extraction in an isolated microVM — The stage you need for burned-in pixel annotation, and the extra decoder stack that comes with it.
- Untrusted document conversion — For the day you discover that a DICOM object can legally contain a PDF.
Related posts
- Parsing Untrusted Binary Protocols: A Length Field You Believed
Text formats fail by doing too much. Binary formats fail by arithmetic. The wire said how long something was and the parser believed it — that is the whole bug class, and the vendor SDK underneath you is not in your fuzzing corpus.
- A Font Is a Program: Rendering User-Uploaded Typefaces Without Trusting Them
The moment you accept a .ttf from a user, you have accepted a program: TrueType carries a stack machine, CFF carries another one, and woff2 puts a Brotli decoder in front of both. Here is why that parser should not run in the process holding your database credentials, and the microVM shape that fixes it.
- Parsing Untrusted XML: The Format That Ships an Interpreter
An external entity is a documented feature that makes the parser read files and open sockets for whoever wrote the document. XXE is not a defect; it is conformance with the wrong author. So stop treating a flag as a boundary.
- Bioinformatics Pipelines on microVMs: Reproducibility, PHI and the 400-Tool Dependency Graph
Two stages of the same pipeline want incompatible htslib builds, the postdoc's R script needs to write into the site library, and the whole thing is bit-for-bit reproducible until somebody upgrades the host kernel. Containers fixed the packaging. The VM boundary fixes the rest.
- One microVM Per Claim: Isolating Insurance Claims Processing Per Carrier
A claims pipeline is an API that invites strangers to upload files they chose, into a parser you did not write, on a machine holding four other carriers' data. One poisoned PDF and your incident report has more than one logo on it.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.