all posts

Running a DICOM Pipeline in a MicroVM: Untrusted Medical Images, One Study at a Time

Ajay Kumar··10 min read

The ticket read: de-identification worker restarted eleven times overnight, backlog one thousand four hundred studies. It was not a memory leak. One object in one study — an ultrasound loop from a scanner model nobody at the hospital could still name — was taking the decoder down with it, and because the worker pulled studies from a shared queue into a shared scratch directory, each restart left half-written frames from three other patients' studies on disk for the next run to trip over.

That is the ordinary shape of a DICOM pipeline failure, and it is worth sitting with how many separate problems are stacked inside it. A native-code image decoder crashed on input it did not choose. The crash took down a process that was also holding three other patients' pixels. The scratch directory outlived the process that owned it. And every byte involved was protected health information, which turns an availability incident into a disclosure question: what exactly was in the core dump, and who can read it?

I build PandaStack, an open-source Firecracker microVM platform, so you can see where this is going. The argument stands without the product though, and the specifics are the useful part: what inside a DICOM file is actually executing, why "malformed" is the normal case rather than the attack case in a real archive, why the de-identifier is itself in the blast radius, and what one-microVM-per-study actually costs to run.

The short version. A DICOM file names its own decoder, and that decoder is usually decades-old C: a libjpeg descendant, OpenJPEG, CharLS, or someone's hand-rolled RLE loop. Hospital archives are full of files written by a decade of vendor encoders, so your parser lives permanently in lenient mode on hostile-shaped input. A container runs that code against the same kernel as everything else you own. A microVM gives it its own kernel, no route to the network, and a lifetime measured in seconds — and then the machine is deleted, along with the crash dump, the temp files and the decoded frames.

A DICOM file names its own decoder

The layout is a product of its era, and knowing it is what makes the rest of this concrete. A file begins with a 128-byte preamble that is allowed to contain anything at all, followed by the four characters `DICM`. Then comes the File Meta Information group — every element with group number `0002` — which is always encoded in Explicit VR Little Endian so that a reader can bootstrap. Inside that group sits one element that decides everything downstream: `(0002,0010)` TransferSyntaxUID.

That UID is not a version number. It is a dispatch table. `1.2.840.10008.1.2` means implicit VR little endian, so the rest of the dataset carries no type information at all. `1.2.840.10008.1.2.4.90` means the pixel data is JPEG 2000 and a wavelet codec is about to run. `1.2.840.10008.1.2.1.99` means deflate, so there is a decompressor in front of the parser. The file is telling your process which several thousand lines of C to execute next, and it is a file you received from outside.

The tag parser, before any codec runs

Under implicit VR there is no value-representation field on each element, so the parser must look the tag up in a data dictionary to learn whether those four bytes are a signed short, a UID string or a nested sequence. Tags with odd group numbers are private — vendor extensions, not in the dictionary — so the parser falls back to `UN`, unknown, and an `UN` element with undefined length is genuinely ambiguous about where it ends.

The standard also requires every value to be even-length, padded with a null or a space. Real files violate this constantly, so every production parser has an odd-length recovery path, and recovery paths are where the interesting bugs live. Sequences make it worse: an `SQ` element may declare undefined length and be terminated by delimiter items — `(FFFE,E000)` for an item, `(FFFE,E00D)` and `(FFFE,E0DD)` to close them. Get one of those wrong and the parser does not fail, it desynchronises, and every element after that point is read at the wrong offset with the wrong type.

Then there is pixel data itself. Element `(7FE0,0010)` can be a flat blob, or it can be encapsulated: a Basic Offset Table followed by fragments, where one frame may span several fragments and the offset table is a set of numbers the file supplies about itself. A multi-frame object whose offset table disagrees with its fragment boundaries is a classic, and the disagreement is attacker-controllable.

Then the codec, which is where the C lives

Image codecs are the most reliably exploited class of code in the history of software, for structural reasons rather than anecdotal ones: they are complex binary parsers, written in unsafe languages for speed, invoked automatically on content nobody vetted, and historically embedded in privileged processes. DICOM hands you most of that family at once, selected by a string in the file.

The transfer syntax chooses the decoder. Which implementation you actually get depends on your stack — pydicom with pylibjpeg, pydicom with GDCM, dcmtk's own bundled codecs — so verify what your install links rather than trusting this table.
Transfer syntaxWhat decodes itTypical implementationFailure mode to plan for
Implicit VR Little EndianThe tag parser itself; no codecPure Python in pydicom, C++ in dcmtk and GDCMNo VR on the wire, so types come from a dictionary and private tags are guessed. A desync here mis-parses everything after it.
Explicit VR Little EndianTag parser; pixel data is rawSameThe safest case, and the one you should transcode to. Still carries odd-length values and undefined-length sequences.
JPEG Baseline and ExtendedA libjpeg-lineage decoderCDecades-old C walking attacker-shaped entropy-coded data. The canonical heap-overflow surface in imaging.
JPEG Lossless and Lossless SV1A rarely-exercised lossless JPEG pathCLow-traffic code inside a high-traffic library, which usually means less fuzzing coverage than the baseline path next to it.
JPEG-LS, lossless and near-losslessCharLSC++Golomb-Rice coding with run modes, plus a quantiser in the near-lossless variant. Small, fast, and not where most fuzzing effort has gone.
JPEG 2000OpenJPEGCWavelets, tiles, precincts, code-blocks and markers: the largest attack surface in the set, with a long and public advisory history. Read it before deciding this is fine.
RLE LosslessA segmented run-length decoderC, or pure Python via pylibjpeg-rleA format simple enough that people hand-roll it. Segment offsets come from the file, so one unchecked bound is an overflow in forty lines of code.
Deflated Explicit VR LEzlib, then the tag parserCA decompressor in front of the parser. Your size caps have to apply after inflation, which is not where people put them.
MPEG-2, MPEG-4, HEVCA video decoderC and C++Yes, a DICOM object can contain H.265. Refuse these by UID unless you consciously decided to run a video decoder on hospital data.

In a real archive, malformed is the normal case

If you have only ever parsed DICOM produced by one modern toolchain, the format feels strict. A PACS archive is the opposite experience. It holds a decade or more of objects written by many vendors' encoders, including ones whose authors are no longer in business, each with its own interpretation of the ambiguous parts and its own private tags carrying duplicates of the patient's name. There are retired big-endian objects, lengths that disagree with reality, sequences closed the wrong way, character sets declared in one place and ignored in another.

The operational consequence is the thing people miss. You cannot ship a strict parser, because a strict parser rejects a meaningful fraction of the archive you were hired to process. So you ship a lenient one — recovery paths on, best-effort length handling, dictionary fallbacks — and you cannot use parse failure as an attack signal either, because parse failures are your Tuesday. A deliberately malformed file arrives wearing exactly the same clothes as a 2009 ultrasound.

Every production DICOM parser is a lenient parser, and leniency is the attack surface. You do not get to choose strictness; you get to choose what the parser is standing inside when it is wrong.

Why "just run it in a container" is the shaky default

The default architecture is a worker pod with `pydicom`, `gdcm` and `dcmtk` installed, pulling study paths off a queue. Done well it has a read-only root filesystem, a non-root user, dropped capabilities, a seccomp profile and a memory limit. All of that is worth having and none of it changes the central fact: the decoder is native code processing hostile input, and it is doing so against the same kernel as every other container on that node. The syscall filter narrows which kernel code is reachable; it does not change which kernel is reachable. A container escape is a kernel bug away, and image-codec bugs are exactly the primitive you would use to go looking for one. The fuller ranking of these boundaries is in The Code Isolation Hierarchy.

A microVM changes the proposition rather than tightening it. The guest gets its own kernel, so the codec bug now buys an attacker root on a machine that holds one study and expires in a minute. Getting further means breaking out of that kernel and then through the VMM — and Firecracker's device model is deliberately tiny, a few virtio devices and a serial port, so there is strikingly little emulated surface to aim at. Two boundaries of different kinds, which is the only reason anyone bothers.

Be fair about when the container is enough, though. If you control the encoder, your inputs come from one scanner you own, volumes are modest and the data is already de-identified, a hardened container is a reasonable place to stop. The microVM argument gets strong when the inputs are a hospital's history, when the codec set is wide, and when the pixels are PHI.

The de-identification paradox: scrubbing is parsing

Here is the part that surprises people who arrive thinking de-identification is the safety step. To remove PHI from a DICOM object you must parse the DICOM object. Your scrubber sits inside the blast radius it was built to shrink, and it is the first code to touch the file. Any pipeline that treats de-identification as a trusted preprocessing stage has the ordering exactly backwards.

The standard's own guidance, PS3.15 Annex E, is a denylist: a long table of attributes with actions to remove, blank, replace or regenerate. Denylists are the wrong shape for this problem, because the identifiers you are chasing also live in private tags, in free-text fields like StudyDescription and SeriesDescription that clinicians type patient details into, in structured-report content, and in vendor duplicates of PatientName with tags no dictionary knows. Every new scanner model is a new chance for the denylist to be incomplete.

Where you can, invert it. If your downstream genuinely needs twelve attributes, build a fresh object containing those twelve rather than deleting several hundred from the original. Reconstruction fails closed; scrubbing fails open and quietly. There is more on that general pattern in Isolating PII Redaction Pipelines in MicroVMs.

Then there is the pixel problem, which has no clean answer. Ultrasound, secondary capture and anything screen-captured off a workstation routinely carry the patient's name burned into the image itself. The `(0028,0301)` BurnedInAnnotation attribute is supposed to tell you, and in practice is absent or wrong often enough to be unusable as a gate. Overlay planes in the `60xx` groups can carry the same text as a separate bitmap. So to find PHI in the pixels you have to read the pixels — which means OCR, which means Tesseract and Leptonica and another stack of image decoders, which is strictly more untrusted-input code, not less. Put that stage in its own VM too; the same shape works, and MicroVM Isolation for AI Invoice OCR and Extraction covers the OCR flavour of it.

One more trapdoor: DICOM has SOP classes for encapsulated PDF, CDA and STL. A perfectly standards-compliant object can contain a PDF, which means your imaging pipeline contains a PDF parser whether that was ever on your architecture diagram or not. Refuse those UIDs explicitly, or handle them with the same care as Document Conversion Is Remote Code Execution With a Progress Bar describes.

The mistake that turns a crash into a breach: letting the de-identifier and the archive credentials live in the same process. If the thing parsing hostile pixel data also holds the PACS service account, the object-store key or the database handle for the re-identification map, then one memory-safety bug in OpenJPEG hands over the whole archive rather than one study. Split them. The parsing VM gets the pixels and no credentials and no route off-link; the orchestrator holds the credentials and never parses anything.

One microVM per study, and nothing to clean up

Scope the VM to the unit your blast radius should be, which is almost always one study, sometimes one series. A poisoned object can then reach the other objects belonging to the same patient — which it could anyway, since they are the thing being processed — and nothing else. There is no previous patient's pixels in the page cache, because there was no previous patient on this machine. There has never been anyone else on this machine.

That reframes cleanup as a non-problem, which matters more than it sounds. In a long-lived worker, every artefact of a crash is a disclosure question you have to answer: the core dump, the files in `/tmp`, decoded frames still resident in page cache, pages that went to host swap, the half-written output of the study that was interrupted. In a disposable VM, teardown deletes the whole machine. Set `core_pattern` empty and `ulimit -c 0` anyway, because defence in depth is cheap, but the structural answer is that the disk and the memory cease to exist. For the snapshots and volumes that do persist, Encryption at Rest for Sandbox Disks and Snapshots is the relevant piece.

The posture to aim for is a strange-sounding sentence that is exactly right: the parsing VM is inside your PHI boundary and outside your network. It has to see the pixels, because it is the thing parsing them. It must not be able to tell anyone. No egress, no credentials, no service-account token, no DNS. The only thing that leaves is the derived artefact, read back by the orchestrator over the control path, on paths the orchestrator chose.

The pipeline, concretely

  1. Ingest at the edge: terminate DIMSE or DICOMweb, write the bytes to object storage, hash them, and do not parse them. Checking for the `DICM` magic at offset 128 is not parsing; it is reading four bytes.
  2. One microVM per study, created from a baked template with your imaging stack already in it.
  3. Take the network away before the parse starts: default-deny in the sandbox's namespace, and delete the default route inside the guest as a second layer.
  4. Validate structurally and refuse rather than repair. Transfer syntax against an allowlist, geometry against your caps, encapsulated-document SOP classes rejected outright.
  5. De-identify by reconstruction, then decode pixels last, after every metadata gate has passed.
  6. Transcode to one boring syntax — explicit VR little endian, or straight out to PNG and a JSON record — so that nothing downstream ever runs a codec again.
  7. Emit only the derived artefacts, on paths the orchestrator named, and then kill the VM.

The host side of that is two short pieces: the network posture, and the fence that bounds the parse. The fence is the part people leave out, so it is written out in full here.

#!/usr/bin/env bash
# Host-side, around one study's sandbox. Two jobs: take the network away
# from the parsing VM, and build the fence that bounds the parse itself.
set -euo pipefail

STUDY=$1                    # a directory of .dcm objects for ONE study
NS=$2                       # this sandbox's netns: ns-<id>, veth + tap0

# 1. Bytes-only triage, on the host, with no parser involved. A DICOM file
#    is a 128-byte preamble then the four characters DICM. Reading four
#    bytes at a fixed offset is not parsing. Everything after this line is,
#    which is exactly why everything after this line happens in the VM.
for f in "$STUDY"/*.dcm; do
  [ "$(dd if="$f" bs=1 skip=128 count=4 2>/dev/null)" = DICM ] \
    || { echo "reject(no-magic): $(basename "$f")" >&2; exit 64; }
done

# 2. Take the network away. Default-deny in the sandbox's own namespace,
#    then exactly one hole: the control path, so the agent's SSH bridge can
#    still reach the guest. The codec has nothing legitimate to say to the
#    internet, so it is given no way to say it. Note the log rule -- the
#    line it writes is the evidence that something tried.
ip netns exec "$NS" nft -f - <<'RULES'
table inet dicomjail {
  chain forward {
    type filter hook forward priority 0; policy drop;
    iifname "vg-*" oifname "tap0" tcp dport 22 ct state new accept
    ct state established,related accept
    iifname "tap0" ct state new log prefix "dicom-egress-denied " drop
  }
}
RULES

# 3. The fence that actually bounds a runaway codec, written out here
#    because it is the load-bearing part. This is the exact shell string
#    the Python below hands to sbx.exec(). One-shot exec() does not
#    reliably enforce a timeout_seconds argument, so the bound is in-guest
#    or it does not exist at all.
#
#    124 = timeout(1) fired. 137 = the process ignored SIGTERM and needed
#    the KILL from --kill-after. Either one means quarantine the study.
read -r -d '' GUEST_CMD <<'FENCE'
ip route del default 2>/dev/null || true   # layer two: no path off-link.
echo > /proc/sys/kernel/core_pattern       # a core file here is PHI on disk.
ulimit -v 2000000                          # 2 GiB of VA inside a 4 GiB VM
ulimit -u 64                               # no fork bombs out of a decoder
ulimit -c 0
cd /work
export PATH=/opt/mise/shims:$PATH
exec timeout --kill-after=10s 120 python3 deident.py
FENCE

# 4. Prove the posture in the run the auditor will ask about, and keep the
#    proof. "We configured default-deny" is a claim; a denied connect from
#    inside the guest, timestamped, is an artefact.
echo "$GUEST_CMD" > "/var/log/dicom/$NS.fence"
ip netns exec "$NS" nft list table inet dicomjail > "/var/log/dicom/$NS.nft"

Then the orchestration. Note what this process does and does not do: it holds credentials and a pseudonym map, it writes bytes in, it reads named artefacts out, and it never parses a DICOM object.

# One study, one microVM. The VM is the unit of cleanup, because there is
# no cleanup -- there is no afterwards.
import hashlib, json, pathlib, secrets
from concurrent.futures import ThreadPoolExecutor
from pandastack import Sandbox

GUEST = pathlib.Path("guest/deident.py").read_text()

# The fence from the bash above, as one string. Everything hostile happens
# on the far side of it.
FENCE = (
    "ip route del default 2>/dev/null; echo > /proc/sys/kernel/core_pattern; "
    "ulimit -v 2000000; ulimit -u 64; ulimit -c 0; cd /work; "
    "export PATH=/opt/mise/shims:$PATH; "
    "exec timeout --kill-after=10s 120 python3 deident.py"
)


def process_study(study: pathlib.Path) -> dict:
    files = sorted(study.glob("*.dcm"))
    if not 1 <= len(files) <= 4096:
        raise ValueError(f"{study.name}: {len(files)} objects, refusing")

    sbx = Sandbox.create(
        template="base",      # 8 vCPU / 4 GiB, baked into the snapshot.
                              # Firecracker cannot change vCPU or RAM at
                              # restore, so a cpu= or memory_mb= on create
                              # is overridden to the baked values. Size the
                              # template, not the request.
        ttl_seconds=600,      # An IDLE clock, not a wall clock. This is the
                              # backstop for an orchestrator that died
                              # holding a VM full of PHI. It is NOT a fence
                              # around the parse -- that is timeout(1).
        metadata={"job": "dicom-deident",
                  "pseudonym": secrets.token_hex(8)},
    )                         # metadata is host-side and searchable, so it
                              # carries a pseudonym and never an MRN. The
                              # pseudonym-to-record map lives in the
                              # orchestrator's store, which this VM has no
                              # route to and no credential for.
    try:
        sbx.exec("mkdir -p /work/in /work/out")
        sbx.filesystem.write("/work/deident.py", GUEST)
        sbx.filesystem.write("/work/job.json", json.dumps(
            {"window": "p1-p99", "max_px": 4096 * 4096, "emit": "png"}))

        # Our code first, then the untrusted bytes, byte-for-byte untouched.
        inputs = {}
        for i, f in enumerate(files):
            raw = f.read_bytes()
            if len(raw) < 132 or raw[128:132] != b"DICM":
                raise ValueError(f"{f.name}: no DICM magic at offset 128")
            name = f"{i:05d}.dcm"
            sbx.filesystem.write(f"/work/in/{name}", raw)
            inputs[name] = hashlib.sha256(raw).hexdigest()

        res = sbx.exec(FENCE)
        if res.exit_code in (124, 137):
            raise TimeoutError(
                f"{study.name}: decoder ran away, study quarantined")
        if res.exit_code != 0:
            # The guest's stderr is deliberately filenames and tag NAMES
            # only. A traceback carrying a PatientName is a disclosure
            # event wearing a log line.
            raise ValueError(
                f"{study.name}: refused -- {res.stderr.strip()[:300]}")

        # Read back EXACTLY the paths we expect and nothing else. No
        # directory walk, no glob: the derived-artefact list is ours, not
        # the guest's, so a compromised guest cannot nominate a file for
        # exfiltration by writing it somewhere we were going to look.
        out = json.loads(sbx.filesystem.read("/work/out/manifest.json"))
        for frame in out["frames"]:
            store_derived(study.name, frame, sbx.filesystem.read(
                f"/work/out/{frame}"))
        return {"study": study.name, "frames": len(out["frames"]),
                "inputs": inputs, "meta": out["meta"]}
    finally:
        sbx.kill()   # Teardown is kill(). There is no sbx.delete().


def run_batch(studies: list[pathlib.Path], workers: int = 64) -> list[dict]:
    # A burst of studies is a burst of VMs. They share a host and a tap
    # driver and nothing else -- no /tmp, no page cache, no kernel.
    with ThreadPoolExecutor(max_workers=workers) as pool:
        return list(pool.map(process_study, studies))

And the guest script, which is the only code here allowed to be wrong in an interesting way. Every gate below runs before `pixel_array` is touched, because that single attribute access is where the C starts executing.

#!/usr/bin/env python3
# /work/deident.py -- runs INSIDE the microVM, with no default route and no
# credentials. Fails closed: anything surprising is a non-zero exit, never a
# best-effort render. The caller treats a refusal as a study to quarantine.
import json, pathlib, sys
import numpy as np
import pydicom
from PIL import Image
from pydicom.uid import (ExplicitVRLittleEndian, ImplicitVRLittleEndian,
                         RLELossless, JPEGLSLossless, JPEG2000Lossless)

# Transfer syntaxes we are willing to hand to a decoder, by UID. Everything
# else is refused unprobed: the retired big-endian syntaxes, the JPEG 2000
# multi-component variants, and -- yes, this is legal DICOM -- MPEG-2 and
# HEVC. If you did not plan to run a video decoder on hospital data, say so
# in a set literal now rather than in a post-mortem later.
ALLOWED = {ExplicitVRLittleEndian, ImplicitVRLittleEndian, RLELossless,
           JPEGLSLossless, JPEG2000Lossless}

# An ALLOWLIST, not a denylist. PS3.15 Annex E enumerates hundreds of
# attributes to remove, and private tags plus vendor duplicates of
# PatientName mean a denylist is a bet you re-place on every new scanner
# model. Downstream needs these twelve, so these twelve are what leaves.
KEEP = ("Modality", "SOPClassUID", "Rows", "Columns", "BitsAllocated",
        "BitsStored", "PhotometricInterpretation", "SamplesPerPixel",
        "PixelSpacing", "SliceThickness", "WindowCenter", "WindowWidth")

ENCAPSULATED = {"1.2.840.10008.5.1.4.1.1.104.1",    # Encapsulated PDF
                "1.2.840.10008.5.1.4.1.1.104.2",    # Encapsulated CDA
                "1.2.840.10008.5.1.4.1.1.104.3"}    # Encapsulated STL

job = json.load(open("/work/job.json"))
MAX_BYTES = int(job["max_px"]) * 2


def die(msg: str) -> None:
    # Filenames and tag NAMES only -- never a tag VALUE. stderr lands in the
    # orchestrator's logs, and your logging stack is not inside the PHI
    # boundary no matter what the architecture diagram says.
    print(msg, file=sys.stderr)
    sys.exit(2)


frames, meta = [], {}
for src in sorted(pathlib.Path("/work/in").glob("*.dcm")):
    raw = src.read_bytes()
    if len(raw) < 132 or raw[128:132] != b"DICM":
        die(f"{src.name}: no DICM magic")
    # force=False: if the file meta group is unreadable we refuse, rather
    # than let the library guess a transfer syntax. Guessing is how a
    # lenient parser ends up feeding wavelet data to a JPEG decoder.
    ds = pydicom.dcmread(src, force=False)
    ts = ds.file_meta.TransferSyntaxUID
    if ts not in ALLOWED:
        die(f"{src.name}: refused transfer syntax {ts}")
    if str(getattr(ds, "SOPClassUID", "")) in ENCAPSULATED:
        die(f"{src.name}: encapsulated document -- a PDF, not a scan")
    if int(getattr(ds, "NumberOfFrames", 1) or 1) != 1 \
            or int(ds.SamplesPerPixel) != 1:
        die(f"{src.name}: multi-frame or colour -- other stage, same jail")
    if int(getattr(ds, "OverlayBitsAllocated", 0) or 0):
        die(f"{src.name}: overlay plane present, may carry burned-in text")
    if str(getattr(ds, "BurnedInAnnotation", "")).upper() == "YES":
        die(f"{src.name}: declares burned-in annotation, send to OCR stage")
    # The allocation bomb. Rows, Columns and BitsAllocated are
    # attacker-controlled integers and numpy will cheerfully try to honour
    # whatever product they describe, before a single pixel is decoded.
    want = int(ds.Rows) * int(ds.Columns) * ((int(ds.BitsAllocated) + 7) // 8)
    if not 0 < want <= MAX_BYTES:
        die(f"{src.name}: declares {want} bytes of pixel data")

    # Decode LAST, after every metadata gate, in the smallest possible
    # scope. This one attribute access is where OpenJPEG, CharLS or a
    # libjpeg descendant runs on bytes a stranger chose.
    arr = ds.pixel_array.astype(np.float32)
    lo, hi = np.percentile(arr, (1, 99))
    img = np.clip((arr - lo) / max(hi - lo, 1e-6) * 255, 0, 255)
    Image.fromarray(img.astype(np.uint8)).save(f"/work/out/{src.stem}.png")

    frames.append(f"{src.stem}.png")
    # Reconstruction, not scrubbing: a fresh dict of permitted scalars. The
    # original dataset is never serialised back out, so nothing we forgot
    # to delete can ride along inside it.
    meta[src.stem] = {k: str(getattr(ds, k, "")) for k in KEEP}

pathlib.Path("/work/out/manifest.json").write_text(
    json.dumps({"frames": frames, "meta": meta, "profile": "allowlist-v3"}))

Throughput: hundreds of short VMs beat a pool you have to trust

The objection to per-study VMs is always cost, and it rests on an assumption about boot time that does not hold on a snapshot-restore platform. On PandaStack a create is a restore of a baked snapshot, not a boot: p50 179 ms, p99 around 203 ms, of which the `/snapshot/load` step is roughly 49 to 80 ms. The first spawn of a template with no baked snapshot is a genuine cold boot at about three seconds, and then the snapshot is baked automatically — so you pay that once per template, not once per study.

Put that next to the work. A 300-slice CT study is one VM, not three hundred, and the decode of those slices is seconds to minutes. Two hundred milliseconds of create against that is rounding error, and it buys you a property no worker pool can offer: the machine is provably clean, because it is new. The alternative is a pool of long-lived workers that you trust to stay clean, where the trust is doing the work and the evidence is doing none.

For fan-out inside one study there is a warmer option. `fork_tree()` snapshots a parent that already has Python up, pydicom imported and your validated job configuration on disk, and boots children from that snapshot so they inherit the parent's running memory. It is capped at sixteen children per call, so wider fan-outs grow the tree breadth-first. Keep one rule absolute: one study per tree. Never fan a single parent out across patients, because the whole point of the parent was that it belongs to exactly one of them. `fork()` is the other call and it is not this one — it clones the disk and the child cold-boots with its own entropy and PIDs, which is the right choice when you want independence rather than warmth.

On timeouts, be precise, because this has bitten people. `ttl_seconds` is an idle clock, not a wall-clock lifetime; a busy sandbox is not reaped. It is the backstop for an orchestrator that died holding a VM full of PHI. It is not a fence around a wedged codec, and one-shot `exec()` does not reliably enforce a timeout argument either. The bound on a runaway decode is `timeout(1)` inside the guest, or it does not exist.

Isolation is a safeguard story, not a certification

This deserves to be said plainly rather than implied. Per-study kernel isolation gives you good answers to real questions an assessor asks — how one patient's data is prevented from reaching another patient's processing run, how you bound what a compromised parser can read, what happens to intermediate artefacts, how you evidence that a run had no network path out. Those are technical-safeguard arguments and they are stronger than the ones a shared worker pool can make.

What it is not: a certification, an attestation, or a business associate agreement. PandaStack does not sign a BAA today, and no architectural property substitutes for one. If you are processing real PHI in production you need a BAA with whoever holds the bytes, plus the administrative and physical safeguards that have nothing to do with hypervisors. Build the technical story on this shape if you like it; do not let it stand in for the paperwork. Running Code Over PHI Without Expanding Your HIPAA Blast Radius goes through that boundary in more detail.

Where I would not use this

We have no GPUs and no GPU passthrough, and medical imaging is the field where that bites hardest. Volumetric segmentation, a 3D reconstruction, or any of the deep-learning inference people actually want to run on a study wants a GPU, and GPU-accelerated JPEG 2000 decode exists for a reason. The honest split is to do parsing, validation and de-identification here — the part that is dangerous rather than compute-heavy — and hand the derived, de-identified volume to a GPU host that never sees an original object. If your workload is GPU-first, this is the wrong platform for that stage and I would rather say so than sell you a workaround.

The guest kernel is 5.10, so tooling that needs something newer will not be happy. The baked template size governs RAM — `base` is 4 GiB, and Firecracker cannot change vCPU or RAM at snapshot restore — so a reconstruction that wants 64 GiB resident is not a thing you squeeze out of this by passing a bigger number at create time. A long-lived DIMSE C-STORE listener is a server rather than a burst job: you can run one in a persistent sandbox, but it does not get the per-study-disposable property, so terminate DIMSE at the edge and fan out behind it. And if your inputs are one vendor's files from one scanner you administer, a hardened container is a defensible place to stop.

The failure you will actually meet, incidentally, is denial of service rather than remote code execution: a decoder that loops, a pixel geometry that asks for 40 GiB, a decompression bomb. Design for that first. The isolation is there for the rarer, worse day, and the nice thing about building for the rare case is that the cheap case comes free — a wedged codec in a disposable VM is one timed-out study and one deleted machine, not a restart loop that drops other patients' work on the floor.

Frequently asked questions

Is pydicom safe because it is written in Python?

Partly, and the part that is not safe is the part that matters. pydicom's dataset and tag parsing is pure Python, so the classic memory-safety bug classes — heap overflow, use-after-free, out-of-bounds read — are largely off the table in the tag parser itself. That is a real advantage over a C++ parser and worth having. The pixel handlers are a different story: pydicom does not decode JPEG, JPEG-LS or JPEG 2000 itself. It delegates to whatever plugin you installed, which is typically pylibjpeg with its openjpeg, libjpeg and rle backends, or GDCM, or Pillow. Those are C and C++ libraries, and the moment you touch `pixel_array` you are executing them on attacker-supplied bytes. Check which handler your environment actually resolves to rather than assuming; it depends on what is installed, and it can differ between your laptop and your production image. Even where the code is pure Python you still have availability and memory problems: unbounded allocation driven by Rows, Columns, NumberOfFrames and BitsAllocated, numpy reshapes sized from those same attacker-controlled integers, deflate bombs in the deflated transfer syntax, and pathological sequence nesting. None of those are remote code execution. All of them take out a shared worker, and in a shared worker that means taking out other patients' jobs too.

Can a library de-identify DICOM reliably, or do I need a human in the loop?

For a research release to third parties, you need a review step, and anyone who tells you otherwise has not audited the output of their own pipeline. For an automated internal pipeline you can get to a defensible place, but only by changing shape. PS3.15 Annex E is a denylist of attributes with actions, and denylists leak: identifiers hide in private tags that no dictionary describes, in free-text fields like StudyDescription and SeriesDescription where clinicians type names and dates, in structured-report content, in UIDs that encode site information, and in vendor-specific duplicates of PatientName. Every new scanner model is a fresh opportunity for your list to be incomplete. The stronger pattern is reconstruction: decide which attributes your downstream genuinely needs, build a new object containing only those, and discard the original rather than editing it. That fails closed. Then there is the pixel problem, which no attribute-level tool solves at all. Burned-in annotation is routine in ultrasound, secondary capture and workstation screen captures, the BurnedInAnnotation attribute is unreliable as a gate, and overlay planes in the 60xx groups can carry the same text as a separate bitmap. Finding it means OCR, which is more untrusted-input C code, and OCR on medical images has a false-negative rate. Practically: allowlist what you emit, refuse modalities where burned-in text is likely unless they have been through an OCR stage, sample the output for human review, and keep the review rate honest rather than aspirational.

One microVM per study or one per file?

Per study is the default, and the reason is that the blast radius should match an existing trust boundary rather than an arbitrary one. Objects within a study belong to the same patient, so a poisoned object reaching its siblings tells you nothing new — they were going to be processed together anyway. Objects from different patients have no business sharing a kernel, a page cache or a scratch directory. Per study also keeps the arithmetic sane: a 300-slice CT is one create at a p50 of 179 ms, against a decode measured in seconds to minutes, so the isolation is close to free. Per file makes sense in two situations. The first is triage of a mixed archive where you have no reliable study grouping yet, because the grouping metadata is the thing you do not trust — then each object genuinely is unrelated to its neighbours. The second is when a specific object is already suspicious: it failed a gate, it came from an unknown source, or it is the one that crashed the decoder last night. Give that one its own machine and keep the artefacts. For fan-out within a study, `fork_tree()` gives you children that inherit a warm parent's memory, capped at sixteen per call. The rule there is one study per tree, always: the parent's whole value is that it belongs to exactly one patient, and fanning it across patients throws that away for a few hundred milliseconds.

Does the microVM boundary actually help here, or is it security theatre on top of a container?

It helps, and it helps for a specific reason rather than a general one: it changes what a single bug buys. In a container, a memory-safety bug in OpenJPEG gives an attacker code execution in a process that shares a kernel with every other workload on the node, so the next step is a kernel bug and the prize is the node. In a microVM the same bug gives code execution inside a guest kernel that owns one study and will be deleted shortly, and the next step is a hypervisor escape through a device model that is deliberately minimal — a handful of virtio devices and a serial port, with no HPET, no sound card, no emulated legacy hardware. That is a much thinner target, and the VMM is Rust rather than decades-accumulated C. It is not magic, and I would rather be specific about the limits. The VMM is still software and still has a published advisory history; read it rather than assuming. Microarchitectural side channels cross the VM boundary — /blog/spectre-meltdown-microvm-isolation covers what that means in practice — so a microVM is not a confidentiality guarantee against a co-tenant with patience. Nothing here protects you from your own orchestrator leaking PHI into logs, which is the far more common incident. And none of it is a substitute for the boring controls: least-privilege credentials, no network egress from the parsing stage, and an emitted-artefact list that your code chose rather than the guest.

Keep reading

Related posts

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.