Bioinformatics Pipelines on microVMs: Reproducibility, PHI and the 400-Tool Dependency Graph
Ask a bioinformatics platform engineer what their hardest problem is and they will not say CPU. They will say that stage four wants `samtools` built against one `htslib`, stage nine wants a `bcftools` built against a different one, the GATK step brings its own bundled Java and a resource bundle keyed to one exact reference build, and somewhere in the DAG a postdoc's R script calls `install.packages()` at runtime because that was the fastest way to get the figure done before the grant deadline.
That is the job. The arithmetic is well understood and somebody published it. The job is a dependency graph with hundreds of nodes where two nodes in the same pipeline want incompatible versions of the same shared library, and where "reproducible" has to mean more than "it ran last March".
I'm Ajay; I build PandaStack, which runs code in Firecracker microVMs. This is for people operating the platform rather than writing the analysis: a hosted variant-calling service, a core facility running Nextflow, Snakemake or WDL for a dozen labs, a biotech SaaS where customers upload their own FASTQ files and sometimes their own code. I want to be precise about what containers solved, what they did not, and where a per-job VM boundary earns its keep.
The dominant engineering cost is the dependency graph, not the compute
Take one ordinary short-read germline workflow: align with `bwa mem`, sort and index with `samtools`, mark duplicates, recalibrate against a known-sites bundle, call variants per interval with GATK, merge the per-interval GVCFs, genotype, filter, annotate, hand the result to an R/Bioconductor stack for the figures. Every step carries a version pin, and the pins are not independent.
- The aligner's index is version-coupled to the aligner. "Which index files are on this machine" is part of the environment, not part of the input.
- `samtools` and `bcftools` are thin tools over `htslib`. Two stages wanting different `htslib` versions want different libraries in the same process namespace — not a packaging failure, just two correct programs with incompatible requirements.
- The reference build is a global invariant that nothing enforces. GRCh37, GRCh38 and T2T-CHM13 are different coordinate systems, and even within one build the contig naming differs between distributions: `chr1` versus `1`. A pipeline that mixes a `chr`-prefixed BAM with a bare interval list does not crash. It produces a VCF with an answer in it.
- GATK-style tools want a sequence dictionary and a FASTA index beside the reference, from a specific tool version, plus a known-sites bundle matching the build. The reference is not a file; it is a bundle with internal consistency requirements.
- The R/Bioconductor layer has its own release train pinned to an R minor version — and at the leaves sit the awk one-liner and the R script that is now load-bearing.
Take the transitive closure and you are managing hundreds of pinned artefacts whose compatibility matrix nobody wrote down. CPU hours you can forecast; the dependency graph is what eats engineer-months, and it is why a field full of people who would rather be doing statistics became experts in package management.
What containers actually solved, and I mean actually
The field's answer was per-process containers, and it was right. Bioconda gave the ecosystem a convention-driven recipe repository; BioContainers turned those recipes into images automatically, including multi-package images for tools that must genuinely coexist. Nextflow's `container` directive and Snakemake's `container:` and `conda:` directives moved the environment declaration out of a wiki page and into the workflow definition, per process. Apptainer made that work on clusters where nobody was ever going to be handed a Docker daemon, by running the image as the invoking user out of a single-file SIF. The consequence: stage four and stage nine declare different images, and the two incompatible `htslib` builds never meet. The conflict stopped being a conflict, and the environment became an artefact you pull by digest.
Three things containers did not solve
Containers solved packaging. They did not solve isolation, privilege or environmental determinism, because those were never properties of the image. They are properties of the kernel the image runs on, and the container shares that kernel with everything else on the box.
Multi-tenancy: you run other people's code against identifiable data
A single-lab pipeline on a single-lab cluster has no multi-tenancy problem worth writing about: everyone on that node sits under the same acceptable-use policy. A core facility is not that. You execute pipeline definitions written by people you have never met, pulling images you have never audited, invoking scripts written in an afternoon, against data that is frequently identifiable human genomic data — the kind where your contract has specific words in it about isolation.
On a shared host kernel, the technical content of that isolation claim is: no local privilege escalation will be found in this kernel, or we will have patched it before anyone here tries. Serious platforms hold that position deliberately. But its failure mode is categorical rather than incremental — the boundary does not degrade, it stops existing — and it is a hard sentence to write into a questionnaire a hospital's compliance office will read.
The second-order version bites more often. Namespaces and cgroups are excellent at stopping accidents and mediocre at stopping resource interference. A pipeline that allocates aggressively or holds tens of thousands of file descriptors through a merge escapes nothing — it just makes somebody else's six-hour alignment take nine, and you are now debugging a performance incident across tenant boundaries with no clean attribution.
Root: a lot of real bioinformatics tooling assumes it owns the machine
This one is invisible in architecture diagrams and dominant in practice. A great deal of genuinely useful scientific software was written by people solving a biology problem on a workstation they administered, and it assumes it can write wherever it likes.
- Tools that write auxiliary outputs next to their inputs, so a read-only reference bundle breaks them in a way the error message does not explain — and R or Python steps calling `install.packages()` or `pip install` mid-pipeline, which need write access to a library path inside the image.
- Tools that want a raised `ulimit -n`, or a sysctl like `vm.max_map_count` lifted before a merge of thousands of files will succeed. Network sysctls are namespaced; the virtual-memory ones are not, so in a container you cannot change them and in your own kernel you can.
- Anything that wants a loop device, a FUSE mount or a filesystem it can `mount` — ordinary requests from software that expects to be root, plus the custom script that opens with `apt-get install -y`, which is not best practice and is also what your customer's pipeline does.
The workarounds are where pipelines go to die: rootless runtimes, user namespaces, subuid ranges, fakeroot modes, bind-mounting a writable overlay over the part of the image the tool wants to write to, then finding the second place it also wanted. Each is tractable; the aggregate is a support queue whose tickets all read "works on my laptop, fails on your platform" — because on the laptop the researcher was root. Inside its own microVM the job is root on a kernel that belongs to it and dies with it, and you stop negotiating with the software about privilege.
Reproducibility: the image is pinned, the kernel is not
This is the part bioinformaticians should care about most, because the field takes reproducibility more seriously than almost any other and has been quietly accepting a hole in it. A container pins user space. It does not pin the kernel, or what the kernel reports about the machine — so pipeline behaviour can depend on the host, specifically and checkably:
- `nproc` inside a container is, by default, the host's CPU count, not the cgroup's quota. Every pipeline written as `-t $(nproc)`, every tool that auto-sizes its thread pool, and R's `detectCores()` therefore run differently on a different host. `/proc/meminfo` leaks the same way for anything sizing a heap or a sort buffer.
- Thread count is not always cosmetic. The common short-read aligner estimates the insert-size distribution from a chunk of reads whose size depends on the thread count, which is why careful pipelines pin the chunk size explicitly. Change the host, change the chunking, change a handful of mate-rescue decisions, change the BAM.
- The instruction set the process sees decides which path a hand-vectorised inner loop takes. Variant callers ship SIMD-accelerated likelihood implementations with a scalar fallback, and the two are not always bit-identical at the margins — a property of the host CPU, not of your image.
- And the kernel itself: syscall availability, directory read order, memory overcommit policy, OOM behaviour under pressure. A pipeline that completes under generous overcommit and gets killed without it is not reproducible, it is lucky.
None of that makes containers bad. It makes the claim narrower than the way it is usually stated: same image, same answer, on hosts that agree about everything the image does not pin — and the hosts are upgraded by someone who is not you.
The pipeline is bit-for-bit reproducible, as long as nobody upgrades the host kernel. Reproducible until Tuesday.
A microVM closes that gap by construction, because the kernel is inside the artefact. PandaStack guests run kernel 5.10, booted from the template's own snapshot, with vCPU count and RAM baked into that snapshot — Firecracker cannot change either at restore, so a create request's size is overridden to match what was baked. So `nproc` reports the guest's vCPUs, `uname -r` reports a kernel you chose and versioned, and neither changes because the fleet underneath you had a maintenance window.
The architecture: run the container inside the microVM
The resolution is not to pick one. Containers and VMs answer different questions. The image is the unit of packaging; the microVM is the unit of isolation and of per-tenant blast radius. Stack them.
- A job arrives: a pipeline definition, a sample sheet, a reference build, a tenant identity. Your control plane validates it and decides what it may touch.
- Create one microVM per job (or per stage group) from a template you baked: workflow engine, container runtime, and the reference bundle already on the rootfs. Every create restores that template's baked Firecracker snapshot, p50 179 ms and p99 203 ms. There is no warm pool of idle VMs; the first-ever spawn of a new template cold-boots in about 3 s and bakes its own snapshot.
- Inside the guest, the workflow engine runs exactly as it does today — the `container` directive, the `conda:` directive, a WDL runner's image pins, all unchanged, because from the engine's point of view the microVM is simply the machine it was given.
- Scope credentials to that one job: this tenant's inputs, this tenant's output prefix, nothing else. The VM boundary does nothing about an over-broad token, and a pipeline handed a bucket-wide credential will eventually use it.
- Stream stdout and stderr out while the stage runs. A six-hour alignment that reports nothing until it finishes is indistinguishable from one that wedged in minute four.
- Write intermediates and outputs to a durable volume or straight to object storage from inside the guest, keyed idempotently.
- Delete the VM. The kernel, page cache, temp directory, half-written sort spill and the `apt-get` the customer's script ran all go together, because they were never anywhere else.
Data gravity is the real constraint, and a 179 ms boot does not fix it
I have to be honest about the economics, because the fast-boot pitch is mostly irrelevant here. A whole-genome FASTQ pair is tens of gigabytes. Alignment is I/O-bound and memory-hungry, sorting spills to disk, and the reference bundle with its indices is itself gigabytes that every job needs before it can do anything. Against that, a sub-200 ms create time is a rounding error. There is exactly one place a microVM genuinely changes the time budget.
Bake the reference bundle into the template rootfs
Build a template whose rootfs already contains the reference FASTA, its indices, its sequence dictionary and the known-sites bundle for the build you support, next to the workflow engine and the container runtime. That rootfs lands on each host once. From then on every job gets it as a local file cloned by XFS reflink — metadata-only copy-on-write, so N concurrent jobs share the same physical blocks until one writes. The per-job fetch becomes a per-host, per-template-version fetch.
One caveat, because people assume otherwise. PandaStack can demand-page a snapshot's guest memory from object storage over ranged HTTP, skipping all-zero chunks via a baked header. That streaming is for memory only: the rootfs must always be a local file, because reflink and dm-snapshot copy-on-write need a local block device. A large reference template really is that large on every host that runs it.
`fork()` for fan-out, with its semantics stated precisely
`fork()` reflinks the parent's disk and then cold-boots the clone. The child inherits the parent's filesystem — reference bundle, tool stack, the sorted BAM the parent just produced — and does not inherit guest memory: fresh kernel boot, own process tree, own PIDs, own entropy pool. Be clear about the cost. A `fork()` child pays roughly a cold boot, the ~3 s shape, and what you buy for it is the disk sharing. The sub-second verb is `fork_tree(count)`, which snapshots the parent once and restores N children from that snapshot, so those children inherit memory as well as disk — same-host restores in 400-750 ms, cross-host 1.2-3.5 s when artefacts must come from object storage first. Both cap at 16 children per call.
For a per-sample or per-chromosome fan-out the slower verb is the right verb, and that is not a consolation prize. What the children need to share is on disk: the reference index, the recalibrated BAM, the interval lists. What you do not want shared is memory, because each shard starts its own JVM, thread pool and temp directory, and a per-interval caller that inherited a sibling's heap is a debugging story nobody wants. Cold-booting also gives each child its own entropy — a fan-out whose children share a pre-seeded PRNG agrees with itself, which looks like reproducibility and is a bug. Three seconds against a multi-minute calling stage is not a number anybody will notice. Save `fork_tree()` for branching live state, which a batch pipeline does not have.
One honest caveat: a cold-booted child does not inherit the parent's warm page cache, so it faults the reference index back in from its own reflinked copy. A page-cache cost rather than a network one, but not free, and the reason fan-out width should follow your per-shard work rather than the contig count.
Scatter-gather, and the worker that dies at hour three
Scatter-gather is the native shape here: split by sample, split by interval, call per interval, merge. The merge is what people under-engineer, because the happy path merges fine. Three rules make a fan-out survivable.
- Stage outputs are idempotent and content-keyed. The output path is a function of everything that can change its bytes: sample id, interval set, tool versions, reference build, the parameters that matter. Re-running a completed shard is then a no-op you detect with one existence check, instead of a question about which pipeline version wrote the file on disk.
- Writes are atomic: temp name on the same filesystem, fsync, rename. A half-written GVCF that looks finished is how a scatter-gather pipeline returns a wrong answer instead of an error, and the gather step cannot tell the difference.
- The gather step tolerates a missing shard. A shard VM can die at hour three — a host fails, a TTL was set too tightly, a step gets OOM-killed inside the guest. The gather must notice the gap and re-dispatch exactly that interval set, which is cheap precisely because the work is interval-scoped and the output content-keyed.
That last rule is why intermediate BAMs do not belong on the ephemeral rootfs. The rootfs is copy-on-write and disposable by design, and when the VM goes so does everything on it. Losing a sort spill is fine; losing a recalibrated BAM that took four hours, because a shard three steps later failed, converts a retry into a re-run. Attach a durable volume and put the stage boundaries on it. If a stage needs relational state — a run table, a dispatcher cursor — a managed Postgres instance is a dedicated microVM with its own durable volume, reachable over TLS, and takes 30-90 s to create because it actually bootstraps a database.
Here is the per-sample scatter. Note what gets streamed and what gets forked.
import hashlib
import json
from pandastack import Sandbox
TEMPLATE = "bio-gatk-grch38" # custom template: tool stack + reference bundle
REF_BUILD = "GRCh38_no_alt" # baked in; a reflinked local path at create time
VOLUME = "bioruns-2026q4" # durable volume: where intermediates live
INTERVALS = [f"chr{i}" for i in range(1, 23)] + ["chrX", "chrY", "chrM"]
SHARDS = 6 # intervals per child, not one VM per contig
def stage_key(sample_id: str, stage: str, intervals: list[str]) -> str:
"""Content key for a stage output. Anything that can change the bytes
goes in: tool versions, reference build, the interval set itself."""
h = hashlib.sha256()
for part in (sample_id, stage, REF_BUILD, "bwa-0.7.17", "gatk-4.5.0", *intervals):
h.update(part.encode() + b"\0")
return h.hexdigest()[:16]
def align_and_call(sample: dict) -> list[str]:
sbx = Sandbox.create(
template=TEMPLATE,
ttl_seconds=8 * 3600, # alignment is measured in hours
volumes=[VOLUME], # /vol survives the VM; the rootfs does not
metadata={"sample": sample["id"], "stage": "align", "ref": REF_BUILD},
)
# The sample sheet goes in as a FILE, not as a shell argument. Sample ids
# come from a LIMS, and a LIMS will eventually hand you one with a space
# in it, or a quote, or a newline.
sheet = f"/work/{sample['id']}.sheet.json"
sbx.filesystem.write(sheet, json.dumps(sample, indent=2))
# exec_stream, not exec: this runs for hours. You want the aligner's and
# sorter's progress lines in your own log pipeline WHILE it happens, not a
# single stdout blob at the end that you may never receive.
sbx.exec_stream(
f"/opt/pipeline/align.sh {sheet}",
on_stdout=lambda line: log(sample["id"], "out", line),
on_stderr=lambda line: log(sample["id"], "err", line),
)
# fork() = reflink the disk + COLD BOOT the child. Children inherit the
# reference index and the sorted BAM on DISK; they do NOT inherit memory.
# Each gets its own kernel, process tree, JVM and entropy pool -- exactly
# right for a per-interval caller, because there is no live state here
# worth branching, only files. It costs about a cold boot (~3 s) and caps
# at 16 children per call; fork_tree() is the sub-second verb and inherits
# memory, which is the opposite of what this stage wants.
groups = [INTERVALS[i::SHARDS] for i in range(SHARDS)]
children = [sbx.fork() for _ in groups]
gvcfs = []
for idx, (child, group) in enumerate(zip(children, groups)):
key = stage_key(sample["id"], "haplotypecaller", group)
out = f"/vol/gvcf/{sample['id']}/{key}.g.vcf.gz"
# call.sh writes out.part, fsyncs, then renames.
child.exec_stream(
f"/opt/pipeline/call.sh --out {out} --intervals {','.join(group)}",
on_stdout=lambda line, i=idx: log(f"{sample['id']}/s{i}", "out", line),
on_stderr=lambda line, i=idx: log(f"{sample['id']}/s{i}", "err", line),
)
gvcfs.append(out)
# Shown sequentially for clarity; in production these go on a thread pool,
# because the whole point of the fan-out is concurrency.
for child in children:
child.delete()
sbx.delete()
# The gather step reads these off the VOLUME. Nothing it needs was ever
# only on a rootfs, so a shard that died at hour three costs one shard.
return gvcfsAnd the stage itself, inside the guest — where the composition becomes obvious. The container is still doing the packaging; the guest kernel is doing the isolation.
#!/usr/bin/env bash
# /opt/pipeline/align.sh -- runs INSIDE the microVM.
#
# The microVM is the tenant boundary. The container is still the packaging
# unit. Composed: the guest runs the same BioContainers images the pipeline
# declares, on a kernel that belongs to this job alone.
set -euo pipefail
SHEET="$1"
SAMPLE=$(jq -r .id "$SHEET")
R1=$(jq -r .fastq_1 "$SHEET")
R2=$(jq -r .fastq_2 "$SHEET")
# Baked into the template rootfs: reflinked at create, never fetched per job.
REF=/ref/GRCh38_no_alt/GRCh38_no_alt.fa
OUT=/vol/bam/$SAMPLE # durable volume -- survives this VM
mkdir -p "$OUT" /work/sort
# nproc here is the GUEST's baked vCPU count, and it is the same number on
# every host in the fleet. Inside a container it would be the HOST's core
# count, which is why '-t $(nproc)' is a reproducibility bug that only shows
# up the day you buy a bigger machine.
THREADS=$(nproc)
# uname -r is part of the artefact now. Log it next to the tool versions;
# future-you will want to know which kernel produced which VCF.
echo "kernel=$(uname -r) threads=$THREADS ref=$REF" >&2
# -K pins the chunk size used to estimate the insert-size distribution. Leave
# it out and the alignment depends on how many threads happened to be
# available, which is a delightful thing to discover mid-study. Pin the FULL
# BioContainers tag including its build suffix -- the suffix is part of the
# pin, not decoration.
apptainer exec --cleanenv --bind /ref:/ref:ro --bind /vol:/vol --bind /work:/work \
docker://quay.io/biocontainers/bwa:0.7.17--BUILD \
bwa mem -t "$THREADS" -K 100000000 \
-R "@RG\tID:$SAMPLE\tSM:$SAMPLE\tPL:ILLUMINA" \
"$REF" "$R1" "$R2" \
| apptainer exec --cleanenv --bind /vol:/vol --bind /work:/work \
docker://quay.io/biocontainers/samtools:1.19.2--BUILD \
samtools sort -@ 4 -T /work/sort/$SAMPLE -o "$OUT/$SAMPLE.sorted.bam.part"
sync
mv "$OUT/$SAMPLE.sorted.bam.part" "$OUT/$SAMPLE.sorted.bam"
# Already have a pipeline? Then none of the above is your code -- the sandbox
# is just the machine the executor was handed, and the workflow is unchanged:
#
# nextflow run <your-pipeline> -r <tag> -profile apptainer \
# --input "$SHEET" --igenomes_base /ref \
# -work-dir /vol/work --max_cpus "$THREADS" -resume
#
# Note -work-dir on the durable volume: that is what makes -resume mean
# something after a worker dies, instead of meaning 'start again'.HPC, Kubernetes, per-job microVM: where each one wins
| Dimension | HPC cluster + Apptainer | Kubernetes + containers | Per-job microVM (PandaStack) |
|---|---|---|---|
| Multi-tenancy | Shared host kernel across every user on the node. The scheduler enforces fair share, not a security boundary — correct for the environment it was designed for, where every user sits inside one institution's AUP | Shared host kernel across pods. Namespaces, cgroups, seccomp and admission policy do the work; strong against accident, and the remaining escape surface is the kernel itself | One Firecracker guest kernel per job, with its own network namespace and veth/tap pair from 16,384 pre-allocated /30 subnets per host agent. A tenant's blast radius is a kernel that belongs to it |
| Root inside the job | Deliberately unavailable. Apptainer keeps the invoking user as themselves and SIF contents read-only, so tools that write into their own prefix need binds or an overlay | Possible, widely discouraged, often forbidden by cluster policy. Rootless runtimes, user namespaces and subuid ranges make it workable and not simple | Root in the guest by default, because it is the guest's own kernel. A non-namespaced sysctl, a loop device, a FUSE mount, an `apt-get install` — all local to this job and gone with it |
| Kernel pinning | Whatever the node runs. Varies across partitions and maintenance windows; `uname -r` is an input you do not control and usually do not record | Whatever the node pool runs, upgraded by the platform on its own cadence. The image is pinned; the kernel, `nproc` and `/proc/meminfo` are not | Guest kernel 5.10 from the template's baked snapshot, with vCPU and RAM baked in too. `nproc` and `/proc/meminfo` report the guest, so the environment is part of the versioned artefact |
| Data locality | The strongest of the three, and the best single reason to stay: a parallel filesystem the reference bundle already lives on, sized by people who do this for a living | A function of the storage class you chose — a CSI volume, an object store with a sidecar, or a per-pod reference download with a node cache you do not control | Reference bundle baked into the template rootfs and reflinked per job, so concurrent jobs share physical blocks until they write. Intermediates go on a durable volume; guest memory can stream from object storage, the rootfs cannot |
| Operational burden | A cluster, a queue, allocations, module files and a team who know it. Mature, well understood, and not something you buy with a credit card this afternoon | A cluster, a CNI, a CSI, an autoscaler and the upgrade treadmill — in exchange for the largest ecosystem in infrastructure | An API call. You own the workflow logic, the retries and the gather step, because nothing dispatches for you and nothing retries for you either |
Read that as a routing function, not a scoreboard. One institution running its own pipelines on its own data, with a parallel filesystem and a sysadmin team, should stay on HPC with Apptainer — the data-locality row is why, and it is not a legacy choice. If you already run Kubernetes well with tenants who are all internal, a per-job kernel is real value that is not urgent. The case sharpens as your tenants become strangers and your inputs become identifiable.
Compliance, without the legal cosplay
I am not a lawyer and PandaStack is not a compliance product. What I can talk about is the architectural property, because that is what survives contact with a security questionnaire. When a hospital's information security office asks how tenant A's genomic data is kept away from tenant B's pipeline, there are two shapes of answer. One: a shared kernel, with namespaces, cgroups, a seccomp profile, admission policy and a patch SLA — defensible, and the reviewer's follow-up will be about kernel CVEs, because that is where the boundary lives. The other: tenant B's pipeline runs on its own kernel, behind its own virtual NIC, and the host-side surface it can address is a hypervisor with a deliberately small device model rather than the full Linux syscall interface. The second is easier to write down and easier for a non-specialist to evaluate, which matters because a security review is a document exercise with a human reading it.
I go through that questionnaire line by line, including the questions where my honest answer is no, at Answering a Security Questionnaire When You Run Customer Code, and the PHI-specific version is at Running Code Over PHI Without Expanding Your HIPAA Blast Radius. What I will not do is dress an architectural property up as a certification. A microVM boundary is not a BAA, not an attestation, not an audit, and not a substitute for access logging, key management, retention policy or the agreements you sign. Per-tenant kernels make the isolation section easier to argue and leave every other section exactly as hard.
What this does not fix
A post that stopped there would be an advert. The real limits, several of which decide whether this suits you at all:
- Data movement still dominates. All of the above improves the reference-bundle half of the I/O budget and nothing about the FASTQ half. If your inputs come from far away, that transfer is your architecture.
- RAM is a template decision, not a request parameter. Firecracker cannot change vCPU or RAM at restore, so the agent overrides a create request's `cpu` and `memory_mb` to the baked values — and alignment and sorting are exactly the stages that want to ask for more memory at submit time. You maintain a small family of templates at the sizes your stages need.
- No GPU. If your pipeline calls variants on an accelerator or runs a deep-learning basecaller, this is the wrong platform, not a compromise to engineer around.
- Egress is open by default. A few targeted DROP rules exist for known-abuse protocols; that is not a default-deny network. If your PHI posture requires that a pipeline cannot reach the internet, that is policy you impose deliberately — and assume a customer's pipeline will try to fetch something from a URL in a config file.
- This is not a workflow engine. No DAG, no retry policy, no backfill, no provenance database. Bring Nextflow, Snakemake, Cromwell, miniwdl or your own dispatcher; a microVM is a good execution unit for an engine to call, not a replacement for one.
- A pinned kernel does not make a non-deterministic tool deterministic. An unseeded PRNG, a timestamp in an output, a dependence on directory read order — none of those were reproducible before and none are now. What you gain is that the environment stops being a variable.
Where to start
Do not migrate the pipeline. Migrate one stage, and pick the one that generates the most support tickets — whichever has needed a privilege workaround. Bake a template with your tool stack, your container runtime and the reference bundle for one build, and run that stage for one sample, streaming the logs, writing output to a durable volume under a content key. Set the TTL like you mean it: a six-hour alignment in a short-TTL sandbox gets reaped at hour one, and the error you chase will be a missing output rather than a timeout.
Then do the thing nobody does and should: log `uname -r`, `nproc`, the full image digests and the reference build beside the output, and run the same sample on two different hosts. Diff the results. If they differ, you have found a host dependency that was in your pipeline the whole time and that nothing in your container story was ever going to catch.
That is the pitch, and it is narrower than "microVMs are faster". You are not replacing containers; the images stay where they are. You are putting a kernel boundary around each tenant, giving each job root over a machine it owns for an hour, and making the environment a versioned part of the artefact rather than a property of whichever host the scheduler picked. The reproducibility you have been claiming becomes the reproducibility you have — including on Tuesday.
Frequently asked questions
Isn't a container enough for reproducibility? We pin every image by digest.
Pinning by digest pins user space, completely and correctly — that was the hard part, so you have done the valuable majority of the work. What it does not pin is the kernel, or what the kernel tells your tools about the machine, and those failures are specific. `nproc` inside a container reports the host's CPU count rather than the cgroup's quota, so anything written as `-t $(nproc)`, any tool that auto-sizes its thread pool, and R's `detectCores()` behave differently on a different host. That is not always cosmetic: the common short-read aligner estimates the insert-size distribution from a chunk whose size depends on thread count, which is why careful pipelines pin the chunk size rather than letting it float. `/proc/meminfo` leaks the same way for anything sizing a heap or sort buffer. Hand-vectorised likelihood code picks a SIMD path or a scalar fallback based on the host CPU, and the two are not always bit-identical at the margins. And the kernel decides syscall availability, directory read order, overcommit policy and OOM behaviour under pressure. So the real claim is "same image, same answer, on hosts that agree about everything the image does not pin" — and the hosts are upgraded by someone who is not you. A microVM pins the guest kernel as part of the template snapshot, which makes `uname -r` and `nproc` properties of the artefact rather than of the fleet.
What about the 30 GB of input data? Doesn't that make a 179 ms boot pointless?
Mostly yes, and I would rather say so than pretend otherwise. A whole-genome FASTQ pair is tens of gigabytes, alignment is I/O-bound and memory-hungry, sorting spills to disk, and against that a sub-200 ms create time is a rounding error. If the job then spends twenty minutes pulling data, boot time was never your problem. Where the microVM does change the budget is the reference bundle — the thing every job needs and that nobody's input actually is. Bake the FASTA, its indices, the sequence dictionary and the known-sites bundle into the template rootfs. That rootfs lands on each host once, and from then on every job gets it as a local file cloned by XFS reflink: metadata-only, so concurrent jobs share the same physical blocks until one writes. You have converted a per-job fetch into a per-host, per-template-version fetch. The honest limit: guest memory can be demand-paged from object storage over ranged HTTP, but the rootfs cannot, because copy-on-write needs a local block device, so a large reference template really is that large on every host that runs it. For the FASTQ side the only answers are boring — keep compute in the same region as the data, stream from object storage inside the guest, and content-key your stage boundaries so you never move the same bytes twice.
Do I have to rewrite my Nextflow or Snakemake pipeline to run on microVMs?
No, and if the answer were yes I would not be recommending it. From the workflow engine's point of view the microVM is simply the machine it was given. The `container` directive, Snakemake's `container:` and `conda:` directives, a WDL runner's image pins — all unchanged, because the images are still the unit of packaging and the engine still pulls them. That is the whole composition: run the container inside the VM. The image handles the dependency graph, which is what the field spent a decade getting right through Bioconda and BioContainers; the VM handles isolation, privilege and kernel pinning, which images were never going to handle. What changes is the boundary your executor talks to — you create a sandbox per job or per stage group from a template that already carries the engine, the container runtime and the reference bundle, then run the pipeline inside it. Two details are worth getting right first time. Put the engine's work directory on a durable volume rather than the ephemeral rootfs, or `-resume` becomes a word with no meaning after a worker dies. And set the sandbox TTL to the real duration of the stage: a six-hour alignment in a short-TTL sandbox gets reaped, and the symptom you chase will be a missing output rather than a timeout.
Where should intermediate BAMs live, and what happens when a shard dies at hour three?
On a durable volume, and the shard should die cheaply. The rootfs is copy-on-write and disposable by design: when the VM goes away so does everything written to it. That is correct for a sort spill and catastrophic for a recalibrated BAM that took four hours, because losing it turns a retry into a full re-run. So put the stage boundaries on a volume and leave genuinely cheap-to-recompute intermediates on the rootfs. Surviving a dead shard then needs three habits. Content-key the stage output: the path is a function of sample id, interval set, tool versions, reference build and the parameters that matter, so re-running a completed shard is a no-op you detect with one existence check rather than a question about which pipeline version wrote the file. Write atomically — temp name on the same filesystem, fsync, rename — because a half-written GVCF that looks finished is how scatter-gather returns a wrong answer instead of an error. And make the gather step tolerate a gap: notice the missing shard, re-dispatch exactly that interval set, carry on. Hosts fail, TTLs get set too tightly, and steps get OOM-killed inside the guest because the tool wanted more than the template was baked with. Interval-scoped, content-keyed work makes every one of those a bounded retry.
Does running each job in its own microVM make us HIPAA compliant?
No, and anyone telling you a hypervisor confers compliance is selling you something. Compliance is an organisational programme: agreements, access logging, key management, retention and deletion policy, training, incident response, and an auditor who reads all of it. No execution boundary does any of that, and PandaStack is not a compliance product. What a per-job microVM gives you is one architectural property, and it happens to be the one the isolation section of a security questionnaire is asking about. When a hospital's infosec office asks how one tenant's genomic data is kept away from another tenant's pipeline, the shared-kernel answer is namespaces, cgroups, a seccomp profile and a patch SLA — defensible, and the next question will be about kernel CVEs, because that is where the boundary lives. The microVM answer is that the other tenant's pipeline runs on its own kernel, behind its own virtual NIC, and the host-side surface it can address is a hypervisor with a deliberately small device model rather than the full Linux syscall interface. That is a simpler claim to write and a simpler one for a non-specialist to evaluate. It makes one section easier to argue and leaves the rest exactly as hard.
Keep reading
- Isolating genomics pipelines per job — The multi-tenancy-first version of this argument: untrusted Nextflow/Snakemake/WDL submissions, resource blast radius and teardown.
- Security questionnaire answers, line by line — The isolation questions a hospital's infosec office actually asks, with the answers I give and the ones where I say no.
- PHI and code execution isolation — The healthcare-data version: what a per-tenant kernel does and does not do for a BAA.
- Persisting data in a sandbox — Ephemeral rootfs versus durable volumes — the mechanism behind the intermediate-BAM advice above.
- Slurm and HPC batch scheduling vs microVM fleets — Why the queue, the allocation model and the parallel filesystem are the strongest case for staying on HPC.
Related posts
- Processing PHI in Per-Job microVMs: Isolation for Regulated Healthcare Data
The senders are hospitals with 1998-era interface engines, the formats are containers full of compressed pixel data and XML, and the parsers are C. Give every PHI job a machine you can afford to lose — and know exactly which parts of compliance that does and does not buy you.
- Notebooks in Production: Parameterised Runs in Disposable microVMs
The quarterly board metric comes out of cell 34, which must be run in order, by Dmitri, on his laptop. The notebook is not the problem. The laptop is. Papermill turns the notebook into a batch artefact; a disposable microVM per run turns it into one you can trust.
- Running academic artifact evaluation on microVMs
An artifact-evaluation reviewer gets a tarball, an install.sh with sudo in it, and two weeks to decide whether a paper's numbers survive contact with a second machine. A snapshotted microVM turns that into a machine you can hand to the next reviewer.
- The Best Multi-Tenant Isolation Platforms in 2026
You are not shopping for a sandbox. You are shopping for an isolation boundary between one paying customer and the next — and the right one depends entirely on what your tenants are allowed to do.
- Isolating Per-Tenant CSV and Bulk Import Pipelines in MicroVMs
Your "Import your data" button accepts arbitrary bytes from anyone with a trial account. Zip bombs, 4GB single lines, formula injection, and a spreadsheet parser made of C. Give it a machine you can throw away.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.