Per-Tenant Payroll and Tax Engines in microVMs: Determinism Is the Product
The email that starts this post arrives in 2029 and has a cheerful subject line: amended return. A tenant is refiling a 2026 period, and their accountant would like two things from you. The original figure, which you have, because you kept the payslips. And an explanation of how it was reached, which you do not have, because what you kept was the source tree.
So you check out the tagged commit, rebuild, feed in the archived inputs, and the engine produces a number eleven cents different for fourteen hundred employees. Now you have a genuinely nasty problem, which is not the eleven cents. It is that nobody in the building can tell you whether 2026 was wrong or 2029 is wrong, and until somebody can, you cannot sign anything.
I build PandaStack, an open-source Firecracker microVM platform, and this failure mode is the most interesting argument for per-tenant isolation that nobody makes, because it is not a security argument. A payroll or tax engine is, on paper, a pure function: employee facts, plus a rule set, plus an effective date, produce gross, deductions and net. In practice it is a function of about fifteen more things, none of which appear in its signature, most of which are not in your repository, and several of which are not even in your container image. This post is about which ones, and about shipping the whole machine as the archival artefact instead of the recipe for one.
The hidden arguments to a pure function
Write down the real signature of a payroll calculation. Not the one in the code — the one the answer actually depends on. It is longer than anyone expects, and every entry on it has changed under somebody's feet in production:
- The jurisdiction rule set, and specifically its version as of an effective date. Rule sets get amended retroactively, which means there are two dates in every calculation: the date the rules apply to, and the date you looked them up.
- The rounding mode. Half-up, half-even, truncation — a per-jurisdiction setting, not a style preference.
- The rounding site. Per line, per period, or per year-to-date. Rounding then summing is a different number from summing then rounding.
- The decimal library and its precision context, including whether anything on the path is a float at all.
- The timezone, which decides which pay period a shift that ends near midnight belongs to.
- The tzdata version, because zone rules are versioned data and updates are corrections — including corrections to history.
- The locale, which changes number PARSING and not merely formatting. This is the one that eats a thousand-fold error without raising.
- The ICU/CLDR version, if you format currency, spell amounts out in words for a cheque field, or sort names for a tie-break.
- The glibc version, which supplies `strtod`, `strftime` and collation to everything above it.
- Iteration and hash order, where a formula sums a dictionary and non-associative arithmetic turns ordering into cents.
- The host kernel, which a container borrows and cannot pin.
- The customer's own earning codes, deduction formulas and benefit accrual rules, if you let them write any — and you will.
A reproducibility strategy is a claim about that list. "We pinned the Docker image" is a claim about roughly four entries on it, and it is weaker than it sounds even about those.
"We pinned the image" pins less than you think
Start with the honest part, because the argument is usually made badly. A digest pin — `ubuntu@sha256:...` rather than `ubuntu:24.04` — genuinely does pin the bytes of that image, and therefore the glibc, ICU and tzdata inside it. If you digest-pin, you are doing better than most teams, and I am not going to pretend otherwise. Most teams pin a tag, and a tag is a mutable pointer: `ubuntu:24.04` in 2029 does not name the bytes it named in 2026, it names whatever the publisher has since decided 24.04 should be. Nobody will tell you it moved.
Now the parts a digest does not cover, in rough order of how often they have ruined somebody's week:
- Build-time fetches. Every `apt-get install`, every `pip install` without hashes, every `curl | sh` in the Dockerfile resolved against the internet on the day you built, and the digest faithfully pins the result exactly once. Rebuild from that Dockerfile in 2029 and you get a different machine with the same instructions, which is the specific thing you were trying to prevent.
- Runtime injection. Bind-mounting `/etc/localtime` from the host is a near-universal convention, and it reaches straight past your carefully pinned tzdata to use the host's. `TZ`, `LANG` and `LC_ALL` arrive from the orchestrator the same way. Your image pinned the data; the platform swapped the configuration.
- Retention. A digest is an immutable name, not an immutable promise that somebody is still serving those bytes. Registry garbage collection, a retention policy, a provider migration or a deleted repository all outlive a five-year audit window with no effort at all.
- The kernel. The container shares the host's, and the host's moves for security reasons on a cadence that has nothing to do with your filing calendar.
Let me be fair about that last one, because it is where this argument usually gets oversold. A newer kernel will almost never change a rounding result. Arithmetic is arithmetic. Where the kernel does matter is narrower and realer: `getrandom` behaviour, timing and clock plumbing, and the sandbox boundary your tenant-code story rests on. The point is not that the kernel will silently restate your tax answer. The point is that it is the one layer a container cannot pin, and a container is ultimately a polite suggestion to the kernel that it keep some processes apart — a suggestion a kernel upgrade is free to reinterpret.
The two entries on the list that will genuinely change your numbers, and that people do not expect to, are tzdata and ICU. Both are updated for correctness, which is exactly what makes them dangerous: a correctness update changes answers by design, and zone histories get revised, so a 2029 tzdata can disagree with a 2026 tzdata about what a 2026 offset was. If you want the full catalogue of how locale, timezone and encoding defaults sabotage short-lived environments specifically, Locale, Timezone and Encoding Bugs That Only Appear in Ephemeral Sandboxes is the dedicated post. Here I only need one claim from it: those are versioned inputs to your calculation, so record the versions you calculated with.
Floating point has no business anywhere near money
This is the oldest advice in the industry and it is still the single most common defect in a payroll pipeline, usually smuggled in by a JSON parser that turns every number into a double before your code sees it. Binary64 cannot represent 0.01. It cannot represent 0.1. `0.1 + 0.2` is `0.30000000000000004`, and that is not a bug in Python, it is the format doing precisely what it promises.
The subtler problem is that a float does not have a rounding policy at all. It has a representation, and the policy becomes an accident of it. In Python, `round(2.675, 2)` returns `2.67`, because `2.675` is stored as `2.67499999999999982...`, so there is no half to round up and the rule a legislature wrote never gets a chance to fire. The same expression on `1234.565` returns `1234.57`, because that one happens to be stored slightly above the half. Your engine will agree with the statute on some lines and disagree on others, with no pattern a reviewer can follow, which is considerably worse than being consistently wrong.
So: a decimal or fixed-point type, constructed from strings and integers and never from a float; an explicit precision context; an explicit rounding mode carried in the rule set as data; and an explicit rounding site. If you never chose a mode, you chose banker's rounding by inheritance — IEEE 754's default is round-to-nearest-even, Python's built-in `round` is half-to-even, and `Decimal`'s default context rounding is `ROUND_HALF_EVEN`. That is a defensible default and it is also simply not what several tax authorities wrote down.
# money.py -- three decisions a tax authority made for you, which your
# code has to state out loud instead of inheriting.
from decimal import Decimal, getcontext, ROUND_HALF_UP, ROUND_HALF_EVEN
CENTS = Decimal("0.01")
# 1. PRECISION is a context, and the default context is not a decision you
# made. Set it from the rule set, not from the interpreter's mood.
getcontext().prec = 28
# 2. The ROUNDING MODE belongs to the jurisdiction, not to the engine.
# Carry it as data so a 2029 replay reads the same byte you read in 2026.
MODES = {"half-up": ROUND_HALF_UP, "half-even": ROUND_HALF_EVEN}
def money(x) -> Decimal:
# Construct from str/int, never from float. Decimal(1234.565) gives you
# 1234.56500000000005456968210637569427490234375 -- the float's error,
# now preserved to 40 digits with perfect fidelity.
if isinstance(x, float):
raise TypeError("a float reached the money path; fix the caller")
return Decimal(x)
# 3. The ROUNDING SITE is a third, separate decision, and it is the one
# people forget. Rounding each line then summing is not the same number
# as summing then rounding, and over a few thousand lines the gap is a
# real figure somebody will ask you to explain.
def total(lines, mode: str, site: str) -> Decimal:
r = MODES[mode]
if site == "line":
return sum((money(v).quantize(CENTS, rounding=r) for v in lines),
Decimal("0"))
if site == "total":
return sum((money(v) for v in lines), Decimal("0")) \
.quantize(CENTS, rounding=r)
raise ValueError(site)
LINES = ["0.005"] * 1000
print(total(LINES, "half-up", "line")) # 10.00 -- each 0.005 -> 0.01
print(total(LINES, "half-even", "line")) # 0.00 -- each 0.005 -> 0.00
print(total(LINES, "half-up", "total")) # 5.00 -- sum first, then round
# Same input, same library, same machine. Three legitimate answers, and the
# difference between the first two is the entire payroll.
One thousand lines of half a cent, one library, one machine: ten pounds, zero pounds, or five pounds, depending on two settings that are usually implicit. That is the whole argument for carrying them as data in the rule set and printing them on the receipt.
The determinism trap, in one script
Here is the environment half of the problem, made concrete. Same input, two environments that any sane reviewer would call identical, different answers. Run it in two containers built from the same Dockerfile six months apart and you will reproduce more of it than you would like.
#!/usr/bin/env bash
# determinism-trap.sh -- one input, two "identical" environments, two
# different answers. Nothing below is exotic. Every line is a default that
# somebody shipped deliberately, for a good reason, in a different context.
set -uo pipefail
AMOUNT='1234.565' # a gross line that lands exactly on a half-cent
TS=1769903940 # 2026-01-31T23:59:00Z -- the last minute of a period
echo '== 1. The decimal separator is a LOCALE setting, and parsers obey it =='
# Under a comma-decimal locale, "." is the THOUSANDS separator. Python's
# locale.atof strips it and hands you a number a thousand times too big.
# No exception, no warning. glibc strtod() in C is differently wrong: it
# stops at the "." and returns 1234.0, quietly losing 56 cents.
for loc in C.UTF-8 de_DE.UTF-8; do
printf '%-14s ' "$loc"
LC_ALL="$loc" python3 -c 'import locale, sys
locale.setlocale(locale.LC_ALL, "")
print(locale.atof(sys.argv[1]))' "$AMOUNT" 2>&1 | tail -1
done
# C.UTF-8 1234.565
# de_DE.UTF-8 1234565.0 <-- this is a payroll run, not a typo
echo '== 2. Which pay period a timestamp belongs to depends on $TZ =='
# Same instant. Three zones. Two different months, therefore two different
# periods, therefore two different rule sets and two different YTD totals.
for tz in UTC America/Los_Angeles Pacific/Auckland; do
printf '%-22s %s\n' "$tz" "$(TZ="$tz" date -d "@$TS" '+%Y-%m-%d %H:%M')"
done
# UTC 2026-01-31 23:59
# America/Los_Angeles 2026-01-31 15:59
# Pacific/Auckland 2026-02-01 12:59 <-- February. Different month.
echo '== 3. ...and the zone rules themselves are versioned data =='
# tzdata ships several releases a year. Updates are CORRECTIONS, including
# corrections to history, so a 2029 tzdata can disagree with a 2026 tzdata
# about what a 2026 offset was. Record the version you calculated with.
cat /usr/share/zoneinfo/+VERSION 2>/dev/null \
|| dpkg-query -W -f='${Version}\n' tzdata
dpkg-query -W -f='${Package} ${Version}\n' 'libicu*' libc6 2>/dev/null
echo '== 4. The thing you actually shipped: an unpinned rounding mode =='
python3 - <<'PY'
from decimal import Decimal, ROUND_HALF_UP, ROUND_HALF_EVEN
cents = Decimal('0.01')
for v in ('1234.565', '2.675'):
print(v,
'float:', round(float(v), 2),
'HALF_UP:', Decimal(v).quantize(cents, rounding=ROUND_HALF_UP),
'HALF_EVEN:', Decimal(v).quantize(cents, rounding=ROUND_HALF_EVEN))
# 1234.565 float: 1234.57 HALF_UP: 1234.57 HALF_EVEN: 1234.56
# 2.675 float: 2.67 HALF_UP: 2.68 HALF_EVEN: 2.68
#
# Read those two rows twice. The float agrees with half-up on one line and
# with half-even on the other, because what it actually did was round the
# nearest binary64 value -- 2.675 is stored as 2.67499999999999982..., so
# "round half up" never got a chance to fire. The float has no rounding
# policy. It has a representation, and the policy is an accident of it.
PY
The first case is the one I would lose sleep over, because it is a parse and not a format. A formatting bug produces an ugly string somebody notices. A parsing bug produces a plausible number that flows into a tax filing. `locale.atof("1234.565")` under a comma-decimal locale returns `1234565.0` — it reads the dot as a thousands separator and removes it — and the C answer is different and also wrong, because `strtod` stops at the dot and returns `1234.0`. Neither raises. Your validation checks that the figure is a positive number, and it is.
A snapshot is not an image. It is a frozen machine
An image is a filesystem plus instructions for becoming a process. A Firecracker snapshot is the process. It captures guest memory, device state and vCPU state at an instant, so the artefact contains the guest kernel as running, glibc as loaded and relocated, the tzdata and ICU tables the process already mapped, the jurisdiction rule set already parsed into whatever structure your engine parses it into, and the interpreter already warm. There is no becoming. There is only resuming.
That difference is the entire point for this workload. Everything on the hidden-arguments list that lives below your code — the kernel, the system libraries, the versioned data tables — stops being something you hope to reconstruct and becomes something you hold. And the restore does not consume it: memory comes back `MAP_PRIVATE` so the kernel copies on write, and the disk is cloned by XFS reflink or dm-snapshot, which means you can replay one archived snapshot a thousand times and the archive is bit-identical afterwards. The replay is non-destructive by construction, which is a property an auditor will care about more than any latency number.
The latency number is good anyway, and it is what makes replay a support action rather than a project. On our platform a create via snapshot-restore is 179 ms p50 and about 203 ms p99 — the `/snapshot/load` step inside that is 49 to 80 ms — against roughly 3 seconds for a full cold boot of a template with no baked snapshot. Restoring a 2026 machine in 2029 is therefore about as expensive as an HTTP request, which changes who is allowed to do it. A support engineer can re-run a disputed calculation while the customer is still on the call. Nobody has to find the 2026 branch.
The archival unit I would actually recommend is one snapshot per jurisdiction per tax year per engine release, named exactly that way, because that tuple is what a dispute is about. Long-term, those snapshots do not have to sit on expensive local disk: our memory streaming path pages guest memory on demand from object storage over HTTP Range GETs in 4 MiB chunks, with a bitmap of which chunks are all zeroes so the empty majority of guest RAM never travels, and a prefetch trace so the hot chunks arrive before the guest faults on them. Honestly: the first restore of a cold archive on a given host pays object-storage latency on its page faults, and later restores of that same snapshot on that host are local-disk fast, because the chunks are cached. Plan the first replay of a five-year-old period to be slower than your steady-state number.
For completeness on the substrate: guest kernel 5.10, Ubuntu 24.04 rootfs, Firecracker v1.16. The mechanics of the restore path are in The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms if you want the step-by-step.
What each reproducibility substrate actually pins
Four honest options, compared on what they pin rather than on what they are marketed as. I have tried to make the row for our own approach the one with the most uncomfortable final column.
| Substrate | Your code | Library versions | glibc / ICU / tzdata | The kernel | What breaks the 2029 replay |
|---|---|---|---|---|---|
| Image tag (ubuntu:24.04) | Yes, if the commit is pinned too | Whatever the tag resolved to that day | No — the base moves under you | No — borrowed from the host | The tag now names different bytes, and nothing tells you it moved. |
| Image digest (sha256 pin) | Yes | Yes as built; build-time apt and pip fetches are not | Yes inside the image — unless /etc/localtime is bind-mounted in | No — borrowed from the host | Registry retention. An immutable name is not a promise that somebody still serves those bytes. |
| Lockfile plus virtualenv | Yes | Yes, if hashes are pinned and the index still serves them | No — system libraries are out of scope by design | No | A yanked artefact, a rebuilt sdist, or a new manylinux glibc floor that will not install on the old base. |
| Nix flake with a pinned nixpkgs | Yes | Yes, down to the derivation hash | Yes — they are derivations too | No, unless you also build the kernel and run it in a VM | Binary cache eviction, and the fact that you are still executing on somebody else's kernel. |
| Firecracker snapshot (PandaStack) | Yes, as the loaded process | Yes, as loaded into memory | Yes, as already mapped and parsed | Yes — the guest kernel is inside the snapshot | Your object-storage bill, and the baked vCPU/RAM: a 2026 snapshot restores at 2026's size and cannot be enlarged. |
Nix is the serious competitor on this table and I would rather say so plainly than pretend otherwise. If your problem is "build the identical artefact from source years later", Nix solves it more cheaply and more verifiably than any snapshot, and it solves it for your developers' laptops too. Our claim is narrower and different: a snapshot pins the running machine including the kernel, requires no rewrite of how you build anything, and restores in 179 ms instead of realising a closure. If you can do both, do both — build the engine with Nix, snapshot the result, and archive the snapshot.
The two things a restored machine lies about
A frozen machine is frozen, which includes the parts you wanted to keep moving. Two of those matter enormously for this workload and I would rather you heard about them from me than from a bank.
The clock is stale, and nothing in the guest objects
On restore, `CLOCK_MONOTONIC` has not advanced across the gap — correctly, from the guest's point of view no time passed — and `CLOCK_REALTIME` reads the wall-clock moment the snapshot was taken, confidently, until something external steps it. `CLOCK_BOOTTIME` is frozen alongside monotonic, because the guest never participated in a suspend it could account for. There is no clock inside the guest that can tell you a gap happened. The full mechanism, including what it does to TLS, is in kvm-clock vs TSC: How a Firecracker Guest Tells the Time, and What a Snapshot Does to It.
Which yields the single most useful rule in this post, and it is worth adopting even if you never touch a microVM: an effective date is an input. Never `date.today()`, never `now()`, never a default argument evaluated at import time, never a database default. Pass it in, record it on the receipt, and make the engine refuse to calculate when it is absent rather than helpfully substituting today. A payroll engine that reads the clock is not reproducible on any substrate, because "today" is the one argument that is guaranteed to differ between the original run and the replay. The frozen clock is what makes this obvious; it is not what makes it true.
Every child from one snapshot shares its entropy
Restore the same snapshot a hundred times and you get a hundred guests with an identical kernel entropy pool and identical seeded userspace PRNGs. For a deterministic calculation this is not merely acceptable, it is the point — you want the same inputs to produce the same bytes. It becomes catastrophic the moment the same code path also mints an identifier: an idempotency key, a payment reference, a BACS or ACH batch id, a UUIDv4 on a payment instruction. Two tenants' runs, restored from one archived engine, can mint the same supposedly unique reference, and downstream systems will resolve that collision in whichever way is worst for you — silently deduplicating a real payment, or paying the same instruction twice. The Snapshot Clone Randomness Problem covers the mechanism and the reseeding options.
The structural fix is a separation worth having anyway: the calculation returns numbers, and the orchestrator assigns references. Nothing inside the deterministic boundary is allowed to generate an identifier, read a clock, or make a network call. If you enforce exactly those three prohibitions, most of this post becomes a non-issue and you have a function you can actually replay.
One API detail follows directly from this. `fork()` clones the disk and the child cold-boots, so it gets its own entropy, its own PIDs and its own clock — but you pay a boot and lose the warm engine. `fork_tree(n)` inherits the parent's running memory, which is exactly what you want for fanning out a warmed engine, and which means every child inherits the parent's RNG state too. Same-host forks run 400 to 750 ms, cross-host 1.2 to 3.5 s, and `fork_tree` is capped at 16 children.
Customer-authored rules are untrusted code wearing a spreadsheet costume
Any payroll product that survives contact with real customers ends up letting them define earning codes, deduction formulas and benefit accrual rules. It always starts as a small expression language, and it never stays one: the first customer asks for a lookup table, the second for a date function, the third for "can it call our HR API", and within two years you are operating a programming language whose users have never heard the phrase "programming language".
Evaluate that in-process and a tenant's formula runs in the same address space as every other tenant's salary data, which is close to the worst possible data to leak and the easiest to notice has leaked. A sandboxed interpreter helps, and I would not sneer at one — but it is one CPython escape, one `__class__.__bases__` walk, one unexpected `__reduce__` away from a cross-tenant disclosure that ends up in a regulatory filing rather than a bug tracker. The second failure is duller and more frequent: a formula with an accidental quadratic over forty thousand employee records. In-process, that stalls the worker and everybody's payroll stops. One microVM per tenant per run and it burns one tenant's 8 vCPU and 2 GiB, hits its in-guest `timeout`, and fails exactly one customer's run.
Be honest about the ladder, though. If you can keep the language genuinely bounded — CEL, Rego, a total expression evaluator with no loops and no I/O — evaluating it in-process is cheaper, faster and easier to reason about, and it is the right answer for a real share of cases. Customer-Authored Policy Code: Rego, CEL, and When a Rules Engine Needs a VM walks that ladder rung by rung. The VM earns its place when the language is Turing-complete, when customers want actual Python because their finance team already writes it, or when you need a resource boundary around a batch that has a statutory deadline. Payroll tends to hit all three.
Payroll is the most honest batch workload in software
The run shape here is almost comically well suited to create-on-demand, and almost comically badly suited to a reserved pool. Payroll is thousands of short, independent, CPU-bound calculations that all want to happen on the 15th and the last working day of the month, with a second spike at year end and a third at every statutory filing deadline, and approximately nothing in between. There is no steady state to size for.
Size a worker pool for the 15th and it idles for most of the month while you pay for it. Size it for the average and you miss a deadline that is a law rather than an SLO. Create at 179 ms p50 and 203 ms p99 and scale to zero in between, and the create cost disappears into the noise of a per-tenant run that takes seconds to minutes anyway — a couple of hundred milliseconds of VM, once, per tenant, per period.
The fan-out pattern that works: one VM per tenant per run for the isolation boundary, then inside a tenant, shard the employee list and use `fork_tree()` to spread a warmed engine — rule set parsed, interpreter hot — across the shards, remembering the cap of 16 and the shared-RNG caveat from above. The per-tenant network boundary is not something you configure per run either: each sandbox gets its own network namespace with a veth pair and a tap device, drawn from 16,384 pre-allocated /30 subnets per host. Pre-allocating that plumbing is the whole trick: building a namespace, a veth pair and the matching firewall rules from cold is one of the most expensive things in a create path, and doing it in advance is what lets a per-tenant network boundary cost nothing per run.
The deliverable is a column, not an architecture
Everything above is infrastructure talk, and infrastructure talk is not what you sell to a finance director. What you sell is a column on the payslip record: the id of the engine snapshot that produced it, sitting next to the rule set version, the effective date, the rounding mode and site, and the hash of the canonical output. That column is the product. It turns "re-run this exact calculation" from a two-week archaeology project involving a 2026 branch and a Dockerfile that no longer builds into a restore, a run and a diff.
It also makes the explanation half trustworthy, which is the part the amended return actually asked for. Engines can emit a per-line trace of which rule fired at which effective date, and plenty do. But a trace you cannot regenerate is a claim, not evidence — it is a log file asserting something about a machine that no longer exists. Regenerate it on demand from the same frozen machine and it becomes evidence. If you are assembling this for a formal audit, SOC 2 for a Platform That Runs Other People's Code and Audit Trails When the Machine Is Gone in 200 Milliseconds cover what the evidence package wants to look like.
# archive_run.py -- the whole pattern. The engine is a pure function; this
# makes the ENVIRONMENT an argument too, by freezing it into an artefact.
import hashlib, json
from pandastack import Sandbox
TENANT, JURISDICTION, TAX_YEAR = "acme", "us-ca", 2026
EFFECTIVE_DATE = "2026-03-15" # an INPUT. Never date.today().
def digest(s: str) -> str:
return hashlib.sha256(s.encode()).hexdigest()
# 1. An engine VM. The template's baked meta.json governs the size --
# code-interpreter is 2 GiB / 8 vCPU -- and Firecracker cannot change
# vCPU or RAM at snapshot restore, so a cpu=/memory_mb= here would be
# overridden to the baked values. Size the ARCHIVAL template on day one.
# ttl_seconds is an IDLE timeout, not a lifetime: a busy run is not reaped.
sbx = Sandbox.create(
template="code-interpreter",
ttl_seconds=900,
metadata={"tenant": TENANT, "jurisdiction": JURISDICTION,
"tax_year": str(TAX_YEAR), "role": "payroll-engine"},
)
# 2. Freeze the locale-shaped inputs explicitly. C.UTF-8 and UTC are not
# "defaults" here; they are a decision you are writing down.
sbx.filesystem.write("/srv/engine.env",
"TZ=UTC\nLC_ALL=C.UTF-8\nPYTHONHASHSEED=0\n")
# 3. The jurisdiction rule set and the tenant's own deduction formula are
# DATA, written in before the snapshot. Nothing is fetched during a
# calculation: a network call at run time is an unpinnable input with a
# different answer every year, and an availability dependency in 2029.
sbx.filesystem.write(f"/srv/rules/{JURISDICTION}-{TAX_YEAR}.json", RULESET)
sbx.filesystem.write(f"/srv/tenant/{TENANT}/deductions.py", TENANT_FORMULA)
def run(effective_date: str, period: str) -> str:
# exec() one-shot does not reliably enforce a client-side timeout, so
# bound it in-guest with timeout(1). A runaway tenant formula then dies
# inside that tenant's VM instead of wedging a shared worker.
cmd = (". /srv/engine.env && export TZ LC_ALL PYTHONHASHSEED && "
"timeout --kill-after=10s 600 python3 /srv/engine/run.py "
f"--ruleset /srv/rules/{JURISDICTION}-{TAX_YEAR}.json "
f"--tenant {TENANT} --effective-date {effective_date} "
f"--period {period} --rounding half-up --round-at line "
"--decimal-prec 28 --out -")
r = sbx.exec(cmd)
if r.exit_code != 0:
raise RuntimeError(r.stderr)
return r.stdout
payslips = run(EFFECTIVE_DATE, "2026-03")
# 4. Snapshot AFTER the rule set is parsed and the interpreter is warm. The
# artefact now contains the guest kernel, glibc as loaded, the tzdata and
# ICU the process already mapped, the parsed rule set and the hot engine
# -- one restorable object, not a recipe for one.
snap = sbx.snapshot()
receipt = {
"tenant": TENANT, "jurisdiction": JURISDICTION, "tax_year": TAX_YEAR,
"effective_date": EFFECTIVE_DATE, "period": "2026-03",
"engine_snapshot": snap, # <-- the column that matters
"rounding": {"mode": "half-up", "site": "line", "prec": 28},
"output_sha256": digest(payslips),
}
ledger.put(receipt) # next to the payslip rows themselves
sbx.kill() # teardown is kill(), not delete()
# ---------------- 2029: an amended return arrives ----------------
# Replay, do not rebuild. No 2026 branch, no Dockerfile archaeology, no
# argument about whether 2026 or 2029 is the wrong one.
rec = ledger.get(payslip_id)
replay = Sandbox.create(from_snapshot=rec["engine_snapshot"])
# The effective date comes from the receipt, NOT from the guest's clock --
# a restored guest's CLOCK_REALTIME still reads the moment of the snapshot.
again = run(rec["effective_date"], rec["period"])
assert digest(again) == rec["output_sha256"], "engine is not deterministic"
replay.kill()
Where I would not reach for this
The most important limit first: this makes your tax logic reproducible, not correct. A reproducibly wrong number is still wrong. It is merely wrong on the record, forever, with a snapshot id attached — which is, to be fair, a considerably better position to be in during a dispute than wrong and unexplainable, but it is not a substitute for a rules team.
Then the platform limits. We have no GPUs and no GPU passthrough, which for once costs you precisely nothing: no payroll calculation in history has wanted a tensor core, and I mention it only so the list is complete. The guest kernel is 5.10, so if your engine depends on something newer — recent io_uring features, a 6.x-only syscall, newer BPF or cgroup behaviour — check that before you build on us. And Firecracker cannot change vCPU or RAM at snapshot restore, which has a sharp archival edge: a snapshot taken in 2026 restores at 2026's baked size, and if your 2029 replay needs more memory than the template had, you cannot give it more. You would have to re-bake, which is the rebuild you were trying to avoid. Size the archival template generously on the first day and never shrink it.
Storage is a real line item and I would price it before committing rather than after. Multiply your jurisdictions by your retention requirement in years by your engine release cadence, and look at the number honestly. Copies on one host share bytes — memory is copy-on-write and disks are reflinked — and the zero-chunk bitmap means the empty majority of guest RAM never travels to object storage, but an archive spread across years and regions does not get that for free. Decide which axis you deduplicate on. One snapshot per jurisdiction per tax year per engine release is the smallest honest unit; one snapshot per customer per run is a storage bill with a project plan attached.
And the case where I would tell you not to bother: if you operate one jurisdiction, ship one rule set, allow no customer-authored logic and run only your own code, then a digest-pinned image, a Nix-built engine, `LC_ALL=C.UTF-8`, `TZ=UTC`, a decimal type and a hard prohibition on reading the clock gets you most of the benefit for a fraction of the operational surface. The microVM argument gets strong in exactly the conditions that make payroll products hard: many jurisdictions, long retention, customer-authored rules, and a burst shape with a legal deadline attached.
The thing I would start with tomorrow, regardless of substrate, is the smallest piece: make the effective date an argument, make the engine refuse to run without it, and write the rounding mode and site into the output alongside the number. Then put a snapshot id next to it, and the 2029 email stops being frightening.
Frequently asked questions
Isn't a digest-pinned container image good enough for payroll reproducibility?
For a lot of teams it is a big improvement and worth doing today, so start there. But it covers fewer of the real inputs than it appears to. A digest pins the bytes of the image, and therefore the glibc, ICU and tzdata inside it — genuinely useful. What it does not pin is anything the Dockerfile fetched at build time, so rebuilding from the same Dockerfile years later produces a different machine from identical instructions; anything the runtime injects, and bind-mounting the host's /etc/localtime over the container's is so common that your pinned tzdata is frequently bypassed in production without anybody noticing; the TZ, LANG and LC_ALL variables the orchestrator supplies; and the kernel, which is borrowed from the host and upgraded on a security cadence unrelated to your filing calendar. There is also a durability question nobody likes: a digest is an immutable name, not a promise that somebody is still serving those bytes in five years. Registry garbage collection, a retention policy change or a provider migration will quietly outlive your audit window. The practical test is to try it. Pull your 2024 digest today, run your archived inputs through it, and compare against what you filed. If it still pulls, still runs and still matches, your pinning works. If any of those three fails, you have learned something cheaply.
How can a tzdata or ICU update change a number I calculated about the past?
Because both are versioned databases of facts, and new versions correct old facts as well as adding new ones. The tz database ships several releases a year, and a meaningful share of each release is historical correction: a researcher establishes that a particular region's offset in a particular year was not what the database said, and the maintainers fix it. Your engine asking "what was the UTC offset for this zone on this date in 2026" can therefore get one answer from 2026's tzdata and a different one from 2029's. That matters whenever a timestamp near midnight decides which pay period a shift falls into, which decides which rule set and which year-to-date totals apply. ICU is the same story with a different surface: it carries CLDR data for formatting and collation, and CLDR changes. If you format currency, spell amounts in words for a cheque field, or sort names to break a tie deterministically, an ICU upgrade can change your output string or your ordering. The defence has two halves. First, minimise exposure: do your arithmetic on dates and decimals rather than on formatted strings, pin TZ=UTC and LC_ALL=C.UTF-8 in the engine, and take the effective date as an explicit argument instead of deriving it from a timestamp. Second, record the versions you calculated with, so when a replay disagrees you can tell whether the engine changed or the world's record of the past did. A snapshot does that recording for you by containing the libraries themselves.
If every sandbox restored from one snapshot has identical entropy, isn't that a bug for a payroll run?
It is simultaneously the feature and the landmine, and which one you get depends entirely on where identifiers are minted. For the calculation itself, identical entropy is exactly right: you want the same inputs to produce the same bytes, every time, in 2026 and in 2029, and a calculation that depends on randomness is not a calculation you can defend. If your engine is deterministic, shared entropy is invisible to it. The danger arrives when the same code path also generates an identifier — an idempotency key, a payment reference, an ACH or BACS batch id, a UUIDv4 attached to a payment instruction. Two runs restored from the same archived engine can produce the same supposedly unique reference, and downstream systems resolve collisions in whichever way hurts most: deduplicating a real payment, or honouring a duplicate. The fix is a boundary rather than a patch. Nothing inside the deterministic calculation may generate an identifier, read a clock, or make a network call; the orchestrator outside it assigns references, stamps times and talks to banks. Enforce those three prohibitions and shared entropy stops being a hazard and becomes a property you rely on. If you genuinely need fresh randomness inside a restored guest, reseed on restore — Firecracker exposes a generation identifier for exactly this — and note that fork() children cold-boot and get their own entropy while fork_tree() children deliberately inherit the parent's running memory, and therefore its RNG state.
Can I let customers write payroll rules in Python without one VM per tenant?
You can run customer Python in-process, and people do, and it works until it does not. The honest ladder has three rungs and you should climb it from the bottom. Rung one is a genuinely bounded language — CEL, Rego, a total expression evaluator with no loops, no imports and no I/O — evaluated in your own process. It is the fastest and cheapest option, it is auditable, and for a real share of deduction and eligibility rules it is sufficient. Rung two is a sandboxed interpreter: restricted builtins, import hooks, an instruction budget. It buys you expressiveness and it is a defence in depth rather than a boundary, because CPython's object graph is reachable in more ways than any allowlist anticipates, and a single escape in a payroll product means one tenant reading another tenant's salaries — about the worst disclosure available to you. Rung three is a hardware boundary, which is where a microVM per tenant per run comes in. Two things push you up to rung three. The first is the language: once customers want real Python because their finance team already writes it, you are hosting a general-purpose runtime and an allowlist is not a boundary. The second is resource behaviour, which people underrate: a formula with an accidental quadratic over forty thousand employees stalls a shared worker and stops everybody's payroll, whereas in a per-tenant VM it burns that tenant's own 8 vCPU and 2 GiB, hits its in-guest timeout and fails exactly one run. If you have a statutory deadline, that blast radius is the argument, not the security story.
How do I estimate the snapshot storage cost of keeping engines for seven years?
Do the multiplication explicitly before committing, because the shape of the answer depends almost entirely on which axis you archive along, and the difference between sensible choices is orders of magnitude. The arithmetic is: number of jurisdictions, times number of tax years you must retain, times your engine release cadence within a year, times the size of a snapshot, which is dominated by the template's baked RAM. One snapshot per jurisdiction per tax year per engine release is the smallest unit I would call honest, because that tuple is what a dispute is actually about — a tenant disagrees about a figure under a specific rule set at a specific effective date. One snapshot per customer per run is the version people drift into, and it is a storage bill with a project plan attached. Two platform facts work in your favour and one against. In your favour: copies living on one host share bytes, since guest memory is restored copy-on-write and disks are cloned by reflink or dm-snapshot rather than copied; and our streaming path carries a bitmap of all-zero memory chunks, so the empty majority of a guest's RAM never travels to object storage at all. Against: an archive spread across years, regions and storage classes does not inherit that sharing automatically, and cold-storage retrieval has both a latency and a billing character you should check against your provider's current terms rather than assume. Also budget for the first replay of an old period being slower than your steady-state restore, because it pays object-storage latency on its page faults before the local chunk cache is warm.
Keep reading
- The snapshot clone randomness problem — Why identical entropy is a feature for a calculation and a hazard the moment the same code mints a payment reference.
- kvm-clock, TSC, and the clock a snapshot freezes — The full mechanism behind "never read the effective date from the guest's clock".
- Locale, timezone and encoding bugs in ephemeral sandboxes — The catalogue of defaults that silently change answers, including the parsing bugs that eat money.
- Customer-authored policy code: Rego, CEL, or a VM — The isolation ladder for tenant-written rules, rung by rung, with the cheap rungs taken seriously.
- SOC 2 audit evidence for code execution — What an evidence package wants to contain once you can regenerate the calculation on demand.
Related posts
- Running Per-Tenant Billing and Usage Metering in MicroVMs
Most isolation bugs leak data. This one mails a customer an invoice built partly from a competitor's usage. The only bug class where the incident review includes your CFO — so make cross-tenant reads structurally impossible, not merely unlikely.
- Running academic artifact evaluation on microVMs
An artifact-evaluation reviewer gets a tarball, an install.sh with sudo in it, and two weeks to decide whether a paper's numbers survive contact with a second machine. A snapshotted microVM turns that into a machine you can hand to the next reviewer.
- Bioinformatics Pipelines on microVMs: Reproducibility, PHI and the 400-Tool Dependency Graph
Two stages of the same pipeline want incompatible htslib builds, the postdoc's R script needs to write into the site library, and the whole thing is bit-for-bit reproducible until somebody upgrades the host kernel. Containers fixed the packaging. The VM boundary fixes the rest.
- Testing Against Ten Toolchains Without Ten Broken Runners
If you ship a library you own a matrix. On a shared runner the legs quietly contaminate each other. A microVM per leg makes the matrix mean what it says.
- Running Fuzzing Harnesses and Crash Reproduction in MicroVMs
A fuzzer is a professional vandal you hired on purpose. Its entire job is to push your parser into states the author never imagined — so don't run it on a kernel you share with anything you care about.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.