Simulating a Fleet of Devices: Digital Twins in MicroVMs
There is a question that arrives at every connected-hardware company at roughly the same moment, usually two weeks before a launch and usually from someone in operations rather than engineering. It sounds like this: what happens when forty thousand of these reconnect at once, after a regional outage, with the clocks they woke up with?
You cannot answer it from a unit test, or from a staging environment with three devices on a bench. Not from a load-test script either, not honestly, because the question is not "can the backend absorb N requests per second" — it is "what does this population of stateful, independently-clocked, individually-enrolled clients do to my backend when they all wake up at once and discover the world moved on without them." So you build a fleet simulator.
I build PandaStack, an open-source Firecracker microVM platform, and I see what people build first. The first fleet simulator is always forty thousand goroutines, or forty thousand asyncio tasks, in one process on one reasonably large box. Good instinct: fast, cheap, produces a number. The problem is that it lies to you in a specific, predictable, structural way — about exactly the failures you built it to find.
The question you cannot answer from a unit test
Start with the honest shape of the failure. A region goes away for eleven minutes. Forty thousand devices lose their broker connection. Their reconnect logic — which you wrote, inherited, or bought with an SDK — has some backoff in it. On the way back, every one of them wants a TLS handshake, a fresh token, a time sync, a config fetch, and a replay of whatever telemetry it buffered while offline. Some have been offline eleven minutes. Some, thanks to a NAT timeout on a path you do not control, have been half-open for an hour and do not know it yet.
What actually breaks is almost never the thing you expected. It is the jitter window being multiplicative with the retry count, so the herd re-synchronises on the third attempt instead of spreading. It is the token endpoint, not the broker, falling over, because every device needs a new token and tokens are expensive to mint. It is the config service returning a 503 that the device SDK treats as "retry immediately" because a 503 is not a 429. It is per-device TLS handshake cost — not the stream cost, the handshake cost — against an endpoint sized for steady state. It is the buffered-telemetry flush arriving as forty thousand simultaneous large bodies, a completely different load shape from the trickle you benchmarked.
This is not firmware emulation, and it is not physics
I have written about two neighbouring problems, and it is worth saying where this one sits. Firmware emulation is about one device, in depth: you take an image a vendor shipped, unpack it with tooling that is itself the dangerous part, boot it under emulation, and watch what its init tries to do. The unit of work is a firmware image and the output is a finding. Robotics simulation is about one robot, in detail: a physics engine, determinism and seeds, whether the sim environment is pinned, with a trajectory or a policy as the output. Both are deep and narrow.
This post is about neither. The unit of work here is not an image and not a robot — it is a population. Twins can be shallow: often a few hundred kilobytes of MQTT client, a state machine, a certificate and a local database. What makes it hard is that there are forty thousand, each slightly different, all pointed at the real backend you are about to ship. Depth is cheap; breadth is the whole cost.
It is also not load testing, though it rhymes. A load-testing fleet exists to saturate something without polluting its own measurement; the workers are interchangeable and stateless, and the design questions are source-address diversity and whether the generator is the bottleneck. A twin fleet exists to reproduce behaviour, and its workers are emphatically not interchangeable — twin 31,402 has a certificate that expires on Tuesday, 400 days of local state, and a clock forty seconds fast. The load generator wants to be identical. The twin wants to be itself.
What one process quietly factors out
The failure mode here is not diffuse inaccuracy. Every shared singleton in that process is a dimension of your test collapsed to a constant, and the failures you need live in exactly those dimensions.
One source address, one conntrack table, one port space
Forty thousand clients from one IP is not forty thousand clients. Your broker's per-IP limit sees one abusive host. Your rate limiter, if it keys on address, protects the backend from the test instead of letting the test discover whether the backend needs protecting. Every middlebox on the path tracks one flow set for one source, so the per-device NAT-timeout interaction behind most mysterious field disconnects is unreachable. And you will exhaust ephemeral ports on the generator long before the backend notices, producing a beautiful graph of a limit you invented.
One clock, and the frozen clock you actually want
Every task in that process reads the same monotonic and wall clock. Real devices do not. A device powered off in a warehouse for three months comes back with a wall clock three months stale, so its TLS validation fails — or, worse and more interestingly, succeeds against a certificate that has since been revoked. A device with a drifting RTC and no reachable NTP signs a request with a timestamp outside your replay window. A fleet in three timezones crosses a day boundary in three waves. Skew is not an edge case; it is a scenario you want to run deliberately, and one process makes it structurally impossible. The best available is a fake clock in your own code, which tests your code and nothing else — not the TLS library's opinion, not the HTTP date header, not the timestamps the twin writes to its own disk.
One TLS session cache, one certificate, one entropy pool
This produces the most confidently wrong result of the four. A shared TLS client keeps a session cache: the first twin does a full handshake, the next 39,999 resume. Your rehearsal measures one handshake and 39,999 resumptions, against a backend that in production sees forty thousand full handshakes from devices that have never met each other. The handshake is where the asymmetric crypto, the OCSP opinion, the SNI routing and the chain validation live. You have tested none of it at scale and your graph is flat and green. And if your product does certificate-per-device — a serious connected product does, or will after the first security review — forty thousand client certificates in one process still share one TLS state machine, one session cache, one entropy source and one set of timers, which is where the interactions you care about live.
One crash domain
This one is the most embarrassing when it happens. Twin 8,221 hits a nil dereference at minute 34 of a 90-minute rehearsal. In production that is one device rebooting and nobody notices. In your simulator it is the process that was simulating its 39,999 siblings, and you have lost the run — or worse, you have not, because somebody wrapped the task body in a broad exception handler, so twin 8,221 is now silently absent from a population your dashboard still reports as forty thousand strong.
That is the dark heart of it. The shared-process fleet simulator does not fail loudly. It passes. It passed for a team I spoke to last year, a perfect green rehearsal, because all forty thousand of their devices shared one TLS session cache and one very optimistic clock. Production found the handshake cliff in about ninety seconds.
What a twin has to own to be a twin
Invert the list. Here is what a faithful twin needs to own privately, and in each case what the shared version silently removes from your test.
- Its own network stack and source address — so per-IP limits, conntrack, ARP, MTU and port exhaustion behave as they will in the field, not as a model of how you assume they behave.
- Its own clock, wall and monotonic — so skew, warehouse drift, day-boundary rollover and timestamp-window rejection are scenarios you run rather than bugs customers report.
- Its own TLS state and client certificate — so a reconnect storm is forty thousand full handshakes, not one handshake and a cache hit, and expiry and rotation are testable per device.
- Its own persistent storage that survives a simulated power cycle — so the twin can accumulate 400 days of state, run out of disk, corrupt its own database on an unclean shutdown, and behave oddly afterwards the way real devices do.
- Its own entropy — so two twins do not share an RNG stream, which matters the moment anything derives a nonce or a jitter interval from it.
- Its own crash domain — so one twin's segfault is one device rebooting, not the end of the rehearsal, and one twin's pathology can never explain away the other 39,999 results.
That list is almost exactly the definition of a machine, which is the uncomfortable conclusion: the list of things a twin must own privately is the list of things a kernel owns. So the twin wants to be a VM.
Network realism: real stacks, not modelled ones
This is where the substrate stops being an implementation detail. On PandaStack every sandbox gets its own pre-allocated Linux network namespace: 16,384 /30 subnets per agent host inside `10.200.0.0/16`, each with its own netns `ns-<id>`, a veth pair (`vh-<id>` in the root namespace, `vg-<id>` guest side) and a `tap0` the microVM attaches to. The agent writes legacy `iptables` rules in the nat table inside each namespace — PREROUTING DNAT per published port, plus two SNAT `--to-source` rules: one out `tap0` to the baked gateway, and one rewriting the shared baked guest IP to that slot's unique veth address. The single MASQUERADE lives in the root namespace for the pool CIDR out the WAN interface. No nftables, no clever overlay.
That second SNAT rule is the mechanism behind everything this post claims. Every twin is restored from the same snapshot and therefore wakes believing it has the same guest IP as its 39,999 siblings; the unique source address is manufactured on the way out, so that root-namespace conntrack can tell sandboxes apart. Faithful per-device addressing is not a property of the guest. It is a rule.
Pre-allocation is a latency decision, not a capacity one: a warm slot is handed out in single-digit milliseconds — about 5 ms — where building one cold is nearer 100 ms. `PANDASTACK_NATID_POOL_SIZE` (default 4) is the warm prebuild depth per template, not a concurrency cap; when the free list drains the agent builds a slot from scratch in roughly 500 ms rather than refusing the create, which matters when your creates arrive in a burst by definition.
The payoff is that the realism is not simulated. Each twin has a genuine source address, so your broker's per-IP limits see forty thousand peers. Each twin's flows get their own conntrack entries, so NAT idle timeout is a property of a real path, not a number in a config file you are pretending to respect. ARP, MTU discovery, port exhaustion, TCP retransmit on a lossy link: two kernels doing their actual jobs. The Firecracker metadata service is wired on this path too — a clean channel for handing each twin its identity at boot without baking secrets into the image.
One sharp edge in that realism is worth naming, because it is a faithful-simulation win and an operational hazard in the same sentence. `net.netfilter.nf_conntrack_max` is exposed per namespace, so it looks like each twin has its own connection budget. It does not. The conntrack hash table is effectively global: the namespace is part of the hash key, the hashsize is one host-wide number settable only from the initial namespace, and entries come from one shared slab. Nothing reconciles thousands of per-netns ceilings against one pool of host RAM. So the twins get genuinely real per-device conntrack behaviour — the whole point — while exhaustion stays a host-wide event one badly-behaved cohort can trigger for everybody, including the twins whose results you were about to believe. A per-namespace maximum is not a budget.
When the shared-process simulator is the right tool
A microVM per twin is sometimes absurd, and I would rather say so than sell a rehearsal you do not need. The hard ceiling on PandaStack is the /16 — 16,384 sandboxes per agent host — and you will never reach it, because host RAM and CPU bind long before the subnet space runs out. Guest RAM is fixed by the baked snapshot, since Firecracker cannot change vCPU or RAM at restore: `base` bakes at 4 GiB and 8 vCPU, `agent` and `code-interpreter` at 2 GiB, `postgres-16` at 1 GiB. The 8 vCPU are burst capacity shared fairly under contention through `cpu.weight`, so CPU oversubscription is fine; memory is the wall. Working-set admission helps — the scheduler admits against what a host can actually back rather than its committed total, charging `max(512, memory_mb × 0.25)` per non-database create — but if your twin is a 200 KB MQTT client, a multi-gigabyte guest is a ratio to be embarrassed by. Bake a lean template, or do not use VMs.
So: use the shared-process simulator for pure protocol load generation — no per-device state, no per-device identity, no crash-isolation requirement, no interest in anything below your application layer. For "can the broker take 200k messages a second" it is correct and a VM fleet is a waste of money.
Use per-VM twins when the per-device dimensions are the point: certificate-per-device enrolment and rotation, OTA rehearsals, firmware-in-the-loop where the twin runs something closer to the real image, chaos drills that break individual devices, anything involving per-device state across power cycles, and — the one that quietly justifies the budget — any rehearsal whose results have to stand up in a review, where one twin's bad behaviour must not be able to explain away the other 39,999.
Four substrates, scored on faithfulness
| Approach | Source IP + path | Clock + TLS state | Crash domain | Per-device disk | Runs real firmware? | Honest cost |
|---|---|---|---|---|---|---|
| 40,000 tasks in one process | One address, one conntrack set, one port space | One clock, one session cache, one RNG | Shared: one panic ends the run | Faked or shared | No | Near zero. Cheapest and least faithful |
| Container per twin | Own netns if you configure it; often shared | Host clock; own process state | Shared kernel: one OOM kill is everyone's problem | Own writable layer | User-space builds only | Low. Density is the whole point |
| MicroVM per twin | Own netns, tap, source IP and real NAT | Own clock, own TLS state and certificate | Own guest kernel: one twin dies alone | Own rootfs, survives a power cycle | Yes, for a Linux image | Guest RAM per twin. This is the real bill |
| Hardware lab | Real radios, real NAT, real everything | Real, including the RTC that drifts | Real, including the one that bricks | Real flash, with real wear | Yes, the actual shipped image | Capex and rack time. 40,000 is not happening |
The hardware lab wins every faithfulness column and loses the only one that decides whether the rehearsal happens. The microVM row is faithful enough to believe and cheap enough to run the week before launch, which is the only week anyone runs it.
Bake one twin, stamp a population
The mechanism that makes a population affordable is that you do not build forty thousand twins. You build one, take it all the way to the state you care about — provisioned, enrolled against the real backend, 400 days of accumulated local history, a database that has been vacuumed a hundred times — and then stamp copies of that state.
That state is the expensive part of a twin and the part a fresh boot cannot give you. Enrolment is a multi-round-trip dance with your backend; backfilling a year of telemetry takes minutes. Doing all that forty thousand times rehearses your provisioning flow, which is a different and also useful test, but not today's.
On PandaStack every create is a restore of a baked Firecracker snapshot — there is no warm pool of idle VMs — which is why creates land at p50 179 ms and p99 203 ms end to end. The first-ever spawn of a template with no snapshot yet is a cold boot at around 3 s, after which the agent bakes the snapshot and every later create takes the restore path. Snapshot load is an `mmap` with `MAP_PRIVATE`, so guest memory pages in lazily and copies on write: twins stamped from one ancestor start out sharing most of their pages and diverge only where they write.
import json
from pandastack import Sandbox
TOTAL = 2000
COHORTS = {"fw-3.2.1": 0.5, "fw-3.3.0": 0.5} # the OTA rehearsal split
# 1. ONE golden twin, taken all the way to the state you care about.
golden = Sandbox.create(
template="base",
persistent=True,
metadata={"role": "golden-twin", "model": "thermostat-v3"},
)
try:
with open("dist/twin-agent.tar", "rb") as fh:
golden.filesystem.write("/tmp/twin-agent.tar", fh.read())
golden.exec("mkdir -p /opt/twin && tar -C /opt/twin -xf /tmp/twin-agent.tar",
check=True)
# Enrol against the REAL backend. This is the expensive state.
golden.exec("/opt/twin/agent enroll --backend https://staging.example.com",
check=True, timeout_seconds=120)
# 400 days of local history, a vacuumed DB, a plausible disk.
golden.exec("/opt/twin/agent backfill --days 400",
check=True, timeout_seconds=900)
golden.exec("sync", check=True)
ancestor = golden.snapshot() # the population's common ancestor
finally:
golden.pause()
# 2. One cohort leader per firmware, restored from that snapshot, pinned.
population = []
for firmware, share in COHORTS.items():
leader = Sandbox.create(
template="base",
from_snapshot=ancestor,
metadata={"role": "twin", "cohort": firmware},
)
leader.exec(f"/opt/twin/agent pin-firmware {firmware}", check=True)
# 3. fork_tree is the MEMORY-INHERITING restore path: 400-750 ms same-host,
# 1.2-3.5 s cross-host. Every child wakes up mid-flight, already
# enrolled, already holding 400 days of state.
# (A bare fork() would be a disk clone plus a COLD BOOT -- the ~3 s
# shape -- and the child draws its own entropy. See the next section.)
# fork_tree takes at most 16 children per call -- the agent rejects more
# outright -- so a population is grown as an actual tree: every child
# you make can parent the next sixteen. That is where the name comes
# from, and it is why breadth costs log(N) rounds and not N.
want = int(TOTAL * share) - 1
siblings, frontier = [], [leader]
while len(siblings) < want and frontier:
kids = frontier.pop(0).fork_tree(min(16, want - len(siblings)))
siblings.extend(kids)
frontier.extend(kids)
population.append(leader)
population.extend(siblings)
print(json.dumps({"twins": len(population), "ancestor": ancestor}))
fork_tree or fork: identical herd, or independent devices
Those two calls look similar and are not, and for a twin population the difference is two scenarios you both want. `fork_tree(count)` is the memory-inheriting path: it restores children from the parent's snapshot, so each child resumes with the parent's memory state. Same-host it lands in 400–750 ms; cross-host it is 1.2–3.5 s, because the memory image has to come down from object storage first. At the instant of the fork the children are genuinely identical — same open connections, same in-memory timers, same RNG state. That is the herd you want for a thundering-herd drill: forty thousand devices that were all in the same state when the region vanished, waking together and racing each other back. One practical limit to design around: a single `fork_tree` call takes at most sixteen children, and the agent refuses a larger count outright. A population of thousands is therefore grown as a genuine tree — each child you make can parent the next sixteen — which costs a logarithmic number of rounds rather than a linear one.
A bare `fork()` is a different animal. It reflinks the rootfs — XFS reflink or dm-snapshot copy-on-write, so the disk clone is cheap and O(metadata) — and then cold-boots a fresh guest on top of it. It does not inherit the parent's memory, and the child draws its own entropy. Its latency shape is the cold-boot shape, around 3 s, not the 400–750 ms restore shape; putting those numbers side by side is how people end up expecting a sub-second cold boot, so plainly: 400–750 ms is `fork_tree` / restore, `fork()` is the ~3 s cold-boot path.
And `fork()`'s properties suit a different scenario: independently-booted devices that must not share a session, a token, a nonce or an RNG stream, each twin's first act being its own TLS handshake from cold, because that is what a device does when it is power-cycled rather than resumed. Build the whole population with `fork_tree` and you get a perfect herd — plus forty thousand twins that inherited one process's RNG state, which, if your twin derives its reconnect jitter from `random` seeded at boot, means you have accidentally built the synchronised herd you were trying to measure. At least that bug fails safe: the rehearsal looks worse than production.
Inside one twin
Everything above is scaffolding. The twin itself is small, and most of its interesting behaviour comes from one number: the keepalive interval, versus the idle timeout of a NAT on the path to your broker. That is the clearest example of something a shared-process simulator cannot show you.
#!/bin/sh
# Runs inside ONE twin, in this sandbox's own netns. The address below is
# the SAME for every twin -- they are all restored from one snapshot -- and
# an SNAT rule on the way out rewrites it to this slot's unique veth IP, so
# the broker sees 40,000 distinct peers, not one very busy one.
set -eu
ip -4 addr show dev eth0 | awk '/inet /{print "baked guest ip:", $2}'
# Per-device identity. Each twin enrolled separately, so each holds its own
# client certificate and its own TLS session state. Nothing is cached across
# twins, which is the entire point.
CA=/etc/ssl/certs/backend-ca.crt
CERT=/var/lib/twin/device.crt
KEY=/var/lib/twin/device.key
DEVICE_ID=$(cat /var/lib/twin/device-id)
# The whole experiment lives in this one variable. -k is the MQTT keepalive
# in seconds: the client promises to speak at least that often, and the
# broker gives up after 1.5x that silence. Stamp one cohort below the path's
# NAT idle timeout and one above it; only one side shows the half-open bug.
KEEPALIVE=240
mosquitto_sub \
-h broker.staging.example.com -p 8883 \
--cafile "$CA" --cert "$CERT" --key "$KEY" \
-i "$DEVICE_ID" \
-k "$KEEPALIVE" \
-q 1 \
-t "cmd/$DEVICE_ID/#" \
>> /var/log/twin-mqtt.log 2>&1 &
echo "$!" > /run/twin-mqtt.pid
# Link realism, IF the guest kernel has sch_netem. Check rather than assume:
# the 5.10 guest kernel is a lean config, and a missing module is much better
# discovered here than at minute 40 of a 40,000-twin run.
if modprobe sch_netem 2>/dev/null; then
tc qdisc add dev eth0 root netem delay 300ms 80ms loss 2%
echo "netem: 300ms +/- 80ms, 2% loss"
else
echo "no sch_netem in this guest kernel -- shape the host path instead" >&2
fi
Two things there are load-bearing. First, the NAT whose idle timeout you are fighting is per-sandbox — this twin's DNAT and SNAT rules in this twin's namespace — so the half-open-socket scenario appears on exactly one side of the keepalive-versus-timeout line, per twin. In one process you get one flow and the scenario is invisible. Second, I check for `sch_netem` rather than assume it: the gap between "works on my laptop's kernel" and "works in a Firecracker guest on 5.10" has eaten more of my time than any other category of surprise.
The OTA rehearsal, and what actually breaks
The reason to build a population rather than a load generator becomes obvious the first time you rehearse an over-the-air update. Half the fleet on the old firmware, half on the new. A staged rollout: 1%, then 10%, then 50%, soaking at each stage. Then — the stage everybody skips and later wishes they had not — a rollback, from the new image to the old, on a cohort that has already written state in the new format.
Here is the thing I most want you to take from this post: what breaks an OTA rollout is almost never the image. The image is the artefact you spent three weeks testing. What breaks is the population's retry behaviour against your artifact server.
Concretely: your download endpoint is a CDN origin that is perfectly happy at 1% and starts returning 503s under concurrency at 10%. Your device SDK's retry has no jitter, or jitter that is a fixed fraction of a backoff that is itself a power of two, so retries re-synchronise. Devices with a partial download retry from byte zero, because resumable downloads are a feature nobody implements until the second outage, so egress is three times image size times fleet. The devices that succeed reboot into the new image at roughly the same moment and reconnect — the thundering herd again, stacked on a rollout. And the rollback, when you reach for it, is a download of the old image by a fleet already hammering that endpoint.
Not one of those is a property of the firmware. All are properties of forty thousand independent clients with independent disks, retry timers and partial-download state. A population reproduces them; a dial labelled RPS does not. And a per-twin disk reproduces the ugliest one: the device that filled its own storage with partial downloads and now behaves oddly for reasons unrelated to the update.
Chaos drills, per-twin audit, unconditional teardown
Once twins have their own crash domains, deliberate cruelty becomes a design tool rather than an accident. Kill 5% of the population mid-handshake. Step one cohort's clock forty seconds fast and watch timestamp-window rejection. Fill one twin's disk. Revoke one twin's certificate while it holds an open connection. Pause a cohort, let the backend decide they are gone, then resume them. Each is one twin misbehaving, and because the twin is a VM it misbehaves alone — the property that makes the other 39,999 results admissible evidence instead of an argument.
Per-twin identity also makes results usable. Each sandbox carries its own metadata, logs and resource accounting: the agent writes `cpu.weight` and `cpu.max` into a per-VM cgroup and reads back `cpu.stat` `usage_usec`, `memory.current` and `io.stat` per VM. So when the rehearsal produces a spike you can ask which twins, not just how many requests. If a twin exposes a status endpoint, its preview URL is tokenless and stable for the sandbox's lifetime at `https://<port>-<sandbox-id>.<suffix>` — the UUID is the credential — a convenient way to poke at twin 31,402 while its siblings run.
And then you tear it down, unconditionally, in a `finally`, because a fleet simulator that leaks is a fleet simulator that bills. Teardown is `sbx.kill()`; there is no `delete()` method, which I mention because I have watched a competent engineer write `sbx.delete()` in a loop, catch the `AttributeError` with a bare `except`, and leave two thousand twins running over a long weekend. The twins were extremely faithful. They simulated a fleet that nobody turns off beautifully.
import json
from concurrent.futures import ThreadPoolExecutor
results, harvest_errors, leaked = {}, [], []
def harvest(sbx):
# Key on the twin's OWN device id, read from inside the guest -- the
# population is not interchangeable, so neither are its results.
dev = sbx.exec("cat /var/lib/twin/device-id", timeout_seconds=30).stdout.strip()
raw = sbx.exec("cat /var/log/twin-result.json", timeout_seconds=30)
if raw.exit_code != 0:
return dev, {"status": "no-result", "stderr": raw.stderr[-500:]}
return dev, json.loads(raw.stdout)
try:
with ThreadPoolExecutor(max_workers=64) as pool:
for fut in [pool.submit(harvest, s) for s in population]:
try:
dev, payload = fut.result()
results[dev] = payload
except Exception as exc: # one twin's bad day, not the run's
harvest_errors.append(repr(exc))
reconnected = sum(1 for r in results.values() if r.get("reconnected"))
print(json.dumps({
"twins": len(population),
"harvested": len(results),
"reconnected": reconnected,
"harvest_errors": len(harvest_errors),
}, indent=2))
finally:
# Unconditional. Teardown is kill(); there is no delete(). A twin you
# forget keeps accruing vCPU-hours and GiB-hours all weekend.
for sbx in population:
try:
sbx.kill()
except Exception as exc:
leaked.append(repr(exc))
if leaked:
raise RuntimeError(f"{len(leaked)} twins did not die: {leaked[:3]}")
The rehearsal is a budget line you can compute first
The last reason this works as a practice rather than a one-off heroic project is that you can price it before you run it, which means you can get it approved. PandaStack charges $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same rate for every class. CPU bills on CPU-seconds actually burned, which suits a twin fleet: a twin sleeps between keepalives, so its vCPU costs almost nothing while idle and real money during the ninety seconds of reconnect storm. Memory bills committed GiB-hours, which is the honest part — guest RAM is reserved for the twin's whole lifetime whether used or not, because Firecracker cannot change RAM at restore. Egress is metered separately, and in an OTA rehearsal it is no rounding error.
I will not invent a total, because the inputs that matter are yours. The shape is twins × baked GiB × hours × $0.0162, plus the CPU-seconds you really burn at $0.054 per vCPU-hour, plus metered egress. Four numbers, in a spreadsheet, before you ask anyone for anything — and no per-request pricing, which matters more here than in most workloads, since a fleet simulator's whole job is emitting an unreasonable number of requests.
The summary
Somebody will eventually ask what happens when the whole fleet comes back at once, and the only honest answer is to build a population and watch it. The naive one — forty thousand tasks in one process — factors out every dimension the real failure lives in, and it does not fail loudly when it is wrong. It passes.
A faithful twin owns its network stack, its clock, its TLS state and certificate, its disk, its entropy and its crash domain — the same list as "the things a kernel owns." So the twin wants to be a microVM: its own netns, tap and source IP out of 16,384 pre-allocated /30s per host, its own guest kernel, its own rootfs that survives a power cycle. Bake one twin to the state you care about and stamp the population from its snapshot — `fork_tree` for the resume-together herd at 400–750 ms same-host, a bare `fork()` for the power-cycle drill with its own entropy and the ~3 s cold-boot shape.
Be honest about the ceiling: 16,384 per host is the subnet limit, host RAM is the real one, guest size is fixed by the snapshot, and a multi-gigabyte guest for a 200 KB MQTT client is a ratio to be ashamed of. For pure protocol load generation, keep the shared-process generator. For certificate-per-device flows, OTA rehearsals, firmware-in-the-loop and chaos drills — anywhere one twin's bad behaviour must not explain away the others' results — the microVM earns its RAM. Then tear the population down in a `finally`, because the only thing worse than a rehearsal that lies to you is one that keeps billing after it has finished.
Frequently asked questions
How many digital twins can I actually run per host?
The hard ceiling on PandaStack is the subnet space: 16,384 pre-allocated /30 networks per agent host inside 10.200.0.0/16, so 16,384 sandboxes per host. You will not get near it. The binding constraint is host RAM, because Firecracker cannot change vCPU or RAM at snapshot restore, so a guest's memory is fixed by whatever the template was baked with — `base` bakes at 4 GiB, `agent` and `code-interpreter` at 2 GiB, `postgres-16` at 1 GiB. Divide your host's usable RAM by your baked guest size and that is your real per-host twin count. CPU is far more forgiving: the 8 vCPU the templates bake are burst capacity shared proportionally under contention through `cpu.weight`, and twins idle between keepalives. Working-set admission improves density by admitting against what a host can actually back rather than its committed total, but it does not change the arithmetic. The lever that matters is baking a lean template, since RAM is the only per-template size knob.
Should a twin population use fork() or fork_tree()?
Both, for different scenarios, and the distinction is easy to get backwards. `fork_tree(count)` is the memory-inheriting path: children are restored from the parent's snapshot and resume with the parent's memory state, at 400-750 ms same-host and 1.2-3.5 s cross-host, the difference being that the memory image has to be fetched first. That gives a genuinely identical herd — same in-memory timers, same open state — the right population for a resume-together drill. A bare `fork()` is different: it reflinks the rootfs with XFS reflink or dm-snapshot copy-on-write and then cold-boots a fresh guest, so it does not inherit the parent's memory and the child draws its own entropy. Its latency is the cold-boot shape, around 3 s, not the 400-750 ms restore shape. Use `fork()` when twins must be independently booted and must not share a session, nonce or RNG stream. Mixing both tells you which shape produced which spike.
Can I run the real firmware in a twin, or only a simulated application?
It depends on what the firmware is. If your device runs embedded Linux and your application is a normal user-space process, you can run something very close to the shipped image inside a microVM: the guest is Ubuntu 24.04 on a 5.10 kernel, so a Linux user-space payload is straightforward, and the twin gets the per-device identity, disk and network stack that make the rehearsal faithful. What you will not get is the real board — no vendor kernel, no device tree for your SoC, no peripheral drivers, no radio, no RTC that drifts on its own. Firecracker deliberately exposes a minimal virtio device model, which is what keeps it fast and small, and that is a direct trade against hardware fidelity. If your firmware is an RTOS image or tightly coupled to silicon, a twin fleet is the wrong tool for the firmware and the right tool for everything above it: protocol, enrolment, retry logic, population behaviour.
How do I give each twin its own clock skew?
Inject it deliberately inside the guest rather than relying on a side effect of snapshot restore. A restored guest does resume with the wall clock the snapshot was taken at, and a frozen clock is a real production bug — it breaks TLS certificate validation, which is why PandaStack resyncs the guest clock on restore, resume and wake after exactly that incident. So the skew you want is not free; the platform is correctly trying to take it away from you. Set it yourself: per cohort, step the clock inside the twin before the run and record the offset in the sandbox metadata so results stay attributable. Because each twin is a VM with its own kernel, the skew applies to everything that matters — the TLS library's notion of now, HTTP date headers, the timestamps the twin writes to its own database — not just your code's fake clock, the ceiling a shared-process simulator hits.
What does a 40,000-twin rehearsal cost?
You can compute it before you run it, from four of your own numbers, which is most of the argument for doing it this way. The rate card is $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same for every class. Memory bills on committed GiB-hours, because the guest's RAM is reserved for its whole lifetime — so twins x baked GiB x hours x $0.0162 is the floor, and it is the term you control, by baking a lean template. CPU bills on CPU-seconds actually burned, which is forgiving here: a twin idles between keepalives and only works during the reconnect storm. Egress is metered separately and is no rounding error in an OTA rehearsal, where you pull firmware images and possibly pull them more than once per twin — itself a cost the rehearsal exists to find. There is no per-request pricing, which matters when the workload's whole purpose is emitting an unreasonable number of requests.
Keep reading
- IoT and embedded firmware emulation in disposable microVMs — The depth-first sibling: one vendor image, unpacked and booted. This post is the breadth-first one.
- Robotics simulation and RL rollouts in isolated microVMs — Physics sim of one machine, where determinism and a pinned environment are the whole problem.
- Building a load-testing fleet on microVMs — The stateless cousin: interchangeable workers whose job is throughput, not per-device identity.
- Fork vs clone vs snapshot, explained — Why fork_tree inherits memory and a bare fork() cold-boots, in more detail than this post has room for.
Related posts
- Chaos Engineering Inside microVMs: Fault Injection Without the Blast Radius
Real chaos experiments need real kernel knobs. A container shares the host kernel, so you either fake the fault at the application layer or you inject it into everyone on the box. A microVM lets you be genuinely nasty to exactly one guest.
- Scheduling sandbox bursts: why 50 creates land on one host
Fifty creates arrive in 200ms. Every one of them reads the same cached capacity snapshot, computes the same winner, and lands on the same host. The scheduler did exactly what it was told and produced exactly the outcome it exists to prevent.
- The best load testing platforms in 2026
The hard part of load testing was never generating requests. It is generating them from enough machines, without saturating the generator, and reading a percentile that means what you think it means.
- Conntrack Table Limits: The Shared Resource Under Your MicroVM Fleet
Your guests have separate kernels, separate namespaces and separate subnets. They share one fixed-size hash table, and one port scanner can fill it for everybody.
- Rehearse Your Data Migration on a Real Copy, Not a Staging Guess
The migration took eight minutes in staging and six hours in production, and somewhere in hour two it took a lock the ORM never mentioned. Staging was never going to tell you. A disposable copy of the real data will.
49ms p50 cold start. Fork, snapshot, and scale to zero.