Blog — page 4 of 38
Slurm and HPC Batch Scheduling vs an On-Demand MicroVM Fleet
Slurm is a queue in front of a machine somebody already bought. Every feature it is famous for exists because the hardware is finite and someone has to arbitrate. An ephemeral microVM fleet inverts that premise, which changes what is easy and what becomes impossible.
GitOps for Ephemeral Compute: Where Reconciliation Stops
Argo CD's entire personality is refusing to let you delete things. That is exactly the right instinct for a database StatefulSet and exactly the wrong one for a sandbox that was always going to die in forty seconds.
The Carbon and Energy Case for Scale-to-Zero Compute
The sustainability argument for scale-to-zero is not that your code got efficient. It is that the machine can be handed to somebody else. Which only counts if somebody actually takes it.
Snapshot the Failure, Not the Log Line
The expensive part of an intermittent bug is not the analysis. It is getting back to the failing state. A log line is a guess someone made in advance about what would matter; a memory snapshot is the state itself.
Private npm and PyPI Mirrors in Front of a Sandbox Fleet
A container fleet warms up. An ephemeral sandbox fleet cannot, because the whole point is that guest number four thousand is byte-identical to guest number one. That property is worth having and it means you will download left-pad four thousand times unless you do something about it.
Mobile CI in MicroVMs: What Runs and What Cannot
Mobile CI is not one workload. About ninety per cent of it is a JVM build that a microVM suits perfectly, and the rest needs hardware you cannot get from a Firecracker guest. Knowing which stage is which saves you a fortnight.
Answering a Security Questionnaire When You Run Customer Code
The spreadsheet was written for a CRM. You run code a customer's model wrote ninety seconds ago. About a dozen rows carry the entire review, and the most important one is not on the sheet at all.
The Best Batch Compute Platforms for Bursty Jobs in 2026
Nine ways to run a few thousand independent jobs, grouped by what they actually are rather than ranked. Most of your decision is made by two facts: whether you need GPUs, and whether you are willing to pay for a cluster that is idle most of the week.
Multi-Tenant WordPress Hosting on Firecracker MicroVMs
WordPress is a plugin-execution engine wearing a CMS costume, which makes shared hosting a multi-tenant remote code execution service with good branding. A look at why the PHP hardening stack is not a boundary, and what changes when every site gets its own kernel.
Giving a MicroVM Access to a Customer's Private Network
The customer's database is in their VPC and your sandbox is not. The naive answer is to hand them your egress IPs and ask them to open a hole; the answer that survives a security review is a WireGuard peer per sandbox, minted after restore, revoked on teardown, and never, ever baked into a snapshot.
Running EDA and Chip-Design Workloads in MicroVMs
In most workloads the compute is worth more than the data. In chip design it is emphatically the other way round: a netlist or a foundry PDK leaking to a co-tenant is a company-ending event, and the NDA you signed has opinions about which kernel your job shares.
CircleCI Self-Hosted Runners on MicroVMs
The moment you move a CircleCI job onto your own machine runner, you quietly trade a fresh VM per job for a box that remembers every build that ever ran on it. That trade is the whole security story, and you do not have to make it.
Rowhammer and the Attacks Below Your Hypervisor
Every isolation guarantee you buy is enforced by software running on hardware that several tenants share. Rowhammer is the clearest example of what that sentence costs: a bit flip in a DRAM row you do not own, achieved by physics rather than by a bug. Here's the honest version — what it takes to land, what ECC and TRR really buy, and the two mitigations that actually change the answer.
Leader Election When Your Compute Can Be Cloned
A restarted process comes up empty and has to ask who the leader is. A restored one never asks — it resumes mid-sentence still believing it holds the lock. Leases, fencing tokens, and why a forked sandbox must revoke leadership before it does anything else.
How to Put a CDN in Front of a Scale-to-Zero App
An origin that is allowed to sleep changes what your cache headers are for. Done properly, the wake happens in a background fetch and lands on nobody's request; done carelessly, one Vary header hands every visitor a cold boot.
Firecracker vs VMware ESXi: the Device Model Decides
Both words mean "hypervisor," and that is the last thing they have in common. ESXi's value is an enormous device model and a datacenter control plane; Firecracker's value is having neither. The whole comparison lives in that one trade.
Vendor Lock-In in Code Execution Infrastructure
Interface lock-in is an adapter and a bad afternoon. Data lock-in is a project. Semantic lock-in — you built on a behaviour nobody else sells — has no exit at all, which is why nobody sells you a mitigation for it. Written by a vendor, so I owe you the same audit of my own product.
Skipping DHCP: the Firecracker boot line you actually ship
Everybody ships the same eleven-token Firecracker kernel command line and nobody measures it. Ours is four tokens plus a generated ip= that skips DHCP entirely — here is how to work out which tokens are buying you anything, and why the answer changes completely once you restore snapshots instead of booting.
virtiofs vs virtio-blk: How Files Actually Get Into a MicroVM
One gives the guest a disk it owns. The other gives it a window onto a directory the host owns. That single difference decides whether you can fork a machine in 400ms, who parses guest-controlled input, and what a multi-tenant escape looks like.
What "persistent" actually means in a sandbox
You wrote the file. You ran cat and saw it. Neither of those facts says the bytes are on a disk. Here is every layer a write passes through inside a microVM, which of them a crash erases, and why the honest answer for an ephemeral rootfs is not "fsync harder" but "get the artifact out".
Interrupts, IRQs, and Where microVM Tail Latency Comes From
Median latency tells you the machine works. p99 tells you how the machine is built. Here is every handoff a virtio interrupt makes inside a Firecracker guest, which of those handoffs are queues, and which ones you can actually do something about.
Why the disk under your sandbox fleet decides your boot time
People pick a sandbox platform on features and then get bitten by storage hardware. A reflink clone is metadata-only and nearly free; every copy-on-write byte afterwards is a real read-modify-write against a real device. Under 50 concurrent restores, that device is either local NVMe or it is your bottleneck.
Knowing Which Customer Costs You Money
If you run code on behalf of customers, your cloud bill arrives as one number and your customers arrive as a list. Splitting the first across the second is a real engineering problem, and nearly everyone gets it wrong the same way: by attributing on wall-clock vCPU, which overcharges the bursty and undercharges the idle.
Can You Run a Sandbox Fleet on Spot Instances?
The discount is real and the question people ask about it is the wrong one. Spot is not a pricing decision, it is a statement about whether your workload can be interrupted and rebuilt — which is a different answer for a 30-second code execution than for a customer's database.