Optimisation Solver as a Service: Isolating Jobs That Run for Hours
Most compute platforms are built on a quiet assumption: that the runtime of a job is a function of the size of its input, and therefore roughly predictable. It is a good assumption. It holds for rendering a page, resizing an image, running a test suite, training for a fixed number of steps. It is what makes a request timeout a reasonable engineering decision rather than a coin toss.
It does not hold for optimisation. Two mixed-integer programs with the same number of variables, the same number of constraints and the same file size can differ in solve time by five orders of magnitude, because what actually determines the runtime is how badly the linear relaxation lies to the branch-and-bound search — and that is a property of the numbers, not the dimensions. A warehouse assignment model that solves in 180 milliseconds in January solves for six hours in March because somebody added one more depot and the bound went soft.
I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. If you are building a product where customers submit optimisation models and you solve them — a scheduling API, a routing service, a planning feature inside a larger SaaS — then this bimodality is not an edge case you handle later. It is the central fact your architecture has to be shaped around, and most platforms are shaped around its opposite.
The runtime is unbounded, and worse, it is bimodal
Branch-and-bound is a search with a pruning rule. Every node either gets solved, gets cut off because its bound is worse than the best known solution, or gets split into children. When the bound is tight, almost everything gets cut off immediately and the whole thing finishes before the HTTP connection has finished establishing. When the bound is loose, the tree grows, and it keeps growing, and there is no result and no error — just a process reporting a gap that shrinks more slowly than you would like.
So the distribution of solve times for a real workload is not a bell curve with a long tail. It is two populations. The vast majority of jobs finish in well under a second; a small fraction run until something intervenes. You cannot classify a submission into one of those populations by looking at it, which means you cannot route the fast ones to a cheap path and the slow ones to an expensive path on arrival. Every job has to be submitted to infrastructure that can survive the six-hour case, even though 99% of them will not need it.
Why this cannot live in a function
The structural argument against running solvers in a serverless function is not that functions are slow or badly isolated. It is that a function platform's execution model is a wall-clock ceiling, and a MIP has no opinion about your ceiling. Every function runtime worth naming has a maximum duration; the numbers differ by vendor and move over time, so check the current limits rather than trusting a blog post, including this one. The point is structural: there is a number, the solver does not know it exists, and when you hit it the result is a killed process with no incumbent solution saved anywhere.
That failure is worse than it first looks, because it is not the honest kind. The solver had a feasible solution in hand — probably a good one, possibly the optimal one, just not yet proven optimal — and the platform threw it away in order to enforce a deadline that the solver would have been happy to respect if anybody had told it. A solver with a time limit returns its best solution. A solver with a SIGKILL returns nothing.
A microVM has no such ceiling by construction: it is a virtual machine running a process, and the process runs until it finishes, is signalled, or is reaped. On our side there is deliberately no write timeout on the agent's HTTP server precisely so that a streaming exec can outlive one, and the control plane sets only a header-read timeout. But "no request timeout" is not the same as "nothing will ever stop your job", and the two things that will stop it are worth knowing before they do.
There is a deeper point hiding in the first trap. Any platform whose liveness signal is "did a request arrive recently" will mis-classify a long compute job as abandoned, because a long compute job is, by definition, silent. Solver workloads are the purest form of that silence: no logs worth streaming, no progress callbacks unless you wire them, nothing to say except a gap number every few minutes. If you build a job runner on top of us, make the job emit something on a schedule — not for the humans, for the reaper.
Memory is the real killer, and the ceiling is the feature
Runtime gets the attention; memory does the damage. Branch-and-bound stores the open nodes of its search tree, and a tree that is growing because the bound is loose is also a data structure that is growing without bound. Gurobi's own guidance on avoiding out-of-memory conditions is explicit about the shape of the problem and about its workaround: `NodefileStart` sets a threshold in gigabytes past which nodes are written to disk instead of held in memory, and `NodefileDir` chooses where. It is also explicit about the limits — the parameter does nothing for continuous models or for a MIP that solves at or near the root node, and if the solver exhausts memory before it has any search nodes to write, the setting is simply ignored.
The same guidance contains the sentence that connects memory to threads, which is the part people miss: each thread in a parallel MIP needs its own copy of the model and several other large data structures. Threads are a memory multiplier, not just a CPU knob. Doubling the thread count on a model that was already near the ceiling is a way of converting a slow solve into a dead one.
A solver that exhausts memory does not degrade politely. Either the solver detects it and errors out, which is the good case, or the allocator fails somewhere less careful, or — on a shared machine with swap — the host starts paging and every other tenant on that box discovers that someone else's planning model has become their incident. This is where the microVM's fixed RAM ceiling stops being a limitation and becomes the product. The guest has exactly the memory it was baked with, there is no balloon device taking it back, and our templates ship without swap. A runaway tree hits the guest kernel's OOM killer, that kills the solver process inside that one guest, and the other sandboxes on the host never find out. One tenant's job dies instead of one machine's worth of tenants' jobs.
Which makes RAM a catalogue decision, not a parameter
Here is the constraint you have to design around. Firecracker cannot change a guest's vCPU count or RAM at snapshot restore, so guest memory is chosen exactly once, at template build time, with `--memory-mb`. Our agent reads the template's metadata on every create and forces the request to match before anything persists — the comment in the source calls a per-request override "a lie that flows straight into the API response, the DB row, and (worst) the billing event." The Python SDK now raises a `DeprecationWarning` if you pass `cpu` or `memory_mb` at all, and `--cpu` on a template build is deprecated and ignored; every template gets 8 burstable vCPU.
So you cannot offer customers a memory slider. What you can do — and what you should do for this workload — is bake a small family of templates at the sizes your models actually need and route submissions to a tier. The build API accepts 128 to 65,536 MiB of guest RAM, so the templates are buildable; what bounds you in practice is the per-sandbox plan ceiling (Free 4 GiB, Pro 16 GiB, Team 64 GiB, Enterprise higher). Two or three tiers is usually enough, because solver memory demand is as bimodal as solver runtime.
#!/usr/bin/env bash
# Two RAM tiers, one Dockerfile. Guest memory is baked into the snapshot at
# BUILD time -- Firecracker cannot change it at restore -- so a RAM tier is a
# catalogue entry you route to, not an argument a caller passes.
set -euo pipefail
cat > Dockerfile <<'DOCKERFILE'
FROM pandastack/base:latest
# Open-source, redistributable solvers: these can be baked into an image you
# clone thousands of times without anyone having to read a contract first.
# HiGHS is MIT, CBC is EPL-2.0, OR-Tools (which bundles CP-SAT) is Apache-2.0.
# Distribution package names for the COIN-OR tools drift between releases --
# check them for the base you are actually on.
RUN apt-get update && apt-get install -y --no-install-recommends \
build-essential cmake coinor-cbc coinor-libcbc-dev \
&& rm -rf /var/lib/apt/lists/*
# Pin exact versions in the requirements file, not here. The solver BUILD is
# part of your reproducibility story -- see the determinism section: "same
# version" is a precondition of every determinism guarantee below.
COPY requirements.txt /srv/requirements.txt
RUN python3 -m pip install --no-cache-dir -r /srv/requirements.txt
# Docker ENV does NOT reach a detached process in a Firecracker guest: we boot
# this rootfs, we do not run the image's config. /etc/environment (PAM) is the
# one that a login shell -- which is what our exec uses -- actually reads.
RUN printf '%s\n' \
'OMP_NUM_THREADS=1' \
'OPENBLAS_NUM_THREADS=1' \
'SOLVER_NODEFILE_DIR=/var/tmp/nodefile' >> /etc/environment \
&& mkdir -p /var/tmp/nodefile
DOCKERFILE
# --memory-mb is the ONLY sizing knob and it is baked. The build API accepts
# 128..65536 MiB; what you can then RUN is capped by the plan (Free 4 GiB,
# Pro 16 GiB, Team 64 GiB per sandbox), so solver-32g needs Team.
# --size-mb is the ROOTFS, and it is where the branch-and-bound node file
# lands. The CLI defaults to 2048 and the build API defaults to 1024 if you
# omit it; both are useless for a node file. Hard cap is 16384 (16 GiB).
# For anything larger, attach a persistent volume and point the node file
# directory at it.
# --cpu is deprecated and ignored: every template gets 8 burstable vCPU.
for mb in 8192 32768; do
pandastack template build \
-n "solver-$((mb / 1024))g" \
-f Dockerfile \
--context . \
--memory-mb "$mb" \
--size-mb 16384
done
Eight burstable vCPU, and exactly one place to parallelise
Every template we bake gets 8 vCPU, and "burstable" is doing real work in that sentence. Each live Firecracker process is placed in its own cgroup with `cpu.weight` proportional to its vCPU entitlement — 100 per vCPU, so 800 for an 8-vCPU guest. Proportional shares only bind under contention, which means that on an idle host your solver genuinely bursts across the physical cores, and on a busy host it gets its weighted fraction. Separately, the free tier gets a hard `cpu.max` ceiling of two physical cores regardless of the eight vCPU the template bakes; paid tiers have no hard ceiling, only the share.
The consequence for a solver is uncomfortable and worth saying plainly: the number of threads you configure is not the number of cores you will get. A solver told to use 8 threads on a contended host runs 8 threads across rather less than 8 cores' worth of time, and it will not notice. Throughput degrades gracefully; what does not degrade gracefully is anything you tuned against a wall clock, because the wall clock is now measuring your neighbours as well as your search.
Given 8 vCPU, there are two sane configurations and a popular insane one. One solver using all 8 threads is right when the job is a single hard model and latency on that one model is the product. Eight concurrent single-threaded solves in one guest is right when the job is a batch of independent models and throughput is the product — and it is strictly better on memory, because single-threaded solves do not each hold their own copy of the model. Eight concurrent jobs each asking for 8 threads is the insane one: 64 threads of tree search on 8 vCPU, each carrying its own model copy, producing a workload that gets slower as you add parallelism and a memory footprint eight times what you budgeted. Pick one level and pin the other to 1.
Licensing, honestly
The reason solver-as-a-service is a smaller category than it should be is mostly licensing, and anybody writing about the architecture without mentioning it is wasting your time. The three best commercial MIP solvers — Gurobi, IBM ILOG CPLEX and FICO Xpress — are licensed commercially, with terms that bear directly on how many places the software may run and how concurrently. A platform whose whole premise is "restore this image a thousand times" is a platform that creates a thousand running instances, and whether that is permitted is a question about your agreement, not about your infrastructure.
The permissively-licensed stack is genuinely good now, and it is the fan-out-friendly path by construction: you can bake it into an image and clone the image without anyone's permission.
| Solver | Licence | Safe to bake into a cloned image? | Determinism as documented |
|---|---|---|---|
| Gurobi | Commercial | Only as your agreement allows — every restore is another running instance | Documents deterministic parallel MIP for the same model, parameters, version, machine and thread count; time-dependent parameters and the default concurrent Method=-1 are the stated exceptions |
| IBM ILOG CPLEX | Commercial | Same question, same answer: ask your agreement | Explicit switch: parallelmode automatic (the default) keeps results deterministic, opportunistic (-1) trades determinism for speed |
| FICO Xpress | Commercial | Same | Check the current manual rather than assuming either way |
| SCIP | Apache-2.0 from 8.0.3; ZIB Academic License up to 8.0.2 | Yes from 8.0.3 — confirm the version you vendor | Verify per release; one thread plus a fixed seed is the portable answer |
| HiGHS | MIT | Yes | Verify per release; one thread plus a fixed seed is the portable answer |
| CBC (COIN-OR) | EPL-2.0 | Yes | Verify per release; same advice |
| OR-Tools / CP-SAT | Apache-2.0 | Yes | One worker is deterministic; multi-worker determinism needs interleave_search, a fixed interleave batch size and clause sharing off, and is documented as usually slower than the non-deterministic default |
| Z3, CaDiCaL, MiniSat | MIT | Yes | Verify per release; single-threaded with a fixed seed is the reproducible configuration |
There is also a product decision buried in that table. If your customers bring their own commercial licence, the isolation story changes shape: you are now running their licensed binary and their credentials in a guest that must not leak either, which is a much better argument for a VM boundary than for a container one. If you ship an open-source solver, you own the whole image and the licence question evaporates — at the cost of solving harder models more slowly, which for a large class of real problems nobody will notice.
Reproducibility: what the solvers actually promise
There is a widely-repeated claim that a multi-threaded MIP is inherently non-deterministic because the search order depends on thread timing. It is a reasonable intuition and it is wrong for the major commercial solvers, which went to considerable trouble to make it wrong. Gurobi's support documentation states that it is deterministic for parallel optimisation with multiple threads: the same model, the same parameters, the same version, the same machine and the same thread count give the same solution values and the same algorithmic path. CPLEX exposes the trade as a parameter rather than a property — `parallelmode` is automatic by default, which the manual describes as applying as much parallelism as possible while still achieving deterministic results, with opportunistic mode available when you would rather have the speed.
What breaks determinism, then, is more interesting than "threads". Three things, in rough order of how often they bite:
- Changing the thread count. Determinism is promised for a fixed number of threads, not across them. A solve that ran with 8 threads and a solve that ran with 4 are two different searches that happen to answer the same question — so if reproducibility matters, the thread count is part of the artefact, pinned in the template, not inferred from the host.
- Setting a wall-clock time limit. Gurobi names time-dependent parameters as an exception to its own determinism guarantee, and this is unavoidable rather than a bug: a limit measured in seconds makes the result a function of how fast the machine was, and on a shared host with proportional CPU shares, how busy your neighbours were. A deterministic run and a bounded run are different things and you have to choose. If you need both, bound the search by node count or iteration count instead of by time.
- Concurrent and portfolio modes. Gurobi's default `Method=-1` can run a non-deterministic concurrent solve; CP-SAT's default multi-worker search is a portfolio, and its own documentation says deterministic multi-worker needs `interleave_search` enabled, a fixed interleave batch size, binary clause sharing turned off — and is usually considerably slower than the non-deterministic version. Single-worker CP-SAT is already deterministic, which is often the cheapest way to buy reproducibility.
The microVM angle here is narrower than the marketing version but it is real. Every determinism guarantee above is conditioned on "the same version" and "the same machine", and both of those are things a baked image pins harder than a lockfile does. A template is the solver binary that was actually built, linked against the BLAS it was actually linked against, on one day, with a fixed vCPU count and a fixed RAM size. If someone asks six months later why a plan came out the way it did, re-running it from the same image and the same parameters is a much stronger claim than re-running it from the same `requirements.txt`.
The snapshot angle: warm the setup, not the search
A solver job is not only the solve. It is an interpreter start, a large shared object being mapped and relocated, a licence handshake if you have one, a model being constructed from data, and then the search. For the 200-millisecond jobs — which is most of them — that preamble is the entire latency budget and the search is noise.
Snapshot-restore is the right tool for exactly that preamble. Warm one sandbox until the interpreter is up, the solver library is mapped and the model skeleton exists, take a full snapshot, and then restore it per job: p50 179 ms, p99 203 ms, of which the snapshot load step itself is around 49 ms. The first create of a template before its snapshot exists is a genuine cold boot at roughly 3 seconds, after which it is auto-baked and every create is on the fast path. You have not made the search faster. You have deleted the part of the job that was never the point.
import json
from pandastack import Sandbox
# ---------------------------------------------------------------- warm, once
warm = Sandbox.create(template="solver-32g", ttl_seconds=3600)
# Warm what a per-job restore should never repeat: imports, the solver's
# shared objects, and the model skeleton if the STRUCTURE is fixed and only
# the data changes. warm.py builds all of that and then blocks on a FIFO, so
# the snapshot catches a process that is ready rather than one mid-import.
#
# If your solver needs a licence handshake, think hard before warming it into
# an image: a snapshot you clone is N instances. Read your agreement.
assert warm.exec_stream("python3 /srv/warm.py", timeout_seconds=900) == 0
# snapshot() returns a snapshot ID *string* and captures memory AND disk.
# It is synchronous and slow on a 32 GiB guest. You do it once.
snap = warm.snapshot()
warm.kill()
# ------------------------------------------------------------------- per job
JOBS = [
{"model": "depot-eu.mps", "mip_gap": 1e-4},
{"model": "depot-us.mps", "mip_gap": 1e-4},
{"model": "depot-apac.mps", "mip_gap": 1e-3},
]
for i, job in enumerate(JOBS):
# Restore, not boot: p50 179 ms. The interpreter is up and the solver
# library is already mapped.
#
# ttl_seconds is an IDLE timeout and the idle clock is bumped when a
# REQUEST ARRIVES, not while one streams -- so a long solve under a
# single exec_stream looks idle. persistent=True takes the sandbox out
# of the idle reaper's scan entirely; belt and braces.
sbx = Sandbox.create(
from_snapshot=snap,
ttl_seconds=8 * 3600,
persistent=True,
metadata={"job": str(i)},
)
# ------------------------------------------------------------------
# RESEED. Every restored child shares the warm process's heap byte for
# byte. A random.Random() or a numpy RandomState your warm-up built
# resumes the SAME stream at the SAME position in all of them. Kernel
# randomness DOES diverge -- the guest CRNG re-mixes RDRAND per extract,
# so getrandom(2) differs per child -- but your library's state does not.
# For a randomised-restart portfolio or a stochastic local search that is
# not N searches; it is one search, run N times, agreeing with itself.
# ------------------------------------------------------------------
sbx.filesystem.write("/srv/params.json", json.dumps({
**job,
"seed": 1_000_003 + i, # derived from the index: replayable
"threads": 8, # ONE job per guest, so take all eight
"node_file_start_gb": 24, # page the tree out at 24 of 32 GiB
}))
print(i, sbx.id)
Two things in that snippet are the actual content and the rest is plumbing. The first is that a warm snapshot is only worth taking if the warm part is genuinely shared — if every customer's model has a different structure, there is nothing to pre-build and you are just paying storage for a slightly warmer interpreter. The second is the reseeding comment, which is the mistake that costs the most and announces itself the least.
Solvers use randomness more than people expect. A seed parameter perturbs tie-breaking and the search path; randomised restart portfolios in SAT and SMT are built on it; stochastic local search is nothing but it. If your warm-up constructs a generator before the freeze, every restored child inherits that generator's exact state, and a portfolio of 64 "independent" restarts becomes 64 copies of one restart. There is no error, no warning, and no symptom except an unusually consistent result — which is, of course, exactly what everyone was hoping for.
Six hours is long enough to lose the machine
A job that runs for six hours will occasionally be running when something happens: a host is drained, a kernel panics, somebody deploys. The honest answer is not a claim about reliability, it is a checkpoint. Solvers make this unusually easy, because they already have the concept: the incumbent solution is a complete, usable answer at every moment after the first feasible point, and most solvers can be handed one back as a warm start — a MIP start — on a subsequent run. Persist the incumbent and its objective value on a callback every few minutes and a lost host costs you the gap you had closed since the last write, not the whole solve.
Write it somewhere that outlives the guest. The rootfs does not: it is a copy-on-write clone that goes away with the sandbox. A persistent volume does — volumes attach as raw block devices that you mount yourself, which is also the right place for a large branch-and-bound node file once 16 GiB of rootfs stops being enough. Shipping the checkpoint out to object storage from inside the job works too and has the advantage of surviving the volume.
Our `hibernate()` and `wake()` pair is a different tool and it is worth not confusing them with checkpointing. `hibernate()` persists the sandbox's memory and disk to a hibernation snapshot and stops the VM; `wake()` brings it back, and the next request to a hibernated sandbox auto-wakes it transparently, so you rarely call `wake()` explicitly. The server-side idle sweeper only does this automatically for persistent sandboxes. That is a genuinely useful primitive for an interactive solver session a human comes back to tomorrow, and it does preserve the running process — but it is a tool for pausing, not for durability, and there is a sharp edge.
import json
from pandastack import Sandbox
# A single long solve, submitted the way a long solve should be submitted.
sbx = Sandbox.create(
template="solver-32g",
ttl_seconds=8 * 3600, # IDLE timeout; set it past the worst case
persistent=True, # and keep the reaper out of it entirely
metadata={"job": "plant-assignment-q3"},
)
# filesystem.write takes ONE file. upload() is write(remote, read_bytes()),
# so passing it a directory raises IsADirectoryError -- tar a tree first.
sbx.filesystem.write("/srv/model.mps", open("plant_assignment.mps").read())
sbx.filesystem.write("/srv/params.json", json.dumps({
"threads": 8,
"seed": 20261004,
"mip_gap": 1e-4,
"node_file_start_gb": 24,
"checkpoint_every_sec": 120, # a callback writes the incumbent
"checkpoint_path": "/mnt/jobs/incumbent.json", # a persistent volume
}))
logs: list[str] = []
# The real deadline belongs in the SHELL. Neither /exec nor /exec/stream
# kills the command server-side on timeout_seconds -- the agent decodes the
# field and runs the command on the request context. What exec_stream's
# timeout_seconds DOES do is raise the client's HTTP read timeout (to at
# least 900 s), which is why a long solve must never go through one-shot
# exec(): that dies at the client's 30 s default with the command still
# running, un-reaped, in the guest.
#
# --signal=INT, not the default TERM: most solvers install an interrupt
# handler that stops the search and returns the incumbent. TERM usually
# takes the answer with it.
rc = sbx.exec_stream(
"timeout --signal=INT 21600 "
"python3 /srv/solve.py --model /srv/model.mps --params /srv/params.json",
on_stdout=logs.append,
on_stderr=logs.append,
timeout_seconds=22000,
)
if rc == 124:
# Not a crash. A budget. coreutils `timeout` exits 124 on expiry.
print("budget reached; taking the best solution found so far")
elif rc != 0:
print(f"solver exited {rc}", "".join(logs[-20:]))
print(sbx.filesystem.read("/mnt/jobs/incumbent.json").decode())
# kill() cascade-deletes this sandbox's snapshots; the idle reaper does not.
# If you snapshotted something you wanted to keep, do not reflexively
# try/finally this call.
sbx.kill()
What this does not fix
- It does not make hard models easy. Isolation, a RAM ceiling and a 179 ms restore are infrastructure properties. If your bound is loose, the tree is large, and no amount of placement cleverness changes that. The fastest thing you can do to a six-hour solve is almost always to reformulate it.
- RAM cannot be asked for at submit time. A tier is a template, a template is a build, and a build is minutes plus a fleet-wide bake. Routing a job to a bigger tier is cheap; inventing a new tier mid-incident is not.
- There is no GPU, which rules out the GPU-accelerated first-order LP solvers that have become genuinely interesting for very large continuous problems. If that is your path, this is the wrong platform rather than a compromise.
- The eight vCPU are shared under contention. A benchmark you ran on an idle host is not the number your customer gets at 3pm, and a solver tuned against a wall clock is tuned against your neighbours.
- A warm snapshot of a 32 GiB guest is a large artefact, and the first restore on each host has to fetch it — which is the difference between a 400–750 ms same-host fork and a 1.2–3.5 s cross-host one. Warm only what is genuinely shared.
- Committed memory is billed memory. A 32 GiB guest sitting on a solve that is mostly waiting on one hard subtree is paying for 32 GiB the whole time. The win from snapshot-restore is on the preamble, not on the RAM.
- Licensing is still the gate, and it is not an engineering problem. No architecture in this post makes a commercial licence permit something it does not permit.
When this shape is right
Reach for one microVM per solve when the jobs are untrusted or cross-tenant, when the memory demand is unpredictable enough that one job can threaten the host, and when the runtime distribution is bimodal enough that a timeout is a lie. That is a precise description of a scheduling API, a routing service, a planning feature where customers upload their own models, and any product where a solver runs inside a boundary you would rather not share.
Do not reach for it when the models are yours, small, and fast. A single long-lived process with an in-memory queue will beat a VM-per-job on latency and operational surface for a workload whose 99th percentile is 300 milliseconds. The architecture in this post is insurance against the tail, and insurance you do not need is just overhead with good branding. Longer-running numerical tenant workloads are an adjacent story with different constraints — Isolating Quant Backtesting Workloads in MicroVMs covers the market-data-and-strategies version, and Slurm and HPC Batch Scheduling vs an On-Demand MicroVM Fleet covers what you give up when you leave a real batch scheduler behind.
A request timeout is a prediction about how long work takes. A solver is a machine for making that prediction wrong.
Design for the six-hour case, charge for the 200-millisecond one, and make sure the thing that kills a runaway job is a boundary you chose rather than a host you share.
Frequently asked questions
Why can't I just run my solver in a serverless function with a long timeout?
Because the mismatch is structural rather than a matter of degree. Every function platform enforces a maximum wall-clock duration — the specific numbers differ by vendor and move, so check the current limits — and a mixed-integer program has no way to know that number exists. When you hit the ceiling the platform kills the process, and the process was holding something valuable: the incumbent solution. A solver given a time limit stops its search and returns the best feasible solution it found, with a bound telling you how good it might be. A solver given a SIGKILL returns nothing at all, which is the strictly worse of the two failures. You can mitigate this by passing the platform's remaining budget to the solver as its own time limit, but then you have made the result non-deterministic — a time-based limit makes the answer a function of how fast the machine happened to be — and you still cannot exceed the ceiling for the genuinely hard models. A VM running a process until that process finishes has neither problem. The cost is that you now own the lifecycle: nothing stops a runaway job except the limits you set and the memory ceiling you baked.
Why is a fixed guest RAM ceiling a feature rather than a limitation for solver workloads?
Because branch-and-bound memory growth is the failure mode that takes down neighbours rather than jobs. The open nodes of the search tree are a data structure that grows when the bound is loose, and the standard mitigation — Gurobi's NodefileStart, which pages nodes to disk past a threshold in gigabytes — has documented holes: it does nothing for continuous models or for a MIP that solves at or near the root, and if memory runs out before there are search nodes to write, the setting is ignored. On a shared machine with swap, a model that outgrows RAM makes the host start paging and converts one tenant's bad formulation into everyone's latency incident. A Firecracker guest has exactly the memory it was baked with, no balloon device reclaiming it, and our templates ship without swap, so a runaway tree hits the guest kernel's OOM killer and kills that one solver process inside that one guest. Nobody else on the host notices. The cost of that property is real and worth stating: guest RAM is chosen at template build time with --memory-mb and cannot be changed at restore, because Firecracker cannot change guest memory when loading a snapshot. Our agent actively overrides any per-request cpu or memory_mb to the baked values, and the Python SDK now raises a DeprecationWarning if you pass them.
Is a multi-threaded MIP solve actually non-deterministic?
Less often than the folklore says, and the real caveats are more useful than the folklore. Gurobi's documentation states it is deterministic for parallel optimisation with multiple threads: the same model, parameters, version, machine and thread count reproduce the same solution values and the same algorithmic path. CPLEX turns the trade into a parameter — parallelmode defaults to automatic, described as applying as much parallelism as possible while still achieving deterministic results, with opportunistic mode available if you would rather have the speed. What actually breaks reproducibility is, first, changing the thread count, since determinism is promised for a fixed count and not across them; second, any wall-clock time limit, which Gurobi names as an exception to its own guarantee and which is unavoidable because a limit in seconds makes the answer depend on how fast and how busy the machine was; and third, concurrent or portfolio modes, including Gurobi's default Method=-1 and CP-SAT's default multi-worker search. CP-SAT is explicit that deterministic multi-worker needs interleave_search on, a fixed interleave batch size and binary clause sharing off, and that this is usually slower than the non-deterministic default — single-worker CP-SAT is already deterministic. Practical upshot: pin the thread count in the template, bound the search by node or iteration count rather than by seconds when you need replay, and fix the seed.
What is the most dangerous mistake when restoring many solver jobs from one warm snapshot?
Not reseeding a random number generator that the warm-up had already constructed. Be precise about where the duplication lives, because the intuitive version of this warning is wrong: kernel randomness diverges across restores perfectly well, since the guest CRNG re-mixes RDRAND per extract, so getrandom(2) and /dev/urandom give each restored child different bytes. What does not diverge is any generator your own code built before the snapshot — a seeded Random instance, numpy's global RandomState, a sampler object your warm-up instantiated. That state lives in the snapshotted heap, byte-identical in every child, resuming the same stream from the same position. Solvers lean on randomness more than people expect: a seed parameter perturbs tie-breaking and the search path, randomised restart portfolios in SAT and SMT are built on it, and stochastic local search is nothing else. Fan out 64 supposedly independent restarts from a warm image and you get 64 runs of the same restart, with no error and no warning — just suspiciously consistent results. The rule has two acceptable halves: either construct no generator before the freeze, or reseed from a per-child source as the first thing the job does. Derive the seed from the job index rather than fresh entropy so the run stays replayable.
How do I keep a six-hour solve from being reaped or lost?
Three separate things, and they fail in different ways. First, the idle reaper: ttl_seconds is an idle timeout with a five-minute default, and the idle clock is bumped by middleware when a request arrives at the agent rather than continuously while one streams — so a long solve under a single exec_stream call looks abandoned minutes in. Set ttl_seconds past the worst case you will tolerate, or set persistent=True, which takes the sandbox out of the reaper's scan entirely. Also note that the free tier applies a hard one-hour wall-clock lifetime to non-persistent sandboxes, independent of activity; paid tiers have no lifetime cap. Second, the client: one-shot exec() dies at the client's 30-second default with the command still running un-reaped in the guest, so use exec_stream, whose timeout_seconds raises the client's HTTP read timeout to at least 900 seconds. For an actual enforced deadline, wrap the command in coreutils timeout with --signal=INT, because most solvers handle SIGINT by stopping the search and returning the incumbent, where SIGTERM takes the answer with it. Third, durability: have a solver callback persist the incumbent solution and its objective every couple of minutes to a persistent volume or object storage, so a lost host costs you the gap you closed since the last write rather than the whole solve. Most solvers can take that solution back as a warm start on the retry.
Keep reading
- Isolating quant backtesting workloads — The closest neighbour: long-running numerical jobs per tenant, but with market data and egress control as the hard problems rather than unbounded runtime.
- Slurm vs an on-demand microVM fleet — What a real batch scheduler gives you that an API does not — queues, fairshare, and arbitration over finite hardware.
- Fanning out Monte Carlo trials from a warm microVM — The fan-out mechanics this post only gestures at: copy-on-write memory, one VM per trial, and the reseeding footgun in full.
- PandaStack sandboxes — Snapshot restore, persistence, volumes and the lifecycle calls this post leans on, in one place.
- PandaStack templates — The first-party catalogue, and how a custom template build like the solver tiers above works.
Related posts
- Per-Tenant Queue Consumers in Isolated microVMs
The shared worker fleet is fine right up until the first customer discovers the bulk import button. Then one tenant's 200k jobs become everyone's queue latency, and one segfault becomes everyone's lost work.
- A Spreadsheet Is a Programming Language Your Users Don't Call One
Your users write programs in your product every day. They spell them =SUMPRODUCT(...) instead of def main(), and that is the only difference. Then someone asks for PDF export and you shell out to an office suite.
- Your API Mock Is a Server Running Customer Code
Static fixtures become templates, templates grow helpers, helpers become a scripting engine — and now your mock server executes customer-authored code with a straight face. Here's the per-tenant microVM shape that makes that fine.
- Hyper-Threading and microVM Isolation: the SMT Decision
The most expensive security decision in multi-tenant compute is one nobody puts in a slide: do you leave SMT on? Two threads on one core share the L1, the TLBs, the branch predictors and the execution ports — below the level where a hypervisor can isolate anything. Here's the topology, the sysfs evidence, and the four rungs of the decision.
- Running Per-Tenant Billing and Usage Metering in MicroVMs
Most isolation bugs leak data. This one mails a customer an invoice built partly from a competitor's usage. The only bug class where the incident review includes your CFO — so make cross-tenant reads structurally impossible, not merely unlikely.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.