all posts

Running MATLAB and Octave Workloads in Isolated microVMs

Ajay Kumar··11 min read

Almost every runtime-isolation problem I deal with is a technical problem. You want user code to run somewhere it cannot reach your database, you pick a boundary, you argue about the boundary, you ship. MATLAB is the exception: the first constraint is not the kernel, the scheduler or the memory ceiling. It is a contract.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This post is about running heavyweight numerical workloads — MATLAB, the MATLAB Runtime, and GNU Octave — one tenant per guest, and it starts with licensing because licensing decides the architecture. If you skip that section and build the fan-out first, you will build the wrong fan-out.

Nothing in this post is legal or licensing advice, and I am not your account manager. MathWorks licence terms vary by product, licence type, region and agreement, and they change. Every statement below about what a licence permits is a description of the general landscape, not a reading of your contract — check your own agreement and MathWorks' current terms before you deploy anything, especially anything that multiplies process counts.

The licence is the architecture

The thing that makes MATLAB different from every other runtime in this blog's back catalogue is that you cannot casually multiply it. Spinning up 200 concurrent Python interpreters is a capacity question. Spinning up 200 concurrent interactive MATLAB sessions is a capacity question and an entitlement question, and the entitlement question has a wrong answer that comes with consequences well outside your monitoring stack.

The landscape, described carefully. Individual and named-user licences are tied to a person and typically to a small number of machines — they are not a fan-out mechanism and treating them as one is the mistake this section exists to prevent. Concurrent, or network, licences are counted centrally: a licence manager daemon hands out seats up to the number you bought, and the count is what you pay for, not the number of machines you own. MathWorks also publishes containerised and cloud-oriented options — reference architectures, container images, hosted offerings, and products specifically aimed at serving compiled MATLAB code at scale. Which of those your organisation has is the single fact that determines whether the rest of this post is a build guide or a research project.

So before any infrastructure work: find out what you actually hold, in writing, and ask MathWorks directly how it counts a short-lived VM. That last question matters more than it sounds. A licence model written for workstations and a workload consisting of eight hundred guests that each live for ninety seconds are not obviously compatible, and the person who can tell you how they interact does not work at your company.

What the licence manager needs from the network

A concurrent licence works by a client reaching a daemon over the network and checking out a seat. That single sentence has a lot of infrastructure hiding inside it, and a per-sandbox network namespace is exactly the environment where it hides worst.

  • Reachability, in the direction people forget. Each PandaStack sandbox gets its own network namespace, its own veth pair and its own /30 out of a pool of 16,384 per host, with NAT'd egress through the host. That is isolation, and isolation is the precise opposite of "can reach a daemon sitting inside your corporate network". We do not ship first-class private-network attachment today: the licence server has to be reachable from the host's egress path, via a tunnel terminated on the host, a proxy you expose on purpose, or a self-hosted deployment that already lives inside your network.
  • A stable hostname and host identity. Licence managers key on hostids — commonly a MAC address or a hostname. Because a restored guest's network identity is frozen at bake time, every sandbox from the same template presents the same MAC and the same guest-side IP, with the host NAT translating each to a distinct /30. Stable is good for anything node-locked and ambiguous for anything counted: a few hundred checkouts arriving from behind one NAT address with identical guest hostnames is a situation your vendor's daemon has an opinion about. Find out what that opinion is from them, not from me.
  • Egress has to be deliberate. The default posture of a sandbox platform is that outbound access is narrow and inbound access does not exist. Punching a hole to one daemon on one port is the correct shape of change here, and it is a change — not a default you can assume.
  • Clocks. A restored guest wakes with a wall clock frozen at bake time unless something resets it; we re-sync on restore, resume and wake, which is a thing we shipped after watching TLS handshakes inside guests fail with "certificate not yet valid" rather than "expired". Licence daemons are also famously unforgiving about time moving in unexpected directions.

There is a joke here and it is at the licence server's expense, not the licence's. You will have spent real effort on a platform where any single host can die without taking a job with it, where the control plane is multi-region, where the scheduler excludes a host whose heartbeat is thirty seconds stale — and then the whole numerical fleet will sit idle one Tuesday because one daemon on one VM in a datacentre nobody has logged into since 2019 stopped answering. The redundancy question for that daemon is a real engineering question and the answer is in your vendor's documentation, which is a sentence I would like you to read as advice and not as sarcasm.

The other licence-manager failure mode — seats orphaned by jobs that died abnormally, reclaimed only when the daemon's timeout expires — is identical to the one commercial EDA tooling has, and I have written it up properly there rather than re-deriving it here. If you are building a fan-out against a counted licence pool, read that section before you write the fan-out: the shell trap, the TTL on every sandbox, and the reconciler that compares checked-out seats against sandboxes that actually exist.

MATLAB Compiler and the MATLAB Runtime: the clean answer for fan-out

If your workload is "run this algorithm over N inputs" rather than "give N people an interactive MATLAB", there is a genuinely clean path and it has existed for a long time. MATLAB Compiler takes your MATLAB code and produces a standalone application; the application runs against the MATLAB Runtime, which MathWorks distributes free of charge and describes as royalty-free to deploy with compiled artefacts. The compile needs a licensed MATLAB with the Compiler product. The execution, on each of your N workers, does not need an interactive MATLAB licence.

That is the whole shape of the win: one licensed build step, arbitrarily many unlicensed executions. It converts a licensing problem into a build-artefact problem, which is a kind of problem infrastructure engineers are actually good at.

Verify this against MathWorks' current documentation before you plan around it. The details that matter in practice — which products and toolboxes are supported for deployment, what the Runtime's own licence permits, whether your use case wants MATLAB Compiler, Compiler SDK or their production-server product — are vendor specifics that move, and the version of them in your head is probably from the last time you read them.

Three operational facts about the Runtime that are easy to find out the hard way. First, Runtime versions are tied to the MATLAB release that compiled the artefact — you install the matching one, which means your template and your build pipeline are now version-coupled and you should pin both. Second, it is a large install; see the sizing note below, because the default rootfs size will not hold it. Third, not every toolbox is deployable. Some products are explicitly excluded from compiled deployment and some have restrictions, so if your algorithm leans on one, check it against MathWorks' current support list before the architecture depends on the answer.

And the Runtime still has a startup cost. A compiled application has to initialise the Runtime, link a lot of shared objects and set up its own environment before your first line of code runs. That cost is identical on every job and it is exactly the sort of thing a memory snapshot deletes, which is the subject of two sections from now.

Octave: the no-licence path, and where it actually breaks

GNU Octave is free software, largely compatible with the MATLAB language, and the correct starting point if what you need is to run user-authored numerical code without an entitlement conversation. For a large class of work — linear algebra, signal processing basics, scripts that are mostly matrices and loops — it is a drop-in. I would rather you knew the compatibility cliff from a post than from a user, so here it is honestly.

  • Toolboxes are the cliff, and they are the whole cliff. Octave's package collection is real and useful, and it is not a reimplementation of MathWorks' toolboxes. Function coverage differs, names differ, numerical behaviour at the edges differs, and "there is an Octave package for signal processing" does not mean "your Signal Processing Toolbox script runs". Test the actual script, not the category.
  • Simulink has no Octave equivalent. Not a partial one — none. If the work is model-based design, graphical block diagrams or anything that compiles out of Simulink, Octave is not on the list of options and no amount of packaging makes it one.
  • `classdef` object-oriented code is the next most common wall. Octave's support has improved over the years and is still not complete; a codebase built around modern MATLAB classes is a poor migration candidate. Old-style `@` class directories fare better, which is a sentence that will cheer up exactly one reader.
  • Data interchange needs testing. Octave reads and writes MATLAB `.mat` files, and the v7.3 format — which is HDF5 underneath — is the one where I would insist on testing your real files rather than trusting a compatibility table. Round-trip a representative file early.
  • MEX compatibility exists but is a build-time dependency on both sides, and a MEX file compiled against MATLAB is not an artefact you can expect to drop into Octave. Octave's own `oct` interface is the native route.
  • Plot output, printing and anything touching graphics are a persistent source of small differences. If the deliverable is a figure that has to look a specific way, that is the thing to validate first.

The useful way to hold this: Octave is an excellent answer when you control the code, or when the code is small, or when a human will fix whatever breaks. It is a risky answer when you are promising compatibility to a third party whose scripts you have never seen. Promising MATLAB compatibility and shipping Octave is a support-ticket generator with a very long tail.

Three ways to run the numerical workload

The three realistic options for executing MATLAB-language workloads per tenant, compared on the axes that actually trade off.
Interactive MATLAB per guestMATLAB Compiler + RuntimeGNU Octave
Per-worker licenceYes — a seat per concurrent sessionNo interactive seat at run time; Compiler licence at build timeNone
Fan-out ceilingYour seat count, and nothing elseCapacity onlyCapacity only
Needs the licence manager reachableYes, from every guestOnly from the build stepNo
Language and toolbox fidelityExact, by definitionExact for what you compiled; deployability varies by productGood for core language; toolboxes and Simulink are the gap
Interactive and exploratory useThe only real optionNo — you ship a fixed artefactYes
What you rebuild when the code changesNothingThe compiled artefact, then the templateNothing
Honest best fitSmall number of named users doing real analysisKnown algorithm, wide fan-out, scored inputsUser-authored code you do not want to license

Most organisations that do this well end up with two of the three. A small pool of interactive seats for the people developing the algorithm, and a compiled artefact — or Octave — for the fan-out that executes it ten thousand times. The architecture mistake is using one mechanism for both jobs.

Why the boundary wants to be a VM

Granted that you have picked a runtime, why per-guest isolation rather than a container per job? The honest answer is that this workload has four properties that each individually argue for a hard boundary, and having all four at once makes the argument fairly one-sided.

  • The code is frequently not yours. A student's coursework, a customer's model, an LLM's second attempt at a solver. Code you did not write and did not review, running on your hardware, is the whole premise — and a numerical runtime is an unusually rich environment to be careless in, because it ships with a compiler, a filesystem and a network stack by default.
  • Native extensions turn user bugs into process deaths. A MEX file — or any compiled extension loaded into the interpreter — is C or Fortran running in-process with no seatbelt. A segfault there takes down the session; a memory stomp there is worse than a crash because it does not crash. You want that contained by something that does not share a kernel with the other tenants.
  • Jobs are long and memory-hungry, which is the hardest shape for soft limits. A numerical job that allocates a large dense matrix is not a misbehaving job, it is a job doing exactly what it was asked, and the difference between that and a runaway is a judgement call you do not want to be making inside a shared kernel's reclaim path at 3am.
  • You need a kill that cannot be argued with. An infinite loop in user code, a wedged MEX file waiting on something that will never happen, a process that has stopped responding to signals — the remedy should be "the VM is gone", not a sequence of escalating requests to a subsystem that is already unhappy. Destroying a Firecracker guest terminates one process on the host and releases its namespace back to the pool.

There is a cost side to this and I will not pretend otherwise: a VM boundary means a fixed RAM allocation per guest, no memory sharing between tenants, and no overcommit cleverness within a single workload. For a web handler that would be an absurd trade. For an hour-long solve on someone else's code it is the trade you want.

The template: bake the runtime, size it deliberately

There is no first-party MATLAB or Octave template — our catalogue is `base`, `code-interpreter`, `agent`, `browser` and `postgres-16` — so this is a custom template build. You hand us a Dockerfile, we build it server-side and bake a Firecracker snapshot from the result. Two flags do more damage when wrong than any other part of this:

  • `--size-mb` is the ext4 rootfs size, and it defaults to 1024 MB in the Python CLI (2048 in the Go CLI). A MATLAB Runtime install will not fit in either, and the failure arrives mid-build rather than as a validation error. Size it for the install you are actually doing, plus headroom for runtime writes.
  • `--memory-mb` is the only place guest RAM is ever chosen. Firecracker cannot change a guest's RAM at snapshot restore, so the value is baked in and the agent silently corrects `memory_mb` on a create to match it. Passing `memory_mb=16384` to `Sandbox.create` against a 4 GiB template is not an error; it is ignored, which is worse. If you need a family of memory sizes, you need a family of templates.
  • `--cpu` is deprecated and ignored. Every template runs 8 burstable vCPUs, shared fairly under contention. This matters in a minute, for reasons involving BLAS.
#!/usr/bin/env bash
# An Octave template: the runtime, the packages, and the thread pinning, all
# paid ONCE at bake time for every sandbox ever created from the result.
set -euo pipefail

cat > Dockerfile <<'DOCKERFILE'
FROM pandastack/base:latest

# octave-cli, not the GUI build: nothing in a headless guest wants an X
# dependency tree. Add the dev headers only if you build .oct files in-guest.
RUN apt-get update && apt-get install -y --no-install-recommends \
      octave octave-dev liboctave-dev \
      libopenblas0 libhdf5-103-1t64 \
 && rm -rf /var/lib/apt/lists/*

# Packages, installed and LOADED at bake time so the load path is resolved.
# Note the honest naming: these are Octave packages, not MathWorks toolboxes.
RUN octave-cli --eval "pkg install -forge -global control signal statistics" \
 && octave-cli --eval "pkg load control signal statistics; disp('packages ok')"

# Docker ENV does NOT reach a detached process inside the Firecracker guest.
# /etc/environment (read by PAM) does. Every env var you need at runtime goes
# here, and this is the single most common source of "it worked in docker run".
RUN printf '%s\n' \
      'OPENBLAS_NUM_THREADS=1' \
      'OMP_NUM_THREADS=1' \
      'MKL_NUM_THREADS=1' \
      'OCTAVE_HISTFILE=/dev/null' >> /etc/environment

# Belt and braces: Octave's own site startup file, because a package can
# reset the thread count underneath you after the environment is read.
RUN mkdir -p /usr/share/octave/site/m/startup \
 && printf '%s\n' \
      'try' \
      '  setenv("OPENBLAS_NUM_THREADS", "1");' \
      '  setenv("OMP_NUM_THREADS", "1");' \
      'catch' \
      'end' > /usr/share/octave/site/m/startup/octaverc

# base ships no locales and no tzdata and sets no LANG or TZ. If your scripts
# parse dates or format numbers, decide that here rather than discovering it.
WORKDIR /work
DOCKERFILE

# --size-mb: DEFAULT IS 1024 in the Python CLI. Octave plus packages plus
#   HDF5 will not fit comfortably; an MCR install needs several times this.
# --memory-mb: baked into the snapshot, the ONLY place guest RAM is chosen.
# --cpu: deprecated and ignored. Every template gets 8 burstable vCPUs.
pandastack template build \
  --name octave-numerics \
  -f Dockerfile \
  --context . \
  --size-mb 8192 \
  --memory-mb 8192

The MATLAB Runtime variant is the same shape with a different middle and one extra constraint: you are installing a large vendor artefact non-interactively, which means reading their current installer documentation rather than copying a snippet from a blog post — including this one. The sketch below is deliberately a sketch.

# A MATLAB Runtime template, sketched. The install mechanics are vendor
# specifics that change between releases: get them from MathWorks' current
# docs, not from here. What IS stable is the shape of the problem.
#
# 1. The Runtime version must MATCH the MATLAB release that compiled your
#    application. Pin both, in the same place, or you will ship a template
#    that cannot run the artefact it was built for.
# 2. It is a large install. --size-mb defaults to 1024. Do the arithmetic
#    before the build, not after it fails nine minutes in.
# 3. A compiled application ships a generated run script that sets up the
#    library path for you. Use it. Hand-rolling LD_LIBRARY_PATH works right
#    up until a release moves a directory.
# 4. Docker ENV is not visible to a detached guest process -- anything the
#    artefact needs at runtime belongs in /etc/environment.

cat > Dockerfile <<'DOCKERFILE'
FROM pandastack/base:latest

ARG MCR_RELEASE=R2024b
RUN apt-get update && apt-get install -y --no-install-recommends \
      unzip libxt6 libxext6 && rm -rf /var/lib/apt/lists/*

# Fetch + silent-install the Runtime. The exact URL, archive layout and
# installer flags (including the licence-agreement flag you must pass
# deliberately) are documented by MathWorks per release. Read theirs.
COPY install-mcr.sh /tmp/install-mcr.sh
RUN sh /tmp/install-mcr.sh "${MCR_RELEASE}" && rm -f /tmp/install-mcr.sh

# Your compiled artefact: built once, on a licensed machine, by MATLAB
# Compiler. This image contains no interactive MATLAB and needs no seat.
COPY dist/solve /opt/solve/solve
COPY dist/run_solve.sh /opt/solve/run_solve.sh
RUN chmod +x /opt/solve/solve /opt/solve/run_solve.sh

# One computational thread per process. See the next section for why this is
# not a micro-optimisation in an 8-vCPU guest.
RUN printf '%s\n' \
      'OMP_NUM_THREADS=1' \
      'MCR_CACHE_ROOT=/tmp/mcr-cache' >> /etc/environment
DOCKERFILE

pandastack template build \
  --name matlab-runtime-solver \
  -f Dockerfile \
  --context . \
  --size-mb 24576 \
  --memory-mb 8192

Thread pinning, in MATLAB and Octave's own terms

This is the highest-value operational note in the post and it is the least exciting. The underlying mechanic — a threaded BLAS starting one thread per core, meeting job-level parallelism, on a guest with a fixed 8 vCPUs — I have covered at length for Julia and R, and the physics is identical here, so I will not re-derive it. What is specific to MATLAB and Octave is the set of knobs, and there are more of them than people expect.

# Four places the thread count gets decided for MATLAB-language workloads.
# Set them in the TEMPLATE, not per job, and then verify rather than assume.

# 1. The environment, which Octave's OpenBLAS and the MCR's threading both
#    read. Remember: /etc/environment, because Docker ENV does not reach a
#    detached process in the guest.
printf '%s\n' 'OPENBLAS_NUM_THREADS=1' 'OMP_NUM_THREADS=1' 'MKL_NUM_THREADS=1' \
  >> /etc/environment

# 2. MATLAB's own startup flag, for an interactive or scripted session. This
#    is the blunt instrument and it is usually the right one for batch work:
matlab -singleCompThread -batch "run('/work/job.m')"

# 3. MATLAB's runtime knob, when you want it set from inside the code. Note
#    that MathWorks has deprecated SETTING it via this function in some
#    releases -- check your release's documentation before relying on the
#    setter, and prefer -singleCompThread plus the environment.
#      n = maxNumCompThreads;        % query
#      maxNumCompThreads(1);         % set (release-dependent; verify)

# 4. Verify. Every layer above can be overridden by a package, a toolbox or a
#    startup file you forgot about, and the symptom of getting it wrong is not
#    an error -- it is a workload that gets SLOWER as you add workers.
octave-cli --eval "disp(getenv('OPENBLAS_NUM_THREADS'))"
matlab -batch "disp(maxNumCompThreads)"

# Then parallelise at EXACTLY ONE level. Eight workers each claiming eight
# BLAS threads on eight vCPUs is not eight times the throughput; it is a
# scheduler being asked to settle an argument with itself.

One extra wrinkle specific to MATLAB: if the code uses `parfor` or a parallel pool, you now have three levels of parallelism in play — BLAS threads, pool workers and your own job-level fan-out — and a parallel pool is itself a licensing question. Decide which single level does the parallelising, and turn the other two down to one. Which level you pick is a performance question with a real answer, and it is nearly always "the outermost one", because that is the level where the work is actually independent.

The snapshot payoff, and the Monte Carlo trap inside it

Here is where the per-guest model stops costing and starts paying. Once a template's snapshot is baked, creating a sandbox from it is a snapshot restore rather than a boot: p50 179 ms, p99 203 ms, of which the `/snapshot/load` step itself is around 49 ms. The very first create of a template before its snapshot exists is a real cold boot at roughly 3 seconds, after which the snapshot is captured and every create lands on the fast path.

Set that against what you were paying. Every job that starts a fresh MATLAB session or initialises the Runtime from cold pays interpreter start-up, shared-library linking, toolbox or package path resolution and — for a network licence — a licence checkout round-trip, before the first line of the user's code executes. That cost is identical on every job and it is pure overhead. Baking it into a template converts it into a file.

You can go further and snapshot a session that has already done your warm-up: a full `snapshot()` captures the guest's memory as well as its disk, so a restore brings back a live process with its loaded paths and constructed data intact. For heavyweight numerical runtimes that is a large win. It also comes with one trap, and for Monte Carlo work the trap is severe.

A warm snapshot captures a live process, including its random number generator state. Kernel randomness diverges across restores — the guest CRNG re-mixes RDRAND per extract, so `getrandom` and `/dev/urandom` give each child different bytes. What does NOT diverge is any generator your warm-up already constructed. An `rng(42)` executed before the snapshot is byte-identical in every restored child, continuing the same stream from the same position. Fan out a thousand trials and you get a thousand copies of one trial, with error bars that will survive review for longer than they should. Reseed per child, as the first thing the job does.

Say that in MATLAB-language terms, because the generic warning is too easy to nod along to. If the warm-up ran `rng(42)` — or `rand('seed', 42)`, or constructed a `RandStream` object, or called Octave's `rand('state', ...)` — that state is in the snapshotted heap. Every child inherits it exactly. The fix is one line at the top of the job, derived from the per-child parameters so the study stays replayable: `rng(cfg.seed)` in MATLAB, `rand('state', cfg.seed); randn('state', cfg.seed)` in Octave — and note that in Octave the generators for `rand` and `randn` are seeded separately, which is the kind of detail that produces a study where half the randomness is independent and half is not. If you built a named stream object before the freeze, reseed it by name; seeding the global generator does not reach it.

Fan-out: parameters in, results out

The mechanics are unglamorous, which is the point. One warm snapshot, N restores, per-job parameters written in as a file, results read back out as a file. No shared filesystem, no message broker, no job framework.

import json
from concurrent.futures import ThreadPoolExecutor
from pandastack import Sandbox

# --- 1. One parent does the expensive shared work, then freezes -------------
# ttl_seconds is an IDLE timeout (default 5 minutes), not a walltime budget:
# the reaper measures time since last activity.
parent = Sandbox.create(template="octave-numerics", ttl_seconds=3600)

# Warm-up has to CALL things, not just load them -- and must construct NO
# random generator state, because anything it builds is inherited identically
# by every child. Resolve paths, touch the hot code, exit.
WARMUP = r"""
set -eu
octave-cli --eval "
  pkg load control signal statistics;
  A = randn(256); B = A' * A;          % force the BLAS path to be resident
  [~] = chol(B + 256*eye(256));
  disp('warm');
"
"""
# Long commands go through exec_stream. NEITHER exec endpoint enforces
# timeout_seconds server-side -- the agent decodes the field and never
# applies it -- so here it only raises the CLIENT HTTP timeout, while a
# one-shot exec() gives up at 30s. The real deadline goes in the shell.
rc = parent.exec_stream(WARMUP, on_stdout=lambda c: print(c, end=""), timeout_seconds=900)
assert rc == 0, f"warm-up failed ({rc})"

# upload() writes a SINGLE file; a directory raises IsADirectoryError. Ship
# the model code as one tarball and unpack it in the guest.
parent.filesystem.write("/work/model.tar.gz", open("model.tar.gz", "rb").read())
parent.exec("tar -xzf /work/model.tar.gz -C /work", timeout_seconds=60, check=True)

# snapshot() returns a snapshot ID *string*, and captures memory AND disk.
snap_id = parent.snapshot()
parent.kill()

# --- 2. N restores, one per parameter set ----------------------------------
SWEEP = [
    {"sigma": 0.15, "horizon": 252},
    {"sigma": 0.22, "horizon": 252},
    {"sigma": 0.30, "horizon": 504},
]

def run_trial(i_params):
    i, params = i_params
    sbx = Sandbox.create(
        from_snapshot=snap_id,
        ttl_seconds=1800,
        metadata={"kind": "octave-sweep", "trial": str(i)},
    )
    try:
        # The per-child seed is the load-bearing line in this whole function.
        # Derived from the trial index, so the study replays; distinct per
        # child, so the trials are actually independent.
        sbx.filesystem.write(
            "/work/trial.json",
            json.dumps({**params, "trial": i, "seed": 20261004 + i}),
        )
        # trial.m reseeds FIRST -- in Octave rand and randn seed separately:
        #   cfg = jsondecode(fileread('/work/trial.json'));
        #   rand('state', cfg.seed); randn('state', cfg.seed);
        rc = sbx.exec_stream(
            "cd /work && octave-cli --eval \"trial\"",
            on_stdout=lambda c, i=i: print(f"[{i}] {c}", end=""),
            timeout_seconds=1800,
        )
        if rc != 0:
            return {"trial": i, "ok": False, "exit_code": rc}
        return json.loads(sbx.filesystem.read("/work/result.json"))
    finally:
        # Release the RAM before the next batch. Note: an explicit kill()
        # cascade-deletes this sandbox's snapshots, so do not kill a guest
        # whose snapshot you still need.
        sbx.kill()

with ThreadPoolExecutor(max_workers=8) as pool:
    results = list(pool.map(run_trial, enumerate(SWEEP)))

failed = [r for r in results if not r.get("ok", True)]
print(f"{len(results) - len(failed)}/{len(results)} trials completed")

Two notes on the fan-out primitives, because picking the wrong one is quiet rather than loud. `fork()` is disk-only: the child cold-boots from a clone of the parent's rootfs, so it keeps the installed runtime and loses the warm session entirely — a sweep built on `fork()` pays the start-up cost N times and still looks like it is working. `fork_tree(count)` does inherit memory and disk, but caps at 16 children pinned to the parent's host. For a wide sweep, take the snapshot yourself and create from it, as above: each create goes through the scheduler and spreads across hosts, with same-host forks at 400–750 ms against 1.2–3.5 s cross-host, the difference being the artefact pull.

What this does not fix

  • It does not get you a licence. Everything above is infrastructure, and infrastructure cannot create entitlements. If the answer from your vendor is that your licence does not support this shape of deployment, the answer is no, and no architecture diagram changes it.
  • There is no GPU. That removes a real slice of numerical work — anything built on `gpuArray`, CUDA-backed toolboxes, GPU-accelerated solvers. For those workloads this is not a compromise to engineer around; it is the wrong platform.
  • RAM is fixed at template-build time. `--memory-mb` is baked into the snapshot and `memory_mb` on a create is silently corrected to match. A solve that needs 32 GiB needs a template baked at 32 GiB, and the number cannot be a per-job parameter. Maintain a small family of sizes rather than pretending this is flexible.
  • No private-network attachment today. Guests egress through the host with NAT. A licence manager reachable only inside a corporate VPN needs a tunnel terminated on the host, a deliberately exposed proxy, or a self-hosted deployment inside your own network. Prove one guest can check out and release a seat before you build anything downstream.
  • Octave is not MATLAB, and the gap is exactly where your users live. Simulink, modern `classdef` code and toolbox-specific functions are the three walls, and you will meet them in that order of severity.
  • A fat warm snapshot is storage and bytes on the wire. Snapshot size grows with the memory you captured; the first restore on each host pulls it, which is why cross-host creates are measurably slower than same-host ones. Warming six gigabytes of data into a session because it made the demo look good is a recurring bill.
  • Warm memory is billed memory. Working-set GiB-hours are charged while a guest is live, so a fat idle guest costs more than a thin idle guest. The win here is on the CPU side — start up once instead of N times — not a free lunch on RAM.

Where to start

  1. Find out what you are licensed for, in writing, and ask MathWorks how their model counts short-lived VMs. Everything else is downstream of that answer.
  2. If the workload is a known algorithm over many inputs, price up MATLAB Compiler plus the Runtime. It is the path that turns a licensing problem into a build problem, which is the trade you want.
  3. If the workload is user-authored code you do not want to license, prototype on Octave and test the real scripts — including a representative `.mat` round-trip and any figure output that has to look a specific way. Do not infer compatibility from a category name.
  4. Build one template with the runtime baked in, sized deliberately. Then measure the create path and the first-line-of-user-code latency, because that difference is the number that justifies the whole exercise.
  5. Pin the thread counts before you measure anything else, or your first benchmark will tell you a confident lie about parallelism.
  6. If you are fanning out against a counted licence pool, write the orphaned-seat reconciler on day one, not after the first time the queue stalls behind seats held by processes that no longer exist.
You can engineer around a cold start, a memory ceiling and a hostile tenant. You cannot engineer around a contract, and the contract is the first thing on the critical path.

Which is an unusual thing for an infrastructure post to conclude, and the reason I put it first rather than last. The microVM part of this problem is solved and fairly boring: bake the runtime, restore in under 200 ms, give each tenant a boundary that a bad MEX file cannot argue with, pin the threads, reseed the generator. The part that will decide whether you ship is a conversation with your vendor — so have it early, while the architecture is still a document rather than a cluster.

Frequently asked questions

Can I just run 200 MATLAB processes in 200 microVMs?

Only if you hold the entitlements for 200 concurrent sessions, and that is a contract question rather than a technical one. Individual and named-user licences are tied to a person and a small number of machines; they are not a fan-out mechanism. Concurrent or network licences are counted centrally by a licence manager daemon that hands out a fixed number of seats, so your fan-out ceiling is your seat count and nothing about the infrastructure changes it. MathWorks also publishes containerised and cloud-oriented options and products aimed specifically at serving MATLAB code at scale, and which of those you hold determines what is available to you. Two things to do before any infrastructure work: find out in writing what you are licensed for, and ask MathWorks directly how their model counts a VM that lives for ninety seconds — a licence model written for workstations and a workload of eight hundred short-lived guests are not obviously compatible, and the person who can answer that does not work at your company. None of this is licensing advice; check your own agreement and their current terms.

What is the difference between MATLAB Compiler, the MATLAB Runtime and an interactive MATLAB licence?

An interactive licence is what a person uses to sit in front of MATLAB and write code, and a concurrent licence of that kind consumes a counted seat for as long as the session lives. MATLAB Compiler is a product that takes your MATLAB code and produces a standalone application. That application runs against the MATLAB Runtime, which MathWorks distributes free of charge and describes as royalty-free to deploy alongside compiled artefacts — so the compile step needs a licensed MATLAB with Compiler, and the N workers that execute the result do not need interactive seats. For a fan-out workload that is the genuinely clean answer, because it converts a licensing problem into a build-artefact problem. Verify the specifics against MathWorks' current documentation before planning around it: which products and toolboxes are supported for compiled deployment, what the Runtime licence actually permits, and whether your case wants Compiler, Compiler SDK or their production-server product. Two practical gotchas: the Runtime version must match the MATLAB release that compiled the artefact, so pin both together, and the install is large enough that the default rootfs size in a template build will not hold it.

Is GNU Octave a drop-in replacement for MATLAB?

For the core language, often yes; for anything involving toolboxes, no, and that gap is where the trouble lives. Octave implements the MATLAB language well and a script that is mostly matrices, loops and basic linear algebra will usually run unchanged. The cliff is that Octave's package collection is its own ecosystem rather than a reimplementation of MathWorks' toolboxes — function coverage, names and edge-case numerical behaviour all differ, so "there is an Octave signal processing package" does not mean your Signal Processing Toolbox script runs. Simulink has no Octave equivalent at all, which rules out model-based design outright. Support for modern `classdef` object-oriented code has improved over the years and is still incomplete, so a codebase built around MATLAB classes is a poor migration candidate. And data interchange needs testing rather than trusting: round-trip a representative `.mat` file early, especially a v7.3 file, which is HDF5 underneath. The useful rule is that Octave is an excellent answer when you control the code or a human will fix what breaks, and a risky one when you are promising MATLAB compatibility to a third party whose scripts you have never seen.

Why does my numerical job get slower when I run more jobs in one sandbox?

Almost certainly BLAS thread oversubscription. A threaded BLAS such as OpenBLAS starts roughly one thread per core by default, and every PandaStack guest runs 8 burstable vCPUs — fixed, every template — so one Octave or MATLAB process will happily claim eight threads of linear algebra. Add job-level parallelism on top, four or eight workers in the same guest, and you are asking for 32 or 64 compute threads on 8 cores. The threads then spend their time contending and migrating rather than computing, and throughput falls as you add workers, which is the most confusing shape a performance problem can take because there is no error anywhere. Fix it in the template, not per job: set `OPENBLAS_NUM_THREADS`, `OMP_NUM_THREADS` and `MKL_NUM_THREADS` in `/etc/environment` (Docker `ENV` does not reach a detached process in the guest), start batch MATLAB with `-singleCompThread`, and check `maxNumCompThreads` for your release before relying on its setter. Then verify what you actually got rather than assuming, and parallelise at exactly one level — normally the outermost one, where the work is genuinely independent. If `parfor` or a parallel pool is in play you have three levels at once, and a pool is its own licensing question.

I snapshot a warm MATLAB session and fan out a Monte Carlo study. What breaks?

Your randomness, silently. A full snapshot captures the guest's memory as well as its disk, which is exactly why it is valuable — the restored child gets a live process with its paths resolved and its data loaded. It also means the random number generator state is captured. Kernel-sourced randomness does diverge across restores, because the guest CRNG re-mixes RDRAND per extract, so `getrandom` and `/dev/urandom` hand each child different bytes. What does not diverge is any generator your warm-up already constructed: an `rng(42)` executed before the snapshot, a `RandStream` object built during the warm call, an Octave `rand('state', ...)` — that state sits in the snapshotted heap, byte-identical in every child, continuing the same stream from the same position. Fan out a thousand trials and you get a thousand runs of the same trial with implausibly tight error bars. The rule has two acceptable halves: construct no generator state before the freeze, or reseed from a per-child source as the job's first action. Derive the seed from the trial index so the study stays replayable. In Octave, note that `rand` and `randn` are seeded separately — seed both — and in either language, seeding the global generator does not reach a named stream object your warm-up created, so reseed that by name too.

Keep reading

Related posts

  • Running a 2009 app in 2026: legacy workloads in microVMs

    Containerising a genuinely old application is where modernisation projects go to die, because a container shares the host kernel and a container image never pins one. A microVM boots its own. Here is the practical shape: rescue, isolate, snapshot, then strangle.

  • Per-Tenant Analytics Queries in Isolated microVMs

    One customer's SELECT * cross join shouldn't take down every other tenant. Run each tenant's user-defined transforms — SQL, pandas, dbt-style models, notebooks — in its own Firecracker microVM, so a runaway query is capped to one VM you can just kill.

  • Testing Against Ten Toolchains Without Ten Broken Runners

    If you ship a library you own a matrix. On a shared runner the legs quietly contaminate each other. A microVM per leg makes the matrix mean what it says.

  • Per-Tenant LLM Fine-Tuning Jobs in Isolated microVMs

    You let customers upload their own training data and run fine-tuning jobs. That job holds a private dataset, a customer-supplied training script, and a credential. Run each one in its own Firecracker microVM so a poisoned script can't read another tenant's data or walk off with your keys.

  • Per-Tenant Fraud Rules in Isolated microVMs

    When customers upload their own scoring rules and those rules run on every checkout, a shared worker pool means one tenant's rule can read another's transaction features — or pin a core and add latency to everyone. Give each tenant's rule its own Firecracker microVM.

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.