all posts

Notebooks in Production: Parameterised Runs in Disposable microVMs

Ajay Kumar··10 min read

Somewhere in your company there is a number on a slide that comes out of cell 34 of a notebook. The notebook must be run in order. Cells 11 and 12 must be skipped because they were for the thing in March. It reads a CSV from a path that exists on exactly one laptop. It is run by Dmitri, who is very good at it, and who is on holiday the week the board meets.

The usual reaction to this is to say the notebook should have been a module, and the usual reaction to that is nothing happening for three years. Which is worth taking seriously rather than sneering at, because the people resisting have a real argument: the notebook is not a draft of the analysis, it is the analysis. The plots are in it. The prose explaining why the outlier in Q2 was a billing migration and not churn is in it. Rewrite it as a well-factored Python package and you have preserved the arithmetic and thrown away the thing the business actually consumes.

Papermill's bet is to refuse that trade. Keep the notebook, inject the parameters, and treat the executed copy as the output. I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs, and I want to make the case that the missing half of that bet is where the notebook runs — because the failure mode in the story above is not the notebook. It is the laptop.

The laptop is the failure mode

Separate two complaints that usually arrive fused together. One is "notebooks encourage bad engineering practice", which is an argument about taste and has been had enough times. The other is "this number is produced by a process that exists in one place, has no record of its inputs, cannot be re-run by anyone else, and fails silently" — which is an argument about operations, and is correct.

Everything in the second complaint is a property of the environment, not of the file format. A notebook executed by a scheduler, in a fresh kernel, against pinned dependencies, with its parameters recorded and its output archived, has none of those problems. A beautifully factored Python module run by hand on Dmitri's laptop has all of them. The format is a red herring; the question is whether the run is reproducible and whether anybody notices when it stops happening.

The thesis: a notebook makes a perfectly good batch job artefact, as long as every run happens somewhere that has never run a notebook before. The format was never the problem. The accumulated state was.

What Papermill actually does

The mechanism is smaller than its reputation. You tag one cell in the notebook `parameters` — an ordinary cell-level tag in the notebook metadata, settable from the JupyterLab property inspector. That cell contains the defaults, written as normal assignments, so the notebook still runs interactively with no special handling.

At execution time Papermill inserts a new cell tagged `injected-parameters` immediately after the `parameters` cell, containing only the values you overrode. Then it executes the notebook and writes an executed copy to the output path. Parameters arrive by `-p key value` for individually typed values, `-r` to force a raw string, `-f` from a YAML file, `-y` as a YAML string, or `-b` base64-encoded — and in a driver you want `-f` or `-y`, because the moment you build `-p` arguments by string concatenation you have invented a quoting bug.

Two consequences of that design are worth internalising before you build on it:

  • If no cell is tagged `parameters`, the injected cell goes to the top of the notebook. It does not error. Your overrides are then shadowed by the author's own assignments further down, and the job runs happily on last quarter's defaults. This is the single most common way a parameterised notebook pipeline lies to you.
  • Injection replaces the variable, not everything derived from it. The docs use the example directly: override `a` from 1 to 9 and a later `twice = a * 2` that was evaluated in the parameters cell still reads 2. Keep the tagged cell to bare inputs and compute derived values below it.
  • The output notebook is both the result and the receipt. Papermill records per-cell `start_time`, `end_time`, `duration`, `exception` and `status` in the notebook metadata, so "which cell was slow" and "which cell raised" are properties of the artefact. On failure you get a notebook with the traceback sitting in the cell that produced it, with every preceding cell's output still visible. That is a materially better debugging experience than a stack trace in a log aggregator.

`nbconvert --execute` is not the same tool

`jupyter nbconvert --execute` will also run a notebook, and people reach for it first because it ships with Jupyter. It executes the input cells and saves the results, and you can raise its per-cell limit with `--ExecutePreprocessor.timeout=600` (the default is 30 seconds, which is the first thing most real notebooks trip over) or let it run to the end through failures with `--ExecutePreprocessor.allow_errors=True`.

What it does not have is parameterisation. There is no supported way to tell `nbconvert` that `region` is `emea` this time; you would have to template the notebook source, which is exactly the hand-rolled fragility Papermill exists to delete. The sane division of labour is Papermill for execution and parameter injection, `nbconvert` for rendering the executed notebook into something a human reads — HTML with `--no-input` if the audience wants the chart and not the code.

Hidden state is the entire problem

Here is the thing that makes notebooks genuinely different from scripts, and it is not the JSON. A notebook has two orderings: the order the cells appear in the file, and the order they were actually executed. The `.ipynb` records the second one, per cell, as `execution_count` — so you can open a notebook and watch it count 1, 2, 3, 7, 4, 19, 5 and know exactly what happened. Someone defined a variable, edited the cell that defined it, never re-ran it, and kept going. The kernel knows the value. The file does not.

The result is a notebook that works and is not runnable. Not broken-and-nobody-noticed: genuinely working, producing correct output, in a kernel whose state cannot be reconstructed from the document. Restart the kernel and run all and it fails on cell 6. The author's experience of the analysis and the artefact's contents have diverged, and nothing in the tooling forced them back together.

The hygiene half of the answer is well known. `nbstripout` as a git filter strips outputs and execution counts on commit, which stops notebook diffs being unreadable and stops large embedded PNGs bloating the repository — and, as a side effect, removes the evidence of out-of-order execution from review. Useful, but it addresses the diff, not the runnability.

The only honest test of "does this notebook run" is a fresh, non-interactive kernel executing it top to bottom. There is no partial credit and no way to fake it: either every cell's dependencies are satisfied by the cells above it, or the run fails. A disposable microVM per run makes that test the default rather than a discipline — you do not have to remember to restart the kernel, because there was never a kernel to restart.

Restart-and-run-all is a thing you have to remember to do. A fresh guest per run is a thing you cannot forget to do.

Why a guest per run, and not a shared kernel farm

The obvious cheaper architecture is a pool of warm kernels that you hand notebooks to. It works until it meets the things real notebooks do, all of which I have watched happen:

  • `!pip install` in a cell. Analysts do this constantly and it is not unreasonable — they found a package that solves their problem. On a shared kernel farm it is a global mutation. One notebook pins `numpy<2` at 03:00 and a different team's notebook produces subtly different floating-point results at 06:00, with no error and no audit trail.
  • Writing to `/tmp` and expecting it next run. This is a cache that works on a laptop and works on a shared worker and then one day the worker is replaced and the notebook has been reading a stale intermediate for six weeks. A guest that dies after every run makes that dependency fail loudly on the first run instead of silently on the hundredth.
  • Leaking memory in cell 12. A notebook that holds a 6 GiB DataFrame and a reference cycle is not a bug anyone is going to fix. On a shared kernel it accumulates until something unrelated gets OOM-killed. In a per-run guest it dies with the VM, every time, by construction.
  • Reading a CSV from the author's home directory. This one is a feature of the disposable model, not a problem with it: the run fails immediately and specifically, and the fix — fetch the input from object storage, parameterised — is the fix you wanted anyway.
  • Being somebody else's code. Notebooks get shared, forked and pasted. If you run notebooks submitted by customers, analysts in another business unit, or an LLM, you are executing arbitrary Python with the credentials of whatever runs it. A kernel-level sandbox is not where you want that boundary.

The historical objection to a VM per run was cost: nobody wanted to pay a multi-second boot to run a 40-second notebook. Snapshot-restore removes that objection. Every create on PandaStack restores a baked Firecracker snapshot rather than booting — p50 179 ms, p99 203 ms, of which the `/snapshot/load` step itself is about 49 ms. The first create of a template before its snapshot exists is a real cold boot at around 3 seconds, and then the snapshot is captured and every create after that is on the fast path.

Four ways to put a notebook into production

The four real options, compared on the axes that actually trade off against each other.
ApproachIsolation between runsDependency driftKeeps plots and narrativeCost of a run
Shared kernel farmNone worth the name: one process, one filesystem, one site-packagesWorst: any `!pip install` mutates everyoneYesLowest per run, highest per incident
Container per runProcess-level; good enough for your own code, thin for someone else'sGood: pinned image, rebuilt deliberatelyYesLow; seconds of start-up, image pull on a cold node
microVM per runA separate kernel and a separate guest, discarded after each runGood: baked template, and a stray install dies with the guestYes~179 ms restore plus the notebook's own runtime
Rewrite as a moduleWhatever your job runner gives you; orthogonal to this choiceNormal packaging problemNo — this is the cost nobody pricesLowest, after the rewrite that keeps not happening

These are not mutually exclusive and the honest answer for most teams is the third row now and the fourth row for the handful of notebooks that turned out to be load-bearing. The point of the table is the fourth column: "rewrite it as a module" is not a free action, and the thing it costs is the reason the notebook existed.

Bake the stack, then restore it per run

A notebook run that begins with `pip install pandas matplotlib scikit-learn` has thrown away the entire latency argument. The dependencies belong in the image, where they are paid for once at bake time.

Our `code-interpreter` template already carries most of a scientific stack baked in — numpy 2.3.5, pandas 2.2.3, scipy, scikit-learn, matplotlib, seaborn, plotly, plus `jupyter-server` and `ipykernel`, on a 12 GiB rootfs. `nbconvert` comes along as a `jupyter-server` dependency. Papermill does not; it is not in the image, so either `pip install papermill` as the first line of your runner (and pay for it on every run) or bake a template that has it. Bake it.

Two numbers decide whether a custom template works, and both are build-time only. `--size-mb` sets the ext4 rootfs size and defaults to 1024 MB on the build API (the bundled Go CLI sends 2048), which a pandas-plus-matplotlib-plus-scikit-learn stack will exhaust without ceremony — size it deliberately. `--memory-mb` is where guest RAM is chosen, and it is the only place: Firecracker cannot change a guest's vCPU count or RAM at snapshot restore, so `cpu=` and `memory_mb=` on a create are silently corrected to the baked snapshot's values. `--cpu` is deprecated and ignored; every first-party template runs 8 burstable vCPU.

`code-interpreter` is baked at 2 GiB of RAM (`base` is 4 GiB). A notebook that loads a few million rows into pandas will be killed by the OOM killer inside the guest, and no argument you pass at create time will help — the agent will correct it. If your notebooks need 16 GiB, that is a template baked with `--memory-mb 16384`, not a per-job parameter. This surprises people doing data work more than any other constraint we have.

One more thing that catches people building images for microVMs rather than containers: Docker `ENV` instructions are not injected into the Firecracker guest. A detached process in the guest gets its environment from `/etc/environment` via PAM, so anything your runner needs — `MPLBACKEND=Agg` so matplotlib does not look for a display, `PYTHONUNBUFFERED=1` so you see output as it happens — should be written there in the Dockerfile, not set with `ENV`.

The run, inside the guest

#!/usr/bin/env bash
# Runs INSIDE the guest. The driver has already written /work/report.ipynb and
# /work/params.yaml; everything below is ordinary papermill, nbconvert and tar.
set -euo pipefail

RUN_ID="${RUN_ID:?set by the driver}"
mkdir -p /work/out

# The notebook needs ONE cell tagged `parameters`. papermill then inserts a new
# cell tagged `injected-parameters` immediately after it, containing only the
# overrides. If no cell carries the tag, the docs are explicit that the injected
# cell goes to the TOP of the notebook -- which means the author's own defaults,
# further down, silently shadow every parameter you thought you set. Tag the cell.
#
#   -f  read parameters from a YAML file, so the driver never shell-quotes a value
#   -k  pick the kernel by name rather than trusting the notebook's recorded one
#   --execution-timeout  is PER CELL, not per notebook. Budget accordingly.
#   --start-timeout      how long to wait for the kernel to come up
#   --log-output         cell stdout goes to OUR stdout, which is how the run gets logs
papermill \
  /work/report.ipynb \
  "/work/out/${RUN_ID}.ipynb" \
  -f /work/params.yaml \
  -k python3 \
  --cwd /work \
  --execution-timeout 900 \
  --start-timeout 120 \
  --log-output \
  --no-progress-bar \
  --stderr-file "/work/out/${RUN_ID}.stderr"

# The executed notebook is the result AND the receipt. papermill records
# start_time, end_time, duration, exception and status per cell in the notebook
# metadata, so "which cell was slow" and "which cell raised" live in the
# artefact rather than in a log you have to correlate against it afterwards.

# Render a copy for the people who asked for a number, not a notebook.
# --no-input excludes the input cells and prompts. ExecutePreprocessor is not
# involved: the notebook is already executed, this is only a conversion.
jupyter nbconvert \
  --to html \
  --no-input \
  --output-dir /work/out \
  "/work/out/${RUN_ID}.ipynb"

# ONE blob out. The filesystem read is a whole-file read -- a directory is an
# error, not a recursive copy -- so tar the output directory and read the tarball.
tar -C /work/out -czf "/tmp/${RUN_ID}.tar.gz" .
echo "packed /tmp/${RUN_ID}.tar.gz"

The driver, outside the guest

The driver's job is unglamorous: one guest per parameter set, write three files in, stream one command, read one blob out. The interesting decisions are all about getting the artefacts out before the guest stops existing.

import json
from pathlib import Path
from pandastack import Sandbox

NOTEBOOK = Path("report.ipynb").read_bytes()      # the analysis, unchanged
RUNNER   = Path("run.sh").read_text()             # the bash above

# One parameter set per run. The notebook is identical in all of them; this is
# the whole bet papermill is making.
GRID = [
    {"region": "emea",   "as_of": "2026-09-30", "currency": "EUR"},
    {"region": "amer",   "as_of": "2026-09-30", "currency": "USD"},
    {"region": "apac",   "as_of": "2026-09-30", "currency": "JPY"},
]

def run_one(params):
    run_id = f"{params['region']}-{params['as_of']}"

    # A fresh guest. Not a reset one, not a cleaned one -- one that has never
    # run anybody else's notebook. Create is a snapshot restore: p50 179 ms.
    # cpu/memory_mb are deliberately absent: the agent silently corrects them
    # to the baked snapshot's values, so passing them only misleads the reader.
    # ttl_seconds is an IDLE timeout, not a walltime budget.
    sbx = Sandbox.create(
        template="papermill-reports",
        ttl_seconds=1800,
        metadata={"run": run_id},
    )
    try:
        # filesystem.write is single-file only and the agent caps the body at
        # 32 MiB. A notebook is fine; a 4 GiB parquet file is not -- pull that
        # from object storage inside the guest instead of pushing it through here.
        sbx.filesystem.write("/work/report.ipynb", NOTEBOOK)
        # Every JSON document is valid YAML, so this needs no yaml dependency.
        sbx.filesystem.write("/work/params.yaml", json.dumps(params))
        sbx.filesystem.write("/work/run.sh", RUNNER)

        # timeout_seconds is CLIENT-side only: the agent decodes it on both exec
        # endpoints and applies it on neither, so exec_stream raises the HTTP
        # timeout while one-shot exec() gives up at 30 s. A notebook run is not
        # a 30-second operation, so stream it -- and bound it in the shell.
        lines = []
        rc = sbx.exec_stream(
            f"RUN_ID={run_id} bash /work/run.sh",
            on_stdout=lines.append,
            on_stderr=lines.append,
            timeout_seconds=1500,
        )

        # Pull the artefacts out BEFORE anything can reap the guest. On failure
        # the executed notebook is the most valuable thing in the building: it
        # has the traceback in the cell that raised it, in context.
        try:
            blob = sbx.filesystem.read(f"/tmp/{run_id}.tar.gz")
            Path(f"out/{run_id}.tar.gz").write_bytes(blob)
        except Exception as e:                 # papermill died before tar ran
            lines.append(f"\n[no artefact: {e}]\n")

        Path(f"out/{run_id}.log").write_text("".join(lines))
        return run_id, rc
    finally:
        # kill() cascade-deletes this sandbox's snapshots as well. Fine here --
        # we took none -- but do not reach for it reflexively in a job whose
        # evidence you just captured as a snapshot.
        sbx.kill()

Path("out").mkdir(exist_ok=True)
results = [run_one(p) for p in GRID]           # serial; see the note on fan-out
for run_id, rc in results:
    print(f"{run_id}: exit {rc}")

That loop is deliberately serial so the file reads as a recipe rather than a concurrency exercise. In production you want these in flight together — a thread pool over `run_one` is enough, since every call is network-bound — and the thing to tune is not the creates, which are cheap, but how many simultaneous 2-GiB guests your account is allowed and how many concurrent readers your source data tolerates. A hundred notebooks hitting one Postgres replica at 06:00 is a thundering herd regardless of how elegantly you isolated them.

Getting the artefacts out

This is the part of every sandbox-based pipeline that gets designed last and then rewritten. The filesystem API is single-file in both directions. `filesystem.read` is a whole-file read — hand it a directory and you get an error, not a recursive copy — and `filesystem.upload` is literally `write(remote, Path(local).read_bytes())`, so passing it a directory raises `IsADirectoryError`. The write path is capped at 32 MiB by the agent.

So: tar the output directory in the guest and read one blob. That is the whole trick, and it has the pleasant property that the tarball is the unit you archive — executed notebook, rendered HTML, figures, stderr, together, named by run. For outputs measured in gigabytes, do not route them through the control plane at all: have the notebook write to object storage from inside the guest and read back only a manifest. The control plane is a good place for a 3 MB notebook and a bad place for a 3 GB parquet file.

Read the artefacts before the guest goes away, and be careful with cleanup. An explicit `kill()` cascade-deletes that sandbox's snapshots; the idle reaper does not. If your failure path snapshots the guest for forensics and then calls `kill()` in a `finally`, you have destroyed the evidence you just captured.

Scheduling the thing

"Every morning at 06:00" is the point of all this, and it is worth being precise about what our schedules surface actually is, because it shapes the architecture. A schedule is a five-field cron expression — minute, hour, day of month, month, day of week, interpreted in UTC — attached to a function id. Not to a sandbox, not to a container image: a function. No seconds field, and no `@daily`-style shorthand; the parser is configured for the five standard fields only.

Each due schedule fires the function in its own fresh microVM, which is deleted afterwards, and the run is recorded with its exit code, stdout and stderr, queryable via the schedule's runs endpoint. Multiple API nodes claim due schedules with `SELECT ... FOR UPDATE SKIP LOCKED`, so a schedule fires once across the fleet rather than once per node. The poll interval is 30 seconds, which means 06:00 means "within half a minute of 06:00" — fine for a report, worth knowing if you were planning to coordinate two schedules by timing.

The constraint that matters for notebooks: a scheduled function run is given a five-minute budget end to end, of which the wait for the sandbox to become ready can consume 30 seconds. A notebook that takes 20 minutes cannot run inside the scheduled function. So the shape is a dispatcher, not a worker — the scheduled function reads the parameter grid, creates the notebook sandboxes, and returns. Track the runs in your own table and let the notebook guests finish on their own clock, with their own idle TTL, writing their tarballs where the next stage expects them.

The other half of scheduling is the half nobody builds: alerting on the run that did not happen. A notebook job that fails loudly is a good day. A notebook job that silently stopped being scheduled in July and whose dashboard has been showing June's number since then is the failure this whole architecture exists to prevent, and it is not prevented by anything in the paragraphs above. Assert on freshness in the consumer, not just on exit codes in the producer.

When you do want a persistent kernel

Everything so far argues for a fresh non-interactive kernel. There is a legitimate opposite case and it is worth naming so the two do not get confused. If you are building something notebook-shaped for a human or an agent — run a cell, look at the chart, run another cell that uses the variable from the first — you want state to persist and you want rich outputs back, not files on disk.

from pandastack import Sandbox

# A persistent kernel, for the INTERACTIVE case. State survives across
# run_code calls and rich outputs come back as results, so there is no
# save-to-disk-then-download dance for a chart.
sbx = Sandbox.create(template="code-interpreter", ttl_seconds=900)
ctx = sbx.create_code_context(language="python")

ctx.run_code("import pandas as pd; df = pd.read_parquet('/work/sales.parquet')")
ex = ctx.run_code(
    "import matplotlib.pyplot as plt\n"
    "df.groupby('region').revenue.sum().plot.bar()\n"
    "plt.show()"
)
png_b64 = ex.results[0].png     # base64 PNG, straight back over the wire

ctx.close()

# This is the WRONG tool for a scheduled run. A batch job wants a kernel with
# no history, because "it worked in my kernel" is the bug being eliminated.
# Reuse a warm kernel across parameter sets and you have rebuilt the shared
# kernel farm, with its state leaks, inside one sandbox.

Use that for the notebook product, the agent that explores a dataset, the human at a keyboard. Do not use it for the 06:00 run. If what you are really building is hosted notebooks for a group of people, that is a different architecture with different concerns — per-user isolation, idle cost, who owns the kernel — and I have written about it separately.

A footnote if you warm-snapshot a kernel

There is a tempting optimisation here: warm a kernel with `import pandas`, `import matplotlib` and your data already loaded, snapshot it, and restore that snapshot per run instead of starting cold. It works, and it has one sharp edge. Kernel-sourced randomness diverges across restores — the guest CRNG re-mixes RDRAND per extract, so `getrandom(2)` and `/dev/urandom` give each restored child different bytes. What does not diverge is any generator your warm-up already constructed. numpy's global `RandomState`, seeded before the snapshot, is part of the snapshotted heap: byte-identical in every child, continuing the same stream from the same position. The kernel protects you; your own library does not. If any of your notebooks sample, bootstrap or shuffle, reseed from a per-run value as the first thing the run does — and derive that value from the parameter set so the run stays replayable.

What this does not fix

A post that stopped at the previous section would be an advert. The real limits:

  • It does not make the notebook good. A notebook with a hardcoded path, a cell that only works in March and no assertions is still all of those things when it runs in a pristine guest. What you gain is that the failures become deterministic and attributable, which is the precondition for fixing them, not the fix.
  • RAM is a template decision. The agent corrects `cpu` and `memory_mb` on create to the baked snapshot's values, because Firecracker cannot change them at restore. Data work is exactly the workload that most wants to ask for more memory at submit time, and it is exactly the thing this architecture cannot give you. You maintain a small family of templates at the sizes you actually need.
  • There is no GPU. If the notebook trains anything on CUDA, stop reading — this is the wrong platform for it, not a compromise to engineer around.
  • Parameterising a notebook does not make it a DAG. Papermill runs one notebook. Dependencies between notebooks, retries, backfills and partial reruns are an orchestration problem, and you will end up with an orchestrator. The useful framing is that Papermill plus a disposable guest is a good execution unit for an orchestrator to call, not a replacement for one.
  • The scheduled function's five-minute budget is a real constraint and the dispatcher pattern is extra moving parts: your own run table, your own timeout policy, your own retry. That is work the architecture requires and does not do for you.
  • A notebook is still a bad review artefact. `nbstripout` makes the diff readable; it does not make a 400-line cell reviewable. If the number matters, the parts of the notebook that compute it want to be a module the notebook imports — which, pleasantly, is the incremental version of the rewrite that never happens.

Where to start

Pick the notebook that produces the number people argue about. Tag its parameters cell. Bake a template with its dependencies and Papermill. Run it in a fresh guest, from the top, with no human present, and find out what breaks — because something will, and that is the first useful information anyone has had about this notebook in two years. Then archive the executed copy, and put the number's freshness on a dashboard that complains.

You have not made the notebook into software. You have made the run into a thing that happens the same way every time, in a place that forgets. Which turns out to be most of what anyone wanted from "productionising" it.

Frequently asked questions

Should I use Papermill or `jupyter nbconvert --execute` to run notebooks non-interactively?

Both execute a notebook; only one parameterises it. `nbconvert --execute` runs the input cells and saves the results, and you will immediately need `--ExecutePreprocessor.timeout=600` or similar because the default per-cell timeout is 30 seconds. But there is no supported way to vary an input per run — you would be templating the notebook source, which is the fragility Papermill exists to remove. Papermill injects a cell tagged `injected-parameters` right after your cell tagged `parameters`, takes values from `-p`, `-r`, `-f` (a YAML file), `-y` (a YAML string) or `-b` (base64 YAML), and writes an executed copy of the notebook as its output. In a driver, prefer `-f` or `-y`: the moment you build `-p` arguments by string concatenation you have invented a quoting bug. The practical division of labour is Papermill for execution and parameters, `nbconvert` for turning the executed notebook into HTML for people who want the chart and not the code — `--no-input` excludes the input cells and prompts.

Why does a fresh VM per notebook run matter if I already pin my dependencies in an image?

Pinning the image handles dependency drift. It does not handle accumulated state, and notebooks are unusually good at accumulating it. The failure that keeps happening is not a version mismatch; it is a notebook that works in a kernel whose state cannot be reconstructed from the document. A `.ipynb` records `execution_count` per cell, so you can open one and watch it count 1, 2, 3, 7, 4, 19 — someone defined a variable, edited the defining cell, never re-ran it, and carried on. The kernel knows the value and the file does not. The only honest test is a fresh non-interactive kernel running the notebook top to bottom, and a disposable guest makes that test the default rather than a discipline you have to remember. The other half is the things notebooks do at runtime: `!pip install` mutating a shared environment, a cache in `/tmp` that silently goes stale, a leaked 6 GiB DataFrame. In a per-run guest each of those dies with the VM. On a shared kernel farm each is somebody else's incident at 03:00.

My notebook needs 16 GiB of RAM for a pandas job. Can I set memory_mb when I create the sandbox?

No, and this is the limitation that most often surprises people doing data work. Firecracker cannot change a guest's vCPU count or RAM at snapshot restore, so whenever a template has a baked snapshot the agent silently corrects the create request's `cpu` and `memory_mb` to the baked values. Passing `memory_mb=16384` is not an error; it is ignored, which is worse. RAM is chosen once, at template build time, with `--memory-mb`. Our first-party `code-interpreter` template is baked at 2 GiB and `base` at 4 GiB, so a notebook that loads a few million rows into pandas will meet the OOM killer inside the guest regardless of what you asked for. The answer is a custom template built with the memory you need, and in practice a small family of templates at the two or three sizes your notebooks actually want. While you are there, note that `--size-mb` — the rootfs size — defaults to 1024 MB, which a pandas plus matplotlib plus scikit-learn stack will exhaust; size it deliberately. `--cpu` is deprecated and ignored, and every first-party template runs 8 burstable vCPU.

How do I get the executed notebook and its figures out of the sandbox?

Tar the output directory inside the guest and read one blob. The filesystem API is single-file in both directions: `filesystem.read` is a whole-file read, so handing it a directory is an error rather than a recursive copy, and `filesystem.upload` is implemented as `write(remote, Path(local).read_bytes())`, so passing a directory raises `IsADirectoryError`. The write path is capped at 32 MiB by the agent. So the run's last step is `tar -czf /tmp/<run>.tar.gz -C /work/out .` and the driver's next call is a single `filesystem.read` of that path — which has the useful side effect that the tarball is the unit you archive: executed notebook, rendered HTML, figures and stderr together, named by run. For outputs measured in gigabytes, do not route them through the control plane at all; have the notebook write to object storage from inside the guest and read back only a manifest. One more thing: read the artefacts before the guest goes away, and remember that an explicit `kill()` cascade-deletes that sandbox's snapshots while the idle reaper does not — a `finally: kill()` can destroy forensics you just captured.

Can PandaStack schedules run my notebook every morning at 06:00 directly?

Almost. A schedule is a five-field cron expression — minute, hour, day of month, month, day of week, in UTC — attached to a function id, so `0 6 * * *` is exactly expressible. There is no seconds field and no `@daily`-style shorthand. Each due schedule fires the function in its own fresh microVM, deletes it afterwards, and records the run with its exit code, stdout and stderr; multiple API nodes claim due schedules with `SELECT ... FOR UPDATE SKIP LOCKED` so a schedule fires once across the fleet, and the poll interval is 30 seconds, so 06:00 means within half a minute of 06:00. The constraint for notebooks is the budget: a scheduled function run gets five minutes end to end, of which waiting for its sandbox to become ready can take 30 seconds. A 20-minute notebook does not fit. So make the scheduled function a dispatcher rather than a worker — read the parameter grid, create the notebook sandboxes, return — and let those guests finish on their own idle TTL, writing their tarballs where the next stage looks for them. You own the run table and the retry policy in that design; the schedule only guarantees the dispatcher fires.

Keep reading

Related posts

  • How to Sandbox Untrusted Jupyter Notebooks Per User

    Every user's notebook runs arbitrary Python. On a shared kernel, one tenant's os.system is everyone's problem. The fix: one disposable microVM per session.

  • Best Jupyter notebook hosting platforms in 2026

    JupyterHub, Colab, Deepnote, Hex, Databricks, SageMaker Studio, Modal, Binder, Codespaces, and PandaStack — judged on who runs the kernel, what it costs while idle, and whether you can safely run somebody else's notebook.

  • Sandboxing LLM Batch Post-Processing at Scale

    Run an LLM over thousands of inputs and each output is a snippet of code to execute. Fan those items out into a pool of throwaway microVMs — one per item or per batch — so item #4,197's `while True: os.fork()` dies with its VM, not your worker fleet.

  • Isolating Batch Jobs and Queue Workers with MicroVMs

    A shared worker runs every tenant's job in one process, on one box, sharing one /tmp. One bad job — an OOM, a fork bomb, a leaked file descriptor — is a murder-suicide pact with every other job on the machine. Give each job its own microVM instead.

  • The Best Batch Compute Platforms for Bursty Jobs in 2026

    Nine ways to run a few thousand independent jobs, grouped by what they actually are rather than ranked. Most of your decision is made by two facts: whether you need GPUs, and whether you are willing to pay for a cluster that is idle most of the week.

More in Code execution · See Code interpreter sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.