all posts

Julia, R and the Cost of the First Import

Ajay Kumar··11 min read

Most of the runtimes a platform engineer deals with are cheap to start and expensive to run. Node boots in milliseconds and then spends the rest of its life doing work. A Go binary has no start cost worth measuring. The whole industry's intuition about ephemeral compute — spin it up, do the thing, throw it away — is calibrated on runtimes shaped like that.

Scientific runtimes are shaped the other way round. The first call is catastrophically expensive and every call after it is nearly free, because the first call is where the compiler runs. Julia's community has a name for this and it is not a flattering one: time to first plot. Put that runtime in a sandbox that starts clean every time and you have built a machine whose entire job is to pay the expensive part, repeatedly, forever.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This is the use case that looks worst for ephemeral sandboxes on paper and turns out to be the single best argument for snapshot-based ones — because a snapshot is a memory image, and a compilation cache is exactly the sort of thing that lives in memory.

Two shapes of cost

Write the cost of a job as setup plus work. For a web-stack runtime, setup is small and work is the bill, so you optimise the work and you are allowed to be careless about start-up. For Julia and R, setup can dominate the bill by an order of magnitude or more, and it is not I/O you can prefetch away — it is a compiler and a linker doing real CPU work that nobody asked for in the problem statement.

That changes what "ephemeral" means economically. An ephemeral container that reinstalls and recompiles a dependency tree on every run is not a lean architecture; it is a CPU furnace with a nice API. And on a rate card where CPU is billed on the active CPU-seconds actually burned — ours is $0.054 per vCPU-hour and $0.0162 per GiB-hour, one card for everything — the compile you repeat N times is a line item you can read off an invoice.

The thesis in one line: you cannot make the compilation faster, so make it happen once, at bake time, for the whole fleet. Everything below is the mechanics of that, including the parts it does not fix.

Julia has three caches and they fail differently

Julia compiles ahead of execution but not ahead of time: method instances are compiled when they are first called with a given set of argument types. On top of that, `Pkg` precompiles packages into a depot so the work is not repeated in the next session. Both the depot path and the code-loading mechanism are controlled by `JULIA_DEPOT_PATH`, whose documentation describes it as controlling where Julia looks for registries, installed packages, named environments, repo clones and cached compiled package images. The default user depot on Linux is `~/.julia`, and the glossary in the Pkg docs lists the subdirectories: `packages` for installed versions, `registries` for registry clones, and `compiled` for the cached compiled package images.

What lands in `compiled` got substantially more valuable in Julia 1.9, which added native code caching through package images. Before that, precompilation largely saved type-inference and lowering work; the native code still had to be generated in-process. The package-image devdocs are blunt about what the artefact now contains: "The code section contains native objects that cache the final output of Julia's LLVM-based compiler." The trade is also documented — precompilation itself got slower and the cache files got bigger, and you can turn the whole mechanism off for a session with `--pkgimages=no`. Check the behaviour for the exact minor version you pin, because this is the area of Julia that has moved fastest.

So there are three distinct layers, and an architecture that confuses them will optimise the wrong one:

  • The installed source tree — `packages` in the depot. Downloaded bytes. Cheap to restore, boring, and the only layer most CI caching advice is actually about.
  • The precompiled cache on disk — `compiled` in the depot. CPU-expensive to produce, survives a process restart, and dies the instant you hand the job a fresh filesystem. This is the layer a container volume cache can save.
  • The compiled code in the live process — whatever the JIT produced for the specific calls this session made, plus every data structure the session built. Dies on every process exit, without exception, and no filesystem cache can hold it. This is the layer only a memory image preserves.

The build-time answer for the first two layers is `PackageCompiler.jl`, which bakes a sysimage containing your packages' compiled code so a fresh process starts with it already linked. It works, it is the right tool for a CLI you ship to users, and it has a cost: a sysimage is a large artefact, building it is slow, and `JULIA_CPU_TARGET` becomes your problem the moment the machine that builds it is not the machine that runs it. There is no version of this story where you get to ignore which CPU features the code was generated for.

What none of it solves is the third layer. A sysimage gets you a process that starts warm; it does not get you a process that is already warm, with your data loaded, your model constructed, and the specific method instances for your specific call sites already resident. That gap is small for a CLI and enormous for a fan-out workload where every worker is about to make the same first call.

R is a different shape, and treating them as one thing will cost you

It is tempting to file Julia and R together under "slow science languages" and apply the same fix. Don't. R's expensive moment is installation, not first call. A large fraction of the CRAN packages that matter contain C, C++ or Fortran, and on Linux the default is to install from source, which means you are running a compiler toolchain at install time and getting a shared object out of it. Nothing is deferred to the first `library()` call in the way Julia defers codegen to the first method call.

That is why the R ecosystem's answer looks completely different. `renv` gives you a project library plus a lockfile and a global package cache — on Linux, `~/.cache/R/renv` — so the same resolved versions land in the next project without re-downloading. And Posit Package Manager exists substantially to serve precompiled binaries for Linux distributions, which is the single biggest install-time lever available to an R user: renv will use the Posit public mirror by default and can transform repository URLs into the binary URLs appropriate for the platform. Verify the current distribution coverage and R-version window against Posit's own documentation before you design around it, because that list changes.

The second thing about R that bites in a sandbox is not CPU at all. A `library(tidyverse)`-shaped dependency tree is a large number of packages and a very large number of small files, so restoring it is a metadata and inode workload more than a bandwidth one. Untarring it, syncing it over a network filesystem, or layering it into an image behaves accordingly: the throughput number on your disk is irrelevant and the per-file cost is everything. A baked block device that you clone, rather than a tree you copy, is a qualitatively different operation here.

BLAS is part of the environment, whether you declared it or not

Which BLAS the image links against changes both performance and, in floating-point edge cases, results — a different library means a different reduction order, and a different thread count means a different reduction order again. R on Linux may be linked against a reference BLAS or against OpenBLAS depending on how the distribution packaged it, and that choice does not appear in your lockfile. Neither does the thread count.

The thread count is also a throughput cliff with a very specific shape. OpenBLAS built with threading starts a thread per core by default. Our guests get 8 burstable vCPUs — that is fixed, every template runs 8 — so a single R or Julia process will happily spawn 8 BLAS threads. Now parallelise at the job level too, with several workers in the guest, and you have 8 times as many BLAS threads as you have cores to run them on. The symptom is not an error; it is a workload that gets slower when you add parallelism, which is the single most confusing performance report a user can file. Pin the BLAS thread count to 1 in the template and parallelise at exactly one level.

Step one: bake the environment into a template

There is no first-party Julia or R template today — our catalogue is `base`, `code-interpreter`, `agent`, `browser` and `postgres-16` — so this is a custom template build. You hand us a Dockerfile, we build the image server-side and bake a Firecracker snapshot from it. Two flags matter more than usual for this workload: `--size-mb` sets the ext4 image size and defaults to 1024, which a Julia depot plus a tidyverse-shaped library will blow through without ceremony, and `--memory-mb` is where the guest's RAM is decided, because it is baked into the snapshot and cannot be changed later.

#!/usr/bin/env bash
# Build a template whose depot and R library are already populated and
# already precompiled. Everything in this file is paid ONCE, at bake time,
# for every sandbox that will ever be created from the result.
set -euo pipefail

cat > Dockerfile <<'DOCKERFILE'
FROM pandastack/base:latest

# --- Julia -----------------------------------------------------------------
# Pin the exact version. Precompile caches are keyed on it, and the native
# code caching behaviour has changed across minor versions (1.9 added
# package images); a floating version means a silently cold depot.
ARG JULIA_VERSION=1.11.3
RUN curl -fsSL "https://julialang-s3.julialang.org/bin/linux/x64/${JULIA_VERSION%.*}/julia-${JULIA_VERSION}-linux-x86_64.tar.gz" \
      | tar -xz -C /opt \
 && ln -s "/opt/julia-${JULIA_VERSION}/bin/julia" /usr/local/bin/julia

# Put the depot somewhere deliberate. The default user depot is ~/.julia,
# which ties the cache to one home directory; a system path is easier to
# reason about when several users or services share the guest.
ENV JULIA_DEPOT_PATH=/opt/julia-depot
ENV JULIA_CPU_TARGET=generic
ENV JULIA_NUM_THREADS=1

COPY Project.toml Manifest.toml /srv/study/

# instantiate = resolve + download. precompile = the expensive part: this is
# where the compiler runs and where `compiled/` in the depot gets filled.
RUN julia --project=/srv/study -e 'using Pkg; Pkg.instantiate(); Pkg.precompile()'

# --- R ---------------------------------------------------------------------
# The expensive moment for R is INSTALL, not first call, so the win is
# prebuilt binaries. Posit Package Manager serves them per distribution --
# check the current distro + R-version coverage in their docs, it moves.
RUN apt-get update && apt-get install -y --no-install-recommends \
      r-base-core r-base-dev libopenblas0 && rm -rf /var/lib/apt/lists/*

COPY renv.lock /srv/study/renv.lock
RUN Rscript -e 'install.packages("renv", repos="https://packagemanager.posit.co/cran/latest")' \
 && Rscript -e 'setwd("/srv/study"); renv::restore(prompt = FALSE)'

# Pin BLAS threading for every process in the guest. See the note below on
# why this matters more in an 8-vCPU guest than on your laptop.
RUN printf '%s\n' \
      'OPENBLAS_NUM_THREADS=1' \
      'OMP_NUM_THREADS=1' \
      'MKL_NUM_THREADS=1' >> /usr/lib/R/etc/Renviron.site
DOCKERFILE

# --size-mb: the DEFAULT IS 1024 MB. A populated Julia depot plus an R
#   library will not fit. Size it for the depot you actually built.
# --memory-mb: baked into the snapshot. Firecracker cannot change guest RAM
#   at restore, so this is the only place the number is chosen.
# --cpu: deprecated and ignored. Every template gets 8 burstable vCPUs.
pandastack template build \
  --name julia-r-study \
  -f Dockerfile \
  --context . \
  --size-mb 12288 \
  --memory-mb 8192

What that buys is the first two layers. Once the template's snapshot is baked, creating a sandbox from it is a snapshot restore rather than a boot: p50 179 ms, p99 203 ms, of which the `/snapshot/load` step itself is around 49 ms. The first create of a template before its snapshot exists is a real cold boot at around 3 seconds, and then the snapshot is captured and every create after that is on the fast path. The depot is present, the R library is present, and nobody runs a compiler.

It also buys a pinned environment that is stronger than a lockfile. The image contains the resolved versions, the precompile caches keyed to that exact Julia build, and the BLAS the packages were linked against. A lockfile pins version strings; this pins the compiled artefacts those strings produced on one machine on one day.

Step two: snapshot a session that has already done the work

Baking the template handles the disk. The in-memory layer needs something else, and this is the part worth being precise about, because it is where most descriptions of this pattern get hand-wavy.

A template bake captures the guest shortly after boot, once it is reachable. Whatever your init has managed to do by then is in the image, and whatever it has not is not — which makes it an unreliable place to put a multi-minute warm-up. The reliable path is explicit: create a sandbox, run the warm-up yourself, wait for it to finish, and take a full snapshot. That snapshot is a Firecracker memory image plus the guest's captured disk, so restoring it gives you back the process, with its compiled method instances, its loaded data and its heap, exactly as they were.

from pandastack import Sandbox

# ttl_seconds is an IDLE timeout (default 5 minutes), not a walltime budget --
# the reaper measures time since last activity. A streamed exec counts as
# activity, but give a long warm-up headroom anyway.
sbx = Sandbox.create(template="julia-r-study", ttl_seconds=3600)

# The warm-up has to actually CALL things. `using DataFrames` loads the
# package; it does not force codegen for the method instances your workload
# will hit. One real call per hot path is the whole trick.
WARMUP = r"""
export JULIA_DEPOT_PATH=/opt/julia-depot
julia --project=/srv/study -e '
  using DataFrames, GLM, StatsBase, Plots
  using LinearAlgebra
  BLAS.set_num_threads(1)          # match the guest, not the core count

  df = DataFrame(x = randn(512), y = randn(512))
  fit = lm(@formula(y ~ x), df)    # compiles the GLM path
  coeftable(fit)
  savefig(plot(df.x, df.y), "/tmp/warm.png")   # compiles the plotting path

  println("julia warm")
'
Rscript -e 'library(dplyr); library(ggplot2);
            invisible(summary(lm(mpg ~ wt, mtcars)));
            cat("r warm\n")'
"""

# Long-running work goes through exec_stream, which honours timeout_seconds.
# One-shot exec() has no server-side deadline past the client default -- for
# a long command use exec_stream, or bound it in the shell with `timeout`.
rc = sbx.exec_stream(WARMUP, on_stdout=lambda c: print(c, end=""), timeout_seconds=1200)
assert rc == 0, f"warm-up failed with {rc}"

# snapshot() returns a snapshot ID *string*, and captures memory AND disk.
# This is the artefact that holds the compiled code. Taking it is slow on a
# multi-GiB guest -- that is fine, you take it once.
snap_id = sbx.snapshot()
print("warm snapshot:", snap_id)
sbx.kill()

Now a restore of `snap_id` is a guest whose Julia process has already compiled the GLM and plotting paths and whose R session already has `dplyr` and `ggplot2` attached. The arithmetic has not changed — the compiler did the same work it always did — but it did it once, and the result is now an artefact you can hand to a thousand machines.

Our two fork primitives are not interchangeable here, and the difference is exactly the disk-versus-memory distinction above. `fork()` is disk-only: the parent is paused, its rootfs is cloned, and the child cold-boots from that copy — so the child keeps the populated depot and loses the warm session. `fork_tree(count)` snapshots the parent's memory and state and restores the children from it, so children DO inherit the live process; it is capped at 16 children and they land on the parent's host. For wider fan-out, take the snapshot yourself and create from it, which goes through the scheduler and spreads across hosts.

Pinning BLAS threads, in both languages

# Three places the thread count gets decided, in increasing order of how
# often people forget them. Set all three in the template, not per job.

# 1. Process environment. OpenBLAS built with threading starts a thread per
#    core by default; in an 8-vCPU guest that is 8 threads per process.
cat >> /usr/lib/R/etc/Renviron.site <<'EOF'
OPENBLAS_NUM_THREADS=1
OMP_NUM_THREADS=1
MKL_NUM_THREADS=1
EOF

# 2. R, at runtime, because a package can re-set it underneath you.
cat >> /usr/lib/R/etc/Rprofile.site <<'EOF'
if (requireNamespace("RhpcBLASctl", quietly = TRUE)) {
  RhpcBLASctl::blas_set_num_threads(1)
  RhpcBLASctl::omp_set_num_threads(1)
}
EOF

# 3. Julia, which manages its own BLAS thread pool independently of the
#    environment variables above.
cat >> /opt/julia-depot/config/startup.jl <<'EOF'
try
    using LinearAlgebra
    BLAS.set_num_threads(1)
catch
end
EOF

# Then parallelise at EXACTLY ONE level. Eight workers each spawning eight
# BLAS threads on eight vCPUs is not eight times the throughput; it is a
# scheduler being asked to resolve a disagreement it cannot win.
# Verify what you actually got, don't assume:
#   Rscript -e 'RhpcBLASctl::blas_get_num_procs()'
#   julia -e 'using LinearAlgebra; println(BLAS.get_num_threads())'

Fan-out: the bit where this pays for itself

A parameter sweep or a Monte Carlo study is N jobs that share one expensive setup, which is the shape this whole pattern was built for. The fan-out mechanics — copy-on-write memory, why N restores of one image do not cost N times the fresh work, how the tree is managed — are covered properly elsewhere and I won't re-derive them here. The point specific to scientific runtimes is narrower and sharper: the thing being shared across the fan-out is compiled code. A cold fan-out cannot share that at all. Every worker in a cold fleet runs the same compiler over the same source to produce the same machine code, independently, and then throws it away.

Two ways to get there. For up to 16 children on one host, `fork_tree(count)` does the snapshot for you and restores them in parallel — handy for an interactive best-of-N. For a real sweep, take the snapshot once yourself and create from it as many times as you like: each create goes through the scheduler, so the work spreads across hosts, and the first restore on each host pulls the snapshot from object storage. That pull is why cross-host is 1.2–3.5 s against 400–750 ms for a same-host fork; a wide sweep pays it once per host and then stops paying it.

import json
from pandastack import Sandbox

# One warm snapshot, N restores. Each restore gets the compiled code for
# free; only the parameters differ.
TRIALS = [
    {"sigma": 0.15, "horizon": 252},
    {"sigma": 0.22, "horizon": 252},
    {"sigma": 0.30, "horizon": 504},
]

workers = []
for i, params in enumerate(TRIALS):
    w = Sandbox.create(from_snapshot=snap_id, metadata={"trial": str(i)})

    # ------------------------------------------------------------------
    # THE FOOTNOTE THAT MATTERS MORE THAN ANY LATENCY NUMBER IN THIS POST
    # ------------------------------------------------------------------
    # The kernel protects you. Your own library does not.
    #
    # Kernel-sourced randomness DIVERGES between restored children: the guest
    # CRNG re-mixes RDRAND per extract on amd64, so getrandom(2) and
    # /dev/urandom give each child different bytes. Fine.
    #
    # What does NOT diverge is any generator your warm-up already built. A
    # MersenneTwister constructed during the warm-up, a Random.seed!(1234)
    # in a startup file, numpy's global RandomState -- that internal state is
    # part of the snapshotted heap, byte-identical in every child, continuing
    # the SAME stream from the SAME position. For a Monte Carlo study that is
    # not a performance bug; it is N replicates that look independent and are
    # the same draw. Reseed per child, as the first thing the trial does.
    w.filesystem.write(
        "/srv/trial.json",
        json.dumps({**params, "seed": 20_251_003 + i, "trial": i}),
    )

    w.exec_stream(
        "julia --project=/srv/study /srv/study/trial.jl",
        on_stdout=lambda c, i=i: print(f"[{i}] {c}", end=""),
        timeout_seconds=1800,
    )
    workers.append(w)

for w in workers:
    print(w.filesystem.read("/srv/result.json").decode())
    w.kill()

It is worth being precise about where the duplication actually lives, because the intuitive version of this warning is wrong in a way that will make you distrust the right version. Randomness that comes from the kernel diverges perfectly well across restores: the guest's CRNG re-mixes RDRAND per extract on amd64, so `getrandom(2)` and `/dev/urandom` hand each restored child different bytes. The kernel protects you. Your own library does not. A generator your warm-up already constructed — a `MersenneTwister` built during the warm call, a `Random.seed!(1234)` in a startup file, numpy's global `RandomState` — has its internal state sitting in the snapshotted heap, byte-identical in every child, continuing the same stream from the same position.

Which gives a rule with two acceptable halves: either the warm-up constructs no RNG at all before the freeze, or every child reseeds from a per-child source as its first action after restore. In practice do both, and be thorough about the second, because "I called the seed function" and "every random stream in this process is now distinct" are not the same statement:

  • In Julia, read the config and call `Random.seed!(cfg.seed)` before any sampling. That reseeds the default global RNG. If you would rather not thread a config file through, `Random.seed!(rand(RandomDevice(), UInt64))` draws from the kernel, which does diverge per child — at the cost of a run you cannot replay.
  • In R, `set.seed(cfg$seed)` immediately after `jsonlite::fromJSON("/srv/trial.json")`. Same position, same reason.
  • Seeding the global RNG does not reach a task-local RNG your code constructed before the snapshot was taken. If the warm-up built a sampler object, reseed that object by name.
  • Nor does it reach RNG state living inside a C or Fortran library the warm session already initialised. Those have their own seeding entry points and their own silence when you skip them.
  • Deriving the seed from the trial index rather than from fresh entropy keeps the study reproducible, which is the whole reason you are doing any of this.
  • The failure mode for every one of these is identical and it is not an error message. It is a study that ran a thousand times and agreed with itself perfectly.

There is a genuine reproducibility argument buried in here that I think is underrated. A pinned snapshot is a stronger reproducibility artefact for a numerical result than a lockfile, because it pins the compiled code and the BLAS rather than the version strings that were supposed to imply them. If a reviewer asks whether your figure reproduces, "here is the manifest" asks them to rebuild your environment and hope; "here is the memory image the figure came out of" does not.

Four ways to not pay for the first call

Approaches to the compile-on-first-use cost, compared on the four axes that actually trade off against each other.
ApproachFirst-call latencyReproducibilityChanging a packageStorage cost
Cold container, install at startWorst: install plus compile on every runLockfile only; the build host still variesTrivial: edit the lockfileLowest: just the image
Container plus cached volume for the depot or libraryBetter: disk cache hits, in-process codegen still paidWeak: cache contents drift and nobody audits themEasy: next run repopulatesModerate: one shared volume
Build-time sysimage, or prebuilt binaries for RGood: compiled ahead of time, process still starts emptyStrong: pinned at build, CPU target is explicitRebuild the artefactModerate: fat image
VM snapshot of a warm sessionBest: compiled code is already resident on restoreStrongest: pins compiled code, BLAS and heap stateRe-bake and re-warm; not a runtime decisionHighest: grows with the warm RAM you captured

These stack rather than compete. The honest configuration for a serious workload is a baked template with prebuilt binaries and a populated depot, plus a warm snapshot on top of it for the hot path. The sysimage row is still the right answer when the artefact has to run somewhere you do not control.

What this does not fix

Everything above is the good news, and a post that stopped there would be an advert. The constraints are real and several of them are structural rather than roadmap items.

  • A package you did not bake costs full price. The user who adds one `using SomeOtherPackage` to their script pays the install, the precompile and the codegen at runtime, inside a guest, on your clock. The warm snapshot does nothing for them and they will reasonably report the platform as slow.
  • The environment is pinned, and that is a real loss. "Use the latest version of the package" becomes a re-bake, not a runtime decision. For production batch work that is a feature. For exploratory analysis — where the whole activity is trying things that were not in the manifest an hour ago — it is friction you have deliberately added, and you should be honest with those users about which mode they are in.
  • RAM is fixed by the baked snapshot. Firecracker cannot change vCPU or RAM at snapshot restore, so `cpu=` and `memory_mb=` on a create are ignored whenever the template has a baked snapshot; the agent overrides them to the baked values. A study that wants 32 GiB needs a template baked at 32 GiB via `--memory-mb`. There is no argument you can pass at create time, and a large in-memory dataset is exactly the workload that most wants to ask.
  • There is no GPU. That rules out a meaningful slice of numerical work outright — anything built on CUDA, most modern deep learning, a good deal of current computational imaging. If that is your workload, this architecture is not a compromise you can engineer around; it is the wrong platform.
  • A fat warm image is not free. Snapshot size grows with the memory you captured, which is storage you pay for and bytes that have to be present on the host doing the restore. Cross-host is where you feel it: a same-host fork runs 400–750 ms while cross-host runs 1.2–3.5 s, and the difference is fetching artefacts over the network. Warming 6 GiB of data into a session because it made the demo look good is a decision with a recurring bill.
  • Warm memory is billed memory. Committed GiB-hours are charged while a guest is live at $0.0162 per GiB-hour, so a fat warm guest that sits idle costs more than a thin one that sits idle. The win here is on the CPU side — compile once instead of N times — not a free lunch on RAM.
  • The two fork primitives behave differently and the docs are the only place that says so. `fork()` carries disk and not memory; `fork_tree()` carries both but caps at 16 children on one host. If you assumed `fork()` gave you the warm session, your sweep is quietly paying the compile N times and still looks like it is working.

When to reach for this, and when not to

Reach for it when the environment is stable and the fan-out is wide: a nightly sweep, a backtest grid, a simulation study, a scoring job that runs the same model over fresh inputs, an agent that runs analysis in a known toolchain. Those are workloads where the environment being pinned is the point and the first call happens thousands of times.

Do not reach for it when the user's job is to install things you have not heard of. An interactive researcher poking at new packages is a workload whose cost is dominated by exactly the thing a baked image cannot pre-pay, and a long-lived machine with a persistent depot serves them better. If the question you are really asking is how to host notebooks for a group of humans, that is a different post and a different architecture — the multi-tenant kernel question is about isolation and per-user state, not about compilation.

You have not made the compiler faster. You have made it run once, somewhere you were not being timed, and turned its output into a file.

Which, when the first call is the whole bill, is the only optimisation that was ever available.

Frequently asked questions

Does a snapshot really preserve Julia's JIT-compiled code, or just the precompile cache on disk?

Both, and the distinction is the entire point. A full snapshot captures the guest's memory image alongside its disk, so restoring it brings back the live Julia process exactly as it was — including every method instance the JIT compiled during the warm-up, the loaded package code, and whatever data structures the session had built. The depot's on-disk precompile cache is preserved too, but that part a container volume could also have done. What only a memory image can hold is the state of a running process, and for a runtime that compiles on first call that state is the expensive thing. The practical consequence is that your warm-up has to actually call the functions you care about, not merely import them: `using DataFrames` loads the package, but the method instances for your specific argument types are compiled when you first call with those types. One representative call per hot path is what converts an import into resident machine code.

Why can't I just use a cached volume for the Julia depot or the renv cache?

You can, and you should — it is a real improvement over installing from scratch, and for R it may be most of the win, because R's expensive moment is installation rather than first call. But it stops at the filesystem. A warm cache means the next process does not have to regenerate the cache; it does not mean the next process starts warm. For Julia in particular, even with a fully populated depot the new process still has to load the package images, link them, and compile whatever your call sites need that the cache did not cover — all of which is per-process cost that recurs on every job. A cached volume also brings its own problems at scale: it is mutable shared state that several jobs write to concurrently, it drifts from whatever you think is in it, and nobody audits it. A baked block device you clone read-mostly, plus a memory image for the process state, is both faster and easier to reason about.

I need 32 GiB of RAM for a large in-memory dataset. Can I pass memory_mb on create?

No, and this is the limitation of our architecture that most often surprises people doing scientific work. Firecracker cannot change a guest's vCPU count or RAM at snapshot restore, so whenever a template has a baked snapshot the agent overrides the create request's `cpu` and `memory_mb` to the baked values. Passing `memory_mb=32768` to a create is not an error; it is silently ignored, which is worse. The place the number is chosen is the template build, with `--memory-mb`, and changing it means baking a new template. In practice that means maintaining a small family of templates at the sizes your workloads actually need rather than treating RAM as a per-job parameter. It is a genuine loss of flexibility, and it is most painful for exactly the data-heavy analysis workloads that would most like to ask for more memory at submit time. CPU is less of an issue: every template runs 8 burstable vCPUs, shared fairly under contention, and the `--cpu` flag is deprecated and ignored.

What is the single most dangerous mistake when forking a warm scientific session?

Not reseeding a random number generator that the warm-up had already constructed. Be precise about where the duplication lives, though, because the intuitive version of this warning is wrong: kernel-sourced randomness diverges across restores just fine, since the guest CRNG re-mixes RDRAND per extract on amd64, so `getrandom(2)` and `/dev/urandom` give each child different bytes. The kernel protects you; your own library does not. A generator built during the warm-up — a MersenneTwister instantiated in your warm call, a seed set in a startup file, numpy's global RandomState — has its internal state in the snapshotted heap, identical in every child, resuming the same stream from the same position. Fan out a thousand workers and you get a thousand runs of the same trial, with implausibly tight error bars that will survive review for longer than they should. The rule: either construct no RNG before the freeze, or reseed from a per-child source as the first action after restore. Derive the seed from the trial index so the study stays replayable, and reseed task-local generators and C-library state by name too.

Why does my R job get slower when I add more parallel workers inside one sandbox?

Almost always BLAS thread oversubscription. OpenBLAS built with threading starts one thread per core by default, and our guests run 8 burstable vCPUs, so a single R process will typically claim 8 BLAS threads. Add job-level parallelism on top — four or eight workers in the same guest — and you are asking for 32 or 64 threads of linear algebra on 8 cores. The threads spend their time contending and migrating rather than computing, and throughput goes down as you add workers, which is the most confusing shape a performance problem can have. The fix is to pin the BLAS thread count to 1 in the template environment and parallelise at exactly one level. Set `OPENBLAS_NUM_THREADS` and `OMP_NUM_THREADS` in `Renviron.site`, call `RhpcBLASctl::blas_set_num_threads(1)` from `Rprofile.site` because packages can change it underneath you, and set `BLAS.set_num_threads(1)` for Julia, which manages its own pool independently of those variables. Then verify what you actually got rather than assuming.

Keep reading

Related posts

More in Snapshots & forking · See Thaw: sub-second cold restore

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.