Cold Start Is Four Problems: JVM Warmup vs the microVM Snapshot
Somebody files a ticket that says "cold start is too slow". Four engineers read it and each of them is thinking about a different problem. One is thinking about the machine not existing yet. One is thinking about Linux booting. One is thinking about the Spring context taking eleven seconds to assemble itself behind an ASCII banner that has become, functionally, a loading screen. And one is thinking about the fact that requests one through four hundred are honestly slow because the JIT has not made up its mind yet.
They are all correct, and they are all describing different layers of the same stack. The useful observation is that each layer has its own snapshot technology, invented independently by people who mostly did not talk to each other, and that the technologies are not substitutes. You can stack them. You can also pick exactly the wrong one and spend a quarter on it.
Cold start is four problems wearing one name
Before anything else, separate the layers. The whole post is an argument about which layer you are actually paying for.
- The machine does not exist yet. Something has to allocate and schedule a container or a VM, pull or clone a filesystem, wire up networking. Measured in hundreds of milliseconds to tens of seconds depending on how much pulling is involved.
- The kernel has not booted. Firmware, kernel decompression, device probing, init, userspace. A general-purpose VM spends a lot of time here; a microVM with a minimal device model spends much less, but it is not free.
- The runtime has not initialised. For the JVM: loading and verifying thousands of classes, running static initialisers, and — in a framework app — reflective component scanning and bean graph construction. This is startup.
- The code is still interpreted. Bytecode runs in the interpreter, then in C1-compiled form, while the JVM collects profiles until C2 decides it is worth generating good machine code. This is warmup, and it is a completely different cost from startup.
Here is the punchline, stated up front so the rest of the post can earn it: a Firecracker snapshot obviously fixes layers 1 and 2, because it replaces "build a machine and boot it" with "map a memory image and resume". What is less obvious is that it also fixes layers 3 and 4, for free, if you take the snapshot at the right moment. A snapshot is a byte-for-byte image of guest RAM. A loaded class lives in RAM. A C2-compiled method lives in RAM. A warm heap with its generational structure already shaped by real traffic lives in RAM. Freeze the machine after warmup and you have frozen all of it.
Startup and warmup are different costs, and nearly everyone conflates them
The JVM is slow to start for reasons that are specific and unglamorous. Every class your application touches has to be found on the classpath, parsed from a class file into HotSpot's internal metadata representation, and verified — the bytecode verifier proves type safety for every method before it is allowed to run. For a trivial program this is tens of classes. For a framework application it is thousands, often ten thousand or more, and the cost is roughly linear in the count.
Then there is the framework. The canonical offender is a classpath-scanning component model: at startup, walk the classpath, read annotations reflectively, build a bean dependency graph, instantiate it in topological order, run post-processors. None of that work is wasted — it is why the programming model is pleasant — but all of it happens at runtime, every single time the process starts, and it produces exactly the same answer every single time. Which is a very loud hint about where the fix belongs.
Warmup is separate and arrives after startup is done. Your service is up, the health check is green, and it is running interpreted or C1-compiled code. The JVM is counting method invocations and loop back-edges, recording which branches are taken and which types actually show up at each call site, and only once a method is hot enough does C2 compile it — using those profiles to inline aggressively, devirtualise calls that are monomorphic in practice, and eliminate bounds checks it can prove are redundant. The result is genuinely excellent code. Getting there takes real traffic and real time.
So the JVM spends twenty seconds being embarrassing and then produces code that can beat the equivalent C, because it had profiling data the C compiler never saw. It is a legitimately great trade and absolutely nobody signed up for it. Everybody signed up for the second half.
AppCDS: the cheapest win on the list, and it does nothing for the JIT
Class Data Sharing attacks layer 3 and only layer 3. The idea is to do the parsing and verification once, write the resulting internal metadata into an archive file, and on subsequent runs memory-map that archive instead of redoing the work. Because it is mapped rather than copied, the pages can be shared read-only across several JVM processes on the same host, which is a nice secondary effect if you run many JVMs per machine.
The mechanics have moved around across releases, so check the flags against the JDK you are actually on. The dynamic-archive form writes the archive when the training run exits. Later JDKs added a flag that creates and refreshes the archive automatically on first run, which removes the two-command dance from your Dockerfile. Project Leyden's AOT cache is the current direction: it stores classes in an already-loaded-and-linked state, and newer JDKs fold method profiles from the training run into the same cache.
#!/usr/bin/env bash
# Layer 3 only: skip re-parsing and re-verifying classes on every start.
# Flags move between JDK releases -- verify against `java -X` / the java(1)
# man page for YOUR JDK before you ship this.
set -euo pipefail
# --- Classic AppCDS: dynamic archive written at JVM exit -------------------
# Training run. Exercise the startup path, then exit cleanly -- the archive is
# written on exit, so a SIGKILL here gives you nothing.
java -XX:ArchiveClassesAtExit=app.jsa \
-Dspring.main.lazy-initialization=false \
-jar app.jar --exit-after-startup
# Production run: map the archive instead of parsing class files.
java -XX:SharedArchiveFile=app.jsa -jar app.jar
# Same thing without the two-step: create it on first run, refresh it when the
# classpath changes. Added in a later JDK than the flags above.
java -XX:+AutoCreateSharedArchive -XX:SharedArchiveFile=app.jsa -jar app.jar
# --- Leyden AOT cache: classes stored already loaded AND linked ------------
# Two-step form (JEP 483, JDK 24):
java -XX:AOTMode=record -XX:AOTConfiguration=app.aotconf -jar app.jar
java -XX:AOTMode=create -XX:AOTConfiguration=app.aotconf -XX:AOTCache=app.aot
# One-step form (JEP 514, JDK 25) -- the launcher splits the invocation itself:
java -XX:AOTCacheOutput=app.aot -jar app.jar
# Production, either way:
java -XX:AOTCache=app.aot -jar app.jar
# What you did NOT get from any of the above: a warm JIT. Request one still
# runs interpreted. Layer 4 is untouched.
Be clear about the boundary. A class-data archive contains class metadata, not compiled code. Your first request after a CDS-accelerated start is interpreted exactly as before. CDS makes the process reach "ready" sooner; it does not make the first few hundred requests faster. It is still the first thing I would turn on, because the change is a flag and a build step rather than an architecture.
Leyden, and the direction of travel
Project Leyden is the OpenJDK effort to shift work earlier in time — out of the production run and into a training run or a build. The AOT cache landed as ahead-of-time class loading and linking, then gained single-command ergonomics and ahead-of-time method profiling, which means a production start can begin with prior knowledge of what gets hot rather than rediscovering it. Caching compiled code itself is the obvious next frontier and the part I would not make promises about.
Treat the specifics here as perishable. Check the JEP status for the JDK you are targeting rather than trusting any blog post, including this one, about what shipped where.
CRaC: the word doing the work is "coordinated"
Coordinated Restore at Checkpoint is the OpenJDK project that attacks layers 3 and 4 together. You run the JVM, let it start, let it warm up under load, then checkpoint the whole process to disk — on Linux, via CRIU underneath. Restoring that checkpoint gives you a JVM that is already initialised and already warm. It is the closest thing in the Java ecosystem to what a VM snapshot does, built one layer up.
The coordination is the interesting part. A running process holds things that cannot be frozen and thawed: open sockets whose peers will be gone, file descriptors pointing at files that may have moved, timers and scheduled executors that will fire late by however long the checkpoint lasted, cached DNS answers, connection pools. So CRaC defines an API — a resource interface with a before-checkpoint and an after-restore callback — and asks the application and its libraries to register, release what cannot survive, and reacquire it on the way back. Frameworks have grown support for this; recent Spring Boot releases ship it, and library support is the thing that determines whether this is a weekend or a quarter.
That requirement is CRaC's cost and also its honesty. It forces you to enumerate every piece of external state your process holds, written down as code that runs at a known moment. Which is precisely the work a correct VM-level snapshot restore silently requires of you too — the difference is that at the VM layer nobody hands you an interface and a compiler error. Nobody asks. The machine simply resumes and your pool starts handing out dead sockets.
Two practical notes. CRaC is not in every mainline JDK build; it has historically arrived through vendor distributions, so your base image matters. And because the mechanism is CRIU, it wants privileges that a locked-down container runtime may not grant you — which is a sentence that has ended more CRaC evaluations than any technical limitation of the API. Verify both against current upstream and vendor documentation before you plan around them.
GraalVM native-image: pay once at build time, pay forever at the margins
The other branch is to stop having a warmup curve at all. Ahead-of-time compilation to a native binary does closed-world analysis over your whole program, compiles it, and emits an executable that starts in milliseconds with a tiny resident set. Layers 3 and 4 largely vanish: there is no class loading to speak of and no JIT to warm, because there is no JIT.
What you give up is real. Dynamic class loading, reflection that is not declared in configuration, dynamic proxies, resource loading by computed name — all of it needs either metadata you supply or a framework that generates that metadata for you, and the failure mode is a runtime error on a code path you did not exercise in testing. The build is slow and memory-hungry enough to be its own capacity planning problem. And for long-running throughput-bound services, peak performance can land below a warmed C2, because C2 had profile data from your actual production traffic and the AOT compiler had a static analysis. This varies enormously by workload and by Graal version, so measure your workload rather than believing anyone's benchmark.
Native-image is not a free win and the people who built it have never claimed it was. It is a trade: build-time cost and dynamism for startup latency and memory footprint. For a function whose whole life is a couple of hundred milliseconds it is close to unarguable. For a service that runs for three weeks it is a question.
The VM-level snapshot: a warm JIT is just pages
Now the hypervisor layer. Firecracker can pause a microVM, write its guest memory, vCPU registers and device state to disk, and later restore that state into a fresh VMM process and execute the next instruction. On PandaStack every create is a restore of a baked template snapshot — there is no warm pool of idle machines — and the end-to-end create sits at a p50 of 179 ms and a p99 of 203 ms, of which the snapshot-load step itself is around 49 ms. The first cold boot of a template, before its snapshot has been baked, takes roughly 3 seconds; after that, nobody ever pays for it again. A same-host fork of an existing machine lands in 400 to 750 ms.
The part that matters for Java is what the memory image contains. Snapshot a JVM that has already loaded its classes, built its bean graph, and been driven hard enough that C2 has compiled the hot paths, and all three of those things are in the image. Restore it and the guest does not know it was paused. There is no participation required from the application, no API to implement, no library compatibility matrix — because the freeze happens below the abstraction the application can perceive.
#!/usr/bin/env bash
# Inside the guest: get the JVM genuinely warm, THEN let the host snapshot it.
# The goal is a memory image that already contains C2-compiled hot paths, not
# just an initialised Spring context.
set -euo pipefail
export MISE_DATA_DIR=/opt/mise MISE_CONFIG_DIR=/opt/mise
export PATH=/opt/mise/shims:$PATH
# Heap sizing note: the baked template fixes guest RAM, so pick -Xmx against
# the template's memory, not against whatever the host has.
nohup java -XX:SharedArchiveFile=app.jsa \
-Xms1g -Xmx2g \
-XX:+PrintCompilation \
-jar /srv/app.jar > /var/log/app.log 2>&1 &
# Wait for layer 3 to finish (startup), which is NOT the same as warm.
for _ in $(seq 1 60); do
curl -fsS http://127.0.0.1:8080/actuator/health >/dev/null && break
sleep 1
done
# Layer 4: drive synthetic traffic over the paths that matter until the JIT
# settles. "Settles" is observable -- compilation events taper off. Hit the
# real endpoints with representative payloads; warming /health warms /health.
for endpoint in /api/orders /api/search /api/quote; do
hey -z 45s -c 16 "http://127.0.0.1:8080${endpoint}" || true
done
# Rough settle check: compilation events in the last 10s of the log.
tail -c 200000 /var/log/app.log | grep -c ' made not entrant\| 4 *java' || true
# Quiesce before the freeze. Anything the guest is holding that the restored
# guest cannot keep holding is better released here than discovered later.
curl -fsS -XPOST http://127.0.0.1:8080/actuator/shutdown-pools || true
sync
# The host takes it from here: pause, write vm.mem + vm.state, resume.
From the SDK side the flow is create, warm, snapshot, then create as many machines as you like from that snapshot. Note the one-shot exec caveat in the code: a long warmup needs a streaming exec or a shell-level timeout, because a plain one-shot exec is bounded by the client's own default and will abandon you at thirty seconds.
from pandastack import Sandbox
# 1. Build the warm machine once.
warm = Sandbox.create(template="base", metadata={"role": "jvm-warm-bake"})
warm.filesystem.upload("./target/app.jar", "/srv/app.jar")
warm.filesystem.upload("./warm.sh", "/srv/warm.sh")
# The warmup takes minutes. A one-shot exec() cannot be given a deadline past
# ~30s today, so stream it (exec_stream DOES honour timeout_seconds) and bound
# the command in the shell as well.
exit_code = warm.exec_stream(
"timeout 900 bash /srv/warm.sh",
on_stdout=lambda chunk: print(chunk, end=""),
timeout_seconds=960,
)
if exit_code != 0:
raise SystemExit(f"warmup failed: exit {exit_code}")
# 2. Freeze it. Guest RAM now holds loaded classes, the built bean graph, a
# shaped heap, and C2-compiled hot methods. snapshot() returns an ID STRING.
snapshot_id = warm.snapshot()
warm.kill()
# 3. Every later machine restores that image instead of starting a JVM.
# cpu/memory_mb are deliberately omitted: Firecracker cannot change vCPU or
# RAM at restore, so the baked template's size (base = 4 GiB / 8 vCPU) wins
# and anything you pass here is overridden.
replicas = [
Sandbox.create(from_snapshot=snapshot_id, metadata={"replica": str(i)})
for i in range(8)
]
for sbx in replicas:
# The guest resumed mid-instruction. Before trusting it, repair the state
# that could not survive the freeze -- see the next section for why.
sbx.exec("curl -fsS -XPOST http://127.0.0.1:8080/actuator/pools/refresh")
print(sbx.id, sbx.preview_url(8080))
What a restored guest wakes up believing
This is the section that decides whether you should trust the rest of the post. A restored guest is not a fresh machine and it is not the machine you froze. It is the machine you froze, dropped into a world that has moved on, and it has no idea.
The clock is wrong. Firecracker resumes with the guest's notion of wall-clock time frozen at the instant of the snapshot, and nothing in the guest corrects it on its own — these microVMs have no RTC on x86 and the templates do not run NTP. The consequence is not abstract. We shipped this bug: the base template's seed was baked on 2026-06-22, github.com rotated its leaf certificate on 2026-07-03, and from then on every app deploy's git clone failed with a certificate verification error, because a guest living in June correctly concluded that a July certificate was not yet valid. The fix is a clock re-sync on every restore-family path — snapshot-restore create, resume, and wake — and it is best-effort by design: a failed sync logs loudly rather than failing the boot, because a stale clock is degraded service and a failed create is an outage. For a JVM this matters doubly, since JWT expiry checks, TLS validation and anything built on currentTimeMillis all read the same lie.
The sockets are dead. Every established TCP connection in the frozen image has a peer that stopped waiting a long time ago. Listening sockets restore fine. Established client connections do not, and the place this bites hardest is the connection pool, because a pool is a cache of exactly the thing that did not survive. If your pool does not validate on borrow — a test query, or at minimum an aggressive idle-eviction policy — then the first request after restore will be served a socket that is already a corpse, and the error it produces will not mention snapshots anywhere.
The cached lookups are from another era. DNS answers past their TTL, a service-discovery registry snapshot naming instances that have since been replaced, a cached OIDC discovery document, a leader election result. All of it restores confidently wrong.
And randomness deserves a precise statement rather than a scary one. The kernel's own CRNG re-mixes a hardware entropy source per extraction on amd64, so kernel-sourced randomness diverges between restores. What does not diverge is anything your application already seeded into its own heap before the freeze: a Random instance constructed at startup, a session ID generator holding its own state, a cached UUID namespace. Fork the same snapshot a hundred times and all hundred share that state exactly. If uniqueness matters, re-seed it after restore, deliberately, in code you wrote.
One more constraint that catches Java teams specifically: Firecracker cannot change vCPU count or RAM at snapshot restore. On PandaStack that means a create against a template with a baked snapshot has its cpu and memory_mb overridden to the baked values — base is 4 GiB and 8 vCPU, the agent and code-interpreter templates 2 GiB, postgres-16 1 GiB. The JVM sizes its heap ergonomically from available memory at startup, so if you snapshot a warm JVM, the heap geometry it chose was chosen against the template's RAM and is now frozen into the image. Set -Xmx explicitly against the template size rather than letting ergonomics guess, and do not write a snippet that asks for 16 GiB, because it will be quietly ignored.
The four technologies, side by side
| Technology | Layers it fixes | App changes needed | Build-time cost | Peak throughput | Other languages |
|---|---|---|---|---|---|
| AppCDS / CDS archive | Startup only (3) | None | A training run per build | Unchanged | Java only |
| Leyden AOT cache | Startup (3), some warmup | None | A training run per build | Unchanged or better | Java only |
| CRaC | Startup and warmup (3-4) | Yes: release and reacquire resources | A checkpoint step, plus CRIU privileges | Warm from the first request | Java only |
| GraalVM native-image | Startup and warmup (3-4) | Yes: reflection config, no dynamic loading | Slow, memory-hungry build | Often below a warmed C2 | A Java toolchain |
| microVM snapshot | All four (1-4) | None in-process; repair state after restore | One template bake, reused forever | Unchanged; warm if you bake warm | Any runtime, no vendor support needed |
The rows are not mutually exclusive and the best configuration is usually a stack. Turn on a class-data archive because it is a flag. Snapshot the machine after warmup because it is a build step. Reach for CRaC when you need warm restore inside a container platform you do not control the hypervisor of. Reach for native-image when startup latency and memory footprint dominate and you can live inside the closed world.
Every runtime has these four layers, and most have fewer tools
Java gets talked about most because its layer-3 and layer-4 costs are the most visible, but the stack is universal. .NET has essentially the same shape: ReadyToRun precompiles IL to native code to cut JIT work at startup, tiered compilation means early code is quick-to-produce and slow-to-run before it is re-jitted hot, and native AOT is the closed-world branch. Node has V8 startup snapshots, which serialise an initialised heap so a process can skip parsing and executing initialisation script — the same idea as a CDS archive, implemented for a different runtime.
Python's version is less celebrated and more annoying. There is barely a JIT to warm, so layer 4 mostly does not exist, but layer 3 is import time: a large site-packages tree is a filesystem walk, a pile of stat calls, bytecode compilation on first import, and module-level side effects. A snapshot taken after imports is, functionally, a cache for all of that — and unlike the Java options, you did not need the runtime's vendor to ship you a feature to get it.
That is the real argument for the VM layer. It is the language-agnostic version of every per-runtime snapshot feature, and it works for the runtime whose maintainers never built one: an old Ruby service, a JRuby monolith, a Scala build, a numerical stack whose startup cost is three gigabytes of shared libraries being mapped and relocated. Nobody will ever ship a checkpoint API for your particular pile. The hypervisor does not need one.
Which one you actually need
- You control the hypervisor and want warm starts for anything, in any language, with no application changes — snapshot at the VM layer, after warmup. This is the broadest tool and the only one that also fixes layers 1 and 2.
- You are on a container platform somebody else runs, and the app is Java — CRaC or a Leyden AOT cache, depending on whether you need warm or merely fast-starting. Check library support before committing.
- Startup latency and resident memory dominate, the app is well-behaved about reflection, and peak throughput is not the binding constraint — native-image.
- You have not measured which layer you are paying for — do that first. Half the "JVM cold start" tickets I have seen were layer 1 in disguise: a multi-gigabyte container image being pulled on a cache miss while everybody blamed Spring.
The summary
Cold start is four stacked problems and the industry has given each of them a different snapshot. AppCDS and the Leyden AOT cache move class loading and linking out of the production run. CRaC checkpoints a warm process and asks the application to be honest about its external state. Native-image deletes layers 3 and 4 by deleting the dynamism they serve. A microVM snapshot sits underneath all of them, restores guest memory, and therefore gets the warm JIT thrown in — p50 179 ms per create on our fleet, with the restore step itself around 49 ms, and no cooperation from the process required.
What none of them delete is the thinking. Something in your process is holding a socket, a timer, a cached lookup or a seeded generator that cannot survive being frozen. CRaC makes you write that down. The hypervisor lets you skip it and find out on the first request after restore instead, which is the more expensive way to learn the same list.
Frequently asked questions
Does a Firecracker snapshot really preserve the JIT-compiled code, or just the loaded classes?
Both, because it preserves memory and both live in memory. A Firecracker snapshot writes guest RAM, vCPU registers and device state; restoring maps that memory back and executes the next instruction. HotSpot's code cache — the region holding C1- and C2-compiled methods — is ordinary process memory inside that guest, as are the loaded class metadata, the profile counters, and the heap in whatever generational shape real traffic gave it. So a snapshot taken after warmup restores a JVM that is already past both startup and warmup. The caveat is that you have to actually warm it before freezing, which means driving representative traffic at the real endpoints rather than hitting a health check in a loop. A snapshot taken the moment the health check turns green captures an initialised but interpreted JVM, which fixes layer 3 and leaves layer 4 exactly where it was. Compilation events tapering off in the JVM's own compilation log is a reasonable, if crude, signal that the profile has settled.
If a VM snapshot does everything CRaC does, why would anyone use CRaC?
Mainly because you often do not control the hypervisor. CRaC works inside a container on a platform somebody else operates, where you have no ability to pause a VM and write its memory to object storage. That is the common case for most teams, and it is a perfectly good reason. There is a second, subtler reason: CRaC's coordination API is a forcing function. It makes the application declare, in code, what it releases before a checkpoint and reacquires after a restore, and the library ecosystem has largely done that work for the obvious resources. At the VM layer the same repairs are necessary and nobody reminds you — the machine simply resumes and a stale connection pool starts handing out dead sockets. Against that, CRaC costs you: it is not in every mainline JDK build, it depends on CRIU and therefore on privileges a hardened container runtime may refuse, and library coverage is uneven outside the mainstream frameworks. Verify all three against current upstream and vendor documentation.
What is the difference between JVM startup and JVM warmup?
Startup is the work before your code can serve anything: finding classes on the classpath, parsing each class file into the VM's internal metadata, running the bytecode verifier over every method, executing static initialisers, and — in a framework application — scanning the classpath reflectively and constructing the bean dependency graph. It ends when the process is ready. Warmup happens afterwards, while the process is already serving traffic. Bytecode starts in the interpreter, moves to quickly-produced C1-compiled code, and only becomes good machine code once the JVM has collected enough profiling data for C2 to compile it with aggressive inlining and devirtualisation based on the types that actually appear at each call site. Startup is fixed by caching class metadata, which is what AppCDS and the Leyden AOT cache do. Warmup is fixed only by preserving the compiled state, which is what CRaC and a VM snapshot do, or by eliminating the JIT, which is what native-image does. Conflating them is why teams turn on CDS, see no change in first-request latency, and conclude it did not work.
What breaks when you restore a snapshot of a running JVM?
Four things, reliably. The wall clock is frozen at the snapshot instant, so TLS certificate validity windows, token expiry and anything reading the current time are wrong until something corrects it — on PandaStack the agent re-syncs the guest clock on every restore, resume and wake, and the failure mode before that existed was a guest rejecting a certificate as not-yet-valid because the certificate was issued after the snapshot was baked. Established TCP connections are dead, which means connection pools are full of corpses unless they validate on borrow. Cached lookups are stale: DNS past its TTL, service-discovery registries, discovery documents. And any randomness your application seeded into its own heap before the freeze is identical in every machine restored from that image — the kernel's CRNG re-mixes hardware entropy per extraction so kernel-sourced randomness diverges, but a Random instance your code constructed at startup does not. Each of these is fixable in a few lines; none of them announces itself, which is the actual hazard.
Can I give a snapshotted JVM more heap by requesting more memory at create time?
No, and this surprises people. Firecracker cannot change the vCPU count or the RAM of a microVM at snapshot restore — the memory image and the machine configuration are the same artefact. On PandaStack that means cpu and memory_mb on a create are overridden to the baked values whenever the template has a baked snapshot: base is 4 GiB and 8 vCPU, the agent and code-interpreter templates 2 GiB, browser 4 GiB, postgres-16 1 GiB. The practical consequence for Java is that heap ergonomics ran against the template's memory at the moment you started the JVM, and if you then snapshot the warm process, that geometry is frozen into the image. Set -Xmx explicitly against the template's RAM rather than relying on ergonomics, leave real headroom for metaspace, code cache, thread stacks and direct buffers, and if you genuinely need a bigger machine, that is a different baked template rather than a different create parameter.
Keep reading
- Snapshot-restore vs AWS Lambda SnapStart — The managed version of the same idea, pitched at the function environment rather than the whole machine.
- Julia, R and the cost of the first import — The same argument for a runtime that compiles on first use instead of warming a JIT — where the cache lives on disk as well as in memory.
- Snapshot and fork, explained — How the restore path actually works, if you want the mechanism rather than the Java-shaped argument.
- Guest clocks and time drift after restore — The full version of the frozen-clock problem, including the expired-certificate incident.
- What happens to network connections across a restore — Why your connection pool is the first thing to break, and what validate-on-borrow is really for.
- Where to run Spring Boot in 2026 — The hosting-shaped version of this question, once you have decided which layer you are paying for.
- PandaStack sandboxes — The create, warm, snapshot, restore primitive this post is built on, with the API surface.
Related posts
- How to Optimize MicroVM Cold Start
A cold start is a tax you pay on every create. The biggest cut isn't a faster boot — it's not booting at all. Here are six techniques to beat the microVM cold-start tax, ranked by impact, with the why behind each.
- How to Benchmark Sandbox Cold Start Honestly
Cold-start numbers are marketing until you define the stopwatch. Here's a rigorous method — where the clock starts and stops, which stages to split out, why p99 is where your users live, and a Python harness you can point at any vendor including mine.
- Snapshot Restore vs Container Image Pull: Two Ways to Start Fast
Your container started in 200ms and then spent nine seconds importing pandas, which is a kind of performance. Lazy pulling attacks the bytes; snapshot restore attacks the part that actually costs you — a process that has already finished initializing.
- Golden Images vs Snapshot Baking
A golden image removes install time. A snapshot removes boot time. Those are different costs, which is why the answer is almost always both — and why the snapshot quietly freezes your RNG, your clock, and anything that was in RAM at bake time.
- Thaw: How We Made a Cold Start Take 164ms
Vercel Fluid pre-warms and shares instances so you rarely hit a cold start. We took the other path: scale fully to zero, then thaw a frozen microVM in ~164ms. Here's how.
More in Snapshots & forking · See Thaw: sub-second cold restore
49ms p50 cold start. Fork, snapshot, and scale to zero.