all posts

One microVM per Attendee: Hands-On Labs for a Whole Room at Once

Ajay Kumar··11 min read

Twenty minutes into every hands-on workshop ever held, the instructor stops and asks the question: "has everyone got a terminal?" Nobody answers, because the people who have one have no reason to speak and the people who do not are busy reading an error message. So the instructor asks the other question — "hands up if you are NOT in yet" — and somewhere between four and forty hands go up, and the next twenty minutes of a ninety-minute slot are spent doing laptop support in front of a room.

This is an infrastructure problem wearing a social costume. A training lab has a load shape that no other workload has, and if you provision it the way you provision anything else, you get one of two bad afternoons.

A cohort is not a queue

Almost every workload a platform is designed for is a queue. Requests arrive according to some distribution, you size for the ninety-ninth percentile, and the arrival of request 200 is statistically independent of the arrival of request 199. CI is a queue. Agent sandboxes are a queue. Even a spiky consumer app is, in the end, a lot of unrelated people deciding to do a thing.

A cohort is not a queue. It is a step function with a human in front of it saying "right, everybody open the link". Four hundred attendees at a conference tutorial are perfectly correlated: they arrive inside the same ninety seconds, they are all on step 3 at the same time, they all hit the database section after the coffee break, and at 5pm the room empties and every single environment becomes garbage simultaneously. There is no smoothing. There is one spike per session, and the whole session is downstream of it.

Two failure modes follow from that, and most lab infrastructure picks one.

  • Pre-provision. You spin up two hundred environments at 08:00 for a 09:30 start, because provisioning is slow and you do not want to find out live. Now you are paying for two hundred idle machines for ninety minutes before a single attendee exists, plus the forty environments belonging to the people who registered and did not come.
  • Provision on demand with something slow. The first forty attendees get a working lab in a reasonable time, and everybody who clicked the link four seconds later gets a spinner while the instructor improvises material they had not planned to deliver. The improvisation is usually fine. The confidence loss in the room is not.

The way out is not a bigger warm pool. It is making the create itself fast enough that the burst stops being a burst and becomes a boring parallel fan-out — and then making the environment handed out be the instructor's finished environment rather than an empty box with a setup script attached.

The thundering herd, and why there is no warm pool

PandaStack has no pool of idle VMs waiting for someone to claim them. Every create restores a baked Firecracker snapshot on demand: allocate a pre-built network slot, reflink the rootfs, fork and exec Firecracker, load the memory snapshot, resume, probe the port. Measured p50 is 179 ms and p99 is 203 ms, and the snapshot load step inside that is about 49 ms.

The reason that matters for a workshop specifically is arithmetic about correlation. A warm pool is a bet on a demand forecast, and a cohort is the one workload where your forecast is both perfect and useless: you know exactly how many people are coming, and they all want their environment in the same ninety seconds, so the pool has to be the full size of the cohort and has to exist before the cohort does. You have simply moved pre-provisioning behind a nicer API. Whereas if a create is sub-second, two hundred creates fanned out across a thread pool is not an event. The pre-allocated network pool is 16,384 /30 subnets per host, so a cohort of a few hundred is nowhere near any structural ceiling; the binding constraint is host memory, which is a capacity-planning conversation rather than a race.

There is exactly one piece of homework, and it is the piece people skip.

The first boot of a template before its snapshot has been baked is a full cold boot — around 3 seconds, not 179 ms. Every restore after that uses the snapshot. If you have not touched your template since you last edited it, attendee #1 pays the cold-boot cost, and because you are running the demo from the lectern, you will be the one watching it happen on a projector while two hundred people watch you watch it. Bake it the day before. Then restore one and throw it away, just to see the fast path with your own eyes.

Fork, do not provision

The fast create is the enabling trick, not the win. The win is that you should never be running setup steps two hundred times in parallel in the first place.

Set the lab up once. Clone the repo, run the installs, pull the container images, apply the migrations, seed the database with the three rows the exercise depends on, start the dev server, leave the editor open on step 1 with the right file focused. Do it slowly and carefully, in one sandbox, the day before, with a cup of tea. Then snapshot it. Every attendee environment is a copy of that snapshot.

The attendee does not get an empty box and a bootstrap script that might fail on their behalf while they watch. They get your finished environment, mid-breath.

What the snapshot actually carries — and the one API footgun

A Firecracker snapshot is a memory image plus device state, and on PandaStack the on-disk rootfs is captured inside the same pause window, so a snapshot is a true point-in-time copy: memory and disk together. That is why a running process survives it. The dev server you started before the snapshot is still listening in the restored guest; it did not restart, it resumed. The shell history is there. The database connection pool is there, which is its own small adventure, and we will come back to it.

Now the footgun, because the two fan-out primitives are not the same and the names do not help:

  • fork_tree(count) snapshots the parent once and restores N children from it in parallel. Children inherit memory and disk. It is capped at 16 children per call, and the children land on the same host as the parent — it is the same-host primitive, and its per-child cost is a snapshot restore, in the 400–750 ms range.
  • fork() is disk-only. It pauses the parent, copies the rootfs, resumes the parent, and each child boots fresh from that disk copy. Dependencies you installed are there; the process you left running is not. It is the right call when all you want is "a machine that already has the deps", and the wrong call when the point of your golden environment is that something is already running in it.
  • Sandbox.create(from_snapshot=snap_id) restores from a snapshot ID and goes through the normal scheduler, so children can land on any host with capacity. This is the one you use for a cohort, because 200 > 16 and because you do not want a whole workshop pinned to one machine.
sbx.snapshot() returns a snapshot ID string, not an object. You restore from it with Sandbox.create(from_snapshot=snap_id). There is no snap.fork() — that method exists only in blog posts written by people who guessed, including, historically, a couple of ours.

Where the cross-host number shows up

A same-host fork is 400–750 ms. A fan-out that lands on a host which does not already hold the artifact is 1.2–3.5 s, because the snapshot has to come down from object storage before it can be restored. For most workloads that number is a footnote. For a cohort it is the number you will actually observe, because two hundred environments do not fit on one host and the scheduler will quite correctly spread them.

Which is fine — a couple of seconds is not the thing that ruins a workshop — but it changes what you do the morning of. The snapshot is mirrored to object storage after it is taken, and the first restore on each host pays the download. So restore a handful of labs during setup, before the room fills: you are not testing your code, you are warming the hosts that will serve your cohort. Think of it as the infrastructure equivalent of arriving early to find out which HDMI adapter the lectern wants.

import concurrent.futures as cf
from pandastack import Sandbox

TEMPLATE = "base"       # 4 GiB / 8 vCPU, mise with Node, Python 3.12, Go, Bun
LAB_PORT = 8080


def build_golden(course: str) -> str:
    """Set the lab up ONCE, the day before, then freeze it.

    Returns a snapshot ID string. Put it in the repo, or in your deck.
    """
    sbx = Sandbox.create(
        template=TEMPLATE,
        persistent=True,   # do not let the idle reaper eat the golden box
        metadata={"kind": "lab-golden", "course": course},
    )

    sbx.exec("git clone --depth 1 https://github.com/acme/fc-101-lab /work/lab")

    # Long setup goes through exec_stream (which honours timeout_seconds) or is
    # bounded in the shell. One-shot exec has no server-side deadline today, so
    # never hand it a 15-minute build and a big timeout and expect it to hold.
    sbx.exec_stream(
        "cd /work/lab && mise install && npm ci && ./scripts/seed-db.sh",
        on_stdout=print,
        on_stderr=print,
        timeout_seconds=900,
    )

    # Leave the lab RUNNING and parked on step 1. It must bind 0.0.0.0, or the
    # preview proxy cannot reach it from outside the guest.
    sbx.exec(
        "cd /work/lab && setsid nohup npm run dev -- --host 0.0.0.0 "
        ">/var/log/lab.log 2>&1 &"
    )
    ready = sbx.exec("timeout 60 sh -c 'until curl -fsS -o /dev/null "
                     "http://127.0.0.1:8080/; do sleep 1; done'")
    assert ready.exit_code == 0, f"golden lab never came up: {ready.stderr}"

    # Memory AND disk, in one pause window. Synchronous and not instant on a
    # multi-GiB guest, which is another reason this is a day-before job.
    snap_id = sbx.snapshot()
    sbx.kill()
    return snap_id


def open_lab(attendee: str, snap_id: str, course: str) -> dict:
    """One attendee, one microVM, restored from the golden snapshot."""
    sbx = Sandbox.create(
        from_snapshot=snap_id,
        ttl_seconds=90 * 60,      # IDLE timeout, not a walltime budget
        metadata={
            "kind": "lab",
            "course": course,
            "attendee": attendee,  # your roster id, not "laptop near the door"
            "golden": snap_id,     # exactly which lab build they were given
        },
    )
    return {
        "attendee": attendee,
        "sandbox_id": sbx.id,
        "url": sbx.preview_url(LAB_PORT),
    }


def open_cohort(snap_id: str, roster: list[str], course: str) -> list[dict]:
    """Two hundred people clicked the link at the same time. Fine."""
    with cf.ThreadPoolExecutor(max_workers=32) as pool:
        futures = {
            pool.submit(open_lab, who, snap_id, course): who for who in roster
        }
        out = []
        for fut in cf.as_completed(futures):
            who = futures[fut]
            try:
                out.append(fut.result())
            except Exception as exc:               # retry once, then escalate
                print(f"[retry] {who}: {exc}")
                out.append(open_lab(who, snap_id, course))
        return out

Note what is not in that function: no git clone per attendee, no npm install per attendee, no database seed per attendee, no waiting for a dev server to boot per attendee. Those all happened once, yesterday, in build_golden. The per-attendee path is a restore and a URL.

Reset is the half that saves your afternoon

At some point in the second hour, attendee 143 will arrive at step 4 — "now edit the config and restart the service" — and will interpret it creatively. Not maliciously. They will simply have a different mental model of what the instruction meant, act on it with the full confidence of root, and produce a machine state that is not on any branch of your lab's decision tree. This is not a failure of the attendee or of the material. It is what a hands-on lab is FOR. Breaking things is the point; the environment is the thing you brought so that breaking them is cheap.

What must not happen is you debugging their box. You have a hundred and ninety-nine other people and eleven minutes of slot left. The recovery is not a diagnosis, it is a replacement: throw the machine away and restore another copy of the golden snapshot. Same URL shape, same starting state, same step 1, thirty seconds of conversation instead of fifteen minutes of shoulder-surfing.

This is also why a per-attendee microVM beats per-attendee accounts on one shared box. On a shared box, "reset attendee 143" is a cleanup script you have to have written in advance, against a failure you did not anticipate, and it runs next to a hundred and ninety-nine other people's work. With one VM each, reset is a delete and a create, and "a hundred percent of the mess is inside one disposable machine" is a property you get for free rather than a script you maintain.

import datetime as dt
from pandastack import Sandbox


def reset_one(attendee: str, snap_id: str, course: str) -> dict:
    """Attendee 143 has discovered `sudo rm -rf /usr`. Do not investigate."""
    for sbx in Sandbox.list():
        md = sbx.metadata or {}
        if md.get("kind") == "lab" and md.get("attendee") == attendee:
            sbx.kill()          # memory, disk, stray processes, regret: gone

    fresh = open_lab(attendee, snap_id, course)
    fresh["reset_at"] = dt.datetime.now(dt.UTC).isoformat()
    return fresh


def reset_room(snap_id: str, course: str) -> list[dict]:
    """After lunch, for the module that starts from a clean slate.

    Collect the roster from the live labs, kill them all, re-open. Cheaper
    than asking two hundred people to undo the morning by hand.
    """
    roster = sorted(
        (sbx.metadata or {}).get("attendee", "")
        for sbx in Sandbox.list()
        if (sbx.metadata or {}).get("course") == course
    )
    roster = [who for who in roster if who]
    for sbx in Sandbox.list():
        if (sbx.metadata or {}).get("course") == course:
            sbx.kill()
    return open_cohort(snap_id, roster, course)

One pragmatic note: send attendees a short redirect link that resolves to their current lab URL, not the raw URL itself. Then a reset does not require the attendee to go back into the chat and find a new link, which — given that they just broke their environment and are already slightly embarrassed — is a kindness worth the half hour of work.

Access: a URL, not an SSH tutorial

The attendee needs a browser tab. Not a key, not a VPN client, not a four-step CLI install that itself becomes a twenty-minute support segment. Any port the guest exposes is reachable at a stable preview host of the form https://PORT-SANDBOXID.your-suffix for the lifetime of that sandbox, which means the thing you hand out is a plain https link and the thing you demo on the projector is the same plain https link.

Be honest about the security model, because it is a tradeoff rather than an accident. Sandbox preview URLs are tokenless. There is no token endpoint and no sign-in; the sandbox UUID in the hostname IS the credential. A random version-4 UUID is not a weak secret in the guessing sense, but it is a bearer credential embedded in a URL, with all the properties that implies.

A UUID pasted into a shared workshop chat is a shared environment. Anyone in that channel can open it, and so can anyone they forward it to. For a two-hour tutorial on a public sample app, that is completely fine and you should say so plainly. For a lab that touches a real customer dataset, a licensed binary, or anything an attendee would be upset to see a stranger typing into, it is not — put your own authenticating proxy in front, or keep that module on infrastructure where the identity story is yours.

The practical workshop version of this: generate the per-attendee links into whatever the attendees already have — the registration email, a per-person row in a sheet, your LMS — and never into the room-wide chat. It costs you ten minutes of scripting and it means attendee 60 cannot accidentally do step 4 inside attendee 12's lab, which is the single funniest and least useful way to spend the back half of a session.

ttl_seconds is an idle timeout, not a walltime budget

This distinction matters more for a lab than for anything else, and it is routinely misread. ttl_seconds does not mean "delete this sandbox after N seconds". It means "delete this sandbox after N seconds with no activity". The reaper compares the current time against the last activity timestamp, not against creation. The default is five minutes, which is correct for a CI job and wrong for a workshop, so set it explicitly.

Read the right way round, this is exactly the semantics a lab wants:

  • An attendee who goes to lunch stops costing you money, and so does the person who registered, got a link, and never opened it.
  • An attendee who is typing, running commands, or loading pages in their lab does not get reaped out from under them, however long the session runs. A four-hour workshop does not need a four-hour TTL; it needs a TTL longer than the longest gap in the room's activity.
  • Activity means requests that actually use the guest — exec, filesystem reads and writes, terminal sessions, and traffic through the preview proxy. Merely looking at a sandbox's status in a dashboard does not count, deliberately: observing a machine must not keep it alive.

And the failure mode, stated plainly so you can design around it: a long lecture segment can reap the room. If you talk for thirty-five minutes with the lab untouched on everyone's second monitor and your TTL is thirty, you come back from theory to two hundred dead environments and a slide that says "now let's try it". Set the TTL to comfortably exceed your longest talking stretch, including the Q&A that overruns. For a full-day course, the simpler answer is to mark the labs persistent, which exempts them from the idle reaper entirely, and tear the whole room down yourself at the end — one scripted loop over the cohort, which you should write before the day anyway, because "delete two hundred VMs by hand at 5:15pm" is how a cost surprise is born.

The arithmetic, and why memory dominates

One rate card, all workload classes: $0.054 per vCPU-hour and $0.0162 per GiB-hour. CPU is billed on active CPU-seconds actually burned; memory is billed on committed GiB-hours. Egress is metered separately. The base template is 4 GiB and 8 vCPU — the 8 vCPU are burst capacity, shared fairly under contention by cgroup weights, not eight cores reserved for one attendee.

So one attendee's lab, running for a six-hour workshop day, costs 4 GiB × 6 h × $0.0162 = $0.389 in memory. The CPU side depends entirely on what the lab does, and here is the thing that makes a training lab different from every other workload on the same rate card: attendees think more than they compute. They read the slide, they type slowly, they ask the person next to them what step 3 meant, they run one build and then discuss it for four minutes. If an attendee's lab burns the equivalent of fifteen minutes of one core across the whole day — an assumption, not a measurement, and yours will differ — that is 0.25 vCPU-hours, or $0.0135.

Memory is therefore something like ninety-six percent of the bill. For a two-hundred-person workshop day: roughly $78 of memory and about $3 of CPU. That ratio is the single biggest difference from a CI workload, where the machine exists specifically to pin every core for four minutes and the CPU term is the one you optimise. In a lab, the only lever that matters is how many GiB-hours of committed memory you are holding, which means the only two cost decisions are the template's RAM size and how long environments exist.

Which also prices the mistakes. Pre-provisioning two hundred labs at 08:00 for a 09:30 start is ninety minutes × 200 × 4 GiB × $0.0162, or about $19 spent on an empty room — not ruinous, but it is a fifth of the whole day's bill for nothing, and it scales with how nervous you are. Leaving the room running overnight because nobody wrote the teardown loop costs more than the workshop did.

Four ways to give two hundred people a lab

Lab provisioning approaches against the four things a cohort actually stresses.
Approach200 arrivals in 90 sReset one attendeeIdle lecture hourBlast radius
Pre-provisioned VMsFine — you have paid since 08:00SSH in and diagnose, liveFull price, 200 times overOne VM each
One big box, 200 accountsInstant, nothing to createA cleanup script you wrote before you knew the bugFull price, onceShared kernel, shared filesystem, shared fate
Container per attendeeFast to startRecreate, then re-run setup per attendeeCheap if you scale to zeroShared host kernel
Bake, snapshot, restore per attendeep50 179 ms per create; 1.2–3.5 s on a host that must pull firstKill and restore from the golden snapshotReaped on idle TTL, or memory-only if persistentOwn kernel per attendee

The container row is deliberately qualitative. Container start time and reset ergonomics depend enormously on your image, your runtime and your orchestrator, and any specific number you read in a blog post — including this one — should be checked against your own platform's current documentation and then measured on your own hardware before you plan a session around it.

What this does not do

The honest paragraph, because a workshop is the worst possible place to discover a platform limit.

  • RAM is fixed by the baked template snapshot. Firecracker cannot change vCPU or memory at snapshot restore, so "give everyone 16 GiB for the Kubernetes module" is not a per-attendee knob. cpu and memory_mb on a create are overridden to the baked values whenever the template has a snapshot — not an error, just ignored, which is worse. If your heavy module needs more memory, you bake a template at the size you need and the whole cohort runs at that size. Decide this before the day, not during it.
  • There is no GPU. No fine-tuning lab, no CUDA module, no "and now we train a small model". Those labs need different infrastructure and there is no amount of cleverness in the snapshot layer that changes it.
  • Taking the golden snapshot is synchronous and not instant on a multi-GiB guest. It is a day-before operation, not something you do at 09:28 after one last fix.
  • fork_tree is capped at 16 children per call and keeps them on the parent's host. For a real cohort you fan out with create(from_snapshot=...) and let the scheduler spread the room.
  • A restored guest resumes with whatever its memory held, which includes sockets. A database connection pool or a long-lived websocket that was open at snapshot time comes back pointing at a connection that no longer exists, and the client has to reconnect. Most frameworks do this on their own; some retry exactly once and then sulk. Test your golden environment by restoring it, waiting a minute, and clicking through the first three steps — not by looking at it immediately after the snapshot, when everything is still warm and nothing has had to reconnect.

The morning-of pre-flight

Run this at 08:45, not at 09:30. It proves the three things that actually go wrong: the snapshot exists and restores, the lab inside it serves traffic, and the preview URL is reachable from outside the guest. Everything else you can improvise.

#!/usr/bin/env bash
# pre-flight.sh — prove the golden lab restores and answers, then throw it away.
set -euo pipefail

: "${PANDASTACK_API_KEY:?export your API key first}"
SNAP="$(cat .lab-golden-snapshot)"   # written by build_golden(), day before

python3 - "$SNAP" <<'PY'
import sys, time, urllib.request
from pandastack import Sandbox

snap = sys.argv[1]
t0 = time.time()
sbx = Sandbox.create(
    from_snapshot=snap,
    ttl_seconds=600,
    metadata={"kind": "lab-preflight"},
)
print(f"restore: {int((time.time() - t0) * 1000)} ms -> {sbx.id}")

# The dev server should already be listening: it was running when we froze it.
print(sbx.exec("pgrep -af 'npm run dev' || echo 'NOTHING RUNNING'").stdout)

url = sbx.preview_url(8080)
for attempt in range(20):
    try:
        with urllib.request.urlopen(url, timeout=5) as resp:
            print(f"lab answered {resp.status} at {url}")
            break
    except Exception as exc:
        print(f"waiting ({attempt}): {exc}")
        time.sleep(1)
else:
    sbx.kill()
    raise SystemExit("golden snapshot does not serve the lab -- fix it NOW")

sbx.kill()
print("pre-flight ok")
PY

And the rest of the checklist, in the order the day happens:

  1. Day before: build the golden environment by hand, leave the lab running on step 1, snapshot it, record the snapshot ID in the repo.
  2. Day before: confirm the template's own snapshot is baked, so nobody pays the ~3 s cold boot. One throwaway create tells you.
  3. Morning: run the pre-flight above, then restore a handful of labs and leave them a few minutes — you are warming hosts, not testing code.
  4. Morning: generate per-attendee links into private channels (registration email, per-person sheet row), never into the room chat.
  5. Set the idle TTL longer than your longest lecture segment, or mark the cohort persistent and own the teardown yourself.
  6. Have reset_one wired to a single command you can run from the lectern without thinking. You will use it. Probably twice.
  7. 5pm: run the teardown loop over the cohort, then list sandboxes and confirm the count is zero. Write this script before the day, not after the invoice.

The same machine, other rooms

Everything above generalises to any synchronised cohort, which is a surprisingly common shape once you notice it: a certification exam with a practical component, a customer onboarding week, a security training day, a sales POC where six people from the same prospect want to click around their own copy at the same time on the same call. The ingredients never change — one golden environment built carefully, a snapshot of it, a restore per person, a URL each, an idle timeout that matches how humans actually behave, and a reset you can run in one command without apologising.

What you get for it is not really a cost saving, although the cost saving is real. It is that the first twenty minutes of your workshop become the first twenty minutes of your workshop. Nobody is doing laptop support. Nobody is improvising. Attendee 143 still breaks their environment on step 4 — they will, that is the whole reason you brought the environment — and the fix is a sentence instead of a diagnosis.

Frequently asked questions

How many attendees can I provision at once?

In practice, as many as you have memory for, fanned out in parallel. Each create is a snapshot restore at p50 179 ms and p99 203 ms, so a few hundred creates through a thread pool is a fan-out rather than an event, and the per-host network pool is 16,384 pre-allocated subnets — nowhere near a structural limit for a cohort. The real constraint is committed memory across the fleet: two hundred labs on the base template is two hundred times 4 GiB, and that has to fit somewhere. Two API-level caveats. First, fork_tree is capped at 16 children per call and places them on the parent's host, so it is the wrong primitive for a whole room; use Sandbox.create(from_snapshot=...) and let the scheduler spread the cohort. Second, creates that land on a host which does not yet hold the snapshot pay a download first, in the 1.2 to 3.5 second range rather than sub-second, which is why warming a few hosts during setup is worth the five minutes.

Will a server I left running in my golden environment still be running in each attendee's lab?

With a snapshot, yes. A PandaStack snapshot captures guest memory, device state and the on-disk rootfs inside the same pause window, so a restore is a true point-in-time copy and a process that was listening when you froze it is still listening when it resumes. It did not restart — it continued. That is the whole reason to snapshot a finished environment rather than ship a setup script. One important exception and one API trap. The exception is sockets: anything that held an open connection at snapshot time comes back pointing at a connection that no longer exists, so database pools and websockets must reconnect, and most libraries do that automatically while a few do not. Test by restoring, waiting a minute, and clicking through the first steps. The trap is that plain fork() is disk-only: children get the filesystem but boot fresh, so the running process does not survive. fork_tree() and create(from_snapshot=...) are the memory-preserving paths.

What happens if the room sits idle during a long lecture segment?

The labs get reaped, if your TTL is shorter than the segment. ttl_seconds is an idle timeout, not a walltime budget: the reaper compares now against the sandbox's last activity, so an untouched lab counts down regardless of how recently it was created, and the default is five minutes. Activity means requests that actually use the guest — exec, filesystem reads and writes, terminal sessions, and traffic through the preview proxy — while purely observational calls like fetching a sandbox's status deliberately do not count, because watching a machine must not keep it alive. The practical rule is to set the TTL comfortably longer than your longest stretch of talking, including the Q&A that overruns and the fire-alarm drill nobody warned you about. For a full-day course the cleaner option is to mark the labs persistent, which exempts them from the idle reaper entirely, and own the teardown yourself with a scripted loop at the end of the day.

Can I give each attendee more memory for a heavier module?

Not per attendee, no. Firecracker cannot change vCPU count or RAM at snapshot restore, so memory is a property of the baked template snapshot rather than of the create request. If you pass cpu or memory_mb to a create on a template that has a baked snapshot, the agent overrides them to the baked values — it does not error, which is arguably worse, because the API response and your invoice will both quietly show the real size. The base template is 4 GiB and 8 vCPU; code-interpreter and agent are 2 GiB, browser is 4 GiB, postgres-16 is 1 GiB, all at 8 vCPU of burst capacity shared by cgroup weight under contention. If a module genuinely needs more memory, the answer is to bake a template at that size and run the whole cohort on it, which also means the decision has to be made before the workshop rather than discovered during it. There is no GPU option at all, so ML-training labs need different infrastructure entirely.

What does a two-hundred-person workshop day actually cost?

On the single rate card — $0.054 per vCPU-hour and $0.0162 per GiB-hour, with CPU billed on active CPU-seconds and memory on committed GiB-hours — one attendee on the 4 GiB base template for a six-hour day is 4 × 6 × $0.0162 = $0.389 of memory. The CPU term depends on your lab, but a training workload is unusual in that attendees think far more than they compute: if a lab burns the equivalent of fifteen minutes of one core across the day, that is $0.0135. So memory is roughly ninety-six percent of the bill, and a two-hundred-person day lands near $78 of memory plus about $3 of CPU, with egress metered separately. The utilisation figure there is an illustrative assumption, not a measurement — yours will differ and you should check it. The useful consequence is that the only two cost levers are the template's RAM size and how long environments exist, which is why pre-provisioning at 08:00 for a 09:30 start burns about $19 on an empty room.

Keep reading

Related posts

More in Snapshots & forking · See Thaw: sub-second cold restore

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.