all posts

Per-Tenant Map Tile Rendering in MicroVMs: When Your Customers Write the Stylesheet

Ajay Kumar··11 min read

The ticket said "maps are slow for everyone". What had actually happened was that one customer on a self-serve plan had saved a style where a single `line-width` expression carried about two thousand zoom stops, and our tile workers were re-evaluating that expression per feature, per tile, at a zoom level where the viewport covered most of a continent. The renderer did not crash. It just got slow, held its heap, and stopped answering, and because tile workers were a shared pool, every other tenant's map stopped drawing too. The customer who broke it never noticed, because their own map was the one tab that was still warm.

That is the shape of multi-tenant tile rendering. You accept a style — a Mapbox GL style JSON, a MapLibre style, or if you have been around a while, CartoCSS compiled down to Mapnik XML — plus the tenant's fonts, their sprite sheet, and increasingly their own data layers. Then you execute that artefact through a native rendering stack on your hardware, inside a hard latency budget, dozens of times per viewport pan. The style is configuration in the same sense that a Dockerfile is configuration.

One pan is forty renders, and the input is a stranger's

Tiles are nastier than most image pipelines for four reasons that compound. The fan-out is enormous: a 1600x900 viewport at 256-pixel tiles is roughly 35 tiles, and a user dragging the map across a city will request several hundred in a few seconds. Each render is short — tens to low hundreds of milliseconds for vector work — so there is no room to amortise setup. The input is customer-controlled, and not just the data: the style, the glyph PBFs, the sprite JSON, the sprite PNG, and sometimes a GDAL dataset descriptor. And the libraries doing the work are Mapnik, GDAL/OGR, PROJ, FreeType, HarfBuzz, Cairo or Skia, libpng and libjpeg — three decades of C and C++ that will cheerfully do exactly what the style told them to.

The thumbnailing version of this problem is well understood — see Per-Tenant MicroVM Isolation for Thumbnail Generation — and tiles differ in one way that decides your architecture. A thumbnail job is a batch unit you can isolate per request. A tile is served from a long-lived renderer holding warm state, and that warm state is the entire reason the latency budget is achievable. You cannot make the boundary per request without throwing away the thing that makes it fast. So the boundary has to be per tenant.

Four ways a stylesheet ruins your afternoon

These are not hypotheticals, they are the realistic failure classes, in roughly ascending order of how badly people take the news.

  1. The expression bomb. Style expressions are a small interpreted language with `interpolate`, `case`, `match`, `step` and arbitrary nesting. A two-thousand-stop interpolation evaluated per feature per tile is not a parse-time cost you can reject — it is a runtime cost that scales with how much data happens to be in that tile. The same style is instant at zoom 18 over a village and pathological at zoom 4 over Europe.
  2. The zoom-0 request over a global dataset. One tile, the whole planet. If the layer is backed by a few hundred million features and the style has no `minzoom`, the renderer dutifully queries, reprojects and draws all of it. This is the request that turns a 200 ms budget into a four-minute transaction holding a database connection.
  3. The raster the style asked for. Between `raster-resampling`, oversized sprite sheets, and a GDAL warp with a generous target resolution, a few kilobytes of config can request a 40000x40000 surface. At RGBA that is about 6.4 GB, requested by a file you accepted over a web form. Cairo and GDAL will both attempt it before they refuse it.
  4. The font with a pathological shaping table. A tenant uploads a typeface. TrueType ships a bytecode interpreter, and HarfBuzz runs the font's own GSUB/GPOS rules; a hostile or merely broken `GSUB` chain can blow up shaping cost or walk somewhere it should not. This one has its own post — A Font Is a Program: Rendering User-Uploaded Typefaces Without Trusting Them — and the summary is that a `.ttf` is a program, so treat it like one.

Then there is the fifth one, the one nobody sees coming, because it does not look like code execution or like a denial of service. It looks like a data layer. GDAL's VRT format is an XML document that describes a dataset by naming its sources, and those source filenames go through GDAL's virtual filesystem layer: `/vsicurl/`, `/vsis3/`, `/vsizip/` and friends. A tenant who can supply a VRT — or a style whose source URL is honoured, or a sprite reference, or a glyphs template — can make your renderer issue an HTTP request to an address of their choosing, from inside your VPC, with your instance's identity. That is a server-side request forgery whose confused deputy is a map renderer, and it is invisible in every log you are currently reading, because to your observability stack it is just GDAL opening a dataset. The same class applies to raw uploads rather than styles, which I covered in Isolating Per-Tenant Geospatial Processing Jobs in MicroVMs; the serving path is worse, because it runs hundreds of times a minute and is the one you tuned for latency rather than paranoia.

If your tile renderer can reach `169.254.169.254`, a cloud metadata credential is one hostile style away, and the exfiltration channel is the rendered tile itself — a text layer that draws the response body into a PNG the attacker then downloads from your own CDN. Egress policy is not a hardening task you schedule for next quarter. It is the fix.

Why a container is a thin boundary for this specific job

I want to be precise, because "containers are not a security boundary" is a slogan and slogans are easy to dismiss. The accurate statement is narrower: for this workload the container boundary does not change the trust domain. The renderer is native code parsing attacker-shaped input, and it is doing that against a kernel shared with every other tenant on the node. A memory limit and a seccomp profile are both real and both worth having, but they are blast-radius accounting, not a different trust domain. They bound how much damage a misbehaving renderer does to you. They do not stop a kernel-level bug reached through a decoder from being a bug in everybody's kernel.

The memory limit deserves its own paragraph, because it is the control people reach for first and the one that behaves least like the mental model. A `memory.max` on a cgroup helps only if the renderer dies cleanly when it hits it. In practice the kernel OOM killer picks a victim inside the cgroup by badness score, and the fattest process in a tile pool is frequently the one holding the shared tile cache or the connection pool rather than the one that misbehaved. C++ allocators also do not return freed arenas promptly, so a renderer that survived one 6 GB request sits on the high-water mark and the next tenant's render is the one that gets killed. You wanted a tenant boundary and you got a priority inversion.

Seccomp has a different problem: you cannot write a tight profile for this stack. GDAL wants `openat`, `mmap`, `mremap`, threads, and depending on the build, `io_uring`; PROJ wants its grid files; Cairo wants shared memory. The profile you can actually ship is wide enough to drive a bus through, and the narrow one breaks a driver you did not know a customer was using. Gated runtimes like gVisor and Kata are a genuinely better answer here than a plain container, and if you are already standing in a Kubernetes cluster they are a much shorter walk than a new platform — verify current capabilities against their docs, because both move. I compared the three in Firecracker vs Kata vs gVisor: three isolation models; the short version of why I built on Firecracker is that I wanted the snapshot primitive as much as I wanted the boundary.

The shape that works: one warm microVM per tenant

Give each tenant a microVM that holds their style, their fonts, their sprites and a read-only mount of their data. Wake it on demand, sleep it when idle. Serve that tenant's tiles only from that machine. Three things fall out of that arrangement, and they are the whole argument.

  • One tenant's pathological style cannot poison another tenant's renderer process, because there is no other tenant in that address space, in that page cache, or on that kernel.
  • A crash costs exactly one tenant's tiles. The failure mode goes from "maps are slow for everyone" to "this customer's map is degraded", which is a support ticket instead of an incident.
  • You can resource-size per tenant. A customer rendering a global dataset gets a bigger machine; the long tail of customers rendering one city gets a small one. In a shared pool that distinction can only be expressed as a quota, and a quota is a promise the OOM killer has not read.

The reason this is affordable rather than merely correct is the snapshot. Bake a per-tenant snapshot after the renderer has started, parsed the style, loaded and shaped the fonts, opened every dataset handle and warmed its block cache. On PandaStack that snapshot restores at p50 179 ms, p99 around 203 ms, and the `/snapshot/load` call inside that is roughly 49-80 ms of it; the rest is network setup, a reflink of the rootfs, and the readiness probe. The mechanics are in The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms. Compare that to a cold renderer process, which has to re-read the whole style, re-shape the glyph set, re-open every dataset and re-warm its caches — and note that the cold renderer is doing that work on hardware where the first tile request is already waiting.

Guest memory also restores lazily, which matters more here than in most workloads. Firecracker maps the memory file `MAP_PRIVATE`, so pages fault in as the guest touches them, and with UFFD streaming they are pulled from object storage in 4 MiB chunks over HTTP range reads, with a zero-chunk bitmap so empty regions cost nothing and a prefetch trace replaying the hot set in the background. A tile renderer is an excellent fit, because its hot set is small and extremely repeatable: the compiled style, the shaped glyph atlas, the dataset headers and the first few index pages. You are not streaming a 4 GiB heap, you are streaming the part the renderer looks at.

Firecracker cannot change vCPU or RAM at snapshot restore, so the baked `meta.json` governs the size — a `cpu=` or `memory_mb=` on create is overridden to the baked values. Per-tenant sizing therefore means per-tenant templates, not a create-time argument. Our `base` template is 4 GiB / 8 vCPU; if a tenant needs 16 GiB you bake a 16 GiB template for them, which is a deliberate, auditable decision rather than a knob a request can turn.

The fan-out trick: fork_tree from a warmed renderer

Now the latency problem. One warm VM serves one tenant's tiles fine at browsing pace, but a viewport pan is a burst, and a burst wants parallelism that a single 8-vCPU guest will eventually run out of. The move is to branch the warm machine rather than start new ones.

`fork_tree(n)` gives you children that inherit the parent's running memory. That is the point: the children come up with the style already parsed, the glyph atlas already shaped, the dataset handles already open and the relevant pages already in the page cache, because they are copy-on-write views of the parent's memory and disk. The rootfs clone is an XFS reflink or a dm-snapshot; the memory is `MAP_PRIVATE` over the same backing. A child that renders a tile and exits has touched a few megabytes of private pages. The cap is 16 children, which is roughly the right number for a viewport batch and definitely not the right number for a traffic spike — for that you want more parent VMs, not deeper trees.

`fork()` is the other primitive and it is deliberately different: it clones the disk and the child cold-boots, with its own entropy, its own PIDs and its own clock. That is the right call for a batch worker that should not inherit anything, and the wrong call here, because a cold boot is exactly the state you spent a snapshot avoiding. Same-host fork lands in 400-750 ms; cross-host is 1.2-3.5 s because the artefacts have to come across. Keep a tile batch on one host. The full distinction is in Snapshots and Forks: Copy-on-Write for Running Machines and it is worth ten minutes, because the two names look like they do the same thing and they do not.

One honest caveat on `fork_tree`: inheriting running memory means inheriting the parent's seeded PRNG state, so every child agrees on the next random number. For deterministic tile output that is a feature — you want the same tile bytes from any child. For anything that mints an identifier inside the child, it is a bug waiting for a support ticket. Generate request ids in the orchestrator, not in the renderer.

import json
from pandastack import Sandbox

TENANT = "acme"
STYLE = json.load(open("acme-style.json"))

sbx = Sandbox.create(
    template="base",
    ttl_seconds=900,  # IDLE timeout; a busy renderer is not reaped
    metadata={"tenant": TENANT, "role": "tile-renderer"},
)

# The tenant's artefacts go in. Nothing else does.
sbx.filesystem.write(
    "/srv/tenant/style.json", json.dumps(STYLE)
)

# exec() does not reliably enforce a timeout argument, so bound
# the render in-guest where the kernel can actually do it.
RENDER = (
    "timeout -k 250ms 1s render-tile"
    " --style /srv/tenant/style.json"
    " --no-network --max-pixels 16777216"
)

r = sbx.exec(f"{RENDER} --z 12 --x 2048 --y 1361 --out /tmp/t.png")
if r.exit_code in (124, 137):
    raise TimeoutError("render blew the 1s budget")
if r.exit_code != 0:
    raise RuntimeError(r.stderr)

png = sbx.filesystem.read("/tmp/t.png")

# Bank the warm state, then branch it for the viewport burst.
warm = sbx.snapshot()          # synchronous; returns a snapshot id
kids = sbx.fork_tree(8)        # inherits RUNNING memory; cap is 16

for child, (z, x, y) in zip(kids, viewport_tiles[:8]):
    out = f"/tmp/{z}-{x}-{y}.png"
    res = child.exec(f"{RENDER} --z {z} --x {x} --y {y} --out {out}")
    if res.exit_code == 0:
        cdn.put(TENANT, z, x, y, child.filesystem.read(out))
    child.kill()               # no sbx.delete(); teardown is kill()

# Next wake skips the whole warm-up.
renderer = Sandbox.create(from_snapshot=warm)

Egress off is not hardening, it is the fix

The parsing and rendering VM should have no outbound network. Not a filtered allowlist, not an egress proxy with a blocklist of link-local addresses — no default route and no interface that can reach one. Do that and the entire GDAL-VRT SSRF class stops existing, along with the sprite-URL variant, the glyphs-URL variant, and the `/vsicurl/` variant you have not thought of yet. There is nothing to filter correctly, because there is nothing to reach.

This is cheap to arrange when every sandbox already has its own network namespace. Each PandaStack sandbox gets one of 16,384 pre-allocated /30 subnets in `10.200.0.0/16`, with its own netns, veth pair and tap device, so "this one has no route out" is a property of a namespace rather than a rule in a shared table that somebody can reorder. The broader treatment is in Controlling Network Egress for Untrusted Code.

The obvious objection is that the renderer needs data. It does, and the orchestrator stages it: the tenant's datasets are mounted read-only into the guest, or materialised into it before the snapshot is baked, by a component that does have network access and does not parse customer config. The split is the point. One process may talk to the network and only reads things you control; the other reads things customers control and cannot talk to anything. Egress-off also kills the two commonest abuse shapes on any platform that executes customer input: exfiltration of the artefacts, and a crypto miner that cannot find a pool.

{
  "version": 8,
  "//": "3 kB on the wire. Both of these are hostile.",
  "sprite": "http://169.254.169.254/latest/meta-data/iam",
  "glyphs": "file:///srv/tenant/fonts/{fontstack}/{range}.pbf",
  "sources": {
    "tenant": { "type": "vector", "url": "file:///srv/t.json" }
  },
  "layers": [
    {
      "id": "roads",
      "type": "line",
      "source": "tenant",
      "source-layer": "transportation",
      "//": "stops continue to ~2000 entries, each re-evaluated",
      "//2": "per feature per tile. No minzoom. Draws at z0.",
      "paint": {
        "line-width": [
          "interpolate", ["linear"], ["zoom"],
          0, 0.01, 0.001, 0.02, 0.002, 0.03
        ]
      }
    }
  ]
}
# The input nobody screens: a VRT is XML that names its own sources.
$ cat tenant-layer.vrt
<VRTDataset rasterXSize="40000" rasterYSize="40000">
  <VRTRasterBand dataType="Byte" band="1">
    <SimpleSource>
      <SourceFilename relativeToVRT="0"
      >/vsicurl/http://169.254.169.254/latest/meta-data/</SourceFilename>
    </SimpleSource>
  </VRTRasterBand>
</VRTDataset>

# The fix is not a better parser. It is an empty routing table.
$ ip netns exec ns-$SBX ip route
# (no output: no default route, nothing to forge a request to)

# Belt and braces inside the guest, because GDAL will still try:
export GDAL_SKIP=VRT                      # drop the driver entirely
export GDAL_DISABLE_READDIR_ON_OPEN=EMPTY_DIR
export CPL_VSIL_CURL_ALLOWED_EXTENSIONS=none
export GDAL_HTTP_TIMEOUT=1
export GDAL_CACHEMAX=256                  # MiB, not "whatever is free"

# And a hard ceiling on what any style may ask for:
ulimit -v 3145728                         # 3 GiB address space
ulimit -t 2                               # 2 s of CPU, then SIGXCPU
exec timeout -k 250ms 1s render-tile --style /srv/tenant/style.json "$@"

The CDN holds the tiles. The VM exists to produce a miss.

Tiles are the most cacheable artefact in web mapping, and getting the layering right is what makes the per-tenant VM model economically boring instead of alarming. The cache key is a style version plus `z/x/y` plus format; nothing in it is user-specific, nothing in it is time-dependent, and a style edit is a new version rather than an invalidation. That means almost every tile request in a healthy system never reaches a renderer at all.

  • CDN edge: every rendered tile, immutable, keyed by style version. This absorbs the overwhelming majority of traffic, including the pan bursts, because a popular area is popular for every viewer.
  • Object storage behind the edge: the durable tile store and an origin shield, so a cold edge is a storage read rather than a render. Pre-render the low zooms (z0 to z8 is a small, bounded pyramid) at style-publish time so no first viewer ever waits for a continent.
  • The per-tenant VM: cache misses only — deep zooms, fresh style versions, and tiles over data that changed. This is the only layer that executes customer config.
  • Inside the VM: the guest page cache over the mounted datasets, plus the renderer's own block cache. This is the state the snapshot preserves, and the reason a wake beats a cold start.

Once the caching is right, the number that governs your bill is not render cost, it is idle cost. If you have two thousand tenants and forty of them have somebody looking at a map right now, a model that keeps two thousand renderers warm is paying for 1,960 machines to hold a style nobody is reading. That is the cost structure that makes people build shared pools in the first place, and it is the one scale-to-zero removes: when a tenant goes idle their VM is deleted, not parked, and the next request wakes a fresh one from the snapshot in object storage. The latency anatomy of that wake is in Where Scale-to-Zero Wake Time Actually Goes, and how to put a CDN in front of it so the first viewer does not feel it is in How to Put a CDN in Front of a Scale-to-Zero App.

The one request that pays the wake is the first cache miss after an idle period, and the trick is to arrange for that request not to be a human waiting on a map. Serve the pre-rendered low-zoom pyramid from the edge so the initial paint is always cached, and let `stale-while-revalidate` cover the first deep-zoom miss. The wake then happens behind a tile the user is already looking at.

What each boundary actually contains

Four boundaries for a tile renderer, compared on the two failures that matter and on what a per-tenant cache can look like.
BoundaryContains a native-code crash?Contains an SSRF from the input?Per-tenant cache shape
In-process threadNo. A segfault in Cairo takes the server and every warm cache with itNo. One network namespace, one set of credentialsOne shared cache, one shared eviction policy, noisiest tenant wins
Worker process per renderPartly. The worker dies; the parent survives if it was written carefullyNo. Same netns, same metadata endpoint, same IAM roleCache lives in the parent, or is rebuilt per render
Container per tenantMostly. Shared kernel, so a bug reached through a decoder is shared tooOnly as well as you wrote the egress policy, and only while it stays writtenPer-tenant, but shared page cache and one OOM killer picking by badness
MicroVM per tenantYes. Own kernel. The blast radius is one tenant's tilesYes, by absence. No default route and no interface that reaches onePer-tenant and private, snapshot-able, restored warm in 179 ms p50

Where I would not use this, and what we do not do

We have no GPUs and no GPU passthrough, and for map rendering that is a real exclusion rather than a footnote. If your product's value is server-side GPU rasterisation of vector tiles — a headless `mapbox-gl-native` on EGL, a Vulkan or Metal-backed Skia pipeline, hardware-accelerated terrain or hillshade — you cannot do it on PandaStack, and you should buy GPU instances from somebody who sells them. The honest framing is that we are a good substrate for CPU-bound rendering and for the parsing that precedes it, and not a substrate for the GPU path at all.

The guest kernel is 5.10 on an Ubuntu 24.04 rootfs. Userspace is modern, so recent GDAL, PROJ and HarfBuzz builds are fine; anything that depends on a newer kernel interface is not. If your renderer leans on recent `io_uring` features or a filesystem that landed after 5.10, check before you port. And a per-tenant VM is pinned to a host, which means a genuinely large tenant is a capacity-planning conversation rather than an autoscaling event — five hundred warm 4 GiB guests is two terabytes of real RAM, and scale-to-zero is what makes the arithmetic survivable rather than what makes it disappear.

There is also a case where none of this is worth it, and I would rather say so than sell you a boundary you do not need: if the style is yours. A first-party cartography product, where you author every style and the only customer input is a bounding box and a layer toggle, does not have an untrusted-input problem. Render it in a container pool behind a CDN and spend the saved complexity elsewhere. Same for pure raster tiles sliced from your own cloud-optimised GeoTIFFs: no customer config, no customer fonts, no VRT, no problem. Per-tenant microVMs start paying for themselves the moment a stranger can save a style, and not one commit earlier. If you want a generic sandbox API rather than this specific topology, E2B and Modal are reasonable places to look too — compare against their current docs, since the category moves.

Where this leaves you

A map style is a program, and the moment customers can write one you have an untrusted-code execution product whether you planned on building one or not. The failure you will actually hit is not a dramatic escape; it is one tenant's expression bomb flattening a shared pool on a Tuesday, or a VRT quietly reading your instance metadata into a PNG. One arrangement answers both: a warm microVM per tenant holding that tenant's artefacts, no route out of it, `fork_tree` for the viewport burst, a snapshot so waking costs 179 ms instead of a warm-up, and a CDN in front doing the actual work. The renderer stops serving your maps and starts filling your cache, which is a much better job for a process you do not fully trust.

Templates, limits and the rest of the platform surface are on features, and what it costs to keep a long tail of tenants asleep is on pricing.

Frequently asked questions

Can I just validate the style JSON before rendering instead of sandboxing?

You should validate, and it will not be enough, and the reason is interesting. Schema validation catches malformed documents and unknown properties, which is genuinely useful — it stops typos and it stops a tenant pointing `sprite` at an `http://` URL if you reject non-local URLs outright. What it cannot do is bound cost, because style expressions are a small interpreted language and the cost of evaluating one depends on data the style does not contain. A two-thousand-stop `interpolate` is cheap over an empty tile and ruinous over a dense one; a layer with no `minzoom` is harmless if the source has a `maxzoom` in its TileJSON and catastrophic if it does not. To know the cost you would have to evaluate the expression against the actual features in the actual tile, which is the render. You have arrived back at needing a bounded execution environment, having written a validator on the way. So do both, in that order: a strict schema pass and a URL allowlist at save time, because they are cheap and they catch the honest mistakes; then a hard runtime boundary, because the schema has no opinion about complexity and the complexity is where the cost lives. The boundary is also where you put the ceilings a schema cannot express — a CPU limit, an address-space limit, a maximum output pixel count, and a wall-clock deadline enforced in-guest by `timeout` rather than by a library you are hoping cooperates.

How do I keep per-tenant VM costs sane with thousands of tenants?

By making the idle case cost nothing, because with a long tail of tenants the idle case is almost all of them. The arrangement is: cache aggressively so most requests never reach a renderer; bake a snapshot per tenant per style version so the warm state is durable artefact rather than a running machine; and delete the VM when the tenant goes idle rather than parking it. On PandaStack that last part is scale-to-zero in the literal sense — the guest is gone, its memory and disk live in object storage, and the next cache miss restores it. Restore is p50 179 ms, p99 around 203 ms, and with UFFD streaming the guest starts answering before its whole memory file has arrived, pulling 4 MiB chunks on demand with a prefetch trace replaying the hot set behind it. The two things that will actually surprise you in the bill are, first, that per-tenant sizing is per-tenant templates because Firecracker cannot change RAM at restore, so a careless default of 4 GiB everywhere is the single biggest lever you have; and second, that a per-tenant VM is pinned to a host, so your capacity planning is about concurrent-awake tenants rather than total tenants. Measure the concurrency, not the roster. If a hundred tenants are ever awake at once out of two thousand, you are buying a hundred machines' worth of RAM plus headroom, not two thousand.

Vector tiles move rendering to the browser. Does that make this moot?

It moves one of the two risks and leaves the other, and it leaves a surprising amount of raster work behind. With MVT the client evaluates the style, so the expression bomb becomes the end user's laptop fan rather than your tile pool — a real improvement, and the main reason the industry went that way. What stays on your side is the tiler: you are still running `ST_AsMVT` over PostGIS, or `tippecanoe`, or `ogr2ogr`, against customer data with customer-supplied layer definitions and filters, which is the same attacker-controlled-parser-dispatch problem wearing a different hat. The VRT and `/vsicurl/` class is entirely unaffected, because that lives in the data path and not the style path. And you will still need server-side raster: static map images for emails and reports, PDF and print export, Open Graph preview images, thumbnails of saved views, and the fallback for clients that cannot run WebGL. In practice most serious mapping products run both paths, which means the honest architecture is a client-rendered interactive map plus a per-tenant server-side renderer for everything that has to come out as pixels. The per-tenant boundary is for the second one, and it is also where the tiler belongs.

Does fork_tree give me per-request isolation between tiles?

No, and it is important not to let the convenience of it blur into a security claim. `fork_tree` children inherit the parent's running memory, which is precisely why they are useful — they come up with the style parsed, the glyphs shaped and the dataset handles open, so a viewport batch gets parallelism without paying warm-up per tile. But inheritance cuts both ways: if the parent's heap was already corrupted by a hostile input, every child inherits the corruption, and if the parent leaked something into memory, every child has it. So treat `fork_tree` as a performance primitive inside one tenant's trust domain, never as a boundary between tenants. The boundary between tenants is a separate VM with a separate kernel. There are two operational details worth knowing: the cap is 16 children, so it sizes a viewport batch and not a traffic spike — for a spike you want more parent VMs rather than a deeper tree. And because children inherit the parent's running memory, they inherit its seeded PRNG state and will agree on their next random numbers, which is harmless for deterministic tile output and a genuine bug if anything in the child mints an identifier. `fork()` is the opposite trade: disk-only clone, child cold-boots with its own entropy and clock, useful for batch work and useless when warm state was the whole point.

Is a 179 ms create fast enough to spin up a VM per tile request?

No, and you should not want it to be. A tile budget is tens to low hundreds of milliseconds end to end, so adding a 179 ms create in front of it is bad arithmetic no matter how good the create is, and a viewport pan would be dozens of those at once. The per-tenant warm model exists specifically so that create is not in the request path. What is in the request path, for a tenant who has been asleep, is a wake — and the same restore mechanics apply, which is why the wake is a few hundred milliseconds rather than the three seconds a cold boot of a fresh template costs before auto-bake has produced a snapshot. The design goal is for that wake to be paid by a request nobody is watching: a pre-rendered low-zoom pyramid served from the edge covers the first paint, `stale-while-revalidate` covers the first deep-zoom miss, and a prefetch of the tiles just outside the viewport warms the machine before the user pans into them. Where a create-per-request shape does make sense is the batch half of a geospatial product — one guest per upload, per clip, per orthomosaic job — where the unit of work is seconds to minutes and the isolation you want is per job rather than per tenant. Two lifecycles, two boundaries, same substrate.

Keep reading

Related posts

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.