A Font Is a Program: Rendering User-Uploaded Typefaces Without Trusting Them
There is a feature request that shows up in every design tool, every certificate generator, every thumbnail service, every e-commerce product customiser and every video title renderer, and it always sounds trivial: let the customer upload their own font. They have a brand typeface. They licensed it. They want their name on the mug in it. The ticket gets pointed at a two, somebody adds a file input and an `ImageFont.truetype()` call, and the sprint ends.
The thing that just landed in your worker is not an image and it is not a config file. It is an executable. TrueType glyphs carry instructions for a stack-based virtual machine — the hinting VM — and FreeType interprets them. CFF/Type 2 charstrings are a second, different interpreter with arithmetic, storage and conditional operators. OpenType layout is a substitution and positioning machine that HarfBuzz executes against your text. And if the upload ended in `.woff2`, all of that sits behind a Brotli decompressor plus a lossless table transform the decoder has to reconstruct before anything else runs.
Font parsers are among the most heavily attacked pieces of code in the history of desktop software. The reason is structural rather than anecdotal: fonts are complex binary formats with nested offset tables and embedded interpreters, they are parsed automatically on untrusted content, and for decades they were parsed inside privileged processes — browsers and, famously, kernel-side font engines. That combination produced a long public history of memory-safety bugs. Both FreeType and HarfBuzz are, to their enormous credit, continuously fuzzed and have been for years. That is why you should take the format seriously, not why you should relax: continuous fuzzing is what maintained, heavily-attacked parsers look like.
I build PandaStack, an open-source Firecracker microVM platform, so the conclusion I am heading for is not a surprise. But the argument stands on its own and the specifics are the useful part: what actually executes inside a font file, why the mitigations everyone reaches for are partial, why the failure you will actually hit is denial of service rather than remote code execution, and what the per-render microVM shape costs — because it is not free, and anyone who tells you it is has not run it.
What is actually inside a font file that executes
A `.ttf`/`.otf` is an sfnt container: a 12-byte header giving a version tag and a table count, then 16 bytes per table — tag, checksum, offset, length — then the tables. Everything interesting is reached through an offset, which is the first thing to notice. A font is a graph of offsets into a blob, and a parser that trusts an offset is a parser you can aim.
The TrueType hinting VM
TrueType glyph outlines live in `glyf`, indexed by `loca`. Each glyph may carry instructions for a stack-based virtual machine whose job is to nudge points onto the pixel grid at small sizes. There is a font-wide program in `fpgm` that defines functions, a per-size pre-program in `prep` that runs at every size change, and a control-value table in `cvt `. The instruction set has arithmetic, storage, function calls, conditionals and loops. FreeType's TrueType bytecode interpreter is a real interpreter for a real instruction set, and the `maxp` table is where the font declares how much stack, storage and how many function definitions it intends to use — numbers an implementation may well size allocations from.
So: a per-glyph program, defined by the uploader, executed by your web worker, every time a character is rendered at a new size.
CFF and Type 2 charstrings: a second interpreter
An OpenType font with the `OTTO` tag stores outlines in a `CFF ` or `CFF2` table as Type 2 charstrings. This is not the same machine as TrueType hinting; it is its own stack-based language, with local and global subroutines, hint masks, and an operator set that includes arithmetic, conditionals and a small storage area. A glyph is a program here too, just in a different language, with a different interpreter, and the subroutine indirection gives you another offset graph to walk. If you were hoping "we only accept OTF" was a mitigation: it swaps one interpreter for another.
OpenType layout: GSUB, GPOS and the shaping machine
Outlines are only half of it. `GSUB` and `GPOS` describe substitutions and positioning as chains of lookups — a ligature lookup replaces a glyph sequence with one glyph, a contextual chaining lookup fires only in a specific neighbourhood, and lookups can cascade. HarfBuzz runs that machine over your text. It is a general enough mechanism that people write intentionally absurd fonts with it for fun, and the shaping run for one short string can expand into a great deal of work depending entirely on tables the uploader wrote.
Variable fonts, colour fonts, and the woff2 front door
Variable fonts add `fvar` (axes), `gvar` (per-glyph point deltas with packed point numbers and inferred intermediate deltas) and the item variation stores behind `HVAR`/`MVAR`. Delta decoding is intricate, bounded by counts in the file, and has to be done before any outline exists. Colour fonts add `COLR` — in its v1 form a graph of paint nodes with gradients and composites, which the spec requires you to cycle-check, meaning the obvious naive implementation loops forever. There is also an `SVG ` table, which embeds actual SVG documents, so a font can hand you an XML parser and a vector renderer as a bonus.
And `woff2` puts a Brotli decompressor in front of all of the above, plus a reversible transform of `glyf`/`loca` that the decoder must reconstruct. That reconstruction is a parser in its own right, and it is the first code to touch the bytes — before any of your table validation gets a turn, unless you decompress in a place you are willing to lose.
Where that parser is running right now
In almost every implementation I have seen, in-process. `ImageFont.truetype(path, size)` in Pillow is FreeType inside your Python interpreter. `registerFont()` in node-canvas is FreeType inside your Node event loop. `fontTools` parses the tables in Python, which at least gets you memory safety for the table walk, and then you hand the file to a rasterizer anyway. resvg, Skia, Cairo, a headless Chromium, a PDF pipeline's embedded font subsetter — all the same shape: a library call, in the request path, in the process.
That process holds the database connection pool. It holds the object-storage credentials for every tenant's bucket prefix, because that is how you wrote the uploader. It holds an internal service token with more scope than you remember granting it, and in a container it very likely holds a Kubernetes service account token mounted at a well-known path. Those are the things in reach of the uploaded program, and "in reach" is doing a lot of work in that sentence — not because exploitation is easy, but because the trust boundary is nowhere in the picture. There is no boundary between the font and the credentials. There is a function call.
A container does not add one. Namespaces and cgroups make the parser see less and bound how much it can allocate, which is genuinely useful and which you should do regardless — but the syscall interface it talks to is the same host kernel every other tenant's pod talks to. A container is a polite suggestion to a shared kernel. If you want the honest framing of what the various options do and do not stop, see MicroVM vs VM vs Container: A 2026 Comparison.
Why the usual mitigations are partial
The first thing a reviewer suggests is to turn the dangerous part off. It helps. It helps less than it sounds, and it is worth being precise about why.
FreeType gives you `FT_LOAD_NO_HINTING`, which skips the bytecode, and `FREETYPE_PROPERTIES=truetype:interpreter-version=40` selects the minimal interpreter that honours only a subset of instructions. You can even build FreeType without `TT_CONFIG_OPTION_BYTECODE_INTERPRETER` at all. Every one of those shrinks the bytecode surface. None of them removes the parsing surface, and the parsing surface is the bigger one: you still parse the sfnt directory, `glyf` and `loca`, `cmap`, `CFF ` charstrings (the bytecode flags do not touch those), `gvar` deltas, and whatever the woff2 decoder reconstructed. Turning off hinting means the font's programs do not run. It does not mean the font is not read.
seccomp-bpf is the right tool for a rasterizer, and the catch is architectural rather than technical. A seccomp filter is applied to a process, and it is only a boundary if the process on the wrong side of it has nothing you care about. That means forking a dedicated renderer that never held your secrets, passing the font and text in over a pipe, getting a bitmap back. If your renderer is a standalone binary you control — a Rust service that does nothing but rasterise — this is excellent and you should do it. If your font call is `ImageFont.truetype()` three stack frames into a Django view, with the ORM, the S3 client and the Stripe key all live in the same heap, you cannot usefully seccomp it, because the thing you would be confining is your application. In a Python or Node monolith that is the normal case, and "isolate the library" turns into a rewrite.
Resource limits are the mitigation people trust most and they are the weakest against this input class. `RLIMIT_AS` and a cgroup memory cap stop an allocation blowup from taking the host, which is real. They do nothing about CPU burned inside a legal computation: a variable font with pathological delta sets, a shaping run with cascading contextual lookups, a `prep` program that runs at every size change. That work is not an allocation. It is your renderer, doing exactly what the file told it to, for a very long time.
The failure you will actually hit is denial of service
I want to be honest about threat ranking, because security posts tend to lead with code execution and then people deploy nothing. Remote code execution through a font parser is the severe outcome. Denial of service is the likely one, and it is likely to the point of being routine, because it does not require a bug at all. It requires a valid font.
- Glyph outlines with absurd point counts. A legal glyph can carry an enormous number of contours and points; scaling and rasterising it is legitimate work that takes as long as it takes.
- An enormous `unitsPerEm`. The `head` table declares the design grid. Combine a large upem with a large requested pixel size and the scaled coordinate space gets big fast — and if any code path sizes a buffer from the font's own metrics rather than from your output dimensions, that is an allocation the uploader chose.
- A colour font with thousands of layers, or a COLRv1 paint graph designed to be walked repeatedly. One character, many composited draws.
- A shaping run that expands. Ligature and contextual chaining lookups can turn a short string into a lot of substitution work, and the expansion factor is a property of the uploaded tables.
- A `prep` program that is slow and runs on every size change. Your thumbnail pipeline renders at six sizes, so you run it six times.
- woff2 as a compression bomb. Brotli expands; the glyf reconstruction then builds tables from the expanded stream. You have to bound the decompressed size, and you have to do it in the decompressor, not after.
In a shared worker, every one of those is a pod-level event rather than a request-level one. The async worker that was handling eight render jobs concurrently now has one job spinning a core and climbing in RSS; the other seven tenants' certificates miss their SLA, the readiness probe fails, the orchestrator restarts the pod, and the queue redelivers the poison font to the next worker, which also dies. I have watched this exact loop in a thumbnail pipeline — different file format, identical shape — and the thing that makes it so unpleasant is that it looks like an infrastructure problem for the first hour. The same dynamic for archives is written up at Extracting Untrusted Archives: Zip Bombs, Zip Slip, Symlink Escape.
The microVM shape: font in, bitmap out, VM gone
The design is boring, which is the point. One VM per render job. Inside it goes our render script and the untrusted font bytes, plus the text and the size. Out of it comes a PNG. Then the VM is deleted, and with it the kernel that parsed the font, the page tables, the heap, the `/tmp` file, and any process the font managed to start.
What makes this practical rather than theoretical is that the create is a snapshot restore, not a boot: p50 179 ms, p99 203 ms, on every create, with no warm pool of idle VMs to pay for. The mechanics are in The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms. The first spawn of a template that has no snapshot yet is a cold boot at around 3 s, and after that you are on the restore path.
# Render one line of text with ONE untrusted font, in its own kernel.
# The only artefact that leaves the VM is a PNG.
import hashlib, json, pathlib
from pandastack import Sandbox
font_bytes = pathlib.Path("uploads/tenant-7f/Brand-Display.woff2").read_bytes()
digest = hashlib.sha256(font_bytes).hexdigest()
# Cheap pre-checks ON THE HOST, on bytes only. We are allowed to look at the
# first four bytes and the length. We are NOT allowed to parse -- every real
# check happens in the guest, because parsing IS the dangerous part and doing
# it here would defeat the whole exercise.
if len(font_bytes) > 8 * 1024 * 1024:
raise ValueError("8 MiB cap: no legitimate text face is bigger")
if font_bytes[:4] not in (b"\\x00\\x01\\x00\\x00", b"OTTO", b"true",
b"wOF2", b"wOFF"):
raise ValueError("not an sfnt or WOFF container")
sbx = Sandbox.create(
template="base", # 8 vCPU / 4 GiB, baked into the snapshot:
# Firecracker cannot change vCPU/RAM at restore,
# so per-create cpu/memory_mb are overridden to
# the baked values. You get 4 GiB whether or not
# you asked for it.
ttl_seconds=120, # IDLE clock, not a wall clock. Guest-touching
# calls reset it; status GETs deliberately do
# not. A font that wedges the renderer gets
# reaped -- the reaper DELETES the VM, which for
# this job is the desired outcome, not a risk.
metadata={"job": "font-render", "tenant": "7f", "font_sha256": digest},
)
try:
# Our code first, then the untrusted bytes, byte-for-byte untouched.
sbx.filesystem.write("/work/render.py",
pathlib.Path("render.py").read_text())
sbx.filesystem.write("/work/font.bin", font_bytes)
sbx.filesystem.write("/work/job.json", json.dumps(
{"text": "Certificate of Completion", "px": 96,
"width": 1600, "height": 400}))
# In production you bake pillow/fonttools/brotli into a snapshot and
# create from_snapshot=..., so this install is not in the request path.
# Shown once so the sample actually runs.
sbx.exec("timeout --kill-after=10s 240 sh -c 'export "
"PATH=/opt/mise/shims:$PATH; pip install -q pillow fonttools "
"brotli'", check=True)
# THE FENCE THAT MATTERS. One-shot exec() does not actually enforce
# timeout_seconds, and exec_stream only widens the HTTP timeout -- so a
# hostile shaping run is bounded in-guest by timeout(1) or not at all.
# 124 = timed out, 137 = ignored SIGTERM and needed --kill-after's KILL.
res = sbx.exec("cd /work && timeout --kill-after=10s 20 sh -c 'export "
"PATH=/opt/mise/shims:$PATH; python3 render.py'")
if res.exit_code in (124, 137):
raise TimeoutError(f"font {digest[:12]}: renderer ran away, rejected")
if res.exit_code != 0:
raise ValueError(f"font rejected: {res.stderr.strip()[:300]}")
png = sbx.filesystem.read("/work/out.png") # bytes, and only bytes
pathlib.Path(f"cache/{digest}-96-certificate.png").write_bytes(png)
finally:
sbx.kill() # teardown is kill(); there is no sbx.delete()
The in-guest script is where the validation lives, and it fails closed: anything surprising is a non-zero exit, not a best-effort render. Every check below exists because something in the format lets the uploader choose a number you would otherwise trust.
#!/usr/bin/env python3
# /work/render.py -- runs INSIDE the microVM. Validate, then rasterise.
# Fails closed: any surprise is a non-zero exit, not a best-effort render.
import io, json, struct, sys
from fontTools.ttLib import TTFont
MAX_TABLES = 64 # real faces ship ~15-30 tables
MAX_GLYPHS = 20000 # a full CJK face is ~20-30k; a Latin brand face ~300
MAX_UPEM = 16384 # 1000 or 2048 is normal; big upem x big px = big coords
MAX_AXES = 8 # variable-font axes
MAX_PIXELS = 4000 * 4000 # OUR cap, from OUR job, never from font metrics
def die(msg):
print(msg, file=sys.stderr)
sys.exit(2)
job = json.load(open("/work/job.json"))
raw = open("/work/font.bin", "rb").read()
# 1. woff2 is Brotli plus a glyf/loca TRANSFORM the decoder reconstructs -- a
# parser in front of the parser. Normalise to a plain sfnt HERE, so that
# FreeType only ever sees the output of that.
if raw[:4] == b"wOF2":
f = TTFont(io.BytesIO(raw))
f.flavor, out = None, io.BytesIO()
f.save(out)
raw = out.getvalue()
# 2. The sfnt directory, by hand, before any library sees it: a font is a
# graph of offsets, and one running past EOF is the oldest trick there is.
if len(raw) < 12:
die("truncated")
tag, num_tables = struct.unpack(">4sH", raw[:6])
if tag not in (b"\x00\x01\x00\x00", b"OTTO", b"true"):
die(f"unsupported sfnt version {tag!r}") # 'ttcf' collections: refused
if not 1 <= num_tables <= MAX_TABLES or len(raw) < 12 + 16 * num_tables:
die(f"bad table directory ({num_tables} tables)")
seen = set()
for i in range(num_tables):
t, _chk, off, ln = struct.unpack(">4sLLL", raw[12 + 16 * i:28 + 16 * i])
if off + ln > len(raw) or off < 12:
die(f"table {t!r} points outside the file")
if t in seen:
die(f"duplicate table {t!r}") # ambiguity is the attack
seen.add(t)
if b"SVG " in seen:
die("SVG table: an embedded XML + vector renderer, not today")
# 3. Now fontTools, lazily, for the counts the rasteriser will act on.
font = TTFont(io.BytesIO(raw), lazy=True, fontNumber=0)
if font["maxp"].numGlyphs > MAX_GLYPHS:
die(f"{font['maxp'].numGlyphs} glyphs")
if not 16 <= font["head"].unitsPerEm <= MAX_UPEM:
die(f"unitsPerEm {font['head'].unitsPerEm}")
if "fvar" in font and len(font["fvar"].axes) > MAX_AXES:
die("too many variation axes")
if "COLR" in font:
die("colour font: the layer/paint-graph walk is unbounded work")
# 4. Our canvas, our numbers. Nothing below is read from the font.
w, h, px = int(job["width"]), int(job["height"]), int(job["px"])
if w <= 0 or h <= 0 or w * h > MAX_PIXELS or not 4 <= px <= 512:
die("job geometry out of range")
open("/work/font.ttf", "wb").write(raw)
# layout_engine=BASIC keeps HarfBuzz out of it -- no GSUB/GPOS machine, no
# shaping expansion. The price is complex scripts: Arabic and Devanagari need
# Layout.RAQM, which means you need this VM more, not less.
from PIL import Image, ImageDraw, ImageFont
face = ImageFont.truetype("/work/font.ttf", px,
layout_engine=ImageFont.Layout.BASIC)
img = Image.new("RGBA", (w, h), (255, 255, 255, 0))
ImageDraw.Draw(img).text((40, h // 3), job["text"][:200], font=face,
fill=(17, 17, 17, 255))
img.save("/work/out.png", optimize=False)
Caching, because a VM per render is not free
Here is the honest economics. A restore is p50 179 ms, and `base` is 8 vCPU with 4 GiB baked in. On the one rate card — $0.054 per vCPU-hour and $0.0162 per GiB-hour — CPU bills on CPU-seconds actually burned, so a render that burns a second or two of CPU costs a small fraction of a cent, while memory bills committed GiB-hours for as long as the sandbox exists. Four GiB for two seconds is about 0.0022 GiB-hours, which is arithmetic rather than a benchmark, and it is cheap. What it is not is free, and 179 ms is still 179 ms on the front of a request. Current numbers are on pricing.
So cache the output, not the font. Key on the tuple that determines the bitmap: the SHA-256 of the exact font bytes, the text, the pixel size, and any variation-axis coordinates. A certificate generator rendering the same customer's name in the same brand face at the same size should hit that cache forever and never start a VM at all. Deduplicating on the font hash across tenants is fine and useful — two tenants who uploaded the same licensed face get the same cached bitmaps — as long as the cache key includes the text, because the text is tenant data and the font hash is not a tenant secret.
The second lever is a warm snapshot. Install Pillow, fontTools and brotli once, import them, let Python and FreeType's shared library be resident, then `snapshot()` and create every render VM `from_snapshot=...`. You are restoring a process that has already done its imports, so the guest-side work is the render and nothing else.
For batch work — five hundred certificates, one brand font — there is a third lever, and it is the one people reach for incorrectly.
# 500 certificates, ONE tenant's font, one warm renderer.
#
# fork_tree() snapshots the parent ONCE and boots the children from that
# snapshot, so each child inherits the parent's MEMORY: Python up, Pillow
# imported, FreeType's shared object resident, the validated font already
# on disk. The parent is untouched (pause -> snapshot -> resume).
#
# fork() is the OTHER call and it is not this one: disk clone only, and the
# child COLD-BOOTS, so every worker re-imports everything from nothing.
# For warm in-memory state it is always fork_tree.
#
# Hard cap: 16 children per fork_tree() call -- more is an error, not a
# clamp -- so a wider fan-out grows the tree breadth-first.
def fan_out(root, total):
frontier, kids = [root], []
while total > 0 and frontier:
batch = frontier.pop(0).fork_tree(min(16, total))
kids += batch
frontier += batch
total -= len(batch)
return kids
# The parent has ALREADY validated and loaded tenant 7f's font. So every
# child is still inside tenant 7f's blast radius and nobody else's -- which
# is the rule: never fan out one parent across tenants. One font per tree.
workers = fan_out(warm, 32)
try:
for w, name in zip(workers, NAMES):
w.filesystem.write("/work/job.json", json.dumps(
{"text": name, "px": 96, "width": 1600, "height": 400}))
if w.exec("cd /work && timeout --kill-after=10s 20 sh -c 'export "
"PATH=/opt/mise/shims:$PATH; python3 render.py'"
).exit_code == 0:
save(name, w.filesystem.read("/work/out.png"))
finally:
for w in workers:
w.kill()
One caveat worth knowing before you use that pattern for anything cryptographic: children of a `fork_tree` inherit the parent's memory, which includes the state of any userspace RNG that was already seeded. For deterministic rasterisation that is irrelevant. For anything that needs unique randomness per child it is not, and the distinction between the two fork calls is covered properly in Snapshot vs restore vs fork vs clone, explained.
Isolation options for this specific job
| Approach | What it stops | What it does not | Honest drawback |
|---|---|---|---|
| In-process library call (Pillow, node-canvas) | Nothing | A parser bug lands in the process holding your DB pool and object-storage keys; one runaway shaping run takes every concurrent job with it | Free, fast, and the trust boundary is a function call |
| Separate process + rlimits | Crashes and allocation blowups stay in the child | The kernel surface; and a child forked from your app inherits the secrets your app had already read | Only a boundary if the child never held anything of yours — which means real IPC, not a fork |
| seccomp-bpf around a standalone rasterizer | Most of what a successful exploit would want to do next | The parse itself, and nothing at all if the font call sits mid-request in a monolith | Excellent when the renderer is a binary you control; usually unavailable in a Django or Express process |
| Container (namespaces + cgroups) | Filesystem and process-tree visibility; cgroups bound the memory blowup | Same kernel as every other tenant's pod; a service-account token is still mounted at a well-known path | The best mitigation that does not move the trust boundary at all |
| gVisor | Most of the host kernel surface, by re-implementing syscalls in userspace | It is not a hardware boundary, and syscall-heavy work pays a tax | Real isolation with a real compatibility and performance cost |
| WASM-compiled rasterizer | Memory-safety escapes — a heap overflow stays inside the linear memory | CPU and time blowups, unless you add fuel metering yourself | Genuinely strong for exactly this job, and you need a WASM build of your whole text stack |
| MicroVM, one per render job | The parser gets its own kernel; the only artefact that leaves is a bitmap, and then the VM is deleted | Egress unless you fence it; and it does not make your own job code correct | A create and a VM lifetime per job, so you cache the output and warm the renderer |
Two of those rows deserve more than a cell. The gVisor trade-off is written up properly in PandaStack vs gVisor: choosing your isolation boundary, and the WASM one in WebAssembly vs MicroVMs for Sandboxed Code — and I will say plainly that a WASM-compiled rasterizer is the most elegant answer to this specific problem if your stack can get there, because memory-safety bugs genuinely cannot escape a linear memory. The reason I do not build on it is coverage: the moment you need HarfBuzz shaping, a colour-font path, a PDF subsetter and an image encoder in the same pipeline, you are maintaining a WASM build of a large C and C++ dependency tree, and the microVM version is the same code you already run. If you want the general framing of the boundary itself, What is a microVM? Firecracker, isolation, and why agents need it is the primer.
What this does not give you
- No GPU. There is no PCI or VFIO passthrough, so GPU-accelerated text rasterization does not run — not Skia on a GPU backend, not a Vulkan path. Software rasterizers work; llvmpipe and lavapipe run, slowly. If your renderer currently depends on a GPU, this is a port, not a config change.
- Egress is open by default. Per-sandbox network namespaces mean no sandbox reaches a sibling's subnet and the cloud metadata range is dropped at the host, but your own VPC, databases and internal APIs are not fenced for you. A font cannot usually make outbound connections on its own — but your render script can, and so can anything that gets code execution through it. Fence it: Controlling Network Egress for Untrusted Code.
- The guest kernel is 5.10. Most text stacks do not care. Find out whether yours does before you build on it — You Ship the Kernel: Firecracker Guest 5.10 vs 6.1 has the detail.
- You cannot shrink the VM per job. vCPU and RAM are baked into the template snapshot and Firecracker cannot change them at restore, so a `base` render VM is 8 vCPU and 4 GiB whether the job needs it or not. That is a floor on the per-job memory bill, and the lever is to shorten the VM's life, not to narrow it.
- A microVM does not validate anything. Everything in the in-guest script above is still your job, and the reason to write it carefully is not that an escape is likely but that garbage output and 20-second renders are certain without it.
- It does not fix the webfont path, which is the mistake I see most. If you also serve that uploaded `.woff2` to end users' browsers, you have handed the font to every visitor's FreeType, DirectWrite or CoreText — a much better-hardened parser than yours, and still not one you should be aiming a stranger's file at on your customers' behalf. Subset and re-emit the font from inside the same VM, or serve rendered images, or accept that you are a font CDN now.
The summary
"Let users upload a font" is a request to execute a stranger's program. Not metaphorically: TrueType glyphs carry bytecode for a stack machine that FreeType interprets, CFF glyphs are programs in a second language with a second interpreter, `GSUB`/`GPOS` are a substitution machine HarfBuzz runs over your text, `gvar` deltas and COLRv1 paint graphs are intricate decoders, and `woff2` puts Brotli plus a table reconstruction in front of all of it. Font parsers have been among the most-attacked code in software for decades for structural reasons, and the fact that FreeType and HarfBuzz are continuously fuzzed is evidence of the pressure, not a reason to relax.
Turning hinting off shrinks the bytecode surface and leaves the parsing surface. seccomp is the right answer if your rasterizer is a standalone binary and no answer at all if it is a library call inside your web process. Resource limits stop the allocation blowups and not the CPU ones, which are the blowups you will actually hit, because a pathological font does not need a bug — it needs to be valid.
So my recommendation, with its limits attached. If you render a fixed set of fonts you shipped yourself, do none of this: validate the upload path that could change them and move on. If you accept fonts from users and render server-side, put the parser in its own kernel — one microVM per render job, validated in-guest, bounded with `timeout --kill-after`, PNG out, VM deleted — and then make it affordable by caching on (font hash, text, size) and restoring a warm renderer rather than installing one. If your whole text pipeline can compile to WASM, that is a genuinely better fit for this one job and you should take it. And if you are at one upload a week from three trusted enterprise customers, a separate process with rlimits and `FT_LOAD_NO_HINTING` is a reasonable place to stop — just write down that you stopped there, because the day somebody ships self-serve font upload, that decision is still in production and nobody will remember making it.
Frequently asked questions
Is a font file really executable code, or is that just a figure of speech?
It is literal. A TrueType glyph can carry instructions for a stack-based virtual machine — the hinting VM — whose purpose is to adjust outline points onto the pixel grid. The instruction set includes arithmetic, storage, conditionals, loops and function calls; functions are defined in the `fpgm` table, a pre-program in `prep` runs on every size change, and FreeType ships an interpreter for all of it. CFF/Type 2 charstrings are a second, different interpreter: each glyph is a little program with local and global subroutines, hint masks and an operator set that includes arithmetic and conditionals. On top of outlines, `GSUB` and `GPOS` describe chains of substitution and positioning lookups that HarfBuzz executes against your text, which is a general enough machine that people build novelty fonts out of it. Variable fonts add delta decoding; COLRv1 adds a paint graph the spec requires you to cycle-check; the `SVG ` table embeds actual SVG documents. And a `.woff2` is all of the above behind a Brotli stream plus a reversible `glyf`/`loca` transform the decoder has to reconstruct first. So when you accept a font upload you have accepted several interpreters and several decoders, all driven by data the uploader controls.
Can I just disable hinting — FT_LOAD_NO_HINTING, or interpreter-version=40 — and skip the sandbox?
You can disable it, and you should know what you kept. `FT_LOAD_NO_HINTING` skips the bytecode. `FREETYPE_PROPERTIES=truetype:interpreter-version=40` selects the minimal interpreter that honours only a subset of instructions, and you can build FreeType without `TT_CONFIG_OPTION_BYTECODE_INTERPRETER` entirely. Each of those shrinks the bytecode surface, which is real and worth doing. None of them removes the parsing surface, which is larger. With hinting fully off you still parse the sfnt table directory and every offset in it, `loca` and `glyf`, `cmap`, `CFF ` charstrings — the TrueType hinting flags do not touch the CFF interpreter at all — `gvar` delta sets, and whatever a woff2 decoder reconstructed before your code got a turn. You also keep the entire denial-of-service class, because that does not need the bytecode: glyphs with enormous point counts, a large `unitsPerEm` against a large pixel size, cascading contextual lookups in shaping, a colour font with thousands of layers. Those are legal fonts doing legal work, slowly, for as long as the uploader designed. Disabling hinting is a good defence-in-depth measure inside the sandbox. It is not the sandbox.
Isn't a container enough? My renderer already runs in its own pod.
A container is a meaningful mitigation and it is not a boundary change. Namespaces mean the font parser sees a small filesystem and a short process list; cgroups mean an allocation blowup hits a memory limit rather than the host; dropping capabilities and running non-root removes a lot of what an exploit would do next. Do all of it. But the syscall interface that process talks to is the host kernel, shared with every other pod on the node, and the historical reason font bugs were so severe is precisely that they reached kernel-side or privileged font code. Within a pod you also typically still have a Kubernetes service account token mounted at a well-known path, a node metadata endpoint unless someone blocked it, and network reach into your own VPC. And a container does nothing for the failure you will actually hit: a valid, pathological font that burns a core for thirty seconds takes out the readiness probe, the pod restarts, the queue redelivers the same font, and you have an outage that looks like an infrastructure problem for the first hour. A microVM changes the substrate — the parser gets its own guest kernel, and a per-job VM that wedges gets deleted rather than restarted.
A microVM per render sounds expensive. What does it actually cost, and how do I make it cheap?
The latency cost is a create, and on PandaStack every create is a snapshot restore rather than a boot: p50 179 ms, p99 203 ms, with no warm pool of idle VMs behind it. The first spawn of a template with no snapshot yet is a cold boot around 3 s, and after that you are on the restore path. The money cost is one rate card — $0.054 per vCPU-hour and $0.0162 per GiB-hour — where CPU bills on CPU-seconds actually burned and memory bills committed GiB-hours for as long as the sandbox exists. The `base` template is 8 vCPU and 4 GiB baked into the snapshot, and Firecracker cannot change those at restore, so a two-second render holds 4 GiB for two seconds: roughly 0.0022 GiB-hours, fractions of a cent. Cheap, not free. Three levers make it comfortable. Cache the rendered bitmap keyed on (font SHA-256, text, pixel size, variation coordinates) so repeat renders never start a VM. Bake your renderer into a snapshot — Pillow and fontTools already imported, FreeType resident — and create `from_snapshot` so the guest does nothing but render. And for batch work, `fork_tree` a warm parent, which inherits memory rather than cold-booting, capped at 16 children per call.
What about the fonts we serve to browsers as webfonts rather than rendering server-side?
That is the case people forget, and it is a different problem with a worse shape. If a tenant uploads a `.woff2` and you serve it from your CDN to every visitor of that tenant's page, you have not isolated a font parser — you have distributed a stranger's file to thousands of browsers and asked FreeType, DirectWrite, CoreText and Skia to parse it. Those are much better-hardened implementations than whatever is in your worker, with sandboxed renderer processes and years of fuzzing behind them, so the practical risk is lower. It is also no longer your risk to accept on your users' behalf, and it is not a decision most teams realise they made. Three reasonable options. Serve rendered images instead of the font, which is the safest and costs you selectable text. Subset and re-emit the font inside the same microVM — run fontTools' subsetter on the untrusted input and publish only the output, which drops `fpgm`, `prep`, the `SVG ` table and anything else you did not ask for, and means your CDN serves a file your code wrote rather than one a user uploaded. Or serve the original and write down that you chose to. The one thing not to do is serve the raw upload because the server-side render happened to succeed; a font that rasterises fine can still carry tables you never touched.
Keep reading
- Per-tenant thumbnail generation in microVMs — The same argument for image decoders, where the decompression-bomb version of this failure is even more routine.
- Processing untrusted PDFs in microVMs — PDFs embed fonts, so a PDF pipeline is a font pipeline with an extra parser in front of it.
- Untrusted archive extraction and zip bombs — The denial-of-service pattern in its purest form: valid input, no bug required, pod gone.
- Controlling network egress for untrusted code — Egress is open by default, and this is the piece of the design that is still your work.
- What is a microVM? — The primer on the boundary this whole post leans on, and how it differs from a container.
Related posts
- Rendering user-authored templates safely: SSTI and the microVM fix
You shipped a text formatter so customers could edit their own emails. Depending on the engine, you may also have shipped them a REPL on your application server.
- Server-Side Rendering Someone Else's React Component
If your product server-renders components your customers wrote, you are running their code in your process, with your env vars and your database socket. node:vm is a sandbox for values, not for a module graph that can require("fs"). Here is the boundary that actually holds, and what it costs.
- Compiling a Stranger's Shader: Isolate the Toolchain, Because You Cannot Isolate a GPU You Do Not Have
If your product accepts shader source from users or from a model, you are running a pile of C++ parsers on hostile input. PandaStack has no GPU and cannot run your shader on one — but glslang, spirv-val, spirv-opt and spirv-cross are all CPU work, and that is the part that actually hurts you.
- Running a Multi-Tenant Render Farm on MicroVMs: A .blend File Is a Program
If strangers send you scene files and render flags, you are not running a render farm. You are running a code-execution platform that happens to produce images, and the boundary you picked for "just rendering" is now load-bearing.
- A Checkpoint Is a Program: Isolating Untrusted Model Weights
You downloaded a checkpoint from a model hub and called torch.load. Congratulations: you ran a stranger's Python as yourself. Here's the threat model, and the disposable-microVM ingest pipeline that turns hostile weights into boring data.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.