Per-Tenant Audio and DSP Rendering in MicroVMs: The Plugin Is the Payload
I build PandaStack, an open-source Firecracker microVM platform, and audio keeps producing the same conversation in a different accent. Someone runs a hosted audio product — stem separation, AI mastering, loudness normalisation against a delivery spec, batch transcode of a back catalogue, a plugin-hosted mixdown service — and describes it the way a render farm describes itself. Customers upload audio, we process it, we hand back a file. Long jobs, CPU-saturating, billed per minute of audio.
Structurally that is a render farm, and most of the render-farm reasoning transfers: files from strangers, a large pile of C doing the parsing, a teardown problem. I wrote that version up at Running a Multi-Tenant Render Farm on MicroVMs: A .blend File Is a Program. Audio has one extra property the 3D world does not, and it decides the architecture.
On a render farm, the dangerous thing a customer uploads is data that happens to be interpretable as a program. On an audio platform it is frequently a program outright, and running it is the feature they pay for. A VST3, CLAP, AU or LV2 plugin is a native shared library. Hosting one means `dlopen`, resolve the module entry point, instantiate, and call into the stranger's code from a tight processing loop. Not a shortcut — the plugin ABI. Your customer's free analog warmth plugin is a shared object you agreed to run in a loop, chewing forty-eight thousand samples a second, in the process that holds your object-store credentials.
One boundary first. This post is offline rendering: read a file, burn CPU, write a file. Not realtime voice, which has a different latency budget and threat model — Isolating Real-Time AI Voice Agents Per Call. Not transcription, where the problem is model loading and audio residue between tenants: Per-Tenant AI Voice Transcription in MicroVMs.
The plugin is the payload, and the payload is the product
Start with the mechanics, because the security argument is an honest reading of them. On Linux a VST3 is a bundle directory with a native shared object inside it; the host loads that object, calls its module init, asks the factory for class metadata, and creates a processor. A CLAP is a shared object exporting a `clap_entry` symbol — simpler, no less native. An LV2 is Turtle manifests next to a `.so`. AU is an Apple framework, which matters later. All resolve to one sentence: a function pointer into memory you mapped from a file somebody else gave you.
Then the host does what a host does. Set the sample rate and maximum block size, activate, and call `process` once per block for the length of the material — offline, as fast as the machine allows, which is millions of calls into third-party native code per job. Along the way the plugin allocates its own memory, spawns its own threads, may read and write files, may open a socket, and shares your address space, your descriptor table and your floating-point control register.
There is a second input channel people miss. Plugin state is an opaque blob: a preset, a saved chain, a session recall. The host hands over bytes and the plugin deserialises them. So even when the binary is one you vetted yourself, that blob is attacker-controlled input to a deserialiser written by someone optimising for reverb tails, not hostile input. A curated plugin list narrows this; it does not close it.
Then licensing, where this stops being theoretical. Many commercial plugins validate a licence at instantiation: some read a file from a fixed path under the user's home directory, some expect a hardware dongle you will never have, and some phone home — an outbound HTTPS request, by a third-party shared library, from inside your infrastructure, as a precondition for multiplying numbers.
Out-of-process hosting: the right instinct, a weaker promise than it looks
Serious DAWs already host plugins out of process, and they are right to: a plugin that segfaults should not take down a four-hour session with unsaved automation in it. The mechanism is a separate process plus a ring buffer — blocks across, processed blocks back, and a crash becomes a disabled plugin instead of a lost afternoon.
But crash isolation and security isolation are different goals sharing one mechanism, and conflating them is how you end up with a confident diagram and a thin boundary. Crash isolation assumes the code in the box is buggy. Security isolation assumes it was chosen specifically so that you would run it.
What a separate process on a shared kernel genuinely gives you: it cannot corrupt your host application's heap directly, you can kill it, fence it with cgroups, run it as another user, point a seccomp filter at it. Worth having, and I would build it inside a VM too. What it does not give you is a smaller kernel. The plugin still makes syscalls into the same kernel every other tenant's work runs on, with exactly one kernel privilege escalation between it and everything on the box. A container is a polite suggestion to the kernel; a bare helper process is a slightly less polite one.
One happy asymmetry, and it is why this argument is easier for offline rendering than for a live DAW. A realtime host cannot afford much of a boundary: a few milliseconds per block, every copy spent against a deadline that clicks audibly when missed. An offline renderer has no deadline. It is throughput work, so it can afford a VM boundary and a copy per block — and the usual "we cannot isolate, latency" objection does not apply.
Even with no plugin, the input path is a decoder zoo
Suppose you never allow customer plugins. You still accept audio, and that path is one of the broadest parsing surfaces in commodity software. A general ingest touches FLAC, MP3, AAC in both ADTS and MP4 framing, Opus, Vorbis, WavPack, ALAC, AIFF, plain WAV with its adventurous chunk layouts, plus Speex and AMR on the long tail. Then the containers — MP4, Matroska, Ogg — and the tags stapled to them: ID3v2 with its own compression and unsynchronisation rules, APE tags, Vorbis comments. Embedded cover art means running an image decoder inside an audio job.
What makes this a security surface rather than a compatibility surface is that the header fields are integers chosen by whoever uploaded the file, and they feed allocation arithmetic: sample rate, channel count, bit depth, block size, frame length. The classic decoder bug has a shape you can recite — read a length field, allocate from it, read that many bytes, interpret them as a structure, recurse. Audio adds local flavours: a declared sample rate of two billion, zero channels, a channel count that disagrees with the layout mask, a resampling ratio that divides by zero.
To be fair to the libraries: ffmpeg, libFLAC, libsndfile and the Opus reference code are fuzzed continuously by people who are very good at this. The problem is volume, not diligence. A very large amount of C, reached by bytes you did not choose, times every tenant, times a year — you are betting on there being no exploitable bug across that surface, forever.
A filter graph is a small programming language
Here is the part that gets waved through in design review. Offline audio products eventually expose processing configuration, and the obvious way is to let advanced users write the chain. Great feature. Also an interpreter. An ffmpeg `-filter_complex` string is a directed graph of filters with labelled edges, parsed from text the customer wrote. Several of those filters take arithmetic expressions evaluated per sample, and an expression evaluator is an interpreter. Several read files from disk; several write them. So "we just let them tweak the EQ" has become "we accept a program, from a stranger, in a process that can see our filesystem."
You can lock this down: allowlist filter names, reject expression-capable filters, forbid anything taking a path, cap node count. Do all of it. But notice what you are doing — writing a validator for someone else's grammar, the activity with the worst historical record in our industry. First layer, not only layer.
A careful invocation, and why it is not the boundary
Here is roughly what I would run inside a per-job guest. Writing it out also shows what it cannot do.
#!/bin/sh
# Inside the per-job microVM: non-root, no credentials in the environment,
# nothing interesting anywhere outside /work.
set -eu
# HOME is where plugins scatter licence files, caches, crash dumps and
# "telemetry". Point it at the work dir so everything a stranger's shared
# library decides to write dies with the VM.
export HOME=/work/home
export XDG_CONFIG_HOME=/work/home/.config
export XDG_CACHE_HOME=/work/home/.cache
export VST3_PATH=/work/plugins/vst3
export CLAP_PATH=/work/plugins/clap
export LV2_PATH=/work/plugins/lv2
mkdir -p "$HOME" "$XDG_CONFIG_HOME" "$XDG_CACHE_HOME" /work/out
# Fences. A pathological input should fail in userspace with an exit code,
# not as a kernel OOM kill that takes the box down with it.
ulimit -t 1800 # CPU seconds
ulimit -v 4194304 # address space, KiB
ulimit -f 4194304 # output file size, blocks
ulimit -c 0 # no core dumps containing somebody's unreleased master
umask 077
# Decode, filter, mix, normalise. -nostdin so nothing can make the graph
# interactive; -xerror so a decoder warning is a failure, not a surprise.
exec timeout --signal=TERM --kill-after=30s 1800 \
ffmpeg -nostdin -hide_banner -xerror \
-analyzeduration 10M -probesize 10M \
-i /work/in/vocal.flac \
-i /work/in/drums.flac \
-filter_complex "[0:a]highpass=f=35[v];[1:a]alimiter=limit=0.97[d];[v][d]amix=inputs=2:normalize=0[m];[m]loudnorm=I=-14:TP=-1.0:LRA=11[out]" \
-map "[out]" -ar 48000 -c:a pcm_s24le \
-y /work/out/mixdown.wav
Pointing `HOME` and the XDG directories at the work tree is the highest-value line: it stops a plugin establishing persistence somewhere you will not think to look — a licence cache, a crash dump containing another tenant's audio, a config file that changes behaviour next run. `ulimit -t` turns a pathological time-stretch ratio from a host-wide memory event into an exit code you can show the customer. `-xerror` converts "warned about a corrupt frame, delivered a truncated master" into a failure you find first.
And none of it is a security boundary. It is configuration applied to a process, by a process, on a kernel that process speaks to directly. The limits constrain what the renderer does on purpose. They have no opinion about what it can be made to do by accident, which is the point of a parser bug.
The reasons that have nothing to do with security
You cannot tell a capacity problem from an escalation
Offline audio rendering is CPU-saturating and long-running by design. A benign, paying mastering job with a convolution reverb and four bands of oversampled dynamics pins every core you give it for minutes and sits at the top of your memory-bandwidth graph. That is what success looks like. A compromised worker running somebody else's arithmetic looks identical from outside the box.
So consider the page. It is 3am, a host is at full CPU, the queue is backing up, one tenant's job just got OOM-killed. Noisy neighbour, or escalation? You cannot tell: the only signal you have is saturated by normal operation. You deleted the ability to tell a capacity problem from a security problem at the architecture stage, months before the pager fired.
Audio has a nastier version than rendering does. On a render farm the tell is a frame that comes back entirely black. Here the output is a plausible file no matter what happened: the abusive tenant's deliverable is forty minutes of near-silence at the correct sample rate, delivered reliably, by a customer who never complains and whose CPU-seconds-per-audio-second ratio nobody checks. A microVM per job does not stop that — isolation is not abuse prevention — but it makes the resource profile belong to one tenant, so the remedy is a delete call and a flag on an org.
A mastering pipeline has to be bit-identical, and float does not cooperate
The second reason is determinism, which in audio is a product requirement rather than an engineering nicety. Customers A/B two masters and expect the difference to be the setting they changed. A label re-renders a delivery six months later and expects the same bytes. Your regression suite compares output hashes, because the alternative is listening to four hundred cases by ear. All of that needs the same input to give the same output on a different machine, and float DSP is startlingly bad at that.
Drift comes from three places. Denormals: a reverb tail decays into denormal numbers, which on some hardware are an order of magnitude slower, so DSP code mitigates by setting flush-to-zero and denormals-are-zero bits in the floating-point control register — per-thread state a plugin can set, after which your code rounds differently than it did in the test. Compiler flags: `-ffast-math` and `-Ofast` permit reassociation and reciprocal approximation, so the same source built with a different toolchain produces different last bits. Runtime CPU dispatch: FFT kernels, resamplers and convolution engines pick an AVX2, AVX-512 or NEON path from what the CPU reports. Hence the conclusion that surprises people: the same container image on a different host is not the same machine for a float pipeline.
A snapshot-restored microVM with a pinned guest userland gets materially closer. On PandaStack every create restores a baked template snapshot, so the whole userland is one artefact — the ffmpeg build, libsndfile, the resampler, the plugin host, Ubuntu 24.04 and guest kernel 5.10 — identical for every job until someone re-bakes. That removes the drift class that bites in production: the host with a different libFLAC minor version, the node still running last quarter's build. General version at Reproducible Builds in Disposable MicroVMs.
Resource shape: real-time factor is the number people actually ask about
Every audio platform's capacity planning, pricing and SLA reduce to one number: real-time factor. Wall-clock seconds of compute per second of audio, or its reciprocal as x-realtime. It is the figure in the sales conversation and the first thing to degrade when density is wrong — and the intuition that audio rendering just wants more cores is only half right.
Audio DSP is more memory-bandwidth and cache bound than its CPU graph suggests. A convolution reverb streams a long impulse response against a long signal. FFT-based time-stretching moves large buffers through cache every hop. Oversampled saturation multiplies the sample count by four or eight before any arithmetic. Streaming kernels compete for bandwidth, not ALU throughput, so your real-time factor on a loaded host is worse than on an idle one at the identical CPU share — the degradation behind the ticket titled "it was faster last week."
What you can control precisely is the per-VM cgroup. Exactly what PandaStack's host agent does, because vague cgroup claims are endemic here:
- It writes `cpu.weight` into a per-VM cgroup named `vm-<id>`: proportional share under contention, kernel ceiling 10000. The knob for "a paid tier should degrade less than a free tier" — free when the host is idle, decisive when it is not.
- It writes `cpu.max`, a hard ceiling over a 100 ms period. Worse for throughput, better for a promise: it is how you quote an x-realtime figure you can hit at p99.
- It reads `cpu.stat` `usage_usec`, `memory.current` and `io.stat` per VM. That is where per-job CPU attribution, working-set measurement and I/O accounting come from — real-time factor per tenant, not per host.
- It never writes `memory.max`: a guest's RAM is fixed by the baked snapshot, because Firecracker cannot change vCPU count or RAM at restore. Guest memory is a template property, and a per-request `memory_mb` is overridden to match what was baked.
- It does write `memory.high`, but only via the pressure ladder: under host memory pressure a controller squeezes it towards roughly 70% of residency on provably cold VMs — never database-class, never hugepage-backed, never CPU-active or touched in the last 60 seconds — then lifts it as pressure falls. A render burning CPU is excluded by definition, so that lever aims at idle guests, not at your mastering job.
- Working-set admission is the fleet default: the scheduler admits against an admittable-memory figure and charges `max(512, memory_mb * 0.25)` per non-database create — how a host runs more 4 GiB guests than it has 4 GiB chunks of RAM. Databases are always charged committed.
The baked-RAM point has a consequence better learned here than at 2am. The `base` template is baked at 4 GiB and 8 vCPU, the 8 vCPU being burst capacity shared by `cpu.weight` rather than a reservation. For most offline audio that is comfortable; for a sixty-four-track session or a heavy separation model it is not. The answer is template tiers, each baked at its own size — not a per-request memory parameter the platform will quietly ignore. The failure you want for a job exceeding its tier is a clean exit against its own `ulimit` plus a retry one tier up.
The architecture
The pipeline then falls out mechanically. Nine steps, and the ordering matters more than any one step.
- Ingest to object storage directly, not through your API. Record a job row: the digest of every stem, the digest of the plugin bundle, the tenant, the declared duration. Nothing has been parsed yet.
- Plan from declared metadata, not from the file. Probing an upload for its duration is a full container parse in a trusted process — the operation you are isolating. If you must probe, probe in a sandbox.
- Treat the plugin bundle as untrusted input: content-addressed, size-capped, pinned by digest to the job. Never configuration, never resolved by name at render time.
- Fetch what the job needs with something that is not the renderer: a boring fetcher with an allowlist, size caps, redirect limits and no credentials, which never decodes audio.
- Create one microVM per job from a baked audio template, with a TTL that is a hard ceiling, not an aspiration.
- Write the work in with no credentials — no object-store token, no queue credential, no database password. A guest holding nothing exfiltrates nothing.
- Render as a non-root user with the fences above, streaming logs out so a failing job says so before the TTL does.
- Read the output back, verify it is a plausible artefact of the right duration and format, and publish under a content-derived key, so retries are idempotent.
- Delete the sandbox unconditionally, in a `finally`, including when your orchestrator crashes. The TTL is the backstop for when your code forgets, not the plan.
Step 4 is the one teams skip; step 6 is the one traded away for convenience. Together they are the whole value of the architecture: the machine running a stranger's DSP code holds nothing worth stealing. The fastest way to lose that is handing the guest a storage token to save two lines.
Here is the orchestrator side.
import hashlib
import pathlib
from pandastack import Sandbox
JOB = "job-9c41"
STAGE = pathlib.Path("/stage") / JOB
STEMS = ["vocal.flac", "drums.flac", "bass.flac"]
# The render script from the previous block. It is versioned with the
# pipeline, not with the job: the customer supplies audio and plugins,
# never the recipe.
RENDER_SH = pathlib.Path("/opt/pipeline/render.sh").read_bytes()
def render_mixdown(timeout_s: int = 1800) -> bytes:
"""Render one customer mixdown in its own microVM. Returns WAV bytes."""
bundle = (STAGE / "plugins.tar").read_bytes()
digest = hashlib.sha256(bundle).hexdigest()[:16]
sbx = Sandbox.create(
template="base", # 4 GiB / 8 vCPU, baked
ttl_seconds=timeout_s + 600, # ceiling ABOVE the in-guest timeout
metadata={"job": JOB, "tenant": "acme", "plugins": digest},
)
try:
sbx.exec("useradd -m -d /work/home render || true",
timeout_seconds=60, check=True)
sbx.exec("mkdir -p /work/in /work/out /work/plugins",
timeout_seconds=60, check=True)
for name in STEMS:
sbx.filesystem.write(f"/work/in/{name}", (STAGE / name).read_bytes())
# The plugin bundle is untrusted input, so it is unpacked without
# inheriting any ownership or mode bits the archive claims.
sbx.filesystem.write("/work/plugins.tar", bundle)
sbx.exec("tar --no-same-owner --no-same-permissions "
"-xf /work/plugins.tar -C /work/plugins",
timeout_seconds=300, check=True)
sbx.filesystem.write("/work/render.sh", RENDER_SH)
sbx.exec("chown -R render:render /work && chmod 0755 /work/render.sh",
timeout_seconds=60, check=True)
# Non-root, and holding no credentials: the stems came in over the
# control channel and the master goes out the same way.
r = sbx.exec("runuser -u render -- sh /work/render.sh",
timeout_seconds=timeout_s + 120)
pathlib.Path(f"/var/log/renders/{JOB}.log").write_text(r.stdout + r.stderr)
if r.exit_code != 0:
raise RenderFailed(JOB, r.exit_code, r.stderr[-4000:])
wav = sbx.filesystem.read("/work/out/mixdown.wav")
if len(wav) < 1024 or wav[:4] != b"RIFF":
raise RenderFailed(JOB, 0, "output is not a plausible WAV")
return wav
finally:
sbx.kill() # the TTL is the backstop, not the plan
Four details are deliberate. The sandbox TTL sits above the in-guest `timeout`, so the normal failure is a clean non-zero exit with usable stderr rather than a VM vanishing mid-write. The plugin digest goes into the metadata, so the audit record answers "which binary did we run." The output is validated in the orchestrator, because a renderer that exits zero and writes a truncated file is not hypothetical. And `sbx.kill()` is in a `finally` — the SDK has no `delete()`, and a forgotten sandbox bills until someone reads an invoice.
On fan-out across a back catalogue, the two branching primitives have confusingly similar names. `sbx.fork()` reflinks the rootfs and cold-boots the child: it inherits the disk but not the memory, draws its own entropy, and has the roughly-three-second cold-boot shape. `sbx.fork_tree(count)` is the memory-inheriting path, restoring children from the parent's snapshot at 400–750 ms same-host and 1.2–3.5 seconds cross-host. For five hundred tracks sharing a big impulse library, `fork()` is usually right: the bytes are shared on disk by reflink, and three seconds of boot is noise against a six-minute master.
Measuring real-time factor from the job's own output
Report the number from the artefacts, not from a wrapper's estimate. This runs inside the guest and emits the figures that diagnose a slow job.
#!/bin/sh
set -eu
# Wall clock, CPU time and peak RSS for the render, from the kernel's
# own counters rather than from a stopwatch around a subprocess.
/usr/bin/time -f "wall=%e user=%U sys=%S maxrss=%M" -o /work/out/time.txt \
sh /work/render.sh
# Seconds of audio we actually produced, from the output we actually wrote.
DUR=$(ffprobe -v error -show_entries format=duration \
-of default=nw=1:nk=1 /work/out/mixdown.wav)
WALL=$(sed -n 's/.*wall=\([0-9.]*\).*/\1/p' /work/out/time.txt)
USR=$(sed -n 's/.*user=\([0-9.]*\).*/\1/p' /work/out/time.txt)
SYS=$(sed -n 's/.*sys=\([0-9.]*\).*/\1/p' /work/out/time.txt)
# rtf < 1 means faster than real time. cores is the average parallelism we
# actually achieved, which is the honest answer to "why is p99 bad?".
awk -v d="$DUR" -v w="$WALL" -v u="$USR" -v s="$SYS" 'BEGIN {
cpu = u + s
printf "audio_s=%.1f wall_s=%.1f cpu_s=%.1f rtf=%.3f x_realtime=%.1f cores=%.2f\n", \
d, w, cpu, w / d, d / w, cpu / w
}' | tee /work/out/rtf.txt
The `cores` figure is the one worth a dashboard. A real-time factor of 0.4 at average parallelism 1.1 is not CPU-starved, it is serialised — a bottleneck stage, or a plugin that refuses to use more than one thread. The same 0.4 at parallelism 7 on an 8 vCPU guest is doing exactly what you built it to do. Without the ratio those look identical in the queue.
The host-side counterpart is the agent reading `cpu.stat` `usage_usec` and `io.stat` for that VM's cgroup, reconciling what the guest thinks it burned against what the host charged. Divergence is steal time, which is a density conversation rather than a code one. The I/O half is at Per-Sandbox Disk I/O Throttling and the Neighbour You Cannot See.
Boundary options for plugin hosting, compared honestly
There are five places to draw the line around a plugin, and each stops something different. The exercise is not picking the strongest, it is being explicit about what you bought.
| Boundary | What it actually stops | What still gets through | Honest cost |
|---|---|---|---|
| Same process (`dlopen` into the renderer) | Nothing. Not a boundary — the plugin ABI working as designed. | Everything: heap, descriptors, credentials, FPU control word, signal handlers. | Zero. Lowest latency, no IPC — and for realtime, the only option. |
| Separate process, shared kernel | Crashes, host heap corruption, accidental interference. Killable, cgroup-fenceable. | The full syscall surface of the one shared kernel, and any kernel LPE. | A ring buffer and a copy per block. The right first move regardless. |
| Container per job | Filesystem and PID views, capability set, cgroup-fenced CPU. | The same shared kernel, through an enormous syscall surface, reached by a stranger's code. | Highest density, fastest teardown, mature tooling. The best fit when every plugin is first-party. |
| Syscall-interception sandbox (gVisor-style) | Direct kernel syscalls, and much of the kernel-LPE surface. A strong middle ground. | The sandbox kernel's own bugs; the plugin's native computation runs regardless. | Syscall-heavy paths slow measurably; compatibility gaps bite exotic native code. |
| MicroVM per job (PandaStack) | A guest kernel compromise stays in a kernel about to be deleted that saw one tenant's audio. Own netns and TAP. | A hypervisor bug, and whatever you handed the guest — hence no credentials inside. | Create 179 ms p50 / 203 ms p99 by snapshot restore; guest RAM fixed by the snapshot; $0.054/vCPU-hour and $0.0162/GiB-hour; no GPU passthrough. |
Read the container row as seriously as the last one. If every plugin was compiled by someone with a badge and every input came from your own ingest, containers are fine and this post is overhead. That assumption expires the moment you ship a plugin upload button or a free tier — and the expiry is almost never written down. The marketplace-shaped version is at Running a Plugin Marketplace Without Getting Owned.
The honest limits
The 179 ms create is almost irrelevant here
I will be blunt, because this is where microVM pitches oversell. A PandaStack create is 179 ms at p50 and 203 ms at p99, because every create restores a baked Firecracker snapshot rather than cold-booting — no warm pool underneath. The first-ever spawn of a template cold-boots in about three seconds and the agent bakes the snapshot for everyone after it. Against a six-minute master those numbers are a rounding error, and anyone selling per-job isolation for offline audio on cold-start latency is selling the wrong axis.
What makes it affordable is the other end of the lifecycle. The rate card is $0.054 per vCPU-hour and $0.0162 per GiB-hour, the same for every class, with CPU charged on CPU-seconds actually burned and memory on committed GiB-hours; egress is metered and there is no per-request pricing. A deleted sandbox accrues nothing. You are paying for the render, not for isolation — which is the only reason it is reasonable to give a free-tier job its own virtual machine.
Some plugins you simply cannot host
AU is an Apple framework and Firecracker guests are Linux, so an AU-only plugin is not a sandboxing problem but a platform one: the honest answer is that the format is unsupported. Same for anything needing a hardware dongle. VST3, CLAP and LV2 all have real Linux stories, but bring-your-own-plugin will produce requests you have to refuse, and publishing that list up front is cheaper than one ticket at a time.
There is no discrete-GPU passthrough either, deliberately — a minimal device model is most of why the hypervisor boundary is worth anything. For CPU DSP that costs nothing. For a GPU-accelerated separation model it is decisive, and that part of your pipeline is a different architecture.
Egress is open by default, and your licence checks will notice
Be precise about the network. On PandaStack, sandbox egress is open by default. There are a few targeted DROP rules for known-abuse protocols, but it is not a default-deny network: if a render job must not reach arbitrary hosts, that is a policy you add rather than a default you assume. The agent writes legacy `iptables` nat rules inside each sandbox's own network namespace — a PREROUTING DNAT per published port plus a POSTROUTING MASQUERADE — so per-job policy has an obvious place to live.
That default is right for the platform's main workloads — a sandbox that cannot install a package is not much use — and exactly wrong for a render job. The structural question is who chooses the URL. If your pipeline resolves whatever a session file or filter graph points at, from inside the render VM, customer input is choosing your outbound traffic: a crafted reference can name an internal service or a metadata endpoint. The fix is step 4 — resolve references outside the render VM, so the renderer needs no egress at all, and "no egress" is much easier to verify than "correct egress."
The summary
A hosted audio platform is a code-execution product in a creative-tools costume, more literally so than a render farm. The scene-file equivalent here is not a document carrying a script — it is a native shared library, uploaded on purpose, that your renderer must `dlopen` and call in a loop because that is how the ABI is specified. Out-of-process hosting is the right first move, but process isolation on a shared kernel assumes buggy code rather than chosen code, and that gap is where the incident lives.
One microVM per job answers the family at once: own guest kernel, own network namespace, a small device model, no credentials inside, a baked userland identical for every job, per-job CPU and I/O attribution you can turn into a real-time factor, and a machine deleted when the job ends. Keep the pitch honest — the 179 ms create is irrelevant against a six-minute render, guest RAM is a template decision, AU plugins and GPU models are out of scope, egress is open until you close it, identical silicon is still your problem. What that buys is isolation you can afford on every job including the free ones: a boundary you stop paying for the instant the music stops.
Frequently asked questions
Can you actually sandbox a VST3 or CLAP plugin, given that hosting means dlopen?
You cannot sandbox it inside the process that loads it — that part is the ABI, not a design choice. A VST3 on Linux is a bundle containing a native shared object, a CLAP is a shared object exporting `clap_entry`, and in both cases the host resolves an entry point and calls into that code from its processing loop. The boundary has to be drawn around the host process, not inside it. The strongest practical arrangement is layered: run the plugin host out of process so a crash is contained and you can fence it with cgroups and seccomp, then run that whole process inside a microVM with its own guest kernel, no credentials, and deletion when the job ends. Out-of-process hosting protects your application from buggy code. The VM boundary protects your other tenants from chosen code, because a separate process still speaks to the same shared kernel every other job runs on, and a kernel privilege escalation does not care how many processes you split your renderer into.
Why is a container not good enough for offline audio rendering?
For first-party code it usually is, and I would not argue otherwise. A container is a process with namespaces and cgroups on your host's single shared kernel, and the isolation is implemented by the same kernel the contained process makes syscalls into. That is a reasonable boundary when the threat model is accidental interference between your own workloads. It gets thin when a stranger uploaded both the input file and, increasingly, the shared library doing the processing, specifically so that you would execute them. Audio makes it thinner twice over. The input path is enormous — a dozen codecs, several containers, tag parsers, an embedded image decoder for cover art — all C reached by attacker-chosen integers. And the workload's own resource signature, saturated CPU and high memory bandwidth for minutes, is indistinguishable from that surface being exploited. On a shared kernel you gave up the ability to tell a capacity problem from a security problem at design time.
How do you get bit-identical output from a floating-point audio pipeline?
By pinning more than you think, then recording what you pinned. Three things cause drift. Denormal handling: reverb tails decay into denormals, DSP code mitigates by setting flush-to-zero and denormals-are-zero bits in the floating-point control register, and those bits are per-thread state a plugin can change before your code runs. Compiler flags: `-ffast-math` and `-Ofast` permit reassociation and reciprocal approximation, so the same source built with a different toolchain produces different last bits, and your dependencies' build flags are not in your repository. Runtime CPU dispatch: FFT kernels, resamplers and convolution engines select an AVX2, AVX-512 or NEON path from what the CPU reports. A snapshot-restored microVM with a pinned guest userland removes the drift class that bites most often — the host with a different library version, the node still running last quarter's build. It does not solve the hardware half: snapshots are baked per architecture, cross-CPU portability is a real constraint, and bit-identical masters still need the host CPU generation pinned or the dispatch path forced, recorded beside the output hash.
What is real-time factor and what actually limits it?
Real-time factor is wall-clock seconds of compute per second of audio produced, often quoted as its reciprocal, x-realtime. It is the number your pricing, queue model and SLA all reduce to, so measure it from the job's own artefacts rather than estimating. The intuition that it is purely CPU-bound is only half right: convolution reverb, FFT-based time and pitch work, and oversampled saturation are streaming kernels competing for memory bandwidth and cache rather than arithmetic throughput, so your real-time factor on a busy host is worse than on an idle one at the identical CPU share. Two cgroup levers matter. `cpu.weight` gives proportional share under contention, so a paid tier degrades less than a free one when the host is loaded, and costs nothing when it is not. `cpu.max` is a hard ceiling over a 100 ms period: worse for throughput, better for a promise, and the only way to quote an x-realtime figure you can hit at p99. Track average parallelism alongside the ratio — a slow job with parallelism near one is serialised, not starved.
Customers' plugins need to phone home for licence checks. How do you handle that?
As a product decision made deliberately, in writing, rather than as a firewall ticket. Many commercial audio plugins validate a licence at instantiation: some read a file from a fixed path under the user's home directory, some expect a hardware dongle you will never have, and some make an outbound HTTPS request as a precondition for processing audio. That last category collides directly with any egress policy, with no clean technical resolution — either a third-party shared library makes outbound requests from inside your infrastructure, or the plugin refuses to run and your customer reports your platform as broken. What you can do is contain the decision. Pre-activate where the vendor's licensing permits, so the render VM needs no network at all. Route licence-requiring plugins to a tier with a named per-destination allowlist and an audit record of every outbound connection. Keep the render VM free of credentials either way. And be willing to refuse: "this binary requires internet access to multiply numbers" is a legitimate reason to say no.
Keep reading
- Running a multi-tenant render farm on microVMs — The 3D sibling: same job shape, but the uploaded program is a scene file rather than a shared library.
- Running a plugin marketplace without getting owned — Versioned snapshots and per-plugin isolation, for when the plugin upload button becomes a marketplace.
- Reproducible build sandboxes — The general version of pinning a whole userland, which is what bit-identical mastering actually requires.
- Per-sandbox disk I/O throttling explained — The I/O side of the resource shape: what io.stat gives you and where throttling helps a batch transcode.
Related posts
- Rendering Untrusted Ad Creatives in Per-Tenant MicroVMs
Your creative-review pipeline renders HTML5 ad units from thousands of advertisers in one shared headless-Chrome farm. That farm is a giant C++ attack surface executing code that strangers paid to have you execute. Give every creative its own microVM.
- Per-Tenant AI Voice Transcription in MicroVMs
Your transcription service holds a customer's support calls, a clinic's dictation, and somebody's deposition in the same Python process, with ffmpeg parsing whatever container format arrived. A malformed .wav is an unsolicited code-execution proposal. One microVM per tenant makes it a boring one.
- Running Per-Tenant Billing and Usage Metering in MicroVMs
Most isolation bugs leak data. This one mails a customer an invoice built partly from a competitor's usage. The only bug class where the incident review includes your CFO — so make cross-tenant reads structurally impossible, not merely unlikely.
- Multi-Tenant WordPress Hosting on Firecracker MicroVMs
WordPress is a plugin-execution engine wearing a CMS costume, which makes shared hosting a multi-tenant remote code execution service with good branding. A look at why the PHP hardening stack is not a boundary, and what changes when every site gets its own kernel.
- Per-Tenant Log Parsing Isolation on microVMs
A tenant-supplied regex with nested quantifiers is a denial-of-service attack you shipped to yourself. Run each tenant's parse pass in its own capped microVM and the blast radius stops at one VM.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.