Customer-Authored Policy Code: Rego, CEL, and When a Rules Engine Needs a VM
Every B2B platform eventually ships the same feature, and it is always framed as a small one. Let the customer write the rule. An approval policy with more than three tiers. A fraud rule only their risk team understands. A routing condition on a field you have never heard of. A pricing formula with a carve-out for one account. An alert predicate. The ticket says "custom rules"; what you have agreed to is running logic written by strangers inside your request path, on behalf of all your other customers simultaneously.
I'm Ajay. I built PandaStack, which runs untrusted code in Firecracker microVMs, so you can guess which answer I am structurally inclined to sell you. This post argues that most teams should stop two rungs below it. A microVM per rule evaluation is the wrong answer for nearly every hot path, and I would rather you heard that here than from a latency graph.
The useful mental model is a ladder with three rungs, and the engineering discipline is climbing exactly as far as your rule shape forces you to and not one rung further. Each rung buys you a stronger boundary and costs you something real: expressiveness, latency, or operational surface.
The ladder, and the rung most teams should stop at
| Rung | Isolation boundary | Evaluation latency | What it cannot stop |
|---|---|---|---|
| 1. Bounded expression language — CEL, JSONLogic, Rego/OPA | The language itself: no unbounded loops, no I/O unless you enable a builtin | Microseconds to low milliseconds, in your own process | An expensive-but-legal query: a comprehension over a large data document, size() on a list the customer sized, a bundle that eats your heap |
| 2. Sandboxed interpreter in-process — Wasm, a JS isolate, RestrictedPython | The host runtime's own sandbox: a Wasm engine's linear memory, an isolate's heap | Sub-millisecond to tens of ms, plus engine instantiation | A bug in the engine, which is a bug in your process; a blocking native call; a rule that waits on the customer's own flaky API |
| 3. microVM — per tenant, per batch, or per evaluation | The kernel: a separate guest kernel, its own network namespace, its own page tables | ~179 ms p50 to create one; microseconds inside a warm one | Anything about your own API design. A VM does not un-hand a customer an SSRF primitive, or un-mount the credential you put in the guest |
Read that right-hand column as the actual decision criterion. You do not climb because a rung feels insecure; you climb because the specific thing you need to stop is not in that rung's vocabulary.
- Rung 1 when the rule is a predicate or a projection over data already in the request. This is the overwhelming majority of real rules, and choosing it is not a compromise — it is correct.
- Rung 2 when customers genuinely need control flow and local computation — a loop, a helper function, string munging — but still no I/O and no filesystem.
- Rung 3 when the rule is actually code, needs to reach the network, needs a real filesystem and process tree, or when the requirement you are satisfying is written in a compliance document that says "kernel" rather than "language".
Rung 1, and the three ways a bounded language still hurts you
Rego and CEL are both genuinely excellent, and they are good at different things. Rego, from OPA, is a query language over documents: it is at its best when the policy decision depends on a body of reference data — org charts, vendor blocklists, role hierarchies — and when you want the policy and the data to be separately deployable. CEL is at its best as a fast, embeddable predicate evaluator with a type checker and a cost model, which is why Kubernetes reached for it for admission and CRD validation. JSONLogic sits lower still: a JSON-serialisable operator tree, trivially storable in a database column and easy to render back into a UI builder, at the price of expressiveness.
package tenant.rule
# An approval policy a customer can plausibly author: bounded, inspectable,
# and evaluated in-process in microseconds. No loops, no I/O, no filesystem.
# `if` and `in` are v1 syntax -- on an older OPA add `import future.keywords`.
default verdict := {"allow": false, "reason": "no matching tier"}
tier_limit := {"ic": 500, "manager": 5000, "director": 50000}
verdict := {"allow": true, "reason": "under tier limit"} if {
input.amount <= tier_limit[input.requester.role]
not input.vendor.id in data.blocked_vendors
}
verdict := {"allow": false, "reason": "second approver required"} if {
input.amount > tier_limit[input.requester.role]
count(input.approvals) < 2
}
# This line is the one to watch in review. It is legal, it is not a loop, and
# it is linear in the size of data.transactions -- on every single request.
recent_spend := sum([t.amount |
some t in data.transactions
t.vendor == input.vendor.id
])
Now the failure modes, because "bounded" gets oversold. Rego is not Turing-complete — there is no unbounded recursion and the evaluator's termination is a language property, not a timeout you configured. That is a real guarantee and it is also narrower than people hear. A terminating query can still be ruinously expensive: the `recent_spend` comprehension in that snippet is linear in the size of `data.transactions`, which is a document you load and the customer does not control the size of — until the day someone decides the policy needs two years of history in it. Termination is not a cost bound. Treat a comprehension over a large document the way you would treat an unindexed `SELECT` in a request handler.
Second, the data plane is memory. OPA holds its `data` documents in process memory, and bundles have a way of growing from a 40 KB blocklist into a multi-hundred-megabyte denormalised export because someone needed one extra join. That memory is per process, multiplied by every replica. If your rule needs a big dataset, the answer is usually to move the lookup out of the policy and pass the already-resolved fact in as `input`, not to ship the warehouse to the evaluator.
Third, and most important: builtins are the hole in the boundary. Rego ships `http.send`, and the moment a customer's rule can make an outbound HTTP call, your microsecond-latency pure-function evaluator has become a request-forgery engine that runs inside your network with your egress identity. The fix is default-deny, not vigilance: generate a baseline with `opa capabilities`, delete every builtin you have not deliberately decided to allow, and pass that file to the evaluator. The same discipline applies to CEL extension functions, which are only as safe as the host functions you registered.
// CEL has no unbounded loops, which is not the same as "cheap". The macros
// iterate: all/exists/map/filter walk the list, and size() on a large list is
// still a linear read of something a customer chose the size of.
request.amount <= limits[request.requester.role]
&& !(request.vendor.id in blockedVendors)
&& request.approvals.all(a, a.actor != request.requester.id)
&& size(request.lineItems) <= 200
// cel-go can estimate an expression's cost at COMPILE time and enforce a
// separate cost limit at RUNTIME. Kubernetes uses exactly that pair for CRD
// validation rules. Turn both on: the compile-time estimate is how you reject
// an expensive expression at save time instead of discovering it in p99.
Rung 2, where the escape surface is your own runtime
When customers need real control flow, you reach for a sandboxed interpreter in your process. The modern defaults are reasonable. A Wasm engine gives you linear memory with no ambient capabilities, and the mature runtimes expose both a fuel-style instruction budget and an interruption mechanism so you can stop a module that will not stop itself; check your engine's own documentation for which knobs its current version exposes, because the names and semantics move. A JavaScript isolate gives you a fresh heap per tenant and an API to terminate execution. RestrictedPython exists, and I will be blunt about it: in-process Python sandboxing has a long and well-documented history of someone finding a path back to the interpreter internals. If your threat model includes a motivated customer, do not put Python in your process.
Two limits define this rung. The first is that the boundary is your own runtime, and your runtime is the process that holds every other tenant's data, your database handle, and whatever credential your pod was issued. A container does not change that calculus much — a container is a polite suggestion to the kernel — because the escape you are worried about does not need to leave the process to be catastrophic.
The second is subtler and bites teams who got the first one right. A 50 ms CPU budget does not stop a customer from writing a rule that calls their own service and waits. Instruction counting measures instructions; it does not measure a blocked host call. You either forbid I/O inside the rule entirely — which is the right default and means the customer's external data has to be fetched by you, before evaluation, with your own timeout and circuit breaker — or you accept that your evaluator's tail latency is now a function of someone else's uptime.
Rung 3: the four honest reasons to climb
- The rule is actually code. Customers are pasting Python, importing a library, calling a model. There is no expression language left to restrict — you are running a program.
- The rule must reach the network, usually because the real decision lives in the customer's own service and they will not export the data to you.
- The rule needs a filesystem and a process tree: it shells out, it writes a temp file, it runs a binary somebody compiled.
- The boundary has to be a kernel boundary because an auditor wrote it down that way, and "our interpreter has a fuel budget" is not an answer that survives a security review at a bank.
What a microVM actually gives you here is specific and worth stating precisely. A different kernel, so a kernel-level bug in the guest is contained by the hypervisor rather than shared with your other tenants. Its own network namespace, veth pair and tap device, so network policy is enforced at an interface rather than inside a library. Its own page tables and memory, so there is no shared heap to read. On PandaStack that costs about 179 ms p50 to create, because every create restores a baked snapshot rather than booting — the first cold boot of a fresh template is around 3 seconds, and after that you are on the snapshot path.
What it does not give you: absolution for your own API design. If you mount a database credential into the guest so the rule can "look things up", the kernel boundary is decorative. If you let the rule call arbitrary hosts, you built the SSRF primitive yourself and then put a hypervisor around it. And there are hard platform edges: the guest kernel is 5.10 with an Ubuntu 24.04 userspace, there is one kernel per host so you cannot swap it per template, and there is no `/dev/kvm` inside a guest — so no nesting a second sandbox inside your sandbox.
Egress policy belongs in the netns, not in the rule language
This is the part worth stealing even if you never run a VM. Because each sandbox gets its own network namespace with its own veth pair and tap device, you write the egress policy once, at the interface, and it applies to every path the guest could possibly take: the policy engine's `http.send` builtin, a `curl` the rule shelled out to, a DNS lookup, a Python `requests` call that follows a redirect to a link-local address. Default-deny outbound, then an explicit per-tenant allowlist of the destinations that tenant's rule is permitted to reach. That allowlist is now a reviewable artifact in your control plane instead of a regex in a code path, and the enforcement point does not care how clever the rule was about constructing the URL.
One related footgun, since people reach for it when exposing an evaluator: PandaStack preview URLs are tokenless — a port is reachable at `https://<port>-<sandbox-id>.<suffix>` and the sandbox UUID is the credential. That is a deliberate trade for sharing a dev preview. A policy evaluator holding tenant decision data is not a dev preview; keep it on private addressing and talk to it through your own authenticated control plane.
The arithmetic that kills the naive design
An in-process CEL evaluation is microseconds. Creating a microVM on our fast path is 179 ms p50 and 203 ms p99. That is roughly five orders of magnitude, and no amount of engineering closes it, because they are not the same kind of operation: one is a function call and the other is a hypervisor restoring a memory image. If your approval check runs on every page load, a VM per evaluation is not a tuning problem, it is a category error.
So the patterns that actually work all move the VM off the per-request path.
- Per-tenant long-lived sandbox with a warm interpreter. You pay the create once per tenant, not once per request, and each evaluation is a round trip plus the microseconds the rule actually takes. The tenant is the isolation unit because the thing you are preventing is tenant A's rule seeing tenant B's context.
- Batch evaluation. Score a thousand events in one call instead of a thousand calls. Nightly fraud re-scoring, bulk approval backfills and alert sweeps are all naturally batch-shaped, and batching is what makes the per-tenant VM affordable.
- Hibernate and wake for idle tenants. Most tenants evaluate nothing at 03:00. A hibernate is a snapshot plus a pause; a wake restores it. You stop paying for a thousand idle evaluators that exist because a thousand tenants each enabled one rule in 2024.
- The pre-flight at save time, which is the one place where latency genuinely does not matter. The customer is clicking Save and waiting; 179 ms or 3 seconds is indistinguishable to them, and you get a full kernel boundary to run their unproven rule inside.
Notice the asymmetry. On the hot path a VM is expensive and a bounded language is nearly free. At admission time a VM is nearly free and a bounded language cannot tell you very much. So use both, in the places where each is cheap. That combination — rung 1 serving traffic, rung 3 guarding the gate — is the design I would defend in a review, and it is not a compromise between the two; it is each tool doing the job it is actually good at.
The pre-flight: smoke-test the rule before you admit it
Admission should prove a short list of things, and it should prove them by execution rather than by inspection. That the rule parses and compiles. That it terminates inside a hard wall clock. That its output has the right shape — a dict with a boolean `allow` and a string `reason`, not a bare string, not `undefined`. That it agrees with a fixture set, including the fixtures you wrote specifically to be denied. That it is deterministic: run it twice on identical input and get identical answers, which is where a rule that reached for the clock or a random source announces itself. And, if your language has one, that its compile-time cost estimate is inside your budget.
import json
from pandastack import Sandbox
# Admission control for a customer-authored rule. This runs ONCE, at save
# time, before the rule goes anywhere near a request. It is the place a
# microVM genuinely earns its keep: the rule is arbitrary until proven
# otherwise, and here we can afford to let it be arbitrary.
RULE = open("tenant-4821-approval.rego").read() # whatever they pasted
FIXTURES = json.load(open("fixtures/approval.json")) # [{"input":..,"expect":..}]
sbx = Sandbox.create(
template="code-interpreter", # 2 GiB RAM, baked at template build time.
ttl_seconds=180, # cpu/memory_mb are deliberately NOT passed:
metadata={ # Firecracker cannot change vCPU or guest RAM
"purpose": "rule-preflight", # at snapshot restore, so whatever you
"tenant": "4821", # pass is silently corrected to the
}, # baked value. Passing it reads like a lie.
)
RUNNER = r"""
import json, subprocess, sys
fixtures = json.load(open("/work/fixtures.json"))
rejects, verdicts = [], []
for i, f in enumerate(fixtures):
try:
p = subprocess.run(
["opa", "eval",
# A capabilities file is how you default-deny builtins. The one
# that matters is http.send: a Rego rule that can call out is an
# SSRF primitive pointed at your own network. Generate a baseline
# with `opa capabilities` and delete what you do not want.
"--capabilities", "/work/caps.json",
"--data", "/work/rule.rego",
"--stdin-input", "--format", "json",
"data.tenant.rule.verdict"],
input=json.dumps(f["input"]),
capture_output=True, text=True, timeout=2,
)
except subprocess.TimeoutExpired:
rejects.append({"fixture": i, "why": "timeout"}); continue
if p.returncode != 0:
rejects.append({"fixture": i, "why": "eval_error",
"detail": p.stderr[-400:]}); continue
res = json.loads(p.stdout).get("result") or []
if not res: # undefined: the rule produced no verdict
rejects.append({"fixture": i, "why": "undefined"}); continue
v = res[0]["expressions"][0]["value"]
# Shape-check before value-check. A rule that returns a string, or a dict
# missing "allow", is a rule your hot path will crash on at 03:00.
if not isinstance(v, dict) or not isinstance(v.get("allow"), bool) \
or not isinstance(v.get("reason"), str):
rejects.append({"fixture": i, "why": "bad_shape", "got": repr(v)[:200]})
continue
if v["allow"] != f["expect"]["allow"]:
rejects.append({"fixture": i, "why": "wrong_verdict",
"want": f["expect"]["allow"], "got": v["allow"]})
verdicts.append(v)
# Determinism check: identical inputs must give identical verdicts. If a rule
# reaches for wall-clock time or randomness, this is where it shows up.
print(json.dumps({"rejects": rejects, "verdicts": verdicts}))
"""
sbx.filesystem.write("/work/rule.rego", RULE)
sbx.filesystem.write("/work/fixtures.json", json.dumps(FIXTURES))
sbx.filesystem.write("/work/runner.py", RUNNER)
# Every hard limit lives in the SHELL. timeout_seconds on exec/exec_stream is
# a CLIENT deadline -- neither exec endpoint enforces it inside the guest -- so
# a spinning rule would just make the HTTP call give up while the guest kept
# burning its 8 burstable vCPUs until the TTL reaped it.
HARNESS = r"""
set -uo pipefail
ulimit -v 1048576 # 1 GiB address space: kills a runaway comprehension
ulimit -f 65536 # 64 MiB of files: a policy has no business writing more
ulimit -u 64 # no fork bombs dressed as a policy bundle
cd /work
timeout --signal=KILL 30s python3 runner.py
"""
res = sbx.exec(HARNESS, timeout_seconds=60)
if res.exit_code == 137:
decision = {"admit": False, "why": "preflight suite killed at 30s wall clock"}
elif res.exit_code != 0:
decision = {"admit": False, "why": f"harness failed: {res.stderr[-400:]}"}
else:
report = json.loads(res.stdout.strip().splitlines()[-1])
decision = ({"admit": True} if not report["rejects"]
else {"admit": False, "rejects": report["rejects"]})
sbx.kill()
print(decision)
# What this proved: the rule parses, terminates inside a hard wall clock,
# returns the right SHAPE, and agrees with your fixtures. What it did NOT
# prove: that it behaves on inputs you did not think of. Pre-flight is a smoke
# test, not a proof -- pair it with a runtime cost limit on the hot path.
The `timeout --signal=KILL` in that harness is not belt-and-braces, it is the actual enforcement. Neither of our exec endpoints enforces `timeout_seconds` server-side — it is a client deadline, so the HTTP call gives up while the guest happily continues burning its eight burstable vCPUs until the TTL reaps it. The same is true of most sandbox APIs in some form, so check yours rather than assuming. Hard limits belong in the shell: `timeout`, `ulimit`, and a TTL on the sandbox as the backstop.
The per-tenant warm evaluator, and parking it
Once a rule is admitted, it needs a home with a warm interpreter in it. The shape below is deliberately boring: one sandbox per tenant, a versioned policy file, a symlink swap for deploys, batch evaluation over an exec, and a hibernate for the tenants who are asleep.
import json
from pandastack import Sandbox
# The per-tenant warm evaluator. One sandbox per tenant, never shared: the
# tenant IS the isolation unit, because the thing you are protecting against
# is tenant A's rule reading tenant B's context.
#
# Cost model: you pay the ~179 ms p50 create once per tenant (not per
# request), then every batch is an exec round trip plus microseconds of actual
# evaluation inside an already-warm interpreter.
class TenantEvaluator:
def __init__(self, tenant_id: str, sandbox_id: str | None = None):
self.tenant_id = tenant_id
self.sandbox_id = sandbox_id # you store this on the tenant row
def _ensure(self) -> Sandbox:
if self.sandbox_id:
sbx = Sandbox.get(self.sandbox_id)
if sbx.status in ("paused", "hibernated"):
sbx.wake() # restore from the snapshot we parked
sbx.refresh()
return sbx
sbx = Sandbox.create(
template="policy-eval", # your own template: OPA + the shims
persistent=True, # exempt from the idle reaper, which
metadata={ # makes parking it YOUR job, not the
"tenant": self.tenant_id, # platform's. That is the trade.
"role": "policy-evaluator",
},
)
# Docker ENV from the template Dockerfile does not reach a detached
# process in the guest -- only /etc/environment (PAM) does -- so the
# evaluator's config gets written, not inherited.
sbx.filesystem.write(
"/etc/pandastack-policy.json",
json.dumps({"tenant": self.tenant_id, "allow_egress_to": []}),
)
self.sandbox_id = sbx.id
return sbx
def deploy_rule(self, rego: str, version: str) -> None:
sbx = self._ensure()
sbx.filesystem.write(f"/policy/{version}.rego", rego)
# Atomic-ish swap: write the new version, then repoint. A half-written
# policy file is a tenant-wide outage you caused, not one they authored.
sbx.exec(f"ln -sfn /policy/{version}.rego /policy/current.rego "
f"&& systemctl reload policy-eval", check=True)
def evaluate_batch(self, events: list[dict], budget_s: int = 10) -> list[dict]:
sbx = self._ensure()
# Note the shape: ONE input document with an events array, not an array
# as the input. `opa eval` evaluates the query once against whatever
# you hand it, so the fan-out has to live in a policy you own:
#
# package platform.batch
# import data.tenant.rule
# verdicts := [v | some e in input.events
# v := rule.verdict with input as e]
#
# That wrapper is also your cost choke point -- it is the one
# comprehension in the system whose length you decide, not the tenant.
sbx.filesystem.write("/work/batch.json", json.dumps({"events": events}))
# Batching is what makes this affordable: one round trip amortised over
# N events, instead of N round trips over a 179 ms create each.
res = sbx.exec(
"set -uo pipefail; ulimit -v 2097152; "
f"timeout --signal=KILL {budget_s}s "
"opa eval --capabilities /policy/caps.json "
"--data /policy/batch.rego --data /policy/current.rego "
"--stdin-input --format json "
"'data.platform.batch.verdicts' < /work/batch.json",
timeout_seconds=budget_s + 5,
)
if res.exit_code == 137:
# The rule blew its budget. Fail closed to your platform default
# and park the VM for inspection rather than retrying into it.
raise TimeoutError(f"tenant {self.tenant_id} rule exceeded {budget_s}s")
if res.exit_code != 0:
raise RuntimeError(res.stderr[-400:])
# Same output shape as the pre-flight: result -> expressions -> value.
return json.loads(res.stdout)["result"][0]["expressions"][0]["value"]
def park(self) -> None:
"""Most tenants evaluate nothing at 03:00. Stop paying for that."""
sbx = self._ensure()
sbx.hibernate() # snapshot + pause; wake() restores it later
def recycle(self) -> None:
"""A poisoned evaluator is not worth debugging in place."""
if self.sandbox_id:
Sandbox.get(self.sandbox_id).kill()
self.sandbox_id = None
self._ensure() # a fresh one costs a snapshot restore, not a boot
Three operational notes on that. `persistent=True` exempts the sandbox from the idle reaper, which means parking it becomes your responsibility rather than the platform's — that is the trade, and forgetting it is how a thousand tenants turn into a thousand idle VMs. A poisoned evaluator is not worth debugging in place: kill it and create a fresh one, because a snapshot restore is not a boot and recycling is cheap. And never, under any amount of density pressure, let two tenants share an evaluator. That line is the entire reason the topology exists.
Determinism, audit, and replaying their own events
A policy decision that you cannot reproduce is a policy decision you cannot defend, which matters the first time a customer asks why an invoice was blocked in March. Make determinism a property of the system rather than a hope: forbid wall-clock access and randomness inside the rule language, and inject `now` as an ordinary field on `input` so that the evaluation is a pure function of data you logged. Then log the triple — a hash of the input document, a hash of the exact rule version, and the verdict — and the question "why did this deny" becomes a re-run instead of an investigation.
The feature this unlocks is the one customers actually fall in love with, and it is not the sandbox: it is replay. When a tenant edits a rule, run the candidate over their own last thirty days of real events, in a sandbox, and diff the verdicts against what the live rule decided. "This change flips 412 of your last 9,000 approvals, here are twelve of them" turns a nervous policy edit into a reviewable diff. It is batch-shaped, it is latency-insensitive, and it is embarrassingly parallel — which is exactly the workload profile where a VM per batch stops being expensive and starts being obvious.
One platform-specific trap if you build that on forks. On PandaStack, `fork()` is a disk-and-memory snapshot restore of the parent, so the guest's RNG state and its clock come back identical in every child. For deterministic replay that is nearly a feature: N evaluators starting from a byte-identical state. Everywhere else it is a hazard — two forks will generate the same "random" identifier and sign with a stale timestamp. Seed your RNG explicitly and re-sync the clock on wake, or your audit log will contain two different decisions with the same UUID and you will spend an afternoon disbelieving your own database.
A rules engine is a promise that your customers' logic cannot hurt your other customers. Every rung of the ladder is a different definition of "cannot".
The summary I would put on a whiteboard: start at rung 1 and stay there as long as the rule shape allows, because a bounded language on the hot path is both the fastest and the simplest answer and most rules never need more. Climb only for a named reason from that list of four. And when you do climb, put the VM where it is cheap — at save time, in batch, and per tenant rather than per request — rather than in the one place its latency is impossible to hide.
Frequently asked questions
Should I just use OPA and Rego for customer-authored policy?
For a large class of rules, yes, and it is the first thing I would try. Rego is a query language over documents, so it fits naturally when the decision depends on reference data — a role hierarchy, a vendor blocklist, an org chart — and when you want policy and data to deploy separately. It is not Turing-complete, so termination is a language guarantee rather than a timeout you remembered to set. Two caveats decide whether it stays the right answer. First, termination is not a cost bound: a comprehension over a large data document is linear in that document and runs on every request, so treat it the way you treat an unindexed query. Second, strip the builtins. Generate a baseline with `opa capabilities`, delete everything you have not explicitly allowed, and in particular remove `http.send` unless you have a network-level egress policy behind it. If your customers need loops, local helper functions and string processing rather than predicates over data, Rego will fight you and a Wasm-hosted language or a real sandbox fits better.
Is one microVM per rule evaluation ever the right design?
Rarely, and only when the evaluation is already slow or already dangerous. Creating a sandbox on our snapshot-restore path is 179 ms p50 and 203 ms p99, against microseconds for an in-process CEL evaluation, so if the rule runs on a page load you have added several orders of magnitude for a boundary you probably did not need. Per-evaluation VMs make sense when the rule is itself long-running — it calls a model, processes a file, talks to the customer's service for seconds — because then the create cost is a rounding error on the work. They also make sense when the input crosses a trust boundary within a single tenant, for instance a rule evaluating an attacker-supplied document, where you want a fresh machine rather than a reused one. For everything else, move the VM off the request path: one warm sandbox per tenant, batch evaluation inside it, hibernate when the tenant is idle, and a fresh sandbox at save time for the pre-flight.
How do I stop a customer's rule from calling my internal services?
Not in the rule language, because that is the layer with the least information. A URL allowlist implemented inside the engine has to defeat redirects, DNS rebinding, IPv6 literals, decimal-encoded addresses and every hostname nobody enumerated, and it only covers the one code path you thought about — not a shelled-out `curl` or a library that resolves differently. Put the control in the network instead. Each sandbox on PandaStack has its own network namespace with its own veth pair and tap device, which gives you one enforcement point covering every syscall path in the guest. Default-deny outbound, then an explicit per-tenant allowlist of destinations, expressed as a reviewable object in your control plane rather than as a regex in a handler. Block link-local and cloud metadata ranges unconditionally. The best option of all is still to forbid I/O inside rules entirely and fetch the external data yourself, before evaluation, with your own timeout and circuit breaker — then the rule stays a pure function and tail latency stays yours to control.
How do I actually enforce a CPU or time limit on a customer's rule in a sandbox?
In the shell inside the guest, not in the SDK call. On PandaStack neither exec endpoint enforces `timeout_seconds` server-side — it is a client deadline, so the HTTP request gives up while the guest keeps running and keeps burning its eight burstable vCPUs until something else reaps it. The pattern that works is layered: wrap the evaluation in `timeout --signal=KILL`, add `ulimit -v` for address space and `ulimit -u` to make fork bombs uninteresting, set a `ttl_seconds` on the sandbox as the backstop that catches everything you forgot, and treat exit code 137 as a first-class outcome in your code rather than an error to retry. If the rule language offers a runtime cost limit, enable that too — it stops an expensive expression before it allocates, whereas `timeout` only notices afterwards. And when a budget trips, kill the sandbox rather than retrying into it; a fresh one is a snapshot restore, not a boot.
Keep reading
- Per-tenant feature-flag evaluation — The same ladder applied to targeting rules, where the hot path is every single request.
- Controlling network egress for untrusted code — The netns-level enforcement this post argues for, in detail: default-deny, allowlists and what still leaks.
- The code isolation hierarchy — The general version of the ladder — process, container, Wasm, microVM — and what each boundary actually promises.
- Wasm vs Firecracker for untrusted code — Rung 2 against rung 3 head to head, including the cases where the Wasm engine is genuinely the better answer.
- Building a multi-tenant audit log — Where the input hash, rule version and verdict triple ends up, and how to make it queryable per tenant.
Related posts
- Per-Tenant Fraud Rules in Isolated microVMs
When customers upload their own scoring rules and those rules run on every checkout, a shared worker pool means one tenant's rule can read another's transaction features — or pin a core and add latency to everyone. Give each tenant's rule its own Firecracker microVM.
- Rendering user-authored templates safely: SSTI and the microVM fix
You shipped a text formatter so customers could edit their own emails. Depending on the engine, you may also have shipped them a REPL on your application server.
- Running Per-Tenant Billing and Usage Metering in MicroVMs
Most isolation bugs leak data. This one mails a customer an invoice built partly from a competitor's usage. The only bug class where the incident review includes your CFO — so make cross-tenant reads structurally impossible, not merely unlikely.
- Running Customer UDFs Safely in a SaaS with microVMs
The moment your product lets a customer write code — a workflow step, a validation hook, a transform — you're hosting untrusted code in your infra. Here's how to run it without betting the company on it.
- A Spreadsheet Is a Programming Language Your Users Don't Call One
Your users write programs in your product every day. They spell them =SUMPRODUCT(...) instead of def main(), and that is the only difference. Then someone asks for PDF export and you shell out to an office suite.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.