PandaStack vs Kubernetes Jobs: Where "Just Create a Job" Stops Being the Right Answer
Someone on your team needs to run code that a user typed into a text box. Maybe it is an agent's generated Python, maybe it is a customer's data transform, maybe it is a build from a pull request opened by an account you have never heard of. There is a cluster. There is a `Job` resource in `batch/v1` that exists precisely to run a thing to completion. The whole conversation takes four minutes and the ticket gets an estimate of two days.
That instinct is correct far more often than my job description would suggest. I build PandaStack, an open-source Firecracker microVM platform, which means I spend a lot of time talking to engineers who are deciding between a Job and a sandbox API — and a meaningful fraction of the time I tell them to stay on Kubernetes. If the code is yours, if the call is asynchronous, if the volume is thousands a day rather than thousands an hour, a Job is a well-understood object running on infrastructure your on-call already knows how to debug at 3am. That is worth a great deal and a feature grid will never show it to you.
What I want to do here is draw the line honestly. Not "Kubernetes is insecure" — it is not, and the people who say so usually have not read Pod Security Admission. The line is narrower and more mechanical than that, and it has four parts: the boundary a pod actually is, the shape of a Job's cold start rather than its median, the gap between a declarative object and a request/response call, and what teardown means when a tenant is adversarial rather than merely sloppy.
If you are on the other side of any of those four, a Job is the right answer and I would rather you kept your money. If you are on this side of all four, you are about to write a distributed system by accident, and I would like to show you its shape before you commit a quarter to it.
The honest case for Kubernetes Jobs
Start with the thing that no sandbox API can give you: it is already there. That sentence does a staggering amount of work in a real engineering org, and it is not a technical property at all.
- The scheduler exists and has been tuned by people who understood your bin-packing problem. The node autoscaler exists and is wired to your cloud account's quotas. The image registry exists, with your pull secrets and your digest-pinning policy in it.
- The secret store, the RBAC model, the network policies and the admission controllers exist, and someone wrote the policy that makes them hang together. Re-deriving that in a second system is not free; it is the most commonly underestimated line in the migration estimate.
- The logging pipeline, the metrics, the dashboards and the alert routes exist. When something hangs at 2am there is a runbook, and the person reading it has debugged this control plane before. A new substrate means a new class of page that nobody has seen yet.
- `Job` and `CronJob` are genuinely good primitives. `backoffLimit` gives you retries, `activeDeadlineSeconds` gives you a wall clock, `ttlSecondsAfterFinished` garbage-collects finished objects, `completions` plus `parallelism` gives you fan-out, and indexed jobs give each worker a stable index so a shard-per-pod map job is a few lines of YAML rather than a queue.
- Resource requests and limits, priority classes, pod disruption budgets, node selectors, taints and tolerations, topology spread constraints: that is real placement and preemption control, and a hosted sandbox API does not expose anything like it. You can pin a workload to a node pool, you can say which work gets evicted first, you can keep two replicas off the same zone. I cannot offer you any of that.
- Per unit of compute it is cheaper, often substantially, when the nodes are already warm and already on a committed-use discount. Running a short job on capacity you have already paid for is close to free at the margin, and that arithmetic is correct.
So before the criticism, here is the Job I would actually write for semi-trusted per-request work. Not the eight-line one from the docs — the one where every knob is set on purpose. Read the comments: several of these lines buy less than people assume they do, and one of them buys more than everything else combined.
# A Job you would not be embarrassed to run a customer's snippet in.
# Every line buys something specific; the comments say what, and also what it
# does NOT buy. The single highest-value line in the file is the one that
# unmounts the service account token.
apiVersion: batch/v1
kind: Job
metadata:
name: run-user-snippet-7f3a1c
namespace: tenant-7f3a1c # a namespace is a NAME, not a boundary:
# same nodes, same kernel, same page cache
spec:
backoffLimit: 0 # retries are a feature for your code and a
# liability for a user's: a retry means the
# side effect happens twice. 0 = one attempt
activeDeadlineSeconds: 120 # the only wall clock Kubernetes gives you.
# Covers the WHOLE pod lifetime -- a slow
# image pull eats your compute budget, and
# you cannot bound the pull separately
ttlSecondsAfterFinished: 300 # stops finished Jobs accumulating in etcd.
# Careful: it deletes the Job AND its pods,
# so it deletes the only copy of your logs
completions: 1
parallelism: 1
template:
metadata:
labels: { app: snippet-runner, tenant: '7f3a1c' }
spec:
restartPolicy: Never # with backoffLimit 0, no silent re-exec
automountServiceAccountToken: false
# THE line. Without it every pod mounts a
# token under /var/run/secrets and anything
# inside can talk to the API server as that
# service account. Most escapes you read
# about start here, not in the kernel
enableServiceLinks: false # stops dozens of SERVICE_* env vars being
# injected, which otherwise hand the snippet
# a map of your cluster's services
hostNetwork: false
hostPID: false
hostIPC: false # all three default to false; declaring them
# is documentation and a reviewable diff
dnsPolicy: Default # the NODE's resolv.conf, so the snippet
# cannot resolve in-cluster services by
# name. This is not a network policy and it
# does not stop an IP literal
nodeSelector: { workload: untrusted }
tolerations:
- key: untrusted
operator: Exists
effect: NoSchedule # with the matching taint, untrusted work is
# pinned to its own node pool. This is the
# most useful mitigation on the page: an
# escape now lands on a node holding nothing
# else of yours
securityContext:
runAsNonRoot: true
runAsUser: 65534
runAsGroup: 65534
fsGroup: 65534
seccompProfile:
type: RuntimeDefault # denies the long tail of syscalls the
# runtime's profile excludes. It is a
# denylist shaped by what containers need,
# not a proof the remainder is safe
containers:
- name: runner
image: registry.internal/runner@sha256:9c1f0b7e # a DIGEST, not a
# tag: mutable tags are a supply-chain hole
# and they also defeat node image caching
imagePullPolicy: IfNotPresent # cache hit is fast; cache miss is a
# pull inside activeDeadlineSeconds
command: ['/usr/bin/timeout', '--kill-after=10s', '60', '/bin/runner']
# an in-container fence that fires BEFORE
# the Job deadline, so you get a readable
# exit code instead of a dead Job object
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true # the snippet cannot write to /. It
# can still write to the emptyDir below,
# which is where it will put the thing you
# did not want written
capabilities:
drop: ['ALL'] # no CAP_SYS_ADMIN, no CAP_NET_RAW. Dropping
# capabilities removes the privileged doors;
# it does not shrink the syscall surface
resources:
requests: { cpu: '500m', memory: '512Mi' }
limits: { cpu: '1', memory: '512Mi' }
# memory limit = cgroup OOM kill, hard.
# cpu limit = CFS throttling, real but
# bursty. Neither covers page cache this pod
# makes the node reclaim from its neighbours
volumeMounts:
- { name: scratch, mountPath: /work }
volumes:
- name: scratch
emptyDir: { sizeLimit: 1Gi, medium: Memory }
# medium: Memory charges the tmpfs to the
# pod's memory limit, so a disk-filling
# snippet OOMs itself instead of filling the
# node's ephemeral storage out from under
# every other pod on it
That manifest is good work. Pod Security Admission at `restricted` would accept it, a reviewer would learn something from it, and it closes the attacks that actually happen most: a mounted service account token, a writable root filesystem, a `root` process, an unbounded runtime. If your threat model is "a customer's transform has a bug" rather than "a customer wants my cluster," you can stop reading here and ship it.
The four places it stops being the right answer
1. A pod is a process group on a shared kernel
Everything in that manifest is defence in depth on one side of a single boundary, and the boundary is the node's kernel. Namespaces, cgroups, seccomp and AppArmor or SELinux are real, valuable and worth configuring. They are also a large, actively-researched attack surface, and it is the same kernel every other tenant's pod on that node is calling into. Nothing in the YAML changes that; it only reduces the number of ways a process gets to reach it.
Be precise about the consequence, because the vague version gets dismissed. A container escape is a node compromise. A node compromise is usually a cluster compromise, not because of a kernel bug but because of plumbing: the kubelet's credentials, the CNI's view of the pod network, whatever secrets other pods on that node have mounted, the node's own cloud instance identity. The taint-and-toleration node pool in my manifest exists exactly to shrink that blast radius, and it is the mitigation I would reach for first — but it is a blast-radius control, not a boundary.
The honest in-cluster answers to this are gVisor and Kata Containers, and they deserve naming rather than hand-waving. gVisor reimplements the Linux syscall surface in userspace so the host kernel sees a much narrower interface. Kata runs each pod in an actual virtual machine — and underneath, Kata is usually Firecracker or Cloud Hypervisor. Which means the argument was never "VM versus Kubernetes." It is "which boundary, and who operates it." The fact that both projects exist, and that both are maintained by people who know Kubernetes extremely well, is the strongest available evidence that the default boundary is not sufficient for this workload class. I walk through the three substrates in Firecracker vs Kata vs gVisor: three isolation models, and the short version is that they make different trades and none of them is free to run.
2. The latency shape, which is variance and not median
A Job's path to running code is: an API server write, then the scheduler binds the pod, then the kubelet on the chosen node pulls the image or hits its cache, then the container runtime starts the container, then your process starts. On a warm node with a cached image, that is fast and you should not let anyone tell you otherwise. I am not going to quote numbers for your cluster because I have not measured your cluster, and neither has whoever wrote the benchmark you are about to read.
What matters is the shape. That path has at least three branches that produce completely different durations: cache hit, cache miss with a node available, and no node available at all. The third branch is a node provision — a cloud API call, a boot, a kubelet registration, a CNI setup, then the image pull — and the number that actually hurts you is the autoscaler's reaction time, which is a control loop and not a latency. If your traffic is request-shaped and bursty, your p99 is that branch, every time, and your p50 will look wonderful in the dashboard while the complaints come in.
The comparison I can make with real numbers is my own side. PandaStack restores a baked Firecracker snapshot on every single create: p50 179 ms, p99 203 ms. The reason those two numbers are close is structural, not clever — there is no warm pool, so there is no warm pool to miss. The one exception is the very first spawn of a template that has no snapshot yet, which is a cold boot plus an auto-bake at around 3 seconds, and then never again. I wrote about why a pool is the wrong shape for this in Snapshot-Restore vs Warm Pools: Two Ways to Kill Sandbox Cold Starts.
Measure your own cluster before you believe anyone, including me. Create a Job per request at your real peak rate, on a node pool sized for your average, and record the distribution rather than the mean. If the tail is acceptable for your product, the latency argument does not apply to you and you should ignore it.
3. Declarative objects versus a request/response call
This is the part that costs the most engineering time and gets the least attention in evaluations. Kubernetes objects are declarative and eventually consistent: you write your intent and controllers converge toward it, reporting progress through status fields and events. That is a beautiful model for a deployment and a slightly hostile one for "run this snippet, hand me stdout, I have an HTTP request waiting."
So you write a shim. The shim creates the Job, watches for completion, scrapes the logs before they vanish, distinguishes "ran and failed" from "never scheduled at all," propagates an exit code up through three layers of abstraction, and garbage-collects whatever is left. That shim is a real distributed system with its own failure modes, and you will own it forever. Here is the shortest honest version I know of.
# The shim. This is the floor, not the ceiling -- everything here is load
# bearing and none of it is the feature your users asked for.
#
# Not shown, and you will need all of it: backoff on 429/409 from the API
# server, leader election so two replicas of this service do not both create
# the Job for one request, metrics, and a reconcile loop that adopts the Jobs
# this process created before it crashed.
import time, uuid
from kubernetes import client, config, watch
from kubernetes.client.rest import ApiException
config.load_incluster_config()
batch, core = client.BatchV1Api(), client.CoreV1Api()
NS = "tenant-7f3a1c"
DEADLINE = 120 # must agree with activeDeadlineSeconds in the manifest
SCHED_GRACE = 30 # how long we wait for a pod to be assigned a node at all
class Unschedulable(RuntimeError):
pass
def run_snippet(snippet: str) -> tuple[int, str]:
name = f"snippet-{uuid.uuid4().hex[:10]}"
batch.create_namespaced_job(NS, build_job(name, snippet)) # the YAML above
try:
pod = wait_for_pod(name)
if pod is None:
# The branch everybody forgets. No pod was ever scheduled: no node
# had room, the pool is mid-scale-up, a taint has no matching
# toleration, the pull secret is wrong, a ResourceQuota is full.
# There is no exit code because nothing ran, and returning 500 is
# a lie when the cause is capacity. Your caller needs to tell the
# difference, so this has to become a distinct error type.
raise Unschedulable(scheduling_reason(name))
exit_code = wait_for_terminal(pod, timeout=DEADLINE + 15)
# Read the logs BEFORE deleting anything. The kubelet's on-disk buffer
# is rotated by size and is destroyed with the pod -- and
# ttlSecondsAfterFinished destroys the pod. kubectl logs is a debugging
# convenience, not a durable output channel. If you need the output to
# survive, that is a log pipeline, which is a second system.
try:
out = core.read_namespaced_pod_log(pod, NS, _request_timeout=10)
except ApiException:
out = "" # already collected; the output is simply gone
return exit_code, out
finally:
# propagation_policy matters. Without Background, the Job is deleted and
# its pods are ORPHANED -- now you have pods nothing owns, holding node
# slots until a human notices.
try:
batch.delete_namespaced_job(
name, NS,
body=client.V1DeleteOptions(propagation_policy="Background"),
)
except ApiException as err:
if err.status != 404:
raise
def wait_for_pod(job_name: str, grace: int = SCHED_GRACE) -> str | None:
deadline = time.monotonic() + grace
w = watch.Watch()
try:
for ev in w.stream(core.list_namespaced_pod, NS,
label_selector=f"job-name={job_name}",
timeout_seconds=grace):
p = ev["object"]
if p.spec.node_name: # bound to a node: it will start
return p.metadata.name
if time.monotonic() > deadline:
break
finally:
w.stop() # leaking watches leaks API server
# connections, which is a fun outage
return None
There is a second cost in that pattern that only shows up at volume. A pod per request is an etcd write per request, plus the watch traffic that every controller and every operator in your cluster now processes. Ten thousand short Jobs an hour is a load pattern Kubernetes will tolerate rather than enjoy, and the symptom is rarely a clean error — it is API server latency creeping up, and your unrelated deployments getting slower, and someone spending a week discovering that the snippet runner is the reason.
4. Teardown is not a transaction, and neighbours are not isolated
Deleting things in Kubernetes is a reconciliation, not a commit. Namespace deletion walks finalizers and can wedge on one that never completes. PersistentVolumes outlive their claims according to a reclaim policy that someone set once. A Job whose process hangs holds its node slot until a deadline fires, and if you got the deadline wrong, until someone pages. These are not bugs; they are the semantics, and they are fine for infrastructure you control and awkward for a tenant you do not.
The neighbour effects are the sharper half. A fork bomb inside a pod spends the node's PID limit, and `pids` limits per pod help only if someone set them. One tenant reading a large working set evicts another tenant's page cache, and no `resources` block in the Pod spec expresses that. Kernel memory a pod makes the node allocate is charged unevenly. The cgroup accounting is genuinely good at the things it accounts for, and the list of things it does not account for is where multi-tenant platform teams spend their quarters.
Compare the microVM shape. The guest's memory ceiling is the VM's, enforced by the hypervisor, and the guest's page cache is inside the guest where it cannot evict anybody else's. A fork bomb spends the guest's PID space and nothing else. Teardown is killing one process and freeing one network namespace plus one copy-on-write disk file, and there are 16,384 pre-allocated subnet slots per host so the network side is a table lookup rather than a provisioning step. That is a smaller, duller set of failure modes — which is the entire pitch.
The comparison, dimension by dimension
| Dimension | Kubernetes Job | PandaStack microVM sandbox |
|---|---|---|
| Isolation boundary | Namespaces, cgroups, seccomp, AppArmor/SELinux on a kernel shared with every other pod on the node. Escape equals node compromise. | Hypervisor. Separate guest kernel (5.10) per sandbox, a handful of emulated virtio devices instead of a full syscall surface. |
| Create-latency shape | Trimodal: cache hit, cache miss with a node free, or node provision plus pull. Variance is dominated by the autoscaler's reaction time. Measure your own. | One path: restore a baked snapshot. p50 179 ms, p99 203 ms on every create. First spawn of an unbaked template is ~3 s, once. |
| Programming model | Declarative, eventually consistent. Create an object, watch status, scrape logs, infer an exit code. You write and own the shim. | Imperative and synchronous. `Sandbox.create()`, `exec()` returns stdout, stderr and an exit code, `kill()` in a `finally`. |
| Teardown semantics | Reconciliation. Finalizers, orphaned pods on the wrong propagation policy, PVs that outlive claims, TTL GC that deletes your logs with the pod. | Kill one process, free one netns and one CoW file. Idle TTL reaper deletes the VM; unpushed work dies with it unless you snapshotted. |
| Resource ceiling enforcement | Memory limit is a hard cgroup OOM; CPU limit is CFS throttling. Page cache pressure, kernel memory and PID exhaustion leak across tenants on a node. | Hypervisor-enforced RAM ceiling; guest page cache is inside the guest. Sizes are baked into the snapshot and cannot change at restore. |
| Fan-out concurrency | Excellent for batch: `parallelism`, `completions`, indexed jobs. Each unit is still an etcd write plus watch traffic for every controller. | Flat per-create cost; 16,384 subnet slots per host, with RAM as the real ceiling. `fork_tree()` shares warm parent memory, capped at 16 per call. |
| State and snapshot story | An image plus whatever volumes you attached. No notion of freezing a running process; warm state means a long-lived pod, which means a long-lived tenant. | Snapshot the running guest and restore it later or in parallel. `fork()` clones the disk and cold-boots; `fork_tree()` inherits live memory. |
| Operational burden | You run it: control plane, node pools, upgrades, CNI, the runtime class, plus the shim. Everything is inspectable and your team already knows it. | I run the hosts; you call an API. Less to debug and less to tune, and a vendor in your request path — which is a real cost, not a footnote. |
| Cost shape | Cheap at the margin on nodes you have already committed to, expensive in idle headroom you must keep to absorb bursts. Billed by node-hour. | $0.054 per vCPU-hour on CPU-seconds actually burned, $0.0162 per GiB-hour of committed memory while the sandbox exists. One rate card; see pricing. |
The same job, as a sandbox call
Here is the other side of the shim. It is short, and the reason it is short is the model, not cleverness: the lifecycle is a function call with a return value instead of an object whose status you poll.
# The same work against the sandbox API. Short because the lifecycle is
# request-shaped, not because the problem got easier.
from pandastack import Sandbox
def run_snippet(snippet: str) -> tuple[int, str]:
sbx = Sandbox.create(
template="code-interpreter", # 2 GiB / 8 vCPU, baked into the
# snapshot. Firecracker cannot resize a
# guest at restore, so per-create cpu
# and memory_mb are overridden to the
# baked values -- pick the template, not
# the numbers
ttl_seconds=300, # an IDLE clock, not a wall clock.
# Guest-touching requests reset it;
# status and metrics GETs deliberately
# do not. This is the backstop for the
# finally below never running at all,
# which is what a SIGKILL on this
# process looks like from my side
metadata={"tenant": "7f3a1c"},
)
try:
sbx.filesystem.write("/work/snippet.py", snippet)
# The fence is IN-GUEST on purpose. One-shot exec() does not actually
# enforce timeout_seconds -- it bounds the HTTP call, not the command --
# so timeout(1) inside the VM is the thing that kills a snippet that
# spins. Same role as the `command:` wrapper in the YAML above.
res = sbx.exec(
"cd /work && timeout --kill-after=10s 60 python3 snippet.py",
timeout_seconds=90,
)
return res.exit_code, res.stdout + res.stderr
finally:
# Unconditional, not inside an except. Teardown is kill(); there is no
# sbx.delete(). The ttl_seconds above covers the case this never runs.
sbx.kill()
# What you traded away, stated plainly:
# - no nodeSelector, no taints, no priority classes, no pod affinity. You do
# not choose the host and you cannot express "evict this one first".
# - a create can still be refused for capacity. A finite fleet is a finite
# fleet; a 503 is a real response and backoff is your code's job.
# - egress from the guest is OPEN by default. Siblings cannot reach each
# other's subnets and the cloud metadata range is dropped at the host, but
# your VPC and your internal APIs are not fenced for you.
# - there is now a vendor in your product's request path. That is a real
# cost. Kubernetes does not have that property and it matters to some orgs
# more than every row in the comparison table combined.
Notice that the in-guest `timeout` fence appears in both versions. That is not a coincidence or a weakness of either platform — it is the correct place for the bound in any architecture, because the thing you want to stop is the process, not the call that started it. Any design that puts the only timeout in the client is a design that leaks runaway work.
Which one, concretely
Use Kubernetes Jobs when
- The code is yours, or your dependencies', and your threat model is bugs rather than adversaries. This covers the large majority of batch work in the world and a Job is simply the right object for it.
- The work is asynchronous and the caller is a queue, a cron schedule or another controller. A declarative object that converges is a better fit than a synchronous call that must not block, and `CronJob` plus `backoffLimit` is less code than any scheduler you would write.
- You need real placement control: a specific node pool, GPUs, a priority class so that batch yields to serving, topology spread, or data locality on a node's local disks. A sandbox API gives you none of this and I am not going to pretend otherwise.
- Your volume is thousands of jobs a day rather than thousands an hour, so the per-object cost in etcd and API server QPS never becomes a topic.
- Compliance or procurement means the compute must stay inside your account and your perimeter, with no third party in the request path. This is a legitimate and often decisive constraint.
- You already have the platform team, the runbooks and the on-call rotation. Operational familiarity is a real asset, and a second substrate spends it.
Use a microVM sandbox when
- The code's author is not on your payroll: a model generated it, a user pasted it, a fork's pull request contains it. `rm -rf` from a language model is funny on a laptop and a different genre of funny on a kernel you share with a customer's production pod.
- The call is synchronous and user-facing, so the tail latency of a node provision is a product problem rather than a graph. The shape you want is one path with a narrow distribution, not three paths with a good median.
- You need the boundary to be something a security reviewer will sign off on in a sentence. "Separate kernel per tenant, enforced by the hypervisor" is that sentence; "seccomp plus a restricted PSA profile plus a dedicated node pool" is a paragraph and an argument.
- You want warm state fanned out: snapshot a process that has finished its expensive setup, then boot many children from it. `fork_tree()` inherits the parent's live memory, with a hard cap of 16 children per call. Kubernetes has no equivalent primitive, because a pod has no freeze.
- Idle must actually cost nothing much. Memory bills as committed GiB-hours only while the sandbox exists, and the idle reaper deletes it, so the lever is non-existence rather than a cheap rate. Node headroom kept warm to absorb bursts is the opposite shape.
- You would rather not write and own the shim from the second code block, and the honest alternative is paying someone else's per-hour rate for the compute.
Most teams run both, and that is not a hedge
The most common real architecture I see is not one or the other. It is a service running on Kubernetes — ordinary `Deployment`, HPA, your service mesh, your existing observability — that calls a sandbox API for the part of its work that executes other people's code. The control plane, the queues, the API handlers, the database access and the business logic all stay exactly where your team is good at operating them. Only the dangerous subprocess moves across a hypervisor boundary.
That split is better than either extreme for a boring reason: the two systems are good at different things and the interface between them is small. A `create`, an `exec`, a `kill`. You do not need to move your platform to move your blast radius, and if you later decide you want to run the sandbox substrate yourself, PandaStack is open source and Kata exists, so neither door is locked.
The failure mode to avoid is the half-migration where some execution paths use Jobs and some use sandboxes and nobody wrote down which. Pick the boundary by property — "did someone outside this company write these bytes" is a good one — and make it a rule rather than a preference.
The bottom line
If I had to compress this into one sentence: Kubernetes Jobs are the right primitive for running your own work to completion on infrastructure you operate, and the wrong primitive for running strangers' code synchronously at request rate. Those are different problems that happen to share the verb "run."
And do not use PandaStack for the first problem. If you want GPUs, there are none — no PCI passthrough, no VFIO, and a software rasterizer is not an answer. If you need placement control, a priority class or data locality, Kubernetes has those and I do not. If your compute must stay inside your perimeter with no vendor in the path, that decides it on its own. If your batch workload already runs happily on committed nodes you have paid for, moving it to a per-second meter is a worse deal and I would tell you so on a call. And the guest kernel is 5.10 on Ubuntu 24.04, which most toolchains do not care about and a few do — test before you plan a migration around it. If you want the substrate background rather than the comparison, What is a microVM? Firecracker, isolation, and why agents need it is the place to start, and Controlling Network Egress for Untrusted Code covers the part of this that neither platform does for you.
The test I would actually run is cheap. Write the hardened Job above, run it per request at your real peak for an afternoon, and record two things: the full latency distribution including the scheduling-failure branch, and how many lines the shim grew by. Then do the same thing through a sandbox API. One of those two numbers will make the decision for you, and it will not be the median.
Frequently asked questions
Can't I just use gVisor or Kata Containers in my cluster?
Yes, genuinely, and if you want to stay on Kubernetes this is the correct answer rather than a consolation prize. You define a `RuntimeClass`, point the untrusted workloads at it, and your pods run with a much stronger boundary while keeping every piece of tooling you already have. gVisor reimplements the syscall surface in userspace; Kata gives each pod a real VM, and underneath that VM is usually Firecracker or Cloud Hypervisor — the same hypervisor I run. So this is not a different philosophy, it is the same boundary with a different operator. What it costs you is the honest part. You are now running a second container runtime: its own version matrix, its own kernel or sentry to patch, failure modes your runbooks do not cover, and a node pool that supports nested virtualisation or bare metal for Kata. Some syscalls and some filesystem behaviour differ under gVisor, so a workload that passes on `runc` can fail there, and you will hear about it from a customer. Performance changes in ways that depend on the workload's syscall and I/O profile. None of that is prohibitive — plenty of teams run it well — but it is a platform-engineering commitment, not a flag. The decision is less "VM or container" and more "do I want to be the team that operates this boundary."
Is ten thousand short Jobs an hour really a problem for Kubernetes?
It is not a wall you hit, which is exactly what makes it dangerous. Kubernetes will accept that load and keep working; the cost is diffuse. Every Job is a write to etcd, plus its pod, plus the events, plus the status updates as it progresses — and every one of those changes is delivered to every controller, operator and admission webhook that watches those resource types. So the cost of your snippet runner is paid by things that have nothing to do with it. The symptom is rarely an error. It is API server latency drifting upward, etcd's write throughput and compaction becoming interesting, your unrelated deployments taking longer to roll, and eventually somebody spending a week with a flame graph to discover that a feature shipped two months ago is the reason. The mitigations are real but they all push you away from the pod-per-request model: batch many units into one Job with `parallelism` and `completions`, keep long-lived worker pods pulling from a queue instead, or set `ttlSecondsAfterFinished` aggressively so finished objects do not accumulate. Notice that the first two mitigations are "stop using a Job per request," which is also my argument. If your volume is in the thousands per day rather than per hour, none of this applies and you should not optimise for it.
How do I get durable output out of a Kubernetes Job?
Not from `kubectl logs`, which is the trap. That reads the kubelet's on-disk buffer for the container: it is rotated by size, so a chatty process loses its own early output, and it is destroyed along with the pod — which means `ttlSecondsAfterFinished`, the field you correctly set to stop etcd filling up, is also the field that deletes the only copy of your results. The race between your log scrape and the garbage collector shows up as intermittently empty output under load, and it is maddening to reproduce. The real options are all "write the output somewhere you own": have the job upload a result object to blob storage and return a key, write a database row the job owns, publish to a queue, or run a node-level log shipper forwarding to a store whose retention you control. Each is correct and each is another moving part. This is the clearest example of the shim problem — what you wanted was a return value, and what the platform offers is an object with a status plus a side channel whose lifetime you do not control. A synchronous sandbox API hands you stdout, stderr and an exit code in the response body. That is not better engineering, just a different place the complexity went: if your process dies before persisting the result, it is equally gone.
Does PandaStack give me the placement control I get from nodeSelector and priority classes?
No, and that is the single biggest thing you give up. There is no `nodeSelector`, no taint, no toleration, no priority class, no pod anti-affinity and no topology spread. You ask for a template and the scheduler places the sandbox using a deterministic score over the hosts' free CPU and admittable memory, with some affinity for hosts that already hold the artifacts a restore needs. That is it. You cannot say "this tenant's work goes on dedicated hardware," you cannot say "preempt batch before serving," and you cannot pin a workload next to data on a local disk. What you get instead of placement control is that placement matters much less: every sandbox has its own kernel, so the "keep the untrusted thing away from the important thing" problem that node pools and taints exist to solve is handled by the boundary rather than by scheduling. There is also a real operational consequence of a finite fleet: a create can be refused for capacity, and a 503 is a response your code must handle with backoff rather than a retry storm. If placement is genuinely load-bearing for you — regulated workloads on specific hosts, GPUs, latency-critical data locality — that is a decisive argument for staying on Kubernetes, and I would make it for you rather than against you.
If my cluster already autoscales, isn't a sandbox API just a more expensive node pool?
Sometimes it is exactly that, and in those cases you should stay put. The arithmetic turns on two things: how much idle headroom your bursts force you to keep, and whether your traffic is request-shaped. An autoscaled pool absorbs a burst by provisioning nodes, and provisioning is a control loop whose reaction time is measured in whatever your cloud's instance boot plus kubelet registration plus image pull takes. To make a user-facing path feel fast you therefore keep headroom warm, and warm headroom is the thing you are actually buying — nodes that exist so that a burst does not wait. That cost does not appear in any per-job calculation. On my side the create path is a snapshot restore at p50 179 ms and p99 203 ms with no warm pool to miss, and the meter is $0.054 per vCPU-hour on CPU-seconds actually burned plus $0.0162 per GiB-hour of committed memory while the sandbox exists. Per unit of compute, that is more than a spot node you already own. Per unit of absorbed burst with no headroom, it is usually less. So run the comparison on your real traffic shape rather than on a steady-state node-hour price, and if your load is smooth and predictable, the node pool wins and you should keep it.
Keep reading
- Firecracker vs Kata vs gVisor — The three in-cluster and out-of-cluster boundaries side by side, including what each costs to operate.
- PandaStack vs gVisor — The narrower head-to-head: a reimplemented syscall surface against a real guest kernel.
- PandaStack vs Modal — The other direction of this comparison — a managed serverless-compute model rather than a cluster you run.
- Docker-in-Docker vs microVMs for CI — The same boundary argument where it bites hardest: privileged builds inside your own pipeline.
- Best multi-tenant isolation platforms in 2026 — The wider field, graded on where the tenant boundary actually sits rather than on the marketing.
Related posts
- Sandboxing User-Written Webhook Transformations
Somebody added a textarea labeled 'Transform (optional)' and shipped it on a Thursday. Congratulations: you are a code-execution company now, and nobody told your threat model.
- Server-Side Rendering Someone Else's React Component
If your product server-renders components your customers wrote, you are running their code in your process, with your env vars and your database socket. node:vm is a sandbox for values, not for a module graph that can require("fs"). Here is the boundary that actually holds, and what it costs.
- Per-Tenant Isolation for User-Defined GraphQL Resolvers
Your headless CMS lets customers write GraphQL resolvers in JS or Python. That's arbitrary tenant code on your request path. Here's how to isolate it per tenant without blowing the latency budget.
- Firecracker vs Bottlerocket: One Is a Hypervisor, One Is a Host OS
These two aren't rivals; they're stacked. Bottlerocket hardens the host you share. Firecracker removes the sharing. If you're running untrusted code, that distinction is the whole ballgame.
- PandaStack vs Namespace Cloud: Two Kinds of "Ephemeral"
Three different things are called "namespace" and only one of them is a company. Here is the disambiguation, and then the honest comparison: a managed-runner-and-cache product against a microVM sandbox API, including the two places the choice is genuinely real.
More in Code execution · See Code interpreter sandboxes on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.