Workload Identity for MicroVMs That Live 90 Seconds
A sandbox restores from a snapshot in about 179 milliseconds, runs somebody's build for ninety seconds, and gets deleted. In that window it needs to write an artifact to an S3 prefix. So it needs a credential — and right now that credential is baked into the template or passed in as an environment variable at create time, you already know both are wrong, and that is presumably why you are here.
I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This is workload identity for things that do not live long enough to have a hostname worth trusting: why the obvious answers fail, one of them in a genuinely non-obvious way, and the three architectures that work. The one-line version is that identity cannot come from inside a VM with no history — it has to be granted from outside, by whatever created the VM and knows what it was created for.
Three ways to get it wrong, in increasing order of subtlety
A key baked into the template
You put the key in the image because at build time the image was the only place that existed. Now the key's lifetime is the artifact's lifetime, and artifacts do not expire — they have generations, and somebody pinned one. Rotating means rebuilding, re-publishing and waiting for every host to pull the new generation: a multi-hour operation you will not perform quarterly, which means you will not perform it.
The deeper problem is that an image is a distribution mechanism: its whole job is to be copied as widely and cheaply as possible, so a secret in one is not a storage mistake but a category error. It produces the artefact this post is named after — the credential with a five-year expiry inside a VM with a five-minute one.
A key passed in as an environment variable at create time
A real improvement: per-sandbox, absent from the artifact, rotatable without touching an image. What it does not fix is who can read it. An environment variable set in the guest's init is inherited by every child, sits in `/proc/<pid>/environ` for anything running as root — which in a single-tenant microVM is everything — and surfaces in crash dumps and error logs. The thing you built the sandbox to contain can read the credential you put in the sandbox, and a model asked to debug its own environment will cheerfully `cat` it into a thirty-day transcript. A scope problem, then, not a storage problem.
The subtle one: a snapshot is a memory dump
Here is the one that catches careful people. A Firecracker snapshot is two artefacts: `vm.state`, the device and vCPU state, and `vm.mem`, a byte-for-byte freeze of the guest's physical RAM. Not a filesystem image. RAM.
Consider what a template bake does: boot the guest, run a bootstrap, warm a package cache, maybe start a server so the first request is fast, snapshot. A token the bootstrap fetched is in the snapshot. A secret it decrypted is in there as plaintext. An SSH private key loaded into an agent is in there as key material, with the session keys of the TLS connection that fetched any of it. All restored identically into every sibling — which, with no warm pool, is every single create.
Then look at how that artefact is handled, because nobody treats it like a secret. It is a multi-gigabyte file in object storage, replicated to every host that might run the template. On PandaStack with streaming restore, `vm.mem` is not even downloaded: a userfaultfd handler serves guest page faults with 4 MiB HTTP Range GETs against GCS, backed by a shared on-disk chunk cache per seed. Lovely engineering, and four more copies of your guest's RAM — bucket, range responses, chunk cache, page cache. A secret in `vm.mem` is not leaked. It is published.
An adjacent sharp edge, and it depends on which call you used. `fork()` clones the parent's disk and then cold-boots the child, so the child gets its own kernel, memory and entropy — nothing to re-seed. `fork_tree(count=N)` restores N children from one snapshot of the parent's memory, and memory includes the kernel's RNG state, so a child that generates a nonce or a key right after being restored can generate the same one as its siblings. If the plan was “the VM mints its own credential material once it starts,” re-seed first — or better, have the thing that created the VM hand the material in, which is the whole argument of this post.
The corollary is useful rather than depressing: public material can be baked, private material cannot. PandaStack's own example is exactly that — the per-host agent owns one ed25519 keypair, the public half goes into the template rootfs as `/root/.ssh/authorized_keys` at bake time so creates can just reflink a prepared rootfs, and the private half stays on the host at mode 0600 and never enters a guest. That asymmetry is the only kind of credential that belongs in an image.
What identity can mean for a VM with no history
So what does a 90-second microVM have to offer as proof of who it is?
- No hardware identity: no TPM, no measured boot to bind a key to, no instance identity document of its own. The host has one; the guest is not the host.
- No stable hostname. It is assigned one at restore, by the host, so trusting it means trusting the host.
- No uniqueness, which kills every clever scheme. It is byte-identical to every sibling restored from the same snapshot, down to the contents of RAM — any secret it can present, every sibling can present.
- No history. It did not exist four seconds ago — no prior authentication, no long-running process that was trusted earlier and could vouch for it.
So nothing inside the guest can be evidence of identity, because anything inside the guest is inside every guest. But something in the system does know: whatever created the VM knows the tenant, the template, the job, the commit, who asked, and how long it may live. That knowledge is the identity, and every architecture below moves it across the guest boundary without converting it into a long-lived secret.
A worked example: how a restored VM learns who it is
On restore, `pandastack-init` races three ways of being told who it is: it listens on a vsock port, it simultaneously dials out over vsock with an aggressive backoff, and it polls Firecracker's MMDS at `http://169.254.169.254/identity`. The host writes this restore's identity — IP, MAC, gateway, hostname — into the MMDS data store while the VM is still paused, before Resume: MMDS state survives snapshot restore, and Firecracker permits the write while paused. First usable answer wins. It delivers network identity rather than credentials, which makes it a good place to watch the pattern without the stakes.
The first thing to steal is the placeholder. The template is baked with an MMDS body of literally `{"identity":{"placeholder":"1"}}`, and the guest polls until it sees something else. The bake-time value is deliberately not a real one: it lets the guest distinguish “not told yet” from “told,” and leaves the artifact holding nothing true. That is the shape you want for a credential — a hole, and a guest that waits for it rather than defaulting.
The second is three channels, for a sharper reason than “redundancy is good.” The two vsock paths share a failure mode — a post-restore wedge where the guest's vsock does not come back cleanly — so they are not independent and do not count as two. MMDS rides virtio-net and survives the same boundary. We learned that from a managed-Postgres create where Firecracker restored and resumed perfectly and every channel then timed out, because that template was baked without an MMDS data store and `PUT /mmds` returned a 400. Which exposes the gotcha: whether the VM has an MMDS device is part of the snapshot's device state, so you cannot add a channel to an existing snapshot.
Shape 1: a broker on the host
The guest asks a host-side endpoint for a credential. The host knows which sandbox is on the other end of that socket — it opened the socket when it created the VM — looks up its tenant and job, and mints a scoped short-lived credential. The cloud metadata-service pattern with better information: your broker knows not just the instance but the job.
The authentication is the socket, and that is the elegant part. With vsock, the host side is a unix socket file belonging to exactly one microVM. The guest presents nothing — no token, no certificate, no claim — because the host already knows, so there is nothing on the client side to steal. Any scheme where the guest instead presents a secret hands you a secret distribution problem, which was the problem.
#!/usr/bin/env bash
# ILLUSTRATIVE: this is the host-broker pattern, not a PandaStack API.
# PandaStack does not ship a credential broker -- the host side here is a
# listener you write inside your own per-host agent. The client half is what
# you bake into the template, and baking THIS is fine, because it contains no
# secret: it is a hundred lines that know where to ask.
set -euo pipefail
# ---------------------------------------------------------------- vsock path
# Preferred, because there is no network involved at all. The guest connects to
# the host CID (2) on an agreed port; Firecracker hands that connection to a
# unix socket on the host that belongs to EXACTLY ONE microVM, because the
# agent created that socket file when it created the VM.
#
# That socket is the authentication. The guest presents nothing -- no token, no
# certificate, no claim about itself -- because the host already knows which
# sandbox is on the other end. There is no credential to steal from the client
# side because the client side has no credential.
#
# socat has had VSOCK-CONNECT since 1.7.4. If your base image's socat is older,
# this is about fifteen lines of Go against AF_VSOCK; do not shell out to nc.
fetch_over_vsock() {
printf 'GET /v1/credentials/artifacts HTTP/1.1\r\nHost: broker\r\nConnection: close\r\n\r\n' \
| socat - VSOCK-CONNECT:2:1024 \
| sed '1,/^\r$/d' # drop headers, keep the JSON body
}
# ------------------------------------------------- link-local variant (worse)
# The other shape you will see, because it is what every cloud does. Inside a
# Firecracker guest, 169.254.169.254 is the VMM's own MMDS: those packets are
# intercepted by the virtio-net device and answered by the hypervisor, so they
# never reach the host's network stack. That makes the store per-VM and
# host-written -- no cross-tenant exposure of the kind a shared metadata host
# would have.
#
# It is still reachable by every process in the guest, and it is still an
# ordinary HTTP GET to a URL, which is the SSRF primitive. If you serve a
# CREDENTIAL (rather than inert config) from a link-local address, require the
# v2 token flow: a PUT to obtain a short-TTL session token, then that token as
# a header on every GET. A bare GET can be performed by anything that can make
# the guest fetch a URL it did not choose.
fetch_over_link_local() {
local tok
tok=$(curl -sf -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-metadata-token-ttl-seconds: 60")
curl -sf -H "X-metadata-token: ${tok}" \
"http://169.254.169.254/credentials/artifacts"
}
CREDS=$(fetch_over_vsock || fetch_over_link_local)
# The response shape is where the design actually lives. Read it for what is
# ABSENT: no account-wide key, no refresh token, nothing that still works
# tomorrow, and no permission that was not needed for this one job.
cat <<'SHAPE' >/dev/null
{
"sandbox_id": "9f1c2ab4-...", # the host filled this in, not the guest
"tenant": "acme",
"aws": {
"access_key_id": "ASIA...", # ASIA = temporary. AKIA here means you failed.
"secret_access_key": "...",
"session_token": "...",
"expiration": "2026-10-06T11:19:00Z" # minutes, not months
},
"scope": {
"bucket": "acme-artifacts",
"prefix": "tenant/acme/job/4172/", # one prefix, this job only
"actions": ["s3:PutObject"] # no ListBucket: listing enumerates customers
}
}
SHAPE
# Now the part people get wrong after doing everything else right. Do not
# `export` this into the init shell: every child process in the guest inherits
# it, it lands in /proc/<pid>/environ for anything running as root (which, in a
# single-tenant microVM, is everything), and it shows up in crash dumps and in
# the helpful log line someone added that prints the environment on error.
#
# AWS SDKs support `credential_process`: the SDK execs a command and reads JSON
# from its stdout, so the credential lives in one process's memory for one call
# chain. ~/.aws/config:
# [default]
# credential_process = /usr/local/bin/broker-creds
# and broker-creds prints exactly:
# {"Version":1,"AccessKeyId":"ASIA...","SecretAccessKey":"...",
# "SessionToken":"...","Expiration":"2026-10-06T11:19:00Z"}
#
# This is hygiene, not a security boundary. Untrusted code running as root in
# the same guest can run broker-creds itself. The boundary is the VM.
# Finally: fail loudly here if the broker handed back something unusable,
# rather than letting the job die forty seconds later with an opaque 403 that
# someone will spend an afternoon blaming on IAM.
jq -e '.aws.session_token and .aws.expiration' >/dev/null <<<"$CREDS" \
|| { echo "broker returned no usable credential" >&2; exit 1; }
169.254.169.254 is the most attacked IP address in the industry: the whole SSRF class is “make the victim fetch a URL it did not choose, point it at the metadata service, read the credentials out.” It is why IMDSv2's token dance exists. There is also a collision at that address worth untangling, because the same four octets mean two different things one layer apart. Inside the guest it is the VMM's own MMDS, intercepted by the virtio-net device and answered by the hypervisor, never reaching the host's network stack. One layer out it is the cloud provider's metadata service, and on GCP that hands out the host VM's service-account token with cloud-platform scope — a token that reads every other customer's snapshots and seeds.
The two are separated by where the packet gets handled, and that separation is a rule you write. PandaStack's host inserts a DROP for the entire 169.254.0.0/16 range from the sandbox pool as one of the first rules in FORWARD, ahead of the egress ACCEPTs, because a guest never legitimately routes link-local off-host. The guest's own MMDS is unaffected — those packets were never going to be forwarded. If you run untrusted code on a cloud VM and have not written that rule, write it today: your host's identity is worth more than every credential your guests hold.
- Prefer vsock to a link-local address: no HTTP client can be tricked into speaking AF_VSOCK to an attacker-supplied URL, because there is no URL scheme for it.
- On a link-local address, require the v2 token flow — a PUT for a short-TTL session token, then that token as a header on every GET. A bare GET is the SSRF primitive.
- Keep the endpoint unreachable from anything the guest proxies for: a headless browser fetching user-supplied URLs, a webhook tester, an agent with a fetch tool.
- Rate-limit and log per sandbox: one asking for credentials four hundred times is telling you something.
And the limitation no hardening removes: a broker eliminates the long-lived secret, but not untrusted code in the guest asking the broker itself. Anything the guest can reach, the guest's worst process can reach. The broker makes the credential short-lived and scoped; the VM boundary makes it tenant-scoped. Both halves are load-bearing.
Shape 2: SPIFFE and short-lived SVIDs
SPIFFE is a specification for workload identity, and the vocabulary is worth learning because the acronyms hide good ideas. A SPIFFE ID is a URI — `spiffe://trust-domain/path` — naming a workload, say `spiffe://sandboxes.example/tenant/acme/job/4172`. An SVID (SPIFFE Verifiable Identity Document) proves a workload holds that ID: either an X509-SVID, a short-lived leaf certificate carrying the ID in a URI SAN, or a JWT-SVID whose `sub` is the ID. The trust bundle is the CA certificates or public keys for a trust domain — all a verifier needs, and public, so distributing it is not a secret-management problem.
The Workload API is what makes this a fit rather than just a standard. An agent on each host serves a unix domain socket; a workload connects, the agent attests it by inspecting properties of the calling process it cannot forge, and returns an SVID — no secret presented to get one. SVIDs are streamed rather than fetched, rotating continuously at short TTLs, so nothing has a rotation schedule. And per-workload beats per-host for an obvious reason: a host certificate says “this is host 17,” useless when host 17 runs two hundred tenants' sandboxes. A per-workload ID names the tenant and the job, and because it is a path you authorize against the path.
The hard part is node attestation. The agent must prove itself before it can issue anything, and the strong attestors want a TPM or a cloud instance identity document, neither of which a microVM has. The realistic shape is a one-time join token: minted at create time, injected into the VM, redeemed exactly once. One-time use is the whole security property — a copy stolen from the snapshot of a VM that already booted buys nothing. Keep its TTL in seconds and treat a second redemption as an incident, not a retry. The cost is a control plane plus every service verifying SPIFFE IDs instead of DNS names.
Shape 3: OIDC federation, where no key exists anywhere
The GitHub Actions pattern, and the first one I would reach for when the consumer is a cloud provider. Your platform becomes an OIDC issuer: publish `/.well-known/openid-configuration` and a JWKS, sign a short-lived JWT per sandbox describing that workload. The provider trusts your issuer and exchanges the JWT for scoped temporary credentials — `AssumeRoleWithWebIdentity` on AWS, workload identity federation on GCP. The property that earns the setup: no long-lived credential exists anywhere — not in the image, the guest, your database or your secrets manager. The only private key is your signing key, which lives in a KMS, is never exportable, and signs a 300-second assertion rather than authenticating to anything.
{
"1_the_token_your_platform_signs": {
"_illustrative": "A JWT your control plane signs per sandbox, TTL in minutes. Nothing here is a PandaStack API; this is the generic OIDC-federation shape.",
"iss": "https://id.example.com",
"sub": "tenant/acme/template/base/job/4172/",
"aud": "sts.amazonaws.com",
"iat": 1791273600,
"nbf": 1791273600,
"exp": 1791273900,
"jti": "01JZ9F1C2AB4K7QX",
"sandbox_id": "9f1c2ab4-6d1e-4f90-9c3a-11a2b3c4d5e6",
"tenant": "acme",
"template": "base",
"git_commit": "9f1c2ab",
"https://aws.amazon.com/tags": {
"principal_tags": { "tenant": ["acme"], "job": ["4172"] },
"transitive_tag_keys": ["tenant"]
}
},
"2_the_aws_role_trust_policy": {
"_illustrative": "Attach the OIDC provider once, then one of these per tenant (or one role plus ABAC on the session tags).",
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::123456789012:oidc-provider/id.example.com"
},
"Action": ["sts:AssumeRoleWithWebIdentity", "sts:TagSession"],
"Condition": {
"StringEquals": {
"id.example.com:aud": "sts.amazonaws.com"
},
"StringLike": {
"id.example.com:sub": "tenant/acme/*"
}
}
}
]
},
"3_the_bug_this_prevents": {
"_note": "A trust policy that names the provider and conditions on nothing accepts EVERY token your issuer will ever sign. Any tenant can then be any tenant, and the failure is silent until someone notices cross-tenant objects.",
"Condition_that_is_a_vulnerability": {},
"Condition_that_is_also_a_vulnerability": {
"StringLike": { "id.example.com:sub": "tenant/acme*" }
},
"why": "Without the trailing delimiter, tenant/acme* also matches tenant/acme-evil/... Always end a prefix match on the delimiter you chose.",
"aws_specific_constraint": "For a generic OIDC provider, IAM exposes only a small set of claims as condition keys (aud, sub, amr) -- your custom tenant claim is NOT one of them. That is precisely why GitHub Actions encodes repo and ref inside sub as repo:org/name:ref:refs/heads/main. Copy the idea: make sub a structured, hierarchical, delimiter-terminated string with the tenant leftmost. Verify the current IAM condition-key list; this set has changed before."
},
"4_the_gcp_equivalent": {
"_note": "GCP workload identity federation can map and condition on arbitrary claims, so the sub-stuffing trick is an AWS-shaped constraint rather than a universal one.",
"attribute_mapping": {
"google.subject": "assertion.sub",
"attribute.tenant": "assertion.tenant"
},
"attribute_condition": "assertion.tenant == 'acme' && assertion.aud == '//iam.googleapis.com/projects/1234/locations/global/workloadIdentityPools/pandastack/providers/sandboxes'",
"then": "Grant the principalSet for attribute.tenant/acme on exactly the bucket prefix it needs, and downscope further with a Credential Access Boundary if the job only touches one object."
}
}
The paragraph everybody skips is the trust policy, and it decides whether this is a security improvement or a cross-tenant vulnerability with excellent documentation. A policy that names your provider and conditions on nothing accepts every token your issuer ever signs: any tenant can be any tenant, silently, until somebody notices objects in the wrong prefix. Pin `aud` with `StringEquals`. Pin `sub` with a prefix that ends on your delimiter, because `tenant/acme*` also matches `tenant/acme-evil/` and an attacker who can influence a tenant slug owns your federation.
There is also a constraint that explains something you may have wondered about. For a generic OIDC provider, IAM exposes only a small set of claims as condition keys — `aud`, `sub`, `amr` — and your custom `tenant` claim is not among them. That is why GitHub Actions stuffs the repository and git ref inside `sub`, producing the famously ugly `repo:org/name:ref:refs/heads/main`. Not a design flaw: the only field the trust policy can see. Copy it: make `sub` structured, hierarchical and delimiter-terminated, tenant leftmost. GCP differs, since attribute mapping and CEL conditions read arbitrary claims.
One extra stops you creating an IAM role per sandbox: AWS reads a claim named `https://aws.amazon.com/tags` and turns it into session tags, which resource policies reference as `aws:PrincipalTag/tenant` — one role plus ABAC instead of a role per tenant.
The options, compared honestly
| Approach | Long-lived secret exists? | Readable by untrusted guest code? | In the snapshot artifact? | Rotation cost | What a leak gets an attacker |
|---|---|---|---|---|---|
| Key baked into the template image | Yes; its lifetime is the artifact's | Yes — on the filesystem of every restore | Yes, and in every host's local copy | Rebuild, re-publish, re-sync the fleet. In practice: never | Everything that key can do, until somebody notices |
| Key passed as an env var at create | Yes, upstream: you still hold the real key | Yes: /proc/<pid>/environ, crash dumps, logs | Only if you snapshot after injecting it | One place to change; re-issue every consumer | Everything it can do, until the upstream rotation |
| Host broker over vsock or link-local | No — minted per sandbox, minutes long | Yes — the guest can re-ask the broker | No, if asked after restore | Nothing to rotate; change the policy | One tenant's scope for the remaining TTL |
| SPIFFE X509 or JWT SVID | No — continuously rotated, never stored | Yes, via the Workload API socket in the guest | No. The join token might be — hence one-time-use | Rotation is the steady state; you run a control plane | One workload's identity until the SVID expires |
| OIDC federation / token exchange | None anywhere; only a KMS signing key | Yes if the guest exchanges; no if the host does | No — minted per sandbox after restore | No key rotation at all | One pinned sub, one role, 15 min. Unpinned sub: every tenant |
| Host-side signing proxy | No | No — it holds a capability, not a credential | No | None | One presigned operation on one object. Usually nothing |
That last row is the quiet winner for more jobs than it gets credit for. If the sandbox only needs to upload one artifact, give it a presigned URL for one object rather than a credential: it then holds a capability rather than a key, and untrusted code stealing it gets to do the upload the job was going to do anyway.
Scope, TTL, and what a leak is actually worth
All three converge on the same discipline, and this is where most of the real risk reduction lives. The architecture decides whether a long-lived secret exists; the scoping decides what the short-lived one costs you when it escapes.
- One tenant's prefix, one bucket, one action set — and look hard at reads, because a credential that can ListBucket enumerates your customers even if it never reads an object.
- Minted per sandbox, not per worker. A credential shared by every job on a box has the blast radius of the box.
- Intersect, do not merely grant: an AWS session policy on the exchange, or a GCP Credential Access Boundary, narrows the credential at mint time.
- Know how to revoke before you need to. On AWS, “revoke sessions” is a deny policy on the role conditioned on aws:TokenIssueTime — put that JSON in a runbook now, not at 3am.
Now the honest tension, because the brochure version does not survive the APIs. `AssumeRoleWithWebIdentity` enforces a floor of 900 seconds on `DurationSeconds` — verify against current STS docs, but it has been stable for years — so a 90-second sandbox cannot be issued an 80-second credential. Its credential outlives it by roughly fourteen minutes however carefully you asked, and deleting the VM revokes nothing: the two lifetimes are separate clocks and only one is yours.
So stop optimising that number and spend the effort on scope. Fourteen surplus minutes of “may PutObject under one prefix” is a nuisance; fourteen of “may read any object in the account” is a Friday evening you will remember. PandaStack's reaper enforces the sandbox TTL, but killing the VM is not killing its credentials — a credential that outlives the VM is a leaked key with extra steps.
Make the audit log name the sandbox
An identity you cannot trace is decoration. The test: when an unexpected API call shows up in your audit log at 2am, can you get from that line to a tenant, a job and a human in under a minute? Two halves, both cheap if wired at the start.
The cloud half is one argument. Set `RoleSessionName` to the sandbox id on the exchange. CloudTrail then records the exchange and every subsequent call made with those credentials under `arn:aws:sts::<account>:assumed-role/<Role>/<sandbox-id>`, so an S3 `GetObject` nobody expected names the VM that made it. Session names allow 2–64 characters of `[\w+=,.@-]`, so a UUID fits and a free-text label may not.
The platform half is `metadata` on create: a string map that travels with the sandbox into your lifecycle events. Put the tenant, job id, commit and requester in it, and the 2am question becomes a lookup rather than an investigation across four systems with no common key.
# The control-plane side. The PandaStack calls are real; mint_workload_token()
# and the unix-socket listener are yours to write -- there is no PandaStack
# identity API and I am not going to invent one.
import json
import boto3
from pandastack import Sandbox
sbx = Sandbox.create(
template="base",
ttl_seconds=180, # enforced by a platform-side reaper, so a crashed
# orchestrator cannot leak this VM past 3 minutes
metadata={ # Dict[str, str] -- this is your audit join key
"tenant": "acme",
"job_id": "4172",
"requested_by": "ci",
"git_commit": "9f1c2ab",
},
)
# A short-lived statement ABOUT this sandbox, signed by your issuer. Note what
# is not happening: the sandbox is never handed a cloud key, and your platform
# does not hold one either. The only private key in the system is the signing
# key, and that should live in a KMS/HSM and never be exportable.
token = mint_workload_token(
subject=f"tenant/acme/template/base/job/{sbx.id}/",
audience="sts.amazonaws.com",
ttl_seconds=300,
claims={"sandbox_id": sbx.id, "tenant": "acme", "template": "base"},
)
# The exchange. Two arguments carry most of the value here.
#
# RoleSessionName: this string is what CloudTrail prints for every subsequent
# call made with these credentials, as
# arn:aws:sts::123456789012:assumed-role/<Role>/<RoleSessionName>. Put the
# sandbox id in it and an S3 PutObject you did not expect names the VM that
# made it. (Session names are limited to 2-64 chars of [\w+=,.@-]; a UUID
# fits, an arbitrary label might not.)
#
# Policy: a session policy intersects with the role's policy, so the role can
# be per-tenant and coarse while the SESSION is per-job and surgical. This is
# how you avoid one IAM role per sandbox.
sts = boto3.client("sts")
assumed = sts.assume_role_with_web_identity(
RoleArn="arn:aws:iam::123456789012:role/sandbox-tenant-acme",
RoleSessionName=sbx.id,
WebIdentityToken=token,
DurationSeconds=900, # 900s is the FLOOR for this API, not a choice.
# Verify against current STS docs.
Policy=json.dumps({
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["s3:PutObject"],
"Resource": "arn:aws:s3:::acme-artifacts/tenant/acme/job/4172/*",
}],
}),
)
# Log the MINT, never the credential. Subject, audience, scope, TTL, sandbox id,
# and the jti so you can correlate a specific token with a specific CloudTrail
# event. Logs are the second most common place credentials leak, right after
# images -- and unlike an image, a log is replicated to three vendors by lunch.
log.info(
"minted workload credential sandbox=%s tenant=%s sub=%s ttl=%ds expires=%s",
sbx.id, "acme", f"tenant/acme/template/base/job/{sbx.id}/", 900,
assumed["Credentials"]["Expiration"].isoformat(),
)
# Hand the TOKEN to the guest over the per-sandbox socket and let the guest do
# its own exchange, or do the exchange here and hand over the credential the
# same way. What you must not do is put either one in metadata (queryable and
# logged), in an env var (inherited by every child), or on an exec command line
# (visible in /proc/<pid>/cmdline to everything in the guest).
One rule for that code path: log the mint, never the credential. Record subject, audience, scope, TTL, sandbox id and the token's `jti` so you can tie an assertion to a cloud event — never the token, never the response. Logs are the second most common place credentials leak after images, and unlike an image a log is replicated to three vendors by lunchtime.
A note on the problem this is not
Two things sound identical and are not. PandaStack's control-plane-to-agent authentication is a node token held on the host (`PANDASTACK_NODE_TOKEN`), presented by a long-lived machine you provisioned. That is a fleet authentication problem — a shared secret on a durable node, with its own rotation story — not a guest proving who it is. Different lifetime, different threat model, different fix. The sibling post below covers it.
What to build, in order
- Audit what is baked — not “is it on the filesystem” but “did any secret enter guest RAM before the snapshot?” Anything that did is published; the fix is rotation.
- Drop 169.254.0.0/16 from guest egress at the host, ahead of the ACCEPTs. One rule stands between untrusted guest code and your host's cloud identity.
- Make every bake-time value a placeholder, and give the guest two delivery channels that do not share a failure mode.
- Pin sub and aud, terminate the prefix on your delimiter, then write the negative test: assume tenant B's role with tenant A's token and assert it fails. Almost nobody writes it, and it is the only test that catches the bug that matters.
- For jobs that write one artifact, skip all of it — presign the object and let the sandbox hold nothing worth stealing.
The slightly uncomfortable conclusion is that a VM living ninety seconds can have better credential hygiene than your laptop, precisely because it never had time to accumulate anything: no key old enough to have been copied, no token old enough to have been pasted into a chat. That is the property to design for, not a limitation to work around. The identity does not come from the VM — it comes from the fact that you created it on purpose, for a reason you can name.
Frequently asked questions
Is it safe to bake a secret into a template image if the image registry is private?
No, and the registry's access control is not the reason. Three problems survive making the image private. Distribution: an image exists to be copied cheaply to as many hosts as possible, so a secret in one has a copy count that grows with your fleet and a lifetime equal to the artifact's. Rotation: changing the key means rebuilding, re-publishing and re-syncing every host, which nobody does on a schedule. And the one people miss: if your build snapshots a running VM, the artifact is not just a filesystem image. A Firecracker snapshot includes vm.mem, a byte-for-byte freeze of guest RAM, so a token your bootstrap fetched and never wrote to disk is in there anyway, along with the TLS session keys of the connection that fetched it. That multi-gigabyte file is then replicated to object storage and cached on every host's local disk. Private or not, you have published it. The only credential that belongs in an image is a public key — PandaStack bakes the host agent's SSH public key into template rootfs images and keeps the private half on the host at mode 0600.
Can untrusted code inside the sandbox steal a credential the guest fetched from a host broker?
Yes, and anyone who tells you otherwise is selling something. If the untrusted code runs as root in the same guest — which in a single-tenant microVM it usually does — it can read /proc/<pid>/environ for every process, read any credential file, attach to a running process, or simply call the broker itself. A broker creates no in-guest trust boundary; the guest has one privilege level that matters. What it genuinely fixes is different and still valuable: it removes the long-lived secret, so what leaks expires in minutes and is scoped to one tenant's prefix rather than working account-wide until somebody notices. The in-guest hygiene measures remain worth doing because they shrink the accident surface — use credential_process so the credential lives in one process's memory rather than an inherited environment variable, keep secrets off exec command lines where they appear in /proc/<pid>/cmdline — but understand them as hygiene, not a boundary. The boundary is the VM — the only line an in-guest compromise cannot step over, and the real argument for one microVM per tenant.
Should I use a host broker, SPIFFE, or OIDC federation?
Decide by who consumes the identity. If it is a cloud provider — S3, GCS, Secrets Manager, a registry — use OIDC federation. It is the only option where no long-lived credential exists anywhere: your platform signs a short assertion with a KMS-held key and the provider exchanges it for temporary credentials, so there is nothing to store and nothing to rotate on a calendar. If it is your own control plane or internal services, and you want mutual TLS where both ends verify an identity rather than a hostname, use SPIFFE: per-workload IDs, continuously rotated SVIDs, and a public trust bundle instead of a shared secret. Budget for a server plus a per-host agent, and solve node attestation deliberately — a microVM has no TPM and no instance identity document, so the realistic path is one-time join tokens with TTLs in seconds. A host broker is the pragmatic middle, and a good answer when a sandbox needs exactly one scoped credential and you do not want to run an identity control plane. They also compose: a broker that performs the OIDC exchange on the host is a good design, because the guest never sees the assertion.
How short can the credential TTL be, and what if the sandbox is shorter than that?
Shorter than you want, and the gap is real rather than a misconfiguration. AssumeRoleWithWebIdentity enforces a floor of 900 seconds on DurationSeconds — verify against current STS docs, but it has been stable for years — so a sandbox that lives 90 seconds cannot be issued an 80-second credential. Its credential outlives it by about fourteen minutes however carefully you asked, and deleting the VM revokes nothing: the two lifetimes are separate clocks and only one is yours. So stop optimising the number and move the effort to scope. Fourteen surplus minutes of 'may PutObject under one prefix' is a nuisance; fourteen of 'may read any object in the account' is an incident. Intersect the role's permissions with a session policy at mint time, use a GCP Credential Access Boundary for the equivalent downscoping, and make sure the credential cannot list anything, because enumeration is a disclosure even without a read. Then make revocation a runbook: on AWS that is a deny policy on the role conditioned on aws:TokenIssueTime, written down before the night you need it. And for jobs that only write one artifact, sidestep the question with a presigned URL for one object.
Keep reading
- The security gotchas of Firecracker snapshots — The long version of the memory-dump problem: what leaks into vm.mem, and the CRNG entropy reuse that comes with it.
- Firecracker MMDS: passing config into a microVM — The mechanism behind the link-local path in this post, including the v1-versus-v2 hardening.
- Authenticating a microVM host fleet — The other problem: node tokens and mTLS between a control plane and long-lived hosts, which is not guest identity.
- Running jobs with customer-supplied cloud credentials — When the credential belongs to your customer rather than to you, and the isolation that implies.
- Managing environment variables and secrets — The practical layer underneath this one: where values actually live and who can read them.
- Secure CI secrets in a microVM — The same architecture applied to build pipelines, where the untrusted code is a pull request.
- Controlling network egress for untrusted code — The other half of blast radius: a scoped credential still needs somewhere it is allowed to go.
Related posts
- Sandboxing AI-Generated Terraform and Infrastructure-as-Code
An agent that writes Terraform and runs `terraform apply` can delete your prod database or exfiltrate your cloud keys — IaC tooling executes arbitrary providers and reads the environment. Run `plan` in a disposable microVM, gate `apply` behind a human review of the diff.
- Sandboxing an AI Agent That Operates Your Kubernetes Clusters
The dangerous thing you hand a Kubernetes agent is not a filesystem, it's a kubeconfig. Isolate the toolchain in a disposable microVM, mint a 10-minute ServiceAccount token per run, and make `kubectl diff` the step that can't be skipped.
- Isolating AI Agents That Publish Packages
Your release token is one `cat ~/.npmrc` away from every postinstall script in your dependency tree. Build in a credential-free microVM, approve the artifact hash, then upload from a second VM that dies afterward.
- Isolating Customer-Managed Key Operations in microVMs
You shipped BYOK to close an enterprise deal. Now your process unwraps tenant A's data key on a heap that also serves tenant B — and a heap dump is a remarkably efficient way to violate a data-processing agreement.
- Isolating Genomics Pipelines in Per-Job microVMs
A 400GB BAM file and a `while True` in someone's cleanup script are both just Tuesday. Run each user-submitted Nextflow/Snakemake/WDL job in its own Firecracker microVM, so a runaway alignment can't OOM another lab's run and one tenant's PHI is never reachable by another.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.