How a Control Plane Authenticates Its Host Fleet: Node Tokens, Rotation, and mTLS
There is an enormous amount written about how your users log in. OAuth flows, JWT validation, session rotation, passkeys — the authentication surface that faces humans is the best-documented part of most systems. Then there is the other authentication surface, the one that faces machines, and the public literature on it is mostly a vendor diagram with a padlock on the arrow. The question it answers badly is this: when `agent-7` — a Linux box with `/dev/kvm`, a Firecracker binary and a few hundred gigabytes of snapshot store — sends a heartbeat to your API, how does the API know it is `agent-7`? And when the API turns around and tells `agent-7` to run a command inside a guest, how does `agent-7` know it is the API?
I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. The control plane is a Go API over Postgres; the fleet is per-host Go agents on KVM boxes. This post is the honest progression of that one credential: a shared bearer token, which is what we actually run and what most fleets should start with; how to rotate it without an outage, which is the part that bites; and mTLS with per-host certificates, which is the next rung and costs more than the diagrams admit. It ends with the thing authentication does not buy you, because an authenticated host can still be useless.
Two directions, two completely different threat models
Draw the system first, because the shape determines everything. A control plane — REST API, Postgres for state, ClickHouse for events — talks to N per-host agents. Three flows matter:
- Agents heartbeat into the control plane every ten seconds or so, carrying capacity: free CPU, free memory, free network slots, which snapshot artifacts they hold locally, whether they support streaming restore. The scheduler reads exactly this to decide where the next sandbox lands.
- The control plane calls agents to create, delete, pause, snapshot and fork microVMs, and proxies the data plane — `exec`, filesystem reads and writes, log streams, interactive PTY sessions — to whichever agent currently holds the sandbox.
- Agents hold leases, renewed on a short interval with a longer timeout, so that when a host dies something can safely conclude it died and reclaim what it owned.
Both directions need authentication, and here is the part people flatten: the two directions have asymmetric consequences on failure. A forged heartbeat is a scheduling attack. Claim 64 free vCPUs and 500 GiB of free memory and you will win every placement decision the scheduler makes — our score is a deterministic function of free capacity, so a liar with the best numbers gets the whole burst. That is a denial of service dressed up as a capacity report, and it is quiet: nothing errors, placements just go to a host that cannot honour them.
A forged call in the other direction is remote code execution on a machine with hardware virtualisation. Read the agent's job description literally: it forks and execs a VMM, writes files into guest root filesystems, reflinks disk images, and runs arbitrary commands inside guests on request. An unauthenticated agent port is not "an internal API" — it is an unauthenticated `exec` endpoint on a host that is also running other tenants' virtual machines. There is no clever mitigation that makes that acceptable; the port either requires a credential or it does not exist.
That asymmetry has one practical design consequence worth copying. Our agent exposes its full router over a Unix socket for the local daemon on the same box, and over TCP only when it needs to be reachable across the VPC. The TCP listener is the dangerous one, so the agent refuses to open it at all if no token is configured — it logs a warning and binds nothing. Fail closed on the path where failing open is a root shell.
The shared bearer token, described honestly
What PandaStack actually does is the simplest thing that is correct against the open internet: a single shared secret, `PANDASTACK_NODE_TOKEN`, present on every agent and on the control plane, sent as an `X-Node-Token` header. The control plane injects it on every call to an agent; the agent's TCP middleware rejects every request that does not carry it, with `/healthz`, `/metrics` and `/version` exempted so that cloud load balancers can do their job. The same token gates the control plane's own node-facing endpoints, so the symmetry is genuine: one secret, both directions.
What that buys is real and should not be sneered at. It is one environment variable, which means it is deployable by every mechanism you already have. It is correct against the thing that will actually probe your ports, which is the internet, continuously, forever. It has no clock dependency, no certificate expiry, no CA to run, and no failure mode that begins "the renewal daemon died six weeks ago". For a fleet of five hosts owned by one team, it is the right answer, and choosing it deliberately is better engineering than half-finishing a PKI.
What it does not buy is identity, and that is the whole limitation in one word. The token proves membership — the caller is somebody who holds the fleet secret — and nothing more. Four consequences follow, and they are not subtle:
- Blast radius of one compromised host is the fleet. Every host holds a credential that can impersonate every other host and can call the control plane's node-scoped endpoints. Compromise one box, and from the control plane's point of view you are now every box.
- There is no revocation. Removing one host's access means changing the secret everywhere, which is the rotation problem below, which is why nobody does it quickly in an incident.
- There is no per-caller authorisation. You cannot say "agent-7 may only touch agent-7's sandboxes" if the credential does not distinguish agent-7 from agent-12. Anything that looks like per-agent authorisation is reading an agent id out of a header or a request body — which is to say, asking the caller who it would like to be.
- Your audit log is weaker than it looks. "Which agent deleted this sandbox?" is answered by a field the caller supplied. That is a hint, not evidence.
And then there is the operational reality of any long-lived shared secret, which is that its true distribution is wider than its intended distribution. It is in your config management, in a secret manager, in the systemd environment file, in a CI variable, in a debug script in somebody's home directory, in the terminal scrollback of the engineer who was on call during the last incident, and in the runbook that four people have now pasted into Slack. None of those are theoretical. One of them is why you are reading the rotation section.
The verifier you want on day one
Write the verifier to accept two values from the very beginning, even while the second one is empty. It costs eight lines now and it is the difference between a rotation being a config change and a rotation being a change-management meeting. The version below also hashes before comparing, which removes the one piece of information a raw constant-time compare still leaks.
// Package nodeauth authenticates the per-host agents to the control plane and
// the control plane to the agents. It is deliberately boring: one shared
// bearer token in a header, verified in constant time, with a SECOND accepted
// value so the token can be rotated without a fleet-wide outage.
package nodeauth
import (
"crypto/sha256"
"crypto/subtle"
"errors"
"net/http"
"strings"
"sync/atomic"
)
// Verifier accepts a primary token and, during a rotation window, a secondary.
// We store DIGESTS, not the strings. subtle.ConstantTimeCompare is constant
// time in the contents but returns 0 immediately when the lengths differ, so
// comparing raw bytes still leaks the token's length. Hashing to a fixed 32
// bytes removes even that, and means the plaintext lives in exactly one place:
// the argument to New.
type Verifier struct {
primary [32]byte
secondary [32]byte
hasSecondary bool
// SecondaryHits is the whole point of step 4 of a rotation. You do not
// delete the old token because the rollout "looked fine". You delete it
// when this counter has been flat for longer than the restart interval of
// the slowest process in your estate -- including the cron job you forgot.
SecondaryHits atomic.Uint64
}
func New(primary, secondary string) (*Verifier, error) {
// Trim. A token read from a systemd EnvironmentFile, a Kubernetes secret or
// `gcloud secrets versions access` very often arrives with a trailing
// newline, and "the token is correct but auth fails" costs an afternoon the
// first time and ten minutes every time after that.
primary = strings.TrimSpace(primary)
secondary = strings.TrimSpace(secondary)
if primary == "" {
// Fail CLOSED. An empty expected token must never degrade to "accept
// anything": on a host with /dev/kvm that is an open exec endpoint.
// Our agent goes further and refuses to bind the TCP listener at all.
return nil, errors.New("nodeauth: primary token empty, refusing to serve")
}
v := &Verifier{primary: sha256.Sum256([]byte(primary))}
if secondary != "" && secondary != primary {
v.secondary = sha256.Sum256([]byte(secondary))
v.hasSecondary = true
}
return v, nil
}
// accepts reports whether the presented token matches either accepted value.
// Both comparisons always run -- no `||` short circuit, no early return -- so
// the time taken does not reveal WHICH token matched. That matters more than it
// sounds: during a rotation window, a timing difference between "primary" and
// "secondary" tells an attacker which value is on its way out.
func (v *Verifier) accepts(presented string) bool {
sum := sha256.Sum256([]byte(strings.TrimSpace(presented)))
new_ := subtle.ConstantTimeCompare(v.primary[:], sum[:])
old := 0
if v.hasSecondary {
old = subtle.ConstantTimeCompare(v.secondary[:], sum[:])
}
if old == 1 {
// Someone is still presenting the value we are trying to retire.
// Export this as a counter and alert on it being NON-zero after the
// rollout window, not on it being non-zero at all.
v.SecondaryHits.Add(1)
}
return (new_ | old) == 1
}
// Middleware gates every route except the ones a cloud load balancer must
// reach unauthenticated. Keep that list short and keep it boring: whatever is
// in it is public to anything that can route to the port, forever. /healthz
// must not grow an "and here are the running sandboxes" field in six months,
// and /metrics must not carry sandbox ids or tenant slugs in label values.
func (v *Verifier) Middleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
switch r.URL.Path {
case "/healthz", "/metrics", "/version":
next.ServeHTTP(w, r)
return
}
if !v.accepts(r.Header.Get("X-Node-Token")) {
// One response for every failure. "wrong length", "expired token"
// and "unknown agent" are three separate oracles, and the person
// they help is never your on-call engineer.
http.Error(w, `{"error":"node token invalid"}`, http.StatusUnauthorized)
return
}
next.ServeHTTP(w, r)
})
}
Rotating the fleet credential without an outage
This is the most useful section in the post, because the naive version is a total outage and it is not obvious in advance how total. Flip the control plane's expected value and every agent still presenting the old token is instantly unauthenticated. Watch what that actually does: heartbeats start failing, so within thirty seconds every host looks stale, so the scheduler has zero candidates, so every create returns a capacity error. Meanwhile every `exec`, every file read, every log stream to a running sandbox 401s. You have not degraded the platform. You have turned it off, and the error message says "no capacity", which will send your first responder to the wrong dashboard.
The invariant to hold is one sentence, and if you internalise nothing else from this post, take this: at every instant, the set of values the verifiers accept must be a superset of the set of values the callers present. Every safe rotation is a sequence of steps that never breaks that containment, which gives you the shape:
- Teach every verifier to accept old OR new. Deploy this everywhere — both the agents and the control plane, since both verify — before any caller changes. Nothing presents the new value yet; you are only widening what is accepted.
- Roll the new value out to every caller. Now the accepted set is {old, new} and the presented set is shrinking from {old} toward {new}. Containment holds throughout, which is why this step can take hours or days and be interrupted safely.
- Prove that nothing presents the old value any more. Not "the rollout script exited 0" — prove it, with a counter and a grep.
- Remove the old value from the verifiers, then verify that a request bearing it now gets a 401. An un-narrowed rotation is not a rotation; it is two valid credentials and a false sense of progress.
#!/usr/bin/env bash
# Rotate the fleet's shared node token with no outage window.
#
# THE INVARIANT: at every instant, the set of values the VERIFIERS accept is a
# superset of the set of values the CALLERS present. Break it for one second
# and every heartbeat 401s, every host looks stale to the scheduler, and every
# create fails with "no capacity" -- which is a lie, and an expensive one,
# because it points the on-call engineer at the capacity dashboard.
set -euo pipefail
SECRET=pandastack-node-token
AGENTS=$(grep -v '^#' /etc/pandastack/fleet.txt) # one host per line
API_REPLICAS=$(grep -v '^#' /etc/pandastack/edges.txt)
# ------------------------------------------------------------- 0. the new value
NEW=$(openssl rand -base64 48 | tr -d '\n')
NEW_VER=$(printf %s "$NEW" \
| gcloud secrets versions add "$SECRET" --data-file=- --format='value(name)')
echo "new token is version ${NEW_VER##*/}"
# PIN THE VERSION NUMBER IN EVERY CONSUMER. A unit that reads ":latest" rotates
# itself at a moment you did not choose -- whenever that particular process next
# restarts, which across a fleet is "continuously, in no order, and differently
# on the host that happened to reboot". You get a rotation with no rollout, no
# window and no way to answer "which hosts have the new one?". A secret named
# `latest` is a secret whose rotation schedule is vibes.
# ------------------------------------- 1. teach every VERIFIER to accept BOTH
# Both sides verify, so both sides need the secondary. Roll this everywhere and
# confirm it everywhere BEFORE step 2. The ordering is the entire trick.
# PANDASTACK_NODE_TOKEN=<old> # still the primary
# PANDASTACK_NODE_TOKEN_SECONDARY=<new> # accepted, not yet presented
for h in $AGENTS $API_REPLICAS; do
ssh "$h" "sudo install -m0600 -o root -g root /dev/stdin \
/etc/pandastack/node-token.secondary" <<<"$NEW"
ssh "$h" 'sudo systemctl reload pandastack-agent.service 2>/dev/null \
|| sudo systemctl reload pandastack-api.service'
done
# ------------------------------------------- 2. roll the NEW value to callers
#
# RESTARTING AN AGENT IS A DATA-PLANE EVENT, NOT A CONTROL-PLANE ONE.
# This agent forks and supervises firecracker processes. If the unit's stop
# timeout expires before they are gone, systemd SIGKILLs the whole cgroup and
# takes running guests -- and their tenants' work -- with it. In order of
# preference:
# a) reload the config in place (SIGHUP / `systemctl reload`): no VM touched.
# This is why the token is read from a FILE and not only from Environment=.
# b) drain: stop accepting creates, let sandboxes finish or move, then restart.
# c) a blind fleet-wide `systemctl restart`: the 14:00-on-a-Tuesday option,
# where the incident report writes itself and the title is your name.
for h in $AGENTS; do
ssh "$h" 'sudo mv /etc/pandastack/node-token.secondary /etc/pandastack/node-token \
&& sudo systemctl reload pandastack-agent.service'
# If the unit genuinely has no reload path, drain first and NEVER more than
# one host at a time -- a fleet-wide restart is also a fleet-wide capacity dip:
# ssh "$h" 'sudo pandastack-agent drain --wait 600 \
# && sudo systemctl restart pandastack-agent.service'
# Prove this host now accepts the new value before touching the next one.
code=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 \
-H "X-Node-Token: ${NEW}" "https://${h}:7070/sandboxes")
[ "$code" = "200" ] || { echo "$h: new token rejected ($code)" >&2; exit 1; }
echo "$h: ok"
done
# ---------------------- 3. PROVE nothing is still presenting the old value
# Three independent signals, because at least one of them will be lying.
# (a) The verifier's own counter, scraped from every agent AND every API replica.
# Flat for longer than the slowest restart interval in your estate -- not
# "flat for five minutes while I watched". Label it so the thing you are
# watching is literally "requests still using the token I am about to
# delete": once step 2 is done, make the OLD value the secondary.
for h in $AGENTS $API_REPLICAS; do
curl -sf "https://${h}:9100/metrics" \
| awk '/^pandastack_node_token_hits_total\{kind="retiring"\}/ {print FILENAME, $0}' \
| sed "s|^|${h} |"
done
# (b) Compare the LIVE value on each host against the new one by digest, so no
# secret ever lands in a shell argument, an ssh command line, or your
# history file. Note the `tr -d` -- the trailing newline is why the digests
# disagree the first time you run this.
NEW_SHA=$(printf %s "$NEW" | sha256sum | cut -d' ' -f1)
for h in $AGENTS; do
live=$(ssh "$h" "sudo tr -d '\n' < /etc/pandastack/node-token | sha256sum | cut -d' ' -f1")
[ "$live" = "$NEW_SHA" ] || echo "STRAGGLER: $h still holds a different token" >&2
done
# (c) grep the config-management repo for REFERENCES, not values. The thing
# still presenting the old token is usually not an agent at all: it is a
# cron job, a Terraform variable, an unmanaged edge VM, a debug script in
# somebody's home directory, or the runbook four people have pasted into
# Slack. Pinned references make this greppable; `latest` makes it unknowable.
if git -C /srv/infra grep -nE "${SECRET}/versions/(latest|[0-9]+)" \
| grep -v "versions/${NEW_VER##*/}"; then
echo "^ references not yet repointed at version ${NEW_VER##*/}" >&2
exit 1
fi
# ------------------------------- 4. narrow: remove the old, then VERIFY the 401
# An un-narrowed rotation is not a rotation. It is two valid credentials and a
# calendar reminder nobody will action.
for h in $AGENTS $API_REPLICAS; do
ssh "$h" 'sudo rm -f /etc/pandastack/node-token.secondary \
&& sudo systemctl reload pandastack-agent.service 2>/dev/null \
|| sudo systemctl reload pandastack-api.service'
code=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 \
-H "X-Node-Token: ${OLD:?export OLD before running step 4}" \
"https://${h}:7070/sandboxes")
[ "$code" = "401" ] || { echo "$h: OLD token still accepted ($code)" >&2; exit 1; }
done
gcloud secrets versions destroy "${NEW_VER%/*}/$((${NEW_VER##*/} - 1))" --quiet
echo "rotation complete; old version destroyed"
Three traps that are not in the diagram
First: pin a specific secret version, never `latest`. This is the one people argue with, so here is the failure precisely. A `latest` reference means the rotation takes effect whenever each consumer next reads the secret, which is whenever that process next restarts, which across a fleet is an unordered trickle spread over however long it takes for a host to reboot, a MIG to replace an instance, or an OOM killer to have an opinion. You cannot answer "which hosts are on the new value?", you cannot stage the change, and you cannot roll back, because rolling back means adding yet another version that happens to contain the old bytes. Pinning turns rotation into a deploy, which is a thing you already know how to do carefully.
Second: know which knobs are owned by instance creation and which by your config management, per knob, in writing. On our fleet the edge API's environment file is written by cloud-init when the VM is created, and config management only hot-swaps the binary. That means appending a value to that file and restarting the service is a perfectly good emergency change — and it silently reverts the next time the managed instance group replaces the instance. A durable change has to go into the thing that runs at creation, or into the code's own default. If you have ever watched a setting you definitely changed come back wrong after an unrelated host replacement, this is the mechanism. It is not drift; it is two systems both being right about a file they both believe they own.
Third, and the expensive one: on a host that supervises virtual machines, restarting the agent is a data-plane event. Our agent ships as a versioned release, and a rolling restart that exceeds its stop timeout gets the whole unit cgroup SIGKILLed — taking running guests with it. So the rotation mechanism has to be a configuration reload, or a drain followed by a restart, one host at a time. Which means the design decision happens long before the rotation: read the token from a file the process can re-read on SIGHUP rather than only from the environment, because an environment variable can only be changed by restarting the process, and restarting the process is the thing you are trying to avoid.
mTLS: one revocable identity per host
The next rung replaces the shared secret with a per-host client certificate signed by a CA you control. Both ends present a certificate, both ends verify the other against the trust bundle, and the identity is now cryptographic rather than declared. What that fixes maps one-to-one onto the four consequences above: the blast radius of one compromised host is one revocable identity rather than the fleet; the control plane can authorise per agent, because it finally knows which agent it is talking to; the audit log records a verified subject instead of a header; and the agent authenticates the server too, so a hijacked DNS record or a misrouted internal hostname does not become a command channel.
The costs are real, and the only honest way to present them is as operations you now own forever.
- A CA. Its private key is now the single most valuable secret in your infrastructure, and it needs to live somewhere better than the host that issues certificates. An intermediate per environment, an offline root, and an answer to "what happens when the root expires in ten years" that is written down rather than assumed.
- A renewal path that does not restart anything. Certificates expire on a schedule you chose, which means renewal is a cron job whose failure mode is a fleet-wide outage at a predictable time. And if renewal requires a process restart, you have coupled certificate expiry to the data-plane event described above.
- A revocation story, and both standard answers are awkward. CRLs are a file that must be distributed and is therefore always a little stale, with an operational question about what you do when fetching it fails. OCSP puts a network dependency in the authentication path — and OCSP stapling, which solves that for server certificates, is poorly supported for client certificates, which is exactly the direction you care about here. Short-lived certificates are usually the better answer: if a credential lives an hour, revocation is mostly "stop renewing it", and the window you are exposed to is bounded by design rather than by a distribution mechanism.
- A clock. Certificate validity is wall-clock, so skew becomes an authentication failure, and short lifetimes make skew worse rather than better: ten minutes of drift against a one-hour certificate is a large fraction of its life.
That last one deserves a paragraph of its own in a post about microVMs, because snapshot-restored machines have a specific and nasty version of it. A guest restored from a snapshot resumes with the clock exactly as it was when the snapshot was taken — which, for a template baked last Tuesday, means it wakes up believing it is last Tuesday. We shipped a real bug of precisely this shape: TLS failing inside restored guests because certificate validity checks were being made against a frozen clock, presenting as inexplicable network problems. The fix is a clock synchronisation step on restore, resume and wake, before anything touches TLS. If you move to short-lived certificates on anything that can be suspended or restored — hosts included — make clock sync a precondition of the credential, not a background nicety.
SPIFFE is the productised version of all of this, and worth naming because it saves you from reinventing it badly. A SPIFFE identity is a URI — `spiffe://trust-domain/path` — carried in an SVID, which in practice is an X.509 certificate with that URI in a SAN (there is a JWT form too, for places where you cannot do mTLS). A node agent attests each workload, issues it a short-lived SVID, rotates it automatically, and distributes the trust bundle so verification does not need a file somebody remembered to copy. If your answer to "short-lived per-workload certificates" is a shell script and a cron entry, you are writing a worse SPIFFE implementation; read theirs first. The sibling post on workload identity for ephemeral microVMs covers the guest-side half of this — how a VM that lives ninety seconds gets an identity at all — which is a harder problem than the host-side one, because hosts at least persist.
#!/usr/bin/env bash
# Issue a short-lived client certificate for one agent host.
#
# The identity that matters is the URI SAN. The control plane reads the agent id
# out of the VERIFIED certificate and ignores any agent id in the request body
# or headers -- that substitution is the entire difference between
# authentication and a self-declared label. If your code still trusts a header
# after you have deployed mTLS, you have bought a CA and kept the bug.
set -euo pipefail
AGENT_ID=${1:?usage: issue-agent-cert <agent-id>}
CA=/etc/pandastack/ca # ca.crt + ca.key, and ca.key is NOT here in
# production -- it lives in a KMS/HSM and
# this step becomes an API call to it.
OUT=/etc/pandastack/tls
HOURS=24 # Short-lived ON PURPOSE. This IS the
# revocation story: CRL distribution is
# always stale and OCSP stapling is not
# meaningfully available for client certs.
umask 077
openssl req -new -nodes -newkey ec -pkeyopt ec_paramgen_curve:P-256 \
-keyout "${OUT}/agent.key.new" -out /tmp/${AGENT_ID}.csr \
-subj "/CN=${AGENT_ID}"
# Minimal extensions. clientAuth only: a host certificate that is also valid
# for serverAuth is a certificate that can impersonate your control plane to
# its own fleet.
cat >/tmp/${AGENT_ID}.ext <<EOF
basicConstraints = critical,CA:FALSE
keyUsage = critical,digitalSignature
extendedKeyUsage = clientAuth
subjectAltName = URI:spiffe://pandastack.internal/agent/${AGENT_ID}
EOF
# -days only has day granularity, so for an hours-long certificate use explicit
# validity (-not_before/-not_after need a recent OpenSSL; check `openssl
# version` before copying this) and BACKDATE the notBefore by a few minutes.
# Every fleet has one host whose clock is slightly behind, and a certificate
# that is not yet valid fails in exactly the same way as one that is forged.
openssl x509 -req -in /tmp/${AGENT_ID}.csr \
-CA "${CA}/ca.crt" -CAkey "${CA}/ca.key" -CAcreateserial \
-not_before "$(date -u -d '-5 min' +%Y%m%d%H%M%SZ)" \
-not_after "$(date -u -d "+${HOURS} hours" +%Y%m%d%H%M%SZ)" \
-sha256 -extfile /tmp/${AGENT_ID}.ext \
-out "${OUT}/agent.crt.new"
rm -f /tmp/${AGENT_ID}.csr /tmp/${AGENT_ID}.ext
# Atomic swap, then signal. Two things to get right:
# 1. rename(2) so a reader never sees a key without its certificate.
# 2. NO RESTART. The server should hold a tls.Config whose GetCertificate /
# GetClientCertificate callback re-reads this pair, so renewal never
# becomes a process restart -- because on a host supervising firecracker
# processes, a restart that overruns its stop timeout kills running guests.
mv -f "${OUT}/agent.key.new" "${OUT}/agent.key"
mv -f "${OUT}/agent.crt.new" "${OUT}/agent.crt"
systemctl reload pandastack-agent.service
# Server side, for completeness -- this is the part that is easy to get subtly
# wrong, because a tls.Config that does not REQUIRE and VERIFY a client cert
# will happily accept a connection with none and leave you feeling secure:
#
# tls.Config{
# ClientAuth: tls.RequireAndVerifyClientCert, // not VerifyClientCertIfGiven
# ClientCAs: fleetPool, // your CA only, never the
# // system roots -- those
# // would let any public CA
# // mint one of your agents
# MinVersion: tls.VersionTLS13,
# }
#
# ...and then, in the handler, derive the agent id from
# r.TLS.PeerCertificates[0].URIs[0] and use THAT for every authorisation
# decision and every audit record.
echo "issued ${AGENT_ID}, valid ${HOURS}h"
Four rungs, compared
Pick a rung deliberately and write down why. The most common mistake is not choosing the wrong one; it is choosing a high rung, implementing two thirds of it, and ending up with the operational burden of mTLS and the security properties of a shared secret.
| Mechanism | Blast radius of one compromised host | Revocation | Rotation cost | Does the agent authenticate the server? | Operational complexity |
|---|---|---|---|---|---|
| One shared bearer token | The whole fleet — every host can impersonate every other host, and call the control plane's node endpoints | None short of rotating everywhere | A staged accept-both rollout across every host and every control-plane replica | Only that the caller holds the fleet secret — which every agent also does | Lowest. One environment variable, no clock dependency, no expiry |
| Per-host bearer tokens | One host, and only if your authorisation is actually scoped to it | Delete one row; takes effect as fast as your verifier's cache | Per host, independently — so no fleet-wide window at all | No. Still a one-way bearer credential | Low–medium. You now need a token store, issuance at provisioning, and a cache you must not trust for a negative |
| mTLS with a private CA | One revocable identity, with authorisation scoped by verified subject | CRL or OCSP, both awkward; short lifetimes are the practical answer | Renewal is continuous and automated, so there is no rotation event — but there is an expiry event if renewal breaks | Yes, mutually, against a trust bundle you control | Medium–high. A CA, a renewal daemon, a trust-bundle distribution path, and a clock that has to be right |
| SPIFFE / SVIDs | One attested workload identity, short-lived by construction | Mostly "stop renewing" — the window is the SVID lifetime | Automatic; rotation stops being a procedure | Yes, with trust-domain semantics and federation across domains | Highest to adopt, lowest to run once adopted. Verify current behaviour against the project's own docs — this ecosystem moves |
Per-host bearer tokens are the underrated middle. They fix the blast radius and give you real per-caller authorisation for a fraction of the mTLS effort — no CA, no expiry, no clock. What they do not give you is mutual authentication or short lifetimes, and they put a database lookup in your hot path, which is where the cache warning in the next section comes from.
What authentication does not solve: liveness and trust-after-auth
Here is the trap on the far side of a successful PKI migration. Authentication answers "who is this?" It says nothing about whether the answer is useful. An authenticated agent can be entirely wrong: its heartbeat can be two minutes old, its lease can have expired while it was unreachable, it can answer your health check in four milliseconds and then fail to start a single VM because it is out of memory, out of store disk, or has a wedged VMM binary. "Cryptographically verified" and "can do the job" are unrelated properties, and conflating them is how you build a scheduler that confidently places work on a corpse.
Liveness is a separate mechanism, and in our case a boring one: heartbeats and leases. Agents heartbeat roughly every ten seconds with their capacity, and the scheduler excludes any host whose heartbeat is more than thirty seconds stale — not because thirty seconds is magic, but because it is a few missed intervals, which distinguishes a dropped packet from a dead host. Leases run on a short refresh with a longer timeout, so ownership of a sandbox can be concluded dead and reclaimed without two hosts both believing they own it. Authentication gates who may write a heartbeat; the heartbeat's freshness decides whether anyone cares.
Which brings me to the lesson I would most like to hand over, because we paid for it. We cache the fleet's capacity view for thirty seconds so the scheduler is not querying Postgres on every create. At some point that cache started being consulted for lease state as well, and a cached entry whose lease looked expired was treated as authoritative: the host was skipped, and creates that should have succeeded returned a capacity error against a fleet with plenty of capacity. The bug is not the cache. The bug is the direction.
The same asymmetry shows up in the data plane. If an agent authenticates, heartbeats, accepts a create, and then cannot produce a working VM, no credential helps. What helps is that the create path returns a specific refusal the scheduler understands — this host is below its disk floor, this host cannot admit that much memory — so placement moves on instead of retrying into the same wall. Authentication is the door. Admission control is whether there is a room behind it.
Authorisation, briefly: scope every mutation by agent id
One more thing, and it is short because the rule is short: an agent's credential should authorise it to act on its own sandboxes and nothing else. With a shared token this is unenforceable at the credential layer, so it has to be enforced in your queries — and "has to be enforced in your queries" is a sentence that should make you check your queries.
This is a real class of bug, not a hypothetical. We shipped a reconcile loop that read the shared sandboxes table, found rows whose VMs it could not see locally, and concluded they were orphans to be cleaned up. During a rolling deploy, with several agents restarting, that logic deleted sandboxes belonging to peer hosts — including a customer's managed database. Nothing was compromised and no credential was misused. The agent was perfectly authenticated. It was also authorised, by construction, to delete absolutely anything, and a scope bug in one `WHERE` clause was all it took.
The fix is unglamorous and total: every mutation an agent performs is scoped by `agent_id` in the statement itself, and that `agent_id` comes from the authenticated identity rather than from the request. Under a shared token, "the authenticated identity" is the agent's own configured id, which is weaker than you want — the enforcement is self-restraint rather than a boundary. That is the strongest practical argument for per-host credentials, stronger than the usual blast-radius one: you cannot scope what you cannot distinguish. A verified certificate subject makes the scope a property of the connection instead of a property of the agent's good behaviour.
A shared secret answers "are you one of us?". Almost every authorisation question you will eventually want to ask needs the answer to "which one of us are you?", and you cannot retrofit that onto a credential everybody holds.
What to build, in order
- A shared token verified in constant time, with the verifier accepting a primary and a secondary value from day one, failing closed on an empty expected value, and refusing to bind a public listener with no credential configured.
- A counter for secondary-token hits, exported from every verifier. This is the only thing that makes step 4 of a rotation a measurement rather than a hope.
- The token read from a file that the process re-reads on SIGHUP, not only from the environment. On a host that supervises VMs, this is what makes rotation a reload instead of a restart — and a restart can be a data-plane event.
- A pinned secret version in every consumer, and a written answer to "who owns this knob: instance creation, or config management?" for each one. A value set only at creation reverts the next time a host is replaced.
- A rehearsed rotation on a calendar. The drill's purpose is to find the straggler — the cron job, the unmanaged VM, the Slack paste — while nothing is wrong.
- Every agent-side mutation scoped by `agent_id` in the statement, sourced from the authenticated identity. Do this before you need it, because the thing it prevents is deleting someone else's database during a deploy.
- Then, and only when the fleet is big enough or multi-tenant enough to need it, per-host identity: per-host tokens for the cheap version, mTLS or SPIFFE SVIDs for the real one, with clock sync as a precondition rather than an afterthought.
The reason this problem stays unfashionable is that nothing about it is visible when it works. There is no feature, no dashboard anybody looks at, no demo. The two moments it becomes visible are the moment somebody rotates a token at 14:00 on a Tuesday and takes the fleet with it, and the moment a single compromised host turns out to have been holding a credential good for every other host in the estate. Both of those are entirely avoidable by an afternoon's work done early: a verifier that accepts two values, a token read from a file, a counter, and a `WHERE agent_id = $1` that was there from the start.
Frequently asked questions
Is a single shared bearer token between a control plane and its hosts ever acceptable?
Yes, and pretending otherwise leads people to half-build a PKI, which is worse than either option. A shared token is correct against the threat it is most often actually facing — the internet, scanning your ports continuously — and it has no clock dependency, no expiry, no CA and no renewal daemon to fail silently. For a single-team fleet of a handful of hosts it is a reasonable, defensible choice, and we run it. What you must be honest about is what it does not give you. It proves membership, not identity: every host holds a credential that can impersonate every other host and call the control plane's node-facing endpoints, so the blast radius of one compromised host is the whole fleet. There is no revocation short of rotating everywhere. There is no per-caller authorisation, because the credential cannot distinguish one agent from another, which means anything resembling per-agent scoping is reading an identifier out of a header the caller supplied. And your audit trail records a self-declared label rather than a verified subject. The practical test is whether you can tolerate those four properties for the time it would take you to move off them in an incident. If the answer is no, or if the fleet is multi-tenant across customers, you want per-host credentials — and per-host bearer tokens get you most of the benefit for a fraction of the mTLS cost.
How do I rotate a fleet-wide shared secret without taking the fleet down?
Hold one invariant: at every instant, the set of values your verifiers accept must be a superset of the set of values your callers present. Every safe rotation is a sequence of steps that never breaks that containment. Concretely, accept-both-then-narrow: first teach every verifier to accept old OR new and roll that everywhere, including the control plane, since both directions verify; then roll the new value out to the callers; then prove nothing presents the old value any more; only then remove the old value, and verify that a request bearing it now returns 401. A naive flip of the server's expected value is a total outage, and a confusingly-labelled one: heartbeats start failing, every host looks stale within about thirty seconds, the scheduler runs out of candidates, and creates begin failing with a capacity error against a fleet that has plenty of capacity. Three details make or break it. Pin a specific secret version in every consumer rather than referencing `latest`, because `latest` means the rotation takes effect whenever each process happens to restart, in no order, with no way to say which hosts are done. Prove step three with both a counter of old-token hits exported from every verifier and a grep of your config-management repo for stale references — the straggler is usually not an agent but a cron job or a debug script. And never rotate with a blind fleet-wide restart: on a host that supervises VMM processes, a restart that overruns its stop timeout kills running guests, so the mechanism must be a config reload or a drain-then-restart, one host at a time.
What actually breaks when you move to mTLS, and is SPIFFE worth the adoption cost?
What breaks is everything with a clock in it, plus everything that assumed restarts were free. Certificate validity is wall-clock, so clock skew becomes an authentication failure, and short lifetimes — which you want, because they are your revocation story — make skew proportionally worse: ten minutes of drift against a one-hour certificate is a large fraction of its life. Anything snapshot-restored is a special hazard, because a machine restored from a snapshot resumes with the clock frozen at snapshot time; we hit exactly this as TLS failures inside restored guests that looked like network problems, and the fix was a clock-sync step on restore, resume and wake. Then renewal: if renewing a certificate requires restarting the process, you have coupled certificate expiry to a restart, and on a host supervising VMs a restart can kill running guests — so the server must re-read its key pair through a tls.Config callback instead. Also check that you used RequireAndVerifyClientCert rather than the permissive variant, and that your client CA pool contains only your CA and not the system roots. On SPIFFE: it is worth it at the point where you would otherwise write a shell script and a cron entry to issue short-lived per-workload certificates, because that script is a worse SPIFFE. It gives you attested identities as URIs, automatic rotation, and trust-bundle distribution. The ecosystem moves, so verify current behaviour against the project's own documentation rather than any blog post, including this one.
If the agents authenticate correctly, why do I still need heartbeats and leases?
Because authentication answers "who is this?" and says nothing about whether the answer is useful. An authenticated agent can have a two-minute-old view of the world, can hold an expired lease, and can pass a health check in four milliseconds and then fail to start a single VM because it is out of memory, out of store disk, or running a wedged VMM binary. Cryptographic verification and operational fitness are unrelated properties, and conflating them produces a scheduler that confidently places work on a dead host. So liveness is a separate mechanism: agents heartbeat their capacity every ten seconds or so, the scheduler excludes any host whose heartbeat is more than thirty seconds stale — a few missed intervals, which distinguishes a dropped packet from a dead host — and leases run on a short refresh with a longer timeout so ownership can be concluded dead and reclaimed without split-brain. One hard-won rule goes with this: never trust a cache for a negative. We cache the fleet capacity view for thirty seconds, and when a cached entry's lease state was treated as authoritative evidence that a host was dead, healthy hosts were skipped and creates failed with a capacity error. A cached positive that is wrong costs one retry; a cached negative that is wrong removes a working host from service for everyone. Positives may be cached. Negatives must be confirmed against the source of truth before you act on them — and that applies to a cached "token invalid" during a rotation window too.
Keep reading
- Leases vs heartbeats for node liveness — The liveness half of this post, in full: why a fresh heartbeat and a valid lease answer different questions.
- Workload identity for ephemeral microVMs — The guest-side half — how a VM that lives ninety seconds gets an identity at all, which is harder than the host-side problem.
- Rotating database credentials without downtime — The same accept-both-then-narrow pattern applied to a credential with a connection pool attached.
- Secrets inside Firecracker snapshots — What happens to a credential that was in guest memory when the snapshot was taken — and to the clock it was validating against.
- How a sandbox scheduler places workloads — What a forged capacity heartbeat would actually poison: the scoring function that decides where every sandbox lands.
Related posts
- Scheduling sandbox bursts: why 50 creates land on one host
Fifty creates arrive in 200ms. Every one of them reads the same cached capacity snapshot, computes the same winner, and lands on the same host. The scheduler did exactly what it was told and produced exactly the outcome it exists to prevent.
- How Many Sandboxes Fit on One Machine?
Everyone wants a number. The real answer depends on which resource binds first — and on most microVM fleets, the host memory metric you'd naturally scale on is actively lying to you.
- Conntrack Table Limits: The Shared Resource Under Your MicroVM Fleet
Your guests have separate kernels, separate namespaces and separate subnets. They share one fixed-size hash table, and one port scanner can fill it for everybody.
- Saying No Correctly: Admission Control for Sandbox Fleets
Almost no capacity bug is 'we ran out of RAM.' It is either a yes you could not fit, or a no you did not have to give — and the second one is more expensive, because it looks like your platform is broken.
- Leader Election When Your Compute Can Be Cloned
A restarted process comes up empty and has to ask who the leader is. A restored one never asks — it resumes mid-sentence still believing it holds the lock. Leases, fencing tokens, and why a forked sandbox must revoke leadership before it does anything else.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.