all posts

Encryption at Rest for Sandbox Disks and Snapshots

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and there is one security-questionnaire row I have come to dread more than the others. Not because the answer is bad. Because the answer is a single checkbox covering four different files with four different stories, and ticking it truthfully still leaves the reviewer with a picture of our system that is wrong.

The row is: is customer data encrypted at rest?

Yes. And that sentence is doing far less work than it appears to. Encryption at rest is a control with a specific, narrow threat model, and the honest version of this answer requires naming the threat first, then walking the four places sandbox state actually lands — one of which is a byte-for-byte copy of a guest's RAM, and none of which PandaStack encrypts itself. There is no dm-crypt, no LUKS, no fscrypt, no cryptsetup anywhere in the platform. I will show you the grep.

This post is the version I would want to read as the engineer on the other side of the questionnaire. It costs us something to write, which is roughly the point.

The short version: full-disk encryption defends against physical media leaving your control and almost nothing else. For sandbox snapshots, retention and access control are better security than cipher choice, because the realistic attacker has credentials rather than a screwdriver. PandaStack relies on provider-level encryption for disks and object storage, encrypts app secret env vars itself with AES-256-GCM, offers no customer-managed keys, and does not encrypt sandbox rootfs images or memory snapshots at the block layer. Here is why.

Encryption at rest is a control in search of a threat

Start with what the mechanism actually does, because every subsequent argument depends on it. Full-disk encryption protects data when the storage medium is separated from the running system and its keys. The canonical scenarios are a decommissioned drive that goes into a skip instead of a shredder, a server stolen out of a colocation rack, a cloud provider recycling physical media between customers, and a disk shipped back to a vendor under warranty.

Those are real threats. Cloud providers encrypt everything at rest by default largely because the drive-reuse scenario is a structural property of their business, and "we do not inventory which customer's bytes were on which platter" is not a sentence anybody wants in an audit.

Now the part the checkbox hides. Disk encryption does essentially nothing against a compromised host process. The process opens the file, the kernel reads the block, the dm-crypt layer decrypts it with a key that is resident in kernel memory on that same machine, and the plaintext arrives in the process's buffer. That is not a weakness in the implementation; it is the implementation working. A running system has to be able to read its own disks, so the key has to be present, so anything that can run code as the right user on that host reads plaintext.

There are three threats people conflate into one row, and separating them is the whole post:

  1. The media leaves your control. A drive, a backup tape, a recycled disk. Encryption at rest is the correct and sufficient answer, and the provider has already given it to you.
  2. Someone with credentials reads the file. An over-broad IAM role, a leaked service-account key, a bucket made public, a support engineer with more access than the diagram says. Encryption at rest does nothing: the credential grants decrypted reads. The controls that matter are authorization, audit logging, and how long the object exists at all.
  3. The host is compromised. Encryption at rest does nothing, because the key is on the host by necessity. The controls that matter are isolation, blast radius, and what was in reach of the compromised process.

When a reviewer asks about encryption at rest, they are usually worried about threats two and three and asking about threat one. That mismatch is not their fault — the questionnaire was written for a CRM with a database, where "encrypted at rest" covers the one place data lives. A sandbox platform has four.

The four places sandbox state lands

Each one has a different file format, a different lifetime, a different reader, and a different honest answer.

1. The rootfs and its copy-on-write clone, on the host's store disk

A template is an ext4 image. Creating a sandbox reflinks that image into a per-sandbox clone — on XFS that is an O(metadata) operation, a few milliseconds, with the data blocks shared until something writes. Everything the guest writes to its filesystem lands in that clone, on a local XFS store on the host's attached disk.

On GCP that disk is a Persistent Disk, and Persistent Disk is encrypted at rest by default with Google-managed keys. That is real encryption, performed competently by people who do it for a living, and it is also the weakest kind of answer available: the party holding the key is the party holding the disk. It covers physical media and the provider's own hardware lifecycle. It does not and cannot cover anything else, because by construction the decryption happens transparently for anyone who can open the file on that machine.

Worth knowing about the end of a sandbox's life: teardown is an unlink. The agent removes the VM directory with an ordinary recursive delete, so blocks return to the filesystem's free pool and get reused by the next sandbox — not overwritten, and with no per-sandbox key we could throw away to make the old bytes unreadable. What you get is allocation behaviour plus the provider's disk encryption: a genuine guarantee against a drive-in-a-skip, not a cryptographic erase. Hold that thought — it returns as the one real argument for per-sandbox encryption.

2. Snapshots and vm.mem in object storage — the uncomfortable one

A Firecracker snapshot is three artifacts: `vm.state` (a few hundred kilobytes of vCPU and device state), the disk, and `vm.mem` — a verbatim dump of guest RAM. For a `base` sandbox that is a 4 GiB file whose contents are exactly what was in the guest's physical memory at the instant of capture.

Not a filesystem. Not a structured export with a schema you can reason about. The unstructured contents of someone's address space. Environment variables as the kernel stored them. A decrypted secret the application pulled from a vault at startup. A database connection string in a driver's connection pool. An SSH private key the agent loaded. TLS session keys. The contents of a buffer the application was about to zero and had not got to yet. The page cache, which means fragments of files the guest read even if it never kept them.

This is not a bug and it is not patchable. It is what snapshot-restore is. The reason a PandaStack create lands at p50 179 ms with no warm pool is that every create restores a baked snapshot rather than booting a kernel — the snapshot is the product. A platform that boots fast almost certainly has files like this, whether or not its security page mentions them.

Snapshots, seeds and streamed memory objects live in GCS, which encrypts objects at rest by default with Google-managed keys. Same story as the disk, same narrow coverage. So for this file, the controls that actually matter are the other two:

Who can read the bucket. The bucket has uniform bucket-level access enabled, which removes per-object ACLs from the picture entirely and makes IAM the single answer to "who can read this." The agent service account holds object admin on it. That is the control doing real work here, and it is an authorization control, not a cryptographic one.

How long the object exists. Lifecycle rules are policy as code and I would rather show them than describe them: objects under the `snapshots/` prefix move to Nearline at 30 days and are deleted at 90; per-app memory seeds under `app-seeds/` are deleted at 180 days as an abandonment backstop; and because the bucket is versioned, a separate rule deletes noncurrent versions 7 days after they stop being current, scoped by prefix. That last rule exists because on a versioned bucket every delete is soft — without it, every overwritten seed and every flipped pointer was retained forever, fully billed and fully readable.

The managed-database prefix is deliberately excluded from that noncurrent-version reaper, so database base backups and WAL keep an indefinite undo window. A correct operational trade — past retention bugs took longer than seven days to notice — and also an answer to "what is your deletion SLA" that is longer than most people expect. Retention policy is where the real security of an at-rest artifact lives, which is why I would rather tell you the rule than the cipher.

And now the detail I would rather you heard from me than found yourself. The idle reaper that deletes an unused sandbox calls a deliberately different delete path from the one your explicit DELETE calls, and the reaper's path does not cascade to the sandbox's snapshots. The code comment says why in a line: they are durable and outlive the sandbox. That is the correct product behaviour — a snapshot is the thing you explicitly saved so you could restore or fork it later, and an automatic cleanup destroying it would be a worse bug than the disk it reclaims. It is also the lifecycle fact most likely to surprise you.

Read it as a lifecycle. A sandbox goes quiet, the reaper notices it is idle past its TTL, the VM is torn down, and nothing in your dashboard suggests anything is outstanding. The verbatim copy of that guest's RAM is still in the bucket with its database row intact, because it was not reaped alongside the machine that made it.

It is bounded, though, and "durable" is not "forever." A second sweep collects orphaned snapshots — ones whose source sandbox row is gone — purging all three layers in a deliberate order: local directory, then the object-storage blobs, then the database row last, so an interrupted purge reappears as an orphan and gets finished instead of leaving bytes with no row pointing at them. It runs every 15 minutes, bounded per tick, and an orphan becomes eligible once the snapshot itself is older than a grace period that defaults to 7 days. So the window is days, measured from when you took the snapshot — not the 90-day bucket rule, which is an outer backstop that in practice never fires for these.

Which is far better than the no-lifecycle-at-all it replaced. But the property worth internalising survives: the VM's death and the secret-bearing file's death are decoupled, and the gap is about a week by default — a window you cannot see in your sandbox list. Two levers if that is too wide: `PANDASTACK_SNAPSHOT_TTL_DAYS=0` expires orphans on the next sweep, a real option for a security-sensitive deployment, and `PANDASTACK_SNAPSHOT_GC_INTERVAL=0` disables the sweep entirely and restores indefinite retention — the setting to check nobody has quietly applied. The cleanest answer needs no knob: delete the sandbox explicitly rather than letting it be reaped, because an explicit delete cascades immediately.

Say it plainly: for a memory snapshot, retention and access control are better security than cipher choice. A 4 GiB AES-encrypted copy of your RAM that twelve service accounts can read and that nothing ever deletes is worse than an unencrypted one that two accounts can read and that expires in a week. The realistic attacker has credentials, not a screwdriver.

One more property of these artifacts, because integrity and confidentiality get conflated too: seeds are published per generation with SHA256 manifests, and the agent verifies them. That protects against corruption and against a wrong-object swap. It is not encryption, and a hash will not stop anyone reading the file.

3. The host-local chunk cache nobody puts on the diagram

There is a fourth location, and I have never seen it on anyone's architecture diagram including, for a while, mine.

When memory is streamed on demand rather than downloaded whole — the userfaultfd path, where guest page faults are served from object storage over range requests — the chunks that get faulted in are written to a persistent cache on the host's local disk. The first restore of a seed generation on a host pays the round trips; every later one is served from local NVMe. That cache is the per-host warm state in a system with no warm pool.

It is also, viewed from a data-at-rest angle, a durable local copy of pieces of a guest's memory image. It is content-addressed by a hash of the bucket and object path, so a re-bake produces a fresh directory rather than a stale overlay. It outlives the VM. And it is evicted on a size budget by least-recently-used order, which means it goes away when the host needs the space — not when you call delete.

There is nothing sinister here — it is a cache, caches are how fast systems are fast, and it inherits the host disk's provider encryption like everything else. I raise it because a reviewer who asks where their data lives deserves four answers and usually gets two.

4. Durable volumes for managed databases

A managed Postgres is a dedicated Firecracker VM plus a durable volume on a separate disk from the ephemeral store, marked persistent so the idle reaper skips it and pinned to its host. Base backups and WAL go to object storage under their own prefix.

Same provider-level story: the volume disk and the backup bucket are encrypted at rest with provider-managed keys, which covers media. Two database-specific additions. In transit, connections are TLS-required and SNI-routed — a different control people fold into the same row. And at rest inside the database, community PostgreSQL 16 has no transparent data encryption; there is no cluster-level knob to turn on. Your options are column-level encryption in the application or with `pgcrypto`, where you hold the key and accept that encrypted columns cannot be indexed usefully, plus whatever the block layer underneath gives you. If a compliance regime means something stronger than "the volume is on an encrypted disk," have that conversation before you build on any managed Postgres, ours included.

The four locations, side by side

Where sandbox state lands, what encrypts it today, and which threats that does and does not address.
Where state landsWhat encrypts it todayThreat that coversThreat it does notWhat you can do
Rootfs + CoW clone (host store disk)Provider disk encryption, provider-managed keys. Nothing from us.Media leaving the provider's control; hardware lifecycle.Anyone who can open the file on that host. Deletion is an unlink, not a shred.Encrypt in the guest before you write, with your key. Keep nothing durable you did not have to.
Snapshots, seeds, vm.mem (object storage)Bucket default encryption, provider-managed keys. SHA256 manifests for integrity only.Media; corruption; wrong-object swaps.A credentialed reader. The file is guest RAM verbatim, cipher or not.Do not hold secrets in memory at capture time. Use short-lived credentials. Delete the sandbox explicitly — an idle reap leaves its snapshots for days.
Host-local streamed-memory chunk cacheHost disk's provider encryption. Content-addressed per seed generation.Media.Host-level access. Eviction is a size budget, not your delete call.Treat snapshot creation as the sensitive act; the cache only holds what a snapshot already held.
Durable volume + database backupsProvider disk and bucket encryption. TLS required in transit.Media; network interception.A credentialed reader. Community Postgres has no transparent data encryption.Column-level encryption with pgcrypto or in-app for the fields that need it.
App secret env vars (control-plane Postgres)AES-256-GCM by us, per-value nonce, owner and key name as additional authenticated data.Media; database read; row-swap between apps; tampering.The plaintext after injection — it is in guest memory, so it is in a snapshot.Use the secret store rather than baking values into a template or a repo.

What PandaStack actually does, stated plainly

No dm-crypt. No LUKS. No fscrypt. No cryptsetup. The grep across the agent, the cloud-init provisioning, the Terraform and the deploy scripts returns zero hits, and I re-ran it the day I wrote this. Sandbox rootfs images, copy-on-write clones, Firecracker snapshots, `vm.mem`, streamed memory chunks and durable volumes are not encrypted by PandaStack.

What exists is provider-level. On GCP, Persistent Disk and Cloud Storage encrypt at rest by default with Google-managed keys; we configure nothing, which is both the honest description and the limitation. In the AWS Terraform module, the artifact bucket sets a server-side encryption configuration with `apply_server_side_encryption_by_default` and `sse_algorithm = "AES256"` — that is S3-managed keys. No `kms_master_key_id` is set anywhere in the configuration, which means no KMS envelope and no customer-managed key. We do not offer CMEK or BYOK today. If your compliance regime requires a key you hold and can revoke, we are not currently the right answer, and I would rather you learn that here than in week six of a procurement.

There is one layer that is genuinely ours, and it is the one I would point a reviewer at first.

The layer we do own: app secret env vars

App environment variables split by sensitivity. Plain variables live in a JSONB column on the app row. Secret variables live only in a separate table, AES-256-GCM encrypted, under a 32-byte master key generated as random bytes by Terraform into Secret Manager and read by the API from an environment variable. The key is never in the repo and never in an image. Stored values carry a version prefix followed by base64 of the nonce concatenated with the sealed ciphertext, so the format can grow a KMS envelope later without ambiguity about what any given row is. Every value gets a fresh 96-bit nonce.

The detail I am actually proud of is the additional authenticated data, because it is the part most platforms skip. AES-GCM lets you bind a ciphertext to context that is not itself encrypted but is covered by the authentication tag. We bind each value to its owner tag and its variable name — the owner being the app id, or a workspace-plus-repository pair for preview environments.

Why that matters is worth spelling out, because "the column is encrypted" sounds like it already covers it. Imagine an attacker who has write access to the control-plane database but not the master key. Without AAD, encryption alone does not stop them copying the ciphertext of someone else's `STRIPE_SECRET_KEY` row into an app they control, deploying it, and reading the plaintext out of their own application. They never broke the cipher; they used your own decrypt oracle as intended, with the row pointed somewhere new. That is a confused-deputy attack, and AAD closes it: a sealed value moved to a different app, or renamed to a different variable, simply fails to open. The authentication tag covers the context, so the context is part of the ciphertext's identity.

The rest of the handling follows from treating decryption as a privilege rather than a utility. One function in the codebase decrypts, and it runs in the deploy path immediately before injection; nothing else calls the decrypt primitive. It returns the plaintext values alongside the env map so the log and comment redactor can scrub them from build output. Errors name the variable key and never the value. The API masks secrets on read with a fixed redaction that does not even reveal length. A production deploy never queries the preview-environment table, so a pull request's throwaway env cannot reach production, and secret lookup is keyed on the app's own id, so a preview app cannot read its parent's secrets. The unit tests assert the properties that matter rather than just the round trip: a tampered ciphertext fails authentication, a ciphertext opened under the wrong owner fails, and the wrong key fails.

And now the ceiling, which is the reason this post exists. Once that value is injected, it is an environment variable in a process inside a guest. It is in guest RAM. If that guest is snapshotted, hibernated or used to bake a seed, the plaintext is in `vm.mem`. Our best at-rest layer and our fastest boot path meet at exactly this point, and the boot path wins, because it is physics rather than policy.

What we did about that, specifically

Since app hosting hibernates idle apps to object storage and wakes them from a per-app memory-and-disk seed, this is not hypothetical for us; it is the hot path. So the seed bake scrubs what it can and documents what it cannot.

Before any seed snapshot, the app's env file is removed from the guest disk. The stored rootfs inside a published seed therefore carries no plaintext secrets, and the deploy re-delivers them on the restore handoff. In the cold variant the app process is stopped before capture, so the memory image holds no in-flight request, no live socket and no plaintext secrets either.

The warm variant is the fast one: the app is left running through the snapshot so a restore resumes an already-listening server instead of paying a JavaScript dev server's ten-to-twenty-second startup. The disk scrub still happens. The secrets, however, remain in the frozen process memory, because that is what freezing a running process means — the same trade AWS SnapStart makes. So a warm seed is treated as sensitive: it stays private to its owner workspace, same as the cold seed, and the age-based reaper deletes it.

That level of detail is the shape of every honest answer here. Not "we encrypt it," but: what we removed, what is structurally still there, how long it lives, and who can read it.

For the complementary read on what else survives a capture — entropy state, TLS session keys, and what a thousand restores of one memory image share — see The Security Gotchas of Firecracker Snapshots (Secrets Frozen in RAM). This post is about the bytes on the disk and who can read them; that one is about the secrets inside the image being reused.

The one-line demonstration, and the shape of the fix

# On a host you own, over a snapshot you took. This is the whole argument.
$ ls -l /var/lib/pandastack/snapshots/$SNAP/
-rw------- 1 root root 4294967296 vm.mem      # guest RAM, verbatim
-rw------- 1 root root     151552 vm.state    # vCPU + device state
-rw------- 1 root root 2147483648 clone.ext4  # the disk

$ strings -n 16 vm.mem |
    grep -aiE 'AWS_SECRET|BEGIN [A-Z ]*PRIVATE KEY|postgres://|xoxb-|Bearer '
AWS_SECRET_ACCESS_KEY=...
postgres://app:...@db.internal:5432/prod
-----BEGIN OPENSSH PRIVATE KEY-----
Authorization: Bearer eyJhbGciOi...

# No cipher fixes that file, because whoever restores it must be able to read it.
# The mitigation is the credential not being there, or not being worth having:
#   1. fetch it AFTER restore, never before capture
#   2. scope it to one job and give it minutes, not months
#   3. scrub the on-disk copy before the snapshot (we do; see the seed bake)
#   4. set a retention rule on the bucket you can state in one sentence
#   5. treat "read a snapshot object" as a privileged action with an audit trail

Point two changes the risk class rather than reducing it. An encrypted snapshot holding a credential that never expires is a time bomb with a lock on it. An unencrypted one holding a credential that expired five minutes after capture is an artifact nobody can do much with. Short-lived credentials convert a confidentiality problem into a timing problem, and timing problems you can win.

The pattern that survives a snapshot

from pandastack import Sandbox
from pandastack.exceptions import CommandFailed

# Mint the credential yourself, scoped to this job, with a short life.
# STS AssumeRole, a Vault lease, an OIDC exchange - whatever you already run.
token = mint_scoped_token(ttl_seconds=300)  # your broker, not ours

# ttl_seconds is an IDLE ttl, not a wall clock: the host reaper deletes the
# sandbox once it has been inactive that long, so it will not bound a sandbox
# you keep touching. The context manager kills it on exit (unless
# persistent=True), so make teardown explicit and do not lean on the TTL.
with Sandbox.create(template="code-interpreter", ttl_seconds=600) as sbx:
    # Keep the credential in ONE process's environment. Not in a file (that
    # lands in the CoW clone), not baked into a template (that lands in every
    # sandbox created from it), not in sandbox metadata.
    # The in-guest `timeout` is the real time bound; exec's own timeout_seconds
    # is capped by the client's HTTP timeout, so keep it at 30 or below and use
    # exec_stream for anything genuinely long.
    r = sbx.exec(
        f"JOB_TOKEN={token} timeout --kill-after=5s 25 python3 /work/pull.py",
        timeout_seconds=30,
    )
    if r.exit_code != 0:
        raise CommandFailed(f"pull failed: {r.stderr[-2000:]}")

    # Do NOT snapshot, hibernate or fork_tree while a live credential is
    # resident - fork_tree is the one that inherits the parent's memory, so
    # every child would get the same token bytes.
    # If you must capture state, finish the credentialed step first:
    snap_id = sbx.snapshot()

# Explicit teardown if you are not using the context manager: sbx.kill().
# There is no sbx.delete().
print(snap_id)

The ordering in that snippet is the whole discipline. Credential in, work done, credential gone, then capture. Most of the snapshot-leaks-a-secret incidents I have seen are not a failure to encrypt; they are a capture that happened while something was hot.

Why per-sandbox dm-crypt is not a free win

The obvious question is why we do not add a block-layer crypto target per sandbox and tick the row with something we own. The answer has three parts, of which only the third is decisive.

First, the create path is fast because of the reflink. A PandaStack create is a snapshot restore every time, with no warm pool of idle VMs, and the disk step is a reflink of the template image — O(metadata), a few milliseconds, data shared until written. A dm-crypt layer sits between the filesystem and the block device, which changes what reflink can do for you: the sharing trick is a filesystem-level property of extents in one XFS filesystem, and a per-sandbox mapper over a per-sandbox encrypted image is not the same shape at all. You would be trading the specific mechanism that makes a 179 ms create possible for a control whose coverage I described at the top of this post.

Second, this lands on the critical path of every single create, not on a background job. The restore path requires the rootfs to be a local file because copy-on-write needs a local block device — userfaultfd streams guest memory from object storage on demand, but it never streams the disk, for exactly that reason. So any crypto layer goes in the hot path of the operation the product is named for, and gets to charge its setup cost on every sandbox you ever start.

Third, and this is the one that settles it: the host holds the key. It must, in order to boot the guest. There is nobody else there at create time. So host-level disk encryption cannot defend against the host being compromised, which is the threat people actually mean when they ask about this. You would add latency and a new failure mode to the create path in exchange for a control that covers the drive-in-a-skip scenario the cloud provider already covers, with keys managed by the same provider either way. That is not a good trade, and I would rather tell you we did not make it than imply we bought you something we did not.

In fairness to the idea, here is what it would genuinely buy, and it is not nothing. A per-sandbox key you destroy at teardown gives you cryptographic erasure: the freed blocks become unreadable immediately rather than merely unallocated, and "deleted" stops depending on filesystem behaviour. That is a real property, it is the strongest argument in the area, and it is why this sits in the "not yet, and here is the cost" column rather than the "never" one.

The version that actually addresses the threat people are worried about is guest-side encryption where the tenant holds the key. Your application encrypts before it writes, the key arrives over the network at runtime and lives only in guest memory, and the platform never sees plaintext at rest. That is a genuinely stronger position than anything the host can offer you, and it costs exactly what you would expect: key management becomes yours, including rotation and the moment you lose a key and the data is gone. And it has one hard edge that follows from everything above — anything in guest RAM is still in the snapshot, so guest-side encryption protects your disk, not your memory image.

What to actually do

  • Keep secrets out of guest memory at capture time. Fetch after restore, not before the snapshot. Know which of your operations are capture points: taking a snapshot, hibernating, baking a template, and `fork_tree` (which inherits the parent's memory — plain `fork` clones the disk and cold-boots, so it does not).
  • Prefer short-lived credentials so a leaked image leaks something already expired. This is the single highest-leverage change in the list, and it is entirely on your side of the boundary.
  • Set retention on anything that holds a memory image, and be able to state the rule in one sentence. If you cannot say how long a snapshot lives, you do not know what your exposure is. Ask your vendor the same question; the answer is more informative than their cipher.
  • Do not assume a TTL cleans up your snapshots. `ttl_seconds` is an idle TTL — it reaps a sandbox that has gone quiet, not one you keep touching — and that reap leaves the sandbox's snapshots for an orphan sweep to collect days later. An explicit delete cascades immediately, so delete on purpose.
  • Treat reading a snapshot object as a privileged action with an audit trail. For a memory image this is the control that maps to the realistic attacker. Enforce it with IAM rather than object ACLs, and alert on reads by anything that is not the restore path.
  • Use the platform's encrypted secret store rather than baking values into a template. A baked template's rootfs is shared by every sandbox created from it, so a secret baked in at build time is on every clone, forever, for everyone who can start one.
  • Encrypt in the guest, with your key, for anything that must be confidential at rest. Accept the key management and the memory-image caveat above.
  • If your compliance regime demands customer-managed keys, ask for CMEK explicitly rather than accepting "encrypted at rest" as an answer — from us or from anybody. We do not offer it today.

The honest limits

  • No CMEK and no BYOK today. The secret-env format carries a version prefix specifically so a KMS envelope can be added without ambiguity, but intent is not a feature. If you need a key you hold and can revoke, we are not the answer yet.
  • No per-sandbox block encryption, so no cryptographic erasure. Deleting a sandbox unlinks its directory; the blocks are freed and reused, not overwritten, and there is no key to throw away. The guarantee you get on deletion is the provider's, not ours.
  • Provider-managed keys mean the provider is inside your trust boundary. GCP holds the keys for Persistent Disk and Cloud Storage; the AWS artifact bucket uses S3-managed keys. If your threat model includes the cloud provider, encryption at rest as configured here does not address it, and saying otherwise would be a lie with a diagram.
  • Snapshot files are guest RAM verbatim, and that is inherent to how snapshot-restore works rather than a bug we can patch. We scrub the on-disk env before a seed bake, and a warm seed still freezes secrets in process memory, because that is what freezing a running process is. The mitigations are retention, scoping and short credentials — not a cipher.
  • A sandbox's death does not take its snapshots with it. The idle reaper deliberately does not cascade, so the artifact holding guest RAM verbatim outlasts the machine until the orphan sweep collects it — by default a 7-day grace from when the snapshot was taken, swept every 15 minutes. Days rather than forever, and still a window you cannot see in your sandbox list. Audit snapshots, not VMs.
  • There is a host-local chunk cache of streamed memory that is evicted on a size budget, not on your delete call. It inherits the host disk's encryption and holds nothing a snapshot did not already hold, but "gone everywhere" and "deleted" are not the same timestamp.
  • Community PostgreSQL has no transparent data encryption, so a managed database's at-rest story is the volume's disk encryption plus whatever you do at the column level. True of every managed Postgres built on community Postgres, which is most of them. Worth checking rather than assuming.
  • "Encrypted at rest" on anyone's security page, ours included, is a weaker statement than most questionnaires assume it to be. If you are the reviewer, the follow-up questions are: whose keys, which artifacts, what retention, and who can read them. If you are the vendor, volunteer those four before you are asked.

The summary

Encryption at rest defends against physical media leaving your control. It does not defend against a credentialed reader, because the credential grants decrypted reads, and it does not defend against a compromised host, because the key has to be on the host for the system to work at all. Every security questionnaire flattens those three threats into one checkbox, and the checkbox answers the one you are least likely to be worried about.

Sandbox state lands in four places: the rootfs and its copy-on-write clone on a provider-encrypted disk; snapshots, seeds and `vm.mem` in a provider-encrypted bucket, where `vm.mem` is a verbatim copy of guest RAM — environment variables, decrypted secrets, key material, page cache; a host-local cache holding streamed pieces of that image on a size budget; and durable database volumes and backups, under a Postgres with no transparent data encryption of its own.

PandaStack adds no block-layer encryption to any of it — no dm-crypt, LUKS, fscrypt or cryptsetup anywhere — and relies on provider defaults with provider-managed keys. We offer no customer-managed keys. The one layer that is ours is app secret env vars, sealed with AES-256-GCM under a master key from Secret Manager, bound by additional authenticated data to the app and variable that own them so a ciphertext cannot be moved or renamed, decrypted at exactly one place in the codebase, and redacted from logs. I think that is a good piece of engineering, and it stops mattering the instant the plaintext reaches guest memory and someone takes a snapshot.

Which is the real lesson. For a platform built on snapshot-restore, data-at-rest security is not primarily a cryptography problem. It is a retention problem, an authorization problem, and a discipline problem about what is resident in memory when the camera goes off. Ask your vendor those questions. If the answer is a cipher name and nothing else, you have learned something anyway.

Frequently asked questions

Is customer data encrypted at rest on PandaStack?

Yes, at the provider level, and I want to be precise about what that covers. Sandbox disks and durable volumes sit on Google Persistent Disk, which is encrypted at rest by default with Google-managed keys. Snapshots, template seeds and memory images sit in Cloud Storage, which is encrypted at rest by default with Google-managed keys. On the AWS path the artifact bucket sets server-side encryption by default with the AES256 algorithm, which is S3-managed keys. PandaStack adds no encryption of its own at the block or filesystem layer: there is no dm-crypt, LUKS, fscrypt or cryptsetup anywhere in the agent, the provisioning, the Terraform or the deploy scripts, and a grep for those terms returns zero hits. One thing is encrypted by us at the application layer — app secret environment variables, with AES-256-GCM under a 32-byte master key held in Secret Manager, each value bound by additional authenticated data to the app and variable name that own it. What all of that covers is physical media leaving the provider's control. What it does not cover is a credentialed reader or a compromised host, because in both of those cases decryption happens transparently for whoever is asking.

Why is a Firecracker memory snapshot more sensitive than a disk image?

Because you can reason about a disk image and you cannot reason about an address space. A disk has a filesystem, so you know where files are, you can decide what not to write, and you can delete something and know roughly what you did. A memory snapshot is a verbatim dump of guest RAM at an instant — for a 4 GiB sandbox, a 4 GiB file whose contents are whatever the guest's physical memory held. That includes environment variables as the kernel stored them, secrets a process decrypted into a heap buffer, database connection strings sitting in a driver's pool, private keys an agent loaded, TLS session keys, buffers an application was about to zero, and page-cache copies of files the guest read and discarded. Nothing in it was placed there by a decision you made about persistence. The uncomfortable part is that this is not a defect to be patched — it is what snapshot-restore is, and it is the mechanism that makes a create a sub-200-millisecond operation instead of a multi-second boot. So the mitigations are all operational rather than cryptographic: do not have the secret resident when the snapshot is taken, scrub what you can off the disk before capture, make credentials short-lived so a leaked image is stale, set a retention rule, and treat reading one of those objects as a privileged action.

Do you support customer-managed encryption keys (CMEK or BYOK)?

No, not today, and I would rather say that plainly than let an "encrypted at rest" answer imply a key hierarchy that does not exist. There is no KMS key wired into the storage configuration — the AWS bucket uses the default AES256 server-side encryption with no kms_master_key_id set, and the GCP buckets and disks use Google's default encryption with Google-managed keys, which we do not configure at all. The one piece of key management we do run is the master key for app secret environment variables: 32 random bytes generated by Terraform into Secret Manager, never in the repository or an image, with a short stable key id derived from its hash so a future rotation sweep can tell which key sealed which row. The ciphertext format carries a version prefix specifically so a KMS envelope can be added without ambiguity about what existing rows are. That is the groundwork for CMEK, not CMEK. If your compliance regime requires a key you hold and can revoke — a hard requirement in some regulated environments, and a reasonable one — ask explicitly, and treat any vendor's "encrypted at rest" as an answer to a different question until they confirm whose keys.

Does encryption at rest protect me if the sandbox host is compromised?

No, and this is the single most important thing to understand about the control. Disk encryption works by holding a key in the running system so that reads are decrypted transparently and writes are encrypted transparently. The system has to be able to read its own disks to function, so the key must be present on the machine, which means any code running with sufficient privilege on that host reads plaintext — it asks the kernel for the file, and the kernel obliges. Adding a per-sandbox crypto layer does not change this, because the host must hold that key too in order to boot the guest from the encrypted image; there is nobody else present at create time to hold it. So the defence against host compromise is not encryption, it is isolation and blast radius: a guest kernel per sandbox rather than a shared one, a small virtual device surface, the VMM boxed by the jailer and constrained by seccomp, and a short-lived sandbox with nothing valuable in reach. If you want data that a compromised host genuinely cannot read, the only shape that works is encrypting it in the guest with a key the platform never sees — and even then, that key is in guest memory while it is in use, so it is in any snapshot taken during that window.

When a sandbox is deleted or reaped, what is actually gone?

The agent removes the sandbox's directory with an ordinary recursive delete, so the copy-on-write clone and the VM's files are unlinked and their blocks return to the filesystem's free pool, to be reused by the next sandbox that needs space. They are not overwritten, and there is no per-sandbox encryption key we could destroy to make the old bytes cryptographically unreadable. In practice the protections you get are that the blocks are only reachable by something with host-level filesystem access, that the underlying disk is encrypted at rest by the provider, and that the space gets reused fairly quickly on a busy host. What you do not get is a cryptographic erase, and that is the strongest honest argument for per-sandbox block encryption — a key destroyed at teardown makes deletion immediate and provable rather than dependent on allocator behaviour. Two things survive, and the first matters more than the disk question. An explicit delete of a sandbox cascades to its snapshots immediately — local bytes, object-storage blobs and the database row. An automatic idle reap deliberately does not, because a snapshot is meant to outlive its sandbox; that is the point of having saved it. A separate sweep then collects orphaned snapshots whose source sandbox is gone, running every 15 minutes and purging those older than a grace period that defaults to 7 days. So a sandbox that quietly aged out leaves a byte-for-byte copy of its RAM behind for days, not forever — and if you want that shorter, the grace is configurable down to zero. The practical advice is to delete sandboxes explicitly rather than letting a TTL reap them, and to audit your snapshot list rather than your sandbox list. The second survivor is smaller: if a memory image was streamed rather than downloaded whole, chunks of it may remain in a host-local cache that is evicted on a size budget rather than on your delete call.

Keep reading

Related posts

  • Running Code Over PHI Without Expanding Your HIPAA Blast Radius

    A snapshot of guest memory is a full-fidelity copy of whatever PHI the process was holding, sitting in a bucket. Here is where protected health information really goes when you run code over it, and the architecture that keeps its residency measured in seconds.

  • Guest page cache: why it bloats microVM snapshots

    A microVM snapshot captures the guest's RAM verbatim — including megabytes of file data Linux cached from a disk that's sitting right there. Drop the cache before you bake, and the memory image shrinks to the live working set.

  • Should You Compress Firecracker Memory Snapshots?

    Compression looks like free money when your snapshot bucket is measured in terabytes. Then you discover that a compressed stream has no byte N — and your lazy 49ms restore turns into reading four gigabytes you were never going to touch.

  • How Firecracker Memory Snapshots Actually Work

    A Firecracker snapshot is the guest's RAM, the VMM device state, and the rootfs. Restore maps the memory copy-on-write so the kernel pages it in lazily — which is exactly why you don't pay for the whole RAM image up front.

  • Firecracker virtio-rng and Guest Entropy Explained

    A freshly booted microVM barely knows any randomness, and a restored one thinks it already does — which is worse. Here's how Firecracker's virtio-rng device seeds the guest, what the entropy rate limiter is for, and why restoring the same memory snapshot into many guests is a genuine cryptographic footgun.

More in Security & isolation · See PandaStack security

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.