Extracting Untrusted Archives: Zip Bombs, Zip Slip, Symlink Escape
The disk alert goes off at 03:12. By 03:20 you have found the culprit, and it is a 42-kilobyte file that a user uploaded with the filename holiday-photos.zip. It is not holiday photos. It is one of the oldest jokes in computing, and your extraction worker has been diligently, obediently writing it to a shared volume for eight minutes, and three other services that happened to share that volume are now also down.
Here is the thing that makes archive extraction different from almost every other untrusted-input problem you have: the attacker chooses the amplification factor. Not you. Parsing a hostile JPEG is dangerous because the parser might have a bug. Extracting a hostile zip is dangerous because the format works exactly as documented. There is no bug. A small file that expands to an enormous one is a feature of compression, and the ratio is a number in the attacker's gift.
Which means every instinct you have about input validation is the wrong instinct here. You cannot inspect your way out of this. Nothing you read out of the archive before extracting tells you the truth, because everything you can read was written by the person attacking you. The only defence that works in the general case is a boundary — a hard ceiling on bytes, inodes, memory and wall-clock time that the extractor cannot talk its way past. And the honest version of a hard ceiling is a machine you are prepared to throw away.
The attacker picks the amplification factor
Start with the arithmetic, because the arithmetic explains the shape of every real attack. A single DEFLATE stream — the compression inside an ordinary zip — has a hard theoretical ceiling of roughly 1032 to 1. So one member of one zip, compressed with the default algorithm, cannot turn 42 KB into a petabyte. It tops out somewhere around 43 megabytes, which is annoying but survivable.
Attackers know this better than defenders do, and everything interesting about decompression bombs is a technique for beating that 1032:1 wall. There are three, and they fail against different defences, which is why a defence that stops one feels like a solution and is not.
Flat bombs: one member, maximum ratio
The boring one. A single member of highly compressible data — a long run of zeroes — at close to the format's maximum ratio. Ten megabytes in, about ten gigabytes out. This is the bomb that every tutorial on the subject stops at, and it is the one that a ratio check genuinely does catch, because the ratio is right there in the central directory and it is enormous.
It is also, for that reason, the one you will almost never see from anyone competent. If your defence is a ratio check, you have defended against the opening move.
Nested bombs: the ratio becomes multiplicative
The famous one. 42.zip is a cultural artefact at this point: roughly forty-two kilobytes of file that, extracted recursively through its layers, reaches about 4.5 petabytes. It does not beat the 1032:1 wall. It multiplies through it. Each layer contains sixteen copies of the next layer, so the ratio of the whole thing is the product of the ratios of every level, and compression at each level is modest and completely unremarkable.
This is where naive scanners die, and the failure mode is specific enough to be worth naming. Any tool that recursively extracts in order to scan — antivirus, content inspection, a build system unpacking vendored dependencies, a search indexer — has made itself the engine of its own destruction. The outer archive's ratio is boring. Forty-two kilobytes expanding to about four megabytes is not a red flag; it is a zip file. Every individual expansion step looks fine. The explosion is in the recursion, and a per-archive ratio check cannot see it because no single archive in the chain is suspicious.
The fix for nesting is a budget that is global across the whole job rather than per-archive, plus a hard depth cap. Both of which are obvious in hindsight and absent from most code I have read, including some of my own.
Quines and overlapping streams: one layer, unbounded output
The elegant ones. A zip quine is an archive that contains itself: extract it and you get a copy of the input, so a recursive extractor never terminates and never notices, because each step's ratio is exactly 1:1. Depth caps stop it. Ratio caps do not — 1:1 is the most innocent number in the file.
More recently, researchers have shown that you can get petabyte-scale expansion out of a single, non-recursive zip by overlapping the members: many entries in the central directory pointing into the same compressed stream, so the declared total output is a large multiple of the bytes actually stored. There is nothing to recurse into. Depth caps do not help; nor does a heuristic that only distrusts archives containing archives. Treat the specifics as the kind of thing to verify against current research rather than against this paragraph, but the structural lesson is stable: the relationship between an archive's size and its output is not bounded by anything you can derive from the format.
Zip Slip: the bug is in your extraction loop, not in the archive
A zip member is a name plus some bytes, and the name is a string. Nothing in the format says the string has to be a relative path that stays put. Put ../../../../etc/cron.d/x in there and a naive extractor will write to /etc/cron.d/x, which on most systems is remote code execution on a five-minute timer.
Snyk named this Zip Slip in 2018 and found it across an enormous swathe of libraries and applications, and the reason it was everywhere is that extraction is the kind of thing people hand-roll. Open archive, loop over members, join the name onto the destination directory, write the file. Four lines, feels complete, ships. The archive-traversal class is far older than the name — the equivalent in Python's tarfile has had a CVE attached to it since 2007 — and it keeps reappearing because every new language's standard library grows a new extraction loop for people to get wrong in the same way.
The correct check is one sentence long and the order of operations is the entire point: join the member name to the destination, resolve the result, and then verify the resolved path is still inside the resolved destination. Resolve, then check. Checking the raw member name for a literal ".." before you join it is the bug, not the fix — it misses absolute paths, it misses encoding tricks, it misses the case where a component of the path is a symlink that something else already put on the disk, and on a case-insensitive or Unicode-normalising filesystem it misses rather a lot more than that.
One piece of genuinely good news: Python's zipfile does sanitise member names on extract, and has for many years. It strips leading separators and drops ".." components. So if you use ZipFile.extract or extractall you are not vulnerable to plain Zip Slip. You are still fully exposed to bombs, and tarfile is a different story entirely.
# Cheap pre-flight. Do all of it. Believe none of it.
# 1. The declared uncompressed size, straight out of the central
# directory -- which the attacker wrote.
zipinfo -t upload.zip
# 1 file, 10813440000 bytes uncompressed, 10485760 bytes compressed: 99.9%
# 2. A ratio cap is better than a size cap, because it compares two
# numbers instead of trusting one. It still loses to nesting: an
# archive of innocent archives has a completely boring outer ratio.
python3 - <<'PY'
import zipfile
z = zipfile.ZipFile("upload.zip")
raw = sum(i.file_size for i in z.infolist())
comp = sum(i.compress_size for i in z.infolist()) or 1
print(f"members={len(z.infolist())} declared_ratio={raw / comp:.1f}")
PY
# 3. Member names and types, before anything touches a disk.
tar -tvf upload.tar
# -rw-r--r-- root/root 12 ../../etc/cron.d/x
# lrwxrwxrwx root/root 0 pwn -> /
# -rw-r--r-- root/root 31 pwn/etc/cron.d/x
# What none of the above gives you is a bound on what the extractor
# will actually write. unzip has no --max-bytes. GNU tar has no
# --max-bytes. There is no flag. That is the whole problem: the two
# most widely deployed extraction tools on earth cannot be told to
# stop, so "run unzip with limits" is not a configuration task.
Symlink escape: why checking member names is not enough
This is the attack that defeats a path check done the wrong way round, and it is the reason I keep insisting on resolve-then-check.
A tar member can be a symlink. So the archive contains two members, in order. The first is a symlink named pwn whose target is /. Its name is perfectly innocent: no slashes, no dots, passes any name-based filter you care to write. The second is a regular file named pwn/etc/cron.d/x. Also innocent-looking: a relative path with no ".." in it.
Extract them in order and the first member creates a symlink to the root filesystem, and the second member writes through it. The extractor never saw a malicious name, because neither name was malicious. The maliciousness was in the state the first member left on the disk before the second member was checked.
Resolving the path fixes this almost by accident, and it is worth understanding why: realpath follows symlinks that already exist on disk. By the time you resolve pwn/etc/cron.d/x, pwn is a real symlink to /, so the resolved path is /etc/cron.d/x, which is plainly not inside your destination, and you refuse it. A check against the raw string resolves nothing and sees nothing. Hardlinks are the same attack with a different syscall: a hardlink member pointing at /etc/shadow gives the archive's owner a readable, writable alias to a file they never had access to.
GNU tar and bsdtar have their own protections against much of this, which is why "just shell out to tar" is often safer than the extraction loop someone wrote by hand. The exposures that keep producing CVEs are in application-level libraries — Python's tarfile, Node's tar package, Go's archive/tar consumers, Java's ZipInputStream — where the standard library hands you a stream of members and leaves the policy entirely to you.
# Runs INSIDE the disposable guest. This layer is not the boundary --
# the VM is. This layer exists so that a hostile archive produces a
# clean error with a reason in it, instead of a dead machine.
import os
import zipfile
MAX_TOTAL_BYTES = 2 * 1024**3 # across every member, and every pass
MAX_MEMBERS = 10_000 # inodes are a finite resource too
MAX_RATIO = 100 # measured on bytes ACTUALLY written
CHUNK = 1 << 20
def resolve_inside(dest: str, member_name: str) -> str:
"""Join, then resolve, then check. In that order, always."""
# realpath() collapses '..', normalises separators, AND follows any
# symlink an earlier member already created on disk -- which is the
# only reason this also stops the tar symlink-escape chain.
# Scanning member_name for '..' before joining is the classic bug.
dest = os.path.realpath(dest)
target = os.path.realpath(os.path.join(dest, member_name))
if target != dest and os.path.commonpath([dest, target]) != dest:
raise ValueError(f"member escapes destination: {member_name!r}")
return target
def extract_zip(src: str, dest: str, budget: int = MAX_TOTAL_BYTES) -> int:
written = 0
with zipfile.ZipFile(src) as z:
infos = z.infolist()
if len(infos) > MAX_MEMBERS:
raise ValueError(f"too many members: {len(infos)}")
compressed = sum(i.compress_size for i in infos) or 1
for info in infos:
target = resolve_inside(dest, info.filename)
if info.is_dir():
os.makedirs(target, exist_ok=True)
continue
os.makedirs(os.path.dirname(target), exist_ok=True)
with z.open(info) as fin, open(target, "wb") as fout:
while chunk := fin.read(CHUNK):
written += len(chunk)
# Enforced HERE, per megabyte, during the stream --
# not after extractall() returns. A post-hoc check
# runs when the disk is already full, which is the
# entire mechanism of the attack. Checking afterwards
# is not a weaker defence; it is not a defence.
if written > budget:
raise ValueError(f"byte budget exceeded at {written}")
if written // compressed > MAX_RATIO:
raise ValueError(f"ratio exceeded: {written}/{compressed}")
fout.write(chunk)
return written
The rest of the member zoo
Bombs and traversal get the names and the CVEs. These get the 3am pages.
- Device nodes and FIFOs. A tar member can be a character device. Extract one as root and you have created a path that, when some later process opens it, is actually /dev/mem or a block device. A FIFO is gentler and funnier: the next process to read it blocks forever, and your worker queue quietly stops.
- Setuid and setgid bits. tar restores permissions by default when run as root. A member that is a setuid-root shell is a complete local privilege escalation that your extractor installed on request. Always extract with the setuid bits cleared, and never extract as root, and do both because you will forget one.
- Absolute paths. Member names beginning with / are legal in tar. GNU tar strips the leading slash unless you ask it not to; a hand-rolled os.path.join silently ignores the destination entirely when the second argument is absolute, which is a Python semantic that has caused real incidents.
- Inode exhaustion. An ext4 filesystem has a fixed inode count chosen at mkfs time. A few million empty files will exhaust it long before the bytes run out, and the resulting failure mode is the confusing one: df reports plenty of free space and every write fails with ENOSPC. A byte budget does not see this coming, which is why the member count cap is a separate number.
- Unicode normalisation collisions. A member named with the decomposed form of a filename you already wrote is a different byte string on Linux and the same file on macOS and on a normalising filesystem. So your allowlist check passes on the name you read, and the write lands on top of a file you had already vetted. If you extract on one platform and consume on another, this is a real difference, not a theoretical one.
- Pathologically deep trees. A directory nested a few thousand levels deep is trivial to generate and blows the stack of any recursive walker, including the one in your post-extraction scanner. It also exceeds PATH_MAX, at which point ordinary tools can no longer address the leaves — rm -rf fails, your cleanup job fails, and removing the tree requires chdir-ing down it in a loop. Ask me how I know.
# tar is the harder format, because a member can be a symlink, a
# hardlink, a FIFO, a char device or a block device -- and a naive
# loop over the members will cheerfully create all five.
import tarfile
# PEP 706 added extraction filters in Python 3.12. The 'data' filter
# refuses absolute paths, refuses '..' components, refuses links that
# point outside the archive, refuses special files, and strips setuid,
# setgid and sticky bits. In 3.12 you have to ask: omitting filter=
# only earns you a DeprecationWarning and the old unsafe behaviour.
# It becomes the default in 3.14. Pass it explicitly anyway, forever,
# because the version of Python this runs on is not your decision.
with tarfile.open("upload.tar.gz") as tf:
tf.extractall("/work/out", filter="data")
# What filter='data' does NOT do is bound anything at all. It will
# happily write four petabytes, one safe, well-named, non-setuid file
# at a time. The budget from the previous block and the throwaway
# machine from the next one are still both required.
Layering the defences, ranked honestly
There is a natural ordering here, and it is worth being blunt about where each layer stops being useful, because the middle of this list is where most teams stop and feel finished.
Size, ratio and count caps: cheap, necessary, insufficient
Do them. They are a few lines of code, they turn the common case into a clean 400 response, and they make your logs legible. But a cap derived from archive metadata is a cap on a number the attacker typed, and a cap on bytes written only works if you check it while writing. The structural limit is that these are all checks inside the same process that the archive is attacking: if the extractor is a library in your API server, a bomb that gets past the check takes your API server with it.
rlimit and cgroups: better, and still sharing a kernel
Running the extractor as a separate process with RLIMIT_FSIZE, RLIMIT_AS and RLIMIT_NOFILE set is a real improvement, because now the kernel enforces something rather than your loop. The gap is that RLIMIT_FSIZE bounds a single file, so ten thousand members of a gigabyte each sail straight past it, and RLIMIT_AS bounds address space rather than the thing you actually care about, which is disk.
cgroup v2 is better again: memory.max gives you an OOM kill scoped to the extraction, pids.max caps the fork-shaped failures, and io limits stop the extraction from starving everything else on the device. This is a genuinely good layer and you should use it. What it does not do is isolate the disk, which needs filesystem project quotas as a separate piece of work, or isolate the kernel. The cgroup OOM kill I remember most clearly did exactly what it was configured to do, and the host still spent the next several minutes unhappy, because the dirty page cache that extraction had generated was still the host's problem to write out, and dirty-writeback thresholds are a global property of the machine. A limit that fires correctly can still ruin the afternoon of every neighbour on the box.
A microVM: the one that holds, because failure is disposable
Extract inside a Firecracker microVM and the ceilings stop being configuration and start being hardware. The guest has its own kernel, so a kernel-level pathology is the guest's kernel. It has its own block device — a copy-on-write ext4 image — so filling the disk means the guest's writes fail with ENOSPC and nothing outside notices. It has its own memory ceiling: on PandaStack's base template that is 4 GiB with 8 vCPU of burst, baked into the template snapshot. And the failure handling is one line, because the correct response to a machine behaving badly is to stop having it.
The create cost is the thing that makes this practical rather than merely correct. A snapshot-restore create on PandaStack runs at a p50 of 179 ms and a p99 of 203 ms, because every create restores a baked snapshot rather than cold-booting — the restore step itself is around 49 ms. A per-upload VM at a fifth of a second is not a rounding error against the time your extraction was going to take anyway. At $0.054 per vCPU-hour and $0.0162 per GiB-hour, with CPU billed on active CPU-seconds actually burned, a thirty-second extraction costs less than the Slack thread about the 3am page.
import threading
from pandastack import Sandbox
# ttl_seconds is an IDLE timeout, not a lifetime: the reaper measures
# time since last activity. An extractor writing flat out never goes
# idle, so the TTL will never fire on the exact workload you are most
# worried about. The wall-clock kill has to be yours -- two of them,
# in fact: one inside the guest's shell, and one on this side for the
# case where the guest is too wedged to honour its own.
DEADLINE = 300
upload_id = "u_8f2c1a04" # whatever the queue handed you
sbx = Sandbox.create(
template="base",
ttl_seconds=120,
metadata={"job": "extract", "upload": upload_id},
)
# Belt to the guest's braces. sbx.kill() is idempotent.
killer = threading.Timer(DEADLINE + 30, sbx.kill)
killer.start()
try:
# upload() is single-file only -- a directory raises
# IsADirectoryError -- which is exactly what we have.
sbx.filesystem.upload("/tmp/upload.zip", "/work/upload.zip")
sbx.filesystem.upload("./safe_extract.py", "/work/safe_extract.py")
# Byte, member and ratio budgets live in safe_extract.py. timeout(1)
# is the wall clock. The guest's 4 GiB and its own ext4 image are
# the ceilings that hold when both of those turn out to be wrong.
# exec_stream honours timeout_seconds; one-shot exec does not
# reliably do so beyond ~30 s, so stream anything long.
code = sbx.exec_stream(
f"cd /work && timeout {DEADLINE} python3 safe_extract.py upload.zip out",
on_stdout=lambda c: print(c, end=""),
on_stderr=lambda c: print(c, end=""),
timeout_seconds=DEADLINE + 15,
)
if code != 0:
# 124 is timeout(1) reporting that it fired. Anything else is a
# budget tripping, or the archive being a crime.
raise RuntimeError(f"extraction refused (exit {code})")
# Only now does anything cross back out. Tar the vetted tree inside
# the guest, because download() is one file at a time too.
sbx.exec("cd /work/out && tar czf /work/clean.tgz .", check=True)
sbx.filesystem.download("/work/clean.tgz", "/tmp/clean.tgz")
finally:
killer.cancel()
sbx.kill()
What each layer actually stops
| Layer | Flat bomb | Nested bomb | Zip Slip | Symlink escape | Inode exhaustion | Blast radius when it fails |
|---|---|---|---|---|---|---|
| Upload size cap | No | No | No | No | No | Whole host |
| Declared-size header check | Sometimes | No | No | No | No | Whole host |
| Ratio cap before extracting | Yes | No | No | No | No | Whole host |
| Streaming byte + member budget | Yes | Only if global | No | No | Yes | Caller process, dirty disk |
| Resolve-then-check paths | No | No | Yes | Yes | No | Caller process |
| Member-type allowlist | No | No | No | Yes | No | Caller process |
| RLIMIT_FSIZE and RLIMIT_AS | Partly | Partly | No | No | No | One process, disk still fills |
| cgroup memory.max, pids.max, io | Yes | Yes | No | No | Only with quotas | Shared kernel and page cache |
| Disposable microVM | Yes | Yes | Contained | Contained | Yes | One VM you were throwing away |
Read the last two columns together, because that is where the argument lives. The middle rows are cheap and you should have all of them. The bottom row is the only one whose failure column describes something you are already happy to lose.
What the microVM does not buy you
It does not make the archive safe. Nothing makes the archive safe. The bytes are still hostile, the bomb still detonates, the symlink is still a symlink. What changes is the consequence: the explosion happens inside a machine whose entire purpose was to be destroyed, with its own kernel, its own page cache, its own disk and its own memory ceiling, and the cleanup is a kill call rather than an incident review.
Three honest caveats, in order of how likely they are to bite you.
First, the attacker still wins a denial of service against one sandbox. They uploaded a file and they burned a VM. If your pipeline is one sandbox per upload, that is a trade you should be delighted with. If it is one shared sandbox per tenant per day, you have rebuilt the shared-volume problem with extra steps and a slower failure mode.
Second, you still need the in-guest budget, and the reason is the one in the code comment above: the idle TTL cannot bound a continuously-busy workload, because it measures time since last activity and an extractor writing at full speed is never inactive. This is not a hypothetical in our codebase; it is a comment in the reaper, next to the hard lifetime cap that exists precisely because an idle timeout once failed to stop something that stayed busy for ten and a half hours. Your wall-clock kill is your job. Put a timeout in the guest and a deadline timer on your side, and assume one of them will be the one that works.
Third, the dangerous moment is not the extraction. It is the handoff. Everything that comes out of that VM is still attacker-controlled: the filenames, the contents, the directory structure, the file types. A sandbox is a boundary around execution, not a laundering service for data. Decide explicitly what is allowed to cross back — a tarball of vetted paths, a manifest, a list of extracted names you have re-validated on the way in — and do that validation on the trusted side, with the same resolve-then-check discipline, because a filename that was safe in the guest's filesystem is not automatically safe in your object store's key namespace.
The order I would build it in
- Stop extracting in the request path. Whatever else you do, the upload handler should write bytes to object storage and enqueue a job. An extraction inside an HTTP request is a bomb with a direct line to your connection pool.
- Put the extractor in its own disposable machine. One per archive. This is the single change that moves the failure mode from incident to log line, and at a sub-second create it costs you almost nothing in latency.
- Give the job a global byte budget, a global member count, and a depth cap — global across every nested pass, not per archive. Enforce the byte budget during the write loop, not after it.
- Use resolve-then-check for every member, and use your language's safe extraction mode where one exists. In Python that means filter="data" on tarfile, explicitly, on every call, including the ones you think are internal.
- Reject member types rather than filtering them. Symlinks, hardlinks, devices and FIFOs should fail the archive, not be silently skipped, because an archive that contains one is telling you something about its author.
- Bound the wall clock twice: timeout in the guest, a deadline timer on the caller. Assume the guest's own limits can be wedged by exactly the workload you are worried about.
- Validate on the way out, on the trusted side. Re-resolve every path, re-check every name against your storage key rules, and cap the total you accept back.
- Then, once all of that works, try it against 42.zip, a tar with a symlink to /, and a million empty files. If any of the three produces an alert rather than a rejected job, you are not finished.
The reason this problem keeps catching good engineers is that it looks like input validation and it is not. Input validation works when you can tell the difference between good input and bad input by looking at it. Here you cannot, because the dangerous property is not in the bytes you can inspect — it is in the arithmetic of what happens when you act on them, and the attacker chose that arithmetic before you ever saw the file.
So stop trying to tell the difference. Give the archive a machine, a budget and a clock, and let it do its worst inside a boundary that cost you a fifth of a second to build and nothing at all to throw away.
Frequently asked questions
Does checking the declared uncompressed size protect me from a zip bomb?
No, for two separate reasons. The first is that the uncompressed size lives in the archive's own central directory and was written by whoever built the archive, so it is a claim rather than a constraint. Extractors disagree about what to do with it: some truncate output to the declared length, others decompress until the stream ends and only then report a CRC mismatch, by which point the bytes are already on your disk. The second reason is more fundamental. Even a scrupulously honest header does not help you, because a correctly-declared four-terabyte member is still four terabytes. Checking the compression ratio instead is a real improvement, since it compares two numbers rather than trusting one, and it does catch flat single-member bombs. But a ratio check is defeated by nesting, because an archive whose members are each unremarkable archives has an entirely boring outer ratio while the product of the ratios through the layers is enormous. The only size check that actually binds is one you enforce on bytes you have already written, incrementally, as you write them.
Is Python's zipfile module safe to use on untrusted archives?
Partly, and the split matters. ZipFile.extract and extractall do sanitise member names: they strip leading separators and drop ".." components, so plain Zip Slip path traversal is handled for you, and that has been true for many years. Python's zipfile also ignores the Unix mode bits that would make a member a symlink or a device node, writing regular files instead, which removes the symlink-escape class as well. What it gives you no protection against at all is resource exhaustion. extractall will write as many bytes as the archive asks for, create as many files as it contains, and nest as deep as the names say, with no cap on any of it. So the honest summary is that zipfile handles the path problems and none of the quantity problems. You still need a streaming byte budget, a member count cap and a depth limit, and you still want it running somewhere disposable. Python's tarfile is the opposite case: historically unsafe on paths, with a traversal CVE dating to 2007, and only fixed if you explicitly pass filter="data" on Python 3.12 or later.
How does a tar symlink escape get past a path traversal check?
Because neither member name is malicious; the combination is. The archive contains a symlink named pwn whose target is /, followed by a regular file named pwn/etc/cron.d/x. Inspect those two names and both look fine: no leading slash, no ".." component, nothing a name-based filter would object to. Extract them in order and the first member creates a symlink pointing at the root filesystem, and the second member writes straight through it to /etc/cron.d/x. The maliciousness lived in the filesystem state that the first member created, not in any string you examined. This is exactly why the correct check is resolve-then-check rather than inspect-then-join: calling realpath on the joined path follows symlinks that already exist on disk, so by the time you resolve the second member the resolution lands outside your destination and you refuse it. Hardlinks are the same attack with a different syscall — a hardlink member targeting a sensitive file hands the archive's author an alias to something they never had access to. The robust policy is to reject symlink, hardlink, device and FIFO members outright rather than attempt to make them safe.
Will a cgroup memory limit stop a decompression bomb?
It stops the memory half of one, and the memory half is usually not the half that hurts. memory.max gives you an OOM kill scoped to the extraction process group, pids.max bounds the fork-shaped failures, and io limits keep the extraction from starving other work on the device. All worth having. But a decompression bomb's primary weapon is disk, not RAM, and cgroups do not bound disk space — that needs filesystem project quotas as a separate piece of work that most teams have not done. There is also a subtler failure. The dirty page cache generated by a large extraction is written back by the kernel under thresholds that are properties of the whole machine, so an extraction that is correctly confined by its cgroup can still make every other writer on that host slow and unhappy while the kernel flushes. The limit fires exactly as configured and the neighbours still suffer. A microVM closes both gaps structurally: the guest's disk is its own image, its page cache is its own kernel's problem, and the recovery action is to destroy the machine.
Can I use ttl_seconds as a hard wall-clock limit on extraction?
No, and this is the most common misreading of the parameter. ttl_seconds is an idle timeout: the reaper compares the current time against the sandbox's last activity, with a default of five minutes, and deletes the sandbox when the gap exceeds the TTL. A workload that is continuously busy never accumulates idle time, so it is never reaped — and an archive extractor writing at full speed is the canonical continuously-busy workload. The idle TTL is therefore least effective against precisely the case you are worried about. Bound the wall clock yourself, in two places. Inside the guest, wrap the extractor in timeout, which returns exit code 124 when it fires so your caller can tell a deadline from a budget rejection. On the caller's side, set a deadline timer that calls kill on the sandbox, because the guest may be too wedged to honour its own limit. Also note that one-shot exec does not reliably honour long timeouts today; for anything long-running use exec_stream, which does, or bound the command in the shell.
Keep reading
- Sandboxing untrusted uploads: ImageMagick, ffmpeg, LibreOffice — The media-parser half of the same problem — memory-safety bugs in decoders rather than archive-format arithmetic.
- Isolating per-tenant CSV and bulk import pipelines — What happens after extraction, when the thing you unpacked turns out to be a hostile spreadsheet.
- Command injection via untrusted filenames — Why the names inside an archive are dangerous for a second, entirely separate reason once they reach a shell.
- The OOM killer inside a Firecracker guest — What actually happens when the bomb exhausts the guest's memory ceiling instead of its disk.
- How to sandbox untrusted code — The general version of the boundary argument, for the case where the hostile input is a program rather than a file.
- PandaStack sandboxes — Snapshot-restore microVMs with their own kernel, disk and memory ceiling — the disposable machine this post keeps asking for.
Related posts
- Per-Tenant Backup and Restore, Isolated by a MicroVM
A backup worker can read everything by design and a restore worker writes bytes a customer chose. Put a hypervisor boundary and a scoped credential around each job, one tenant at a time.
- A Spreadsheet Is a Programming Language Your Users Don't Call One
Your users write programs in your product every day. They spell them =SUMPRODUCT(...) instead of def main(), and that is the only difference. Then someone asks for PDF export and you shell out to an office suite.
- Per-Tenant MicroVM Isolation for Thumbnail Generation
A JPEG should not be able to read your AWS credentials. It can, though, if your thumbnailer shares a worker with every other tenant's uploads.
- Compiling User-Submitted LaTeX Without Handing Over a Shell
You added LaTeX because the typography is beautiful. You also added a macro language with a documented primitive for running shell commands, and pointed it at strangers.
- Isolating Per-Tenant Geospatial Processing Jobs in MicroVMs
A customer uploads a shapefile and your worker decides, based on those bytes, which of GDAL's hundred-plus drivers to run. That's not a data pipeline — that's attacker-controlled parser dispatch with a database credential in the room.
More in Security & isolation · See PandaStack security
49ms p50 cold start. Fork, snapshot, and scale to zero.