all posts

Analysing APKs Nobody on Your Team Wrote: One MicroVM Per Sample

Ajay Kumar··10 min read

There is a category of pipeline that quietly ingests binaries nobody on the team wrote and runs tooling over them. An app-store or enterprise-MDM review queue that must say yes or no before a build reaches ten thousand managed phones. A mobile AppSec scanner in the MobSF shape, pointed at whatever a customer uploads. A brand-protection service crawling third-party app markets for repackaged clones of a bank's app. A dependency and SBOM service that needs the real library versions, not the ones the vendor claims. A malware research team. Increasingly, an agent-driven triage loop that reads a decompiler's output and decides what to look at next.

All of these have the same shape: an untrusted file arrives, a pile of parsers chew it, and a structured finding comes out. And all of them get built the same way the first time — a worker process on a CI runner, or a container, that shells out to `apktool`, `jadx`, `aapt2` and `unzip` in sequence and writes rows to a findings database.

Then you run it against real submissions for a month. The incident that actually happens is almost never a shell on your host. It is a `jadx` run that pins every core and every spare gigabyte of RAM decompiling one crafted class file. It is an `apktool` invocation that writes several hundred thousand small files and exhausts inodes on the volume your findings database also lives on. It is a zip entry named `../../etc/cron.d/x`. It is a tool that shells out with an attacker-controlled filename and a build-config path that was never meant to be reachable from a scan. The pipeline does not get owned; it gets wedged, and then it gets wedged again, and the on-call rotation becomes a job about restarting scanners.

I build PandaStack, an open-source Firecracker microVM platform, and this is one of the workloads I find easiest to reason about, because the right answer falls straight out of an honest reading of what the file is and what the tools are. So let us read it honestly first and get to the infrastructure second.

The short version: an APK is a ZIP full of untrusted binary formats, and your scanner is a pile of unhardened parsers in four languages. The boundary you need is a kernel you are willing to lose, created per sample and deleted after it. A microVM gives you that for roughly the cost of a container start — but the resource fence that actually saves you still has to be written in-guest, and no amount of isolation gets you a full Android emulator inside a Firecracker guest.

What is actually inside an APK, and which parser eats each part

"Scan the APK" sounds like one operation. It is a dispatch table over half a dozen unrelated binary formats, each handled by a different program with a different pedigree.

  • The container is a ZIP. Not ZIP-like, not inspired by ZIP: an APK is a ZIP archive with a local file header, a central directory and an end-of-central-directory record, and an IPA is a ZIP too. Everything in the archive-parsing threat model is therefore already in scope before you have looked at a single line of Dalvik bytecode.
  • `AndroidManifest.xml` is Android binary XML, not text. If you open it in an editor you get mojibake. Decoding AXML means a resource-chunk parser walking length-prefixed, offset-chasing structures — `apktool`, `aapt2` and `androguard` each have their own, and they disagree about malformed input, which is itself a documented repackaging technique.
  • `resources.arsc` is the compiled resource table: string pools, type specs and configuration variants, all offset-addressed. It is the single most quietly dangerous structure in the file, because it is a large graph of attacker-chosen offsets and counts that every tool parses eagerly to resolve any string at all.
  • `classes.dex`, and `classes2.dex` through `classesN.dex` for multi-DEX builds, are Dalvik bytecode. These are what `baksmali`, `jadx`, `androguard` and every commercial decompiler consume. Decompilation is not a linear pass over the input; it is control-flow reconstruction, which is exactly the kind of analysis you can make blow up superlinearly on purpose.
  • `lib/<abi>/*.so` are native ELF shared objects, usually for several ABIs at once. Analysing those means `file`, `readelf`, `nm`, `strings`, `radare2` or Ghidra — a completely separate ecosystem of C and C++ parsers, several of which are fuzzing targets in their own right.
  • `assets/` and `res/raw/` are arbitrary bytes the developer chose: embedded SQLite databases, archives inside the archive, fonts, video, serialised model files, more native blobs. If your pipeline recurses into them, your threat model is now the union of every format you recurse into.
  • META-INF signature files plus, for v2 and v3 signing, an APK Signing Block sitting between the last entry and the central directory. Signature verification is the one parser in the list people tend to think of as security-relevant, and historical repackaging attacks worked precisely by exploiting disagreements between the verifier's view of the archive and the loader's.

So the sentence "we run static analysis, so we do not execute anything" is not quite right. You are not executing the app's code, but you are feeding seven attacker-controlled binary formats to seven parsers, most of which were written to handle the output of a cooperating build system. That is the actual risk surface, and it is a lot bigger than the one on the whiteboard.

The ZIP comes first, and the ZIP is already enough

Before a single decompiler starts, something has to open the archive, and archive extraction is a solved-and-then-unsolved problem that I have written up at length in Extracting Untrusted Archives: Zip Bombs, Zip Slip, Symlink Escape. The short form for this context: entry names are attacker-controlled strings that your extractor turns into filesystem paths, so `../../` segments, absolute paths and Windows-style separators all have to be rejected rather than normalised; symlink entries can point out of the extraction root and a later entry can then write through them; the central directory's declared sizes are a claim, not a fact, so a ratio check based on them is a cheap pre-filter and never a guarantee; and entry counts are a resource in their own right, because a few hundred thousand tiny entries exhaust inodes and dentry cache long before they exhaust bytes.

What makes the mobile case worse than the generic archive case is that you cannot simply refuse weird archives. A malformed central directory, duplicate entry names, an oversized signing block — in a file-upload feature those are a 400 response. In an app-review or brand-protection pipeline they are the finding. The whole point is to look at the samples other people's tools choke on, which means your extractor has to survive inputs it is explicitly forbidden from rejecting.

Resource exhaustion is the weekly event, not the yearly one

If you run one of these pipelines you already know this, and if you are about to build one, this is the paragraph to believe. Remote code execution through a decompiler is a real risk and it is not your common failure. Your common failure is a tool that consumes every resource on the box while behaving entirely within spec. Concrete shapes, all of which are cheap to produce and none of which require a vulnerability:

  • Flat and nested compression bombs. A high-ratio entry expands to orders of magnitude more than it cost to ship, and an archive inside an archive multiplies that per layer. Your disk fills, and because the scratch volume is usually the same volume as something that matters, the blast radius is not the scan.
  • DEX bombs. Crafted bytecode whose control-flow graph, exception-handler nesting or switch structure drives a decompiler's analysis into pathological time and memory. `jadx` is a JVM program doing graph reconstruction: give it something adversarial and it will allocate until the JVM heap limit, and if there is no heap limit, until the kernel's OOM killer picks a victim — which on a shared host is not always the scanner.
  • Resource-table bombs. Enormous string pools, absurd configuration-variant counts, or offsets that make a parser build a very large in-memory graph from a very small file. This one is the best value-for-bytes attack in the whole format, because `resources.arsc` is parsed eagerly by almost every tool in the chain.
  • Entry-count and filename abuse. `apktool d` writes one file per decoded resource and one smali file per class, so an archive engineered for breadth produces hundreds of thousands of files. Deeply nested directory names hit path-length limits mid-write and leave half-extracted trees behind.
  • Fork storms from tools that shell out. Several analysis tools invoke other binaries, and some do it with filenames derived from archive entries. A tool that shells out in a loop over attacker-chosen names is a fork bomb with extra steps, and process-table exhaustion takes down everything sharing that namespace.
  • Path traversal, still. An entry named `../../etc/cron.d/x` is not a clever attack in 2026, it is a regression test, and it keeps working because people keep writing new extraction code for new formats.

Notice what all six have in common: they are not stopped by syscall filtering, signature verification, running as a non-root user, or a read-only root filesystem. They are stopped by limits — on CPU time, address space, file size, file count and process count — and by making the thing that gets exhausted something you do not mind destroying.

Dynamic analysis changes the question completely. If your pipeline ever runs the sample — a Gradle or build-config path, a plugin, an instrumented harness, a tool that execs with an attacker-controlled argument — you are no longer arguing about parser robustness, you are hosting someone else's code. Treat the boundary accordingly and assume the sample is network-aware, environment-aware and looking for your credentials.

Why container plus seccomp plus cgroups is a partial answer

I want to be fair to containers here, because the usual version of this argument is lazy. A container with a read-only rootfs, a dropped capability set, a seccomp profile, no network and strict cgroup limits is a genuinely meaningful reduction in risk, and it is much better than the CI runner most of these pipelines start on. If that is what you have, it is not nothing.

But look at the specific workload. It is a pile of parsers in four languages — a JVM for `jadx` and `apktool`, native C and C++ for `unzip` and `radare2`, Python for `androguard`, plus whatever the sample's own assets drag in — running as one pipeline against hostile input, on a kernel shared with every other tenant on the host. A seccomp profile that is tight enough to be interesting breaks the JVM; a profile loose enough to run the JVM leaves a large reachable surface. Syscall filtering also does precisely nothing about the six exhaustion shapes above. And cgroup limits, which do address exhaustion, are the thing people get wrong most often: a memory limit without a swap limit, a CPU quota set per-container while the kernel's reclaim path is shared, no `pids` limit at all because nobody expected a fork storm. Meanwhile a scanning fleet is one of the most credential-rich targets in your estate — it holds submissions from your customers, write access to an artifact store, and a database of unreleased findings. This is the shape where "a container is a polite suggestion to the kernel" stops being a joke and starts being a design review. I compare the substrates properly in Firecracker vs Docker: which one runs untrusted code? and MicroVM vs VM vs Container: A 2026 Comparison.

Isolation options for a per-sample mobile-binary analysis job, with the honest drawback of each.
OptionWhat the boundary actually isHonest drawback
Bare CI runnerUnix permissions on a long-lived, credential-rich hostOne traversal entry writes into your runner's filesystem, and the cloud role attached to that box is one metadata request away. State from sample N is still there for sample N+1.
ContainerNamespaces and cgroups over a shared host kernelEvery parser bug lands on the host kernel. Exhaustion is only bounded if you got every cgroup knob right, including the pids limit nobody sets, and the OOM killer's victim selection is host-wide.
Container plus seccomp or gVisorFiltered or re-implemented syscall surfaceA real reduction with a real cost: compatibility gaps land exactly in the JVM and native-tooling corner this workload lives in, and syscall filtering does not bound CPU, memory, inodes or processes at all.
Dedicated VM per sampleHardware virtualisation, a full guest OS per sampleCorrect boundary, wrong economics. Minutes to boot and gigabytes of image per sample, so you end up queueing samples through a small pool of long-lived VMs and reusing dirty state.
MicroVM per sampleHardware virtualisation, restored from a snapshot in about 179 msStill a guest kernel you have to patch, no GPU, and no nested virtualisation — so a full Android emulator inside the guest is out, and true dynamic analysis needs a device or emulator farm instead.

The shape that works: one VM per sample, findings out, VM destroyed

The arrangement I would build, and the one I would defend in a design review: the sample goes in, a structured findings JSON comes out, and the guest kernel that parsed the sample is deleted. No reuse between samples, no shared scratch volume, no long-lived scanner process accumulating state from a thousand submissions. Every sample gets a kernel nobody else has touched, and the worst case for a crafted input is that one disposable VM dies.

That only works if creating the VM is cheap, which is the whole reason microVMs are interesting here. On PandaStack every create is a snapshot restore rather than a boot — p50 179 ms, p99 203 ms, with no warm pool of idle VMs behind it — so per-sample isolation costs about what a container start costs and you stop having to justify reuse. The first spawn of a template with no snapshot yet is a cold boot plus an automatic bake, roughly 3 seconds, and after that you are on the restore path. The mechanics are in The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms if you want them.

# One microVM per sample. The APK goes in, findings.json comes out, the
# kernel that parsed it is destroyed. Nothing is reused between samples.

import hashlib, json, pathlib
from pandastack import Sandbox

SAMPLE = pathlib.Path("/srv/submissions/com.acme.banking-1.4.2.apk")
blob = SAMPLE.read_bytes()
digest = hashlib.sha256(blob).hexdigest()

# A custom template baked with the toolchain pinned: a specific apktool,
# a specific jadx, a specific aapt2. This is also what makes a finding from
# March re-derivable in October -- see the reproducibility section.
# RAM is chosen at bake time and BAKED INTO THE SNAPSHOT: Firecracker cannot
# change vCPU or memory at restore, so a per-create memory_mb is overridden
# to the baked value. If jadx needs more heap, re-bake the template; do not
# pass a bigger number at create time and expect it to take.
sbx = Sandbox.create(
    template="apk-analysis",
    ttl_seconds=20 * 60,        # IDLE clock, not a wall clock. exec traffic
                                # resets it; status and log GETs deliberately
                                # do not. This is the backstop that reaps a
                                # VM whose analysis wedged, NOT the timeout.
    metadata={"sample_sha256": digest,
              "pipeline": "apk-static-v7",
              "source": "mdm-review-queue"},
)

try:
    sbx.exec("mkdir -p /samples /out", check=True)

    # filesystem.write() takes bytes or str. Note that upload() is SINGLE
    # FILES ONLY -- handing it a directory raises IsADirectoryError -- so if
    # you ever need to ship a tree of YARA rules, tar it first.
    sbx.filesystem.write(f"/samples/{digest}.apk", blob)
    sbx.filesystem.write("/opt/analyze.sh",
                         pathlib.Path("analyze.sh").read_text())
    sbx.exec("chmod 0755 /opt/analyze.sh", check=True)

    # The in-guest fence is doing the real work here, and I want to be blunt
    # about why. One-shot exec() does not actually enforce timeout_seconds,
    # and exec_stream() only widens the HTTP timeout -- so a Python-side
    # timeout aborts YOUR REQUEST while jadx keeps burning eight cores inside
    # the guest. timeout(1) in-guest is the only thing that kills the process,
    # and --kill-after turns a SIGTERM a wedged JVM ignores into a SIGKILL.
    cmd = (
        f"timeout --kill-after=10s 900 "
        f"/opt/analyze.sh /samples/{digest}.apk /out"
    )
    rc = sbx.exec_stream(
        cmd,
        on_stdout=lambda line: print("[guest]", line, end=""),
        on_stderr=lambda line: print("[guest:err]", line, end=""),
    )

    # 124 is timeout(1) saying it fired. That is a finding about the sample,
    # not an error in your pipeline: "this binary defeats our decompiler" is
    # exactly what a review queue wants to know.
    findings = {"sha256": digest, "exit_code": rc, "timed_out": rc == 124}
    if sbx.filesystem.exists("/out/findings.json"):
        findings["analysis"] = json.loads(sbx.filesystem.read("/out/findings.json"))

    print(json.dumps(findings)[:400])

finally:
    # Teardown is kill(); there is no sbx.delete(). In a finally block because
    # an exception on the client side must not leave a VM holding a sample.
    sbx.kill()

Two details in there are worth pulling out. The first is that a timeout is a finding: an app whose bytecode reliably defeats your decompiler is interesting, and a pipeline that logs that as an infrastructure error throws away a signal. The second is that `ttl_seconds` is an idle clock rather than a wall clock, so it is a backstop for a wedged analysis and not a substitute for a timeout. If the guest is grinding away with no API traffic, the idle reaper will eventually delete the VM — and "eventually" is the wrong granularity for a scanner, which is why the fence is in-guest. The full semantics are in How to control sandbox lifetime: TTL, idle, and cleanup.

The in-guest fence, with a reason for every limit

This is the script that actually stops the six failure shapes. Every limit here exists because something specific happened without it.

#!/bin/sh
# /opt/analyze.sh <sample.apk> <outdir>  -- runs INSIDE the microVM.
# Defensive discipline first, analysis second. The VM is the blast wall; these
# limits are what stop one sample from wasting a whole scheduling slot.
set -u
SAMPLE="$1"; OUT="$2"
mkdir -p "$OUT" "$OUT/unpacked" "$OUT/logs"

# ---- the fence -------------------------------------------------------------
ulimit -t 600        # CPU SECONDS. Catches the tool that is busy rather than
                     # slow: a resource-table parser spinning on crafted
                     # offsets never trips a wall clock until it is too late.
ulimit -f 2097152    # max file size, in 1K blocks (~2 GiB). Stops a
                     # decompiler emitting one absurd .java file and stops a
                     # compression bomb inflating a single entry forever.
ulimit -u 192        # MAX USER PROCESSES -- the fork-storm cap, and the limit
                     # people forget. Caveat: this is per-UID, so run this
                     # script as a dedicated unprivileged user or you have
                     # also capped your own shell.
ulimit -n 1024       # file descriptors. An extractor opening every entry at
                     # once is a real pattern.
ulimit -c 0          # no core dumps. A 4 GiB core from a crashed parser is
                     # how you fill the disk while debugging a crash.
# Deliberately NOT using `ulimit -v` for the JVM steps: the JVM reserves a
# large virtual address space up front and simply refuses to start under a
# tight -v. Bound the JVM with -Xmx instead, and let the VM's own baked RAM be
# the outer wall. `ulimit -v` is still worth setting around the native tools.

# ---- inspect the archive BEFORE trusting any of it --------------------------
zipinfo -1 "$SAMPLE" > "$OUT/entries.txt" 2>"$OUT/logs/zipinfo.err" || {
  echo "verdict=unreadable_zip"; exit 2; }

ENTRIES=$(wc -l < "$OUT/entries.txt")
if [ "$ENTRIES" -gt 25000 ]; then
  echo "verdict=entry_count_exceeded entries=$ENTRIES"; exit 3
fi

# Zip Slip and friends. Reject, never normalise: a normaliser is one more
# parser with its own disagreements. Absolute paths, dot-dot segments, and
# backslashes (which some extractors helpfully treat as separators).
if grep -q '^/' "$OUT/entries.txt" \
   || grep -Eq '(^|/)[.][.](/|$)' "$OUT/entries.txt" \
   || grep -Fq '\' "$OUT/entries.txt"; then
  echo "verdict=hostile_entry_name"; exit 4
fi

# The declared ratio is a CHEAP PRE-FILTER, not a guarantee: these numbers
# come from the central directory, which the attacker wrote. The real
# guarantees are `ulimit -f` above and the fact that the disk is disposable.
unzip -Zt "$SAMPLE" > "$OUT/logs/totals.txt" 2>&1 || true

# ---- analysis, each step fenced independently ------------------------------
# Separate timeouts, because one slow stage should not consume the budget of
# the stage that would have produced the finding you care about.
run() { name="$1"; secs="$2"; shift 2
  timeout --kill-after=10s "$secs" "$@" \
    > "$OUT/logs/$name.out" 2> "$OUT/logs/$name.err"
  echo "$name=$?" >> "$OUT/logs/rc"
}

run badging 60  aapt2 dump badging "$SAMPLE"
run apktool 420 apktool d -f --no-debug-info -o "$OUT/unpacked" "$SAMPLE"

# jadx: bound the HEAP, not the address space, and accept a non-zero exit --
# partial decompilation with errors is the normal outcome on obfuscated code.
JAVA_TOOL_OPTIONS="-Xmx2g -XX:+ExitOnOutOfMemoryError" \
  run jadx 600 jadx --no-res --no-imports -d "$OUT/src" "$SAMPLE"

find "$OUT/unpacked/lib" -name '*.so' -type f 2>/dev/null \
  | head -200 > "$OUT/logs/natives.txt"

# ---- serialise ONE json, without shell-interpolating attacker strings ------
# -I: isolated mode, so a planted module in the extraction tree cannot be
# imported by the serialiser. Run it from a directory the sample never wrote.
cd / && python3 -I - "$OUT" <<'PY' > "$OUT/findings.json"
import json, os, sys
out = sys.argv[1]
def read(p, cap=200000):
    f = os.path.join(out, p)
    if not os.path.exists(f): return ""
    with open(f, errors="replace") as fh: return fh.read(cap)
json.dump({
    "entries": len(read("entries.txt").splitlines()),
    "badging": read("logs/badging.out"),
    "zip_totals": read("logs/totals.txt"),
    "return_codes": read("logs/rc"),
    "native_libs": read("logs/natives.txt").splitlines(),
}, sys.stdout)
PY
echo "verdict=complete"

Read the limits as a list of incidents. `ulimit -t` is the parser that spins without allocating. `ulimit -f` is the decompiler that writes one enormous file. `ulimit -u` is the tool that shells out in a loop. `ulimit -c` is the 4 GiB core dump that filled the disk during the investigation of a crash. The `ulimit -v` comment is the one I would most like people to take away, because "cap the address space" is standard advice that quietly does not work for JVM tooling, and a scanner whose `jadx` step has been silently failing to start for three weeks is a worse outcome than one with no limit at all.

Network policy, which matters more here than almost anywhere

Mobile samples phone home. Even in a static pipeline you will trip outbound traffic: a Gradle or build-config path that resolves dependencies, a tool that fetches a symbol server, an embedded SDK's first-run beacon, a packer that pulls its payload on unpack. If the sample is part of a campaign, that beacon tells the operator that a scanning fleet looked at their build, from which network, at what time. For a brand-protection or malware-research pipeline that is a tipping-off problem, not just a hygiene one.

On PandaStack each sandbox gets its own Linux network namespace with a veth pair and tap device, pre-allocated — 16,384 per-sandbox subnets per agent host, which is also the hard ceiling on sandboxes per host before RAM and CPU become the real limit. Sibling sandboxes cannot reach each other's subnets, and the cloud metadata IP range is dropped at the host, so the classic container-escalation-to-cloud-role path is closed. But I have to be straight about the default that matters most for this workload: egress to the internet is open by default. Your own VPC, your artifact store, your findings database and your internal APIs are not protected by anything I ship — writing those deny rules is still your work, and for sample analysis the correct default is almost certainly deny-all with a tiny allowlist, or no egress at all. Controlling Network Egress for Untrusted Code is the long version. One related trap: preview URLs are tokenless, so a guest port is reachable at a predictable host for the sandbox's lifetime with the UUID as the only credential. Do not expose a port on a sample-analysis VM for convenience.

Batching samples, and making a March finding re-derivable in October

Two problems with one solution. The batching problem is that you have ten thousand queued samples and do not want to install a toolchain ten thousand times. The reproducibility problem is harder and more important: a finding is an assertion about a binary produced by a specific version of a specific decompiler, and when a customer disputes it six months later, "we re-ran it and got something else" is not an answer. Pin the toolchain into a snapshot and both problems collapse into one mechanism.

# Bake the toolchain once, restore it per sample. The snapshot id IS the
# provenance record: store it on every finding and you can re-derive that
# finding months later against the exact same apktool and jadx builds.

from pandastack import Sandbox

warm = Sandbox.create(template="apk-analysis", ttl_seconds=15 * 60)
warm.exec_stream(
    "timeout --kill-after=10s 300 sh -c "
    "'apktool --version && jadx --version && aapt2 version'",
    on_stdout=print,
)
TOOLCHAIN = warm.snapshot()   # synchronous; writes the full guest memory image
warm.kill()
print("toolchain snapshot:", TOOLCHAIN)

def analyse(digest, blob):
    # Every sample gets its own restore of the SAME frozen toolchain. A restore
    # on every create, p50 179 ms, with no warm pool of idle VMs behind it.
    sbx = Sandbox.create(
        from_snapshot=TOOLCHAIN,
        ttl_seconds=20 * 60,
        metadata={"sample_sha256": digest, "toolchain": TOOLCHAIN},
    )
    try:
        sbx.exec("mkdir -p /samples /out", check=True)
        sbx.filesystem.write(f"/samples/{digest}.apk", blob)
        rc = sbx.exec_stream(
            f"timeout --kill-after=10s 900 /opt/analyze.sh /samples/{digest}.apk /out",
            on_stdout=print,
        )
        return rc, sbx.filesystem.read("/out/findings.json")
    finally:
        sbx.kill()

# On fan-out, and on the one API pair people get wrong:
#
#   from_snapshot=   restores the frozen toolchain into a FRESH VM. This is the
#                    right call for sample analysis -- each sample wants its own
#                    JVM anyway, so there is no warm in-memory state to inherit.
#
#   fork_tree(n)     snapshots the parent once and boots n children that inherit
#                    the parent's RUNNING MEMORY. Use it when a warm process is
#                    the expensive thing. HARD CAP of 16 children per call --
#                    more is an error, not a clamp -- so wider fan-outs grow the
#                    tree breadth-first.
#
#   fork()           is NOT the warm one. It clones the DISK and the child
#                    COLD-BOOTS: 400-750 ms same-host, 1.2-3.5 s cross-host,
#                    with its own entropy and its own PIDs. Reaching for fork()
#                    when you wanted fork_tree() silently costs you the warm
#                    state you were paying for.
#
# One caveat worth knowing if you ever add randomised fuzzing to this pipeline:
# N guests restored from one snapshot start from the same RNG state. That is a
# feature for reproducibility and a bug for sampling diversity.

Storing the snapshot id alongside every finding is the cheapest audit improvement available to a scanning pipeline. It turns "our scanner said X" into "this exact frozen toolchain said X, and here it is." Snapshots are durable and outlive the VM that made them, and an idle auto-reap does not cascade-delete them, so you can let every analysis VM die on its own schedule while the toolchain you pinned stays put. Snapshots and Forks: Copy-on-Write for Running Machines covers the primitives.

What this cannot do, stated plainly

The most useful thing I can tell you about this architecture is where it stops.

  • No Android emulator, and this is the big one. Nested virtualisation is not available inside a Firecracker guest, so you cannot run KVM-accelerated anything in there — which means no Android Emulator, no hardware-accelerated system image, no instrumented dynamic analysis of the app as an app. Full software emulation through a TCG-mode QEMU is theoretically startable and practically unusable for a modern Android image. If you need true dynamic analysis, the answer is a real-device farm or an emulator farm on bare metal with KVM, and this pipeline is the static half that feeds it.
  • No GPU. There is no PCI or VFIO GPU passthrough, so anything GPU-accelerated does not run. Software rasterisers such as llvmpipe and lavapipe do work, slowly, which is enough for a headless render but not for anything resembling a frame budget.
  • Guest kernel 5.10 on Ubuntu 24.04 userland. Modern tooling almost always works, but if an analysis tool depends on a newer kernel interface, that is a real constraint rather than a detail. You Ship the Kernel: Firecracker Guest 5.10 vs 6.1 has the specifics.
  • vCPU and RAM are baked into the template snapshot. Firecracker cannot change them at restore, so a per-create memory_mb is overridden to the baked value when a snapshot exists. Giving jadx more heap means re-baking the template, not passing a bigger number at create time.
  • It is still a kernel. A microVM is a much smaller and much better-defended boundary than a shared kernel, but it is a boundary with a CVE feed. You still patch, you still watch Firecracker advisories, and you still assume a determined researcher is a possibility rather than an impossibility.

On cost, since per-sample isolation sounds expensive and mostly is not: one rate card, $0.054 per vCPU-hour and $0.0162 per GiB-hour, where CPU bills on CPU-seconds actually burned and memory bills committed GiB-hours for as long as the sandbox exists. For this workload the CPU term dominates, because a decompiler really does burn cores, and that is the correct incentive — it means a tight in-guest timeout is a cost control as well as a safety control, and it means the ten thousand samples that unpack in four seconds each cost roughly what they consume rather than what you provisioned.

What I would actually do

If you are running mobile binaries through a toolchain today on a CI runner or a container, the single highest-value change is not a tighter seccomp profile. It is making the unit of analysis disposable: one kernel per sample, a hard in-guest fence on CPU time, file size, file count and process count, deny-all egress, a pinned toolchain snapshot whose id you record on every finding, and teardown in a `finally`. That is a weekend of work and it converts your most common incident from a page into a row in a table that says `timed_out: true`.

And the honest counter-case, because it exists. If what you need is dynamic analysis — running the app, driving the UI, watching the traffic it generates, hooking its runtime — PandaStack is the wrong tool and I would rather tell you now than after you have built half of it: no nested virtualisation means no Android emulator in the guest, and no GPU means no accelerated rendering. Go get a device farm. If your samples are first-party builds from your own CI and the pipeline is a formality, a container is fine and you should spend the effort elsewhere. And if your corpus is small enough that a human looks at every sample anyway, you do not have an infrastructure problem yet. The microVM-per-sample shape earns its keep exactly when the samples are numerous, untrusted, and written by people who would be delighted to learn what your scanner does with a crafted DEX. Everything here is open source and the claims are checkable: AI agent sandboxes for the primitives, pricing for the rate card, What is a microVM? Firecracker, isolation, and why agents need it if you want the substrate argument from first principles.

Frequently asked questions

Can I actually run the app, not just analyse it? Can I put an Android emulator inside the microVM?

No, and this is the clearest limit in the whole architecture, so I will not soften it. Nested virtualisation is not available inside a Firecracker guest. The Android Emulator needs hardware acceleration — KVM on Linux for x86 system images — and it simply cannot get it from inside a microVM, because the guest has no KVM of its own to hand to a nested hypervisor. You can start a fully software-emulated QEMU in TCG mode and it will technically execute instructions, but a modern Android system image under full emulation is not a usable dynamic-analysis platform; it is a very slow way to watch a boot animation. There is also no GPU: no PCI or VFIO passthrough, so no hardware-accelerated rendering, and llvmpipe or lavapipe will carry a headless render but not an interactive UI. The practical division of labour is that microVMs are excellent for the static half of a mobile pipeline — unpack, decode AXML and resources.arsc, decompile DEX, triage native libraries, run YARA and secret scanners, read the signing block — and the dynamic half belongs on a real-device farm or an emulator farm running on bare metal with KVM available. Build the static pipeline here, feed the interesting samples there, and keep the two cleanly separated. Pretending one substrate does both is how you end up with a dynamic analysis stage that silently never ran.

Is a hardened container not enough? I already drop capabilities, use a read-only rootfs and apply a seccomp profile.

It is a genuine improvement and far better than the bare CI runner most of these pipelines start on, so keep doing it. But look at the two things it does not cover. First, the kernel: this workload is a JVM plus native C and C++ parsers plus Python, all consuming attacker-chosen binary formats, and in a container every one of those parser bugs reaches a kernel shared with every other tenant on the host. A seccomp profile tight enough to be interesting breaks the JVM, and a profile loose enough to run the JVM leaves a large reachable surface. Second, and more immediately relevant, syscall filtering does nothing at all about resource exhaustion, which is your actual weekly failure. A compression bomb, a DEX bomb, a resource-table bomb and a fork storm are all perfectly ordinary syscall sequences. Those need cgroup limits, and cgroup limits are where I see the most mistakes in practice: a memory limit with no swap limit, a CPU quota set per-container while kernel reclaim stays shared, and almost always no pids limit, because nobody designed for a tool that shells out in a loop over archive entry names. The microVM answer is not that containers are useless. It is that a kernel per sample makes the exhaustion question boring: the thing that gets exhausted is a disposable guest, and the recovery is a delete.

How do I stop a sample from phoning home during analysis?

Deliberately, because it will not happen by accident. On PandaStack every sandbox gets its own Linux network namespace with a veth pair and tap device, pre-allocated, so sibling sandboxes cannot reach each other's subnets and the cloud metadata IP range is dropped at the host — which closes the escalate-to-cloud-role path that makes container escapes so profitable. What is not closed for you is the general internet: egress is open by default, and your own VPC, artifact store, findings database and internal APIs are protected by nothing I ship. For sample analysis specifically I would default to no egress at all, and add a tiny allowlist only if a stage genuinely needs one — a symbol server, a reputation API. Expect to trip outbound traffic even in a static pipeline: build-config paths resolve dependencies, SDKs beacon on first run, packers fetch payloads on unpack. The tipping-off risk is real for brand-protection and malware-research work, because a beacon tells the operator which network looked at their build and when. Two related traps. Preview URLs are tokenless, so a guest port is reachable at a predictable host for the VM's lifetime with the UUID as the only credential — do not expose a port on an analysis VM for convenience. And do not pass real credentials into the guest as environment variables just because it is easy; the sample's tooling can read the environment.

What timeouts and resource limits should I actually set, and where?

In the guest, always, and this is the single most common mistake I see. The client-side knob looks like the answer and is not: one-shot exec() does not actually enforce timeout_seconds, and exec_stream only widens the HTTP timeout, so a Python-side timeout aborts your request while jadx keeps pinning eight cores inside the guest. Wrap every untrusted or long-running command in timeout --kill-after=10s N sh -c '...' so a SIGTERM a wedged JVM ignores becomes a SIGKILL. Then add ulimits inside the analysis script, one per failure mode: ulimit -t for CPU seconds, which catches a parser that is busy rather than slow; ulimit -f for maximum file size, which stops a decompiler emitting one absurd file and a bomb inflating a single entry; ulimit -u as a fork cap, remembering it is per-UID so you need a dedicated unprivileged user or you have capped your own shell; ulimit -n for descriptors; ulimit -c 0 so a crashed parser does not fill the disk with a multi-gigabyte core. One real gotcha: do not use ulimit -v around JVM tooling. The JVM reserves a large virtual address space up front and refuses to start under a tight -v, so you get a silently skipped decompilation stage rather than a bounded one. Bound the JVM with -Xmx and let the VM's baked RAM be the outer wall. Treat the idle TTL as a backstop for a wedged VM, not as a timeout — it is an idle clock, not a wall clock.

How do I make a finding from March re-derivable in October?

Pin the toolchain into a snapshot and record the snapshot id on the finding. A finding is not a property of a binary; it is a property of a binary as seen by a specific build of a specific decompiler, with a specific rule set and a specific set of heuristics. Six months later, when a customer disputes a result, re-running against whatever apktool and jadx resolve to today is not a reproduction — it is a new experiment. So bake a custom template with the versions pinned, create one VM, verify the versions, call snapshot(), and then create every per-sample VM with from_snapshot pointed at that id. Every sample is analysed by a bit-identical toolchain, and the id on the finding row is your provenance record. Snapshots are durable and outlive the VM that made them, and an idle auto-reap deliberately does not cascade-delete a sandbox's snapshots, so analysis VMs can die on their own schedule while the pinned toolchain stays put. Two caveats. A pinned toolchain goes stale against new obfuscators, so you will maintain several generations and re-run a sample corpus when you cut a new one — which is itself much easier when restores are cheap. And if you add randomised components such as fuzzing, remember that N guests restored from one snapshot start with the same RNG state: excellent for reproducibility, quietly terrible for sampling diversity.

Keep reading

Related posts

  • Scanning Package Postinstall Scripts in a microVM

    The only reliable way to find out what an install script does is to run it — which is the exact thing you were trying to avoid. So run it somewhere disposable, with the network turned off, and let the package tell on itself.

  • Testing Browser Extensions with AI Agents in MicroVMs

    An extension gets content-script access to every page you visit, your cookie jar, and a background worker that can phone home. Installing one to 'just test it' on your laptop is a trust decision. Do it in a microVM you throw away.

  • Server-Side Rendering Someone Else's React Component

    If your product server-renders components your customers wrote, you are running their code in your process, with your env vars and your database socket. node:vm is a sandbox for values, not for a module graph that can require("fs"). Here is the boundary that actually holds, and what it costs.

  • Rendering Untrusted Ad Creatives in Per-Tenant MicroVMs

    Your creative-review pipeline renders HTML5 ad units from thousands of advertisers in one shared headless-Chrome farm. That farm is a giant C++ attack surface executing code that strangers paid to have you execute. Give every creative its own microVM.

  • How to Vet a Code Execution Vendor's Security

    If your AI agent runs model-generated code, you've outsourced a security boundary. Here are the questions worth asking a vendor, why SOC 2 answers almost none of them, and what the honest answers sound like.

More in Code execution · See Code interpreter sandboxes on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.