all posts

nftables vs iptables for Per-Sandbox Networking: Why I Still Write iptables

Ajay Kumar··11 min read

I build PandaStack, an open-source Firecracker microVM platform, and the per-host agent writes legacy `iptables` rules. Four of them, in the nat table, inside each sandbox's own network namespace: a DNAT per published port, a catch-all DNAT, and two SNATs. Thirteen more in the root namespace, shared by every sandbox on the box. That is it. No nftables, no eBPF datapath, no clever map lookups.

Every time someone reads that part of the code I get the same note, phrased politely: you know iptables is deprecated, right? And the honest answer is that I went and checked, several times, and each time I came back having learned something that made the choice look better rather than worse. Not because iptables is great. Because the thing nftables is genuinely better at is a problem we do not have, and the problems we do have are not solved by either of them.

The short version: nftables wins on atomic rule-set replacement and on set/map lookups instead of linear chain traversal. Both of those are arguments about large, shared, frequently-rewritten rule sets. If your architecture gives every workload its own namespace with a handful of rules in it, the second argument mostly evaporates and the first one becomes a nicety rather than a necessity. The thing that actually bites at sixteen thousand namespaces is conntrack, the neighbour table, and NAT tuple space — none of which care which front end you used.

Two front ends, one kernel, and the shim nobody mentions

Start with what these things are, because the vocabulary is a trap. People say "iptables vs nftables" as if it were a choice between two kernel subsystems. On almost every machine you will touch in 2026, it is not.

Classic iptables is a userspace tool over `ip_tables`, a matcher framework in the kernel with four largely independent tables — `filter`, `nat`, `mangle`, `raw` — each with its own chains hooked into netfilter at fixed points. Rules are matched in order, and extensions arrive as `-m` modules. The userspace side has a famous design wart: adding one rule means reading the whole table blob out of the kernel, splicing your rule into the copy, and writing the whole blob back. One rule, two full table transfers.

nftables replaced that with one unified framework. A single register-based virtual machine in the kernel evaluates bytecode; there are no built-in tables, you declare your own with explicit families (`ip`, `ip6`, `inet`, `arp`, `bridge`, `netdev`), chains with explicit hooks and numeric priorities, typed sets and maps, named counters, and — the big one — an atomic transaction model. An `nft -f` file is applied as one unit or not at all.

Here is the part that changes the whole conversation. Since roughly Debian 10 and Ubuntu 20.04, the binary called `iptables` on your host is usually not classic iptables. It is `iptables-nft`, a compatibility front end that parses iptables syntax and speaks the nftables kernel API underneath, creating nft tables with compat markers. The genuinely legacy backend still exists as `iptables-legacy`, and on most distributions it is one entry in an alternatives symlink that nobody has looked at in years.

# The one command that settles the argument on any given host.
$ iptables -V
iptables v1.8.10 (nf_tables)

# "(nf_tables)" means: iptables syntax, nftables kernel API.
# "(legacy)" would mean the real ip_tables matcher framework.

# Which binaries are actually installed, and which one the name points at:
$ update-alternatives --display iptables 2>/dev/null | head -4
$ ls -l /usr/sbin/iptables

# The rules you "wrote in iptables" are visible from the other front end:
$ nft list ruleset | head -20
table ip nat {
	chain PREROUTING {
		type nat hook prerouting priority dstnat; policy accept;
		meta l4proto tcp ip daddr 10.200.0.2 tcp dport 22 counter packets 4 bytes 240 \
			# xt_DNAT
	}
}

# And the inverse does NOT hold: iptables-save cannot see nft-native rules.
# That asymmetry is the single most common cause of "my DROP rule is gone."

So when I say PandaStack writes iptables, the precise statement is: the agent shells out to the `iptables` binary with legacy syntax, and on our hosts that binary is `iptables-nft`, so the kernel is running nftables. Both sentences are true. "We use iptables" and "we are running nftables" are not in tension; they describe different layers. Anyone who tells you the migration is urgent because the old code path is unmaintained has usually not run `iptables -V` on the machine in question.

This matters beyond pedantry: most performance arguments people make about iptables do not apply to the thing they are actually running. The whole-table-blob rewrite is a property of `iptables-legacy`; under `iptables-nft` each add is an incremental netlink transaction. The linear-traversal problem does survive the shim, because ordered evaluation is what iptables semantics mean.

Where nftables genuinely wins

I want to make the strongest version of this case, because I am about to argue against it and a weak version would be dishonest.

Atomic rule-set replacement. This is a correctness property, not a convenience. With `nft -f`, the kernel commits your entire file as one transaction. There is no instant at which half your rules are live. With iptables, `iptables-restore` is atomic per-table, but a sequence of individual `-A`/`-I` calls — which is what every orchestrator that builds rules programmatically actually does — has a visible window per rule. If your rule set is a security boundary, there is a moment during every reload where it is a different boundary. I have debugged a production incident whose entire causal chain was "the DROP had not landed yet." That is worth something.

Sets and maps instead of rules. In nftables, "these ten thousand IPs are blocked" is one rule referencing a typed set with a hash or interval backing, and the lookup is O(1)-ish rather than ten thousand sequential comparisons. "This port maps to that address" is a map entry, not a rule. This is the difference between expressing policy as data and expressing it as control flow. The canonical demonstration is Kubernetes: kube-proxy's iptables mode generates rules proportional to services times endpoints, which is why a large cluster's rule set runs to tens of thousands of rules and why both IPVS and nftables modes exist. Not a vendor failing — the linear-chain model meeting a workload it was never shaped for.

One tool for v4 and v6. The `inet` family means one rule set, one policy, one review. With iptables you maintain `iptables` and `ip6tables` as parallel universes, and the bug is always in the one you forgot.

Fewer round-trips for a big rule set. One `nft -f` is one transaction regardless of how many rules are in the file. Building the same thing with N `iptables` invocations is N process spawns and N netlink conversations.

Observability that was designed in. `nft monitor` streams rule-set and conntrack events as they happen, which beats polling. Named counters attach to the thing you care about instead of making you infer from per-rule counters. And `nft -j` emits JSON, so your agent stops parsing human-readable text — the sort of thing that only bites during an incident.

Per-tenant policy as data. A verdict map keyed on a tenant identifier, whose values are verdicts, replaces a rule per tenant. Adding a tenant becomes `nft add element`, which is a single small transaction against a set, rather than an insert into a chain at a position you have to reason about.

That is a real list. If I were designing a host-global policy engine from scratch today, I would not consider iptables.

Where iptables still wins, including at 3am

Now the other side, and I will separate the architectural arguments from the human ones, because conflating them is how people end up with a rule set only one person on the team can read.

`-C` makes idempotent application cheap. This is the one that actually drove our design. `iptables -C` checks whether an identical rule exists and exits non-zero if it does not. That gives you check-then-add in two cheap calls, with no parsing, no diffing, and no state file. Our agent's `ensureRoot` helper is literally that: build the `-C` form, run it, and if it exits non-zero, run the `-A` form. The shared root rules are re-asserted on every create and the operation is a no-op after the first. The nftables equivalent is rebuilding the table atomically from a template — better engineering, more code, and it requires you to own the whole table rather than cooperate with whatever else is on the host.

The ecosystem is enormous. Every `-m` match, every target, every piece of third-party tooling, every Stack Overflow answer from the last twenty years, every cloud provider's documentation, and — importantly — every other daemon on your host that writes firewall rules. Docker writes iptables. Many cloud agents write iptables. `fail2ban` writes iptables. If anything else on the box is managing rules, you are already in a mixed world, and choosing nftables for your own rules does not fix that; it makes you the second opinion in an argument.

It is greppable from a shell, by anyone. This is not an architectural property and I will not pretend it is. It is an operational one, and those are real. At 3am the question is "is traffic being dropped, and by what," and `iptables -t nat -L -n -v` with its packet counters answers it in a form every engineer can read without looking anything up. `nft list ruleset` is more precise, more complete, and reads like a configuration language — exactly the wrong genre of document at 3am. I have watched a competent engineer lose four minutes to nftables syntax mid-incident. Four minutes is a long time.

And the shim means you get a lot of nftables anyway. Because our `iptables` is `iptables-nft`, we already get the incremental-netlink behaviour and the modern kernel code path. What we give up is the expressive surface: sets, maps, atomic files, the `inet` family, JSON output. We keep the familiar front end. That is a real trade and it is not obviously the wrong side of it.

The three of them, side by side

iptables-legacy, iptables-nft and native nftables on the properties that matter for a per-workload NAT path. Kernel and distribution behaviour moves between releases -- verify against your own kernel and `iptables -V` before you rely on any row.
Front endRule-set updateLookup shapeIPv4 + IPv6Per-namespace fitFamiliarityWhat breaks when you mix it
iptables-legacyWhole-table blob read + write per rule; `-restore` atomic per tableStrictly linear per chainTwo separate tools and two rule setsFine; cost is per-rule process spawnUniversalAttaches to the same hooks by a different mechanism, so both rule sets evaluate and neither tool shows the other's rules
iptables-nftIncremental netlink transaction per rule; no blob rewriteStill linear -- iptables semantics are orderedStill two tools, though one kernel subsystemGood; cheap `-C` idempotency, tiny rule setsUniversal (same syntax)Visible to `nft list ruleset` but NOT the reverse -- `iptables-save` silently omits nft-native rules
nftablesAtomic: `nft -f` commits the whole file or nothingSets and maps give hash/interval lookups; chains still orderedOne `inet` rule set, one policyGood; needs you to own the table, and `nft` may be absent on minimal imagesPatchy outside network teamsSame-priority chains from two sources have creation-order semantics you should never depend on

Read the last column twice. The failure mode of a half-finished migration is not slowness; it is a rule you believe is enforced being evaluated after something else already accepted the packet.

What PandaStack actually writes into each namespace

Concrete now. PandaStack pre-allocates 16,384 `/30` subnets per agent out of `10.200.0.0/16`. Each sandbox gets its own network namespace `ns-<id>`, a veth pair with `vh-<id>` in the root namespace and `vg-<id>` inside, and a `tap0` that Firecracker attaches to. The guest's IP, MAC and gateway are frozen at template-bake time and never change, which is the whole trick: because each VM lives in its own namespace, every sandbox from the same template can be handed the same baked guest IP with no collision, so a restored snapshot finds the network exactly as it was when the snapshot was taken.

That design decision is what creates the NAT. Here is the real rule set.

#!/bin/sh
# Per-sandbox netns. NS=ns-abc123, veth guest side vg-abc123.
# 10.200.0.2  = this slot's UNIQUE veth-guest IP (root-reachable)
# 172.20.6.118 = the BAKED guest IP, identical in every sandbox of this template
# 172.20.6.117 = the baked gateway the guest believes in
set -eu
NS=ns-abc123

# 1. Explicit per-port DNAT. PREROUTING is ordered, so this is matched
#    before the catch-all below and port 22 keeps its own destination.
ip netns exec "$NS" iptables -t nat -A PREROUTING \
  -d 10.200.0.2 -p tcp --dport 22 \
  -j DNAT --to-destination 172.20.6.118:22

# 2. Catch-all DNAT, so the agent can reach any port the user listens on
#    (3000 for Next.js, 8000 for uvicorn) without touching the firewall
#    on every request. ORDER IS LOAD-BEARING: this must come second.
ip netns exec "$NS" iptables -t nat -A PREROUTING \
  -d 10.200.0.2 -p tcp \
  -j DNAT --to-destination 172.20.6.118

# 3. Inbound replies must appear to come from the baked gateway, or the
#    guest's default route does not match and it answers into the void.
ip netns exec "$NS" iptables -t nat -A POSTROUTING \
  -o tap0 -j SNAT --to-source 172.20.6.117

# 4. Egress: rewrite the SHARED baked source to this slot's UNIQUE IP
#    before the packet reaches the root namespace. Without this, two
#    sandboxes of the same template present identical tuples upstream
#    and conntrack in the root namespace cannot tell them apart.
ip netns exec "$NS" iptables -t nat -A POSTROUTING \
  -s 172.20.6.116/30 -o vg-abc123 -j SNAT --to-source 10.200.0.2

# Four nat rules. No filter rules. Constant, whether the host holds one
# sandbox or sixteen thousand. The shared rules live in the ROOT namespace:
# one nat MASQUERADE for 10.200.0.0/16 out the WAN interface, two FORWARD
# ACCEPTs, a pool->pool DROP so tenants cannot scan each other, a
# 169.254.0.0/16 DROP so nobody curls the cloud metadata service, and a
# DROP per Stratum mining port. Added once per host, guarded by `-C`.

Look at rule 4 for a second, because it is the post in miniature. That rule does not exist for policy reasons. It exists because conntrack in the root namespace would otherwise see two different sandboxes as the same flow. The hardest thing in this architecture was never the rule set; it was connection tracking state, and the front end has no opinion about it.

The same policy in nftables, which is genuinely nicer

I wrote the nftables version as an exercise and I want to be fair about how it came out: better. Not faster — better to read, and better to change.

#!/usr/sbin/nft -f
# ip netns exec ns-abc123 nft -f /etc/pandastack/slot.nft
#
# The create-then-delete-then-define prelude makes the whole file one
# atomic transaction: the slot's NAT policy is never half-applied.
table ip psnat
delete table ip psnat

define veth_guest = 10.200.0.2
define guest      = 172.20.6.118
define gateway    = 172.20.6.117
define guest_net  = 172.20.6.116/30

table ip psnat {
  # Published ports as DATA, not as a rule each. Adding a port later is
  # `nft add element ip psnat published { 8080 : 172.20.6.118 . 8080 }`
  # -- one small transaction against a hash, no chain position to reason
  # about and no ordering assumption to get wrong.
  map published {
    type inet_service : ipv4_addr . inet_service
    elements = { 22 : $guest . 22 }
  }

  chain prerouting {
    type nat hook prerouting priority dstnat; policy accept;

    # A map lookup that MISSES simply does not match, so this rule and
    # the catch-all below are no longer order-dependent on each other.
    # That is the part I actually envy.
    ip daddr $veth_guest dnat to tcp dport map @published
    ip daddr $veth_guest meta l4proto tcp dnat to $guest
  }

  chain postrouting {
    type nat hook postrouting priority srcnat; policy accept;
    oifname "tap0" snat to $gateway
    ip saddr $guest_net oifname "vg-abc123" snat to $veth_guest
  }
}

Two things improved and one did not. The published-port table became data, so growing it is an element insert rather than a rule insert at a computed position. And the ordering dependency between the explicit DNAT and the catch-all went away, because a map miss does not match — which removes a load-bearing assumption from the code, and removing load-bearing assumptions is most of what senior engineering is. What did not improve: performance. Four sequential comparisons became roughly three plus a hash lookup, on a path where conntrack short-circuits matching after the first packet of each flow anyway. That is not a number you can measure.

Why the linear-traversal argument mostly evaporates per-namespace

The sets-and-maps case is an argument about rule count. It says: if you have ten thousand rules in a chain, you are doing ten thousand comparisons, and a hash lookup is better. That is unarguable. It is also a statement about a shared rule set, because the only way to accumulate ten thousand rules is for many things to write into one namespace.

A namespace-per-workload architecture inverts the growth curve. Our per-namespace rule count is four, plus one per additional published port, and it does not grow with fleet size. The thousandth sandbox on a host has the same four rules the first one had. The root namespace set is thirteen rules and also constant — it is pool-wide, written once, re-asserted idempotently. There is no chain in this system whose length is a function of how many tenants we have, which means there is no chain whose traversal cost we need to amortise with a hash.

The namespaces themselves do cost something — each is a network stack with its own routing tables, conntrack view and sysctls, which is real memory and real setup time. But it is not rule lookup cost, and swapping the front end does not change it by a byte. And if you decide namespaces-per-workload is too expensive, the alternative is one shared namespace with policy keyed on source address — which moves you from four rules times N namespaces to four-times-N rules in one chain. Now the linear chain is your problem and sets are your answer.

So the honest framing is not "nftables is better." It is: the shape of your rule set is a consequence of your isolation architecture, and it determines which front end you need. Pick the isolation model first. The firewall syntax falls out of it.

What actually bites at sixteen thousand namespaces

Here is the part I wish someone had told me, because I spent the front-end question's worth of attention on the wrong layer. The things that break on a dense NAT'd fleet are shared, and most of them are only partially namespaced, which is worse than not namespaced at all because it looks solved.

Connection tracking. Be precise about what is per-namespace here, because the half-truths are dangerous. The conntrack hash table is global — since the netns became part of the hash key, there is one table for the whole host, not one per namespace. `nf_conntrack_buckets` is a single host-wide number you can only change from the initial namespace. The entries come from one shared slab cache, so the memory is host-wide. What is per-namespace is the view: each namespace has its own count, its own expectations, and its own `nf_conntrack_max` sysctl.

That last one is the trap. A per-namespace `nf_conntrack_max` looks like a budget and is not one. Nothing reconciles them. Sixteen thousand namespaces each inheriting a default ceiling sum to several times more entries than the host has RAM for, and nothing is responsible for noticing. The binding constraint is host memory and average hash-bucket chain length, and you discover you crossed it as `nf_conntrack: table full, dropping packet` in `dmesg` — a host-wide event presenting as a few tenants' connections mysteriously failing.

And a NAT'd sandbox fleet is specifically a conntrack-pressure workload, because of double counting. A single guest TCP connection to the outside world is tracked twice: once in the sandbox's own namespace, where the egress SNAT runs, and once in the root namespace, where the pool MASQUERADE runs. Two entries in the global slab per guest flow. Multiply by the number of concurrent sandboxes times their connection fan-out — `npm ci` opens a lot of sockets — and you get a number that arrives much sooner than the rule-count number anyone was worrying about.

#!/bin/sh
# The knobs that actually matter on a dense NAT host, and how to watch
# them before dmesg tells you about them in the past tense.

# Host-wide, read from the initial namespace:
sysctl -n net.netfilter.nf_conntrack_max       # ceiling (per-netns sysctl, global slab)
sysctl -n net.netfilter.nf_conntrack_buckets   # hash buckets -- ONE table for the host
cat /proc/sys/net/netfilter/nf_conntrack_count # entries in THIS namespace

# Rule of thumb: keep max/buckets near 1-4. A high ratio means long hash
# chains, which turns lookup into a linear scan -- the cost people wrongly
# attribute to iptables rule traversal.
sysctl -w net.netfilter.nf_conntrack_buckets=262144
sysctl -w net.netfilter.nf_conntrack_max=1048576

# Timeouts do more for headroom than raising max does. The default
# established timeout of 432000s -- five days -- is a choice from a
# different era. Short-lived sandboxes do not need five-day memory.
sysctl -w net.netfilter.nf_conntrack_tcp_timeout_established=3600
sysctl -w net.netfilter.nf_conntrack_tcp_timeout_time_wait=30
sysctl -w net.netfilter.nf_conntrack_tcp_timeout_close_wait=30
sysctl -w net.netfilter.nf_conntrack_udp_timeout=20

# The neighbour table is the limit nobody budgets for: one entry per veth
# peer in the root namespace. Defaults top out around 1024 -- two orders of
# magnitude below a fully populated slot pool.
sysctl -w net.ipv4.neigh.default.gc_thresh1=8192
sysctl -w net.ipv4.neigh.default.gc_thresh2=32768
sysctl -w net.ipv4.neigh.default.gc_thresh3=65536

# Watch, don't guess. Scrape these; alert on utilisation, not on dmesg.
conntrack -C                                   # current count
conntrack -S                                   # per-CPU insert_failed / drop / early_drop
dmesg -T | grep -i 'nf_conntrack: table full'  # the page you want to prevent
dmesg -T | grep -i 'neighbour table overflow'  # the one you will not expect

The neighbour table is the second one, and it is the one that surprised me. Every veth peer in the root namespace is a neighbour, and the ARP cache garbage-collection thresholds default to numbers chosen for a machine with a handful of interfaces. `neighbour table overflow` in `dmesg` is a fleet-scale symptom that reads like a hardware fault. No firewall syntax has an opinion about it.

NAT tuple space is the third. The root MASQUERADE picks a unique source port per destination, so many sandboxes hammering one well-known endpoint — a package registry, an object store — contend for roughly 64k tuples against it, and the symptom reads as "the registry is rate-limiting us." These are the limits that scale with your fleet, and none of them is a front-end choice.

The one place where front-end cost touches us is the create path, and it is worth being exact because this is where the architecture earned its keep.

Building a slot from nothing is roughly a dozen `ip` invocations — namespace, veth pair, two addresses, a tap, two link-ups, two sysctls — plus four `iptables` invocations, each a process spawn, a namespace entry and a netlink conversation. End to end that is about 500 ms: not because netlink is slow, but because sixteen fork-exec-and-wait cycles with namespace transitions in the middle add up, and the firewall calls are a minority of them.

You cannot pay 500 ms on a create path whose entire budget is 179 ms at p50. So the agent does not. It pre-builds slots and parks them on a free list, and a create claims one. The claim is a MAC patch and a route adjustment — single-digit milliseconds, against roughly 100 ms if you did the namespace plumbing cold. `PANDASTACK_NATID_POOL_SIZE` (default 4) is the warm prebuild depth per template identity, and it is worth saying clearly because people read it as a concurrency cap: it is not. When the free list drains, the agent builds a slot from scratch and eats the ~500 ms. It does not return a 503. The hard ceiling is the `/16` itself — 16,384 sandboxes per agent — and in practice you run out of host RAM long before you run out of subnets.

That is the answer to the netlink-round-trip argument. Pre-allocation moves the entire cost of rule installation off the request path and into a background prewarmer, where it is nobody's latency. An `nft -f` that installs all four rules in one transaction instead of four would shave a few milliseconds off an operation that already happens when no user is waiting. It would be a real improvement to a number that does not appear in any SLO.

Worth knowing if you are copying this design: PandaStack execs `firecracker` directly inside the per-sandbox namespace and does not use the Firecracker jailer. The jailer gives you chroot, cgroup and capability setup for free and is the right default for most people; we do that work ourselves because the restore path needs control the jailer's model does not expose. If you are building from scratch, start with the jailer and only leave it when you can name the reason.

Default-deny is a rule-set shape, and that is where nftables pays

I need to say this plainly because it is the most load-bearing honest disclosure in the post. Egress on PandaStack is open by default. There is no default-deny. A sandbox can reach the internet. What exists is a small set of targeted DROPs in the root namespace's FORWARD chain: pool-to-pool traffic is dropped so tenants cannot scan each other, the entire `169.254.0.0/16` link-local range is dropped so nobody can curl the cloud metadata service and walk away with the host's service-account token, and the well-known Stratum mining-pool ports are dropped because a free tier plus a miner is a recurring and tedious fact of life.

That is a denylist. It is the right default for a platform whose users need `pip install` and `git clone` to work without filing a ticket, and it is the wrong default if you are running genuinely adversarial code and need to know exactly where it can talk.

And here is where the front-end question stops being academic. If you want default-deny, you are not adding rules; you are changing the shape of the rule set. Your policy stops being "four NAT rules plus a handful of DROPs" and becomes "a per-tenant allowlist of destinations, which changes when customers change it." That rule set grows with your tenant count and churns with your product. In iptables that is a chain per tenant and a rule per destination, inserted at positions you have to reason about. In nftables it is one interval set and a verdict map keyed on the source, updated with `nft add element` and replaced atomically.

That is not a marginal win. That is the difference between a feature you can ship and a feature that becomes a source of 3am pages. If default-deny egress is on your roadmap, do not reach for iptables; that is the one place in this entire comparison where I think the answer is unambiguous, and it is the reason I keep the migration plan written down rather than filed under "someday."

What a migration would actually take

So here is the plan, in the order I would actually run it, with the risks named.

  1. Find out what you are really running. `iptables -V` on every host class, including the images your autoscaler creates. If any host says `(legacy)` while others say `(nf_tables)`, you have a fleet inconsistency that will make the migration's canary results meaningless. Fix that first, separately.
  2. Inventory every writer of firewall rules on the host. Container runtime, cloud guest agent, intrusion-prevention tooling, your own agent, anything installed by a compliance package. Each one is a party to the ordering contract. If you cannot enumerate them, you cannot reason about the result, and `nft list ruleset` is the tool that shows you all of them regardless of which front end wrote them.
  3. Migrate the per-namespace rules first. They are the easy half and it is not close. A sandbox namespace is created fresh and destroyed whole, so there is no migration — new slots get an `nft -f`, old slots keep their iptables rules, and both work because they never interact. Gate it behind a flag, prewarm a few nft slots on one host, run the test suite against them, compare create latency, and let it sit for a week.
  4. Leave the root-namespace rules until last, and treat them as a separate project. These are shared, long-lived, and ordering-sensitive in a way that is genuinely load-bearing: the pool-to-pool DROP and the metadata DROP must be evaluated before the egress ACCEPTs or they do nothing at all, which is why our agent inserts them at the top of the chain rather than appending. Re-expressing that in nftables means owning the whole chain at a chosen priority — which is cleaner, and which also means you are now responsible for not shadowing somebody else's rules.
  5. Decide what happens during the overlap, and write it down. Two front ends on one host is not a failure state, but same-hook same-priority chains from two sources have creation-order semantics, and depending on creation order is how you get a rule set that is correct on every host except the one that rebooted in a different sequence. Pick explicit, distinct priorities. Document them next to the code.
  6. Keep a rollback that does not require a reboot. `nft flush ruleset` plus re-running your iptables path should restore the old world in seconds. Test the rollback before you need it, on a host with live sandboxes on it, because the version where you test it on an empty host is the version that works.

What you gain, concretely: atomic replacement of a slot's whole NAT policy; the published-port map replacing a rule per port; one rule set for v4 and v6 when we get around to IPv6; `nft -j` so the agent parses JSON instead of text; and the foundation for per-tenant egress policy as data, which is the actual prize.

What you risk: the PREROUTING ordering assumption, which is explicit and commented in our code and would be silently preserved-or-not by a careless translation; the mixed-front-end ordering surprise; `nft` being absent from a minimal host image at exactly the wrong moment; and a team whose incident reflexes are all in the other syntax. The last one is not a joke. Migration cost is mostly paid in human attention during incidents, months after the code change, by whoever is on call.

The recommendation, without hedging

If you are building a fleet where every workload gets its own network namespace and a handful of rules in it: either front end works, the performance difference is unmeasurable, and iptables' familiarity is worth real money. Use whichever your team can debug half-asleep. Spend the attention you saved on conntrack sizing, neighbour-table thresholds and pre-allocating your namespaces, because those are the things that will actually page you.

If your rules are host-global and numbered in the thousands, or if they grow with your tenant count, or if your policy is default-deny with per-tenant allowlists: nftables is not a preference, it is the answer. Sets and maps are not a nicer syntax for the same thing; they are a different complexity class, and no amount of careful iptables will get you there.

And if you are not sure which you are, the diagnostic is one question: does your rule count grow when you add a customer? If yes, you need nftables and you need it before the growth arrives, because migrating a rule set is much harder than starting with the right one. If no, you have an architecture that made the question small, and the correct move is to notice that and go work on something that matters.

For us the answer is no, so the agent still writes `iptables` and the kernel still runs nftables underneath because that is what the distribution decided. Both sentences will probably still be true next year. The place I expect to change my mind is default-deny egress: the moment that becomes a product feature rather than a thing I say we could do, the rule set changes shape and the comparison flips. I have the migration written down for that day, and I am in no hurry to arrive at it.

Frequently asked questions

Is iptables deprecated? Do I have to migrate to nftables?

Not in the sense most people mean. The `iptables` command is actively maintained and, on virtually every modern distribution, is really `iptables-nft` — a front end that parses iptables syntax and speaks the nftables kernel API underneath. So if you are running Debian 10 or later, Ubuntu 20.04 or later, or a comparable RHEL release, you are already using nftables in the kernel while typing iptables commands. Run `iptables -V`: output ending in `(nf_tables)` confirms it. What is genuinely legacy is the `ip_tables` matcher framework, reachable as `iptables-legacy`, and even that is still in the kernel. So there is no deadline forcing a migration. There is a good reason to migrate — the expressive surface, specifically sets, maps and atomic rule-set replacement — and that reason applies when your rule set is large, shared, or grows with your tenant count. It does not apply much when your rules are few and scoped to a namespace that is created and destroyed as a unit.

Does nftables actually perform better than iptables for a per-sandbox NAT path?

For a per-namespace NAT path with a handful of rules, no — not measurably. The nftables performance argument is about lookup complexity: a set or map gives you a hash or interval lookup instead of walking a chain linearly, which matters enormously when the chain holds thousands of rules. PandaStack's per-namespace rule set is four nat rules and no filter rules, and it stays four no matter how many sandboxes the host is running, because the growth is in the number of namespaces rather than the length of any chain. Going from four sequential comparisons to three plus a hash lookup is not a number you can measure on a path where connection tracking short-circuits matching after the first packet of each flow anyway. The real costs on this path are elsewhere: process spawns during namespace setup, conntrack memory, and neighbour-table pressure. If your rules are host-global and numbered in the thousands, the answer reverses completely and nftables wins on complexity class, not on constant factors.

What breaks if iptables and nftables rules exist on the same host?

Nothing crashes, but you lose the ability to reason about evaluation order, which is worse. The specific hazards are three. First, visibility is asymmetric: `nft list ruleset` shows rules written through `iptables-nft`, but `iptables-save` does not show nft-native rules at all, so an operator reading iptables output during an incident sees a partial picture and concludes a rule is missing. Second, two chains attached to the same hook at the same priority from different sources have creation-order semantics, which means a rule set that is correct on every host can be wrong on the one that rebooted its daemons in a different sequence. Third, `iptables-legacy` attaches to the same netfilter hooks by a different mechanism, so both rule sets evaluate and neither tool shows the other's. The practical mitigations: standardise one front end per host class, assign explicit distinct chain priorities and document them next to the code, and inventory every daemon that writes rules — container runtimes and cloud guest agents usually do.

If rule count is not the limit on a dense NAT'd microVM fleet, what is?

Shared state that is only partly namespaced. Connection tracking is the main one, and the details matter: the conntrack hash table is global for the host, since the namespace is part of the hash key rather than there being one table per namespace, and entries come from one shared slab, so memory exhaustion is a host-wide event. What is per-namespace is the view — each namespace has its own count and its own `nf_conntrack_max`, which looks like a budget but is not, because nothing reconciles thousands of per-namespace ceilings against one pool of RAM. A NAT'd fleet also double-counts: each guest flow to the outside world is tracked once in the sandbox namespace where the egress SNAT runs and once in the root namespace where the pool MASQUERADE runs. Then there is the neighbour table, which holds one entry per veth peer in the root namespace against garbage-collection thresholds that default to around a thousand, and NAT tuple space, where many sandboxes contending for source ports against one popular destination looks exactly like that destination rate-limiting you.

Does PandaStack default-deny sandbox egress?

No, and that is a deliberate default worth stating plainly rather than discovering. Sandbox egress is open: a sandbox can reach the internet, because `pip install` and `git clone` working without a support ticket is table stakes for the product. What exists instead is a small denylist in the root namespace's FORWARD chain — pool-to-pool traffic is dropped so one tenant cannot scan another's guests, the whole `169.254.0.0/16` link-local range is dropped so nobody can reach the cloud metadata service and take the host's service-account credentials, and the well-known Stratum mining-pool ports are dropped because free-tier cryptomining is a recurring nuisance. If you need default-deny with a per-tenant allowlist, you are not adding rules, you are changing the rule set's shape to one that grows with your customer count — and that is the single clearest case in this comparison for nftables, because an allowlist is a typed set you update atomically rather than a chain you re-order.

Keep reading

Related posts

  • Firecracker Networking Explained: TAP, netns, NAT

    The guest thinks it has a normal network card. It has a polite illusion of one. Here's the full path — virtio-net to TAP to a per-sandbox network namespace to NAT'd egress — and why isolating the network per sandbox is the difference between a sandbox and a liability.

  • Best Firecracker Networking Tools & Approaches (2026)

    Firecracker hands you a TAP device and a firm handshake; the rest of the network is your problem. This is a 2026 field guide to the building blocks that solve it — TAP + bridge, per-VM netns + veth + iptables NAT, CNI plugins, vhost-net for throughput, the built-in rate limiter, MMDS for metadata, egress firewalling, and pre-allocated netns pools for fast create — with a 'reach for this when…' per approach.

  • IP Address Planning for a MicroVM Fleet

    Subnet exhaustion is the outage that gives you no warning and then all of it at once. Here's the address math for a microVM fleet, the leak that drained our pool, and the durable allocator that fixed it.

  • IPv6 for Firecracker MicroVM Fleets

    Every microVM networking guide is IPv4, TAP, NAT, a /30 per guest. That works until you run out of address space — or until someone notices your carefully written egress allowlist only speaks one address family. Here's the IPv6 version, including the parts snapshots make awkward.

  • Per-Tenant Fraud Rules in Isolated microVMs

    When customers upload their own scoring rules and those rules run on every checkout, a shared worker pool means one tenant's rule can read another's transaction features — or pin a core and add latency to everyone. Give each tenant's rule its own Firecracker microVM.

More in Internals · See PandaStack benchmarks

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.