virtio-net Offloads and Why Your microVM's Network Is Slower Than the Wire
The report that started this was two sentences long. "Bulk upload out of the sandbox is fine. Our RPC benchmark gets about a third of what the same binary does on a bare EC2 box." Both halves were true, both were measured correctly, and the gap between them is the entire subject of this post. The sandbox was moving large payloads at a rate nobody complained about, and dying on a test that did sixty-four bytes out, sixty-four bytes back, as fast as it could, forever.
I build PandaStack, an open-source Firecracker microVM platform. Our guests reach the network through a virtio-net device, a host TAP file descriptor, a per-sandbox network namespace, a veth pair and a NAT. Every one of those hops has an opinion about packet size, and the features we call "offloads" are the mechanism by which those opinions get reconciled. Get them right and bulk throughput is a non-issue. Get them subtly wrong and you get the signature failure: throughput looks fine, and the host's CPU is on fire.
The path a packet takes to leave a guest
Start with the full itinerary, because almost every wrong intuition about microVM networking comes from imagining a shorter one. A packet leaving a Firecracker guest goes through this, in order:
- The guest's `virtio_net` driver places a descriptor on the transmit virtqueue. The descriptor points at guest memory; the payload is not copied here.
- The driver kicks the device by writing to its notification register. That write traps out of the guest -- a VM exit -- and KVM turns it into a signal on an eventfd the VMM registered in advance.
- Firecracker's net device thread wakes on that eventfd, reads the descriptor chain, and writes the frame into the TAP file descriptor. This is where the bytes get copied out of guest memory and into a host socket buffer.
- The host kernel receives the frame on `tap0`, which lives inside that sandbox's own network namespace.
- It traverses that namespace's routing table and its netfilter chains -- for us, a wildcard DNAT rule set on the way in and an SNAT on the way out.
- It crosses the veth pair (`vg-<id>` to `vh-<id>`) into the root namespace, which is a queue handoff rather than another copy.
- In the root namespace it hits the shared chains: the cross-tenant DROP, the metadata-service DROP, the blocked-egress-port DROPs, then the pool-wide MASQUERADE.
- Finally it is handed to the physical NIC, which DMAs it onto the wire.
Two things in that list cost money. The first is the notification in step 2, which is a boundary crossing with a fixed price regardless of how many bytes it carries. The second is steps 4 through 7, which are per-packet host kernel work -- an skb, a routing lookup, two netfilter traversals, two NAT translations, two conntrack lookups -- also largely independent of packet size. Add them up and you get a fixed per-packet tax. A 64 KiB transfer amortises it to nothing. A 64-byte request pays it in full, twice, once in each direction. That is the whole of the customer's bug report, and no tuning flag makes it go away.
The copies deserve precision, because "virtualised networking does lots of copies" is folklore that survives mostly because nobody counts. On transmit there is essentially one real copy, at the TAP boundary in step 3. On receive there is one in the other direction, from the host skb into the buffers the guest posted. The veth crossing does not copy. What is genuinely expensive is not the copying, it is the per-packet bookkeeping and the notification -- and that is exactly what offloads attack.
Nothing here is offloaded to silicon
The word "offload" comes from physical NICs, where it means something literal: a chip on the card does the work, and the host CPU genuinely does not. Carry that meaning into a VM and you will mis-model your own system, because inside a guest none of these features move work onto hardware. They are a negotiation between the guest driver and a virtual device about who does the segmentation and when.
TSO -- TCP segmentation offload -- means the guest hands the device a single buffer of up to 64 KiB with a note attached saying "this is really a TCP stream, the MSS is 1448, please cut it up." The guest pays one descriptor, one notification, one trip across the boundary for forty-odd packets' worth of payload. The segmentation still happens. Someone still runs it. It just happens later, in bigger batches, on a different CPU, on the other side of a boundary your guest's own CPU accounting cannot see. GSO is the same trick confined to the guest's own stack: the super-frame stays whole through the layers above the driver and is only segmented at the last moment, which saves guest stack traversals even when the device offers no offload at all. On receive, GRO and LRO are the mirror image -- many wire frames coalesced into one large buffer handed up as a unit. Checksum offload is the smallest and most literal of the family: the guest leaves a field blank and records, in the virtio header, where the blank is and what range it covers.
That virtio header is the entire mechanism, and it is worth keeping in your head as a concrete object rather than an abstraction. It is a small struct in front of every frame -- flags, a GSO type, a header length, a segment size, a checksum start and offset, and, with merged receive buffers, a count of how many descriptors the frame spans. "Offloading" is filling in those fields instead of doing the work. The work is deferred, not deleted.
There is exactly one place in this path where real hardware offload can still happen, and it is the last hop: if the large buffer survives all the way to the physical NIC with its GSO metadata intact, the NIC's own TSO engine can do the segmentation in silicon. Whether it survives depends on the MTU of every hop and the capability of every device in between. That is why MTU mismatches are not a separate topic from offloads. An MTU mismatch is the thing that forces the kernel to segment in software, in softirq context, on your host, once per frame, for every sandbox on the box.
| Feature | What work moves | Who pays the CPU | Helps bulk | Helps small packets |
|---|---|---|---|---|
| TX checksum (`VIRTIO_NET_F_CSUM`) | Guest leaves the L4 checksum blank; the virtio header says where | Host kernel on the forward path, or the NIC if the frame reaches it intact | Yes, modestly | Barely -- one checksum over 64 bytes is not your problem |
| RX checksum (`VIRTIO_NET_F_GUEST_CSUM`) | Device may hand up frames the guest has not verified | Nobody, if the host already validated it on the way in | Yes | Marginal |
| Guest TX TSO (`HOST_TSO4`/`TSO6`) | Segmentation of up to a 64 KiB buffer leaves the guest entirely | Host device thread and host softirq -- or genuinely the NIC | Yes -- the single biggest lever there is | No. There is nothing to segment |
| Guest RX GRO (`GUEST_TSO4`/`TSO6`) | Coalescing many wire frames into one big buffer, done host-side | Host receive softirq | Yes | No, and a coalescer that waits costs you latency |
| Merged RX buffers (`MRG_RXBUF`) | Nothing; it lets one frame span several descriptors | Neither side -- it is a buffer layout change | Enables the row above | Neutral |
| In-guest GSO (software, no feature bit) | Segmentation moves to the last layer before the driver | Guest vCPU, but in fewer and bigger steps | Yes, even with no device offload at all | No |
Feature bits, and why you must read ethtool instead of a blog post
virtio is a negotiation protocol. The device advertises a bitmap of features it can support; the driver reads it, selects a subset it also supports, and writes that subset back. The intersection is what you get. Neither side's wish list is observable from the other side afterwards -- only the result is.
- `VIRTIO_NET_F_CSUM` (bit 0): the DEVICE can accept a frame with a partial checksum. This is what lets the guest skip computing TCP and UDP checksums on transmit.
- `VIRTIO_NET_F_GUEST_CSUM` (bit 1): the DRIVER can accept a frame with a partial checksum. The receive-side mirror.
- `VIRTIO_NET_F_GUEST_TSO4` / `GUEST_TSO6` (bits 7 and 8): the DRIVER can receive buffers larger than the MTU, carrying GSO metadata. This is what permits host-side GRO to be delivered as one large frame rather than being re-segmented first.
- `VIRTIO_NET_F_HOST_TSO4` / `HOST_TSO6` (bits 11 and 12): the DEVICE can accept oversized buffers with GSO metadata. This is guest TSO on transmit, split by address family.
- `VIRTIO_NET_F_MRG_RXBUF` (bit 15): the device may satisfy one received frame from several chained descriptors. Without it, receiving a 64 KiB coalesced frame would mean posting 64 KiB per descriptor, so in practice this bit is the precondition for large receive.
Notice the naming trap, because it has cost people whole afternoons: HOST_TSO is the feature that makes the GUEST's transmit fast, and GUEST_TSO is the feature that makes the guest's receive fast. The name describes who performs the segmentation, not who benefits. Read every one of these as "X is allowed to hand the other side something oversized."
The exact set a given Firecracker version advertises changes between releases, and I am not going to pretend otherwise by printing a list here. The cautionary tale is UFO, UDP fragmentation offload: it was in the virtio spec, host kernels stopped supporting it, and the feature bit went away from VMMs that had offered it for years. Anyone who had written "UFO is available" in a runbook was wrong without having changed anything. So: run `ethtool -k eth0` inside your own guest, on your own build, and believe that. We run Firecracker v1.16 with a 5.10 guest kernel on an Ubuntu 24.04 rootfs, and I still check.
# ---------------------------------------------------------------------------
# PART 1 -- IN THE GUEST. What was actually negotiated, not what you hoped.
# ---------------------------------------------------------------------------
ethtool -i eth0 # driver: virtio_net. If it says anything else, stop.
ethtool -k eth0 # THE answer. Everything below is interpretation.
# Read the output like this:
# rx-checksumming / tx-checksumming -> VIRTIO_NET_F_GUEST_CSUM / _F_CSUM
# tcp-segmentation-offload -> VIRTIO_NET_F_HOST_TSO4 / _TSO6
# tx-tcp-segmentation -> the v4 half
# tx-tcp6-segmentation -> the v6 half (often absent)
# generic-receive-offload (GRO) -> coalescing; needs GUEST_TSO4 + MRG_RXBUF
# to arrive as one big buffer
# large-receive-offload -> usually "off [fixed]" on virtio-net
# generic-segmentation-offload (GSO) -> pure guest software, no feature bit
# "[fixed]" means the driver will not let you change it. That is a negotiation
# outcome, not a bug, and `ethtool -K` will refuse you.
# The feature bitmap itself, straight from the virtio bus. LSB first, one
# character per bit. Cross-reference against the virtio spec numbers:
# 0 CSUM 1 GUEST_CSUM 7 GUEST_TSO4 8 GUEST_TSO6
# 11 HOST_TSO4 12 HOST_TSO6 15 MRG_RXBUF
cat /sys/class/net/eth0/device/features
cat /sys/class/net/eth0/mtu
# ---------------------------------------------------------------------------
# PART 2 -- ON THE HOST. The tap has its own opinion, and it wins ties.
# ---------------------------------------------------------------------------
SBX=4f1c9a7e-... # sandbox id
NS=ns-$SBX # per-sandbox network namespace
PID=$(pgrep -f "fc-$SBX.socket") # the Firecracker process for that VM
ip netns exec "$NS" ethtool -k tap0 | grep -E 'segmentation|checksum'
ip netns exec "$NS" ip -d link show tap0 # look for vnet_hdr and the MTU
ip netns exec "$NS" ip -s link show tap0 # drops here are not the guest's
# The VMM, not you, is the authority: it opened the tap with IFF_VNET_HDR and
# set TUN_F_CSUM/TUN_F_TSO4/... with a TUNSETOFFLOAD ioctl when the device was
# created. `ethtool -k` on the tap REPORTS that decision. `ethtool -K` on a
# tap under a running VM is how you manufacture a disagreement -- the guest
# still believes TSO was negotiated, so it keeps sending 64 KiB buffers, and
# the host starts segmenting every one of them in software.
# ---------------------------------------------------------------------------
# PART 3 -- THE MTU WALK. Every hop, in order, in one screen.
# ---------------------------------------------------------------------------
ip netns exec "$NS" ip -br link show # tap0 and vg-$SBX
ip -br link show # vh-$SBX and the WAN iface
ip route get 1.1.1.1 | head -1 # which iface really carries egress
# Then stop reading numbers and test the path end to end, from the guest:
# 1472 payload + 20 IP + 8 ICMP = 1500. -M do forbids fragmentation, so a
# hop with a smaller MTU must reply "frag needed" or silently black-hole you.
ping -M do -s 1472 -c 3 1.1.1.1
ping -M do -s 1372 -c 3 1.1.1.1 # succeeds where 1472 fails => ~100 B of
# tunnel overhead somewhere in the path
tracepath -n 1.1.1.1 # prints the MTU it discovers per hop
The TAP-side half everybody forgets
The guest's view is only one end of the agreement. A TAP device has its own offload configuration, set by whoever opened it -- the VMM, with an ioctl, at device-creation time. The VMM also has to open the TAP with the vnet-header flag, or the virtio header carrying all that GSO metadata has nowhere to travel and the whole scheme collapses to one-packet-per-descriptor.
When the two ends agree, large buffers flow through. When they disagree, nothing breaks loudly, which is the problem. The guest believes TSO was negotiated, so it keeps handing over 64 KiB buffers. The host accepts them, discovers that the next hop cannot carry a frame that size, and calls into the software segmentation path, once per buffer, in softirq context. Your throughput graph barely moves. Your host CPU graph changes shape. This is the canonical "throughput is fine, CPU is on fire" failure, and on a dense host it shows up first as unrelated sandboxes getting slower.
MTU is the other half of the same coin, and it is the one that produces the most confusing bug reports. The chain is guest `eth0`, then `tap0`, then `vg-<id>`, then `vh-<id>`, then the host's WAN interface -- and if your host's egress sits inside any kind of tunnel, the real path MTU is 1500 minus the encapsulation overhead. Get one hop wrong by fifty bytes and you get a network that works perfectly for everything except the one large request: DNS resolves, health checks pass, `git clone` of a small repo succeeds, and a large POST or a TLS handshake with a fat certificate chain hangs until it times out. The diagnostic is not reading MTU values -- it is `ping -M do` with a payload that fills the frame, from inside the guest, which turns a silent black hole into an immediate failure. There is a whole post on that failure mode at Firecracker Guest MTU and Network Tuning Explained.
Measure both halves, or you have measured nothing
Here is the discipline, and it is the most portable thing in this post. A throughput number without a CPU number beside it is not a measurement, it is an anecdote with a unit attached. And a single throughput number is not even one measurement, because there are two independent questions and they have different answers.
The bulk question is bytes per notification. You answer it with `iperf3`, in both directions separately, because transmit and receive use different halves of the feature set and it is entirely normal for one to be healthy and the other not. Offloads dominate this number. If it is bad, the fault is in the feature set or the MTU, and the previous two sections are your checklist.
The small-packet question is notifications per second. `iperf3` cannot answer it -- it measures bytes, and you are asking about round trips. You need a request/response test: one tiny request, one tiny reply, strictly serialised, so the result is the reciprocal of your latency and nothing else. `netperf -t TCP_RR` is the classic; `sockperf ping-pong` does the same job. The VM exit and the host forwarding path dominate this number, and offloads do approximately nothing for it, because there is nothing large to batch.
Then attribute the CPU, in two places people routinely miss. On the host, the network cost of a Firecracker VM lands on its net device thread, not spread across the process, so you need per-thread accounting -- a per-process number averages the device thread's work across the vCPU threads and hides it. And the forwarding half -- routing, netfilter, NAT, conntrack, any software segmentation you accidentally requested -- lands in softirq context, which belongs to no process at all. If you only watch the VMM process you will conclude the network is free. It is not free; it is just billed to a different line.
#!/usr/bin/env bash
# The honest benchmark: two different questions, and the CPU bill for both.
# A throughput number with no CPU number beside it is an anecdote.
set -euo pipefail
SBX=4f1c9a7e-...
NS=ns-$SBX
PID=$(pgrep -f "fc-$SBX.socket")
TARGET=10.0.0.9 # a machine whose OWN capacity you have verified
# --- ON THE HOST: start the accounting before any traffic moves -------------
grep -E 'NET_RX|NET_TX' /proc/softirqs > /tmp/si.before
# Per-THREAD CPU for the VMM. -t matters: Firecracker runs one device thread
# per virtio device plus one thread per vCPU, and the network cost lands on
# the device thread, which is invisible in a per-process number.
pidstat -t -p "$PID" 1 70 > /tmp/vmm.cpu &
# Host-wide: %soft is the forwarding path -- routing, netfilter, veth, and any
# software segmentation you accidentally asked for.
mpstat -P ALL 1 70 > /tmp/host.cpu &
# If your host exposes the KVM tracepoints, count the exits too. This is the
# notification bill in its rawest form.
perf stat -e kvm:kvm_exit -p "$PID" -- sleep 70 2> /tmp/exits.txt &
# --- IN THE GUEST, test 1: BULK. Bytes per notification. ---------------------
# This is the question offloads answer. Expect TX and RX to differ; they use
# different halves of the feature set.
iperf3 -c "$TARGET" -t 30 --format m # guest -> target (TSO path)
sleep 2
iperf3 -c "$TARGET" -t 30 -R --format m # target -> guest (GRO path)
# --- IN THE GUEST, test 2: SMALL PACKET. Notifications per second. -----------
# iperf3 cannot answer this: it measures bytes, and you are asking about round
# trips. Use a request/response test -- one tiny request, one tiny reply,
# strictly serialised, so the number you get is 1/latency and nothing else.
netperf -H "$TARGET" -t TCP_RR -l 30 -- -r 64,64
# Equivalent, if you prefer it:
# sockperf ping-pong --tcp -i "$TARGET" -t 30 -m 64
# Do NOT run this concurrently with the bulk test. You will measure neither.
wait
# --- ON THE HOST: read the bill --------------------------------------------
grep -E 'NET_RX|NET_TX' /proc/softirqs > /tmp/si.after
diff /tmp/si.before /tmp/si.after # softirqs burned per unit of traffic
awk '/firecracker|fc_/ {print}' /tmp/vmm.cpu | tail -20
cat /tmp/exits.txt
# What you are looking for:
# * bulk fine, RR terrible -> per-packet path. Exits, rules, conntrack.
# * bulk terrible, RR fine -> offloads or MTU. Go back to PART 2 above.
# * both fine, host %soft high -> you are paying for someone, probably in
# software segmentation. Nobody will bill you for it until the host melts.
One more measurement rule, learned the embarrassing way: verify the target. Half of the "microVM networking is slow" results I have been sent were a correct measurement of something else -- a t3.micro with a burst credit balance, a load generator sharing a core with the thing it was generating load for, or an endpoint behind a CDN that was rate-limiting a single source IP. Point your test at something whose own capacity you have confirmed independently, from a machine that is not the one under test.
The version of this I actually keep is a script, because the answer changes when an image is rebuilt and I would rather a CI job notice than a customer. It walks every template we ship, prints what each one negotiated, and runs both halves of the benchmark against a fixed target.
# Inventory the negotiated network offloads across every template we ship.
# "The platform supports TSO" is not a fact about a platform -- it is a fact
# about a guest kernel, a rootfs and a Firecracker version, and ours are baked
# into images that get rebuilt. So re-measure, in CI, on the images you run.
import shlex
from pandastack import Sandbox
PROBE = r"""
set -eu
echo "mtu=$(cat /sys/class/net/eth0/mtu)"
echo "features=$(cat /sys/class/net/eth0/device/features)"
ethtool -k eth0 | grep -E 'checksumming|segmentation-offload|receive-offload'
"""
# Note what is NOT in this script: a sandbox-to-sandbox iperf3. On our fleet
# the host FORWARD chain drops pool-to-pool traffic outright -- a tenant must
# never reach a neighbour's guest -- so the only honest targets are external
# endpoints you control.
TARGET = "10.0.0.9"
for tmpl in ("base", "code-interpreter", "agent", "browser"):
# cpu=/memory_mb= would be ignored here: Firecracker cannot change vCPU or
# RAM at snapshot restore, so the baked meta.json governs the size. base is
# 4 GiB / 8 vCPU, code-interpreter and agent are 2 GiB / 8 vCPU.
sbx = Sandbox.create(
template=tmpl,
ttl_seconds=600, # IDLE timeout, not a wall clock. A
# sandbox mid-iperf3 is not idle.
metadata={"probe": "net-offloads"},
)
try:
print("=" * 20, tmpl)
r = sbx.exec("sh -c " + shlex.quote(PROBE))
print(r.stdout.strip() or r.stderr.strip())
# Bound long commands IN THE GUEST with timeout(1). One-shot exec()
# does not reliably enforce a timeout_seconds argument -- do not pass
# one and assume you are covered.
bulk = sbx.exec(
f"timeout --kill-after=5s 60 iperf3 -c {TARGET} -t 20 --format m"
)
print(bulk.stdout.strip()[-400:] or bulk.stderr.strip())
# And the other half, which is the one people skip.
rr = sbx.exec(
f"timeout --kill-after=5s 60 netperf -H {TARGET} -t TCP_RR -l 20"
" -- -r 64,64"
)
print(rr.stdout.strip()[-400:] or rr.stderr.strip())
finally:
sbx.kill() # teardown is kill(); there is no delete()
Why this becomes a fleet property
On a single VM, per-packet cost is a curiosity. On a host running many sandboxes it is an architectural property, and this is the part that makes microVM networking genuinely different from tuning one virtual machine.
Our agents pre-allocate 16,384 /30 subnets out of 10.200.0.0/16, each with its own network namespace, veth pair and TAP device. Building a namespace, a veth pair, a TAP and the iptables rules from cold costs on the order of a hundred milliseconds, which is most of a create budget we measure in hundreds of milliseconds end to end; pre-allocating takes that down to a few milliseconds of warm work, mostly patching a MAC address. That is a real win and it is worth being clear about what it does not do: it removes the setup cost from create and has precisely no effect on the steady-state per-packet cost. The namespace is cheaper to make. It is not cheaper to route through.
And routing through it is per-sandbox work that aggregates. Every packet from every guest traverses its own namespace's chains -- our DNAT rules inbound, an SNAT outbound -- and then the shared root chains, where it is matched against the cross-tenant DROP, the metadata-service DROP, the blocked-egress-port DROPs and the pool-wide MASQUERADE. Three consequences follow, and all three are counter-intuitive if you are used to thinking per-VM:
- A rule is not a one-time cost, it is a multiplier. Netfilter traversal is linear in the rules examined before a verdict, so a rule set that is obviously fine for one sandbox is a per-packet tax paid by every sandbox on the box. The fix is structural -- jump to a chain, or match against a set, so the common case is a hash lookup rather than a walk -- not "be sparing with rules."
- Two NATs means two conntrack lookups, and a conntrack entry per flow in each namespace it transits. Conntrack accounting is per-namespace; the hash table underneath is not partitioned the way the namespaces suggest, so one guest opening ten thousand connections is a fleet event. Conntrack Table Limits: The Shared Resource Under Your MicroVM Fleet is the whole story.
- Softirq CPU is a host-global resource with no tenant label on it. A guest that generates a million small packets per second is spending host CPU that no per-sandbox accounting attributes to it, and the symptom lands on its neighbours.
This is also why the "offloads help bulk, not small packets" distinction matters more to us than it would to someone running one VM. Offloads reduce the number of times the per-packet path runs. On a fleet, the per-packet path is the shared resource. Negotiating TSO properly is not a throughput optimisation; it is a density optimisation that happens to show up on a throughput graph.
Where a microVM behind a TAP is the wrong substrate
Plainly, because I would rather you find this out here than in week three. We do not do SR-IOV, and we do not do PCI device passthrough of any kind. There is no virtual function you can attach, no way to hand a guest a slice of a real NIC, and therefore no way to bypass the TAP and the host forwarding path. That is a deliberate consequence of the isolation model -- passthrough means guest-controlled DMA and a much larger blast radius -- but it is still a ceiling, and it is a ceiling you cannot buy your way past on this architecture. No amount of money moves it.
The guest kernel is 5.10, which bounds what you can do above the device as well. Features that raise the segmentation ceiling above 64 KiB arrived in considerably later kernels, so on our guests 64 KiB per super-frame is the arithmetic, full stop. We also have no GPUs and no GPU passthrough, which is not a networking limit but is the other thing people ask for in the same breath.
So: if your workload is a line-rate packet processor -- a DPDK or XDP pipeline, a software router, a load balancer whose job is forwarding rather than computing, anything where packets per second IS the product -- a microVM behind a TAP is the wrong substrate and no tuning in this post will save you. Buy bare metal with a real NIC and a kernel bypass. If instead your workload is agents, builds, tests, scrapers, code execution and request/response services, where the network is a means and not the end, then this path is fine and the only things you need from this post are: negotiate the offloads, walk the MTU, and measure the CPU next to the throughput.
One closing note on expectations. The reason the per-packet cost is a fixed overhead rather than a tunable is that it buys something: the boundary that costs you a VM exit is the same boundary that means a tenant's kernel bug is a tenant's problem. If you want the full picture of why we keep virtio-net in VMM userspace rather than pushing it into the host kernel for speed, that trade is written up at Firecracker Doesn't Use vhost-net (On Purpose), and the queue-level mechanics are at Firecracker virtio-net RX/TX Queues Explained. Nothing in this post is offloaded to silicon. It is all just someone else's CPU, later, in bigger batches -- and on a fleet, "someone else" is you.
Frequently asked questions
Should I turn TSO and GSO off to debug a network problem in a microVM?
As a bisect, yes, and it is a good one -- but know what you are actually changing, and change it back. Disabling segmentation offload in the guest multiplies your boundary crossings by roughly the ratio of your super-frame size to your MTU, so a forty-fold increase in notifications and per-packet host work is normal. That is why "I disabled TSO to rule it out and everything got dramatically slower" is such a common story: the slowdown is the expected result of the experiment, not a second bug. What the experiment is genuinely good for is distinguishing a correctness problem from a performance problem. If a transfer that hangs or corrupts starts working with offloads off, you have learned something real: somewhere in the path a large frame is being mishandled, which usually means an MTU mismatch or a disagreement between the TAP's offload configuration and the guest's negotiated features. If instead the transfer still fails, offloads were never the variable and you should stop touching them. The one thing not to do is leave them off in a template because it made a symptom go away. You will have traded a visible bug for a permanent tax on every packet every sandbox ever sends, and the next person to look at the host's CPU graph will have no idea why it looks like that.
Why is my bulk throughput fine while my small-request benchmark is terrible?
Because they are measurements of two different quantities, and only one of them is improved by anything in the offload family. Bulk throughput is bytes per notification: with segmentation offload negotiated, one descriptor and one boundary crossing can carry up to 64 KiB, so the fixed per-crossing cost is amortised almost to zero and you end up measuring the wire and the host's forwarding capacity. A small request/response test is notifications per second: each tiny request is its own descriptor, its own notification, its own VM exit, its own skb allocation, its own routing lookup, its own two netfilter traversals and its own conntrack lookups -- then the same again for the reply. There is nothing to batch, so offloads cannot help, and the fixed overhead that bulk traffic hides is now the entire number. This is structural, not a misconfiguration, and the fix is not in the platform. It is in the application: fewer round trips, pipelining, keepalive and connection reuse, batching at the protocol level, and moving chatty exchanges inside the guest rather than across its network boundary. If your service is genuinely latency-bound on tiny messages and you cannot batch, measure it honestly against bare metal before you commit, because a microVM's network boundary will cost you something and you want to know what before it is in production.
Does Firecracker use vhost-net, and would it be faster if it did?
It does not, and yes, it probably would be -- that is the trade, made deliberately. Firecracker implements virtio-net in the VMM's own userspace: a device thread reads descriptors out of guest memory and writes frames to a TAP file descriptor. The kernel's vhost-net alternative moves that loop into the host kernel, which removes a userspace round trip per batch and generally does improve throughput and latency. The cost is attack surface in the worst possible place. A userspace device model means a bug reachable from guest input lands inside a process you can confine with seccomp, run unprivileged, and put a jailer around; the same bug in a kernel data path is a host kernel bug. For a platform whose entire proposition is running untrusted code from strangers, that is not a close call, and it is the same reasoning behind the deliberately tiny device model -- a serial port, virtio-net, virtio-block, vsock, a balloon and an RNG, and no HPET or emulated RTC. The practical upshot for you: the per-packet overhead described in this post is not a bug waiting to be optimised away, it is a property you are buying on purpose. Verify the current device and feature details against Firecracker's own documentation for your version rather than trusting a summary, including this one.
How many netfilter rules is too many on a dense microVM host?
The honest answer is that the count is the wrong unit, and that is more useful than a number would have been. Netfilter cost is proportional to how many rules a packet is examined against before something returns a verdict, so ten rules that all match early can be cheaper than four rules whose match conditions are expensive, and a hundred rules reached through a jump that skips ninety-five of them can be cheaper still. What you want is structural: hoist the common case to the front, use a jump into a per-class chain so irrelevant rules are never examined, and replace long runs of near-identical rules with a single set lookup, which is a hash rather than a walk. Our root chains are short on purpose for exactly this reason -- a handful of security DROPs and one pool-wide masquerade -- and each one earns its place because it is paid for by every packet from every sandbox on the host. As for measuring rather than reasoning: the number to watch is softirq CPU per unit of traffic, not rule count. Capture the NET_RX and NET_TX softirq counters before and after a known traffic pattern, change the rule set, and run exactly the same pattern again. If the softirq bill per megabyte moved, you changed something real. If it did not, you have a tidier rule set and the same performance, which is still worth having.
Can I get a dedicated NIC, SR-IOV or a virtual function for a sandbox?
No, and there is no roadmap item that changes that answer, so plan around it rather than waiting. PCI device passthrough -- which is what SR-IOV virtual functions require -- means giving a guest direct, DMA-capable access to a piece of real hardware. That is a fundamentally different isolation posture from the one we sell: it widens the boundary from "a small userspace device model plus KVM" to "a device, its driver, an IOMMU configuration and whatever firmware is on the card," and it ties a sandbox to a specific physical host in a way that breaks snapshot-restore and fork entirely. Everything we are good at -- creating a sandbox in a couple of hundred milliseconds by restoring a snapshot, forking one into several, moving a workload to whichever host has capacity -- depends on the guest's hardware view being synthetic and identical everywhere. A passed-through NIC is the opposite of that. What you can do instead is reduce how much traffic has to cross the boundary at all: keep chatty components inside one sandbox, use vsock rather than TCP for host-to-guest control traffic, batch at the protocol level, and make sure the offloads you do have are actually negotiated so that the bulk path is paying the fixed cost once per 64 KiB rather than once per 1500 bytes. If after all that packets per second is still your product, the right answer is bare metal with a real NIC, and I would rather tell you that than sell you a ceiling.
Keep reading
- virtio-net RX/TX queue tuning — The queue-level mechanics under this post: ring depth, one queue pair, and what is genuinely tunable.
- Firecracker doesn't use vhost-net (on purpose) — Why the device model stays in VMM userspace, and what that choice costs you per packet.
- Guest MTU and network tuning — The failure where small requests work and one large response hangs forever.
- Conntrack table limits under a microVM fleet — The shared, fixed-size table that two NATs per packet keep writing into.
- Firecracker networking: TAP, netns, NAT — Start here if the data path in section one was new to you.
Related posts
- The I/O Scheduler Inside a MicroVM Is Almost Always Wrong
Your guest's block layer is sorting requests by sector number to optimise seeks on a platter that is a file on somebody else's filesystem. Why `none` is the answer, why it still merges, and why readahead is the knob that actually moves the number.
- Guest Writeback Throttling: Why dirty_ratio Is Wrong for a 2 GiB MicroVM
The guest is extremely confident about a disk that does not exist. The host is also confident, about different things. Between those two confidences live your tail latencies.
- Your benchmark ran at a different clock speed than production
The same core does not run at the same speed twice. Governor, turbo bin, how many neighbours are busy, thermal headroom and ramp latency all move it — which is why the first sandbox on a quiet host looks fast and the fiftieth on a busy one gets blamed on the platform.
- Best Firecracker Networking Tools & Approaches (2026)
Firecracker hands you a TAP device and a firm handshake; the rest of the network is your problem. This is a 2026 field guide to the building blocks that solve it — TAP + bridge, per-VM netns + veth + iptables NAT, CNI plugins, vhost-net for throughput, the built-in rate limiter, MMDS for metadata, egress firewalling, and pre-allocated netns pools for fast create — with a 'reach for this when…' per approach.
- Interrupts, IRQs, and Where microVM Tail Latency Comes From
Median latency tells you the machine works. p99 tells you how the machine is built. Here is every handoff a virtio interrupt makes inside a Firecracker guest, which of those handoffs are queues, and which ones you can actually do something about.
More in Internals · See PandaStack benchmarks
49ms p50 cold start. Fork, snapshot, and scale to zero.