all posts

Shadow Traffic and Replay: Testing a Rewrite Against Real Requests

Ajay Kumar··11 min read

You are replacing a service. Not refactoring it — replacing it, in a different language or a different framework or a different data model, because the old one has reached the end of whatever it was. And the specification for the new one is a single sentence that nobody will write down in the ticket: whatever the old one does.

Including the bugs. Especially the bugs. Somewhere out there is a client that depends on your endpoint returning `200` with an empty array where a `404` would have been correct, and a mobile app three versions back that parses the typo in your error string, and a nightly batch job at a customer you have never spoken to that works entirely by accident. None of that is in your test suite. Most of it is not in anyone's head. It is only in one place: the running code, and the traffic that proves what the running code does.

I'm Ajay; I built PandaStack, which runs code in Firecracker microVMs. This post is about using production traffic as the test oracle for a rewrite — mirroring it or replaying it at two implementations, diffing the responses, and dealing with the fact that half those requests were trying to change something.

The old service is the spec, and the traffic is the only proof

The useful framing is that you have no test oracle. A unit test is an oracle you wrote; it encodes what you believed the behaviour was on the day you wrote it. For a rewrite, the thing you need to preserve is the behaviour you do not believe in and have never articulated. There is exactly one artefact that has never been wrong about what the old service does, and it is the old service.

So: run both implementations against the same request and compare the answers. This is an old pattern with two named ancestors worth reading, because both of them encode a warning in their design.

GitHub's `scientist` describes itself as "A Ruby library for carefully refactoring critical paths." You wrap the old path in a `use` block and the new one in a `try` block; it randomises the order they run in, always returns the control's value to the caller, swallows and records any exception the candidate raised, and publishes the mismatches. The design is careful in a specific direction — the candidate can be completely broken and your users cannot tell. And the README does not hedge about the limit: Scientist is only safe for wrapping methods that aren't changing data. Its advice for writes is not "use an experiment"; it is to write to both systems wherever writes happen and verify at read time.

Twitter's `diffy` — "Find potential bugs in your services with Diffy" — attacks the harder half, which is not duplication but comparison. It runs three instances of your service, not two: the candidate, a primary running your last known-good code, and a secondary running the same known-good code as the primary. The secondary is there because responses are not deterministic and you need to know how non-deterministic before you can read a diff. Rather than filtering noise, it compares disagreement rates: primary against secondary tells you the ambient disagreement, primary against candidate tells you the total, and only the excess is a finding. For safety, `POST`, `PUT` and `DELETE` are ignored by default and you have to opt into side effects with a flag.

Both ancestors independently landed on the same two conclusions: your control needs a control group, and the write path is not covered. Every design decision in the rest of this post is a consequence of one of those two.

Mirror or replay: the decision that shapes everything downstream

There are two ways to get the same request into two implementations, and they are not variants of one technique. They have different failure modes, different compliance stories, and different answers to "can you re-run that?"

Mirroring: duplicate the live request

A proxy in front of the primary copies each request to a second cluster in real time. In Envoy that is `request_mirror_policies` on a route; in nginx it is `ngx_http_mirror_module`. The input is perfectly fresh — it is literally this second's traffic, referencing rows that exist right now — and you store nothing, which makes the privacy conversation enormously easier.

The thing people discover a week in is that both implementations are fire-and-forget by design. Envoy's documentation says so in as many words: it "will not wait for the shadow cluster to respond before returning the response from the primary cluster." nginx's is shorter: "Responses to mirror subrequests are ignored." The proxy gives you duplication. The comparison is yours to build, and it is the part with all the engineering in it.

# Envoy: duplicate a slice of GET traffic to the candidate. The primary's
# response is what the client gets; the shadow's response is thrown away.
#
# Straight from the RequestMirrorPolicy docs: "The current implementation is
# 'fire and forget,' meaning Envoy will not wait for the shadow cluster to
# respond before returning the response from the primary cluster."
#
# Read that twice before you plan a diff around it. Envoy hands you the
# TRAFFIC. It does not hand you the COMPARISON -- nothing here records either
# response. You build that part yourself, and it is most of the work.
route_config:
  name: local_route
  virtual_hosts:
    - name: orders
      domains: ["*"]
      routes:
        # Mutating methods get NO mirror policy. This is the cheap, boring,
        # safe tier, and it is where you start -- not where you finish.
        - match:
            prefix: "/v1/orders"
            headers:
              - name: ":method"
                string_match:
                  exact: "POST"
          route:
            cluster: orders_primary

        - match: { prefix: "/v1/orders" }
          route:
            cluster: orders_primary
            request_mirror_policies:
              - cluster: orders_candidate
                # Start at 1%. The shadow has to keep up with whatever you
                # set here, and a shadow that falls behind drops requests
                # non-uniformly -- which looks exactly like a bug in the
                # candidate until you work out it was a queue.
                runtime_fraction:
                  default_value: { numerator: 1, denominator: HUNDRED }
                  runtime_key: shadow.orders.fraction
                # During shadowing Envoy alters the host/authority header so
                # that "-shadow" is appended (cluster1 -> cluster1-shadow).
                # The docs present this as a logging convenience. It is also
                # the candidate's ONLY in-band way to know it is a shadow,
                # which makes it the natural gate for every side effect in
                # the process. disable_shadow_host_suffix_append turns it
                # off. Do not turn it off.

  # Not optional. Without a circuit breaker the shadow cluster is a new,
  # untested dependency sharing a connection pool with your live path.
clusters:
  - name: orders_candidate
    connect_timeout: 0.25s
    circuit_breakers:
      thresholds:
        - priority: DEFAULT
          max_connections: 64
          max_pending_requests: 64
          max_retries: 0

One detail in there is worth more than the rest of the config. Envoy alters the host/authority header during shadowing so that `-shadow` is appended — `cluster1` becomes `cluster1-shadow` — and documents it as useful for logging. It is better than that: it is the candidate process's only in-band way to know it is a shadow, which makes it the natural place to gate every side effect in the codebase. There is a field to disable the suffix. Leave it alone.

# nginx does the same thing with ngx_http_mirror_module, which has exactly
# two directives: `mirror` and `mirror_request_body`.
location /v1/orders {
    mirror /shadow;

    # DEFAULT IS ON, and you want it on for a real diff -- a mirror without
    # bodies cannot test a POST. It is not free on the LIVE path, though.
    # The docs are explicit: with this on, the client request body is read
    # BEFORE the mirror subrequests are created, which disables unbuffered
    # client request body proxying set by proxy_request_buffering (and the
    # fastcgi/scgi/uwsgi equivalents).
    #
    # So the "read-only observability change" you just shipped altered how
    # your production proxy buffers uploads. Measure that before you argue
    # with anyone about whether mirroring is risk-free.
    mirror_request_body on;

    proxy_pass http://orders_primary;
}

location = /shadow {
    internal;
    proxy_pass http://orders_candidate$request_uri;

    # "Responses to mirror subrequests are ignored." Same shape as Envoy:
    # duplication, not comparison. The candidate has to record its own
    # answer, keyed on something you can join against the primary's log --
    # $request_id is right there and is the cheapest join key you will find.
    proxy_set_header X-Shadow     1;
    proxy_set_header X-Replay-Id  $request_id;

    # The subrequest is not promised to be free. Bound it so a wedged
    # candidate cannot become a latency story on the live path.
    proxy_connect_timeout 250ms;
    proxy_read_timeout      2s;
}
nginx's `mirror_request_body` defaults to on, and the docs state that the client request body is read before the mirror subrequests are created — which disables unbuffered client request body proxying set by `proxy_request_buffering` and its fastcgi/scgi/uwsgi equivalents. Turning on a read-only observability feature just changed how your production proxy handles uploads. This is the single most common way a shadow-traffic project causes its first incident, and the incident is on the primary.

Replay: capture a corpus, run it whenever you like

The other shape is to record requests to a file and run them later. GoReplay is the obvious tool here and describes itself as "An open-source tool for capturing and replaying live HTTP traffic into a test environment in order to continuously test your system with real data" — `--input-raw` listens on an interface (raw sockets, so root), `--output-file` writes a capture, and `--input-file` plus `--output-http` plays it back somewhere else. There is a commercial GoReplay PRO alongside the open-source build, so check which features you are reading about before you plan around one.

What replay buys is repeatability, and repeatability is what turns a diff into an investigation. You can run the same 50,000 requests against two commits. You can bisect. You can re-run after a normaliser change and see whether the diff count moved for the reason you intended. A mirror can never do any of that, because yesterday's traffic is gone.

What it costs is freshness, in a way that is worse than it sounds. A corpus goes stale in content, not just in time: the order it references gets cancelled, the feature flag it was captured under gets flipped, the user is deleted, the coupon expires. Replay that request against a world where the row no longer exists and you get a `404` from both sides, which is a pass that proves nothing, or a `404` from one side, which is a diff that means nothing. Captures also frequently lack fields you did not know you needed — the header your auth middleware reads, the body of the request your proxy was not buffering.

One correction worth making because the advice circulates: `vegeta` is not a replay tool. It reads targets from a file in its plain-text HTTP format or as JSON lines and fires them at the rate you set with `-rate`; there is no notion of replaying at the captured timing. Pointing it at a corpus gives you a load test with realistic URLs, which is a genuinely useful but completely different experiment — it proves the candidate survives the shapes, not that it answers them the same way.

The side-effect problem is the whole post

Here is the sentence that ends most shadow-traffic projects: a mirrored `POST /v1/charges` charges the customer twice. Not in theory. The request is byte-identical, the candidate has the same code path, and nothing in Envoy or nginx knows that one of these two is pretend.

There is no configuration that makes this safe. There are four layers of answer, and they trade honesty for coverage in a very specific order.

Layer 1: filter to read-only, and admit what you gave up

Mirror `GET` and `HEAD` only. This is what diffy does by default and what scientist tells you to do, it takes ten minutes, and it is the right first move. It also tests the boring half. The interesting logic in most services is behind a mutating verb, and the regressions you are most afraid of — the rewrite that writes the wrong enum, the migration that loses a timezone — are write-path regressions by construction.

Worse, "read-only" is a lie at the HTTP layer. Plenty of `GET` handlers write: a last-seen timestamp, a view counter, a lazy backfill on first read, a cache entry, an audit row, an analytics event. Grep for writes in your `GET` handlers before you call this tier safe, and be prepared to be annoyed by what you find.

Layer 2: a fresh database per shadow run

Give the candidate its own Postgres, restored from a sanitised dump. Now a mirrored `POST` writes a row nobody will ever read, and you can finally diff the write path. This is the first tier that tests the half you cared about.

It introduces the dominant source of noise in any write-enabled shadow, which is state skew. A mirrored request referencing a row created in production ninety seconds ago will `404` in your shadow, because your dump is from this morning. That is a diff. It is not a bug. You will chase several of them before you accept that the only fixes are a very fresh copy, or a corpus filtered to requests whose referenced entities existed at dump time — and that filter quietly biases your corpus towards older, calmer data.

Layer 3: stub the providers that offer a stub

Payment providers have test keys. Object stores have a local emulator. Use them, and notice that you have just written a second implementation of somebody else's semantics. When your diff says "the candidate handled the declined card differently," the first question is not about the candidate — it is whether the stub declined it the same way twice.

Layer 4: the downstreams you do not own, which is where it actually bites

Everything above is tractable. This is the list that gets you, because each item is a system you cannot give the shadow its own copy of:

  • The email provider — a mirrored password-reset request sends a real email to a real customer, who now has two reset links and a reasonable question about your security posture.
  • Outbound webhooks — the shadow fires your `order.created` hook into a customer's endpoint. They received it twice, their handler was not idempotent, and the resulting conversation is not about your rewrite.
  • Analytics and product events — your funnel doubles on one endpoint. Someone writes a note about a growth inflection. Three weeks later someone else finds the cause, and your event stream has a permanent scar that every cohort query now has to know about.
  • Third-party rate limits, which is the sharp one. The shadow spends the same API budget as production, from the same key. You hit the limit, and the path that gets the `429` is the production path. Your read-only experiment caused a customer-facing outage, and the proximate cause is a feature whose entire selling point was that it does not affect users.
  • Message topics — the shadow publishes to the same Kafka topic and a consumer owned by a different team acts on it. Their service was correct. Your experiment was not contained. Nobody involved had the information to predict it.
  • Idempotency keys — reuse the production key and the provider dedupes the shadow's call, which saves you by accident. Mint a fresh one and you charge twice. Whichever happens, you did not decide it; your HTTP client did, months ago.
  • Shared caches — the shadow writes a key the production path then reads. A read-only shadow has just poisoned production through Redis, and the request that broke was not one of the mirrored ones.
  • Egress itself — you roughly doubled outbound call volume to every dependency. Some of them bill per call, and the bill is the only part of this experiment that was never in doubt.
The only containment that holds is an environment that cannot reach the thing. "We were careful" is not a control; it is a narrative you will deliver at the incident review.

Which is the real argument for default-deny egress on the shadow and an unroutable stub for everything you did not explicitly allow. If the candidate cannot resolve your email provider, nobody has to remember that the password-reset handler sends mail.

Why a shadow run wants a whole world, created and destroyed

Follow that containment requirement to its conclusion and the unit you actually need is not a service — it is a world. The candidate, its database, its cache, its outbound network policy, all created from a definition and all destroyed afterwards, sharing nothing with production that a mistake could travel along.

The usual version of this is a long-lived staging environment, and the usual complaint about staging is that it drifts. It drifts because it persists: somebody patched a config in March to unblock a demo, somebody's test data is load-bearing, and the dump is from a schema version ago. A per-run environment cannot drift, for the uninteresting reason that it does not live long enough to.

That is only affordable if creating one is cheap. On our stack a sandbox create is a Firecracker snapshot restore rather than a boot — p50 179 ms, p99 203 ms, with the `/snapshot/load` step itself around 49 ms. The first create of a template before its snapshot exists is a real cold boot at roughly 3 seconds, and then the snapshot is baked and every create after that is on the fast path. At 179 ms the natural unit of isolation stops being "the staging environment" and becomes "this diff investigation."

The database is the honest part of the budget. A managed Postgres is a dedicated microVM with a durable volume, and it takes 30–90 seconds of wall clock to be usable — the REST API returns `202` with `status: "provisioning"` immediately and the client polls, so nothing blocks on an HTTP request, but the thirty seconds are real. Amortise it: one database per investigation, not one per request.

import io
import json
import tarfile
from pathlib import Path

import pandastack
from pandastack import Sandbox

client = pandastack.Client()          # reads PANDASTACK_API_KEY

# --------------------------------------------------------------------------
# 1. A whole-world copy. The candidate, plus a database nobody will read.
# --------------------------------------------------------------------------
# POST /v1/databases answers 202 immediately with status "provisioning"; the
# SDK polls GET /databases/{id} until it reports running, so this call does
# block for the real 30-90s Postgres takes to come up. Budget for it once per
# investigation, not once per request.
#
# size is a RAM tier -- "1g" (default), "4g", "16g" -- and it is FIXED at
# create time, because a snapshot-restored microVM cannot be resized. To
# change it later you clone with a new size.
db = client.databases.create(label="shadow-orders-run-41", size="1g")

# `base` is the language-agnostic apps runtime: 4 GiB of guest RAM, 8
# burstable vCPU, mise with Node/Python/Go/Bun already warm. Creating it is a
# snapshot restore, not a boot -- p50 179 ms, p99 203 ms.
#
# ttl_seconds is an IDLE timeout (default 5 minutes), not a walltime budget:
# the reaper measures time since last activity. A running replay keeps the
# guest alive. You staring at a diff for twenty minutes does not.
sbx = Sandbox.create(
    template="base",
    ttl_seconds=1800,
    metadata={"role": "shadow-candidate", "run": "41"},
)

# --------------------------------------------------------------------------
# 2. Get the bundle in. filesystem.write handles SINGLE FILES -- upload() is
#    literally write(remote, Path(local).read_bytes()), so handing it a
#    directory raises IsADirectoryError. Tar the tree first.
# --------------------------------------------------------------------------
buf = io.BytesIO()
with tarfile.open(fileobj=buf, mode="w:gz") as tar:
    tar.add("candidate", arcname="candidate")        # the rewrite
    tar.add("replay.py", arcname="replay.py")        # the in-guest driver
    tar.add("corpus.jsonl", arcname="corpus.jsonl")  # captured, scrubbed
sbx.filesystem.write("/srv/bundle.tar.gz", buf.getvalue())

# Docker ENV from the template's Dockerfile does NOT reach a Firecracker
# guest; only /etc/environment (via PAM) reaches a detached process. And the
# export has to happen in the launch shell BEFORE setsid, or the detached
# child never inherits it. Both of those have cost me an afternoon.
sbx.filesystem.write("/srv/start.sh", f"""#!/bin/sh
set -eu
tar -xzf /srv/bundle.tar.gz -C /srv
cd /srv/candidate
export DATABASE_URL='{db["connection_url"]}'
export SHADOW=1
setsid nohup ./server --port 8080 >/var/log/candidate.log 2>&1 </dev/null &
for i in $(seq 1 60); do
  curl -fsS -o /dev/null http://127.0.0.1:8080/healthz && exit 0
  sleep 1
done
echo "candidate never came up" >&2; tail -40 /var/log/candidate.log >&2; exit 1
""")

# Keep one-shot exec() short. No exec endpoint enforces timeout_seconds
# server-side -- the agent decodes it and applies it on neither -- and the
# client kills the call at 30s by default, so a long command needs
# exec_stream (where the value raises the CLIENT timeout) plus a real
# `timeout` in the shell.
sbx.exec("sh /srv/start.sh", timeout_seconds=90, check=True)

# --------------------------------------------------------------------------
# 3. Replay INSIDE the guest, against 127.0.0.1. Two reasons, both load-
#    bearing: there is no ingress to configure, and the guest's own network
#    namespace is the containment. The candidate cannot email your customer
#    because it cannot reach your email provider -- not because you promised
#    yourself you had filtered that route.
# --------------------------------------------------------------------------
rc = sbx.exec_stream(
    "python3 /srv/replay.py /srv/corpus.jsonl /srv/out.jsonl",
    on_stdout=lambda c: print(c, end=""),
    timeout_seconds=3600,
)
assert rc == 0, f"replay exited {rc}"

# --------------------------------------------------------------------------
# 4. Pull the evidence out BEFORE teardown, and note what is NOT here: a
#    try/finally around sbx.kill(). An explicit kill() cascade-deletes that
#    sandbox's snapshots; the idle reaper does not. The reflexive
#    `finally: sbx.kill()` is therefore also the line that destroys the
#    snapshot you took to investigate the diff with.
# --------------------------------------------------------------------------
Path("out.jsonl").write_bytes(sbx.filesystem.read("/srv/out.jsonl"))

# Worth keeping if the diff is interesting: snapshot() captures memory AND
# disk and returns a snapshot ID *string*. That string is a reproducible
# starting point for everyone who is about to argue about this diff.
if json.loads(Path("out.jsonl").read_text().splitlines()[-1]).get("diffs"):
    print("snapshot for the investigation:", sbx.snapshot())

sbx.kill()
client.databases.delete(db["id"])

A note on fanning out, because the two fork primitives are not interchangeable and the difference matters for exactly this workload. `fork()` is disk-only: the child gets a copy-on-write clone of the parent's rootfs and then cold-boots from it, with no parent memory and no running processes. For a replay shard that is often precisely what you want — the installed candidate and the populated database files, with a server that starts clean. `fork_tree(count)` does inherit memory and disk, so the children come up with the parent's live process, but it caps at 16 children and pins them to the parent's host. For a replay wide enough to need more than one host, take a `snapshot()` — which captures memory and disk and returns an ID string — and then `Sandbox.create(from_snapshot=...)` as many times as you like, so each create goes through the scheduler and spreads.

Diffing responses is much harder than it looks

Your first unnormalised diff will come back at close to 100%, and the reason will be the `Date` header. Then you fix that and it comes back at 94% because the rewrite changed language and JSON key ordering with it. Here is the list, roughly in the order it ambushes people:

  • `Date`, `Server`, `X-Request-Id`, `Set-Cookie` session identifiers, `X-Envoy-Upstream-Service-Time` — free, obvious, fixed in five minutes.
  • `Content-Length`, which changes if the new serialiser emits a space after a colon. Nothing semantic moved and every single response differs.
  • Key ordering. Go's `encoding/json` sorts map keys on marshal; a struct marshals in field-declaration order; Python's `json.dumps` preserves insertion order. Change language and every object differs.
  • Timestamps in bodies, at whatever precision the new driver happens to render — and the timezone suffix, where `+00:00` and `Z` are the same instant and not the same string.
  • Generated UUIDs, which differ per request by definition and are the field most likely to be blanked carelessly. See below.
  • Float formatting. Shortest-round-trip representation versus six fixed decimals is a formatting difference; a different accumulation order in a sum is not, and both arrive as a changed digit.
  • Integer precision. An int64 id serialised by a runtime that parses JSON numbers as doubles loses fidelity past 2**53. This is real corruption wearing the costume of a formatting diff.
  • Pagination cursors, which are opaque by design and encode internal state by construction, so they differ as soon as the internals do.
  • Result ordering with ties. `ORDER BY status` over equal keys returns ties in whatever order the plan produced, the old service never promised an order, and every client that depended on one is about to find out.
  • Error bodies. Message strings, field names, and whether a stack trace leaks — the last of which is a finding even when it is not a regression.
# The normaliser is the most dangerous file in a replay harness, because its
# job is to make diffs go away and it is extremely good at its job.

VOLATILE_HEADERS = {
    "date", "server", "x-request-id", "x-trace-id", "set-cookie",
    "content-length",          # whitespace in the serialiser changes this
    "age", "via", "x-envoy-upstream-service-time",
}

VOLATILE_FIELDS = {"created_at", "updated_at", "request_id", "trace_id"}


def normalise(status, headers, body_text):
    """Return a comparable form of one response. Everything removed here is
    a claim that it could not possibly have been the bug."""
    h = {k.lower(): v for k, v in headers.items() if k.lower() not in VOLATILE_HEADERS}

    try:
        body = json.loads(body_text)
    except ValueError:
        # Not JSON. Compare bytes and accept that a changed error page is a
        # diff you will have to read with your eyes.
        return status, h, body_text

    def scrub(node):
        if isinstance(node, dict):
            # sort_keys at dump time handles ordering. Worth knowing WHY you
            # need it: Go's encoding/json sorts map keys on marshal, Python's
            # json.dumps preserves insertion order, and a struct marshals in
            # field-declaration order. Rewrite the service in another
            # language and every single object differs on ordering alone --
            # a 100% diff rate that means nothing.
            return {k: ("<ts>" if k in VOLATILE_FIELDS else scrub(v))
                    for k, v in node.items()}
        if isinstance(node, list):
            return [scrub(v) for v in node]
        return node

    return status, h, json.dumps(scrub(body), sort_keys=True, separators=(",", ":"))


# ---------------------------------------------------------------------------
# AND HERE IS THE LINE YOU ARE ABOUT TO ADD, WHICH WILL HIDE A REAL BUG
# ---------------------------------------------------------------------------
# UUIDs differ per request, so the obvious next move is to blank every id:
#
#     if k.endswith("_id"):
#         node[k] = "<id>"
#
# Add that and the diff goes quiet. It also goes quiet about the fact that
# the Node rewrite returned order_id 9007199254740993 as 9007199254740992,
# because an int64 primary key does not survive being parsed as a double past
# 2**53. That is silent data corruption, it surfaced in exactly this field,
# and the three lines you wrote to reduce noise deleted the only evidence.
#
# Normalise the SHAPE, not the VALUE. For an id field, assert that both sides
# produced an integer and compare the RAW JSON TEXT of the number, never the
# parsed value -- your own json.loads has the same 2**53 problem the bug does.
#
# The general rule, which is not a rule you can automate: every entry in
# VOLATILE_* is a hypothesis that this field cannot carry a regression. Write
# them down as hypotheses, keep the list short, and re-read it the next time
# a customer reports something the replay said was fine.

That is the real hazard and it does not have a tooling fix. Every entry in a normaliser's ignore list is an unstated hypothesis that the field cannot carry a regression, and the incentives are all wrong: a longer list makes your diff report look better, your dashboard greener, and your rewrite more finished. Keep the list short, write each entry down as a claim rather than a config value, and re-read it the first time a customer reports something the replay said was fine.

Diffy's structural answer is better than any ignore list and is worth copying even if you never run their tool: compare the known-good against itself. Replay the corpus at the old implementation twice and record where it disagrees with itself. A field that disagrees with itself four percent of the time is not evidence of anything at four percent, and you now know that without having to guess which fields are volatile. It costs one extra world — the old implementation, stood up twice, replayed against the same corpus — and the sandbox half of that is 179 ms. The database is where the thirty seconds go, which is the argument for one per investigation rather than one per run.

Production traffic is production data

Replaying captured traffic moves real customer data into a system that did not have it before. That is a new processing location, a new retention obligation, a new breach surface and possibly a new subprocessor. It is a compliance event, not an implementation detail, and the fact that it is technically easy is exactly why it gets done without anyone being asked.

  • Scrub at capture time, not at replay time. A capture store that has ever held raw bodies is in scope from that moment, and "we delete it later" is a plan, not a control.
  • Drop `Authorization` and every cookie at capture. You almost never need the customer's real credential — you need a token that resolves to an equivalent principal in the shadow. Keeping live session tokens in a file on a laptop is its own incident waiting for a date.
  • Give the capture store an actual expiry, and do not make it your logging pipeline. A corpus that lands in the same place as your logs has just extended your log retention policy to raw request bodies.
  • If the shadow runs in a different account, region or provider than production, the mirror is a data-transfer path. Check that it is one your agreements already cover before the config merges, not after.
  • The genuinely good property of ephemeral infrastructure here: a per-run guest that is destroyed afterwards is a bounded, short-lived copy with a known end. The common alternative is a staging database seeded from a production dump two years ago that nobody has the authority to delete, and that is strictly the worse artefact on every axis a regulator cares about.

Reading the results: shape, not count

Say your run comes back with a diff on 0.3% of requests — a plausible shape for a serious rewrite with a decent normaliser, and emphatically not a target. The instinct is to drive it to zero, and the instinct is wrong, because most of that 0.3% is noise and a percentage cannot tell you which part. A diff rate is a single number with no actionable content in it.

Group instead by `(endpoint, JSON path, status pair)` and sort by count. The shape of the top entries tells you what to do:

  • One field, across every endpoint: a serialiser or driver difference. One fix takes most of your diff volume with it, and you should do it before reading anything else.
  • One endpoint, most of its fields: that handler is wrong. This is the finding you ran the whole exercise for.
  • Scattered, and gone on a re-replay: noise, ordering, or state skew. Confirm it against the control-against-itself run rather than arguing about it.
  • Only requests from one client or one user-agent: the old code has a compatibility branch nobody documented, and you just found the documentation.
  • Only requests past some size or page depth: a boundary condition — a limit, an overflow, a pagination edge the new implementation reads differently.
  • Only requests from one tenant: either their data has a shape yours does not, or there is per-tenant behaviour in the old code that was never a feature.

And then the category nobody plans for: the candidate is right and the old service was wrong. The replay has handed you a real bug in production, several clients that have been compensating for it, and a decision that is not yours. Preserve the bug and ship, or fix it and break them on a schedule you announce. Finding it was the entire value of the exercise. It will not tell you which to pick.

Mirror, replay, canary, contract tests

These get argued about as though you have to pick one. You do not, and the argument is usually a proxy for a different disagreement about how much a rewrite is allowed to cost. Graded on the four axes that actually trade against each other:

Four ways to gain confidence in a rewrite. They are not substitutes — the usual right answer is contract tests in CI, a replay corpus at merge, and a canary at ship.
TechniqueInput freshnessRepeatabilitySide-effect riskWhat it actually proves
Live mirroring (Envoy, nginx)Perfect: this second's traffic, against rows that exist nowNone. Yesterday is gone and cannot be re-runHighest: every mirrored write happens for real, immediately, with no undoThat the candidate answers today's real distribution identically — but only if you built the recording and joining yourself, because the proxy is fire-and-forget
Captured replay (a corpus on disk)Decays: referenced rows get deleted, flags flip, the capture ages badlyTotal. Same corpus, two commits, bisectable, and you can re-run after changing the normaliserControllable: the corpus is a file you can filter, inspect and scrub before anything executesThat the candidate answers that corpus identically. A claim about the past, which is the only claim you can make twice
Canary (a small share of real users)Perfect, and it is genuinely in productionNone, and the next canary is a different populationReal but bounded by the share — and the blast radius is users, not duplicated writesThat the candidate does not fall over under production conditions. Measured on error rate and latency, not on response equality — a subtly wrong answer passes a canary
Contract tests (Pact and relatives)None: the inputs are hand-written examplesTotal, and they run in CI in seconds with no infrastructureZeroThat both sides agree about the expectations somebody wrote down. Silent about every behaviour nobody encoded — which is precisely the behaviour a rewrite is trying to preserve

What this does not fix

A post that stopped at the table would be an advert. The limits here are structural, not roadmap items.

  • Stateful sequences. A corpus of independent requests replays cleanly. A corpus where request N reads the row request N−1 created does not, unless you preserve order, concurrency and the state as of capture start. Most of the genuinely interesting bugs live in exactly those sequences, and the honest workaround is to replay whole sessions as units and accept a much smaller corpus.
  • Timing and concurrency. A replay is a different interleaving by construction. Races do not reproduce, lock-ordering bugs do not reproduce, and a clean replay is not evidence about either.
  • Load. A matching response at your replay rate says nothing about behaviour at peak. That is a different experiment with different tooling, and conflating the two is how a rewrite passes every diff and then falls over on the first Monday morning.
  • The write path, fundamentally. There is no setting that makes a mirrored `POST` safe. You filter it, or you fake the world around it, or you accept the consequence. Any document that implies otherwise is selling you something.
  • Traffic is not coverage. Production traffic exercises what customers actually do, which systematically excludes error paths, admin endpoints and the one integration an enterprise calls once a quarter. A 100% match rate is proof of equivalence on the distribution you captured, and the quarterly endpoint is the one that will page you.
  • Anything that is not request/response. Websockets, gRPC streaming, long-poll: a mirror module duplicates a request, and a bidirectional stream has no clean duplicate. You can replay the messages; you cannot mirror the connection.
  • Guest RAM is a build-time decision. Firecracker cannot change a guest's vCPU or RAM at snapshot restore, so a template's memory is chosen at build time with `--memory-mb` and the agent silently corrects whatever `memory_mb` you pass on create. `base` is 4 GiB; if your candidate needs more, that is a template, not a parameter.

What you get for all of it is narrow and worth having. Not a proof of correctness — a list, ordered by frequency, of the places where your new implementation disagrees with the only specification that was ever true. Everything on that list is either a regression you were going to ship, a bug you did not know you had, or a hypothesis in your normaliser that deserves a second look.

Three outcomes, all of them better than finding out from a customer.

Frequently asked questions

Is there any safe way to mirror POST requests?

Not in the proxy, no. A mirrored POST is a byte-identical request arriving at code that was written to do things, and neither Envoy nor nginx has any concept that one of the two copies is pretend. You have three real options and they are all work. Filter mutating methods out entirely — what diffy does by default and what scientist's README tells you to do — and accept that you are testing the read half. Or build the candidate a world of its own: its own database, stubbed providers with test credentials, default-deny egress so the things you forgot about cannot be reached, and a run that is destroyed afterwards. Or gate side effects inside the application, using the fact that Envoy appends `-shadow` to the host/authority header during shadowing, so the process can tell. The third is the most elegant and the least trustworthy, because it depends on every future contributor remembering the gate. The second is the one that holds, because it does not depend on anybody remembering anything.

Should I start with live mirroring or captured replay?

Start with a captured corpus, almost always. Mirroring looks cheaper — it is a few lines of proxy config and no storage — but it is unrepeatable, which means every diff is a one-shot observation you cannot investigate, re-run after fixing your normaliser, or bisect to a commit. A corpus turns each diff into something you can work on. It is also the safer first step operationally: a file on disk can be inspected, filtered and scrubbed before a single request executes, while a mirror policy is live the moment it merges. Mirroring earns its place later, once your normaliser is mature and your containment is real, because it is the only one of the two that tests against today's data instead of this morning's dump. One genuine caution on the mirror side regardless of ordering: nginx's `mirror_request_body` defaults to on and the docs note it disables unbuffered request-body proxying, so enabling mirroring changes behaviour on your live path, not only on the shadow.

How do I keep my response normaliser from hiding the bug I am looking for?

Treat every ignore rule as a written hypothesis that the field cannot carry a regression, and keep the list embarrassingly short. The classic failure is blanking everything matching `*_id` because UUIDs differ per request — which also silently absorbs an int64 primary key being mangled past 2**53 by a runtime that parses JSON numbers as doubles. That is real data corruption presenting as a formatting diff, and three lines of noise reduction deleted the only evidence of it. Normalise shape rather than value: assert that both sides produced an integer and compare the raw JSON text of the number, since your own parser has the same precision problem the bug does. Then copy diffy's structural trick, which is stronger than any ignore list: replay the corpus at the known-good implementation twice and record where it disagrees with itself. That gives you a measured per-field noise floor instead of a guessed one, and it means you only escalate disagreement that exceeds what the old code already exhibits against itself.

Is replaying production traffic into a test environment a compliance problem?

Yes, and it is worth saying out loud before the capture starts rather than after. Production traffic contains production data, so a capture file is a new copy of customer personal data in a new location, with a new retention obligation, a new breach surface and possibly a new subprocessor if the shadow runs somewhere production does not. The controls are not exotic: scrub at capture time rather than replay time, because a store that has ever held raw bodies is in scope from that moment; drop `Authorization` and cookies at capture, since you need a token that resolves to an equivalent principal, not the customer's live session; put a real expiry on the capture store and keep it out of your logging pipeline. Ephemeral infrastructure genuinely helps with the last mile here — a per-run guest that is destroyed afterwards is a bounded copy with a known end date, which is a much easier thing to describe to an auditor than the staging database seeded from a two-year-old production dump that nobody has the authority to delete.

My replay matched 100%. Is the rewrite safe to ship?

It is safer than it was, and it is not proven. A 100% match is a statement about the distribution you captured, and production traffic is a biased sample of your API by construction: it over-represents whatever customers do constantly and under-represents error paths, admin endpoints, migration routes and the one integration an enterprise calls at quarter end. It also says nothing about concurrency, because a replay is a different interleaving than the original, so races and lock-ordering bugs do not reproduce; nothing about load, because matching at your replay rate is not matching at peak; and nothing about anything that is not request/response, since a mirror duplicates requests and a websocket has no clean duplicate. Treat a clean replay as having retired one specific risk — unintended behaviour change on the traffic you see most — and then ship behind a canary anyway, watching error rate and latency, which is the dimension response diffing is blind to.

Keep reading

Related posts

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.