all posts

Lighthouse and axe-core Audits, One MicroVM Each

Ajay Kumar··11 min read

The accepted methodology at most companies goes like this. You push a branch. The performance budget fails on largest contentful paint by a hundred and forty milliseconds. You stare at your diff, which changed a copy string and a border radius. You click Re-run job. It passes. You merge, and the green run becomes the true one by a process best described as selection.

Nobody in that story is being dishonest. Re-running until green is not cheating when your measurement is noise — it is sampling. The problem is that you then report the sample you liked, and the next person inherits a trend line made of samples somebody liked. The bug is not in CI. It is in the experiment.

Here is the framing that fixes it, and it is the reason this is a microVM post rather than a Docker post. A performance audit is a measurement. Lighthouse is an instrument. Instruments need controlled conditions, and a shared, contended CI host is the opposite of a controlled condition. When forty Lighthouse runs share a kernel, a page cache and a CPU, every one of them reports a worse LCP than the truth, and the variance between them is larger than the regression you were trying to detect. One microVM per audit is not primarily a security boundary. It is a scientific control that happens to also be a security boundary.

Why your LCP moved when nothing shipped

Before fixing anything, it is worth being specific about where the noise comes from, because the fixes are different for each source and some of them are not fixable at all.

  • CPU contention — the big one. Lighthouse's main-thread work (parse, compile, execute, layout) is measured in wall-clock time on a core that something else also wants. Total blocking time is almost a direct readout of how busy your runner was.
  • Page cache sharing — the second run of the day finds Chrome's binary, the font files and the Node module tree already resident in the host's page cache. The first run paid for those reads. Those are different experiments with the same name.
  • Frequency and thermals — turbo budget is shared across a socket. A neighbour saturating four cores drops the clock your single-threaded parse was relying on, and nothing in your metrics mentions it.
  • Steal time — on a shared cloud instance your guest can be ready to run and simply not scheduled. The guest's own clock keeps ticking through it, so the time lands in your LCP and not in your CPU utilisation graph.
  • Network variance — DNS, TLS handshake, which CDN edge answered, whether the origin had just autoscaled. The path to the site under test is part of the measurement whether you modelled it or not.
  • Browser state — a leftover profile directory, a service worker from the previous run, an extension somebody installed six months ago, a Chrome that silently updated itself between Tuesday and Wednesday.

Lighthouse's throttling makes this worse, not better

The instinctive answer is "but Lighthouse throttles the CPU and the network, so the host does not matter." It is the opposite. Lighthouse's default simulated throttling loads the page at full speed and then models what the load would have looked like on a slower device and a slower link. That model has to be calibrated against the machine it is running on, so Lighthouse computes a synthetic CPU score at run time and records it in the report as the benchmark index.

Follow the consequence. On a contended host the underlying trace is already degraded, and the calibration constant used to simulate a slow device is measured during the same contention. You are not getting a clean simulation of a slow phone. You are getting a simulation of a slow phone layered on top of a machine that was genuinely slow, with the simulator's own idea of "fast" distorted by the same cause. The two error terms are not independent, which is the worst possible property for an error term to have. Switching to DevTools-based throttling gives you real request-level network shaping, but the CPU slowdown multiplier is still applied on top of a CPU you are sharing.

The benchmark index in the Lighthouse report is the most under-used field in the whole JSON. It is a direct statement about the machine, taken at the moment of the run. Record it with every result, and treat a run whose benchmark index falls below your fleet's normal floor as invalid data rather than as a slow page. You will throw away a surprising number of "regressions" that way.

The architecture: one microVM per audit

The shape is boring, which is the point. Every audit run gets its own Firecracker microVM, restored from a baked snapshot of the `browser` template: Ubuntu 24.04, Node 24, Playwright's Chromium, xvfb, and 4 GiB of RAM with 8 vCPU of burst. The audit runs. The JSON comes back. The VM dies. Nothing is reused except the snapshot.

Three properties follow from that, and the third is the one a container cannot give you.

  1. Its own kernel. Your run's scheduler, page cache, dirty-page writeback and network stack belong to your run. A noisy co-tenant on the same host still exists — see the honest section below — but it is no longer inside your kernel making decisions about your threads.
  2. Its own memory, sized by the template. The 4 GiB and 8 vCPU are baked into the snapshot, so every run of every audit for the next six months has exactly the same resource envelope. Nobody can accidentally tighten a limit in a YAML file and shift your whole trend line.
  3. A byte-identical starting state, including RAM. A snapshot restore does not just give you the same filesystem. It gives you the same memory image — the same page cache contents, the same already-initialised process state, the same empty /tmp, the same font set, the same zero extensions. Run number one and run number four hundred begin from the same bytes.

That third point deserves a sentence of elaboration because it is easy to under-rate. A container shares the host's page cache, so "cold" means whatever the host happened to be caching, which depends on what the previous tenant did. A snapshot restore defines cold. And the font set is not a trivial detail: installed fonts change text metrics, which changes line breaks, which changes which element is the largest contentful paint. A container inheriting a slightly different font package from a base-image bump will quietly start measuring a different element and report it as a regression.

Restores are cheap enough for this to be the default rather than a luxury: a create on a baked template has a p50 of 179 ms and a p99 of 203 ms, with the snapshot load itself around 49 ms. A per-audit VM costs you roughly a fifth of a second of setup, which is noise against a Lighthouse run and very much not noise against the variance it removes.

Pin the browser, or you are charting Chrome releases

Lighthouse is not baked into the stock `browser` template — Playwright's Chromium is, Lighthouse is not. The tempting fix is to install it at the top of every run, which means every audit fetches whatever npm is serving that minute. Combine an unpinned Lighthouse with a Chromium that updates itself and your performance dashboard becomes a changelog of other people's releases, rendered as a line chart, with your own commits as noise on top.

Pin both explicitly, and once you care about the trend line, bake them into a custom template so the versions are a property of the snapshot rather than a property of today's registry. Then treat a template re-bake as a deliberate break in the time series: annotate the date, and do not compare medians across it. A Chrome major version is a step change in your numbers, not a regression, and if you cannot see where the steps are on your chart you will spend a sprint bisecting one.

The parts people get wrong

Isolation buys you the right to run a real experiment. It does not run one for you. Four methodology errors account for most of the bad performance dashboards I have seen.

Take the median, never the mean. Load-time distributions are right-skewed: most runs cluster, and a few land far out because of a garbage collection pause, a cold CDN object or a retried request. The mean chases those outliers, which is exactly the behaviour you do not want from a regression detector. Run the same commit at least five times and take the median of the metric you care about. Three is the minimum at which a median means anything at all; one run is a coin flip with extra steps.

Audit a pinned commit, not "production". If a deploy lands between run two and run three, your five runs are not five samples of one thing. Point the audit at an immutable artefact — a specific commit's preview deployment — and record the commit SHA in the result.

Cold-cache and warm-cache runs are different experiments. Lighthouse clears storage before a run by default, which measures a first visit. Disabling that reset measures a repeat visit with the HTTP cache and any service worker in play. Both are legitimate and they answer different questions. Averaging them together answers neither, and a config change that flips the default will look exactly like a performance win.

The self-inflicted one: auditing a sleeping app

This is the error I have actually made. If the app under test scales to zero — a preview environment, a per-branch deployment, anything with an idle timeout — the first request of your audit pays for waking it up. Your time to first byte is then a measurement of your own hosting platform, not of your frontend, and it arrives in the one metric engineers are most likely to treat as gospel.

You will then spend an afternoon optimising a server response that was never slow. The fix is one line of protocol: before the audit, send a throwaway request to the target, assert it came back 200, and only then start measuring. If you are auditing against a PandaStack preview URL, remember that `ttl_seconds` is an idle timeout rather than a wall-clock budget and defaults to five minutes — so a target that sat untouched through a long build may well be asleep by the time your audit job starts.

The Lighthouse invocation that is actually repeatable

What follows is the script that runs inside the guest. Most of its value is in the flags it sets and in the two flags it deliberately does not set.

#!/bin/sh
# run-audit.sh -- one Lighthouse run, inside one microVM.
#
# Runs on the `browser` template: Ubuntu 24.04, Node 24, Playwright's Chromium,
# 4 GiB / 8 vCPU baked into the snapshot. Lighthouse itself is NOT in the stock
# template, so it is installed here -- pinned. An unpinned `npx lighthouse`
# turns your performance dashboard into a chart of npm's release schedule.
# Once the trend line matters, bake this into a custom template instead:
#   pandastack template build -f Dockerfile -n audit-runner
set -eu

URL="$1"
LH_VERSION="12.2.1"

npm i -g --silent "lighthouse@${LH_VERSION}"

# Point Lighthouse at the Chromium that Playwright installed at template build
# time, rather than letting it search the filesystem and find something else.
CHROME_PATH="$(node -e "console.log(require('playwright').chromium.executablePath())")"
export CHROME_PATH

# A fresh profile per run. Carried-over profile state (a service worker, a
# cached DNS entry, a prerender hint) is the quietest source of drift there is.
PROFILE="$(mktemp -d)"

# `timeout` in the shell, not a long client-side deadline: it is the guest's own
# kernel enforcing the bound, which is the only place it is reliable.
timeout 600 lighthouse "$URL" \
  --output=json --output-path=/workspace/lhr.json \
  --only-categories=performance \
  --preset=desktop \
  --throttling-method=devtools \
  --max-wait-for-load=45000 \
  --quiet \
  --chrome-flags="--headless=new \
    --user-data-dir=${PROFILE} \
    --disable-extensions \
    --disable-background-networking \
    --disable-component-update \
    --no-first-run --no-default-browser-check"

# Deliberately ABSENT: --no-sandbox and --disable-dev-shm-usage.
#
# Both are container workarounds that most CI recipes copy without asking why.
# --no-sandbox exists because a container often cannot let Chrome set up its own
# namespaces; a microVM has its own kernel, so Chrome's renderer sandbox works
# and there is no reason to switch off the thing standing between a hostile page
# and the rest of the guest.
# --disable-dev-shm-usage exists because a container's /dev/shm defaults to
# 64 MiB and Chrome crashes when it runs out. A full guest kernel mounts
# /dev/shm as tmpfs sized from guest RAM -- half of it by default, so roughly
# two gigabytes here -- and Chrome can use shared memory the way it does on a
# real machine. Keeping the flag means measuring Chrome's slower fallback path.

# Only the performance category. Accessibility is collected separately with
# axe-core, where you control exactly when in the page lifecycle it runs.

The two omitted flags are worth dwelling on, because they are the clearest illustration of the general principle. Every container workaround you carry into your audit environment is a thing you are silently measuring around. `--disable-dev-shm-usage` does not make Chrome correct; it makes Chrome use a slower path for shared memory so it stops crashing in a constrained namespace. If your production users are on real machines with real shared memory, you have been benchmarking a configuration nobody runs.

Fanning out: one VM per URL per run

With the per-run script settled, the orchestration is unremarkable — which is the goal. Create a sandbox, push the script, stream it, pull the JSON back, kill the sandbox.

import json
import shlex
import statistics
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path

from pandastack import Sandbox

URLS = [
    "https://preview-9f3c21a.example.dev/",
    "https://preview-9f3c21a.example.dev/pricing",
    "https://preview-9f3c21a.example.dev/docs/getting-started",
]
RUNS_PER_URL = 5          # 5 so the median means something. 1 is a coin flip.
COMMIT = "9f3c21a"

SCRIPT = Path("run-audit.sh").read_text()


def one_audit(url: str, run: int) -> dict:
    # Note what is NOT passed: cpu= and memory_mb=. The `browser` template has a
    # baked snapshot, and Firecracker cannot change vCPU or RAM at restore, so
    # the agent overrides them with the baked 4 GiB / 8 vCPU regardless. Sizing
    # is a property of the template. Print sbx.memory_mb and you will see 4096
    # whatever you asked for -- which is exactly the property you want from a
    # measurement rig.
    #
    # ttl_seconds is an IDLE timeout (default 5 minutes), not a wall-clock
    # budget. It is the backstop for a wedged run. The kill() in the finally
    # block is what actually stops the meter.
    sbx = Sandbox.create(
        template="browser",
        ttl_seconds=120,
        metadata={"job": "audit", "commit": COMMIT, "url": url, "run": str(run)},
    )
    try:
        sbx.filesystem.write("/workspace/run-audit.sh", SCRIPT)

        # exec_stream, not a long deadline on one-shot exec: one-shot exec has no
        # server-side deadline today and the client gives up at 30 seconds, so a
        # multi-minute audit needs the streaming call. exec_stream honours the
        # deadline, and the `timeout 600` inside the script is the second belt.
        log: list[str] = []
        code = sbx.exec_stream(
            f"sh /workspace/run-audit.sh {shlex.quote(url)}",
            on_stdout=log.append,
            on_stderr=log.append,
            timeout_seconds=900,
        )
        if code != 0:
            return {"url": url, "run": run, "ok": False, "log": "".join(log)[-2000:]}

        lhr = json.loads(sbx.filesystem.read("/workspace/lhr.json"))
        audits = lhr["audits"]
        env = lhr["environment"]
        return {
            "url": url,
            "run": run,
            "ok": True,
            "commit": COMMIT,
            "lighthouse": lhr["lighthouseVersion"],
            "user_agent": env["hostUserAgent"],
            # The validity signal. Keep it next to the metrics, not in a log.
            "benchmark_index": env["benchmarkIndex"],
            "lcp_ms": audits["largest-contentful-paint"]["numericValue"],
            "tbt_ms": audits["total-blocking-time"]["numericValue"],
            "cls": audits["cumulative-layout-shift"]["numericValue"],
            "si_ms": audits["speed-index"]["numericValue"],
        }
    finally:
        sbx.kill()


# Warm the target FIRST. If it scales to zero, the first request of the first
# run pays for the wake, and that lands in TTFB as if it were your frontend.
with Sandbox.create(template="browser", ttl_seconds=60) as warmer:
    probe = warmer.exec(f"curl -s -o /dev/null -w '%{{http_code}}' {shlex.quote(URLS[0])}")
    assert probe.stdout.strip() == "200", f"target not ready: {probe.stdout!r}"

jobs = [(u, r) for u in URLS for r in range(RUNS_PER_URL)]
results = []

# Width here is a fleet-throughput decision, not a measurement decision: each
# audit has its own kernel, so widening the pool does not change what any single
# run sees inside its guest. It does change how busy the HOSTS are, which is why
# every result carries its benchmark index.
with ThreadPoolExecutor(max_workers=12) as pool:
    futures = [pool.submit(one_audit, u, r) for u, r in jobs]
    for fut in as_completed(futures):
        results.append(fut.result())

Path(f"audits-{COMMIT}.json").write_text(json.dumps(results, indent=2))

The aggregation step is where the validity gate earns its place. Group by URL, discard runs whose benchmark index came in below your floor, and only then take medians. Calibrate the floor from your own fleet — collect benchmark indices over a week of normal runs, take something like the tenth percentile of that distribution, and use it. Do not copy a number out of a blog post, including this one; the value is a property of your hardware and will be wrong for mine.

import json
import statistics
from collections import defaultdict
from pathlib import Path

# Calibrate this from YOUR fleet: collect benchmarkIndex across a week of runs
# on an unloaded fleet and take roughly the 10th percentile. A run below the
# floor is not a slow page, it is a bad measurement, and averaging it in is how
# a busy Tuesday afternoon becomes a performance regression in a review comment.
BENCHMARK_FLOOR = 1100
MIN_VALID_RUNS = 3

results = json.loads(Path("audits-9f3c21a.json").read_text())
baseline = json.loads(Path("audits-main.json").read_text())


def medians(rows: list[dict]) -> dict:
    by_url: dict[str, list[dict]] = defaultdict(list)
    for r in rows:
        if not r.get("ok"):
            continue
        if r["benchmark_index"] < BENCHMARK_FLOOR:
            print(f"DISCARD {r['url']} run {r['run']}: "
                  f"benchmarkIndex {r['benchmark_index']:.0f} below floor")
            continue
        by_url[r["url"]].append(r)

    out = {}
    for url, runs in by_url.items():
        if len(runs) < MIN_VALID_RUNS:
            # Not a pass and not a fail. There is no measurement here.
            print(f"INCONCLUSIVE {url}: only {len(runs)} valid runs")
            continue
        out[url] = {
            # Median, not mean: load-time distributions are right-skewed and the
            # mean chases the one run that hit a cold CDN object.
            "lcp_ms": statistics.median(r["lcp_ms"] for r in runs),
            "tbt_ms": statistics.median(r["tbt_ms"] for r in runs),
            "cls": statistics.median(r["cls"] for r in runs),
            "spread_ms": max(r["lcp_ms"] for r in runs) - min(r["lcp_ms"] for r in runs),
            "valid_runs": len(runs),
        }
    return out


head, base = medians(results), medians(baseline)
failed = False

for url, m in sorted(head.items()):
    b = base.get(url)
    if b is None:
        print(f"NEW {url}: lcp {m['lcp_ms']:.0f}ms (no baseline)")
        continue
    delta = m["lcp_ms"] - b["lcp_ms"]
    # Gate on the DELTA against the same URL's baseline, not on an absolute
    # score. Absolute numbers move with every Chrome release; a delta measured
    # the same day on the same template does not.
    #
    # And refuse to call a regression smaller than the observed run-to-run
    # spread. If your spread is 180ms, a 90ms "regression" is a sampling story.
    noise = max(m["spread_ms"], b["spread_ms"])
    verdict = "FAIL" if delta > max(noise, 50) else "ok"
    if verdict == "FAIL":
        failed = True
    print(f"{verdict:4s} {url}: lcp {b['lcp_ms']:.0f} -> {m['lcp_ms']:.0f}ms "
          f"(delta {delta:+.0f}, noise floor {noise:.0f}, n={m['valid_runs']})")

raise SystemExit(1 if failed else 0)

That last comparison is the bit that stops the re-run ritual. A gate that refuses to call anything smaller than the measured run-to-run spread a regression cannot be satisfied by clicking Re-run, because there is nothing to click past. If the spread is wide enough to hide real regressions, the gate tells you that too, and the fix is to reduce the spread rather than to lower the threshold.

Accessibility auditing has a completely different cost model

It would be convenient to claim that performance and accessibility auditing are the same problem, and plenty of tooling pitches them that way because they ship in the same CLI. They are not the same problem, and pretending otherwise leads you to over-provision one and under-build the other.

axe-core runs as JavaScript inside the page, walks the accessibility tree and evaluates a rule set against the DOM. It is cheap, it is fast, and critically it is deterministic given the same DOM. A contended CPU makes an axe run take longer; it does not make it report a different answer. Nobody has ever filed a flaky-colour-contrast bug, because contrast is a ratio between two numbers rather than a race against a clock.

The two halves of a web audit pipeline have different bottlenecks, and the isolation buys you different things.
Performance (Lighthouse)Accessibility (axe-core)
Nature of the resultA timing measurementA rule evaluation over a DOM
Sensitive to CPU contentionYes — it is the outputNo — only to total runtime
DeterministicNo, needs repeats and a medianYes, given the same DOM
Runs per URLFive or moreOne is enough
URLs per microVMOne — sharing reuses page cacheMany — reuse the browser
Real bottleneckCPU time and run countCrawling, auth state and route coverage
Why isolateMeasurement validityUntrusted third-party JavaScript
Main failure modeNoise read as a regressionAuditing a page before it settled

So build it differently. For accessibility the bottleneck is not CPU, it is URL discovery: finding every route, getting past login, reaching the states that only exist after an interaction. Run many URLs inside one longer-lived microVM, reuse the browser process, and spend your engineering effort on the crawler and on authentication rather than on parallelism you do not need.

Authenticated crawls are where a snapshot earns its keep on this side too. Log in once, snapshot the sandbox with the session cookies and local storage live in guest memory, then fork that snapshot for each crawl worker. A same-host fork lands in 400 to 750 milliseconds, cross-host in 1.2 to 3.5 seconds, and each worker starts already logged in without replaying a login flow that will eventually break because somebody added a CAPTCHA.

#!/bin/sh
# crawl-and-axe.sh -- MANY urls, ONE microVM. The opposite allocation to the
# Lighthouse script, because the cost model is the opposite.
#
# axe-core is deterministic given the same DOM, so repeats buy nothing. What
# buys something is route coverage, which means a crawler and a login.
set -eu

AXE_VERSION="4.10.2"
npm i -g --silent "@axe-core/cli@${AXE_VERSION}"

CHROME_PATH="$(node -e "console.log(require('playwright').chromium.executablePath())")"
export CHROME_PATH

mkdir -p /workspace/axe

# One axe process, many URLs. --tags scopes the rule set to the conformance
# level you are actually claiming; running every rule and then arguing about
# which ones count is how an accessibility programme stalls.
while IFS= read -r url; do
  slug="$(printf '%s' "$url" | tr -cs 'a-zA-Z0-9' '-' | cut -c1-80)"
  timeout 180 axe "$url" \
    --chrome-path "$CHROME_PATH" \
    --tags wcag2a,wcag2aa,wcag21a,wcag21aa \
    --save "/workspace/axe/${slug}.json" \
    --exit || echo "violations: $url" >> /workspace/axe/failed.txt
done < /workspace/urls.txt

# Honest caveat for whoever reads the report: automated rules cover a minority
# of WCAG success criteria. Keyboard trap, focus order that makes sense, alt
# text that is actually descriptive, heading structure that matches the visual
# hierarchy -- none of those are decidable by a rule engine. Check Deque's
# current documentation for their own coverage figure rather than trusting a
# number in a blog post, and budget for manual testing regardless.

The caveat in that comment is not decoration. Automated accessibility testing catches a minority of WCAG failure modes — the published estimates vary and you should check the tool vendor's current documentation rather than a figure quoted second-hand — and the failures it cannot see are disproportionately the ones that make a site genuinely unusable. A green axe run means you have no machine-detectable violations. It does not mean the site is accessible, and a dashboard that implies otherwise is worse than no dashboard.

Making the network condition uniform too

Isolating the CPU and the starting state fixes most of the variance. The network is the part people forget, because Lighthouse's throttling makes it feel handled and it is not. Lighthouse shapes the simulated client link. It does nothing about the real path from your audit VM to the origin: which host the VM landed on, what that host's NIC was doing, how far it is from the origin, and whether some other tenant on the same box was pulling a container image at that moment.

  • Shape the sandbox's own egress so every run sees the same ceiling. A token-bucket qdisc on the host side of the sandbox's veth pair gives you a hard, uniform bandwidth and latency floor, which is a different and stronger guarantee than Lighthouse's model of a slow link.
  • Keep the whole audit fleet in one region, so round-trip time to the origin is comparable between runs rather than between continents.
  • Record the host each run landed on. If one agent produces systematically worse numbers, that is a fleet fact worth knowing and not a property of your frontend.
  • Audit a stable artefact, not a moving target. A CDN that was cold for run one and warm for run five has given you two experiments.

What this does not fix

A microVM is not a dedicated physical core. The 8 vCPU on the `browser` template are burst capacity, and under contention the host's cgroup weights share real cores between guests. If another audit VM on the same host is mid-Lighthouse, your run is still competing for silicon. What the microVM removes is everything inside the boundary: the shared kernel, the shared page cache, the shared /tmp, the shared Chrome profile, the shared font set, the neighbour's dirty-page writeback stalling your read.

That distinction decides what you can honestly claim. If you need absolute numbers — "our LCP is better than our competitor's" — you need dedicated hardware and a carefully specified client profile, and even then the claim is fragile. If you need to detect a regression between commit A and commit B, you need identical conditions, and identical conditions are precisely what a snapshot plus a per-run kernel plus a recorded benchmark index gives you. Relative, same-day, same-template comparison is the question this architecture answers well. Absolute truth about the real web is not, and no CI pipeline has ever answered it.

A performance budget that passes on the second attempt was never a budget. It was a lottery with a pass/fail interface.

The security angle, stated as secondary

This post is about measurement, so the security property goes last — but it should not go unmentioned, because people build audit pipelines without noticing what they have built.

Pointing a real browser at a URL means executing that URL's JavaScript, with the full attack surface of a modern browser engine, on your infrastructure, with your credentials and your cloud metadata endpoint one request away. If the URLs come from your own repository that is a modest risk. If they come from customers — a site-audit product, an SEO tool, a monitoring service, anything with a "paste your URL" box — then you are running untrusted third-party code at scale, whether or not you ever wrote that sentence down in a design document.

The ugly detail is that the standard CI recipe makes this worse. `--no-sandbox` appears in approximately every Dockerised Chrome snippet on the internet, because the container would not let Chrome set up its own namespaces, and switching off the renderer sandbox was easier than fixing the container. That turns a renderer compromise directly into code running as your CI user, inside a kernel shared with every other job on the box. In a microVM you generally do not need the flag at all, and if you keep it anyway you still have a hypervisor boundary underneath — which is the difference between a bad afternoon and an incident report.

What it costs, and the mistake that doubles it

The arithmetic is simple enough to do in your head. The `browser` template commits 4 GiB, billed at $0.0162 per GiB-hour, so memory costs you a touch under seven cents per VM-hour. CPU is billed on active CPU-seconds actually burned at $0.054 per vCPU-hour, and a Lighthouse run is a short burst of real work rather than a steady draw — so for a per-audit VM the bill is dominated by how long the VM exists, not by how hard it worked. Egress is metered separately, which matters if your audits are pulling heavy pages at volume.

Which leads to the one operational mistake that reliably doubles the bill: letting the idle reaper clean up. `ttl_seconds` measures time since last activity, defaults to five minutes, and will dutifully hold a finished audit VM open for those five minutes while it waits to be sure. Thirty URLs times five runs times five idle minutes is twelve and a half VM-hours of committed memory spent on nothing. Call `kill()` in a `finally` block and keep the TTL purely as a backstop for a wedged run.

The honest closing position is that none of this is exotic. It is one VM per measurement, a pinned browser, five runs and a median, a validity field you were already being handed and probably ignoring, and a gate that compares deltas instead of absolutes. The reason it is worth doing is not that it makes your numbers prettier — it will often make them slightly worse, because the contended host was flattering your simulated throttling in ways you did not want. It is that when the gate fails, you will believe it, and when it passes you will not feel the urge to click Re-run.

Frequently asked questions

Why do Lighthouse scores vary between runs on the same code?

Because the score is a composite of timing measurements, and timings depend on the machine. Four causes dominate. CPU contention: Lighthouse's main-thread metrics, especially total blocking time, are close to a direct readout of how busy the host was. Page cache state: a second run finds binaries, fonts and modules already resident, so it genuinely starts faster than the first. Network variance: DNS, TLS, which CDN edge answered, whether the origin had just scaled. And browser state: a leftover profile, a service worker, an extension, or a Chrome that updated itself overnight. There is also a subtler cause that catches people out. Lighthouse's default simulated throttling models a slower device, and it calibrates that model against a synthetic CPU benchmark it runs on the host at measurement time. On a contended host the trace is already degraded and the calibration constant is measured during the same contention, so the two error sources are correlated. The practical response is to isolate each run, repeat it at least five times, take the median rather than the mean, and record the benchmark index alongside every result so you can discard runs taken while the machine was busy.

Why run each audit in a microVM instead of a container?

Because a container shares the host's kernel and page cache, and both of those are inputs to a performance measurement. Concretely: your run's scheduling decisions are made by a kernel also serving other tenants; "cold cache" means whatever the host happened to be caching, which depends on what ran before you; and a neighbour's dirty-page writeback can stall your reads with no trace in your metrics. A microVM restored from a snapshot gives you something a container structurally cannot — a byte-identical starting state that includes memory, so run one and run four hundred begin from the same bytes: same cold page cache, same empty temp directory, same font set, same zero extensions. That last one matters more than it sounds: installed fonts change text metrics, which changes line breaks, which can change which element is the largest contentful paint. A base-image bump pulling a different font package looks exactly like a regression. The honest limit is that a microVM is not a dedicated physical core, so another guest on the same host still competes for silicon. What it removes is everything inside the isolation boundary, which is where most of the variance lives.

How many Lighthouse runs do I need, and should I use the mean or the median?

Median, and at least five runs. Load-time distributions are right-skewed: most runs cluster and a few land far out because of a garbage collection pause, a cold CDN object or a retried request. The mean follows those outliers, which is exactly the wrong behaviour for a regression detector — one unlucky run moves the number and you go bisecting a commit that changed nothing. The median ignores them. Three runs is the minimum at which a median means anything; five is a reasonable default; more helps if your spread is wide. Two further rules matter as much as the count. First, compare medians taken the same day on the same template and the same browser version, not against a number from last month, because absolute scores move with Chrome releases. Second, gate on the delta against a baseline rather than on an absolute threshold, and refuse to call anything smaller than your measured run-to-run spread a regression. If the spread is 180 ms, a 90 ms movement is a sampling story, and a gate that fails on it will simply train the team to click Re-run until it passes.

Can I run Lighthouse and axe-core in the same pipeline stage?

You can, but you should not treat them as the same problem, because their cost models are opposites. Lighthouse produces a timing measurement, so it is sensitive to CPU contention, needs several repeats per URL to be meaningful, and wants one isolated microVM per run. axe-core evaluates a rule set against a DOM, so it is deterministic given the same DOM, needs one run per URL, and happily does many URLs inside one longer-lived VM with the browser process reused. Contention makes an axe run slower, not wrong. So the accessibility half's bottleneck is not CPU at all — it is route discovery, authentication and reaching states that only exist after an interaction. Spend the effort on the crawler, not on parallelism. A snapshot helps here in a different way: log in once, snapshot the sandbox with the session live in guest memory, then fork it for each crawl worker so every worker starts already authenticated. Same-host forks land in 400 to 750 milliseconds. Keep the two halves as separate jobs with separate gates, and keep the accessibility gate honest about coverage.

Is a green axe-core run the same as being WCAG compliant?

No, and treating it as such is the most common failure in automated accessibility programmes. Automated rules cover a minority of WCAG success criteria — the published coverage estimates vary, so check the tool vendor's current documentation rather than a figure quoted in a blog post — and the criteria they cannot evaluate are disproportionately the ones that decide whether a site is actually usable. A rule engine can tell you a contrast ratio is below threshold or that an image has no alt attribute. It cannot tell you whether the alt text is descriptive, whether focus order follows the visual hierarchy, whether a custom widget traps the keyboard, whether headings reflect the real document structure, or whether an error message is announced in a way a screen reader user can act on. A green run means you have no machine-detectable violations in the states you crawled, which is a useful and worth-automating floor. Report it as that. Scope the rule set with explicit conformance tags so you are claiming a specific level, keep a record of which URLs and which interaction states were actually visited, and budget for manual and assistive-technology testing on top.

Keep reading

Related posts

More in CI & ephemeral environments · See Ephemeral CI runners on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.