all posts

kvm-clock vs TSC: How a Firecracker Guest Tells the Time, and What a Snapshot Does to It

Ajay Kumar··10 min read

I build PandaStack, an open-source Firecracker microVM platform, and the strangest bug report I have ever had to triage read like this: every HTTPS request out of a restored sandbox fails with `certificate is not yet valid`. Not expired — not yet valid. The certificate chain was fine. The CA bundle was fine. DNS was fine. The clock was not.

The guest had been restored from a snapshot taken a week earlier, and it woke up convinced the week had not happened. Every TLS peer it contacted presented a certificate issued after the guest's idea of `now`, and OpenSSL did exactly what it is supposed to do with a certificate from the future: refuse it. That was a real incident on our platform, not a thought experiment, and the fix was not in the TLS stack.

To understand why a restore breaks time you have to know how a Linux guest tells the time in the first place, and that turns out to be a small, genuinely elegant subsystem with a rating contest, a shared memory page, and a performance cliff hidden inside it. So: how a guest reads a clock, what each clocksource actually costs per read, what a snapshot does to all of it, and which of the available fixes are real versus which merely make the bug intermittent.

The short version. Linux picks a clocksource by rating, and on a KVM guest `kvm-clock` wins: the host publishes a scale and an offset into a guest page, the guest reads the TSC and applies it, and `clock_gettime` never leaves userspace. A snapshot freezes all of that. On restore, `CLOCK_MONOTONIC` has not advanced and `CLOCK_REALTIME` is simply stale, and the only mechanism that genuinely fixes it is the layer that knows the real time — the host — telling the guest. Everything else either hides the bug or turns it into a race.

The clocksource subsystem: a rating contest you can watch

Linux does not have "the clock". It has a list of registered clocksources, and a clocksource is a deliberately boring object: a function that returns a monotonically increasing count of ticks, a mask saying how many bits of that count are real before it wraps, a `mult` and `shift` pair for converting ticks to nanoseconds with integer arithmetic, and a `rating`. The timekeeping core sorts the list by rating and uses the best one. Everything else in the kernel's notion of time — `ktime_get`, the scheduler's view of elapsed time, hrtimer deadlines, the seconds in your log lines — is that one counter plus bookkeeping.

Two sysfs files expose the result, and they are the first thing to read when time looks wrong inside a guest. `available_clocksource` lists the registered candidates in rating order, best first. `current_clocksource` names the winner, and writing a name into it switches clocksources at runtime. On a Firecracker guest running the 5.10 kernel we ship, you will typically see exactly two entries — `kvm-clock` and `tsc` — with `kvm-clock` current. That short list is not an accident, and the reason it is short is the most useful thing in this section.

kvm-clock: the host does the arithmetic for you

`kvm-clock` is a paravirtual clocksource, which means the guest and the host cooperate instead of the guest guessing. At boot the guest allocates a page, registers its physical address with the host by writing a KVM-specific MSR, and from then on the host keeps a small structure in that page up to date: a reference TSC value, the system time in nanoseconds that corresponds to it, a multiplier and shift to convert TSC ticks into nanoseconds, a flags field, and a version counter used as a seqlock so the guest can tell it read a torn update and retry.

Reading it is then: bump through the seqlock, execute `RDTSC`, subtract the reference TSC, scale the delta by the published multiplier and shift, add the published system time. That is a handful of instructions and a few loads from a page that is already warm. There is no VM exit, because `RDTSC` is not trapped in the normal configuration and the page is ordinary guest memory. A separate MSR supplies a wall-clock base once, at boot, which is how the guest learns what century it is.

The reason `kvm-clock` outranks a raw TSC read is correctness under things the guest cannot see. If the host migrates the vCPU to a core whose TSC is offset differently, or rescales the TSC, or the guest is restored on a machine with a different TSC frequency, the host can republish the transform and the guest's nanoseconds stay coherent. The guest never learns that anything happened; it just keeps reading the page. That is also exactly the property that makes the snapshot problem so slippery — `kvm-clock` is built to make discontinuities invisible, and a week-long gap is a discontinuity.

tsc: faster, and correct only if several things are true

The `tsc` clocksource skips the paravirtual layer and reads the Time Stamp Counter directly, applying a mult/shift the kernel calibrated at boot. It is the cheapest possible clock read: one instruction plus arithmetic. It is also the one with preconditions, and the preconditions are CPU features rather than configuration.

You want `constant_tsc` (the counter ticks at a fixed rate regardless of the core's actual frequency) and `nonstop_tsc` (it keeps ticking through deep C-states). Without both, a TSC read is a measurement of something that changes underneath you, and the kernel is right not to trust it. Those flags arrive through CPUID, which under Firecracker is filtered — the VMM masks and normalises what the guest sees, which is the right call for snapshot portability and also means the flag set in the guest is a platform decision, not a hardware fact. Firecracker CPUID masking, explained covers that machinery.

There is a second mechanism worth knowing: the clocksource watchdog. Periodically the kernel reads the current clocksource and a lower-rated reference and checks they agree. If they disagree beyond a threshold it prints `Clocksource tsc unstable` and demotes to the next candidate. In a guest with only two clocksources this is thin protection, and if you pin the TSC with `tsc=reliable` you have switched the watchdog off and taken responsibility for the outcome yourself.

acpi_pm and hpet: the clocks that cost a VM exit

On a conventional x86 machine the fallbacks below the TSC are the ACPI power-management timer and the HPET. `acpi_pm` is a 24-bit counter at a fixed 3.579545 MHz read through an I/O port. `hpet` is a memory-mapped counter read through MMIO. Both are hardware the guest does not own, so in a virtual machine every single read is a trap out to the VMM, handled, and returned — the mechanics of which are in VM exits: the actual currency of virtualization overhead.

That is a disaster at any serious read rate, and "serious" is lower than people think: a request-logging middleware, a metrics library sampling a histogram, a Go runtime scheduling goroutines, a Python app calling `time.monotonic()` in a loop. A workload that was CPU-bound on arithmetic becomes CPU-bound on the exit path, and nothing in your profile points at the clock.

The good news under Firecracker is that this failure mode is mostly unavailable to you, because the device model is deliberately tiny. Firecracker does not emulate an HPET, and the x86_64 machine it builds has no conventional CMOS RTC either — there is a serial port, a keyboard controller stub used for reset, and virtio devices, and that is approximately the list. Fewer emulated timers is fewer places for a guest to find a slow clock, and fewer lines of VMM code reachable from the guest. Check the device list for your own Firecracker version rather than trusting my summary, but the shape of it is: do not write code that assumes an RTC chip is there to re-read.

How each clocksource is read, what it costs, and how it behaves across a snapshot or migration.
ClocksourceHow a read worksCost per readvDSO fast pathAcross snapshot / migration
kvm-clockSeqlock read of a host-written page, RDTSC, scale and offsetA few instructions and warm loads; no VM exitYes — userspace reads the pvclock page directlyDesigned to stay coherent: the host republishes the transform. Monotonic time is continuous, which is why a long gap is invisible rather than obvious.
tscA single RDTSC plus the kernel's calibrated mult/shiftCheapest available; one instructionYesDepends entirely on the host restoring the TSC offset and frequency. Coherent in practice on a single host; the frequency question is real across heterogeneous hosts.
acpi_pmI/O port read of a fixed-rate 24-bit counterA VM exit on every readNo — falls back to a real syscallStateless and therefore unbothered, which is the only nice thing about it. Still a fixed-rate counter with no idea what day it is.
hpetMMIO read of a memory-mapped counterA VM exit on every readNo — falls back to a real syscallDevice state is part of the snapshot. Not emulated by Firecracker at all, so under Firecracker this row is a thing you will not see.

Three POSIX clocks, three different ways to be wrong

Above the clocksource sits the interface your code actually calls, and it matters which clock id you asked for, because a restore damages them differently.

  • `CLOCK_REALTIME` is wall-clock time: seconds since the Unix epoch, settable, subject to NTP steps and leap-second smearing. This is what `date`, `gettimeofday()`, log timestamps, certificate validity checks and JWT `exp` claims read. After a restore it is simply stale — it says the snapshot's time, and it is confidently wrong.
  • `CLOCK_MONOTONIC` counts from an arbitrary origin, never goes backwards, and is not affected by anyone setting the wall clock. It is the correct clock for measuring a duration, a timeout or a backoff. After a restore it has not advanced across the gap — which is arguably the right answer, because from the guest's point of view no time passed, and that is also why it is useless for noticing that a week did.
  • `CLOCK_BOOTTIME` is `CLOCK_MONOTONIC` plus time spent suspended, and people reach for it expecting it to cover this case. It does not. A Firecracker snapshot restore is not a suspend/resume cycle that the guest participated in; there is no suspend event to account for, so `BOOTTIME` is frozen alongside `MONOTONIC`.
  • `CLOCK_MONOTONIC_RAW` skips NTP's frequency adjustment and is the one to use when you are measuring the clock itself rather than measuring with it — which is exactly what you are doing while debugging this.

And now the part that makes clocksource choice a performance decision and not a trivia question: on Linux, `clock_gettime` is normally not a syscall. The vDSO maps a page of kernel-provided code and timekeeping data into every process, and the libc wrapper reads the clocksource and does the arithmetic entirely in userspace. No `syscall` instruction, no ring transition, nothing. That fast path exists only for clocksources the kernel marks as vDSO-capable, which on x86 means the TSC and the paravirtual clock. Pick a clocksource outside that set and `clock_gettime` quietly becomes a real syscall that performs a real VM exit — the same function call, now a trap, with no diff and no log line to tell you.

That is the actual cost model to carry around. The difference between a good and a bad guest clocksource is not "a bit slower". It is a change in the class of operation, from arithmetic on a warm page to a round trip through the hypervisor, applied to one of the most frequently called functions in your process.

#!/usr/bin/env bash
# Run this INSIDE the guest. Five minutes of reading beats an afternoon of
# guessing which clock your VM is on and what it costs you.
set -euo pipefail

CS=/sys/devices/system/clocksource/clocksource0

# 1. Who is registered, and who won? available_clocksource is printed in
#    rating order, best first. On a Firecracker guest this list is SHORT --
#    typically just "kvm-clock tsc" -- because Firecracker does not emulate
#    an HPET or an ACPI PM timer. A short list is a feature.
cat "$CS/available_clocksource"      # e.g. kvm-clock tsc
cat "$CS/current_clocksource"        # e.g. kvm-clock

# 2. Is a direct TSC read even defensible on this guest? These two flags are
#    the preconditions: constant_tsc = fixed rate regardless of core
#    frequency, nonstop_tsc = keeps ticking through deep C-states. They come
#    through CPUID, which Firecracker filters -- so what you see here is a
#    platform decision, not a hardware fact.
grep -m1 '^flags' /proc/cpuinfo | tr ' ' '\n' \
  | grep -E '^(constant_tsc|nonstop_tsc|tsc_known_freq|tsc_reliable|tsc_adjust|tsc_deadline_timer)$' \
  | sort

# 3. Did the kernel ever distrust its clock? The watchdog cross-checks the
#    current clocksource against a lower-rated one and demotes on
#    disagreement. These lines are gold during an incident.
dmesg 2>/dev/null | grep -iE 'clocksource|tsc:|kvm-clock' | tail -20 || true

# 4. Switch at runtime (root). You may only write a name that appears in
#    available_clocksource. This is the cheap way to A/B the cost below.
#    Permanently: add clocksource=tsc to the kernel command line. Related
#    boot params worth knowing: no-kvmclock (disable the paravirtual clock),
#    tsc=reliable (pin the TSC AND switch the watchdog off -- you now own
#    the consequences), clocksource=hpet (only if an HPET exists at all).
# echo tsc > "$CS/current_clocksource"

# 5. The vDSO demonstration, which is the whole point. clock_gettime on a
#    vDSO-capable clocksource (tsc, kvm-clock) never enters the kernel.
#    strace counts syscalls: expect ZERO clock_gettime calls here.
strace -c -f -e trace=clock_gettime \
  python3 -c 'import time
for _ in range(2_000_000): time.monotonic()' 2>&1 | tail -5

# 6. And the wall-clock cost, for the A/B. Run it on kvm-clock, write a
#    different name into current_clocksource, run it again.
for src in $(cat "$CS/available_clocksource"); do
  echo "$src" > "$CS/current_clocksource" 2>/dev/null || { echo "skip $src"; continue; }
  printf '%-12s ' "$src"
  # timeout(1) because a pathological clocksource makes this loop crawl, and
  # you want a bounded experiment rather than a hung shell.
  timeout --kill-after=5s 60 python3 -c 'import time
t0=time.perf_counter_ns()
for _ in range(2_000_000): time.monotonic()
print((time.perf_counter_ns()-t0)//2_000_000, "ns per read")'
done

# If your guest only offers kvm-clock and tsc you cannot reproduce the slow
# case at all. That is not a failed experiment -- it is the device model
# refusing to hand you a footgun.

A snapshot freezes time, and nothing in the guest objects

A Firecracker snapshot captures guest memory and device state at an instant. Time is state. The pvclock page is in guest memory. The vCPU's TSC value is in the saved register state. The guest's wall-clock base — seeded once at boot and never re-read — is in a kernel data structure in that same memory. All of it is frozen together, consistently, which is precisely the problem: nothing is inconsistent enough for the guest to notice.

When you restore, the guest resumes as though a pause button were pressed and released. `CLOCK_MONOTONIC` continues from exactly where it stopped, which is defensible — no time passed for the guest. `CLOCK_REALTIME` continues from the snapshot's wall-clock value, which is not defensible at all, and is the one everything reads. The restore mechanics, including why we restore rather than cold-boot on every create, are in The Snapshot-Restore Boot Path: Every Sandbox in Under 200ms.

The consequences are not evenly distributed, and the order matters because it determines what your error messages will say:

  • TLS breaks first and loudest. A restored guest whose clock sits a week in the past sees perfectly valid certificates as not-yet-valid; one restored from a months-old snapshot sees them as expired. Either way the error names the certificate, so the first three hours of debugging go into the CA bundle. This is the one that bit us in production, and it is the reason the rest of this post exists.
  • JWT and OIDC validation fails on `nbf` and `exp`. The guest rejects freshly minted tokens as not-yet-valid, or — worse, quieter — accepts tokens that revoked an hour ago because to the guest they have not expired yet. One of those is an outage and the other is a security finding.
  • Everything with a window computes nonsense: cache TTLs, rate-limit buckets, lease expiries, retry backoff, circuit-breaker timers. Most of these are built on `CLOCK_MONOTONIC` and so behave "correctly" relative to a guest clock that spans a week of real time — your 30-second health-check interval is still 30 guest-seconds, and the lease you were renewing expired six days ago on the server that actually counts.
  • Timers bunch. Anything armed against an absolute wall-clock deadline — `TIMER_ABSTIME` on `CLOCK_REALTIME`, a `timerfd` with `TFD_TIMER_CANCEL_ON_SET`, a cron-shaped scheduler — becomes due all at once the moment anyone steps the clock forward. The fix for the clock is the trigger for the stampede.
  • `systemd-timesyncd` or `chrony` eventually corrects it, which is the worst outcome available. The bug does not disappear; it becomes a race between your application's first outbound TLS handshake and the time daemon's first successful sync. You will reproduce it about one time in nine, on someone else's machine, during a demo.
  • Log timestamps from a restored guest are wrong and do not line up with the host's, so the one artefact you would use to reconstruct the incident is itself a victim of the incident.
If your platform lets a time daemon fix this asynchronously after resume, you have not fixed the bug — you have converted a deterministic failure into an intermittent one, and moved it from your test suite into production. Any correction has to land before the guest's first request, which means it has to be part of the resume path rather than a service that starts during it.
/* threeclocks.c -- the entire lesson in three numbers.
 *
 *   cc -O2 -o threeclocks threeclocks.c
 *
 * Run it, snapshot the guest, wait, restore, run it again. REALTIME will be
 * stale by the length of the gap. MONOTONIC and BOOTTIME will have advanced
 * only by however long the guest actually ran -- the gap is simply absent
 * from both, which is why neither can be used to detect it.
 */
#include <stdio.h>
#include <time.h>

static void show(const char *name, clockid_t id) {
    struct timespec ts;
    if (clock_gettime(id, &ts) != 0) { perror(name); return; }
    printf("%-22s %10lld.%09ld\n", name, (long long)ts.tv_sec, ts.tv_nsec);
}

int main(void) {
    /* CLOCK_REALTIME: seconds since the epoch. Settable. Everything
     * human-facing reads this -- date(1), certificate validity, JWT exp.
     * After a restore this is the stale one. */
    show("CLOCK_REALTIME", CLOCK_REALTIME);

    /* CLOCK_MONOTONIC: arbitrary origin, never steps, immune to anyone
     * setting the wall clock. Correct for durations and timeouts. Does NOT
     * advance across a snapshot gap. */
    show("CLOCK_MONOTONIC", CLOCK_MONOTONIC);

    /* CLOCK_BOOTTIME: MONOTONIC plus time spent suspended. People expect
     * this one to cover a snapshot restore. It does not: the guest never
     * participated in a suspend, so there is no suspended interval to add. */
    show("CLOCK_BOOTTIME", CLOCK_BOOTTIME);

    /* MONOTONIC_RAW skips NTP's frequency correction -- the clock to use
     * when you are measuring the clock rather than measuring with it. */
    show("CLOCK_MONOTONIC_RAW", CLOCK_MONOTONIC_RAW);

    /* Print the clocksource too, so the three numbers above come with the
     * name of the thing that produced them. */
    FILE *f = fopen("/sys/devices/system/clocksource/clocksource0/current_clocksource", "r");
    if (f) {
        char buf[64] = {0};
        if (fgets(buf, sizeof buf, f)) printf("%-22s %s", "clocksource", buf);
        fclose(f);
    }
    return 0;
}

The fixes, ranked by how much you should trust them

There is exactly one mechanism that fully solves this, and it is structural rather than clever: the layer that knows the real time has to tell the guest, as part of resuming it, before the guest does anything. Everything else is a partial measure, and it is worth being clear about which is which.

  1. The host pushes the time in on resume. The VMM and its control plane know the wall-clock time; the guest does not and cannot work it out. Have the orchestrator set the guest clock as a step in the restore path, synchronously, before the guest is reachable. This is the only option with no race in it, and it is what PandaStack does on restore, resume and wake.
  2. Re-base the paravirtual clock from the host. Because `kvm-clock`'s transform lives in a page the host writes, the host can republish it — this is the same machinery that keeps a guest coherent across vCPU migration. It is the most architecturally correct place to intervene; it is also hypervisor-version-dependent and the kind of thing to verify on your own build rather than assume from a blog post.
  3. Step the clock from inside the guest with a value handed in from outside. `clock_settime(CLOCK_REALTIME, ...)` or `settimeofday()` from a tiny agent in the guest, fed the host's time over vsock or an exec channel, is unglamorous and works. `hwclock --hctosys` is the traditional incantation and is exactly what you should not reach for under Firecracker, because there may be no RTC to read from.
  4. Make the time daemon step rather than slew on startup. `chrony` slews small offsets by design, which is correct for drift and useless for a week-long jump. `makestep` with a large threshold and a small limit lets it step hard on the first few updates; `chronyd -s` and the `initstepslew`-style behaviour in older configurations exist for the same reason. This makes the correction fast. It does not make it synchronous, so it remains a race with your first request — use it as a backstop, never as the fix.
  5. `ptp_kvm`. The guest kernel can expose a PTP hardware clock device backed by a hypercall that reads the host's clock, and a time daemon can then use it as a local reference with no network involved — which is genuinely attractive. It depends on both the guest kernel module and the hypercall being available to your guest, so treat it as something to test on your exact kernel and VMM version rather than a configuration you can copy.

If you run Firecracker yourself, items one through three are your responsibility and nobody else's. Firecracker's job is to restore the guest faithfully, and it does that correctly — the TSC offset is reprogrammed so monotonic time is continuous, which is the right behaviour for a resume. No layer of that stack has any idea what today's date is. Wall-clock correction is an orchestration concern, and if your orchestrator does not do it, the guest will boot into the past and the first thing you will hear about it is a certificate error.

What we do, and what we still do not fix

Every PandaStack create is a snapshot restore — there is no warm pool of idle VMs, so the restore path is the normal path, at p50 179 ms and p99 203 ms. That makes the frozen clock not an edge case but something that would happen on every single sandbox, which is how we found it the hard way. The guest clock is now force-synced on restore, on resume and on wake, as part of bringing the VM up rather than as a service that starts inside it.

Two honest limits. First, this is a step, not a rewrite of history: `CLOCK_MONOTONIC` in a restored guest still does not account for the gap, because nothing in the guest's state records one. If your application derives anything durable from monotonic time across a snapshot boundary, it is still wrong and we cannot fix that for you. Second, the clock is one member of a family of state that encodes `now`, and the others are still yours to handle — which is the next section.

# Demonstrate the frozen-clock problem and the fix, end to end.
#
# Run this, go and have lunch, then run it again with the printed snapshot id
# in SNAP. The three clocks before and after tell the whole story.

import os, sys
from pandastack import Sandbox

THREE = (
    "python3 -c \"import time;"
    "print('realtime ', time.time());"
    "print('monotonic', time.monotonic());"
    "print('boottime ', time.clock_gettime(time.CLOCK_BOOTTIME))\""
)
CLOCKSRC = "cat /sys/devices/system/clocksource/clocksource0/current_clocksource"

SNAP = os.environ.get("SNAP")

if not SNAP:
    # --- Phase 1: bake a snapshot with a known clock reading inside it ------
    sbx = Sandbox.create(
        template="base",         # 4 GiB / 8 vCPU, baked into the snapshot:
                                 # Firecracker cannot change vCPU or RAM at
                                 # restore, so cpu=/memory_mb= would be
                                 # overridden to the baked values anyway.
        ttl_seconds=15 * 60,     # IDLE clock, not a wall clock. Irrelevant
                                 # here because we tear down explicitly.
        metadata={"demo": "clocksource"},
    )
    try:
        print("clocksource:", sbx.exec(CLOCKSRC).stdout.strip())  # kvm-clock
        print("--- before snapshot ---")
        print(sbx.exec(THREE).stdout)

        # snapshot() is synchronous and writes the full guest memory image,
        # so budget tens of seconds on a 4 GiB guest. Time stops here.
        snap = sbx.snapshot()
        print("snapshot:", snap)
        print("now wait a while, then re-run with SNAP=" + snap)
    finally:
        sbx.kill()               # teardown is kill(); there is no .delete()
    sys.exit(0)

# --- Phase 2: restore it later and read the clocks again --------------------
later = Sandbox.create(from_snapshot=SNAP, metadata={"demo": "clocksource"})
try:
    print("--- after restore ---")
    print(later.exec(THREE).stdout)
    # realtime  : correct, because PandaStack force-syncs the guest clock on
    #             restore/resume/wake. On raw Firecracker this would read the
    #             snapshot's wall-clock time and you would be in the past.
    # monotonic : has NOT advanced by the gap -- and never will. The guest has
    #             no record that any time passed, which is arguably correct
    #             and definitely not what your lease renewal assumed.
    # boottime  : frozen alongside monotonic. There was no suspend event for
    #             it to account for, so it does not rescue you here.

    # The TLS canary. This succeeds here BECAUSE the clock was stepped before
    # the guest became reachable. Without that step, a week-old snapshot gets
    # "certificate is not yet valid" and a months-old one gets "expired" --
    # and the error names the certificate, not the clock, which is why this
    # costs people an afternoon.
    #
    # timeout(1) in-guest because one-shot exec() does not actually enforce
    # timeout_seconds -- bound anything that can hang on the guest side.
    r = later.exec(
        "timeout --kill-after=5s 20 curl -sS -o /dev/null -w '%{http_code} %{time_total}s\\n' "
        "https://api.github.com/zen"
    )
    print("tls:", r.stdout.strip() or r.stderr.strip(), "exit", r.exit_code)

    # Prove the clock really is the variable: break it deliberately and watch
    # the same curl fail in the way the incident failed.
    later.exec("date -s '2019-01-01 00:00:00' >/dev/null 2>&1 || true")
    bad = later.exec("timeout --kill-after=5s 20 curl -sS https://api.github.com/zen")
    print("tls with a 2019 clock:", (bad.stderr or bad.stdout).strip()[:160])
finally:
    later.kill()

A snapshot is a time machine, and the clock is not the only passenger

Here is the general lesson, and it is worth more than the clock-specific fix. A snapshot is a time machine, and every piece of state inside it that encodes "now" or "this instance is unique" is a bug waiting for a restore. The clock is simply the loudest member of the family, because TLS fails immediately and in public. The quiet ones are worse.

  • Entropy and RNG state. Every child restored from one snapshot starts with the same kernel entropy pool and the same seeded userspace PRNGs — so two sandboxes can generate the same session token. Firecracker exposes a generation identifier on restore precisely so the guest can reseed; the full treatment is The Snapshot Clone Randomness Problem. Note the asymmetry in our own API: `fork()` clones the disk and the child cold-boots with its own entropy, while `fork_tree()` inherits the parent's running memory and therefore inherits this problem.
  • Open TCP connections. The peer closed them hours ago, or a NAT table forgot them, or the sequence numbers are now meaningless. The guest holds sockets it believes are established and will wait out a retransmit timeout to find out otherwise.
  • DHCP leases and anything else with a negotiated expiry. The lease in the guest expired during the gap; the guest thinks it has hours left.
  • Cached DNS with TTLs. A resolver cache restored from a snapshot is a set of answers with TTLs measured against a monotonic clock that did not move, so it will cheerfully serve a record that has been stale for a week. The resolution path inside a microVM guest is covered in Firecracker Guest DNS Resolution Explained.
  • Process start times, uptime, and anything derived from them. `/proc/uptime` is monotonic-based, so your restored guest reports an uptime that excludes the gap, which makes it look healthier than it is.
  • Anything that derived an identifier from the time: UUIDv1 and UUIDv7, Snowflake-style ids, time-bucketed cache keys, idempotency keys. Restore the same snapshot twice and you can mint the same "unique" id twice.
  • Cached credentials: signed URLs, presigned uploads, short-lived cloud tokens, session cookies. All of them carry an expiry the guest will misjudge in whichever direction is least convenient.

The practical response is a short list of things a guest should re-derive on resume rather than trust from its own memory:

  1. Wall-clock time — ideally handed in by the platform before you run, verified by your own code if the platform does not promise it.
  2. Entropy: reseed the kernel pool and any userspace PRNG on the restore signal, before issuing a single token or key.
  3. Network identity: re-resolve DNS, re-acquire the DHCP lease, drop the connection pool and reconnect rather than discovering the sockets are dead one timeout at a time.
  4. Credentials and leases: treat every cached token as expired on resume and refresh it, because the cost of an unnecessary refresh is a request and the cost of a wrong one is an outage.
  5. Caches with a TTL: flush, or store absolute wall-clock deadlines rather than monotonic ones so a restore invalidates them automatically.

The bottom line

Check `current_clocksource` in your guests, because it is a one-line read that tells you whether `clock_gettime` is arithmetic or a trap, and the difference is invisible in every profile you will take. Expect `kvm-clock`, understand that it is a coherency mechanism rather than a synchronisation service, and accept that nothing in the TSC/pvclock stack has any idea what today's date is.

Then assume your snapshots are time machines and audit what rides in them. If you run Firecracker yourself, make a synchronous clock step part of your resume path today, before your first certificate rotation makes the lesson expensive. If you let a time daemon handle it asynchronously, you have built a race that will fire in front of a customer rather than in CI.

Where I would not reach for PandaStack: if you need a GPU, we have none and no passthrough, so anything accelerated is off the table. If you need a guest kernel newer than 5.10 for a feature you depend on, check You Ship the Kernel: Firecracker Guest 5.10 vs 6.1 before you build on us. And if your workload needs one long-lived machine with a stable identity rather than thousands of disposable ones, a plain VM is a simpler purchase and nobody will have to explain clocksources to you. What snapshot-restore buys is a 179 ms create on every single call — and the price of that bargain is that you now have to think about what `now` means inside a guest, which is a trade I would make again, having made it badly once.

Frequently asked questions

Should I switch my guests from kvm-clock to tsc for performance?

Almost certainly not, and the reasoning is worth having straight because the measurement looks like it favours `tsc`. A direct TSC read is a single instruction, while a `kvm-clock` read is a seqlock check plus an `RDTSC` plus a multiply, shift and add against values loaded from a page. So yes, `tsc` is cheaper per read. But both are vDSO-capable, which means both stay entirely in userspace, so you are comparing one cheap thing to another cheap thing — and the cliff you actually care about is the one between either of them and a clocksource that traps to the VMM, which is a different order of magnitude entirely. What you give up by pinning `tsc` is the host's ability to republish the transform. `kvm-clock` exists so the guest's nanoseconds stay coherent when the hypervisor migrates a vCPU, rescales the TSC, or restores the VM somewhere whose TSC frequency differs. Pin `tsc` and you have asserted that none of those will ever happen to this guest, which is a strong claim on a platform whose entire premise is snapshotting and restoring VMs. If you additionally pass `tsc=reliable` you have also disabled the watchdog that would have told you when you were wrong. Measure first; if `clock_gettime` genuinely shows up in your profile, the answer is usually to call it less, not to change what it reads.

Why does CLOCK_MONOTONIC not advance across a snapshot restore, and is that a bug?

It is not a bug, it is the definition, and the definition happens to be inconvenient here. `CLOCK_MONOTONIC` measures time elapsed for this kernel from an arbitrary origin, and a snapshot restore does not involve any elapsed time from the guest's point of view — the vCPU state, including the TSC value the monotonic clock is derived from, is restored exactly as it was saved. Firecracker reprograms the TSC offset so that monotonic time is continuous across the restore, which is the correct behaviour for a resume: you do not want the guest's scheduler, its hrtimers or its RCU grace periods to observe a sudden multi-day jump, because large parts of the kernel treat that as a fault condition rather than information. People then reach for `CLOCK_BOOTTIME`, which is `CLOCK_MONOTONIC` plus time spent suspended, and discover it is frozen too — correctly, because the guest never went through a suspend/resume cycle it could account for. The practical consequence is that there is no clock inside the guest that can tell you a gap occurred. The guest needs an external signal: the platform stepping `CLOCK_REALTIME`, or the generation identifier that the hypervisor exposes on restore. If your code must detect a restore, watch for one of those rather than looking for a discontinuity, because by design there is not one.

Can I just run chrony or systemd-timesyncd in the guest and be done with it?

You can, and it will usually work, and that is exactly what makes it dangerous. A time daemon will eventually correct a restored guest's wall clock, so the failure disappears from your manual testing and reappears as an intermittent production incident. The problem is ordering: the daemon has to start, resolve a time server, complete an exchange and apply a correction, and your application has to not make an outbound TLS call before that finishes. On a platform where restore is the normal create path and takes a couple of hundred milliseconds, your application wins that race most of the time. There are two configuration details worth knowing if you use a daemon as a backstop, which you should. First, chrony slews by default — it adjusts the clock's rate rather than jumping it — which is right for drift and hopeless for a week-long offset; `makestep` with a generous threshold and a small count lets it step hard on the first updates, and `chronyd -s` plus `initstepslew`-style behaviour exist for the same reason in older setups. Second, a step forward makes every absolute wall-clock timer in the guest due simultaneously, so expect a burst when the correction lands. The real fix is still synchronous and external: the platform knows the time, so the platform should set it before the guest serves anything.

Why can't the guest just read the hardware clock after a restore?

Because under Firecracker there may not be one, and this surprises people coming from QEMU. Firecracker's device model is deliberately minimal — a serial port, a keyboard-controller stub used for reset signalling, and virtio devices for networking, block storage, vsock, ballooning and entropy. There is no HPET. On x86_64 there is no conventional CMOS/RTC chip to re-read, so `hwclock --hctosys`, the traditional way to recover wall-clock time after a suspend, has nothing to talk to. That minimalism is a security property first: every emulated device is VMM code reachable from guest input, and Firecracker's small device count is a large part of why its attack surface compares favourably to a general-purpose hypervisor. It is a timekeeping property second — fewer emulated timers means fewer ways for a guest to end up on a clocksource that traps on every read, which is why `available_clocksource` in a Firecracker guest is a two-item list rather than a menu. Verify the exact device list for your Firecracker version rather than trusting this summary, since the device model does change between releases. But do not design a recovery path that assumes an RTC is sitting there waiting to be consulted.

What else besides the clock breaks when I restore the same snapshot many times?

Entropy is the one that should worry you more than the clock, because it fails silently and the consequences are security-shaped rather than operational. Every guest restored from a single snapshot starts with an identical kernel entropy pool and identical seeded userspace PRNGs, so two independent sandboxes can generate the same session token, the same temporary filename or the same key. Firecracker exposes a generation identifier on restore specifically so a guest can notice and reseed, and there is a dedicated post on it. Our own API has a useful asymmetry here: `fork()` clones the disk and the child cold-boots, so it gets its own entropy and its own PIDs, while `fork_tree()` deliberately inherits the parent's running memory — which is the point when you want warm in-memory state fanned out, and which means it inherits this problem too. Beyond entropy, audit anything with a negotiated lifetime. Open TCP connections whose peers have long since hung up. DHCP leases that expired during the gap. DNS caches holding records whose TTLs were measured against a monotonic clock that did not move. Signed URLs, presigned uploads and short-lived cloud credentials. Identifiers derived from the time, including UUIDv1 and UUIDv7 and Snowflake-style ids, where restoring one snapshot twice can mint the same supposedly unique value twice. The rule that covers all of it: on resume, re-derive rather than trust.

Keep reading

Related posts

  • Guest Clocks and the TSC After a Firecracker Restore

    Restore a snapshot a day later and the guest wakes up convinced it's still last Tuesday — and every HTTPS handshake now disagrees. Here's how guest time actually works, why it freezes on restore, and how to unfreeze it.

  • Snapshot, Restore, and the Connections You Left Open

    A snapshot captures the guest's entire opinion about its network: socket table, sequence numbers, retransmit timers, ARP cache, TLS sessions. The one participant never consulted is the machine on the other end of every one of those connections.

  • Guest Clock Drift After a Firecracker Snapshot Restore

    Restore a snapshot from last Tuesday and the guest still thinks it's last Tuesday. That frozen wall clock quietly breaks TLS, JWTs, cron, and rate limiters — until you step it back to now on restore.

  • Golden Images vs Snapshot Baking

    A golden image removes install time. A snapshot removes boot time. Those are different costs, which is why the answer is almost always both — and why the snapshot quietly freezes your RNG, your clock, and anything that was in RAM at bake time.

  • The Firecracker VMGenID Device, Explained

    Two VMs, one entropy pool, zero good outcomes. When you restore or fork a snapshot, the guest wakes up sure it's the same running system with the same RNG state. The VM Generation ID device is how it finds out it was forked — and reseeds before it hands out a duplicate nonce.

More in Firecracker & microVMs · See Firecracker microVM sandboxes

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.