all posts

Locale, Timezone and Encoding Bugs That Only Appear in Ephemeral Sandboxes

Ajay Kumar··11 min read

There is a bug that only reproduces for Jonas in Hamburg. The suite is green on your laptop, green on everyone else's laptop, green in the nightly, and red on his machine with a stack trace about a comma. You pair on it for forty minutes, read the same eleven lines twice, and find nothing — because there is nothing in the code to find. His LC_NUMERIC is de_DE.UTF-8 and yours is not.

This family of bug has a second home, and it is the minimal base image. A developer laptop is a maximalist environment: every locale generated or faked by the OS, a real /etc/localtime, a full ICU data set, fonts, and a timezone database some updater refreshed last month without telling you. A minimal image is the opposite by design. It ships forty megabytes less, and most of what it dropped was opinions — about which bytes are letters, about what order strings go in, about which character separates a decimal, about what hour it is.

The code did not change. The environment's opinions changed. And these particular opinions are read from the process environment at startup, cached in the process, and then baked into behaviour — which is why a snapshot makes them permanent in a way a container restart does not. A restarted container re-reads the world. A restored snapshot resumes a process that already decided.

This post is about locale, encoding and collation as a correctness class. The guest clock is one member of that family and gets a section here, but why a microVM's clock drifts and which clock source it reads is a separate subject with its own post — linked at the bottom. Here the clock is a symptom, not the topic.

Your laptop is maximalist; your image is minimalist

It is worth being concrete about how little a stock Ubuntu 24.04 rootfs believes — not a hardened distroless thing, the ordinary base layer half the world builds on.

  • locale -a returns exactly three entries: C, C.utf8, POSIX. en_US.UTF-8 is not one of them until something installs the locales package and runs locale-gen.
  • LANG and LC_ALL are both unset, so locale reports LC_CTYPE="POSIX" and locale charmap answers ANSI_X3.4-1968 — which is the ISO registry's name for US-ASCII, a fact that has confused every engineer who has ever grepped for it.
  • There is no /etc/localtime. There is no /usr/share/zoneinfo at all. The directory does not exist.
  • ldconfig finds no libicuuc, so there is no ICU data for anything that wanted it.
  • None of this is a bug. It is the correct outcome of shipping a smaller image, and it is why the image is smaller.

Our own general-purpose template is in this group, and pretending otherwise would be silly: the base Dockerfile installs neither locales nor tzdata and sets no LANG, LC_ALL or TZ. The postgres-16 template does the opposite — it sets LANG and LC_ALL to en_US.UTF-8 and TZ to UTC, installs locales and tzdata, and generates en_US.UTF-8 — because a database cannot afford to be vague about collation. The difference between those two templates is the whole post.

So here is the shape of every failure below. A library asks the environment a question at startup; a maximalist environment gives a rich answer, a minimalist one a conservative fallback, a misconfigured one a confident wrong answer. Then it caches that for the rest of the process's life. All three cases at once:

# Same binaries. Same input. Three different answers. Only the environment's
# opinion changed. Verified on Ubuntu 24.04, glibc 2.39 -- run it in YOUR image.

$ printf 'B\na\nC\nb\n_x\n' | LC_ALL=C sort | tr '\n' ' '
B C _x a b
$ printf 'B\na\nC\nb\n_x\n' | LC_ALL=en_US.UTF-8 sort | tr '\n' ' '
a b B C _x
#   ^ pipe either into sha256sum and you have two different "reproducible" builds

$ printf 'naive\nnaïve\n' | LC_ALL=C grep -c '^[[:alpha:]]*$'
1
$ printf 'naive\nnaïve\n' | LC_ALL=en_US.UTF-8 grep -c '^[[:alpha:]]*$'
2
#   ^ LC_CTYPE decides which bytes count as letters, so it decides what
#     [[:alpha:]] matches -- and therefore what your regex quietly filters out.

$ LC_ALL=de_DE.UTF-8 printf '%.2f\n' 3.14
bash: printf: 3.14: invalid number        # <- stderr. Your pipeline drops stderr.
3,00
#   ^ not 3,14. bash's BUILTIN printf parses the input with strtod as well, so in
#     a comma locale the literal 3.14 fails to parse and you get 3,00 -- the
#     cents are gone. /usr/bin/printf (coreutils) accepts it and prints 3,14.
#     Same command name, two different answers, decided by what your shell found.

$ LC_ALL=C printf '%.2f\n' 3,14
bash: printf: 3,14: invalid number
3.00
#   ^ and that is the round trip: write in de_DE, read in C, lose the fraction.

$ TZ=Europe/Berlin date '+%Y-%m-%d %H:%M %Z'   # image WITH tzdata
2026-10-03 09:43 CEST
$ TZ=Europe/Berlin date '+%Y-%m-%d %H:%M %Z'   # same image, no tzdata package
2026-10-03 07:43 Europe
#   ^ no error, no warning. glibc could not find the zone file, fell back to
#     parsing "Europe/Berlin" as a POSIX TZ STRING, took "Europe" as the zone
#     abbreviation and zero as the offset. That is UTC wearing a continent's name.

Three things in that output deserve to annoy you. The sort order changed, which changes checksums. The printf case lost the fraction, not just the separator, and told you only on stderr. And the timezone case produced a zone abbreviation of "Europe", silently, with no error and no indication that the setting you carefully provided was discarded.

LANG unset, LC_ALL=POSIX, and the ASCII default

The mechanism: glibc's setlocale reads LC_ALL, then the specific LC_* category, then LANG, in that order. In the C or POSIX locale, nl_langinfo(CODESET) returns ANSI_X3.4-1968, and anything that asks the platform what encoding to use gets ASCII.

Python is the version-sensitive one, and the received wisdom about it is now mostly wrong, so be specific about your interpreter rather than trusting a blog post — including this one. On CPython 3.12 with LANG and LC_ALL both unset, sys.flags.utf8_mode is 1 and locale.getpreferredencoding(False) returns utf-8. Opening a UTF-8 file works fine. PEP 538 coerces the C locale to C.UTF-8 and PEP 540 auto-enables UTF-8 mode, both since 3.7, and PEP 686 makes UTF-8 mode the default outright from 3.15 — at which point the default stops depending on the environment at all.

So where does the classic UnicodeDecodeError still live? Set PYTHONCOERCECLOCALE=0 and PYTHONUTF8=0 with LC_ALL=POSIX and you get utf8_mode 0, getpreferredencoding ANSI_X3.4-1968, and reading a file containing one accented character fails with "'ascii' codec can't decode byte 0xc3 in position 3". That is not a contrived configuration: it is what you get from an interpreter older than 3.7 (a vendored one, or a stray python2 script still load-bearing in your pipeline), from a locale that is set but not UTF-8 — coercion only fires for C and POSIX, never for a real non-UTF-8 locale — or from a base image whose maintainer pinned those variables for reasons lost to history.

The single most useful tool here is not a fix, it is a flag. Run with -X warn_default_encoding (or PYTHONWARNDEFAULTENCODING=1, 3.10 and later) and every place in your code that relied on the platform default emits an EncodingWarning naming the file and line. You will be surprised how many there are, and they are all one encoding= argument away from never being a problem again.

One widely repeated detail is worth correcting while we are here. The accented-filename version of this failure is usually not a UnicodeDecodeError: Python decodes undecodable filenames with surrogateescape, so os.listdir succeeds and hands back a str containing lone surrogates, and the crash arrives later as a UnicodeEncodeError when something prints or logs that name. Reading a UTF-8 file's contents in an ASCII locale is the decode error. Same root cause, two different exception names — which is exactly why grepping CI logs for one of them misses the other.

The same bug in four other languages

  • Java — file.encoding was derived from the platform locale, so a C-locale container gave an ASCII or Latin-1 default and new FileReader(f) mangled UTF-8 silently. JEP 400 made UTF-8 the default charset from JDK 18, which fixed the common case and MOVED the bug: a JDK 18+ process now reads as UTF-8 a file an older one wrote in the platform default. -Dfile.encoding=COMPAT restores the old behaviour if you need the old bug back.
  • Java again, for numbers — String.format("%.2f", x) with no explicit Locale uses Locale.getDefault(). A JVM that picked up a comma locale formats currency with a comma, in your API response, and the method signature gives no hint that it consulted the environment.
  • Ruby — Encoding.default_external comes from the locale, so the C locale makes it US-ASCII and the first non-ASCII byte yields "invalid byte sequence in US-ASCII" from deep inside a regex. RUBYOPT=-EUTF-8, or an explicit external encoding on each read.
  • Perl — does no implicit decoding at all, so the symptom inverts: "Wide character in print" and double-encoded output. use open qw(:std :encoding(UTF-8)) at the top of anything handling text.
  • Shell tools — sed, grep, awk and wc switch between byte and character semantics on LC_CTYPE. The [[:alpha:]] count above is one face of it; "invalid multibyte sequence" from sed and a wrong wc -m are the others.

LC_COLLATE: the one that actually breaks reproducible builds

In the C locale, strcoll is strcmp: pure byte order, so uppercase sorts before lowercase and underscore sorts before every lowercase letter. In en_US.UTF-8, glibc uses weight tables derived from ISO 14651: case folds to a lower-priority weight, punctuation weighs below letters. Same input, different order, as shown above — BC_xab becomes abBC_x.

Now follow where a sort order goes. Into a lockfile. Into a generated header or client. Into tar member order, a concatenated bundle, the checksum of any of those, a golden-file test that diffs output against a committed expectation. None of those artefacts mention collation anywhere, and all of them change when it does.

My honest claim, offered as experience rather than measurement: most investigations of "why isn't this build byte-reproducible" end at collation, timestamps or directory-read order. The compiler is the first suspect and almost always innocent. Collation is the one nobody checks, because the idea that sort is a pure function is assumed so deeply that testing it does not occur to anyone.

The fix is to pin byte collation — LANG=C.UTF-8, or LC_ALL=C if you can live with the ctype consequences — for everything that orders. Not en_US.UTF-8, which is pinned to a glibc version's opinion rather than to bytes: glibc 2.28 rebuilt its collation tables wholesale, a change invisible in a diff of your own repository that alters the output of every sort you run. If you need human-friendly ordering, do it at the presentation layer with an explicit, versioned collator, not in the pipeline that produces artefacts.

LC_ALL=C also sets LC_CTYPE to C, which turns multibyte handling off in sed, grep and awk. The correct pairing is usually LC_COLLATE=C with a UTF-8 LC_CTYPE — which is exactly what LANG=C.UTF-8 gives you, and which almost nobody sets deliberately.

LC_NUMERIC: a data-corruption bug wearing a formatting bug's clothes

LC_NUMERIC sets the decimal point and the thousands separator for everything that goes through the C library's number conversion. On the output side that means printf emits 3,14 instead of 3.14. On the input side it means strtod, atof and scanf expect a comma and stop at a period. Those two halves are where the damage happens, because they are rarely in the same process.

The verified behaviour is nastier than the usual summary. bash's builtin printf parses its arguments with strtod too, so LC_ALL=de_DE.UTF-8 printf '%.2f' 3.14 does not print 3,14 — it writes "invalid number" to stderr and prints 3,00. The fraction is gone, and the only warning went to the stream your pipeline redirects to /dev/null. Meanwhile /usr/bin/printf from coreutils accepts the same input and prints 3,14. Same command name, two different answers, decided by whether your shell resolved a builtin or a binary.

Run the round trip and you have silent corruption: write 3,14 in a comma locale, read it back with LC_ALL=C, and strtod stops at the comma and hands you 3.00. No exception, no non-zero exit from the parsing itself, errno untouched in the common path. If that number was a price, you just billed someone three euros.

Then it travels. A comma in a CSV is a field separator, so one locale-formatted float turns one column into two and shifts every column after it. A comma in a JSON-ish payload is a parse error if you are lucky and a truncated number if the producer quoted it. A comma in a generated SQL literal is a syntax error in the best case and a two-argument function call in the worst.

Two useful asymmetries. gawk's printf output is not locale-dependent by default — but gawk --posix turns the locale decimal point ON, so the POSIX-compliance flag is what introduces the bug. And several modern languages ignore the locale for numbers entirely by design: Go's strconv, Rust, and Python's float formatting and float() are locale-independent, with the locale-aware path opt-in via locale.format_string and locale.atof. A polyglot pipeline therefore has a locale-sensitive half and a locale-insensitive half, and the corruption lives exactly at the boundary.

Timezone: the bug hides because UTC is correct

With TZ unset, glibc reads /etc/localtime. If that file is missing — and in a stock minimal image it is, along with the entire zone database — the process runs in UTC. Which is correct. That is exactly why this one hides: logs look right, timestamps are monotone, comparisons work, nothing is obviously broken. The failure needs a trigger.

  • Something formats a local time for a human. An invoice, a reminder email, a calendar entry, a "last seen" string.
  • Something crosses a DST boundary. The twice-yearly test failure that reverts itself a day later and nobody files.
  • A test asserts on a date derived from "today". A job running at 23:30 in Berlin lands on tomorrow inside a UTC guest, and the test is green for twenty-three and a half hours a day.
  • A period boundary — billing, a daily rollup, a cron at "midnight" — silently shifts by the offset you did not set.

Then the silent-failure mode, which is the one that cost me an afternoon: setting TZ correctly in an image with no tzdata. glibc cannot find the zone file, falls back to interpreting the value as a POSIX TZ string, takes "Europe" as the zone abbreviation and zero as the offset, and prints the UTC time labelled "Europe". No error. No warning. You configured it, it was discarded, and the output is plausible enough to pass review.

The sandbox-flavoured version: a clock that jumps

Firecracker's snapshot captures CLOCK_REALTIME along with everything else, so a restored guest resumes believing it is the moment the snapshot was taken. Correct that, and the guest experiences a jump rather than a drift — which breaks a different set of things than drift does. Anything that cached a timezone offset, a formatted date, a token expiry or a certificate validity window at process start is now holding a value computed for a different instant. glibc caches the parsed TZ per process; Java caches TimeZone.getDefault(); Python needs an explicit time.tzset() to re-read TZ. None of them expect the ground to move.

PandaStack re-syncs the guest wall clock on every restore-family ready path — snapshot-restore create, resume and wake — setting it from host time over the guest exec bridge, best-effort with one retry, never failing the boot. It exists because of a live incident, and the symptom is worth stating precisely because everybody guesses it backwards. The guest's clock was behind, frozen at its seed's bake date. An upstream rotated onto a leaf certificate whose notBefore was after that date, so the guest rejected a perfectly valid certificate as not yet valid — not expired, the opposite — and every git clone inside a restored guest died with "server certificate verification failed". Everyone's first instinct was a missing CA bundle, because that is what the message mentions. The CA bundle was fine. The clock was in June.

A golden image is a slowly rotting timezone database

IANA ships several tzdata releases a year because governments change the rules, usually with very little notice. Mexico abolished DST in 2022. Iran dropped it the same year. Lebanon moved its 2023 spring change days before it was due to happen. A baked template pins a tzdata version, so a long-lived golden image is a photograph of what the world's timezone rules were on bake day. The symptom is an off-by-one-hour bug, on a specific future date, in a specific country, affecting a subset of users — the worst possible shape for a bug to have. Not reproducible today; perfectly reproducible on the 26th of October.

The part people miss: there are several copies of tzdata on a typical filesystem. The JVM bundles its own. Node reads zones through ICU. A browser has its own. Postgres has its own unless it was built against the system copy. So "we updated tzdata" is a statement about one of four or five databases on that disk, and the one you updated may not be the one your application reads.

Practically: print the tzdata release in the preflight below and treat it as part of the image's identity, so a bump is a deliberate re-bake rather than an accident of a cache miss. Re-bake on a schedule rather than on demand. And for anything that computes a future local time, do not cache the offset — compute it from the zone name at use time, because the rule may change between now and then.

ICU, and the honest version of the big-image problem

A full ICU data set is tens of megabytes, which is an enormous fraction of an image you intended to keep small. So minimal images ship without it, with a stub, or against a system copy of a version nobody chose. The consequences are real and the details are version-specific enough that you should not take them from a blog post.

  • Node — builds can be configured with full ICU, English-only small-icu, or a system ICU. Official and distro builds have differed historically; Alpine builds in particular vary. Check your binary rather than assuming.
  • .NET — globalization-invariant mode is commonly switched on in slim images because ICU is absent. In that mode every culture behaves like the invariant culture and string comparison becomes ordinal, which changes sorting and equality results, not merely display.
  • Postgres — collations come from either the libc provider or the ICU provider, and the ICU provider carries its own version. Which one your database used is recorded, not inferred.
  • Everything else that formats a date, a number or a name per locale — if ICU is how it does that, ICU's absence is a behaviour change rather than an error.

The honest position is that the correct detail depends on the exact image and version, and anything specific I write here has a shelf life measured in months. So the deliverable is not a fact, it is a probe. Ask Node to format a German month name: a German month means you have the data, "January" means you have a fallback pretending to be a locale. One line, worth more than a paragraph of received wisdom — including this one.

Postgres collation: an index that disagrees with its own comparison

This is the nastiest member of the family, because it corrupts query results rather than crashing. A btree index on a text column is physically ordered by the collation in force when it was built, and the planner assumes that order still holds. A glibc upgrade can change what strcoll answers; if the index was built before and queried after, the index's order no longer matches the comparison function.

The consequences are the quiet kind: an equality lookup served from the index can miss rows present in the heap, and a unique index can admit a duplicate. Neither raises an error. Postgres 13 and later record the collation version — pg_collation.collversion and pg_database.datcollversion — and warn when it changes. ALTER ... REFRESH COLLATION VERSION silences the warning; only REINDEX fixes the index. Refreshing without reindexing is the fast way to turn a warning into silence while keeping the corruption. Which means a base-image bump underneath a Postgres is a schema-affecting change even though nothing in your migrations directory moved — an unusual sentence, and true.

The check worth running before and after any image bump is two queries: select datname, pg_encoding_to_char(encoding), datcollate and datcollversion from pg_database, then select collname, collprovider and collversion from pg_collation where collversion differs from pg_collation_actual_version(oid). Rows from the second one mean the library's answer has moved away from what your indexes were built against. Put it in the same preflight as everything else — that version string is the one thing here your database already knows and has never told you.

The version-stable alternative is to stop asking for linguistic ordering you do not need. C and C.UTF-8 collations are byte order: no version to drift, no table for a glibc release to rebuild. Most application indexes — ids, slugs, emails, enum-ish strings — never needed a human collation. Postgres 17 added a builtin provider aimed squarely at this; check your version's docs for the current spelling rather than copying mine.

Our own answer, since this is what a managed database is actually buying you: the postgres-16 template sets LANG and LC_ALL to en_US.UTF-8 and TZ to UTC in the image, installs locales and tzdata, generates en_US.UTF-8, initialises the cluster with --locale=en_US.UTF-8 --encoding=UTF8, and creates every database with ENCODING 'UTF8' LC_COLLATE 'en_US.utf8' LC_CTYPE 'en_US.utf8'. Two databases from the same template generation agree by construction, and a clone or point-in-time restore lands in the image it came from. The honest limit: that collation is glibc's, so re-baking postgres-16 onto a newer glibc is a collation change for everything created after it, and carrying a cluster across that boundary is a REINDEX conversation rather than a non-event. Pinning the image makes the problem schedulable, not eternal — anyone claiming their managed database has solved collation drift forever is describing a pin, not a cure.

Which brings us to ephemeral per-pull-request databases, where this gets expensive. The entire value of a throwaway database per PR is that it behaves like production. If production is en_US.UTF-8 on one glibc and the throwaway is C.UTF-8 on another, then your ORDER BY test, your unique-constraint test and your keyset-pagination test all passed against a database that sorts differently from the one your users will hit. The test ran. It proved nothing, and it will keep proving nothing, confidently, until a customer reports that search results are in a strange order.

The failure classes, in one table

Each row is one question the environment answers at process start. The last column is what belongs in a boot-time preflight.
Failure classVariableWhere it bitesBoot-time probe
Default text encodingLC_CTYPE, LANG, LC_ALLReading a UTF-8 file as ASCII; mojibake written to a database that cannot be undonelocale charmap equals UTF-8
Character classesLC_CTYPEgrep, sed and awk match different lines; wc -m counts bytes, not charactersgrep -c on a known accented line
Sort orderLC_COLLATEChecksums, lockfiles, tar member order, golden-file diffs — reproducible buildssort a fixed list, compare the exact string
Decimal separatorLC_NUMERICprintf emits a comma; strtod reads 3,14 as 3.00 and says nothingprintf '%.2f' 3.14 equals 3.14
Missing zone databaseno tzdata, no /etc/localtimeEverything runs in UTC and looks fine until a local time is formatteddate +%z and the existence of a zone file
TZ silently ignoredTZ set, tzdata absentYou configured Europe/Berlin and got UTC labelled "Europe"date +%Z matches the zone you asked for
Stale tzdatatzdata release in the imageOff-by-one-hour on a specific future date in one countryprint the tzdata release as image identity
Missing or stubbed ICUimage contents, build flagsLocale formatting silently falls back to English; .NET comparison turns ordinalformat a German month name and read it
Collation version driftglibc or ICU versionPostgres index disagrees with its comparison: missed rows, admitted duplicatescompare datcollversion with the actual version
Frozen then jumping clocksnapshot restoreCached offsets, token expiries and certificate validity computed for another instantcompare guest UTC time against the host's

Assert on the environment instead of hoping

Every item above has the same fix, and it is unglamorous: at boot, print what the environment believes, compare it with what you decided, and fail loudly when they differ. Not a document describing the intended configuration — a script that exits non-zero.

The argument for failing at boot rather than at use is about where the information is. At boot the whole environment is in front of you and nothing has happened yet, so the message can say "collation is not byte order" and name one variable. At use you have a UnicodeDecodeError forty frames deep in a dependency, in a job that already wrote half its output, at two in the morning, and somebody has to reconstruct the chain from one byte offset. Same bug, two orders of magnitude difference in time-to-understand.

The useful script prints two kinds of thing: what is set, and what the libraries actually do with it. The second matters more. LANG=en_US.UTF-8 tells you somebody's intention; sort returning a specific string tells you what will happen to your build.

#!/usr/bin/env bash
# /usr/local/bin/env-fingerprint
# Print what this environment actually believes, then exit non-zero if it
# disagrees with what this image decided. Runs at bake time, at guest boot,
# and as the first step of every job. Cheap enough to run three times.
set -uo pipefail

kv() { printf '%s=%s\n' "$1" "${2:-<unset>}"; }

# --- what is SET -------------------------------------------------------------
kv LANG       "${LANG-}"
kv LC_ALL     "${LC_ALL-}"
kv LC_COLLATE "$(locale 2>/dev/null | sed -n 's/^LC_COLLATE=//p' | tr -d '"')"
kv LC_NUMERIC "$(locale 2>/dev/null | sed -n 's/^LC_NUMERIC=//p' | tr -d '"')"
kv charmap    "$(locale charmap 2>/dev/null)"
kv locales    "$(locale -a 2>/dev/null | tr '\n' ',')"
kv TZ         "${TZ-}"
kv localtime  "$(readlink -f /etc/localtime 2>/dev/null || echo '<missing>')"

# --- tzdata version ----------------------------------------------------------
# /usr/share/zoneinfo/+VERSION does NOT exist on Debian or Ubuntu. tzdata.zi
# carries "# version <release>" and is the portable answer; package metadata is
# the fallback. Treat this string as part of the image's identity.
kv tzdata "$(sed -n 's/^# version //p' /usr/share/zoneinfo/tzdata.zi 2>/dev/null \
             || dpkg-query -W -f='${Version}' tzdata 2>/dev/null \
             || rpm -q --qf '%{VERSION}' tzdata 2>/dev/null \
             || echo '<no tzdata>')"

# --- what the libraries DO with it (probes, not settings) --------------------
kv probe_sort  "$(printf 'B\na\nC\nb\n_x\n' | sort | tr -d '\n')"
kv probe_float "$(/usr/bin/printf '%.2f' 3.14 2>/dev/null)"
kv probe_alpha "$(printf 'na\xc3\xafve\n' | grep -c '^[[:alpha:]]*$')"
kv probe_zone  "$(date '+%Z')"
kv probe_off   "$(date '+%z')"
kv probe_now   "$(date -u '+%Y-%m-%dT%H:%M:%SZ')"
kv probe_icu   "$(ldconfig -p 2>/dev/null | grep -c libicuuc)"
command -v node >/dev/null 2>&1 && kv probe_intl \
  "$(node -p "new Intl.DateTimeFormat('de-DE',{month:'long'}).format(new Date())" 2>/dev/null)"
command -v python3 >/dev/null 2>&1 && kv probe_py \
  "$(python3 -c 'import sys,locale;print(sys.version.split()[0],sys.flags.utf8_mode,locale.getpreferredencoding(False),sep="/")')"

# --- assertions: loud at boot beats subtle at 02:00 --------------------------
fail=0
chk() { # chk <label> <got> <want>
  [ "$2" = "$3" ] || { printf 'FAIL %s: got %q want %q\n' "$1" "$2" "$3" >&2; fail=1; }
}
chk charmap     "$(locale charmap)"                                   'UTF-8'
chk collation   "$(printf 'B\na\nC\nb\n_x\n' | sort | tr -d '\n')"    'BC_xab'
chk decimal     "$(/usr/bin/printf '%.2f' 3.14)"                      '3.14'
chk utc_offset  "$(date '+%z')"                                       '+0000'
[ -e /usr/share/zoneinfo/UTC ] || { echo 'FAIL no zone database on disk' >&2; fail=1; }
exit "$fail"

Then run it as a gate rather than as documentation. This is the whole job wrapper: create a sandbox, run the fingerprint, compare it with the values this pipeline committed to, and refuse to build on any mismatch.

from pandastack import Sandbox

# The fingerprint script from above, shipped with the pipeline so the GATE is
# versioned alongside the code it protects.
PREFLIGHT = open("env-fingerprint.sh", "rb").read()

# What this pipeline has DECIDED. Not what it hopes to find. Every value here
# was produced by running the probe once and writing down the answer.
EXPECTED = {
    "charmap": "UTF-8",
    "LC_COLLATE": "C.UTF-8",
    "LC_NUMERIC": "C",
    "probe_sort": "BC_xab",     # byte order: stable across glibc releases
    "probe_float": "3.14",      # a period, in every service, forever
    "probe_zone": "UTC",
    "probe_off": "+0000",
    "tzdata": "2026c",          # bump deliberately, with a re-bake
}


def fingerprint(sbx: Sandbox) -> dict[str, str]:
    """Run the preflight inside the guest and parse its key=value output."""
    sbx.filesystem.write("/usr/local/bin/env-fingerprint", PREFLIGHT)
    res = sbx.exec("bash /usr/local/bin/env-fingerprint")
    env = dict(
        line.split("=", 1) for line in res.stdout.splitlines() if "=" in line
    )
    if res.exit_code != 0:
        # The script's own assertions failed -- the image disagrees with itself.
        raise EnvironmentError(
            f"preflight failed in guest:\n{res.stderr.strip()}\nfingerprint: {env}"
        )
    return env


def main() -> int:
    with Sandbox.create(
        template="base",
        ttl_seconds=600,                       # IDLE timeout, not a walltime cap
        metadata={"job": "build", "gate": "env-preflight"},
    ) as sbx:
        env = fingerprint(sbx)

        drift = {k: (v, env.get(k)) for k, v in EXPECTED.items() if env.get(k) != v}
        if drift:
            # Fail the JOB, with one legible line per surprise -- not 400 tests
            # downstream of it, and not a UnicodeDecodeError at 02:00 on a Sunday.
            for key, (want, got) in sorted(drift.items()):
                print(f"env drift: {key} want={want!r} got={got!r}")
            print(f"{len(drift)} environment mismatch(es); refusing to build")
            return 1

        # Only now does the build get to run. Bound long commands in the shell or
        # use exec_stream: one-shot exec has no server-side deadline today, and
        # the client's own default cuts in at 30s.
        return sbx.exec_stream(
            "cd /work && timeout 1800 ./build.sh",
            on_stdout=print,
            on_stderr=print,
            timeout_seconds=1800,
        )


if __name__ == "__main__":
    raise SystemExit(main())

This belongs in the template bake, not in each job

An export in a job script protects that job. It does not protect the processes you did not write — the package manager's subprocess, the test runner's child, a dev server somebody launched detached, the language server an agent started on its own initiative. Those inherit whatever the image decided, and the image decided at bake time.

Which is where the mechanism that made the problem sticky becomes the thing that fixes it. A Firecracker snapshot captures the booted machine: generated locale archive, zone files, environment, all inside the artefact you restore. Every create restores that baked snapshot — p50 179 ms end to end, the snapshot-load step itself around 49 ms; only the first spawn before a snapshot exists is a real cold boot, roughly 3 seconds. There is no warm pool of long-lived machines to drift apart, because there are no long-lived machines. Set the answers once, bake them, and every guest for the life of that template generation gives the same ones.

# templates/<name>/Dockerfile -- set it once here, snapshot it, and every
# restored guest inherits the same answers. Packages BEFORE locale-gen, because
# locale-gen compiles its output from sources the `locales` package installs.

FROM ubuntu:24.04

RUN apt-get update && apt-get install -y --no-install-recommends \
      locales tzdata \
 && locale-gen C.UTF-8 en_US.UTF-8 \
 && ln -sf /usr/share/zoneinfo/UTC /etc/localtime \
 && echo UTC > /etc/timezone \
 && rm -rf /var/lib/apt/lists/*
# Do NOT `rm -rf /usr/share/i18n` to slim the image: that deletes the SOURCES
# locale-gen reads, so the next layer that wants a locale silently gets nothing.
# Deleting /usr/share/locale (translated program messages) is fine -- the
# compiled locale data lives in /usr/lib/locale/locale-archive, not there.

# C.UTF-8, not en_US.UTF-8, as the default: UTF-8 ctype with BYTE collation, so
# `sort` output does not move when glibc rebuilds its collation tables. Set LANG
# (a default each LC_* category falls back to) rather than LC_ALL (an override
# nothing downstream can escape), so a job that genuinely needs a linguistic
# collator can still ask for one explicitly. LC_NUMERIC is pinned separately
# because nothing good has ever come of a locale-dependent decimal separator.
ENV LANG=C.UTF-8 \
    LC_NUMERIC=C \
    TZ=UTC

# Firecracker boots this rootfs under systemd, and Docker's ENV is NOT part of
# the flattened filesystem -- it lives in image metadata the microVM never
# reads. /etc/environment IS read by PAM for every guest session, login or not,
# so this is the line that actually reaches a detached build process.
RUN printf 'LANG=C.UTF-8\nLC_NUMERIC=C\nTZ=UTC\n' >> /etc/environment

COPY env-fingerprint /usr/local/bin/env-fingerprint
RUN chmod +x /usr/local/bin/env-fingerprint && /usr/local/bin/env-fingerprint
#   ^ the BAKE fails here if the image disagrees with itself. A check that fails
#     a build is a control; a sentence in a README is a description.
The /etc/environment line is not optional padding. Docker's ENV lives in image metadata, and a Firecracker microVM boots the flattened filesystem — it never reads that metadata. We learned this the slow way on our own base template: mise shims were inert for detached build processes until the same values were written to /etc/environment, which PAM reads for every session, login or not. An ENV you set in a Dockerfile and then verified with docker run can be entirely absent from the microVM built from the same image.

What this does not fix

  • Data that already has the wrong bytes in it. Pinning the environment stops new mojibake; it does not un-mojibake a column written by three years of CP-1252 guesses. That is a migration, and it needs a human deciding what the original text was.
  • A dependency that bundles its own tzdata or ICU. You pinned the system copy; it brought its own, pinned to its release cycle rather than yours.
  • A library that consults Locale.getDefault() deep inside a formatting call. You can set the default to something sane, but the call site still gives no hint that it asked the environment a question.
  • The genuine tradeoff. C.UTF-8 is the right default for pipelines and the wrong default for a product that sorts names for Swedish users. Linguistic collation needs an explicit, versioned collator at a specific layer — more work than a global variable, and the only thing that actually works.
  • The first incident in a class you have not met. A preflight only compares what you thought to check, so its first version is always incomplete: something breaks, you understand it, you add a row. The value compounds; the starting point does not impress anyone.

"It works on my machine" is literally true

That is the whole problem with the sentence. It is not a dismissal; it is an accurate report of a true fact about a machine nobody else has. It is also useless, because it names an outcome without naming one of the fifteen or so environmental opinions that produced it. The useful version is longer and much less satisfying to say out loud: my machine's charmap is UTF-8, its collation is byte order, its decimal separator is a period, its offset is +0000, its tzdata release is 2026c — here they are, printed at boot, and the build refuses to run if yours differ. That sentence takes forty lines of bash and it ends the argument.

Jonas, for the record, was right. His machine was the only honest one in the building.

Frequently asked questions

Why does my Python script raise UnicodeDecodeError only in CI?

Because your laptop's locale says UTF-8 and the CI image's says ASCII, and Python asked the platform rather than you. Opening a text file without an explicit encoding= argument uses the platform default, which in the C or POSIX locale resolves to ANSI_X3.4-1968 — the registry's name for US-ASCII. The first non-ASCII byte then raises something like "'ascii' codec can't decode byte 0xc3". Be version-specific before you start fixing: on CPython 3.7 and later, PEP 538 coerces the C locale to C.UTF-8 and PEP 540 auto-enables UTF-8 mode, so a plain LANG-unset environment usually self-heals, and PEP 686 makes UTF-8 mode the default from 3.15. The failure survives where a locale is set but is not UTF-8 (coercion only fires for C and POSIX), where PYTHONCOERCECLOCALE=0 or PYTHONUTF8=0 are pinned in the image, and on anything older than 3.7. Check with python3 -c 'import sys,locale;print(sys.version,sys.flags.utf8_mode,locale.getpreferredencoding(False))'. Then run your suite with -X warn_default_encoding, which names every call site that relied on the default, and pass encoding= at each one.

Should I set LC_ALL=C, C.UTF-8, or en_US.UTF-8 in my container?

For a build or data pipeline, LANG=C.UTF-8 with LC_NUMERIC=C is the best default, and you should prefer LANG over LC_ALL. C.UTF-8 gives you UTF-8 character handling with byte collation, so text is decoded correctly while sort stays stable across glibc releases — it has no collation tables for a library upgrade to rebuild. Plain LC_ALL=C also gives byte collation but turns off multibyte handling in sed, grep and awk, which produces a different set of bugs. en_US.UTF-8 is the worst choice for a pipeline: its sort order is a specific glibc version's opinion, and glibc 2.28 rebuilt those tables wholesale. Prefer LANG because it is a default each LC_* category falls back to, whereas LC_ALL is an override nothing downstream can escape — a job that genuinely needs a linguistic collator can still ask for one. One practical note: C.UTF-8 is built into recent glibc, so it works without installing the locales package, while en_US.UTF-8 requires locales plus a locale-gen run. Verify with locale -a in your actual image.

Does a missing tzdata package matter if my application only uses UTC?

Mostly no, and the exception is the one that bites. With no zone database the process runs in UTC, which is what you wanted, so logs and comparisons are fine. Two things still break. First, if anything sets TZ to a zone name, glibc cannot find the file and falls back to parsing the value as a POSIX TZ string — TZ=Europe/Berlin yields the UTC time labelled with a zone abbreviation of "Europe", with no error and no warning. You configured it, it was silently discarded, and the output looks plausible. Second, any code path that eventually formats a local time for a human, computes a billing or reporting boundary, or runs a schedule at "midnight" is implicitly asking for a zone, and it will get UTC whether that is right or not. The failure arrives at the boundary of a reporting period or a DST change rather than at startup. If you are certain you are UTC-only, say so explicitly: install tzdata, link /etc/localtime to UTC, set TZ=UTC, and assert date +%z is +0000 at boot. A deliberate UTC costs a megabyte and removes the whole class.

What happens to Postgres indexes when a base image upgrades glibc?

Potentially the worst kind of bug: wrong answers with no error. A btree index on a text column is physically ordered by the collation in force when it was built, and the planner trusts that order. A glibc upgrade can change what strcoll returns — glibc 2.28 rebuilt its collation tables from ISO 14651 and changed orderings broadly. After such a change the index's order no longer matches the comparison function, so an index-only equality lookup can miss rows that exist in the heap, and a unique index can admit a duplicate. Postgres 13 and later record collation versions in pg_collation.collversion and pg_database.datcollversion and warn when they change. Be careful here: ALTER DATABASE ... REFRESH COLLATION VERSION only updates the bookkeeping and stops the warning. REINDEX is what actually fixes the index. The structural implication is that a base-image bump under a database is a schema-affecting change even when no migration moved, and the cheapest long-term answer for application indexes that never needed linguistic ordering is a byte-order collation, which has no version to drift.

Does a Firecracker snapshot restore bring back a stale clock?

Yes — a snapshot captures CLOCK_REALTIME with everything else, so a restored guest resumes believing it is the instant the snapshot was taken. On PandaStack the agent re-syncs the guest's wall clock from host time over the guest exec bridge on every restore-family ready path: snapshot-restore create, resume and wake. It is best-effort with one retry and never fails a boot. That code exists because of a real incident worth describing precisely, since everyone guesses the symptom backwards. The guest's clock was behind, frozen at its seed's bake date. An upstream rotated onto a leaf certificate whose notBefore was after that date, so the guest rejected a perfectly valid certificate as not yet valid — not expired — and every git clone inside a restored guest failed with "server certificate verification failed". The error text mentions the CA file, so everyone looked there first; the CA bundle was fine. The residual risk after the fix is a jump rather than a drift: a process that cached a timezone offset, a formatted date or a token expiry at startup still holds a value computed for the wrong instant.

Keep reading

Related posts

  • Your Firecracker Snapshot Restore Failed: A Field Guide

    A restore either works in tens of milliseconds or fails with an error that tells you almost nothing. Here are the seven classes of Firecracker snapshot restore failure, what each looks like, and how to bisect your way to the real cause.

  • Running Nix Builds Inside a Disposable VM

    Nix's sandbox is a hermeticity fence, not a hypervisor — a hostile derivation still runs on your kernel. Here's how to put Nix inside a disposable VM without a cold /nix/store making every build miserable.

  • Firecracker vs the Nix build sandbox: different jobs

    "Firecracker vs the Nix sandbox" is a category mix-up worth untangling: Nix's sandbox makes builds pure and reproducible; Firecracker makes untrusted code hardware-isolated. They solve different problems — and for untrusted reproducible builds you want both.

  • The best Java and Spring Boot hosting platforms in 2026

    The JVM does not behave like the runtimes most platforms were designed around. It wants a big chunk of memory it can keep, several seconds to warm up, and a process that stays alive. Pick a host that offers all three.

  • Building a Minimal Firecracker Guest Kernel

    Firecracker boots an uncompressed vmlinux with a deliberately tiny config. Strip the drivers, filesystems, and subsystems a microVM never sees; keep virtio, the console, and the KVM guest bits. Here's the config, the boot args, and why the kernel is pinned into every snapshot.

More in Ephemeral databases · See Ephemeral Postgres databases on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.