Top 10 Throwaway psql Platforms for Engineers (2026)
Every roundup of throwaway PostgreSQL grades the same number: how fast you get a connection string. It is a real number and about a quarter of the right one. A database you can create in two seconds and then have to load forty gigabytes into is a twenty-minute database, and your test suite has never once cared which phase of those twenty minutes the delay came from.
The part engineers actually wait on is the seed. Watch any CI job that touches a real schema and the shape is the same: a few hundred milliseconds of provisioning, a few hundred milliseconds of migrations, then minutes of fixtures being inserted, indexes being maintained row by row, and a planner with no statistics making confidently terrible decisions. The vendor optimised the first bar on that chart. The tall ones are yours.
So the axis I want to grade these ten on is not latency. It is the provisioning verb -- what the platform actually does when you ask for a database. There are four verbs in this entire market. Fresh-and-empty runs your migrations and your seed every single time. Template-clone copies the bytes of a database you already built. A copy-on-write branch hands you a reference to a parent's pages and copies nothing until you write. A snapshot restore brings up a whole machine that already has the data inside it. Only two of those four are independent of how much data you have, and that single property reorders the entire shortlist.
I'm Ajay. I build PandaStack, which runs Firecracker microVMs behind an API and offers managed Postgres as one microVM per database. It is entry ten, and the awkward part goes here rather than in a footnote: it has the slowest create on this board by two orders of magnitude. Thirty to ninety seconds. Whether that disqualifies it or is irrelevant depends on which of the four verbs your problem needs.
The provisioning verb decides everything else
Here is the test. Multiply your production dataset by ten and ask what happens to time-to-first-query. "Roughly nothing" means the verb is size-independent and seeding has stopped being a capacity problem. "It goes up by about ten" means you bought a fast create and an expensive seed.
Verb 1: fresh-and-empty
An empty database, then your migrations, then your fixtures. Free, portable, where everyone starts, and the right answer for most unit tests, which need eleven rows and a foreign key constraint.
Its cost structure is the problem. Time-to-first-query is t_create plus t_schema plus t_data, and t_data is linear in rows with a constant far worse than people expect, because inserting into a table that already has indexes maintains every one of them per row. The fixture load is also the step most likely to be wrong: a seed script nineteen engineers have appended to over four years is not a specification of your data, it is an archaeology site. The failure mode is a suite that is green against eleven rows and a production query plan nothing has ever exercised.
Verb 2: template-clone
Plain PostgreSQL, no vendor required, and still the most underused trick in this market. You pay t_schema and t_data once into a template database, and every subsequent database is `CREATE DATABASE test_7 TEMPLATE app_template` -- a file copy inside the cluster rather than a replay of SQL statements.
It is still O(bytes); it genuinely copies, so it is not size-independent. But the constant is one to two orders of magnitude better than a logical restore, because there is no SQL parsing, no per-row index maintenance and no WAL for the data pages. For a parallel test matrix against a non-trivial fixture this is usually the highest-leverage change available, and it costs an afternoon rather than a procurement cycle. The sharp edge is exclusivity: no other session may be connected to the template while you clone it, which in a parallel runner surfaces as a flaky test rather than as the lock it is.
Verb 3: copy-on-write branch
The verb that created this product category. The storage layer hands the new database a reference to the parent's pages instead of a copy of them, so create is a metadata operation whose cost is independent of whether the parent holds four megabytes or four terabytes. Divergence is paid lazily, per page you write. This is the honest winner on time-to-first-query, and the only verb that makes "a database per pull request, with real data in it" economically sane.
Two caveats the marketing does not foreground. A free create does not make the first query free: a fresh branch has a cold page cache, and where compute is separated from storage the first touch of a page is a network round trip, so early queries on a new branch can be meaningfully slower than the same queries on the parent. And the verb needs a storage engine built for it, so you are not running the same PostgreSQL build on the same kind of disk as production.
Verb 4: snapshot restore of a whole machine
Restore a machine image that already contains the data directory, resume the process, serve queries. The logical size of the dataset is not in the critical path at all: you map a disk and start a process, you do not replay statements. A thousand rows or a billion, the restore does the same work.
Size-independent like verb 3, but with a floor: there is a machine to provision underneath, and that floor is tens of seconds in every implementation I know of, including mine. Which makes verb 4 flatly wrong for a per-test fixture and uniquely right for a narrow set of jobs -- rehearsing a destructive migration against real data, reproducing a bug that only exists at production scale, and answering "what did this table look like before the deploy" at three in the morning. Verbs 1 and 2 have nothing to offer those questions. Not less -- nothing.
#!/usr/bin/env bash
# Two provisioning verbs side by side. The point is not which is faster today.
# The point is whose clock runs faster as your dataset grows -- because that is
# the term that decides this in eighteen months, not today's benchmark.
set -euo pipefail
# --------------------------------------------- VERB 1: fresh-and-empty
# initdb, migrate, seed. The honest baseline, and where everybody starts.
# t_create is tiny. t_schema is bounded by your migration count. t_data is
# linear in rows AND pays for every index on every table you load into.
time psql -q -c 'CREATE DATABASE app_test'
time psql -d app_test -f schema.sql # ~hundreds of ms
time psql -d app_test -f seed_40gb.sql # go get a coffee, or two
# Why t_data hurts more than the row count suggests: COPY into a table that
# already has indexes maintains every one of them per row. Load first, index
# after, and the same bytes land several times faster:
# pg_restore --section=pre-data -d app_test dump.pgc
# pg_restore --section=data -d app_test dump.pgc # no indexes yet
# pg_restore --section=post-data -d app_test dump.pgc # build them once
# Then ANALYZE, or the planner decides based on nothing and your "slow query in
# CI" is a statistics bug wearing a performance costume.
# --------------------------------------------- VERB 2: template-clone
# Pay t_schema + t_data ONCE into a template, then every later database is a
# file copy inside the same cluster. This is plain PostgreSQL. No vendor.
psql -q -c 'CREATE DATABASE app_template'
psql -d app_template -f schema.sql
psql -d app_template -f seed_40gb.sql
psql -q -c "UPDATE pg_database SET datallowconn = false WHERE datname = 'app_template'"
# The per-worker path. Still O(bytes) -- it is a copy, not a reference -- but a
# sequential local copy with no SQL parsing, no per-row index maintenance and
# no WAL for the data pages, which in practice is one to two orders of
# magnitude off the seed_40gb.sql line above.
for w in 1 2 3 4 5 6 7 8; do
psql -q -c "CREATE DATABASE test_w${w} TEMPLATE app_template"
done
# The gotcha that costs everyone one afternoon: CREATE DATABASE ... TEMPLATE
# requires that NOBODY else is connected to the template. One straggling
# connection from a previous worker and you get
# ERROR: source database "app_template" is being accessed by other users
# which in a parallel matrix reads as a flaky test rather than as a lock.
# --------------------------------------------- VERB 4: snapshot restore
# Restore a whole machine that already has the data in it. The logical size of
# the dataset is not in the critical path: you map a disk and resume a
# process, you do not replay statements.
#
# RDS, roughly:
aws rds restore-db-instance-from-db-snapshot \
--db-instance-identifier repro-bug-4471 \
--db-snapshot-identifier prod-nightly-2026-10-07
#
# PandaStack, roughly -- a clone is a NEW database id built from the source's
# archive, and target_time makes it a point-in-time restore:
# client.databases.clone(src_id, label="repro-bug-4471",
# target_time="2026-10-07T22:14:00Z")
#
# What verb 4 buys is a time-to-first-query that does not care how big the
# database is. What it costs is a floor: there is a machine underneath, and
# that floor is tens of seconds. Wrong for a per-test fixture; right for the
# 3am "what did the data look like before the migration" question, where verbs
# 1 and 2 have nothing whatsoever to offer.
Time-to-first-query, decomposed properly
Measure four terms separately. They respond to different interventions, and one number hides which is eating you.
- t_create -- from the API call to a process that accepts a connection. The only term vendors publish, and typically the smallest in the sum.
- t_schema -- your migrations. Bounded by how many you have and whether anyone has ever squashed them. Six hundred migrations replayed from zero on every create is a self-inflicted wound no platform can heal.
- t_data -- fixtures, dumps, restores. The term that scales with your dataset, and on most teams the largest of the four by a wide margin.
- t_warm -- statistics and cache. ANALYZE, plus whatever it takes for the pages your first queries touch to be somewhere faster than cold storage. Routinely ignored, and the reason a "slow query in CI" investigation so often ends at a missing ANALYZE rather than a missing index.
Stop your clock on a real query through a real index, not on a `SELECT 1` and absolutely not on the platform's own readiness field. A database answering `SELECT 1` has told you a process is listening. It has told you nothing about whether the planner has statistics, whether the buffer cache holds one useful page, or whether your migrations finished -- and I once watched a team chase flaky CI for a fortnight that was entirely a health check going green before the last migration committed.
#!/usr/bin/env bash
# seed-slope.sh -- measure the term nobody publishes.
#
# You are not measuring time-to-first-query. You are measuring its SLOPE
# against dataset size, which is the only thing that tells you whether a
# platform's quoted create latency survives contact with your 2027 database.
#
# Run it at three sizes and read the shape, not the absolute numbers. Flat
# means the verb is size-independent (copy-on-write and snapshot restore are
# roughly flat; template-clone is linear but cheap). Linear means you bought a
# fast create and an expensive seed, which is the same wall clock with better
# marketing on the front of it.
set -euo pipefail
SIZES=${SIZES:-"100000 1000000 10000000"} # rows in the fat table
for ROWS in $SIZES; do
# 1. Build the fixture ONCE per size, outside the clock. A harness that
# generates data inside the timed region is benchmarking your own
# generator, which is a real result about the wrong thing.
[ -f "fixture-${ROWS}.pgc" ] || {
psql -q -c "DROP DATABASE IF EXISTS fixgen"
psql -q -c "CREATE DATABASE fixgen"
psql -q -d fixgen -c "CREATE TABLE events AS
SELECT g AS id,
(random()*1000)::int AS actor_id,
repeat('x', 200) AS payload,
now() - (g || ' seconds')::interval AS at
FROM generate_series(1, ${ROWS}) g"
psql -q -d fixgen -c "CREATE INDEX ON events (actor_id, at DESC)"
pg_dump -Fc -d fixgen -f "fixture-${ROWS}.pgc"
}
start=$(date +%s.%N)
# 2. Whatever "give me a throwaway database" is on the platform under test.
# Replace these two lines and nothing else; that is the whole experiment.
DSN=$(your-platform create --quiet)
pg_restore -d "$DSN" --no-owner --jobs 4 "fixture-${ROWS}.pgc"
# 3. Stop the clock on a query that touches the data through the index you
# just built -- not on SELECT 1, and never on the platform's own "ready"
# field. A database answering SELECT 1 has told you a process is
# listening. It has told you nothing about whether the planner has
# statistics or the buffer cache holds one useful page.
psql -q -d "$DSN" -c "ANALYZE events"
psql -q -d "$DSN" -c "SELECT count(*) FROM events
WHERE actor_id = 7 AND at > now() - interval '1 day'" >/dev/null
end=$(date +%s.%N)
printf 'rows=%-10s ttfq=%.2fs\n' "$ROWS" "$(echo "$end - $start" | bc)"
your-platform delete "$DSN" >/dev/null
done
# Read the three lines as a slope. If ttfq grows ~10x when rows grow 10x, your
# platform's provisioning verb is fresh-and-empty whatever the marketing page
# calls it -- and the interesting optimisation is not a different vendor, it is
# moving your seed out of the create path entirely.
The ten, by verb
1. Testcontainers -- the honest baseline
Verb: fresh-and-empty, with an escape hatch. A real `postgres` container started and torn down by your test process, isolated by kernel namespaces around a database you were going to destroy anyway. It belongs in a roundup of platforms because it is better than most of them at most of what people buy them for, and it is free.
Its honest catch is the seed, which is to say it is this whole post. Every container starts empty and pays t_schema plus t_data per run. But both cheap verbs live inside it: bake schema and fixtures into a derived image so they arrive as a layer rather than as SQL, or start one container per job and use `CREATE DATABASE ... TEMPLATE` for per-worker isolation. Teams who do neither and then conclude they need a database platform are buying out of a problem they built themselves.
2. Docker Compose plus pg_dump -- the hand-rolled version
Verb: fresh-and-empty, restored from a dump -- how a surprising share of the industry actually gets a development database with real-shaped data in it. Free, debuggable, and the dump is a file you can reason about.
Two catches. The technical one: `pg_restore` time is dominated by index builds, so restoring in sections -- pre-data, data, post-data -- builds each index once at the end instead of per row, and `--jobs` often halves what is left. The organisational one is worse: the dump ages. Within a quarter, asking in Slack whether anyone has a fresher dump has quietly become load-bearing infrastructure.
3. Neon -- the reference implementation of verb 3
Verb: copy-on-write branch, in a storage layer built for it, with compute separated from storage. This is the architecture the rest of the branching business is copying or competing with, and it matters here for exactly this post's reason: branch cost does not depend on parent size. A terabyte branch and a megabyte branch are the same metadata operation, which is what makes a per-pull-request database carrying real data affordable at all.
Two things to verify in their current docs. A branch is cheap to create but not free to hold -- branches accumulate divergence, and the thing that made them cheap made them uncountable. And a fresh branch has a cold cache, so measure the queries you care about on a new branch rather than a warm one, or your instant database will be instant and then slow in a way your benchmark never saw.
4. Supabase branching -- a branch of the environment
Verb: depends, and establishing which is the first thing to do. Supabase's branching is positioned around a preview environment for a pull request -- not only a database but the project surface attached to it -- and historically a preview branch is produced by running your migrations and your seed, which is verb 1 with excellent ergonomics on top. It is also the entry most likely to have moved since I wrote this, so check what a branch carries today.
The case is strong when Supabase is already your platform: a branch is the whole environment, auth and storage and policies included, which is far more useful to a reviewer clicking a link than a bare endpoint. The catch is the verb. If a branch replays your seed, time-to-first-query grows with your fixtures -- which matters enormously at a hundred open pull requests and not at all at five.
5. PlanetScale -- branch-and-deploy-request, now with Postgres
Verb: branch, inside a workflow built around schema change rather than data volume. PlanetScale's lasting contribution here is cultural: the deploy request -- a schema change reviewed, tested on a branch, then applied to the parent as a controlled operation -- is the only mainstream product that has treated a migration like a code review. They have since added a Postgres offering alongside the MySQL one, and that is the thing to verify before shortlisting, because branching semantics and engine internals are the product.
It fits teams whose pain is schema change on a large table under uptime obligations, not teams whose pain is seeding. The catch is workflow shape: branch-and-deploy-request is deliberate and human-in-the-loop, and if you wanted four hundred anonymous ephemeral databases a day for a CI matrix, you are using a governance tool as a fixture factory.
6. Nile -- the throwaway unit is a tenant
Verb: this one reframes the question, which is why it earns a slot. Nile is Postgres built around tenancy as a first-class concept -- a tenant is an object the platform understands, so one tenant's data can be addressed and isolated as a unit. For a multi-tenant SaaS the thing you want to throw away is frequently not a database at all; it is one customer's worth of data inside one.
That matters for seeding more than it first looks. The expensive part of a per-pull-request database for a multi-tenant app is usually that you copied every tenant to test against one -- slow, and the worst available version of the compliance problem below. Where the tenant is the unit, t_data is one customer's rows. The catch is scope: a strong fit for a genuinely tenant-partitioned schema, awkward outside one, and a younger project than the incumbents here.
7. Xata -- copies of production, aimed at being safe to hand over
Verb: copy-on-write branch on Postgres, with the interesting part being what the product does about the fact that a branch of production is production data. Xata has reshaped itself substantially over the past couple of years -- the sort of thing to check rather than trust a blog post about -- and its current positioning pairs instant branching with tooling meant to make a branch safe to hand a developer: anonymisation and masking as part of how the copy is made, rather than a script run afterwards.
That is the right place to put it, and it is what most teams get wrong: anonymising after the copy exists means there was a window, usually a long one, in which the unanonymised copy sat in a developer-accessible environment. The catch is the usual one for a product that has changed direction -- verify what it does today, and how masking handles your free-text columns, because a rule that does not understand them leaks through them.
8. AWS RDS snapshot restore -- verb 4 at enterprise weight
Verb: snapshot restore, in the version almost every large organisation already owns. Restore a snapshot into a new instance, or point-in-time restore onto a timestamp, and you get a whole server with your data in it and none of your seed script in the critical path. For a migration rehearsal or a forensic reproduction it is often already approved, already inside the VPC and already in the audit trail -- which beats a better tool procurement has never heard of.
The catches are cost shape and latency floor. A restore produces a full-size instance with full-size storage, billed as such, and the most common failure is not technical: it is the instance from last quarter's incident that nobody deleted. Restore is not instantaneous either, and a freshly restored volume loads blocks lazily, so early queries can be slow in a way people misread as a configuration problem.
9. Prisma Postgres -- the shortest path from the ORM
Verb: fresh-and-empty, with the best ergonomics in that category for a team already living in Prisma. The database is reachable from the same tooling that defines your schema, `migrate` and `seed` are first-class rather than bolted on, and the distance from "I need a database for this branch" to a working DSN is about as short as developer experience gets.
The catch is this post's thesis, said in the friendliest way available: if create is fast and the seed is your own script, your time-to-first-query is your own script, and no platform polish changes that. The optimisation that helps is not a different provider -- it is getting the seed out of the create path. Verify what the architecture underneath gives you for branching and idle today, because that part has been moving.
10. PandaStack managed Postgres -- a microVM per database
Mine, so weigh it accordingly. Verb: snapshot restore of a whole machine -- a Firecracker microVM with a durable volume attached, rather than a schema in a shared cluster. One database equals one guest kernel, one filesystem, one set of connection limits and a `shared_buffers` nobody else is competing for. RAM comes in tiers -- `1g` by default, `4g`, `16g` -- and that is the only size knob, because Firecracker cannot change guest memory at snapshot restore. Managed databases are always charged committed memory and never overcommitted: a deliberately worse deal for me, and the only honest way to sell a database.
Create is 30 to 90 seconds -- bootstrap plus a readiness check, the slowest number in this post, and I am not dressing it up: for a unit-test fixture, use a container and a template database instead. What the shape buys is the two jobs the cheap verbs cannot do at all. `clone` with a `target_time` builds a new database id from the source's archive at a point in time, source untouched and still serving -- the 3am "what did this table look like before the migration" question as an API call instead of an incident. And idle needs no cron job: auto-suspend when nothing is talking to it, wake on the next connection, the wake paid on the connect path. The endpoint is real -- SNI-routed TLS at `<id>.db.pandastack.ai:5432` -- so psql, pgbench and your DBA's GUI work with no proxy shim.
The rest of the catches, because a roundup where the author's own entry has one catch is an advertisement. The unit is coarse: one VM per database, so a hundred throwaway databases is a hundred VMs and none of the per-test-database economics a shared cluster gives you free. The clone replays WAL to reach your timestamp, so size-independence holds for restoring the snapshot and not for catching it up -- a large database is minutes. And there is no TTL, deliberately, because a database that deletes itself on a timer is a liability; the deletion is yours to remember. The most expensive object in this post is the one somebody forgot to delete, and that is true of every row below.
| Option | Provisioning verb | Time-to-first-query grows with dataset size? | Idle behaviour | Isolation boundary | The honest catch |
|---|---|---|---|---|---|
| Testcontainers | Fresh-and-empty; bake or template to escape | Yes, unless the seed is baked into the image | Dies with the test process | Container on your runner | Your seed script is the latency |
| Compose + pg_dump | Fresh-and-empty, restored from a dump | Yes -- dominated by index builds | Runs until you stop it | Container on a shared host | The dump ages; Slack becomes the data pipeline |
| Neon | Copy-on-write branch | No -- a branch is a metadata operation | Compute scales down; branches persist | Branch in a managed service | Branches accumulate; a new branch has a cold cache |
| Supabase branching | Usually migrations plus seed, per branch | Yes if the branch replays your seed -- verify | Tied to the preview environment | Project and branch in a platform | The whole environment, with verb 1's cost shape |
| PlanetScale | Branch, inside a deploy-request workflow | No for the branch; the workflow is the gate | Per current plan behaviour -- verify | Branch in a managed service | Governance tool, not a fixture factory; check the engine |
| Nile | Branch or copy, with the tenant as the unit | No, if you copy one tenant and not all of them | Per current docs -- verify | Tenant, then database | Fits tenant-partitioned schemas; awkward outside them |
| Xata | Copy-on-write branch, masked at copy time | No for the branch; masking is the variable | Per current docs -- verify | Branch in a managed service | Product changed shape recently; verify everything |
| AWS RDS snapshot restore | Snapshot restore of a whole instance | No for the restore; blocks load lazily | Keeps running and keeps billing | Dedicated instance in your VPC | Full-size instance, full-size bill, nobody deletes it |
| Prisma Postgres | Fresh-and-empty, driven from the ORM | Yes -- migrate and seed are the latency | Per current docs -- verify | Database in a managed service | Great ergonomics around the expensive verb |
| PandaStack | Snapshot restore of a microVM plus durable volume | No for the restore; a point-in-time clone replays WAL | Auto-suspends, wakes on connect | Firecracker microVM, own guest kernel | 30-90s create; one VM per database; no TTL, so you delete it |
Four ways to make t_data go away
If you take one actionable thing from this post, take this list rather than a vendor name. The goal is not a faster seed. The goal is for the seed not to be in the create path at all, and there are exactly four ways to get there.
- Bake it into the artifact. Schema and fixtures become part of the image, the template or the snapshot, built once in CI and referenced by digest or generation thereafter, so the create path resolves nothing and executes no SQL. Highest leverage, and no purchase required.
- Template it. Build one database with the data in it, then clone it per worker with CREATE DATABASE ... TEMPLATE. Still linear in bytes, but with a constant so much smaller than a logical restore that it feels like a different category.
- Reference it instead of copying it. Copy-on-write branching makes create independent of parent size and pays for divergence lazily -- the only approach that makes a per-pull-request database carrying real data genuinely affordable.
- Shrink it. Subset production to a referentially-consistent slice -- one tenant, one month, the thousand rows that exercise your query plans -- or synthesise data with production's statistical shape. A 2 GB subset that preserves cardinality and skew tells you more about your plans than a 400 GB copy nobody can wait for, and it sidesteps most of the next section.
The first two are free and the second two are products, and that order matters: teams who evaluate platforms before exhausting one and two buy a solution to a problem they created themselves. I have been on the receiving end of that sales call, and the honest answer is "go template your fixtures first, and if you still have a problem in a month, come back".
#!/usr/bin/env python3
"""Verb 4, concretely: a Postgres that is a machine, created and destroyed.
Nothing here is a fixture for a unit test. Create is 30-90 seconds, which is
disqualifying for the inner loop and completely irrelevant for the two jobs
this shape is actually good at: rehearsing a migration against real data, and
answering "what did this table look like at 22:14 last night".
"""
from pandastack import Client
client = Client()
# A database is a Firecracker microVM with a durable volume attached, not a
# schema inside somebody else's cluster. `size` is the RAM tier -- 1g (the
# default), 4g, 16g -- and it is the only size knob, because Firecracker cannot
# change guest RAM at snapshot restore. Managed databases are always charged
# committed memory, never overcommitted: you get the whole tier.
db = client.databases.create(size="4g", label="migration-rehearsal-4471")
try:
# 30-90 seconds: bootstrap plus a readiness check, and the slowest create
# in this entire post. Budget for it in your orchestration rather than
# discovering it inside a 15-second CI step timeout.
client.databases.wait_until_ready(db["id"], timeout=180.0)
# SNI-routed TLS at <id>.db.pandastack.ai:5432 -- a real wire-protocol
# endpoint, so psql, pgbench, your ORM and your DBA's GUI all work with no
# proxy shim in front of them.
conn = client.databases.connection(db["id"])
print(conn["connection_url"])
# -------- the part that makes verb 4 worth its floor ------------------
# Point-in-time clone into a NEW database id. The source is untouched and
# stays serving. target_time is RFC3339 and must be at least two minutes
# in the past, because the archive trails live writes -- ask for "now" and
# you are asking for WAL that has not been shipped yet.
repro = client.databases.clone(
db["id"],
label="before-the-migration",
target_time="2026-10-07T22:14:00Z",
size="16g", # land the repro on a bigger tier than the source
)
client.databases.wait_until_ready(repro["id"], timeout=900.0)
# Caveat stated here rather than in a footnote: the clone replays WAL to
# reach your timestamp, so a large database is minutes. Verb 4's
# independence from dataset size holds for RESTORING a snapshot. It does
# not hold for catching that snapshot up to an arbitrary point in time.
# Idle needs no cron job: the database auto-suspends when nothing is
# talking to it and wakes on the next connection. The wake is paid on the
# connect path, so the first client back from a quiet weekend absorbs it --
# set your pool's connect timeout accordingly instead of finding out via a
# 3am alert about a health check.
client.databases.delete(repro["id"])
finally:
# The joke writes itself: the expensive object in this post is the one
# somebody forgot to delete. There is no TTL on a managed database,
# because a database that deletes itself on a timer is a liability, so the
# deletion is yours.
client.databases.delete(db["id"])
"Throwaway" is not a compliance category
Here is the trap nobody budgets for, and it is the direct consequence of succeeding at everything above. The moment your throwaway database carries real data it becomes a copy of production with a short intended lifespan, and every obligation attached to those bytes comes along: your DPA with your own customers, whatever GDPR or HIPAA or PCI scope the original sat in, and whatever your auditor believes about where personal data may live. Obligations attach to bytes, not intentions, and no regulator has ever been moved by the argument that you were planning to delete it.
How many copies exist, and who can enumerate them?
This is where cheap branching bites. The property that makes verb 3 wonderful -- a copy costs nothing -- also makes copies uncountable, and an inventory that cannot enumerate the copies of a dataset is not an inventory. One branch per pull request, a quarter of ignored pull requests, a few personal branches named after a bug number from April, and your register of processing activities is a sentence in a wiki page while the real number lives in somebody's dashboard. Make it a query: tag every ephemeral database with an owner and a reason at create time, and report anything older than a week to a human.
What makes the bytes unrecoverable, and how long does it take?
"Deleted" is doing a lot of work in most people's heads. Dropping a copy-on-write branch may mark pages unreferenced while the storage layer garbage-collects on its own schedule. Deleting a cloud database instance leaves automated backups and manual snapshots behind on purpose, because that is almost always what you want -- right up until the data you are obliged to erase is in one. If you have a deletion obligation with a deadline, get the vendor's real answer to "when are the bytes gone" in writing, and find out whether it covers their backups of your throwaway.
Where does the connection string end up?
The most likely path from a compliant environment to an incident is not an exotic attack. It is the DSN. A preview database's connection string gets printed by a migration step, posted by a bot into a pull request comment so a reviewer can click it, cached in a build log readable by the whole organisation for ninety days, and pasted into a Slack channel with four hundred members when something breaks. None of those is an access-controlled system, and real customer data behind a URL in a readable build log is a breach with a short timeline and a very boring root cause. Rotate credentials per ephemeral database, scope them to one network path, and treat the log as a hostile surface, because operationally it is.
Which region is it in, and who classified it?
A branch generally lands in the region its parent lives in, which is usually fine and occasionally a cross-border transfer nobody filed; a snapshot copied to a cheaper region for a reproduction definitely is one. The structural problem underneath both: an ephemeral database has no owner, so nobody classifies it. Your production database has a classification, an owner, a retention policy and a row in a spreadsheet; the forty copies in your preview environments have none of those, which is precisely why they cause the incident. Anonymise or subset at the point the copy is made, and inherit the classification along with the data.
What to actually pick
Map the verb to the job, not the logo to the shortlist.
- Unit and integration tests you control both ends of: Testcontainers or a service container, with fixtures baked into a derived image and CREATE DATABASE ... TEMPLATE for per-worker isolation. Spend the afternoon you were going to spend evaluating vendors making your seed cheap instead -- the win is bigger.
- A database per pull request that has to contain real data: copy-on-write branching -- the only verb whose economics work at a hundred open pull requests, and worth treating as a storage-architecture decision rather than a feature you switch on.
- A multi-tenant application whose tests only need one customer's data: copy the tenant, not the database. Tenant-aware platform or your own subsetting job, it is the cheapest and safest answer at once, which is rare.
- Rehearsing a destructive migration, or reproducing a bug that only exists at production scale: snapshot restore. Already-owned RDS restore if it is in your VPC and your audit trail; a microVM per database if you want a point-in-time clone into a fresh id as one API call.
- A human needs something production-shaped to poke at for a week: almost anything here will do, and the real question is who holds the credential and when it expires.
If you are evaluating platforms right now, run the slope harness above against two of them at three dataset sizes before you read another comparison table, including this one. Three numbers from your own schema beat every vendor benchmark combined, because the only question that matters is whether their fast create survives contact with your data -- and that is a property of your data, not of their marketing.
A database you can create in two seconds and must then load forty gigabytes into is a twenty-minute database. Grade the verb, not the create.
The reason I keep hammering the verb is that it is the one property that does not change when a vendor ships a new pricing page. Latency numbers rot, free tiers appear and vanish, product shapes get rewritten. But whether a platform replays your SQL, copies your bytes, references your pages or restores your machine is architecture, and architecture is what you are actually buying.
Frequently asked questions
What is the fastest way to get a Postgres database with real data in it?
There are two champions, depending on what you mean by fast. For the lowest time-to-first-query regardless of dataset size, it is a copy-on-write branch: the storage layer hands you a reference to the parent's pages, so branching a terabyte costs roughly what branching a megabyte costs, and divergence is paid lazily per page you write. For the fastest thing available with tools you already own, it is a template database: build one database with the schema and fixtures in it, then CREATE DATABASE x TEMPLATE y per worker. That is still a byte copy, but it skips SQL parsing, per-row index maintenance and WAL for the data pages -- one to two orders of magnitude off a logical restore. What is never fastest is running your seed script on every create, which is what most teams do. Measure your four terms separately: if t_data dominates, a faster create buys you almost nothing.
Is copy-on-write branching always better than a snapshot restore?
No, and conflating them is how people end up disappointed by both. Both are independent of dataset size, which is the property that matters most, but they differ in granularity and in floor. A copy-on-write branch is a database-level object with a near-zero create cost, which makes it right for high-cardinality use: a branch per pull request, hundreds live. A snapshot restore brings up a whole machine, which costs tens of seconds in every implementation I know of including mine, so it is wrong for a per-test fixture. What the machine buys is everything that is a property of the machine rather than the database: your own kernel, your own connection limits, your own shared_buffers with no neighbours, your own extension set, and a point-in-time restore that lands a whole server on a timestamp. For four hundred ephemeral databases a day it absolutely is not the right unit.
Why is my throwaway database slow for the first few minutes?
Two causes, routinely mistaken for each other and for a third thing that is usually not happening. The first is statistics: a freshly loaded database has no useful planner statistics until something gathers them, so queries get plans chosen from defaults, which can be spectacularly wrong on a large table. Run ANALYZE as part of provisioning rather than waiting for autovacuum, and a surprising number of slow-query-in-CI mysteries evaporate. The second is cache: a new branch or a freshly restored volume has a cold page cache, and where compute is separated from storage the first touch of a page is a network round trip. So the first queries against a brand-new database can be far slower than identical queries against the parent, which makes benchmarks run on warm databases flattering to the vendor. The third explanation, which people reach for first and which is rarely right, is a misconfigured instance.
Does deleting an ephemeral database satisfy a data deletion obligation?
Usually not on its own, and this is a question to ask vendors in writing rather than infer from a marketing page. Dropping a copy-on-write branch may mark pages unreferenced while the storage layer garbage-collects on its own schedule. Deleting a managed instance commonly leaves automated backups and manual snapshots behind by design, which is exactly what you want until the data you are obliged to erase is inside one. None of that is a vendor behaving badly -- durability and erasure are genuinely opposed requirements -- but it means "we deleted the preview database" is not the sentence a deletion request requires. Get the vendor's real answer on when bytes become unrecoverable, and design so you rarely have to ask: anonymise or subset at the moment the copy is made, so the ephemeral copies never held the data you would later have to chase.
Keep reading
- Seeding test data in ephemeral databases — The t_data term from this post taken much further -- baking fixtures, subsetting production, and why the seed script is the latency.
- Branching Postgres for a pull request — Verb 3 in practice, including the lifecycle problem of branches that outlive the pull requests that created them.
- Top 9 temporary PostgreSQL platforms — The sibling roundup, graded by how the database is produced -- a process, a schema, a branch, or a whole server.
- Managed Postgres on Firecracker microVMs — How the one-VM-per-database shape works: durable volumes, RAM tiers, SNI-routed TLS and the point-in-time clone.
Related posts
- Top 7 Disposable Postgres Platforms for Developers (2026)
Three of the seven ways to get a disposable Postgres are not products you buy. Here are all seven shapes graded on time-to-DSN, isolation boundary, whether they branch real data, and the failure mode each one actually has.
- Top 5 Scratch Postgres Platforms for CI (2026)
Before you shop for a disposable-Postgres product: a tuned postgres:16 container plus CREATE DATABASE ... TEMPLATE beats every platform on the board for unit and most integration tests. Five options for the cases where it genuinely doesn't, including the one where my own platform is the slowest thing here.
- Top 6 Throwaway Postgres Platforms for Testing in 2026
A database for a human to poke at and a database for forty parallel CI jobs are not the same product. The axes that decide the automated case are parallel isolation, seed-cost amortisation, and what leaks when a run is cancelled at 40%.
- Top 8 Ephemeral Development Environment Platforms (2026)
Feature grids do not decide this. Two numbers do: how fast a fresh environment can exist, and what happens to it when nobody is looking. Eight platforms graded on both, with the substrate each one actually isolates with and the catch I would want before the purchase order.
- Your Backup Is a Hypothesis: DR Drills in Disposable microVMs
Almost nobody tests their restore path, because testing it properly means restoring production data somewhere real — and the available venues are production itself and a staging environment that drifted two years ago. Here is a third one.
More in Ephemeral databases · See Ephemeral Postgres databases on PandaStack
49ms p50 cold start. Fork, snapshot, and scale to zero.