all posts

Top 5 Scratch Postgres Platforms for CI (2026)

Ajay Kumar··10 min read

This is a roundup with a precondition. The job is narrow and specific: a Postgres that a CI job or a test suite creates, migrates, hammers with a few thousand assertions, and then destroys without a moment's sentiment. Not an analytics warehouse, not your production primary, not a dev database that has been alive since 2023 and has four indexes nobody can explain. A scratch database.

I run a platform that sells managed Postgres, so you can discount what follows accordingly. But I would rather be useful than persuasive, and the useful thing to say first is that most people reading this should not buy anything. For unit tests and the large majority of integration tests, a `postgres:16` container started inside the CI job, tuned to be unsafe on purpose, plus one template database cloned per test worker, is faster and cheaper than every platform in this post, including mine. It has no network hop, no provisioning wait, no bill, and no vendor whose status page you will come to know intimately.

So the first section is that. Then the five conditions under which it stops being enough, which are real and specific. Then the platforms.

Do this first: the container you already have

Two ideas carry almost all of the weight here, and neither is a product.

The first is that a test database does not need durability. Postgres spends real effort making sure your data survives a power cut, and in a CI job a power cut is indistinguishable from a successful teardown. Turning `fsync` and `full_page_writes` off, putting the datadir on a tmpfs, and using unlogged tables for the throwaway bits removes most of the I/O your tests were paying for. This is the single largest free speedup available to a test suite that talks to a real database.

The second is that `CREATE DATABASE ... TEMPLATE` is a file copy. Seeding is the expensive step in almost every suite I have profiled — not the queries, the seed. If you pay for the schema plus the seed once, into a template database, every worker afterwards gets a genuinely independent database for the cost of copying its files. This is the closest thing to copy-on-write branching you can have without leaving your CI runner, and it is approximately free.

# The baseline. Do this before you shop for a platform.
# One Postgres per CI JOB, one DATABASE per test worker, cloned from a
# template that already has the schema and the seed in it.

docker run -d --name pgtest \
  -e POSTGRES_PASSWORD=test -e POSTGRES_DB=template_app \
  --tmpfs /var/lib/postgresql/data:rw,size=2g \
  -p 5432:5432 postgres:16 \
  -c fsync=off \
  -c full_page_writes=off \
  -c synchronous_commit=off \
  -c wal_level=minimal -c max_wal_senders=0 \
  -c max_connections=200 \
  -c shared_buffers=512MB

# WARNING, and I mean it: fsync=off and full_page_writes=off tell Postgres it
# may lose or corrupt your data on a crash. That is *correct* for a database
# whose entire lifetime is one CI job and whose datadir is a tmpfs that
# evaporates with the container. It is data loss waiting for a power cut on
# anything you care about. Never copy these flags into a staging or production
# config. There is no "but it is only staging".

until pg_isready -h localhost -U postgres; do sleep 0.2; done

# Pay for the schema + seed ONCE, into a template database.
psql "postgres://postgres:test@localhost/template_app" -f schema.sql
psql "postgres://postgres:test@localhost/template_app" -f seed.sql
psql "postgres://postgres:test@localhost/postgres" \
  -c "UPDATE pg_database SET datistemplate = true WHERE datname = 'template_app'"

# Then each worker gets a real, independent database for the price of a
# file copy. CREATE DATABASE ... TEMPLATE copies the source's files; it
# needs no other session connected to the template, which is why the
# seeding above happens before the workers start, not during.
for i in $(seq 1 8); do
  psql "postgres://postgres:test@localhost/postgres" \
    -c "CREATE DATABASE test_$i TEMPLATE template_app"
done

# pytest-xdist, jest --maxWorkers, go test -p: each worker reads its index
# from the environment and connects to test_$index. Teardown is one DROP
# DATABASE per worker, or just docker rm -f pgtest, which is what actually
# happens in CI.
`fsync=off` and `full_page_writes=off` are an instruction to Postgres that it may corrupt your data on an unclean shutdown. They belong only to a database whose entire lifetime is one CI job on a tmpfs. Every real incident I have read about these flags starts with somebody copying a CI config into a staging environment, and staging is where people put data they later discover was the only copy.

Testcontainers is the same mechanism with better ergonomics and a lifecycle manager: it starts the container from your test code, waits for readiness properly, and tears it down even when the suite panics. If you are in Java, Go, Node or Python and you want this without writing the bash, use it. The constraint it inherits is nested Docker — your CI runner has to be able to start containers, which on some hosted runners means a privileged daemon, and on others means a long argument with your platform team.

The five conditions that justify a platform

A container in the job is the default. Here is what breaks it. If none of these describe you, close the tab and go tune your seed.

  • You need a real network endpoint. A preview deploy, a Vercel or Render build, a mobile client, a Playwright run on a hosted browser grid, or a partner's webhook has to reach the database. Something living inside a CI job's network namespace cannot be reached from outside it, and tunnelling to it is a worse problem than provisioning a database.
  • You need a branch of production-shaped DATA, not an empty schema. A query plan on 40 rows is not a query plan. Pagination bugs, missing indexes, N+1s, and anything involving a statistics-driven plan choice simply do not reproduce on a seed script, and writing a seed that is statistically honest is a project, not a ticket.
  • You need point-in-time state. "Reproduce the state at 22:14 last night, just before the migration" is a thing you ask about once a quarter and cannot do at all with a container.
  • You need many of them concurrently. Fifty containers on one runner is not isolation, it is a memory-pressure experiment with a test suite attached. Once your matrix is wide, the per-database overhead has to live somewhere other than the runner.
  • You need it to persist past the job. A review environment that a human clicks through on Thursday cannot be backed by a database that died when the Tuesday build finished.

Notice that only one of those five is about speed. If your reason for shopping is "the container takes four seconds to start", you are optimising the wrong four seconds; the seed is almost certainly costing you more.

The five

Numbered for the headline, not ranked — the right answer inverts depending on which of the five conditions above is the binding one. I have ordered these from least to most machinery, and the first entry is deliberately not a product. Everything I say about other vendors is qualitative on purpose: pricing, limits and feature matrices move faster than blog posts, so verify every claim here against their current documentation before you commit a quarter to it.

1. A container in the job (Testcontainers, or your CI provider's service block)

What it is: `postgres:16` started by your CI provider's service block, or by Testcontainers from inside your test process. Ready in a couple of seconds, zero marginal cost, and the full real Postgres — extensions, `COMMIT`, `LISTEN`/`NOTIFY`, the actual planner. Genuinely good at: unit and integration tests, migration tests, anything where you control both ends of the connection. The catch: no reachable endpoint for anything outside the job, the seed cost is yours forever, it dies with the job, and concurrency is bounded by the runner's RAM. This is the right answer more often than any vendor will tell you.

2. A branching managed Postgres (Neon-style copy-on-write branches)

What it is: a managed Postgres whose storage layer is copy-on-write, so creating a branch of a database — data included — is a metadata operation rather than a copy. Neon is the reference implementation of the idea; several others now offer a variant of it. Genuinely good at: exactly the problem a seed script cannot solve. If your binding condition is "production-shaped data per pull request", this architecture is the one that is actually designed for it, and branch creation is fast enough to put in a workflow step. The catch is that you are adopting a storage architecture, not just a database: understand how it behaves under sustained write load, what a branch costs once it diverges, how long branches live before something reaps them, and what the connection story is (poolers are usually mandatory at CI concurrency). Read their current docs — this is the fastest-moving category on this list.

3. A general app platform's database add-on (Render, Railway, Fly, Heroku-style)

What it is: you already deploy preview environments on an app platform, and that platform will attach a Postgres to each one. Genuinely good at: the review-environment case, because the database arrives wired to the app with no work from you, and the platform's own preview lifecycle is what deletes it. If your binding condition is "a human needs to click through this on Thursday", the integration is worth more than any individual feature. The catch: provisioning is typically tens of seconds to a few minutes, branching production data is usually not part of the deal, and the per-preview database is frequently the line item that makes someone ask what preview environments cost. Check how aggressively the platform reaps them, because an orphaned add-on database bills quietly and forever.

4. The local-first options (Supabase CLI, embedded-postgres, pg_tmp)

What it is: a family of tools that give you a Postgres on a developer's laptop or a CI runner with no container daemon in the way. `pg_tmp` starts a throwaway cluster on a socket and reaps it after a timeout. The embedded-postgres libraries (Java, Go, Node and others) download a real Postgres binary and run it in-process. The Supabase CLI stands up a local stack if that is the shape of your app. Genuinely good at: making the local test run identical to the CI run, and working where nested Docker is forbidden. The catch: you are now managing a Postgres build and its version matrix yourself, `initdb` is not instant, and these are emphatically local tools — no endpoint, no data branching, no persistence. Also, every one of them has its own idea of where the socket lives, which is its own small genre of CI failure.

5. PandaStack managed Postgres (a dedicated microVM per database)

What it is: a `postgres-16` template booted as a dedicated Firecracker microVM with a durable volume, reachable over a TLS-required `postgres://` connection string. Each database is its own kernel and its own machine, not a schema in a shared cluster — a noisy test suite cannot touch another database's page cache or steal its CPU, and `ALTER SYSTEM`, superuser-ish work and extension installation are yours because the instance is yours. The template is baked at 1 GiB of RAM with 4 GiB and 16 GiB tiers available; 8 burstable vCPUs either way. Clone and point-in-time restore produce a NEW database id from the source's archive, leaving the source untouched.

The catch, stated plainly, because it is the number that matters for this post's job: create takes 30 to 90 seconds. That is not competitive with a container start and it never will be, because the time is not platform overhead — it is PostgreSQL bootstrapping itself. A real `initdb`, a real cluster coming up, a real readiness check. We boot the microVM from a snapshot in well under a second; Postgres then takes the rest. If your binding condition is per-test-worker provisioning, use entry one. Where this shape wins is when you want a real endpoint that survives the job, hard isolation between concurrent databases, and the clone-to-a-timestamp path for the bug that only exists in production's data.

The five, on the axes that decide a CI suite. Verify every vendor claim against their current docs before you commit.
OptionWhat it isReady inBest forThe catch
Container in the jobpostgres:16 or Testcontainers on the runnerSecondsUnit + most integration testsNo reachable endpoint, dies with the job, seed cost is yours
Branching managed PostgresCopy-on-write storage, branch = metadata opFast enough for a workflow stepProduction-shaped data per pull requestYou adopt a storage architecture; branch sprawl and pooling need thought
App platform add-onA Postgres attached to each preview envTens of seconds to minutesReview environments a human clicks throughRarely branches real data; orphaned add-ons bill quietly
Local-first (pg_tmp, embedded)A real Postgres with no container daemonSeconds, after initdbParity between laptop and CI, no nested DockerYou own the version matrix; local only, no endpoint
PandaStack managed PostgresDedicated Firecracker microVM + durable volume30-90 s (Postgres bootstrap, not us)Real TLS endpoint, hard isolation, clone to a timestampFar too slow to provision per test worker

Schema per test, database per test, instance per test

Before you pick a platform, pick your isolation unit, because that decision determines how many databases you need and therefore which platform is even plausible.

Schema per test is cheap and fast: `CREATE SCHEMA`, set `search_path`, run, drop. It breaks in three specific ways that people rediscover every year. Code that hard-codes `public.` ignores your `search_path` and reaches straight into the shared schema. Sequences and extensions are schema-scoped objects but your migrations probably installed them once into `public`, so two tests end up sharing a sequence and your test on "the new user gets id 1" becomes order-dependent. And `CREATE EXTENSION` is generally a per-database act with a schema target, so pgvector or PostGIS tests do not cleanly partition this way.

Database per test (or per worker, which is what you actually want) gives you real `COMMIT`, independent sequences, independent extensions, and a `TEMPLATE` clone that costs a file copy. This is the sweet spot for most suites. What it cannot test is anything cluster-wide: roles and permissions, `ALTER SYSTEM`, replication, `pg_stat_statements` across the instance.

Instance per test is the only unit where "the suite corrupted the database" has no blast radius and where cluster-level behaviour is testable at all. It is also the unit that makes provisioning time a first-class cost, which is where platforms earn their keep. Per job: reasonable. Per test: a special kind of masochism.

max_connections is the thing that fails at 3am

Here is the failure mode nobody plans for. Your suite goes parallel. Eight workers, each with a connection pool of ten, is eighty connections before a single test asserts anything — and your ORM probably opens a second pool for migrations, and your app code opens one per process, and the default `max_connections` on a stock Postgres is 100. The symptom is not a clean error, it is a flaky test that fails differently every run and only when the matrix is full.

Three things to do about it. Raise `max_connections` on your test instance and give it `shared_buffers` to match, because each connection is a backend process with real memory behind it. Set the pool size per worker explicitly, usually to two or three, instead of inheriting a production-shaped default — a test worker runs one query at a time. And on any managed platform, find out before you commit whether you are talking to a pooler, in what mode, and what that mode breaks: transaction-mode pooling will cheerfully sever session state, so prepared statements, `SET` statements, advisory locks and `LISTEN`/`NOTIFY` stop behaving the way your tests assume.

Every disposable-database system leaks

The orphaned test database is the most reliably reproduced bug in this entire category. Jobs get cancelled. Runners get preempted mid-test. Someone force-pushes while a workflow is running. A teardown hook runs inside the process that just segfaulted. The result is a database nobody owns, which on a container is free and forgotten, and on a managed platform is a line item that grows monotonically until finance asks a question.

Write the delete in a `finally` block, and then do not trust it. Put a run identifier in the label of every scratch database you create, and run a sweeper that lists them and deletes anything whose run has finished or whose age exceeds a threshold you are comfortable defending. If the platform supports a TTL on the resource, set it: a server-side deadline is the only cleanup that survives your client being killed. The money the industry spends on databases created for a test that failed in 2024 is genuinely funny, right up until it is your invoice.

import os
import subprocess
import pandastack

client = pandastack.Client(api_key=os.environ["PANDASTACK_API_KEY"])

# create() returns only when the thing is actually ready. Under the hood the
# API answers 202 immediately and the SDK polls GET /v1/databases/{id} until
# status == "running", because the honest number here is 30-90 SECONDS: a real
# PostgreSQL 16 initdb + bootstrap inside a dedicated Firecracker microVM on a
# durable volume. That is two orders of magnitude slower than starting a
# container and I am not going to dress it up.
#
# size picks the RAM tier: "1g" (the default -- the postgres-16 template is
# baked at 1 GiB), "4g", or "16g". It is fixed at create time. Firecracker
# cannot change guest RAM at snapshot restore, so a running database cannot be
# resized; the supported resize is clone(size=...) into a new database id.
db = client.databases.create(
    label=f"ci-{os.environ.get('GITHUB_RUN_ID', 'local')}",
    size="4g",
    timeout=180.0,
)

try:
    # TLS is required on the wire. The connection_url is the credential, so
    # treat it like one -- do not echo it into a build log.
    url = db["connection_url"]  # postgres://pandastack:<pw>@<id>.db.pandastack.ai:5432/pandastack
    env = {**os.environ, "DATABASE_URL": url + "?sslmode=require"}

    subprocess.run(["alembic", "upgrade", "head"], check=True, env=env)
    subprocess.run(["pytest", "-n", "8", "tests/integration"], check=True, env=env)

finally:
    # The whole reason to use an API instead of a dashboard. A database that
    # outlives the job it was created for is the single most common way these
    # bills get surprising, and "the job was cancelled at 40%" is not a rare
    # event -- it is Tuesday. Put the delete in a finally block AND run a
    # sweeper over client.databases.list() that kills anything whose label
    # names a finished run.
    client.databases.delete(db["id"])

If the data came from production, that is a legal question first

The strongest argument for data branching is also where the whole approach gets dangerous. The moment you clone production into a test environment, every obligation attached to that data comes with it — GDPR, HIPAA, PCI scope, your own DPA with your own customers, and whatever your SOC 2 auditor believes about where personal data is allowed to live. A copy in a CI environment is a copy. It is reachable by every engineer who can read a build log, and build logs are not an access-controlled system.

So do the legal question before the technical one, and when you get to the technical one, be suspicious of your own anonymisation. Masking emails and names is the easy 80% and the useless 80%: free-text fields leak identities, foreign keys and timestamps re-identify people by correlation, and a "scrubbed" dataset where one customer has 400x everyone else's volume is not anonymous to anybody who knows the business. Either anonymise at the point the copy is made, never afterwards and never by hand, or synthesise data with production's statistical shape and skip the problem. The second option ages better.

# The forensic path: branch the DATA, not the schema. clone() builds a NEW
# database id from the source's archive -- the source is untouched, and it
# works against running, hibernated, even failed databases.
#
# target_time is a point-in-time restore: RFC3339, at least 2 minutes in the
# past (backups trail live writes), no later than the last archived WAL.
clone = client.databases.clone(
    production_db_id,
    label="repro-bug-4471",
    target_time="2026-10-04T22:14:00Z",
    size="16g",      # land the repro on a bigger tier than the source
    wait=False,      # 202 now; poll yourself if you have other work to do
)
ready = client.databases.wait_until_ready(clone["id"], timeout=600.0)

# It replays WAL, so a large database takes minutes, not seconds. This is a
# debugging tool and a "what did the migration actually do" tool. It is not a
# per-test fixture, and anything you clone out of production is production
# data with production obligations attached until you have anonymised it.

What to actually pick

If you are writing unit tests or integration tests you control both ends of: a tuned `postgres:16` container in the job, one template database with the schema and seed baked in, one real database per worker cloned from it. Spend the afternoon you were going to spend evaluating vendors on making your seed cheap instead. You will get a bigger win.

If your tests fail on real data and pass on seeds, you want copy-on-write branching, and that is a storage-architecture decision worth taking seriously rather than a feature you bolt on.

If a human or an external system needs to reach the database, you want a real endpoint, and then the question is whether to wire it into a preview-environment lifecycle (an app platform's add-on) or to have it isolated and durable with a clone-to-a-timestamp escape hatch (what I built). Mine is the slowest thing on this list to create and I would still choose it for the 3am "what did the data look like before the migration" question, where a container has nothing to offer at all.

And whichever you choose: label everything with the run that made it, and write the sweeper on day one. The platform is not the hard part. The cleanup is.

Frequently asked questions

Is a postgres container really faster than a managed ephemeral database for tests?

For the test case, almost always yes, and by a margin that surprises people. A `postgres:16` container on the CI runner is ready in a couple of seconds, there is no network hop between your test process and the database, and cloning a prepared template database per worker with `CREATE DATABASE ... TEMPLATE` costs a file copy rather than a re-run of your seed script. Any managed platform has to provision something and then put a network between you and it. Our own managed Postgres takes 30 to 90 seconds to create, because that time is PostgreSQL's own initdb and bootstrap rather than platform overhead, and no amount of fast VM booting changes it. The reasons to pay for a platform are structural rather than speed: you need an endpoint something outside the job can reach, you need a branch of real production data, you need point-in-time state, you need high concurrency without loading it onto one runner, or you need the database to outlive the job.

Is it safe to run Postgres with fsync=off in CI?

In CI specifically, yes, and it is one of the largest free speedups available to a suite that uses a real database. `fsync=off`, `full_page_writes=off` and `synchronous_commit=off` tell Postgres it does not have to guarantee that committed data survives a crash. For a database whose whole life is one CI job, on a tmpfs datadir that disappears with the container, a crash and a successful teardown have the same outcome, so you are trading away a guarantee you were never going to collect on. Unlogged tables and a generous `shared_buffers` help further. The danger is entirely about blast radius: these settings mean a crash can leave a corrupted cluster, not just lose the last few transactions. Every real incident involving them traces back to a config being copied out of CI into a staging or production environment. Keep them in a file whose name makes that mistake obvious, and never in a shared base config.

Why do my parallel tests fail with connection errors?

Because `max_connections` is almost certainly lower than your worker count multiplied by your pool size, plus whatever your migration tooling opens. Eight workers with a pool of ten each is eighty connections before any test runs, against a stock default of 100, and the failure looks like flakiness rather than a clean limit error — it only happens when the matrix is full, which is exactly when you are least able to debug it. Fix it from both ends. Set the per-worker pool size explicitly to two or three, since a test worker issues one query at a time and a production-shaped pool default is pure waste there. Raise `max_connections` on the test instance and increase `shared_buffers` with it, because each connection is a real backend process with memory attached. If you are on a managed platform, establish whether a pooler sits in front of you and in what mode: transaction-mode pooling breaks session state, so prepared statements, `SET`, advisory locks and `LISTEN`/`NOTIFY` will not behave as your tests expect.

Schema per test or database per test?

Database per worker, in most suites. Schema per test is cheaper but breaks in three specific ways that cost more to debug than the time it saves. Code that hard-codes `public.` bypasses your `search_path` and writes into shared state. Sequences and extensions usually got installed once into `public` by your migrations, so tests end up sharing a sequence and any assertion about generated ids becomes order-dependent. And `CREATE EXTENSION` is effectively per-database, so pgvector or PostGIS suites do not partition cleanly by schema at all. A database per worker gives you real `COMMIT` semantics, independent sequences and extensions, and a template clone that costs a file copy. What it still cannot test is cluster-wide behaviour — roles, `ALTER SYSTEM`, replication, instance-level statistics — and that is the only case where you genuinely need an instance per job.

Can I clone production data into a test database legally?

That is a question for whoever owns data protection at your company, and it should be answered before you build the pipeline, not after. A copy of production is production data with all of its obligations intact: GDPR, HIPAA, PCI scope, your contractual commitments to your own customers, and your auditor's view of where personal data may live. A test environment is usually reachable by every engineer who can read a build log, which is not an access-controlled system. If the answer is yes with conditions, anonymise at the moment the copy is made rather than afterwards, and be sceptical of your own masking — free-text fields leak identities, timestamps and foreign keys permit re-identification by correlation, and a dataset where one account has 400 times everyone else's volume is not anonymous to anyone who knows the business. Synthesising data with production's statistical shape is more work up front and causes fewer problems later.

Keep reading

Related posts

  • Top 6 Throwaway Postgres Platforms for Testing in 2026

    A database for a human to poke at and a database for forty parallel CI jobs are not the same product. The axes that decide the automated case are parallel isolation, seed-cost amortisation, and what leaks when a run is cancelled at 40%.

  • Ephemeral Databases for AI Agents

    Give an agent a run_sql tool and it will use it — including the DELETE it emits while debugging its own step. The fix isn't a better statement filter; it's a database that exists for one task and is deleted when the task ends.

  • Contract testing when you can run the real services

    Pact and its relatives were designed around a constraint: you can't run twelve services in CI. That constraint has weakened. What survives is the part about who owns the expectation.

  • Testing LLM-Generated Database Migrations Safely in a Sandbox

    An AI agent writing a migration is one hallucinated DROP TABLE away from ruining your night. Apply it to a disposable Postgres VM first, diff the schema, and throw the VM away.

  • Run Flaky Parallel Tests in Isolated MicroVMs

    Flaky integration tests are usually shared state in disguise — same DB, same ports, same /tmp. Give each test its own throwaway microVM forked from a seeded base, and the flake just… stops.

More in Ephemeral databases · See Ephemeral Postgres databases on PandaStack

Run code in a microVM in one API call.

49ms p50 cold start. Fork, snapshot, and scale to zero.

Start free
Written by Ajay Kumar, Founder, PandaStack.