Testing & benchmarks
Test commands, SQL checks, browser coverage, and benchmark workloads.
One set of focused runners covers fast correctness checks, real-browser behavior, performance regressions, and the benchmark harness behind the live benchmarks page. This page is the source of truth for running and maintaining them.
Treat a failure as a defect to explain. Expected answers come from SQL semantics, independent engines, or an independently authored model; never record Minnow's failing output as its new expected answer. An oracle correction needs evidence from that independent contract. Preserve failing inputs and reduce them into regressions. Do not loosen assertions, deadlines, coverage floors, or resource budgets to hide a failure, and do not treat a successful retry as a clean run. Storage and browser-runtime regressions must run against real browser APIs and durable stores; Node models and injected faults are useful additional checks, not substitutes for that gate.
The typecheck, build, and browser prerequisites force a fresh TypeScript project-reference build.
They do not trust an existing .tsbuildinfo, so a release gate proves the current sources even
after generated declarations or asynchronous signatures change.
Choose the smallest runner that proves the change
| Command | What it runs | Use it for |
|---|---|---|
npm test | Vitest plus the committed fast SQLLogicTest profile | Normal implementation work |
npm run test:unit | Vitest unit, generated, simulator, comparison, and harness tests | A focused inner loop |
npm run test:sql:standard | 6,345 result-checked upstream queries; committed, offline, about fifteen seconds | Every SQL or optimizer change |
npm run test:sql:full | The full supported profile, including exhaustive multi-way join permutations | Before a release |
npm run test:coverage | Vitest with the coverage floors enforced | Before pushing |
npm run soak | The generative suites on fresh random seeds, to find new failures | Hunting bugs rather than pinning |
npm run simulate | One replayable deterministic concurrency, fault, and storage simulation | Storage and transaction changes |
npm run test:simulator:full | 4,000 concurrent mutations plus deterministic fault and maintenance checkpoints, then three long interaction plans over the memory, IndexedDB, and OPFS stores | Before a release |
npm run simulate:interactions | One replayable interaction plan: DDL, DML, transactions, concurrent rounds, faults, reopens, and maintenance against a shadow model | Engine, index, and transaction changes |
npm run fixture:format | Freezes a database this build wrote, as a compatibility fixture | Before changing a format version |
npm run test:browser:library | Core IndexedDB, OPFS, transaction, SQLLogicTest, and multi-tab interaction-plan tests in Chromium, Firefox, and WebKit | Storage, engine, or browser-runtime changes |
npm run test:browser:conformance | Full core browser suite, full SQLLogicTest, and 800-step interactions over native IndexedDB and OPFS in all three browsers | Release and nightly storage correctness |
npm run test:browser:site | Public-site examples in Chromium, Firefox, and WebKit | Docs, examples, and site changes |
npm run test:browser | Both browser runners, in sequence | Cross-browser regression checks |
npm run test:consumer | Packed libraries installed, typed, bundled, and run outside the workspace | Package and public-entry changes |
npm run benchmark:gate | Seeded Node performance ratios against the checked-in baseline | Query-executor performance changes |
npm run benchmark:sizes | Download size of Minnow, SQLite Wasm, and PGlite, from the installed packages | Dependency or public-entry changes |
npm run api:check | The public API surface against the committed contract snapshot | Public-entry or declaration changes |
npm run check | Format, types, lint, build, API contract, coverage, and fast SQLLogicTest | The local merge gate |
npm run check:release | The local gate, full SQL/simulator, performance, and browser runners | Release candidates |
Install the browser binaries once before using a browser runner:
npx playwright install chromium firefox webkitThe merge gate and fast SQLLogicTest profile run in .github/workflows/ci.yml on every push and pull
request. Browser jobs run independently: Chromium and Firefox on Linux, and WebKit on macOS.
Each runs the library and site suites; Chromium also runs the packed consumer. The Real browsers
check requires every engine to pass. Splitting the jobs preserves their 45-minute deadlines and
leaves time to save failure artifacts. A test that passes only on retry still fails the gate.
.github/workflows/sql-conformance.yml runs the full SQL profile and long simulator
nightly and on demand. Both it and the release workflow also run
test:browser:conformance: the entire core browser suite, including the full pinned SQL corpus
and longer multi-tab plans over native IndexedDB and OPFS. WebKit runs on macOS so worker OPFS is available; a missing worker OPFS API
fails this profile instead of skipping its tests. Publishing requires the exact commit's successful
CI, SQL, simulator, and browser conformance gates. Retry a failed release through GitHub's rerun
control; there is no manual trigger that can bypass CI. .github/workflows/performance.yml runs the performance gate
nightly. The performance gate is deliberately kept off the merge path — it measures a machine as
much as it measures the code, and a merge gate that fails on a noisy runner is one people learn to
re-run rather than read. When a completed run crosses a ratio threshold, the gate starts a fresh
process and fails only if that same workload and comparison engine cross it again. Setup errors,
missing profiles, and incomplete runs still fail immediately.
Each browser runner owns its test directory, server, port, and build prerequisites. Shared browser
defaults live in playwright.shared.mjs; runner-specific configuration stays in its named
playwright.*.config.ts file. This keeps every runner independently callable without hiding which
application it starts. The library server disables file watching and hot reload, so editing a
shared checkout cannot navigate a page in the middle of a conformance test. Finish edits before
starting a release run so its workers and pages all load the same build.
The tarball-shape test packs an isolated copy: its prepack hook cannot rewrite the compiler
output while concurrent bundle and declaration tests read it.
Browser runs save an HTML report in playwright-report/, including attached seed plans and
result summaries. CI, release, and nightly jobs retain reports for passing and failing browser
runs for 14 days; failed tests also retain traces. Each engine's library report is saved before
the site runner can replace it. Successful attempts delete their persistent browser profiles;
failed attempts retain and upload the profile with the trace, including a failed first attempt
whose retry passes. Because macOS WebKit otherwise stores OPFS outside Playwright's profile
directory, its storage fixture redirects that browser-owned data into the per-test profile. The
failure fixture also copies every file from the isolated test origin through the browser File
System API. Its manifest records each original path, size, captured byte count, hash, and capture
error, so incomplete evidence is explicit. Evidence collection has a separate 20-second,
20,000-file, and 512 MiB cap and marks the manifest incomplete if any limit is reached. Library
traces retain action timing, errors, sources, and attachments;
the static runner pages omit continuous screenshots and DOM snapshots. A console log alone is not
a replay artifact.
Focused native recovery tests force real checkpoint and WAL boundaries, terminate the worker, and compare the reopened manifest list with the list acknowledged before termination. IndexedDB cleanup tests check integrity between bounded pages and after reopening, including a legacy cleanup marker, concurrent pruning, and deliberately unrelated gaps. Worker lifecycle tests check slow disposal, its absolute timeout cap, and cleanup after an already-reported fatal error. These reductions run in all three engines alongside the longer seeded plans.
The native SQL comparison runner initializes PGlite, SQLite Wasm, and Minnow sequentially to reduce overlapping startup allocations. Its trace records each engine's start and readiness, so a page-process failure can be distinguished from a SQL comparison failure. Every oracle, query, and result check still runs, and a test that passes only on retry still fails CI.
Site tests distinguish editor text from executed output. TypeScript errors must appear in the compiler diagnostics alert, and the guarded-upsert example must log exactly zero updated rows and one skipped row in the results panel. A matching word in the source editor cannot satisfy either check.
Unit files are capped at four concurrent workers (or the available CPU count, if lower). Each file can start complete databases and comparison engines, so unrestricted file concurrency can produce heap pressure and scheduling timeouts unrelated to the engine. The maintenance stress file runs alone after the other unit files, so the growing simulator seed corpus cannot consume its CPU time. Its own writer and background-maintenance concurrency remains intact, and coverage from both execution groups is combined. Test workloads, deadlines, assertions, and coverage floors remain unchanged. Browser runners likewise cap independent scenarios at two workers (one on a two-CPU runner); each concurrency scenario still creates its full set of tabs.
What each layer proves
- Unit and language tests stay beside the source they exercise. They cover predictable
behavior, supported SQL and mutations, comparisons with other engines, fault injection, and the
benchmark generator and answer-checking rules. Reads and writes are both checked
differentially: the read conformance suite diffs a generated query corpus against SQLite and
PGlite, and the mutation conformance suite replays a seeded script of inserts, updates,
deletes, upserts, and atomic write scopes against the same two oracles — comparing statement
outcomes, affected-row counts,
RETURNINGrows, trigger-written audit rows, and the full table state after every step. The fixture's triggers are declared in each oracle's own dialect (plpgsql trigger functions for PGlite) with identical bodies, andMERGE— which SQLite lacks — runs verbatim on PGlite and as an equivalent statement pair on SQLite.npm testpicks upapps/site/benchalongsidepackages/**, so the dataset generator, the oracles, and the suite definitions are checked on every unit run. - Library browser tests exercise IndexedDB, OPFS, and transaction behavior that a Node substitute cannot prove, and drive a database through a real module worker: the package entry booting, transferred buffers arriving intact, and a second worker reopening what the first one wrote. Storage fixtures use persistent browser profiles in every engine so durable adapters reach their disk-backed browser implementations. They also run the SQL engine itself where it ships: the committed SQLLogicTest profile executes through the published worker against real IndexedDB and OPFS in all three browsers, and a seeded interaction plan runs with each plan connection being a real tab, so the same properties the Node simulator checks are judged against real workers, real storage, and real crashes.
- The packed-consumer test builds every public package, installs the resulting tarballs in an OS-temporary Vite application outside the npm workspace, and resolves every exported entry with both Bundler and NodeNext TypeScript resolution. Chromium then runs the emitted worker against IndexedDB, migrates and queries through Kysely, observes a live update, streams exports, mounts the React and devtools adapters, exports a snapshot, restores it into OPFS, and reopens it. The Minnow archives and direct tool and peer versions are pinned to the build; npm resolves their ordinary transitive dependencies into a private temporary cache as it would for a new project.
- The public API contract assigns every package subpath a stable audience and snapshots the
declaration syntax exported by every typed entry point, including whether each name is
type-only or available at runtime.
npm run api:checkfails on an unclassified subpath or an added, removed, or changed public declaration;npm run api:reportprints the surface for review. After classifying compatibility and documenting the release impact,npm run api:updaterecords an intentional surface. Coverage has a separate non-regression floor for every published package, so the core suite cannot hide a drop in devtools or an adapter. The v1 support policy defines the audiences and guarantee. - Site browser tests execute the public examples and drive the benchmarks page itself: one
case runs a suite end to end in each browser and asserts it verified against the oracles,
another asserts the
/benchmarksroute is cross-origin isolated. Documentation cannot drift into non-running sample code, and the harness cannot drift from the page that runs it. Comparison suites require one result per declared engine/workload and report expected, attempted, supported, verified and failed counts. A missing live driver on an engine that declares subscriptions fails; the SQLite and PGlite comparison drivers' absent subscription support is declared explicitly rather than inferred from a missing driver method. Timings show medians. The exportedp95Msis an empirical percentile of five or seven timed windows, which is their maximum, rather than an estimate of production tail latency. - Semantic path equivalence generates NULLs, signed zero, exact decimals, invalid text, casts and empty/nonempty membership. It compares optimized and unoptimized row execution, vector execution, literal and parameter input, memoized results and worker RPC. Values and result domains agree exactly; refusals retain their error class. PostgreSQL independently checks acceptance and values without converting NUMERIC text through Float64. The same corpus runs over native IndexedDB and OPFS workers. Interaction runs also check that the bounded background history agrees with observer evidence.
- The performance gate detects regressions on stable seeded query shapes, reads and writes alike. Read-only engines calibrate their repeat counts independently so every fast path gets a measurable sample. A regression is reported only when the full ratio range observed across both engines' samples crosses the checked-in threshold, then repeats in a fresh process. Every stateful mutation window executes the same fixed 12-operation sequence on all three engines and finishes with a cross-engine state check. The suite includes exact-NUMERIC arithmetic, composite-key lookup, and the Kysely compilation/driver path as well as direct SQL. It is a guardrail, not a published cross-engine benchmark: the comparison engines run without indexes, which is fine for noticing that Minnow got slower and useless as a fair cross-engine claim.
Plain COUNT(*) reads use the snapshot's surviving row count without preparing column schemas
or index scans. The regression suite checks this path after deletes, updates, upserts, and
compaction, including old snapshots, concurrent connections, memory budgets, and cancellation.
Native-browser SQL comparisons
The library runner executes every SQL feature-matrix entry through Minnow's built module worker over both native IndexedDB and OPFS in Chromium, Firefox, and WebKit. SQLite Wasm and PGlite run in those same real browsers as independent comparison engines. Background maintenance stays enabled, and unexpected worker or page errors fail the campaign even when query answers match.
The 370 matrix entries are accounted for explicitly: 306 supported entries execute, and 64
unsupported entries must reject with their recorded error while leaving the fixture unchanged.
Portable reads compare columns and values with PostgreSQL. Compatible mutations compare affected
counts, RETURNING, and their actual target table's state, with domain checks for volatile values.
The report separately counts cross-engine comparisons and features with fixed behavior probes.
Every supported entry must have one of those semantic checks; acceptance alone cannot pass the
matrix. The category totals must cover the entire matrix.
An additional 137 independently authored behavior steps cover 74 matrix features, with two more
PostgreSQL state comparisons. They check TRUNCATE and UPSERT post-images, stored defaults and runtime
values, trigger effects and removal, constraint enforcement, generated columns, views, sequences,
identities, ALTER backfill, and transactions. Minnow's documented minimum-value ANY_VALUE result
has a fixed assertion even though PostgreSQL may choose another value. Runtime defaults must be a
valid timestamp, a finite random value in range, and a version-4 UUID. Negative probes require the
specific error type and message. The remaining DDL forms have observable schema checks, and all
28 nonportable reads have fixed answers or documented value-domain assertions. A new entry in a
category without cross-engine result comparison must provide a behavior probe, or the suite fails.
Single-argument AGE must agree with its documented two-argument form using the same statement's
CURRENT_DATE, including NULL propagation. BM25 checks matching rows and finite positive scores;
it does not invent an exact public scoring formula. Index probes check catalog behavior and query
results, while the separate performance suites cover execution cost.
Each seed also runs fixed edge cases, 42 generated queries, and four mutations with immediate post-state checks against SQLite and PostgreSQL: 117 independent comparisons per storage adapter and browser. Cases include NULL logic, joins, aggregates, windows, correlated subqueries, Unicode, JSON, numeric boundaries, and a literal question mark alongside a bound parameter. Numeric values compare exactly unless a particular aggregate or cross-runtime math expression declares a floating-point precision. JSON and numeric normalization follows result-domain metadata; ordinary text is never parsed as JSON.
The campaign replays the committed sql-conformance seeds and accepts MINNOW_SEED for new or
reproduced inputs. Its report records the seed and each coverage category.
SQLLogicTest: broad, strict SQL correctness
Minnow ships a native runner for SQLite's database-neutral
SQLLogicTest format. The parser and
answer checker live in @minnowdb/core/testing; the command-line runner streams one record at a
time and creates one empty database per file. A large corpus therefore does not become an equally
large in-memory syntax tree.
The runner implements the original format's statement ok/error, typed query, nosort,
rowsort, valuesort, result labels, MD5 result hashes, hash-threshold, skipif, onlyif
(with the trailing # comment the upstream corpus writes after an engine name), and halt
directives. Malformed records, wrong value counts, wrong types, duplicate labels with
different answers, and unknown directives are failures with exact source lines. Conditions are
reported as skips; they are never silently discarded.
The ordinary local profile is committed and needs no network:
npm run test:sql:standard
# Iterate on one source file or its first 200 records.
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/standard-select3.test
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/standard-select4.test --stop-after 200
# Show broad-suite progress, or trace every record from an isolated boundary.
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/full-select5.test --progress-every 100
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/full-select5.test --trace --trace-from 800It runs 1,822 statements and 6,345 queries, checking 385,389 scalar results across 8,169 source
records. The same committed files also run inside Chromium, Firefox, and WebKit
(sqllogictest.spec.ts), through the published module worker against real IndexedDB and OPFS with
automatic maintenance left on, so a browser engine's Date, regular-expression, or number-formatting
difference, a result that does not survive the worker boundary, or a durable store's commit path
under thousands of statements is checked against the corpus's recorded answers rather than
against Node's. Each run attaches its per-file timings to the report. Every thirty seconds,
native runs also log and attach their phase, completed statement and query counts, accumulated
timings, and a bounded prefix of the active SQL. This preserves progress when the unchanged
ten-minute per-file deadline expires. The fast set includes representative five-table and seventeen-table join-order families
that guard against pathological Cartesian work. On the development machine used to measure it,
the whole profile took about fifteen seconds. That makes the broad suite part of
npm test, npm run check, and every pull request, instead of a suite agents discover only after
they finish a change.
The supported full profile adds the exhaustive multi-way join families:
npm run test:sql:fullIt is fully runnable locally, reports progress every 500 records, and contains all 10,708 records
from the five selected upstream files: 1,822 statements, 8,884 queries, and 5,093,664 checked
scalar results. There are no all-profile exclusions. standard-exclusions.json still records the
2,539 supported exhaustive-join records that move from the fast profile to the full one solely
because of runtime.
The corpus is pinned to immutable revision
572273bf5076176e17f47aea75115a225e70fef8. upstream.json records its URL, date, 115,244,234-byte
archive size, archive SHA-256, five source hashes, and expected 622 test files. Regeneration first
verifies all of those inputs:
npm run test:sql:fetch
npm run test:sql:prepare-standardThe raw 622-file corpus remains locally runnable for boundary discovery. It is expected to expose
unsupported SQL; use --keep-going when auditing rather than treating it as a release profile:
npm run test:sql:upstream -- --stop-after 500
npm run test:sql:upstream -- --keep-going --max-failures 25Adding support means removing its exact exclusion, regenerating the committed files, and adding a focused regression test. Adding an exclusion without a concrete reason makes the ledger diff visible in review; there is no wildcard skip list.
When adding a runner, give it one responsibility and make it independently runnable. When adding a suite case, derive smoke-test counts from the suite itself instead of copying totals into tests or HTML.
Seeds, and how a soak failure becomes a permanent test
The generated suites — SQL support, write support, the deterministic simulator, the
columnar-versus-row comparison, and the compaction and multi-tab soaks — build their corpora from
a seeded generator. A committed run is deterministic: it uses the checked-in seeds plus every seed
that has ever failed, listed in packages/core/regression-seeds.json. Each recorded seed runs as
its own test case next to the suite's default, so a seed in that file is executed on every
commit, not merely listed. That makes the suite a reliable regression net, and on its own it
would be nothing else, because the questions it asks never change.
npm run soak is the other half. It runs the same suites on seeds nobody has tried and stops at
the first failure, printing the seed:
npm run soak -- --rounds 200A failing seed is the whole artifact. Replay it directly:
MINNOW_SEED=1476318588 npx vitest run packages/core/src/engine/sql-conformance.test.tsThen add it to regression-seeds.json under the suite that failed, where every future run picks
it up. Never remove one: a seed in that file is a bug that used to exist, and the entry is what
stops it coming back. The explored space only grows, and it grows by exactly what the soak found.
Format compatibility
A browser database's data lives in the user's browser, so it outlives every deploy. There is no migration window and no way to reach back and rewrite it: if a format change makes yesterday's blocks unreadable, the first anyone hears is a user whose application will not open.
packages/core/format-fixtures/ holds one frozen database per released block/snapshot pair — a
framed snapshot, which carries the raw block bytes verbatim, plus the answers that build gave to a
fixed set of queries. Each manifest records the exact @minnowdb/core version that wrote it.
format-compatibility.test.ts rejects a fixture that claims a future writer, reproduces the
current writer byte for byte, opens every fixture, checks the answers still hold, and checks that
writes into a restored database still work. Frozen raw vectors independently cover every
block-format-2 physical type.
The directory also holds one native fixture for every locked OPFS layout. The test requires a
fixture for the current writer and checks its recorded package version, then reopens each fixture's
format marker, checkpoint, WAL tail, acknowledgement file, and packed extent through the current
adapter before committing and recovering another write. Every supported older fixture must
upgrade automatically through ordinary open; an older-layout refusal cannot pass this test.
The ordered OPFS upgrade registry must cover every version between the first locked layout and
the current one. upgrades.test.ts injects failure at every preparation/publication I/O boundary,
models both tab death and power loss, exercises partial transfers and concurrent openers, and
keeps concurrent openers waiting through a slow live conversion. A resurrected ready record
cannot roll back later acknowledged writes, even with a missing/torn marker and missing/corrupt
completion receipt. Damaged native proof after resetting a nonempty old WAL is refused rather
than repaired from stale staging. The native browser
worker suite discovers the same frozen fixtures and checks automatic upgrade and new writes
across restart. A pinned test-only alias of the released core 0.10.0 package proves that the actual
layout-6 reader opens the old fixture and refuses the upgraded tree before mutation, in Node
and native workers. A second alias, of core 0.12.1, is the last IndexedDB schema-2 and OPFS
layout-7 writer: marker-upgrade.test.ts and indexeddb-schema3.test.ts have it leave a
merge-v1 fold in flight, then upgrade, finish that fold, write, reopen, and check that 0.12.1
refuses the result without a byte changing. A third alias, of core 0.13.1, is the last IndexedDB
schema-3 and OPFS layout-8 writer: marker-upgrade.test.ts upgrades its database to layout 9,
logs a commit whose frame spans continuation frames and whose index changes exceed one layout-8
chunk, reopens from the log and from a checkpoint, and checks that 0.13.1 then refuses the
database; indexeddb-schema4.test.ts does the same for schema 4 with every older writer. The
marker upgrade is cut at each of its eight I/O boundaries by tab death and by power loss, from
both older layouts. Unsupported pre-contract prototypes are refused without changing their bytes.
The layout-7 tests also zero or truncate acknowledged WAL bytes and
damage either acknowledgement slot; recovery must report corruption rather than roll back.
It also fails when the current BLOCK_FORMAT_VERSION or SNAPSHOT_FORMAT_VERSION has no fixture
behind it. That failure is the useful one, because it fires before the damage:
npm run fixture:formatThe native OPFS fixture has its own writer:
npm run fixture:opfs-formatRun the relevant writer on the build that still writes the old format, commit the fixture, and only then change the version. A fixture can only be produced by the build that writes it — once that code is gone, the format it wrote cannot be regenerated. For the same reason, never delete a released, locked fixture. Block format 2 is the first such lock; the unsupported experimental block-1 fixture was replaced before v1.
Internal orchestration and adapter transitions
Catalog mutations, per-connection writer admission, and query preparation/execution have separate internal owners. The public database API delegates to them; SQL kernels and adapter atomic publication remain shared with the ordinary paths. Refactoring these owners must preserve cancellation, reader leases, compaction turn lending, shutdown draining, and original failure reporting. The lifecycle, admission, semantic equivalence, and native interaction suites exercise those boundaries. Caller-provided clock and ID callbacks retain their original database receiver through the extracted catalog and query paths, with a lifecycle regression on all three adapters.
storage/commit-admission.test.ts runs the same multi-table commit refusals on memory, IndexedDB,
and OPFS. Invalid coverage, duplicate limits, invalid limits, and segment backpressure must leave
the manifest, transaction journal, and staged bytes unchanged. A corrected retry on that same
transaction must publish every changed table.
Record-core hardening tests verify bulk unique-key removal without per-key ordered-index work, canonical snapshot export and restore after mixed mutations, and exact batched block lengths and checksums after refusal, retry, and placement-only adapter replay. Snapshot ordering is lazy; constraint membership stays immediate. IndexedDB hardening also verifies that independent commit reads are queued together and a refused commit leaves the journal and staged bytes unchanged for a corrected retry. The controller extraction and storage changes add about 4 KiB to the raw OPFS bundle allowance while keeping the compressed download budgets unchanged.
Matched-index performance comparisons
npm run benchmark:compare runs the Node suite with primary-key indexes on every engine's
ordinary, mutation, and settled tables. It verifies query and final-state results before timing.
Use this command and the live browser comparison for competitive speed claims.
The existing benchmark:gate preserves its historical schema and runtime-specific thresholds,
including unindexed ordinary-key comparator tables. Its purpose is detecting changes against
that same workload. Its ordinary key ratios cannot support competitive speed claims. A comparison
run cannot update those thresholds; changing that gate's schema requires fresh calibration on
every recorded runtime, including Linux.
Running out of quota
quota.test.ts covers what happens when the browser refuses a write because the origin is out of
space — the characteristic way a browser database fails, and the one an application most needs to
handle deliberately. It pins four things: the write fails rather than half-landing, the error
keeps its QuotaExceededError identity so an application can branch on it, everything committed
earlier stays readable, and the same write succeeds once there is room, with no repair step.
Property-based tests
block-format/properties.test.ts describes the format as properties and checks them
against generated inputs rather than chosen ones: a value written and read back is the same value,
both codecs agree, a zone map contains everything it summarizes, and a flipped byte is detected
rather than decoded into plausible rows. The generators reach for what breaks encoders — negative
zero, the extremes of the double range, subnormals, empty and astral-plane strings, all-null and
empty columns. A failure shrinks to a minimal case and prints the seed and shrinking path. The same property
campaign runs in dedicated workers in Chromium, Firefox, and WebKit, using each engine's native
CompressionStream, DecompressionStream, UTF-8 conversion, dates, and typed arrays. Every
browser replays the same committed seeds; set MINNOW_SEED to reproduce a reported failure.
The block-format-properties entry in regression-seeds.json and the soak runner keep discovered
failures in the permanent regression set.
Soaks and deterministic simulation
Four suites cover accumulation rather than a single operation, which is where a fold that loses a row, a scheduler interleaving, or a compactor that stops folding actually shows up:
compaction-soak.test.tsruns 1,500 interleaved mutations against a referenceMap, compacting at checkpoints, and compares the whole table to the reference each time. Compaction is bounded and incremental — one call folds a limited number of blocks — so the test expects steady progress, not a finished table after one call.concurrency-simulation.test.tsruns twelve independent databases over one shared store, the shape of twelve browser tabs on one origin, issuing a seeded random schedule. It checks that the outcome is explicable: every visible row was written by some tab, acknowledged writes are present, no key appears twice, the write the store acknowledged last is the one visible, and every tab sees the same database afterwards.auto-compaction-soak.test.tsis the same kind of history with the background maintenance left on, as it ships: 1,800 seeded mutations and reads while folds and collection passes land underneath them, the table checked against the reference at every checkpoint, then settled and held to convergence bounds — a scan reads the bounded folded partitions and only a level-zero tail below both background fold thresholds, and the store holds the data rather than the history.simulator.test.tscontrols the completion order of every asynchronousBlockStorecall across independent database clients. A referenceMapchecks commuting concurrent mutations at each checkpoint; scheduled crash points must reopen to exactly the state before or after the write; compaction and garbage collection run throughout. The final quiescent store is bounded by the key space rather than the number of rounds across blocks, segment records, transaction records, and manifests, and must have no leaked leases. The diagnostic trace is a fixed-size tail, so the simulator itself does not grow with the run.
Run or save an exact plan directly:
npm run simulate -- --seed 24301 --rounds 500 --clients 8 --key-space 128
npm run simulate -- --seed 24301 --write-plan /tmp/minnow-plan.json
npm run simulate -- --plan /tmp/minnow-plan.jsonThe seed fixes both workload generation and storage completion choices. A failure prints its
replay command. npm run soak -- --suite deterministic-simulator explores fresh seeds, and a
failure joins packages/core/regression-seeds.json so every future unit run replays it.
Interaction plans
The storage simulator above drives one table with keyed upserts and deletes. The interaction-plan
simulator (interaction-simulator.ts, exported from @minnowdb/core/testing) drives the SQL
surface an application uses, in the spirit of Turso's simulator: a seeded plan of interactions
over several connections that share one database — CREATE TABLE with random column types and
nullability, multi-row inserts by literal and by parameter, predicate UPDATE and DELETE,
selects with three-valued predicates, LIMIT and ORDER BY, secondary indexes, DROP TABLE,
SQL transactions that commit or roll back, rounds of concurrent keyed writes with readers in
flight, storage faults and worker crashes, reopens, compaction and collection, and checkpoints.
The plan is JSON, so a failing seed replays exactly and can be checked in.
Generated plans preserve the randomized sequence and add a fresh-row insert for each configured fault point before the final checkpoint. These keys lie outside the random key range, so each probe can change the database and reach its write hook. An update or delete of a missing row can otherwise make a random fault step do no work. The fresh probes supplement those random cases; unsupported fault points are still reported as skipped by drivers that cannot inject them.
Every interaction is executed through the real SQL entry points and judged against a shadow model
that evaluates the predicates with SQL's three-valued logic. Beyond "the rows match", the runner
checks properties a database must honour whatever its internals do: a rejected insert leaves none
of its rows behind; reported row counts equal the model's; nothing matching a delete survives it;
LIMIT returns exactly the first rows in order; COUNT(p) + COUNT(NOT p) + COUNT(p IS NULL)
equals COUNT(*); a UNION ALL has the sum of its members' cardinalities; uncommitted rows are
visible on their own connection and invisible to another; a read taken during a round of
commuting writes shows every untouched key exactly and every touched key at its old or new
value; a mutation interrupted by a fault lands entirely or not at all and a reported success is
durable; and at every checkpoint every connection reads the same database as the model.
Two refusals inside a transaction are engine behaviour rather than defects, so the runner counts
them and continues: a COMMIT that lost the race to another connection's data commit, and a
transaction the engine rolled back after transactionIdleTimeoutMs of idleness. Both discard the
whole transaction, and an expired one is acknowledged with a ROLLBACK before the plan goes on.
See transactions for the rules themselves.
The runner also validates result shape: projected columns must match in order, each row must
have exactly its declared fields, and a missing field cannot stand in for NULL. Its expected-error
paths check the engine's specific error identity or exact documented refusal. An unrelated error
does not become acceptable just because a statement was meant to fail or a crash was injected.
Custom drivers must preserve error names and causes across their transport, including
DatabaseWorkerOutcomeUnknownError for a mutation interrupted by a worker crash.
The runner is driver-agnostic. In Vitest (interaction-simulator.test.ts) each connection is a
MinnowDatabase over one shared memory, IndexedDB, or OPFS store with the storage fault points
armed. In Playwright (interaction.spec.ts) each connection is a real tab holding the published
worker over a real IndexedDB or OPFS database, and the plan's faults terminate a worker with a
statement in flight — the tab then reports the loss to its client at once, so the in-flight call
fails as an unknown outcome instead of waiting out the client's silence deadline. The same seed
produces the same plan in both.
The browser driver recovers a crash the way an application must: by reloading the document. A new
worker in the same document is enough on Chromium and Firefox, but WebKit keeps the dead worker's
IndexedDB connection registered until its document goes away and blocks every connection to that
database meanwhile — so a reload is the remedy the engine's StorageUnresponsiveError names, and
the driver uses it. A deliberate WebKit/IndexedDB crash campaign that meets an
unresponsive store records it in transientsAccepted and stops, with interactions counting the
steps it judged. This exception is limited to that browser, store, and deliberate crash campaign.
Every browser/store pair also runs a separate campaign without fault steps that must complete
the entire plan; a crash exception cannot replace that coverage. Other early stops fail.
The standard browser campaigns generate 220 steps with three connections. The conformance profile
generates 800 with four connections. Each run attaches its exact generated plan before executing
it. The nightly workflow varies a logged MINNOW_BROWSER_INTERACTION_SEED_BASE input; reproduce that frontier with:
MINNOW_SEED=123456 MINNOW_BROWSER_INTERACTION_SEED_BASE=123456 npm run test:browser:conformanceThis fixes the generated plan, while actual browser scheduling remains uncontrolled. Node's IndexedDB and OPFS simulations are additional diagnostic checks, not proof of native durability.
npm run simulate:interactions -- --seed 24301 --length 300 --connections 4 --store indexeddb
npm run simulate:interactions -- --seed 24301 --write-plan /tmp/minnow-interactions.json
npm run simulate:interactions -- --plan /tmp/minnow-interactions.json --store opfs
npm run soak -- --suite interaction-simulatorA failure names the interaction, the property, the model's rows, and the SQL leading up to it.
The first soaks of this simulator found seven defects the existing suites had missed — an
index-pruned scan that resurrected deleted rows, a composite index that skipped rows with a NULL
trailing column, a stale intermediate value served by block-level pruning, a fold that refused
any transaction writing two segments to one table, a zone-pruned delete that made a reinserted
key fail every keyed range read, a compaction refused by a background index build, and an
IndexedDB key cache that let two tabs write one primary key twice — each now pinned in
simulator-regressions.test.ts. The last one only shows with a separate adapter instance per
connection, which is why the Node driver opens one per connection for IndexedDB and OPFS. Its
first real-browser runs then added two more: WebKit keeping a terminated worker's IndexedDB
connection and open transaction alive until its document goes away, which wedged every tab
(now bounded by StorageUnresponsiveError), and a crashed tab's pending call waiting out the
client's silence deadline long enough for another tab's idle transaction rollback to fire.
Background maintenance, size, and memory
auto-maintenance.test.ts pins what autoCompact and autoCollect guarantee, as bounds rather
than timings: a burst of writes folds back down and, a quiet minute later, collects to a store no
bigger than twice the data; repeated bursts settle to the same place instead of each leaving a
residue; versions from a moment ago stay readable and old ones are pruned; the switches off do
nothing; the buffer pool stays inside its budget under many distinct queries.
automatic-compaction-fit.test.ts holds automatic folds to the same standard on tables whose
folds do not fit the default memory budget as planned: a wide table refreshed wholesale keeps
folding inside the budget, a fold whose distinct keys outgrow it is cut, the smallest fold gets
the memory it needs, and a refresh written in an unrelated order folds with a job record a few
kilobytes long. A fold resumed by another engine recomputes its replay, and one whose replay no
longer matches its plan is abandoned unwritten. A single delta as large as its table folds with
no further write, including a lone one over already-folded partitions. Each case fails with its
part of the fix removed.
large-delta-reads.test.ts refreshes 400,000- and 1,000,000-row tables with one full-table
upsert under default settings, on memory, OPFS, and IndexedDB stores, and checks exact results
at once — while the fold runs — and after it; it also folds such a table with a compactTable
that names no budget. overlay-fallback.test.ts forces the replay's two bounded paths — patches
replayed per range of rows, and touched keys replayed in hash partitions — through projections,
aggregates, ordering, index-selected rows, layered partial updates, deletes, and reads of older
versions, and requires the in-memory replay's exact rows in its exact order; one case reaches
both paths through a real small budget. delta-scan.test.ts runs its SQLite-checked mutation
script through both paths as well.
event-loop-stalls.test.ts runs a 1 ms watchdog through the workloads that once held every
other query for hundreds of milliseconds or minutes — a fold of reordered wide upserts, live
aggregates patched after large commits, scans over upsert-heavy tables, secondary-index builds
and their first lookups, a full-text index build, and a 600,000-row write and upsert on OPFS —
and fails if any stretch the watchdog
cannot run exceeds 250 ms (750 ms in CI). Each case fails against the code before 0.13.0.
checkpoint-slicing.test.ts pins that a multi-slice OPFS checkpoint gives the event loop turns
between its slot writes, publishes exactly the bytes a one-step checkpoint would, and survives
power loss or a refused write at every one of its file operations, with the refused checkpoint
retried by later writes; wire.test.ts checks the sliced encoder against the one-step encoder
byte for byte, and the sliced checkpoint parser against JSON.parse, each with a property test
over arbitrary JSON-shaped states and, for the parser, malformed input it must refuse.
large-commit.test.ts pins on the memory and OPFS stores that a commit large enough to check its
keys in slices still refuses a conflicting key without changing anything, that its
ahead-encoded log frame replays after a crash, and that an index change of 70,000 postings —
more than one stored chunk — commits, replays from the OPFS log, and loads from a checkpoint.
wal.test.ts writes payloads that span continuation frames and checks that replay joins them,
treats pieces with nothing after them as an unwritten tail wherever the log stops, takes pieces
back when an append fails, and refuses a corrupt piece inside the acknowledged log.
follower-large-commit.test.ts commits 400,000 indexed rows from an OPFS follower — past the
message's structural limit — and checks that a re-sent large commit runs once under its first
delivery's identity. index-reads.test.ts pins that an OPFS index lookup answers while a large
commit holds the leader's queue, and that a chunk read failing partway runs again on the queue.
lease-lane.test.ts pins that OPFS reader leases run between the slices of a large commit and a
checkpoint, with every power-loss and torn-write boundary of such a checkpoint recovered.
unique-index-staged-build.test.ts pins, on the memory, IndexedDB, and OPFS stores, that
creating a UNIQUE index stages its keys in ordered chunks of at most 4,096 and never sends
them in one catalog update, enforces the index across a reopen, handles an empty table, and
releases staged keys when publication fails. index-capacity.test.ts builds a secondary index
over 600,000 distinct values and a full-text index over 450,000 documents, both of which the
chunk limit once refused, and pins that at most one chunk per block of rows covers a term, so a
lookup reads no more chunks than before.
npm run benchmark:stalls measures the real numbers for those shapes and a few more, during the
writes and during the maintenance that follows; -- --store opfs runs them over the native
OPFS store.
The heap itself is checked by the memory soak, which the unit suites cannot do honestly:
npm run soak:memory # node --expose-gc; four workloads, eight rounds each
npm run soak:memory -- --rounds 12 --workload writesEach workload — queries, writes with background maintenance, live query sets opened and closed, databases opened and closed over one store — runs in rounds with a forced collection and a heap reading after each. The assertion is a plateau: the last rounds may not sit more than a bounded margin above the early ones. The database runs with a small buffer pool so the pool fills during the warm-up rounds, which are excluded; a climb that continues past a full pool is a leak.
Concurrent writes
write-coordinator.test.ts covers the writer queue itself against a controlled lock manager:
arrival order, one pending lock request per context, prompt cancellation of a queued or
lock-waiting writer, one stall report per unchanging holder with no bypass, silence while
holders keep changing, and separate queues for separate databases.
It also refuses to invent a stalled holder from an empty or unavailable lock snapshot and
proves that writers behind the local queue do not inspect remote locks. The browser corpus
retains every console error's structured arguments in its report while still failing on errors.
write-contention.test.ts covers writers that do and do not take turns. Coordinated writers all
land; a store marked uncoordinated through the test seam (markStoreUncoordinatedForTests, the
stand-in for an older build in another tab) accepts only maxCommitRetries + 1 of a synchronized
IndexedDB burst, and that budget does not grow with concurrency. Rejected writes publish nothing
and acknowledged writes remain fully present. The same seam is what every test of the defensive
compare-and-swap path uses — a commit landing under an open scope or transaction, a compaction
publishing without a turn, an index becoming ready under a stale writer — so those paths stay
covered without a public optimistic mode.
write-scope-admission.test.ts loads 45 tables at concurrency 1, 6, and 12 through direct
engines and worker clients over memory, IndexedDB, and OPFS. With commit retries disabled, every
scope must run once and retain every row, and a shared counter read and rewritten by every scope
must see every value exactly once. Mixed batch and SQL writes, rollback, shutdown, a scope
cancelled through its signal, and the stall report for a callback that awaits another write on
the same database have separate checks. The native-browser companion loads 39,015 rows across
45 tables with 12 writers and automatic maintenance enabled, then reopens the store and verifies
counts and sums; a second case opens three tabs over one database, bursts scopes, batch writes,
SQL statements, and BEGIN transactions from all three while one of them adds a column and an
index, requires zero conflicts with zero retries, and then holds a scope open in one tab to prove
that another tab's write waits rather than fails and that closing the waiting tab lets go
without waiting for the holder. storage-probe.spec.ts drives the queue against real Web Lock
holders: a progressing queue is waited out without a report, a holder that never lets go is
reported once and never bypassed, and a waiter that gives up leaves at once. The interaction
simulator asserts zero lost commit races in Node and, for complete campaigns, across real tabs.
The fault sweep
packages/core/src/testing/fault-sweep.ts runs a fixed workload once without faults, then
interrupts every reached block-read, block-write, and transaction-commit boundary in turn.
fault-sweep.test.ts runs it in Node; fault-sweep.spec.ts runs the same campaign in dedicated
workers over native IndexedDB and OPFS in Chromium, Firefox, and WebKit.
Recovery must match an exact complete table state. Every acknowledged statement must survive; only the interrupted statement may be wholly present or wholly absent. Partial batches, lost acknowledged rows, duplicate keys, changed values, and an unreadable database all fail. A failed read cannot change the allowed durable state. The uninterrupted baseline must also survive a cold adapter reopen with every expected row.
The campaign records each injected error and rejects unrelated engine errors instead of swallowing them. Every fault point must be reached, and every scheduled injection must fire. Automatic maintenance is disabled here so the operation counts are deterministic; separate maintenance and crash suites test background work and abrupt worker termination. Each trial closes the engine and adapter, reopens the database, and removes its browser storage afterwards.
Benchmark workload model
The benchmarks page keeps four workload classes apart. They are never combined into one score:
| Workload | Read measurement | Write measurement |
|---|---|---|
| OLTP | Selective key, small-set, and bounded-range latency | Point and small-batch inserts, updates, and upserts (1–100 rows) |
| OLAP | Scan, join, window, and aggregate execution | Bulk ingestion and mutation throughput (10,000–100,000 rows) |
A read workload keeps every query visible rather than collapsing to one figure: each cell is the median of the timed windows for that query, after one untimed warm-up. Writes keep every operation and batch size visible the same way. Nothing is reported unless it was verified — every read must match an independent JavaScript oracle and every write is read back and compared row for row, and a query an engine got wrong or could not run prints a dash instead of a timing. Statement deallocation and session cleanup are awaited outside the timing windows. An unexpected cleanup failure invalidates the measurement or suite; a paired execution and cleanup failure retains both original errors.
Two further measurements cover what an application pays above the engine, so a regression in the worker client or the live-query layer shows on the page instead of hiding behind engine numbers:
- The cached columns (
Minnow (cached),Minnow (OPFS, cached)) repeat each statement with the default result memo enabled. They measure repeat-query caching separately from execution. The read table times engine calls directly inside the benchmark worker; it does not show separate client RPC columns. - The live-query suite registers 1 to 100 subscriptions through that client, commits one row,
and times the commit until the last affected subscription's
onChangehas fired with its rows in hand. One case keeps 99 of 100 subscriptions on a table the commit never touches, which measures the cost of ruling them out. Every affected subscription must fire exactly once per commit and end on the row count the commits imply, and no other may fire at all. The suite owns its tables and never writes the dataset's, so the read oracles stay true after it. Only Minnow takes part in this harness. PGlite has a live-query extension, but this harness has no driver for it; an unmeasured driver is reported as unsupported here, not as a missing engine capability.
There is nothing to publish or regenerate. The benchmarks page has no checked-in
numbers: it builds every engine and runs the suites in the visitor's browser, on the engines,
suites, and dataset size they pick, and explains the measurement methodology, storage differences,
caching, and memory caveats as it reports them. The suites themselves live in apps/site/bench
and the page that drives them in apps/site/app/benchmarks.
To iterate on the harness, run the unit tests (npm test) for the generator, answer checkers, and
suite rules, then npm run test:browser:site to run a suite in real browsers. To try a change by
hand, npm run site:dev and open /benchmarks.
Performance after repeated writes
Most gate shapes read a table as it was loaded. The settled-* shapes read it as it is after a
day of use: the same rows with a thousand point updates applied and the background maintenance
allowed to finish, timed against SQLite and PGlite with the same updates. Their costs do not move
with history, so the ratio is the engine's own cost of having been written to — a read that
slows down as a table is used, or maintenance that stops keeping up, moves one of these numbers.
settle-after-updates times the updates and the settle itself, once, so a fold or a collection
pass that grows expensive shows up too. Its waiter checks active jobs first, reads segment
visibility only between jobs, and measures the full storage footprint only after the table is no
longer due for a fold. This keeps the observer from competing with the maintenance it times; the
sample ends only after ten consecutive 50ms observations see no job and the same footprint.
The report separates the time to the first confirmed-stable observation from the following
confirmation delay. Observation resolution is 50ms. The gate and comparison row keep the full
elapsed time, including confirmation; SQLite and PGlite time their updates without this window.
Update the performance-gate baseline
After an intentional executor change, inspect the performance-gate output before updating its thresholds:
npm run benchmark:gate -- --updateTreat that update as a code change: explain it in review and run the gate again without --update.
Refresh the download-size comparison
The benchmarks page opens with what a browser downloads to run each engine. It is measured from the installed packages rather than quoted, so refresh it whenever a dependency version or the public entry point changes:
npx tsc -b packages/core --force
npm run benchmark:sizesEach measured engine's browser entry is bundled and minified with identical esbuild settings, and
the WebAssembly and data files it fetches at run time are added at their shipped size. The script
measures Minnow, SQLite Wasm, and PGlite. The result lands in
apps/site/components/bench/library-sizes.json, which is a scratch output — nothing imports it.
The figure each measured engine shows on the page is the download field in
apps/site/components/bench/config.ts, updated by hand from that run.
The unit suite also compares fresh bundles against the installed SQLite Wasm loader plus its Wasm file. The engine with OPFS must stay smaller after gzip. The generic worker with all stores plus its separately bundled client has a tight 465 KiB combined budget; those two downloads repeat some code and currently slightly exceed SQLite's single loader-plus-Wasm measurement. The report and test share the same measurement code. Per-entry budgets separately bound the main engine and IndexedDB worker. Exact JSON parsing and bounded regex matching add code, and those intentional additions are recorded alongside the budgets. This guards compressed download size; it does not claim a smaller uncompressed bundle or prove that every possible behavior is regression-free.
OPFS browser cases check for the storage API inside a dedicated worker before running. Some Linux WebKit builds expose it only in the page; those cases are skipped for that missing capability in the ordinary browser runner, while IndexedDB cases still run. The conformance profile requires worker OPFS and fails if it is absent. Worker startup errors and probe timeouts remain failures. Contention correctness tests allow queued durable writes a 30-second RPC deadline; the separate timeout and sustained-latency tests retain their own limits.
Browser interaction campaigns retain every worker diagnostic in their attached evidence. A typed ownership or revision refusal from an unpublished index build or compaction job can occur when another tab wins the same maintenance work. Those specific maintenance events are classified by their operation, error type, and identity/revision fields. Generic errors, I/O failures, corruption, changed posting chunks, and window or unhandled errors still fail the campaign; message text never suppresses a failure.