Reference

Testing & benchmarks

Test commands, SQL checks, browser coverage, and benchmark workloads.

One set of focused runners covers fast correctness checks, real-browser behavior, performance regressions, and the benchmark harness behind the live benchmarks page. This page is the source of truth for running and maintaining them.

Treat a failure as a defect to explain. Expected answers come from SQL semantics, independent engines, or an independently authored model; never record Minnow's failing output as its new expected answer. An oracle correction needs evidence from that independent contract. Preserve failing inputs and reduce them into regressions. Do not loosen assertions, deadlines, coverage floors, or resource budgets to hide a failure, and do not treat a successful retry as a clean run. Storage and browser-runtime regressions must run against real browser APIs and durable stores; Node models and injected faults are useful additional checks, not substitutes for that gate.

The typecheck, build, and browser prerequisites force a fresh TypeScript project-reference build. They do not trust an existing .tsbuildinfo, so a release gate proves the current sources even after generated declarations or asynchronous signatures change.

Choose the smallest runner that proves the change

CommandWhat it runsUse it for
npm testVitest plus the committed fast SQLLogicTest profileNormal implementation work
npm run test:unitVitest unit, generated, simulator, comparison, and harness testsA focused inner loop
npm run test:sql:standard6,345 result-checked upstream queries; committed, offline, about fifteen secondsEvery SQL or optimizer change
npm run test:sql:fullThe full supported profile, including exhaustive multi-way join permutationsBefore a release
npm run test:coverageVitest with the coverage floors enforcedBefore pushing
npm run soakThe generative suites on fresh random seeds, to find new failuresHunting bugs rather than pinning
npm run simulateOne replayable deterministic concurrency, fault, and storage simulationStorage and transaction changes
npm run test:simulator:full4,000 concurrent mutations plus deterministic fault and maintenance checkpoints, then three long interaction plans over the memory, IndexedDB, and OPFS storesBefore a release
npm run simulate:interactionsOne replayable interaction plan: DDL, DML, transactions, concurrent rounds, faults, reopens, and maintenance against a shadow modelEngine, index, and transaction changes
npm run fixture:formatFreezes a database this build wrote, as a compatibility fixtureBefore changing a format version
npm run test:browser:libraryCore IndexedDB, OPFS, transaction, SQLLogicTest, and multi-tab interaction-plan tests in Chromium, Firefox, and WebKitStorage, engine, or browser-runtime changes
npm run test:browser:conformanceFull core browser suite, full SQLLogicTest, and 800-step interactions over native IndexedDB and OPFS in all three browsersRelease and nightly storage correctness
npm run test:browser:sitePublic-site examples in Chromium, Firefox, and WebKitDocs, examples, and site changes
npm run test:browserBoth browser runners, in sequenceCross-browser regression checks
npm run test:consumerPacked libraries installed, typed, bundled, and run outside the workspacePackage and public-entry changes
npm run benchmark:gateSeeded Node performance ratios against the checked-in baselineQuery-executor performance changes
npm run benchmark:sizesDownload size of Minnow, SQLite Wasm, and PGlite, from the installed packagesDependency or public-entry changes
npm run api:checkThe public API surface against the committed contract snapshotPublic-entry or declaration changes
npm run checkFormat, types, lint, build, API contract, coverage, and fast SQLLogicTestThe local merge gate
npm run check:releaseThe local gate, full SQL/simulator, performance, and browser runnersRelease candidates

Install the browser binaries once before using a browser runner:

npx playwright install chromium firefox webkit

The merge gate and fast SQLLogicTest profile run in .github/workflows/ci.yml on every push and pull request. Browser jobs run independently: Chromium and Firefox on Linux, and WebKit on macOS. Each runs the library and site suites; Chromium also runs the packed consumer. The Real browsers check requires every engine to pass. Splitting the jobs preserves their 45-minute deadlines and leaves time to save failure artifacts. A test that passes only on retry still fails the gate. .github/workflows/sql-conformance.yml runs the full SQL profile and long simulator nightly and on demand. Both it and the release workflow also run test:browser:conformance: the entire core browser suite, including the full pinned SQL corpus and longer multi-tab plans over native IndexedDB and OPFS. WebKit runs on macOS so worker OPFS is available; a missing worker OPFS API fails this profile instead of skipping its tests. Publishing requires the exact commit's successful CI, SQL, simulator, and browser conformance gates. Retry a failed release through GitHub's rerun control; there is no manual trigger that can bypass CI. .github/workflows/performance.yml runs the performance gate nightly. The performance gate is deliberately kept off the merge path — it measures a machine as much as it measures the code, and a merge gate that fails on a noisy runner is one people learn to re-run rather than read. When a completed run crosses a ratio threshold, the gate starts a fresh process and fails only if that same workload and comparison engine cross it again. Setup errors, missing profiles, and incomplete runs still fail immediately.

Each browser runner owns its test directory, server, port, and build prerequisites. Shared browser defaults live in playwright.shared.mjs; runner-specific configuration stays in its named playwright.*.config.ts file. This keeps every runner independently callable without hiding which application it starts. The library server disables file watching and hot reload, so editing a shared checkout cannot navigate a page in the middle of a conformance test. Finish edits before starting a release run so its workers and pages all load the same build. The tarball-shape test packs an isolated copy: its prepack hook cannot rewrite the compiler output while concurrent bundle and declaration tests read it.

Browser runs save an HTML report in playwright-report/, including attached seed plans and result summaries. CI, release, and nightly jobs retain reports for passing and failing browser runs for 14 days; failed tests also retain traces. Each engine's library report is saved before the site runner can replace it. Successful attempts delete their persistent browser profiles; failed attempts retain and upload the profile with the trace, including a failed first attempt whose retry passes. Because macOS WebKit otherwise stores OPFS outside Playwright's profile directory, its storage fixture redirects that browser-owned data into the per-test profile. The failure fixture also copies every file from the isolated test origin through the browser File System API. Its manifest records each original path, size, captured byte count, hash, and capture error, so incomplete evidence is explicit. Evidence collection has a separate 20-second, 20,000-file, and 512 MiB cap and marks the manifest incomplete if any limit is reached. Library traces retain action timing, errors, sources, and attachments; the static runner pages omit continuous screenshots and DOM snapshots. A console log alone is not a replay artifact.

Focused native recovery tests force real checkpoint and WAL boundaries, terminate the worker, and compare the reopened manifest list with the list acknowledged before termination. IndexedDB cleanup tests check integrity between bounded pages and after reopening, including a legacy cleanup marker, concurrent pruning, and deliberately unrelated gaps. Worker lifecycle tests check slow disposal, its absolute timeout cap, and cleanup after an already-reported fatal error. These reductions run in all three engines alongside the longer seeded plans.

The native SQL comparison runner initializes PGlite, SQLite Wasm, and Minnow sequentially to reduce overlapping startup allocations. Its trace records each engine's start and readiness, so a page-process failure can be distinguished from a SQL comparison failure. Every oracle, query, and result check still runs, and a test that passes only on retry still fails CI.

Site tests distinguish editor text from executed output. TypeScript errors must appear in the compiler diagnostics alert, and the guarded-upsert example must log exactly zero updated rows and one skipped row in the results panel. A matching word in the source editor cannot satisfy either check.

Unit files are capped at four concurrent workers (or the available CPU count, if lower). Each file can start complete databases and comparison engines, so unrestricted file concurrency can produce heap pressure and scheduling timeouts unrelated to the engine. The maintenance stress file runs alone after the other unit files, so the growing simulator seed corpus cannot consume its CPU time. Its own writer and background-maintenance concurrency remains intact, and coverage from both execution groups is combined. Test workloads, deadlines, assertions, and coverage floors remain unchanged. Browser runners likewise cap independent scenarios at two workers (one on a two-CPU runner); each concurrency scenario still creates its full set of tabs.

What each layer proves

  • Unit and language tests stay beside the source they exercise. They cover predictable behavior, supported SQL and mutations, comparisons with other engines, fault injection, and the benchmark generator and answer-checking rules. Reads and writes are both checked differentially: the read conformance suite diffs a generated query corpus against SQLite and PGlite, and the mutation conformance suite replays a seeded script of inserts, updates, deletes, upserts, and atomic write scopes against the same two oracles — comparing statement outcomes, affected-row counts, RETURNING rows, trigger-written audit rows, and the full table state after every step. The fixture's triggers are declared in each oracle's own dialect (plpgsql trigger functions for PGlite) with identical bodies, and MERGE — which SQLite lacks — runs verbatim on PGlite and as an equivalent statement pair on SQLite. npm test picks up apps/site/bench alongside packages/**, so the dataset generator, the oracles, and the suite definitions are checked on every unit run.
  • Library browser tests exercise IndexedDB, OPFS, and transaction behavior that a Node substitute cannot prove, and drive a database through a real module worker: the package entry booting, transferred buffers arriving intact, and a second worker reopening what the first one wrote. Storage fixtures use persistent browser profiles in every engine so durable adapters reach their disk-backed browser implementations. They also run the SQL engine itself where it ships: the committed SQLLogicTest profile executes through the published worker against real IndexedDB and OPFS in all three browsers, and a seeded interaction plan runs with each plan connection being a real tab, so the same properties the Node simulator checks are judged against real workers, real storage, and real crashes.
  • The packed-consumer test builds every public package, installs the resulting tarballs in an OS-temporary Vite application outside the npm workspace, and resolves every exported entry with both Bundler and NodeNext TypeScript resolution. Chromium then runs the emitted worker against IndexedDB, migrates and queries through Kysely, observes a live update, streams exports, mounts the React and devtools adapters, exports a snapshot, restores it into OPFS, and reopens it. The Minnow archives and direct tool and peer versions are pinned to the build; npm resolves their ordinary transitive dependencies into a private temporary cache as it would for a new project.
  • The public API contract assigns every package subpath a stable audience and snapshots the declaration syntax exported by every typed entry point, including whether each name is type-only or available at runtime. npm run api:check fails on an unclassified subpath or an added, removed, or changed public declaration; npm run api:report prints the surface for review. After classifying compatibility and documenting the release impact, npm run api:update records an intentional surface. Coverage has a separate non-regression floor for every published package, so the core suite cannot hide a drop in devtools or an adapter. The v1 support policy defines the audiences and guarantee.
  • Site browser tests execute the public examples and drive the benchmarks page itself: one case runs a suite end to end in each browser and asserts it verified against the oracles, another asserts the /benchmarks route is cross-origin isolated. Documentation cannot drift into non-running sample code, and the harness cannot drift from the page that runs it. Comparison suites require one result per declared engine/workload and report expected, attempted, supported, verified and failed counts. A missing live driver on an engine that declares subscriptions fails; the SQLite and PGlite comparison drivers' absent subscription support is declared explicitly rather than inferred from a missing driver method. Timings show medians. The exported p95Ms is an empirical percentile of five or seven timed windows, which is their maximum, rather than an estimate of production tail latency.
  • Semantic path equivalence generates NULLs, signed zero, exact decimals, invalid text, casts and empty/nonempty membership. It compares optimized and unoptimized row execution, vector execution, literal and parameter input, memoized results and worker RPC. Values and result domains agree exactly; refusals retain their error class. PostgreSQL independently checks acceptance and values without converting NUMERIC text through Float64. The same corpus runs over native IndexedDB and OPFS workers. Interaction runs also check that the bounded background history agrees with observer evidence.
  • The performance gate detects regressions on stable seeded query shapes, reads and writes alike. Read-only engines calibrate their repeat counts independently so every fast path gets a measurable sample. A regression is reported only when the full ratio range observed across both engines' samples crosses the checked-in threshold, then repeats in a fresh process. Every stateful mutation window executes the same fixed 12-operation sequence on all three engines and finishes with a cross-engine state check. The suite includes exact-NUMERIC arithmetic, composite-key lookup, and the Kysely compilation/driver path as well as direct SQL. It is a guardrail, not a published cross-engine benchmark: the comparison engines run without indexes, which is fine for noticing that Minnow got slower and useless as a fair cross-engine claim.

Plain COUNT(*) reads use the snapshot's surviving row count without preparing column schemas or index scans. The regression suite checks this path after deletes, updates, upserts, and compaction, including old snapshots, concurrent connections, memory budgets, and cancellation.

Native-browser SQL comparisons

The library runner executes every SQL feature-matrix entry through Minnow's built module worker over both native IndexedDB and OPFS in Chromium, Firefox, and WebKit. SQLite Wasm and PGlite run in those same real browsers as independent comparison engines. Background maintenance stays enabled, and unexpected worker or page errors fail the campaign even when query answers match.

The 370 matrix entries are accounted for explicitly: 306 supported entries execute, and 64 unsupported entries must reject with their recorded error while leaving the fixture unchanged. Portable reads compare columns and values with PostgreSQL. Compatible mutations compare affected counts, RETURNING, and their actual target table's state, with domain checks for volatile values. The report separately counts cross-engine comparisons and features with fixed behavior probes. Every supported entry must have one of those semantic checks; acceptance alone cannot pass the matrix. The category totals must cover the entire matrix.

An additional 137 independently authored behavior steps cover 74 matrix features, with two more PostgreSQL state comparisons. They check TRUNCATE and UPSERT post-images, stored defaults and runtime values, trigger effects and removal, constraint enforcement, generated columns, views, sequences, identities, ALTER backfill, and transactions. Minnow's documented minimum-value ANY_VALUE result has a fixed assertion even though PostgreSQL may choose another value. Runtime defaults must be a valid timestamp, a finite random value in range, and a version-4 UUID. Negative probes require the specific error type and message. The remaining DDL forms have observable schema checks, and all 28 nonportable reads have fixed answers or documented value-domain assertions. A new entry in a category without cross-engine result comparison must provide a behavior probe, or the suite fails. Single-argument AGE must agree with its documented two-argument form using the same statement's CURRENT_DATE, including NULL propagation. BM25 checks matching rows and finite positive scores; it does not invent an exact public scoring formula. Index probes check catalog behavior and query results, while the separate performance suites cover execution cost.

Each seed also runs fixed edge cases, 42 generated queries, and four mutations with immediate post-state checks against SQLite and PostgreSQL: 117 independent comparisons per storage adapter and browser. Cases include NULL logic, joins, aggregates, windows, correlated subqueries, Unicode, JSON, numeric boundaries, and a literal question mark alongside a bound parameter. Numeric values compare exactly unless a particular aggregate or cross-runtime math expression declares a floating-point precision. JSON and numeric normalization follows result-domain metadata; ordinary text is never parsed as JSON.

The campaign replays the committed sql-conformance seeds and accepts MINNOW_SEED for new or reproduced inputs. Its report records the seed and each coverage category.

SQLLogicTest: broad, strict SQL correctness

Minnow ships a native runner for SQLite's database-neutral SQLLogicTest format. The parser and answer checker live in @minnowdb/core/testing; the command-line runner streams one record at a time and creates one empty database per file. A large corpus therefore does not become an equally large in-memory syntax tree.

The runner implements the original format's statement ok/error, typed query, nosort, rowsort, valuesort, result labels, MD5 result hashes, hash-threshold, skipif, onlyif (with the trailing # comment the upstream corpus writes after an engine name), and halt directives. Malformed records, wrong value counts, wrong types, duplicate labels with different answers, and unknown directives are failures with exact source lines. Conditions are reported as skips; they are never silently discarded.

The ordinary local profile is committed and needs no network:

npm run test:sql:standard

# Iterate on one source file or its first 200 records.
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/standard-select3.test
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/standard-select4.test --stop-after 200

# Show broad-suite progress, or trace every record from an isolated boundary.
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/full-select5.test --progress-every 100
npm run test:sql:logic -- --file packages/core/testdata/sqllogictest/full-select5.test --trace --trace-from 800

It runs 1,822 statements and 6,345 queries, checking 385,389 scalar results across 8,169 source records. The same committed files also run inside Chromium, Firefox, and WebKit (sqllogictest.spec.ts), through the published module worker against real IndexedDB and OPFS with automatic maintenance left on, so a browser engine's Date, regular-expression, or number-formatting difference, a result that does not survive the worker boundary, or a durable store's commit path under thousands of statements is checked against the corpus's recorded answers rather than against Node's. Each run attaches its per-file timings to the report. Every thirty seconds, native runs also log and attach their phase, completed statement and query counts, accumulated timings, and a bounded prefix of the active SQL. This preserves progress when the unchanged ten-minute per-file deadline expires. The fast set includes representative five-table and seventeen-table join-order families that guard against pathological Cartesian work. On the development machine used to measure it, the whole profile took about fifteen seconds. That makes the broad suite part of npm test, npm run check, and every pull request, instead of a suite agents discover only after they finish a change.

The supported full profile adds the exhaustive multi-way join families:

npm run test:sql:full

It is fully runnable locally, reports progress every 500 records, and contains all 10,708 records from the five selected upstream files: 1,822 statements, 8,884 queries, and 5,093,664 checked scalar results. There are no all-profile exclusions. standard-exclusions.json still records the 2,539 supported exhaustive-join records that move from the fast profile to the full one solely because of runtime.

The corpus is pinned to immutable revision 572273bf5076176e17f47aea75115a225e70fef8. upstream.json records its URL, date, 115,244,234-byte archive size, archive SHA-256, five source hashes, and expected 622 test files. Regeneration first verifies all of those inputs:

npm run test:sql:fetch
npm run test:sql:prepare-standard

The raw 622-file corpus remains locally runnable for boundary discovery. It is expected to expose unsupported SQL; use --keep-going when auditing rather than treating it as a release profile:

npm run test:sql:upstream -- --stop-after 500
npm run test:sql:upstream -- --keep-going --max-failures 25

Adding support means removing its exact exclusion, regenerating the committed files, and adding a focused regression test. Adding an exclusion without a concrete reason makes the ledger diff visible in review; there is no wildcard skip list.

When adding a runner, give it one responsibility and make it independently runnable. When adding a suite case, derive smoke-test counts from the suite itself instead of copying totals into tests or HTML.

Seeds, and how a soak failure becomes a permanent test

The generated suites — SQL support, write support, the deterministic simulator, the columnar-versus-row comparison, and the compaction and multi-tab soaks — build their corpora from a seeded generator. A committed run is deterministic: it uses the checked-in seeds plus every seed that has ever failed, listed in packages/core/regression-seeds.json. Each recorded seed runs as its own test case next to the suite's default, so a seed in that file is executed on every commit, not merely listed. That makes the suite a reliable regression net, and on its own it would be nothing else, because the questions it asks never change.

npm run soak is the other half. It runs the same suites on seeds nobody has tried and stops at the first failure, printing the seed:

npm run soak -- --rounds 200

A failing seed is the whole artifact. Replay it directly:

MINNOW_SEED=1476318588 npx vitest run packages/core/src/engine/sql-conformance.test.ts

Then add it to regression-seeds.json under the suite that failed, where every future run picks it up. Never remove one: a seed in that file is a bug that used to exist, and the entry is what stops it coming back. The explored space only grows, and it grows by exactly what the soak found.

Format compatibility

A browser database's data lives in the user's browser, so it outlives every deploy. There is no migration window and no way to reach back and rewrite it: if a format change makes yesterday's blocks unreadable, the first anyone hears is a user whose application will not open.

packages/core/format-fixtures/ holds one frozen database per released block/snapshot pair — a framed snapshot, which carries the raw block bytes verbatim, plus the answers that build gave to a fixed set of queries. Each manifest records the exact @minnowdb/core version that wrote it. format-compatibility.test.ts rejects a fixture that claims a future writer, reproduces the current writer byte for byte, opens every fixture, checks the answers still hold, and checks that writes into a restored database still work. Frozen raw vectors independently cover every block-format-2 physical type.

The directory also holds one native fixture for every locked OPFS layout. The test requires a fixture for the current writer and checks its recorded package version, then reopens each fixture's format marker, checkpoint, WAL tail, acknowledgement file, and packed extent through the current adapter before committing and recovering another write. Every supported older fixture must upgrade automatically through ordinary open; an older-layout refusal cannot pass this test. The ordered OPFS upgrade registry must cover every version between the first locked layout and the current one. upgrades.test.ts injects failure at every preparation/publication I/O boundary, models both tab death and power loss, exercises partial transfers and concurrent openers, and keeps concurrent openers waiting through a slow live conversion. A resurrected ready record cannot roll back later acknowledged writes, even with a missing/torn marker and missing/corrupt completion receipt. Damaged native proof after resetting a nonempty old WAL is refused rather than repaired from stale staging. The native browser worker suite discovers the same frozen fixtures and checks automatic upgrade and new writes across restart. A pinned test-only alias of the released core 0.10.0 package proves that the actual layout-6 reader opens the old fixture and refuses the upgraded tree before mutation, in Node and native workers. A second alias, of core 0.12.1, is the last IndexedDB schema-2 and OPFS layout-7 writer: marker-upgrade.test.ts and indexeddb-schema3.test.ts have it leave a merge-v1 fold in flight, then upgrade, finish that fold, write, reopen, and check that 0.12.1 refuses the result without a byte changing. A third alias, of core 0.13.1, is the last IndexedDB schema-3 and OPFS layout-8 writer: marker-upgrade.test.ts upgrades its database to layout 9, logs a commit whose frame spans continuation frames and whose index changes exceed one layout-8 chunk, reopens from the log and from a checkpoint, and checks that 0.13.1 then refuses the database; indexeddb-schema4.test.ts does the same for schema 4 with every older writer. The marker upgrade is cut at each of its eight I/O boundaries by tab death and by power loss, from both older layouts. Unsupported pre-contract prototypes are refused without changing their bytes. The layout-7 tests also zero or truncate acknowledged WAL bytes and damage either acknowledgement slot; recovery must report corruption rather than roll back.

It also fails when the current BLOCK_FORMAT_VERSION or SNAPSHOT_FORMAT_VERSION has no fixture behind it. That failure is the useful one, because it fires before the damage:

npm run fixture:format

The native OPFS fixture has its own writer:

npm run fixture:opfs-format

Run the relevant writer on the build that still writes the old format, commit the fixture, and only then change the version. A fixture can only be produced by the build that writes it — once that code is gone, the format it wrote cannot be regenerated. For the same reason, never delete a released, locked fixture. Block format 2 is the first such lock; the unsupported experimental block-1 fixture was replaced before v1.

Internal orchestration and adapter transitions

Catalog mutations, per-connection writer admission, and query preparation/execution have separate internal owners. The public database API delegates to them; SQL kernels and adapter atomic publication remain shared with the ordinary paths. Refactoring these owners must preserve cancellation, reader leases, compaction turn lending, shutdown draining, and original failure reporting. The lifecycle, admission, semantic equivalence, and native interaction suites exercise those boundaries. Caller-provided clock and ID callbacks retain their original database receiver through the extracted catalog and query paths, with a lifecycle regression on all three adapters.

storage/commit-admission.test.ts runs the same multi-table commit refusals on memory, IndexedDB, and OPFS. Invalid coverage, duplicate limits, invalid limits, and segment backpressure must leave the manifest, transaction journal, and staged bytes unchanged. A corrected retry on that same transaction must publish every changed table.

Record-core hardening tests verify bulk unique-key removal without per-key ordered-index work, canonical snapshot export and restore after mixed mutations, and exact batched block lengths and checksums after refusal, retry, and placement-only adapter replay. Snapshot ordering is lazy; constraint membership stays immediate. IndexedDB hardening also verifies that independent commit reads are queued together and a refused commit leaves the journal and staged bytes unchanged for a corrected retry. The controller extraction and storage changes add about 4 KiB to the raw OPFS bundle allowance while keeping the compressed download budgets unchanged.

Matched-index performance comparisons

npm run benchmark:compare runs the Node suite with primary-key indexes on every engine's ordinary, mutation, and settled tables. It verifies query and final-state results before timing. Use this command and the live browser comparison for competitive speed claims.

The existing benchmark:gate preserves its historical schema and runtime-specific thresholds, including unindexed ordinary-key comparator tables. Its purpose is detecting changes against that same workload. Its ordinary key ratios cannot support competitive speed claims. A comparison run cannot update those thresholds; changing that gate's schema requires fresh calibration on every recorded runtime, including Linux.

Running out of quota

quota.test.ts covers what happens when the browser refuses a write because the origin is out of space — the characteristic way a browser database fails, and the one an application most needs to handle deliberately. It pins four things: the write fails rather than half-landing, the error keeps its QuotaExceededError identity so an application can branch on it, everything committed earlier stays readable, and the same write succeeds once there is room, with no repair step.

Property-based tests

block-format/properties.test.ts describes the format as properties and checks them against generated inputs rather than chosen ones: a value written and read back is the same value, both codecs agree, a zone map contains everything it summarizes, and a flipped byte is detected rather than decoded into plausible rows. The generators reach for what breaks encoders — negative zero, the extremes of the double range, subnormals, empty and astral-plane strings, all-null and empty columns. A failure shrinks to a minimal case and prints the seed and shrinking path. The same property campaign runs in dedicated workers in Chromium, Firefox, and WebKit, using each engine's native CompressionStream, DecompressionStream, UTF-8 conversion, dates, and typed arrays. Every browser replays the same committed seeds; set MINNOW_SEED to reproduce a reported failure. The block-format-properties entry in regression-seeds.json and the soak runner keep discovered failures in the permanent regression set.

Soaks and deterministic simulation

Four suites cover accumulation rather than a single operation, which is where a fold that loses a row, a scheduler interleaving, or a compactor that stops folding actually shows up:

  • compaction-soak.test.ts runs 1,500 interleaved mutations against a reference Map, compacting at checkpoints, and compares the whole table to the reference each time. Compaction is bounded and incremental — one call folds a limited number of blocks — so the test expects steady progress, not a finished table after one call.
  • concurrency-simulation.test.ts runs twelve independent databases over one shared store, the shape of twelve browser tabs on one origin, issuing a seeded random schedule. It checks that the outcome is explicable: every visible row was written by some tab, acknowledged writes are present, no key appears twice, the write the store acknowledged last is the one visible, and every tab sees the same database afterwards.
  • auto-compaction-soak.test.ts is the same kind of history with the background maintenance left on, as it ships: 1,800 seeded mutations and reads while folds and collection passes land underneath them, the table checked against the reference at every checkpoint, then settled and held to convergence bounds — a scan reads the bounded folded partitions and only a level-zero tail below both background fold thresholds, and the store holds the data rather than the history.
  • simulator.test.ts controls the completion order of every asynchronous BlockStore call across independent database clients. A reference Map checks commuting concurrent mutations at each checkpoint; scheduled crash points must reopen to exactly the state before or after the write; compaction and garbage collection run throughout. The final quiescent store is bounded by the key space rather than the number of rounds across blocks, segment records, transaction records, and manifests, and must have no leaked leases. The diagnostic trace is a fixed-size tail, so the simulator itself does not grow with the run.

Run or save an exact plan directly:

npm run simulate -- --seed 24301 --rounds 500 --clients 8 --key-space 128
npm run simulate -- --seed 24301 --write-plan /tmp/minnow-plan.json
npm run simulate -- --plan /tmp/minnow-plan.json

The seed fixes both workload generation and storage completion choices. A failure prints its replay command. npm run soak -- --suite deterministic-simulator explores fresh seeds, and a failure joins packages/core/regression-seeds.json so every future unit run replays it.

Interaction plans

The storage simulator above drives one table with keyed upserts and deletes. The interaction-plan simulator (interaction-simulator.ts, exported from @minnowdb/core/testing) drives the SQL surface an application uses, in the spirit of Turso's simulator: a seeded plan of interactions over several connections that share one database — CREATE TABLE with random column types and nullability, multi-row inserts by literal and by parameter, predicate UPDATE and DELETE, selects with three-valued predicates, LIMIT and ORDER BY, secondary indexes, DROP TABLE, SQL transactions that commit or roll back, rounds of concurrent keyed writes with readers in flight, storage faults and worker crashes, reopens, compaction and collection, and checkpoints. The plan is JSON, so a failing seed replays exactly and can be checked in.

Generated plans preserve the randomized sequence and add a fresh-row insert for each configured fault point before the final checkpoint. These keys lie outside the random key range, so each probe can change the database and reach its write hook. An update or delete of a missing row can otherwise make a random fault step do no work. The fresh probes supplement those random cases; unsupported fault points are still reported as skipped by drivers that cannot inject them.

Every interaction is executed through the real SQL entry points and judged against a shadow model that evaluates the predicates with SQL's three-valued logic. Beyond "the rows match", the runner checks properties a database must honour whatever its internals do: a rejected insert leaves none of its rows behind; reported row counts equal the model's; nothing matching a delete survives it; LIMIT returns exactly the first rows in order; COUNT(p) + COUNT(NOT p) + COUNT(p IS NULL) equals COUNT(*); a UNION ALL has the sum of its members' cardinalities; uncommitted rows are visible on their own connection and invisible to another; a read taken during a round of commuting writes shows every untouched key exactly and every touched key at its old or new value; a mutation interrupted by a fault lands entirely or not at all and a reported success is durable; and at every checkpoint every connection reads the same database as the model.

Two refusals inside a transaction are engine behaviour rather than defects, so the runner counts them and continues: a COMMIT that lost the race to another connection's data commit, and a transaction the engine rolled back after transactionIdleTimeoutMs of idleness. Both discard the whole transaction, and an expired one is acknowledged with a ROLLBACK before the plan goes on. See transactions for the rules themselves.

The runner also validates result shape: projected columns must match in order, each row must have exactly its declared fields, and a missing field cannot stand in for NULL. Its expected-error paths check the engine's specific error identity or exact documented refusal. An unrelated error does not become acceptable just because a statement was meant to fail or a crash was injected. Custom drivers must preserve error names and causes across their transport, including DatabaseWorkerOutcomeUnknownError for a mutation interrupted by a worker crash.

The runner is driver-agnostic. In Vitest (interaction-simulator.test.ts) each connection is a MinnowDatabase over one shared memory, IndexedDB, or OPFS store with the storage fault points armed. In Playwright (interaction.spec.ts) each connection is a real tab holding the published worker over a real IndexedDB or OPFS database, and the plan's faults terminate a worker with a statement in flight — the tab then reports the loss to its client at once, so the in-flight call fails as an unknown outcome instead of waiting out the client's silence deadline. The same seed produces the same plan in both.

The browser driver recovers a crash the way an application must: by reloading the document. A new worker in the same document is enough on Chromium and Firefox, but WebKit keeps the dead worker's IndexedDB connection registered until its document goes away and blocks every connection to that database meanwhile — so a reload is the remedy the engine's StorageUnresponsiveError names, and the driver uses it. A deliberate WebKit/IndexedDB crash campaign that meets an unresponsive store records it in transientsAccepted and stops, with interactions counting the steps it judged. This exception is limited to that browser, store, and deliberate crash campaign. Every browser/store pair also runs a separate campaign without fault steps that must complete the entire plan; a crash exception cannot replace that coverage. Other early stops fail.

The standard browser campaigns generate 220 steps with three connections. The conformance profile generates 800 with four connections. Each run attaches its exact generated plan before executing it. The nightly workflow varies a logged MINNOW_BROWSER_INTERACTION_SEED_BASE input; reproduce that frontier with:

MINNOW_SEED=123456 MINNOW_BROWSER_INTERACTION_SEED_BASE=123456 npm run test:browser:conformance

This fixes the generated plan, while actual browser scheduling remains uncontrolled. Node's IndexedDB and OPFS simulations are additional diagnostic checks, not proof of native durability.

npm run simulate:interactions -- --seed 24301 --length 300 --connections 4 --store indexeddb
npm run simulate:interactions -- --seed 24301 --write-plan /tmp/minnow-interactions.json
npm run simulate:interactions -- --plan /tmp/minnow-interactions.json --store opfs
npm run soak -- --suite interaction-simulator

A failure names the interaction, the property, the model's rows, and the SQL leading up to it. The first soaks of this simulator found seven defects the existing suites had missed — an index-pruned scan that resurrected deleted rows, a composite index that skipped rows with a NULL trailing column, a stale intermediate value served by block-level pruning, a fold that refused any transaction writing two segments to one table, a zone-pruned delete that made a reinserted key fail every keyed range read, a compaction refused by a background index build, and an IndexedDB key cache that let two tabs write one primary key twice — each now pinned in simulator-regressions.test.ts. The last one only shows with a separate adapter instance per connection, which is why the Node driver opens one per connection for IndexedDB and OPFS. Its first real-browser runs then added two more: WebKit keeping a terminated worker's IndexedDB connection and open transaction alive until its document goes away, which wedged every tab (now bounded by StorageUnresponsiveError), and a crashed tab's pending call waiting out the client's silence deadline long enough for another tab's idle transaction rollback to fire.

Background maintenance, size, and memory

auto-maintenance.test.ts pins what autoCompact and autoCollect guarantee, as bounds rather than timings: a burst of writes folds back down and, a quiet minute later, collects to a store no bigger than twice the data; repeated bursts settle to the same place instead of each leaving a residue; versions from a moment ago stay readable and old ones are pruned; the switches off do nothing; the buffer pool stays inside its budget under many distinct queries.

automatic-compaction-fit.test.ts holds automatic folds to the same standard on tables whose folds do not fit the default memory budget as planned: a wide table refreshed wholesale keeps folding inside the budget, a fold whose distinct keys outgrow it is cut, the smallest fold gets the memory it needs, and a refresh written in an unrelated order folds with a job record a few kilobytes long. A fold resumed by another engine recomputes its replay, and one whose replay no longer matches its plan is abandoned unwritten. A single delta as large as its table folds with no further write, including a lone one over already-folded partitions. Each case fails with its part of the fix removed.

large-delta-reads.test.ts refreshes 400,000- and 1,000,000-row tables with one full-table upsert under default settings, on memory, OPFS, and IndexedDB stores, and checks exact results at once — while the fold runs — and after it; it also folds such a table with a compactTable that names no budget. overlay-fallback.test.ts forces the replay's two bounded paths — patches replayed per range of rows, and touched keys replayed in hash partitions — through projections, aggregates, ordering, index-selected rows, layered partial updates, deletes, and reads of older versions, and requires the in-memory replay's exact rows in its exact order; one case reaches both paths through a real small budget. delta-scan.test.ts runs its SQLite-checked mutation script through both paths as well.

event-loop-stalls.test.ts runs a 1 ms watchdog through the workloads that once held every other query for hundreds of milliseconds or minutes — a fold of reordered wide upserts, live aggregates patched after large commits, scans over upsert-heavy tables, secondary-index builds and their first lookups, a full-text index build, and a 600,000-row write and upsert on OPFS — and fails if any stretch the watchdog cannot run exceeds 250 ms (750 ms in CI). Each case fails against the code before 0.13.0. checkpoint-slicing.test.ts pins that a multi-slice OPFS checkpoint gives the event loop turns between its slot writes, publishes exactly the bytes a one-step checkpoint would, and survives power loss or a refused write at every one of its file operations, with the refused checkpoint retried by later writes; wire.test.ts checks the sliced encoder against the one-step encoder byte for byte, and the sliced checkpoint parser against JSON.parse, each with a property test over arbitrary JSON-shaped states and, for the parser, malformed input it must refuse. large-commit.test.ts pins on the memory and OPFS stores that a commit large enough to check its keys in slices still refuses a conflicting key without changing anything, that its ahead-encoded log frame replays after a crash, and that an index change of 70,000 postings — more than one stored chunk — commits, replays from the OPFS log, and loads from a checkpoint. wal.test.ts writes payloads that span continuation frames and checks that replay joins them, treats pieces with nothing after them as an unwritten tail wherever the log stops, takes pieces back when an append fails, and refuses a corrupt piece inside the acknowledged log. follower-large-commit.test.ts commits 400,000 indexed rows from an OPFS follower — past the message's structural limit — and checks that a re-sent large commit runs once under its first delivery's identity. index-reads.test.ts pins that an OPFS index lookup answers while a large commit holds the leader's queue, and that a chunk read failing partway runs again on the queue. lease-lane.test.ts pins that OPFS reader leases run between the slices of a large commit and a checkpoint, with every power-loss and torn-write boundary of such a checkpoint recovered. unique-index-staged-build.test.ts pins, on the memory, IndexedDB, and OPFS stores, that creating a UNIQUE index stages its keys in ordered chunks of at most 4,096 and never sends them in one catalog update, enforces the index across a reopen, handles an empty table, and releases staged keys when publication fails. index-capacity.test.ts builds a secondary index over 600,000 distinct values and a full-text index over 450,000 documents, both of which the chunk limit once refused, and pins that at most one chunk per block of rows covers a term, so a lookup reads no more chunks than before. npm run benchmark:stalls measures the real numbers for those shapes and a few more, during the writes and during the maintenance that follows; -- --store opfs runs them over the native OPFS store.

The heap itself is checked by the memory soak, which the unit suites cannot do honestly:

npm run soak:memory                     # node --expose-gc; four workloads, eight rounds each
npm run soak:memory -- --rounds 12 --workload writes

Each workload — queries, writes with background maintenance, live query sets opened and closed, databases opened and closed over one store — runs in rounds with a forced collection and a heap reading after each. The assertion is a plateau: the last rounds may not sit more than a bounded margin above the early ones. The database runs with a small buffer pool so the pool fills during the warm-up rounds, which are excluded; a climb that continues past a full pool is a leak.

Concurrent writes

write-coordinator.test.ts covers the writer queue itself against a controlled lock manager: arrival order, one pending lock request per context, prompt cancellation of a queued or lock-waiting writer, one stall report per unchanging holder with no bypass, silence while holders keep changing, and separate queues for separate databases. It also refuses to invent a stalled holder from an empty or unavailable lock snapshot and proves that writers behind the local queue do not inspect remote locks. The browser corpus retains every console error's structured arguments in its report while still failing on errors.

write-contention.test.ts covers writers that do and do not take turns. Coordinated writers all land; a store marked uncoordinated through the test seam (markStoreUncoordinatedForTests, the stand-in for an older build in another tab) accepts only maxCommitRetries + 1 of a synchronized IndexedDB burst, and that budget does not grow with concurrency. Rejected writes publish nothing and acknowledged writes remain fully present. The same seam is what every test of the defensive compare-and-swap path uses — a commit landing under an open scope or transaction, a compaction publishing without a turn, an index becoming ready under a stale writer — so those paths stay covered without a public optimistic mode.

write-scope-admission.test.ts loads 45 tables at concurrency 1, 6, and 12 through direct engines and worker clients over memory, IndexedDB, and OPFS. With commit retries disabled, every scope must run once and retain every row, and a shared counter read and rewritten by every scope must see every value exactly once. Mixed batch and SQL writes, rollback, shutdown, a scope cancelled through its signal, and the stall report for a callback that awaits another write on the same database have separate checks. The native-browser companion loads 39,015 rows across 45 tables with 12 writers and automatic maintenance enabled, then reopens the store and verifies counts and sums; a second case opens three tabs over one database, bursts scopes, batch writes, SQL statements, and BEGIN transactions from all three while one of them adds a column and an index, requires zero conflicts with zero retries, and then holds a scope open in one tab to prove that another tab's write waits rather than fails and that closing the waiting tab lets go without waiting for the holder. storage-probe.spec.ts drives the queue against real Web Lock holders: a progressing queue is waited out without a report, a holder that never lets go is reported once and never bypassed, and a waiter that gives up leaves at once. The interaction simulator asserts zero lost commit races in Node and, for complete campaigns, across real tabs.

The fault sweep

packages/core/src/testing/fault-sweep.ts runs a fixed workload once without faults, then interrupts every reached block-read, block-write, and transaction-commit boundary in turn. fault-sweep.test.ts runs it in Node; fault-sweep.spec.ts runs the same campaign in dedicated workers over native IndexedDB and OPFS in Chromium, Firefox, and WebKit.

Recovery must match an exact complete table state. Every acknowledged statement must survive; only the interrupted statement may be wholly present or wholly absent. Partial batches, lost acknowledged rows, duplicate keys, changed values, and an unreadable database all fail. A failed read cannot change the allowed durable state. The uninterrupted baseline must also survive a cold adapter reopen with every expected row.

The campaign records each injected error and rejects unrelated engine errors instead of swallowing them. Every fault point must be reached, and every scheduled injection must fire. Automatic maintenance is disabled here so the operation counts are deterministic; separate maintenance and crash suites test background work and abrupt worker termination. Each trial closes the engine and adapter, reopens the database, and removes its browser storage afterwards.

Benchmark workload model

The benchmarks page keeps four workload classes apart. They are never combined into one score:

WorkloadRead measurementWrite measurement
OLTPSelective key, small-set, and bounded-range latencyPoint and small-batch inserts, updates, and upserts (1–100 rows)
OLAPScan, join, window, and aggregate executionBulk ingestion and mutation throughput (10,000–100,000 rows)

A read workload keeps every query visible rather than collapsing to one figure: each cell is the median of the timed windows for that query, after one untimed warm-up. Writes keep every operation and batch size visible the same way. Nothing is reported unless it was verified — every read must match an independent JavaScript oracle and every write is read back and compared row for row, and a query an engine got wrong or could not run prints a dash instead of a timing. Statement deallocation and session cleanup are awaited outside the timing windows. An unexpected cleanup failure invalidates the measurement or suite; a paired execution and cleanup failure retains both original errors.

Two further measurements cover what an application pays above the engine, so a regression in the worker client or the live-query layer shows on the page instead of hiding behind engine numbers:

  • The cached columns (Minnow (cached), Minnow (OPFS, cached)) repeat each statement with the default result memo enabled. They measure repeat-query caching separately from execution. The read table times engine calls directly inside the benchmark worker; it does not show separate client RPC columns.
  • The live-query suite registers 1 to 100 subscriptions through that client, commits one row, and times the commit until the last affected subscription's onChange has fired with its rows in hand. One case keeps 99 of 100 subscriptions on a table the commit never touches, which measures the cost of ruling them out. Every affected subscription must fire exactly once per commit and end on the row count the commits imply, and no other may fire at all. The suite owns its tables and never writes the dataset's, so the read oracles stay true after it. Only Minnow takes part in this harness. PGlite has a live-query extension, but this harness has no driver for it; an unmeasured driver is reported as unsupported here, not as a missing engine capability.

There is nothing to publish or regenerate. The benchmarks page has no checked-in numbers: it builds every engine and runs the suites in the visitor's browser, on the engines, suites, and dataset size they pick, and explains the measurement methodology, storage differences, caching, and memory caveats as it reports them. The suites themselves live in apps/site/bench and the page that drives them in apps/site/app/benchmarks.

To iterate on the harness, run the unit tests (npm test) for the generator, answer checkers, and suite rules, then npm run test:browser:site to run a suite in real browsers. To try a change by hand, npm run site:dev and open /benchmarks.

Performance after repeated writes

Most gate shapes read a table as it was loaded. The settled-* shapes read it as it is after a day of use: the same rows with a thousand point updates applied and the background maintenance allowed to finish, timed against SQLite and PGlite with the same updates. Their costs do not move with history, so the ratio is the engine's own cost of having been written to — a read that slows down as a table is used, or maintenance that stops keeping up, moves one of these numbers. settle-after-updates times the updates and the settle itself, once, so a fold or a collection pass that grows expensive shows up too. Its waiter checks active jobs first, reads segment visibility only between jobs, and measures the full storage footprint only after the table is no longer due for a fold. This keeps the observer from competing with the maintenance it times; the sample ends only after ten consecutive 50ms observations see no job and the same footprint. The report separates the time to the first confirmed-stable observation from the following confirmation delay. Observation resolution is 50ms. The gate and comparison row keep the full elapsed time, including confirmation; SQLite and PGlite time their updates without this window.

Update the performance-gate baseline

After an intentional executor change, inspect the performance-gate output before updating its thresholds:

npm run benchmark:gate -- --update

Treat that update as a code change: explain it in review and run the gate again without --update.

Refresh the download-size comparison

The benchmarks page opens with what a browser downloads to run each engine. It is measured from the installed packages rather than quoted, so refresh it whenever a dependency version or the public entry point changes:

npx tsc -b packages/core --force
npm run benchmark:sizes

Each measured engine's browser entry is bundled and minified with identical esbuild settings, and the WebAssembly and data files it fetches at run time are added at their shipped size. The script measures Minnow, SQLite Wasm, and PGlite. The result lands in apps/site/components/bench/library-sizes.json, which is a scratch output — nothing imports it. The figure each measured engine shows on the page is the download field in apps/site/components/bench/config.ts, updated by hand from that run.

The unit suite also compares fresh bundles against the installed SQLite Wasm loader plus its Wasm file. The engine with OPFS must stay smaller after gzip. The generic worker with all stores plus its separately bundled client has a tight 465 KiB combined budget; those two downloads repeat some code and currently slightly exceed SQLite's single loader-plus-Wasm measurement. The report and test share the same measurement code. Per-entry budgets separately bound the main engine and IndexedDB worker. Exact JSON parsing and bounded regex matching add code, and those intentional additions are recorded alongside the budgets. This guards compressed download size; it does not claim a smaller uncompressed bundle or prove that every possible behavior is regression-free.

OPFS browser cases check for the storage API inside a dedicated worker before running. Some Linux WebKit builds expose it only in the page; those cases are skipped for that missing capability in the ordinary browser runner, while IndexedDB cases still run. The conformance profile requires worker OPFS and fails if it is absent. Worker startup errors and probe timeouts remain failures. Contention correctness tests allow queued durable writes a 30-second RPC deadline; the separate timeout and sustained-latency tests retain their own limits.

Browser interaction campaigns retain every worker diagnostic in their attached evidence. A typed ownership or revision refusal from an unpublished index build or compaction job can occur when another tab wins the same maintenance work. Those specific maintenance events are classified by their operation, error type, and identity/revision fields. Generic errors, I/O failures, corruption, changed posting chunks, and window or unhandled errors still fail the campaign; message text never suppresses a failure.

On this page