Storage

OPFS

The file-based durable adapter — how leadership, the write-ahead log, and speed work.

import { OpfsBlockStore } from "@minnowdb/core/storage/opfs";

const store = await OpfsBlockStore.open({
  name: "shop",
  durability: "strict",
});
OptionDefaultEffect
name—The database directory name. Two stores with the same name are the same database.
durability"strict""relaxed" allows deferred final flushes for more write throughput.
rootthe origin'sA FileSystemDirectoryHandle to use instead, for tests.
onDiagnostic—Hears background failures and acknowledgement flush failures. Without a hook, errors go to console.error.

The store lives in the Origin Private File System and is built around one especially fast browser feature: a synchronous access handle that stays open, where reads and writes cost microseconds. Those handles need no special headers and no cross-origin isolation — only a dedicated worker, which is where the engine already runs. Pages with embedded payment or third-party components keep working unchanged.

Where it runs

Anything with OPFS synchronous access handles: current Chrome, Firefox, Safari 16.4+, and Edge. Two gaps to know about: Safari's private browsing has no OPFS at all, and the store must live in a dedicated worker ({ kind: "opfs", name } through the worker client just works; the main thread cannot hold synchronous handles). Where OPFS is unavailable, IndexedDB is the durable fallback, and the worker client can make that choice itself:

const store = { kind: "auto", name: "shop" }; // Reuses existing data; prefers OPFS for a new name

The choice is remembered per database name, so a database never silently reopens empty on the other store; see choosing the store at runtime.

One leader, held handles

At any moment, exactly one connection — the leader — holds the database's file handles: the write-ahead log, its acknowledgement file, two checkpoint slots, and the packed data files. Leadership is the log file's own exclusive handle, a lock the browser enforces against the actual resource and releases the instant its holder dies. There is no lease to time out and no election protocol to trust: becoming leader is opening the file, and failover is the next connection's open succeeding.

The leader does every operation at held-handle speed. A commit validates in memory and appends one checksummed frame to the log. In strict mode it also flushes a separate checksummed acknowledgement boundary before replying. A large commit prepares its keys and index changes a slice at a time first. When its frame is larger than one piece of a few megabytes, it writes all but the last piece as continuation frames, then applies itself and appends the last piece in one short step. A commit whose frame could never fit the log's headroom (about 960 MiB) is refused with StorageResourceLimitError ("log frame byte") before anything changes. Reads answer from memory and from synchronous reads of the data files. And because the leader is provably the only writer, it never checks anything for freshness — the per-operation probes and version checks other designs pay simply do not exist. That is why the store is fast: on the benchmarks page both Minnow columns run the identical engine and workload, and the results show the cost of each store in the browser running them.

A reader's lease — the pin that stops collection from removing the version a query reads — is logged like any write, because a reader in another tab must keep its pin when the leader crashes or hands over. A lease is a tiny log frame, though, so it does not wait for a large commit or a checkpoint to finish: the leader logs it between their slices, ahead of the commit it overtook.

Other tabs are followers: their operations travel a BroadcastChannel to the leader and are acknowledged only after the leader's log holds them. Requests and answers are addressed — an operation goes to the leader's own channel and its answer to the asking tab's — so a block read's bytes are copied into the tab that asked and no other; only leadership traffic is heard by every tab. Request ids keep delivery retries from applying a mutation twice while the same leader is serving. In-flight mutations are never evicted from deduplication; completed answers use a bounded 512-entry, 128 MiB cache that retains request fingerprints rather than request payloads. Each follower admits at most 512 pending requests or 128 MiB of request data. The leader admits at most 1,024 requests in total, with separate 128 MiB mutation and 256 MiB read ceilings; an individual message cannot exceed 66 MiB. A follower's commit whose keys and index changes are too large for that sends them ahead as record JSON in 8 MiB pieces, and the leader joins them back into the commit. One such request is admitted even past the byte ceilings when nothing else is waiting. Admitted reads wait for response capacity, reserving the maximum response size only while they execute. Small metadata bursts therefore queue behind the read memory limit. Request-count and request-byte overload is still rejected before creating more work, so a stalled tab cannot create an unbounded message queue. Request identity includes the follower, request id, method, and a canonical SHA fingerprint of the arguments — for a commit sent in pieces, of the pieces' bytes. Reusing an identity with different arguments is refused. After failover, an operation keeps looking for a leader — discovering, calling, electing — for ten seconds. That budget restarts whenever a connection holding the handles answers that it is still recovering the log or checkpointing on its way out, so a long recovery of a large database is waited out while a leader frozen with the handles is not. Exhausting the budget throws OpfsCoordinationError with reason: "leader-unavailable". Admission overload uses "leader-queue-full", "mutation-queue-full", or "follower-queue-full". The error includes the store method and is exported from @minnowdb/core and the OPFS entry point; instanceof works through both follower and database worker RPC. Established live subscriptions retain their last good result and retry these errors with backoff. Initial subscription and explicit reads can reject with the typed error; boot code can retry, including a failed migrate() catalog read.

A follower gives up on a leader only after a full second of silence, and a leader is never silent while it holds a request: it posts a keepalive every half-timeout for each admitted request, and announces a hold before each checkpoint, sized from the previous one. A connection that is asked to run a request while it is not leading — after a handover, after closing, or when leadership moved while the request waited — answers with a decline, which proves the request never ran, so the follower simply finds the leader and sends it again. A leader that hands over or closes says goodbye first, and answers or declines everything it still holds; its own operations queued behind the handover go to whoever leads next, never failed. Even when a leader does fall silent, a mutation is not given up on at once: the follower pings, and a leader that answers the ping is alive and still holds the request. When the leader really is gone, the request goes to whoever leads next, with the same identity and the time it was first sent, and the log decides. Every frame a follower's mutation appended carries that request's identity, and each checkpoint carries a ledger of recently served requests with the values they returned, so a recovering leader answers a re-sent request from the log, or runs it fresh when the log proves it never ran. Three things are left to report as OpfsUncertainOutcomeError: a request first sent before the recovered ledger's coverage begins (its entries are kept for ten minutes, 65,536 requests, or 8 MiB of returned values), a request whose frame was durable but whose leader died in the instant before it recorded the value it returned, and a request whose returned value was larger than 64 KiB — a staging call answers with the whole transaction record, and the ledger records that such a request ran without keeping what it answered. The engine's transaction layer recovers a lost staging or commit acknowledgement by reading the record back, so that last case never reaches an application through write(). The mutation may have committed in any of these cases; reconcile the operation's stable ids or revisions before deciding whether to retry. Correctness never rides on a message; the channel affects how fast multi-tab work moves, never whether it is right. Temporary spill writes are reserved in the leader's durable ledger before a follower writes the file, so all tabs share the same page, run, and byte limits.

The store also prefers to put leadership where the work is: the worker client reports page visibility, and a background leader yields to a foreground tab that asks. A leader that goes hidden while another tab stays visible announces so, and the visible tab asks. A leader yields at most once every three seconds; a bid that arrives sooner is honored when the cooldown ends, so tabs switched back and forth still settle on the visible one. After yielding, the old leader stays out of elections for one and a half seconds so the bidder can win the handles, then takes them back only if nobody did. A hidden leader with other connections around and nothing to do for fifteen seconds lets go of its handles on its own, because a browser may freeze a hidden tab, worker and all, and a frozen leader can neither serve nor yield; its own next operation simply elects again. The tab the user is looking at is normally the tab holding the fast path.

The pieces used by the leader — the record engine, write-ahead log, packed data files, and checkpoint encoding — are published as @minnowdb/core/storage/toolkit, so an adapter for another file-based store can be built from the same parts.

What it creates

One directory per database under minnowdb/ in the origin's private file system: format.json (the native-layout version), wal (the write-ahead log), checkpoint-a and checkpoint-b (alternating full-state snapshots), extents/ (packed, append-only files holding column blocks and index postings), temp/ (query spill), and snapshots-v1/ (checksummed, append-only ledgers for active or recently completed framed-snapshot sessions). Opening fails closed when the format marker is missing or malformed beside storage artifacts; a torn marker is repaired only in an otherwise empty database directory. The locked marker is canonical: extra fields, duplicate keys, or alternate JSON spellings are corruption, not extension points. Posting row locators use a compact delta-varint encoding rather than JSON, so a secondary index does not spend most of its space spelling out bigint values. Deleting the directory deletes the database — deleteOpfsDatabase({ name }) does exactly that, but only when no participating connection is open. It requires Web Locks, holds an exclusive deletion lock, and refuses with OpfsDatabaseInUseError before removing files if a leader or follower holds a shared connection lock. Openers wait until deletion finishes. Close every client and store first. A connection releases its lock only after its shutdown has closed its handles, which finishes shortly after close() returns, so deletion waits up to one second for a lock that is being let go before it refuses. A refusal therefore means a connection is still open.

Close tabs and workers running Minnow versions before 0.9.0 before deletion: those clients do not participate in this lock protocol. Normal reads and writes can still run without Web Locks; deletion is refused in that environment. Delete only data you have deliberately decided is disposable. A migration rejection does not authorize deleting the database.

Collection keeps every retained segment’s block references intact, including retired history. Upgrading prevents the previous block-before-segment cleanup bug from recurring, but does not automatically repair a store already affected by missing referenced blocks.

Layout 9 is the current native OPFS contract. Like layout 8, it stores layout 7's bytes. Layout 8 admits compaction jobs whose merge plan is recomputed from its sources rather than stored cell by cell, which a layout-7 build cannot read. Layout 9 admits a commit too large for one WAL frame: its frame is written as continuation frames, a few megabytes each, completed by an ordinary frame, and recovery joins them. A checkpoint may also hold one commit's index changes at any size. A layout-8 build cannot read either. Layouts 7 and 8 upgrade automatically on open by rewriting the format marker alone: a small upgrade-8-9 witness is written and flushed first, so a crash that tears the marker is finished on the next open instead of refused, and the witness is removed once the new marker is durable. An older build then refuses the database without changing it.

Layout 7 added a 48-byte acknowledgement file with two checksummed slots, alongside the structural schemaEpoch and served-request ledger. Recovery requires both slots to be intact and refuses a missing file, damaged slot, or WAL ending before its acknowledged boundary. A zero-filled acknowledged tail is corruption, never an empty store. An interrupted acknowledgement update can therefore prevent reopening; this trades availability for refusing an uncertain rollback. Keep snapshots for recovery from damaged storage.

Layout 6 upgrades automatically on open. Minnow holds exclusive storage ownership, validates a temporary copy, flushes the existing payloads, and publishes current-format checkpoints and the acknowledgement file before resetting the old WAL. Interrupted publication resumes on the next open. A completion receipt prevents an old conversion record from overwriting later writes. Recovery also checks native history when the marker is missing or torn. Stale conversion records cannot override newer checkpoints or repair a damaged installation proof after resetting a nonempty old WAL. The temporary upgrade-6-7/ tree is removed only after current recovery and full integrity validation; the small upgrade-6-7-complete receipt remains. Allow temporary quota for a native copy. No export/import or separate migration call is needed.

Close older connections if they block exclusive ownership, then retry opening. Quota and I/O failures report errors and leave the upgrade retryable. Missing payloads, corrupt old checkpoints and ambiguous incomplete legacy WALs are refused; conversion cannot recreate acknowledgement evidence that layout 6 never stored. Unknown layouts still raise StorageFormatVersionError without replay, automatic repair or deletion. See versioning.

Every extent placement records its offset, length, and a CRC over all stored bytes. The checksum belongs to the placement layer rather than a block codec, so recovery verifies column blocks, postings, and opaque BlockStore payloads uniformly with one linear scan and no decompression.

Collection deletes a sealed extent when it becomes empty and repacks one once dead payloads take more than half its file. Sealed extent bytes therefore stay below twice their live payload, plus the single append tail. A reverse extent index finds only the payloads in the extent being moved, and relocation is split into WAL frames of at most 256 placements. Each batch survives a crash without changing block IDs, and the source extent is deleted only after all of its live placements have moved.

The WAL is capped at 1 GiB and either checkpoint at 256 MiB. The log is folded into a checkpoint slot well before its bound; one commit's frame may take all but 64 MiB of it. The per-resource limits below share the checkpoint bound: everything they admit must still encode into one checkpoint slot, and a write that would push the encoded metadata past it is refused with StorageResourceLimitError (resource: "checkpoint byte"), while aborts, rollbacks, and collection steps still go through, so the database can always be brought back under the bound; getStorageStats() reports the failed checkpoint meanwhile. Relaxed payload writes are flushed first, then the new checkpoint is written and flushed, and only then is the log reset. A checkpoint encodes the whole metadata state, so checkpoints — when the log fills, before garbage collection deletes a drained data file, and at shutdown — encode and write a slice at a time, handing the thread back every few milliseconds. They still hold the write queue, so the state they publish cannot change underneath them, while reads keep answering from memory between slices. Their bytes, slots, and order are the same as a checkpoint written in one step. Reader leases are the one exception: they are logged between slices too, after the state the checkpoint captured, so a checkpoint that logged leases keeps the log instead of resetting it. Recovery skips the frames that checkpoint covers and replays the leases after them. The log is reset once it has been quiet for a second, or else by the next checkpoint, which makes leases wait for it so that it can reset; a drained data file waiting for deletion waits for that reset too. A crash at any moment therefore leaves either a truncated tail frame — detectably incomplete and invisible — or a torn slot with the other slot and the un-reset log still holding everything. A power loss that persists a file's new length ahead of its bytes leaves a zero-filled tail where the unacknowledged frame was to go, or a slot whose magic landed ahead of its header. A zero-filled WAL suffix is accepted as unwritten only when it is beyond the independently recorded acknowledged boundary; missing acknowledged bytes stop recovery. An incomplete checkpoint can fall back to its intact mirror and WAL. Opening the database is the recovery; there is no separate repair step. Recovery runs in steps too: it decodes one slot and reuses that work when the mirror holds the same bytes, parses the checkpoint a bounded piece at a time, builds each table's key membership a slice at a time, and replays the log with turns in between, so opening a database of a million keyed rows holds the thread for under 40 ms at a time on a laptop. Recovery reclaims only a proven unpublished suffix or an artifact absent from the fully validated current layout. A recognized but unsupported envelope or checkpoint version stops before redundant-slot fallback and before any truncation or deletion.

A shutdown checkpoint can advance its generation without adding a WAL entry. If shutdown ends between mirror writes, recovery accepts adjacent generations only when their WAL sequences agree. It still rejects a sequence gap without the log that bridges it. Staged index builds can renew their leases across checkpoints; recovery validates the current lease interval, not the total age of the build. Low-level renewal and chunk-append calls must place the expiration after both expiresAtCutoff and updatedAt, within the one-hour lease limit of each. Invalid intervals are rejected before writing durable state.

Removing pruned manifest summaries advances from the oldest stored version in bounded steps. Every intermediate checkpoint retains a contiguous predecessor chain, and replay selects the same prefix from durable state. Cleanup never depends on a cursor held only in worker memory. This prevents new gaps from cleanup. It does not repair a checkpoint that already contains an invalid manifest chain: recovery still rejects it rather than guessing which missing records were safe to delete. Such a database needs a known-good snapshot.

Durability

strict, the default, flushes each payload, the WAL frame that publishes it, and the separate acknowledgement boundary before the operation resolves. A missing or corrupt durable payload is reported as storage corruption rather than silently rolled back. Use this mode whenever losing an acknowledged write is unacceptable.

An acknowledgement flush failure is reported through onDiagnostic (or console.error) and returns OpfsUncertainOutcomeError to low-level callers. The WAL may already contain the mutation; the engine reconciles its transaction before deciding whether it can return success. Never retry an uncertain mutation blindly. A throwing diagnostic hook is contained and its error is logged.

relaxed completes the payload and WAL writes before acknowledging, but lets the operating system schedule their final flush. A tab, worker, or browser-process failure normally recovers the acknowledged writes already held by the operating system; a machine or device power loss can lose an acknowledged suffix. Recovery still accepts only a consecutive WAL prefix whose extent ranges and checksums verify. The extra strict-mode flushes cost write throughput, so select relaxed only when every acknowledged write can be reconstructed, replayed, or recovered from another durable source, and that loss window is an accepted performance trade-off. A later checkpoint makes the writes it includes durable, but relaxed mode provides no per-write promise about when that checkpoint or an operating-system flush will happen.

Quota

The same origin quota as everything else, and the same two calls to watch it:

import { ensureOriginPersistence } from "@minnowdb/core/storage/persistence";

const { quota, usage } = await navigator.storage.estimate();
await ensureOriginPersistence("required");

A refused write escapes as the browser's own QuotaExceededError, with everything committed before it intact and the same write succeeding once space frees. Origin persistence protects against automatic quota eviction when the browser grants it; it cannot prevent the user from clearing site data. Applications whose local rows are authoritative should require both strict durability and origin persistence. When acknowledged writes must survive deliberate site-data clearing or device loss too, export or synchronize an independent durable copy. getLogicalStorageBytes() reports the database's logical footprint, including its control log, newest checkpoint, and temporary spill.

Use getStorageStats() when diagnosing growth. It reports actual bytes across extents, WAL, checkpoint slots, and temporary files alongside live and obsolete logical bytes, plus checkpoint and cleanup failure health, cleanup debt and its backpressure limit, and the WAL safety limit. checkIntegrity({ mode: "full" }) reads every live extent placement and verifies both its extent checksum and its block or postings envelope; metadata mode checks checkpoint, WAL, reference, and live-byte accounting without paying the complete payload scan.

Failed file deletion or truncation becomes explicit cleanup debt rather than a forgotten promise. The leader retries it before later growth; 64 MiB of unresolved debt applies write backpressure so repeated filesystem failures cannot grow packed files without limit. Statistics report the debt, last cleanup error, and ceiling. An unrecognized zero-byte temporary file still contributes 64 KiB of modeled debt, so file-count growth cannot bypass the byte ceiling.

Durable resource limits

Low-level storage calls have hard global limits as well as bounded batches. A database admits at most 4,096 active transactions, 4,096 leases, 1,024 temporary owners, 1,024 active compaction jobs, one active garbage-collection job, 1,024 active UNIQUE builds, 128 active full-text builds, 128 active secondary-index builds, and 4,096 catalog table/view records, including a table owned by an unpublished transaction. Active transaction journals are capped in aggregate at 1,048,576 staged blocks, 1,048,576 staged segments, and 512 MiB; a single transaction's journal has no length limit of its own. Staged UNIQUE, full-text, and secondary-index generations share a 1 GiB and 16,777,216-entry ceiling. The catalog also has a 64 MiB aggregate ceiling measured as its exact UTF-8 record-wire bytes. Manifest history holds at most 65,536 summaries or 64 MiB of reservation bytes, and segment metadata holds at most 1,048,576 records or 512 MiB of exact segment-wire bytes. An unpruned manifest reserves the exact canonical summary bytes it will occupy after its 24-byte UTC prunedAt tombstone is added. Pruning is therefore byte-neutral even at the ceiling; removing the tombstone later releases the reservation. Collecting or dropping a segment releases its exact reservation. The counts, bytes, and records change inside one leader operation before its WAL frame is acknowledged. Checkpoint load and reopen recompute the exact accounting from the records and reject an over-limit or inconsistent state before accepting new growth.

Temporary spill is bounded across every tab by a durable reservation ledger: 65,536 runs, 262,144 pages, and 1 GiB per database, with per-owner limits of 1,024 runs, 16,384 pages, and 512 MiB. Overwriting a page reserves only its byte delta; deleting a run or expired owner releases the exact reservation. Reopening streams the temporary directory, reconciles missing and orphaned files, rebuilds the counters, and refuses a write before creating a file if a limit would be crossed.

Durable history is bounded too. A reader or maintenance pin may lag the current manifest by at most 4,096 versions and may retain at most 65,536 retired blocks or 512 MiB. Total retired payload history is capped at 1 GiB, and retained terminal metadata is capped at 65,536 transactions, 4,096 compaction jobs, and 1,024 completed collection jobs. These limits are checked before the WAL or an extent changes, so refusal leaves the previous durable state intact.

Framed snapshots

OPFS exports one semantic metadata item per bounded snapshot frame and one raw block per block frame. Imports durably stage consecutive frames and publish the catalog, generations, and blocks in one final operation; replay compares exact bytes, so a lost acknowledgement cannot duplicate or silently replace an import. Final validation shares the durable accelerator ceiling of 1 GiB and 16,777,216 retained entries across combined UNIQUE and posting generations; a larger framed import is refused before publication, leaving the destination logically empty and the staged session resumable or cancellable instead of risking an origin-sized validation allocation. Closing, cancelling, or replacing a session checkpoints the resolved state and resets WAL history before deleting its ledger, so every crash point retains either the old resumable session or the new durable state.

Ready posting accelerators remain ready when their exact merged generation fits the ordered-read fuse: at most 64 MiB and 65,536 retained row IDs in OPFS. A larger accelerator is deliberately omitted from the portable snapshot, restored as invalid, and rebuilt in the background; queries continue to scan correctly while that rebuild runs. UNIQUE membership remains authoritative and enforced.

Multiple tabs

Several tabs may open the same database at once; every operation from every tab lands in one ordered log, so readers never see half of anything. Reads wait for recovery after a refused WAL append instead of observing its tentative state. The leader orders individual operations; a transaction's whole read, stage, and commit sequence is ordered one level up, where every writer in every tab takes its turn through the Web Lock minnowdb-write:minnowdb-live:opfs:<name> before reading the state it depends on (see writer turns), so tabs never conflict with each other and compare-and-swap in the log stays a defensive check. A tab paused by the browser inside its turn holds the others until it resumes or is discarded; the wait is reported through onBackgroundError after ten seconds and never bypassed, and closing a waiting engine or client cancels its wait at once. The moving parts differ from the IndexedDB adapter in one honest way: multi-tab responsiveness uses BroadcastChannel (universal since 2022), while correctness rests entirely on the storage lock and the checksummed log. A browser with no channel at all would still be safe — one tab at a time would hold the database.

Browsers can suspend a hidden leader while it still owns exclusive file handles. Foreground leadership preference is cooperative: the old leader must run and yield before another worker can take over, which is why an idle hidden leader releases its handles ahead of time. A leader frozen while busy can still delay other tabs or time out a request; durability does not guarantee uninterrupted availability. Resume or terminate that owner to release its handles, then reconcile any unknown write outcome before retrying. Background maintenance is stepped and resumable for the same reason.

Choosing between OPFS and IndexedDB

Both are durable, both are safe across tabs, and both hold the same block format — a snapshot moves a database between them. The benchmarks page measures both live in your browser. Synchronous file access can reduce storage overhead, but relative latency depends on the workload, browser, durability, and competing tabs. Measure the queries and writes your application uses. Choose IndexedDB where OPFS is unavailable (older browsers, Safari private windows).

On this page