mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 00:48:18 +00:00
v0.42.75.0 fix(pglite): in-place WAL auto-repair for the macOS Aborted() startup crash (#2575, #223, #1670) (#3901)
* fix(pglite): in-place WAL auto-repair for the Aborted() startup crash (#223, #1670, #2575) The 'macOS 26.x WASM bug' was a misdiagnosis: an unclean shutdown (typically the OS-upgrade reboot) tears the data dir's WAL, and every subsequent open fails WAL replay inside WASM with an opaque RuntimeError: Aborted(). This ports the pg_resetwal recovery upstream rejected (electric-sql/pglite#994, by @yestheboxer) and wires it into connect() as bounded auto-repair: - src/core/pglite-resetwal.ts: pg_resetwal for PG17 NodeFS dirs, fail-closed layout validation, atomic+durable writes (tmp+fsync+rename), idempotent. - src/core/pglite-repair.ts: whole-pg_wal-dir rename backup (zero transient disk), overwrite-order restore with mtime guard, cooldown sidecar + episode-scoped backup retention (newest 3 episodes), and a never-throws engine seam. Kill-switch: GBRAIN_PGLITE_WAL_REPAIR=off. - pglite-engine.ts: verdict rename macos-26-3 -> wasm-abort, classifier now matches the real production message (it previously fell to 'unknown'), corrupt-beats-wasm precedence preserved, honest per-outcome error copy incl. the failed-not-restored arm, and repair only under a cleanly-acquired lock (new LockHandle.reaped provenance; never after reaping a holder). - gbrain pglite-repair: manual dry-run/repair command (validate-before-lock, serve/reaped refusals, no --force by design). - doctor: pglite_data_dir fs-check with recurrence escalation and backup inventory when a PGLite brain fails to connect. - reinit-pglite: embedding flags default from file-only config so the recovery ladder's rebuild rung works bare mid-outage. - stringifyPgliteInitError: message-less Emscripten ErrnoError objects no longer surface as [object Object]. Regression-tested against real brains: corrupt every WAL segment (truncate and garbage variants), reopen, auto-repair fires, original rows readable, process.exitCode stays contained (#2084). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(pglite): replace the macOS-26.x misdiagnosis with the corrupt-WAL recovery ladder README + INSTALL.md shipped (via #1671) the claim that PGLite is incompatible with macOS 26.x and that a Bun/WASM fix would restore it. The real cause is torn WAL state from the upgrade reboot, now auto-repaired in place. Rewrites those sections around the recovery ladder (auto-repair -> gbrain pglite-repair -> reinit-pglite -> engine switch; native-Postgres recipe kept, credit @roysaurav), adds the ENGINES.md troubleshooting section, updates the KEY_FILES.md entries to current truth, files the two follow-up TODOs (SIGTERM engine-close extension; pglite upgrade blocker), and regenerates the llms bundles. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pglite): harden WAL auto-repair (pre-landing + adversarial review) Review-army (security/testing/maintainability/perf) + Claude & Codex adversarial passes on the WAL-repair wave. Correctness + safety hardening, no behavior change to the happy path: - Live-writer safety: repair refuses any reaped lock acquisition, a corrupt (unknowable-liveness) reap writes a cross-process quarantine marker that gates auto-repair AND the manual command for 10 min, isProcessAlive treats only ESRCH as dead (EPERM/malformed-pid read as alive), and a live postmaster.pid (native Postgres) is refused. Lock heartbeat + initial write are atomic (tmp+rename) so a torn read can't misclassify a healthy holder; an in-flight acquisition is no longer mistaken for corrupt. - resetWal verifies the stored pg_control CRC before trusting/re-signing it — a damaged control file routes to rebuild instead of laundering corrupt checkpoint counters under a fresh CRC. Atomic 'wx' writes (no symlink follow), whole-pg_wal-dir rename backup, 64MB seg-size cap. - Honest failure reporting: repairPgliteWal threads the real restore result out via WalRepairError so the 'failed-restored' vs 'failed-not-restored' message never lies; the not-restored copy names the correct restore paths. - Episode lifecycle: episodes close on the next healthy connect (not just on a verified repair), a gutted (restored) backup loses its pin, stale (>24h) episode backups aren't reused, and the cooldown also caps repaired-only crash loops. Empty backup dirs are pruned on refusal. - Command: rejects unknown flags and valueless --path (a destructive command must not silently mis-parse), confirm prompt goes to stderr (stdout stays clean for --json), embedding-flag defaults come from the config file only. - Symlink confinement extended to global/; sidecar reuse path validated (prefix + no '..' + must still hold pg_wal); sidecar writes atomic. - doctor recurrence escalation counts all attempts; data dir absolutized. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(pglite): current-state KEY_FILES + WAL-repair follow-up TODOs KEY_FILES.md pglite entries updated to the hardened truth (reap marker + quarantine, atomic writes, CRC gate, global-symlink refusal, WalRepairError, episode lifecycle). TODOS.md files the deferred judgment-call follow-ups (unclean-shutdown gate on auto-repair; non-gbrain pglite consumer boundary; mixed-version torn-lock double-read). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v0.42.75.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
0b47afbf40
commit
f15480b9d0
@@ -2,6 +2,37 @@
|
||||
|
||||
All notable changes to GBrain will be documented in this file.
|
||||
|
||||
## [0.42.75.0] - 2026-08-08
|
||||
|
||||
**The "PGLite crashes on macOS 26" era is over: gbrain now repairs a torn brain in place, automatically, with your data preserved.**
|
||||
|
||||
The dreaded `RuntimeError: Aborted()` at startup — the one that made zero-config brains unusable after a macOS upgrade and pushed people onto Homebrew Postgres — was never a macOS or WASM bug. An unclean shutdown (typically the upgrade reboot) tears the write-ahead log inside the data dir, and every open after that dies replaying it. gbrain now detects that failure on any command, backs up the WAL state to a sibling directory, resets it in place (the pg_resetwal recovery Postgres has shipped for decades, ported to run against PGLite data dirs), and reopens your brain — pages, embeddings, and history intact. Transactions that never reached a checkpoint may be lost; that is the standard trade for a database that would otherwise not open at all.
|
||||
|
||||
### Added
|
||||
- **Automatic WAL repair on startup.** A torn-WAL abort self-heals on the next gbrain command: backup → in-place reset → retry, with a loud notice naming the backup and recommending `gbrain doctor`. Disable with `GBRAIN_PGLITE_WAL_REPAIR=off`.
|
||||
- **`gbrain pglite-repair`** — the deliberate version: `--dry-run` gives a read-only diagnosis of the data dir; `--yes` runs the same in-place repair manually. Refuses to operate while any live process holds the brain, and never force-removes another process's lock.
|
||||
- **`gbrain doctor` diagnoses unopenable PGLite brains.** A new `pglite_data_dir` check reads the data dir from disk when connect fails, names the right recovery rung (repair vs rebuild), inventories repair backups, and escalates when repairs keep recurring — the signal that something is still killing gbrain mid-write.
|
||||
- **Recovery guardrails throughout:** repair runs only under a cleanly-acquired lock (never after taking over another process's lock, with a quarantine window when a lock's holder couldn't be verified); a live database — including a native Postgres one — is refused by a `postmaster.pid` liveness check; repeated attempts inside one corruption episode reuse one backup instead of stacking copies (newest three episodes retained); a cooldown stops repair loops from silently eating data on machines where crashes keep recurring; and every restore path reports honestly whether your original files are back in place or waiting in the backup.
|
||||
|
||||
### Changed
|
||||
- **`gbrain reinit-pglite` works bare.** The embedding model and dimensions now default from your config file, so the rebuild rung of the recovery ladder is one command mid-outage (explicit flags still win; environment overrides are deliberately ignored so a stale shell export can't change the rebuild target).
|
||||
- **Honest error messages.** The startup-abort hint now names the real cause (torn WAL after an unclean shutdown), states exactly what auto-repair did or why it stood down, and lays out the full ladder: repair → rebuild → engine switch. The docs that claimed PGLite is "incompatible with macOS 26.x" have been rewritten (README, INSTALL, ENGINES) — thanks @roysaurav for the original native-Postgres walkthrough, which remains the engine-switch rung.
|
||||
- Message-less WASM error objects no longer surface as `[object Object]`.
|
||||
|
||||
### Fixed
|
||||
- The classifier that routes startup failures now matches the abort message PGLite actually produces (it previously fell through to a generic hint), while catalog corruption keeps routing to rebuild — WAL repair is never suggested for damage it cannot fix.
|
||||
- Lock-file reads can no longer misclassify a healthy live holder as corrupt (writes are atomic now), a holder owned by another user is treated as alive, and an in-flight acquisition is no longer mistaken for a corrupt lock.
|
||||
|
||||
Credit where due: @yang1996202-cpu (#2575), @AndreLYL (#223), and @roysaurav (#1670) for reports and diagnosis, the #223 thread contributors whose recoveries proved the root cause, and @yestheboxer, whose rejected upstream recovery PR (electric-sql/pglite#994) this port builds on.
|
||||
|
||||
### To take advantage of v0.42.75.0
|
||||
|
||||
```bash
|
||||
gbrain upgrade
|
||||
```
|
||||
|
||||
If your brain currently won't open, that's it — the next command repairs it. If you'd rather look first: `gbrain pglite-repair --dry-run`.
|
||||
|
||||
## [0.42.74.0] - 2026-08-07
|
||||
|
||||
**Two fixes for agents that reach a brain over the network: takes-holder visibility now works the way you set it, and the voice recipe is safe by default.**
|
||||
|
||||
@@ -317,7 +317,7 @@ Data flowing into the brain. Each integration is a recipe — markdown + setup h
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**`gbrain init --pglite` crashes on macOS 26.x (Tahoe)?** PGLite's embedded WASM engine is incompatible with macOS 26.x on Apple Silicon. The fix is to use native Homebrew PostgreSQL + pgvector instead. Full step-by-step setup in [`docs/INSTALL.md` — Troubleshooting: PGLite crashes on macOS 26.x](docs/INSTALL.md#pglite-crashes-on-macos-26x-tahoe).
|
||||
**PGLite crashes at startup with `RuntimeError: Aborted()` (often right after a macOS upgrade)?** Not a macOS incompatibility — the OS-upgrade reboot killed gbrain mid-write and tore the data dir's WAL. gbrain now repairs this automatically on the next command (data preserved, backup kept); if auto-repair is disabled or skipped, run `gbrain pglite-repair --dry-run` to diagnose and `gbrain pglite-repair --yes` to repair in place. Full recovery ladder (repair → rebuild → engine switch) in [`docs/ENGINES.md` — Troubleshooting: startup abort](docs/ENGINES.md#troubleshooting-startup-abort-runtimeerror-aborted) and [`docs/INSTALL.md`](docs/INSTALL.md#pglite-crashes-on-macos-26x-tahoe).
|
||||
|
||||
**`gbrain import` fails with `expected N dimensions, not M`?** Run `gbrain doctor`. It will print the exact `gbrain config set ...` or `gbrain retrieval-upgrade` command to repair the mismatch. You should not need to delete `~/.gbrain`. Fresh `gbrain init --pglite` auto-detects your embedding provider from API keys in your environment: set `OPENAI_API_KEY` (or `ZEROENTROPY_API_KEY` / `VOYAGE_API_KEY`) before running init, or pass `--embedding-model <provider>:<model>` explicitly. With multiple keys set, init fires an interactive picker. In non-TTY contexts (CI, Docker) with no keys, init exits 1 with a paste-ready setup hint; pass `--no-embedding` to defer setup until runtime. See [`docs/integrations/embedding-providers.md`](docs/integrations/embedding-providers.md) for the full provider matrix and [`docs/operations/headless-install.md`](docs/operations/headless-install.md) for Docker/CI sequencing.
|
||||
|
||||
|
||||
@@ -1,5 +1,54 @@
|
||||
# TODOS
|
||||
|
||||
## WAL-repair wave follow-ups (#223/#1670/#2575)
|
||||
|
||||
- [ ] **P2 — gate auto-repair on an unclean-shutdown marker (adversarial F7).** The classifier
|
||||
deliberately over-matches (`RuntimeError`/`unreachable` → `wasm-abort`). If an unclean
|
||||
shutdown leaves a REPLAYABLE WAL tail (normal crash recovery would restore those committed
|
||||
txns) and the reopen then fails on a transient WASM error (OOM), auto-repair fires, layout
|
||||
validation can't tell torn from replayable, and resetWal discards the tail while the notice
|
||||
says "data preserved." Bounded today (backup always taken + restore + honest failure + repeated
|
||||
attempts capped), but a false-positive-with-successful-retry silently drops committed data.
|
||||
Fix direction (probe-verified): PGLite removes `postmaster.pid` on clean close, so gate AUTO
|
||||
repair (not the manual command) on `postmaster.pid` presence — a clean dir that aborts is not
|
||||
torn-WAL. Requires making the serial regression test stamp a `postmaster.pid` before corrupting
|
||||
(it currently clean-disconnects then corrupts, which the red-team flagged as unfaithful anyway).
|
||||
Needs a recall/precision call before landing.
|
||||
- [ ] **P3 — live non-gbrain PGLite consumer not caught by the postmaster.pid liveness guard
|
||||
(adversarial F8).** PGLite writes a sentinel `postmaster.pid` of `-42`; the liveness refusal in
|
||||
`validateWalRepairTarget` requires `pid > 0`, so it protects native Postgres dirs but not a
|
||||
non-gbrain pglite app that has the dir open (such an app writes no `.gbrain-lock`). Deliberate
|
||||
misuse of `pglite-repair --path <foreign pglite dir>` required. Option: refuse when
|
||||
postmaster.pid holds pid ≤ 0 with a very recent mtime, or document the boundary.
|
||||
- [ ] **P3 — mixed-version torn-lock read (adversarial F10 residual).** The heartbeat + initial
|
||||
lock writes are atomic (tmp+rename) now, but an OLD gbrain binary writing heartbeats IN PLACE
|
||||
while a NEW binary poll-reads can still catch a torn read → corrupt-lock verdict → a live
|
||||
holder's lock reaped → two writers (the #2348 class, version-skew-triggered). The reap marker
|
||||
quarantines repair, not the concurrent open. Cheap hardening: double-read the lock file (~50ms
|
||||
apart) before declaring it corrupt.
|
||||
|
||||
|
||||
- [ ] **P2 — graceful PGLite close on SIGTERM for the remaining long-running paths.**
|
||||
The torn-WAL genesis this wave repairs is an unclean shutdown: `src/core/process-cleanup.ts`
|
||||
releases locks on SIGTERM but never closes the PGlite handle, so `serve` / `jobs work` /
|
||||
`sync` killed mid-write (macOS-upgrade reboot, `systemctl stop`) leave the WAL torn.
|
||||
Autopilot already ships the pattern (d2fd1f29, #3178/#1872: `registerCleanup('autopilot-engine-close', ...)`
|
||||
— abort in-flight work → ≤2s bounded wait inside the 3s cleanup deadline →
|
||||
`engine.disconnect()`, double-call safe; rationale comment at autopilot.ts:438-452).
|
||||
Extend that exact pattern to the remaining long-running PGLite paths (register in
|
||||
connect()/command scope; dedupe so autopilot doesn't double-close), pinned by a serial
|
||||
lifecycle test. Interacts with #2084 exitCode containment + #1337 close ordering — read
|
||||
those comments in pglite-engine.ts first. Auto-repair makes recurrence self-healing
|
||||
meanwhile, so this is prevention, not recovery.
|
||||
- [ ] **P3 — pglite upgrade blocker tracker.** Two couplings make a "routine" pglite bump a
|
||||
breaking change: (a) pglite ≥0.5 removes the `@electric-sql/pglite/vector` export that
|
||||
`pglite-engine.ts` imports (verified against npm); (b) the pg_resetwal port
|
||||
(`src/core/pglite-resetwal.ts`) is coupled to the PG17 pg_control layout
|
||||
(`PG_CONTROL_VERSION` 1700 — guarded at runtime by `WalResetUnsupportedError`, so a
|
||||
mismatched bump makes the repair tool refuse every dir rather than corrupt, but it still
|
||||
means the repair feature silently dies). Any future pglite upgrade wave must revisit BOTH
|
||||
together and re-derive the ControlFileData offset table for the new PG major.
|
||||
|
||||
## serve --http takes-holders + agent-voice hardening follow-ups (filed v0.42.74.0)
|
||||
|
||||
Deferred from the #2529/#2477 security-fix wave (plan-eng-review + codex outside
|
||||
|
||||
@@ -221,6 +221,66 @@ live in `test/postgres-engine-rls-scope.test.ts`.
|
||||
|
||||
**Migration:** `gbrain migrate --to supabase` exports everything (pages, chunks, embeddings, links, tags, timeline) and imports into Supabase. `gbrain migrate --to pglite` goes the other direction. Bidirectional, lossless.
|
||||
|
||||
### Troubleshooting: startup abort (`RuntimeError: Aborted()`)
|
||||
|
||||
**Symptom:** every PGLite-touching command dies at startup with
|
||||
`PGLite failed to initialize its WASM runtime … Aborted(). Build with
|
||||
-sASSERTIONS for more info.` — commonly first seen right after a macOS
|
||||
upgrade.
|
||||
|
||||
**Real root cause:** corrupt WAL/checkpoint state in the data dir after an
|
||||
unclean shutdown (the OS-upgrade reboot kills gbrain mid-write and tears the
|
||||
write-ahead log; every subsequent open fails WAL replay inside WASM and
|
||||
Emscripten surfaces only the opaque abort). It is **not** a macOS/WASM
|
||||
incompatibility — the same signature reproduces across macOS versions and on
|
||||
Linux, and rebuilding the data dir on the same OS fixes it. No pglite or Bun
|
||||
version bump changes it.
|
||||
|
||||
**Recovery ladder** (top rung first):
|
||||
|
||||
1. **Auto-repair (default).** `PGLiteEngine.connect()` detects the abort,
|
||||
backs up `pg_wal/` + `pg_control` into a sibling
|
||||
`<dataDir>.wal-repair-backup-<ts>/` dir, resets the WAL in place
|
||||
(pg_resetwal semantics — data files preserved; transactions not
|
||||
checkpointed before the corruption may be lost), and retries once. On
|
||||
success it prints a loud stderr notice naming the backup and recommending
|
||||
`gbrain doctor`. Safety bounds: repair only runs under a cleanly-acquired
|
||||
data-dir lock (never after reaping another process's lock), skips for a
|
||||
cooldown window after a failed attempt
|
||||
(`GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS`, default 3600), reuses one
|
||||
backup per corruption episode (newest 3 episodes retained), and restores
|
||||
the original files if the retry still fails. Kill-switch:
|
||||
`GBRAIN_PGLITE_WAL_REPAIR=off`.
|
||||
2. **Manual repair.** `gbrain pglite-repair --dry-run` diagnoses the data dir
|
||||
(read-only); `gbrain pglite-repair --yes` runs the same in-place WAL reset
|
||||
deliberately. Refuses when another gbrain process holds the brain (a live
|
||||
`gbrain serve` is named explicitly) and never force-removes `.gbrain-lock`.
|
||||
3. **Rebuild.** `gbrain reinit-pglite` (embedding model/dimensions default
|
||||
from your config) wipes and re-creates the brain from your brain repo, or
|
||||
manually: back up `~/.gbrain`, move `brain.pglite` aside,
|
||||
`gbrain init --pglite`, re-add sources, `gbrain sync`, `gbrain embed`.
|
||||
Required for *catalog* corruption (58P01 / pgvector load failure) — WAL
|
||||
repair cannot fix that class.
|
||||
4. **Switch engines.** `gbrain init --supabase`, or native Postgres +
|
||||
pgvector (recipe below, contributed by @roysaurav):
|
||||
|
||||
```bash
|
||||
brew install postgresql@17
|
||||
brew services start postgresql@17
|
||||
createdb gbrain
|
||||
cd /tmp && git clone --branch v0.8.0 https://github.com/pgvector/pgvector.git
|
||||
cd pgvector && make && make install
|
||||
psql gbrain -c "CREATE EXTENSION IF NOT EXISTS vector;"
|
||||
# ~/.gbrain/config.json: { "engine": "postgres",
|
||||
# "database_url": "postgresql://localhost:5432/gbrain" }
|
||||
gbrain apply-migrations --yes && gbrain doctor
|
||||
```
|
||||
|
||||
`gbrain doctor` runs a `pglite_data_dir` check whenever a PGLite brain fails
|
||||
to connect: it diagnoses the dir from disk, names the repair command, reports
|
||||
retained repair backups, and escalates when repairs keep recurring (that
|
||||
means the unclean-shutdown genesis is still active — see the ladder's rung 4).
|
||||
|
||||
## JSONB writes: never double-encode (the #2339 trap)
|
||||
|
||||
Writing a JS value into a `jsonb` column has exactly two correct forms. Get this
|
||||
|
||||
+15
-4
@@ -117,7 +117,20 @@ If anything's yellow, `gbrain doctor` names the fix command in the message. Most
|
||||
|
||||
### PGLite crashes on macOS 26.x (Tahoe)
|
||||
|
||||
PGLite's embedded WASM engine is incompatible with macOS 26.x (Tahoe) on Apple Silicon. If `gbrain init --pglite` crashes during engine initialization, switch to native Homebrew PostgreSQL:
|
||||
This crash (`RuntimeError: Aborted()` at engine startup, typically first seen
|
||||
after a macOS upgrade) is **not** a macOS/WASM incompatibility. The upgrade
|
||||
reboot kills gbrain mid-write and tears the data dir's write-ahead log; every
|
||||
subsequent open then fails WAL replay. Recovery ladder:
|
||||
|
||||
1. **Auto-repair (default):** just run any gbrain command — gbrain detects the
|
||||
abort, resets the WAL in place (data preserved; a backup of the pre-repair
|
||||
state is kept next to the data dir), and continues. Then run `gbrain doctor`.
|
||||
2. **Manual repair:** `gbrain pglite-repair --dry-run` to diagnose,
|
||||
`gbrain pglite-repair --yes` to repair in place.
|
||||
3. **Rebuild:** `gbrain reinit-pglite` (wipes and re-creates the brain from
|
||||
your brain repo; embedding settings default from your config).
|
||||
4. **Switch engines** — if you prefer a server database anyway, native
|
||||
Homebrew PostgreSQL works great and supports multiple concurrent agents:
|
||||
|
||||
```bash
|
||||
# Install PostgreSQL + pgvector
|
||||
@@ -144,6 +157,4 @@ gbrain apply-migrations --yes
|
||||
gbrain doctor
|
||||
```
|
||||
|
||||
All 102 migrations run on first try. Once `gbrain doctor` shows green, the brain works identically to PGLite — same commands, same skills, same data model. The only difference is the storage backend.
|
||||
|
||||
> **Note:** This workaround is temporary. When the upstream WASM runtime fix ships (likely via a Bun update), `--pglite` will work on Tahoe again.
|
||||
Once `gbrain doctor` shows green, the brain works identically to PGLite — same commands, same skills, same data model. The only difference is the storage backend (plus multi-connection support: several agents can share one Postgres brain, which PGLite's single-process lock doesn't allow).
|
||||
|
||||
File diff suppressed because one or more lines are too long
+61
-1
@@ -1831,7 +1831,7 @@ Data flowing into the brain. Each integration is a recipe — markdown + setup h
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**`gbrain init --pglite` crashes on macOS 26.x (Tahoe)?** PGLite's embedded WASM engine is incompatible with macOS 26.x on Apple Silicon. The fix is to use native Homebrew PostgreSQL + pgvector instead. Full step-by-step setup in [`docs/INSTALL.md` — Troubleshooting: PGLite crashes on macOS 26.x](docs/INSTALL.md#pglite-crashes-on-macos-26x-tahoe).
|
||||
**PGLite crashes at startup with `RuntimeError: Aborted()` (often right after a macOS upgrade)?** Not a macOS incompatibility — the OS-upgrade reboot killed gbrain mid-write and tore the data dir's WAL. gbrain now repairs this automatically on the next command (data preserved, backup kept); if auto-repair is disabled or skipped, run `gbrain pglite-repair --dry-run` to diagnose and `gbrain pglite-repair --yes` to repair in place. Full recovery ladder (repair → rebuild → engine switch) in [`docs/ENGINES.md` — Troubleshooting: startup abort](docs/ENGINES.md#troubleshooting-startup-abort-runtimeerror-aborted) and [`docs/INSTALL.md`](docs/INSTALL.md#pglite-crashes-on-macos-26x-tahoe).
|
||||
|
||||
**`gbrain import` fails with `expected N dimensions, not M`?** Run `gbrain doctor`. It will print the exact `gbrain config set ...` or `gbrain retrieval-upgrade` command to repair the mismatch. You should not need to delete `~/.gbrain`. Fresh `gbrain init --pglite` auto-detects your embedding provider from API keys in your environment: set `OPENAI_API_KEY` (or `ZEROENTROPY_API_KEY` / `VOYAGE_API_KEY`) before running init, or pass `--embedding-model <provider>:<model>` explicitly. With multiple keys set, init fires an interactive picker. In non-TTY contexts (CI, Docker) with no keys, init exits 1 with a paste-ready setup hint; pass `--no-embedding` to defer setup until runtime. See [`docs/integrations/embedding-providers.md`](docs/integrations/embedding-providers.md) for the full provider matrix and [`docs/operations/headless-install.md`](docs/operations/headless-install.md) for Docker/CI sequencing.
|
||||
|
||||
@@ -2216,6 +2216,66 @@ live in `test/postgres-engine-rls-scope.test.ts`.
|
||||
|
||||
**Migration:** `gbrain migrate --to supabase` exports everything (pages, chunks, embeddings, links, tags, timeline) and imports into Supabase. `gbrain migrate --to pglite` goes the other direction. Bidirectional, lossless.
|
||||
|
||||
### Troubleshooting: startup abort (`RuntimeError: Aborted()`)
|
||||
|
||||
**Symptom:** every PGLite-touching command dies at startup with
|
||||
`PGLite failed to initialize its WASM runtime … Aborted(). Build with
|
||||
-sASSERTIONS for more info.` — commonly first seen right after a macOS
|
||||
upgrade.
|
||||
|
||||
**Real root cause:** corrupt WAL/checkpoint state in the data dir after an
|
||||
unclean shutdown (the OS-upgrade reboot kills gbrain mid-write and tears the
|
||||
write-ahead log; every subsequent open fails WAL replay inside WASM and
|
||||
Emscripten surfaces only the opaque abort). It is **not** a macOS/WASM
|
||||
incompatibility — the same signature reproduces across macOS versions and on
|
||||
Linux, and rebuilding the data dir on the same OS fixes it. No pglite or Bun
|
||||
version bump changes it.
|
||||
|
||||
**Recovery ladder** (top rung first):
|
||||
|
||||
1. **Auto-repair (default).** `PGLiteEngine.connect()` detects the abort,
|
||||
backs up `pg_wal/` + `pg_control` into a sibling
|
||||
`<dataDir>.wal-repair-backup-<ts>/` dir, resets the WAL in place
|
||||
(pg_resetwal semantics — data files preserved; transactions not
|
||||
checkpointed before the corruption may be lost), and retries once. On
|
||||
success it prints a loud stderr notice naming the backup and recommending
|
||||
`gbrain doctor`. Safety bounds: repair only runs under a cleanly-acquired
|
||||
data-dir lock (never after reaping another process's lock), skips for a
|
||||
cooldown window after a failed attempt
|
||||
(`GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS`, default 3600), reuses one
|
||||
backup per corruption episode (newest 3 episodes retained), and restores
|
||||
the original files if the retry still fails. Kill-switch:
|
||||
`GBRAIN_PGLITE_WAL_REPAIR=off`.
|
||||
2. **Manual repair.** `gbrain pglite-repair --dry-run` diagnoses the data dir
|
||||
(read-only); `gbrain pglite-repair --yes` runs the same in-place WAL reset
|
||||
deliberately. Refuses when another gbrain process holds the brain (a live
|
||||
`gbrain serve` is named explicitly) and never force-removes `.gbrain-lock`.
|
||||
3. **Rebuild.** `gbrain reinit-pglite` (embedding model/dimensions default
|
||||
from your config) wipes and re-creates the brain from your brain repo, or
|
||||
manually: back up `~/.gbrain`, move `brain.pglite` aside,
|
||||
`gbrain init --pglite`, re-add sources, `gbrain sync`, `gbrain embed`.
|
||||
Required for *catalog* corruption (58P01 / pgvector load failure) — WAL
|
||||
repair cannot fix that class.
|
||||
4. **Switch engines.** `gbrain init --supabase`, or native Postgres +
|
||||
pgvector (recipe below, contributed by @roysaurav):
|
||||
|
||||
```bash
|
||||
brew install postgresql@17
|
||||
brew services start postgresql@17
|
||||
createdb gbrain
|
||||
cd /tmp && git clone --branch v0.8.0 https://github.com/pgvector/pgvector.git
|
||||
cd pgvector && make && make install
|
||||
psql gbrain -c "CREATE EXTENSION IF NOT EXISTS vector;"
|
||||
# ~/.gbrain/config.json: { "engine": "postgres",
|
||||
# "database_url": "postgresql://localhost:5432/gbrain" }
|
||||
gbrain apply-migrations --yes && gbrain doctor
|
||||
```
|
||||
|
||||
`gbrain doctor` runs a `pglite_data_dir` check whenever a PGLite brain fails
|
||||
to connect: it diagnoses the dir from disk, names the repair command, reports
|
||||
retained repair backups, and escalates when repairs keep recurring (that
|
||||
means the unclean-shutdown genesis is still active — see the ladder's rung 4).
|
||||
|
||||
## JSONB writes: never double-encode (the #2339 trap)
|
||||
|
||||
Writing a JS value into a `jsonb` column has exactly two correct forms. Get this
|
||||
|
||||
+1
-1
@@ -148,7 +148,7 @@
|
||||
"bun": ">=1.3.10"
|
||||
},
|
||||
"license": "MIT",
|
||||
"version": "0.42.74.0",
|
||||
"version": "0.42.75.0",
|
||||
"overrides": {
|
||||
"@hono/node-server": "^2.0.5",
|
||||
"fast-uri": "^3.1.5",
|
||||
|
||||
+11
-1
@@ -55,7 +55,7 @@ export function bigintToStringReplacer(_key: string, value: unknown): unknown {
|
||||
}
|
||||
|
||||
// CLI-only commands that bypass the operation layer
|
||||
export const CLI_ONLY = new Set(['init', 'reinit-pglite', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'sync', 'extract', 'extract-conversation-facts', 'enrich', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'skillpack', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'maintain', 'sources', 'mounts', 'dream', 'check-resolvable', 'routing-eval', 'skillify', 'smoke-test', 'providers', 'storage', 'repos', 'code-def', 'code-refs', 'reindex', 'reindex-code', 'reindex-frontmatter', 'code-callers', 'code-callees', 'reconcile-links', 'frontmatter', 'auth', 'friction', 'claw-test', 'book-mirror', 'takes', 'think', 'salience', 'anomalies', 'calibration', 'transcripts', 'models', 'remote', 'recall', 'forget', 'edges-backfill', 'cache', 'ze-switch', 'retrieval-upgrade', 'founder', 'brainstorm', 'lsd', 'schema', 'capture', 'onboard', 'conversation-parser', 'status', 'connect', 'skillopt', 'quarantine', 'self-upgrade', 'advisor', 'watch', 'reindex-search-vector', 'pages', 'bench', 'backfill']);
|
||||
export const CLI_ONLY = new Set(['init', 'reinit-pglite', 'pglite-repair', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'sync', 'extract', 'extract-conversation-facts', 'enrich', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'skillpack', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'maintain', 'sources', 'mounts', 'dream', 'check-resolvable', 'routing-eval', 'skillify', 'smoke-test', 'providers', 'storage', 'repos', 'code-def', 'code-refs', 'reindex', 'reindex-code', 'reindex-frontmatter', 'code-callers', 'code-callees', 'reconcile-links', 'frontmatter', 'auth', 'friction', 'claw-test', 'book-mirror', 'takes', 'think', 'salience', 'anomalies', 'calibration', 'transcripts', 'models', 'remote', 'recall', 'forget', 'edges-backfill', 'cache', 'ze-switch', 'retrieval-upgrade', 'founder', 'brainstorm', 'lsd', 'schema', 'capture', 'onboard', 'conversation-parser', 'status', 'connect', 'skillopt', 'quarantine', 'self-upgrade', 'advisor', 'watch', 'reindex-search-vector', 'pages', 'bench', 'backfill']);
|
||||
// CLI-only commands whose handlers print their own --help text. These are
|
||||
// excluded from the generic short-circuit so detailed per-command and
|
||||
// per-subcommand usage stays reachable.
|
||||
@@ -101,6 +101,9 @@ const CLI_ONLY_SELF_HELP = new Set([
|
||||
// the generic short-circuit so the destructive-action warning text
|
||||
// reaches the user.
|
||||
'reinit-pglite',
|
||||
// WAL-repair wave: pglite-repair ships its own --help with the
|
||||
// dry-run/repair semantics + the un-checkpointed-tail caveat.
|
||||
'pglite-repair',
|
||||
// v0.40.6.0 Schema Cathedral v3 — `gbrain schema --help` should hit
|
||||
// schema.ts printHelp() with the full 22+ verb taxonomy, not the
|
||||
// generic short-circuit's one-line stub.
|
||||
@@ -1315,6 +1318,13 @@ async function handleCliOnly(command: string, args: string[]) {
|
||||
await runReinitPglite(args);
|
||||
return;
|
||||
}
|
||||
// WAL-repair wave (#223/#1670/#2575): in-place torn-WAL recovery. Never
|
||||
// connects an engine — the whole point is that the DB won't open.
|
||||
if (command === 'pglite-repair') {
|
||||
const { runPgliteRepair } = await import('./commands/pglite-repair.ts');
|
||||
setCliExitVerdict(await runPgliteRepair(args));
|
||||
return;
|
||||
}
|
||||
if (command === 'auth') {
|
||||
const { runAuth } = await import('./commands/auth.ts');
|
||||
await runAuth(args);
|
||||
|
||||
@@ -4651,6 +4651,96 @@ export async function checkCycleFreshness(
|
||||
* - `progress` reporter writes to stderr (heartbeats per check)
|
||||
* - `engine.executeRaw` / handler-leaf calls (the actual probe work)
|
||||
*/
|
||||
// ≥2 failed repair attempts inside 7 days = the corruption keeps regenerating.
|
||||
const REPAIR_RECURRENCE_WINDOW_MS = 7 * 24 * 3600 * 1000;
|
||||
const REPAIR_RECURRENCE_THRESHOLD = 2;
|
||||
|
||||
/**
|
||||
* WAL-repair wave (#223/#1670/#2575): when the DB failed to connect on a
|
||||
* PGLite brain, diagnose the data dir from the FILESYSTEM (the connect error
|
||||
* itself was swallowed by doctor's fs-only fallback — this check re-derives
|
||||
* the state from disk). Pure: interprets an `inspectPgliteDataDir` diagnosis
|
||||
* into a Check; exported so `test/doctor-pglite-datadir.test.ts` drives it
|
||||
* directly (same convention as computeWorkerOomLoopCheck). Returns a Check
|
||||
* always — the call site only runs it when connect already failed, so even a
|
||||
* healthy-looking dir warrants a pointer at the repair tooling.
|
||||
*
|
||||
* Recurrence escalation (eng-review 2A): repeated failed repair attempts on
|
||||
* record mean the corruption keeps regenerating (unclean-shutdown genesis) —
|
||||
* escalate to the engine-switch ladder instead of letting the brain silently
|
||||
* lose a WAL tail per cycle. Backup-dir inventory rides along (same
|
||||
* disk-visibility class as orphan_clones).
|
||||
*/
|
||||
export function computePgliteDataDirCheck(
|
||||
dataDir: string,
|
||||
diagnosis: import('../core/pglite-repair.ts').PgliteDirDiagnosis,
|
||||
): Check {
|
||||
const backupNote = diagnosis.backupDirs.length > 0
|
||||
? ` ${diagnosis.backupDirs.length} repair backup dir(s) on disk (newest: ${diagnosis.backupDirs[0]}) — delete old ones to reclaim space once the brain is healthy.`
|
||||
: '';
|
||||
// Count BOTH outcomes (adversarial review F12): a >1h-period crash loop where
|
||||
// each repair "succeeds" discards a WAL tail per cycle with zero FAILED
|
||||
// attempts on record — escalation must still fire.
|
||||
const recentAttempts = diagnosis.recentAttempts.filter(
|
||||
(a) => Date.now() - a.ts < REPAIR_RECURRENCE_WINDOW_MS,
|
||||
).length;
|
||||
const recurrence = recentAttempts >= REPAIR_RECURRENCE_THRESHOLD
|
||||
? ` Auto-repair has run ${recentAttempts}x this week — the corruption keeps regenerating (likely an unclean-shutdown loop). Consider switching engines (docs/ENGINES.md: \`gbrain init --supabase\` or native Postgres).`
|
||||
: '';
|
||||
|
||||
switch (diagnosis.verdict) {
|
||||
case 'locked':
|
||||
return {
|
||||
name: 'pglite_data_dir',
|
||||
status: 'warn',
|
||||
message:
|
||||
`Could not connect, and the PGLite data-dir lock is held by live PID ${diagnosis.lockHolderPid} — ` +
|
||||
`another gbrain process (often \`gbrain serve\`) has the brain open. Stop it and re-run.${backupNote}`,
|
||||
remediation_status: 'human_only',
|
||||
};
|
||||
case 'missing':
|
||||
return {
|
||||
name: 'pglite_data_dir',
|
||||
status: 'warn',
|
||||
message: `No PGLite data dir at ${dataDir}. Run \`gbrain init --pglite\` to create one.`,
|
||||
remediation_status: 'human_only',
|
||||
};
|
||||
case 'unsupported-layout':
|
||||
return {
|
||||
name: 'pglite_data_dir',
|
||||
status: 'fail',
|
||||
message:
|
||||
`PGLite data dir at ${dataDir} is not repairable in place (${diagnosis.detail}). ` +
|
||||
`Rebuild from your brain repo: \`gbrain reinit-pglite\` (or back up ~/.gbrain, move the dir aside, ` +
|
||||
`\`gbrain init --pglite\`, re-add sources + sync + embed).${backupNote}${recurrence}`,
|
||||
remediation_status: 'human_only',
|
||||
};
|
||||
case 'wal-corruption-likely':
|
||||
return {
|
||||
name: 'pglite_data_dir',
|
||||
status: 'fail',
|
||||
message:
|
||||
`PGLite failed to open and the data dir shows unclean-shutdown state (${diagnosis.detail}). ` +
|
||||
`This is the torn-WAL class behind issue #223 — repairable in place, data preserved: ` +
|
||||
`\`gbrain pglite-repair --dry-run\` to diagnose, \`gbrain pglite-repair --yes\` to repair.${backupNote}${recurrence}`,
|
||||
remediation_status: 'human_only',
|
||||
};
|
||||
case 'looks-healthy':
|
||||
default:
|
||||
return {
|
||||
name: 'pglite_data_dir',
|
||||
status: 'fail',
|
||||
message:
|
||||
`PGLite failed to open but the data dir layout validates (${diagnosis.detail}). ` +
|
||||
`IF the connect error mentions \`Aborted()\` this is likely torn WAL state — ` +
|
||||
`\`gbrain pglite-repair --dry-run\` to diagnose, \`gbrain pglite-repair --yes\` to repair in place ` +
|
||||
`(repair discards the un-checkpointed WAL tail — don't run it for lock-contention or ` +
|
||||
`catalog-corruption errors; 58P01/pgvector load failures need \`gbrain reinit-pglite\` instead).${backupNote}${recurrence}`,
|
||||
remediation_status: 'human_only',
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* issue #1685 (GAP A) — the single authoritative "worker is OOM-looping" signal.
|
||||
*
|
||||
@@ -5947,6 +6037,27 @@ export async function buildChecks(
|
||||
// Filesystem read failure is non-fatal.
|
||||
}
|
||||
|
||||
// 3d. PGLite data-dir diagnosis (WAL-repair wave). Only meaningful when the
|
||||
// connect already FAILED on a PGLite brain (engine === null): the connect
|
||||
// error was swallowed by the fs-only fallback, so this check re-derives the
|
||||
// dir state from disk and names the repair ladder. Skipped under --fast
|
||||
// (connect wasn't attempted, so "engine === null" proves nothing there).
|
||||
if (!fastMode && !engine) {
|
||||
try {
|
||||
const cfg = loadConfig();
|
||||
if (cfg?.engine === 'pglite') {
|
||||
const { inspectPgliteDataDir } = await import('../core/pglite-repair.ts');
|
||||
const { resolve } = await import('node:path');
|
||||
// Absolutize: a RELATIVE database_path would make the sidecar/backup
|
||||
// lookups resolve against doctor's cwd instead of the engine's.
|
||||
const pgliteDataDir = resolve(cfg.database_path || gbrainPath('brain.pglite'));
|
||||
checks.push(computePgliteDataDirCheck(pgliteDataDir, inspectPgliteDataDir(pgliteDataDir)));
|
||||
}
|
||||
} catch {
|
||||
// Best-effort: an unreadable config or fs failure must not stop doctor.
|
||||
}
|
||||
}
|
||||
|
||||
// --- DB checks (skip if --fast or no engine) ---
|
||||
|
||||
if (fastMode || !engine) {
|
||||
|
||||
@@ -0,0 +1,340 @@
|
||||
/**
|
||||
* `gbrain pglite-repair` — diagnose and repair a torn-WAL PGLite data dir
|
||||
* in place (#223 / #1670 / #2575 recovery, the manual surface for the
|
||||
* auto-repair that `PGLiteEngine.connect()` runs on `wasm-abort` failures).
|
||||
*
|
||||
* Never connects an engine (the whole point is that the DB won't open), so it
|
||||
* works when auto-repair is disabled (GBRAIN_PGLITE_WAL_REPAIR=off) or was
|
||||
* skipped. `--dry-run` is strictly read-only. The real run:
|
||||
*
|
||||
* validate (BEFORE locking — `acquireLock` mkdirs the data dir, and a
|
||||
* typo'd --path must not create directories) → TTY confirm unless --yes →
|
||||
* acquireLock (refuses a reaped acquisition: a corrupt-lock reap cannot
|
||||
* prove the holder is dead, and WAL surgery under a possibly-live writer is
|
||||
* never correct; a live `gbrain serve` holder fast-fails via
|
||||
* LiveServeLockError) → re-validate under the lock → repair → receipt.
|
||||
*
|
||||
* There is deliberately NO --force: force-removing `.gbrain-lock` while the
|
||||
* holder is alive would reopen exactly the concurrent-writer corruption hole
|
||||
* #2348 closed. For catalog corruption (the `corrupt` classifier verdict) WAL
|
||||
* repair does not help — `gbrain reinit-pglite` is the rebuild path.
|
||||
*/
|
||||
|
||||
import { createInterface } from 'readline';
|
||||
import { loadConfig, gbrainPath } from '../core/config.ts';
|
||||
import { acquireLock, releaseLock, LiveServeLockError, msSinceLastReap } from '../core/pglite-lock.ts';
|
||||
import {
|
||||
inspectPgliteDataDir,
|
||||
listRepairBackups,
|
||||
readRepairSidecar,
|
||||
recordRepairAttempt,
|
||||
repairPgliteWal,
|
||||
validateWalRepairTarget,
|
||||
WalRepairError,
|
||||
} from '../core/pglite-repair.ts';
|
||||
|
||||
interface RepairCmdOpts {
|
||||
dryRun: boolean;
|
||||
yes: boolean;
|
||||
jsonOutput: boolean;
|
||||
customPath: string | null;
|
||||
help: boolean;
|
||||
}
|
||||
|
||||
class UnknownFlagError extends Error {}
|
||||
|
||||
function parseArgs(args: string[]): RepairCmdOpts {
|
||||
const opts: RepairCmdOpts = { dryRun: false, yes: false, jsonOutput: false, customPath: null, help: false };
|
||||
for (let i = 0; i < args.length; i++) {
|
||||
const a = args[i];
|
||||
if (a === '--dry-run') opts.dryRun = true;
|
||||
else if (a === '--yes' || a === '-y') opts.yes = true;
|
||||
else if (a === '--json') opts.jsonOutput = true;
|
||||
else if (a === '--path') {
|
||||
const val = args[++i];
|
||||
// F1: a valueless --path (typo / shell mangling) must NOT fall through
|
||||
// to the configured default brain and run surgery on the wrong dir.
|
||||
if (val === undefined || val.startsWith('-')) throw new UnknownFlagError('--path requires a directory argument');
|
||||
opts.customPath = val;
|
||||
}
|
||||
else if (a === '--help' || a === '-h') opts.help = true;
|
||||
// Reject unknown args on a DESTRUCTIVE command (codex): silently ignoring
|
||||
// a typo like `--dry-rnu` would run a real WAL reset instead of a dry run.
|
||||
else throw new UnknownFlagError(`unknown argument: ${a}`);
|
||||
}
|
||||
return opts;
|
||||
}
|
||||
|
||||
function printHelp(): void {
|
||||
console.log(`gbrain pglite-repair — repair a torn-WAL PGLite data dir in place
|
||||
|
||||
The default gbrain engine (PGLite) can fail to open after an unclean shutdown
|
||||
(commonly a macOS-upgrade reboot) with "RuntimeError: Aborted()". The cause is
|
||||
torn WAL/checkpoint state on disk, not a macOS WASM bug. This command resets
|
||||
the WAL in place (pg_resetwal semantics): data files are preserved;
|
||||
transactions not checkpointed before the corruption may be lost. The
|
||||
pre-repair pg_wal + pg_control are kept in a sibling backup directory.
|
||||
|
||||
Usage:
|
||||
gbrain pglite-repair --dry-run [--json] [--path <dir>] diagnose only
|
||||
gbrain pglite-repair [--yes] [--json] [--path <dir>] repair (confirm on TTY)
|
||||
|
||||
Flags:
|
||||
--dry-run Read-only diagnosis of the data dir. Mutates nothing.
|
||||
--yes, -y Skip the confirmation prompt (required in non-TTY runs).
|
||||
--json Machine-readable output on stdout.
|
||||
--path <dir> Repair a specific data dir (default: the configured brain).
|
||||
|
||||
Notes:
|
||||
Auto-repair runs on ordinary commands by default; disable it with
|
||||
GBRAIN_PGLITE_WAL_REPAIR=off and use this command deliberately.
|
||||
Catalog corruption (58P01 / pgvector load failure) is NOT repairable in
|
||||
place — use \`gbrain reinit-pglite\` for that class.`);
|
||||
}
|
||||
|
||||
function emitError(jsonOutput: boolean, code: string, message: string): void {
|
||||
if (jsonOutput) {
|
||||
console.log(JSON.stringify({ status: 'error', code, message }));
|
||||
} else {
|
||||
console.error(`Error (${code}): ${message}`);
|
||||
}
|
||||
}
|
||||
|
||||
async function promptYesNo(question: string): Promise<boolean> {
|
||||
// Prompt on stderr: stdout stays clean for --json payloads.
|
||||
const rl = createInterface({ input: process.stdin, output: process.stderr });
|
||||
return new Promise((resolve) => {
|
||||
rl.question(`${question} [y/N] `, (answer) => {
|
||||
rl.close();
|
||||
resolve(/^y(es)?$/i.test(answer.trim()));
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
export async function runPgliteRepair(args: string[]): Promise<number> {
|
||||
let opts: RepairCmdOpts;
|
||||
try {
|
||||
opts = parseArgs(args);
|
||||
} catch (err) {
|
||||
if (err instanceof UnknownFlagError) {
|
||||
const jsonOut = args.includes('--json');
|
||||
emitError(jsonOut, 'unknown_flag', `${err.message}. Run \`gbrain pglite-repair --help\`.`);
|
||||
return 2;
|
||||
}
|
||||
throw err;
|
||||
}
|
||||
if (opts.help) {
|
||||
printHelp();
|
||||
return 0;
|
||||
}
|
||||
|
||||
// Resolve dir: --path > config > default brain path. With an explicit
|
||||
// --path we skip the engine check (repairing an arbitrary dir is the point).
|
||||
let dataDir = opts.customPath;
|
||||
if (!dataDir) {
|
||||
const cfg = loadConfig();
|
||||
if (cfg?.engine !== 'pglite') {
|
||||
emitError(
|
||||
opts.jsonOutput,
|
||||
'not_pglite',
|
||||
`gbrain pglite-repair is for PGLite brains (current engine: ${cfg?.engine || 'none'}). ` +
|
||||
'Pass --path <dir> to repair a specific data dir.',
|
||||
);
|
||||
return 1;
|
||||
}
|
||||
dataDir = cfg.database_path || gbrainPath('brain.pglite');
|
||||
}
|
||||
|
||||
// Read-only diagnosis first — BEFORE any lock (acquireLock mkdirs the data
|
||||
// dir; a typo'd --path must produce a clean refusal with zero side effects).
|
||||
const validation = validateWalRepairTarget(dataDir);
|
||||
const diagnosis = inspectPgliteDataDir(dataDir);
|
||||
|
||||
if (opts.dryRun) {
|
||||
if (opts.jsonOutput) {
|
||||
console.log(JSON.stringify({
|
||||
status: 'ok',
|
||||
action: 'dry-run',
|
||||
data_dir: dataDir,
|
||||
validation,
|
||||
diagnosis,
|
||||
}));
|
||||
} else {
|
||||
console.log(`PGLite data dir: ${dataDir}`);
|
||||
console.log(` Verdict: ${diagnosis.verdict} — ${diagnosis.detail}`);
|
||||
console.log(` PG_VERSION: ${diagnosis.pgVersion ?? '(unreadable)'} pg_control: ${diagnosis.pgControlOk ? 'ok (8192 bytes)' : 'BAD'}`);
|
||||
console.log(` WAL segments: ${diagnosis.walSegments.length} stale postmaster.pid: ${diagnosis.postmasterPid ? 'YES' : 'no'}`);
|
||||
console.log(` Lock: ${diagnosis.lockHeld ? `HELD by live PID ${diagnosis.lockHolderPid}` : 'free'}`);
|
||||
if (diagnosis.recentAttempts.length > 0) {
|
||||
console.log(` Repair attempts on record: ${diagnosis.recentAttempts.map((a) => `${a.outcome}@${new Date(a.ts).toISOString()}`).join(', ')}`);
|
||||
}
|
||||
if (diagnosis.backupDirs.length > 0) {
|
||||
console.log(` Repair backups on disk: ${diagnosis.backupDirs.join(', ')}`);
|
||||
}
|
||||
if (!validation.ok) {
|
||||
console.log(` Repairable: NO — ${validation.detail}`);
|
||||
} else if (diagnosis.verdict === 'looks-healthy') {
|
||||
console.log(' Repairable: yes — but no unclean-shutdown markers found; repair is likely');
|
||||
console.log(' unnecessary. Run `gbrain pglite-repair --yes` ONLY if PGLite fails to open');
|
||||
console.log(' with `RuntimeError: Aborted()` (repair discards the un-checkpointed WAL tail).');
|
||||
} else {
|
||||
console.log(' Repairable: yes — run `gbrain pglite-repair --yes` to reset the WAL in place.');
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
if (!validation.ok) {
|
||||
emitError(opts.jsonOutput, `refused_${validation.reason}`, validation.detail);
|
||||
return 1;
|
||||
}
|
||||
if (diagnosis.lockHeld) {
|
||||
emitError(
|
||||
opts.jsonOutput,
|
||||
'refused_locked',
|
||||
`another gbrain process (PID ${diagnosis.lockHolderPid}) is using this brain — stop it, then re-run.`,
|
||||
);
|
||||
return 1;
|
||||
}
|
||||
const sinceReap = msSinceLastReap(dataDir);
|
||||
const REAP_QUARANTINE_MS = 10 * 60 * 1000;
|
||||
if (sinceReap !== null && sinceReap >= 0 && sinceReap < REAP_QUARANTINE_MS) {
|
||||
emitError(
|
||||
opts.jsonOutput,
|
||||
'refused_reap_quarantine',
|
||||
`a lock on this brain was reaped ${Math.round(sinceReap / 1000)}s ago from a holder whose ` +
|
||||
'liveness could not be verified — that process may still be writing. Confirm no gbrain ' +
|
||||
`process is running (\`pgrep -af gbrain\`), wait ${Math.ceil((REAP_QUARANTINE_MS - sinceReap) / 60000)} more minute(s), then re-run.`,
|
||||
);
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (!opts.yes) {
|
||||
if (!process.stdin.isTTY) {
|
||||
emitError(opts.jsonOutput, 'no_tty_no_yes', 'Non-TTY environment requires --yes to confirm the WAL reset.');
|
||||
return 1;
|
||||
}
|
||||
console.error(`About to reset the WAL of ${dataDir} in place.`);
|
||||
console.error('Data files are preserved; un-checkpointed transactions may be lost.');
|
||||
console.error('The current pg_wal + pg_control are kept in a sibling backup directory.');
|
||||
const confirmed = await promptYesNo('Repair now?');
|
||||
if (!confirmed) {
|
||||
if (opts.jsonOutput) {
|
||||
console.log(JSON.stringify({ status: 'aborted', reason: 'user_declined' }));
|
||||
} else {
|
||||
console.log('Aborted. Data dir untouched.');
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
|
||||
// Short timeout: the diagnosis said the lock is free; if we still can't get
|
||||
// it quickly, someone raced us — refuse rather than queue behind them.
|
||||
let lock;
|
||||
try {
|
||||
lock = await acquireLock(dataDir, { timeoutMs: 5_000 });
|
||||
} catch (err) {
|
||||
if (err instanceof LiveServeLockError) {
|
||||
emitError(
|
||||
opts.jsonOutput,
|
||||
'refused_live_serve',
|
||||
`a live \`gbrain serve\` (MCP) process holds this brain — stop \`gbrain serve\` first, then re-run. ${String((err as Error).message)}`,
|
||||
);
|
||||
return 1;
|
||||
}
|
||||
emitError(opts.jsonOutput, 'refused_lock_timeout', String((err as Error)?.message ?? err));
|
||||
return 1;
|
||||
}
|
||||
try {
|
||||
if (!lock.acquired) {
|
||||
emitError(opts.jsonOutput, 'refused_locked', 'could not acquire the PGLite data-dir lock — another gbrain process is using this brain.');
|
||||
return 1;
|
||||
}
|
||||
if (lock.reaped) {
|
||||
// A reaped acquisition (dead-PID or corrupt-lock-file reap) cannot prove
|
||||
// the prior holder is gone. WAL surgery under a possibly-live writer is
|
||||
// never correct — no --force by design.
|
||||
emitError(
|
||||
opts.jsonOutput,
|
||||
'refused_reaped_lock',
|
||||
'the data-dir lock was acquired by reaping a prior holder’s lock — ' +
|
||||
'another gbrain process may still be using this brain. Confirm no gbrain ' +
|
||||
'process is running, then re-run (a cleanly-acquired lock enables repair).',
|
||||
);
|
||||
return 1;
|
||||
}
|
||||
|
||||
// Cheap re-validate under the lock (TOCTOU window between diagnosis and
|
||||
// lock acquisition).
|
||||
const revalidation = validateWalRepairTarget(dataDir);
|
||||
if (!revalidation.ok) {
|
||||
emitError(opts.jsonOutput, `refused_${revalidation.reason}`, revalidation.detail);
|
||||
return 1;
|
||||
}
|
||||
|
||||
process.stderr.write(`Repairing WAL of ${dataDir} in place…\n`);
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
// F4: only reuse a FRESH (<24h) episode backup — a stale pin may predate
|
||||
// real data (same bound as the auto seam's episodeFresh).
|
||||
const episodeFresh =
|
||||
sidecar.episodeStartedAt !== null &&
|
||||
Date.now() - sidecar.episodeStartedAt >= 0 &&
|
||||
Date.now() - sidecar.episodeStartedAt < 24 * 3600 * 1000;
|
||||
let receipt;
|
||||
try {
|
||||
receipt = await repairPgliteWal(dataDir, {
|
||||
reuseBackupPath: episodeFresh ? sidecar.episodeBackupPath ?? undefined : undefined,
|
||||
});
|
||||
} catch (err) {
|
||||
if (err instanceof WalRepairError) {
|
||||
// Reset failed after the backup was taken — report the restore's REAL
|
||||
// outcome so the user knows whether the dir is back or in reset state.
|
||||
recordRepairAttempt(dataDir, 'failed', err.receipt.backupPath);
|
||||
emitError(
|
||||
opts.jsonOutput,
|
||||
'repair_failed',
|
||||
err.message + (err.restore.restored
|
||||
? ` (data dir restored to its pre-repair state; backup kept at ${err.receipt.backupPath})`
|
||||
: ` (RESTORE ALSO FAILED — the dir is in a reset state; your pre-repair files are intact at ${err.receipt.backupPath}: ${err.restore.detail ?? ''})`),
|
||||
);
|
||||
return 1;
|
||||
}
|
||||
recordRepairAttempt(dataDir, 'failed', sidecar.episodeBackupPath);
|
||||
emitError(opts.jsonOutput, 'repair_failed', String((err as Error)?.message ?? err));
|
||||
return 1;
|
||||
}
|
||||
// "repaired" here means the reset completed; the next connect PROVES it.
|
||||
// Record a FAILED attempt (not repaired-with-closeEpisode:false, which
|
||||
// leaves the episode null when none was open — codex): this opens/keeps an
|
||||
// episode pinned to this backup so a later healthy connect closes it and
|
||||
// prunes, and repeated manual runs during one incident reuse the pinned
|
||||
// backup instead of deleting the pre-damage forensic copy.
|
||||
recordRepairAttempt(dataDir, 'failed', receipt.backupPath);
|
||||
|
||||
if (opts.jsonOutput) {
|
||||
console.log(JSON.stringify({
|
||||
status: 'ok',
|
||||
action: 'repaired',
|
||||
data_dir: receipt.dataDir,
|
||||
backup_path: receipt.backupPath,
|
||||
backed_up: receipt.backedUpFiles,
|
||||
reused_episode_backup: receipt.reusedEpisodeBackup,
|
||||
reset_segment: receipt.resetSegment,
|
||||
timeline_id: receipt.timelineId,
|
||||
wal_seg_size: receipt.walSegSize,
|
||||
repaired_at: receipt.repairedAt,
|
||||
backups_on_disk: listRepairBackups(dataDir),
|
||||
}));
|
||||
} else {
|
||||
console.log('WAL reset complete.');
|
||||
console.log(` Data dir: ${receipt.dataDir}`);
|
||||
console.log(` Backup: ${receipt.backupPath}${receipt.reusedEpisodeBackup ? ' (reused this episode’s existing backup)' : ''}`);
|
||||
console.log(` Reset segment: ${receipt.resetSegment} (timeline ${receipt.timelineId}, ${receipt.walSegSize / (1024 * 1024)}MB segments)`);
|
||||
console.log(' Data files were preserved; un-checkpointed transactions may be lost.');
|
||||
console.log(' Next: run any gbrain command to reopen the brain, then `gbrain doctor`.');
|
||||
}
|
||||
return 0;
|
||||
} finally {
|
||||
await releaseLock(lock);
|
||||
}
|
||||
}
|
||||
@@ -18,7 +18,7 @@
|
||||
* output via `--json` for scripted callers.
|
||||
*/
|
||||
|
||||
import { existsSync, renameSync, statSync } from 'fs';
|
||||
import { existsSync, renameSync, statSync, rmSync } from 'fs';
|
||||
import { dirname } from 'path';
|
||||
import { loadConfig, loadConfigFileOnly, gbrainPath } from '../core/config.ts';
|
||||
|
||||
@@ -122,6 +122,13 @@ export async function runReinitPglite(args: string[]): Promise<void> {
|
||||
|
||||
try {
|
||||
renameSync(dbPath, bakPath);
|
||||
// WAL-repair state travels with the OLD brain (red-team: a fresh brain at
|
||||
// the same path must not inherit the old brain's open repair episode,
|
||||
// cooldown, or reap quarantine — a stale episodeBackupPath would be reused
|
||||
// over the NEW brain's WAL).
|
||||
for (const sibling of [`${dbPath}.wal-repair-attempt.json`, `${dbPath}.lock-reap.json`]) {
|
||||
try { rmSync(sibling, { force: true }); } catch { /* best-effort */ }
|
||||
}
|
||||
} catch (e: unknown) {
|
||||
fail(
|
||||
opts.jsonOutput,
|
||||
@@ -210,21 +217,59 @@ function parseArgs(args: string[]): ReinitOpts {
|
||||
const dimsIdx = args.indexOf('--embedding-dimensions');
|
||||
const pathIdx = args.indexOf('--path');
|
||||
|
||||
if (modelIdx < 0 || modelIdx === args.length - 1) {
|
||||
fail(jsonOutput, 'missing_model', '--embedding-model <provider:model> is required.');
|
||||
// Default omitted flags from the config FILE. Deliberately
|
||||
// `loadConfigFileOnly()`, NOT `loadConfig()`: loadConfig merges the
|
||||
// GBRAIN_EMBEDDING_MODEL / GBRAIN_EMBEDDING_DIMENSIONS env overrides,
|
||||
// and a transient outage-shell export must not silently change the
|
||||
// rebuild target. Precedence: explicit flag > config-file value >
|
||||
// hard-fail (the original missing_model/missing_dims errors).
|
||||
const fileCfg = (modelIdx < 0 || dimsIdx < 0) ? loadConfigFileOnly() : null;
|
||||
|
||||
let embeddingModel: string;
|
||||
if (modelIdx >= 0) {
|
||||
if (modelIdx === args.length - 1) {
|
||||
fail(jsonOutput, 'missing_model', '--embedding-model <provider:model> is required.');
|
||||
}
|
||||
embeddingModel = args[modelIdx + 1];
|
||||
} else if (fileCfg?.embedding_model) {
|
||||
embeddingModel = fileCfg.embedding_model;
|
||||
console.error(`--embedding-model defaulted from config: ${embeddingModel}`);
|
||||
} else {
|
||||
fail(
|
||||
jsonOutput,
|
||||
'missing_model',
|
||||
'--embedding-model <provider:model> is required (no embedding_model in the config file to default from).',
|
||||
);
|
||||
}
|
||||
if (dimsIdx < 0 || dimsIdx === args.length - 1) {
|
||||
fail(jsonOutput, 'missing_dims', '--embedding-dimensions <N> is required.');
|
||||
|
||||
let dimsStr: string;
|
||||
let dimsFromConfig = false;
|
||||
if (dimsIdx >= 0) {
|
||||
if (dimsIdx === args.length - 1) {
|
||||
fail(jsonOutput, 'missing_dims', '--embedding-dimensions <N> is required.');
|
||||
}
|
||||
dimsStr = args[dimsIdx + 1];
|
||||
} else if (fileCfg?.embedding_dimensions !== undefined && fileCfg?.embedding_dimensions !== null) {
|
||||
dimsStr = String(fileCfg.embedding_dimensions);
|
||||
dimsFromConfig = true;
|
||||
} else {
|
||||
fail(
|
||||
jsonOutput,
|
||||
'missing_dims',
|
||||
'--embedding-dimensions <N> is required (no embedding_dimensions in the config file to default from).',
|
||||
);
|
||||
}
|
||||
|
||||
const dimsStr = args[dimsIdx + 1];
|
||||
const dims = parseInt(dimsStr, 10);
|
||||
if (!Number.isInteger(dims) || dims <= 0) {
|
||||
fail(jsonOutput, 'invalid_dims', `--embedding-dimensions must be a positive integer (got: ${dimsStr}).`);
|
||||
}
|
||||
if (dimsFromConfig) {
|
||||
console.error(`--embedding-dimensions defaulted from config: ${dims}`);
|
||||
}
|
||||
|
||||
return {
|
||||
embeddingModel: args[modelIdx + 1],
|
||||
embeddingModel,
|
||||
embeddingDimensions: dims,
|
||||
yes,
|
||||
jsonOutput,
|
||||
@@ -240,9 +285,16 @@ Wipe the PGLite brain and re-init with new embedding model/dimensions.
|
||||
This is the canonical path for switching embedding providers on PGLite
|
||||
because pgvector (WASM) cannot ALTER vector column types in place.
|
||||
|
||||
Required:
|
||||
Embedding target (each defaults from the config file when omitted):
|
||||
--embedding-model <provider:model> New embedding model (e.g. openai:text-embedding-3-large).
|
||||
Defaults to embedding_model in ~/.gbrain/config.json.
|
||||
--embedding-dimensions <N> New dimension count (e.g. 1280, 1536, 2048).
|
||||
Defaults to embedding_dimensions in ~/.gbrain/config.json.
|
||||
|
||||
Defaults read the config FILE only; GBRAIN_EMBEDDING_MODEL /
|
||||
GBRAIN_EMBEDDING_DIMENSIONS env overrides are deliberately ignored so a
|
||||
transient shell export cannot change the rebuild target. If neither the
|
||||
flag nor the config file provides a value, the command fails.
|
||||
|
||||
Optional:
|
||||
--path <path> Active brain path (default: ~/.gbrain/brain.pglite).
|
||||
@@ -258,6 +310,9 @@ Examples:
|
||||
gbrain reinit-pglite --embedding-model openai:text-embedding-3-large \\
|
||||
--embedding-dimensions 1536 --no-sync
|
||||
|
||||
# Rebuild with the model/dimensions already in the config file:
|
||||
gbrain reinit-pglite --yes
|
||||
|
||||
The old brain is preserved as \`<path>.bak\`. To roll back, mv it back.
|
||||
|
||||
See also:
|
||||
|
||||
@@ -153,6 +153,7 @@ export const OPS_CHECK_NAMES: ReadonlySet<string> = new Set([
|
||||
'oauth_confidential_client_health',
|
||||
'orphan_clones',
|
||||
'pgbouncer_prepare',
|
||||
'pglite_data_dir',
|
||||
'pgvector',
|
||||
'pool_budget',
|
||||
'progressive_batch_audit_health',
|
||||
|
||||
+202
-18
@@ -46,6 +46,9 @@ import { DELETE_BATCH_SIZE } from './engine-constants.ts';
|
||||
import { SOURCE_CONFIG_OBJECT_SQL } from './source-config-sql.ts';
|
||||
import { MARKDOWN_CHUNKER_VERSION } from './chunkers/recursive.ts';
|
||||
import { acquireLock, releaseLock, type LockHandle } from './pglite-lock.ts';
|
||||
// Engine-live path (#3596): static import, never a lazy `import()` in the
|
||||
// connect() catch. No cycle: pglite-repair.ts imports nothing from this file.
|
||||
import { attemptWalRepairAndRetry, closeRepairEpisodeIfOpen, type WalRepairReceipt } from './pglite-repair.ts';
|
||||
import { getFtsLanguage } from './fts-language.ts';
|
||||
import type {
|
||||
Page, PageInput, PageFilters, PageType,
|
||||
@@ -173,23 +176,49 @@ export function computeSnapshotSchemaHash(
|
||||
* payload. Fix: `bun upgrade` (newer Bun versions mount the vfs
|
||||
* writable) or run via Node.
|
||||
*
|
||||
* `macos-26-3` — the pre-existing #223 hint signature (early macOS
|
||||
* 26.3 builds shipped a broken WASM runtime).
|
||||
* `corrupt` — catalog/pgvector corruption (#2348): 58P01 /
|
||||
* internal_load_library / missing vector type or core relation.
|
||||
* WAL reset cannot fix this class; routes to `gbrain reinit-pglite`.
|
||||
* MUST stay matched BEFORE the wasm arm — a `58P01 … Aborted()`
|
||||
* message is catalog corruption, not a WAL tear.
|
||||
*
|
||||
* `wasm-abort` — the Emscripten runtime abort (`Aborted(). Build with
|
||||
* -sASSERTIONS…`, `RuntimeError: unreachable`, and the legacy #223
|
||||
* signatures). Root cause is almost always corrupt WAL/checkpoint
|
||||
* state after an unclean shutdown (historically misdiagnosed as a
|
||||
* "macOS 26.3 WASM bug" — see #223); this verdict is the trigger
|
||||
* for the in-place WAL auto-repair (`pglite-repair.ts`).
|
||||
*
|
||||
* `unknown` — falls through to a generic hint that names the doctor
|
||||
* command; the macOS 26.3 link is offered only on darwin (#2674).
|
||||
* command; the #223 pointer is offered only on darwin (#2674).
|
||||
*
|
||||
* Regex tightened per Codex eng-review finding #9: don't match
|
||||
* generic `pglite.data` substring (could fire on unrelated PGLite
|
||||
* errors). Match the literal `$$bunfs` marker OR ENOENT+pglite.data
|
||||
* co-occurrence.
|
||||
*/
|
||||
export type PgliteInitFailure = 'bunfs' | 'macos-26-3' | 'corrupt' | 'unknown';
|
||||
export type PgliteInitFailure = 'bunfs' | 'wasm-abort' | 'corrupt' | 'unknown';
|
||||
|
||||
// #2674: non-Error rejections (Emscripten aborts can throw plain objects)
|
||||
// used to stringify as "[object Object]" — prefer .message when present.
|
||||
// WAL-repair wave: Emscripten's FS layer also throws message-LESS objects
|
||||
// (e.g. `ErrnoError { name: 'ErrnoError', errno: 20 }` when the data dir is a
|
||||
// symlink NODEFS refuses to mount) — surface name+errno / JSON instead of the
|
||||
// useless "[object Object]".
|
||||
export function stringifyPgliteInitError(err: unknown): string {
|
||||
return String((err as { message?: unknown })?.message ?? err);
|
||||
const message = (err as { message?: unknown })?.message;
|
||||
if (message != null) return String(message);
|
||||
if (typeof err === 'object' && err !== null) {
|
||||
const name = (err as { name?: unknown }).name;
|
||||
const errno = (err as { errno?: unknown }).errno;
|
||||
if (typeof name === 'string' && errno != null) return `${name} (errno ${errno})`;
|
||||
try {
|
||||
const json = JSON.stringify(err);
|
||||
if (json && json !== '{}') return typeof name === 'string' ? `${name}: ${json}` : json;
|
||||
} catch { /* circular — fall through */ }
|
||||
if (typeof name === 'string') return name;
|
||||
}
|
||||
return String(err);
|
||||
}
|
||||
|
||||
export function classifyPgliteInitError(message: string): PgliteInitFailure {
|
||||
@@ -202,18 +231,79 @@ export function classifyPgliteInitError(message: string): PgliteInitFailure {
|
||||
if (/58P01|internal_load_library|type "?vector"? does not exist|relation "?content_chunks"? does not exist/i.test(message)) {
|
||||
return 'corrupt';
|
||||
}
|
||||
if (/abort.*runtime|macos.*26\.3|wasm.*runtime/i.test(message)) {
|
||||
return 'macos-26-3';
|
||||
// Broadened (v0.42.x WAL-repair wave): the REAL production message is
|
||||
// `Aborted(). Build with -sASSERTIONS for more info.` — no "runtime" in it,
|
||||
// so the legacy arms alone let the primary crash fall through to 'unknown'.
|
||||
// Deliberately over-matches (RuntimeError/unreachable are generic WASM
|
||||
// traps); the repair path downstream is bounded by layout validation, the
|
||||
// reaped-lock gate, and restore-on-failure.
|
||||
if (/aborted\s*\(\)|RuntimeError|unreachable|abort.*runtime|macos.*26\.3|wasm.*runtime/i.test(message)) {
|
||||
return 'wasm-abort';
|
||||
}
|
||||
return 'unknown';
|
||||
}
|
||||
|
||||
/**
|
||||
* What the auto-repair path did (or why it didn't run) for a `wasm-abort`
|
||||
* failure — folded into the user-facing error so the message never lies about
|
||||
* the state of the data dir. `'failed-not-restored'` is the arm that matters
|
||||
* most: repair ran, PGLite still failed, AND the automatic restore failed —
|
||||
* the dir is in a reset state and the user must restore from the backup.
|
||||
*/
|
||||
export interface PgliteInitRepairContext {
|
||||
repair:
|
||||
| 'not-attempted'
|
||||
| 'in-memory'
|
||||
| 'disabled'
|
||||
| 'skipped-validation'
|
||||
| 'skipped-live-writer'
|
||||
| 'skipped-cooldown'
|
||||
| 'failed-restored'
|
||||
| 'failed-not-restored';
|
||||
backupPath?: string;
|
||||
detail?: string;
|
||||
}
|
||||
|
||||
function repairContextLine(ctx: PgliteInitRepairContext): string {
|
||||
switch (ctx.repair) {
|
||||
case 'in-memory':
|
||||
return ' This engine is in-memory (no data dir), so there is no stored state to\n' +
|
||||
' repair — this is an environment/runtime failure, not data corruption.';
|
||||
case 'disabled':
|
||||
return ' Auto-repair is disabled (GBRAIN_PGLITE_WAL_REPAIR=off). Run\n' +
|
||||
' `gbrain pglite-repair` to repair manually.';
|
||||
case 'skipped-validation':
|
||||
return ` Auto-repair skipped: ${ctx.detail ?? 'the data dir did not validate as a PG17 pglite layout'}.`;
|
||||
case 'skipped-live-writer':
|
||||
return ` Auto-repair skipped: ${ctx.detail ?? 'the data-dir lock was acquired by reaping a prior holder'}`;
|
||||
case 'skipped-cooldown':
|
||||
return ` Auto-repair skipped: ${ctx.detail ?? 'a recent attempt failed (cooldown active)'}`;
|
||||
case 'failed-restored':
|
||||
return ' Auto-repair ran but PGLite still failed to start. The data dir was\n' +
|
||||
` RESTORED to its pre-repair state (backup kept at ${ctx.backupPath ?? '<dataDir>.wal-repair-backup-*'}).` +
|
||||
(ctx.detail ? `\n Detail: ${ctx.detail}` : '');
|
||||
case 'failed-not-restored':
|
||||
return ' Auto-repair ran, PGLite still failed to start, AND the automatic restore\n' +
|
||||
' itself failed — the data dir is currently in a RESET state. Your\n' +
|
||||
` pre-repair files are intact in the backup at ${ctx.backupPath ?? '<dataDir>.wal-repair-backup-*'};\n` +
|
||||
' restore manually: move the backup\'s `pg_wal` dir back to `<dataDir>/pg_wal`\n' +
|
||||
' and its `pg_control` file back to `<dataDir>/global/pg_control`.' +
|
||||
(ctx.detail ? `\n Detail: ${ctx.detail}` : '');
|
||||
case 'not-attempted':
|
||||
default:
|
||||
return ' Auto-repair was not attempted.';
|
||||
}
|
||||
}
|
||||
|
||||
export function buildPgliteInitErrorMessage(
|
||||
verdict: PgliteInitFailure,
|
||||
original: string,
|
||||
// #2674: threaded (defaulted) so tests can exercise both branches without
|
||||
// monkey-patching process.platform.
|
||||
platform: NodeJS.Platform = process.platform,
|
||||
// WAL-repair wave: what auto-repair did for a wasm-abort, so the hint tells
|
||||
// the truth about the current state of the data dir.
|
||||
ctx?: PgliteInitRepairContext,
|
||||
): string {
|
||||
const header = 'PGLite failed to initialize its WASM runtime.';
|
||||
let hint: string;
|
||||
@@ -226,17 +316,31 @@ export function buildPgliteInitErrorMessage(
|
||||
' does not help, run via Node: `node src/cli.ts` or install gbrain\n' +
|
||||
' using the Node-based path. See #1340 for details.';
|
||||
break;
|
||||
case 'macos-26-3':
|
||||
case 'wasm-abort':
|
||||
hint =
|
||||
' This is most commonly the macOS 26.3 WASM bug:\n' +
|
||||
' https://github.com/garrytan/gbrain/issues/223';
|
||||
' Most common cause: corrupt WAL/checkpoint state after an unclean\n' +
|
||||
' shutdown (often a macOS-upgrade reboot killing gbrain mid-write) —\n' +
|
||||
' NOT a macOS WASM bug, despite the historical diagnosis in\n' +
|
||||
' https://github.com/garrytan/gbrain/issues/223.\n' +
|
||||
repairContextLine(ctx ?? { repair: 'not-attempted' }) + '\n' +
|
||||
' Recovery ladder:\n' +
|
||||
' 1. gbrain pglite-repair --dry-run (diagnose, mutates nothing)\n' +
|
||||
' gbrain pglite-repair --yes (in-place WAL repair, data preserved)\n' +
|
||||
' 2. Rebuild from your brain repo: `gbrain reinit-pglite` (or manually:\n' +
|
||||
' back up ~/.gbrain, move brain.pglite aside, `gbrain init --pglite`,\n' +
|
||||
' re-add sources + `gbrain sync` + `gbrain embed`).\n' +
|
||||
' 3. Switch engines (docs/ENGINES.md): `gbrain init --supabase` or\n' +
|
||||
' native Postgres.\n' +
|
||||
' Run `gbrain doctor` for a full diagnosis.';
|
||||
break;
|
||||
case 'corrupt':
|
||||
hint =
|
||||
' Your PGLite store looks corrupted (the catalog or the pgvector\n' +
|
||||
' extension cannot load). This happens when two processes opened the\n' +
|
||||
' same brain at once — now prevented (#2348), but an already-damaged\n' +
|
||||
' store cannot be repaired in place. Recover:\n' +
|
||||
' store cannot be repaired in place (WAL repair does not fix catalog\n' +
|
||||
' corruption; `gbrain pglite-repair --dry-run` can still report the\n' +
|
||||
' state of the data dir). Recover:\n' +
|
||||
' 1. Restore a backup of the brain.pglite directory if you have one, OR\n' +
|
||||
' 2. Rebuild from your brain repo:\n' +
|
||||
' gbrain reinit-pglite --embedding-model <id> --embedding-dimensions <N>\n' +
|
||||
@@ -245,21 +349,42 @@ export function buildPgliteInitErrorMessage(
|
||||
break;
|
||||
case 'unknown':
|
||||
default:
|
||||
// #2674: only blame the macOS 26.3 WASM bug on macOS. On other
|
||||
// platforms, point at the causes that are actually plausible there.
|
||||
// #2674: name the plausible causes per platform. The darwin branch keeps
|
||||
// the #223 pointer (readers arrive from that issue), reframed to the
|
||||
// real root cause behind those reports: torn WAL from unclean shutdown.
|
||||
hint = platform === 'darwin'
|
||||
? ' Possible cause: the macOS 26.3 WASM bug\n' +
|
||||
' (https://github.com/garrytan/gbrain/issues/223).\n' +
|
||||
' Run `gbrain doctor` for a full diagnosis.'
|
||||
? ' Possible cause: corrupt WAL/checkpoint state after an unclean\n' +
|
||||
' shutdown — the failure class behind\n' +
|
||||
' https://github.com/garrytan/gbrain/issues/223.\n' +
|
||||
' Try `gbrain pglite-repair --dry-run` to diagnose the data dir, and\n' +
|
||||
' run `gbrain doctor` for a full diagnosis.'
|
||||
: ' Possible causes: another gbrain process holding the database\n' +
|
||||
' (lock contention), or a damaged PGLite data directory.\n' +
|
||||
' Run `gbrain doctor` for a full diagnosis; if the data dir is\n' +
|
||||
' Try `gbrain pglite-repair --dry-run` to diagnose the data dir, and\n' +
|
||||
' run `gbrain doctor` for a full diagnosis; if the data dir is\n' +
|
||||
' damaged, `gbrain reinit-pglite` rebuilds it from your brain repo.';
|
||||
break;
|
||||
}
|
||||
return `${header}\n${hint}\n Original error: ${original}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* The loud stderr notice printed when connect() auto-repaired the data dir in
|
||||
* place. Exported for the serial regression test.
|
||||
*/
|
||||
export function buildWalRepairNotice(receipt: WalRepairReceipt): string {
|
||||
return [
|
||||
'⚠️ gbrain repaired this brain\'s PGLite WAL in place.',
|
||||
` Data dir: ${receipt.dataDir}`,
|
||||
` Cause: torn WAL/checkpoint state from an unclean shutdown (issue #223 class).`,
|
||||
` Data files were preserved; transactions not checkpointed before the`,
|
||||
` corruption may be lost (the standard pg_resetwal caveat).`,
|
||||
` Pre-repair backup: ${receipt.backupPath}`,
|
||||
` Recommended: run \`gbrain doctor\` to verify brain integrity.`,
|
||||
` Disable auto-repair with GBRAIN_PGLITE_WAL_REPAIR=off.`,
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
/**
|
||||
* #2084 — PGLite's Emscripten runtime hijacks `process.exitCode` as ITS status
|
||||
* channel: instantiation REPLACES the property with an accessor whose getter
|
||||
@@ -297,6 +422,12 @@ export class PGLiteEngine implements BrainEngine {
|
||||
// PGlite.create(loadDataDir), initSchema is a no-op (schema is already
|
||||
// present + migrations already applied). Saves ~1-3s per fresh test PGLite.
|
||||
private _snapshotLoaded = false;
|
||||
/**
|
||||
* Set when connect() auto-repaired the data dir's WAL in place (mirrors
|
||||
* upstream PR #994's `repairedDataDir`). Null on every non-repaired connect.
|
||||
* Test seam + programmatic callers can surface the receipt.
|
||||
*/
|
||||
walRepairReceipt: WalRepairReceipt | null = null;
|
||||
|
||||
get db(): PGLiteDB {
|
||||
if (!this._db) throw new Error('PGLite not connected. Call connect() first.');
|
||||
@@ -306,6 +437,7 @@ export class PGLiteEngine implements BrainEngine {
|
||||
// Lifecycle
|
||||
async connect(config: EngineConfig): Promise<void> {
|
||||
this._savedConfig = config; // #2034: remember for reconnect()
|
||||
this.walRepairReceipt = null; // per-connect: stale receipts must not survive reconnect()
|
||||
const dataDir = config.database_path || undefined; // undefined = in-memory
|
||||
|
||||
// Acquire file lock to prevent concurrent PGLite access (crashes with Aborted())
|
||||
@@ -343,6 +475,11 @@ export class PGLiteEngine implements BrainEngine {
|
||||
extensions: { vector, pg_trgm },
|
||||
}),
|
||||
);
|
||||
// Healthy open: close any repair episode left open by a prior failed
|
||||
// attempt (red-team: episodes otherwise stayed open forever — doctor
|
||||
// kept reporting corruption-likely and a weeks-stale episode backup
|
||||
// could be reused over much newer data). Cheap no-op without a sidecar.
|
||||
if (dataDir) closeRepairEpisodeIfOpen(dataDir);
|
||||
} catch (err) {
|
||||
// v0.13.1: any PGLite.create() failure becomes actionable. v0.41.8.0
|
||||
// (#1340): the previous error hint hardcoded the macOS 26.3 link, but
|
||||
@@ -352,7 +489,54 @@ export class PGLiteEngine implements BrainEngine {
|
||||
// users get the right next step.
|
||||
const original = stringifyPgliteInitError(err); // #2674
|
||||
const verdict = classifyPgliteInitError(original);
|
||||
const wrapped = new Error(buildPgliteInitErrorMessage(verdict, original));
|
||||
let ctx: PgliteInitRepairContext = { repair: 'not-attempted' };
|
||||
|
||||
// WAL-repair wave (#223/#1670/#2575): a wasm-abort on a PERSISTENT data
|
||||
// dir is almost always torn WAL/checkpoint state from an unclean
|
||||
// shutdown — repairable in place. The seam NEVER throws (its failure
|
||||
// modes fold into `ctx`), so every non-repaired path still funnels
|
||||
// through the single lock-release-then-throw site below.
|
||||
if (verdict === 'wasm-abort') {
|
||||
if (!dataDir) {
|
||||
ctx = { repair: 'in-memory' };
|
||||
} else {
|
||||
const attempt = await attemptWalRepairAndRetry(
|
||||
dataDir,
|
||||
() => preservingProcessExitCode(() =>
|
||||
// No loadDataDir on the retry: the snapshot path is
|
||||
// in-memory-only (see above), and dataDir is persistent here.
|
||||
PGlite.create({
|
||||
dataDir,
|
||||
extensions: { vector, pg_trgm },
|
||||
}),
|
||||
),
|
||||
{ reaped: this._lock?.reaped },
|
||||
);
|
||||
if (attempt.status === 'repaired') {
|
||||
this._db = attempt.db;
|
||||
this.walRepairReceipt = attempt.receipt;
|
||||
console.warn(buildWalRepairNotice(attempt.receipt));
|
||||
return; // success: lock stays held, normal connect contract
|
||||
}
|
||||
if (attempt.status === 'skipped') {
|
||||
const reasonToCtx = {
|
||||
'disabled': 'disabled',
|
||||
'validation-failed': 'skipped-validation',
|
||||
'possibly-live-writer': 'skipped-live-writer',
|
||||
'recently-failed': 'skipped-cooldown',
|
||||
} as const;
|
||||
ctx = { repair: reasonToCtx[attempt.reason], detail: attempt.detail };
|
||||
} else {
|
||||
ctx = {
|
||||
repair: attempt.restored ? 'failed-restored' : 'failed-not-restored',
|
||||
backupPath: attempt.receipt?.backupPath,
|
||||
detail: attempt.repairError,
|
||||
};
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const wrapped = new Error(buildPgliteInitErrorMessage(verdict, original, process.platform, ctx));
|
||||
// Release the lock so a fresh process can try again; leaking the lock
|
||||
// here turns a recoverable init error into a stuck-brain state.
|
||||
if (this._lock?.acquired) {
|
||||
|
||||
+89
-12
@@ -14,7 +14,7 @@
|
||||
* try { ... } finally { await releaseLock(lock); }
|
||||
*/
|
||||
|
||||
import { mkdirSync, existsSync, readFileSync, writeFileSync, rmSync, statSync } from 'fs';
|
||||
import { mkdirSync, existsSync, readFileSync, writeFileSync, rmSync, statSync, renameSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { parseGlobalFlags } from './cli-options.ts';
|
||||
|
||||
@@ -25,7 +25,7 @@ const LOCK_FILE = 'lock';
|
||||
// LIVE holder (embed jobs run for many minutes) is never mistaken for stale.
|
||||
const HEARTBEAT_INTERVAL_MS = 30_000;
|
||||
|
||||
class LiveServeLockError extends Error {}
|
||||
export class LiveServeLockError extends Error {}
|
||||
|
||||
function isServeCommand(lockData: { subcommand?: unknown; command?: unknown }): boolean {
|
||||
// New lock files store the command after the same global-flag parsing used
|
||||
@@ -71,6 +71,15 @@ export interface LockHandle {
|
||||
* the NEW owner's live lock and re-open the concurrent-writer hole).
|
||||
*/
|
||||
ownerToken?: string;
|
||||
/**
|
||||
* WAL-repair gate (#223 auto-repair): true when this acquisition reaped a
|
||||
* prior holder's lock — dead-PID reap or corrupt-lock-file removal. A
|
||||
* corrupt lock file cannot prove its holder is dead, and even a dead-PID
|
||||
* verdict can be wrong under PID reuse, so auto WAL surgery refuses to run
|
||||
* on a reaped acquisition (`'possibly-live-writer'`) and asks for a clean
|
||||
* re-run instead. Never set for in-memory engines.
|
||||
*/
|
||||
reaped?: boolean;
|
||||
}
|
||||
|
||||
/** The on-disk lock identity, used to detect "we were reaped and replaced". */
|
||||
@@ -97,13 +106,51 @@ function startHeartbeat(lockPath: string, ownerToken: string): ReturnType<typeof
|
||||
return;
|
||||
}
|
||||
raw.refreshed_at = Date.now();
|
||||
writeFileSync(lockPath, JSON.stringify(raw), { mode: 0o644 });
|
||||
// Atomic tmp+rename (security review): waiting acquirers poll-read this
|
||||
// file every second — an in-place write can be caught mid-flight and a
|
||||
// torn read misclassifies a HEALTHY live holder as a corrupt lock,
|
||||
// getting it reaped. rename makes every read see old-or-new, never torn.
|
||||
const tmpPath = `${lockPath}.tmp-${process.pid}`;
|
||||
writeFileSync(tmpPath, JSON.stringify(raw), { mode: 0o644 });
|
||||
renameSync(tmpPath, lockPath);
|
||||
} catch { /* best-effort — file removed or transient FS error */ }
|
||||
}, HEARTBEAT_INTERVAL_MS);
|
||||
(timer as { unref?: () => void }).unref?.();
|
||||
return timer;
|
||||
}
|
||||
|
||||
/**
|
||||
* Persisted reap marker (security review): written ONLY for corrupt-lock-file
|
||||
* reaps, where the holder's liveness is UNKNOWABLE (the PID can't be read).
|
||||
* The in-process `reaped` flag dies with the acquisition — so the reaper
|
||||
* destroys a possibly-live holder's lock, exits, and the NEXT process
|
||||
* acquires "cleanly" and would run WAL surgery under a live writer. The
|
||||
* marker makes that reap visible across processes: `attemptWalRepairAndRetry`
|
||||
* refuses auto-repair while a recent unknowable-liveness reap is on record.
|
||||
* Dead-PID reaps (affirmative ESRCH verdict) deliberately do NOT write it —
|
||||
* the dead-holder recovery cost stays at one failed command + one re-run.
|
||||
*/
|
||||
function reapMarkerPath(dataDir: string): string {
|
||||
return `${dataDir}.lock-reap.json`;
|
||||
}
|
||||
|
||||
function recordReap(dataDir: string): void {
|
||||
try {
|
||||
writeFileSync(reapMarkerPath(dataDir), JSON.stringify({ ts: Date.now(), by: process.pid }), { mode: 0o644 });
|
||||
} catch { /* best-effort — a marker write failure must not block acquisition */ }
|
||||
}
|
||||
|
||||
/** Milliseconds since the last recorded reap on this data dir, or null. */
|
||||
export function msSinceLastReap(dataDir: string | undefined): number | null {
|
||||
if (!dataDir) return null;
|
||||
try {
|
||||
const raw = JSON.parse(readFileSync(reapMarkerPath(dataDir), 'utf-8')) as { ts?: unknown };
|
||||
return typeof raw.ts === 'number' && Number.isFinite(raw.ts) ? Date.now() - raw.ts : null;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
function getLockDir(dataDir: string | undefined): string {
|
||||
// Use the parent of the data dir for the lock, or a temp location for in-memory
|
||||
if (!dataDir) {
|
||||
@@ -114,13 +161,17 @@ function getLockDir(dataDir: string | undefined): string {
|
||||
return join(dataDir, LOCK_DIR_NAME);
|
||||
}
|
||||
|
||||
function isProcessAlive(pid: number): boolean {
|
||||
export function isProcessAlive(pid: number): boolean {
|
||||
// Only ESRCH (no such process) is affirmative proof of death. EPERM means
|
||||
// the process EXISTS under another user; ERR_INVALID_ARG_TYPE / a malformed
|
||||
// or non-finite pid means we can't tell — all of which must read as ALIVE,
|
||||
// because a false "dead" reaps a live holder's lock (security/codex review).
|
||||
if (!Number.isInteger(pid) || pid <= 0) return true;
|
||||
try {
|
||||
// Sending signal 0 checks existence without actually sending a signal
|
||||
process.kill(pid, 0);
|
||||
process.kill(pid, 0); // signal 0 = existence check, no signal delivered
|
||||
return true;
|
||||
} catch {
|
||||
return false;
|
||||
} catch (err) {
|
||||
return (err as NodeJS.ErrnoException)?.code !== 'ESRCH';
|
||||
}
|
||||
}
|
||||
|
||||
@@ -172,6 +223,7 @@ export async function acquireLock(dataDir: string | undefined, opts?: { timeoutM
|
||||
|
||||
const timeoutMs = opts?.timeoutMs ?? 30_000; // 30 second default timeout
|
||||
const startTime = Date.now();
|
||||
let reaped = false; // see LockHandle.reaped
|
||||
|
||||
while (Date.now() - startTime < timeoutMs) {
|
||||
// Check for stale lock first
|
||||
@@ -187,7 +239,12 @@ export async function acquireLock(dataDir: string | undefined, opts?: { timeoutM
|
||||
// heartbeat" is NOT evidence of death — only a dead PID is.
|
||||
const alive = isProcessAlive(lockPid);
|
||||
if (!alive) {
|
||||
// Holder process is gone — reap and try to acquire.
|
||||
// Holder process is gone — reap and try to acquire. This verdict is
|
||||
// affirmative (kill-0 threw ESRCH; EPERM reads as alive), so no
|
||||
// cross-process quarantine marker: the same-acquisition `reaped`
|
||||
// flag alone gates repair, keeping the dead-holder recovery cost at
|
||||
// one failed command + one re-run.
|
||||
reaped = true;
|
||||
try { rmSync(lockDir, { recursive: true, force: true }); } catch { /* race condition, try again */ }
|
||||
} else {
|
||||
if (isServeCommand(lockData)) {
|
||||
@@ -209,7 +266,22 @@ export async function acquireLock(dataDir: string | undefined, opts?: { timeoutM
|
||||
// A live MCP server is not a stale or corrupt lock. Surface the useful
|
||||
// explanation without touching the lock it still owns.
|
||||
if (err instanceof LiveServeLockError) throw err;
|
||||
// Corrupt lock file — remove it
|
||||
// ENOENT = acquisition in flight (a concurrent acquirer did mkdir but
|
||||
// hasn't written the lock file yet) — reaping HERE would destroy a
|
||||
// LIVE acquirer's lock and put two writers on one dir (red-team).
|
||||
// Give the writer a grace window keyed on the lock dir's age.
|
||||
if ((err as NodeJS.ErrnoException)?.code === 'ENOENT') {
|
||||
let lockDirAgeMs = Infinity;
|
||||
try { lockDirAgeMs = Date.now() - statSync(lockDir).mtimeMs; } catch { /* dir gone — retry loop handles */ }
|
||||
if (lockDirAgeMs < 10_000) {
|
||||
await new Promise(r => setTimeout(r, 200));
|
||||
continue;
|
||||
}
|
||||
}
|
||||
// Corrupt lock file — remove it. The holder's liveness is UNKNOWABLE
|
||||
// here (unreadable PID), so this counts as a reap for the repair gate.
|
||||
reaped = true;
|
||||
recordReap(dataDir as string);
|
||||
try { rmSync(lockDir, { recursive: true, force: true }); } catch { /* race condition */ }
|
||||
}
|
||||
}
|
||||
@@ -221,16 +293,21 @@ export async function acquireLock(dataDir: string | undefined, opts?: { timeoutM
|
||||
// the heartbeat so this holder reads as alive-and-working to others.
|
||||
const lockPath = join(lockDir, LOCK_FILE);
|
||||
const now = Date.now();
|
||||
writeFileSync(lockPath, JSON.stringify({
|
||||
// Atomic tmp+rename, same torn-read protection as the heartbeat: a
|
||||
// concurrent poll-reader must see the file complete or absent, never
|
||||
// mid-write (a torn read classifies a LIVE holder as corrupt).
|
||||
const initTmp = `${lockPath}.tmp-${process.pid}`;
|
||||
writeFileSync(initTmp, JSON.stringify({
|
||||
pid: process.pid,
|
||||
acquired_at: now,
|
||||
refreshed_at: now,
|
||||
command: process.argv.slice(1).join(' '),
|
||||
subcommand: parseGlobalFlags(process.argv.slice(2)).rest[0] ?? null,
|
||||
}), { mode: 0o644 });
|
||||
renameSync(initTmp, lockPath);
|
||||
|
||||
const ownerToken = tokenOf({ pid: process.pid, acquired_at: now });
|
||||
return { lockDir, acquired: true, lockPath, ownerToken, heartbeat: startHeartbeat(lockPath, ownerToken) };
|
||||
return { lockDir, acquired: true, lockPath, ownerToken, reaped, heartbeat: startHeartbeat(lockPath, ownerToken) };
|
||||
} catch (e: unknown) {
|
||||
// mkdir failed — someone else grabbed it between our check and mkdir
|
||||
// This is fine, we'll retry
|
||||
|
||||
@@ -0,0 +1,771 @@
|
||||
/**
|
||||
* PGLite WAL-repair orchestrator (#223 / #1670 / #2575).
|
||||
*
|
||||
* Wraps the pg_resetwal port (`pglite-resetwal.ts`) with everything that makes
|
||||
* it safe to run automatically from `PGLiteEngine.connect()`:
|
||||
*
|
||||
* validate (read-only, fail-closed) → back up (rename, not copy) →
|
||||
* resetWal → retry create() once → restore on failure.
|
||||
*
|
||||
* Safety posture (eng-review 1A/2A/3A/4A + codex round):
|
||||
* - WAL surgery only runs under a CLEANLY-acquired data-dir lock. A reaped
|
||||
* acquisition (dead-PID or corrupt-lock-file reap — the only reaps that
|
||||
* exist post-#2348) refuses with `'possibly-live-writer'`; this module
|
||||
* never force-removes `.gbrain-lock`.
|
||||
* - Backup is a whole-`pg_wal/`-directory rename into a sibling dir (O(1),
|
||||
* zero extra disk — a real brain's pg_wal is ~144MB and a copy would
|
||||
* transiently double it, ENOSPC-ing exactly on disk-pressure machines);
|
||||
* only the 8KB `global/pg_control` is copied (it is mutated in place).
|
||||
* - Restore never deletes anything and never leaves the dir without a valid
|
||||
* pg_control: control is restored first (atomic tmp+rename), then the reset
|
||||
* pg_wal is renamed ASIDE into the backup dir and the original renamed back.
|
||||
* - A cooldown sidecar + episode-scoped backups bound the reconnect loops
|
||||
* (autopilot ~10s tick under launchd KeepAlive; minion supervisor): repeated
|
||||
* attempts inside one corruption episode reuse the episode's first backup
|
||||
* (the pre-damage forensic state) instead of stacking new ones, and a
|
||||
* recently-failed attempt skips repair entirely for the cooldown window.
|
||||
*
|
||||
* `attemptWalRepairAndRetry` NEVER throws — `connect()`'s catch consumes the
|
||||
* discriminated union, so no new code path can bypass the engine's single
|
||||
* lock-release-then-throw site.
|
||||
*
|
||||
* Env knobs (incident escape hatches, env-only by design):
|
||||
* GBRAIN_PGLITE_WAL_REPAIR=off disable auto-repair
|
||||
* GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS default 3600
|
||||
*/
|
||||
import {
|
||||
existsSync, lstatSync, readdirSync, readFileSync, statSync, writeFileSync,
|
||||
mkdirSync, rmSync, renameSync,
|
||||
} from 'node:fs';
|
||||
import { readFile, rename } from 'node:fs/promises';
|
||||
import { basename, dirname, join } from 'node:path';
|
||||
import { resetWal, writeFileAtomicSynced, WalResetUnsupportedError, PG_CONTROL_FILE_SIZE, isWalSegmentName } from './pglite-resetwal.ts';
|
||||
import { msSinceLastReap, isProcessAlive } from './pglite-lock.ts';
|
||||
|
||||
// A recent reap on this data dir — by ANY process — means a holder that may
|
||||
// still be alive lost its lock; auto WAL surgery stays off until the window
|
||||
// clears (security review: the in-process `reaped` flag alone let the NEXT
|
||||
// acquirer look clean while the reaped holder was still writing).
|
||||
const REAP_QUARANTINE_MS = 10 * 60 * 1000;
|
||||
|
||||
const BACKUP_DIR_MARKER = '.wal-repair-backup-';
|
||||
const SIDECAR_SUFFIX = '.wal-repair-attempt.json';
|
||||
const MAX_SIDECAR_ATTEMPTS = 10;
|
||||
const KEEP_EPISODES = 3;
|
||||
const DEFAULT_COOLDOWN_SECONDS = 3600;
|
||||
|
||||
export interface WalRepairReceipt {
|
||||
dataDir: string;
|
||||
/** Sibling dir holding the pre-repair state: `<dataDir>.wal-repair-backup-<ts>/` */
|
||||
backupPath: string;
|
||||
/** Relative paths preserved in the backup (e.g. 'pg_wal/', 'postmaster.pid', 'global/pg_control'). */
|
||||
backedUpFiles: string[];
|
||||
/** True when this attempt reused an open episode's existing backup. */
|
||||
reusedEpisodeBackup: boolean;
|
||||
resetSegment: string;
|
||||
timelineId: number;
|
||||
walSegSize: number;
|
||||
repairedAt: string; // ISO
|
||||
}
|
||||
|
||||
export type WalRepairValidation =
|
||||
| { ok: true }
|
||||
| {
|
||||
ok: false;
|
||||
reason: 'missing-dir' | 'not-pglite-layout' | 'unsupported-pg-version' | 'bad-pg-control';
|
||||
detail: string;
|
||||
};
|
||||
|
||||
export interface RestoreResult {
|
||||
restored: boolean;
|
||||
steps: string[];
|
||||
detail?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Thrown by `repairPgliteWal` when the reset failed AFTER the backup was
|
||||
* taken. Carries the receipt and the result of the best-effort restore so the
|
||||
* seam can report `restored` HONESTLY instead of assuming the restore worked
|
||||
* (the `failed-not-restored` message arm depends on this being truthful).
|
||||
*/
|
||||
export class WalRepairError extends Error {
|
||||
constructor(
|
||||
message: string,
|
||||
readonly receipt: WalRepairReceipt,
|
||||
readonly restore: RestoreResult,
|
||||
) {
|
||||
super(message);
|
||||
this.name = 'WalRepairError';
|
||||
}
|
||||
}
|
||||
|
||||
export type WalRepairAttempt<T> =
|
||||
| { status: 'repaired'; db: T; receipt: WalRepairReceipt }
|
||||
| {
|
||||
status: 'skipped';
|
||||
reason: 'disabled' | 'validation-failed' | 'possibly-live-writer' | 'recently-failed';
|
||||
detail: string;
|
||||
}
|
||||
| { status: 'failed'; receipt: WalRepairReceipt | null; restored: boolean; repairError: string };
|
||||
|
||||
export interface PgliteDirDiagnosis {
|
||||
exists: boolean;
|
||||
postmasterPid: boolean;
|
||||
pgControlOk: boolean;
|
||||
pgVersion: string | null;
|
||||
walSegments: string[];
|
||||
lockHeld: boolean;
|
||||
lockHolderPid: number | null;
|
||||
/** Sibling `*.wal-repair-backup-*` dirs, newest first. */
|
||||
backupDirs: string[];
|
||||
/** Recent repair attempts from the sidecar, newest last. */
|
||||
recentAttempts: Array<{ ts: number; outcome: 'repaired' | 'failed' }>;
|
||||
verdict: 'looks-healthy' | 'wal-corruption-likely' | 'locked' | 'missing' | 'unsupported-layout';
|
||||
detail: string;
|
||||
}
|
||||
|
||||
interface RepairSidecar {
|
||||
/** ts of the first failed attempt of the open episode; null = no open episode. */
|
||||
episodeStartedAt: number | null;
|
||||
/** The open episode's (first) backup dir — the pre-damage forensic state. */
|
||||
episodeBackupPath: string | null;
|
||||
attempts: Array<{ ts: number; outcome: 'repaired' | 'failed'; backupPath: string | null }>;
|
||||
}
|
||||
|
||||
function sidecarPath(dataDir: string): string {
|
||||
return `${dataDir}${SIDECAR_SUFFIX}`;
|
||||
}
|
||||
|
||||
export function readRepairSidecar(dataDir: string): RepairSidecar {
|
||||
try {
|
||||
const raw = JSON.parse(readFileSync(sidecarPath(dataDir), 'utf-8')) as Partial<RepairSidecar>;
|
||||
return {
|
||||
episodeStartedAt: typeof raw.episodeStartedAt === 'number' ? raw.episodeStartedAt : null,
|
||||
episodeBackupPath: typeof raw.episodeBackupPath === 'string' ? raw.episodeBackupPath : null,
|
||||
attempts: Array.isArray(raw.attempts)
|
||||
? raw.attempts.filter(
|
||||
(a): a is RepairSidecar['attempts'][number] =>
|
||||
!!a && typeof a.ts === 'number' && (a.outcome === 'repaired' || a.outcome === 'failed'),
|
||||
)
|
||||
: [],
|
||||
};
|
||||
} catch {
|
||||
return { episodeStartedAt: null, episodeBackupPath: null, attempts: [] };
|
||||
}
|
||||
}
|
||||
|
||||
function writeRepairSidecar(dataDir: string, sidecar: RepairSidecar): void {
|
||||
try {
|
||||
// Atomic tmp+rename: a kill/power-loss mid-write must not truncate the
|
||||
// sidecar to invalid JSON (readRepairSidecar would then silently reset the
|
||||
// episode/cooldown state — codex review). rename is atomic; a torn tmp is
|
||||
// discarded on the next write.
|
||||
const tmp = `${sidecarPath(dataDir)}.tmp-${process.pid}`;
|
||||
writeFileSync(tmp, JSON.stringify(sidecar), { mode: 0o644 });
|
||||
renameSync(tmp, sidecarPath(dataDir));
|
||||
} catch { /* best-effort — a sidecar write failure must never block recovery */ }
|
||||
}
|
||||
|
||||
/**
|
||||
* Record a real repair attempt (repaired|failed) and manage episode state:
|
||||
* a `failed` attempt opens an episode (if none is open) pinning its backup as
|
||||
* the episode backup; a VERIFIED `repaired` attempt closes the episode and
|
||||
* prunes retained backups to the newest KEEP_EPISODES. The manual command
|
||||
* passes `closeEpisode: false` — its "repaired" is unverified (the next
|
||||
* connect proves it), and closing+pruning on unverified success let repeated
|
||||
* manual runs delete the pre-damage forensic backup (red-team finding); the
|
||||
* episode instead closes on the next successful connect
|
||||
* (`closeRepairEpisodeIfOpen`).
|
||||
*
|
||||
* Re-pin rule (red-team finding): a restore MOVES pg_wal back out of the
|
||||
* backup, gutting it — if a later failed attempt took a FRESH backup while an
|
||||
* episode pinned a gutted dir, the pin moves to the fresh backup so the
|
||||
* episode's protected copy is always one that still holds pg_wal.
|
||||
*/
|
||||
export function recordRepairAttempt(
|
||||
dataDir: string,
|
||||
outcome: 'repaired' | 'failed',
|
||||
backupPath: string | null,
|
||||
opts?: { closeEpisode?: boolean },
|
||||
): void {
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
sidecar.attempts.push({ ts: Date.now(), outcome, backupPath });
|
||||
if (sidecar.attempts.length > MAX_SIDECAR_ATTEMPTS) {
|
||||
sidecar.attempts = sidecar.attempts.slice(-MAX_SIDECAR_ATTEMPTS);
|
||||
}
|
||||
if (outcome === 'failed') {
|
||||
if (sidecar.episodeStartedAt === null) {
|
||||
sidecar.episodeStartedAt = Date.now();
|
||||
sidecar.episodeBackupPath = backupPath;
|
||||
} else if (
|
||||
backupPath &&
|
||||
backupPath !== sidecar.episodeBackupPath &&
|
||||
(!sidecar.episodeBackupPath || !existsSync(join(sidecar.episodeBackupPath, 'pg_wal')))
|
||||
) {
|
||||
sidecar.episodeBackupPath = backupPath;
|
||||
}
|
||||
} else if (opts?.closeEpisode !== false) {
|
||||
sidecar.episodeStartedAt = null;
|
||||
sidecar.episodeBackupPath = null;
|
||||
}
|
||||
writeRepairSidecar(dataDir, sidecar);
|
||||
if (outcome === 'repaired' && opts?.closeEpisode !== false) {
|
||||
pruneRepairBackups(dataDir);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Close any open repair episode after a HEALTHY connect (red-team finding: a
|
||||
* plain successful open never touched the sidecar, so an episode stayed open
|
||||
* forever — doctor kept reporting corruption-likely, and a weeks-stale
|
||||
* episode backup could be reused over much newer data). Cheap no-op when no
|
||||
* sidecar exists. Called by PGLiteEngine.connect() on every non-repaired
|
||||
* success and by the seam's 'repaired' arm via recordRepairAttempt.
|
||||
*/
|
||||
export function closeRepairEpisodeIfOpen(dataDir: string): void {
|
||||
try {
|
||||
if (!existsSync(sidecarPath(dataDir))) return;
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
if (sidecar.episodeStartedAt === null) return;
|
||||
sidecar.episodeStartedAt = null;
|
||||
sidecar.episodeBackupPath = null;
|
||||
writeRepairSidecar(dataDir, sidecar);
|
||||
pruneRepairBackups(dataDir);
|
||||
} catch { /* best-effort */ }
|
||||
}
|
||||
|
||||
/**
|
||||
* Cooldown: true when the last FAILED attempt is inside the cooldown window.
|
||||
* DELIBERATE: a later successful repair does NOT clear the cooldown — repeated
|
||||
* corruption right after a "success" usually means the unclean-shutdown
|
||||
* genesis is still active, and looping surgery would silently eat a WAL tail
|
||||
* per cycle. The manual `gbrain pglite-repair` command bypasses the cooldown.
|
||||
*/
|
||||
export function repairCooldownActive(dataDir: string): { active: boolean; detail: string } {
|
||||
const seconds = Number(process.env.GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS ?? DEFAULT_COOLDOWN_SECONDS);
|
||||
const windowMs = (Number.isFinite(seconds) && seconds >= 0 ? seconds : DEFAULT_COOLDOWN_SECONDS) * 1000;
|
||||
if (windowMs === 0) return { active: false, detail: 'cooldown disabled (0s)' };
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
// Repaired-loop guard (red-team finding): a crash loop where every reopen
|
||||
// aborts but repair "succeeds" each time would silently discard a WAL tail
|
||||
// per cycle with no failed attempt ever recorded. Two successful repairs
|
||||
// inside one window = the corruption genesis is active — stop auto-repair
|
||||
// and let doctor escalate.
|
||||
const repairedInWindow = sidecar.attempts.filter(
|
||||
(a) => a.outcome === 'repaired' && Date.now() - a.ts >= 0 && Date.now() - a.ts < windowMs,
|
||||
).length;
|
||||
if (repairedInWindow >= 2) {
|
||||
return {
|
||||
active: true,
|
||||
detail:
|
||||
`auto-repair already ran ${repairedInWindow}x in the last ${windowMs / 1000}s — repeated ` +
|
||||
'corruption means the unclean-shutdown genesis is still active; refusing to silently ' +
|
||||
'discard another WAL tail. Run `gbrain doctor`, or `gbrain pglite-repair` manually.',
|
||||
};
|
||||
}
|
||||
const lastFailed = [...sidecar.attempts].reverse().find((a) => a.outcome === 'failed');
|
||||
if (!lastFailed) return { active: false, detail: 'no prior failed attempt' };
|
||||
const ageMs = Date.now() - lastFailed.ts;
|
||||
// Clock skew (unclean-reboot recovery is exactly when clocks step): a
|
||||
// negative age means the recorded ts is in the future — treat as expired
|
||||
// rather than suppressing auto-repair until wall-clock catches up.
|
||||
if (ageMs >= 0 && ageMs < windowMs) {
|
||||
return {
|
||||
active: true,
|
||||
detail:
|
||||
`last auto-repair attempt failed ${Math.round(ageMs / 1000)}s ago ` +
|
||||
`(cooldown ${windowMs / 1000}s — set GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS=0 to bypass, ` +
|
||||
`or run \`gbrain pglite-repair\` manually)`,
|
||||
};
|
||||
}
|
||||
return { active: false, detail: 'cooldown expired' };
|
||||
}
|
||||
|
||||
/** Sibling `*.wal-repair-backup-*` dirs for this data dir, newest first. */
|
||||
export function listRepairBackups(dataDir: string): string[] {
|
||||
try {
|
||||
const parent = dirname(dataDir);
|
||||
const prefix = `${basename(dataDir)}${BACKUP_DIR_MARKER}`;
|
||||
return readdirSync(parent)
|
||||
.filter((name) => name.startsWith(prefix))
|
||||
.sort()
|
||||
.reverse()
|
||||
.map((name) => join(parent, name));
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Keep the newest KEEP_EPISODES backups; never prune the open episode's backup.
|
||||
* Runs only after a successful repair (episode close) — never mid-incident.
|
||||
*/
|
||||
export function pruneRepairBackups(dataDir: string): void {
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
const backups = listRepairBackups(dataDir); // newest first
|
||||
const keep = new Set(backups.slice(0, KEEP_EPISODES));
|
||||
if (sidecar.episodeBackupPath) keep.add(sidecar.episodeBackupPath);
|
||||
for (const dir of backups) {
|
||||
if (!keep.has(dir)) {
|
||||
try { rmSync(dir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
export function walRepairEnabled(): boolean {
|
||||
return process.env.GBRAIN_PGLITE_WAL_REPAIR !== 'off';
|
||||
}
|
||||
|
||||
function isSymlink(path: string): boolean {
|
||||
try {
|
||||
return lstatSync(path).isSymbolicLink();
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Read-only, fail-closed: does this look like a PG17 pglite data dir we know
|
||||
* how to repair? Tolerates the `.gbrain-lock` entry (the lock lives INSIDE the
|
||||
* data dir). Refuses symlinked components (codex 14.8 — a symlinked `pg_wal`
|
||||
* or `pg_control` could redirect the backup/restore renames at unrelated
|
||||
* files; same confinement discipline as the v0.42.55.0 security wave).
|
||||
*/
|
||||
export function validateWalRepairTarget(dataDir: string): WalRepairValidation {
|
||||
if (!dataDir) return { ok: false, reason: 'missing-dir', detail: 'no data dir configured (in-memory engine)' };
|
||||
if (!existsSync(dataDir)) return { ok: false, reason: 'missing-dir', detail: `${dataDir} does not exist` };
|
||||
if (
|
||||
isSymlink(dataDir) ||
|
||||
isSymlink(join(dataDir, 'pg_wal')) ||
|
||||
// `global/` itself must be checked too: lstat on global/pg_control follows
|
||||
// the INTERMEDIATE symlink, so a symlinked global/ would pass and surgery
|
||||
// would write a forged pg_control through it into a foreign directory
|
||||
// (security review finding).
|
||||
isSymlink(join(dataDir, 'global')) ||
|
||||
isSymlink(join(dataDir, 'global', 'pg_control'))
|
||||
) {
|
||||
return { ok: false, reason: 'not-pglite-layout', detail: 'data dir, pg_wal, global, or pg_control is a symlink — refusing to run rename-based repair through symlinks' };
|
||||
}
|
||||
let pgVersion: string;
|
||||
try {
|
||||
pgVersion = readFileSync(join(dataDir, 'PG_VERSION'), 'utf-8').trim();
|
||||
} catch {
|
||||
return { ok: false, reason: 'not-pglite-layout', detail: `no readable PG_VERSION in ${dataDir}` };
|
||||
}
|
||||
if (pgVersion !== '17') {
|
||||
return { ok: false, reason: 'unsupported-pg-version', detail: `PG_VERSION is ${pgVersion}, this repair understands 17 only` };
|
||||
}
|
||||
if (!existsSync(join(dataDir, 'base'))) {
|
||||
return { ok: false, reason: 'not-pglite-layout', detail: `no base/ directory in ${dataDir}` };
|
||||
}
|
||||
// Live-postmaster refusal (red-team finding): real pg_resetwal refuses when
|
||||
// postmaster.pid exists. A LIVE native Postgres 17 data dir passes every
|
||||
// layout check here — without this guard, `gbrain pglite-repair --path` at
|
||||
// such a dir would rename pg_wal out from under a running postmaster the
|
||||
// gbrain lock cannot see. A stale pid file (dead process) stays repairable.
|
||||
try {
|
||||
const pidRaw = readFileSync(join(dataDir, 'postmaster.pid'), 'utf-8').split('\n')[0]?.trim();
|
||||
const pid = Number(pidRaw);
|
||||
if (Number.isInteger(pid) && pid > 0 && isProcessAlive(pid)) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'not-pglite-layout',
|
||||
detail: `postmaster.pid names a LIVE process (PID ${pid}) — refusing WAL surgery on a possibly-running database`,
|
||||
};
|
||||
}
|
||||
} catch { /* no postmaster.pid or unreadable — fine */ }
|
||||
const controlPath = join(dataDir, 'global', 'pg_control');
|
||||
try {
|
||||
const size = statSync(controlPath).size;
|
||||
if (size !== PG_CONTROL_FILE_SIZE) {
|
||||
return { ok: false, reason: 'bad-pg-control', detail: `pg_control is ${size} bytes, expected ${PG_CONTROL_FILE_SIZE}` };
|
||||
}
|
||||
} catch {
|
||||
return { ok: false, reason: 'bad-pg-control', detail: `no readable ${controlPath}` };
|
||||
}
|
||||
return { ok: true };
|
||||
}
|
||||
|
||||
/** Read-only diagnosis for `gbrain doctor` and `pglite-repair --dry-run`. */
|
||||
export function inspectPgliteDataDir(dataDir: string): PgliteDirDiagnosis {
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
const base: Omit<PgliteDirDiagnosis, 'verdict' | 'detail'> = {
|
||||
exists: !!dataDir && existsSync(dataDir),
|
||||
postmasterPid: !!dataDir && existsSync(join(dataDir, 'postmaster.pid')),
|
||||
pgControlOk: false,
|
||||
pgVersion: null,
|
||||
walSegments: [],
|
||||
lockHeld: false,
|
||||
lockHolderPid: null,
|
||||
backupDirs: listRepairBackups(dataDir),
|
||||
recentAttempts: sidecar.attempts.map(({ ts, outcome }) => ({ ts, outcome })),
|
||||
};
|
||||
if (!base.exists) {
|
||||
return { ...base, verdict: 'missing', detail: `${dataDir || '(in-memory)'} does not exist` };
|
||||
}
|
||||
try {
|
||||
base.pgVersion = readFileSync(join(dataDir, 'PG_VERSION'), 'utf-8').trim();
|
||||
} catch { /* leave null */ }
|
||||
try {
|
||||
base.pgControlOk = statSync(join(dataDir, 'global', 'pg_control')).size === PG_CONTROL_FILE_SIZE;
|
||||
} catch { /* leave false */ }
|
||||
try {
|
||||
base.walSegments = readdirSync(join(dataDir, 'pg_wal')).filter(isWalSegmentName).sort();
|
||||
} catch { /* leave empty */ }
|
||||
try {
|
||||
const lockData = JSON.parse(readFileSync(join(dataDir, '.gbrain-lock', 'lock'), 'utf-8')) as { pid?: number };
|
||||
if (typeof lockData.pid === 'number' && isProcessAlive(lockData.pid)) {
|
||||
base.lockHeld = true;
|
||||
base.lockHolderPid = lockData.pid;
|
||||
}
|
||||
} catch { /* no lock / unreadable — not held */ }
|
||||
|
||||
if (base.lockHeld) {
|
||||
return { ...base, verdict: 'locked', detail: `data dir lock held by live PID ${base.lockHolderPid}` };
|
||||
}
|
||||
const validation = validateWalRepairTarget(dataDir);
|
||||
if (!validation.ok) {
|
||||
return { ...base, verdict: 'unsupported-layout', detail: validation.detail };
|
||||
}
|
||||
if (base.postmasterPid || sidecar.episodeStartedAt !== null) {
|
||||
return {
|
||||
...base,
|
||||
verdict: 'wal-corruption-likely',
|
||||
detail: base.postmasterPid
|
||||
? 'stale postmaster.pid present — an unclean shutdown left WAL/checkpoint state torn'
|
||||
: 'an unresolved repair episode is open (a prior auto-repair attempt failed)',
|
||||
};
|
||||
}
|
||||
return { ...base, verdict: 'looks-healthy', detail: 'layout validates; no unclean-shutdown markers on disk' };
|
||||
}
|
||||
|
||||
/**
|
||||
* Mechanical repair: (backup unless reusing an episode backup) → resetWal.
|
||||
* Backup = rename the ENTIRE pg_wal/ dir + postmaster.pid into the backup dir
|
||||
* (covers archive_status/summaries too — restore is truly byte-identical) and
|
||||
* COPY the 8KB pg_control. If resetWal throws after the backup was taken, a
|
||||
* best-effort restore runs before the error propagates — this function never
|
||||
* leaves the dir backed-up-but-unrepaired without attempting to put it back.
|
||||
*/
|
||||
export async function repairPgliteWal(
|
||||
dataDir: string,
|
||||
opts?: { reuseBackupPath?: string },
|
||||
): Promise<WalRepairReceipt> {
|
||||
const validation = validateWalRepairTarget(dataDir);
|
||||
if (!validation.ok) {
|
||||
throw new WalResetUnsupportedError(`refusing repair: ${validation.detail}`);
|
||||
}
|
||||
|
||||
// Defense-in-depth (security + red-team reviews): the reuse path comes from
|
||||
// the user-writable sidecar JSON — only honor it when it is a real,
|
||||
// non-symlink, traversal-free sibling backup dir of THIS data dir that
|
||||
// STILL CONTAINS pg_wal (a restore MOVES pg_wal back out, gutting the
|
||||
// backup; reusing a gutted backup would let resetWal unlink the only
|
||||
// surviving WAL copy in place). Anything else gets a fresh backup.
|
||||
const safeReusePath =
|
||||
opts?.reuseBackupPath &&
|
||||
opts.reuseBackupPath.startsWith(`${dataDir}${BACKUP_DIR_MARKER}`) &&
|
||||
!opts.reuseBackupPath.includes('..') &&
|
||||
existsSync(opts.reuseBackupPath) &&
|
||||
!isSymlink(opts.reuseBackupPath) &&
|
||||
existsSync(join(opts.reuseBackupPath, 'pg_wal')) &&
|
||||
!isSymlink(join(opts.reuseBackupPath, 'pg_wal'))
|
||||
? opts.reuseBackupPath
|
||||
: undefined;
|
||||
const backedUpFiles: string[] = [];
|
||||
const reusedEpisodeBackup = !!safeReusePath;
|
||||
|
||||
let backupPath: string;
|
||||
if (reusedEpisodeBackup) {
|
||||
// Episode reuse: the open episode's backup still holds pg_wal (enforced by
|
||||
// safeReusePath — a restore MOVES pg_wal back out and guts the backup; a
|
||||
// gutted backup must never be reused or resetWal would unlink the only
|
||||
// surviving WAL copy in place). resetWal's own deletion loops clear the
|
||||
// current segments.
|
||||
backupPath = safeReusePath!;
|
||||
} else {
|
||||
// Fresh backup dir: mkdir with recursive:false and fail-closed on
|
||||
// collision (red-team: a predictable pre-existing dir or symlink-to-dir
|
||||
// would silently receive the renames, and restore would later read
|
||||
// pg_control bytes back OUT of it). Retry with a suffix, then verify we
|
||||
// created a real directory.
|
||||
backupPath = `${dataDir}${BACKUP_DIR_MARKER}${Date.now()}`;
|
||||
for (let attempt = 0; ; attempt++) {
|
||||
try {
|
||||
mkdirSync(backupPath, { recursive: false });
|
||||
break;
|
||||
} catch (err) {
|
||||
if ((err as NodeJS.ErrnoException)?.code === 'EEXIST' && attempt < 5) {
|
||||
backupPath = `${dataDir}${BACKUP_DIR_MARKER}${Date.now()}-${attempt + 1}`;
|
||||
continue;
|
||||
}
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
if (isSymlink(backupPath) || !statSync(backupPath).isDirectory()) {
|
||||
throw new WalResetUnsupportedError(`backup path ${backupPath} is not a real directory`);
|
||||
}
|
||||
}
|
||||
// Track whether the backup dir received anything, so refusal paths can prune
|
||||
// an empty leftover (adversarial review F11: doctor would otherwise inventory
|
||||
// an empty `<dataDir>.wal-repair-backup-*` as a real backup).
|
||||
const pruneEmptyBackup = () => {
|
||||
if (reusedEpisodeBackup) return;
|
||||
try { if (readdirSync(backupPath).length === 0) rmSync(backupPath, { recursive: true, force: true }); } catch { /* best-effort */ }
|
||||
};
|
||||
|
||||
const receipt: WalRepairReceipt = {
|
||||
dataDir,
|
||||
backupPath,
|
||||
backedUpFiles,
|
||||
reusedEpisodeBackup,
|
||||
resetSegment: '',
|
||||
timelineId: 0,
|
||||
walSegSize: 0,
|
||||
repairedAt: new Date().toISOString(),
|
||||
};
|
||||
|
||||
if (!reusedEpisodeBackup) {
|
||||
// Backup phase. Once the FIRST rename lands, any failure here must run a
|
||||
// restore and surface via WalRepairError — a generic throw would read as
|
||||
// "dir never touched" while pg_wal is actually sitting in the backup dir
|
||||
// (red-team finding).
|
||||
let backupStarted = false;
|
||||
try {
|
||||
const walDir = join(dataDir, 'pg_wal');
|
||||
if (existsSync(walDir)) {
|
||||
await rename(walDir, join(backupPath, 'pg_wal'));
|
||||
backupStarted = true;
|
||||
backedUpFiles.push('pg_wal/');
|
||||
}
|
||||
const pidFile = join(dataDir, 'postmaster.pid');
|
||||
if (existsSync(pidFile)) {
|
||||
await rename(pidFile, join(backupPath, 'postmaster.pid'));
|
||||
backupStarted = true;
|
||||
backedUpFiles.push('postmaster.pid');
|
||||
}
|
||||
const control = await readFile(join(dataDir, 'global', 'pg_control'));
|
||||
await writeFileAtomicSynced(backupPath, 'pg_control', Buffer.from(control));
|
||||
backedUpFiles.push('global/pg_control');
|
||||
} catch (err) {
|
||||
if (!backupStarted) { pruneEmptyBackup(); throw err; } // dir genuinely untouched
|
||||
const restore = await restoreWalBackup(receipt);
|
||||
throw new WalRepairError(String((err as Error)?.message ?? err), receipt, restore);
|
||||
}
|
||||
}
|
||||
|
||||
try {
|
||||
const result = await resetWal(dataDir);
|
||||
receipt.resetSegment = result.resetSegment;
|
||||
receipt.timelineId = result.timelineId;
|
||||
receipt.walSegSize = result.walSegSize;
|
||||
} catch (err) {
|
||||
// Best-effort restore — never leave backed-up-but-unrepaired. The result
|
||||
// is THREADED OUT via WalRepairError so callers can report `restored`
|
||||
// honestly (review finding: discarding it let 'failed-restored' lie).
|
||||
const restore = await restoreWalBackup(receipt);
|
||||
throw new WalRepairError(String((err as Error)?.message ?? err), receipt, restore);
|
||||
}
|
||||
return receipt;
|
||||
}
|
||||
|
||||
/**
|
||||
* Put the data dir back to the backed-up state. Overwrite order (eng-review
|
||||
* 3A): pg_control FIRST (atomic tmp+rename — no instant leaves the dir without
|
||||
* a valid control file), then the pg_wal dir swap (reset dir renamed ASIDE
|
||||
* into the backup dir — nothing is ever deleted during restore), postmaster.pid
|
||||
* deliberately NOT restored (it was stale by definition). Mtime guard
|
||||
* (eng-review 1A): refuses when the current pg_wal contains segments newer
|
||||
* than the backup that this repair did not write — a live writer advanced the
|
||||
* dir; renaming it away would destroy real WAL.
|
||||
* Never throws — reports `{restored:false, detail}` instead.
|
||||
*/
|
||||
export async function restoreWalBackup(receipt: WalRepairReceipt): Promise<RestoreResult> {
|
||||
const steps: string[] = [];
|
||||
try {
|
||||
const { dataDir, backupPath } = receipt;
|
||||
const walDir = join(dataDir, 'pg_wal');
|
||||
const backupWal = join(backupPath, 'pg_wal');
|
||||
const backupControl = join(backupPath, 'pg_control');
|
||||
|
||||
// Symlinked backup CHILDREN would let restore read attacker-chosen
|
||||
// pg_control bytes or rename a foreign pg_wal into the data dir
|
||||
// (red-team) — the top-level checks don't cover them.
|
||||
if (isSymlink(backupWal) || isSymlink(backupControl)) {
|
||||
return { restored: false, steps, detail: `backup at ${backupPath} contains symlinked components — refusing restore` };
|
||||
}
|
||||
|
||||
// Mtime guard: any foreign WAL segment newer than this repair's start?
|
||||
const backupTs = Date.parse(receipt.repairedAt);
|
||||
if (existsSync(walDir)) {
|
||||
for (const f of readdirSync(walDir)) {
|
||||
if (!isWalSegmentName(f) || f === receipt.resetSegment) continue;
|
||||
try {
|
||||
if (statSync(join(walDir, f)).mtimeMs > backupTs) {
|
||||
return {
|
||||
restored: false,
|
||||
steps,
|
||||
detail: `mtime-guard: ${f} in pg_wal is newer than the backup — a live writer may have advanced this dir; refusing to swap WAL back`,
|
||||
};
|
||||
}
|
||||
} catch { /* statable race — ignore */ }
|
||||
}
|
||||
}
|
||||
|
||||
if (existsSync(backupControl)) {
|
||||
const control = await readFile(backupControl);
|
||||
await writeFileAtomicSynced(join(dataDir, 'global'), 'pg_control', Buffer.from(control));
|
||||
steps.push('pg_control restored');
|
||||
}
|
||||
|
||||
if (existsSync(backupWal)) {
|
||||
if (existsSync(walDir)) {
|
||||
const aside = join(backupPath, `pg_wal.reset-aside-${Date.now()}`);
|
||||
await rename(walDir, aside);
|
||||
steps.push(`reset pg_wal set aside at ${aside}`);
|
||||
}
|
||||
await rename(backupWal, walDir);
|
||||
steps.push('pg_wal restored');
|
||||
}
|
||||
if (steps.length === 0) {
|
||||
// A missing/empty backup means nothing was put back — never claim
|
||||
// restoration that did not happen (the 'failed-not-restored' honesty arm).
|
||||
return { restored: false, steps, detail: `nothing to restore from backup at ${backupPath} (missing or empty)` };
|
||||
}
|
||||
return { restored: true, steps };
|
||||
} catch (err) {
|
||||
return {
|
||||
restored: false,
|
||||
steps,
|
||||
detail: `restore failed after [${steps.join(', ') || 'nothing'}]: ${String((err as Error)?.message ?? err)}`,
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The engine seam. NEVER throws. Gates (in order): kill-switch → live-writer
|
||||
* (reaped lock) → layout validation → cooldown. Then: repair (reusing the open
|
||||
* episode's backup when present) → retry create() ONCE → on failure, restore
|
||||
* and record. Prints a repair-start stderr line the instant surgery begins so
|
||||
* a timeout-killed attempt is self-explaining (eng-review 4A).
|
||||
*/
|
||||
export async function attemptWalRepairAndRetry<T>(
|
||||
dataDir: string,
|
||||
retryCreate: () => Promise<T>,
|
||||
opts?: { reaped?: boolean },
|
||||
): Promise<WalRepairAttempt<T>> {
|
||||
try {
|
||||
if (!walRepairEnabled()) {
|
||||
return { status: 'skipped', reason: 'disabled', detail: 'GBRAIN_PGLITE_WAL_REPAIR=off' };
|
||||
}
|
||||
if (opts?.reaped) {
|
||||
return {
|
||||
status: 'skipped',
|
||||
reason: 'possibly-live-writer',
|
||||
detail:
|
||||
'this process acquired the data-dir lock by reaping a prior holder — ' +
|
||||
'another gbrain process may still be using this brain. Stop it (or confirm ' +
|
||||
'none is running), then re-run; a cleanly-acquired lock enables auto-repair.',
|
||||
};
|
||||
}
|
||||
const sinceReap = msSinceLastReap(dataDir);
|
||||
// `>= 0` guard (adversarial review F5): a future-dated marker (clock step
|
||||
// during the unclean-reboot recovery this feature exists for) yields a
|
||||
// negative age; treat it as expired rather than quarantining forever —
|
||||
// same policy as repairCooldownActive.
|
||||
if (sinceReap !== null && sinceReap >= 0 && sinceReap < REAP_QUARANTINE_MS) {
|
||||
return {
|
||||
status: 'skipped',
|
||||
reason: 'possibly-live-writer',
|
||||
detail:
|
||||
`a lock on this brain was reaped ${Math.round(sinceReap / 1000)}s ago (possibly from a ` +
|
||||
'still-live process) — auto-repair stays off for ' +
|
||||
`${REAP_QUARANTINE_MS / 60000} minutes after any reap. Confirm no gbrain process is ` +
|
||||
'running, then re-run or use `gbrain pglite-repair`.',
|
||||
};
|
||||
}
|
||||
const validation = validateWalRepairTarget(dataDir);
|
||||
if (!validation.ok) {
|
||||
return { status: 'skipped', reason: 'validation-failed', detail: validation.detail };
|
||||
}
|
||||
const cooldown = repairCooldownActive(dataDir);
|
||||
if (cooldown.active) {
|
||||
return { status: 'skipped', reason: 'recently-failed', detail: cooldown.detail };
|
||||
}
|
||||
|
||||
try {
|
||||
process.stderr.write(
|
||||
`gbrain: PGLite failed to open ${dataDir} — attempting automatic WAL repair ` +
|
||||
`(backup at ${dataDir}${BACKUP_DIR_MARKER}*). If this command times out, run ` +
|
||||
`\`gbrain pglite-repair\` to finish. Disable auto-repair with GBRAIN_PGLITE_WAL_REPAIR=off.\n`,
|
||||
);
|
||||
} catch { /* EPIPE under a closed-pipe daemon parent must not read as surgery failure */ }
|
||||
|
||||
const sidecar = readRepairSidecar(dataDir);
|
||||
// Stale-episode bound (red-team): an episode left open for a long time
|
||||
// means the pinned backup may predate real data — take a fresh backup.
|
||||
const episodeFresh =
|
||||
sidecar.episodeStartedAt !== null &&
|
||||
Date.now() - sidecar.episodeStartedAt >= 0 &&
|
||||
Date.now() - sidecar.episodeStartedAt < 24 * 3600 * 1000;
|
||||
let receipt: WalRepairReceipt;
|
||||
try {
|
||||
receipt = await repairPgliteWal(dataDir, {
|
||||
reuseBackupPath: episodeFresh ? sidecar.episodeBackupPath ?? undefined : undefined,
|
||||
});
|
||||
} catch (err) {
|
||||
if (err instanceof WalRepairError) {
|
||||
// Reset failed AFTER the backup was taken; report the best-effort
|
||||
// restore's REAL outcome (hardcoding restored:true here made the
|
||||
// 'failed-restored' message lie when the restore itself failed).
|
||||
recordRepairAttempt(dataDir, 'failed', err.receipt.backupPath);
|
||||
return {
|
||||
status: 'failed',
|
||||
receipt: err.receipt,
|
||||
restored: err.restore.restored,
|
||||
repairError: err.message +
|
||||
(err.restore.restored ? '' : ` [restore: ${err.restore.detail}]`),
|
||||
};
|
||||
}
|
||||
// Pre-backup refusal (validation) — the dir was never touched, so there
|
||||
// is nothing to restore and `restored: true` reads as "dir intact".
|
||||
recordRepairAttempt(dataDir, 'failed', sidecar.episodeBackupPath);
|
||||
return {
|
||||
status: 'failed',
|
||||
receipt: null,
|
||||
restored: true,
|
||||
repairError: String((err as Error)?.message ?? err),
|
||||
};
|
||||
}
|
||||
|
||||
try {
|
||||
const db = await retryCreate();
|
||||
recordRepairAttempt(dataDir, 'repaired', receipt.backupPath);
|
||||
return { status: 'repaired', db, receipt };
|
||||
} catch (retryErr) {
|
||||
const restore = await restoreWalBackup(receipt);
|
||||
recordRepairAttempt(dataDir, 'failed', receipt.backupPath);
|
||||
return {
|
||||
status: 'failed',
|
||||
receipt,
|
||||
restored: restore.restored,
|
||||
repairError: String((retryErr as Error)?.message ?? retryErr) +
|
||||
(restore.restored ? '' : ` [restore: ${restore.detail}]`),
|
||||
};
|
||||
}
|
||||
} catch (err) {
|
||||
// The seam's never-throw contract is load-bearing (single lock-release site
|
||||
// in connect()'s catch) — any unexpected error degrades to 'failed'.
|
||||
// `restored: true` here is honest: repairPgliteWal/restoreWalBackup handle
|
||||
// their own mutation failures via WalRepairError above, so a throw landing
|
||||
// HERE happened outside surgery and the dir is untouched (red-team: the
|
||||
// old restored:false told users to manually restore a nonexistent backup).
|
||||
try { recordRepairAttempt(dataDir, 'failed', null); } catch { /* best-effort — cooldown still engages when possible */ }
|
||||
return {
|
||||
status: 'failed',
|
||||
receipt: null,
|
||||
restored: true,
|
||||
repairError: `unexpected repair-path error: ${String((err as Error)?.message ?? err)}`,
|
||||
};
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,354 @@
|
||||
/**
|
||||
* pg_resetwal for PGLite NodeFS data dirs, in TypeScript.
|
||||
*
|
||||
* Ported from electric-sql/pglite PR #994 by @yestheboxer (Apache-2.0,
|
||||
* https://github.com/electric-sql/pglite/pull/994 — closed upstream as
|
||||
* "should be a separate tool"; gbrain is that tool). The byte surgery is
|
||||
* regression-tested upstream (create DB → insert → corrupt WAL → reopen →
|
||||
* row readable) and deliberately NOT "improved" here.
|
||||
*
|
||||
* What it does: for a data dir whose WAL/checkpoint state is torn (unclean
|
||||
* shutdown — the #223/#1670/#2575 class), rewrite `global/pg_control` with a
|
||||
* fresh shutdown checkpoint and emit one replacement WAL segment containing
|
||||
* that checkpoint record, so Postgres-in-WASM can start without replaying the
|
||||
* torn WAL. Data files are preserved; transactions not checkpointed before
|
||||
* the corruption are lost (the standard pg_resetwal caveat).
|
||||
*
|
||||
* LAYOUT COUPLING: the `OFF` table below is the PostgreSQL 17 ControlFileData
|
||||
* layout (`PG_CONTROL_VERSION` 1700) that @electric-sql/pglite 0.4.x ships.
|
||||
* Any pglite bump PAST PG17 must revisit this file together with the
|
||||
* `./vector` export blocker — see the TODOS.md "pglite upgrade blocker" entry.
|
||||
* Every unsupported shape throws `WalResetUnsupportedError` (fail-closed).
|
||||
*
|
||||
* gbrain adaptation on top of the upstream port: both file writes (the new
|
||||
* WAL segment AND pg_control) go through tmp-file + fsync(tmp) + rename +
|
||||
* fsync(parent dir) instead of upstream's in-place `writeFileSynced` — a
|
||||
* mid-write kill (SIGKILL, `process.exit(124)` from a read-only command
|
||||
* timeout) leaves old-or-new per file, never a torn file. Write ORDER stays
|
||||
* upstream's (segment first, control last): the pair is not atomic across
|
||||
* files, but a kill between them leaves pg_control pointing at old state, so
|
||||
* startup still fails and the next repair attempt re-runs this (idempotent).
|
||||
* A torn pair can never claim success.
|
||||
*/
|
||||
import { existsSync } from 'node:fs';
|
||||
import { mkdir, open, readdir, readFile, rename, unlink } from 'node:fs/promises';
|
||||
import { join } from 'node:path';
|
||||
|
||||
// Exported: pglite-repair.ts validates against the same layout literals — the
|
||||
// PG17 coupling the TODOS "pglite upgrade blocker" entry says moves together.
|
||||
export const PG_CONTROL_FILE_SIZE = 8192;
|
||||
const WAL_SEGMENT_RE = /^[0-9A-F]{24}(?:\.partial)?$/;
|
||||
/** Is this filename a WAL segment (incl. `.partial`)? Shared layout predicate. */
|
||||
export function isWalSegmentName(name: string): boolean {
|
||||
return WAL_SEGMENT_RE.test(name);
|
||||
}
|
||||
const PG_CONTROL_VERSION = 1700;
|
||||
const DB_SHUTDOWNED = 1;
|
||||
const XLOG_BLCKSZ = 8192;
|
||||
const MIN_WAL_SEG_SIZE = 1024 * 1024;
|
||||
// Postgres-general max is 1GB, but this port targets pglite (ships 16MB
|
||||
// segments). A corrupt-but-plausible control field must not be able to drive
|
||||
// a 1GB zero-fill allocation + write on the repair path (perf review) — cap
|
||||
// at 64MB and fail closed above it.
|
||||
const MAX_WAL_SEG_SIZE = 64 * 1024 * 1024;
|
||||
const SIZE_OF_XLOG_LONG_PHD = 40;
|
||||
const SIZE_OF_XLOG_RECORD = 24;
|
||||
const SIZE_OF_CHECKPOINT = 88;
|
||||
const XLOG_PAGE_MAGIC = 0xd116;
|
||||
const XLP_LONG_HEADER = 0x0002;
|
||||
const XLOG_CHECKPOINT_SHUTDOWN = 0x00;
|
||||
const XLR_BLOCK_ID_DATA_SHORT = 255;
|
||||
const RM_XLOG_ID = 0;
|
||||
|
||||
// PostgreSQL 17 ControlFileData offsets (pglite 0.4.x layout).
|
||||
const OFF = {
|
||||
systemIdentifier: 0,
|
||||
pgControlVersion: 8,
|
||||
state: 16,
|
||||
time: 24,
|
||||
checkPoint: 32,
|
||||
checkPointCopy: 40,
|
||||
checkPointCopyRedo: 40,
|
||||
checkPointCopyThisTimeLineID: 48,
|
||||
checkPointCopyTime: 104,
|
||||
minRecoveryPoint: 136,
|
||||
minRecoveryPointTLI: 144,
|
||||
backupStartPoint: 152,
|
||||
backupEndPoint: 160,
|
||||
backupEndRequired: 168,
|
||||
walLevel: 172,
|
||||
walLogHints: 176,
|
||||
maxConnections: 180,
|
||||
maxWorkerProcesses: 184,
|
||||
maxWalSenders: 188,
|
||||
maxPreparedXacts: 192,
|
||||
maxLocksPerXact: 196,
|
||||
trackCommitTimestamp: 200,
|
||||
xlogBlcksz: 224,
|
||||
xlogSegSize: 228,
|
||||
crc: 288,
|
||||
} as const;
|
||||
|
||||
/** The dir does not look like a PG17 pglite layout — refuse to touch it. */
|
||||
export class WalResetUnsupportedError extends Error {
|
||||
constructor(message: string) {
|
||||
super(message);
|
||||
this.name = 'WalResetUnsupportedError';
|
||||
}
|
||||
}
|
||||
|
||||
export interface WalResetResult {
|
||||
/** Filename of the replacement WAL segment written (24 hex chars). */
|
||||
resetSegment: string;
|
||||
timelineId: number;
|
||||
walSegSize: number;
|
||||
}
|
||||
|
||||
const crcTable = new Uint32Array(256);
|
||||
for (let i = 0; i < 256; i++) {
|
||||
let crc = i;
|
||||
for (let j = 0; j < 8; j++) {
|
||||
crc = crc & 1 ? (crc >>> 1) ^ 0x82f63b78 : crc >>> 1;
|
||||
}
|
||||
crcTable[i] = crc >>> 0;
|
||||
}
|
||||
|
||||
export function crc32c(chunks: Uint8Array[]): number {
|
||||
let crc = 0xffffffff;
|
||||
for (const chunk of chunks) {
|
||||
for (const byte of chunk) {
|
||||
crc = (crc >>> 8) ^ crcTable[(crc ^ byte) & 0xff]!;
|
||||
}
|
||||
}
|
||||
return (crc ^ 0xffffffff) >>> 0;
|
||||
}
|
||||
|
||||
function readUInt64LE(buf: Buffer, offset: number): bigint {
|
||||
return buf.readBigUInt64LE(offset);
|
||||
}
|
||||
|
||||
function writeUInt64LE(buf: Buffer, value: bigint, offset: number): void {
|
||||
buf.writeBigUInt64LE(value, offset);
|
||||
}
|
||||
|
||||
export function parseWalSegNo(fileName: string, walSegSize: number): bigint | null {
|
||||
if (!/^[0-9A-F]{24}$/.test(fileName)) return null;
|
||||
const log = BigInt(`0x${fileName.slice(8, 16)}`);
|
||||
const seg = BigInt(`0x${fileName.slice(16, 24)}`);
|
||||
return log * (0x100000000n / BigInt(walSegSize)) + seg;
|
||||
}
|
||||
|
||||
export function xlogFileName(tli: number, segNo: bigint, walSegSize: number): string {
|
||||
const segmentsPerXlogId = 0x100000000n / BigInt(walSegSize);
|
||||
const log = segNo / segmentsPerXlogId;
|
||||
const seg = segNo % segmentsPerXlogId;
|
||||
return [
|
||||
tli.toString(16).toUpperCase().padStart(8, '0'),
|
||||
log.toString(16).toUpperCase().padStart(8, '0'),
|
||||
seg.toString(16).toUpperCase().padStart(8, '0'),
|
||||
].join('');
|
||||
}
|
||||
|
||||
async function unlinkIfExists(path: string): Promise<void> {
|
||||
await unlink(path).catch((error: NodeJS.ErrnoException) => {
|
||||
if (error.code !== 'ENOENT') throw error;
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Atomic + durable single-file write: tmp in the same dir → fsync(tmp) →
|
||||
* rename over the target → fsync(parent dir). A kill at any instant leaves
|
||||
* either the old file or the new file, never a torn one; the dir fsync makes
|
||||
* the rename itself survive power loss. (Directory fsync is best-effort —
|
||||
* some filesystems refuse it; the rename is still atomic without it.)
|
||||
*/
|
||||
export async function writeFileAtomicSynced(dir: string, name: string, data: Buffer): Promise<void> {
|
||||
const tmpName = `.${name}.tmp-${process.pid}`;
|
||||
const tmpPath = join(dir, tmpName);
|
||||
// 'wx' (exclusive create, never follows an existing symlink) after clearing
|
||||
// any stale tmp: a pre-planted symlink at the predictable tmp path must not
|
||||
// redirect the write (security review).
|
||||
await unlinkIfExists(tmpPath);
|
||||
const file = await open(tmpPath, 'wx');
|
||||
try {
|
||||
await file.writeFile(data);
|
||||
await file.sync();
|
||||
} finally {
|
||||
await file.close();
|
||||
}
|
||||
await rename(tmpPath, join(dir, name));
|
||||
try {
|
||||
const dirHandle = await open(dir, 'r');
|
||||
try {
|
||||
await dirHandle.sync();
|
||||
} finally {
|
||||
await dirHandle.close();
|
||||
}
|
||||
} catch {
|
||||
// Best-effort: some platforms/filesystems reject fsync on a directory fd.
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reset WAL in place for a PG17 pglite NodeFS data dir. Preserves data files;
|
||||
* discards the (torn) WAL tail. Throws `WalResetUnsupportedError` when the
|
||||
* layout does not match what this port understands; throws raw fs errors on
|
||||
* I/O failure. Idempotent: re-running after a partial attempt converges.
|
||||
*/
|
||||
export async function resetWal(rootDir: string): Promise<WalResetResult> {
|
||||
await unlinkIfExists(join(rootDir, 'postmaster.pid'));
|
||||
|
||||
let pgVersion: string;
|
||||
try {
|
||||
pgVersion = (await readFile(join(rootDir, 'PG_VERSION'), 'utf8')).trim();
|
||||
} catch {
|
||||
throw new WalResetUnsupportedError(`No readable PG_VERSION in ${rootDir}`);
|
||||
}
|
||||
if (pgVersion !== '17') {
|
||||
throw new WalResetUnsupportedError(`Cannot reset WAL for unsupported PG_VERSION ${pgVersion}`);
|
||||
}
|
||||
|
||||
const controlPath = join(rootDir, 'global', 'pg_control');
|
||||
let control: Buffer;
|
||||
try {
|
||||
control = Buffer.from(await readFile(controlPath));
|
||||
} catch {
|
||||
throw new WalResetUnsupportedError(`No readable global/pg_control in ${rootDir}`);
|
||||
}
|
||||
if (control.length !== PG_CONTROL_FILE_SIZE) {
|
||||
throw new WalResetUnsupportedError(`Unexpected pg_control size ${control.length}`);
|
||||
}
|
||||
if (control.readUInt32LE(OFF.pgControlVersion) !== PG_CONTROL_VERSION) {
|
||||
throw new WalResetUnsupportedError('Unsupported pg_control version');
|
||||
}
|
||||
// Verify the STORED CRC before trusting (and re-signing) the checkpoint copy
|
||||
// (adversarial review F6): a torn pg_control with an intact version field
|
||||
// but garbage checkpoint counters (nextXid/nextOid/...) would otherwise be
|
||||
// preserved verbatim and laundered under a fresh valid CRC — Postgres then
|
||||
// starts and corrupts silently (xid-wraparound class). Real pg_resetwal
|
||||
// refuses on CRC mismatch; so do we → the caller falls to the rebuild rung.
|
||||
const storedCrc = control.readUInt32LE(OFF.crc);
|
||||
if (crc32c([control.subarray(0, OFF.crc)]) !== storedCrc) {
|
||||
throw new WalResetUnsupportedError(
|
||||
'pg_control CRC mismatch — the control file itself is damaged; WAL reset ' +
|
||||
'would launder corrupt checkpoint counters. Rebuild the brain instead ' +
|
||||
'(`gbrain reinit-pglite`).',
|
||||
);
|
||||
}
|
||||
|
||||
const walSegSize = control.readUInt32LE(OFF.xlogSegSize);
|
||||
const xlogBlcksz = control.readUInt32LE(OFF.xlogBlcksz);
|
||||
if (
|
||||
walSegSize < MIN_WAL_SEG_SIZE ||
|
||||
walSegSize > MAX_WAL_SEG_SIZE ||
|
||||
(walSegSize & (walSegSize - 1)) !== 0 ||
|
||||
0x100000000 % walSegSize !== 0
|
||||
) {
|
||||
throw new WalResetUnsupportedError(`Unsupported WAL segment size ${walSegSize}`);
|
||||
}
|
||||
if (xlogBlcksz !== XLOG_BLCKSZ) {
|
||||
throw new WalResetUnsupportedError(`Unsupported WAL block size ${xlogBlcksz}`);
|
||||
}
|
||||
|
||||
const tli = control.readUInt32LE(OFF.checkPointCopyThisTimeLineID);
|
||||
let newSegNo = readUInt64LE(control, OFF.checkPointCopyRedo) / BigInt(walSegSize);
|
||||
const walDir = join(rootDir, 'pg_wal');
|
||||
// Recreates pg_wal/archive_status when the whole pg_wal dir was renamed
|
||||
// away into the repair backup (pglite-repair.ts) — resetWal then starts
|
||||
// from an empty WAL dir and numbers the fresh segment off pg_control alone.
|
||||
await mkdir(join(walDir, 'archive_status'), { recursive: true });
|
||||
for (const file of await readdir(walDir)) {
|
||||
const segNo = parseWalSegNo(file, walSegSize);
|
||||
if (segNo !== null && segNo > newSegNo) {
|
||||
newSegNo = segNo;
|
||||
}
|
||||
}
|
||||
newSegNo += 1n;
|
||||
|
||||
const redo = newSegNo * BigInt(walSegSize) + BigInt(SIZE_OF_XLOG_LONG_PHD);
|
||||
const now = BigInt(Math.floor(Date.now() / 1000));
|
||||
|
||||
writeUInt64LE(control, redo, OFF.checkPointCopyRedo);
|
||||
writeUInt64LE(control, now, OFF.checkPointCopyTime);
|
||||
control.writeInt32LE(DB_SHUTDOWNED, OFF.state);
|
||||
writeUInt64LE(control, now, OFF.time);
|
||||
writeUInt64LE(control, redo, OFF.checkPoint);
|
||||
writeUInt64LE(control, 0n, OFF.minRecoveryPoint);
|
||||
control.writeUInt32LE(0, OFF.minRecoveryPointTLI);
|
||||
writeUInt64LE(control, 0n, OFF.backupStartPoint);
|
||||
writeUInt64LE(control, 0n, OFF.backupEndPoint);
|
||||
control.writeUInt8(0, OFF.backupEndRequired);
|
||||
control.writeInt32LE(0, OFF.walLevel);
|
||||
control.writeUInt8(0, OFF.walLogHints);
|
||||
control.writeInt32LE(100, OFF.maxConnections);
|
||||
control.writeInt32LE(8, OFF.maxWorkerProcesses);
|
||||
control.writeInt32LE(10, OFF.maxWalSenders);
|
||||
control.writeInt32LE(0, OFF.maxPreparedXacts);
|
||||
control.writeInt32LE(64, OFF.maxLocksPerXact);
|
||||
control.writeUInt8(0, OFF.trackCommitTimestamp);
|
||||
control.writeUInt32LE(crc32c([control.subarray(0, OFF.crc)]), OFF.crc);
|
||||
|
||||
for (const file of await readdir(walDir)) {
|
||||
if (isWalSegmentName(file)) {
|
||||
await unlink(join(walDir, file));
|
||||
}
|
||||
}
|
||||
|
||||
const archiveStatusDir = join(walDir, 'archive_status');
|
||||
if (existsSync(archiveStatusDir)) {
|
||||
for (const file of await readdir(archiveStatusDir)) {
|
||||
if (/^[0-9A-F]{24}(?:\.partial)?\.(?:ready|done)$/.test(file)) {
|
||||
await unlink(join(archiveStatusDir, file));
|
||||
}
|
||||
}
|
||||
}
|
||||
const walSummaryDir = join(walDir, 'summaries');
|
||||
if (existsSync(walSummaryDir)) {
|
||||
for (const file of await readdir(walSummaryDir)) {
|
||||
if (/^[0-9A-F]{40}\.summary$/.test(file)) {
|
||||
await unlink(join(walSummaryDir, file));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const wal = Buffer.alloc(walSegSize);
|
||||
wal.writeUInt16LE(XLOG_PAGE_MAGIC, 0);
|
||||
wal.writeUInt16LE(XLP_LONG_HEADER, 2);
|
||||
wal.writeUInt32LE(tli, 4);
|
||||
writeUInt64LE(wal, redo - BigInt(SIZE_OF_XLOG_LONG_PHD), 8);
|
||||
wal.writeUInt32LE(0, 16);
|
||||
writeUInt64LE(wal, readUInt64LE(control, OFF.systemIdentifier), 24);
|
||||
wal.writeUInt32LE(walSegSize, 32);
|
||||
wal.writeUInt32LE(XLOG_BLCKSZ, 36);
|
||||
|
||||
const recordOffset = SIZE_OF_XLOG_LONG_PHD;
|
||||
const recordTotalLength = SIZE_OF_XLOG_RECORD + 2 + SIZE_OF_CHECKPOINT;
|
||||
wal.writeUInt32LE(recordTotalLength, recordOffset);
|
||||
wal.writeUInt32LE(0, recordOffset + 4);
|
||||
writeUInt64LE(wal, 0n, recordOffset + 8);
|
||||
wal.writeUInt8(XLOG_CHECKPOINT_SHUTDOWN, recordOffset + 16);
|
||||
wal.writeUInt8(RM_XLOG_ID, recordOffset + 17);
|
||||
wal.writeUInt16LE(0, recordOffset + 18);
|
||||
wal.writeUInt8(XLR_BLOCK_ID_DATA_SHORT, recordOffset + SIZE_OF_XLOG_RECORD);
|
||||
wal.writeUInt8(SIZE_OF_CHECKPOINT, recordOffset + SIZE_OF_XLOG_RECORD + 1);
|
||||
control.copy(
|
||||
wal,
|
||||
recordOffset + SIZE_OF_XLOG_RECORD + 2,
|
||||
OFF.checkPointCopy,
|
||||
OFF.checkPointCopy + SIZE_OF_CHECKPOINT,
|
||||
);
|
||||
|
||||
const record = wal.subarray(recordOffset, recordOffset + recordTotalLength);
|
||||
const recordCrc = crc32c([
|
||||
record.subarray(SIZE_OF_XLOG_RECORD),
|
||||
record.subarray(0, 20),
|
||||
]);
|
||||
wal.writeUInt32LE(recordCrc, recordOffset + 20);
|
||||
|
||||
const resetSegment = xlogFileName(tli, newSegNo, walSegSize);
|
||||
// Segment FIRST, control LAST (upstream order — see header comment).
|
||||
await writeFileAtomicSynced(walDir, resetSegment, wal);
|
||||
await writeFileAtomicSynced(join(rootDir, 'global'), 'pg_control', control);
|
||||
|
||||
return { resetSegment, timelineId: tli, walSegSize };
|
||||
}
|
||||
@@ -0,0 +1,240 @@
|
||||
/**
|
||||
* `doctor` pglite_data_dir check (#223 WAL-repair wave) — parallel-safe unit
|
||||
* coverage of:
|
||||
*
|
||||
* - `computePgliteDataDirCheck` — the PURE verdict → Check mapping in
|
||||
* src/commands/doctor.ts (synthetic PgliteDirDiagnosis inputs, no fs).
|
||||
* - `OPS_CHECK_NAMES` carrying 'pglite_data_dir' (category routing).
|
||||
* - `inspectPgliteDataDir` — the read-only diagnoser in
|
||||
* src/core/pglite-repair.ts, exercised against synthetic on-disk layouts
|
||||
* in hermetic mkdtemp dirs.
|
||||
*
|
||||
* No process.env writes, no PGLite cold starts — safe for the parallel unit
|
||||
* shards. The command-level (real repair) coverage lives in
|
||||
* test/pglite-repair-command.serial.test.ts.
|
||||
*/
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, symlinkSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
|
||||
import { computePgliteDataDirCheck } from '../src/commands/doctor.ts';
|
||||
import { OPS_CHECK_NAMES } from '../src/core/doctor-categories.ts';
|
||||
import { inspectPgliteDataDir } from '../src/core/pglite-repair.ts';
|
||||
import type { PgliteDirDiagnosis } from '../src/core/pglite-repair.ts';
|
||||
|
||||
function tmp(prefix: string): string {
|
||||
return mkdtempSync(join(tmpdir(), prefix));
|
||||
}
|
||||
|
||||
/** Synthetic diagnosis with healthy-looking defaults; override per case. */
|
||||
function makeDiagnosis(
|
||||
overrides: Partial<PgliteDirDiagnosis> & { verdict: PgliteDirDiagnosis['verdict'] },
|
||||
): PgliteDirDiagnosis {
|
||||
return {
|
||||
exists: true,
|
||||
postmasterPid: false,
|
||||
pgControlOk: true,
|
||||
pgVersion: '17',
|
||||
walSegments: [],
|
||||
lockHeld: false,
|
||||
lockHolderPid: null,
|
||||
backupDirs: [],
|
||||
recentAttempts: [],
|
||||
detail: 'synthetic diagnosis',
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* A minimal fake PG17 pglite layout that passes `validateWalRepairTarget`:
|
||||
* PG_VERSION '17', base/ dir, 8192-byte global/pg_control, empty pg_wal/.
|
||||
*/
|
||||
function makeFakeLayout(dir: string): void {
|
||||
mkdirSync(join(dir, 'base'), { recursive: true });
|
||||
mkdirSync(join(dir, 'global'), { recursive: true });
|
||||
mkdirSync(join(dir, 'pg_wal'), { recursive: true });
|
||||
writeFileSync(join(dir, 'PG_VERSION'), '17\n');
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), Buffer.alloc(8192));
|
||||
}
|
||||
|
||||
describe('computePgliteDataDirCheck — verdict → Check mapping', () => {
|
||||
const DIR = '/synthetic/brain.pglite';
|
||||
|
||||
test('every verdict maps to a check named pglite_data_dir with human_only remediation', () => {
|
||||
const verdicts: PgliteDirDiagnosis['verdict'][] = [
|
||||
'looks-healthy',
|
||||
'wal-corruption-likely',
|
||||
'locked',
|
||||
'missing',
|
||||
'unsupported-layout',
|
||||
];
|
||||
for (const verdict of verdicts) {
|
||||
const check = computePgliteDataDirCheck(DIR, makeDiagnosis({ verdict }));
|
||||
expect(check.name).toBe('pglite_data_dir');
|
||||
expect(check.remediation_status).toBe('human_only');
|
||||
}
|
||||
});
|
||||
|
||||
test('wal-corruption-likely → fail, names pglite-repair --dry-run and #223', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({
|
||||
verdict: 'wal-corruption-likely',
|
||||
postmasterPid: true,
|
||||
detail: 'stale postmaster.pid present',
|
||||
}),
|
||||
);
|
||||
expect(check.status).toBe('fail');
|
||||
expect(check.message).toContain('pglite-repair --dry-run');
|
||||
expect(check.message).toContain('#223');
|
||||
expect(check.message).toContain('gbrain pglite-repair');
|
||||
});
|
||||
|
||||
test('looks-healthy (but connect failed) → fail, points at gbrain pglite-repair', () => {
|
||||
const check = computePgliteDataDirCheck(DIR, makeDiagnosis({ verdict: 'looks-healthy' }));
|
||||
expect(check.status).toBe('fail');
|
||||
expect(check.message).toContain('gbrain pglite-repair');
|
||||
});
|
||||
|
||||
test('unsupported-layout → fail', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({ verdict: 'unsupported-layout', pgControlOk: false, pgVersion: null }),
|
||||
);
|
||||
expect(check.status).toBe('fail');
|
||||
expect(check.remediation_status).toBe('human_only');
|
||||
});
|
||||
|
||||
test('locked → warn, names the live holder PID', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({ verdict: 'locked', lockHeld: true, lockHolderPid: 12345 }),
|
||||
);
|
||||
expect(check.status).toBe('warn');
|
||||
expect(check.message).toContain('12345');
|
||||
});
|
||||
|
||||
test('missing → warn', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({ verdict: 'missing', exists: false, pgControlOk: false, pgVersion: null }),
|
||||
);
|
||||
expect(check.status).toBe('warn');
|
||||
expect(check.message).toContain(DIR);
|
||||
});
|
||||
|
||||
test('recurrence: >=2 failed attempts within 7 days → message points at docs/ENGINES.md', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({
|
||||
verdict: 'wal-corruption-likely',
|
||||
postmasterPid: true,
|
||||
recentAttempts: [
|
||||
{ ts: Date.now() - 1000, outcome: 'failed' },
|
||||
{ ts: Date.now() - 2000, outcome: 'failed' },
|
||||
],
|
||||
}),
|
||||
);
|
||||
expect(check.status).toBe('fail');
|
||||
expect(check.message).toContain('docs/ENGINES.md');
|
||||
});
|
||||
|
||||
test('single recent failed attempt does NOT trigger the engine-switch recurrence note', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({
|
||||
verdict: 'wal-corruption-likely',
|
||||
postmasterPid: true,
|
||||
recentAttempts: [{ ts: Date.now() - 1000, outcome: 'failed' }],
|
||||
}),
|
||||
);
|
||||
expect(check.message).not.toContain('docs/ENGINES.md');
|
||||
});
|
||||
|
||||
test('non-empty backupDirs → message mentions the backup(s)', () => {
|
||||
const check = computePgliteDataDirCheck(
|
||||
DIR,
|
||||
makeDiagnosis({
|
||||
verdict: 'wal-corruption-likely',
|
||||
postmasterPid: true,
|
||||
backupDirs: [`${DIR}.wal-repair-backup-1700000000000`],
|
||||
}),
|
||||
);
|
||||
expect(check.message).toContain('backup');
|
||||
expect(check.message).toContain(`${DIR}.wal-repair-backup-1700000000000`);
|
||||
});
|
||||
});
|
||||
|
||||
describe('doctor-categories', () => {
|
||||
test('OPS_CHECK_NAMES contains pglite_data_dir', () => {
|
||||
expect(OPS_CHECK_NAMES.has('pglite_data_dir')).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('inspectPgliteDataDir — synthetic on-disk layouts', () => {
|
||||
test('(a) nonexistent path → missing', () => {
|
||||
const parent = tmp('gbrain-inspect-a-');
|
||||
const diagnosis = inspectPgliteDataDir(join(parent, 'does-not-exist.pglite'));
|
||||
expect(diagnosis.verdict).toBe('missing');
|
||||
expect(diagnosis.exists).toBe(false);
|
||||
});
|
||||
|
||||
test('(b) dir without PG_VERSION → unsupported-layout', () => {
|
||||
const dir = tmp('gbrain-inspect-b-');
|
||||
const diagnosis = inspectPgliteDataDir(dir);
|
||||
expect(diagnosis.verdict).toBe('unsupported-layout');
|
||||
expect(diagnosis.pgVersion).toBeNull();
|
||||
});
|
||||
|
||||
test('(c) full fake layout + stale postmaster.pid → wal-corruption-likely', () => {
|
||||
const dir = join(tmp('gbrain-inspect-c-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
writeFileSync(join(dir, 'postmaster.pid'), '99999\n');
|
||||
const diagnosis = inspectPgliteDataDir(dir);
|
||||
expect(diagnosis.verdict).toBe('wal-corruption-likely');
|
||||
expect(diagnosis.postmasterPid).toBe(true);
|
||||
expect(diagnosis.pgVersion).toBe('17');
|
||||
expect(diagnosis.pgControlOk).toBe(true);
|
||||
});
|
||||
|
||||
test('(d) full fake layout, no postmaster.pid → looks-healthy', () => {
|
||||
const dir = join(tmp('gbrain-inspect-d-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
const diagnosis = inspectPgliteDataDir(dir);
|
||||
expect(diagnosis.verdict).toBe('looks-healthy');
|
||||
expect(diagnosis.postmasterPid).toBe(false);
|
||||
expect(diagnosis.lockHeld).toBe(false);
|
||||
});
|
||||
|
||||
test('(e) fake layout + live .gbrain-lock holder → locked with the holder PID', () => {
|
||||
const dir = join(tmp('gbrain-inspect-e-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
mkdirSync(join(dir, '.gbrain-lock'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(dir, '.gbrain-lock', 'lock'),
|
||||
JSON.stringify({
|
||||
pid: process.pid, // this test process — provably alive
|
||||
acquired_at: Date.now(),
|
||||
refreshed_at: Date.now(),
|
||||
command: 'gbrain embed',
|
||||
subcommand: 'embed',
|
||||
}),
|
||||
);
|
||||
const diagnosis = inspectPgliteDataDir(dir);
|
||||
expect(diagnosis.verdict).toBe('locked');
|
||||
expect(diagnosis.lockHeld).toBe(true);
|
||||
expect(diagnosis.lockHolderPid).toBe(process.pid);
|
||||
});
|
||||
|
||||
test('(f) symlinked data dir → unsupported-layout (rename-based repair refuses symlinks)', () => {
|
||||
const parent = tmp('gbrain-inspect-f-');
|
||||
const real = join(parent, 'real.pglite');
|
||||
makeFakeLayout(real);
|
||||
const link = join(parent, 'link.pglite');
|
||||
symlinkSync(real, link);
|
||||
const diagnosis = inspectPgliteDataDir(link);
|
||||
expect(diagnosis.verdict).toBe('unsupported-layout');
|
||||
expect(diagnosis.detail).toContain('symlink');
|
||||
});
|
||||
});
|
||||
@@ -33,6 +33,7 @@
|
||||
|
||||
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
||||
import { spawn, spawnSync } from 'child_process';
|
||||
import { randomBytes } from 'crypto';
|
||||
import {
|
||||
cpSync,
|
||||
mkdirSync,
|
||||
@@ -96,9 +97,12 @@ beforeAll(() => {
|
||||
delete runEnv.ANTHROPIC_API_KEY;
|
||||
delete runEnv.GOOGLE_API_KEY;
|
||||
|
||||
// NOTE: init grew strict flag validation (#2201); `--repo`/`--yes` were
|
||||
// never real init flags (previously silently ignored). The repo is wired
|
||||
// through `sync --repo` below, matching the other e2e suites.
|
||||
const initResult = spawnSync(
|
||||
SHIM_PATH,
|
||||
['init', '--pglite', '--repo', repoSourceDir, '--no-embedding', '--yes'],
|
||||
['init', '--pglite', '--no-embedding', '--non-interactive'],
|
||||
{
|
||||
cwd: REPO_ROOT,
|
||||
env: runEnv,
|
||||
@@ -409,6 +413,49 @@ describe('#2084 — explicit-exit teardown: every swept site exits clean, exit c
|
||||
}, 30_000);
|
||||
});
|
||||
|
||||
describe('WAL-repair wave — corrupt persistent brain, auto-repair off: owned exit 1 (#2084 class)', () => {
|
||||
test('gbrain status on a torn-WAL brain with GBRAIN_PGLITE_WAL_REPAIR=off exits 1 (not 0, not 99)', async () => {
|
||||
// Fixture: a fake-but-layout-valid PG17 pglite data dir whose control +
|
||||
// WAL state is garbage, so PGlite.create aborts. With auto-repair
|
||||
// disabled the CLI must fail LOUDLY through the owned verdict channel:
|
||||
// real process exit 1 — never 0 (silent success over a broken brain),
|
||||
// never 99 (Emscripten's hijacked process.exitCode, the #2084 class).
|
||||
const corruptHome = mkdtempSync(join(tmpdir(), 'gbrain-pglite-corrupt-'));
|
||||
try {
|
||||
const dataDir = join(corruptHome, 'brain.pglite');
|
||||
// GBRAIN_HOME is a parent dir: config lives at <home>/.gbrain/config.json.
|
||||
mkdirSync(join(corruptHome, '.gbrain'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(corruptHome, '.gbrain', 'config.json'),
|
||||
JSON.stringify({ engine: 'pglite', database_path: dataDir }, null, 2) + '\n',
|
||||
'utf-8',
|
||||
);
|
||||
mkdirSync(join(dataDir, 'base'), { recursive: true });
|
||||
mkdirSync(join(dataDir, 'global'), { recursive: true });
|
||||
mkdirSync(join(dataDir, 'pg_wal'), { recursive: true });
|
||||
writeFileSync(join(dataDir, 'PG_VERSION'), '17\n', 'utf-8');
|
||||
writeFileSync(join(dataDir, 'global', 'pg_control'), randomBytes(8192));
|
||||
writeFileSync(join(dataDir, 'pg_wal', '000000010000000000000001'), randomBytes(1024));
|
||||
|
||||
const { code, stdout, stderr, durationMs } = await runWithTimeout(
|
||||
['status'],
|
||||
30_000,
|
||||
{ GBRAIN_HOME: corruptHome, GBRAIN_PGLITE_WAL_REPAIR: 'off' },
|
||||
);
|
||||
if (code !== 1) {
|
||||
throw new Error(
|
||||
`expected exit 1, got ${code}; duration=${durationMs}ms\n` +
|
||||
`STDOUT:\n${stdout}\nSTDERR:\n${stderr}`,
|
||||
);
|
||||
}
|
||||
expect(code).toBe(1);
|
||||
expect(stdout + stderr).toContain('PGLite failed to initialize');
|
||||
} finally {
|
||||
rmSync(corruptHome, { recursive: true, force: true });
|
||||
}
|
||||
}, 60_000);
|
||||
});
|
||||
|
||||
describe('v0.41.8.0 — daemon survival (regression guard for narrow force-exit)', () => {
|
||||
test('gbrain serve --http stays alive past the timeout window', async () => {
|
||||
// Pick a likely-free ephemeral port. We're testing "still alive
|
||||
|
||||
@@ -247,7 +247,41 @@ describe('v0.41.8.0 #1340 — PGLite WASM init classifier', () => {
|
||||
test('pglite-engine.ts connect catch block routes through the classifier', () => {
|
||||
const src = readFileSync('src/core/pglite-engine.ts', 'utf8');
|
||||
expect(src).toMatch(/classifyPgliteInitError\(original\)/);
|
||||
expect(src).toMatch(/buildPgliteInitErrorMessage\(verdict, original\)/);
|
||||
// WAL-repair wave: the call gained platform + repair-context args, so pin
|
||||
// only the (verdict, original, …) prefix — the routing seam, not the arity.
|
||||
expect(src).toMatch(/buildPgliteInitErrorMessage\(verdict, original/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('WAL-repair wave structural pins (#223/#2575)', () => {
|
||||
test('connect() catch wires the WAL auto-repair seam', () => {
|
||||
const src = readFileSync('src/core/pglite-engine.ts', 'utf8');
|
||||
expect(src).toMatch(/attemptWalRepairAndRetry\(/);
|
||||
});
|
||||
|
||||
test('the repair-retry lambda stays inside the #2084 exitCode guard', () => {
|
||||
// The retry re-runs PGlite.create; unguarded, Emscripten would hijack
|
||||
// process.exitCode on the retry path exactly as it did on the first
|
||||
// attempt (the #2084 class). Pin the wrap at the seam call-site.
|
||||
const src = readFileSync('src/core/pglite-engine.ts', 'utf8');
|
||||
expect(src).toMatch(/attemptWalRepairAndRetry\([\s\S]{0,300}preservingProcessExitCode/);
|
||||
});
|
||||
|
||||
test('no bare PGlite.create outside the wrapped engine sites', () => {
|
||||
// The repair/resetwal modules take the retry as a callback — if either
|
||||
// grew its own PGlite.create call it would bypass BOTH the exitCode
|
||||
// guard and the single-writer lock.
|
||||
const repair = readFileSync('src/core/pglite-repair.ts', 'utf8');
|
||||
const resetwal = readFileSync('src/core/pglite-resetwal.ts', 'utf8');
|
||||
expect(repair).not.toMatch(/PGlite\.create/);
|
||||
expect(resetwal).not.toMatch(/PGlite\.create/);
|
||||
});
|
||||
|
||||
test('pglite-resetwal.ts carries the upstream attribution', () => {
|
||||
// The reset-WAL sequence mirrors upstream electric-sql/pglite PR #994;
|
||||
// the pointer is the audit trail for future divergence.
|
||||
const resetwal = readFileSync('src/core/pglite-resetwal.ts', 'utf8');
|
||||
expect(resetwal).toContain('electric-sql/pglite/pull/994');
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
@@ -31,9 +31,33 @@ describe('classifyPgliteInitError', () => {
|
||||
expect(classifyPgliteInitError(msg)).toBe('bunfs');
|
||||
});
|
||||
|
||||
test('macos-26-3 verdict for the existing #223 signature', () => {
|
||||
test('wasm-abort verdict for the existing #223 signature', () => {
|
||||
const msg = 'abort() called from wasm runtime on macOS 26.3 build';
|
||||
expect(classifyPgliteInitError(msg)).toBe('macos-26-3');
|
||||
expect(classifyPgliteInitError(msg)).toBe('wasm-abort');
|
||||
});
|
||||
|
||||
// WAL-repair wave: THE real production message from a torn-WAL Emscripten
|
||||
// abort — no "runtime"/"wasm" in it, so the legacy arms let it fall
|
||||
// through to 'unknown' (which is exactly how #223 got misdiagnosed).
|
||||
test('wasm-abort verdict for the bare Emscripten Aborted() message', () => {
|
||||
expect(
|
||||
classifyPgliteInitError('Aborted(). Build with -sASSERTIONS for more info.'),
|
||||
).toBe('wasm-abort');
|
||||
});
|
||||
|
||||
test('wasm-abort verdict for RuntimeError-prefixed Aborted()', () => {
|
||||
expect(
|
||||
classifyPgliteInitError('RuntimeError: Aborted(). Build with -sASSERTIONS for more info.'),
|
||||
).toBe('wasm-abort');
|
||||
});
|
||||
|
||||
test('wasm-abort verdict for the generic RuntimeError: unreachable trap', () => {
|
||||
expect(classifyPgliteInitError('RuntimeError: unreachable')).toBe('wasm-abort');
|
||||
});
|
||||
|
||||
test('bunfs still wins when a wasm-abort marker co-occurs (bunfs arm is first)', () => {
|
||||
const msg = "RuntimeError: Aborted(). ENOENT open '/$$bunfs/root/pglite.data'";
|
||||
expect(classifyPgliteInitError(msg)).toBe('bunfs');
|
||||
});
|
||||
|
||||
test('unknown verdict for generic / unrecognized errors', () => {
|
||||
@@ -64,8 +88,9 @@ describe('classifyPgliteInitError', () => {
|
||||
});
|
||||
|
||||
test('corrupt verdict beats the wasm-runtime match (58P01 wins over "wasm runtime")', () => {
|
||||
// A message mentioning both must classify as corrupt, not macos-26-3 —
|
||||
// recovery guidance, not the wrong macOS-WASM hint.
|
||||
// A message mentioning both must classify as corrupt, not wasm-abort —
|
||||
// recovery guidance, not the WAL-repair hint (WAL repair cannot fix
|
||||
// catalog corruption).
|
||||
expect(classifyPgliteInitError('wasm runtime: 58P01 internal_load_library')).toBe('corrupt');
|
||||
});
|
||||
});
|
||||
@@ -82,14 +107,67 @@ describe('buildPgliteInitErrorMessage — hint routing', () => {
|
||||
expect(msg).not.toContain('issues/223');
|
||||
});
|
||||
|
||||
test('macos-26-3 verdict surfaces the #223 link AND original error', () => {
|
||||
const msg = buildPgliteInitErrorMessage('macos-26-3', original);
|
||||
test('wasm-abort verdict names torn WAL as the cause, keeps the #223 link, AND original error', () => {
|
||||
const msg = buildPgliteInitErrorMessage('wasm-abort', original);
|
||||
// The re-diagnosis is the load-bearing copy: corrupt WAL after an unclean
|
||||
// shutdown, explicitly NOT the historical macOS-WASM attribution.
|
||||
expect(msg).toContain('NOT a macOS WASM bug');
|
||||
expect(msg).toContain('https://github.com/garrytan/gbrain/issues/223');
|
||||
expect(msg).toContain('macOS 26.3');
|
||||
expect(msg).toContain(original);
|
||||
// Full recovery ladder: in-place repair → rebuild → switch engines.
|
||||
expect(msg).toContain('gbrain pglite-repair --dry-run');
|
||||
expect(msg).toContain('reinit-pglite');
|
||||
expect(msg).toContain('docs/ENGINES.md');
|
||||
expect(msg).toContain('gbrain doctor');
|
||||
expect(msg).toContain(`Original error: ${original}`);
|
||||
expect(msg).not.toContain('Bun vfs');
|
||||
});
|
||||
|
||||
// WAL-repair wave: the 4th param folds what auto-repair did (or why it
|
||||
// didn't run) into the hint so the message never lies about the state of
|
||||
// the data dir.
|
||||
test('wasm-abort + {repair: disabled} names the off switch', () => {
|
||||
const msg = buildPgliteInitErrorMessage('wasm-abort', original, 'darwin', { repair: 'disabled' });
|
||||
expect(msg).toContain('GBRAIN_PGLITE_WAL_REPAIR=off');
|
||||
expect(msg).toContain(original);
|
||||
});
|
||||
|
||||
test('wasm-abort + {repair: failed-restored} says RESTORED and names the backup path', () => {
|
||||
const msg = buildPgliteInitErrorMessage('wasm-abort', original, 'darwin', {
|
||||
repair: 'failed-restored',
|
||||
backupPath: '/x/b',
|
||||
});
|
||||
expect(msg).toContain('RESTORED');
|
||||
expect(msg).toContain('/x/b');
|
||||
expect(msg).toContain(original);
|
||||
});
|
||||
|
||||
test('wasm-abort + {repair: failed-not-restored} says RESET state, backup path, restore manually', () => {
|
||||
const msg = buildPgliteInitErrorMessage('wasm-abort', original, 'darwin', {
|
||||
repair: 'failed-not-restored',
|
||||
backupPath: '/x/b',
|
||||
});
|
||||
expect(msg).toContain('RESET state');
|
||||
expect(msg).toContain('/x/b');
|
||||
expect(msg.toLowerCase()).toContain('restore manually');
|
||||
expect(msg).toContain(original);
|
||||
});
|
||||
|
||||
test('wasm-abort + {repair: in-memory} says there is no stored state to repair', () => {
|
||||
const msg = buildPgliteInitErrorMessage('wasm-abort', original, 'darwin', { repair: 'in-memory' });
|
||||
expect(msg).toContain('in-memory');
|
||||
expect(msg).toContain(original);
|
||||
});
|
||||
|
||||
test('wasm-abort + {repair: skipped-live-writer} surfaces the skip detail verbatim', () => {
|
||||
const detail = 'the data-dir lock was reaped from pid 4242 (SENTINEL-LIVE-WRITER)';
|
||||
const msg = buildPgliteInitErrorMessage('wasm-abort', original, 'darwin', {
|
||||
repair: 'skipped-live-writer',
|
||||
detail,
|
||||
});
|
||||
expect(msg).toContain(detail);
|
||||
expect(msg).toContain(original);
|
||||
});
|
||||
|
||||
// #2674: the unknown-verdict hint is platform-gated. The macOS 26.3
|
||||
// attribution (#223) only appears on darwin; elsewhere the hint names
|
||||
// the causes that are actually plausible off-macOS.
|
||||
@@ -97,6 +175,12 @@ describe('buildPgliteInitErrorMessage — hint routing', () => {
|
||||
const msg = buildPgliteInitErrorMessage('unknown', original, 'darwin');
|
||||
expect(msg).toContain('gbrain doctor');
|
||||
expect(msg).toContain('issues/223');
|
||||
// WAL-repair wave: the darwin branch is reframed to the real root cause
|
||||
// behind the #223 reports (torn WAL from unclean shutdown) and offers the
|
||||
// mutation-free diagnosis command.
|
||||
expect(msg).toContain('corrupt WAL/checkpoint state');
|
||||
expect(msg).toContain('unclean');
|
||||
expect(msg).toContain('gbrain pglite-repair --dry-run');
|
||||
expect(msg).toContain(original);
|
||||
});
|
||||
|
||||
@@ -115,12 +199,15 @@ describe('buildPgliteInitErrorMessage — hint routing', () => {
|
||||
const msg = buildPgliteInitErrorMessage('corrupt', original);
|
||||
expect(msg).toContain('gbrain reinit-pglite');
|
||||
expect(msg).toContain('corrupted');
|
||||
// WAL-repair wave: the dry-run diagnosis is offered (report-only — WAL
|
||||
// repair cannot fix catalog corruption, and the copy says so).
|
||||
expect(msg).toContain('gbrain pglite-repair --dry-run');
|
||||
expect(msg).toContain(original);
|
||||
expect(msg).not.toContain('issues/223');
|
||||
});
|
||||
|
||||
test('all verdicts produce the canonical header line', () => {
|
||||
for (const v of ['bunfs', 'macos-26-3', 'corrupt', 'unknown'] as const) {
|
||||
for (const v of ['bunfs', 'wasm-abort', 'corrupt', 'unknown'] as const) {
|
||||
const msg = buildPgliteInitErrorMessage(v, original);
|
||||
expect(msg.startsWith('PGLite failed to initialize its WASM runtime.')).toBe(true);
|
||||
}
|
||||
@@ -145,6 +232,27 @@ describe('stringifyPgliteInitError — non-Error rejections (#2674)', () => {
|
||||
expect(stringifyPgliteInitError(null)).toBe('null');
|
||||
expect(stringifyPgliteInitError(undefined)).toBe('undefined');
|
||||
});
|
||||
|
||||
// WAL-repair wave: Emscripten's FS layer throws message-LESS objects (e.g.
|
||||
// `ErrnoError { name: 'ErrnoError', errno: 20 }` when the data dir is a
|
||||
// symlink NODEFS refuses to mount) — never "[object Object]".
|
||||
test('message-less ErrnoError-shaped object yields name + errno', () => {
|
||||
expect(stringifyPgliteInitError({ name: 'ErrnoError', errno: 20 })).toBe('ErrnoError (errno 20)');
|
||||
});
|
||||
|
||||
test('message-less nameless object with other props yields its JSON', () => {
|
||||
expect(stringifyPgliteInitError({ code: 'ENOENT' })).toBe('{"code":"ENOENT"}');
|
||||
});
|
||||
|
||||
test('message-less object with a name and serializable props yields name-prefixed JSON', () => {
|
||||
expect(stringifyPgliteInitError({ name: 'Weird' })).toBe('Weird: {"name":"Weird"}');
|
||||
});
|
||||
|
||||
test('circular object with a name falls back to the bare name (JSON.stringify throws)', () => {
|
||||
const c: Record<string, unknown> = { name: 'Circ' };
|
||||
c.self = c;
|
||||
expect(stringifyPgliteInitError(c)).toBe('Circ');
|
||||
});
|
||||
});
|
||||
|
||||
describe('#1340 reproducer — exact reporter error string maps to bunfs', () => {
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
import { describe, test, expect, beforeEach, afterEach } from 'bun:test';
|
||||
import { mkdirSync, rmSync, existsSync, readFileSync, writeFileSync } from 'fs';
|
||||
import { mkdirSync, mkdtempSync, rmSync, existsSync, readFileSync, writeFileSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { tmpdir } from 'os';
|
||||
import { acquireLock, releaseLock, type LockHandle } from '../src/core/pglite-lock';
|
||||
@@ -271,3 +271,98 @@ describe('pglite-lock #2058 heartbeat + steal-grace', () => {
|
||||
expect(existsSync(join(TEST_DIR, '.gbrain-lock'))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('pglite-lock reap classification (WAL-repair wave)', () => {
|
||||
// Unique per-test tmpdirs: the reap marker lands at `${dataDir}.lock-reap.json`
|
||||
// — a SIBLING of the data dir — so each test gets its own parent to rm.
|
||||
function freshDataDir(): { parent: string; dataDir: string } {
|
||||
const parent = mkdtempSync(join(tmpdir(), 'gbrain-lock-reap-'));
|
||||
return { parent, dataDir: join(parent, 'data') };
|
||||
}
|
||||
|
||||
/**
|
||||
* A PID that provably belongs to no live process: spawn a short-lived child,
|
||||
* wait for it (spawnSync reaps it), then verify kill(pid, 0) throws. Retries
|
||||
* to dodge instant PID reuse.
|
||||
*/
|
||||
function deadPid(): number {
|
||||
for (let attempt = 0; attempt < 5; attempt++) {
|
||||
const proc = Bun.spawnSync(['bash', '-c', 'exit 0']);
|
||||
const pid = proc.pid;
|
||||
try {
|
||||
process.kill(pid, 0); // still alive/visible → PID reused, try again
|
||||
} catch {
|
||||
return pid;
|
||||
}
|
||||
}
|
||||
throw new Error('could not obtain a provably-dead PID after 5 spawns');
|
||||
}
|
||||
|
||||
test('corrupt lock file: reaped acquisition + persisted .lock-reap.json marker', async () => {
|
||||
const { parent, dataDir } = freshDataDir();
|
||||
try {
|
||||
const lockDir = join(dataDir, '.gbrain-lock');
|
||||
mkdirSync(lockDir, { recursive: true });
|
||||
writeFileSync(join(lockDir, 'lock'), 'not json {{{'); // holder liveness UNKNOWABLE
|
||||
|
||||
const lock = await acquireLock(dataDir, { timeoutMs: 5000 });
|
||||
try {
|
||||
expect(lock.acquired).toBe(true);
|
||||
expect(lock.reaped).toBe(true);
|
||||
// Unknowable-liveness reap is persisted cross-process for the repair gate.
|
||||
expect(existsSync(`${dataDir}.lock-reap.json`)).toBe(true);
|
||||
const marker = JSON.parse(readFileSync(`${dataDir}.lock-reap.json`, 'utf-8'));
|
||||
expect(typeof marker.ts).toBe('number');
|
||||
expect(marker.by).toBe(process.pid);
|
||||
} finally {
|
||||
await releaseLock(lock);
|
||||
}
|
||||
} finally {
|
||||
rmSync(parent, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('clean acquisition: reaped falsy, no .lock-reap.json marker', async () => {
|
||||
const { parent, dataDir } = freshDataDir();
|
||||
try {
|
||||
const lock = await acquireLock(dataDir, { timeoutMs: 5000 });
|
||||
try {
|
||||
expect(lock.acquired).toBe(true);
|
||||
expect(lock.reaped).toBeFalsy();
|
||||
expect(existsSync(`${dataDir}.lock-reap.json`)).toBe(false);
|
||||
} finally {
|
||||
await releaseLock(lock);
|
||||
}
|
||||
} finally {
|
||||
rmSync(parent, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('dead-PID lock: reaped acquisition but NO marker (affirmative ESRCH verdict)', async () => {
|
||||
const { parent, dataDir } = freshDataDir();
|
||||
try {
|
||||
const lockDir = join(dataDir, '.gbrain-lock');
|
||||
mkdirSync(lockDir, { recursive: true });
|
||||
const now = Date.now();
|
||||
writeFileSync(join(lockDir, 'lock'), JSON.stringify({
|
||||
pid: deadPid(),
|
||||
acquired_at: now - 60_000,
|
||||
refreshed_at: now - 60_000,
|
||||
command: 'gbrain embed',
|
||||
subcommand: 'embed',
|
||||
}));
|
||||
|
||||
const lock = await acquireLock(dataDir, { timeoutMs: 5000 });
|
||||
try {
|
||||
expect(lock.acquired).toBe(true);
|
||||
expect(lock.reaped).toBe(true);
|
||||
// Dead-PID reaps deliberately do NOT quarantine the next acquirer.
|
||||
expect(existsSync(`${dataDir}.lock-reap.json`)).toBe(false);
|
||||
} finally {
|
||||
await releaseLock(lock);
|
||||
}
|
||||
} finally {
|
||||
rmSync(parent, { recursive: true, force: true });
|
||||
}
|
||||
}, 30_000);
|
||||
});
|
||||
|
||||
@@ -0,0 +1,418 @@
|
||||
/**
|
||||
* `gbrain pglite-repair` command surface (#223 WAL-repair wave) — SERIAL:
|
||||
* the happy path does a real persistent-PGLite cold start (create → corrupt
|
||||
* pg_wal → repair in place → reconnect), which is too heavy + lock-contended
|
||||
* for the parallel unit shards.
|
||||
*
|
||||
* `runPgliteRepair` is imported directly (no process spawns); stdout/stderr
|
||||
* are captured by spying console.log/console.error per test and restored in
|
||||
* finally. Every case uses `--path <hermetic tmpdir>` so the user's real
|
||||
* brain and config are never touched, and `--json` so assertions parse a
|
||||
* machine receipt instead of prose.
|
||||
*
|
||||
* Refusal-order note (cases 5-7): the fake layout deliberately PASSES
|
||||
* `validateWalRepairTarget` (PG_VERSION 17 + base/ + 8192-byte pg_control) so
|
||||
* the command reaches its lock gates; an actual repair on the garbage
|
||||
* pg_control would fail in resetWal, but all three cases must refuse BEFORE
|
||||
* repair — asserting the refused_* codes proves the ordering.
|
||||
*/
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import {
|
||||
existsSync, mkdirSync, mkdtempSync, readdirSync, rmSync, writeFileSync,
|
||||
} from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { basename, join } from 'node:path';
|
||||
|
||||
import { runPgliteRepair } from '../src/commands/pglite-repair.ts';
|
||||
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
||||
import { withEnv } from './helpers/with-env.ts';
|
||||
|
||||
function tmp(prefix: string): string {
|
||||
return mkdtempSync(join(tmpdir(), prefix));
|
||||
}
|
||||
|
||||
/** Fake PG17 layout that passes validateWalRepairTarget (see file header). */
|
||||
function makeFakeLayout(dir: string): void {
|
||||
mkdirSync(join(dir, 'base'), { recursive: true });
|
||||
mkdirSync(join(dir, 'global'), { recursive: true });
|
||||
mkdirSync(join(dir, 'pg_wal'), { recursive: true });
|
||||
writeFileSync(join(dir, 'PG_VERSION'), '17\n');
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), Buffer.alloc(8192));
|
||||
}
|
||||
|
||||
function writeLockFile(dir: string, lock: Record<string, unknown>): string {
|
||||
const lockDir = join(dir, '.gbrain-lock');
|
||||
mkdirSync(lockDir, { recursive: true });
|
||||
writeFileSync(join(lockDir, 'lock'), JSON.stringify(lock), { mode: 0o644 });
|
||||
return lockDir;
|
||||
}
|
||||
|
||||
interface Captured {
|
||||
logs: string[];
|
||||
errors: string[];
|
||||
restore: () => void;
|
||||
}
|
||||
|
||||
/** Spy console.log/console.error; caller MUST call restore() in finally. */
|
||||
function captureConsole(): Captured {
|
||||
const logs: string[] = [];
|
||||
const errors: string[] = [];
|
||||
const origLog = console.log;
|
||||
const origErr = console.error;
|
||||
console.log = (...args: unknown[]) => { logs.push(args.map(String).join(' ')); };
|
||||
console.error = (...args: unknown[]) => { errors.push(args.map(String).join(' ')); };
|
||||
return {
|
||||
logs,
|
||||
errors,
|
||||
restore: () => {
|
||||
console.log = origLog;
|
||||
console.error = origErr;
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
/** Last parseable JSON object line from captured console.log output. */
|
||||
function parseJsonLine(logs: string[]): Record<string, any> {
|
||||
for (let i = logs.length - 1; i >= 0; i--) {
|
||||
const line = logs[i].trim();
|
||||
if (!line.startsWith('{')) continue;
|
||||
try {
|
||||
return JSON.parse(line) as Record<string, any>;
|
||||
} catch { /* not this line — keep looking */ }
|
||||
}
|
||||
throw new Error(`no JSON line in captured output: ${JSON.stringify(logs)}`);
|
||||
}
|
||||
|
||||
/**
|
||||
* A PID that provably belongs to no live process: spawn a short-lived child,
|
||||
* wait for it (spawnSync reaps it), then verify kill(pid, 0) throws. Retries
|
||||
* to dodge instant PID reuse.
|
||||
*/
|
||||
function deadPid(): number {
|
||||
for (let attempt = 0; attempt < 5; attempt++) {
|
||||
const proc = Bun.spawnSync(['bash', '-c', 'exit 0']);
|
||||
const pid = proc.pid;
|
||||
try {
|
||||
process.kill(pid, 0); // still alive/visible → PID reused, try again
|
||||
} catch {
|
||||
return pid;
|
||||
}
|
||||
}
|
||||
throw new Error('could not obtain a provably-dead PID after 5 spawns');
|
||||
}
|
||||
|
||||
const backupDirsBeside = (dir: string): string[] =>
|
||||
readdirSync(join(dir, '..')).filter((n) => n.startsWith(`${basename(dir)}.wal-repair-backup-`));
|
||||
|
||||
describe('gbrain pglite-repair — dry-run is strictly read-only', () => {
|
||||
test('1. dry-run on a corrupt-ish layout: exit 0, JSON diagnosis, zero mutation', async () => {
|
||||
const parent = tmp('gbrain-repair-dry-');
|
||||
const dir = join(parent, 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
writeFileSync(join(dir, 'postmaster.pid'), '99999\n'); // unclean-shutdown marker
|
||||
const parentBefore = readdirSync(parent).sort();
|
||||
const dirBefore = readdirSync(dir).sort();
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--dry-run', '--path', dir, '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
|
||||
expect(rc).toBe(0);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('ok');
|
||||
expect(out.action).toBe('dry-run');
|
||||
expect(out.data_dir).toBe(dir);
|
||||
expect(out.validation.ok).toBe(true);
|
||||
expect(out.diagnosis.verdict).toBeDefined();
|
||||
expect(out.diagnosis.verdict).toBe('wal-corruption-likely');
|
||||
|
||||
// Read-only: no backup dirs, no sidecar, nothing added or removed.
|
||||
expect(backupDirsBeside(dir)).toEqual([]);
|
||||
expect(existsSync(`${dir}.wal-repair-attempt.json`)).toBe(false);
|
||||
expect(readdirSync(parent).sort()).toEqual(parentBefore);
|
||||
expect(readdirSync(dir).sort()).toEqual(dirBefore);
|
||||
expect(existsSync(join(dir, 'postmaster.pid'))).toBe(true);
|
||||
});
|
||||
|
||||
test('2. dry-run on a missing path: exit 0, validation.ok false, path NOT created', async () => {
|
||||
const missing = join(tmp('gbrain-repair-dry-missing-'), 'never-created.pglite');
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--dry-run', '--path', missing, '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
|
||||
expect(rc).toBe(0);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('ok');
|
||||
expect(out.action).toBe('dry-run');
|
||||
expect(out.validation.ok).toBe(false);
|
||||
expect(out.diagnosis.verdict).toBe('missing');
|
||||
expect(existsSync(missing)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('gbrain pglite-repair — refusals (validate before lock, never mkdir a typo)', () => {
|
||||
test('3. non-dry-run on a missing path: exit 1, refused_missing-dir, path NOT created', async () => {
|
||||
const missing = join(tmp('gbrain-repair-missing-'), 'typo.pglite');
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', missing, '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
|
||||
expect(rc).toBe(1);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('error');
|
||||
expect(out.code).toBe('refused_missing-dir');
|
||||
// validate-before-lock: acquireLock would have mkdir'd the dir.
|
||||
expect(existsSync(missing)).toBe(false);
|
||||
});
|
||||
|
||||
test('5. live lock holder: exit 1, refused_locked, no repair attempted', async () => {
|
||||
const dir = join(tmp('gbrain-repair-locked-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
const lockDir = writeLockFile(dir, {
|
||||
pid: process.pid, // this test process — provably alive
|
||||
acquired_at: Date.now(),
|
||||
refreshed_at: Date.now(),
|
||||
command: 'gbrain embed',
|
||||
subcommand: 'embed',
|
||||
});
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', dir, '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
rmSync(lockDir, { recursive: true, force: true });
|
||||
}
|
||||
|
||||
expect(rc).toBe(1);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('error');
|
||||
expect(out.code).toBe('refused_locked');
|
||||
// Refused BEFORE repair: no backup dir, no sidecar.
|
||||
expect(backupDirsBeside(dir)).toEqual([]);
|
||||
expect(existsSync(`${dir}.wal-repair-attempt.json`)).toBe(false);
|
||||
});
|
||||
|
||||
test('6. reaped (dead-PID) lock: exit 1, refused_reaped_lock', async () => {
|
||||
const dir = join(tmp('gbrain-repair-reaped-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
writeLockFile(dir, {
|
||||
pid: deadPid(), // provably dead — acquireLock reaps it, then refuses
|
||||
acquired_at: Date.now() - 60_000,
|
||||
refreshed_at: Date.now() - 60_000,
|
||||
command: 'gbrain embed',
|
||||
subcommand: 'embed',
|
||||
});
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', dir, '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
|
||||
expect(rc).toBe(1);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('error');
|
||||
expect(out.code).toBe('refused_reaped_lock');
|
||||
// Refused BEFORE repair: no backup dir, no sidecar.
|
||||
expect(backupDirsBeside(dir)).toEqual([]);
|
||||
expect(existsSync(`${dir}.wal-repair-attempt.json`)).toBe(false);
|
||||
}, 30_000);
|
||||
|
||||
test('7. live serve holder: exit 1, refused_locked names the PID', async () => {
|
||||
const dir = join(tmp('gbrain-repair-serve-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
const lockDir = writeLockFile(dir, {
|
||||
pid: process.pid, // alive — the pre-lock diagnosis catches it as 'locked'
|
||||
acquired_at: Date.now(),
|
||||
refreshed_at: Date.now(),
|
||||
command: 'gbrain serve',
|
||||
subcommand: 'serve',
|
||||
});
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', dir, '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
rmSync(lockDir, { recursive: true, force: true });
|
||||
}
|
||||
|
||||
expect(rc).toBe(1);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('error');
|
||||
expect(out.code).toBe('refused_locked');
|
||||
expect(out.message).toContain(String(process.pid));
|
||||
expect(backupDirsBeside(dir)).toEqual([]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('gbrain pglite-repair — TTY + config gates', () => {
|
||||
test('8. non-TTY without --yes refuses: exit 1, no_tty_no_yes, zero mutation', async () => {
|
||||
const dir = join(tmp('gbrain-repair-notty-'), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
|
||||
// Pin stdin to non-TTY: under `bun test` in a terminal stdin can still be
|
||||
// a TTY, which would route into the interactive confirm instead.
|
||||
const origTty = Object.getOwnPropertyDescriptor(process.stdin, 'isTTY');
|
||||
Object.defineProperty(process.stdin, 'isTTY', { value: false, configurable: true });
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', dir, '--json']); // no --yes
|
||||
} finally {
|
||||
cap.restore();
|
||||
if (origTty) Object.defineProperty(process.stdin, 'isTTY', origTty);
|
||||
else delete (process.stdin as unknown as Record<string, unknown>).isTTY;
|
||||
}
|
||||
|
||||
expect(rc).toBe(1);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('error');
|
||||
expect(out.code).toBe('no_tty_no_yes');
|
||||
// Refused BEFORE any surgery: no backup dir, no sidecar.
|
||||
expect(backupDirsBeside(dir)).toEqual([]);
|
||||
expect(existsSync(`${dir}.wal-repair-attempt.json`)).toBe(false);
|
||||
});
|
||||
|
||||
test('9. no --path with a non-pglite configured engine: exit 1, not_pglite', async () => {
|
||||
// Hermetic GBRAIN_HOME (same convention as apply-migrations-pglite-spawn):
|
||||
// configDir() appends '.gbrain', so the config lands at <home>/.gbrain/.
|
||||
const home = tmp('gbrain-repair-home-');
|
||||
mkdirSync(join(home, '.gbrain'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(home, '.gbrain', 'config.json'),
|
||||
JSON.stringify({ engine: 'postgres', database_url: 'postgresql://localhost:5432/x' }) + '\n',
|
||||
);
|
||||
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await withEnv({ GBRAIN_HOME: home }, () => runPgliteRepair(['--yes', '--json']));
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
|
||||
expect(rc).toBe(1);
|
||||
const out = parseJsonLine(cap.logs);
|
||||
expect(out.status).toBe('error');
|
||||
expect(out.code).toBe('not_pglite');
|
||||
expect(out.message).toContain('--path');
|
||||
});
|
||||
});
|
||||
|
||||
describe('gbrain pglite-repair — the happy path (real PGLite)', () => {
|
||||
test('4. corrupt pg_wal → repair in place → data survives, no auto-repair on reconnect', async () => {
|
||||
const dir = join(tmp('gbrain-repair-happy-'), 'brain.pglite');
|
||||
|
||||
// 1) Real brain with a probe row, closed cleanly.
|
||||
const engine = new PGLiteEngine();
|
||||
await engine.connect({ database_path: dir });
|
||||
try {
|
||||
await engine.executeRaw('CREATE TABLE repair_probe(id int)');
|
||||
await engine.executeRaw('INSERT INTO repair_probe VALUES (42)');
|
||||
} finally {
|
||||
await engine.disconnect();
|
||||
}
|
||||
|
||||
// 2) Corrupt every WAL segment with garbage.
|
||||
const walDir = join(dir, 'pg_wal');
|
||||
const segments = readdirSync(walDir).filter((f) => /^[0-9A-F]{24}$/.test(f));
|
||||
expect(segments.length).toBeGreaterThan(0); // sanity: there IS WAL to corrupt
|
||||
for (const seg of segments) {
|
||||
writeFileSync(join(walDir, seg), Buffer.alloc(1024, 0xff));
|
||||
}
|
||||
|
||||
// 3) Repair in place.
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', dir, '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
expect(rc).toBe(0);
|
||||
const receipt = parseJsonLine(cap.logs);
|
||||
expect(receipt.status).toBe('ok');
|
||||
expect(receipt.action).toBe('repaired');
|
||||
expect(receipt.data_dir).toBe(dir);
|
||||
expect(receipt.reset_segment).toMatch(/^[0-9A-F]{24}$/);
|
||||
expect(existsSync(receipt.backup_path)).toBe(true);
|
||||
expect(existsSync(join(receipt.backup_path, 'pg_wal'))).toBe(true);
|
||||
|
||||
// 4) The repaired dir opens WITHOUT auto-repair firing, data intact.
|
||||
const engine2 = new PGLiteEngine();
|
||||
await engine2.connect({ database_path: dir });
|
||||
try {
|
||||
expect(engine2.walRepairReceipt).toBeNull();
|
||||
const rows = await engine2.executeRaw<{ id: number }>('SELECT id FROM repair_probe');
|
||||
expect(rows.length).toBe(1);
|
||||
expect(rows[0].id).toBe(42);
|
||||
} finally {
|
||||
await engine2.disconnect();
|
||||
}
|
||||
}, 180_000);
|
||||
});
|
||||
|
||||
describe('gbrain pglite-repair — argument + quarantine hardening (adversarial fixes)', () => {
|
||||
test('unknown flag is rejected (exit 2), not silently ignored on a destructive command', async () => {
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--dry-rnu', '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
expect(rc).toBe(2);
|
||||
expect(parseJsonLine(cap.logs).code).toBe('unknown_flag');
|
||||
});
|
||||
|
||||
test('--path with no value is rejected (does NOT retarget the default brain)', async () => {
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--yes', '--json', '--path']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
expect(rc).toBe(2);
|
||||
expect(parseJsonLine(cap.logs).code).toBe('unknown_flag');
|
||||
});
|
||||
|
||||
test('the command honors the cross-process reap quarantine (F3: a second --yes cannot bypass it)', async () => {
|
||||
const dir = join(mkdtempSync(join(tmpdir(), 'pgrepaircmd-')), 'brain.pglite');
|
||||
makeFakeLayout(dir);
|
||||
// A fresh corrupt-lock reap marker from a prior run — the possibly-live
|
||||
// writer it protects must not be repaired under.
|
||||
writeFileSync(`${dir}.lock-reap.json`, JSON.stringify({ ts: Date.now(), by: 999999 }), { mode: 0o644 });
|
||||
const cap = captureConsole();
|
||||
let rc: number;
|
||||
try {
|
||||
rc = await runPgliteRepair(['--path', dir, '--yes', '--json']);
|
||||
} finally {
|
||||
cap.restore();
|
||||
}
|
||||
expect(rc).toBe(1);
|
||||
expect(parseJsonLine(cap.logs).code).toBe('refused_reap_quarantine');
|
||||
// No surgery: no backup dir created.
|
||||
expect(readdirSync(join(dir, '..')).some((f) => f.includes('.wal-repair-backup-'))).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,548 @@
|
||||
/**
|
||||
* Unit tests for the WAL-repair orchestrator (src/core/pglite-repair.ts):
|
||||
* validation, rename-based backup, overwrite-order restore + mtime guard,
|
||||
* cooldown sidecar, episode-scoped retention, and the never-throws engine
|
||||
* seam (attemptWalRepairAndRetry) with injected retryCreate — no real PGLite.
|
||||
*
|
||||
* The real-engine regression (corrupt a real brain → connect() auto-repairs →
|
||||
* row readable) lives in test/pglite-wal-repair.serial.test.ts.
|
||||
*/
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import {
|
||||
mkdtempSync, mkdirSync, writeFileSync, existsSync, readFileSync, readdirSync,
|
||||
symlinkSync, rmSync, utimesSync,
|
||||
} from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { withEnv } from './helpers/with-env.ts';
|
||||
import { xlogFileName, crc32c } from '../src/core/pglite-resetwal.ts';
|
||||
import {
|
||||
validateWalRepairTarget,
|
||||
inspectPgliteDataDir,
|
||||
repairPgliteWal,
|
||||
restoreWalBackup,
|
||||
attemptWalRepairAndRetry,
|
||||
readRepairSidecar,
|
||||
recordRepairAttempt,
|
||||
repairCooldownActive,
|
||||
listRepairBackups,
|
||||
pruneRepairBackups,
|
||||
WalRepairError,
|
||||
closeRepairEpisodeIfOpen,
|
||||
} from '../src/core/pglite-repair.ts';
|
||||
|
||||
const SEG_SIZE = 1024 * 1024;
|
||||
|
||||
function makeControl(): Buffer {
|
||||
const control = Buffer.alloc(8192);
|
||||
control.writeBigUInt64LE(0x1122334455667788n, 0); // systemIdentifier
|
||||
control.writeUInt32LE(1700, 8); // pg_control version
|
||||
control.writeUInt32LE(1, 48); // timeline
|
||||
control.writeUInt32LE(8192, 224); // xlogBlcksz
|
||||
control.writeUInt32LE(SEG_SIZE, 228); // xlogSegSize
|
||||
control.writeBigUInt64LE(3n * BigInt(SEG_SIZE) + 40n, 40); // redo → seg 3
|
||||
control.writeUInt32LE(crc32c([control.subarray(0, 288)]), 288); // valid CRC
|
||||
return control;
|
||||
}
|
||||
|
||||
/** A synthetic PG17 layout that resetWal fully accepts. */
|
||||
function makeLayout(opts?: { segments?: string[]; postmasterPid?: boolean }): string {
|
||||
const parent = mkdtempSync(join(tmpdir(), 'pgrepair-'));
|
||||
const dir = join(parent, 'brain.pglite');
|
||||
mkdirSync(dir, { recursive: true });
|
||||
writeFileSync(join(dir, 'PG_VERSION'), '17\n');
|
||||
mkdirSync(join(dir, 'global'), { recursive: true });
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), makeControl());
|
||||
mkdirSync(join(dir, 'base'), { recursive: true });
|
||||
mkdirSync(join(dir, 'pg_wal', 'archive_status'), { recursive: true });
|
||||
for (const seg of opts?.segments ?? [xlogFileName(1, 3n, SEG_SIZE)]) {
|
||||
writeFileSync(join(dir, 'pg_wal', seg), Buffer.alloc(2048, 0xaa));
|
||||
}
|
||||
if (opts?.postmasterPid) writeFileSync(join(dir, 'postmaster.pid'), '12345\n');
|
||||
return dir;
|
||||
}
|
||||
|
||||
/**
|
||||
* Overwrite pg_control with an 8192-byte buffer carrying a WRONG control
|
||||
* version: it PASSES validateWalRepairTarget (size-only check) but FAILS
|
||||
* resetWal's version check — the fixture for the reset-fails-AFTER-backup
|
||||
* (WalRepairError) path.
|
||||
*/
|
||||
function poisonControlVersion(dir: string): void {
|
||||
const control = makeControl();
|
||||
control.writeUInt32LE(1600, 8);
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), control);
|
||||
}
|
||||
|
||||
describe('validateWalRepairTarget', () => {
|
||||
test('accepts a full PG17 layout, tolerating .gbrain-lock inside it', () => {
|
||||
const dir = makeLayout();
|
||||
mkdirSync(join(dir, '.gbrain-lock'), { recursive: true });
|
||||
writeFileSync(join(dir, '.gbrain-lock', 'lock'), '{}');
|
||||
expect(validateWalRepairTarget(dir)).toEqual({ ok: true });
|
||||
});
|
||||
|
||||
test('refusal matrix: missing dir / no PG_VERSION / wrong version / no base / bad control', () => {
|
||||
expect(validateWalRepairTarget('')).toMatchObject({ ok: false, reason: 'missing-dir' });
|
||||
expect(validateWalRepairTarget('/nope/never/exists')).toMatchObject({ ok: false, reason: 'missing-dir' });
|
||||
|
||||
const noVersion = mkdtempSync(join(tmpdir(), 'pgrepair-'));
|
||||
expect(validateWalRepairTarget(noVersion)).toMatchObject({ ok: false, reason: 'not-pglite-layout' });
|
||||
|
||||
const v16 = makeLayout();
|
||||
writeFileSync(join(v16, 'PG_VERSION'), '16\n');
|
||||
expect(validateWalRepairTarget(v16)).toMatchObject({ ok: false, reason: 'unsupported-pg-version' });
|
||||
|
||||
const noBase = makeLayout();
|
||||
rmSync(join(noBase, 'base'), { recursive: true });
|
||||
expect(validateWalRepairTarget(noBase)).toMatchObject({ ok: false, reason: 'not-pglite-layout' });
|
||||
|
||||
const badControl = makeLayout();
|
||||
writeFileSync(join(badControl, 'global', 'pg_control'), Buffer.alloc(100));
|
||||
expect(validateWalRepairTarget(badControl)).toMatchObject({ ok: false, reason: 'bad-pg-control' });
|
||||
});
|
||||
|
||||
test('refuses symlinked dataDir and symlinked pg_wal (codex 14.8)', () => {
|
||||
const real = makeLayout();
|
||||
const link = join(mkdtempSync(join(tmpdir(), 'pgrepair-')), 'link.pglite');
|
||||
symlinkSync(real, link);
|
||||
expect(validateWalRepairTarget(link)).toMatchObject({ ok: false, reason: 'not-pglite-layout' });
|
||||
|
||||
const dir = makeLayout();
|
||||
const walBackup = join(dir, 'pg_wal_real');
|
||||
rmSync(join(dir, 'pg_wal'), { recursive: true });
|
||||
mkdirSync(walBackup);
|
||||
symlinkSync(walBackup, join(dir, 'pg_wal'));
|
||||
expect(validateWalRepairTarget(dir)).toMatchObject({ ok: false, reason: 'not-pglite-layout' });
|
||||
});
|
||||
|
||||
test('refuses a symlinked global/ dir (security review — lstat on pg_control follows the intermediate link)', () => {
|
||||
const dir = makeLayout();
|
||||
// A foreign dir holding a perfectly valid 8192-byte pg_control: without the
|
||||
// global/ lstat check, surgery would write a forged control THROUGH the
|
||||
// link into this directory.
|
||||
const foreign = mkdtempSync(join(tmpdir(), 'pgrepair-foreign-'));
|
||||
writeFileSync(join(foreign, 'pg_control'), makeControl());
|
||||
rmSync(join(dir, 'global'), { recursive: true });
|
||||
symlinkSync(foreign, join(dir, 'global'));
|
||||
const result = validateWalRepairTarget(dir);
|
||||
expect(result).toMatchObject({ ok: false, reason: 'not-pglite-layout' });
|
||||
if (!result.ok) expect(result.detail).toContain('symlink');
|
||||
});
|
||||
});
|
||||
|
||||
describe('repairPgliteWal — rename-based backup', () => {
|
||||
test('backs up the WHOLE pg_wal dir + postmaster.pid (rename) and copies pg_control', async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg], postmasterPid: true });
|
||||
writeFileSync(join(dir, 'pg_wal', 'archive_status', `${seg}.ready`), '');
|
||||
const originalControl = readFileSync(join(dir, 'global', 'pg_control'));
|
||||
|
||||
const receipt = await repairPgliteWal(dir);
|
||||
|
||||
expect(receipt.backupPath.includes('.wal-repair-backup-')).toBe(true);
|
||||
expect(receipt.backedUpFiles).toEqual(['pg_wal/', 'postmaster.pid', 'global/pg_control']);
|
||||
// Backup holds the ORIGINAL bytes, archive_status entries included.
|
||||
expect(readFileSync(join(receipt.backupPath, 'pg_wal', seg), 'utf-8')).toBe(Buffer.alloc(2048, 0xaa).toString());
|
||||
expect(existsSync(join(receipt.backupPath, 'pg_wal', 'archive_status', `${seg}.ready`))).toBe(true);
|
||||
expect(existsSync(join(receipt.backupPath, 'postmaster.pid'))).toBe(true);
|
||||
expect(readFileSync(join(receipt.backupPath, 'pg_control')).equals(originalControl)).toBe(true);
|
||||
// Data dir: fresh pg_wal with exactly the reset segment; pid gone.
|
||||
expect(existsSync(join(dir, 'postmaster.pid'))).toBe(false);
|
||||
const segs = readdirSync(join(dir, 'pg_wal')).filter((f) => /^[0-9A-F]{24}$/.test(f));
|
||||
expect(segs).toEqual([receipt.resetSegment]);
|
||||
});
|
||||
|
||||
test('refuses (typed) on an invalid layout without touching anything', async () => {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'pgrepair-'));
|
||||
await expect(repairPgliteWal(dir)).rejects.toThrow(/refusing repair/);
|
||||
expect(listRepairBackups(dir)).toEqual([]);
|
||||
});
|
||||
|
||||
test('throws WalRepairError after backup and restores the dir when resetWal fails', async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg] });
|
||||
poisonControlVersion(dir); // passes validation, fails resetWal
|
||||
|
||||
let caught: unknown;
|
||||
try {
|
||||
await repairPgliteWal(dir);
|
||||
} catch (e) {
|
||||
caught = e;
|
||||
}
|
||||
expect(caught).toBeInstanceOf(WalRepairError);
|
||||
const err = caught as WalRepairError;
|
||||
// The best-effort restore ran and is reported HONESTLY on the error.
|
||||
expect(err.restore.restored).toBe(true);
|
||||
expect(existsSync(join(dir, 'pg_wal'))).toBe(true);
|
||||
expect(existsSync(join(dir, 'pg_wal', seg))).toBe(true); // original segment back
|
||||
expect(existsSync(err.receipt.backupPath)).toBe(true); // forensic backup kept
|
||||
});
|
||||
});
|
||||
|
||||
describe('restoreWalBackup — overwrite order + guards', () => {
|
||||
test('byte-identical restore: control first, dir swap, nothing deleted', async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg] });
|
||||
const originalControl = readFileSync(join(dir, 'global', 'pg_control'));
|
||||
|
||||
const receipt = await repairPgliteWal(dir);
|
||||
const result = await restoreWalBackup(receipt);
|
||||
|
||||
expect(result.restored).toBe(true);
|
||||
// Original WAL + control are back, byte-identical.
|
||||
expect(readFileSync(join(dir, 'pg_wal', seg), 'utf-8')).toBe(Buffer.alloc(2048, 0xaa).toString());
|
||||
expect(readFileSync(join(dir, 'global', 'pg_control')).equals(originalControl)).toBe(true);
|
||||
// The reset-state pg_wal was set ASIDE inside the backup dir, not deleted.
|
||||
const asides = readdirSync(receipt.backupPath).filter((f) => f.startsWith('pg_wal.reset-aside-'));
|
||||
expect(asides.length).toBe(1);
|
||||
expect(existsSync(join(receipt.backupPath, asides[0]!, receipt.resetSegment))).toBe(true);
|
||||
// The dir still has a valid 8192-byte pg_control at every observable point.
|
||||
expect(readFileSync(join(dir, 'global', 'pg_control')).length).toBe(8192);
|
||||
});
|
||||
|
||||
test('mtime guard: refuses when a foreign WAL segment is newer than the backup', async () => {
|
||||
const dir = makeLayout();
|
||||
const receipt = await repairPgliteWal(dir);
|
||||
// A "live writer" drops a fresh segment into the (reset) pg_wal.
|
||||
const foreign = xlogFileName(1, 99n, SEG_SIZE);
|
||||
writeFileSync(join(dir, 'pg_wal', foreign), 'live-writer-bytes');
|
||||
const future = new Date(Date.now() + 60_000);
|
||||
utimesSync(join(dir, 'pg_wal', foreign), future, future);
|
||||
|
||||
const result = await restoreWalBackup(receipt);
|
||||
expect(result.restored).toBe(false);
|
||||
expect(result.detail).toContain('mtime-guard');
|
||||
// Nothing was swapped or deleted; backup remains intact.
|
||||
expect(existsSync(join(receipt.backupPath, 'pg_wal'))).toBe(true);
|
||||
});
|
||||
|
||||
test('missing/empty backup never claims restoration (8A honesty)', async () => {
|
||||
const dir = makeLayout();
|
||||
const receipt = await repairPgliteWal(dir);
|
||||
rmSync(receipt.backupPath, { recursive: true, force: true });
|
||||
const result = await restoreWalBackup(receipt);
|
||||
expect(result.restored).toBe(false);
|
||||
expect(result.detail).toContain('nothing to restore');
|
||||
});
|
||||
});
|
||||
|
||||
describe('cooldown sidecar + episode retention', () => {
|
||||
test('recordRepairAttempt opens an episode on failure, closes on success, caps history', () => {
|
||||
const dir = makeLayout();
|
||||
// Real backup dirs: the re-pin rule inspects them for pg_wal (a gutted
|
||||
// pinned backup — restore moved its pg_wal back — must lose the pin).
|
||||
const backupOne = `${dir}.wal-repair-backup-1001`;
|
||||
const backupTwo = `${dir}.wal-repair-backup-1002`;
|
||||
const backupThree = `${dir}.wal-repair-backup-1003`;
|
||||
mkdirSync(join(backupOne, 'pg_wal'), { recursive: true });
|
||||
mkdirSync(join(backupTwo, 'pg_wal'), { recursive: true });
|
||||
mkdirSync(join(backupThree, 'pg_wal'), { recursive: true });
|
||||
|
||||
recordRepairAttempt(dir, 'failed', backupOne);
|
||||
let sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeStartedAt).not.toBeNull();
|
||||
expect(sidecar.episodeBackupPath).toBe(backupOne);
|
||||
// Second failure does NOT re-pin while the pinned backup still holds pg_wal.
|
||||
recordRepairAttempt(dir, 'failed', backupTwo);
|
||||
sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeBackupPath).toBe(backupOne);
|
||||
// …but a GUTTED pinned backup loses the pin to the fresh one (red-team:
|
||||
// the episode's protected copy must always be one that still has pg_wal).
|
||||
rmSync(join(backupOne, 'pg_wal'), { recursive: true });
|
||||
recordRepairAttempt(dir, 'failed', backupThree);
|
||||
sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeBackupPath).toBe(backupThree);
|
||||
|
||||
recordRepairAttempt(dir, 'repaired', backupThree);
|
||||
sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeStartedAt).toBeNull();
|
||||
expect(sidecar.episodeBackupPath).toBeNull();
|
||||
for (let i = 0; i < 15; i++) recordRepairAttempt(dir, 'repaired', null);
|
||||
expect(readRepairSidecar(dir).attempts.length).toBeLessThanOrEqual(10);
|
||||
});
|
||||
|
||||
test('unverified success (closeEpisode:false) keeps the episode open; closeRepairEpisodeIfOpen closes it', () => {
|
||||
const dir = makeLayout();
|
||||
const backup = `${dir}.wal-repair-backup-2001`;
|
||||
mkdirSync(join(backup, 'pg_wal'), { recursive: true });
|
||||
recordRepairAttempt(dir, 'failed', backup);
|
||||
// The manual command's unverified "repaired" must NOT close/prune.
|
||||
recordRepairAttempt(dir, 'repaired', backup, { closeEpisode: false });
|
||||
let sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeStartedAt).not.toBeNull();
|
||||
expect(existsSync(backup)).toBe(true);
|
||||
// A healthy connect closes it.
|
||||
closeRepairEpisodeIfOpen(dir);
|
||||
sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeStartedAt).toBeNull();
|
||||
expect(sidecar.episodeBackupPath).toBeNull();
|
||||
});
|
||||
|
||||
test('repairCooldownActive: active after a recent failure, respects the env knob', async () => {
|
||||
// Pin a known baseline (default cooldown, repair enabled): an ambient
|
||||
// GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS=0 would flip the assertions.
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR: undefined, GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: undefined }, async () => {
|
||||
const dir = makeLayout();
|
||||
expect(repairCooldownActive(dir).active).toBe(false);
|
||||
recordRepairAttempt(dir, 'failed', null);
|
||||
expect(repairCooldownActive(dir).active).toBe(true);
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: '0' }, async () => {
|
||||
expect(repairCooldownActive(dir).active).toBe(false);
|
||||
});
|
||||
// A success clears nothing retroactively, but cooldown keys on the LAST
|
||||
// failed attempt — still inside the window here.
|
||||
recordRepairAttempt(dir, 'repaired', null);
|
||||
expect(repairCooldownActive(dir).active).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
test('pruneRepairBackups keeps the newest 3 and never the open episode backup', () => {
|
||||
const dir = makeLayout();
|
||||
const parentBackups: string[] = [];
|
||||
for (let i = 1; i <= 5; i++) {
|
||||
const b = `${dir}.wal-repair-backup-${1000 + i}`;
|
||||
mkdirSync(b, { recursive: true });
|
||||
parentBackups.push(b);
|
||||
}
|
||||
// Pin the OLDEST as the open episode's backup.
|
||||
recordRepairAttempt(dir, 'failed', parentBackups[0]!);
|
||||
pruneRepairBackups(dir);
|
||||
const kept = listRepairBackups(dir);
|
||||
// Newest 3 + the protected episode backup.
|
||||
expect(kept).toContain(parentBackups[0]!);
|
||||
expect(kept).toContain(parentBackups[4]!);
|
||||
expect(kept).toContain(parentBackups[3]!);
|
||||
expect(kept).toContain(parentBackups[2]!);
|
||||
expect(kept).not.toContain(parentBackups[1]!);
|
||||
});
|
||||
});
|
||||
|
||||
describe('attemptWalRepairAndRetry — the never-throws engine seam', () => {
|
||||
test('repaired: retryCreate succeeds → db returned, episode closed, notice printed', async () => {
|
||||
const dir = makeLayout();
|
||||
const stderrChunks: string[] = [];
|
||||
const origWrite = process.stderr.write.bind(process.stderr);
|
||||
process.stderr.write = ((chunk: string | Uint8Array) => {
|
||||
stderrChunks.push(String(chunk));
|
||||
return true;
|
||||
}) as typeof process.stderr.write;
|
||||
try {
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => 'the-db-handle');
|
||||
expect(attempt.status).toBe('repaired');
|
||||
if (attempt.status === 'repaired') {
|
||||
expect(attempt.db).toBe('the-db-handle');
|
||||
expect(attempt.receipt.resetSegment).toMatch(/^[0-9A-F]{24}$/);
|
||||
}
|
||||
} finally {
|
||||
process.stderr.write = origWrite;
|
||||
}
|
||||
expect(stderrChunks.join('')).toContain('gbrain pglite-repair');
|
||||
expect(readRepairSidecar(dir).episodeStartedAt).toBeNull();
|
||||
expect(readRepairSidecar(dir).attempts.at(-1)?.outcome).toBe('repaired');
|
||||
});
|
||||
|
||||
test('failed + restored: retryCreate keeps throwing → dir restored, episode opened', async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg] });
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => {
|
||||
throw new Error('Aborted(). still broken');
|
||||
});
|
||||
expect(attempt.status).toBe('failed');
|
||||
if (attempt.status === 'failed') {
|
||||
expect(attempt.restored).toBe(true);
|
||||
expect(attempt.receipt).not.toBeNull();
|
||||
expect(attempt.repairError).toContain('still broken');
|
||||
}
|
||||
// Original segment is back in place.
|
||||
expect(existsSync(join(dir, 'pg_wal', seg))).toBe(true);
|
||||
const sidecar = readRepairSidecar(dir);
|
||||
expect(sidecar.episodeStartedAt).not.toBeNull();
|
||||
expect(sidecar.attempts.at(-1)?.outcome).toBe('failed');
|
||||
});
|
||||
|
||||
test('failed + restored:false (8A): restore blocked by the mtime guard is reported honestly', async () => {
|
||||
const dir = makeLayout();
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => {
|
||||
// Simulate a live writer advancing pg_wal between repair and restore.
|
||||
const foreign = xlogFileName(1, 99n, SEG_SIZE);
|
||||
writeFileSync(join(dir, 'pg_wal', foreign), 'live-writer-bytes');
|
||||
const future = new Date(Date.now() + 60_000);
|
||||
utimesSync(join(dir, 'pg_wal', foreign), future, future);
|
||||
throw new Error('Aborted(). still broken');
|
||||
});
|
||||
expect(attempt.status).toBe('failed');
|
||||
if (attempt.status === 'failed') {
|
||||
expect(attempt.restored).toBe(false);
|
||||
expect(attempt.repairError).toContain('mtime-guard');
|
||||
}
|
||||
});
|
||||
|
||||
test('guards: disabled / reaped lock / validation-failed — no backup dir is ever created', async () => {
|
||||
const before = process.env.GBRAIN_PGLITE_WAL_REPAIR;
|
||||
const dir = makeLayout();
|
||||
// Pin a known baseline (repair enabled, default cooldown) so an ambient
|
||||
// GBRAIN_PGLITE_WAL_REPAIR=off can't turn every arm into 'disabled'.
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR: undefined, GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: undefined }, async () => {
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR: 'off' }, async () => {
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => 'x');
|
||||
expect(attempt).toMatchObject({ status: 'skipped', reason: 'disabled' });
|
||||
});
|
||||
const reaped = await attemptWalRepairAndRetry(dir, async () => 'x', { reaped: true });
|
||||
expect(reaped).toMatchObject({ status: 'skipped', reason: 'possibly-live-writer' });
|
||||
const invalid = await attemptWalRepairAndRetry('/nope/never', async () => 'x');
|
||||
expect(invalid).toMatchObject({ status: 'skipped', reason: 'validation-failed' });
|
||||
expect(listRepairBackups(dir)).toEqual([]);
|
||||
});
|
||||
// withEnv restored whatever the ambient value was (including "unset").
|
||||
expect(process.env.GBRAIN_PGLITE_WAL_REPAIR).toBe(before);
|
||||
});
|
||||
|
||||
test('cooldown skip + episode backup reuse across attempts', async () => {
|
||||
// Pin a known baseline: an ambient COOLDOWN_SECONDS=0 would break the
|
||||
// 'recently-failed' gate assertion; an ambient WAL_REPAIR=off breaks all.
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR: undefined, GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: undefined }, async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg] });
|
||||
// Attempt 1 fails → episode opens with backup #1.
|
||||
const first = await attemptWalRepairAndRetry(dir, async () => { throw new Error('Aborted()'); });
|
||||
expect(first.status).toBe('failed');
|
||||
const backupsAfterFirst = listRepairBackups(dir);
|
||||
expect(backupsAfterFirst.length).toBe(1);
|
||||
|
||||
// Immediate retry is cooldown-gated…
|
||||
const gated = await attemptWalRepairAndRetry(dir, async () => 'x');
|
||||
expect(gated).toMatchObject({ status: 'skipped', reason: 'recently-failed' });
|
||||
|
||||
// …and with the cooldown off, the retry takes a FRESH backup: attempt 1's
|
||||
// restore MOVED pg_wal back out of its backup, so reusing that gutted dir
|
||||
// would let resetWal destroy the only surviving WAL copy (red-team).
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: '0' }, async () => {
|
||||
const second = await attemptWalRepairAndRetry(dir, async () => 'db');
|
||||
expect(second.status).toBe('repaired');
|
||||
if (second.status === 'repaired') {
|
||||
expect(second.receipt.reusedEpisodeBackup).toBe(false);
|
||||
expect(second.receipt.backupPath).not.toBe(backupsAfterFirst[0]!);
|
||||
}
|
||||
});
|
||||
expect(listRepairBackups(dir).length).toBe(2);
|
||||
expect(readRepairSidecar(dir).episodeStartedAt).toBeNull(); // episode closed
|
||||
});
|
||||
});
|
||||
|
||||
test('episode backup IS reused when it still holds pg_wal (restore was blocked)', async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg] });
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR: undefined, GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: '0' }, async () => {
|
||||
// Attempt 1: retry fails AND restore is blocked by the mtime guard
|
||||
// (a foreign future-dated segment appears mid-attempt) — the backup
|
||||
// KEEPS pg_wal.
|
||||
const first = await attemptWalRepairAndRetry(dir, async () => {
|
||||
const foreign = xlogFileName(1, 99n, SEG_SIZE);
|
||||
writeFileSync(join(dir, 'pg_wal', foreign), 'live-writer-bytes');
|
||||
const future = new Date(Date.now() + 60_000);
|
||||
utimesSync(join(dir, 'pg_wal', foreign), future, future);
|
||||
throw new Error('Aborted(). still broken');
|
||||
});
|
||||
expect(first.status).toBe('failed');
|
||||
if (first.status === 'failed') expect(first.restored).toBe(false);
|
||||
const episodeBackup = readRepairSidecar(dir).episodeBackupPath!;
|
||||
expect(existsSync(join(episodeBackup, 'pg_wal'))).toBe(true);
|
||||
// Clear the foreign segment so attempt 2's surgery isn't re-blocked.
|
||||
rmSync(join(dir, 'pg_wal'), { recursive: true, force: true });
|
||||
mkdirSync(join(dir, 'pg_wal', 'archive_status'), { recursive: true });
|
||||
const second = await attemptWalRepairAndRetry(dir, async () => 'db');
|
||||
expect(second.status).toBe('repaired');
|
||||
if (second.status === 'repaired') {
|
||||
expect(second.receipt.reusedEpisodeBackup).toBe(true);
|
||||
expect(second.receipt.backupPath).toBe(episodeBackup);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
test('seam reports honest restored from WalRepairError (reset fails after backup)', async () => {
|
||||
const seg = xlogFileName(1, 3n, SEG_SIZE);
|
||||
const dir = makeLayout({ segments: [seg] });
|
||||
poisonControlVersion(dir); // backup succeeds, resetWal throws → WalRepairError
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => 'x');
|
||||
expect(attempt.status).toBe('failed');
|
||||
if (attempt.status === 'failed') {
|
||||
// `restored` is threaded from WalRepairError.restore — not hardcoded.
|
||||
expect(attempt.restored).toBe(true);
|
||||
expect(attempt.receipt).not.toBeNull();
|
||||
expect(attempt.repairError).toContain('pg_control version');
|
||||
}
|
||||
// Restore actually happened: original segment is back.
|
||||
expect(existsSync(join(dir, 'pg_wal', seg))).toBe(true);
|
||||
});
|
||||
|
||||
test('poisoned sidecar episodeBackupPath outside the backup prefix is IGNORED — fresh backup taken', async () => {
|
||||
const dir = makeLayout();
|
||||
// An existing dir that fails the `${dataDir}.wal-repair-backup-` prefix
|
||||
// check: the user-writable sidecar must not be able to point repair's
|
||||
// renames at an arbitrary target.
|
||||
const evil = mkdtempSync(join(tmpdir(), 'pgrepair-evil-'));
|
||||
writeFileSync(`${dir}.wal-repair-attempt.json`, JSON.stringify({
|
||||
episodeStartedAt: Date.now(),
|
||||
episodeBackupPath: evil,
|
||||
attempts: [],
|
||||
}));
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: '0' }, async () => {
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => 'db');
|
||||
expect(attempt.status).toBe('repaired');
|
||||
if (attempt.status === 'repaired') {
|
||||
expect(attempt.receipt.reusedEpisodeBackup).toBe(false);
|
||||
expect(attempt.receipt.backupPath.startsWith(`${dir}.wal-repair-backup-`)).toBe(true);
|
||||
}
|
||||
});
|
||||
// The poisoned target was never renamed into or written through.
|
||||
expect(existsSync(evil)).toBe(true);
|
||||
expect(readdirSync(evil)).toEqual([]);
|
||||
});
|
||||
|
||||
test('reap quarantine gates the seam; a marker older than the window does not', async () => {
|
||||
const dir = makeLayout();
|
||||
const marker = `${dir}.lock-reap.json`;
|
||||
writeFileSync(marker, JSON.stringify({ ts: Date.now(), by: 1 }));
|
||||
|
||||
const gated = await attemptWalRepairAndRetry(dir, async () => 'x');
|
||||
expect(gated).toMatchObject({ status: 'skipped', reason: 'possibly-live-writer' });
|
||||
if (gated.status === 'skipped') expect(gated.detail).toContain('reaped');
|
||||
// Gated BEFORE any surgery: no backup dir was created.
|
||||
expect(listRepairBackups(dir)).toEqual([]);
|
||||
|
||||
// Marker older than the 10-minute quarantine → the seam proceeds.
|
||||
writeFileSync(marker, JSON.stringify({ ts: Date.now() - 11 * 60 * 1000, by: 1 }));
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS: '0' }, async () => {
|
||||
const attempt = await attemptWalRepairAndRetry(dir, async () => 'db');
|
||||
expect(attempt.status).toBe('repaired');
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe('inspectPgliteDataDir', () => {
|
||||
test('verdicts: missing / unsupported / wal-corruption-likely / looks-healthy', () => {
|
||||
expect(inspectPgliteDataDir('/nope/never').verdict).toBe('missing');
|
||||
expect(inspectPgliteDataDir(mkdtempSync(join(tmpdir(), 'pgrepair-'))).verdict).toBe('unsupported-layout');
|
||||
const withPid = makeLayout({ postmasterPid: true });
|
||||
expect(inspectPgliteDataDir(withPid).verdict).toBe('wal-corruption-likely');
|
||||
const clean = makeLayout();
|
||||
expect(inspectPgliteDataDir(clean).verdict).toBe('looks-healthy');
|
||||
});
|
||||
|
||||
test('locked verdict for a live-PID lock; open episode reads as corruption-likely', () => {
|
||||
const dir = makeLayout();
|
||||
mkdirSync(join(dir, '.gbrain-lock'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(dir, '.gbrain-lock', 'lock'),
|
||||
JSON.stringify({ pid: process.pid, acquired_at: Date.now(), refreshed_at: Date.now(), command: 'gbrain embed', subcommand: 'embed' }),
|
||||
);
|
||||
const diag = inspectPgliteDataDir(dir);
|
||||
expect(diag.verdict).toBe('locked');
|
||||
expect(diag.lockHolderPid).toBe(process.pid);
|
||||
rmSync(join(dir, '.gbrain-lock'), { recursive: true });
|
||||
|
||||
recordRepairAttempt(dir, 'failed', null); // opens an episode
|
||||
expect(inspectPgliteDataDir(dir).verdict).toBe('wal-corruption-likely');
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,190 @@
|
||||
/**
|
||||
* Unit tests for the pg_resetwal port (src/core/pglite-resetwal.ts).
|
||||
*
|
||||
* Everything here runs on SYNTHETIC PG17 layouts (hand-built pg_control
|
||||
* buffers) — fast, parallel-safe, no PGLite. The real-engine proof (corrupt a
|
||||
* real brain's WAL → reopen → row readable) lives in
|
||||
* test/pglite-wal-repair.serial.test.ts.
|
||||
*/
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, existsSync, readFileSync, readdirSync, statSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import {
|
||||
resetWal,
|
||||
crc32c,
|
||||
parseWalSegNo,
|
||||
xlogFileName,
|
||||
WalResetUnsupportedError,
|
||||
} from '../src/core/pglite-resetwal.ts';
|
||||
|
||||
const SEG_SIZE = 1024 * 1024; // 1MB — valid (power of two, divides 2^32), fast to write
|
||||
|
||||
const OFF = {
|
||||
systemIdentifier: 0,
|
||||
pgControlVersion: 8,
|
||||
state: 16,
|
||||
checkPoint: 32,
|
||||
checkPointCopyRedo: 40,
|
||||
checkPointCopyThisTimeLineID: 48,
|
||||
xlogBlcksz: 224,
|
||||
xlogSegSize: 228,
|
||||
crc: 288,
|
||||
} as const;
|
||||
|
||||
function makeControl(opts?: { version?: number; segSize?: number; blcksz?: number; tli?: number; redoSegNo?: bigint }): Buffer {
|
||||
const control = Buffer.alloc(8192);
|
||||
control.writeBigUInt64LE(0x1122334455667788n, OFF.systemIdentifier);
|
||||
control.writeUInt32LE(opts?.version ?? 1700, OFF.pgControlVersion);
|
||||
control.writeUInt32LE(opts?.tli ?? 1, OFF.checkPointCopyThisTimeLineID);
|
||||
control.writeUInt32LE(opts?.blcksz ?? 8192, OFF.xlogBlcksz);
|
||||
control.writeUInt32LE(opts?.segSize ?? SEG_SIZE, OFF.xlogSegSize);
|
||||
const redoSegNo = opts?.redoSegNo ?? 3n;
|
||||
control.writeBigUInt64LE(redoSegNo * BigInt(opts?.segSize ?? SEG_SIZE) + 40n, OFF.checkPointCopyRedo);
|
||||
control.writeUInt32LE(crc32c([control.subarray(0, OFF.crc)]), OFF.crc); // valid CRC
|
||||
return control;
|
||||
}
|
||||
|
||||
function makeLayout(opts?: Parameters<typeof makeControl>[0] & { pgVersion?: string; segments?: string[] }): string {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'resetwal-'));
|
||||
writeFileSync(join(dir, 'PG_VERSION'), `${opts?.pgVersion ?? '17'}\n`);
|
||||
mkdirSync(join(dir, 'global'), { recursive: true });
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), makeControl(opts));
|
||||
mkdirSync(join(dir, 'base'), { recursive: true });
|
||||
mkdirSync(join(dir, 'pg_wal', 'archive_status'), { recursive: true });
|
||||
for (const seg of opts?.segments ?? []) {
|
||||
writeFileSync(join(dir, 'pg_wal', seg), Buffer.alloc(1024, 0xaa));
|
||||
}
|
||||
return dir;
|
||||
}
|
||||
|
||||
describe('resetWal — validation refusals (fail-closed)', () => {
|
||||
test('refuses PG_VERSION 16', async () => {
|
||||
const dir = makeLayout({ pgVersion: '16' });
|
||||
await expect(resetWal(dir)).rejects.toThrow(WalResetUnsupportedError);
|
||||
});
|
||||
|
||||
test('refuses missing PG_VERSION', async () => {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'resetwal-'));
|
||||
await expect(resetWal(dir)).rejects.toThrow(WalResetUnsupportedError);
|
||||
});
|
||||
|
||||
test('refuses wrong pg_control size', async () => {
|
||||
const dir = makeLayout();
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), Buffer.alloc(100));
|
||||
await expect(resetWal(dir)).rejects.toThrow(/pg_control size/);
|
||||
});
|
||||
|
||||
test('refuses wrong pg_control version', async () => {
|
||||
const dir = makeLayout({ version: 1600 });
|
||||
await expect(resetWal(dir)).rejects.toThrow(/pg_control version/);
|
||||
});
|
||||
|
||||
test('refuses non-power-of-two WAL segment size', async () => {
|
||||
const dir = makeLayout({ segSize: 3 * 1024 * 1024 });
|
||||
await expect(resetWal(dir)).rejects.toThrow(/segment size/);
|
||||
});
|
||||
|
||||
test('refuses unsupported WAL block size', async () => {
|
||||
const dir = makeLayout({ blcksz: 4096 });
|
||||
await expect(resetWal(dir)).rejects.toThrow(/block size/);
|
||||
});
|
||||
|
||||
test('refuses a pg_control whose stored CRC does not verify (F6: no laundering corrupt counters)', async () => {
|
||||
const dir = makeLayout();
|
||||
// A structurally-valid control (right size/version/seg/block) but with a
|
||||
// damaged checkpoint copy and a STALE crc — real pg_resetwal refuses this.
|
||||
const control = readFileSync(join(dir, 'global', 'pg_control'));
|
||||
control.writeBigUInt64LE(0xdeadbeefn, 56); // trash a checkpointCopy field
|
||||
// leave the old CRC in place → mismatch
|
||||
writeFileSync(join(dir, 'global', 'pg_control'), control);
|
||||
await expect(resetWal(dir)).rejects.toThrow(/CRC mismatch/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('resetWal — byte surgery on a synthetic PG17 layout', () => {
|
||||
test('resets WAL: pid removed, old segments gone, fresh checkpoint segment + CRC-valid control', async () => {
|
||||
const oldSegs = [xlogFileName(1, 3n, SEG_SIZE), xlogFileName(1, 4n, SEG_SIZE)];
|
||||
const dir = makeLayout({ redoSegNo: 3n, segments: oldSegs });
|
||||
writeFileSync(join(dir, 'postmaster.pid'), '12345\n');
|
||||
writeFileSync(join(dir, 'pg_wal', 'archive_status', `${oldSegs[0]}.ready`), '');
|
||||
mkdirSync(join(dir, 'pg_wal', 'summaries'), { recursive: true });
|
||||
writeFileSync(join(dir, 'pg_wal', 'summaries', `${'0'.repeat(40)}.summary`), '');
|
||||
|
||||
const result = await resetWal(dir);
|
||||
|
||||
// Stale run state + old WAL removed.
|
||||
expect(existsSync(join(dir, 'postmaster.pid'))).toBe(false);
|
||||
for (const seg of oldSegs) {
|
||||
expect(existsSync(join(dir, 'pg_wal', seg))).toBe(false);
|
||||
}
|
||||
expect(readdirSync(join(dir, 'pg_wal', 'archive_status'))).toEqual([]);
|
||||
expect(readdirSync(join(dir, 'pg_wal', 'summaries'))).toEqual([]);
|
||||
|
||||
// newSegNo = max(redo=3, existing max=4) + 1 = 5.
|
||||
expect(result.resetSegment).toBe(xlogFileName(1, 5n, SEG_SIZE));
|
||||
expect(result.timelineId).toBe(1);
|
||||
expect(result.walSegSize).toBe(SEG_SIZE);
|
||||
|
||||
const segPath = join(dir, 'pg_wal', result.resetSegment);
|
||||
expect(existsSync(segPath)).toBe(true);
|
||||
expect(statSync(segPath).size).toBe(SEG_SIZE);
|
||||
const wal = readFileSync(segPath);
|
||||
expect(wal.readUInt16LE(0)).toBe(0xd116); // XLOG_PAGE_MAGIC
|
||||
expect(wal.readUInt16LE(2) & 0x0002).toBe(0x0002); // XLP_LONG_HEADER
|
||||
|
||||
// Control: shutdown state + self-consistent CRC32C over bytes 0..288.
|
||||
const control = readFileSync(join(dir, 'global', 'pg_control'));
|
||||
expect(control.length).toBe(8192);
|
||||
expect(control.readInt32LE(OFF.state)).toBe(1); // DB_SHUTDOWNED
|
||||
expect(control.readUInt32LE(OFF.crc)).toBe(crc32c([control.subarray(0, OFF.crc)]));
|
||||
// checkPoint points into the new segment.
|
||||
const checkPoint = control.readBigUInt64LE(OFF.checkPoint);
|
||||
expect(checkPoint / BigInt(SEG_SIZE)).toBe(5n);
|
||||
|
||||
// No torn tmp files left behind (atomic-write hygiene).
|
||||
expect(readdirSync(join(dir, 'pg_wal')).filter((f) => f.includes('.tmp-'))).toEqual([]);
|
||||
expect(readdirSync(join(dir, 'global')).filter((f) => f.includes('.tmp-'))).toEqual([]);
|
||||
});
|
||||
|
||||
test('works on an EMPTY pg_wal (the whole-dir-rename backup path) and numbers off pg_control alone', async () => {
|
||||
const dir = makeLayout({ redoSegNo: 7n, segments: [] });
|
||||
const result = await resetWal(dir);
|
||||
// No existing segments — newSegNo = redo(7) + 1.
|
||||
expect(result.resetSegment).toBe(xlogFileName(1, 8n, SEG_SIZE));
|
||||
expect(existsSync(join(dir, 'pg_wal', 'archive_status'))).toBe(true);
|
||||
});
|
||||
|
||||
test('is idempotent: a second run converges (numbering keeps moving forward)', async () => {
|
||||
const dir = makeLayout({ redoSegNo: 3n });
|
||||
const first = await resetWal(dir);
|
||||
const second = await resetWal(dir);
|
||||
const firstNo = parseWalSegNo(first.resetSegment, SEG_SIZE)!;
|
||||
const secondNo = parseWalSegNo(second.resetSegment, SEG_SIZE)!;
|
||||
expect(secondNo).toBeGreaterThan(firstNo);
|
||||
// Exactly one segment remains after each run.
|
||||
const segs = readdirSync(join(dir, 'pg_wal')).filter((f) => /^[0-9A-F]{24}$/.test(f));
|
||||
expect(segs).toEqual([second.resetSegment]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('WAL segment name helpers', () => {
|
||||
test('parseWalSegNo / xlogFileName round-trip', () => {
|
||||
for (const segNo of [0n, 1n, 255n, 4096n, 0x1_0000_0000n / BigInt(SEG_SIZE) + 7n]) {
|
||||
const name = xlogFileName(1, segNo, SEG_SIZE);
|
||||
expect(name).toMatch(/^[0-9A-F]{24}$/);
|
||||
expect(parseWalSegNo(name, SEG_SIZE)).toBe(segNo);
|
||||
}
|
||||
});
|
||||
|
||||
test('parseWalSegNo rejects non-segment names', () => {
|
||||
expect(parseWalSegNo('archive_status', SEG_SIZE)).toBeNull();
|
||||
expect(parseWalSegNo('000000010000000000000001.partial', SEG_SIZE)).toBeNull();
|
||||
expect(parseWalSegNo('lowercase0000000000000001', SEG_SIZE)).toBeNull();
|
||||
});
|
||||
|
||||
test('crc32c matches a known vector', () => {
|
||||
// CRC-32C of ASCII "123456789" is 0xE3069283 (Castagnoli test vector).
|
||||
expect(crc32c([Buffer.from('123456789')])).toBe(0xe3069283);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,200 @@
|
||||
/**
|
||||
* The ported upstream regression (electric-sql/pglite PR #994) against a REAL
|
||||
* PGLite brain: create → insert → clean shutdown → corrupt the WAL → reopen
|
||||
* through PGLiteEngine.connect() → auto-repair fires → the original row is
|
||||
* still readable. Plus the kill-switch, gate-level negatives, and the #2084
|
||||
* exitCode pin.
|
||||
*
|
||||
* .serial: real PGLite cold starts + process.env writes (docs/TESTING.md R1).
|
||||
*/
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, existsSync, readdirSync, readFileSync, truncateSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join, dirname, basename } from 'node:path';
|
||||
import { withEnv } from './helpers/with-env.ts';
|
||||
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
||||
import type { EngineConfig } from '../src/core/types.ts';
|
||||
|
||||
const COLD_START_TIMEOUT = 120_000;
|
||||
|
||||
function engineConfig(dir: string): EngineConfig {
|
||||
return { engine: 'pglite', database_path: dir } as EngineConfig;
|
||||
}
|
||||
|
||||
/** Build a real brain with one probe row, cleanly shut down. */
|
||||
async function buildRealBrain(): Promise<string> {
|
||||
const dir = join(mkdtempSync(join(tmpdir(), 'walrepair-')), 'brain.pglite');
|
||||
const engine = new PGLiteEngine();
|
||||
await engine.connect(engineConfig(dir));
|
||||
await engine.db.exec('CREATE TABLE repair_probe (id int); INSERT INTO repair_probe VALUES (42);');
|
||||
await engine.disconnect();
|
||||
return dir;
|
||||
}
|
||||
|
||||
function walSegments(dir: string): string[] {
|
||||
return readdirSync(join(dir, 'pg_wal')).filter((f) => /^[0-9A-F]{24}$/.test(f));
|
||||
}
|
||||
|
||||
function backupDirs(dir: string): string[] {
|
||||
return readdirSync(dirname(dir)).filter((f) => f.startsWith(`${basename(dir)}.wal-repair-backup-`));
|
||||
}
|
||||
|
||||
function corruptAllSegments(dir: string, mode: 'truncate' | 'garbage'): void {
|
||||
const segs = walSegments(dir);
|
||||
expect(segs.length).toBeGreaterThan(0);
|
||||
for (const seg of segs) {
|
||||
const p = join(dir, 'pg_wal', seg);
|
||||
if (mode === 'truncate') {
|
||||
truncateSync(p, 1024);
|
||||
} else {
|
||||
// Overwrite the whole segment with garbage, keeping its size.
|
||||
const size = readFileSync(p).length;
|
||||
writeFileSync(p, Buffer.alloc(size, 0xff));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
async function connectExpectingRepair(dir: string): Promise<PGLiteEngine> {
|
||||
const engine = new PGLiteEngine();
|
||||
const warns: string[] = [];
|
||||
const origWarn = console.warn;
|
||||
console.warn = (...args: unknown[]) => { warns.push(args.join(' ')); };
|
||||
try {
|
||||
await engine.connect(engineConfig(dir));
|
||||
} finally {
|
||||
console.warn = origWarn;
|
||||
}
|
||||
// #2084 pin: the retry create's exit-status scribble is contained.
|
||||
expect(Number(process.exitCode ?? 0)).toBe(0);
|
||||
expect(engine.walRepairReceipt).not.toBeNull();
|
||||
expect(warns.join('\n')).toContain('repaired');
|
||||
return engine;
|
||||
}
|
||||
|
||||
describe('WAL auto-repair — real-brain regression (#223/#1670/#2575)', () => {
|
||||
test('case A (truncated WAL): connect() auto-repairs and the row survives', async () => {
|
||||
const dir = await buildRealBrain();
|
||||
corruptAllSegments(dir, 'truncate');
|
||||
|
||||
const engine = await connectExpectingRepair(dir);
|
||||
try {
|
||||
const receipt = engine.walRepairReceipt!;
|
||||
expect(existsSync(receipt.backupPath)).toBe(true);
|
||||
expect(existsSync(join(receipt.backupPath, 'pg_wal'))).toBe(true);
|
||||
const rows = await engine.db.query('SELECT id FROM repair_probe');
|
||||
expect((rows.rows[0] as { id: number }).id).toBe(42);
|
||||
} finally {
|
||||
await engine.disconnect();
|
||||
}
|
||||
}, COLD_START_TIMEOUT);
|
||||
|
||||
test('case B (garbage-overwritten WAL): connect() auto-repairs and the row survives', async () => {
|
||||
const dir = await buildRealBrain();
|
||||
corruptAllSegments(dir, 'garbage');
|
||||
|
||||
const engine = await connectExpectingRepair(dir);
|
||||
try {
|
||||
const rows = await engine.db.query('SELECT id FROM repair_probe');
|
||||
expect((rows.rows[0] as { id: number }).id).toBe(42);
|
||||
// A healthy reconnect afterwards does NOT re-fire repair.
|
||||
await engine.disconnect();
|
||||
const engine2 = new PGLiteEngine();
|
||||
await engine2.connect(engineConfig(dir));
|
||||
expect(engine2.walRepairReceipt).toBeNull();
|
||||
await engine2.disconnect();
|
||||
} finally {
|
||||
try { await engine.disconnect(); } catch { /* already disconnected */ }
|
||||
}
|
||||
}, COLD_START_TIMEOUT);
|
||||
|
||||
test('kill-switch: GBRAIN_PGLITE_WAL_REPAIR=off → honest error, no backup, lock released', async () => {
|
||||
const dir = await buildRealBrain();
|
||||
corruptAllSegments(dir, 'garbage');
|
||||
const backupsBefore = backupDirs(dir).length;
|
||||
|
||||
await withEnv({ GBRAIN_PGLITE_WAL_REPAIR: 'off' }, async () => {
|
||||
const engine = new PGLiteEngine();
|
||||
let message = '';
|
||||
try {
|
||||
await engine.connect(engineConfig(dir));
|
||||
throw new Error('connect unexpectedly succeeded');
|
||||
} catch (err) {
|
||||
message = String((err as Error).message);
|
||||
}
|
||||
expect(message).toContain('PGLite failed to initialize');
|
||||
expect(message).toContain('GBRAIN_PGLITE_WAL_REPAIR=off');
|
||||
expect(message).toContain('gbrain pglite-repair');
|
||||
expect(message).toContain('Original error:');
|
||||
});
|
||||
// No surgery happened…
|
||||
expect(backupDirs(dir).length).toBe(backupsBefore);
|
||||
// …and the lock was released: repair works on the next (enabled) connect.
|
||||
const engine = await connectExpectingRepair(dir);
|
||||
const rows = await engine.db.query('SELECT id FROM repair_probe');
|
||||
expect((rows.rows[0] as { id: number }).id).toBe(42);
|
||||
await engine.disconnect();
|
||||
}, COLD_START_TIMEOUT);
|
||||
|
||||
test('gate negative: a REAL wasm-abort with the cooldown gate active refuses repair — no new backup, honest skip reason', async () => {
|
||||
// A genuinely corrupt brain (create() aborts with the production
|
||||
// `RuntimeError: Aborted()` signature) whose sidecar records a fresh
|
||||
// failed attempt — the classifier says wasm-abort, but the cooldown gate
|
||||
// must refuse BEFORE any surgery. Pins the skip path end-to-end: verdict
|
||||
// fired, gate refused, zero backup dirs created, honest message.
|
||||
const { recordRepairAttempt } = await import('../src/core/pglite-repair.ts');
|
||||
const dir = await buildRealBrain();
|
||||
corruptAllSegments(dir, 'garbage');
|
||||
recordRepairAttempt(dir, 'failed', null);
|
||||
|
||||
const engine = new PGLiteEngine();
|
||||
let message = '';
|
||||
try {
|
||||
await engine.connect(engineConfig(dir));
|
||||
throw new Error('connect unexpectedly succeeded');
|
||||
} catch (err) {
|
||||
message = String((err as Error).message);
|
||||
}
|
||||
expect(message).toContain('PGLite failed to initialize');
|
||||
expect(message).toContain('Auto-repair skipped');
|
||||
expect(message).toContain('gbrain pglite-repair');
|
||||
expect(backupDirs(dir).length).toBe(0);
|
||||
}, COLD_START_TIMEOUT);
|
||||
|
||||
test('gate negative: symlinked data dir — Emscripten refuses the mount with a NAMED error, repair never fires', async () => {
|
||||
// PGLite's NODEFS cannot mount through a symlinked data dir: it throws a
|
||||
// message-less `ErrnoError { errno: 20 }`. Two pins: (a) the error
|
||||
// stringifier surfaces name+errno instead of "[object Object]" (#2674
|
||||
// class), (b) no repair surgery runs on either path.
|
||||
const { symlinkSync } = await import('node:fs');
|
||||
const real = await buildRealBrain();
|
||||
corruptAllSegments(real, 'garbage');
|
||||
const link = join(mkdtempSync(join(tmpdir(), 'walrepair-')), 'link.pglite');
|
||||
symlinkSync(real, link);
|
||||
|
||||
const engine = new PGLiteEngine();
|
||||
let message = '';
|
||||
try {
|
||||
await engine.connect(engineConfig(link));
|
||||
throw new Error('connect unexpectedly succeeded');
|
||||
} catch (err) {
|
||||
message = String((err as Error).message);
|
||||
}
|
||||
expect(message).toContain('PGLite failed to initialize');
|
||||
expect(message).not.toContain('[object Object]');
|
||||
expect(message).toContain('ErrnoError (errno 20)');
|
||||
expect(backupDirs(real).length).toBe(0);
|
||||
expect(backupDirs(link).length).toBe(0);
|
||||
}, COLD_START_TIMEOUT);
|
||||
|
||||
test('gate shape: the seam only runs for wasm-abort + persistent dataDir (structural pin)', () => {
|
||||
const src = readFileSync('src/core/pglite-engine.ts', 'utf-8');
|
||||
expect(src).toMatch(/if \(verdict === 'wasm-abort'\)/);
|
||||
expect(src).toMatch(/if \(!dataDir\) \{\s*\n\s*ctx = \{ repair: 'in-memory' \}/);
|
||||
// The seam call sits INSIDE the wasm-abort branch (no call site outside it).
|
||||
const firstSeamCall = src.indexOf('await attemptWalRepairAndRetry(');
|
||||
const gate = src.indexOf("if (verdict === 'wasm-abort')");
|
||||
expect(gate).toBeGreaterThan(-1);
|
||||
expect(firstSeamCall).toBeGreaterThan(gate);
|
||||
expect(src.indexOf('await attemptWalRepairAndRetry(', firstSeamCall + 1)).toBe(-1);
|
||||
});
|
||||
});
|
||||
@@ -424,20 +424,151 @@ describe('reinit-pglite — backup + reinit', () => {
|
||||
expect(exits).toContain(1);
|
||||
});
|
||||
|
||||
test('refuses when missing required --embedding-model / --embedding-dimensions', async () => {
|
||||
// ── Flag defaulting from the config FILE (eng-review 6A + codex 14.10) ──
|
||||
// Omitted --embedding-model / --embedding-dimensions default from
|
||||
// loadConfigFileOnly() (NOT loadConfig(): a transient outage-shell
|
||||
// GBRAIN_EMBEDDING_* export must not silently change the rebuild target).
|
||||
// Precedence: explicit flag > config-file value > missing_model/missing_dims.
|
||||
|
||||
/**
|
||||
* Run runReinitPglite with process.exit stubbed (same throw-on-exit
|
||||
* pattern as the tests above) and console.log/console.error captured,
|
||||
* so the plan output + defaulting notes are assertable.
|
||||
*/
|
||||
async function captureRun(args: string[]): Promise<{ exits: number[]; logs: string[]; errs: string[] }> {
|
||||
const { runReinitPglite } = await import('../src/commands/reinit-pglite.ts');
|
||||
const origExit = process.exit;
|
||||
const origLog = console.log;
|
||||
const origErr = console.error;
|
||||
const exits: number[] = [];
|
||||
const logs: string[] = [];
|
||||
const errs: string[] = [];
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
(process as any).exit = ((code?: number) => { exits.push(code ?? 0); throw new Error('exit:' + (code ?? 0)); });
|
||||
console.log = (...a: unknown[]) => { logs.push(a.map(String).join(' ')); };
|
||||
console.error = (...a: unknown[]) => { errs.push(a.map(String).join(' ')); };
|
||||
try {
|
||||
await runReinitPglite(['--json']);
|
||||
await runReinitPglite(args);
|
||||
} catch (e) {
|
||||
expect((e as Error).message).toMatch(/^exit:/);
|
||||
} finally {
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
(process as any).exit = origExit;
|
||||
console.log = origLog;
|
||||
console.error = origErr;
|
||||
}
|
||||
return { exits, logs, errs };
|
||||
}
|
||||
|
||||
test('no flags: defaults BOTH from the config file', async () => {
|
||||
// Sentinel: pre-create the .bak so the run halts at bak_exists AFTER
|
||||
// parseArgs + the plan print — proving the defaulting resolved from
|
||||
// the file without invoking the real (destructive) init+sync path.
|
||||
writeFileSync(join(tmpHome, '.gbrain', 'brain.pglite.bak'), 'sentinel');
|
||||
|
||||
const { exits, logs, errs } = await captureRun(['--yes']);
|
||||
|
||||
// Halted at the sentinel — parseArgs did NOT fail missing_model/missing_dims.
|
||||
expect(exits).toContain(1);
|
||||
const err = errs.join('\n');
|
||||
expect(err).toContain('Backup already exists');
|
||||
// The plan shows the config-file values.
|
||||
const out = logs.join('\n');
|
||||
expect(out).toContain('New embedding model: openai:text-embedding-3-large');
|
||||
expect(out).toMatch(/New dimensions:\s+1536/);
|
||||
// One stderr note per defaulted flag.
|
||||
expect(err).toContain('--embedding-model defaulted from config: openai:text-embedding-3-large');
|
||||
expect(err).toContain('--embedding-dimensions defaulted from config: 1536');
|
||||
});
|
||||
|
||||
test('no flags + config missing the values: still fails missing_model / missing_dims', async () => {
|
||||
const cfgPath = join(tmpHome, '.gbrain', 'config.json');
|
||||
|
||||
// Neither value in the file → missing_model (checked first).
|
||||
writeFileSync(cfgPath, JSON.stringify({
|
||||
engine: 'pglite',
|
||||
database_path: join(tmpHome, '.gbrain', 'brain.pglite'),
|
||||
}));
|
||||
const noModel = await captureRun(['--json']);
|
||||
expect(noModel.exits).toContain(1);
|
||||
const noModelPayload = JSON.parse(noModel.logs[noModel.logs.length - 1]);
|
||||
expect(noModelPayload.status).toBe('error');
|
||||
expect(noModelPayload.reason).toBe('missing_model');
|
||||
|
||||
// Model present but no dimensions → missing_dims.
|
||||
writeFileSync(cfgPath, JSON.stringify({
|
||||
engine: 'pglite',
|
||||
database_path: join(tmpHome, '.gbrain', 'brain.pglite'),
|
||||
embedding_model: 'openai:text-embedding-3-large',
|
||||
}));
|
||||
const noDims = await captureRun(['--json']);
|
||||
expect(noDims.exits).toContain(1);
|
||||
const noDimsPayload = JSON.parse(noDims.logs[noDims.logs.length - 1]);
|
||||
expect(noDimsPayload.reason).toBe('missing_dims');
|
||||
});
|
||||
|
||||
test('flag present but valueless still fails missing_model (no silent config fallback)', async () => {
|
||||
// A malformed explicit flag is a typo, not an omission — it must not
|
||||
// silently rebuild against whatever the config file happens to hold.
|
||||
const { exits, logs } = await captureRun(['--json', '--embedding-model']);
|
||||
expect(exits).toContain(1);
|
||||
const payload = JSON.parse(logs[logs.length - 1]);
|
||||
expect(payload.reason).toBe('missing_model');
|
||||
});
|
||||
|
||||
test('explicit flags win over config-file values', async () => {
|
||||
writeFileSync(join(tmpHome, '.gbrain', 'brain.pglite.bak'), 'sentinel');
|
||||
|
||||
const { exits, logs, errs } = await captureRun([
|
||||
'--embedding-model', 'zeroentropyai:zembed-1',
|
||||
'--embedding-dimensions', '1280',
|
||||
'--yes',
|
||||
]);
|
||||
|
||||
expect(exits).toContain(1); // bak_exists sentinel
|
||||
const out = logs.join('\n');
|
||||
expect(out).toContain('New embedding model: zeroentropyai:zembed-1');
|
||||
expect(out).toMatch(/New dimensions:\s+1280/);
|
||||
expect(out).not.toContain('openai:text-embedding-3-large');
|
||||
// No defaulting note when both values came from flags.
|
||||
expect(errs.join('\n')).not.toContain('defaulted from config');
|
||||
});
|
||||
|
||||
test('env poisoning: GBRAIN_EMBEDDING_* env is ignored — config FILE values win', async () => {
|
||||
writeFileSync(join(tmpHome, '.gbrain', 'brain.pglite.bak'), 'sentinel');
|
||||
|
||||
await withEnv({
|
||||
GBRAIN_EMBEDDING_MODEL: 'voyage:poisoned-model',
|
||||
GBRAIN_EMBEDDING_DIMENSIONS: '9999',
|
||||
}, async () => {
|
||||
const { exits, logs, errs } = await captureRun(['--yes']);
|
||||
|
||||
expect(exits).toContain(1); // bak_exists sentinel
|
||||
const out = logs.join('\n');
|
||||
expect(out).toContain('New embedding model: openai:text-embedding-3-large');
|
||||
expect(out).toMatch(/New dimensions:\s+1536/);
|
||||
expect(out).not.toContain('voyage:poisoned-model');
|
||||
// Scoped to the plan line — the tmpdir's random suffix in the path
|
||||
// lines could otherwise collide with a bare '9999' substring check.
|
||||
expect(out).not.toMatch(/New dimensions:\s+9999/);
|
||||
const err = errs.join('\n');
|
||||
expect(err).toContain('--embedding-model defaulted from config: openai:text-embedding-3-large');
|
||||
expect(err).toContain('--embedding-dimensions defaulted from config: 1536');
|
||||
});
|
||||
});
|
||||
|
||||
test('invalid_dims validation applies to the config-sourced value too', async () => {
|
||||
const cfgPath = join(tmpHome, '.gbrain', 'config.json');
|
||||
writeFileSync(cfgPath, JSON.stringify({
|
||||
engine: 'pglite',
|
||||
database_path: join(tmpHome, '.gbrain', 'brain.pglite'),
|
||||
embedding_model: 'openai:text-embedding-3-large',
|
||||
embedding_dimensions: -5,
|
||||
}));
|
||||
|
||||
const { exits, logs } = await captureRun(['--json']);
|
||||
expect(exits).toContain(1);
|
||||
const payload = JSON.parse(logs[logs.length - 1]);
|
||||
expect(payload.reason).toBe('invalid_dims');
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user