mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 00:48:18 +00:00
v0.42.75.0 fix(pglite): in-place WAL auto-repair for the macOS Aborted() startup crash (#2575, #223, #1670) (#3901)
* fix(pglite): in-place WAL auto-repair for the Aborted() startup crash (#223, #1670, #2575) The 'macOS 26.x WASM bug' was a misdiagnosis: an unclean shutdown (typically the OS-upgrade reboot) tears the data dir's WAL, and every subsequent open fails WAL replay inside WASM with an opaque RuntimeError: Aborted(). This ports the pg_resetwal recovery upstream rejected (electric-sql/pglite#994, by @yestheboxer) and wires it into connect() as bounded auto-repair: - src/core/pglite-resetwal.ts: pg_resetwal for PG17 NodeFS dirs, fail-closed layout validation, atomic+durable writes (tmp+fsync+rename), idempotent. - src/core/pglite-repair.ts: whole-pg_wal-dir rename backup (zero transient disk), overwrite-order restore with mtime guard, cooldown sidecar + episode-scoped backup retention (newest 3 episodes), and a never-throws engine seam. Kill-switch: GBRAIN_PGLITE_WAL_REPAIR=off. - pglite-engine.ts: verdict rename macos-26-3 -> wasm-abort, classifier now matches the real production message (it previously fell to 'unknown'), corrupt-beats-wasm precedence preserved, honest per-outcome error copy incl. the failed-not-restored arm, and repair only under a cleanly-acquired lock (new LockHandle.reaped provenance; never after reaping a holder). - gbrain pglite-repair: manual dry-run/repair command (validate-before-lock, serve/reaped refusals, no --force by design). - doctor: pglite_data_dir fs-check with recurrence escalation and backup inventory when a PGLite brain fails to connect. - reinit-pglite: embedding flags default from file-only config so the recovery ladder's rebuild rung works bare mid-outage. - stringifyPgliteInitError: message-less Emscripten ErrnoError objects no longer surface as [object Object]. Regression-tested against real brains: corrupt every WAL segment (truncate and garbage variants), reopen, auto-repair fires, original rows readable, process.exitCode stays contained (#2084). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(pglite): replace the macOS-26.x misdiagnosis with the corrupt-WAL recovery ladder README + INSTALL.md shipped (via #1671) the claim that PGLite is incompatible with macOS 26.x and that a Bun/WASM fix would restore it. The real cause is torn WAL state from the upgrade reboot, now auto-repaired in place. Rewrites those sections around the recovery ladder (auto-repair -> gbrain pglite-repair -> reinit-pglite -> engine switch; native-Postgres recipe kept, credit @roysaurav), adds the ENGINES.md troubleshooting section, updates the KEY_FILES.md entries to current truth, files the two follow-up TODOs (SIGTERM engine-close extension; pglite upgrade blocker), and regenerates the llms bundles. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pglite): harden WAL auto-repair (pre-landing + adversarial review) Review-army (security/testing/maintainability/perf) + Claude & Codex adversarial passes on the WAL-repair wave. Correctness + safety hardening, no behavior change to the happy path: - Live-writer safety: repair refuses any reaped lock acquisition, a corrupt (unknowable-liveness) reap writes a cross-process quarantine marker that gates auto-repair AND the manual command for 10 min, isProcessAlive treats only ESRCH as dead (EPERM/malformed-pid read as alive), and a live postmaster.pid (native Postgres) is refused. Lock heartbeat + initial write are atomic (tmp+rename) so a torn read can't misclassify a healthy holder; an in-flight acquisition is no longer mistaken for corrupt. - resetWal verifies the stored pg_control CRC before trusting/re-signing it — a damaged control file routes to rebuild instead of laundering corrupt checkpoint counters under a fresh CRC. Atomic 'wx' writes (no symlink follow), whole-pg_wal-dir rename backup, 64MB seg-size cap. - Honest failure reporting: repairPgliteWal threads the real restore result out via WalRepairError so the 'failed-restored' vs 'failed-not-restored' message never lies; the not-restored copy names the correct restore paths. - Episode lifecycle: episodes close on the next healthy connect (not just on a verified repair), a gutted (restored) backup loses its pin, stale (>24h) episode backups aren't reused, and the cooldown also caps repaired-only crash loops. Empty backup dirs are pruned on refusal. - Command: rejects unknown flags and valueless --path (a destructive command must not silently mis-parse), confirm prompt goes to stderr (stdout stays clean for --json), embedding-flag defaults come from the config file only. - Symlink confinement extended to global/; sidecar reuse path validated (prefix + no '..' + must still hold pg_wal); sidecar writes atomic. - doctor recurrence escalation counts all attempts; data dir absolutized. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(pglite): current-state KEY_FILES + WAL-repair follow-up TODOs KEY_FILES.md pglite entries updated to the hardened truth (reap marker + quarantine, atomic writes, CRC gate, global-symlink refusal, WalRepairError, episode lifecycle). TODOS.md files the deferred judgment-call follow-ups (unclean-shutdown gate on auto-repair; non-gbrain pglite consumer boundary; mixed-version torn-lock double-read). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v0.42.75.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
0b47afbf40
commit
f15480b9d0
@@ -221,6 +221,66 @@ live in `test/postgres-engine-rls-scope.test.ts`.
|
||||
|
||||
**Migration:** `gbrain migrate --to supabase` exports everything (pages, chunks, embeddings, links, tags, timeline) and imports into Supabase. `gbrain migrate --to pglite` goes the other direction. Bidirectional, lossless.
|
||||
|
||||
### Troubleshooting: startup abort (`RuntimeError: Aborted()`)
|
||||
|
||||
**Symptom:** every PGLite-touching command dies at startup with
|
||||
`PGLite failed to initialize its WASM runtime … Aborted(). Build with
|
||||
-sASSERTIONS for more info.` — commonly first seen right after a macOS
|
||||
upgrade.
|
||||
|
||||
**Real root cause:** corrupt WAL/checkpoint state in the data dir after an
|
||||
unclean shutdown (the OS-upgrade reboot kills gbrain mid-write and tears the
|
||||
write-ahead log; every subsequent open fails WAL replay inside WASM and
|
||||
Emscripten surfaces only the opaque abort). It is **not** a macOS/WASM
|
||||
incompatibility — the same signature reproduces across macOS versions and on
|
||||
Linux, and rebuilding the data dir on the same OS fixes it. No pglite or Bun
|
||||
version bump changes it.
|
||||
|
||||
**Recovery ladder** (top rung first):
|
||||
|
||||
1. **Auto-repair (default).** `PGLiteEngine.connect()` detects the abort,
|
||||
backs up `pg_wal/` + `pg_control` into a sibling
|
||||
`<dataDir>.wal-repair-backup-<ts>/` dir, resets the WAL in place
|
||||
(pg_resetwal semantics — data files preserved; transactions not
|
||||
checkpointed before the corruption may be lost), and retries once. On
|
||||
success it prints a loud stderr notice naming the backup and recommending
|
||||
`gbrain doctor`. Safety bounds: repair only runs under a cleanly-acquired
|
||||
data-dir lock (never after reaping another process's lock), skips for a
|
||||
cooldown window after a failed attempt
|
||||
(`GBRAIN_PGLITE_WAL_REPAIR_COOLDOWN_SECONDS`, default 3600), reuses one
|
||||
backup per corruption episode (newest 3 episodes retained), and restores
|
||||
the original files if the retry still fails. Kill-switch:
|
||||
`GBRAIN_PGLITE_WAL_REPAIR=off`.
|
||||
2. **Manual repair.** `gbrain pglite-repair --dry-run` diagnoses the data dir
|
||||
(read-only); `gbrain pglite-repair --yes` runs the same in-place WAL reset
|
||||
deliberately. Refuses when another gbrain process holds the brain (a live
|
||||
`gbrain serve` is named explicitly) and never force-removes `.gbrain-lock`.
|
||||
3. **Rebuild.** `gbrain reinit-pglite` (embedding model/dimensions default
|
||||
from your config) wipes and re-creates the brain from your brain repo, or
|
||||
manually: back up `~/.gbrain`, move `brain.pglite` aside,
|
||||
`gbrain init --pglite`, re-add sources, `gbrain sync`, `gbrain embed`.
|
||||
Required for *catalog* corruption (58P01 / pgvector load failure) — WAL
|
||||
repair cannot fix that class.
|
||||
4. **Switch engines.** `gbrain init --supabase`, or native Postgres +
|
||||
pgvector (recipe below, contributed by @roysaurav):
|
||||
|
||||
```bash
|
||||
brew install postgresql@17
|
||||
brew services start postgresql@17
|
||||
createdb gbrain
|
||||
cd /tmp && git clone --branch v0.8.0 https://github.com/pgvector/pgvector.git
|
||||
cd pgvector && make && make install
|
||||
psql gbrain -c "CREATE EXTENSION IF NOT EXISTS vector;"
|
||||
# ~/.gbrain/config.json: { "engine": "postgres",
|
||||
# "database_url": "postgresql://localhost:5432/gbrain" }
|
||||
gbrain apply-migrations --yes && gbrain doctor
|
||||
```
|
||||
|
||||
`gbrain doctor` runs a `pglite_data_dir` check whenever a PGLite brain fails
|
||||
to connect: it diagnoses the dir from disk, names the repair command, reports
|
||||
retained repair backups, and escalates when repairs keep recurring (that
|
||||
means the unclean-shutdown genesis is still active — see the ladder's rung 4).
|
||||
|
||||
## JSONB writes: never double-encode (the #2339 trap)
|
||||
|
||||
Writing a JS value into a `jsonb` column has exactly two correct forms. Get this
|
||||
|
||||
+15
-4
@@ -117,7 +117,20 @@ If anything's yellow, `gbrain doctor` names the fix command in the message. Most
|
||||
|
||||
### PGLite crashes on macOS 26.x (Tahoe)
|
||||
|
||||
PGLite's embedded WASM engine is incompatible with macOS 26.x (Tahoe) on Apple Silicon. If `gbrain init --pglite` crashes during engine initialization, switch to native Homebrew PostgreSQL:
|
||||
This crash (`RuntimeError: Aborted()` at engine startup, typically first seen
|
||||
after a macOS upgrade) is **not** a macOS/WASM incompatibility. The upgrade
|
||||
reboot kills gbrain mid-write and tears the data dir's write-ahead log; every
|
||||
subsequent open then fails WAL replay. Recovery ladder:
|
||||
|
||||
1. **Auto-repair (default):** just run any gbrain command — gbrain detects the
|
||||
abort, resets the WAL in place (data preserved; a backup of the pre-repair
|
||||
state is kept next to the data dir), and continues. Then run `gbrain doctor`.
|
||||
2. **Manual repair:** `gbrain pglite-repair --dry-run` to diagnose,
|
||||
`gbrain pglite-repair --yes` to repair in place.
|
||||
3. **Rebuild:** `gbrain reinit-pglite` (wipes and re-creates the brain from
|
||||
your brain repo; embedding settings default from your config).
|
||||
4. **Switch engines** — if you prefer a server database anyway, native
|
||||
Homebrew PostgreSQL works great and supports multiple concurrent agents:
|
||||
|
||||
```bash
|
||||
# Install PostgreSQL + pgvector
|
||||
@@ -144,6 +157,4 @@ gbrain apply-migrations --yes
|
||||
gbrain doctor
|
||||
```
|
||||
|
||||
All 102 migrations run on first try. Once `gbrain doctor` shows green, the brain works identically to PGLite — same commands, same skills, same data model. The only difference is the storage backend.
|
||||
|
||||
> **Note:** This workaround is temporary. When the upstream WASM runtime fix ships (likely via a Bun update), `--pglite` will work on Tahoe again.
|
||||
Once `gbrain doctor` shows green, the brain works identically to PGLite — same commands, same skills, same data model. The only difference is the storage backend (plus multi-connection support: several agents can share one Postgres brain, which PGLite's single-process lock doesn't allow).
|
||||
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user