* docs(designs): agent-bootstrap plan + design docs (normative, review-absorbed) The scrubbed, in-repo sources of truth for the gbrain bootstrap wave: AGENT_BOOTSTRAP_DESIGN.md (product scope/sequencing) and AGENT_BOOTSTRAP_PLAN.md (implementation; all review-finding IDs inlined). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): format spec, question bank, identity templates, bundled assets agent.json manifest (format_version 1, initialized sentinel) + machine-local install receipt [CX2-1, CX2-12]; 12-question/6-required interview bank with consent keys and a persist:false sink for the optional provider key [CX2-13]; ten {{TOKEN}} identity templates (generic, adapted to gbrain ops — gates call recall/query/put_page, write-through-ops rule, keyless agent-authored facts, silence contract); assets embedded compiled-binary-safe via file-type imports [ENG-6]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): runbook, README paste block, bootstrap guide, TODOS entries BOOTSTRAP_FOR_AGENTS.md (agent-driven install runbook: CLI phase list is the source of truth, never-invent rules, Codex approvals preflight, keyless posture, failure-modes table, version stamp for the skew check); README gains the full-agent paste block pinned to latest-stable inside the Claude Code/Codex quick start (memory-only tier stays); docs/guides/bootstrap.md carries the full install/security/consent/degradation/uninstall contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(designs): spike instrument for the bootstrap wave (build order 0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): interview + render engines Interview gate with read-back confirm-hash (any later answer change clears the confirmation — the hostile single-batch case is structurally impossible), per-answer provenance, caps + escaping at set time, config-sink routing for the provider key; renderer with hard-fail token sweep, subordinate fencing of principal input, never-clobber + backups, deterministic minimal mode for the template repo, scaled byte floors. 58 unit tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): private-repo lifecycle — repo create, attach, uninstall, run lock gh-gated private repo creation with API-verified privacy (rate-limit distinct from public), refuse-foreign-origin with attach as the sanctioned path, atomic bootstrap mutex (pid liveness + age + token), receipt-keyed uninstall that never wholesale-deletes the gbrain home and only offers --delete-brain for a brain it created; read-only PGLite lock probe (never opens the engine). 54 unit tests, injectable exec seam throughout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(release): latest-stable ref, template-repo publish job, bootstrap CI guards release.yml advances the latest-stable tag only after assets publish (the paste block's permanent ref — copies in the wild never rot) and gains a PAT-gated publish-template job verified against the vendored tree; two skip-graceful guards (sanctioned-ref + runbook stamp; template/token bijection + placeholder assertion + generator byte-diff) wired into verify; README + runbook re-admitted to the CI cache hash; vendored deterministic template tree generated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(context): IPC v2 turn_context + 8KB assembly + visibility resolver + session identity Discriminated-union IPC with handler map, protocol echo (stale-serve detection), shared-secret gate, server-side source binding, per-kind budgets; turn-context assembly (reflex pointers + volunteered pages + world-only hot facts) under a data-not-instructions envelope trimmed to the harness's 10KB hook-output cap; facts.default_visibility resolved through one helper at all four sites (explicit caller wins, typos fail closed); typed sessionId threads _meta.session_id into the hot-memory cache key. 50 new tests; 180 adjacent tests confirmed green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(persistence): secret-scan, gbrain sources push, durability unification Pattern secret scanner (own runtime allowlist, redacted previews, corpus-write redaction mode); sources push runs the whole scan→stage→commit→pull→push sequence under one cross-platform lock (mkdir-atomic, pid+age+token) with a deny-glob backstop, commit-first divergence-safe pull, refuse-unverifiable visibility, and push-status telemetry; gbrain-home choke point unifies GBRAIN_HOME semantics with config (0700); durability is parent-repo-aware and rotates its push log at 0600. 35 new tests; 200 existing green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sources): harden/pull gates accept sources inside a parent git repo The bootstrap workspace registers brain/ (a subdirectory) as the source; the durability core already resolves the repo root, so the command gates now check inside-a-repo rather than .git-right-here [CX2-3]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(serve): resident maintenance sweep + keyless capability probe The lock-owning serve process now closes the persistence loop: startup (3s post-connect, best-effort, unref'd) and idle (10-min quiet intervals through the injectable timer seam) sweeps run facts-fence reconciliation, deterministic link/timeline extraction over recent workspace pages, and spend-gated corpus ingest (skipped keyless — agent-authored fences cover it). gbrain sweep --once is the trusted CLI seam bootstrap verify uses. Capability probe renders the honest keyless/keyed report. Full reuse of the cycle extractor + extract cores; 26 new tests, neighbors green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(hooks): engine-free gbrain hook command, settings writers, transcript parser Four hook events (session-start digest + crashed-session recovery push, user-prompt turn-context injection under an 800ms deadline and the 10KB cap, stop buffers, session-end corpus write with redaction/retention/dedup + best-effort push); structural JSON settings merger keyed by a _gbrain marker (foreign hooks and permissions survive); dated host-spec registry; Claude Code .jsonl parser as a spec-target with a scrubbed 7-shape fixture. Heartbeat is counters-only by construction. 59 tests; zero engine modules in the import graph. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(bootstrap): cross-link the full-agent path from the connection docs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): dispatcher, verify, status — the command assembled gbrain bootstrap {status,interview,render,repo,hooks,verify,uninstall,attach}: engine-free except verify (owns its engine, in-process sweep — no live-serve conflict); phase list is the TS source of truth with install.jsonl telemetry and the support blob; verify's fail-soft check suite covers the real write path (put_page → write-through file → sweep → graph floor → recall), passes keyless, persists snapshots, and ends with the first-run tour. cli.ts wired per the three-touchpoint rule; doctor gains the bootstrap check group (silent on machines with no bootstrap state). 28 new tests; 353 adjacent green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: KEY_FILES bootstrap cluster + CLAUDE.md dispatcher row (+ build:llms) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): e2e pins — hook-under-live-serve, attach, degraded modes, compiled binary, Docker harness The permanent pins: a real serve holds the PGLite lock while the engine-free hook completes (and a direct engine open provably throws LiveServeLockError); stale-socket fail-open; machine-2 attach with marker-keyed hook repair; decline-everything installs verify green with every degradation named; the compiled binary renders bundled templates in an empty cwd. Offline Docker harness (networkless, read-only) gated into heavy-tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bootstrap): register doctor check categories + system-of-record allow comments The six bootstrap doctor checks join OPS_CHECK_NAMES; the sweep's batch link/ timeline inserts carry the explicit extract-path allow comments (the sweep IS the extraction path for workspace pages). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard wedge cap tracks suite growth (1500s -> 1800s) At ~9000 tests a healthy shard finished at 1466s and two progressing shards were false-killed at the old cap; 1800s restores ~25% headroom over the slowest observed healthy shard. Real hangs still hit it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ci): cache-hash policy — README + runbook edits must invalidate [C2] The old deny-list assertion predates the paste block; README.md and BOOTSTRAP_FOR_AGENTS.md are policy-doc re-admissions now, so their edits must change the hash (a paste-block edit shipping under a cached green was the C2 hole). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): classify post-suite exit-hangs as warn-pass; file the leak forensics A shard killed by the wedge watchdog with every assigned file started and zero fail markers did all its work and leaked a handle at exit — pre-existing and master-reproducible (P1 TODO carries the full bisect forensics). Bun's per-test timeout turns a hung test into a (fail), so the classifier cannot mask one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard cap 2400s — the count-balanced heavy shard needs it under contention Observed: the heavy shard still progressing 22s before an 1800s kill while siblings finish at 1150-1550s (split balances file count, not weight). Filed the load-sensitive WAL-repair flake (pre-existing, master's v0.42.75.0 wave). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): quarantine env-mutating suites to the serial lane check-test-isolation R1: six new files mutate GBRAIN_HOME/env at module scope — the serial lane (one process per file) is the guard's prescribed home for them. All 114 tests pass post-rename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): per-harness install sections — Codex, Claude Code, then OpenClaw/Hermes Each harness gets its own complete paste-block section (desktop app first, terminal noted — Claude Code CLI is the identical harness; Codex CLI works pull-based today); the OpenClaw/Hermes platform path keeps equal weight with its one-click deploys and INSTALL_FOR_AGENTS block intact; memory-only and remote-connect tiers consolidated under 'Lighter ways in'. Supersedes the review's D5 ordering by user direction; stale heading references updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): Codex as the recommended first step; OpenClaw/Hermes framed as-intended, high-cost The install section now routes newcomers explicitly: Codex first (subscription-priced, nothing to deploy), OpenClaw/Hermes as GBrain used the way it was designed — always on, at real server + API cost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: doc-audit code chasers — broken recovery hints, stale op description, auto_link key Four small code fixes surfaced by the markdown accuracy audit: - doctor's auto-RLS recovery hint pointed at `apply-migrations --force-retry 35`, which cannot work (--force-retry targets the vX.Y.Z orchestrator registry, not the numeric schema MIGRATIONS array). Hint now points at the recreate SQL in docs/guides/rls-and-you.md; test pins against regression. - v0_11_0 migration printed the same broken-mechanism class of hint (`config set minion_mode` writes DB config nothing reads); now names `apply-migrations --mode` + preferences.json, the real setter. - submit_job's op description hardcoded a stale handler list; now points at registerBuiltinHandlers as the source plus the --follow discovery trick. - `auto_link` added to KNOWN_CONFIG_KEYS: read by link-extraction, reconcile-links, and sweep, and documented as the off-switch in brain-ops/maintain, but the allowlist rejected `config set auto_link false`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: repo-wide accuracy + MECE reform from the 9-bucket markdown audit A code-grounded audit of every markdown file (root, architecture, guides, mcp, tutorials, docs-root, operations/eval/designs, skills, recipes) followed by a fix wave with per-bucket ownership. Four classes of change: Accuracy — every documented command/flag verified against src/ before writing: dead commands replaced with working ones (pages purge-deleted, jobs watch --follow, gbrain restore, import-based Obsidian flow, space-separated --scopes, real thin-client recipes, working isolation verification, real supervisor restart procedure, curl-based ngrok health check, real minion_mode setter); count drift fixed with rot-proof phrasing (100+ ops, 50+ bundled skills via skills/manifest.json, 140+ engine methods, KNOBS_HASH_VERSION pointer instead of hardcoded versions); stale claims corrected (search-mode defaults, RETRIEVAL pipeline order incl. autocut, sentinel rules, refusal-list mechanism, engine snapshot, shard cap 2400s + EXIT-HANG classifier in TESTING.md, latest-stable + publish-template documented in RELEASING.md as release.yml promises). MECE — one home per concept, pointers elsewhere: test isolation → TESTING.md; OAuth registration + --bind/--public-url lore → DEPLOY.md; mode bundles → guides/search-modes.md (the home the CLAUDE.md dispatcher always promised); merge contract → schema-packs.md; WAL ladder → ENGINES.md; quiet-hours → quiet-hours.md; capture taxonomy → entity-detection.md; person-page taxonomy → compiled-truth.md; brain-first protocol → brain-first-lookup.md; refresh semantics → refresh-algorithm.md; KEY_FILES.md deduplicated (58 extension entries merged, one entry per file); infra-layer.md rewritten as a pointer page. Privacy — placeholder sweep across guides, docs, skills, and recipes per the iron rule; per-release narration stripped from reference docs (current-state prose only). Bootstrap coverage — AGENTS.md pointer, RESOLVER routing row, INSTALL.md path, tutorial cross-links, keyless-mode sections in spend-controls/headless-install. skills.lock.json regenerated; llms.txt/llms-full.txt rebuilt. Gates: verify 36/36, typecheck clean, doctor 96/96, skills-integrity + resolver + build-llms + config-set + migrations all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): refresh production-brain stats to current brain-repo counts 155,795 pages / 24,589 people / 5,340 companies, counted from the brain repo's current HEAD; the "100K-page brain" framing moves to 150K to match. Cron-fleet count unchanged (its store lives on the deployment host, not in the repos available for verification). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scan,push): modern OpenAI/Voyage key patterns, scan staged blobs not disk - secret-scan matches sk-proj-/sk-svcacct-/sk-None- and pa- Voyage keys (the bare sk- pattern missed every current OpenAI key format). - workspacePush stages first, then scans the staged index blobs via git cat-file, closing the scan-then-stage TOCTOU where a file changed between snapshot and commit shipped unscanned. - shared binary-sniff helper, memoized glob regexes, atomic push-status write, and tests for pull_conflict + gitignored deny-match paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): regenerate flag registry for new commands, harden shard classifier + release token - cli-flag-registry.generated.ts regenerated: bootstrap/hook/sweep and sources push --message/--allow-unverified-remote were missing, so the strict #2185 validator rejected real invocations and skipped the new commands entirely. - EXIT-HANG shard classifier now requires every assigned file to have started before warn-passing a watchdog kill (was fail-open). - release.yml passes TEMPLATE_REPO_PAT via http.extraheader, off the argv. - compiled-binary e2e fails loud in CI instead of a silent permanent skip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hook,sweep,ipc): non-blocking hook pushes, bounded sweep + cache, source-bound resolve - session-start/session-end no longer run synchronous git + inline push inside their self-deadline; a detached child does the push and the hook returns immediately (blocked Claude Code startup for minutes on a dirty tree before). - serve sweep drops the unbounded listAllPageRefs, resolves only candidate targets, claims corpus files atomically (no double-LLM-spend race), and caps the fence LIKE scan; heartbeat writes are O_APPEND with rare compaction. - hot-memory cache evicts expired entries and bounds entry count (the key is caller-controlled via _meta.session_id). - v1 resolve IPC honors boundSourceId like turn_context; turn-context runs its arms concurrently. doctor reads push/heartbeat thresholds from hook.ts. - new tests: doctor bootstrap checks, hook push-gate + deadline, concurrent sweep claims, cache eviction, bound-source resolve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bootstrap): origin-ownership gate, world visibility, collision-safe source id, consent + templates - repo adoption requires an exact receipt repo_url match or authed-owner check (undefined repo_url was a wildcard); create verifies privacy BEFORE the first push. - verify sets facts.default_visibility=world if unset, so agent-authored facts surface in per-turn context (they defaulted private before). - source_id derives a path-hash suffix when 'workspace' is taken by another checkout; every consumer reads manifest.source_id. - skipped HOOKS_CONSENT now declines (was falling through to default yes); --minimal refuses on an initialized manifest; tilde fences escaped. - MCP registration pins --surface full; status hard-fails a public origin (template door); receipt writers guard against newer/corrupt receipts; uninstall only claims brain-deleted after a real rm. - templates ship jobs disabled + provider-consent + support-relay lines; soul-audit re-runs over the shared interview bank. TODOS: 11 follow-ups. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): close adversarial-review findings — scan fails closed, whole-PEM redaction, bound repo push Cross-model adversarial pass (Claude + Codex) on the bootstrap wave: - secret scan fails CLOSED: an unreadable, oversized, or binary staged blob now blocks the push (blocked_unscannable, exit 5) instead of committing unscanned; only a confirmed staged deletion is skipped. This was the headline "block secrets before they leave the machine" property failing open. - private-key redaction spans the whole PEM block (header+body+footer), not just the header line — the base64 body no longer survives into the corpus the sweep sends to an extraction provider. - bootstrap repo commits the workspace (secret-scan-gated) before the first push and verifies the remote actually received it, so a push-fail retry can't adopt an empty remote as success. - privacy verify is re-bound to origin immediately before push (a concurrent origin rewrite between verify and push is refused). - session-end corpus write is atomic and clears the stale ingested/in-progress sidecars so a resumed session's appended transcript is re-ingested. - public-origin refusal enforced at render (not only status); MCP "already registered" is verified to target this workspace, not blessed blindly; verify probe cleanup scopes deletes to its own slugs, not a token substring; allowlist fingerprint floor raised 8→16 hex. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code Turns a Codex or Claude Code session into a persistent personal agent: interview-rendered identity files, a local PGLite brain, per-turn context via serve IPC (Claude Code hooks / Codex pull protocol), session-triggered persistence, and a private GitHub repo as the agent's portable body. Keyless- first (the harness model is the LLM; one optional key adds embeddings + extraction). New `gbrain bootstrap` command family + `gbrain hook` + `gbrain sweep`; doctor bootstrap health checks; latest-stable distribution ref + template-repo publish job. Opt-in, additive — existing installs untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: regenerate flag registry for security-fix flags; drop fabricated gbrain capabilities doc ref CI caught two real failures under the merged state: - the flag registry lagged the blocked_unscannable/exit-5 flags the security round added, tripping the #2185 freshness guard. - headless-install.md described the keyless capability report as a `gbrain capabilities` command, which the #3502 doc-command resolver rejects — reworded to prose (the real surface is bootstrap verify's report). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: sync KEY_FILES + bootstrap plan to security-fix behavior Cross-referenced the security-fix round against the reference docs and corrected the drift those commits introduced: - workspace-push.ts entry: stage-FIRST-then-scan order (the TOCTOU fix), fail-closed blocked_unscannable, and the sources-push status -> exit-code map. - hooks.ts entry: MCP registration pins `serve --surface full`. - hook.ts entry: session-start/session-end pushes run in a detached child (non-blocking); atomic corpus write clears stale sidecars. - bootstrap.ts entry: render hard-refuses a public origin (template door). - verify.ts entry: source_id collision resolution (workspace-<path-hash>). - AGENT_BOOTSTRAP_PLAN as-shipped delta note for the scan/stage reorder. llms bundle unchanged (KEY_FILES is link-only); build:llms and test/build-llms.test.ts green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): silence SC2016 on the intentional askpass literal in release.yml The one-shot GIT_ASKPASS script must contain literal $1 and $TEMPLATE_REPO_PAT so they expand when /bin/sh runs it at git's credential prompt, not when the outer shell writes the file — single quotes are correct. Add a scoped shellcheck disable so actionlint passes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills): declare 'bootstrap my data' trigger in cold-start frontmatter The doc reform added 'bootstrap my data' to cold-start's RESOLVER.md row (to disambiguate data-bootstrap from agent-bootstrap) but not to the skill's own frontmatter triggers, tripping the RESOLVER↔frontmatter round-trip contract (resolver.test.ts). Declare it; regenerate skills.lock.json. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): default per-turn hooks + search mode ON without a prompt Installing gbrain for your coding agent IS the consent for the behaviors that make it work, so stop re-litigating them with install-time questions whose "no" defeats the product: - Per-turn hooks (Claude Code) install ON by default — no prompt. Off-ramps: `--no-hooks` at install, `GBRAIN_HOOKS=0` at runtime, `bootstrap uninstall`. The "hooks installed" line now surfaces the kill switch so default-on is never silent. A persisted HOOKS_CONSENT=no (interview --skip) still declines. - Search mode defaults to `balanced` silently (nobody knows the modes at install; `gbrain search modes` changes it any time). - MCP scope stays the ONE deliberate prompt — project vs user is a real cross-repo privacy choice, not friction. Marks the two consents `silent: true` in the question bank (new QuestionSpec field), rewrites the runbook phases so the agent no longer asks them, adds the `--no-hooks` flag (+ registry regen), and adds default-on / opt-out tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): real end-to-end coverage — cross-session recall, per-turn content, Codex door, realistic corpus Closes the seven e2e gaps a coverage audit surfaced: the plumbing was well-unit-tested but the product claims ("Codex works, context shows up every turn with real content, it remembers across restarts, machine two recovers, Postgres works") were unproven end to end. Test-only wave — zero src changes. - Hermetic synthetic corpus (test/fixtures/bootstrap-corpus/ + a loader helper): 12 interlinked pages (52 edges, timelines), 12 world/private beliefs, 8 gold queries — curated from the gbrain-evals synthetic corpora, 100% placeholder names, so recall is asserted on a real multi-entity brain instead of a 2-node self-planted probe. - GAP1 magic moment: author a fact via the real write path, disconnect the engine, reopen against the same DB, recall it — a real session boundary, not verify.ts's same-connection SQL read-back. Plus a source-isolation assertion. - GAP2 per-turn content: hook-under-serve Pin 1 now seeds a known fact and asserts its text lands in the injected block AND private beliefs never do (was: empty brain, empty_block accepted as a pass). - GAP3 Codex door: assert the rendered AGENTS.md carries the Gate-3 brain-first pull protocol; make the fake codex shim implement `mcp get` so the [FIX7] target-verification can actually fail; the Docker cold-machine harness now exercises the hooks/MCP registration step instead of skipping it. - GAP4 corpus recall: turn-context + verify graph-floor/qrels run on the real multi-entity brain with real edges. - GAP5 attach: machine-two now re-ingests the cloned brain/ into a fresh DB and recalls a fact authored only on machine one — the multi-device payoff. - GAP6 keyed + Postgres (env-gated): real embeddings prove semantic recall a paraphrase query can reach but keyless BM25 cannot; bootstrap verify drives a real Postgres engine (skipIf DATABASE_URL/keys absent). - GAP7 persistence: session-end runs the REAL push (not the mocked seam) to a local bare remote and the remote receives the content; a planted secret is blocked at the gate; the 15-min cron installs and fires a scan-gated push. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): real-agent e2e — drive the actual claude + codex binaries end to end Closes the audit's biggest gap ("no real harness ever drives a turn"). Adapts gstack's PTY/headless agent harness to prove the bootstrap install + smoke work against the REAL binaries, not PATH shims. Test/CI/docs only — zero src changes. - test/helpers/agent-harness.ts: hermetic clean-room child env (ported from gstack; drops CONDUCTOR_/CLAUDE_/GSTACK_/MCP_/GBRAIN_, promotes GSTACK_ANTHROPIC_API_KEY→ANTHROPIC_API_KEY), real-binary resolvers + auth probes, headless `claude -p --output-format stream-json` and `codex exec --json` turn runners, a gbrain stdio MCP-config writer, and a keyless brain seeder. + a fixture-parse unit test (no binary needed). - test/e2e/bootstrap-real-claude.serial.test.ts: real `gbrain bootstrap` install → REAL `claude mcp add` (verified via `claude mcp get`) → verify exit 0 → a real `claude -p --mcp-config --strict-mcp-config` turn that invokes mcp__gbrain__search and answers from the brain (proven: toolCalls include mcp__gbrain__search, final text carries the seeded fact). - test/e2e/bootstrap-real-codex.serial.test.ts: same install with REAL `codex mcp add` into a real ~/.codex/config.toml + Gate-3 pull-protocol assertion, then a real `codex exec --json` turn surfacing the fact (MCP or the pull- protocol shell path). Bounded retry absorbs codex's occasional MCP-call cancellation without softening the fact-requiring assertion. - Everything hermetic (temp HOME/CLAUDE_CONFIG_DIR/CODEX_HOME/GBRAIN_HOME; real ~/.codex auth copied read-only) and skipIf-gated so it self-skips cleanly where the binaries/auth are absent. - heavy-tests.yml: gated `real-agent-e2e` job (nightly/label, never the PR shard; no-op on a runner without authed binaries). - TODOS: compiled `gbrain` binary can't serve a PGLite brain (bun compile omits the WASM/extension payloads); harness falls back to `bun run` serve. Verified against live claude 4.6 + codex 0.147.0: 15 pass / 0 fail; verify 36/36; typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): real-agent-e2e job — bash array + --timeout (actionlint SC2086 + bun-test-timeout guard) The real-agent-e2e job's file loop used an unquoted $FILES (SC2086) and ran `bun test` without --timeout (check-bun-test-timeout guard). Switch to a bash array and add --timeout=600000 (real-agent turns are slow; the tests self-skip without authed binaries so it's a no-op elsewhere). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pglite): embed WASM + extension assets so the compiled binary can serve A `bun build --compile` gbrain binary could not `serve` a PGLite brain: the compile bundles JS but not PGLite's runtime payload (pglite.wasm, initdb.wasm, pglite.data, vector/pg_trgm tarballs), so `serve` on PGLite died with a bunfs/ENOENT. Now the assets ride inside the binary. - src/core/pglite-embedded-assets.ts: embeds the five assets via `import … with { type: 'file' }` (the ENG-6 idiom) and exposes getEmbeddedPgliteOptions() → { pgliteWasmModule, initdbWasmModule, fsBundle, extensions:{vector,pg_trgm} }. WASM/fsBundle are consumed as bytes; the two extension tarballs are materialized to a content-addressed temp file (atomic, size-verified reuse) because PGLite reads them via fs.createReadStream, which cannot read a /$bunfs path. Unconditional (works in bun-run and compiled), so no fragile mode branch. - src/core/pglite-engine.ts: static-import getEmbeddedPgliteOptions (engine path stays static per the engine-dynamic-import invariant); spread into both PGlite.create sites (initial + WAL-repair retry). The bunfs classifier stays as a backstop but no longer fires for a correct binary. - scripts/check-pglite-embedded.sh (+ smoketest): compiles a focused binary and asserts it boots PGLite, CREATE EXTENSION vector/pg_trgm, and round-trips a page — wired into `bun run verify` (now 37 checks), check:all, and check:pglite-embedded. Fail-soft only when compile is unavailable. - agent-harness.ts probeCompiledPglite now passes → the real-agent e2e uses the fast compiled MCP server. TODOS: the P2 "can't serve PGLite" item is closed. Verified: fresh compiled binary ran `search`/`query` against a PGLite brain and returned the seeded row (no bunfs/ENOENT); verify 37/37; pglite-engine 120/0 source-mode; typecheck clean; engine-dynamic-import + parity guards pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
54 KiB
Testing (gbrain repo)
On-demand reference (see CLAUDE.md Reference map). Current behavior + invariants only.
test/e2e/serve-http-oauth.test.ts additionally pins confidential POST/Basic revocation, public-client SDK fallthrough, malformed/mixed authentication rejection, cross-client isolation, unknown-token opacity, metadata auth methods, no-store responses, strict post-revoke 401, and retryable backend 503 semantics.
Test command tiers
Seven test command tiers, each with a clear scope:
| Command | What it runs | Wallclock | When to use |
|---|---|---|---|
bun run test |
Parallel unit-test fast loop. Sharded fan-out via scripts/run-unit-parallel.sh (default 4 shards — CPU-detected, clamped to a max of 8; 4 matches CI's fan-out and avoids PGLite WASM-init contention), then a serial pass over *.serial.test.ts. Excludes *.slow.test.ts and test/e2e/*. No pre-checks, no typecheck. Memory-safe by default: total concurrency (shards × intra-shard files) is capped to available memory at GBRAIN_TEST_MEM_PER_FILE_MB (default 1536 — a PGLite WASM instance) per concurrent file, and two phantom-failure classes are automatically re-run serially (the rescue pass): failures carrying the WASM out-of-memory signature, and shards killed externally (SIGTERM/SIGKILL well before the shard timeout — sibling workspaces' process cleanup, memory jetsam). Phantoms pass serially and the run goes green with an oom_rescued note; real failures fail again serially and stay red. Knobs: GBRAIN_TEST_NO_MEM_ADAPT=1, GBRAIN_TEST_NO_OOM_FALLBACK=1, GBRAIN_TEST_MAX_CONCURRENCY (intra-shard, default 4), GBRAIN_TEST_SHARD_TIMEOUT / GBRAIN_TEST_SHARD_KILL_AFTER, plus --shards N / --max-concurrency N / --dry-run script args. |
a few minutes on a Mac dev box | Inner edit loop. Default. |
bun run verify |
CI's authoritative pre-test gate set, fanned out in parallel by scripts/run-verify-parallel.sh: the full check:* battery (privacy, jsonb, progress, source-id, test-isolation, wasm, …) plus bun run typecheck. The CHECKS array in that script is the single source of truth — CI literally calls bun run verify in a dedicated job. |
~16s (parallel; typecheck dominates) | Before pushing; before /ship. |
bun run test:full |
verify && bun run test && bun run test:slow && [smart e2e]. The local equivalent of "everything CI runs." Smart e2e: runs e2e only when DATABASE_URL is set; else loud skip notice to stderr. |
~3-5min depending on slow + e2e | Pre-merge sanity, before opening a PR. |
bun run test:slow |
Just the *.slow.test.ts set (intentional cold-path correctness checks). |
seconds-to-minutes | When touching slow-path code. |
bun run test:serial |
Just the *.serial.test.ts set (cross-file-contention quarantine; one bun process per file for true module-registry isolation). |
~1s per quarantined file | Debugging a specific quarantined file. |
bun run test:e2e |
Real Postgres E2E. Requires Docker + DATABASE_URL. Sequential. |
~5-10min | Pre-ship; nightly. |
bun run check:all |
The historical pre-check scripts (chained sequentially in package.json). Overlaps verify heavily but is NOT a superset — verify's CHECKS array in scripts/run-verify-parallel.sh is the authoritative gate; check:all keeps a few local-only extras (trailing-newline, exports-count, no-legacy-getconnection). |
~10s | Local-only sweep for the extras. |
Shell dispatch and Windows
All four of test, verify, ci:local and test:e2e hand off to shell scripts
under scripts/, so every check:* entry in package.json invokes its script as
bash scripts/<name>.sh instead of relying on the shebang — bun on Windows cannot
exec a .sh directly. Add a new shell-script check with that same prefix. The
scripts/*.ts entries run under bun and take no prefix.
The scripts must also be on disk with Unix line endings. A strict bash (WSL, Linux
CI, macOS) rejects CRLF and dies on the script's first meaningful line; the Cygwin
bash that ships with Git for Windows tolerates it, so a green local run is not by
itself evidence that a script is CRLF-clean.
The root .gitattributes pins *.sh text eol=lf, which overrides the
core.autocrlf=true default that Git for Windows installs. It pins *.md the
same way, because the frontmatter readers anchor on a --- fence followed by a
Unix line ending and a CRLF checkout makes a document parse as having no
frontmatter, silently. Working copies cloned
before those pins need a one-time git rm --cached -r . -q && git reset --hard to
pick them up; see the Windows section of CONTRIBUTING.md.
Wallclock figures in the table above are from a Mac dev box. Windows is
substantially slower because each check pays full process-creation cost, and three
tree-walking checks (check:privacy, check:test-names, check:test-isolation)
plus typecheck can exceed the 120s per-check cap in run-verify-parallel.sh
there even though they pass on Linux and macOS.
CI vs local: intentionally divergent file sets
- CI matrix (
.github/workflows/test.yml) runsscripts/test-shard.shacross 10 matrix shards partitioned by weight-aware LPT bin-packing (scripts/sharding.ts) and INCLUDES*.slow.test.ts(the two outlier slow files run as dedicated jobs alongside the matrix). CI EXCLUDES*.serial.test.tsfrom the shards and runs them in a dedicated job viabun run test:serial, one bun process per file — keeping serial files out of the shard processes is what preserves themock.modulequarantine (a top-level mock in one file leaks into every other file sharing its process).bun run verifygets its own job too, as does the BrainBench memory-conformance gate (brainbenchjob →scripts/ci-brainbench-gate.sh, hermetic in-memory PGLite, ~15s), which compares HEAD's fresh run against master's committed baseline (evals/brainbench/baselines/main.json) — thetest-statusaggregate checks its result explicitly. CI is the ground truth for "did everything pass." - Local fast loop (
scripts/run-unit-shard.shvia the parallel wrapper) uses round-robin-by-index sharding and EXCLUDES*.slow.test.tsAND*.serial.test.ts. Local trades coverage for inner-loop speed; CI catches what local skips.
This divergence is intentional. Don't try to make them equal — the two scripts deliberately solve different problems. The regression test at test/scripts/run-unit-shard.test.ts pins what the local fast loop should and shouldn't include; test/scripts/run-unit-parallel.test.ts pins the wrapper's memory-adaptive concurrency and the OOM/external-kill serial rescue pass.
Failure-first logging
When bun run test finds any failure, the wrapper:
- Writes failure blocks (each prefixed with
--- shard N: <test name> ---) to.context/test-failures.log(workspace-local, gitignored). On systems without a writable.context/, falls back to/tmp/gbrain-test-failures.log. - Prints a loud stderr banner with the absolute log path, plus the last 30 lines of the failure log inlined. Banner survives
| head/| tail/ agent-side log truncation. - Writes a one-line-per-shard summary to
.context/test-summary.txt(shard N/M: pass=X fail=Y skip=Z rc=W). - Exits non-zero. Empty failure log + non-zero exit = infrastructure problem (wedged shard, killed child); the banner says so.
If a shard hits the per-shard GBRAIN_TEST_SHARD_TIMEOUT cap (default 3000s — sized so the heaviest count-balanced shard finishes under 4-way contention; GBRAIN_TEST_SHARD_KILL_AFTER sets the grace after TERM before KILL, default 30s), the wrapper classifies the kill one of two ways:
- EXIT-HANG → warn-pass. If the shard's log had been silent for ≥300s at kill time AND shows zero
(fail)markers, the shard finished all its work, leaked a handle, and never exited (a pre-existing, master-reproducible PGLite-adjacent leak — see TODOS.md "unit-shard exit hang"). The wrapper prints a⚠️ shard N/M: EXIT-HANG ... Treating as pass-with-warningbanner, writesEXIT-HANG (idle Ns, 0 fails) ... warn-passto the summary, and does NOT fail the run. Its pass counts are undercounted (bun never printed its final summary). Bun's per-test--timeoutturns a genuinely hung TEST into a printed(fail)— new output — so this classification cannot mask a hung test; the residual maskable case is a file-level import hang in the very last file, which the banner keeps visible. - WEDGED → hard failure. Anything else (failures present, or the log was still growing) writes
--- shard N: WEDGED after ${SHARD_TIMEOUT}s ---to the failure log with the last 50 lines of the shard log, marks the run failed, and proceeds with other shards' results.
Triage rule: a warn-pass EXIT-HANG line in .context/test-summary.txt is NOT a test failure — don't burn time bisecting it; a WEDGED line is.
File taxonomy
*.test.ts→ fast loop (parallel up-to-4-shard fan-out, memory-adaptive).*.slow.test.ts→ run viabun run test:slowonly (intentional cold-path tests; would dominate the fast loop's wallclock).*.serial.test.ts→ run viabun run test:serialafter the parallel pass completes; one bun process per file (--max-concurrency=1within a shared process is not enough — the module registry still leaksmock.module). Quarantine for tests that share file-wide state and race when run alongside other files in the samebun testprocess. Several dozen files, discovered by the*.serial.test.tsglob — no list to maintain. Typical residents:mock.module(...)users (top-level mocks leak across files in a shard process, e.g.test/embed.serial.test.ts), env-coupled files (e.g.test/brain-registry.serial.test.ts), and process-lifecycle suites that assert onprocess.exitCode(e.g.test/pglite-engine-disconnect.serial.test.ts). Do not put the parallelism back on a serial file unless you've fixed the contention root cause (it just re-introduces the flake).test/e2e/*.test.ts→ real-Postgres E2E. Skipped whenDATABASE_URLis unset.tests/heavy/*.sh→ ops-shape shell scripts. Cost minutes per run; NOT in defaultbun test. Run viabun run test:heavyor scheduled nightly via.github/workflows/heavy-tests.yml. Examples: pg_upgrade matrix (boot legacy brain → walk to head), RSS budget gate (measure peak worker RSS vs committed baseline), read-latency-under-sync (p50/p95/p99 under concurrent writer load), sync lock regression (N concurrent syncs assert 1 winner + N-1 lock-busy + zero leakedgbrain_cycle_locksrows). Seetests/heavy/README.mdfor when to add a script here vs*.slow.test.ts. Files prefixed with_(e.g.tests/heavy/_build_legacy_fixtures.sh) are helpers/libs invoked by sibling tests — the runner skips them.test/fuzz/*.test.ts→ property-based fuzz harness. Pure-validator targets inpure-validators.test.tsare guarded byscripts/check-fuzz-purity.sh(inbun run verify), whichbun build --target=bunbundles each target and greps the resulting bundle for banned transitive imports (node:fs,node:child_process, engine modules). Anything that fails the guard moves tomixed-validators.test.ts(still property-tested, but no purity guarantee) orfilesystem-validators.test.ts(fs-backed, uses temp dirs). Fuzz tests run in the defaultbun testloop because they're fast (~3s for ~12 properties × 1000 runs each).
Skills-manifest freshness guard
skills/skills.lock.json is a committed sha256 inventory of every bundled file under
skills/ (tamper evidence, not signatures — see src/core/skills-integrity.ts).
Any change under skills/ must regenerate it: bun run scripts/generate-skills-manifest.ts.
scripts/check-skills-manifest-fresh.sh (bun run check:skills-manifest, wired into
bun run verify) regenerates to a tmp file and diffs, failing CI on drift; at runtime
gbrain doctor reports the same drift as a warn-only skills_manifest_integrity check.
Test-isolation lint and helpers
This section is the canonical home of the test-isolation discipline — CONTRIBUTING.md and other docs link here rather than restating the rules.
The cross-file flake class is enforced statically by scripts/check-test-isolation.sh, wired into bun run verify and bun run check:all. Rules (non-serial unit files only; *.serial.test.ts and test/e2e/* are skipped):
| Rule | What it bans | Fix |
|---|---|---|
| R1 | process.env.X = ..., bracket assignment, delete process.env.X, Object.assign(process.env, ...), Reflect.set(process.env, ...) |
Use withEnv() from test/helpers/with-env.ts, OR rename file to *.serial.test.ts |
| R2 | mock.module(...) anywhere in the file |
Rename file to *.serial.test.ts (no DI on production code for testability) |
| R3 | new PGLiteEngine( outside ~50 lines after a beforeAll( line |
Use the canonical block (below) inside beforeAll( |
| R4 | Files creating new PGLiteEngine( without engine.disconnect( inside an afterAll( block |
Add afterAll(() => engine.disconnect()) |
Files that violated these rules at the isolation-lint baseline are listed in scripts/check-test-isolation.allowlist. The allow-list MUST shrink over time — never add new entries.
Canonical PGLite block (R3 + R4 compliant)
Every test file that needs a PGLite engine should use this exact pattern:
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await resetPgliteState(engine);
});
Why this exact shape: beforeAll creates a single engine per file (PGLite WASM cold-start + initSchema is ~20s); beforeEach truncates user data via resetPgliteState ("two orders of magnitude faster" than fresh-engine-per-test); afterAll disconnects so the engine doesn't leak across file boundaries within a shard process.
withEnv pattern (R1 fix)
import { withEnv } from './helpers/with-env.ts';
test('reads OPENAI_API_KEY', async () => {
await withEnv({ OPENAI_API_KEY: 'sk-test' }, async () => {
expect(loadConfig().openai_key).toBe('sk-test');
});
});
// Delete a var (override is undefined):
await withEnv({ GBRAIN_HOME: undefined }, fn);
// Multiple keys:
await withEnv({ A: '1', B: '2', C: undefined }, fn);
withEnv saves the prior value of every key it touches and restores via try/finally — including when the callback throws. It is cross-test safe but NOT intra-file concurrent-safe. process.env is process-global; two test.concurrent() calls in the same file both touching the same key will race. Files using withEnv stay outside the test.concurrent() codemod's eligibility filter.
When to quarantine instead of fix
Rename to *.serial.test.ts when:
- The file uses
mock.module(...)(R2 — there's no clean fix without changing production code). - The file is genuinely env-coupled (e.g.
gbrain-home-isolation.test.ts,claw-test-cli.test.ts) — module-load env readers + ESM caching defeat dynamic-import-after-env tricks. - The file's tests intentionally share state across
it()boundaries.
The quarantine has grown to dozens of files — treat it as debt: every addition needs a reason from the list above, and prefer fixing the contention root cause when one exists.
Unit test inventory
bun test runs all tests without a database. E2E tests skip gracefully when DATABASE_URL is not set.
Unit tests and what they cover:
test/markdown.test.ts— frontmatter parsing;splitBodysentinel precedence, horizontal-rule preservation,inferTypewiki subtypes.test/chunkers/recursive.test.ts— chunking.test/parity.test.ts— operations contract parity.test/cli.test.ts— CLI structure.test/cli-finish-teardown.test.ts— the #2084 teardown contract:computeTeardownDeadlineMsformula/floor/live-registry scaling +GBRAIN_TEARDOWN_DEADLINE_MSoverride (garbage/zero/negative values fall back to the formula);finishCliTeardownclean path (drain BEFORE disconnect, no exit, no warn), backstop on hung drain or disconnect (honors an errored op's exit code), throwing drain/disconnect warned + swallowed; the gbrain-owned verdict channel is immune to PGLite WASMprocess.exitCodewrites;flushThenExitunit coverage with mocked streams (exits once after both stream callbacks, non-TTY aliveness grace, blocked-pipe guard, EPIPE-safe,GBRAIN_FLUSH_GRACE_MSoverride).test/flush-then-exit-harness.test.ts— real spawned-Bun pipe semantics forflushThenExit(fixture:test/fixtures/flush-then-exit-harness.ts): a 4MB piped stdout payload arrives byte-complete with the exit code even with a late reader, small output survives exit with a concurrent reader, and the fence resolves promptly (wall time well under the guard + grace ceiling).test/cli-should-force-exit.test.ts—shouldForceExitAfterMaindaemon-survival gate:serve(stdio and--http) never force-exits, including with preceding global flags; op commands / empty / flag-only argv do; the #2084 case that space-separated global-flag VALUES can't fake a command (--timeout 30s serveresolves to theservedaemon, not a30scommand).test/cli-exit-verdict-pin.test.ts— #2084 structural class pin: grepssrc/so the NEXT rawprocess.exitCode =write fails CI (a raw write bypasses the gbrain-owned verdict channel and gets silently zeroed by the deliberate flush-exit — the bug that made doctor's FAIL path exit 0). Runtime variants live intest/cli-finish-teardown.test.ts; this is the review-time guard.test/cli-pipe-truncation.test.ts— real-CLI pipe completeness (the #1959 incident class), implementation-agnostic: the actual CLI run the way agents run it (piped stdout) produces complete, parseable, byte-stable--tools-jsonoutput and exits deliberately, well under the teardown backstop. Synthetic flush-mechanism coverage stays intest/flush-then-exit-harness.test.ts.test/volunteer-context.test.ts— push-based context core (#2095), hermetic in-memory PGLite:parseWindowlenientuser:/assistant:parsing, multi-turn window extraction, confidence-gated volunteering (arm confidences, multi-turn/newest-turn boosts,min_confidencegate, max-pages cap), slug-only suppression, privacy (rationales are deterministic templates; synopses pass the takes/facts fence), and the approximate usage-stats join.test/watch-command.test.ts—gbrain watchpush transport (#2095): streaming loop, rolling window, session dedupe,--jsonJSONL shape,channel: 'watch'event logging, clean EOF return. Hermetic PGLite + injected line/write deps (no subprocess, no real stdin).test/watch-sigint.serial.test.ts—gbrain watchSIGINT lifecycle against a real spawned CLI subprocess with a tmpdir brain. SERIAL: parallel unit shards flake on concurrent subprocess spawns (same rationale asapply-migrations-pglite-spawn.serial.test.ts).test/cli-format-volunteer.test.ts—formatResult'svolunteer_contexthuman rendering: pointer lines with confidence/arm/rationale, the empty-result message, the approximate stats summary.test/config.test.ts— config redaction.test/files.test.ts— MIME/hash.test/import-file.test.ts— import pipeline.test/upgrade.test.ts— schema migrations.test/file-migration.test.ts— file migration.test/file-resolver.test.ts— file resolution.test/import-resume.test.ts— import checkpoints.test/migrate.test.ts— migration: v8/v9 helper-btree-index SQL structural assertions; 1000-row wall-clock fixtures guarding the O(n²)→O(n log n) fix; v12/v13 SQL shape;sqlFor+transaction:falserunner semantics; themax_stalled DEFAULT 1regression guard; v24sqlFor.pglite: ''no-op assertion; v117context_volunteer_events(named + idempotent entry, documented columns + both source-scoped indexes afterinitSchema, insert + 90-daypurgeStaleVolunteerEventsround-trip).test/bootstrap.test.ts— bootstrap contract: no-op on fresh install, idempotent across twoinitSchema()calls, no-op on modern brain that already has every probed column, full bootstrap path on a simulated legacy brain, fresh-install regression guard, legacylinksshape coverage.test/schema-bootstrap-coverage.test.ts— CI guard.REQUIRED_BOOTSTRAP_COVERAGElists every forward reference inPGLITE_SCHEMA_SQL; the test fails loudly ifapplyForwardReferenceBootstrapskips one (extend both arrays when adding a column-with-index to the embedded schema blob). Also parsessrc/core/migrate.tssource text for everyALTER TABLE ... ADD COLUMN(top-levelsql:,sqlFor.{postgres,pglite}overrides, AND handler-bodyengine.runMigration(N, \ALTER TABLE ...`)) and asserts each (table, column) pair is covered by the bootstrap OR by the schema blob's CREATE TABLE bodies — catching the column-only forward-reference class (e.g.sources.archived,oauth_clients.source_id) that a CREATE INDEX parser alone can't see.parseBaseTableColumns` strips SQL line + block comments before identifying column names so commented-out lines don't hide adjacent columns.test/helpers/schema-diff.ts+test/helpers/schema-diff.test.ts+test/e2e/schema-drift.test.ts— cross-engine schema parity gate. Helper exports puresnapshotSchema(query)/diffSnapshots(pg, pglite, opts)/formatDiffForFailure(diff)/isCleanDiff(diff)over a four-tuple per column (data_type,udt_name,is_nullable,column_default). E2E test spins up fresh PGLite + Postgres, runsengine.initSchema()on each, snapshotsinformation_schema.columns, then diffs. 2-table allowlist (files,file_migration_ledger) — every other Postgres table must reach PGLite viaPGLITE_SCHEMA_SQLor a migration'ssqlFor.pglitebranch. Sentinels foroauth_clients,mcp_request_log,access_tokens,eval_candidatesgive tighter blame messages. Skips withoutDATABASE_URL. Wired intoscripts/e2e-test-map.tsso changes tosrc/schema.sql,src/core/pglite-schema.ts, orsrc/core/migrate.tstrigger it. The failure message names every drift with a paste-ready hint pointing atsrc/core/pglite-schema.ts.test/setup-branching.test.ts— setup flow.test/slug-validation.test.ts— slug validation.test/storage.test.ts— storage backends.test/supabase-admin.test.ts— Supabase admin.test/yaml-lite.test.ts— YAML parsing.test/check-update.test.ts— version check + update CLI.test/pglite-engine.test.ts— PGLite engine, all BrainEngine methods includingaddLinksBatch/addTimelineEntriesBatch(empty batch, missing optionals, within-batch dedup via ON CONFLICT, missing-slug rows dropped by JOIN, half-existing batch, batch of 100) plusconnect()error-wrap assertion (original error nested, #223 link in message, lock released).test/links-timeline-jsonb-poison.test.ts— gbrain#1861 PGLite half (always-on, noDATABASE_URL). Locks thejsonb_to_recordsetbatch-insert path for links/timeline/takes against free-text "poison" payloads (commas, quotes, backslashes, braces, em-dashes) and asserts NUL is stripped from free-text body fields but rejected in identity fields. gbrain#2011 adds lone-UTF-16-surrogate cases: every free-text field (link context; timeline summary/detail/source; take claim/source) well-forms to U+FFFD across batch + scalar write paths, while a surrogate in an identity field (slug) still fail-closed rejects the batch. The Postgres lane (test/e2e/jsonb-batch-poison-postgres.test.ts) is the one that actually reproduced the original crash.test/engine-factory.test.ts— engine factory + dynamic imports.test/integrations.test.ts— recipe parsing, CLI routing, recipe validation.test/publish.test.ts— content stripping, encryption, password generation, HTML output.test/backlinks.test.ts— entity extraction, back-link detection, timeline entry generation.test/lint.test.ts— LLM artifact detection, code fence stripping, frontmatter validation.test/report.test.ts— report format, directory structure.test/skills-conformance.test.ts— skill frontmatter + required sections validation.test/resolver.test.ts— RESOLVER.md coverage, routing validation; round-trip that every quoted RESOLVER.md trigger matches a frontmattertriggers:entry in the target skill, and everyname="<word>"reference in any SKILL.md resolves to a declared op insrc/core/operations.tsor a Minions handler inPROTECTED_JOB_NAMES.test/search.test.ts— RRF normalization, compiled truth boost, cosine similarity, dedup key.test/sql-ranking.test.ts— source-boost helpers: longest-prefix-match in SQL CASE,detail=hightemporal-bypass, three-meta-char LIKE escape (%,_,\), single-quote SQL-literal doubling, env override parsing forGBRAIN_SOURCE_BOOST+GBRAIN_SEARCH_EXCLUDE,resolveBoostMap/resolveHardExcludesmerge semantics.test/dedup.test.ts— source-aware dedup, compiled truth guarantee, layer interactions.test/intent.test.ts— query intent classification: entity/temporal/event/general.test/eval.test.ts— retrieval metrics:precisionAtK,recallAtK,mrr,ndcgAtK,parseQrels.test/brainbench-fixtures.test.ts/test/brainbench-generator.test.ts/test/brainbench-metrics.test.ts/test/brainbench-continuity.test.ts/test/brainbench-writeback.test.ts/test/brainbench-adapters.test.ts/test/brainbench-scoreboard.test.ts— the BrainBench memory-conformance unit suites (src/eval/brainbench/): fixture loader/validator + the sealed-gold seal (agoldkey inside a fixture must reject) and committed-corpus integrity; generator determinism (the committed corpus is exactly whatgen.tsproduces, holdout discipline, category counts); metric formulas over hand-built turn rows (zero should-retrieve turns, empty injections, acceptable-vs-gold asymmetry, micro-averaging); cross-harness continuity (writer's decision persists through the production write-back pipeline, reader recalls on the SAME brain); write-back grading the PRODUCTION conversation→facts pipeline via the injected gold extractor; adapter seam contracts over hermetic PGLite (budget caps, suppression modes); scoreboard + gate governance (baseline determinism, count-aware gating, corpus-bless modes, justification flow, isolation gates-at-zero).test/eval-brainbench-e2e.test.ts— BrainBench CLI end-to-end via subprocess against a small tmp corpus: the literal exit codes (0 pass / 1 regression / 2 error-or-inconclusive — the CI product),--outartifact validity incl._meta.metric_glossary, byte-deterministic--update-baseline, anti-vacuous-pass, and theeval run-allin-process wiring.test/check-resolvable.test.ts— resolver reachability, MECE overlap, gap detection, proximity-based DRY detection,extractDelegationTargetscoverage.test/dry-fix.test.ts— auto-fix: three shape-aware expander pure-function tests; five guards (working-tree-dirty, no-git-backup, inside-code-fence, already-delegated within 40 lines, ambiguous-multi-match, block-is-callout).test/doctor-fix.test.ts—gbrain doctor --fixCLI integration: dry-run preview, apply path, JSON output shape.test/backoff.test.ts— load-aware throttling, concurrency limits, active hours.test/fail-improve.test.ts— deterministic/LLM cascade, JSONL logging, test generation, rotation.test/transcription.test.ts— provider detection, format validation, API key errors.test/enrichment-service.test.ts— entity slugification, extraction, tier escalation.test/data-research.test.ts— recipe validation, MRR/ARR extraction, dedup, tracker parsing, HTML stripping.test/minions.test.ts— Minions job queue: CRUD, state machine, backoff, stall detection, dependencies, worker lifecycle, lock management, claim mechanics, depth/child-cap, timeouts, cascade kill, idempotency,child_doneinbox, attachments, removeOnComplete/Fail,max_stalledclamp/default/plumbing coverage.test/extract.test.ts— link extraction, timeline extraction, frontmatter parsing, directory type inference.test/extract-db.test.ts—gbrain extract --source db: typed link inference, idempotency,--typefilter,--dry-runJSON output.test/extract-fs.test.ts—gbrain extract --source fs: first-run inserts + second-run reports zero, dry-run dedups candidates across files, second-run perf regression guard for the N+1 dedup bug.test/link-extraction.test.ts— canonicalextractEntityRefsboth formats,extractPageLinksdedup,inferLinkTypeheuristics,parseTimelineEntriesdate variants,isAutoLinkEnabledconfig.test/graph-query.test.ts— direction in/out/both, type filter, indented tree output.test/features.test.ts— feature scanning, brain_score calculation, CLI routing, persistence.test/file-upload-security.test.ts— symlink traversal, cwd confinement, slug + filename allowlists, remote vs local trust.test/query-sanitization.test.ts— prompt-injection stripping, output sanitization, structural boundary.test/search-limit.test.ts—clampSearchLimitdefault/cap behavior acrosslist_pagesandget_ingest_log.test/repair-jsonb.test.ts— JSONB repair: TARGETS list, idempotency, engine-awareness.test/migrations-v0_12_2.test.ts— JSONB-repair orchestrator phases: schema → repair → verify → record.test/orphans.test.ts— orphans command: detection, pseudo filtering, text/json/count outputs, MCP op.test/postgres-engine.test.ts—statement_timeoutscoping:sql.begin+SET LOCALshape, source-level grep guardrail against a reintroduced bareSET statement_timeout.test/sync.test.ts— sync logic + regression guard asserting top-levelengine.transactionis not called.test/sync-pull-failed-anchor.serial.test.ts— #3068 regression: a failed internalgit pull(local-path origin vsprotocol.file.allow=never) with zero imports returnspartial/pull_failed(notup_to_date), freezeslast_commit+last_sync_at, recovers after a manual pull; fall-through import of local commits preserved. Serial: pinsGBRAIN_HOMEto a temp dir for the whole file.test/sync-concurrency.test.ts—autoConcurrency()thresholds + PGLite-forces-serial + explicit-override clamping;shouldRunParallel()explicit-bypasses-floor contract;parseWorkers()validation rejecting'0'/'-3'/'foo'/'1.5'/trailing chars.test/sync-parallel.test.ts— PGLite-routed coverage of the bookmark gate under concurrency, head-drift gate, vanished-file failure capture, PGLite-stays-serial, and thegbrain-syncwriter-lock contract.test/sync-all-missing-path.test.ts—sync --all --missing-path <fail|skip>pure helpers:parseMissingPathMode(default fail, explicit values, loud rejection of bad/dangling values, never swallows a following flag) andpartitionMissingPathSources(classification driven only by the injected pathExists predicate — no fs; nulllocal_pathpasses through runnable; order preserved).test/sync-failures.test.ts—classifyErrorCoderegex coverage for all 12 codes against literal production message strings frommarkdown.tsandimport-file.ts;summarizeFailuresByCodesort + pre-classified-honor;recordSyncFailurescode-field persistence;acknowledgeSyncFailuresAcknowledgeResultshape + backfill on legacy entries.test/doctor.test.ts— doctor command; assertions thatjsonb_integrityscans the four JSONB write sites andmarkdown_body_completenessis present.test/utils.test.ts— shared SQL utilities +tryParseEmbeddingnull-return and single-warn semantics.test/build-llms.test.ts—llms.txt/llms-full.txtgenerator: path resolution, idempotence, spec shape, regen-drift guard, content contract, AGENTS.md install-path mirror, size-budget enforcement.test/oauth.test.ts— OAuth 2.1 provider: register, getClient,client_credentialsgrant exchange,authorization_codeflow with PKCE challenge/verifier, refresh token rotation,verifyAccessTokenwith both OAuth + legacyaccess_tokensfallback,revokeToken,sweepExpiredTokens; contract test assertingscope+localOnlyannotations on all operations;coerceTimestampunit cases (null/undefined/string/number/throw-on-NaN); NULL-expires_at-as-expired contract for both refresh + access token paths; cascade-delete contract assertingrevoke-clientpurgesoauth_tokens+oauth_codesvia FK CASCADE; cross-client isolation (wrong-client attempt MUST reject AND rightful owner MUST still succeed atomically afterward); empty-stringredirect_uribypass guard; PKCE DCR public-client gate (token_endpoint_auth_method: "none"returns noclient_secret, defaultclient_secret_postclients get the one-time-reveal secret,getClientNULL→undefined normalization, full PKCE/authorize→/tokenround-trip against a public client).test/mcp-dispatch-summarize.test.ts—summarizeMcpParamsinvariants: declared-keys allow-list intersection, attacker-key-name leak guard (unknown keys counted not named), 1KB byte bucketing for size-probe defense, missing op falls through to fully-redacted shape, declared-keys sorted for deterministic output.test/trust-boundary-contract.test.ts— fail-closed trust semantics under cast bypass:ctx.remote === undefinedtreated as remote/untrusted at every flipped call site;as anyandPartial<>spreads can't downgrade trust by accident.test/check-resolvable-cli.test.ts— CLI wrapper: exit codes, JSON envelope shape, AGENTS.md fallback chain.test/regression-v0_16_4.test.ts—findRepoRootregression guard, hermetic startDir parameterization.test/repo-root.test.ts—findRepoRootwalk semantics + default-arg parity; the 4-tierautoDetectSkillsDirfallback chain ($OPENCLAW_WORKSPACE→~/.openclaw/workspace→ repo-root →./skills); RESOLVER.md/AGENTS.md filename precedence; explicit-env-wins-over-repo-root; tier-0$GBRAIN_SKILLS_DIRvalid/invalid/precedence-over-OPENCLAW_WORKSPACE; the install-path walk inautoDetectSkillsDirReadOnly; no-drift on primary success;AUTO_DETECT_HINT+AUTO_DETECT_HINT_READ_ONLYcontent; regression guard asserting the sharedautoDetectSkillsDirMUST NEVER return'install_path'source (how the read-path/write-path split stays safe).test/resolver-merge.test.ts— multi-file resolver merge:findAllResolverFilesempty / RESOLVER.md-only / AGENTS.md-only / both-present (RESOLVER.md first);checkResolvablemerge semantics acrossskills/RESOLVER.md+../AGENTS.mdfor the OpenClaw layout where the skillpack ships a thin RESOLVER.md and the real dispatcher lives at the workspace root; dedup byskillPath(first occurrence wins); AGENTS.md-at-workspace-root works alone.test/filing-audit.test.ts— filing audit:writes_pages/writes_tofrontmatter, filing-rules JSON validation.test/skill-brain-first.test.ts— shared frontmatter parser;analyzeSkillBrainFirstcompliance ladder across 9 fixtures undertest/fixtures/brain-first-skills/(compliant-callout, compliant-phase, compliant-position, exempt-frontmatter, missing-brain-first, multi-pattern, negation-prose, no-external, typo-frontmatter); offset helpers; external-lookup regex shape; audit snapshot+diff transition logic;FORMERLY_HARDCODED_EXEMPTregression absorption.test/routing-eval.test.ts— fixture parsing, structural routing,ambiguous_with, Haiku tie-break layer.test/skill-manifest.test.ts— skill manifest parser: drift detection, managed-block markers.test/skillify-scaffold.test.ts—gbrain skillify scaffoldstubs: SKILL.md, script, tests, routing-eval fixtures.test/skillpack-install.test.ts—gbrain skillpack installmanaged-block install / update / no-clobber semantics.test/skillpack-sync-guard.test.ts— sync-guard: bundled skills stay byte-identical toskills/source.test/http-transport.test.ts— HTTP transport: bearer auth + missing/no-Bearer/unknown/revoked +/healthbypass; dispatch.ts round-trip; invalid_params; application/json response shape (not SSE); CORS default-deny + allowlist; body cap on Content-Length AND chunked; two-bucket rate limit (refill, exhaust+Retry-After, LRU eviction, TTL prune, pre-auth IP fires before DB);mcp_request_logaudit on success + auth_failed.test/restart-sweep.test.ts—recipes/restart-sweep.mdinlined script: sentinel-anchored fenced-block extraction with salted tmp filenames to bypass ESM cache; constructor-time env reads (proves no module-load snapshot); idempotency layer load/save/atomic-tmp-rename/corrupt-JSON-recovery/30-day-prune;(sessionKey, lastAlertedAt)cooldown gate with 6h threshold; AGGRESSIVE-gate two-state tests; execFile argv shape proving shell metachars inOPENCLAW_TELEGRAM_GROUPcannot reach/bin/sh; real-\n-not-literal alert formatting;GBRAIN_HOMEstate path override.test/eval-longmemeval.test.ts— LongMemEval harness, hermetic with noDATABASE_URLand no API keys: PGLite create + reset over runtime-enumeratedpg_tables, infrastructure-table preservation across resets, JSONL question parsing, retrieval-only and answer-gen modes via stubbedThinkLLMClient,--limitcutoff,--keyword-onlyvs hybrid, default--expansion=offbehavior, perf gate (p50 < 30ms / p99 < 50ms warm reset+import+search on Apple Silicon),--helpworks without a configured brain, fixture round-trip viatest/fixtures/longmemeval-mini.jsonl.test/longmemeval-sanitize.test.ts— sanitization parity pinning thatINJECTION_PATTERNSfromsrc/core/think/sanitize.tsis the single source of truth (adding a pattern there must cover both<take>framing and<chat_session>framing, no per-surface regex drift).test/openai-compat-multimodal.test.ts— gateway's openai-compatible multimodal path: happy-path single + multi-input embedding, unauthenticated proxy mode, dimension-mismatch guard (throwsAIConfigErrorwith model id + observed + expected pre-storage), default-dim fallback when recipe declaresdefault_dims, HTTP 401 / 400 / malformed-JSON / non-array error paths, regression that the existing Voyage/multimodalembeddingsrecipe still routes through its dedicated path. Hermetic via the__setEmbedTransportForTestsseam.test/serve-stdio-lifecycle.test.ts—MCP_STDIO=1env guard: stdin EOF does NOT trigger shutdown when the env is set, SIGTERM still does (guard scope is correct), unset env preserves the CLI lifecycle. Exercises theServeOptions.mcpStdio?: booleantest seam directly so tests don't mutateprocess.env.
E2E test inventory
E2E tests live in test/e2e/ and run against real Postgres+pgvector (require DATABASE_URL), except where noted as PGLite in-memory (no DATABASE_URL needed).
bun run test:e2eruns Tier 1 (mechanical, all operations, no API keys). Includes dedicated cases for the postgres-engineaddLinksBatch/addTimelineEntriesBatchbind path — postgres-js's JSONB bind (jsonb_to_recordset(($1::jsonb)->'rows')) differs from PGLite's and gets its own coverage.test/e2e/search-quality.test.ts— search quality against PGLite (no API keys, in-memory).test/e2e/graph-quality.test.ts— knowledge graph pipeline (auto-link via put_page, reconciliation, traversePaths) against PGLite in-memory.test/e2e/jsonb-batch-poison-postgres.test.ts— gbrain#1861 regression, the engine that actually crashed. Seeds free-text "poison" context (Zoom URL with?pwd=, commas, quotes, Windows backslash path, braces, em-dash) and asserts the links/timeline/takes batch writers no longer error with "malformed array literal"; also asserts NUL is stripped from free-text bodies (context/summary/detail/claim) and still rejected in identity fields. gbrain#2011 adds the lone-surrogate crash lock: a lone UTF-16 surrogate in free text (the value that abortedextract --stalewith22P02on Supabase) well-forms to U+FFFD across batch + scalar paths (incl. timeline + takesource), while a surrogate in an identity field still rejects the batch.DATABASE_URL-gated.test/e2e/postgres-jsonb.test.ts— round-trips all 5 JSONB write sites (pages.frontmatter,raw_data.data,ingest_log.pages_updated,files.metadata,page_versions.frontmatter) against real Postgres and assertsjsonb_typeof='object'plus->>'key'returns the expected scalar. Guards against the double-encode bug.test/e2e/integrity-batch.test.ts— parity forscanIntegrity's batch-load fast path vs sequential. Cases (dedup, hits, validate, topPages) seed a fixture and assert both paths return identical results. Dedup case uses raw SQL viagetConn().unsafe()to seed a(test-source-2, people/alice)row alongside the default-source row, sinceengine.putPagedoesn't take asource_id. Pins multi-source overcounting; the "multi-source duplicate slugs scan once" case expects both batch + sequential paths to report 2.test/e2e/jsonb-roundtrip.test.ts— companion regression against the 4 doctor-scanned JSONB sites. Assertion-level overlap withpostgres-jsonb.test.tsis intentional defense-in-depth: if doctor's scan surface drifts from the actual write surface, one of these tests catches it.test/e2e/sync.test.ts—--skip-failedfailure-loop test alongside happy-path tests: broken file →performSyncreturnsblocked_by_failureswith grouped breakdown →performSync({skipFailed: true})advances bookmark and returnsAcknowledgeResultwith code summary → second broken file → second cycle. Saves and restores the user's real~/.gbrain/sync-failures.jsonlso the test is hermetic. Asserts bookmark gating, JSONL state, dedup across paths, summary aggregation, and the literal doctor-rendering string format.test/e2e/upgrade.test.ts— check-update against real GitHub API (network required).test/e2e/minions-shell-pglite.test.ts— PGLite--followinline shell-job path (in-memory, noDATABASE_URLrequired) — the path the minion-orchestrator skill documents for dev use.test/e2e/pglite-cli-exit.serial.test.ts— real spawned-CLI exit behavior on PGLite (in-memory, noDATABASE_URL): read commands (search/get/query) exit 0 promptly; CLI_ONLYcaptureexits clean and frees the single-writer lock; the#2084describes pin every swept disconnect site — a failed op exits 1 with the error on stderr, and the dashboard, read-only-timeout, doctor, anddream --dry-runpaths all exit with no force-exit banner.test/e2e/pgbouncer-teardown.test.ts— PgBouncer TRANSACTION-mode teardown (#2084 / the #1972→#2015→#2084 class). Pins the bug CLASS, not timings: a CLI op against a txn-mode pooled URL exits 0 with intact stdout and does NOT ride the 10s hard-deadline backstop (theengine.disconnect() did not returnbanner is the smoking gun — pre-#2084 it printed on 100% of query-shaped ops). Gated byGBRAIN_PGBOUNCER_URL+GBRAIN_PGBOUNCER_DIRECT_URL(NOTDATABASE_URL) — set automatically bybun run ci:local'spgbouncercompose service; skips gracefully elsewhere. Uses a DEDICATEDgbrain_pgbouncerdatabase so it never races thegbrain_testTRUNCATE fixtures.test/e2e/volunteer-context-postgres.test.ts—volunteer_contexton REAL Postgres (#2095; engine parity beyond the hermetic PGLite unit suite): resolution arms through the actual op handler, the fire-and-forget volunteer-event sink landing rows, the stats join, and the RLS pin thatcontext_volunteer_eventshas ROW LEVEL SECURITY enabled (keeps the v35 auto-RLS event trigger honest for migration-created tables).DATABASE_URL-gated.test/e2e/openclaw-reference-compat.test.ts—check-resolvable+skillpack installagainst a minimal AGENTS.md workspace fixture (test/fixtures/openclaw-reference-minimal/), regression guard for the OpenClaw deployment shape.test/e2e/search-swamp.test.ts— reproduces the source-swamp case. Seeds a curatedoriginals/talks/article-outline-fat-codepage against two<fork>/chat/pages stuffed with the same multi-word phrase. Asserts the article wins keyword AND vector ranking, thatdetail=highlets the chat swamp re-surface, and thatsource_idpasses through the two-stage CTE intact. PGLite in-memory.test/e2e/search-exclude.test.ts—test/+archive/pages hidden by default,include_slug_prefixesopts back in, caller-suppliedexclude_slug_prefixesadds to defaults. Both keyword and vector search paths.test/e2e/engine-parity.test.ts— Postgres ↔ PGLite top-result and result-set parity forsearchKeyword+searchVector(Postgres ranks pages then picks best chunk while PGLite returns chunks directly, so the source-boost behavior needs parity coverage). Skips withoutDATABASE_URL.test/e2e/postgres-bootstrap.test.ts— exercisesPostgresEngine.initSchema()directly against a fresh real Postgres database. Asserts the bootstrap path is no-op on fresh installs and that SCHEMA_SQL replays cleanly through the engine path (not via the standalonedb.initSchemafromsrc/core/db.ts).test/e2e/http-transport.test.ts—gbrain serve --httpend-to-end against real Postgres: bearer auth round-trip,last_used_atSQL-level debounce,mcp_request_logrow insertion on success and auth_failed paths,/healthDB-down → 503 (DB-probing health check), and the dispatch round-trip with a real operation. Skips withoutDATABASE_URL.test/e2e/serve-http-oauth.test.ts— real-Postgres E2E againstgbrain serve --httpwith full OAuth 2.1. Spawns a subprocess server, registers a client via the CLI, mintsclient_credentialstokens, exercises the/mcpJSON-RPC pipeline. Real DCR/registerHTTP-level response-shape test (assertstypeof body.client_id_issued_at === 'number'over the wire, RFC 7591 §3.2.1); real CLI subprocess test forrevoke-client(registers → mints token → revokes viaexecSync→ asserts token rejected at/mcp→ asserts re-run exits 1); server fixture flips on--enable-dcrso/registeris reachable. bun execSync env-inheritance contract: bun'sexecSyncdoes NOT inherit env mutations done viaprocess.env.X = ..., only OS-level env from before bun started. helpers.ts loads.env.testingand setsDATABASE_URLviaprocess.envmutation, which is invisible to subprocesses unlessenv: { ...process.env }is passed explicitly — every subprocess call in this file passesenv: { ...process.env }. Reference fix for the same failure mode in sibling sync/cycle/dream/claw-test E2Es.afterAllcleanup is guarded onclientId(won't throw ifbeforeAllfailed before registration); cleanup errors surface to stderr without throwing so real test failures aren't masked. Also covers the trust-boundary fix: an HTTP MCPsubmit_jobforname: "shell"MUST reject with a permission error (request handler setsremote: trueandsubmit_job's protected-name guard fires), and the same guard rejects subagent submission. Skips withoutDATABASE_URL.test/e2e/sync-parallel.test.ts—DATABASE_URL-gated. 60-file Postgres sync at concurrency=4 imports all + no connection leak (probespg_stat_activitybefore/after to confirm worker engines disconnected). 120-file serial-vs-parallel benchmark printsSYNC_PARALLEL_BENCH N files | serial=Xms | parallel(4)=Yms | speedup=Zx. Asserts parallel ≤ serial × 1.5 (CI-noise tolerant; not a strict speedup gate).test/e2e/multi-source-bug-class.test.ts— PGLite in-memory regression suite pinning every multi-source bug site:listAllPageRefsordering by(source_id, slug),getPagewith sourceId picks the right(source, slug)row,extract-takesprocesses both overlappingpeople/alicerows independently,listPagesfilters correctly withPageFilters.sourceId,addLinksBatchwithfrom/to_source_idtargets the right rows,validateSourceIdrejects path traversal, reverse-write disk layout usesbrainDir/.sources/<id>/<slug>.mdfor non-default sources,copyMigrationSourceslands source metadata before overlapping-slug pages. NoDATABASE_URLneeded. Wired intoscripts/e2e-test-map.tsso changes to extract-takes / patterns / synthesize / embed / extract / migrate-engine auto-trigger it.test/e2e/migrate-engine-sources-postgres.test.ts—DATABASE_URL-gated companion forgbrain migrate --to: migrates a PGLite brain carrying two non-default sources with overlapping slugs into real Postgres and assertscopyMigrationSourcescreated everysourcesFK parent (config JSONB intact, not double-encoded) before any page write. Unit-level manifest identity (crash manifest resumes only against the SAME target; legacy engine-only manifests start fresh) istest/migrate-engine-resume.test.ts.test/e2e/facts-fence-reconcile-postgres.test.ts—DATABASE_URL-gated round-trip for the escape-aware fence parser: renders a## Factsfence whose cells carry literal pipes, backslashes (Windows paths), and empty cells viarenderFactsTable, runs the wipe-and-reinsert reconcile (runExtractFacts) on real Postgres, and asserts every cell survives byte-identically with no column shift.test/e2e/source-isolation-pglite.test.ts— PGLite in-memory regression suite pinning the source-isolation seal at two layers. Engine layer:searchKeyword/searchVector/searchKeywordChunks/listPages/getPage/traverseGraph/traversePathsapplysourceId(scalar fast path) andsourceIds(array path) correctly across both engines. Op-handler layer: routes throughsourceScopeOpts(ctx)so aread+write-scoped OAuth client bound to--source dept-xcannot see rows from neighboring sources viasearch,query,list_pages,get_page, orfind_experts. Covers bothctx.sourceId(single-source clients) andctx.auth.allowedSources(federated_read clients) precedence; federated array wins over scalar wins over nothing. NoDATABASE_URLneeded.test/e2e/think-source-isolation-pglite.test.ts— PGLite in-memory suite pinning thethinkgather stage's source scope: seeds three sources with cross-source links and embedded takes, then assertsrunGatherunder a federatedsourceIdsgrant (and under a scalarsourceId) keeps every stream — hybrid retrieval, takes keyword + vector (searchTakes/searchTakesVector), and thetraversePathsgraph walk — inside the grant while still reaching authorized neighboring sources. NoDATABASE_URLneeded.test/e2e/skill-brain-first.test.ts— doctor reportsskill_brain_firstcheck with structured issues;--fix --dry-runpreviews insertion without writing;--fixapplies the canonical Convention callout idempotently;brain_first: exemptfrontmatter resolves the warn;brain_first_typosurfaces a paste-ready hint; audit JSONL recordsdetected/resolved/fixedtransitions; stable brain emits 0 audit lines/run.- Tier 2 (
test/e2e/skills.test.ts) requires OpenClaw + API keys, runs nightly in CI. - If
.env.testingdoesn't exist in this directory, check sibling worktrees:find ../ -maxdepth 2 -name .env.testing -print -quitand copy it here if found. - Run E2E tests without asking permission. When you want to verify behavior, there's a relevant E2E test, or you're shipping anything covered by an E2E suite — spin up the test DB, run the tests, tear down. Don't ask, don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL unset" is silent regression, not caution.
API keys and running ALL tests
ALWAYS source the user's shell profile before running tests:
source ~/.zshrc 2>/dev/null || true
This loads OPENAI_API_KEY and ANTHROPIC_API_KEY. Without these, Tier 2 tests
skip silently. Do NOT skip Tier 2 tests just because they require API keys — load
the keys and run them.
When asked to "run all E2E tests" or "run tests", that means ALL tiers:
- Tier 1:
bun run test:e2e(mechanical, sync, upgrade — no API keys needed) - Tier 2:
test/e2e/skills.test.ts(requires OpenAI + Anthropic + openclaw CLI) - Always spin up the test DB, source zshrc, run everything, tear down.
E2E test DB lifecycle (ALWAYS follow this)
You are responsible for spinning up and tearing down the test Postgres container. Do not leave containers running after tests. Do not skip E2E tests, do not ask permission to run them — see the "run without asking" rule above.
- Check for
.env.testing— if missing, copy from sibling worktree. Read it to get the DATABASE_URL (it has the port number). - Check if the port is free:
docker ps --filter "publish=PORT"— if another container is on that port, pick a different port (try 5435, 5436, 5437) and start on that one instead. - Start the test DB:
Wait for ready:
docker run -d --name gbrain-test-pg \ -e POSTGRES_USER=postgres -e POSTGRES_PASSWORD=postgres \ -e POSTGRES_DB=gbrain_test \ -p PORT:5432 pgvector/pgvector:pg16docker exec gbrain-test-pg pg_isready -U postgres - Bootstrap the schema (required — fresh containers have no
oauth_clients,mcp_request_log,pagesetc.; tests likeserve-http-oauth.test.tswill fail withrelation "oauth_clients" does not existif you skip this):DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \ bun run src/cli.ts doctor --json > /dev/null 2>&1gbrain doctortriggersinitSchema()on first connect, which is the canonical way to bring a fresh DB to head.apply-migrations --yesalone does NOT seed the base schema — it runs ALTER-style migrations on top ofinitSchema. Tests that bypass the engine (rawexecSync-spawnedauth register-client) hit the schema directly and need this step to have run first. - Run E2E tests:
DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e - Tear down immediately after tests finish (pass or fail):
docker stop gbrain-test-pg && docker rm gbrain-test-pg
Never leave gbrain-test-pg running. If you find a stale one from a previous run,
stop and remove it before starting a new one.