mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 00:48:18 +00:00
* docs(designs): agent-bootstrap plan + design docs (normative, review-absorbed) The scrubbed, in-repo sources of truth for the gbrain bootstrap wave: AGENT_BOOTSTRAP_DESIGN.md (product scope/sequencing) and AGENT_BOOTSTRAP_PLAN.md (implementation; all review-finding IDs inlined). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): format spec, question bank, identity templates, bundled assets agent.json manifest (format_version 1, initialized sentinel) + machine-local install receipt [CX2-1, CX2-12]; 12-question/6-required interview bank with consent keys and a persist:false sink for the optional provider key [CX2-13]; ten {{TOKEN}} identity templates (generic, adapted to gbrain ops — gates call recall/query/put_page, write-through-ops rule, keyless agent-authored facts, silence contract); assets embedded compiled-binary-safe via file-type imports [ENG-6]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): runbook, README paste block, bootstrap guide, TODOS entries BOOTSTRAP_FOR_AGENTS.md (agent-driven install runbook: CLI phase list is the source of truth, never-invent rules, Codex approvals preflight, keyless posture, failure-modes table, version stamp for the skew check); README gains the full-agent paste block pinned to latest-stable inside the Claude Code/Codex quick start (memory-only tier stays); docs/guides/bootstrap.md carries the full install/security/consent/degradation/uninstall contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(designs): spike instrument for the bootstrap wave (build order 0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): interview + render engines Interview gate with read-back confirm-hash (any later answer change clears the confirmation — the hostile single-batch case is structurally impossible), per-answer provenance, caps + escaping at set time, config-sink routing for the provider key; renderer with hard-fail token sweep, subordinate fencing of principal input, never-clobber + backups, deterministic minimal mode for the template repo, scaled byte floors. 58 unit tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): private-repo lifecycle — repo create, attach, uninstall, run lock gh-gated private repo creation with API-verified privacy (rate-limit distinct from public), refuse-foreign-origin with attach as the sanctioned path, atomic bootstrap mutex (pid liveness + age + token), receipt-keyed uninstall that never wholesale-deletes the gbrain home and only offers --delete-brain for a brain it created; read-only PGLite lock probe (never opens the engine). 54 unit tests, injectable exec seam throughout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(release): latest-stable ref, template-repo publish job, bootstrap CI guards release.yml advances the latest-stable tag only after assets publish (the paste block's permanent ref — copies in the wild never rot) and gains a PAT-gated publish-template job verified against the vendored tree; two skip-graceful guards (sanctioned-ref + runbook stamp; template/token bijection + placeholder assertion + generator byte-diff) wired into verify; README + runbook re-admitted to the CI cache hash; vendored deterministic template tree generated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(context): IPC v2 turn_context + 8KB assembly + visibility resolver + session identity Discriminated-union IPC with handler map, protocol echo (stale-serve detection), shared-secret gate, server-side source binding, per-kind budgets; turn-context assembly (reflex pointers + volunteered pages + world-only hot facts) under a data-not-instructions envelope trimmed to the harness's 10KB hook-output cap; facts.default_visibility resolved through one helper at all four sites (explicit caller wins, typos fail closed); typed sessionId threads _meta.session_id into the hot-memory cache key. 50 new tests; 180 adjacent tests confirmed green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(persistence): secret-scan, gbrain sources push, durability unification Pattern secret scanner (own runtime allowlist, redacted previews, corpus-write redaction mode); sources push runs the whole scan→stage→commit→pull→push sequence under one cross-platform lock (mkdir-atomic, pid+age+token) with a deny-glob backstop, commit-first divergence-safe pull, refuse-unverifiable visibility, and push-status telemetry; gbrain-home choke point unifies GBRAIN_HOME semantics with config (0700); durability is parent-repo-aware and rotates its push log at 0600. 35 new tests; 200 existing green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sources): harden/pull gates accept sources inside a parent git repo The bootstrap workspace registers brain/ (a subdirectory) as the source; the durability core already resolves the repo root, so the command gates now check inside-a-repo rather than .git-right-here [CX2-3]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(serve): resident maintenance sweep + keyless capability probe The lock-owning serve process now closes the persistence loop: startup (3s post-connect, best-effort, unref'd) and idle (10-min quiet intervals through the injectable timer seam) sweeps run facts-fence reconciliation, deterministic link/timeline extraction over recent workspace pages, and spend-gated corpus ingest (skipped keyless — agent-authored fences cover it). gbrain sweep --once is the trusted CLI seam bootstrap verify uses. Capability probe renders the honest keyless/keyed report. Full reuse of the cycle extractor + extract cores; 26 new tests, neighbors green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(hooks): engine-free gbrain hook command, settings writers, transcript parser Four hook events (session-start digest + crashed-session recovery push, user-prompt turn-context injection under an 800ms deadline and the 10KB cap, stop buffers, session-end corpus write with redaction/retention/dedup + best-effort push); structural JSON settings merger keyed by a _gbrain marker (foreign hooks and permissions survive); dated host-spec registry; Claude Code .jsonl parser as a spec-target with a scrubbed 7-shape fixture. Heartbeat is counters-only by construction. 59 tests; zero engine modules in the import graph. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(bootstrap): cross-link the full-agent path from the connection docs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): dispatcher, verify, status — the command assembled gbrain bootstrap {status,interview,render,repo,hooks,verify,uninstall,attach}: engine-free except verify (owns its engine, in-process sweep — no live-serve conflict); phase list is the TS source of truth with install.jsonl telemetry and the support blob; verify's fail-soft check suite covers the real write path (put_page → write-through file → sweep → graph floor → recall), passes keyless, persists snapshots, and ends with the first-run tour. cli.ts wired per the three-touchpoint rule; doctor gains the bootstrap check group (silent on machines with no bootstrap state). 28 new tests; 353 adjacent green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: KEY_FILES bootstrap cluster + CLAUDE.md dispatcher row (+ build:llms) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): e2e pins — hook-under-live-serve, attach, degraded modes, compiled binary, Docker harness The permanent pins: a real serve holds the PGLite lock while the engine-free hook completes (and a direct engine open provably throws LiveServeLockError); stale-socket fail-open; machine-2 attach with marker-keyed hook repair; decline-everything installs verify green with every degradation named; the compiled binary renders bundled templates in an empty cwd. Offline Docker harness (networkless, read-only) gated into heavy-tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bootstrap): register doctor check categories + system-of-record allow comments The six bootstrap doctor checks join OPS_CHECK_NAMES; the sweep's batch link/ timeline inserts carry the explicit extract-path allow comments (the sweep IS the extraction path for workspace pages). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard wedge cap tracks suite growth (1500s -> 1800s) At ~9000 tests a healthy shard finished at 1466s and two progressing shards were false-killed at the old cap; 1800s restores ~25% headroom over the slowest observed healthy shard. Real hangs still hit it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ci): cache-hash policy — README + runbook edits must invalidate [C2] The old deny-list assertion predates the paste block; README.md and BOOTSTRAP_FOR_AGENTS.md are policy-doc re-admissions now, so their edits must change the hash (a paste-block edit shipping under a cached green was the C2 hole). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): classify post-suite exit-hangs as warn-pass; file the leak forensics A shard killed by the wedge watchdog with every assigned file started and zero fail markers did all its work and leaked a handle at exit — pre-existing and master-reproducible (P1 TODO carries the full bisect forensics). Bun's per-test timeout turns a hung test into a (fail), so the classifier cannot mask one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard cap 2400s — the count-balanced heavy shard needs it under contention Observed: the heavy shard still progressing 22s before an 1800s kill while siblings finish at 1150-1550s (split balances file count, not weight). Filed the load-sensitive WAL-repair flake (pre-existing, master's v0.42.75.0 wave). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): quarantine env-mutating suites to the serial lane check-test-isolation R1: six new files mutate GBRAIN_HOME/env at module scope — the serial lane (one process per file) is the guard's prescribed home for them. All 114 tests pass post-rename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): per-harness install sections — Codex, Claude Code, then OpenClaw/Hermes Each harness gets its own complete paste-block section (desktop app first, terminal noted — Claude Code CLI is the identical harness; Codex CLI works pull-based today); the OpenClaw/Hermes platform path keeps equal weight with its one-click deploys and INSTALL_FOR_AGENTS block intact; memory-only and remote-connect tiers consolidated under 'Lighter ways in'. Supersedes the review's D5 ordering by user direction; stale heading references updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): Codex as the recommended first step; OpenClaw/Hermes framed as-intended, high-cost The install section now routes newcomers explicitly: Codex first (subscription-priced, nothing to deploy), OpenClaw/Hermes as GBrain used the way it was designed — always on, at real server + API cost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: doc-audit code chasers — broken recovery hints, stale op description, auto_link key Four small code fixes surfaced by the markdown accuracy audit: - doctor's auto-RLS recovery hint pointed at `apply-migrations --force-retry 35`, which cannot work (--force-retry targets the vX.Y.Z orchestrator registry, not the numeric schema MIGRATIONS array). Hint now points at the recreate SQL in docs/guides/rls-and-you.md; test pins against regression. - v0_11_0 migration printed the same broken-mechanism class of hint (`config set minion_mode` writes DB config nothing reads); now names `apply-migrations --mode` + preferences.json, the real setter. - submit_job's op description hardcoded a stale handler list; now points at registerBuiltinHandlers as the source plus the --follow discovery trick. - `auto_link` added to KNOWN_CONFIG_KEYS: read by link-extraction, reconcile-links, and sweep, and documented as the off-switch in brain-ops/maintain, but the allowlist rejected `config set auto_link false`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: repo-wide accuracy + MECE reform from the 9-bucket markdown audit A code-grounded audit of every markdown file (root, architecture, guides, mcp, tutorials, docs-root, operations/eval/designs, skills, recipes) followed by a fix wave with per-bucket ownership. Four classes of change: Accuracy — every documented command/flag verified against src/ before writing: dead commands replaced with working ones (pages purge-deleted, jobs watch --follow, gbrain restore, import-based Obsidian flow, space-separated --scopes, real thin-client recipes, working isolation verification, real supervisor restart procedure, curl-based ngrok health check, real minion_mode setter); count drift fixed with rot-proof phrasing (100+ ops, 50+ bundled skills via skills/manifest.json, 140+ engine methods, KNOBS_HASH_VERSION pointer instead of hardcoded versions); stale claims corrected (search-mode defaults, RETRIEVAL pipeline order incl. autocut, sentinel rules, refusal-list mechanism, engine snapshot, shard cap 2400s + EXIT-HANG classifier in TESTING.md, latest-stable + publish-template documented in RELEASING.md as release.yml promises). MECE — one home per concept, pointers elsewhere: test isolation → TESTING.md; OAuth registration + --bind/--public-url lore → DEPLOY.md; mode bundles → guides/search-modes.md (the home the CLAUDE.md dispatcher always promised); merge contract → schema-packs.md; WAL ladder → ENGINES.md; quiet-hours → quiet-hours.md; capture taxonomy → entity-detection.md; person-page taxonomy → compiled-truth.md; brain-first protocol → brain-first-lookup.md; refresh semantics → refresh-algorithm.md; KEY_FILES.md deduplicated (58 extension entries merged, one entry per file); infra-layer.md rewritten as a pointer page. Privacy — placeholder sweep across guides, docs, skills, and recipes per the iron rule; per-release narration stripped from reference docs (current-state prose only). Bootstrap coverage — AGENTS.md pointer, RESOLVER routing row, INSTALL.md path, tutorial cross-links, keyless-mode sections in spend-controls/headless-install. skills.lock.json regenerated; llms.txt/llms-full.txt rebuilt. Gates: verify 36/36, typecheck clean, doctor 96/96, skills-integrity + resolver + build-llms + config-set + migrations all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): refresh production-brain stats to current brain-repo counts 155,795 pages / 24,589 people / 5,340 companies, counted from the brain repo's current HEAD; the "100K-page brain" framing moves to 150K to match. Cron-fleet count unchanged (its store lives on the deployment host, not in the repos available for verification). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scan,push): modern OpenAI/Voyage key patterns, scan staged blobs not disk - secret-scan matches sk-proj-/sk-svcacct-/sk-None- and pa- Voyage keys (the bare sk- pattern missed every current OpenAI key format). - workspacePush stages first, then scans the staged index blobs via git cat-file, closing the scan-then-stage TOCTOU where a file changed between snapshot and commit shipped unscanned. - shared binary-sniff helper, memoized glob regexes, atomic push-status write, and tests for pull_conflict + gitignored deny-match paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): regenerate flag registry for new commands, harden shard classifier + release token - cli-flag-registry.generated.ts regenerated: bootstrap/hook/sweep and sources push --message/--allow-unverified-remote were missing, so the strict #2185 validator rejected real invocations and skipped the new commands entirely. - EXIT-HANG shard classifier now requires every assigned file to have started before warn-passing a watchdog kill (was fail-open). - release.yml passes TEMPLATE_REPO_PAT via http.extraheader, off the argv. - compiled-binary e2e fails loud in CI instead of a silent permanent skip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hook,sweep,ipc): non-blocking hook pushes, bounded sweep + cache, source-bound resolve - session-start/session-end no longer run synchronous git + inline push inside their self-deadline; a detached child does the push and the hook returns immediately (blocked Claude Code startup for minutes on a dirty tree before). - serve sweep drops the unbounded listAllPageRefs, resolves only candidate targets, claims corpus files atomically (no double-LLM-spend race), and caps the fence LIKE scan; heartbeat writes are O_APPEND with rare compaction. - hot-memory cache evicts expired entries and bounds entry count (the key is caller-controlled via _meta.session_id). - v1 resolve IPC honors boundSourceId like turn_context; turn-context runs its arms concurrently. doctor reads push/heartbeat thresholds from hook.ts. - new tests: doctor bootstrap checks, hook push-gate + deadline, concurrent sweep claims, cache eviction, bound-source resolve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bootstrap): origin-ownership gate, world visibility, collision-safe source id, consent + templates - repo adoption requires an exact receipt repo_url match or authed-owner check (undefined repo_url was a wildcard); create verifies privacy BEFORE the first push. - verify sets facts.default_visibility=world if unset, so agent-authored facts surface in per-turn context (they defaulted private before). - source_id derives a path-hash suffix when 'workspace' is taken by another checkout; every consumer reads manifest.source_id. - skipped HOOKS_CONSENT now declines (was falling through to default yes); --minimal refuses on an initialized manifest; tilde fences escaped. - MCP registration pins --surface full; status hard-fails a public origin (template door); receipt writers guard against newer/corrupt receipts; uninstall only claims brain-deleted after a real rm. - templates ship jobs disabled + provider-consent + support-relay lines; soul-audit re-runs over the shared interview bank. TODOS: 11 follow-ups. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): close adversarial-review findings — scan fails closed, whole-PEM redaction, bound repo push Cross-model adversarial pass (Claude + Codex) on the bootstrap wave: - secret scan fails CLOSED: an unreadable, oversized, or binary staged blob now blocks the push (blocked_unscannable, exit 5) instead of committing unscanned; only a confirmed staged deletion is skipped. This was the headline "block secrets before they leave the machine" property failing open. - private-key redaction spans the whole PEM block (header+body+footer), not just the header line — the base64 body no longer survives into the corpus the sweep sends to an extraction provider. - bootstrap repo commits the workspace (secret-scan-gated) before the first push and verifies the remote actually received it, so a push-fail retry can't adopt an empty remote as success. - privacy verify is re-bound to origin immediately before push (a concurrent origin rewrite between verify and push is refused). - session-end corpus write is atomic and clears the stale ingested/in-progress sidecars so a resumed session's appended transcript is re-ingested. - public-origin refusal enforced at render (not only status); MCP "already registered" is verified to target this workspace, not blessed blindly; verify probe cleanup scopes deletes to its own slugs, not a token substring; allowlist fingerprint floor raised 8→16 hex. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code Turns a Codex or Claude Code session into a persistent personal agent: interview-rendered identity files, a local PGLite brain, per-turn context via serve IPC (Claude Code hooks / Codex pull protocol), session-triggered persistence, and a private GitHub repo as the agent's portable body. Keyless- first (the harness model is the LLM; one optional key adds embeddings + extraction). New `gbrain bootstrap` command family + `gbrain hook` + `gbrain sweep`; doctor bootstrap health checks; latest-stable distribution ref + template-repo publish job. Opt-in, additive — existing installs untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: regenerate flag registry for security-fix flags; drop fabricated gbrain capabilities doc ref CI caught two real failures under the merged state: - the flag registry lagged the blocked_unscannable/exit-5 flags the security round added, tripping the #2185 freshness guard. - headless-install.md described the keyless capability report as a `gbrain capabilities` command, which the #3502 doc-command resolver rejects — reworded to prose (the real surface is bootstrap verify's report). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: sync KEY_FILES + bootstrap plan to security-fix behavior Cross-referenced the security-fix round against the reference docs and corrected the drift those commits introduced: - workspace-push.ts entry: stage-FIRST-then-scan order (the TOCTOU fix), fail-closed blocked_unscannable, and the sources-push status -> exit-code map. - hooks.ts entry: MCP registration pins `serve --surface full`. - hook.ts entry: session-start/session-end pushes run in a detached child (non-blocking); atomic corpus write clears stale sidecars. - bootstrap.ts entry: render hard-refuses a public origin (template door). - verify.ts entry: source_id collision resolution (workspace-<path-hash>). - AGENT_BOOTSTRAP_PLAN as-shipped delta note for the scan/stage reorder. llms bundle unchanged (KEY_FILES is link-only); build:llms and test/build-llms.test.ts green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): silence SC2016 on the intentional askpass literal in release.yml The one-shot GIT_ASKPASS script must contain literal $1 and $TEMPLATE_REPO_PAT so they expand when /bin/sh runs it at git's credential prompt, not when the outer shell writes the file — single quotes are correct. Add a scoped shellcheck disable so actionlint passes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills): declare 'bootstrap my data' trigger in cold-start frontmatter The doc reform added 'bootstrap my data' to cold-start's RESOLVER.md row (to disambiguate data-bootstrap from agent-bootstrap) but not to the skill's own frontmatter triggers, tripping the RESOLVER↔frontmatter round-trip contract (resolver.test.ts). Declare it; regenerate skills.lock.json. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): default per-turn hooks + search mode ON without a prompt Installing gbrain for your coding agent IS the consent for the behaviors that make it work, so stop re-litigating them with install-time questions whose "no" defeats the product: - Per-turn hooks (Claude Code) install ON by default — no prompt. Off-ramps: `--no-hooks` at install, `GBRAIN_HOOKS=0` at runtime, `bootstrap uninstall`. The "hooks installed" line now surfaces the kill switch so default-on is never silent. A persisted HOOKS_CONSENT=no (interview --skip) still declines. - Search mode defaults to `balanced` silently (nobody knows the modes at install; `gbrain search modes` changes it any time). - MCP scope stays the ONE deliberate prompt — project vs user is a real cross-repo privacy choice, not friction. Marks the two consents `silent: true` in the question bank (new QuestionSpec field), rewrites the runbook phases so the agent no longer asks them, adds the `--no-hooks` flag (+ registry regen), and adds default-on / opt-out tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): real end-to-end coverage — cross-session recall, per-turn content, Codex door, realistic corpus Closes the seven e2e gaps a coverage audit surfaced: the plumbing was well-unit-tested but the product claims ("Codex works, context shows up every turn with real content, it remembers across restarts, machine two recovers, Postgres works") were unproven end to end. Test-only wave — zero src changes. - Hermetic synthetic corpus (test/fixtures/bootstrap-corpus/ + a loader helper): 12 interlinked pages (52 edges, timelines), 12 world/private beliefs, 8 gold queries — curated from the gbrain-evals synthetic corpora, 100% placeholder names, so recall is asserted on a real multi-entity brain instead of a 2-node self-planted probe. - GAP1 magic moment: author a fact via the real write path, disconnect the engine, reopen against the same DB, recall it — a real session boundary, not verify.ts's same-connection SQL read-back. Plus a source-isolation assertion. - GAP2 per-turn content: hook-under-serve Pin 1 now seeds a known fact and asserts its text lands in the injected block AND private beliefs never do (was: empty brain, empty_block accepted as a pass). - GAP3 Codex door: assert the rendered AGENTS.md carries the Gate-3 brain-first pull protocol; make the fake codex shim implement `mcp get` so the [FIX7] target-verification can actually fail; the Docker cold-machine harness now exercises the hooks/MCP registration step instead of skipping it. - GAP4 corpus recall: turn-context + verify graph-floor/qrels run on the real multi-entity brain with real edges. - GAP5 attach: machine-two now re-ingests the cloned brain/ into a fresh DB and recalls a fact authored only on machine one — the multi-device payoff. - GAP6 keyed + Postgres (env-gated): real embeddings prove semantic recall a paraphrase query can reach but keyless BM25 cannot; bootstrap verify drives a real Postgres engine (skipIf DATABASE_URL/keys absent). - GAP7 persistence: session-end runs the REAL push (not the mocked seam) to a local bare remote and the remote receives the content; a planted secret is blocked at the gate; the 15-min cron installs and fires a scan-gated push. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): real-agent e2e — drive the actual claude + codex binaries end to end Closes the audit's biggest gap ("no real harness ever drives a turn"). Adapts gstack's PTY/headless agent harness to prove the bootstrap install + smoke work against the REAL binaries, not PATH shims. Test/CI/docs only — zero src changes. - test/helpers/agent-harness.ts: hermetic clean-room child env (ported from gstack; drops CONDUCTOR_/CLAUDE_/GSTACK_/MCP_/GBRAIN_, promotes GSTACK_ANTHROPIC_API_KEY→ANTHROPIC_API_KEY), real-binary resolvers + auth probes, headless `claude -p --output-format stream-json` and `codex exec --json` turn runners, a gbrain stdio MCP-config writer, and a keyless brain seeder. + a fixture-parse unit test (no binary needed). - test/e2e/bootstrap-real-claude.serial.test.ts: real `gbrain bootstrap` install → REAL `claude mcp add` (verified via `claude mcp get`) → verify exit 0 → a real `claude -p --mcp-config --strict-mcp-config` turn that invokes mcp__gbrain__search and answers from the brain (proven: toolCalls include mcp__gbrain__search, final text carries the seeded fact). - test/e2e/bootstrap-real-codex.serial.test.ts: same install with REAL `codex mcp add` into a real ~/.codex/config.toml + Gate-3 pull-protocol assertion, then a real `codex exec --json` turn surfacing the fact (MCP or the pull- protocol shell path). Bounded retry absorbs codex's occasional MCP-call cancellation without softening the fact-requiring assertion. - Everything hermetic (temp HOME/CLAUDE_CONFIG_DIR/CODEX_HOME/GBRAIN_HOME; real ~/.codex auth copied read-only) and skipIf-gated so it self-skips cleanly where the binaries/auth are absent. - heavy-tests.yml: gated `real-agent-e2e` job (nightly/label, never the PR shard; no-op on a runner without authed binaries). - TODOS: compiled `gbrain` binary can't serve a PGLite brain (bun compile omits the WASM/extension payloads); harness falls back to `bun run` serve. Verified against live claude 4.6 + codex 0.147.0: 15 pass / 0 fail; verify 36/36; typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): real-agent-e2e job — bash array + --timeout (actionlint SC2086 + bun-test-timeout guard) The real-agent-e2e job's file loop used an unquoted $FILES (SC2086) and ran `bun test` without --timeout (check-bun-test-timeout guard). Switch to a bash array and add --timeout=600000 (real-agent turns are slow; the tests self-skip without authed binaries so it's a no-op elsewhere). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pglite): embed WASM + extension assets so the compiled binary can serve A `bun build --compile` gbrain binary could not `serve` a PGLite brain: the compile bundles JS but not PGLite's runtime payload (pglite.wasm, initdb.wasm, pglite.data, vector/pg_trgm tarballs), so `serve` on PGLite died with a bunfs/ENOENT. Now the assets ride inside the binary. - src/core/pglite-embedded-assets.ts: embeds the five assets via `import … with { type: 'file' }` (the ENG-6 idiom) and exposes getEmbeddedPgliteOptions() → { pgliteWasmModule, initdbWasmModule, fsBundle, extensions:{vector,pg_trgm} }. WASM/fsBundle are consumed as bytes; the two extension tarballs are materialized to a content-addressed temp file (atomic, size-verified reuse) because PGLite reads them via fs.createReadStream, which cannot read a /$bunfs path. Unconditional (works in bun-run and compiled), so no fragile mode branch. - src/core/pglite-engine.ts: static-import getEmbeddedPgliteOptions (engine path stays static per the engine-dynamic-import invariant); spread into both PGlite.create sites (initial + WAL-repair retry). The bunfs classifier stays as a backstop but no longer fires for a correct binary. - scripts/check-pglite-embedded.sh (+ smoketest): compiles a focused binary and asserts it boots PGLite, CREATE EXTENSION vector/pg_trgm, and round-trips a page — wired into `bun run verify` (now 37 checks), check:all, and check:pglite-embedded. Fail-soft only when compile is unavailable. - agent-harness.ts probeCompiledPglite now passes → the real-agent e2e uses the fast compiled MCP server. TODOS: the P2 "can't serve PGLite" item is closed. Verified: fresh compiled binary ran `search`/`query` against a PGLite brain and returned the seeded row (no bunfs/ENOENT); verify 37/37; pglite-engine 120/0 source-mode; typecheck clean; engine-dynamic-import + parity guards pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
820 lines
42 KiB
Bash
Executable File
820 lines
42 KiB
Bash
Executable File
#!/usr/bin/env bash
|
||
# scripts/run-unit-parallel.sh — fast unit-test loop, parallel fan-out.
|
||
#
|
||
# Spawns N parallel `bun test` processes, each running a hash-disjoint shard
|
||
# of the unit-test set (files only — no e2e, no .slow, no .serial). After
|
||
# all shards complete, runs serial-only files (*.serial.test.ts) with
|
||
# --max-concurrency=1. Failure-first logging: extracts failure blocks from
|
||
# each shard's log, writes to .context/test-failures.log with --- shard $i:
|
||
# prefixes, prints loud stderr banner if any failures, exit non-zero.
|
||
#
|
||
# Usage:
|
||
# bash scripts/run-unit-parallel.sh [--shards N] [--max-concurrency N] [--dry-run]
|
||
#
|
||
# Env overrides:
|
||
# SHARDS=N same as --shards
|
||
# GBRAIN_TEST_SHARD_TIMEOUT per-shard wallclock cap, seconds (default 3000)
|
||
# GBRAIN_TEST_SHARD_KILL_AFTER grace after TERM before KILL (default 30)
|
||
# GBRAIN_TEST_MAX_CONCURRENCY passed through to bun test (default 4)
|
||
# GBRAIN_TEST_MEM_PER_FILE_MB memory budget per concurrent test file used by
|
||
# the adaptive sizing below (default 1536 — a
|
||
# PGLite WASM instance reserves ~1-1.5GB)
|
||
# GBRAIN_TEST_NO_MEM_ADAPT=1 disable memory-aware concurrency reduction
|
||
# GBRAIN_TEST_NO_OOM_FALLBACK=1 disable the serial OOM-rescue pass
|
||
#
|
||
# Memory safety (two layers; both default-on):
|
||
# 1. ADAPTIVE SIZING — before spawning, total concurrency (shards ×
|
||
# intra-shard --max-concurrency) is capped to what available memory can
|
||
# hold at GBRAIN_TEST_MEM_PER_FILE_MB per concurrent file. Concurrent
|
||
# Conductor workspaces running their own suites shrink the budget
|
||
# automatically instead of OOMing each other.
|
||
# 2. SERIAL PHANTOM RESCUE — two phantom classes are re-run serially
|
||
# (--max-concurrency 1) after the parallel pass: (a) failures whose
|
||
# shard log carries the PGLite WASM out-of-memory signature, and
|
||
# (b) shards killed EXTERNALLY (SIGTERM/SIGKILL well before the shard
|
||
# timeout — sibling Conductor workspaces' process cleanup, macOS memory
|
||
# jetsam). Phantoms pass serially and the run goes green with an
|
||
# oom_rescued note; real failures fail again and stay red. Plain
|
||
# assertion failures never match either signature.
|
||
#
|
||
# Output files (workspace-local; falls back to /tmp if .context/ unwritable):
|
||
# .context/test-failures.log failure blocks (cleared at start)
|
||
# .context/test-summary.txt per-shard pass/fail/skip/duration (cleared at start)
|
||
# .context/test-shards/ per-shard logs + exit codes (cleared at start)
|
||
|
||
set -uo pipefail
|
||
|
||
cd "$(dirname "$0")/.."
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# CPU detection: Apple Silicon perf cores → Mac total physical → nproc → 4.
|
||
# Returns a single positive integer.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
detect_cpus() {
|
||
local n=""
|
||
n=$(sysctl -n hw.perflevel0.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
|
||
n=$(sysctl -n hw.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
|
||
n=$(nproc 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
|
||
echo 4
|
||
}
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Available-memory detection (MB). macOS: vm_stat free + inactive +
|
||
# speculative + purgeable pages (inactive/purgeable are reclaimable on
|
||
# pressure, which is exactly the scenario we size for). Linux: MemAvailable.
|
||
# Unknown platform → 0, and the caller skips adaptation entirely.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
detect_available_mem_mb() {
|
||
if command -v vm_stat >/dev/null 2>&1; then
|
||
vm_stat 2>/dev/null | awk '
|
||
/page size of/ { psize = $8 }
|
||
/Pages free/ { free = $NF }
|
||
/Pages inactive/ { inactive = $NF }
|
||
/Pages speculative/ { spec = $NF }
|
||
/Pages purgeable/ { purge = $NF }
|
||
END {
|
||
gsub(/\./, "", free); gsub(/\./, "", inactive)
|
||
gsub(/\./, "", spec); gsub(/\./, "", purge)
|
||
if (psize == 0) psize = 16384
|
||
printf "%d\n", (free + inactive + spec + purge) * psize / 1048576
|
||
}'
|
||
return
|
||
fi
|
||
if [ -r /proc/meminfo ]; then
|
||
awk '/MemAvailable/ { printf "%d\n", $2 / 1024; found = 1 } END { if (!found) print 0 }' /proc/meminfo
|
||
return
|
||
fi
|
||
echo 0
|
||
}
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Argument parsing. --shards N override wins over $SHARDS; both are clamped.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
SHARDS_OVERRIDE=""
|
||
MAX_CONCURRENCY_OVERRIDE=""
|
||
DRY_RUN=0
|
||
while [ $# -gt 0 ]; do
|
||
case "$1" in
|
||
--shards) SHARDS_OVERRIDE="$2"; shift 2 ;;
|
||
--shards=*) SHARDS_OVERRIDE="${1#*=}"; shift ;;
|
||
--max-concurrency) MAX_CONCURRENCY_OVERRIDE="$2"; shift 2 ;;
|
||
--max-concurrency=*) MAX_CONCURRENCY_OVERRIDE="${1#*=}"; shift ;;
|
||
--dry-run) DRY_RUN=1; shift ;;
|
||
*) echo "ERROR: unknown arg: $1" >&2; exit 2 ;;
|
||
esac
|
||
done
|
||
|
||
N="${SHARDS_OVERRIDE:-${SHARDS:-$(detect_cpus)}}"
|
||
if ! printf '%s' "$N" | grep -qE '^[0-9]+$' || [ "$N" -lt 1 ]; then
|
||
echo "ERROR: invalid shard count: $N" >&2; exit 2
|
||
fi
|
||
# v0.40.10 flake-hardening: clamp default to 4 (was 8) to match CI's
|
||
# test-shard.sh fan-out. At 8-shard parallel on Apple Silicon we observed
|
||
# shard 5 SIGKILL during source-health.test.ts's PGLite migration replay —
|
||
# 8 parallel PGLite WASM inits contend severely on the lockfile, and the
|
||
# 92-migration replay × 8 simultaneous can wedge past even 900s. CI uses
|
||
# 4 and is stable. Trade ~2x wallclock for reliability + parity with CI's
|
||
# fan-out. Override via --shards N or SHARDS=N (still capped at 8).
|
||
[ "$N" -gt 8 ] && N=8
|
||
if [ -z "${SHARDS_OVERRIDE:-}" ] && [ -z "${SHARDS:-}" ] && [ "$N" -gt 4 ]; then
|
||
N=4
|
||
fi
|
||
|
||
INTRA_CONC="${MAX_CONCURRENCY_OVERRIDE:-${GBRAIN_TEST_MAX_CONCURRENCY:-4}}"
|
||
# v0.40.10 flake-hardening: bump per-shard cap 600 → 1500 (was 900). At
|
||
# 4-shard default each shard runs 159 files / ~2420 tests with internal
|
||
# wallclock 960-1020s. The 900s value (sized for 8-shard's ~80 files /
|
||
# 1100 tests at 620-770s) false-killed shard 1 at 900s even though it
|
||
# had completed in 968s. The cap must track suite growth: the suite roughly
|
||
# tripled since the 1500s cap was set (June: ~3900 tests, 92-migration PGLite
|
||
# replay; now: 13k+ tests with the agent-bootstrap wave, 120-migration replay
|
||
# per PGLite init). The split balances file COUNT, not weight — the heaviest
|
||
# count-balanced shard is still making steady per-test progress at 1800s under
|
||
# 4-way contention while its siblings finish at 1150-1550s. 3000s keeps the
|
||
# ~55%-headroom doctrine over observed wallclock; genuinely hung TESTS still
|
||
# die at bun's per-test timeout, mid-run stalls still hit this cap, and
|
||
# post-completion exit-hangs are classified separately (see the EXIT-HANG
|
||
# block below). Override via GBRAIN_TEST_SHARD_TIMEOUT=N.
|
||
SHARD_TIMEOUT="${GBRAIN_TEST_SHARD_TIMEOUT:-3000}"
|
||
SHARD_KILL_AFTER="${GBRAIN_TEST_SHARD_KILL_AFTER:-30}"
|
||
if ! printf '%s' "$SHARD_KILL_AFTER" | grep -qE '^[0-9]+$' || [ "$SHARD_KILL_AFTER" -lt 1 ]; then
|
||
echo "ERROR: invalid shard kill-after: $SHARD_KILL_AFTER" >&2; exit 2
|
||
fi
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Memory-aware concurrency (layer 1). Total concurrent test files =
|
||
# N shards × INTRA_CONC; each concurrent file can hold a PGLite WASM
|
||
# instance (~1-1.5GB reserved). 4×4 = 16 concurrent instances OOM'd on a
|
||
# 128GB machine when other Conductor workspaces ran their suites at the
|
||
# same time — every PGLite connect across every shard failed at once
|
||
# ("Out of memory" at PGlite.create). Cap total concurrency to what's
|
||
# actually available, keeping a 4GB reserve for the OS + bun itself.
|
||
# Applies to explicit --shards overrides too (an operator who wants an
|
||
# over-committed run sets GBRAIN_TEST_NO_MEM_ADAPT=1).
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
MEM_PER_FILE_MB="${GBRAIN_TEST_MEM_PER_FILE_MB:-1536}"
|
||
MEM_NOTE=""
|
||
if [ "${GBRAIN_TEST_NO_MEM_ADAPT:-0}" != "1" ]; then
|
||
AVAIL_MB=$(detect_available_mem_mb)
|
||
if [ "${AVAIL_MB:-0}" -gt 0 ] 2>/dev/null; then
|
||
BUDGET_MB=$((AVAIL_MB - 4096))
|
||
[ "$BUDGET_MB" -lt "$MEM_PER_FILE_MB" ] && BUDGET_MB="$MEM_PER_FILE_MB"
|
||
MAX_TOTAL=$((BUDGET_MB / MEM_PER_FILE_MB))
|
||
[ "$MAX_TOTAL" -lt 1 ] && MAX_TOTAL=1
|
||
ORIG_N="$N"; ORIG_INTRA="$INTRA_CONC"
|
||
# Shed shards before intra-shard concurrency: fewer bun processes frees
|
||
# more than narrower ones (each process carries its own heap + WASM).
|
||
while [ $((N * INTRA_CONC)) -gt "$MAX_TOTAL" ]; do
|
||
if [ "$N" -gt 1 ]; then N=$((N - 1))
|
||
elif [ "$INTRA_CONC" -gt 1 ]; then INTRA_CONC=$((INTRA_CONC - 1))
|
||
else break
|
||
fi
|
||
done
|
||
if [ "$N" != "$ORIG_N" ] || [ "$INTRA_CONC" != "$ORIG_INTRA" ]; then
|
||
# Fewer shards → more files per shard → each shard legitimately runs
|
||
# longer. Scale the per-shard cap by the shed ratio so adaptation
|
||
# doesn't convert memory safety into false WEDGED verdicts.
|
||
if [ "$N" -lt "$ORIG_N" ]; then
|
||
SHARD_TIMEOUT=$((SHARD_TIMEOUT * ORIG_N / N))
|
||
fi
|
||
MEM_NOTE=" | mem-adapted ${ORIG_N}x${ORIG_INTRA}→${N}x${INTRA_CONC} (avail=${AVAIL_MB}MB, ${MEM_PER_FILE_MB}MB/file, timeout→${SHARD_TIMEOUT}s)"
|
||
else
|
||
MEM_NOTE=" | mem-ok (avail=${AVAIL_MB}MB)"
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Output directories. Prefer workspace-local .context/, fall back to /tmp.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
LOG_DIR=""
|
||
if mkdir -p .context/test-shards 2>/dev/null; then
|
||
LOG_DIR=".context/test-shards"
|
||
FAILURES_LOG=".context/test-failures.log"
|
||
SUMMARY_FILE=".context/test-summary.txt"
|
||
else
|
||
LOG_DIR="/tmp/gbrain-test-shards-$$"
|
||
FAILURES_LOG="/tmp/gbrain-test-failures.log"
|
||
SUMMARY_FILE="/tmp/gbrain-test-summary.txt"
|
||
mkdir -p "$LOG_DIR" || { echo "ERROR: cannot create log dir" >&2; exit 2; }
|
||
fi
|
||
# Clear from prior run.
|
||
rm -f "$LOG_DIR"/shard-*.log "$LOG_DIR"/shard-*.exit "$LOG_DIR"/shard-*.wedged "$LOG_DIR"/shard-*.lastkb "$LOG_DIR"/shard-*.lastprogress "$LOG_DIR"/shard-*.start "$LOG_DIR"/shard-*.end 2>/dev/null
|
||
: > "$FAILURES_LOG"
|
||
: > "$SUMMARY_FILE"
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Resolve `timeout` command. macOS without coreutils has neither; we degrade
|
||
# to bg-pid + sleep cap. For now, prefer gtimeout (brew coreutils) → timeout.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
TIMEOUT_BIN=""
|
||
if command -v gtimeout >/dev/null 2>&1; then TIMEOUT_BIN="gtimeout"
|
||
elif command -v timeout >/dev/null 2>&1; then TIMEOUT_BIN="timeout"
|
||
fi
|
||
|
||
START_TS=$(date +%s)
|
||
echo "[unit-parallel] N=$N shards | --max-concurrency=$INTRA_CONC | timeout=${SHARD_TIMEOUT}s | kill-after=${SHARD_KILL_AFTER}s | logs=$LOG_DIR${MEM_NOTE}" >&2
|
||
|
||
if [ "$DRY_RUN" = "1" ]; then
|
||
echo "[unit-parallel] dry-run: would spawn $N shards with the above settings."
|
||
for i in $(seq 1 "$N"); do
|
||
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null \
|
||
| sed "s|^| [s$i] |"
|
||
done
|
||
exit 0
|
||
fi
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Spawn shards. Each child captures its own exit code into a sentinel file
|
||
# so $? is recoverable per-shard (we never trust `wait`'s aggregate value).
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
SHARD_PIDS=()
|
||
for i in $(seq 1 "$N"); do
|
||
(
|
||
SHARD_LOG="$LOG_DIR/shard-$i.log"
|
||
date +%s > "$LOG_DIR/shard-$i.start"
|
||
if [ -n "$TIMEOUT_BIN" ]; then
|
||
"$TIMEOUT_BIN" --signal=TERM --kill-after="${SHARD_KILL_AFTER}s" "${SHARD_TIMEOUT}s" \
|
||
env SHARD="$i/$N" \
|
||
bash scripts/run-unit-shard.sh --max-concurrency="$INTRA_CONC" \
|
||
> "$SHARD_LOG" 2>&1
|
||
rc=$?
|
||
else
|
||
env SHARD="$i/$N" \
|
||
bash scripts/run-unit-shard.sh --max-concurrency="$INTRA_CONC" \
|
||
> "$SHARD_LOG" 2>&1 &
|
||
pid=$!
|
||
( sleep "$SHARD_TIMEOUT" && kill -TERM "$pid" 2>/dev/null && \
|
||
sleep "$SHARD_KILL_AFTER" && kill -KILL "$pid" 2>/dev/null ) &
|
||
cap_pid=$!
|
||
wait "$pid" 2>/dev/null
|
||
# Capture the shard's exit code from ITS `wait`, before any watchdog
|
||
# teardown runs. The teardown commands below overwrite $? — the killed
|
||
# watchdog reports 143 — which used to get stamped into every shard's
|
||
# sentinel on machines with no gtimeout/timeout: every run "failed"
|
||
# with rc=143 summaries even when all tests passed.
|
||
rc=$?
|
||
# Reap the watchdog's `sleep` child too (pkill -P), then the watchdog.
|
||
# Killing only the subshell leaves the sleep orphaned until
|
||
# $SHARD_TIMEOUT elapses — same quirk the heartbeat cleanup below works
|
||
# around; CI's orphan-process sweep flags those.
|
||
pkill -P "$cap_pid" 2>/dev/null
|
||
kill "$cap_pid" 2>/dev/null
|
||
wait "$cap_pid" 2>/dev/null
|
||
fi
|
||
date +%s > "$LOG_DIR/shard-$i.end"
|
||
echo "$rc" > "$LOG_DIR/shard-$i.exit"
|
||
{ [ "$rc" = "124" ] || [ "$rc" = "137" ]; } && echo "WEDGED" > "$LOG_DIR/shard-$i.wedged"
|
||
) &
|
||
SHARD_PIDS+=($!)
|
||
done
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Heartbeat: every 10s, print per-shard progress to stderr by tailing logs
|
||
# and counting Bun's `(pass)` / `(fail)` / `(skip)` markers. Read-only.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# grep_count: returns 0 (single integer) if file is missing or zero matches,
|
||
# otherwise the match count. Avoids the `grep -c | echo 0` double-output bug
|
||
# where 0 matches produces a 2-line "0\n0" string that breaks arithmetic.
|
||
grep_count() {
|
||
local pattern="$1"; local file="$2"
|
||
if [ ! -f "$file" ]; then echo 0; return; fi
|
||
local n
|
||
n=$(grep -cE "$pattern" "$file" 2>/dev/null) || n=0
|
||
echo "${n:-0}"
|
||
}
|
||
|
||
# bun_summary_count: parses Bun's summary lines (one per `bun test` invocation
|
||
# inside a shard — there's only one when we pass an explicit file list).
|
||
# Looks for ` N pass` / ` N fail` / ` N skip` patterns and sums them across
|
||
# all summary blocks the shard emitted. `bun test` prints these near the end
|
||
# of its output. Format: leading whitespace + integer + space + label.
|
||
bun_summary_count() {
|
||
local label="$1"; local file="$2"
|
||
if [ ! -f "$file" ]; then echo 0; return; fi
|
||
awk -v label="$label" '
|
||
$1 ~ /^[0-9]+$/ && $2 == label { total += $1 }
|
||
END { print total + 0 }
|
||
' "$file"
|
||
}
|
||
|
||
# shard_total_files: parse the "[unit-shard N/M] running X files" line that
|
||
# run-unit-shard.sh echoes before invoking bun test. Returns the file count
|
||
# the shard was given, or 0 if the line isn't there yet (shard still
|
||
# bootstrapping). Uses sed-then-grep so it's portable to macOS awk (BSD awk
|
||
# doesn't support `match($0, /re/, arr)` with the array sink — that's gawk-only).
|
||
shard_total_files() {
|
||
local file="$1"
|
||
[ -f "$file" ] || { echo 0; return; }
|
||
local n
|
||
n=$(sed -n 's/^\[unit-shard [0-9][0-9]*\/[0-9][0-9]*\] running \([0-9][0-9]*\) files.*/\1/p' "$file" 2>/dev/null | head -1)
|
||
echo "${n:-0}"
|
||
}
|
||
|
||
# shard_pglite_init_count: count "Schema version" lines as a proxy for "test
|
||
# files initialized so far." Each PGLite-using test file's beforeAll triggers
|
||
# one initSchema() which prints this. Undercounts because not every test file
|
||
# opens a PGLite engine, but it's the only real-time progress signal bun's
|
||
# default reporter leaves in the log (bun has no per-file progress markers,
|
||
# only a final shard-end summary).
|
||
shard_pglite_init_count() {
|
||
local file="$1"
|
||
[ -f "$file" ] || { echo 0; return; }
|
||
grep -cE 'Schema version [0-9]+ → [0-9]+' "$file" 2>/dev/null || echo 0
|
||
}
|
||
|
||
# log_size_kb: total stderr+stdout written by the shard so far. Strictly
|
||
# monotonic — useful as a "definitely alive" signal when other heuristics
|
||
# read 0 (e.g. very early in shard startup before initSchema fires).
|
||
log_size_kb() {
|
||
local file="$1"
|
||
[ -f "$file" ] || { echo 0; return; }
|
||
local b
|
||
b=$(wc -c < "$file" 2>/dev/null | tr -d ' ')
|
||
echo $(( ${b:-0} / 1024 ))
|
||
}
|
||
|
||
# fmt_elapsed: pretty-print seconds → "Mm:SS" or "SSs" for short.
|
||
fmt_elapsed() {
|
||
local s=$1
|
||
if [ "$s" -ge 60 ]; then
|
||
printf '%dm%02ds' $((s / 60)) $((s % 60))
|
||
else
|
||
printf '%ds' "$s"
|
||
fi
|
||
}
|
||
|
||
heartbeat() {
|
||
local hb_start=$(date +%s)
|
||
while true; do
|
||
sleep 10
|
||
local line=""
|
||
local now; now=$(date +%s)
|
||
local hb_elapsed=$((now - hb_start))
|
||
for i in $(seq 1 "$N"); do
|
||
if [ -f "$LOG_DIR/shard-$i.exit" ]; then
|
||
local rc; rc=$(cat "$LOG_DIR/shard-$i.exit" 2>/dev/null || echo "?")
|
||
local status="✓"
|
||
[ "$rc" != "0" ] && status="✗"
|
||
local f
|
||
f=$(bun_summary_count "fail" "$LOG_DIR/shard-$i.log")
|
||
local p
|
||
p=$(bun_summary_count "pass" "$LOG_DIR/shard-$i.log")
|
||
line="$line [s$i: done $status ${p}p ${f}f]"
|
||
else
|
||
local lf="$LOG_DIR/shard-$i.log"
|
||
if [ -f "$lf" ]; then
|
||
# Bun's default reporter has no per-file progress markers, only a
|
||
# final shard-end summary, so we surface three complementary signals
|
||
# mid-run: (1) PGLite initSchema() count as a "files started" proxy,
|
||
# (2) total files this shard was assigned (from the runner banner),
|
||
# (3) log size in KB as a strictly-monotonic liveness signal.
|
||
local total; total=$(shard_total_files "$lf")
|
||
local pglite; pglite=$(shard_pglite_init_count "$lf")
|
||
local kb; kb=$(log_size_kb "$lf")
|
||
local et; et=$(fmt_elapsed "$hb_elapsed")
|
||
# Progress stamp for the exit-hang classifier: any log growth counts
|
||
# as progress. A wedged shard whose log went silent (≥ idle window)
|
||
# with zero fails did its work and hung at exit.
|
||
local prev_kb=""
|
||
[ -f "$LOG_DIR/shard-$i.lastkb" ] && prev_kb=$(cat "$LOG_DIR/shard-$i.lastkb" 2>/dev/null)
|
||
if [ "$kb" != "$prev_kb" ]; then
|
||
echo "$kb" > "$LOG_DIR/shard-$i.lastkb"
|
||
echo "$now" > "$LOG_DIR/shard-$i.lastprogress"
|
||
fi
|
||
if [ "$total" -gt 0 ]; then
|
||
line="$line [s$i: ~${pglite}/${total}f ${kb}KB ${et}]"
|
||
else
|
||
line="$line [s$i: starting ${kb}KB ${et}]"
|
||
fi
|
||
else
|
||
line="$line [s$i: spawning]"
|
||
fi
|
||
fi
|
||
done
|
||
printf '[heartbeat] %s\n' "$line" >&2
|
||
done
|
||
}
|
||
heartbeat &
|
||
HB_PID=$!
|
||
# v0.41.11.0 cleanup: pkill children FIRST, then kill heartbeat. If we
|
||
# kill the heartbeat shell first, its current `sleep 10` is reparented
|
||
# to init/launchd and pkill -P can no longer find it (orphan). Order:
|
||
# children first while the parent PID is still findable, then parent.
|
||
# Known bash quirk: SIGTERM to a shell sleeping inside `sleep` doesn't
|
||
# propagate to the sleep child before the wait returns. Without this,
|
||
# each invocation of this script leaks ONE orphan sleep; CI's "orphan
|
||
# process cleanup" at end-of-job reports them as (unnamed) test failures.
|
||
# Seen on the garrytan/port-pr-1406 PR, 2 CI runs in a row, 6 orphans
|
||
# matching the 6 invocations in test/scripts/run-unit-parallel.test.ts.
|
||
trap 'pkill -P "$HB_PID" 2>/dev/null; kill "$HB_PID" 2>/dev/null; wait "$HB_PID" 2>/dev/null' EXIT
|
||
|
||
# Wait for every shard. Don't care about wait's exit code.
|
||
for pid in "${SHARD_PIDS[@]}"; do wait "$pid" 2>/dev/null || true; done
|
||
|
||
pkill -P "$HB_PID" 2>/dev/null
|
||
kill "$HB_PID" 2>/dev/null
|
||
wait "$HB_PID" 2>/dev/null
|
||
trap - EXIT
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Aggregate failures (single writer; serial; never concurrent).
|
||
# Bun failure block format: from `(fail) ...` line through next `(pass)`,
|
||
# `(skip)`, blank line, or `__bun_test_summary__` marker.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
TOTAL_FAILURES=0
|
||
TOTAL_PASS=0
|
||
TOTAL_SKIP=0
|
||
TOTAL_RC=0
|
||
|
||
# Layer 2 state (serial OOM rescue). A shard whose log carries the WASM
|
||
# out-of-memory signature gets its failing files queued for a serial re-run;
|
||
# NON_OOM_FAIL records that at least one failure exists that the rescue lane
|
||
# must NOT absolve (plain assertion failures, wedges without the signature).
|
||
OOM_RE='Out of memory|WebAssembly\.Memory|RuntimeError: [Aa]borted|Aborted\(\)'
|
||
OOM_RESCUE_LIST="$LOG_DIR/oom-rescue-files.txt"
|
||
: > "$OOM_RESCUE_LIST"
|
||
NON_OOM_FAIL=0
|
||
# Set when any shard was killed externally — killed-midrun shards leave lock/
|
||
# state residue that can poison the LATER serial pass, so serial failures are
|
||
# only rescue-eligible under this flag (or their own OOM signature). A flaky
|
||
# serial test in an otherwise-clean run must stay red.
|
||
EXTERNAL_KILL_ANY=0
|
||
|
||
# failing_files_in_log: attribute each `(fail)` block to the test file whose
|
||
# `path.test.ts:` header most recently preceded it in bun's output. Under
|
||
# GITHUB_ACTIONS the shard wraps each file section as `::group::path.test.ts:`
|
||
# — strip that prefix or the rescue pass feeds bun literal `::group::...`
|
||
# non-paths that match zero test files (CI-only; local runs have no groups).
|
||
failing_files_in_log() {
|
||
local file="$1"
|
||
[ -f "$file" ] || return 0
|
||
awk '
|
||
/^(::group::)?[^ ].*\.test\.ts:$/ {
|
||
current = $0
|
||
sub(/^::group::/, "", current)
|
||
current = substr(current, 1, length(current) - 1)
|
||
next
|
||
}
|
||
/^\(fail\) / && current != "" { print current }
|
||
' "$file" | sort -u
|
||
}
|
||
|
||
# shard_unstarted_files: completion evidence for the EXIT-HANG classifier.
|
||
# Prints every file assigned to shard $1 (same deterministic split the shard
|
||
# itself used, via --dry-run-list) whose started file-header never appeared
|
||
# in the shard log $2. Bun prints `path.test.ts:` as each file starts; under
|
||
# GITHUB_ACTIONS that header is wrapped as `::group::path.test.ts:` — both
|
||
# forms count as started. Fail-closed: an underivable assigned list or a
|
||
# missing log emits markers so the caller treats the shard as WEDGED rather
|
||
# than warn-passing without evidence.
|
||
shard_unstarted_files() {
|
||
local shard_idx="$1" log="$2"
|
||
local assigned
|
||
assigned=$(SHARD="$shard_idx/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null)
|
||
if [ -z "$assigned" ]; then
|
||
echo "(assigned-file-list-underivable)"
|
||
return
|
||
fi
|
||
if [ ! -f "$log" ]; then
|
||
printf '%s\n' "$assigned"
|
||
return
|
||
fi
|
||
local af
|
||
while IFS= read -r af; do
|
||
[ -n "$af" ] || continue
|
||
if ! grep -qxF "${af}:" "$log" && ! grep -qxF "::group::${af}:" "$log"; then
|
||
printf '%s\n' "$af"
|
||
fi
|
||
done <<< "$assigned"
|
||
}
|
||
|
||
for i in $(seq 1 "$N"); do
|
||
SHARD_LOG="$LOG_DIR/shard-$i.log"
|
||
EXIT_FILE="$LOG_DIR/shard-$i.exit"
|
||
WEDGED_FILE="$LOG_DIR/shard-$i.wedged"
|
||
rc=1
|
||
[ -f "$EXIT_FILE" ] && rc=$(cat "$EXIT_FILE" 2>/dev/null || echo 1)
|
||
|
||
pass_count=$(bun_summary_count "pass" "$SHARD_LOG")
|
||
fail_count=$(bun_summary_count "fail" "$SHARD_LOG")
|
||
skip_count=$(bun_summary_count "skip" "$SHARD_LOG")
|
||
TOTAL_PASS=$((TOTAL_PASS + pass_count))
|
||
TOTAL_FAILURES=$((TOTAL_FAILURES + fail_count))
|
||
TOTAL_SKIP=$((TOTAL_SKIP + skip_count))
|
||
|
||
shard_oom=0
|
||
if [ "$rc" != "0" ] && [ "${GBRAIN_TEST_NO_OOM_FALLBACK:-0}" != "1" ] \
|
||
&& [ -f "$SHARD_LOG" ] && grep -qE "$OOM_RE" "$SHARD_LOG"; then
|
||
shard_oom=1
|
||
fi
|
||
|
||
# External-kill detection: rc 143 (SIGTERM) / 137 (SIGKILL) with the shard
|
||
# dying before 80% of the shard timeout means something OUTSIDE the runner
|
||
# killed it — sibling Conductor workspaces' process cleanup and macOS
|
||
# memory jetsam both present exactly this way (observed: 3 shards TERM'd +
|
||
# 1 KILL'd at ~700s under a 3000s cap, all mid-progress). A REAL wedge is
|
||
# killed BY the runner at ~SHARD_TIMEOUT and stays red. Externally-killed
|
||
# shards are phantoms: queue for the serial rescue lane like OOM.
|
||
shard_external_kill=0
|
||
if [ "$shard_oom" = "0" ] && [ "${GBRAIN_TEST_NO_OOM_FALLBACK:-0}" != "1" ] \
|
||
&& { [ "$rc" = "143" ] || [ "$rc" = "137" ]; }; then
|
||
s_start=$(cat "$LOG_DIR/shard-$i.start" 2>/dev/null) || s_start=""
|
||
s_end=$(cat "$LOG_DIR/shard-$i.end" 2>/dev/null) || s_end=""
|
||
if [ -n "$s_start" ] && [ -n "$s_end" ]; then
|
||
s_elapsed=$((s_end - s_start))
|
||
if [ "$s_elapsed" -lt $((SHARD_TIMEOUT * 80 / 100)) ]; then
|
||
shard_external_kill=1
|
||
EXTERNAL_KILL_ANY=1
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
if [ -f "$WEDGED_FILE" ]; then
|
||
# EXIT-HANG classifier (pre-existing PGLite-adjacent leak, TODOS.md
|
||
# "unit-shard exit hang"): a shard killed by the watchdog whose log shows
|
||
# every assigned file STARTED and zero (fail) markers did all its work and
|
||
# then failed to exit (a leaked ref'd handle; reproduces on master with
|
||
# the same file combination). Bun's per-test --timeout turns a genuinely
|
||
# hung TEST into a (fail), so this cannot mask one — the residual
|
||
# maskable case is a file-level import hang in the very last file, which
|
||
# the loud banner keeps visible. Classified shards warn instead of
|
||
# red-Xing the run; their pass counts are undercounted (bun never printed
|
||
# its final summary before the kill).
|
||
inline_fails=$(grep_count '^\(fail\) ' "$SHARD_LOG")
|
||
# Idle window: the log stopped growing this long before the kill. Bun's
|
||
# per-test --timeout turns a hung TEST into a printed (fail) — new output —
|
||
# so a silent-with-zero-fails shard was done with its work.
|
||
idle_secs=-1
|
||
if [ -f "$LOG_DIR/shard-$i.lastprogress" ] && [ -f "$WEDGED_FILE" ]; then
|
||
last_prog=$(cat "$LOG_DIR/shard-$i.lastprogress" 2>/dev/null || echo 0)
|
||
kill_ts=$(stat -f %m "$WEDGED_FILE" 2>/dev/null || stat -c %Y "$WEDGED_FILE" 2>/dev/null || echo 0)
|
||
[ "$kill_ts" -gt 0 ] && [ "$last_prog" -gt 0 ] && idle_secs=$((kill_ts - last_prog))
|
||
fi
|
||
# Warn-pass gate: rescue-eligible kills (OOM signature / external kill)
|
||
# are excluded so they reach the serial rescue queue below instead of
|
||
# being absolved without a re-run.
|
||
if [ "$fail_count" = "0" ] && [ "$inline_fails" = "0" ] && [ "$idle_secs" -ge 300 ] \
|
||
&& [ "$shard_oom" = "0" ] && [ "$shard_external_kill" = "0" ]; then
|
||
# Completion evidence (fail-closed): warn-pass additionally requires
|
||
# every assigned file to have STARTED (its file-header appears in the
|
||
# log). A silent idle window can also mean the shard wedged before
|
||
# reaching its last files — that stays a hard WEDGE.
|
||
unstarted=$(shard_unstarted_files "$i" "$SHARD_LOG")
|
||
if [ -z "$unstarted" ]; then
|
||
{
|
||
echo "⚠️ shard $i/$N: EXIT-HANG after ${SHARD_TIMEOUT}s — log silent for ${idle_secs}s with 0 failures"
|
||
echo " and every assigned file started; the process finished its work, leaked a handle, and"
|
||
echo " never exited (pre-existing, master-reproducible; see TODOS.md 'unit-shard exit hang')."
|
||
echo " Treating as pass-with-warning."
|
||
} >&2
|
||
echo "shard $i/$N: EXIT-HANG (idle ${idle_secs}s, 0 fails, all files started) rc=$rc — warn-pass" >> "$SUMMARY_FILE"
|
||
continue
|
||
fi
|
||
unstarted_count=$(printf '%s\n' "$unstarted" | grep -c .)
|
||
{
|
||
echo "⚠️ shard $i/$N: watchdog-killed with 0 fails and idle ${idle_secs}s, but ${unstarted_count} assigned"
|
||
echo " file(s) never started — classifying WEDGED, not EXIT-HANG:"
|
||
printf '%s\n' "$unstarted" | sed 's/^/ /'
|
||
} >&2
|
||
fi
|
||
TOTAL_RC=1
|
||
if [ "$shard_external_kill" = "1" ]; then
|
||
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
|
||
echo "shard $i/$N: KILLED externally after ${s_elapsed}s (rc=$rc, well before ${SHARD_TIMEOUT}s cap — queued for serial rescue)" >> "$SUMMARY_FILE"
|
||
elif [ "$shard_oom" = "1" ]; then
|
||
# Wedged UNDER memory pressure: we can't attribute failures, so queue
|
||
# the shard's entire file list for the serial rescue pass.
|
||
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
|
||
echo "shard $i/$N: WEDGED after ${SHARD_TIMEOUT}s (rc=$rc, OOM signature — queued for serial rescue)" >> "$SUMMARY_FILE"
|
||
else
|
||
NON_OOM_FAIL=1
|
||
echo "shard $i/$N: WEDGED after ${SHARD_TIMEOUT}s (rc=$rc)" >> "$SUMMARY_FILE"
|
||
fi
|
||
{
|
||
echo "--- shard $i: WEDGED after ${SHARD_TIMEOUT}s ---"
|
||
[ -f "$SHARD_LOG" ] && tail -50 "$SHARD_LOG"
|
||
echo ""
|
||
} >> "$FAILURES_LOG"
|
||
continue
|
||
fi
|
||
|
||
if [ "$rc" != "0" ]; then
|
||
if [ "$shard_oom" = "1" ]; then
|
||
# One scan, reused for both the queue append and the emptiness check.
|
||
shard_failing_files=$(failing_files_in_log "$SHARD_LOG")
|
||
if [ -n "$shard_failing_files" ]; then
|
||
printf '%s\n' "$shard_failing_files" >> "$OOM_RESCUE_LIST"
|
||
else
|
||
# OOM signature but no attributable files (e.g. bun died before any
|
||
# file header) → rescue the whole shard.
|
||
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
|
||
fi
|
||
elif [ "$shard_external_kill" = "1" ]; then
|
||
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
|
||
echo "shard $i/$N: KILLED externally after ${s_elapsed}s (rc=$rc — queued for serial rescue)" >> "$SUMMARY_FILE"
|
||
else
|
||
NON_OOM_FAIL=1
|
||
fi
|
||
fi
|
||
|
||
echo "shard $i/$N: pass=$pass_count fail=$fail_count skip=$skip_count rc=$rc" >> "$SUMMARY_FILE"
|
||
|
||
if [ "$rc" != "0" ]; then
|
||
TOTAL_RC=1
|
||
if [ "$fail_count" -gt 0 ] && [ -f "$SHARD_LOG" ]; then
|
||
# Extract each (fail) block: from `(fail)` line through next `(pass)`,
|
||
# `(skip)`, blank line, or `__bun_test_summary__`. Single awk pass.
|
||
awk -v shard="$i" '
|
||
/^\(fail\) / { in_block=1; print "--- shard " shard ": " $0; next }
|
||
in_block {
|
||
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
|
||
print $0
|
||
}
|
||
' "$SHARD_LOG" >> "$FAILURES_LOG"
|
||
elif [ -f "$SHARD_LOG" ]; then
|
||
# Non-zero rc but no (fail) line found — extraction couldn't pinpoint.
|
||
# Dump the full shard log so we never silently lose the failure cause.
|
||
{
|
||
echo "--- shard $i: rc=$rc, no (fail) markers — full log follows ---"
|
||
cat "$SHARD_LOG"
|
||
echo ""
|
||
} >> "$FAILURES_LOG"
|
||
fi
|
||
fi
|
||
done
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Print each shard's full output to stdout (developer expects to scroll
|
||
# through it). Print summary file last for one-glance overview.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
for i in $(seq 1 "$N"); do
|
||
SHARD_LOG="$LOG_DIR/shard-$i.log"
|
||
echo ""
|
||
echo "════════════ shard $i/$N ════════════"
|
||
[ -f "$SHARD_LOG" ] && cat "$SHARD_LOG"
|
||
done
|
||
echo ""
|
||
echo "════════════ summary ════════════"
|
||
cat "$SUMMARY_FILE"
|
||
echo ""
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Serial pass: any *.serial.test.ts files run after parallel pass.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
SERIAL_RC=0
|
||
SERIAL_FILES_COUNT=0
|
||
SERIAL_FILES_COUNT=$(find test -name '*.serial.test.ts' -not -path 'test/e2e/*' 2>/dev/null | wc -l | tr -d ' ')
|
||
if [ "$SERIAL_FILES_COUNT" -gt 0 ]; then
|
||
echo "════════════ serial pass ($SERIAL_FILES_COUNT files) ════════════"
|
||
bash scripts/run-serial-tests.sh > "$LOG_DIR/serial.log" 2>&1
|
||
SERIAL_RC=$?
|
||
cat "$LOG_DIR/serial.log"
|
||
if [ "$SERIAL_RC" != "0" ]; then
|
||
TOTAL_RC=1
|
||
if [ "${GBRAIN_TEST_NO_OOM_FALLBACK:-0}" != "1" ] \
|
||
&& { grep -qE "$OOM_RE" "$LOG_DIR/serial.log" || [ "$EXTERNAL_KILL_ANY" = "1" ]; }; then
|
||
# Serial failures are rescue-eligible ONLY with their own OOM signature
|
||
# or when an externally-killed shard ran earlier in this invocation
|
||
# (killed-midrun shards leave lock/state residue that poisons the serial
|
||
# pass). A merely-OOM'd sibling shard is NOT grounds — a flaky serial
|
||
# test must stay red rather than get silently absolved.
|
||
failing_files_in_log "$LOG_DIR/serial.log" >> "$OOM_RESCUE_LIST"
|
||
else
|
||
NON_OOM_FAIL=1
|
||
fi
|
||
s_fail=$(bun_summary_count "fail" "$LOG_DIR/serial.log")
|
||
TOTAL_FAILURES=$((TOTAL_FAILURES + s_fail))
|
||
if [ "$s_fail" -gt 0 ]; then
|
||
awk '
|
||
/^\(fail\) / { in_block=1; print "--- shard serial: " $0; next }
|
||
in_block {
|
||
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
|
||
print $0
|
||
}
|
||
' "$LOG_DIR/serial.log" >> "$FAILURES_LOG"
|
||
else
|
||
{
|
||
echo "--- shard serial: rc=$SERIAL_RC, no (fail) markers — full log follows ---"
|
||
cat "$LOG_DIR/serial.log"
|
||
echo ""
|
||
} >> "$FAILURES_LOG"
|
||
fi
|
||
echo "serial: rc=$SERIAL_RC fail=$s_fail" >> "$SUMMARY_FILE"
|
||
else
|
||
s_pass=$(bun_summary_count "pass" "$LOG_DIR/serial.log")
|
||
TOTAL_PASS=$((TOTAL_PASS + s_pass))
|
||
echo "serial: pass=$s_pass rc=0" >> "$SUMMARY_FILE"
|
||
fi
|
||
fi
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Layer 2: serial OOM rescue. Re-run every file that failed inside an
|
||
# OOM-signature shard, one at a time (1 shard, --max-concurrency 1), after
|
||
# the parallel fan-out has fully drained. Phantom failures (the WASM ran out
|
||
# of memory because 16 instances were up at once) pass here and the run goes
|
||
# green with an oom_rescued note; real failures fail again and stay red.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
OOM_RESCUED=0
|
||
OOM_RESCUE_NOTE=""
|
||
sort -u "$OOM_RESCUE_LIST" -o "$OOM_RESCUE_LIST" 2>/dev/null
|
||
# grep -c exits 1 on zero matches — assign in two steps so an empty rescue
|
||
# list yields a single "0" (the grep_count double-output bug, same class).
|
||
RESCUE_COUNT=$(grep -c . "$OOM_RESCUE_LIST" 2>/dev/null) || RESCUE_COUNT=0
|
||
if [ "$TOTAL_RC" != "0" ] && [ "${RESCUE_COUNT:-0}" -gt 0 ]; then
|
||
echo "════════════ OOM rescue pass ($RESCUE_COUNT files, serial) ════════════"
|
||
echo "[unit-parallel] OOM signature detected — re-running $RESCUE_COUNT failing file(s) at --max-concurrency 1" >&2
|
||
RESCUE_LOG="$LOG_DIR/oom-rescue.log"
|
||
# 60s-per-file floor with the shard cap as a minimum, and 2x the shard cap
|
||
# as a CEILING: a wedged shard queueing its whole file list must not turn
|
||
# `bun run test` into an unbounded multi-hour serial re-run — hitting the
|
||
# ceiling reads as a red rescue, not silence.
|
||
RESCUE_TIMEOUT=$((RESCUE_COUNT * 60))
|
||
[ "$RESCUE_TIMEOUT" -lt "$SHARD_TIMEOUT" ] && RESCUE_TIMEOUT="$SHARD_TIMEOUT"
|
||
[ "$RESCUE_TIMEOUT" -gt $((SHARD_TIMEOUT * 2)) ] && RESCUE_TIMEOUT=$((SHARD_TIMEOUT * 2))
|
||
# Split the queue: *.serial.test.ts files require one bun PROCESS per file
|
||
# (run-serial-tests.sh's isolation contract — top-level mock.module leaks
|
||
# across files in a shared registry); the remainder batches in one process.
|
||
# Both lanes mirror the shard invocation's --timeout=60000 — bun's default
|
||
# 5s per-test timeout would re-fail PGLite phantoms (120-migration replay)
|
||
# and mislabel them 'confirmed real'.
|
||
grep -v '\.serial\.test\.ts$' "$OOM_RESCUE_LIST" > "$LOG_DIR/oom-rescue-batch.txt" || true
|
||
grep '\.serial\.test\.ts$' "$OOM_RESCUE_LIST" > "$LOG_DIR/oom-rescue-serial.txt" || true
|
||
RESCUE_RC=0
|
||
: > "$RESCUE_LOG"
|
||
run_rescue() { # $1 = per-invocation timeout seconds; rest = test-file args
|
||
local t="$1"; shift
|
||
if [ -n "$TIMEOUT_BIN" ]; then
|
||
"$TIMEOUT_BIN" --signal=TERM --kill-after="${SHARD_KILL_AFTER}s" "${t}s" \
|
||
bun test --max-concurrency 1 --timeout=60000 "$@" >> "$RESCUE_LOG" 2>&1
|
||
else
|
||
bun test --max-concurrency 1 --timeout=60000 "$@" >> "$RESCUE_LOG" 2>&1
|
||
fi
|
||
}
|
||
if [ -s "$LOG_DIR/oom-rescue-batch.txt" ]; then
|
||
# shellcheck disable=SC2046
|
||
run_rescue "$RESCUE_TIMEOUT" $(cat "$LOG_DIR/oom-rescue-batch.txt") || RESCUE_RC=1
|
||
fi
|
||
if [ -s "$LOG_DIR/oom-rescue-serial.txt" ]; then
|
||
while IFS= read -r serial_file; do
|
||
[ -n "$serial_file" ] || continue
|
||
run_rescue 300 "$serial_file" || RESCUE_RC=1
|
||
done < "$LOG_DIR/oom-rescue-serial.txt"
|
||
fi
|
||
cat "$RESCUE_LOG"
|
||
r_pass=$(bun_summary_count "pass" "$RESCUE_LOG")
|
||
r_fail=$(bun_summary_count "fail" "$RESCUE_LOG")
|
||
if [ "$RESCUE_RC" = "0" ] && [ "$NON_OOM_FAIL" = "0" ]; then
|
||
# Every failure in the run was OOM-phantom and every rescued file passed
|
||
# serially: the run is green. Adjust the headline numbers so they reflect
|
||
# the rescue verdict, and mark the earlier failure blocks superseded.
|
||
TOTAL_RC=0
|
||
OOM_RESCUED=1
|
||
# Do NOT fold r_pass into TOTAL_PASS — the failing shard's own summary
|
||
# already counted the rescued files' passing tests, so folding would
|
||
# double-count. Rescue results ride in the note instead.
|
||
TOTAL_FAILURES=0
|
||
OOM_RESCUE_NOTE=" | oom_rescued=${RESCUE_COUNT}files(${r_pass}p serial)"
|
||
{
|
||
echo "--- OOM rescue: all $RESCUE_COUNT file(s) passed serially (${r_pass} tests) ---"
|
||
echo "--- failure blocks above were WASM out-of-memory phantoms, superseded ---"
|
||
} >> "$FAILURES_LOG"
|
||
echo "oom-rescue: $RESCUE_COUNT files pass=$r_pass rc=0 (phantom OOM failures superseded)" >> "$SUMMARY_FILE"
|
||
else
|
||
# Real failures confirmed serially (or a non-OOM failure exists anyway).
|
||
OOM_RESCUE_NOTE=" | oom_rescue_failed=${r_fail}real"
|
||
awk '
|
||
/^\(fail\) / { in_block=1; print "--- oom-rescue (serial, confirmed real): " $0; next }
|
||
in_block {
|
||
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
|
||
print $0
|
||
}
|
||
' "$RESCUE_LOG" >> "$FAILURES_LOG"
|
||
echo "oom-rescue: $RESCUE_COUNT files pass=$r_pass fail=$r_fail rc=$RESCUE_RC (real failures confirmed)" >> "$SUMMARY_FILE"
|
||
fi
|
||
fi
|
||
|
||
END_TS=$(date +%s)
|
||
ELAPSED=$((END_TS - START_TS))
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Loud banner if anything failed. To stderr so it survives `| head`/`| tail`.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
if [ "$TOTAL_RC" != "0" ]; then
|
||
ABS_FAIL=$(cd "$(dirname "$FAILURES_LOG")" && pwd)/$(basename "$FAILURES_LOG")
|
||
{
|
||
echo ""
|
||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||
echo "❌ $TOTAL_FAILURES TEST FAILURES — full details:"
|
||
echo " $ABS_FAIL"
|
||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||
tail -30 "$FAILURES_LOG"
|
||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||
echo "[unit-parallel] elapsed=${ELAPSED}s | pass=$TOTAL_PASS fail=$TOTAL_FAILURES skip=$TOTAL_SKIP${OOM_RESCUE_NOTE}"
|
||
} >&2
|
||
exit 1
|
||
fi
|
||
|
||
echo "[unit-parallel] elapsed=${ELAPSED}s | pass=$TOTAL_PASS fail=$TOTAL_FAILURES skip=$TOTAL_SKIP${OOM_RESCUE_NOTE}" >&2
|
||
exit 0
|