Files
gbrain/scripts/run-unit-parallel.sh
T
Garry TanandClaude Fable 5 d35c9c9e44 v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code (#3975)
* docs(designs): agent-bootstrap plan + design docs (normative, review-absorbed)

The scrubbed, in-repo sources of truth for the gbrain bootstrap wave:
AGENT_BOOTSTRAP_DESIGN.md (product scope/sequencing) and
AGENT_BOOTSTRAP_PLAN.md (implementation; all review-finding IDs inlined).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): format spec, question bank, identity templates, bundled assets

agent.json manifest (format_version 1, initialized sentinel) + machine-local
install receipt [CX2-1, CX2-12]; 12-question/6-required interview bank with
consent keys and a persist:false sink for the optional provider key [CX2-13];
ten {{TOKEN}} identity templates (generic, adapted to gbrain ops — gates call
recall/query/put_page, write-through-ops rule, keyless agent-authored facts,
silence contract); assets embedded compiled-binary-safe via file-type imports
[ENG-6].

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): runbook, README paste block, bootstrap guide, TODOS entries

BOOTSTRAP_FOR_AGENTS.md (agent-driven install runbook: CLI phase list is the
source of truth, never-invent rules, Codex approvals preflight, keyless posture,
failure-modes table, version stamp for the skew check); README gains the
full-agent paste block pinned to latest-stable inside the Claude Code/Codex
quick start (memory-only tier stays); docs/guides/bootstrap.md carries the full
install/security/consent/degradation/uninstall contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(designs): spike instrument for the bootstrap wave (build order 0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): interview + render engines

Interview gate with read-back confirm-hash (any later answer change clears the
confirmation — the hostile single-batch case is structurally impossible),
per-answer provenance, caps + escaping at set time, config-sink routing for the
provider key; renderer with hard-fail token sweep, subordinate fencing of
principal input, never-clobber + backups, deterministic minimal mode for the
template repo, scaled byte floors. 58 unit tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): private-repo lifecycle — repo create, attach, uninstall, run lock

gh-gated private repo creation with API-verified privacy (rate-limit distinct
from public), refuse-foreign-origin with attach as the sanctioned path, atomic
bootstrap mutex (pid liveness + age + token), receipt-keyed uninstall that
never wholesale-deletes the gbrain home and only offers --delete-brain for a
brain it created; read-only PGLite lock probe (never opens the engine).
54 unit tests, injectable exec seam throughout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(release): latest-stable ref, template-repo publish job, bootstrap CI guards

release.yml advances the latest-stable tag only after assets publish (the paste
block's permanent ref — copies in the wild never rot) and gains a PAT-gated
publish-template job verified against the vendored tree; two skip-graceful
guards (sanctioned-ref + runbook stamp; template/token bijection + placeholder
assertion + generator byte-diff) wired into verify; README + runbook re-admitted
to the CI cache hash; vendored deterministic template tree generated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(context): IPC v2 turn_context + 8KB assembly + visibility resolver + session identity

Discriminated-union IPC with handler map, protocol echo (stale-serve detection),
shared-secret gate, server-side source binding, per-kind budgets; turn-context
assembly (reflex pointers + volunteered pages + world-only hot facts) under a
data-not-instructions envelope trimmed to the harness's 10KB hook-output cap;
facts.default_visibility resolved through one helper at all four sites (explicit
caller wins, typos fail closed); typed sessionId threads _meta.session_id into
the hot-memory cache key. 50 new tests; 180 adjacent tests confirmed green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(persistence): secret-scan, gbrain sources push, durability unification

Pattern secret scanner (own runtime allowlist, redacted previews, corpus-write
redaction mode); sources push runs the whole scan→stage→commit→pull→push
sequence under one cross-platform lock (mkdir-atomic, pid+age+token) with a
deny-glob backstop, commit-first divergence-safe pull, refuse-unverifiable
visibility, and push-status telemetry; gbrain-home choke point unifies
GBRAIN_HOME semantics with config (0700); durability is parent-repo-aware and
rotates its push log at 0600. 35 new tests; 200 existing green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(sources): harden/pull gates accept sources inside a parent git repo

The bootstrap workspace registers brain/ (a subdirectory) as the source; the
durability core already resolves the repo root, so the command gates now check
inside-a-repo rather than .git-right-here [CX2-3].

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(serve): resident maintenance sweep + keyless capability probe

The lock-owning serve process now closes the persistence loop: startup (3s
post-connect, best-effort, unref'd) and idle (10-min quiet intervals through
the injectable timer seam) sweeps run facts-fence reconciliation, deterministic
link/timeline extraction over recent workspace pages, and spend-gated corpus
ingest (skipped keyless — agent-authored fences cover it). gbrain sweep --once
is the trusted CLI seam bootstrap verify uses. Capability probe renders the
honest keyless/keyed report. Full reuse of the cycle extractor + extract cores;
26 new tests, neighbors green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(hooks): engine-free gbrain hook command, settings writers, transcript parser

Four hook events (session-start digest + crashed-session recovery push,
user-prompt turn-context injection under an 800ms deadline and the 10KB cap,
stop buffers, session-end corpus write with redaction/retention/dedup +
best-effort push); structural JSON settings merger keyed by a _gbrain marker
(foreign hooks and permissions survive); dated host-spec registry; Claude Code
.jsonl parser as a spec-target with a scrubbed 7-shape fixture. Heartbeat is
counters-only by construction. 59 tests; zero engine modules in the import
graph.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(bootstrap): cross-link the full-agent path from the connection docs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): dispatcher, verify, status — the command assembled

gbrain bootstrap {status,interview,render,repo,hooks,verify,uninstall,attach}:
engine-free except verify (owns its engine, in-process sweep — no live-serve
conflict); phase list is the TS source of truth with install.jsonl telemetry
and the support blob; verify's fail-soft check suite covers the real write path
(put_page → write-through file → sweep → graph floor → recall), passes keyless,
persists snapshots, and ends with the first-run tour. cli.ts wired per the
three-touchpoint rule; doctor gains the bootstrap check group (silent on
machines with no bootstrap state). 28 new tests; 353 adjacent green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: KEY_FILES bootstrap cluster + CLAUDE.md dispatcher row (+ build:llms)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): e2e pins — hook-under-live-serve, attach, degraded modes, compiled binary, Docker harness

The permanent pins: a real serve holds the PGLite lock while the engine-free
hook completes (and a direct engine open provably throws LiveServeLockError);
stale-socket fail-open; machine-2 attach with marker-keyed hook repair;
decline-everything installs verify green with every degradation named; the
compiled binary renders bundled templates in an empty cwd. Offline Docker
harness (networkless, read-only) gated into heavy-tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): register doctor check categories + system-of-record allow comments

The six bootstrap doctor checks join OPS_CHECK_NAMES; the sweep's batch link/
timeline inserts carry the explicit extract-path allow comments (the sweep IS
the extraction path for workspace pages).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): shard wedge cap tracks suite growth (1500s -> 1800s)

At ~9000 tests a healthy shard finished at 1466s and two progressing shards
were false-killed at the old cap; 1800s restores ~25% headroom over the
slowest observed healthy shard. Real hangs still hit it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): cache-hash policy — README + runbook edits must invalidate [C2]

The old deny-list assertion predates the paste block; README.md and
BOOTSTRAP_FOR_AGENTS.md are policy-doc re-admissions now, so their edits must
change the hash (a paste-block edit shipping under a cached green was the C2
hole).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): classify post-suite exit-hangs as warn-pass; file the leak forensics

A shard killed by the wedge watchdog with every assigned file started and zero
fail markers did all its work and leaked a handle at exit — pre-existing and
master-reproducible (P1 TODO carries the full bisect forensics). Bun's per-test
timeout turns a hung test into a (fail), so the classifier cannot mask one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): shard cap 2400s — the count-balanced heavy shard needs it under contention

Observed: the heavy shard still progressing 22s before an 1800s kill while
siblings finish at 1150-1550s (split balances file count, not weight). Filed
the load-sensitive WAL-repair flake (pre-existing, master's v0.42.75.0 wave).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): quarantine env-mutating suites to the serial lane

check-test-isolation R1: six new files mutate GBRAIN_HOME/env at module scope —
the serial lane (one process per file) is the guard's prescribed home for them.
All 114 tests pass post-rename.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): per-harness install sections — Codex, Claude Code, then OpenClaw/Hermes

Each harness gets its own complete paste-block section (desktop app first,
terminal noted — Claude Code CLI is the identical harness; Codex CLI works
pull-based today); the OpenClaw/Hermes platform path keeps equal weight with
its one-click deploys and INSTALL_FOR_AGENTS block intact; memory-only and
remote-connect tiers consolidated under 'Lighter ways in'. Supersedes the
review's D5 ordering by user direction; stale heading references updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): Codex as the recommended first step; OpenClaw/Hermes framed as-intended, high-cost

The install section now routes newcomers explicitly: Codex first
(subscription-priced, nothing to deploy), OpenClaw/Hermes as GBrain used the
way it was designed — always on, at real server + API cost.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: doc-audit code chasers — broken recovery hints, stale op description, auto_link key

Four small code fixes surfaced by the markdown accuracy audit:
- doctor's auto-RLS recovery hint pointed at `apply-migrations --force-retry 35`,
  which cannot work (--force-retry targets the vX.Y.Z orchestrator registry, not
  the numeric schema MIGRATIONS array). Hint now points at the recreate SQL in
  docs/guides/rls-and-you.md; test pins against regression.
- v0_11_0 migration printed the same broken-mechanism class of hint
  (`config set minion_mode` writes DB config nothing reads); now names
  `apply-migrations --mode` + preferences.json, the real setter.
- submit_job's op description hardcoded a stale handler list; now points at
  registerBuiltinHandlers as the source plus the --follow discovery trick.
- `auto_link` added to KNOWN_CONFIG_KEYS: read by link-extraction, reconcile-links,
  and sweep, and documented as the off-switch in brain-ops/maintain, but the
  allowlist rejected `config set auto_link false`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: repo-wide accuracy + MECE reform from the 9-bucket markdown audit

A code-grounded audit of every markdown file (root, architecture, guides,
mcp, tutorials, docs-root, operations/eval/designs, skills, recipes) followed
by a fix wave with per-bucket ownership. Four classes of change:

Accuracy — every documented command/flag verified against src/ before writing:
dead commands replaced with working ones (pages purge-deleted, jobs watch
--follow, gbrain restore, import-based Obsidian flow, space-separated --scopes,
real thin-client recipes, working isolation verification, real supervisor
restart procedure, curl-based ngrok health check, real minion_mode setter);
count drift fixed with rot-proof phrasing (100+ ops, 50+ bundled skills via
skills/manifest.json, 140+ engine methods, KNOBS_HASH_VERSION pointer instead
of hardcoded versions); stale claims corrected (search-mode defaults, RETRIEVAL
pipeline order incl. autocut, sentinel rules, refusal-list mechanism, engine
snapshot, shard cap 2400s + EXIT-HANG classifier in TESTING.md, latest-stable +
publish-template documented in RELEASING.md as release.yml promises).

MECE — one home per concept, pointers elsewhere: test isolation → TESTING.md;
OAuth registration + --bind/--public-url lore → DEPLOY.md; mode bundles →
guides/search-modes.md (the home the CLAUDE.md dispatcher always promised);
merge contract → schema-packs.md; WAL ladder → ENGINES.md; quiet-hours →
quiet-hours.md; capture taxonomy → entity-detection.md; person-page taxonomy →
compiled-truth.md; brain-first protocol → brain-first-lookup.md; refresh
semantics → refresh-algorithm.md; KEY_FILES.md deduplicated (58 extension
entries merged, one entry per file); infra-layer.md rewritten as a pointer page.

Privacy — placeholder sweep across guides, docs, skills, and recipes per the
iron rule; per-release narration stripped from reference docs (current-state
prose only).

Bootstrap coverage — AGENTS.md pointer, RESOLVER routing row, INSTALL.md path,
tutorial cross-links, keyless-mode sections in spend-controls/headless-install.

skills.lock.json regenerated; llms.txt/llms-full.txt rebuilt. Gates: verify
36/36, typecheck clean, doctor 96/96, skills-integrity + resolver + build-llms
+ config-set + migrations all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): refresh production-brain stats to current brain-repo counts

155,795 pages / 24,589 people / 5,340 companies, counted from the brain
repo's current HEAD; the "100K-page brain" framing moves to 150K to match.
Cron-fleet count unchanged (its store lives on the deployment host, not in
the repos available for verification).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scan,push): modern OpenAI/Voyage key patterns, scan staged blobs not disk

- secret-scan matches sk-proj-/sk-svcacct-/sk-None- and pa- Voyage keys
  (the bare sk- pattern missed every current OpenAI key format).
- workspacePush stages first, then scans the staged index blobs via
  git cat-file, closing the scan-then-stage TOCTOU where a file changed
  between snapshot and commit shipped unscanned.
- shared binary-sniff helper, memoized glob regexes, atomic push-status
  write, and tests for pull_conflict + gitignored deny-match paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): regenerate flag registry for new commands, harden shard classifier + release token

- cli-flag-registry.generated.ts regenerated: bootstrap/hook/sweep and
  sources push --message/--allow-unverified-remote were missing, so the
  strict #2185 validator rejected real invocations and skipped the new
  commands entirely.
- EXIT-HANG shard classifier now requires every assigned file to have
  started before warn-passing a watchdog kill (was fail-open).
- release.yml passes TEMPLATE_REPO_PAT via http.extraheader, off the argv.
- compiled-binary e2e fails loud in CI instead of a silent permanent skip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(hook,sweep,ipc): non-blocking hook pushes, bounded sweep + cache, source-bound resolve

- session-start/session-end no longer run synchronous git + inline push
  inside their self-deadline; a detached child does the push and the hook
  returns immediately (blocked Claude Code startup for minutes on a dirty
  tree before).
- serve sweep drops the unbounded listAllPageRefs, resolves only candidate
  targets, claims corpus files atomically (no double-LLM-spend race), and
  caps the fence LIKE scan; heartbeat writes are O_APPEND with rare compaction.
- hot-memory cache evicts expired entries and bounds entry count (the key is
  caller-controlled via _meta.session_id).
- v1 resolve IPC honors boundSourceId like turn_context; turn-context runs
  its arms concurrently. doctor reads push/heartbeat thresholds from hook.ts.
- new tests: doctor bootstrap checks, hook push-gate + deadline, concurrent
  sweep claims, cache eviction, bound-source resolve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): origin-ownership gate, world visibility, collision-safe source id, consent + templates

- repo adoption requires an exact receipt repo_url match or authed-owner
  check (undefined repo_url was a wildcard); create verifies privacy BEFORE
  the first push.
- verify sets facts.default_visibility=world if unset, so agent-authored
  facts surface in per-turn context (they defaulted private before).
- source_id derives a path-hash suffix when 'workspace' is taken by another
  checkout; every consumer reads manifest.source_id.
- skipped HOOKS_CONSENT now declines (was falling through to default yes);
  --minimal refuses on an initialized manifest; tilde fences escaped.
- MCP registration pins --surface full; status hard-fails a public origin
  (template door); receipt writers guard against newer/corrupt receipts;
  uninstall only claims brain-deleted after a real rm.
- templates ship jobs disabled + provider-consent + support-relay lines;
  soul-audit re-runs over the shared interview bank. TODOS: 11 follow-ups.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): close adversarial-review findings — scan fails closed, whole-PEM redaction, bound repo push

Cross-model adversarial pass (Claude + Codex) on the bootstrap wave:

- secret scan fails CLOSED: an unreadable, oversized, or binary staged blob
  now blocks the push (blocked_unscannable, exit 5) instead of committing
  unscanned; only a confirmed staged deletion is skipped. This was the
  headline "block secrets before they leave the machine" property failing open.
- private-key redaction spans the whole PEM block (header+body+footer), not
  just the header line — the base64 body no longer survives into the corpus
  the sweep sends to an extraction provider.
- bootstrap repo commits the workspace (secret-scan-gated) before the first
  push and verifies the remote actually received it, so a push-fail retry
  can't adopt an empty remote as success.
- privacy verify is re-bound to origin immediately before push (a concurrent
  origin rewrite between verify and push is refused).
- session-end corpus write is atomic and clears the stale ingested/in-progress
  sidecars so a resumed session's appended transcript is re-ingested.
- public-origin refusal enforced at render (not only status); MCP "already
  registered" is verified to target this workspace, not blessed blindly;
  verify probe cleanup scopes deletes to its own slugs, not a token substring;
  allowlist fingerprint floor raised 8→16 hex.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code

Turns a Codex or Claude Code session into a persistent personal agent:
interview-rendered identity files, a local PGLite brain, per-turn context via
serve IPC (Claude Code hooks / Codex pull protocol), session-triggered
persistence, and a private GitHub repo as the agent's portable body. Keyless-
first (the harness model is the LLM; one optional key adds embeddings +
extraction). New `gbrain bootstrap` command family + `gbrain hook` + `gbrain
sweep`; doctor bootstrap health checks; latest-stable distribution ref +
template-repo publish job. Opt-in, additive — existing installs untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: regenerate flag registry for security-fix flags; drop fabricated gbrain capabilities doc ref

CI caught two real failures under the merged state:
- the flag registry lagged the blocked_unscannable/exit-5 flags the security
  round added, tripping the #2185 freshness guard.
- headless-install.md described the keyless capability report as a
  `gbrain capabilities` command, which the #3502 doc-command resolver
  rejects — reworded to prose (the real surface is bootstrap verify's report).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: sync KEY_FILES + bootstrap plan to security-fix behavior

Cross-referenced the security-fix round against the reference docs and
corrected the drift those commits introduced:

- workspace-push.ts entry: stage-FIRST-then-scan order (the TOCTOU fix),
  fail-closed blocked_unscannable, and the sources-push status -> exit-code map.
- hooks.ts entry: MCP registration pins `serve --surface full`.
- hook.ts entry: session-start/session-end pushes run in a detached child
  (non-blocking); atomic corpus write clears stale sidecars.
- bootstrap.ts entry: render hard-refuses a public origin (template door).
- verify.ts entry: source_id collision resolution (workspace-<path-hash>).
- AGENT_BOOTSTRAP_PLAN as-shipped delta note for the scan/stage reorder.

llms bundle unchanged (KEY_FILES is link-only); build:llms and
test/build-llms.test.ts green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): silence SC2016 on the intentional askpass literal in release.yml

The one-shot GIT_ASKPASS script must contain literal $1 and
$TEMPLATE_REPO_PAT so they expand when /bin/sh runs it at git's credential
prompt, not when the outer shell writes the file — single quotes are correct.
Add a scoped shellcheck disable so actionlint passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): declare 'bootstrap my data' trigger in cold-start frontmatter

The doc reform added 'bootstrap my data' to cold-start's RESOLVER.md row (to
disambiguate data-bootstrap from agent-bootstrap) but not to the skill's own
frontmatter triggers, tripping the RESOLVER↔frontmatter round-trip contract
(resolver.test.ts). Declare it; regenerate skills.lock.json.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): default per-turn hooks + search mode ON without a prompt

Installing gbrain for your coding agent IS the consent for the behaviors that
make it work, so stop re-litigating them with install-time questions whose
"no" defeats the product:

- Per-turn hooks (Claude Code) install ON by default — no prompt. Off-ramps:
  `--no-hooks` at install, `GBRAIN_HOOKS=0` at runtime, `bootstrap uninstall`.
  The "hooks installed" line now surfaces the kill switch so default-on is
  never silent. A persisted HOOKS_CONSENT=no (interview --skip) still declines.
- Search mode defaults to `balanced` silently (nobody knows the modes at
  install; `gbrain search modes` changes it any time).
- MCP scope stays the ONE deliberate prompt — project vs user is a real
  cross-repo privacy choice, not friction.

Marks the two consents `silent: true` in the question bank (new QuestionSpec
field), rewrites the runbook phases so the agent no longer asks them, adds the
`--no-hooks` flag (+ registry regen), and adds default-on / opt-out tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): real end-to-end coverage — cross-session recall, per-turn content, Codex door, realistic corpus

Closes the seven e2e gaps a coverage audit surfaced: the plumbing was
well-unit-tested but the product claims ("Codex works, context shows up every
turn with real content, it remembers across restarts, machine two recovers,
Postgres works") were unproven end to end. Test-only wave — zero src changes.

- Hermetic synthetic corpus (test/fixtures/bootstrap-corpus/ + a loader helper):
  12 interlinked pages (52 edges, timelines), 12 world/private beliefs, 8 gold
  queries — curated from the gbrain-evals synthetic corpora, 100% placeholder
  names, so recall is asserted on a real multi-entity brain instead of a
  2-node self-planted probe.
- GAP1 magic moment: author a fact via the real write path, disconnect the
  engine, reopen against the same DB, recall it — a real session boundary, not
  verify.ts's same-connection SQL read-back. Plus a source-isolation assertion.
- GAP2 per-turn content: hook-under-serve Pin 1 now seeds a known fact and
  asserts its text lands in the injected block AND private beliefs never do
  (was: empty brain, empty_block accepted as a pass).
- GAP3 Codex door: assert the rendered AGENTS.md carries the Gate-3 brain-first
  pull protocol; make the fake codex shim implement `mcp get` so the [FIX7]
  target-verification can actually fail; the Docker cold-machine harness now
  exercises the hooks/MCP registration step instead of skipping it.
- GAP4 corpus recall: turn-context + verify graph-floor/qrels run on the real
  multi-entity brain with real edges.
- GAP5 attach: machine-two now re-ingests the cloned brain/ into a fresh DB and
  recalls a fact authored only on machine one — the multi-device payoff.
- GAP6 keyed + Postgres (env-gated): real embeddings prove semantic recall a
  paraphrase query can reach but keyless BM25 cannot; bootstrap verify drives a
  real Postgres engine (skipIf DATABASE_URL/keys absent).
- GAP7 persistence: session-end runs the REAL push (not the mocked seam) to a
  local bare remote and the remote receives the content; a planted secret is
  blocked at the gate; the 15-min cron installs and fires a scan-gated push.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): real-agent e2e — drive the actual claude + codex binaries end to end

Closes the audit's biggest gap ("no real harness ever drives a turn"). Adapts
gstack's PTY/headless agent harness to prove the bootstrap install + smoke work
against the REAL binaries, not PATH shims. Test/CI/docs only — zero src changes.

- test/helpers/agent-harness.ts: hermetic clean-room child env (ported from
  gstack; drops CONDUCTOR_/CLAUDE_/GSTACK_/MCP_/GBRAIN_, promotes
  GSTACK_ANTHROPIC_API_KEY→ANTHROPIC_API_KEY), real-binary resolvers + auth
  probes, headless `claude -p --output-format stream-json` and `codex exec
  --json` turn runners, a gbrain stdio MCP-config writer, and a keyless brain
  seeder. + a fixture-parse unit test (no binary needed).
- test/e2e/bootstrap-real-claude.serial.test.ts: real `gbrain bootstrap`
  install → REAL `claude mcp add` (verified via `claude mcp get`) → verify
  exit 0 → a real `claude -p --mcp-config --strict-mcp-config` turn that
  invokes mcp__gbrain__search and answers from the brain (proven: toolCalls
  include mcp__gbrain__search, final text carries the seeded fact).
- test/e2e/bootstrap-real-codex.serial.test.ts: same install with REAL `codex
  mcp add` into a real ~/.codex/config.toml + Gate-3 pull-protocol assertion,
  then a real `codex exec --json` turn surfacing the fact (MCP or the pull-
  protocol shell path). Bounded retry absorbs codex's occasional MCP-call
  cancellation without softening the fact-requiring assertion.
- Everything hermetic (temp HOME/CLAUDE_CONFIG_DIR/CODEX_HOME/GBRAIN_HOME;
  real ~/.codex auth copied read-only) and skipIf-gated so it self-skips
  cleanly where the binaries/auth are absent.
- heavy-tests.yml: gated `real-agent-e2e` job (nightly/label, never the PR
  shard; no-op on a runner without authed binaries).
- TODOS: compiled `gbrain` binary can't serve a PGLite brain (bun compile
  omits the WASM/extension payloads); harness falls back to `bun run` serve.

Verified against live claude 4.6 + codex 0.147.0: 15 pass / 0 fail; verify
36/36; typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): real-agent-e2e job — bash array + --timeout (actionlint SC2086 + bun-test-timeout guard)

The real-agent-e2e job's file loop used an unquoted $FILES (SC2086) and ran
`bun test` without --timeout (check-bun-test-timeout guard). Switch to a bash
array and add --timeout=600000 (real-agent turns are slow; the tests self-skip
without authed binaries so it's a no-op elsewhere).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pglite): embed WASM + extension assets so the compiled binary can serve

A `bun build --compile` gbrain binary could not `serve` a PGLite brain: the
compile bundles JS but not PGLite's runtime payload (pglite.wasm, initdb.wasm,
pglite.data, vector/pg_trgm tarballs), so `serve` on PGLite died with a
bunfs/ENOENT. Now the assets ride inside the binary.

- src/core/pglite-embedded-assets.ts: embeds the five assets via
  `import … with { type: 'file' }` (the ENG-6 idiom) and exposes
  getEmbeddedPgliteOptions() → { pgliteWasmModule, initdbWasmModule, fsBundle,
  extensions:{vector,pg_trgm} }. WASM/fsBundle are consumed as bytes; the two
  extension tarballs are materialized to a content-addressed temp file (atomic,
  size-verified reuse) because PGLite reads them via fs.createReadStream, which
  cannot read a /$bunfs path. Unconditional (works in bun-run and compiled),
  so no fragile mode branch.
- src/core/pglite-engine.ts: static-import getEmbeddedPgliteOptions (engine
  path stays static per the engine-dynamic-import invariant); spread into both
  PGlite.create sites (initial + WAL-repair retry). The bunfs classifier stays
  as a backstop but no longer fires for a correct binary.
- scripts/check-pglite-embedded.sh (+ smoketest): compiles a focused binary and
  asserts it boots PGLite, CREATE EXTENSION vector/pg_trgm, and round-trips a
  page — wired into `bun run verify` (now 37 checks), check:all, and
  check:pglite-embedded. Fail-soft only when compile is unavailable.
- agent-harness.ts probeCompiledPglite now passes → the real-agent e2e uses the
  fast compiled MCP server. TODOS: the P2 "can't serve PGLite" item is closed.

Verified: fresh compiled binary ran `search`/`query` against a PGLite brain and
returned the seeded row (no bunfs/ENOENT); verify 37/37; pglite-engine 120/0
source-mode; typecheck clean; engine-dynamic-import + parity guards pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 16:42:11 -07:00

820 lines
42 KiB
Bash
Executable File
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env bash
# scripts/run-unit-parallel.sh — fast unit-test loop, parallel fan-out.
#
# Spawns N parallel `bun test` processes, each running a hash-disjoint shard
# of the unit-test set (files only — no e2e, no .slow, no .serial). After
# all shards complete, runs serial-only files (*.serial.test.ts) with
# --max-concurrency=1. Failure-first logging: extracts failure blocks from
# each shard's log, writes to .context/test-failures.log with --- shard $i:
# prefixes, prints loud stderr banner if any failures, exit non-zero.
#
# Usage:
# bash scripts/run-unit-parallel.sh [--shards N] [--max-concurrency N] [--dry-run]
#
# Env overrides:
# SHARDS=N same as --shards
# GBRAIN_TEST_SHARD_TIMEOUT per-shard wallclock cap, seconds (default 3000)
# GBRAIN_TEST_SHARD_KILL_AFTER grace after TERM before KILL (default 30)
# GBRAIN_TEST_MAX_CONCURRENCY passed through to bun test (default 4)
# GBRAIN_TEST_MEM_PER_FILE_MB memory budget per concurrent test file used by
# the adaptive sizing below (default 1536 — a
# PGLite WASM instance reserves ~1-1.5GB)
# GBRAIN_TEST_NO_MEM_ADAPT=1 disable memory-aware concurrency reduction
# GBRAIN_TEST_NO_OOM_FALLBACK=1 disable the serial OOM-rescue pass
#
# Memory safety (two layers; both default-on):
# 1. ADAPTIVE SIZING — before spawning, total concurrency (shards ×
# intra-shard --max-concurrency) is capped to what available memory can
# hold at GBRAIN_TEST_MEM_PER_FILE_MB per concurrent file. Concurrent
# Conductor workspaces running their own suites shrink the budget
# automatically instead of OOMing each other.
# 2. SERIAL PHANTOM RESCUE — two phantom classes are re-run serially
# (--max-concurrency 1) after the parallel pass: (a) failures whose
# shard log carries the PGLite WASM out-of-memory signature, and
# (b) shards killed EXTERNALLY (SIGTERM/SIGKILL well before the shard
# timeout — sibling Conductor workspaces' process cleanup, macOS memory
# jetsam). Phantoms pass serially and the run goes green with an
# oom_rescued note; real failures fail again and stay red. Plain
# assertion failures never match either signature.
#
# Output files (workspace-local; falls back to /tmp if .context/ unwritable):
# .context/test-failures.log failure blocks (cleared at start)
# .context/test-summary.txt per-shard pass/fail/skip/duration (cleared at start)
# .context/test-shards/ per-shard logs + exit codes (cleared at start)
set -uo pipefail
cd "$(dirname "$0")/.."
# ──────────────────────────────────────────────────────────────────────────
# CPU detection: Apple Silicon perf cores → Mac total physical → nproc → 4.
# Returns a single positive integer.
# ──────────────────────────────────────────────────────────────────────────
detect_cpus() {
local n=""
n=$(sysctl -n hw.perflevel0.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
n=$(sysctl -n hw.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
n=$(nproc 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
echo 4
}
# ──────────────────────────────────────────────────────────────────────────
# Available-memory detection (MB). macOS: vm_stat free + inactive +
# speculative + purgeable pages (inactive/purgeable are reclaimable on
# pressure, which is exactly the scenario we size for). Linux: MemAvailable.
# Unknown platform → 0, and the caller skips adaptation entirely.
# ──────────────────────────────────────────────────────────────────────────
detect_available_mem_mb() {
if command -v vm_stat >/dev/null 2>&1; then
vm_stat 2>/dev/null | awk '
/page size of/ { psize = $8 }
/Pages free/ { free = $NF }
/Pages inactive/ { inactive = $NF }
/Pages speculative/ { spec = $NF }
/Pages purgeable/ { purge = $NF }
END {
gsub(/\./, "", free); gsub(/\./, "", inactive)
gsub(/\./, "", spec); gsub(/\./, "", purge)
if (psize == 0) psize = 16384
printf "%d\n", (free + inactive + spec + purge) * psize / 1048576
}'
return
fi
if [ -r /proc/meminfo ]; then
awk '/MemAvailable/ { printf "%d\n", $2 / 1024; found = 1 } END { if (!found) print 0 }' /proc/meminfo
return
fi
echo 0
}
# ──────────────────────────────────────────────────────────────────────────
# Argument parsing. --shards N override wins over $SHARDS; both are clamped.
# ──────────────────────────────────────────────────────────────────────────
SHARDS_OVERRIDE=""
MAX_CONCURRENCY_OVERRIDE=""
DRY_RUN=0
while [ $# -gt 0 ]; do
case "$1" in
--shards) SHARDS_OVERRIDE="$2"; shift 2 ;;
--shards=*) SHARDS_OVERRIDE="${1#*=}"; shift ;;
--max-concurrency) MAX_CONCURRENCY_OVERRIDE="$2"; shift 2 ;;
--max-concurrency=*) MAX_CONCURRENCY_OVERRIDE="${1#*=}"; shift ;;
--dry-run) DRY_RUN=1; shift ;;
*) echo "ERROR: unknown arg: $1" >&2; exit 2 ;;
esac
done
N="${SHARDS_OVERRIDE:-${SHARDS:-$(detect_cpus)}}"
if ! printf '%s' "$N" | grep -qE '^[0-9]+$' || [ "$N" -lt 1 ]; then
echo "ERROR: invalid shard count: $N" >&2; exit 2
fi
# v0.40.10 flake-hardening: clamp default to 4 (was 8) to match CI's
# test-shard.sh fan-out. At 8-shard parallel on Apple Silicon we observed
# shard 5 SIGKILL during source-health.test.ts's PGLite migration replay —
# 8 parallel PGLite WASM inits contend severely on the lockfile, and the
# 92-migration replay × 8 simultaneous can wedge past even 900s. CI uses
# 4 and is stable. Trade ~2x wallclock for reliability + parity with CI's
# fan-out. Override via --shards N or SHARDS=N (still capped at 8).
[ "$N" -gt 8 ] && N=8
if [ -z "${SHARDS_OVERRIDE:-}" ] && [ -z "${SHARDS:-}" ] && [ "$N" -gt 4 ]; then
N=4
fi
INTRA_CONC="${MAX_CONCURRENCY_OVERRIDE:-${GBRAIN_TEST_MAX_CONCURRENCY:-4}}"
# v0.40.10 flake-hardening: bump per-shard cap 600 → 1500 (was 900). At
# 4-shard default each shard runs 159 files / ~2420 tests with internal
# wallclock 960-1020s. The 900s value (sized for 8-shard's ~80 files /
# 1100 tests at 620-770s) false-killed shard 1 at 900s even though it
# had completed in 968s. The cap must track suite growth: the suite roughly
# tripled since the 1500s cap was set (June: ~3900 tests, 92-migration PGLite
# replay; now: 13k+ tests with the agent-bootstrap wave, 120-migration replay
# per PGLite init). The split balances file COUNT, not weight — the heaviest
# count-balanced shard is still making steady per-test progress at 1800s under
# 4-way contention while its siblings finish at 1150-1550s. 3000s keeps the
# ~55%-headroom doctrine over observed wallclock; genuinely hung TESTS still
# die at bun's per-test timeout, mid-run stalls still hit this cap, and
# post-completion exit-hangs are classified separately (see the EXIT-HANG
# block below). Override via GBRAIN_TEST_SHARD_TIMEOUT=N.
SHARD_TIMEOUT="${GBRAIN_TEST_SHARD_TIMEOUT:-3000}"
SHARD_KILL_AFTER="${GBRAIN_TEST_SHARD_KILL_AFTER:-30}"
if ! printf '%s' "$SHARD_KILL_AFTER" | grep -qE '^[0-9]+$' || [ "$SHARD_KILL_AFTER" -lt 1 ]; then
echo "ERROR: invalid shard kill-after: $SHARD_KILL_AFTER" >&2; exit 2
fi
# ──────────────────────────────────────────────────────────────────────────
# Memory-aware concurrency (layer 1). Total concurrent test files =
# N shards × INTRA_CONC; each concurrent file can hold a PGLite WASM
# instance (~1-1.5GB reserved). 4×4 = 16 concurrent instances OOM'd on a
# 128GB machine when other Conductor workspaces ran their suites at the
# same time — every PGLite connect across every shard failed at once
# ("Out of memory" at PGlite.create). Cap total concurrency to what's
# actually available, keeping a 4GB reserve for the OS + bun itself.
# Applies to explicit --shards overrides too (an operator who wants an
# over-committed run sets GBRAIN_TEST_NO_MEM_ADAPT=1).
# ──────────────────────────────────────────────────────────────────────────
MEM_PER_FILE_MB="${GBRAIN_TEST_MEM_PER_FILE_MB:-1536}"
MEM_NOTE=""
if [ "${GBRAIN_TEST_NO_MEM_ADAPT:-0}" != "1" ]; then
AVAIL_MB=$(detect_available_mem_mb)
if [ "${AVAIL_MB:-0}" -gt 0 ] 2>/dev/null; then
BUDGET_MB=$((AVAIL_MB - 4096))
[ "$BUDGET_MB" -lt "$MEM_PER_FILE_MB" ] && BUDGET_MB="$MEM_PER_FILE_MB"
MAX_TOTAL=$((BUDGET_MB / MEM_PER_FILE_MB))
[ "$MAX_TOTAL" -lt 1 ] && MAX_TOTAL=1
ORIG_N="$N"; ORIG_INTRA="$INTRA_CONC"
# Shed shards before intra-shard concurrency: fewer bun processes frees
# more than narrower ones (each process carries its own heap + WASM).
while [ $((N * INTRA_CONC)) -gt "$MAX_TOTAL" ]; do
if [ "$N" -gt 1 ]; then N=$((N - 1))
elif [ "$INTRA_CONC" -gt 1 ]; then INTRA_CONC=$((INTRA_CONC - 1))
else break
fi
done
if [ "$N" != "$ORIG_N" ] || [ "$INTRA_CONC" != "$ORIG_INTRA" ]; then
# Fewer shards → more files per shard → each shard legitimately runs
# longer. Scale the per-shard cap by the shed ratio so adaptation
# doesn't convert memory safety into false WEDGED verdicts.
if [ "$N" -lt "$ORIG_N" ]; then
SHARD_TIMEOUT=$((SHARD_TIMEOUT * ORIG_N / N))
fi
MEM_NOTE=" | mem-adapted ${ORIG_N}x${ORIG_INTRA}${N}x${INTRA_CONC} (avail=${AVAIL_MB}MB, ${MEM_PER_FILE_MB}MB/file, timeout→${SHARD_TIMEOUT}s)"
else
MEM_NOTE=" | mem-ok (avail=${AVAIL_MB}MB)"
fi
fi
fi
# ──────────────────────────────────────────────────────────────────────────
# Output directories. Prefer workspace-local .context/, fall back to /tmp.
# ──────────────────────────────────────────────────────────────────────────
LOG_DIR=""
if mkdir -p .context/test-shards 2>/dev/null; then
LOG_DIR=".context/test-shards"
FAILURES_LOG=".context/test-failures.log"
SUMMARY_FILE=".context/test-summary.txt"
else
LOG_DIR="/tmp/gbrain-test-shards-$$"
FAILURES_LOG="/tmp/gbrain-test-failures.log"
SUMMARY_FILE="/tmp/gbrain-test-summary.txt"
mkdir -p "$LOG_DIR" || { echo "ERROR: cannot create log dir" >&2; exit 2; }
fi
# Clear from prior run.
rm -f "$LOG_DIR"/shard-*.log "$LOG_DIR"/shard-*.exit "$LOG_DIR"/shard-*.wedged "$LOG_DIR"/shard-*.lastkb "$LOG_DIR"/shard-*.lastprogress "$LOG_DIR"/shard-*.start "$LOG_DIR"/shard-*.end 2>/dev/null
: > "$FAILURES_LOG"
: > "$SUMMARY_FILE"
# ──────────────────────────────────────────────────────────────────────────
# Resolve `timeout` command. macOS without coreutils has neither; we degrade
# to bg-pid + sleep cap. For now, prefer gtimeout (brew coreutils) → timeout.
# ──────────────────────────────────────────────────────────────────────────
TIMEOUT_BIN=""
if command -v gtimeout >/dev/null 2>&1; then TIMEOUT_BIN="gtimeout"
elif command -v timeout >/dev/null 2>&1; then TIMEOUT_BIN="timeout"
fi
START_TS=$(date +%s)
echo "[unit-parallel] N=$N shards | --max-concurrency=$INTRA_CONC | timeout=${SHARD_TIMEOUT}s | kill-after=${SHARD_KILL_AFTER}s | logs=$LOG_DIR${MEM_NOTE}" >&2
if [ "$DRY_RUN" = "1" ]; then
echo "[unit-parallel] dry-run: would spawn $N shards with the above settings."
for i in $(seq 1 "$N"); do
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null \
| sed "s|^| [s$i] |"
done
exit 0
fi
# ──────────────────────────────────────────────────────────────────────────
# Spawn shards. Each child captures its own exit code into a sentinel file
# so $? is recoverable per-shard (we never trust `wait`'s aggregate value).
# ──────────────────────────────────────────────────────────────────────────
SHARD_PIDS=()
for i in $(seq 1 "$N"); do
(
SHARD_LOG="$LOG_DIR/shard-$i.log"
date +%s > "$LOG_DIR/shard-$i.start"
if [ -n "$TIMEOUT_BIN" ]; then
"$TIMEOUT_BIN" --signal=TERM --kill-after="${SHARD_KILL_AFTER}s" "${SHARD_TIMEOUT}s" \
env SHARD="$i/$N" \
bash scripts/run-unit-shard.sh --max-concurrency="$INTRA_CONC" \
> "$SHARD_LOG" 2>&1
rc=$?
else
env SHARD="$i/$N" \
bash scripts/run-unit-shard.sh --max-concurrency="$INTRA_CONC" \
> "$SHARD_LOG" 2>&1 &
pid=$!
( sleep "$SHARD_TIMEOUT" && kill -TERM "$pid" 2>/dev/null && \
sleep "$SHARD_KILL_AFTER" && kill -KILL "$pid" 2>/dev/null ) &
cap_pid=$!
wait "$pid" 2>/dev/null
# Capture the shard's exit code from ITS `wait`, before any watchdog
# teardown runs. The teardown commands below overwrite $? — the killed
# watchdog reports 143 — which used to get stamped into every shard's
# sentinel on machines with no gtimeout/timeout: every run "failed"
# with rc=143 summaries even when all tests passed.
rc=$?
# Reap the watchdog's `sleep` child too (pkill -P), then the watchdog.
# Killing only the subshell leaves the sleep orphaned until
# $SHARD_TIMEOUT elapses — same quirk the heartbeat cleanup below works
# around; CI's orphan-process sweep flags those.
pkill -P "$cap_pid" 2>/dev/null
kill "$cap_pid" 2>/dev/null
wait "$cap_pid" 2>/dev/null
fi
date +%s > "$LOG_DIR/shard-$i.end"
echo "$rc" > "$LOG_DIR/shard-$i.exit"
{ [ "$rc" = "124" ] || [ "$rc" = "137" ]; } && echo "WEDGED" > "$LOG_DIR/shard-$i.wedged"
) &
SHARD_PIDS+=($!)
done
# ──────────────────────────────────────────────────────────────────────────
# Heartbeat: every 10s, print per-shard progress to stderr by tailing logs
# and counting Bun's `(pass)` / `(fail)` / `(skip)` markers. Read-only.
# ──────────────────────────────────────────────────────────────────────────
# grep_count: returns 0 (single integer) if file is missing or zero matches,
# otherwise the match count. Avoids the `grep -c | echo 0` double-output bug
# where 0 matches produces a 2-line "0\n0" string that breaks arithmetic.
grep_count() {
local pattern="$1"; local file="$2"
if [ ! -f "$file" ]; then echo 0; return; fi
local n
n=$(grep -cE "$pattern" "$file" 2>/dev/null) || n=0
echo "${n:-0}"
}
# bun_summary_count: parses Bun's summary lines (one per `bun test` invocation
# inside a shard — there's only one when we pass an explicit file list).
# Looks for ` N pass` / ` N fail` / ` N skip` patterns and sums them across
# all summary blocks the shard emitted. `bun test` prints these near the end
# of its output. Format: leading whitespace + integer + space + label.
bun_summary_count() {
local label="$1"; local file="$2"
if [ ! -f "$file" ]; then echo 0; return; fi
awk -v label="$label" '
$1 ~ /^[0-9]+$/ && $2 == label { total += $1 }
END { print total + 0 }
' "$file"
}
# shard_total_files: parse the "[unit-shard N/M] running X files" line that
# run-unit-shard.sh echoes before invoking bun test. Returns the file count
# the shard was given, or 0 if the line isn't there yet (shard still
# bootstrapping). Uses sed-then-grep so it's portable to macOS awk (BSD awk
# doesn't support `match($0, /re/, arr)` with the array sink — that's gawk-only).
shard_total_files() {
local file="$1"
[ -f "$file" ] || { echo 0; return; }
local n
n=$(sed -n 's/^\[unit-shard [0-9][0-9]*\/[0-9][0-9]*\] running \([0-9][0-9]*\) files.*/\1/p' "$file" 2>/dev/null | head -1)
echo "${n:-0}"
}
# shard_pglite_init_count: count "Schema version" lines as a proxy for "test
# files initialized so far." Each PGLite-using test file's beforeAll triggers
# one initSchema() which prints this. Undercounts because not every test file
# opens a PGLite engine, but it's the only real-time progress signal bun's
# default reporter leaves in the log (bun has no per-file progress markers,
# only a final shard-end summary).
shard_pglite_init_count() {
local file="$1"
[ -f "$file" ] || { echo 0; return; }
grep -cE 'Schema version [0-9]+ → [0-9]+' "$file" 2>/dev/null || echo 0
}
# log_size_kb: total stderr+stdout written by the shard so far. Strictly
# monotonic — useful as a "definitely alive" signal when other heuristics
# read 0 (e.g. very early in shard startup before initSchema fires).
log_size_kb() {
local file="$1"
[ -f "$file" ] || { echo 0; return; }
local b
b=$(wc -c < "$file" 2>/dev/null | tr -d ' ')
echo $(( ${b:-0} / 1024 ))
}
# fmt_elapsed: pretty-print seconds → "Mm:SS" or "SSs" for short.
fmt_elapsed() {
local s=$1
if [ "$s" -ge 60 ]; then
printf '%dm%02ds' $((s / 60)) $((s % 60))
else
printf '%ds' "$s"
fi
}
heartbeat() {
local hb_start=$(date +%s)
while true; do
sleep 10
local line=""
local now; now=$(date +%s)
local hb_elapsed=$((now - hb_start))
for i in $(seq 1 "$N"); do
if [ -f "$LOG_DIR/shard-$i.exit" ]; then
local rc; rc=$(cat "$LOG_DIR/shard-$i.exit" 2>/dev/null || echo "?")
local status="✓"
[ "$rc" != "0" ] && status="✗"
local f
f=$(bun_summary_count "fail" "$LOG_DIR/shard-$i.log")
local p
p=$(bun_summary_count "pass" "$LOG_DIR/shard-$i.log")
line="$line [s$i: done $status ${p}p ${f}f]"
else
local lf="$LOG_DIR/shard-$i.log"
if [ -f "$lf" ]; then
# Bun's default reporter has no per-file progress markers, only a
# final shard-end summary, so we surface three complementary signals
# mid-run: (1) PGLite initSchema() count as a "files started" proxy,
# (2) total files this shard was assigned (from the runner banner),
# (3) log size in KB as a strictly-monotonic liveness signal.
local total; total=$(shard_total_files "$lf")
local pglite; pglite=$(shard_pglite_init_count "$lf")
local kb; kb=$(log_size_kb "$lf")
local et; et=$(fmt_elapsed "$hb_elapsed")
# Progress stamp for the exit-hang classifier: any log growth counts
# as progress. A wedged shard whose log went silent (≥ idle window)
# with zero fails did its work and hung at exit.
local prev_kb=""
[ -f "$LOG_DIR/shard-$i.lastkb" ] && prev_kb=$(cat "$LOG_DIR/shard-$i.lastkb" 2>/dev/null)
if [ "$kb" != "$prev_kb" ]; then
echo "$kb" > "$LOG_DIR/shard-$i.lastkb"
echo "$now" > "$LOG_DIR/shard-$i.lastprogress"
fi
if [ "$total" -gt 0 ]; then
line="$line [s$i: ~${pglite}/${total}f ${kb}KB ${et}]"
else
line="$line [s$i: starting ${kb}KB ${et}]"
fi
else
line="$line [s$i: spawning]"
fi
fi
done
printf '[heartbeat] %s\n' "$line" >&2
done
}
heartbeat &
HB_PID=$!
# v0.41.11.0 cleanup: pkill children FIRST, then kill heartbeat. If we
# kill the heartbeat shell first, its current `sleep 10` is reparented
# to init/launchd and pkill -P can no longer find it (orphan). Order:
# children first while the parent PID is still findable, then parent.
# Known bash quirk: SIGTERM to a shell sleeping inside `sleep` doesn't
# propagate to the sleep child before the wait returns. Without this,
# each invocation of this script leaks ONE orphan sleep; CI's "orphan
# process cleanup" at end-of-job reports them as (unnamed) test failures.
# Seen on the garrytan/port-pr-1406 PR, 2 CI runs in a row, 6 orphans
# matching the 6 invocations in test/scripts/run-unit-parallel.test.ts.
trap 'pkill -P "$HB_PID" 2>/dev/null; kill "$HB_PID" 2>/dev/null; wait "$HB_PID" 2>/dev/null' EXIT
# Wait for every shard. Don't care about wait's exit code.
for pid in "${SHARD_PIDS[@]}"; do wait "$pid" 2>/dev/null || true; done
pkill -P "$HB_PID" 2>/dev/null
kill "$HB_PID" 2>/dev/null
wait "$HB_PID" 2>/dev/null
trap - EXIT
# ──────────────────────────────────────────────────────────────────────────
# Aggregate failures (single writer; serial; never concurrent).
# Bun failure block format: from `(fail) ...` line through next `(pass)`,
# `(skip)`, blank line, or `__bun_test_summary__` marker.
# ──────────────────────────────────────────────────────────────────────────
TOTAL_FAILURES=0
TOTAL_PASS=0
TOTAL_SKIP=0
TOTAL_RC=0
# Layer 2 state (serial OOM rescue). A shard whose log carries the WASM
# out-of-memory signature gets its failing files queued for a serial re-run;
# NON_OOM_FAIL records that at least one failure exists that the rescue lane
# must NOT absolve (plain assertion failures, wedges without the signature).
OOM_RE='Out of memory|WebAssembly\.Memory|RuntimeError: [Aa]borted|Aborted\(\)'
OOM_RESCUE_LIST="$LOG_DIR/oom-rescue-files.txt"
: > "$OOM_RESCUE_LIST"
NON_OOM_FAIL=0
# Set when any shard was killed externally — killed-midrun shards leave lock/
# state residue that can poison the LATER serial pass, so serial failures are
# only rescue-eligible under this flag (or their own OOM signature). A flaky
# serial test in an otherwise-clean run must stay red.
EXTERNAL_KILL_ANY=0
# failing_files_in_log: attribute each `(fail)` block to the test file whose
# `path.test.ts:` header most recently preceded it in bun's output. Under
# GITHUB_ACTIONS the shard wraps each file section as `::group::path.test.ts:`
# — strip that prefix or the rescue pass feeds bun literal `::group::...`
# non-paths that match zero test files (CI-only; local runs have no groups).
failing_files_in_log() {
local file="$1"
[ -f "$file" ] || return 0
awk '
/^(::group::)?[^ ].*\.test\.ts:$/ {
current = $0
sub(/^::group::/, "", current)
current = substr(current, 1, length(current) - 1)
next
}
/^\(fail\) / && current != "" { print current }
' "$file" | sort -u
}
# shard_unstarted_files: completion evidence for the EXIT-HANG classifier.
# Prints every file assigned to shard $1 (same deterministic split the shard
# itself used, via --dry-run-list) whose started file-header never appeared
# in the shard log $2. Bun prints `path.test.ts:` as each file starts; under
# GITHUB_ACTIONS that header is wrapped as `::group::path.test.ts:` — both
# forms count as started. Fail-closed: an underivable assigned list or a
# missing log emits markers so the caller treats the shard as WEDGED rather
# than warn-passing without evidence.
shard_unstarted_files() {
local shard_idx="$1" log="$2"
local assigned
assigned=$(SHARD="$shard_idx/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null)
if [ -z "$assigned" ]; then
echo "(assigned-file-list-underivable)"
return
fi
if [ ! -f "$log" ]; then
printf '%s\n' "$assigned"
return
fi
local af
while IFS= read -r af; do
[ -n "$af" ] || continue
if ! grep -qxF "${af}:" "$log" && ! grep -qxF "::group::${af}:" "$log"; then
printf '%s\n' "$af"
fi
done <<< "$assigned"
}
for i in $(seq 1 "$N"); do
SHARD_LOG="$LOG_DIR/shard-$i.log"
EXIT_FILE="$LOG_DIR/shard-$i.exit"
WEDGED_FILE="$LOG_DIR/shard-$i.wedged"
rc=1
[ -f "$EXIT_FILE" ] && rc=$(cat "$EXIT_FILE" 2>/dev/null || echo 1)
pass_count=$(bun_summary_count "pass" "$SHARD_LOG")
fail_count=$(bun_summary_count "fail" "$SHARD_LOG")
skip_count=$(bun_summary_count "skip" "$SHARD_LOG")
TOTAL_PASS=$((TOTAL_PASS + pass_count))
TOTAL_FAILURES=$((TOTAL_FAILURES + fail_count))
TOTAL_SKIP=$((TOTAL_SKIP + skip_count))
shard_oom=0
if [ "$rc" != "0" ] && [ "${GBRAIN_TEST_NO_OOM_FALLBACK:-0}" != "1" ] \
&& [ -f "$SHARD_LOG" ] && grep -qE "$OOM_RE" "$SHARD_LOG"; then
shard_oom=1
fi
# External-kill detection: rc 143 (SIGTERM) / 137 (SIGKILL) with the shard
# dying before 80% of the shard timeout means something OUTSIDE the runner
# killed it — sibling Conductor workspaces' process cleanup and macOS
# memory jetsam both present exactly this way (observed: 3 shards TERM'd +
# 1 KILL'd at ~700s under a 3000s cap, all mid-progress). A REAL wedge is
# killed BY the runner at ~SHARD_TIMEOUT and stays red. Externally-killed
# shards are phantoms: queue for the serial rescue lane like OOM.
shard_external_kill=0
if [ "$shard_oom" = "0" ] && [ "${GBRAIN_TEST_NO_OOM_FALLBACK:-0}" != "1" ] \
&& { [ "$rc" = "143" ] || [ "$rc" = "137" ]; }; then
s_start=$(cat "$LOG_DIR/shard-$i.start" 2>/dev/null) || s_start=""
s_end=$(cat "$LOG_DIR/shard-$i.end" 2>/dev/null) || s_end=""
if [ -n "$s_start" ] && [ -n "$s_end" ]; then
s_elapsed=$((s_end - s_start))
if [ "$s_elapsed" -lt $((SHARD_TIMEOUT * 80 / 100)) ]; then
shard_external_kill=1
EXTERNAL_KILL_ANY=1
fi
fi
fi
if [ -f "$WEDGED_FILE" ]; then
# EXIT-HANG classifier (pre-existing PGLite-adjacent leak, TODOS.md
# "unit-shard exit hang"): a shard killed by the watchdog whose log shows
# every assigned file STARTED and zero (fail) markers did all its work and
# then failed to exit (a leaked ref'd handle; reproduces on master with
# the same file combination). Bun's per-test --timeout turns a genuinely
# hung TEST into a (fail), so this cannot mask one — the residual
# maskable case is a file-level import hang in the very last file, which
# the loud banner keeps visible. Classified shards warn instead of
# red-Xing the run; their pass counts are undercounted (bun never printed
# its final summary before the kill).
inline_fails=$(grep_count '^\(fail\) ' "$SHARD_LOG")
# Idle window: the log stopped growing this long before the kill. Bun's
# per-test --timeout turns a hung TEST into a printed (fail) — new output —
# so a silent-with-zero-fails shard was done with its work.
idle_secs=-1
if [ -f "$LOG_DIR/shard-$i.lastprogress" ] && [ -f "$WEDGED_FILE" ]; then
last_prog=$(cat "$LOG_DIR/shard-$i.lastprogress" 2>/dev/null || echo 0)
kill_ts=$(stat -f %m "$WEDGED_FILE" 2>/dev/null || stat -c %Y "$WEDGED_FILE" 2>/dev/null || echo 0)
[ "$kill_ts" -gt 0 ] && [ "$last_prog" -gt 0 ] && idle_secs=$((kill_ts - last_prog))
fi
# Warn-pass gate: rescue-eligible kills (OOM signature / external kill)
# are excluded so they reach the serial rescue queue below instead of
# being absolved without a re-run.
if [ "$fail_count" = "0" ] && [ "$inline_fails" = "0" ] && [ "$idle_secs" -ge 300 ] \
&& [ "$shard_oom" = "0" ] && [ "$shard_external_kill" = "0" ]; then
# Completion evidence (fail-closed): warn-pass additionally requires
# every assigned file to have STARTED (its file-header appears in the
# log). A silent idle window can also mean the shard wedged before
# reaching its last files — that stays a hard WEDGE.
unstarted=$(shard_unstarted_files "$i" "$SHARD_LOG")
if [ -z "$unstarted" ]; then
{
echo "⚠️ shard $i/$N: EXIT-HANG after ${SHARD_TIMEOUT}s — log silent for ${idle_secs}s with 0 failures"
echo " and every assigned file started; the process finished its work, leaked a handle, and"
echo " never exited (pre-existing, master-reproducible; see TODOS.md 'unit-shard exit hang')."
echo " Treating as pass-with-warning."
} >&2
echo "shard $i/$N: EXIT-HANG (idle ${idle_secs}s, 0 fails, all files started) rc=$rc — warn-pass" >> "$SUMMARY_FILE"
continue
fi
unstarted_count=$(printf '%s\n' "$unstarted" | grep -c .)
{
echo "⚠️ shard $i/$N: watchdog-killed with 0 fails and idle ${idle_secs}s, but ${unstarted_count} assigned"
echo " file(s) never started — classifying WEDGED, not EXIT-HANG:"
printf '%s\n' "$unstarted" | sed 's/^/ /'
} >&2
fi
TOTAL_RC=1
if [ "$shard_external_kill" = "1" ]; then
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
echo "shard $i/$N: KILLED externally after ${s_elapsed}s (rc=$rc, well before ${SHARD_TIMEOUT}s cap — queued for serial rescue)" >> "$SUMMARY_FILE"
elif [ "$shard_oom" = "1" ]; then
# Wedged UNDER memory pressure: we can't attribute failures, so queue
# the shard's entire file list for the serial rescue pass.
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
echo "shard $i/$N: WEDGED after ${SHARD_TIMEOUT}s (rc=$rc, OOM signature — queued for serial rescue)" >> "$SUMMARY_FILE"
else
NON_OOM_FAIL=1
echo "shard $i/$N: WEDGED after ${SHARD_TIMEOUT}s (rc=$rc)" >> "$SUMMARY_FILE"
fi
{
echo "--- shard $i: WEDGED after ${SHARD_TIMEOUT}s ---"
[ -f "$SHARD_LOG" ] && tail -50 "$SHARD_LOG"
echo ""
} >> "$FAILURES_LOG"
continue
fi
if [ "$rc" != "0" ]; then
if [ "$shard_oom" = "1" ]; then
# One scan, reused for both the queue append and the emptiness check.
shard_failing_files=$(failing_files_in_log "$SHARD_LOG")
if [ -n "$shard_failing_files" ]; then
printf '%s\n' "$shard_failing_files" >> "$OOM_RESCUE_LIST"
else
# OOM signature but no attributable files (e.g. bun died before any
# file header) → rescue the whole shard.
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
fi
elif [ "$shard_external_kill" = "1" ]; then
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null >> "$OOM_RESCUE_LIST"
echo "shard $i/$N: KILLED externally after ${s_elapsed}s (rc=$rc — queued for serial rescue)" >> "$SUMMARY_FILE"
else
NON_OOM_FAIL=1
fi
fi
echo "shard $i/$N: pass=$pass_count fail=$fail_count skip=$skip_count rc=$rc" >> "$SUMMARY_FILE"
if [ "$rc" != "0" ]; then
TOTAL_RC=1
if [ "$fail_count" -gt 0 ] && [ -f "$SHARD_LOG" ]; then
# Extract each (fail) block: from `(fail)` line through next `(pass)`,
# `(skip)`, blank line, or `__bun_test_summary__`. Single awk pass.
awk -v shard="$i" '
/^\(fail\) / { in_block=1; print "--- shard " shard ": " $0; next }
in_block {
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
print $0
}
' "$SHARD_LOG" >> "$FAILURES_LOG"
elif [ -f "$SHARD_LOG" ]; then
# Non-zero rc but no (fail) line found — extraction couldn't pinpoint.
# Dump the full shard log so we never silently lose the failure cause.
{
echo "--- shard $i: rc=$rc, no (fail) markers — full log follows ---"
cat "$SHARD_LOG"
echo ""
} >> "$FAILURES_LOG"
fi
fi
done
# ──────────────────────────────────────────────────────────────────────────
# Print each shard's full output to stdout (developer expects to scroll
# through it). Print summary file last for one-glance overview.
# ──────────────────────────────────────────────────────────────────────────
for i in $(seq 1 "$N"); do
SHARD_LOG="$LOG_DIR/shard-$i.log"
echo ""
echo "════════════ shard $i/$N ════════════"
[ -f "$SHARD_LOG" ] && cat "$SHARD_LOG"
done
echo ""
echo "════════════ summary ════════════"
cat "$SUMMARY_FILE"
echo ""
# ──────────────────────────────────────────────────────────────────────────
# Serial pass: any *.serial.test.ts files run after parallel pass.
# ──────────────────────────────────────────────────────────────────────────
SERIAL_RC=0
SERIAL_FILES_COUNT=0
SERIAL_FILES_COUNT=$(find test -name '*.serial.test.ts' -not -path 'test/e2e/*' 2>/dev/null | wc -l | tr -d ' ')
if [ "$SERIAL_FILES_COUNT" -gt 0 ]; then
echo "════════════ serial pass ($SERIAL_FILES_COUNT files) ════════════"
bash scripts/run-serial-tests.sh > "$LOG_DIR/serial.log" 2>&1
SERIAL_RC=$?
cat "$LOG_DIR/serial.log"
if [ "$SERIAL_RC" != "0" ]; then
TOTAL_RC=1
if [ "${GBRAIN_TEST_NO_OOM_FALLBACK:-0}" != "1" ] \
&& { grep -qE "$OOM_RE" "$LOG_DIR/serial.log" || [ "$EXTERNAL_KILL_ANY" = "1" ]; }; then
# Serial failures are rescue-eligible ONLY with their own OOM signature
# or when an externally-killed shard ran earlier in this invocation
# (killed-midrun shards leave lock/state residue that poisons the serial
# pass). A merely-OOM'd sibling shard is NOT grounds — a flaky serial
# test must stay red rather than get silently absolved.
failing_files_in_log "$LOG_DIR/serial.log" >> "$OOM_RESCUE_LIST"
else
NON_OOM_FAIL=1
fi
s_fail=$(bun_summary_count "fail" "$LOG_DIR/serial.log")
TOTAL_FAILURES=$((TOTAL_FAILURES + s_fail))
if [ "$s_fail" -gt 0 ]; then
awk '
/^\(fail\) / { in_block=1; print "--- shard serial: " $0; next }
in_block {
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
print $0
}
' "$LOG_DIR/serial.log" >> "$FAILURES_LOG"
else
{
echo "--- shard serial: rc=$SERIAL_RC, no (fail) markers — full log follows ---"
cat "$LOG_DIR/serial.log"
echo ""
} >> "$FAILURES_LOG"
fi
echo "serial: rc=$SERIAL_RC fail=$s_fail" >> "$SUMMARY_FILE"
else
s_pass=$(bun_summary_count "pass" "$LOG_DIR/serial.log")
TOTAL_PASS=$((TOTAL_PASS + s_pass))
echo "serial: pass=$s_pass rc=0" >> "$SUMMARY_FILE"
fi
fi
# ──────────────────────────────────────────────────────────────────────────
# Layer 2: serial OOM rescue. Re-run every file that failed inside an
# OOM-signature shard, one at a time (1 shard, --max-concurrency 1), after
# the parallel fan-out has fully drained. Phantom failures (the WASM ran out
# of memory because 16 instances were up at once) pass here and the run goes
# green with an oom_rescued note; real failures fail again and stay red.
# ──────────────────────────────────────────────────────────────────────────
OOM_RESCUED=0
OOM_RESCUE_NOTE=""
sort -u "$OOM_RESCUE_LIST" -o "$OOM_RESCUE_LIST" 2>/dev/null
# grep -c exits 1 on zero matches — assign in two steps so an empty rescue
# list yields a single "0" (the grep_count double-output bug, same class).
RESCUE_COUNT=$(grep -c . "$OOM_RESCUE_LIST" 2>/dev/null) || RESCUE_COUNT=0
if [ "$TOTAL_RC" != "0" ] && [ "${RESCUE_COUNT:-0}" -gt 0 ]; then
echo "════════════ OOM rescue pass ($RESCUE_COUNT files, serial) ════════════"
echo "[unit-parallel] OOM signature detected — re-running $RESCUE_COUNT failing file(s) at --max-concurrency 1" >&2
RESCUE_LOG="$LOG_DIR/oom-rescue.log"
# 60s-per-file floor with the shard cap as a minimum, and 2x the shard cap
# as a CEILING: a wedged shard queueing its whole file list must not turn
# `bun run test` into an unbounded multi-hour serial re-run — hitting the
# ceiling reads as a red rescue, not silence.
RESCUE_TIMEOUT=$((RESCUE_COUNT * 60))
[ "$RESCUE_TIMEOUT" -lt "$SHARD_TIMEOUT" ] && RESCUE_TIMEOUT="$SHARD_TIMEOUT"
[ "$RESCUE_TIMEOUT" -gt $((SHARD_TIMEOUT * 2)) ] && RESCUE_TIMEOUT=$((SHARD_TIMEOUT * 2))
# Split the queue: *.serial.test.ts files require one bun PROCESS per file
# (run-serial-tests.sh's isolation contract — top-level mock.module leaks
# across files in a shared registry); the remainder batches in one process.
# Both lanes mirror the shard invocation's --timeout=60000 — bun's default
# 5s per-test timeout would re-fail PGLite phantoms (120-migration replay)
# and mislabel them 'confirmed real'.
grep -v '\.serial\.test\.ts$' "$OOM_RESCUE_LIST" > "$LOG_DIR/oom-rescue-batch.txt" || true
grep '\.serial\.test\.ts$' "$OOM_RESCUE_LIST" > "$LOG_DIR/oom-rescue-serial.txt" || true
RESCUE_RC=0
: > "$RESCUE_LOG"
run_rescue() { # $1 = per-invocation timeout seconds; rest = test-file args
local t="$1"; shift
if [ -n "$TIMEOUT_BIN" ]; then
"$TIMEOUT_BIN" --signal=TERM --kill-after="${SHARD_KILL_AFTER}s" "${t}s" \
bun test --max-concurrency 1 --timeout=60000 "$@" >> "$RESCUE_LOG" 2>&1
else
bun test --max-concurrency 1 --timeout=60000 "$@" >> "$RESCUE_LOG" 2>&1
fi
}
if [ -s "$LOG_DIR/oom-rescue-batch.txt" ]; then
# shellcheck disable=SC2046
run_rescue "$RESCUE_TIMEOUT" $(cat "$LOG_DIR/oom-rescue-batch.txt") || RESCUE_RC=1
fi
if [ -s "$LOG_DIR/oom-rescue-serial.txt" ]; then
while IFS= read -r serial_file; do
[ -n "$serial_file" ] || continue
run_rescue 300 "$serial_file" || RESCUE_RC=1
done < "$LOG_DIR/oom-rescue-serial.txt"
fi
cat "$RESCUE_LOG"
r_pass=$(bun_summary_count "pass" "$RESCUE_LOG")
r_fail=$(bun_summary_count "fail" "$RESCUE_LOG")
if [ "$RESCUE_RC" = "0" ] && [ "$NON_OOM_FAIL" = "0" ]; then
# Every failure in the run was OOM-phantom and every rescued file passed
# serially: the run is green. Adjust the headline numbers so they reflect
# the rescue verdict, and mark the earlier failure blocks superseded.
TOTAL_RC=0
OOM_RESCUED=1
# Do NOT fold r_pass into TOTAL_PASS — the failing shard's own summary
# already counted the rescued files' passing tests, so folding would
# double-count. Rescue results ride in the note instead.
TOTAL_FAILURES=0
OOM_RESCUE_NOTE=" | oom_rescued=${RESCUE_COUNT}files(${r_pass}p serial)"
{
echo "--- OOM rescue: all $RESCUE_COUNT file(s) passed serially (${r_pass} tests) ---"
echo "--- failure blocks above were WASM out-of-memory phantoms, superseded ---"
} >> "$FAILURES_LOG"
echo "oom-rescue: $RESCUE_COUNT files pass=$r_pass rc=0 (phantom OOM failures superseded)" >> "$SUMMARY_FILE"
else
# Real failures confirmed serially (or a non-OOM failure exists anyway).
OOM_RESCUE_NOTE=" | oom_rescue_failed=${r_fail}real"
awk '
/^\(fail\) / { in_block=1; print "--- oom-rescue (serial, confirmed real): " $0; next }
in_block {
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
print $0
}
' "$RESCUE_LOG" >> "$FAILURES_LOG"
echo "oom-rescue: $RESCUE_COUNT files pass=$r_pass fail=$r_fail rc=$RESCUE_RC (real failures confirmed)" >> "$SUMMARY_FILE"
fi
fi
END_TS=$(date +%s)
ELAPSED=$((END_TS - START_TS))
# ──────────────────────────────────────────────────────────────────────────
# Loud banner if anything failed. To stderr so it survives `| head`/`| tail`.
# ──────────────────────────────────────────────────────────────────────────
if [ "$TOTAL_RC" != "0" ]; then
ABS_FAIL=$(cd "$(dirname "$FAILURES_LOG")" && pwd)/$(basename "$FAILURES_LOG")
{
echo ""
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
echo "❌ $TOTAL_FAILURES TEST FAILURES — full details:"
echo " $ABS_FAIL"
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
tail -30 "$FAILURES_LOG"
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
echo "[unit-parallel] elapsed=${ELAPSED}s | pass=$TOTAL_PASS fail=$TOTAL_FAILURES skip=$TOTAL_SKIP${OOM_RESCUE_NOTE}"
} >&2
exit 1
fi
echo "[unit-parallel] elapsed=${ELAPSED}s | pass=$TOTAL_PASS fail=$TOTAL_FAILURES skip=$TOTAL_SKIP${OOM_RESCUE_NOTE}" >&2
exit 0