Compare commits

...
Author SHA1 Message Date
Garry TanandClaude Fable 5 8bf23abf71 v0.46.6.0 fix(minions): verify-before-evict lock renewal + per-job leases (#4145) (#4170)
* minions: classify lock-renewal failure causes + starvation telemetry (#4145)

The incident's forensics cost ~8h because the eviction log lines could not
say WHY renewal failed. This commit makes every renewal fault self-explaining
without changing abort semantics:

- Named RenewalCallTimeoutError so cause classification (call-timeout vs
  refused vs fenced-lost) is name-based, never message-sniffing.
- Tick lateness (now - lastTickFiredAt - intervalMs) as the primary
  local-starvation signal — interval callbacks coalesce under a blocked
  loop, so a missed-tick counter cannot measure starvation; lateness can.
  overlap_skips counts tickInFlight re-entrancy skips only.
- Elapsed-time arithmetic now binds deps.now to performance.now()
  (monotonic): wall-clock jumps can no longer distort the deadline math.
  Date.now remains only in log/audit timestamps.
- loadSnapshot dep (raw loadavg[0] + cached core count), try/caught at
  every call site — telemetry must never re-open the unhandledRejection
  class this module exists to close.
- Worker-level event-loop-delay histogram (perf_hooks.monitorEventLoopDelay,
  fail-open when the runtime lacks it), RESET on every successful renewal
  so an eviction-time sample attributes to the exact window in which
  renewal was failing; sampled into the abort + grace-evict log lines.
- Per-launch abortMeta stash so the grace-evict line (which fires 30s
  after the abort) reports cause/lateness/load instead of just the Error
  string that made healthy evictions read like orphan leaks.
- Audit events gain additive optional fields (cause, lateness_ms,
  overlap_skips, load1, cores, via, deadline_deferred) behind a
  back-compatible optional trailing ctx param; the 4-outcome contract and
  pre-upgrade JSONL readback are unchanged (pinned by new tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* minions: inFlight generation-safety — lockToken-conditional deletes (#4145)

Force-evict and the handler's finally both deleted the inFlight entry by
bare job.id. When a force-evicted job is requeued and re-claimed by the
SAME worker while the old handler is still alive, the old execution's
late delete removed the NEW execution's entry — concurrency undercount
and lost tracking for the replacement run. Every claim mints a unique
lockToken and the entry already stores it, so the token is the
generation: both deletes now only remove the entry when it is still
their own.

Pre-existing bug surfaced by the #4145 outside-voice review (R2-1);
fixed in its own commit because the eviction path is exactly what this
wave modifies. Pinned by a deterministic same-worker reclaim test
(stale execution A's finally leaves execution B's entry intact).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* minions: stall-sweep reclaim grace (#4145)

handleStalled reclaimed on a bare lock_until < now(). When a CPU-starved
worker's event loop unblocks, its coalesced renewal tick and the stall
sweep fire in the same burst — if the sweep's UPDATE lands first it
steals the OWNER'S live job and discards its in-flight work. All three
sweep predicates now carry a reclaim grace (default 15s, env
GBRAIN_MINION_STALL_RECLAIM_GRACE_MS, 0 = exact legacy behavior).

The grace is a head-start for the owner's recovery renewal, not a
guarantee: it covers starvation bursts shorter than the grace; a healthy
second worker's sweep still wins beyond it. Cost: dead-worker recovery
becomes lock_until + grace + up to stalledInterval. This is the minion
analog of the cycle-lock steal grace, adapted because minion_jobs has no
last_refreshed_at column.

Existing stall tests move their synthetic lock_until offsets from 1s to
30s past (they pin stall mechanics on the production default path); new
tests pin within-grace hold, beyond-grace reclaim, grace=0 legacy, and
env resolution incl. warn-once fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* minions: verify-before-evict + hard eviction deadline (#4145)

The root-cause fix. The only abort path was the throw branch's local
arithmetic: one timed-out renewal on a starvation-delayed tick past
lockDuration - safetyMargin evicted a healthy job — under load the
renewal UPDATE may even have LANDED server-side while the local race
timeout won. 2,571 subagent jobs submitted / 24 done in 24h.

New contract (ports the cycle-lock fencing doctrine to Minion job locks):

- Fenced-false is the only CERTAIN eviction signal; a throw is not
  evidence of loss. When the NEXT tick would land past the soft deadline
  (cadence-aware: sinceLastSuccess + intervalMs >= deadline — the bare
  >= gate is unreachable under cadence quantization for long leases),
  the tick runs ONE bounded VERIFY renewal. renewLock fences on
  lock_token and deliberately ignores lock_until, so an
  expired-but-unstolen lease revives: fenced-true → starved-but-ours,
  keep working (the incident-saving path); fenced-false → certain loss,
  abort (stall detector requeues, no attempt burned); verify unreachable
  → defer + reconnect-once, aborting only past the hardEvictMs backstop
  (default 2×lease, env GBRAIN_LOCK_RENEWAL_HARD_EVICT_MS, floored to
  the soft deadline) — a LOCAL decision under uncertainty that bounds
  blind external side effects during a total outage.
- The verify is a synchronous, cancelled()-guarded, callTimeoutMs-bounded
  call inside the tick's own flow — reconciled in-code with queue.ts's
  no-background-retry rationale (both UPDATEs are same-token idempotent
  lease extensions; a fenced row cannot gain two holders).
- Best-effort cancellation: the race timeout now aborts the in-flight
  renewLock via AbortSignal threaded to executeRawDirect; both engines'
  entry points gained an already-aborted preflight BEFORE dispatch/pool
  acquisition. The fence stays the correctness authority.
- Relational knob validation: margin < lease/2, callTimeout <= cadence,
  hardEvict >= soft deadline — clamped with warn-once; positive-integer
  parsing alone could silently re-break the deadline math.
- Doctrine comments rewritten (tick header + worker grace-evict) — the
  old abort-at-deadline prose actively misled.

Tests: incident replay (starved renewal times out, verify succeeds, job
survives — the exact #4145 shape), fenced-false via verify, deferral +
hard-backstop timelines, CDX-4 quantization pin (300s/60s verifies at
240s with lease left), first-tick verify at the 30s default,
mid-verify cancellation, reconnect-on-deferral, relational clamps, and
a DATABASE_URL-gated e2e pinning the DB-level foundations (expired-
but-unstolen revival, grace hold, post-reclaim fenced-false, and a
blocked-event-loop worker completing with zero stall bounces).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* minions: per-job lock_duration_ms end-to-end (#4145)

A single worker-global 30s lockDuration cannot serve both 2s shell jobs
and 173s-average LLM subagent jobs — the structural half of the #4145
incident. The lease is now a per-job column resolved through the same
three-layer model as timeout_ms (explicit submit → handler-type map →
worker default), claim-stamped so it survives worker restarts.

- Migration v129 + the 3 schema copies: nullable lock_duration_ms with
  the positive CHECK added via the idempotent drop-then-add pattern (v7
  precedent) — no fresh-vs-migrated asymmetry, no backfill (NULL = worker
  default = pre-#4145 behavior; claim COALESCE owns all defaulting).
- HANDLER_DEFAULT_LOCK_DURATION_MS beside the timeout map (long LLM
  handlers 300s, single-call LLM handlers 120s, shell deliberately absent
  for fast dead-worker reclaim) under one rewritten header explaining why
  lease and wall-clock budget are different quantities.
- Claim derives lock_until from COALESCE(row, map, worker default) and
  stamps the resolved lease; the map binds as a RAW object (jsonb
  double-encode rule). Wall-clock null-fallback becomes
  COALESCE(lock_duration_ms, worker default) so an explicit lease on an
  unmapped handler isn't killed at the old bound (NULL rows pinned to
  exact legacy behavior).
- Worker consumes the effective per-job lease at all three launch sites;
  renewal cadence clamps to min(lease/2, 60s) — a 300s lease renews 5x
  per window instead of every 150s; ≤120s leases keep legacy /2 exactly.
  Derived knob defaults cap at 15s call-timeout / 30s margin so long
  leases don't inherit wedge-inducing values.
- Shared clampLockDurationMs [5s, 1h] used by queue.add, the new
  gbrain jobs submit --lock-duration-ms flag, and the MCP submit_job
  param (handler-side clamp; ParamDef has no min/max — wrong types are
  rejected by the existing number validation). INSERT-only on idempotent
  re-submit, matching the max_stalled footgun rule.
- jobs get shows the lease line alongside the timeout line.

Tests: migration v129 structure/idempotency/CHECK/pre-shape re-run;
claim-stamps-from-map + lock_until horizon; explicit-wins + clamp bounds;
idempotent-resubmit immutability; wall-clock NULL-row regression pin +
leased-row survival; lifetime max_stalled accumulation pin (the known
coverage gap); per-lease knob derivation caps.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(minions): lock-renewal knobs, eviction-forensics guide, verify-before-evict current-state (#4145)

- queue-operations-runbook.md gains the incident-reading section the
  #4145 forensics lacked: how to read a gave_up/eviction line (cause,
  lateness_ms, load1/cores, overlap_skips, deadline_deferred,
  event-loop-delay), the was-the-DB-down-or-the-worker-starved decision
  table, the full env-knob table (CALL_TIMEOUT / SAFETY_MARGIN /
  HARD_EVICT / MAX_FAILURES / STALL_RECLAIM_GRACE) with relational-clamp
  semantics and legacy escape hatches, and the honest zombie caveat
  (eviction is cooperative until the kill/reap follow-up lands).
- minions-deployment.md's 'what can still bite' section rewritten for
  verify-before-evict + per-type leases + reclaim grace; documents the
  mixed-version-fleet degradation (old workers = legacy behavior, no
  drain needed) and adds --lock-duration-ms to the per-job tuning list.
- KEY_FILES.md entries (lock-renewal-tick.ts ×2 duplicated entries,
  worker.ts, queue.ts, handler-timeouts.ts) updated to current state.
- TODOS.md: filed the kill/reap-evicted-handler follow-up (P2), the
  worker-level --lock-duration flag (P3), and the TODO-LR-2 note that
  its doctor-check inputs now exist in the audit events.
- llms bundles regenerated (unchanged — these guides aren't inlined).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: scrub leaked GBRAIN_HOME from doctor-minions-check subprocess env

The test seeds $HOME/.gbrain/migrations fixtures and spawns gbrain doctor
with HOME overridden — but doctor resolves its home via resolveGbrainHome,
which prefers GBRAIN_HOME over HOME. Sibling test files in the same bun
process (preferences, friction, bootstrap-*, and several .serial files'
beforeEach hooks) set process.env.GBRAIN_HOME; a value captured by the
{...process.env} spread makes the fixture invisible, doctor finds nothing,
and the expected FAIL exit code never happens. Scrub GBRAIN_HOME exactly
like the DATABASE_URL variables already scrubbed two lines up.

Surfaced while triaging a non-blessed whole-suite invocation during the
#4145 wave (reproduced identically on master); the blessed sharded runner
can also co-schedule a GBRAIN_HOME-mutating file into this shard, so the
scrub closes a real flake vector, not just a synthetic one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): verify probe cleanup hard-deletes instead of leaving soft-delete tombstones

verify's end-of-run probe cleanup invoked the delete_page OP, which since
v0.26.5 is a SOFT delete — every verify run left two probe tombstone rows
in the user's brain (visible to include_deleted readers) until the 72h
purge. Cleanup now hard-deletes via engine.deletePage(slug, {sourceId}),
the same primitive sweepProbeLeftovers already uses on both engines, with
the warning-capture semantics preserved. Pinned by the (previously
failing) probe-residue assertion in the Postgres bootstrap-verify e2e.

Surfaced by the e2e fix wave: CI runs only 6 of 185 test/e2e files
(.github/workflows/e2e.yml), so the developer-lane files rot silently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(search): deterministic relationalFanout path pick on equal-depth multi-seed ties

The fanout's representative-path pick ordered only by (depth, path
length); a node reachable at the same depth from multiple seeds had NO
tie-break, so the winner was plan/heap-order dependent — a fresh PGLite
and a lived-in Postgres heap could disagree, violating both engine parity
and the documented deterministic relational-retrieval contract. Appended
a lexicographic final tie-break to the array_agg ORDER BY in BOTH engines
(lockstep). The engine-parity e2e's multi-seed fanout case now passes on
real Postgres; its stale-page arm also moves off client wall-clock stamps
onto per-row updated_at_iso (the #1768 production semantics) so VM clock
drift can't skew the count.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e): per-file outer-timeout override for LLM-bound Tier-2 files

run-e2e.sh's hard 180s-per-file gtimeout SIGKILLed skills.test.ts mid-run
(the ingest skill alone has been observed at ~131s of real provider
round-trips), producing a mystery failure with no assertion output. CI
runs the Tier-2 keyed files in their own job WITHOUT this wrapper, so the
cap only ever bit local runs. The cap is now GBRAIN_E2E_FILE_TIMEOUT
(default 180) with 4x for skills.test.ts + zeroentropy-live.test.ts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e): un-rot the developer-lane files — 27 failures across 10 files, all green on both lanes

CI runs only 6 of the 185 test/e2e files; the rest run solely through the
developer-machine lane (bun run test:e2e / ci:local) and had rotted as src
moved deliberately underneath them. Every fix pins CURRENT intended
behavior with the causing commit cited in-file; no assertion was weakened
and several were strengthened. Root causes:

- sync.test.ts (13): the #2114 global-anchor ownership guard (636628fdb)
  refuses anchor writes when the default source's local_path names another
  repo; setupDB truncates config but not sources, so residue from earlier
  files vetoed every bookmark write. The test now resets the default
  source identity in beforeAll.
- v0_29-mcp-dispatch (2): the #4096 WP1/D7 locality backstop dispatches
  localOnly ops only on transport 'stdio'; tests now dispatch with the
  real stdio shape, plus a NEW fail-closed unknown_tool pin for unset
  transport markers across both trust values.
- extract-atoms-discovery-sql (4): PR #2615 widened discovery to the
  schema pack's extractable:true types ('note' included); tests re-pin
  with 'person' as the non-extractable control + wider seed cleanup.
- pglite-cli-exit (1) + bootstrap-harness-lifecycle (1): the wrapper's
  ambient DATABASE_URL leaked into subprocess/in-process env, silently
  retargeting PGLite tests onto shared Postgres (#801 env-override
  precedence); both files now scrub it at their boundaries.
- embedding-column-pglite (1): #3554 changed resetGateway() to restore
  the test-preload baseline; the #3461 fallback test now uses
  __unconfigureGatewayForTests() so the unconfigured path really fires.
- openclaw-plugin-load-real (1): the Retrieval Reflex import chain pulls
  PGLite WASM assets into the bundle — bun build --outfile cannot emit
  multi-output builds; switched to --outdir with entry naming + a
  version-robust runtime inspection helper.
- phantom-redirect (1): halfvec migration (v40) + the #2932 idempotent
  reconcile; the string-shape guard now accepts both legitimate vector
  types.
- serve-http-oauth (1): sql.array() on a fresh connection races the
  async typeArrayMap fetch and binds text instead of text[]; seed uses a
  plain-array bind (the same untyped approach production pgArray() uses).
- type-unification-full-flow (1): checkPackUpgradeAvailable reads the
  operator's real ~/.gbrain config; the test now isolates GBRAIN_HOME and
  pins both the warn and ok arms hermetically.

Verified: full e2e lane 186/186 files, 1279/1279 tests on a pristine
pgvector container; the 14 previously-failing files re-verified green on
the residue-carrying locally-configured database as well.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(minions): ship-review hardening — histogram lifecycle, cadence single-home, consumption re-clamp, DRY telemetry (#4145)

Fix-First batch from /ship's specialist + coverage reviews (21 findings,
all informational; the mechanical ones applied, the rest pinned by tests):

- The event-loop-delay histogram now disables in stop() and (re-)enables
  in start() — each cycled worker instance previously leaked a ~50Hz
  native sampling timer for process lifetime (embedding hosts and the
  test suite cycle workers constantly).
- renewalIntervalFor()/RENEWAL_INTERVAL_CAP_MS: ONE home for the
  min(lease/2, 60s) cadence formula, used by the worker's timer AND as
  resolveLockRenewalKnobs' default intervalMs — previously the knob
  default (lease/2 uncapped) diverged from production cadence for leases
  over 120s, silently weakening the CDX-10 relational validation for
  callers that omit intervalMs.
- Defense-in-depth re-clamp at consumption: launchJob clamps a row
  lease through clampLockDurationMs — the exposed submit surfaces clamp,
  but a writer bypassing add() could stamp a 1ms lease (renewal-storm
  interval) or a ~25-day one (weeks-long dead-worker pin).
- formatAbortMeta(): one formatter for the classified abort telemetry;
  the three log sites had already drifted (load1: vs load1_at_abort:).
- Knob docstrings updated to the capped defaults (min(lease/3, 15s) /
  min(lease/6, 30s)); KEY_FILES dedup + stale-count scrub; dead test
  binding removed; run-e2e.sh validates GBRAIN_E2E_FILE_TIMEOUT
  digits-only before arithmetic/interpolation.

New pins: worker renews with the per-job lease (launchJob wiring, GAP-1);
claim precedence row-beats-map at the claim UPDATE; MCP submit_job
clamp round-trip incl. the 0→default boundary; jobs get lease lines
(all three states); bootstrap-verify tombstone-proof pages count on
PGLite; relationalFanout lexicographic WINNER (not just parity);
grace-env warn-once dedupe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(minions): adversarial-review hardening — claim-time clamp, grace cap, range CHECK, honest dry-run echo (#4145)

Two independent adversarial passes (Codex + Claude) converged on the same
lease-bounds gaps; this batch closes them:

- Claim SQL clamps the resolved row/map lease to [5s,1h] before deriving
  lock_until, so a bypass-written out-of-range value can't produce a
  pathological lease at claim time. The worker-default path ($2) passes
  through untouched (tests pin a 1ms worker lease).
- Migration v130 + all 3 schema copies upgrade the CHECK to a full range
  constraint (>= 5000 AND <= 3600000) — the DB bound now matches the
  app-side clamp instead of only enforcing positivity.
- GBRAIN_MINION_STALL_RECLAIM_GRACE_MS is capped at 600s with a warn-once
  (an absurd env value silently disabled stall reclaim fleet-wide).
- Hard-evict comparison recomputes elapsed AFTER the bounded verify, so
  the backstop can't defer one extra cadence past its advertised bound;
  the should_abort payload carries the recomputed value.
- Verify-failure telemetry re-samples loadavg at its own failure instead
  of reusing a snapshot stale by the verify's duration; audit note that
  the attempt counter advances by 2 per at-deadline tick.
- racedRenewLock/attemptReconnectOnce clear the losing race timer on the
  win path (no stray late abort against a settled query).
- jobs submit --dry-run echoes the CLAMPED lease (annotated when it
  differs from the raw input) instead of echoing a value add() won't store.

Tests: migrations-v130 range-CHECK rejection loop + bypass-clamp claim
test; grace-cap pin; dry-run echo cases. 301 pass / 0 fail targeted;
typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v0.46.6.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: document-release sweep for v0.46.6.0 — lease-bound enforcement, grace cap, e2e file timeout

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): stop run-child-entry SIGTERM test from broadcasting to leaked process listeners

The SIGTERM-semantics test fired a bare process.emit('SIGTERM') in the
shared bun test process. When an earlier file in the same shard had
installed process-cleanup.ts's signal handlers (module-global, never
uninstalled), the broadcast reached its handler, whose cleanup pass ends
in process.exit(143) — killing the entire shard mid-suite. The runner
then misclassified rc=143 as an external kill ("sibling workspace pkill /
memory jetsam") and queued a rescue pass that died the same way when it
reached the same test. Whether it fired depended on file interleaving;
under heavy host load the schedule made it deterministic (three
consecutive suite runs died at the ~3-minute mark with zero test
failures).

The test now snapshots the SIGTERM listener set before invoking
runChildJobEntry and fires ONLY the listener(s) the entry registered —
same wiring under test, no global broadcast.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: changelog entry for the unit-suite shard self-kill fix

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): best-effort tmpdir cleanup in run-child-entry afterAll — EFAULT rmSync flake reds the CI shard

CI shard 3 failed with an "(unnamed)" test: bun treats a throwing afterAll
as a failed test, and the hook's recursive rmSync EFAULT'd (bun 1.3.13,
ubuntu-24.04) immediately after the PGLite WASM engine teardown. All 8
real tests passed. Cleanup is now try/retry/warn — a tmpdir the OS reaps
anyway must never red the suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 22:06:05 -07:00
Garry TanandClaude Fable 5 11ad23e346 v0.46.5.0 perf(test,ci,eval): CI in half — pooled serial lane, snapshot-in-CI, hermetic retrieval canary (#4154)
* perf(pglite): memoize snapshot loading — read the 42MB tar once per process, not per engine

tryLoadSnapshot re-read the snapshot tar and re-hashed all migration handler
sources on every engine construction (600+ per full suite, ~84MB transient
allocation each). The schema hash and the (versionLines, blob) pair are now
memoized per (path, process); missing/stale/torn paths memoize a terminal
null. The dims/model shape gate stays per-call so mid-process gateway
reconfiguration (the zembed/1280 class) still falls back to cold init —
pinned by new memo tests in snapshot-shape-guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(test): pool the serial-test runner + wire the PGLite snapshot into every CI-facing runner

The serial lane ran 140 per-file bun processes strictly one-at-a-time (8.5
min in CI) — but the quarantine contract only requires per-PROCESS isolation.
run-serial-tests.sh now runs a pool (min(cpus,4), memory-adaptive, 120s
per-test budget, 300s SIGTERM-then-SIGKILL wall clock per file), keeps two
machine-global files on a sequential EXCLUSIVE lane (launchd/cron), treats a
missing exit sentinel as failure, and prints sorted per-file durations.
First full run: 140/140 in 145s at pool=4.

The snapshot fast-path (previously local-only) is now shared via
scripts/lib/test-env.sh (detect_cpus + mem detection + ensure_pglite_snapshot,
one implementation across the runner family) and wired into test-shard.sh
(+ --max-concurrency, mirroring the local runner), run-slow-tests.sh, and
run-e2e.sh's env-scrub keep-list. Serial lane also gains the #3485 ambient
DB-URL scrub its sibling lanes already had.

Guards: pool-behavior tests (fail → exit 1, full log; hang → timeout kill;
dry-run-list), EXCLUSIVE_FILES growth guard (≤3, justification comments),
missing-sentinel source pin, sandbox staging carries scripts/lib/.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(test): snapshot tar cache, one bun-cache saver, gitleaks tarball cache, shallow brainbench fetch

- PGLite snapshot (~42MB tar) cached across jobs keyed on schema inputs:
  serial-tests saves, the 10-shard matrix + slow jobs restore-only; the
  runner's own hash check stays authoritative (stale restores are rebuilt,
  never trusted). Slow jobs export the snapshot env via the shared lib.
- Bun install cache: verify becomes the single saver (admin/bun.lock joins
  its key); every other job restores with restore-keys so a bun.lock touch
  no longer cold-installs all six jobs. All installs are --frozen-lockfile.
- gitleaks: the release tarball is cached; its published checksum is
  fetched fresh and re-verified on every run, so a cache restore is never
  trusted.
- brainbench: fetch-depth 1 + a depth-1 fetch of the master ref replaces
  the full-history clone (the gate reads one file via git show).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(e2e,heavy,release,osv): parallel e2e tiers behind a spend gate, content-hash skip, cache/timeout hygiene

e2e.yml: tier2 no longer waits ~2min behind tier1 (separate DBs — never
shared state); it now gates on jsonb-parity (~40s), keeping a fast
broken-build spend gate in front of the job that burns real provider
tokens. New e2e-cache-check/e2e-cache-write (e2e-pass-<hash> namespace)
skip the whole suite on doc-only pushes; scheduled nightly runs are
exempt so the live-provider drift check always fires. e2e-status is the
stable aggregate name. Bun caches + --frozen-lockfile on all tiers.

heavy-tests: bun cache restores on all 4 jobs. release: timeout-minutes
on all 4 jobs (was 360-min default), bun caches, frozen installs.
osv-scanner: concurrency group cancels superseded PR scans.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(verify): bounded worker pool, longest-first ordering, chronicle gate, two revived guards

run-verify-parallel.sh fanned out 44 checks unbounded — two cp -R src +
bun build --compile builds, the admin vite build, tsc, and ~40 greps all
simultaneous on a 4-vCPU runner, feeding the 120s per-check timeout flake
class. The spawn loop is now a pool (default detect_cpus; escape hatch
GBRAIN_VERIFY_MAX_PARALLEL) with the heavy checks ordered first
(LPT-style makespan). 47 checks green in 38s locally.

New checks: check:eval-chronicle ($0 deterministic eval, exit-0-only-on-
perfect — first CLI-level CI gate for it), plus the two registered-but-
never-executed guards check:pagetype-exhaustive and check:pg-url-redaction
(the latter's marker now works inside block comments; its one legitimate
hit in the redactor's own docs carries the marker). A new registration⇒
execution coverage test closes the dead-guard class: every guards-manifest
row must be reachable from CHECKS or carry an explicit exemption reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): collect evals/ tests into the matrix behind a keyless allowlist

evals/functional-area-resolver/harness-runner.test.ts (47 keyless tests)
was collected by NO runner — real tests that never executed anywhere.
test-shard.sh now finds test/ AND evals/; a new allowlist guard asserts
every collected evals file is keyless-verified (this repo's eval harnesses
are key-requiring by default, so unlisted growth would silently spend
tokens in CI). The isolation (R1-R4) and real-names lints extend their
scan roots to evals/ so everything CI executes carries the same hygiene bar.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(test): shrink migrate dedup perf-gates to 200 rows; one engine for chunk-grain-fts

The two v8/v9 dedup gates inserted 1000 rows one-at-a-time each — ~15-25s
of row traffic per test that adds no discriminating power (the O(n²) shape
they guard is minutes-vs-sub-second at 200 rows; the full v7→current chain
replay they also pay is unchanged and still exercised).

chunk-grain-fts.test.ts booted three describe-scoped engines for 11 tests;
now one file-level engine + resetPgliteState per data-bearing describe
(the reset is required, not hygiene: the searchKeyword corpus would
pollute the searchKeywordChunks expectations).

The planned resetPgliteState pg_tables caching is deliberately NOT done:
the pre-agreed DDL check found 7 files that create/alter tables between
resets on a shared engine — a cached table list would skip truncating
mid-file tables and leak rows across tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(eval): hermetic CLI retrieval canary — deterministic embedder for eval gate, W1 mandate satisfied

gbrain eval gate gains an embedder option (deterministic) scoped to the
qrels correctness gate: query vectors come from a fixture-keyed basis
embedder (src/eval/deterministic-embed.ts, shared with the hermetic test)
through a new additive HybridSearchOpts.queryEmbedFn seam — the full RRF
pipeline (keyword/FTS, title, alias, relational arms + fusion) runs for
real with zero API keys. Bare hybridSearch is cache-free by construction
(lookup + writeback live only in hybridSearchCached), so deterministic
runs cannot poison query_cache. Behavior is unchanged when the seam is
absent. Flag registry regenerated.

scripts/run-eval-canary.ts seeds a throwaway PGLite brain from the qrels
corpus (expected pages visible to page-grain FTS via timeline — pages
FTS deliberately excludes compiled_truth) and spawns the real CLI:
check mode (CI, writes nothing) is wired as check:eval-canary in verify;
record mode appends the committed .gbrain-evals/eval-results.jsonl ledger.

Measured, deterministic across processes and keyless: recall@10 1.0000,
first_relevant 1.0000, expected_top1 0.8333 vs floors 0.70/0.60/0.50.
The FIX_WAVE_BASELINES W0 retrieval-canary mandate is now PASS (recorded
with honest scope: synthetic vectors gate the ranking pipeline; semantic
embedding regressions remain the keyed suites' job). Spike-first design
per the plan's OV2-1 respec.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(test): current-state TESTING.md for the pooled/pool-bounded runners; TODO ledger updates

TESTING.md: pooled serial lane (+EXCLUSIVE lane, knobs, timings), verify
worker pool + eval gates, snapshot-in-CI, evals/ collection, e2e cache-skip
+ e2e-status. ci-local.sh: stale "36 E2E files" comments (actual 181).
TODOS: deeper-speedup entry closed by this pass; test.concurrent P0
downgraded to P3 with stale-premise rationale; eval-gate baseline entry
narrowed to the sibling-repo regression half; 7 pass deferrals filed with
context (sleep-to-poll, e2e lanes, per-shape snapshots, persistent-engine
snapshot, engine-consolidation audit, verify double-spawn, image-decoders).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): re-mine shard weights post-snapshot; p75 fallback; balance test asserts what CI runs

Weights re-mined from the branch's green Test run (31893516029): 1200/1212
files covered (was 663/1204 — 45% of the corpus rode a 30ms median fallback
while really averaging seconds), total mined suite time 986s vs 3185s
pre-snapshot. Rebalanced 10-shard totals: ~99s each.

sharding.ts missing-file fallback median → p75 (the distribution is
right-skewed and unweighted files skew heavy — new integration tests land
unweighted more often than pure-unit ones).

The balance regression test previously recomputed shard totals with the
same weights map + same fallback the partitioner used (max/min ≈ 1.0 by
construction) and asserted 4/6-shard splits while CI runs 10. It now:
asserts the 10-shard split with the shard count cross-checked against
test.yml's matrix (js-yaml parse, not a format-brittle regex), gates
weights coverage ≥70% of matrix-eligible files (the anti-rot forcing
function the regen cadence never had), and fails on weight keys naming
untracked files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): external-kill rescue pass for the serial pool; brain-repo-durability goes exclusive

Two contention classes surfaced by the first pooled CI runs:

1. Stray SIGTERM/SIGKILL from outside the runner (sibling-workspace process
   cleanup, memory jetsam) killed 1-12s-old bun processes with exit 143 and
   truncated logs — the exact class run-unit-parallel.sh already rescues.
   The serial runner now queues exit 143/137 and missing-sentinel files for
   ONE sequential rescue re-run: phantoms stay green with a rescue note,
   real failures fail again and stay red. Pinned by a self-SIGTERM-once
   fixture test.

2. brain-repo-durability.serial.test.ts: hardenBrainRepo's scaffolding
   commit fires the just-installed post-commit hook (background push) which
   races the synchronous push-probe on the same bare remote — "cannot lock
   ref" lands in needs_attention. Near-deterministic on a contended 4-vCPU
   runner, never observed locally. Moved to the sequential EXCLUSIVE lane
   (third justified entry; growth guard capped at 3) until the probe learns
   to retry ref-lock contention.

Also fixes a duplicated word in the runner's timeout header line.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(test): document the serial pool's external-kill rescue pass

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: pre-landing review fixes

From the ship-stage specialist review (testing/maintainability/security/
performance):
- serial runner: exclusive-lane files keep their no-kill contract on rescue
  re-runs; exit 137 at full duration is classified as our timeout's SIGKILL
  escalation (real failure), not an external kill — no ~315s re-hang in rescue
- snapshot memo: tar read deferred until the first shape-MATCHING caller —
  a process that only ever refuses (zembed/1280 class) now reads zero bytes
- e2e.yml: workflow_dispatch joins schedule in the cache-skip exemption (a
  manual dispatch is an explicit ask for a live run)
- test.yml: verify job restores the snapshot cache like its siblings;
  cache-key homes cross-referenced
- eval gate: the four embedder-flag validation exits are now asserted;
  legacy-qrels parsing deduped into deterministic-embed.ts
- canary test: git-status invariance scoped to touchable paths; outer spawn
  budget strictly above the inner CLI timeout
- doc/comment rot: TESTING.md exclusive-lane count, runner header, sharding
  test comments, ci-local.sh sentence, CI_SHARDS in a test name
- TODOs: snapshot-tar digest verification (P3) + eval-ledger redaction (P2)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): serial pool re-emits a bun-format pass aggregate; ledger gets merge=union

The pooled runner's compressed per-file lines starved run-unit-parallel.sh's
headline counter (awk wants " N pass") — bun run test's pass=N banner
silently dropped the entire serial suite. The runner now emits one
aggregate " N pass" line in bun's own summary format. Failing files' logs
stream raw, so fail counting was never affected and stays single-counted.

.gbrain-evals/eval-results.jsonl (append-only tracked ledger) takes
merge=union so concurrent workspaces recording runs don't conflict.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: bump version and changelog (v0.46.1.0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update project documentation for v0.46.1.0

- eval-bench.md: document the hermetic --embedder deterministic correctness-gate
  mode and the check:eval-canary CI gate (scripts/run-eval-canary.ts, --record)
- KEY_FILES.md: current-state updates — eval-gate hermetic mode +
  deterministic-embed.ts + canary runner in the eval-loop entry, queryEmbedFn
  seam in the hybrid.ts entry, per-process snapshot memoization + the shared
  scripts/lib/test-env.sh helper in the snapshot entry, guard count 45→46
- TESTING.md: snapshot activation now via ensure_pglite_snapshot across five
  runners (was two callers), honest per-file speedup figures (~3.5x), serial
  runner output/knob semantics (GBRAIN_SERIAL_POOL=N width), CI shard
  --max-concurrency bound, p75 missing-weight fallback
- CONTRIBUTING.md: drop the orphaned duplicate guard-checks line
- CHANGELOG.md (voice only): "every CI runner" -> "the CI test runners";
  note the p75 fallback in the shard-rebalance clause

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: re-bump to v0.46.2.0 (0.46.1.0 claimed; user-pinned)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 21:51:52 -07:00
Garry TanandClaude Fable 5 48890e35be v0.46.4.0 feat(opencode): full-parity client support — bootstrap, harness, connect, claw-test, real-binary e2e door (#4162)
* docs(mcp): pin observed opencode CLI behavior (OPENCODE-CLI-PIN.md) + registration guide

Phase-0 hermetic observation of opencode-ai@1.18.18 (npm wrapper + platform
payload integrities pinned). Load-bearing observations: keyless anonymous
free tier answers headless runs AND drives MCP tool calls without --auto
(nonce SMOKE proven end-to-end against a real gbrain serve --surface verbs);
mcp list is the honest discriminator (spawns servers, exit 0 regardless —
parse the text); mcp add takes '-- command' (undocumented in --help) but
always writes user-global opencode.jsonc; project-defined local servers
spawn with NO trust gate (drives the user-global bootstrap default); JSONC
parses in .json-named files and both filenames merge; OPENCODE_CONFIG* env
vars observed inert (docs-contradiction, called out); OPENCODE=1 set in bash
children (detectHarness probe); AGENTS.md loads, CLAUDE.md not double-loaded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): opencode-json managed config writer + opencode-2026-08 host spec

- src/core/bootstrap/opencode-json.ts: comment-preserving JSONC writer
  (jsonc-parser surgical edits — opencode's own mcp add preserves comments,
  the writer matches that bar). Ownership is a 4-state structural
  fingerprint (ours-same-source | ours-other-source | foreign | absent)
  keyed on GBRAIN_SOURCE EQUALITY ([FIX7] parity), never a marker key.
  Distinct read-failure classes (ENOENT create / empty-as-{} / unreadable
  refuse); foreign refusal on write AND remove; post-render validation
  (our entry round-trips, every other key survives) keeps the original on
  failure; 0600 + .bak-0600 only for inline-bearer entries; bearer
  recovery helper for harness --status.
- host-specs.ts: TARGETS['opencode-2026-08'] (verified 2026-08-15 against
  a hermetic opencode-ai@1.18.18) + opencodeConfigDir/GlobalConfigPath/
  ProjectConfigPath (XDG-only — OPENCODE_CONFIG* observed INERT in
  1.18.18, honoring them would be a silent no-op install) +
  OPENCODE_HAS_HOOKS=false.
- atomic-write.ts: rule-of-three extraction of the symlink-resolving,
  mode-inheriting atomic writer; codex-toml.ts + hooks.ts ported onto it
  (behavior pinned by their existing suites).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): opencode workspace lane — hooks --harness opencode, scope-aware direct-writer registration, channels, templates

The Harness union widening is typecheck-SILENT at every existing
'claude-code ? A : B' ternary, so the hot sites now dispatch exhaustively
(HARNESSES satisfies anchor; exec-lane bin map returns null for opencode —
its registrations go through the JSONC writer whose fingerprint IS the
[FIX7] check, never through <host> mcp get).

Scope INVERSION for opencode: default user-global (opencode spawns
project-config-defined MCP servers with NO trust prompt — verified; a
committed project entry would auto-execute on every collaborator machine).
MCP_SCOPE=project is an explicit opt-in that writes the workspace
opencode.json with a PATH-resolved command (committed-candidate file: no
absolute machine paths, no fail-open analog exists) and prints the sharing
warning + enabled:false opt-out. detectHarness probes OPENCODE/OPENCODE_PID
(observed 1.18.18). Ownership: a remote-type mcp.gbrain in the global
config makes the stdio lane step aside (codexBlockOwnsName analog); foreign
entries refuse. Verification: writer post-render parse-back is
authoritative; best-effort 'opencode mcp list --pure' probe (skipped on
plugin-bearing configs — mcp list is a code-execution surface).

Atomic with this commit (each-commit-green): questions.json MCP_SCOPE +
SURFACE_PRIMARY copy, AGENTS/GITHUB template pull-protocol generalization,
BOOTSTRAP_FOR_AGENTS.md scope guidance + opencode wiring bullet,
status.ts interview/wire resume hints, check-bootstrap-templates.sh §(e)
pins (now 'Claude Code and opencode' + the 'NO trust prompt' spawn-gate
rationale pin), guard-test fixtures, the status-test hint pin, the vendored
template-repo regen, the offline docker opencode leg, and uninstall's
receipt-keyed opencode removal. Channels: 'opencode' joins
VOLUNTEER_CHANNELS + HARNESS_CHANNELS (reserved attribution slot, codex
precedent).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): opencode harness-mode target — managed remote entry with inline bearer, remove/status/rollback

HarnessSelector gains 'opencode' (forced-wire like codex: the JSONC writer
needs no opencode CLI). Wiring is one managed mcp.<name> remote entry with
the inline Authorization bearer in the user-global opencode config, 0600,
under the [X11] lock ordering (config-dir → opencode-dir). Ownership [C8]:
idempotent re-runs match on the serve url; rotation across a url change
recognizes the old entry via the PRIOR receipt's url; anything else under
the name refuses inside the writer. Failed-smoke rollback restores the .bak
or removes a fresh entry, and the fresh mint is revoked (impostor-guard
economics hold). --remove classifies against the receipt url and skips
not-ours entries with a note; --status recovers the bearer from the entry
(url-matched — a foreign entry's credential is never transmitted). Consent
copy: per-host numbered item, joined-list reach statement (a fourth harness
can no longer silently mislabel the ternary tree), opencode off-ramp.

Fixture hygiene: the harness serial fixture now injects opencodeConfig +
detectOpencode — the default path resolution reaches the operator's REAL
~/.config/opencode (the claudeUserSettingsPath lesson, caught live when the
registrar-mode test wrote a fixture token there; cleaned up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(connect): --agent opencode — env-interpolated bearer, direct-writer --install

buildOpencodeMcpAddArgv pins the validated one-liner: the --header value
carries opencode's {env:GBRAIN_REMOTE_TOKEN} interpolation LITERALLY, so the
token never enters argv, the config file, or --json output. The print block
mirrors codexBlock (export line + one-liner + restart note). --install goes
through a new ConnectDeps.writeOpencodeRemoteEntry member (the existing
injectable seam, connect.ts:ConnectDeps) wrapping the JSONC writer in env
token mode — no opencode binary required, idempotent re-runs, foreign
same-name entries refuse with the writer's message (token-redacted), and
the D4 probe smoke-tests the credential end to end.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(claw-test): opencode runner — detection, pinned one-shot invoke, multi-provider env allowlist

OpencodeRunner (4th AgentRunner): detectBinary('OPENCODE_BIN','opencode');
argv pinned to 'run <brief> --format default' (explicit format so an
upstream default flip cannot silently change the transcript shape; NO
--auto — MCP tool calls fire in run mode without it, verified). Env
allowlist = BASE + an EXPLICIT multi-provider delta (XAI / Google / Gemini
/ OpenRouter keys — BASE carries only Anthropic+OpenAI, and a live-lane
operator on other providers would otherwise see a misleading auth failure)
+ XDG dirs + OPENCODE_CONFIG(_DIR) + OPENCODE_DISABLE_AUTOUPDATE;
OPENCODE_CONFIG_CONTENT (the inline config-shadow channel) deliberately
absent. Bare-semver version preamble (the SST-vs-claimant discriminator)
+ global-config mcp.gbrain contamination tripwire (JSONC-tolerant, checks
BOTH merged filenames). --list-agents pin moves to all four runners with
the openclaw<opencode ordering note.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(test): door-family extraction in agent-harness — shared resolver/childEnv/spawn core; grok+hermes ported; opencode first consumer

The 'P3 — Door-adapter extraction + CI-tail composite action' TODO armed
this at 'the NEXT door agent (4th)' — opencode is the 4th. Test-side only
(the composite CI-tail action stays deferred until the first green
grok-door AND opencode-door dispatches; workflow yaml can't be proven
locally).

- makeBinaryResolver: one shape (fail-closed $*_BIN > which > landing
  spots + nvm/PATH sweeps); claude/codex/hermes/grok resolvers become
  factory products with identical candidate lists.
- makeAgentChildEnv + GITHUB_STEP_META_KEYS: hermeticChildEnv + per-agent
  overrides + key deletion + the step-metadata scrub + binDir prepend.
  hermesChildEnv GAINS the GITHUB_* deletion via the factory (the filed P2
  backport; truth-table extended).
- runOneShotSpawn: shared timeout/kill/kill-9-escalation/bounded-drain core;
  hermesOneShotTurn gains the escalation + bounded drain (strictly safer,
  nothing pinned the old unbounded wait); grokOneShotTurn is now a thin argv
  builder over it.
- 5a-opencode family (first consumer): resolveOpencodeBinary (fail-closed
  OPENCODE_BIN), hasOpencodeAuth (PAID-leg-only gate — the keyless free
  tier carries the core SMOKE), opencodeChildEnv (HOME + BOTH XDG dirs,
  anthropic re-admission, other-provider + OPENCODE_CONFIG* shadow-trio
  deletes), seedOpencodeConfig (config half of the double autoupdate kill),
  opencodeOneShotTurn, and parseOpencodeJsonl (event shapes pinned from the
  live v1.18.18 observation — {type,part} with part.text / part.tool).

EV1 gate (local keyless grok-door): 4 pass / paid-skip in 20.5s BEFORE and
AFTER the port, against a hermetically npm-pinned @xai-official/grok@1.0.4.
The hermes door self-skips without a binary — its port is pinned by the
unit truth-tables (stated honestly, per the plan).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e): opencode door — split-gated real-binary e2e on the extracted family (keyless SMOKE included)

The door goes a step beyond grok's split gating: opencode's anonymous free
tier drives MCP tool calls with zero credentials (observed, load-bearing),
so even the nonce SMOKE runs keyless. Tiers: T1 bare-semver version pin
(the SST-vs-claimant discriminator), T2 INSTALL via the documented
'mcp add … -- gbrain serve --surface verbs' shape + the honest 'mcp list'
discriminator (spawns servers; ✓/✗ text asserted — exit code is 0 even on
failure), T2b spawn-gate CANARY (a project-config decoy is spawn-attempted
with no trust prompt — if this ever gates, the bootstrap user-global
default rationale changed: re-observe), T3 writer parity (gbrain's
opencode-json.ts output handshakes through the real binary; cross-tool
preservation both ways incl. the autoupdate seed), T4 keyless SMOKE
(per-run nonce + STRUCTURAL gbrain_* tool_use proof via parseOpencodeJsonl,
list preflight before any turn, 2 attempts), and the paid T5 anthropic leg
(hasOpencodeAuth-gated; self-validating models-gate pins the model id
BEFORE any spend). Hermeticity: HOME + both XDG dirs per child, tmp cwds,
config/credential tripwire over the operator's real opencode state,
checkout guard, --pure on every probe (mcp list autoloads plugins), --pure
placed BEFORE the '--' separator (a trailing append lands inside the server
command — caught live). run-e2e.sh scrubs the OPENCODE_ prefix.

Verified live: 6/6 pass in 35.8s (keyless tier + paid anthropic leg)
against the hermetically pinned opencode-ai@1.18.18.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(heavy): opencode-door job (day-one full posture, keyless SMOKE) + canary leg + pin guards + hermes installer re-pin

opencode-door takes hermes-door's triggers (nightly + labels + dispatch;
cadence policy: nightly for the NEWEST door agent) with grok-door's
internals — keyless-first ordering, secretless npm provisioning with
wrapper AND per-platform integrity pre-checks, pass-count + paid sentinels,
mid-job version-drift tripwire, evidence scrub RE-KEYED to
ANTHROPIC_API_KEY + auth.json (not XAI/mcp_credentials), unconditional
credential removal. No dedicated dispatch input (any workflow_dispatch
already passes the non-PR arm — an input would be dead yaml). The keyless
tier includes the nonce SMOKE (free tier), so the core door needs NO
secret; the paid anthropic leg rides the secret hermes-door already
consumes. opencode-door-canary lands IN-WAVE (schedule-only,
continue-on-error, unpinned latest): opencode ships near-continuously — a
red canary is a pin-refresh signal, never a gate. real-agent-e2e adds the
opencode door file + env pins.

Guards: check-opencode-pin.sh (stamp↔workflow parity, job-block anchored so
the UNPINNED canary leg cannot satisfy it; fail-closed when the door exists
without the pin doc) and check-pin-doc-privacy.sh (placeholder discipline
for ALL docs/mcp/*-CLI-PIN.md — no operator home paths, no key-shaped
material outside sha512 pins, no non-example emails), both in bun run
verify + guards-manifest, both with fixture-tree bun tests.

Maintenance: hermes-door installer pin refreshed (upstream install.sh
drifted past the prior digest — last two nightlies red; reviewed: the
--commit payload-pin path is intact and the payload pins are unchanged).
docs/TESTING.md gains the opencode door entry + the door cadence policy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: opencode across the install/testing/architecture surface

README client roster + remote-connect bullet (with the not-OpenClaw
disambiguation and the bootstrap-supported banner), INSTALL_FOR_AGENTS
'If you are opencode' block (routes bootstrap-capable readers to the
runbook; brain-only registration otherwise), docs/INSTALL per-client list,
MEMORY_VERBS register snippet, bootstrap guide (degradation-matrix row
naming the INVERTED scope default + rationale; harness-mode opencode
bullet; dx-explore scenario line), ambient-recall/push-context harness
mentions, and KEY_FILES current-state entries (opencode-json.ts,
atomic-write.ts, connect/harness/hooks/claw-test entry refreshes).
llms bundles regenerated (build:llms chaser).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* dx(explore): opencode-install TTY scenario — bootstrap paste block under the real interactive TUI

Unlike grok's brain-only prompt, opencode gets the FULL bootstrap paste
block (it is a bootstrap-supported harness) under a hermetic HOME + both
XDG dirs, the double autoupdate kill (config seed + env), BROWSER=false so
a first-run can never bounce the operator's browser, and auth.json
pre-registered for the secret scrub. Keyless posture INVERTS the grok
scenario: the anonymous free tier means a --keyless run should COMPLETE
the flow — a sign-in wall here is itself a pin-refresh signal, and the
generic early-stop in runInstallSession records it as friction if it ever
appears.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(todos): file opencode-wave follow-ups + retire fired triggers by title

Door-adapter extraction (test-side) and cadence policy: DONE — the 4th-door
trigger fired. CI-tail composite action re-filed with the sharpened trigger
(first GREEN grok-door AND opencode-door dispatches). hermesChildEnv
GITHUB_* backport: DONE via the shared factory. PIN-doc privacy guard:
DONE (check-pin-doc-privacy.sh in verify). New follow-ups: first-dispatch
watch, plugin/event-system wiring (ambient-recall lane), BrainBench
adapter (with hermes+grok), connect --oauth authorization-code lane,
OPENCODE_CONFIG* re-observation on bumps, opencode-install PTY promotion.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: coverage pins for the opencode channel widenings

Ship-audit additions: hook.ts --harness opencode flag-parse attribution
end-to-end, and 'opencode' membership in VOLUNTEER_CHANNELS +
isHarnessChannel (a regression here silently rebadges opencode deliveries
as claude-code).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): close two plan-audit gaps — ACCESS_POLICY opencode scope paragraph + doctor host:opencode pin

Plan-completion audit (ship Step 8) flagged both as PARTIAL: the
ACCESS_POLICY template's MCP-scope section didn't state opencode's
inverted default (user-global; project spawns with NO trust prompt),
and bootstrap_harness_health had no named pin proving an opencode
receipt flows through the host-generic filter. Template-repo regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: pre-landing review fixes — review-army + red-team wave

Security: opencode error-path snippets render <paste-token-here> instead
of the live bearer; the harness opencode catch redacts like the claude
lane; test-harness child envs unconditionally drop GITHUB_TOKEN/ACTIONS_*.
Correctness: bootstrap-lock coverage for every shared-config writer
(hooks/connect/uninstall); uninstall + step-aside gate sweep BOTH global
opencode filenames (merged namespace); uninstall passes a sourceId
expectation and skips other-workspace entries; cross-kind fingerprint
matches classify ours-other-source instead of silently replacing;
dangling-symlink writes preserve the link; failed-smoke rollback restores
atomically; registration probe pins OPENCODE_DISABLE_AUTOUPDATE, a 20s
cap, ANSI-stripped exact-name matching. Guard: check-opencode-pin now
cross-checks per-platform integrities + every OPENCODE_VERSION copy.
Plus deny-path/uninstall/rollback/truth-table/symlink test coverage,
help-prose cosmetics, downgrade doc note, 3 P3 TODOs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: adversarial-review fix wave — cross-model (Claude + Codex) findings

P0: the registration probe no longer executes from the invoking cwd
(mkdtemp cwd for user scope; project scope skips the live probe —
parse-back is authoritative), so a cloned repo's committed opencode.json
can't gain code execution during bootstrap. Probe timeouts now kill the
child (SIGTERM→SIGKILL, bounded drain) instead of abandoning it over the
PGLite lock. Global writes reconcile mcp.<name> across BOTH merged
global filenames. Stale-target cleanup takes the target config-dir lock.
Backups are unique per operation; rollback is content-guarded and
remove-path backups tighten to 0600 when token-bearing. connect --force
now works on the opencode lane with url-appropriate refusal copy.
atomic-write cleans tmp litter and survives the exists/realpath race.
Bun-lane fingerprint is fail-closed on gbrain-less args. Scope answers
trim. bounded() clears its drain-cap timer. CI installs opencode from
byte-verified tarballs. Consent-semantics + hermetic-live-runner
follow-ups filed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: regen cli-flag-registry after review-fix waves

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.46.2.0 feat(opencode): full-parity client support — bootstrap, harness, connect, claw-test, e2e door

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: fold the post-review opencode fix-wave behaviors into KEY_FILES

Three current-state completions the fix agents didn't carry into the
per-file index: connect --agent opencode --install's --force semantics
(maps to the writer's allowReplaceOtherSource — ours-at-old-url
replaceable, foreign still refuses), removeOpencodeMcpEntry's
skipOtherSource option, and bootstrap uninstall's expectation-keyed
opencode sweep (both merged global filenames under the config-dir lock
plus the project file; other-workspace entries skipped with a note).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: cross-model doc-review fixes — opencode roster + probe/remote accuracy

Findings from the standard post-ship Codex doc review, verified against
the shipped code: the bootstrap guide's intro, install table, and door
inventory still described a two-client product (opencode added to all
three); INSTALL_FOR_AGENTS' grok section said the personal-agent path is
Claude Code/Codex only; OPENCODE.md called its recipe the bootstrap
"manual equivalent" (bootstrap additionally pins GBRAIN_SOURCE + full
surface), lacked the mcp-list trust caution the pin doc carries, and
never documented the connect --install / --force remote lane; the pin
doc's provisioning bullet now names the pack-verify-install posture the
CI job actually runs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.46.4.0 chore(release): re-bump 0.46.2.0 → 0.46.4.0 (version slots claimed by in-flight PRs)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 21:43:31 -07:00
154 changed files with 12340 additions and 1666 deletions
+1
View File
@@ -0,0 +1 @@
{"schema_version":3,"run_id":"f2b40f7ef-retrieval-canary-na-0","ran_at":"2026-08-15T15:37:16.659Z","suite":"retrieval-canary","mode":"n/a","commit":"f2b40f7ef","seed":0,"params":{"qrels":"test/fixtures/eval-baselines/qrels-search.json","embedder":"deterministic","k":10,"metrics":{"mean_recall_at_k":1,"first_relevant_hit_rate":1,"expected_top1_hit_rate":0.8333333333333334,"expected_top1_denominator":12,"queries_run":12,"queries_total":12},"floors":{"recall_at_k":0.7,"first_relevant_hit":0.6,"expected_top1":0.5}},"status":"completed","duration_ms":2290}
+1
View File
@@ -27,3 +27,4 @@
# Markdown it does not own), but pinning this repo's own .md checkout to LF
# removes the whole class for anyone working here.
*.md text eol=lf
/.gbrain-evals/eval-results.jsonl merge=union
+126 -8
View File
@@ -20,6 +20,44 @@ concurrency:
cancel-in-progress: true
jobs:
# ──────────────────────────────────────────────────────────────────────
# e2e-cache-check: same content-hash skip as test.yml's cache-check, in
# its own key namespace (e2e-pass-<hash>). Doc-only pushes previously
# provisioned 3 pgvector services and spent real OpenAI/Anthropic/
# ZeroEntropy tokens in tier2; now they skip. SCHEDULED runs are exempt
# below — the nightly is a drift check against live providers and must
# run even when the tree is unchanged.
# ──────────────────────────────────────────────────────────────────────
e2e-cache-check:
runs-on: ubuntu-latest
timeout-minutes: 10
outputs:
hit: ${{ steps.lookup.outputs.cache-hit }}
hash: ${{ steps.compute.outputs.hash }}
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Compute content hash
id: compute
run: |
HASH=$(bash scripts/ci-cache-hash.sh --verbose 2>/tmp/cache-diag)
cat /tmp/cache-diag
echo "Computed cache hash: $HASH"
echo "hash=$HASH" >> "$GITHUB_OUTPUT"
- name: Lookup actions/cache for e2e-pass-<hash>
id: lookup
uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
key: e2e-pass-${{ steps.compute.outputs.hash }}
path: .e2e-cache-marker
lookup-only: true
- name: Cache status
run: |
if [ "${{ steps.lookup.outputs.cache-hit }}" = "true" ]; then
echo "✓ e2e cache HIT for hash ${{ steps.compute.outputs.hash }} — e2e jobs will skip (unless scheduled)"
else
echo "✗ e2e cache MISS for hash ${{ steps.compute.outputs.hash }} — e2e suite will run"
fi
jsonb-parity:
# Dedicated required guard for the JSONB double-encode bug-class (#2339).
# PGLite parses a double-encoded jsonb string silently, so this assertion can
@@ -28,6 +66,8 @@ jobs:
# Postgres and HARD-FAILS if DATABASE_URL is missing, so the guard can never
# silently skip.
name: JSONB parity (#2339 regression guard)
needs: e2e-cache-check
if: needs.e2e-cache-check.outputs.hit != 'true' || github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
timeout-minutes: 15
services:
@@ -49,7 +89,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
- name: Require DATABASE_URL (no silent skip)
env:
DATABASE_URL: postgresql://postgres:postgres@localhost:5432/gbrain_test
@@ -73,6 +118,8 @@ jobs:
tier1:
name: Tier 1 (Mechanical)
needs: e2e-cache-check
if: needs.e2e-cache-check.outputs.hit != 'true' || github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
timeout-minutes: 20
services:
@@ -94,7 +141,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
- name: Run Tier 1 E2E tests
# job-isolation rides tier1 deliberately: e2e.yml runs only explicitly
# NAMED files (no glob) — an unwired e2e file is silent coverage loss.
@@ -106,12 +158,15 @@ jobs:
tier2:
name: Tier 2 (LLM Skills)
# Runs on every push/PR (promoted from schedule-only in v0.19.0), in
# PARALLEL with tier1 (own postgres service — the old `needs: tier1`
# serialized ~2min for no shared state). The jsonb-parity gate (~40s)
# stays in front as the broken-build SPEND gate: this job burns real
# OpenAI/Anthropic/ZeroEntropy tokens and must not fire when the build
# can't even pass the cheapest DB guard.
needs: [e2e-cache-check, jsonb-parity]
if: needs.e2e-cache-check.outputs.hit != 'true' || github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
# Runs on every push/PR now (promoted from schedule-only in v0.19.0).
# Tier 1 must pass first; Tier 2 uses OPENAI_API_KEY + ANTHROPIC_API_KEY
# from repo/org secrets. Nightly + manual triggers still supported via
# the workflow-level `on:` list.
needs: tier1
timeout-minutes: 30
services:
postgres:
@@ -132,7 +187,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
- name: Install OpenClaw
# Bound + retry the install: a transient npm/registry stall here used to
# hang unbounded and (since the v0.42.50.0 job timeout) burn the entire
@@ -179,3 +239,61 @@ jobs:
# zeroEntropyCompatFetch response-rewriter + URL rewrite + flexible
# dim handling + gateway.rerank against the real provider.
ZEROENTROPY_API_KEY: ${{ secrets.ZEROENTROPY_API_KEY }}
# ──────────────────────────────────────────────────────────────────────
# e2e-cache-write: seals e2e-pass-<hash> only when every gated job
# succeeded (writing earlier would bless states the suite never proved).
# Scheduled runs may also write: a nightly green at an unchanged hash is
# the same proof a push green is.
# ──────────────────────────────────────────────────────────────────────
e2e-cache-write:
needs: [e2e-cache-check, jsonb-parity, tier1, tier2]
if: success() && needs.e2e-cache-check.outputs.hit != 'true'
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Create cache marker
run: |
mkdir -p .e2e-cache-marker
echo "${{ needs.e2e-cache-check.outputs.hash }}" > .e2e-cache-marker/hash
echo "$GITHUB_SHA" > .e2e-cache-marker/sha
echo "$GITHUB_REF" > .e2e-cache-marker/ref
- uses: actions/cache/save@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
key: e2e-pass-${{ needs.e2e-cache-check.outputs.hash }}
path: .e2e-cache-marker
# ──────────────────────────────────────────────────────────────────────
# e2e-status: the single stable "did E2E pass?" name (mirror of
# test.yml's test-status). Succeeds when the cache hit on a non-scheduled
# run, or when every gated job succeeded.
# ──────────────────────────────────────────────────────────────────────
e2e-status:
needs: [e2e-cache-check, jsonb-parity, tier1, tier2]
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: Aggregate result
run: |
HIT="${{ needs.e2e-cache-check.outputs.hit }}"
JSONB="${{ needs.jsonb-parity.result }}"
TIER1="${{ needs.tier1.result }}"
TIER2="${{ needs.tier2.result }}"
EVENT="${{ github.event_name }}"
echo "e2e-cache-check.hit=$HIT event=$EVENT"
echo "jsonb-parity=$JSONB tier1=$TIER1 tier2=$TIER2"
# schedule AND workflow_dispatch always run the real suite — a
# manual dispatch is an explicit ask for a live run, so a cache
# hit must not report green-without-running for either.
if [ "$HIT" = "true" ] && [ "$EVENT" != "schedule" ] && [ "$EVENT" != "workflow_dispatch" ]; then
echo "✓ e2e cache HIT for hash ${{ needs.e2e-cache-check.outputs.hash }} — E2E green"
exit 0
fi
for r in "$JSONB" "$TIER1" "$TIER2"; do
if [ "$r" != "success" ]; then
echo "✗ gated e2e job did not succeed (got $r) — E2E fail"
exit 1
fi
done
echo "✓ all e2e jobs succeeded — E2E green"
+298 -8
View File
@@ -64,7 +64,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
- name: Run heavy tests
env:
@@ -109,8 +114,9 @@ jobs:
retention-days: 14
if-no-files-found: ignore
# Real-agent door e2e: drives the ACTUAL `claude` + `codex` + `hermes`
# binaries (no PATH shims) against a real gbrain over MCP. These pay real API
# Real-agent door e2e: drives the ACTUAL `claude` + `codex` + `hermes` +
# `grok` + `opencode` binaries (no PATH shims) against a real gbrain over
# MCP. These pay real API
# cost and need the binaries installed + authed, which a stock GitHub runner
# does NOT have — so the tests self-SKIP (describe.skipIf on binary/auth) and
# the job is a clean no-op here. It exists so a self-hosted /
@@ -133,10 +139,13 @@ jobs:
# a grok binary, which a stock runner does not have.
GBRAIN_REAL_HERMES_E2E: '1'
GBRAIN_REAL_GROK_E2E: '1'
GBRAIN_REAL_OPENCODE_E2E: '1'
# Pin so a provisioned runner's grok version-shape test asserts against
# the supported version (and a colliding community `grok` binary fails
# loud instead of running the keyless tier confusingly).
GROK_VERSION: "1.0.4"
# Same posture for opencode: a provisioned runner's version pin.
OPENCODE_VERSION: "1.18.18"
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
@@ -144,7 +153,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
# Reference the door tests; run only the ones present (a door may land
# in a sibling PR). Missing binary/auth → the file self-skips, so a
@@ -156,7 +170,8 @@ jobs:
test/e2e/bootstrap-real-claude.serial.test.ts \
test/e2e/bootstrap-real-codex.serial.test.ts \
test/e2e/install-real-hermes.serial.test.ts \
test/e2e/install-real-grok.serial.test.ts; do
test/e2e/install-real-grok.serial.test.ts \
test/e2e/install-real-opencode.serial.test.ts; do
[ -f "$f" ] && files+=("$f")
done
if [ "${#files[@]}" -eq 0 ]; then
@@ -195,7 +210,7 @@ jobs:
HERMES_VERSION: "0.20.0"
HERMES_GIT_TAG: "v2026.8.3"
HERMES_GIT_COMMIT: "3c27eb6234bf91b8ceee9e9071591b31e9b148cb"
HERMES_INSTALL_SHA256: "c118ff31618dc70339049ce71061b8f1351a1c70d9c2a236ed50d8a2550c550d"
HERMES_INSTALL_SHA256: "868ed3a91e0fabbff6d7418b3ede82bf4833652ec4e77196a42852fb35a9e5b9"
GBRAIN_REAL_HERMES_E2E: '1'
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -204,7 +219,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
# `runner.temp` is not an allowed context in job-level env, so the
# evidence dir is derived here and exported for every later step (the
@@ -406,7 +426,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
- name: Prepare evidence dir
run: |
@@ -595,3 +620,268 @@ jobs:
# hermetic homes carry no key file (env-only auth) but may hold
# grok-derived credentials once the authed inventory lands.
rm -rf /tmp/gb-grok-* 2>/dev/null || true
# opencode door e2e (SST opencode): PROVISIONS the real opencode binary via
# the pinned npm package (wrapper + per-platform payload integrities
# verified — both pins live in docs/mcp/OPENCODE-CLI-PIN.md, enforced
# against this file by scripts/check-opencode-pin.sh in `bun run verify`).
#
# DAY-ONE FULL POSTURE (a step past grok's pre-secret gating, deliberate):
# opencode's anonymous free tier drives MCP tool calls keyless (observed,
# load-bearing — OPENCODE-CLI-PIN.md §One-shot), so the ENTIRE core door —
# including the nonce SMOKE — runs with no secret; and the paid anthropic
# leg rides the ANTHROPIC_API_KEY secret that already exists (hermes-door
# consumes it). So this job takes the hermes-door triggers (nightly +
# labels + dispatch, cadence policy: nightly for the NEWEST door agent)
# with grok-door's internals (keyless-first ordering, secretless pinned
# provisioning, sentinels, scrub triple, unconditional credential removal).
# No dedicated dispatch input: any workflow_dispatch already passes the
# non-PR arm, so an input would be dead yaml.
opencode-door:
name: opencode door e2e (real binary, keyless SMOKE)
if: |
github.event_name != 'pull_request' ||
contains(github.event.pull_request.labels.*.name, 'real-agent-e2e') ||
contains(github.event.pull_request.labels.*.name, 'heavy-tests')
runs-on: ubuntu-latest
# Measured local door wall-time: full 6-test run 35.8s + one-time
# compiled gbrain build (~2-4 min) + npm install (~15s); free-tier +
# paid turn budgets 2 x 240s each. 20 min = measured + >50% headroom.
timeout-minutes: 20
env:
# Pin values documented in docs/mcp/OPENCODE-CLI-PIN.md — update them
# together, deliberately, after reviewing upstream changes
# (scripts/check-opencode-pin.sh fails `bun run verify` on drift).
OPENCODE_VERSION: "1.18.18"
OPENCODE_NPM_PACKAGE: "opencode-ai"
OPENCODE_NPM_INTEGRITY: "sha512-J+5HFq8tf+wPBBpBpMPSNjSytF2/EkNWYfFZh4si1d9auFbQriqDyqZv+vFUsLWERfdMU32Eajwuiq3rKBvZLQ=="
# Per-platform payload pins: the wrapper's integrity covers only the
# wrapper tarball; the binary that EXECUTES is the platform sub-package.
OPENCODE_NPM_LINUX_X64_INTEGRITY: "sha512-WmeUnhljYJ252wywKTiW4bNDzsas2njpjPUEh0jM6HKNI4vFxJtREtzaWViY4AKEAcOkLWT8Ll17ixvcHz3AnA=="
OPENCODE_NPM_LINUX_ARM64_INTEGRITY: "sha512-e8D3g0qJEIzawEg2+ygW3vkZjAYL2ssyAx4GbihjwXwZFvlZZy5zRWWzdz5KLBoHSTl0FB73vNtnNeXONyHpVQ=="
GBRAIN_REAL_OPENCODE_E2E: '1'
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
persist-credentials: false
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- name: Prepare evidence dir
run: |
echo "GBRAIN_E2E_EVIDENCE_DIR=$RUNNER_TEMP/opencode-door-evidence" >> "$GITHUB_ENV"
mkdir -p "$RUNNER_TEMP/opencode-door-evidence"
# Compile gbrain ONCE for both bun test invocations below.
- name: Build gbrain (compile once for both door runs)
run: |
bun build --compile --outfile "$RUNNER_TEMP/gbrain-door-bin" src/cli.ts
echo "GBRAIN_COMPILED_BIN=$RUNNER_TEMP/gbrain-door-bin" >> "$GITHUB_ENV"
# SECRETLESS provisioning, pack-verify-install: `npm pack` DOWNLOADS
# each artifact and reports the integrity of the BYTES it wrote, so the
# asserts below cover the tarballs actually held — closing the
# view-then-install TOCTOU (two registry round-trips a payload-swapping
# registry could split). The wrapper then installs FROM the verified
# local tarball, not a fresh registry resolve of the name. Payload
# resolution, honestly: that install still fetches the platform
# sub-package (opencode-linux-*) over the network; after the pack step
# byte-confirms the registry's payload artifact matches its pin, npm
# validates the install-time fetch against the same packument
# integrity. No --ignore-scripts: opencode-ai's postinstall places the
# platform binary (verified locally — with the flag the CLI refuses to
# run). Version assert lives here too — before any secret-bearing step.
- name: Install opencode (pinned npm package, pack-verify-install)
timeout-minutes: 10
run: |
packdir=$(mktemp -d)
read_integrity() {
node -e 'let d;try{d=JSON.parse(require("fs").readFileSync(0,"utf8"))}catch{d=[]}process.stdout.write((Array.isArray(d)&&d[0]&&d[0].integrity)||"")'
}
pushd "$packdir" >/dev/null
served=$(npm pack "$OPENCODE_NPM_PACKAGE@$OPENCODE_VERSION" --json 2>/dev/null | read_integrity || true)
if [ "$served" != "$OPENCODE_NPM_INTEGRITY" ]; then
echo "::error::opencode npm integrity drift for $OPENCODE_NPM_PACKAGE@$OPENCODE_VERSION — packed tarball integrity '$served', pinned '$OPENCODE_NPM_INTEGRITY'. Re-pin deliberately: update the stamps in docs/mcp/OPENCODE-CLI-PIN.md + this workflow after reviewing upstream (see the pin doc's re-observation checklist)." >&2
exit 1
fi
arch=$(uname -m)
case "$arch" in
x86_64) plat_pkg="opencode-linux-x64"; plat_pin="$OPENCODE_NPM_LINUX_X64_INTEGRITY" ;;
aarch64|arm64) plat_pkg="opencode-linux-arm64"; plat_pin="$OPENCODE_NPM_LINUX_ARM64_INTEGRITY" ;;
*) echo "::error::unsupported runner arch for the opencode payload pin: $arch" >&2; exit 1 ;;
esac
plat_served=$(npm pack "$plat_pkg@$OPENCODE_VERSION" --json 2>/dev/null | read_integrity || true)
if [ "$plat_served" != "$plat_pin" ]; then
echo "::error::opencode platform payload integrity drift for $plat_pkg@$OPENCODE_VERSION — packed tarball integrity '$plat_served', pinned '$plat_pin'. Re-pin deliberately (OPENCODE-CLI-PIN.md stamps + this workflow)." >&2
exit 1
fi
npm install -g ./opencode-ai-*.tgz
popd >/dev/null
rm -rf "$packdir"
if ! command -v opencode >/dev/null 2>&1; then
echo "::error::opencode did not resolve on PATH after npm install" >&2
exit 1
fi
version_output=$(opencode --version)
echo "$version_output"
# Observed shape: BARE semver (`1.18.18` — no name, no hash); the
# SST-vs-claimant discriminator (OPENCODE-CLI-PIN.md §Pin).
if [ "$(printf '%s' "$version_output" | tr -d '[:space:]')" != "$OPENCODE_VERSION" ]; then
echo "::error::opencode version drift — expected bare '$OPENCODE_VERSION', got: $version_output (see docs/mcp/OPENCODE-CLI-PIN.md triage table)" >&2
exit 1
fi
# KEYLESS TIER FIRST — and on opencode that includes the nonce SMOKE
# (free tier). ANTHROPIC_API_KEY is absent from this step by
# construction, so the paid describe self-skips.
- name: Run opencode door tests (keyless tier — SMOKE included)
run: |
EXIT=0
bun test --timeout=600000 test/e2e/install-real-opencode.serial.test.ts > door-keyless.txt 2>&1 || EXIT=$?
tail -40 door-keyless.txt
cp door-keyless.txt "$GBRAIN_E2E_EVIDENCE_DIR/" 2>/dev/null || true
if [ "$EXIT" -ne 0 ]; then
exit "$EXIT"
fi
# Exact expected shape for this tier: 5 keyless tests pass (T1, T2,
# T2b, T3, T4-SMOKE), the 1 paid test skips. Zero/partial-pass
# refuses green.
pass_count=$(grep -Eo '[0-9]+ pass' door-keyless.txt | tail -1 | grep -Eo '^[0-9]+' || true)
if [ -z "$pass_count" ] || [ "$pass_count" -lt 5 ]; then
echo "::error::opencode door keyless tier expected 5 passing tests, summary shows '${pass_count:-none}' — refusing to go green (see docs/mcp/OPENCODE-CLI-PIN.md triage table)" >&2
exit 1
fi
- name: Preconditions (secret present)
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
if [ -z "$ANTHROPIC_API_KEY" ]; then
echo "::error::ANTHROPIC_API_KEY secret is empty — the keyless tier above already ran (its coverage, including the SMOKE, is banked); the paid anthropic leg needs the secret hermes-door already consumes. Fork PRs get no secrets from GitHub." >&2
exit 1
fi
# Full run (paid anthropic leg included). The T5 models-gate inside the
# suite is the named bad-pin tripwire: it validates the pinned model id
# against the AUTHED `opencode models` list BEFORE any spend.
- name: Run opencode door tests (full — paid anthropic leg included)
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
EXIT=0
bun test --timeout=600000 test/e2e/install-real-opencode.serial.test.ts > door.txt 2>&1 || EXIT=$?
tail -40 door.txt
cp door.txt "$GBRAIN_E2E_EVIDENCE_DIR/" 2>/dev/null || true
if [ "$EXIT" -ne 0 ]; then
exit "$EXIT"
fi
# PAID-SENTINEL: with the key present, a skipping paid tier must
# never read as green (the split-gating false-green class). The
# grep target is the suite's literal skip log — mirrored in
# test/e2e/install-real-opencode.serial.test.ts (change together).
if grep -q 'SKIP paid tier' door.txt; then
echo "::error::opencode door paid tier skipped despite a present ANTHROPIC_API_KEY — hasOpencodeAuth() gate drift; refusing to go green" >&2
exit 1
fi
pass_count=$(grep -Eo '[0-9]+ pass' door.txt | tail -1 | grep -Eo '^[0-9]+' || true)
if [ -z "$pass_count" ] || [ "$pass_count" -lt 6 ]; then
echo "::error::opencode door full run expected 6 passing tests, summary shows '${pass_count:-none}'" >&2
exit 1
fi
# Auto-update tripwire: the DOUBLE kill (config seed + env var) is the
# whole defense — a version that MOVED mid-job means it failed and the
# pins above are no longer what just ran.
- name: Version re-check (mid-job drift tripwire)
if: always()
run: |
if command -v opencode >/dev/null 2>&1; then
version_output=$(opencode --version || true)
if [ "$(printf '%s' "$version_output" | tr -d '[:space:]')" != "$OPENCODE_VERSION" ]; then
echo "::error::opencode version moved mid-job — auto-update kill failed (expected '$OPENCODE_VERSION', got: $version_output)" >&2
exit 1
fi
fi
- name: Scrub credentials from evidence (defensive)
if: failure() && env.GBRAIN_E2E_EVIDENCE_DIR != ''
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
# Same triple as the sibling doors, RE-KEYED for this lane: the
# credential file candidate is opencode's auth.json and the content
# grep sweeps ANTHROPIC_API_KEY (not XAI). Auth is env-only here —
# the content grep is the layer that matters for opencode-written
# logs on the failure path.
find "$GBRAIN_E2E_EVIDENCE_DIR" -type f \( -name '.env' -o -name '*.env' -o -name 'auth.json' \) -exec rm -f {} + 2>/dev/null || true
find "$GBRAIN_E2E_EVIDENCE_DIR" -type l -delete 2>/dev/null || true
if [ -n "$ANTHROPIC_API_KEY" ]; then
grep -rlF "$ANTHROPIC_API_KEY" "$GBRAIN_E2E_EVIDENCE_DIR" 2>/dev/null | while IFS= read -r f; do
echo "::warning::removing evidence file containing the API key: ${f#"$GBRAIN_E2E_EVIDENCE_DIR"/}" >&2
rm -f "$f"
done
fi
- name: Upload opencode door evidence
if: failure() && env.GBRAIN_E2E_EVIDENCE_DIR != ''
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: opencode-door-evidence
path: ${{ env.GBRAIN_E2E_EVIDENCE_DIR }}
retention-days: 14
if-no-files-found: ignore
# Auth travels env-only, but a future login flow would persist
# auth.json — remove the known candidate unconditionally so nothing
# outlives the job even on a future self-hosted runner.
- name: Remove opencode credentials (unconditional)
if: always()
run: |
rm -f ~/.local/share/opencode/auth.json
rm -rf /tmp/gb-opencode-* 2>/dev/null || true
# opencode canary: latest-version leg (schedule-scoped, continue-on-error,
# own timeout — landed IN-WAVE, reversing the grok-style deferral, because
# opencode ships near-continuously and a frozen pin goes stale in weeks;
# the pinned lane above stays the deterministic gate while this tracks
# what users actually run). Keyless tier only (incl. the free-tier SMOKE);
# no secret ever reaches this job. A red here is a PIN-REFRESH SIGNAL
# (OPENCODE-CLI-PIN.md §Pin-refresh cadence), never a gate.
opencode-door-canary:
name: opencode door canary (latest, keyless, non-gating)
if: github.event_name == 'schedule'
runs-on: ubuntu-latest
timeout-minutes: 20
continue-on-error: true
env:
GBRAIN_REAL_OPENCODE_E2E: '1'
# Deliberately NO OPENCODE_VERSION pin: T1 asserts the bare-semver
# SHAPE only, and the suite runs against whatever `latest` is today.
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
persist-credentials: false
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- name: Build gbrain
run: |
bun build --compile --outfile "$RUNNER_TEMP/gbrain-door-bin" src/cli.ts
echo "GBRAIN_COMPILED_BIN=$RUNNER_TEMP/gbrain-door-bin" >> "$GITHUB_ENV"
- name: Install opencode@latest (unpinned — the whole point)
timeout-minutes: 10
run: |
npm install -g opencode-ai@latest
command -v opencode >/dev/null 2>&1
echo "canary version: $(opencode --version)"
- name: Run opencode door tests (keyless tier against latest)
run: |
EXIT=0
bun test --timeout=600000 test/e2e/install-real-opencode.serial.test.ts > door-canary.txt 2>&1 || EXIT=$?
tail -40 door-canary.txt
if [ "$EXIT" -ne 0 ]; then
echo "::warning::opencode canary red against latest — pin-refresh signal (OPENCODE-CLI-PIN.md §Pin-refresh cadence); the pinned lane is the gate."
exit "$EXIT"
fi
+5
View File
@@ -19,6 +19,11 @@ on:
permissions:
contents: read
# Rapid pushes to the same PR previously queued duplicate scans.
concurrency:
group: osv-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
osv-scan:
permissions:
+17 -2
View File
@@ -35,6 +35,7 @@ concurrency:
jobs:
version:
runs-on: ubuntu-latest
timeout-minutes: 10
outputs:
version: ${{ steps.v.outputs.version }}
exists: ${{ steps.v.outputs.exists }}
@@ -80,6 +81,7 @@ jobs:
target: bun-linux-x64
artifact: gbrain-linux-x64
runs-on: ${{ matrix.os }}
timeout-minutes: 30
permissions:
contents: read
id-token: write # for attest-build-provenance (Sigstore OIDC)
@@ -89,7 +91,12 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- run: bun install
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
# No test re-run here: the Test workflow already gated this exact SHA at
# merge (10 shards + E2E). Re-running the whole suite serially on the
# release runner is a flakier duplicate gate — it blocked the first
@@ -116,6 +123,7 @@ jobs:
needs: [version, build]
if: needs.version.outputs.exists == 'false'
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: write # create the tag + release (scoped to this job only)
steps:
@@ -182,6 +190,7 @@ jobs:
needs: [version, release]
if: needs.version.outputs.exists == 'false'
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: read
env:
@@ -207,7 +216,13 @@ jobs:
with:
bun-version: 1.3.13
- if: steps.gate.outputs.publish == 'true'
run: bun install
uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
- if: steps.gate.outputs.publish == 'true'
run: bun install --frozen-lockfile
- if: steps.gate.outputs.publish == 'true'
name: Generate template tree and byte-diff against the vendored copy
run: |
+90 -21
View File
@@ -91,16 +91,24 @@ jobs:
# now enforces a paid GITLEAKS_LICENSE (fails the job with "missing
# gitleaks license" for accounts it can't validate). The CLI is free, uses
# the committed .gitleaks.toml allowlist, and scans the same commit range.
- name: Cache gitleaks tarball
uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: /tmp/gitleaks-dl
key: gitleaks-8.30.1-linux-x64
- name: Install gitleaks (pinned + checksum-verified)
run: |
set -euo pipefail
VER=8.30.1
BASE="gitleaks_${VER}_linux_x64.tar.gz"
URL="https://github.com/gitleaks/gitleaks/releases/download/v${VER}"
curl -fsSL -o "/tmp/${BASE}" "${URL}/${BASE}"
mkdir -p /tmp/gitleaks-dl
[ -f "/tmp/gitleaks-dl/${BASE}" ] || curl -fsSL -o "/tmp/gitleaks-dl/${BASE}" "${URL}/${BASE}"
# Checksums fetched fresh EVERY run: a cache-restored tarball is
# re-verified against the published digest, never trusted.
curl -fsSL -o /tmp/gitleaks_checksums.txt "${URL}/gitleaks_${VER}_checksums.txt"
( cd /tmp && grep " ${BASE}\$" gitleaks_checksums.txt | sha256sum -c - )
tar -xzf "/tmp/${BASE}" -C /tmp gitleaks
( cd /tmp/gitleaks-dl && grep " ${BASE}\$" /tmp/gitleaks_checksums.txt | sha256sum -c - )
tar -xzf "/tmp/gitleaks-dl/${BASE}" -C /tmp gitleaks
install /tmp/gitleaks /usr/local/bin/gitleaks
gitleaks version
- name: Scan for secrets (gitleaks CLI, .gitleaks.toml)
@@ -134,11 +142,25 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
# This job is the ONE designated saver of the bun cache (the others
# restore-only, so 5 redundant post-job save attempts disappear).
# admin/bun.lock is in the key because verify's check:admin-build
# installs from it.
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- run: bun install
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock', 'admin/bun.lock') }}
restore-keys: bun-cache-${{ runner.os }}-
# verify's runner sources test-env.sh and builds the snapshot for its
# PGLite-booting eval checks — restore the cache so it's the ~40ms
# freshness check, not a cold build ahead of all ~47 checks.
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: |
test/fixtures/pglite-snapshot.tar
test/fixtures/pglite-snapshot.version
key: pglite-snapshot-${{ runner.os }}-${{ hashFiles('src/core/migrate.ts', 'src/core/pglite-schema.ts', 'test/helpers/legacy-embedding-config.ts', 'scripts/build-pglite-snapshot.ts') }}
- run: bun install --frozen-lockfile
- run: bun run verify
# Guard: no bare `bun test` in workflows/scripts — bun ignores
# bunfig.toml's timeout, and hooks (beforeAll/afterAll) get the 5s
@@ -147,10 +169,11 @@ jobs:
- run: bash scripts/check-bun-test-timeout.sh
serial-tests:
# *.serial.test.ts at --max-concurrency=1. Lives in its own runner so
# the matrix shards aren't carrying the serial-pass tail (the old shape
# stuffed this into `test (1)` after the matrix work, which compounded
# shard 1's overload).
# *.serial.test.ts — one bun process per file (module-registry isolation),
# POOLED across files by scripts/run-serial-tests.sh (was strictly
# sequential: an 8.5-minute job whose serialization the quarantine
# contract never required). Lives in its own runner so the matrix shards
# aren't carrying the serial tail.
needs: cache-check
if: needs.cache-check.outputs.hit != 'true'
runs-on: ubuntu-latest
@@ -160,11 +183,27 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- run: bun install
restore-keys: bun-cache-${{ runner.os }}-
# PGLite schema snapshot (~42MB): the runner builds it when absent or
# stale (its runtime hash is authoritative — a stale restore is rebuilt,
# never trusted). Cached so the build is paid once per schema change,
# not once per job per run. This job SAVES; verify + matrix + slow jobs
# restore-only. The key is an approximation on purpose: it only has to
# be a superset-trigger of real schema changes.
# KEY HAS 5 HOMES in this file (this save + 4 restores: verify, matrix,
# slow-eval, slow-perf) — edit all together, or drift shows up only as
# silent rebuild cost.
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: |
test/fixtures/pglite-snapshot.tar
test/fixtures/pglite-snapshot.version
key: pglite-snapshot-${{ runner.os }}-${{ hashFiles('src/core/migrate.ts', 'src/core/pglite-schema.ts', 'test/helpers/legacy-embedding-config.ts', 'scripts/build-pglite-snapshot.ts') }}
- run: bun install --frozen-lockfile
- run: bun run test:serial
slow-eval-longmemeval:
@@ -185,11 +224,20 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- run: bun install
restore-keys: bun-cache-${{ runner.os }}-
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: |
test/fixtures/pglite-snapshot.tar
test/fixtures/pglite-snapshot.version
key: pglite-snapshot-${{ runner.os }}-${{ hashFiles('src/core/migrate.ts', 'src/core/pglite-schema.ts', 'test/helpers/legacy-embedding-config.ts', 'scripts/build-pglite-snapshot.ts') }}
- run: bun install --frozen-lockfile
- name: Ensure PGLite snapshot (build-or-validate, non-fatal)
run: bash -c '. scripts/lib/test-env.sh && ensure_pglite_snapshot slow-eval && echo "GBRAIN_PGLITE_SNAPSHOT=${GBRAIN_PGLITE_SNAPSHOT:-}" >> "$GITHUB_ENV"'
- run: bun test test/eval-longmemeval-e2e.slow.test.ts --timeout=60000
brainbench:
@@ -206,16 +254,19 @@ jobs:
timeout-minutes: 10 # ~15s hermetic run; matches the per-job-timeout hardening (#2254)
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
fetch-depth: 0 # the gate needs origin/master's baseline
- name: Fetch origin/master baseline ref (shallow)
# The gate reads ONE file via `git show origin/master:...` — a depth-1
# fetch of the master ref replaces the previous full 3700-commit clone.
run: git fetch --no-tags --depth=1 origin +refs/heads/master:refs/remotes/origin/master
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- run: bun install
restore-keys: bun-cache-${{ runner.os }}-
- run: bun install --frozen-lockfile
- run: bash scripts/ci-brainbench-gate.sh
env:
BRAINBENCH_OUT: ${{ runner.temp }}/brainbench-result.json
@@ -242,11 +293,20 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- run: bun install
restore-keys: bun-cache-${{ runner.os }}-
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: |
test/fixtures/pglite-snapshot.tar
test/fixtures/pglite-snapshot.version
key: pglite-snapshot-${{ runner.os }}-${{ hashFiles('src/core/migrate.ts', 'src/core/pglite-schema.ts', 'test/helpers/legacy-embedding-config.ts', 'scripts/build-pglite-snapshot.ts') }}
- run: bun install --frozen-lockfile
- name: Ensure PGLite snapshot (build-or-validate, non-fatal)
run: bash -c '. scripts/lib/test-env.sh && ensure_pglite_snapshot slow-perf && echo "GBRAIN_PGLITE_SNAPSHOT=${GBRAIN_PGLITE_SNAPSHOT:-}" >> "$GITHUB_ENV"'
- run: bun test test/entity-resolve-perf.slow.test.ts --timeout=300000
# MEMORY_VERBS v1 (Cathedral 1): the entity() p99 < 100ms contract gate
# (20K-page corpus + ratio guard) shares this runner — same perf-job
@@ -299,11 +359,20 @@ jobs:
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: 1.3.13
- uses: actions/cache@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: ~/.bun/install/cache
key: bun-cache-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- run: bun install
restore-keys: bun-cache-${{ runner.os }}-
# Restore-only: test-shard.sh validates the snapshot's runtime hash and
# rebuilds when stale (the serial-tests job is the designated saver).
- uses: actions/cache/restore@5a3ec84eff668545956fd18022155c47e93e2684 # v4.2.3
with:
path: |
test/fixtures/pglite-snapshot.tar
test/fixtures/pglite-snapshot.version
key: pglite-snapshot-${{ runner.os }}-${{ hashFiles('src/core/migrate.ts', 'src/core/pglite-schema.ts', 'test/helpers/legacy-embedding-config.ts', 'scripts/build-pglite-snapshot.ts') }}
- run: bun install --frozen-lockfile
- name: Run test shard ${{ matrix.shard }}/10
run: scripts/test-shard.sh ${{ matrix.shard }} 10
+25 -12
View File
@@ -1,16 +1,17 @@
<!-- gbrain-runbook-stamp: 0.46.3.0 -->
<!-- gbrain-runbook-stamp: 0.46.6.0 -->
<!-- This stamp must equal the VERSION file at every release; CI enforces it
(scripts/check-bootstrap-tag.sh). `gbrain bootstrap status` compares it to
the installed binary and warns on skew. -->
# BOOTSTRAP_FOR_AGENTS.md — become your human's persistent agent
You are an AI agent (Claude Code or Codex) reading this because a human pasted a
block asking you to set yourself up as their persistent personal agent, with gbrain
as your memory. This runbook is the judgment layer; the deterministic work lives in
`gbrain bootstrap` subcommands with exit codes. Follow it top to bottom.
You are an AI agent (Claude Code, Codex, or opencode) reading this because a human
pasted a block asking you to set yourself up as their persistent personal agent, with
gbrain as your memory. This runbook is the judgment layer; the deterministic work lives
in `gbrain bootstrap` subcommands with exit codes. Follow it top to bottom.
**Scope note:** this path is for Claude Code and Codex (desktop apps or CLIs).
**Scope note:** this path is for Claude Code, Codex, and opencode (desktop apps or
CLIs; opencode = the SST terminal agent, opencode.ai — not OpenClaw).
Running OpenClaw or Hermes? Use `INSTALL_FOR_AGENTS.md` instead.
**End state:** this folder is your workspace — identity files rendered from your
@@ -96,12 +97,16 @@ you needed; report the count at the end (it feeds the install-time measurement).
3. **Interview.** `gbrain bootstrap interview --init`, then ask the questions from
the bank (the CLI prints them) in three batches, recording each answer verbatim
with `--set KEY "value"`. Push once on vague answers to the required questions.
Claude Code only: with the final batch, also ask the ONE operational consent —
MCP scope. It is not one of the 12 interview questions; consents ride alongside
the bank. The choice: project (recommended — any other repo you open cannot
read your brain) vs user (your agent everywhere, but any repo you open can
reach it — read and write — and two open sessions contend for the database).
Record it with
Claude Code and opencode: with the final batch, also ask the ONE operational
consent — MCP scope. It is not one of the 12 interview questions; consents ride
alongside the bank. On Claude Code the choice: project (recommended — any other
repo you open cannot read your brain) vs user (your agent everywhere, but any
repo you open can reach it — read and write — and two open sessions contend for
the database). On opencode the recommendation INVERTS: user-global is the
default and the sharing-safe choice (opencode spawns project-config-defined
servers with NO trust prompt, so a committed project entry executes on every
collaborator's machine) — offer project only as a deliberate opt-in and state
that consequence. Record it with
`gbrain bootstrap interview --set MCP_SCOPE <project|user>` BEFORE the
read-back, so the confirmation covers it. On Codex, skip this question
entirely — the wiring step states the Codex reality instead.
@@ -133,6 +138,14 @@ you needed; report the count at the end (it feeds the install-time measurement).
on this machine can reach the brain (read and write) through its MCP
tools; the off-ramps are `codex mcp remove gbrain` (registration only) or
`gbrain bootstrap uninstall` (full teardown).
- opencode: writes the MCP entry directly into opencode's JSONC config (no
CLI exec needed) and relies on the AGENTS.md protocol, which opencode loads
natively — say plainly that opencode gets pull-based context, not per-turn
push. Scope follows the recorded MCP_SCOPE answer (user-global default; a
project answer writes the committed-candidate `opencode.json` and the CLI
prints the sharing warning). Restart opencode after wiring — it reads config
at session start. Off-ramps: the entry's `"enabled": false`, or
`gbrain bootstrap uninstall`.
7. **Private repo.** `gbrain bootstrap repo` — creates a PRIVATE GitHub repo from
the workspace, verifies the privacy bit through the API, pushes. If the human
started from a repo they created themselves (create-repo-first: an EMPTY private
+180
View File
@@ -2,6 +2,186 @@
All notable changes to GBrain will be documented in this file.
## [0.46.6.0] - 2026-08-15
**A busy machine can no longer make the job queue evict its own healthy
work.** ([#4145](https://github.com/garrytan/gbrain/issues/4145)) Under
sustained CPU load, a long-running background job (a subagent averaging
~3 minutes) could miss one lock-renewal window and get force-evicted
mid-inference — the queue would churn for hours while completions stayed
near zero, and the logs read like an orphan leak. Lock renewal is now
**verify-before-evict**: a slow or failed renewal is never treated as
loss; the worker asks the database the one authoritative question (a
fenced re-check) and keeps the job whenever the lease is still its own.
### Added
- **Per-job lock leases.** Long LLM handlers (subagent, autopilot-cycle,
embed-backfill, …) now hold a 300s lease by default instead of the
global 30s; single-call LLM handlers get 120s; short jobs keep 30s for
fast dead-worker recovery. Override per submission with
`gbrain jobs submit --lock-duration-ms N` (clamped to 5s1h; also an
MCP `submit_job` param). Stored on the job row (migration v130), so it
survives worker restarts, and renewed at a `min(lease/2, 60s)` cadence.
The bound is enforced end-to-end — at submit, again on the resolved
lease at claim, and by a database range constraint — and
`--dry-run` echoes the clamped value that will actually be stored.
- **Self-explaining eviction forensics.** Every renewal fault now logs and
audits WHY it failed (call-timeout vs refused vs fenced-lost), how late
the renewal timer fired vs its own cadence (the "was the worker starved
or was the database down?" discriminator), host load, and an
event-loop-delay sample scoped to the failing window. The ops runbook
gained a table for reading these plus the full env-knob reference
(`GBRAIN_LOCK_RENEWAL_*`, `GBRAIN_MINION_STALL_RECLAIM_GRACE_MS`).
- **Stall-sweep reclaim grace.** A lease that lapsed within the last 15s
is not reclaimed — a just-recovered worker's own renewal wins the race
against the sweep instead of having its live job stolen (env-tunable,
capped at 10 minutes with a warn-once clamp; `0` restores the previous
behavior).
### Changed
- **Eviction requires evidence.** The renewal state machine aborts a job
only on a fenced miss (the row was genuinely reclaimed — requeued with
no attempt burned) or after a hard backstop (default 2× the lease)
during a total database outage. Wall-clock jumps can no longer distort
the decision (elapsed-time math runs on a monotonic clock), and a
renewal timeout now also cancels the in-flight query so it stops
holding a pool slot.
- **`gbrain jobs get`** shows the job's lock lease alongside its
wall-clock budget, including the default that will stamp at claim.
### Fixed
- **The unit-suite's SIGTERM-semantics test no longer kills its own
shard.** A bare in-process signal broadcast could reach a leaked
cleanup handler from an earlier test file and exit the whole test
process mid-suite, misreading as an external kill; the test now fires
only the listeners it registered.
- **`gbrain verify` no longer leaves probe tombstones in your brain.**
The end-of-run probe cleanup previously soft-deleted its two probe
pages; every verify run left residue visible to `include_deleted`
readers until the 72h purge. Cleanup now hard-deletes.
- **Relational retrieval is deterministic on ties.** When a graph node is
reachable at the same depth from multiple seeds, the reported path was
plan/heap-order dependent (and could differ between engines); a
lexicographic tie-break restores the documented determinism.
- **A worker slot can no longer lose track of a re-claimed job.** After a
force-evict, the stale execution's cleanup could delete the tracking
entry of the SAME job re-claimed by the same worker; both cleanup sites
now verify generation (lock token) before deleting.
- **Developer e2e lane un-rotted.** 29 test failures across 13 files in
the developer-machine e2e lane (which CI does not run) were fixed:
ten rotted test files re-pinned to current intended behavior, plus the
wrapper's per-file timeout now accommodates the LLM-bound Tier-2 files
(`GBRAIN_E2E_FILE_TIMEOUT`).
To take advantage of v0.46.6.0: upgrade and restart your worker
(`gbrain jobs supervisor stop && gbrain jobs supervisor start --detach`).
Existing queues need no migration steps — the new lease column defaults
every existing row to its handler's lease at next claim. If you tuned
around the old eviction behavior (e.g. very high `--max-stalled`), you can
likely lower it now. Mixed fleets are safe: an old worker simply keeps the
old 30s behavior until restarted.
## [0.46.5.0] - 2026-08-15
**CI in half, evals actually gating.** The Test workflow ran 89.5 minutes on
every push; its long pole was a serial-test job that executed ~140 per-file bun
processes strictly one-at-a-time even though the quarantine only ever required
per-process isolation. This release pools that lane (8.5 min → ~4 min in CI,
~2.5 min locally), wires the PGLite schema-snapshot fast-path into the CI test
runners (it previously existed but only the local loop used it), and rebalances
the 10-shard matrix on freshly mined weights (new files without a mined weight
now fall back to the p75 file weight instead of the median) — a measured branch
run landed the whole workflow at 255s. E2E stops spending real provider tokens on doc-only
pushes (content-hash skip with nightly + manual-dispatch exemptions) and runs
its tiers in parallel behind a fast broken-build spend gate.
Retrieval quality now has a hermetic CLI canary: `gbrain eval gate` accepts a
deterministic embedder option that drives the full hybrid/RRF pipeline with
zero API keys, gated in CI on every run (`check:eval-canary`, alongside the new
`check:eval-chronicle` gate) with its run ledger committed to
`.gbrain-evals/eval-results.jsonl`. Two registered-but-never-executed guards
came alive, a registration⇒execution coverage test closes that class for good,
and 47 orphaned eval-harness tests joined the CI matrix behind a keyless
allowlist. Test reliability hardening rounds it out: externally-killed serial
files get a sequential rescue re-run (never a silent pass), machine-global
files live on a growth-guarded exclusive lane, and the shard-balance test now
asserts the matrix CI actually runs instead of recomputing its own inputs.
**To take advantage of v0.46.5.0:** nothing to configure — CI and the local
loops (`bun run test`, `bun run test:serial`, `bun run verify`) are just
faster. New knobs if you need them: `GBRAIN_SERIAL_POOL=1` restores the old
fully-sequential serial lane, `GBRAIN_VERIFY_MAX_PARALLEL` bounds verify's
worker pool, `GBRAIN_NO_SNAPSHOT=1` opts any runner out of the snapshot
fast-path. Run the retrieval canary yourself with
`bun run scripts/run-eval-canary.ts` (add `--record` to append the committed
ledger).
## [0.46.4.0] - 2026-08-15
**opencode joins the supported-client roster — at full parity from day one.**
(opencode is opencode.ai, SST's terminal agent — not OpenClaw.) Unlike earlier
clients that started with a manual recipe, opencode lands with every install
lane gbrain has: the paste-in workspace bootstrap, machine-level harness
wiring, `gbrain connect`, a claw-test runner, and a real-binary e2e door in
CI. Every asserted flag, config shape, and quirk was observed against a
pinned install (opencode 1.18.18), recorded in a machine-checked pin
document, and exercised against the real binary — including the part that
makes opencode special: its keyless anonymous free tier drives MCP tool
calls, so the end-to-end proof needs zero secrets.
### Added
- **`gbrain bootstrap hooks --harness opencode`** — workspace-lane MCP
registration via direct, comment-preserving JSONC writes (never a CLI
exec, works offline). MCP scope is honored with a deliberately INVERTED
default: user-global, because opencode spawns project-config servers with
no trust prompt; project scope is an explicit opt-in that prints a sharing
warning. A structural ownership fingerprint refuses to touch entries
gbrain didn't write.
- **`gbrain bootstrap harness --harness opencode`** — machine-level remote
MCP wiring with an inline bearer written 0600, token rotation across URL
changes, content-guarded rollback on failed smoke, `--status` and
`--remove`.
- **`gbrain connect --agent opencode [--install]`** — env-interpolated
bearer (`{env:GBRAIN_REMOTE_TOKEN}`): the token never enters the config
file. `--force` replaces a registration whose endpoint moved.
- **`gbrain claw-test --agent opencode`** and a split-gated real-binary e2e
door in CI: keyless tier (version pin, install + `mcp list` handshake,
spawn-gate canary, writer parity, MCP SMOKE on the free tier) plus a paid
Anthropic leg that model-gates before spending; npm supply-chain
provisioning verifies the actual downloaded tarball bytes against pinned
integrities; a schedule-only canary tracks the latest upstream release.
- **Docs:** `docs/mcp/OPENCODE.md` install guide,
`docs/mcp/OPENCODE-CLI-PIN.md` observation pin (with a verify-time drift
guard and a pin-refresh cadence), roster updates across README / INSTALL /
bootstrap guides. opencode reads the rendered AGENTS.md pull-protocol
contract natively.
### Changed
- The bootstrap config writers (Claude hooks JSON, Codex TOML, opencode
JSONC) now share one atomic-write helper; symlinked configs — including
dangling dotfile-manager links — survive writes as links.
- The door-test family (binary resolution, hermetic child envs, one-shot
spawns) extracted into shared factories; the hermes and grok runners were
ported onto them, hermes child envs gained the GitHub step-metadata scrub,
and the hermes installer pin was refreshed (its nightly door had gone red
on upstream installer drift).
- A new pin-doc privacy guard asserts every agent pin document ships with
placeholder paths and no key material.
### Fixed
- Security and robustness hardening from the pre-landing cross-model review
pass: registration verification probes run isolated and time-bounded, and
a hung probe is killed instead of abandoned; global config writes
reconcile both opencode global filenames under the bootstrap lock; config
backups are unique per operation with content-guarded restore; error
paths never echo credentials; test-harness child processes drop CI
credentials before spawning third-party binaries.
### To take advantage of v0.46.4.0
opencode users: run `gbrain bootstrap hooks --harness opencode` in your
brain workspace (or paste the standard bootstrap block into an opencode
session). The keyless free tier is enough to verify the wiring end to end —
`opencode mcp list` should show `✓ gbrain connected`. Existing installs:
nothing changes; this release adds a client, it doesn't modify brain
behavior.
## [0.46.3.0] - 2026-08-15
**ZeroEntropy is shutting down on 2026-09-04 — gbrain now gets you off it
-2
View File
@@ -127,8 +127,6 @@ the database name must carry "test" as a word segment (like `gbrain_test`
above) or destructive tests refuse to run — opt a differently-named database
in one-shot with `GBRAIN_E2E_ALLOW_DB=<name>`.
Use `bun run verify` before pushing. It runs 19+ guard checks in parallel
Use `bun run verify` before pushing. It runs 40+ guard checks in parallel
(`scripts/run-verify-parallel.sh`), including: banned fork-name leaks
(`scripts/check-privacy.sh`), `JSON.stringify(x)::jsonb` interpolation
+16 -1
View File
@@ -239,10 +239,25 @@ grok mcp add gbrain -e "GBRAIN_HOME=$HOME" -- gbrain serve --surface verbs
The add is lazy (exit 0 without connecting) — verify with
`grok mcp doctor gbrain`, which spawns the server and must report
`7 tools discovered`. This is the brain-only install; the `gbrain bootstrap`
personal-agent path does not support Grok yet (Claude Code/Codex only).
personal-agent path does not support Grok yet (Claude Code, Codex, and opencode only).
Verified against Grok Build v1.0.4. Full reference:
[docs/mcp/GROK.md](docs/mcp/GROK.md).
**If you are opencode** (the SST terminal agent, opencode.ai — not OpenClaw):
you are a bootstrap-supported harness — for the full persistent-personal-agent
install, follow `BOOTSTRAP_FOR_AGENTS.md` instead of this page. For the
brain-only MCP registration:
```bash
opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs
```
The add is lazy (exit 0 without connecting) — verify with `opencode mcp list`,
which spawns the server and must show `✓ gbrain connected` (the exit code is 0
even on failure; read the output). Restart opencode afterwards — it reads
config at session start. Verified against opencode v1.18.18. Full reference:
[docs/mcp/OPENCODE.md](docs/mcp/OPENCODE.md).
Whether you scaffolded or not, read `skills/RESOLVER.md` (in your workspace, or the
bundled copy at `~/gbrain/skills/RESOLVER.md` when running from the cloned repo). It's
the skill dispatcher — tells you which skill to read for any task. Save this to your
+1
View File
@@ -174,6 +174,7 @@ GBrain exposes nearly all of its 100+ operations as MCP tools (stdio and HTTP; a
- **[Cursor / Windsurf / any stdio MCP client](docs/mcp/CLAUDE_CODE.md)** — same shape, add `{"command": "gbrain", "args": ["serve"]}` to your MCP config.
- **[Hermes](docs/mcp/HERMES.md)** — `printf 'Y\n' | hermes mcp add gbrain --env GBRAIN_HOME=$HOME --connect-timeout 60 --command $(which gbrain) --args serve`. Keep `--args` last, and verify with `hermes mcp test gbrain` (the add exits 0 even on failure).
- **[Grok Build](docs/mcp/GROK.md)** — `grok mcp add gbrain -e "GBRAIN_HOME=$HOME" -- gbrain serve --surface verbs`. The add is lazy (exit 0 without connecting) — verify with `grok mcp doctor gbrain`, which spawns the server and reports `7 tools discovered`. Verified against Grok Build v1.0.4.
- **[opencode](docs/mcp/OPENCODE.md)** (opencode.ai / SST — not OpenClaw) — `opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs`, or let `gbrain bootstrap hooks --harness opencode` write the config for you (opencode is a bootstrap-supported harness — it reads AGENTS.md natively). The add is lazy — verify with `opencode mcp list`, which spawns the server (`✓ gbrain connected`). Remote: `gbrain connect https://your-host/mcp --token gbrain_xxx --agent opencode [--install]` — the config stores only the `{env:GBRAIN_REMOTE_TOKEN}` interpolation. Verified against opencode v1.18.18.
- **[OpenClaw](docs/mcp/OPENCLAW.md)** — the ClawHub bundle plugin registers gbrain automatically (`openclaw.plugin.json` ships in this repo), or add `{"command": "gbrain", "args": ["serve"]}` to `~/.openclaw/config.json`'s `mcpServers`.
- **[Claude Desktop (Cowork)](docs/mcp/CLAUDE_DESKTOP.md)** — Settings → Integrations → add the URL of your HTTP server. Remote only; the local `claude_desktop_config.json` does not work for remote servers.
- **[Claude Cowork (team plan)](docs/mcp/CLAUDE_COWORK.md)** — org Owner adds the connector under Organization Settings → Connectors.
+223 -19
View File
@@ -1,5 +1,37 @@
# TODOS
## #4145 lock-renewal wave follow-ups (filed 2026-08-15)
- [ ] **P2 — Kill or reap the force-evicted handler process.** **What:** when the
grace-evict fires for a handler that ignores its AbortSignal, actually
terminate the handler's work (LLM loop cancellation vs shell child-tree
kill differ per handler class) or track it as a zombie instead of only
freeing the inFlight slot. **Why:** today the evicted handler keeps
burning CPU/spend on an already-saturated host while the worker claims
new work — the #4145 amplification loop — and the duplicate-external-
side-effect window during an asymmetric outage is bounded only by
handler cooperation, not by `hardEvictMs`. **Context:** deliberately
scoped out of the #4145 wave (grace-evict at
`src/core/minions/worker.ts` frees the slot; the alternative — retaining
the slot until handler exit — re-opens the wedged-slot class D8b closed).
Eviction frequency collapsed with verify-before-evict, so this is
hygiene, not the incident driver. Kill semantics need their own review.
**Effort:** M (human) / S (CC). **Priority:** P2.
- [ ] **P3 — Worker-level `--lock-duration` flag on `jobs work` + supervisor
passthrough.** **What:** a CLI flag for the worker-global default lease,
threaded through `buildWorkerArgs` (`src/core/minions/supervisor.ts`).
**Why:** convenience only — per-job/per-type leases
(`HANDLER_DEFAULT_LOCK_DURATION_MS`, `--lock-duration-ms`) plus the
`GBRAIN_LOCK_RENEWAL_*` env knobs already cover every incident-tuning
case shipped in the #4145 wave. **Context:** requested shape existed in
the issue; deferred because no production caller overrides
`lockDuration` and env wins for incident response. **Effort:** S.
**Priority:** P3.
- [ ] **Note for TODO-LR-2 (doctor `lock_renewal_health`, already filed
below):** the #4145 wave shipped exactly its inputs — audit events now
carry `cause`, `lateness_ms`, `overlap_skips`, `load1`/`cores`, `via`,
`deadline_deferred` — so the doctor check can classify starved-worker
vs DB-outage windows without new plumbing.
## v0.47 SEPTEMBER REMOVAL — ZeroEntropy (filed v0.46.3.0; TARGET: ship 2026-09-04..2026-09-08)
ZeroEntropy's hosted API dies 2026-09-04. v0.46.3.0 deprecated it (split-default:
@@ -177,14 +209,77 @@ fix-wave plan; the wave series (W0.5W9, 3.4, 3.6) tracks its own scope there.
- [ ] **Legacy Anthropic-SDK subagent loop deletion.** **Priority: P2.** One
release after W8 flips `agent.use_gateway_loop` default ON (flag stays as
the revert path for that release).
- [ ] **Deeper test-suite speedup** beyond the W0 snapshot default-on (which
already cut the full parallel suite ~4,900s → ~490s). **Priority: P3.**
Revisit with post-W0 timing data; diminishing returns until measured.
- [x] **Deeper test-suite speedup** beyond the W0 snapshot default-on
LANDED in the test/eval/CI speedup pass (serial pool 8.5min → ~2.5min,
snapshot in every CI runner + memoized loader, verify worker pool,
perf-gate row shrink, chunk-grain engine consolidation). Remaining
long-tail items are filed in "Test/eval/CI speedup pass deferrals" below.
- [ ] **PGLite schema build-time derivation** from SCHEMA_SQL via a named
transform list. **Priority: P3.** Only if W3's schema drift TEST proves
annoying in practice — the test alone kills the drift bug class (Codex
D4.8/D5.23: fresh-schema equivalence ≠ upgrade correctness; old-shape
bootstrap fixtures + replay coverage stay regardless).
## Test/eval/CI speedup pass deferrals (filed with the pass; plan: ~/.claude/plans/system-instruction-you-are-working-iterative-hopcroft.md)
Each was explicitly deferred in the pass's CEO/eng/outside-voice reviews.
- [ ] **Sleep-to-poll conversions.** **What:** replace ~49.5s of hard-coded
`setTimeout` waits with event/poll-based waits; no fake timers exist in the
suite. Worst offenders: test/minions.test.ts (12.2s across 43 sites),
test/process-cleanup.test.ts (5.0s), test/worker-lock-renewal-e2e.serial.test.ts
(4.0s), test/e2e/worker-abort-recovery.test.ts (3.6s), test/e2e/zombie-reaping.test.ts
(3.3s). **Why deferred:** careful per-site work against flake-hardened timings;
~50s ceiling. **Effort:** M. **Priority:** P3.
- [ ] **E2E: PGLite-only parallel lane + default SHARD.** **What:** run-e2e.sh runs
181 files sequentially (one bun cold start each); ~42 PGLite-only files need no
Postgres and no TRUNCATE-race protection — run them in a parallel lane; default
the existing SHARD support (only ci-local uses it). Fold into the Postgres
template-database entry below in this file (CREATE DATABASE … TEMPLATE, ~50ms).
**Why deferred:** e2e is off the CI critical path after the workflow restructure;
ci-local + nightly benefit only. **Effort:** M. **Priority:** P2.
- [ ] **Second PGLite snapshot keyed by dims/model.** **What:** ~34 test files
configure zembed/1280 and always cold-init (the snapshot's shape gate correctly
refuses the 1536 fixture). Bake a second snapshot per shape; the version-file
format already carries dims/model. **Why deferred:** moderate effort, small win,
and it interacts with the shape gate the memoized loader deliberately keeps hot.
**Effort:** M. **Priority:** P3.
- [ ] **Persistent-engine snapshot.** **What:** the snapshot fast-path only covers
in-memory engines (`!dataDir` gate at pglite-engine.ts). ~58 files pass
database_path and pay full cold init (~121s weighted). Needs tar-extract-into-
dataDir (or PGlite loadDataDir with a dataDir) design. **Effort:** M. **Priority:** P3.
- [ ] **Engine consolidation audit: doctor/bootstrap/migrations-v0_19_0.** **What:**
33 files construct 95 engines; chunk-grain-fts was consolidated in-pass, but
doctor.test.ts (9 engines), bootstrap.test.ts (9), migrations-v0_19_0.test.ts (7)
need a per-file audit — migration-from-old-schema tests structurally cannot share
a current-schema engine or use the snapshot. **Effort:** M. **Priority:** P3.
- [ ] **Verify per-check double-spawn removal.** **What:** each CHECKS entry costs a
`bun run <key>` startup before its bash script; invoking scripts directly from a
manifest would drop ~47 bun startups. **Why deferred:** micro-win; touches the
package.json-scripts-as-API convention. **Effort:** S. **Priority:** P3.
- [ ] **Snapshot-tar digest verification (defense-in-depth).** **What:** the CI
actions/cache for `test/fixtures/pglite-snapshot.tar` validates only the
schema-hash/dims lines in the sidecar `.version` — which travels in the SAME
cache entry, so both are forgeable together by anyone with cache write access.
Record a sha256 of the tar bytes in the version file at build time and have
`tryLoadSnapshot` verify it (mirror of the gitleaks fetch-fresh-digest
pattern). **Why deferred:** exploitability bounded by GitHub cache scoping
(fork caches isolated; poisoning needs push access) and impact is test-DB
contents only. **Effort:** S. **Priority:** P3.
- [ ] **Redact provider/DB strings in eval ledger writes.** **What:**
`EvalRunRecord.error` (free text) is persisted unredacted by
`persistRunRecord` (eval-run-all) and the canary's record mode into the now-
TRACKED `.gbrain-evals/eval-results.jsonl` — a failed keyed run whose error
embeds a connection string would ride a later commit into the public repo.
Route `record.error` + provider-derived params through
`redactConnectionInfo`/`redactPgUrl` before append; optionally add
`.gbrain-evals/` to the fixture-privacy scan surface. **Effort:** S.
**Priority:** P2.
- [ ] **check-image-decoders-embedded.sh into verify CHECKS.** **What:** the guard
runs its own `bun build --compile` (~60s) — too heavy per-verify. Revisit if the
binary-embed bug class recurs; guards-manifest.tsv carries the exemption note,
and the registration⇒execution coverage test allowlists it explicitly.
**Effort:** S. **Priority:** P3.
## Jobs fix-wave follow-ups (filed v0.45.15.0 — upstream issues #2/#3/#4)
- [ ] **P2 — `jobs submit --max-pending` public flag.** maxPending stays an
@@ -3011,8 +3106,13 @@ outside-voice triage on the reshaped plan.
- [ ] **v0.42+: ship the coordinated `gbrain-evals/baselines/v0.41-launch.baseline.ndjson`
+ `gbrain-evals/qrels/v0.41-launch.qrels.json` (hermetic-synthetic per D9).**
Generate locally via `gbrain bench publish --from <hermetic-test-corpus>` then
commit to the sibling gbrain-evals repo. Gives `gbrain eval gate` a canonical
baseline target so users don't have to bootstrap their own immediately.
commit to the sibling gbrain-evals repo. PARTIALLY SUPERSEDED by the test/eval/CI
speedup pass: an in-repo canonical qrels target now exists (`gbrain eval gate`
with the deterministic embedder option against `test/fixtures/eval-baselines/
qrels-search.json`; runner `scripts/run-eval-canary.ts`, CI-gated via
check:eval-canary, ledger `.gbrain-evals/eval-results.jsonl`). What remains
here is only the sibling-repo REGRESSION baseline (.baseline.ndjson for the
jaccard/top1 gate) — the correctness-gate half is done.
## v0.40.7.0 Schema Cathedral v3 follow-ups (v0.40.7+)
@@ -4088,7 +4188,12 @@ verify Voyage adapter integration in `src/core/ai/recipes/voyage.ts`).
## test infra (v0.26.4 follow-up — intra-file parallelism)
### Sweep cross-file shared-state contention; enable `bun test --concurrent` for another 2-3x speedup
**Priority:** P0
**Priority:** P3 (downgraded from P0 in the test/eval/CI speedup pass — premises stale:
the entry says "~58 PGLiteEngine instantiations", the suite now has 600+; the serial
quarantine grew from 4 files to ~140, and the pass's pooled serial runner + CI snapshot
+ verify pool delivered a comparable multiple for hours of work instead of the 1-2
weeks this sweep estimates. Re-scope against post-pass timing data before spending
anything here; `test.concurrent` adoption remains at zero.)
**Status:** v0.26.7 shipped foundation slice (helpers + lint + mock.module quarantine). v0.26.8 (env sweep) and v0.26.9 (PGLite sweep + codemod + measurement) carry the rest.
**What:** v0.26.4 shipped file-level parallel fan-out (8 shards) and got `bun run test` from 18 minutes to ~85s — a 12x speedup. The next layer is **intra-file** parallelism via Bun's `--concurrent` flag (or per-test `test.concurrent()` markers). This requires every test file to be safe under concurrent execution within the same `bun test` process.
@@ -6056,6 +6161,39 @@ respective shapes. Small, mechanical; pinned by `test/init-embed-check.test.ts`
(`skillpack status`/`sync`, doctor `skill_currency`) already keeps the brain's skill
set current on upgrade; this item is purely about semantic retrieval of skills.
## opencode wave follow-ups (filed at build time)
- [ ] **P2 — Watch the first opencode-door + canary dispatches.** The job is
day-one full posture (nightly + labels; keyless SMOKE + paid anthropic leg
on the existing secret) — after the wave merges, confirm the first nightly
run goes green end-to-end and the canary leg's latest-version result, then
update OPENCODE-CLI-PIN.md §Pending auth with anything the authed CI run
observes (exact `opencode models` output, per-turn cost note). Effort: S.
- [ ] **P3 — Wire opencode's plugin/event system** (the ambient-recall lane).
opencode ships a JS plugin system with lifecycle events; `OPENCODE_HAS_HOOKS
= false` in host-specs.ts marks the gap. Needs its own observation pass
(plugin API shapes, event timing, context-injection surface) before design —
would upgrade opencode from pull-protocol to per-turn push, above codex.
Effort: M/L.
- [ ] **P3 — BrainBench opencode adapter.** `src/eval/brainbench/adapters/` +
`ALL_HARNESSES` entry — build together with the already-filed hermes + grok
adapters (three pending; one eval wave). Effort: M.
- [ ] **P3 — connect `--agent opencode --oauth`.** opencode's `mcp auth` is an
authorization-code OAuth flow (not client-credentials) — a connect lane for
it needs the interactive-grant plumbing the current `--oauth`
(perplexity/generic client-credentials) path does not model. Effort: M.
- [ ] **P3 — Re-observe the OPENCODE_CONFIG* env trio on version bumps.**
Observed INERT in 1.18.18 (docs-contradiction pinned in OPENCODE-CLI-PIN.md
§Path seams); host-specs resolves via XDG only. If a future release
activates them, `opencodeConfigDir()` and the hermetic child-env deletes
must move together. The pin doc's re-observation checklist carries the
probe. Effort: S.
- [ ] **P3 — opencode-install PTY promotion.** Same criterion as grok-install:
2 consecutive stable dx-scenario runs ≥1 month apart with unchanged
boot/first-run copy → promote to a PTY assertion test. opencode's keyless
free tier means the scenario should COMPLETE the bootstrap, making it a
stronger promotion candidate than grok's sign-in-wall early-stop. Effort: M.
## Transcripts-import follow-ups (filed from cathedral-4, `gbrain transcripts ingest`)
Scoped OUT of the cathedral-4 PR by the CEO review's cherry-pick ceremony and the
@@ -6097,25 +6235,36 @@ covers DEAD logs; go-forward capture beyond Claude Code is deliberately absent.
CLAUDE_CODE.md + GROK.md now recommend `--surface verbs`. Update the
register one-liner + Direct config block (+ INSTALL_FOR_AGENTS hermes
block) and re-verify against the pinned hermes. Effort: S.
- [ ] **P2 — Backport the GITHUB_ENV/GITHUB_PATH/GITHUB_OUTPUT/GITHUB_STATE
- [x] **P2 — Backport the GITHUB_ENV/GITHUB_PATH/GITHUB_OUTPUT/GITHUB_STATE
deletion from `grokChildEnv` to `hermesChildEnv`** (and consider narrowing
the `GITHUB_` ALLOW_PREFIX to the read-only metadata names) — the prefix
rule forwards writable CI step-metadata files to untrusted agent children.
Unit truth-table exists for the grok side to clone. Effort: S.
DONE (opencode-support wave): `hermesChildEnv` now rides `makeAgentChildEnv`,
which scrubs the GITHUB_* step-metadata files for every door agent; truth-table
extended in `test/helpers/agent-harness.unit.test.ts`.
- [ ] **P3 — Grok bootstrap-harness target.** `gbrain bootstrap` personal-agent
support for Grok Build: `HarnessSelector` + `parseHarnessArgs`, a dated
`TARGETS` spec in `host-specs.ts`, a `wireGrok` branch + TOML writer (grok
config schema pinned; `codex-toml.ts` is the precedent), receipt/rollback/
status handling, and the INSTALL_FOR_AGENTS honest-classification flip.
Docs currently state "bootstrap does not support Grok yet". Effort: M.
- [ ] **P3 — Door-adapter extraction + CI-tail composite action.** Trigger: the
NEXT door agent (4th). Extract the agent-harness door family shape
(resolve/auth/seed/childEnv/pin/turn) and hoist the shared workflow tail
(evidence prep / scrub triple / upload / zero-pass grep / cred cleanup)
into a composite action; port grok-door as first consumer. Until then the
hermes-door/grok-door scrub blocks carry cross-reference comments. Also
adopt a door CADENCE policy: nightly for the newest/most-churning agent,
label-only after 2 stable monthly cycles per agent. Effort: M.
- [x] **P3 — Door-adapter extraction (test-side) + door cadence policy.**
Trigger FIRED at the 4th door agent (opencode, the opencode-support wave):
`makeBinaryResolver`/`makeAgentChildEnv`/`runOneShotSpawn` extracted in
`test/helpers/agent-harness.ts`, grok+hermes ported (hermes gained the
GITHUB_* scrub + bounded drain), opencode landed as first consumer; the
cadence policy is adopted in `docs/TESTING.md` (nightly for the newest
agent, label-only after 2 stable monthly cycles).
- [ ] **P3 — Door CI-tail composite action.** Trigger: the FIRST GREEN
grok-door AND opencode-door dispatches (workflow yaml cannot be proven
locally, and refactoring never-run jobs compounds risk — grok-door has
never dispatched: its XAI_API_KEY secret does not exist yet). Hoist the
shared workflow tail (evidence prep / scrub triple / upload / pass-count +
paid sentinels / version re-check / cred cleanup) from
hermes-door/grok-door/opencode-door into a composite action; port
opencode-door as first consumer (it is the freshest copy). Until then the
three doors' scrub blocks carry cross-reference comments. Effort: M.
- [ ] **P3 — Promote grok-install to a PTY assertion test.** Criterion: 2
consecutive stable runs ≥1 month apart of the dx scenario (pre-ship ritual
on grok-touching waves) with unchanged boot/sign-in copy. Would be the
@@ -6139,11 +6288,66 @@ covers DEAD logs; go-forward capture beyond Claude Code is deliberately absent.
registry shape generalizes. Unify into one data-driven table AFTER the
door-adapter extraction lands (earn it — don't freeze hermes-isms in).
Effort: L.
- [ ] **P3 — PIN-doc privacy-guard candidate.** GROK-CLI-PIN/HERMES-CLI-PIN
carry verbatim observation transcripts; consider extending check-privacy.sh
(or a dedicated check) to assert pin docs use `<tmp>`/placeholder paths and
never carry key material or account ids. Effort: S.
- [x] **P3 — PIN-doc privacy guard.** DONE (opencode-support wave):
`scripts/check-pin-doc-privacy.sh` (in `bun run verify` + guards-manifest,
fixture-tested) asserts every `docs/mcp/*-CLI-PIN.md` uses placeholder paths
and carries no key-shaped material or non-example emails.
- [x] **P3 — opencode-door npm view-vs-install TOCTOU.** DONE (adversarial-review
fix wave): the door job's install step is now pack-verify-install — `npm pack
<pkg>@<ver> --json` downloads the artifact and reports the integrity of the
BYTES written; both the wrapper and the platform payload are asserted against
their pins before `npm install -g ./opencode-ai-*.tgz` installs from the
verified local tarball (no fresh registry resolve of the name; the payload's
install-time fetch is npm-validated against the same byte-confirmed packument).
Verified locally on darwin-arm64 (wrapper integrity == pin; `--ignore-scripts`
breaks opencode's postinstall binary placement, so it is deliberately absent).
- [x] **P3 — `opencode mcp list` probe spawns project-config servers.** DONE
(adversarial-review fix wave): the user-scope probe spawns from a fresh EMPTY
mkdtemp cwd (no project config can load), project scope SKIPS the live probe
entirely with a printed note (parse-back is authoritative), and the probe now
holds the real process handle so the 20s timeout actually kills the child
(SIGTERM → SIGKILL) instead of abandoning it.
- [ ] **P3 — dedupe the opencode read→parse→classify dance.** The
read-config → parseOpencodeConfig → opencodeEntryKind sequence is spelled
three times (bootstrap.ts runHooks pre-check, harness.ts apply expectUrl
fallback, harness.ts remove ownership check); extract a
`classifyOpencodeEntryAt(path, name, expect)` helper and drop the
double-printed other-source warning (the caller AND the writer note it).
Effort: S.
## opencode adversarial-review fix-wave follow-ups (filed at fix time)
- [ ] **P2 — per-harness MCP-scope consent key.** An interview MCP_SCOPE answer
recorded for Claude Code (where 'project' is the privacy-SAFE default)
currently authorizes opencode's INVERTED-risk scopes without fresh
confirmation ('project' on opencode = committed file that auto-spawns on
every collaborator machine, no trust gate), and an ABSENT answer defaults
opencode to user-global exposure (any repo on the machine reaches the
brain). Design a harness-specific consent confirm — either per-harness
answer keys (MCP_SCOPE_OPENCODE) or a one-time "your recorded scope means
something riskier here — confirm" gate on the opencode lane. Relates to the
agent-bootstrap A8 consent-semantics TODO. Effort: M.
- [ ] **P3 — opencodeEntryKind remote ownership: normalize the url compare.**
Ownership uses exact string equality on the entry url vs the receipt/expect
url — trailing-slash and host-case variants misclassify in BOTH directions
(ours read as foreign → orphaned entry; a variant-url foreign endpoint
never matches, fine, but the asymmetry is accidental). Consider URL
normalization (scheme/host case-fold, trailing-slash) plus an
Authorization-shape check before comparing. Effort: S.
- [ ] **P2 — claw-test --live runners inherit real HOME/XDG.** The grok /
hermes / opencode --live runners run against the operator's real
HOME/XDG config surface and only WARN on a pre-existing global gbrain
entry; a scripted run can mutate or exercise the operator's live wiring.
Consider a fail-closed flag (refuse when a global gbrain registration
exists unless --allow-live-config) or hermetic-by-default across the
runner family. Effort: M.
- [ ] **P3 — fixed-name `.bak` parity: codex-toml.ts + hooks.ts writers.**
opencode-json.ts now takes UNIQUE `.bak-<hex>` backups per operation
(overlapping runs can't clobber each other's snapshot; harness restores
from the returned path and unlinks on success). The codex TOML writer and
the hooks settings writers still use fixed-name backups with the same
theoretical overlap window — port the unique-backup pattern (and the
restore-guard compare) for parity. Effort: S/M.
## Dream triage cascade follow-ups (#4152, filed at implementation)
- [ ] **P2 — Incremental submit-drain + deadline threading in synthesize
+1 -1
View File
@@ -1 +1 @@
0.46.3.0
0.46.6.0
+3
View File
@@ -27,6 +27,7 @@
"gray-matter": "^4.0.3",
"heic-decode": "^2.1.0",
"js-yaml": "^3.15.1",
"jsonc-parser": "^3.3.1",
"marked": "^18.0.2",
"openai": "^4.0.0",
"pgvector": "^0.2.0",
@@ -469,6 +470,8 @@
"json-schema-typed": ["json-schema-typed@8.0.2", "", {}, "sha512-fQhoXdcvc3V28x7C7BMs4P5+kNlgUURe2jmUT1T//oBRMDrqy1QPelJimwZGo7Hg9VPV3EQV5Bnq4hbFy2vetA=="],
"jsonc-parser": ["jsonc-parser@3.3.1", "", {}, "sha512-HUgH65KyejrUFPvHFPbqOY0rsFip3Bo5wb4ngvdi1EpCYWUQDC5V+Y7mZws+DLkr4M//zQJoanu1SP+87Dv1oQ=="],
"kind-of": ["kind-of@6.0.3", "", {}, "sha512-dcS1ul+9tmeD95T+x28/ehLgd9mENa3LsvDTtzm3vyBEO7RPptvAD+t44WVXaUjTBRcrpFeFlC8WCruUR456hw=="],
"libheif-js": ["libheif-js@1.19.8", "", {}, "sha512-vQJWusIxO7wavpON1dusciL8Go9jsIQ+EUrckauFYAiSTjcmLAsuJh3SszLpvkwPci3JcL41ek2n+LUZGFpPIQ=="],
+1
View File
@@ -103,6 +103,7 @@ Per-client setup guides live in [`docs/mcp/`](mcp/):
- [`docs/mcp/PERPLEXITY.md`](mcp/PERPLEXITY.md)
- [`docs/mcp/HERMES.md`](mcp/HERMES.md) — Hermes (Nous Research CLI)
- [`docs/mcp/GROK.md`](mcp/GROK.md) — Grok Build (xAI CLI)
- [`docs/mcp/OPENCODE.md`](mcp/OPENCODE.md) — opencode (opencode.ai / SST terminal agent)
- [`docs/mcp/OPENCLAW.md`](mcp/OPENCLAW.md) — OpenClaw (bundle plugin or stdio)
- [`docs/mcp/CLAUDE_COWORK.md`](mcp/CLAUDE_COWORK.md) — Claude Cowork (team plan)
- [`docs/mcp/DEPLOY.md`](mcp/DEPLOY.md) — production deploy patterns
+20 -11
View File
@@ -11,11 +11,11 @@ Six test command tiers, each with a clear scope:
| Command | What it runs | Wallclock | When to use |
|---|---|---|---|
| `bun run test` | Parallel unit-test fast loop. Sharded fan-out via `scripts/run-unit-parallel.sh` (default 4 shards — CPU-detected, clamped to a max of 8; 4 matches CI's fan-out and avoids PGLite WASM-init contention), then a serial pass over `*.serial.test.ts`. Excludes `*.slow.test.ts` and `test/e2e/*`. No pre-checks, no typecheck. Builds/refreshes the PGLite schema snapshot BEFORE the shard fan-out and exports `GBRAIN_PGLITE_SNAPSHOT` so PGLite-booting files restore a baked schema instead of replaying every migration (~10x wallclock on a full run; see "PGLite schema snapshot" below). Opt out: `GBRAIN_NO_SNAPSHOT=1`. Memory-safe by default: total concurrency (shards × intra-shard files) is capped to available memory at `GBRAIN_TEST_MEM_PER_FILE_MB` (default 1536 — a PGLite WASM instance) per concurrent file, and two phantom-failure classes are automatically re-run serially (the rescue pass): failures carrying the WASM out-of-memory signature, and shards killed externally (SIGTERM/SIGKILL well before the shard timeout — sibling workspaces' process cleanup, memory jetsam). Phantoms pass serially and the run goes green with an `oom_rescued` note; real failures fail again serially and stay red. Knobs: `GBRAIN_TEST_NO_MEM_ADAPT=1`, `GBRAIN_TEST_NO_OOM_FALLBACK=1`, `GBRAIN_TEST_MAX_CONCURRENCY` (intra-shard, default 4), `GBRAIN_TEST_SHARD_TIMEOUT` / `GBRAIN_TEST_SHARD_KILL_AFTER`, plus `--shards N` / `--max-concurrency N` / `--dry-run` script args. | a few minutes on a Mac dev box | Inner edit loop. Default. |
| `bun run verify` | CI's authoritative pre-test gate set, fanned out in parallel by `scripts/run-verify-parallel.sh`: the full `check:*` battery (privacy, jsonb, progress, source-id, test-isolation, wasm, …) plus `bun run typecheck`. The `CHECKS` array in that script is the single source of truth — CI literally calls `bun run verify` in a dedicated job. | ~16s (parallel; typecheck dominates) | Before pushing; before `/ship`. |
| `bun run test` | Parallel unit-test fast loop. Sharded fan-out via `scripts/run-unit-parallel.sh` (default 4 shards — CPU-detected, clamped to a max of 8; 4 matches CI's fan-out and avoids PGLite WASM-init contention), then a serial pass over `*.serial.test.ts`. Excludes `*.slow.test.ts` and `test/e2e/*`. No pre-checks, no typecheck. Builds/refreshes the PGLite schema snapshot BEFORE the shard fan-out and exports `GBRAIN_PGLITE_SNAPSHOT` so PGLite-booting files restore a baked schema instead of replaying every migration (~3.5x per booting file; see "PGLite schema snapshot" below). Opt out: `GBRAIN_NO_SNAPSHOT=1`. Memory-safe by default: total concurrency (shards × intra-shard files) is capped to available memory at `GBRAIN_TEST_MEM_PER_FILE_MB` (default 1536 — a PGLite WASM instance) per concurrent file, and two phantom-failure classes are automatically re-run serially (the rescue pass): failures carrying the WASM out-of-memory signature, and shards killed externally (SIGTERM/SIGKILL well before the shard timeout — sibling workspaces' process cleanup, memory jetsam). Phantoms pass serially and the run goes green with an `oom_rescued` note; real failures fail again serially and stay red. Knobs: `GBRAIN_TEST_NO_MEM_ADAPT=1`, `GBRAIN_TEST_NO_OOM_FALLBACK=1`, `GBRAIN_TEST_MAX_CONCURRENCY` (intra-shard, default 4), `GBRAIN_TEST_SHARD_TIMEOUT` / `GBRAIN_TEST_SHARD_KILL_AFTER`, plus `--shards N` / `--max-concurrency N` / `--dry-run` script args. | a few minutes on a Mac dev box | Inner edit loop. Default. |
| `bun run verify` | CI's authoritative pre-test gate set, fanned out by `scripts/run-verify-parallel.sh` through a bounded worker pool (default `detect_cpus`; override `GBRAIN_VERIFY_MAX_PARALLEL`) with the heavy checks ordered first (typecheck, the two compile-embed checks, admin build, fuzz bundles, guard self-tests, the PGLite-booting eval checks, whole-tree greps). The battery includes the deterministic `check:eval-chronicle` and `check:eval-canary` eval gates. The `CHECKS` array in that script is the single source of truth — CI literally calls `bun run verify` in a dedicated job. | ~40s (pool-bounded; longest check dominates) | Before pushing; before `/ship`. |
| `bun run test:full` | `verify && bun run test && bun run test:slow && [smart e2e]`. The local equivalent of "everything CI runs." Smart e2e: runs e2e only when `DATABASE_URL` is set; else loud skip notice to stderr. | ~3-5min depending on slow + e2e | Pre-merge sanity, before opening a PR. |
| `bun run test:slow` | Just the `*.slow.test.ts` set (intentional cold-path correctness checks). | seconds-to-minutes | When touching slow-path code. |
| `bun run test:serial` | Just the `*.serial.test.ts` set (cross-file-contention quarantine; one bun process per file for true module-registry isolation). | ~1s per quarantined file | Debugging a specific quarantined file. |
| `bun run test:serial` | Just the `*.serial.test.ts` set (cross-file-contention quarantine; one bun process per file for true module-registry isolation), run through a POOL of concurrent per-file processes — the isolation is per-process, not per-machine. Pool defaults to `min(detect_cpus, 4)` then memory-adapts (same doctrine as the parallel runner); a small growth-guarded set of files (machine-global state or contention-critical timing — see the justified `EXCLUSIVE_FILES` list in `scripts/run-serial-tests.sh`, capped at 3 by `test/scripts/serial-files.test.ts`) runs on a sequential EXCLUSIVE lane after the pool. Per-test timeout 120s (pooled contention headroom); each pooled file is wall-clock-killed at 300s (`timeout -k`, exit-hang containment). Externally-killed files (exit 143/137 or a missing exit sentinel — sibling-workspace cleanup, memory jetsam) get ONE sequential rescue re-run, mirroring the parallel runner's doctrine: phantoms stay green with a rescue note, real failures stay red. Prints per-file PASS lines plus a top-10 slowest-files list. Knobs: `GBRAIN_SERIAL_POOL=N` (explicit pool width — bypasses the memory clamp; `1` restores fully-sequential), `GBRAIN_SERIAL_FILE_TIMEOUT`. | ~2.5min for all ~140 files at pool=4 (was ~8.5min sequential) | Debugging quarantined files; CI's serial-tests job. |
| `bun run test:e2e` | Real Postgres E2E. Requires Docker + `DATABASE_URL`. Sequential. | ~5-10min | Pre-ship; nightly. |
There is no `check:all` script anymore — it was a second, hand-synced guard
@@ -32,11 +32,17 @@ self-test" below).
post-`initSchema()` PGLite data dir into `test/fixtures/pglite-snapshot.tar`
plus a version file; `PGLiteEngine.initSchema()` restores the tar instead of
replaying the embedded schema + all migrations when the env var
`GBRAIN_PGLITE_SNAPSHOT` points at it. Both `bun run test`
(`scripts/run-unit-parallel.sh`, before the shard fan-out) and
`scripts/ci-local.sh` call the builder unconditionally and export the env var.
Measured effect: a full parallel suite run drops ~10x (PGLite-booting files go
~1.63s → ~0.91s each). Properties:
`GBRAIN_PGLITE_SNAPSHOT` points at it. Runners activate it through the shared
`ensure_pglite_snapshot` helper in `scripts/lib/test-env.sh` (also home of
`detect_cpus` and `detect_available_mem_mb`), sourced by
`run-unit-parallel.sh`, `test-shard.sh`, `run-slow-tests.sh`,
`run-serial-tests.sh`, and `run-verify-parallel.sh`; `scripts/ci-local.sh`
calls the builder directly. The helper builds/refreshes the snapshot and
exports the env var, no-ops on `GBRAIN_NO_SNAPSHOT=1` or an already-inherited
path, and is non-fatal on build failure — tests fall back to cold init, with
a one-line "active" echo so a silent fallback stays visible in CI logs.
Measured effect: ~3.5x per PGLite-booting file (a cold boot replays every
migration, ~3.1s each on a CI shard). Properties:
- **Idempotent.** A hash short-circuit exits in ~40ms when the snapshot is
fresh, and REBUILDS a stale one. The hash covers `PGLITE_SCHEMA_SQL`, every
@@ -106,7 +112,7 @@ there even though they pass on Linux and macOS.
### CI vs local: intentionally divergent file sets
- **CI matrix** (`.github/workflows/test.yml`) runs `scripts/test-shard.sh` across 10 matrix shards partitioned by weight-aware LPT bin-packing (`scripts/sharding.ts`) and INCLUDES `*.slow.test.ts` (the two outlier slow files run as dedicated jobs alongside the matrix). CI EXCLUDES `*.serial.test.ts` from the shards and runs them in a dedicated job via `bun run test:serial`, one bun process per file — keeping serial files out of the shard processes is what preserves the `mock.module` quarantine (a top-level mock in one file leaks into every other file sharing its process). `bun run verify` gets its own job too, as does the BrainBench memory-conformance gate (`brainbench` job → `scripts/ci-brainbench-gate.sh`, hermetic in-memory PGLite, ~15s), which compares HEAD's fresh run against master's committed baseline (`evals/brainbench/baselines/main.json`) — the `test-status` aggregate checks its result explicitly. CI is the ground truth for "did everything pass."
- **CI matrix** (`.github/workflows/test.yml`) runs `scripts/test-shard.sh` across 10 matrix shards partitioned by weight-aware LPT bin-packing (`scripts/sharding.ts`; files with no mined weight fall back to the p75 file weight so a new unweighted file can't silently unbalance a shard) and INCLUDES `*.slow.test.ts` (the two outlier slow files run as dedicated jobs alongside the matrix) plus `evals/**/*.test.ts` (keyless-allowlist-gated — `test/scripts/evals-collection.test.ts`). Each shard's bun process is bounded by `--max-concurrency` (`GBRAIN_TEST_MAX_CONCURRENCY`, default 4). Every bun-test job — matrix shards, serial-tests, verify, the slow/eval jobs — activates the PGLite schema snapshot (built in-runner via `scripts/lib/test-env.sh`; the brainbench gate brings its own in-memory PGLite and skips it; the ~42MB tar is also cached across jobs via actions/cache, with the runner's own hash check staying authoritative). CI EXCLUDES `*.serial.test.ts` from the shards and runs them in the pooled `serial-tests` job via `bun run test:serial` — one bun process per file preserves the `mock.module` quarantine; the pool runs those processes concurrently. `bun run verify` gets its own job too, as does the BrainBench memory-conformance gate (`brainbench` job → `scripts/ci-brainbench-gate.sh`, hermetic in-memory PGLite, ~15s), which compares HEAD's fresh run against master's committed baseline (`evals/brainbench/baselines/main.json`) — the `test-status` aggregate checks its result explicitly. E2E (`.github/workflows/e2e.yml`) mirrors the content-hash skip in its own `e2e-pass-<hash>` namespace (scheduled nightly runs are exempt and always run), runs tier1 and tier2 in parallel with the jsonb-parity job in front of tier2 as the token-spend gate, and aggregates through `e2e-status`. CI is the ground truth for "did everything pass."
- **Local fast loop** (`scripts/run-unit-shard.sh` via the parallel wrapper) uses round-robin-by-index sharding and EXCLUDES `*.slow.test.ts` AND `*.serial.test.ts`. Local trades coverage for inner-loop speed; CI catches what local skips.
This divergence is intentional. Don't try to make them equal — the two scripts deliberately solve different problems. The regression test at `test/scripts/run-unit-shard.test.ts` pins what the local fast loop should and shouldn't include; `test/scripts/run-unit-parallel.test.ts` pins the wrapper's memory-adaptive concurrency and the OOM/external-kill serial rescue pass.
@@ -131,8 +137,8 @@ Triage rule: a `warn-pass` EXIT-HANG line in `.context/test-summary.txt` is NOT
- `*.test.ts` → fast loop (parallel up-to-4-shard fan-out, memory-adaptive).
- `*.slow.test.ts` → run via `bun run test:slow` only (intentional cold-path tests; would dominate the fast loop's wallclock).
- `*.serial.test.ts` → run via `bun run test:serial` after the parallel pass completes; one bun process per file (`--max-concurrency=1` within a shared process is not enough — the module registry still leaks `mock.module`). Quarantine for tests that share file-wide state and race when run alongside other files in the same `bun test` process. Several dozen files, discovered by the `*.serial.test.ts` glob — no list to maintain. Typical residents: `mock.module(...)` users (top-level mocks leak across files in a shard process, e.g. `test/embed.serial.test.ts`), env-coupled files (e.g. `test/brain-registry.serial.test.ts`), and process-lifecycle suites that assert on `process.exitCode` (e.g. `test/pglite-engine-disconnect.serial.test.ts`). **Do not put the parallelism back on a serial file unless you've fixed the contention root cause** (it just re-introduces the flake).
- `test/e2e/*.test.ts` → real-Postgres E2E. Skipped when `DATABASE_URL` is unset. One out-of-directory file rides this lane: `test/phantom-redirect-engine-parity.test.ts` (lives in `test/` for its PGLite arm, but its Postgres arm is only reachable through a DATABASE_URL-bearing lane — the unit wrappers strip the URL per #3485, so `run-e2e.sh`'s no-args list and CI's parity job carry it).
- `*.serial.test.ts` → run via `bun run test:serial` after the parallel pass completes; one bun process per file (`--max-concurrency=1` within a shared process is not enough — the module registry still leaks `mock.module`), with those per-file processes POOLED (per-process isolation never required one-at-a-time execution). Files touching machine-global state (launchd/cron) live on the sequential `EXCLUSIVE_FILES` lane inside `scripts/run-serial-tests.sh` — growth-guarded to ≤3 entries with justification comments. Quarantine for tests that share file-wide state and race when run alongside other files in the same `bun test` process. Several dozen files, discovered by the `*.serial.test.ts` glob — no list to maintain. Typical residents: `mock.module(...)` users (top-level mocks leak across files in a shard process, e.g. `test/embed.serial.test.ts`), env-coupled files (e.g. `test/brain-registry.serial.test.ts`), and process-lifecycle suites that assert on `process.exitCode` (e.g. `test/pglite-engine-disconnect.serial.test.ts`). **Do not put the parallelism back on a serial file unless you've fixed the contention root cause** (it just re-introduces the flake).
- `test/e2e/*.test.ts` → real-Postgres E2E. Skipped when `DATABASE_URL` is unset. One out-of-directory file rides this lane: `test/phantom-redirect-engine-parity.test.ts` (lives in `test/` for its PGLite arm, but its Postgres arm is only reachable through a DATABASE_URL-bearing lane — the unit wrappers strip the URL per #3485, so `run-e2e.sh`'s no-args list and CI's parity job carry it). `run-e2e.sh` wraps each file in a hard outer timeout (default 180s; `GBRAIN_E2E_FILE_TIMEOUT=<seconds>` overrides) because a synchronously-blocking PGLite WASM call can outlive bun's timer-based `--timeout`; LLM-bound Tier-2 files (`skills.test.ts`, `zeroentropy-live.test.ts`) automatically get 4× the cap since real provider round-trips legitimately run past 180s.
- `tests/heavy/*.sh` → ops-shape shell scripts. Cost minutes per run; NOT in default `bun test`. Run via `bun run test:heavy` or scheduled nightly via `.github/workflows/heavy-tests.yml`. Examples: pg_upgrade matrix (boot legacy brain → walk to head), RSS budget gate (measure peak worker RSS vs committed baseline), read-latency-under-sync (p50/p95/p99 under concurrent writer load), sync lock regression (N concurrent syncs assert 1 winner + N-1 lock-busy + zero leaked `gbrain_cycle_locks` rows). See `tests/heavy/README.md` for when to add a script here vs `*.slow.test.ts`. Files prefixed with `_` (e.g. `tests/heavy/_build_legacy_fixtures.sh`) are helpers/libs invoked by sibling tests — the runner skips them.
- `test/fuzz/*.test.ts` → property-based fuzz harness. Pure-validator targets in `pure-validators.test.ts` are guarded by `scripts/check-fuzz-purity.sh` (in `bun run verify`), which `bun build --target=bun` bundles each target and greps the resulting bundle for banned transitive imports (`node:fs`, `node:child_process`, engine modules). Anything that fails the guard moves to `mixed-validators.test.ts` (still property-tested, but no purity guarantee) or `filesystem-validators.test.ts` (fs-backed, uses temp dirs). Fuzz tests run in the default `bun test` loop because they're fast (~3s for ~12 properties × 1000 runs each).
@@ -420,6 +426,9 @@ E2E tests live in `test/e2e/` and run against real Postgres+pgvector (require `D
- `test/e2e/workspace-generic-compat.test.ts` — always-on (PGLite, no binary): pins the INSTALL_FOR_AGENTS.md "any repo with a workspace" contract against `test/fixtures/generic-agents-workspace/` (Hermes is the motivating consumer): `cwd_walk_up` detection, the `GBRAIN_SKILLS_DIR` override, `check-resolvable` on a root AGENTS.md, and scaffold additivity + refuse-overwrite. The real Hermes-behavior proof is the door suite below.
- `test/e2e/install-real-hermes.serial.test.ts` — the hermes "door": real `hermes` binary + real `hermes mcp add` handshake (full-catalog tool discovery; the count tracks the op catalog, so the test asserts discovery happened, not a number) + a paid `hermes -z` recall turn against a seeded brain. Triple-gated: `GBRAIN_REAL_HERMES_E2E=1` (explicit opt-in — run-e2e.sh scrubs GBRAIN_*, so it can never fire under `bun run test:e2e`) + resolvable binary + non-empty ANTHROPIC key (anthropic-pinned on purpose: a second provider key flips hermes provider-auto into a mis-routed 401). Hermetic HOME + HERMES_HOME with a tripwire on the operator's real config; evidence copies to `GBRAIN_E2E_EVIDENCE_DIR` for CI upload. Venue: heavy-tests.yml (`real-agent-e2e` + `hermes-door` jobs).
- `test/e2e/install-real-grok.serial.test.ts` — the grok "door" (xAI Grok Build; every asserted shape observed against the pin in `docs/mcp/GROK-CLI-PIN.md`). SPLIT-GATED, a deliberate divergence from the hermes door: grok's `mcp add/list/doctor` run keyless, so the compat tier (version-shape pin, documented-shape `grok mcp add gbrain -- gbrain serve --surface verbs` via a PATH-staged bin dir, saved-TOML asserts via `Bun.TOML.parse`, `mcp doctor` handshake proving the seven-verb surface, vendor-fallback provenance guard, direct-TOML surface) needs only `GBRAIN_REAL_GROK_E2E=1` + a resolvable binary; the paid SMOKE additionally needs a non-empty `XAI_API_KEY` and asserts a PER-RUN NONCE fact (grok has fs/shell tools — the committed fact is greppable, so recall of it proves nothing) with web search disabled. `mcp add` is lazy (exit 0 always) — `mcp doctor <name> --json` is the honest discriminator (exit 0/1 observed). Hermetic HOME + GROK_HOME + tmp cwd on every spawn (grok reads vendor MCP configs for trusted folders and loads `.envrc` from cwd); bounded tripwire over the operator's real `~/.grok` config/credential files (volatile paths excluded — grok rewrites logs/sessions/bin/docs every run) + a checkout guard that no `.grok/`/`.mcp.json` appeared in the repo root. Venue: heavy-tests.yml (`real-agent-e2e` + `grok-door` jobs); run directly via `GBRAIN_REAL_GROK_E2E=1 bun test test/e2e/install-real-grok.serial.test.ts`.
- `test/e2e/install-real-opencode.serial.test.ts` — the opencode "door" (SST opencode; every asserted shape observed against the pin in `docs/mcp/OPENCODE-CLI-PIN.md`). SPLIT-GATED a step past the grok door: opencode's anonymous FREE TIER drives MCP tool calls keyless, so even the nonce SMOKE runs in the keyless tier — T1 bare-semver version pin (the SST-vs-claimant discriminator), T2 documented-shape `opencode mcp add gbrain --env … -- gbrain serve --surface verbs` + the honest `opencode mcp list` discriminator (it SPAWNS every server; `✓/✗` text is the assertion surface — exit code is 0 even on failure, and `mcp debug` is OAuth-only), T2b spawn-gate CANARY (a project-config decoy is spawn-attempted with NO trust prompt — if this ever gates, the bootstrap user-global scope default's rationale changed: re-observe), T3 writer parity (gbrain's `opencode-json.ts` output handshakes through the real binary; cross-tool preservation both ways), T4 keyless SMOKE (per-run nonce + STRUCTURAL `gbrain_*` tool_use proof via `parseOpencodeJsonl`, `--format json`). The paid T5 anthropic leg additionally needs a non-empty `ANTHROPIC_API_KEY` and self-validates the pinned model id against the authed `opencode models` list BEFORE any spend. Hermetic HOME + both XDG dirs + tmp cwd on every spawn; `--pure` on every probe (`mcp list` autoloads plugins — a code-execution surface); bounded tripwire over the operator's real opencode configs/auth.json + a repo-root checkout guard. Venue: heavy-tests.yml (`real-agent-e2e` + `opencode-door` jobs, plus the schedule-only `opencode-door-canary` latest-version leg — continue-on-error, a pin-refresh signal, never a gate); run directly via `GBRAIN_REAL_OPENCODE_E2E=1 bun test test/e2e/install-real-opencode.serial.test.ts`.
**Door cadence policy** (adopted with the 4th door agent): the NEWEST door agent runs at nightly/schedule cadence (currently opencode, whose canary leg also tracks `latest`); a door drops to label-only (`real-agent-e2e`) after 2 stable monthly cycles with unchanged pins. Rationale: churn concentrates in the newest integration; steady-state doors pay for themselves on demand, not nightly.
- `test/helpers/tty-harness.ts` + `test/tty-harness.test.ts` — the DX real-PTY harness (`Bun.spawn({terminal:})`): pure text/timing helpers unit-tested with zero subprocesses, plus three live PTY smokes against `sh` guarded by `describe.skipIf(!ptySupported())`. The harness itself is a dev instrument surface — its consumer `scripts/dx-explore.ts` never runs in CI (transcripts land in gitignored `.context/dx-runs/`); see `docs/guides/bootstrap.md` for the scenario runbook.
- `test/e2e/search-swamp.test.ts` — reproduces the source-swamp case. Seeds a curated `originals/talks/article-outline-fat-code` page against two `<fork>/chat/` pages stuffed with the same multi-word phrase. Asserts the article wins keyword AND vector ranking, that `detail=high` lets the chat swamp re-surface, and that `source_id` passes through the two-stage CTE intact. PGLite in-memory.
- `test/e2e/search-exclude.test.ts``test/` + `archive/` pages hidden by default, `include_slug_prefixes` opts back in, caller-supplied `exclude_slug_prefixes` adds to defaults. Both keyword and vector search paths.
File diff suppressed because one or more lines are too long
+15
View File
@@ -64,6 +64,21 @@ CI. Production retrieval differs via the query cache, salience freshness,
expansion, etc. The gate measures retrieval quality with a fixed pipeline;
your users may see different results when the cache is warm.
For a fully hermetic run (CI canaries, keyless environments), add
`--embedder deterministic` to the correctness gate: query embeddings come
from the qrels fixture's basis-vector dims (`src/eval/deterministic-embed.ts`)
instead of the gateway, so the gate runs with no API keys and no network.
Correctness-gate-only — it is rejected together with `--baseline` (replay
re-embeds captured queries via the gateway) and requires `--qrels`. Bare
`hybridSearch` never reads or writes the semantic query cache, so a
deterministic run cannot poison cached production results. This is what CI's
`check:eval-canary` gate runs via `scripts/run-eval-canary.ts`: a throwaway
PGLite brain seeded from the qrels fixture, gating the hybrid ranking
pipeline (keyword/title/alias arms + RRF) with synthetic vectors. Honest
scope: semantic-embedding regressions remain the keyed eval suites' job.
Reproduce locally with `bun run scripts/run-eval-canary.ts` (`--record`
appends to the `.gbrain-evals/eval-results.jsonl` ledger).
### `.qrels.json` shape
Two equivalent representations per entry:
+8 -2
View File
@@ -51,8 +51,14 @@ Test infra: PGLite snapshot default-on for `bun run test`. Per-PGLite-file:
Full-suite wall-clock (post-snapshot): recorded in the W0 ship notes — see
the run banner of the W0 PR's `bun run test` evidence.
Retrieval canary: NOT RUN at W0 (production brain locked by live serve; W0
touches no search paths). REQUIRED before W1 lands.
Retrieval canary: PASS @ f2b40f7ef (hermetic deterministic-embedder CLI run;
recall@10=1.0000 first_relevant=1.0000 expected_top1=0.8333 vs floors
0.70/0.60/0.50; run `bun run scripts/run-eval-canary.ts` to reproduce, ledger:
.gbrain-evals/eval-results.jsonl). Honest scope: the canary gates the hybrid
ranking pipeline (keyword/title/alias arms + RRF against gold qrels) with
synthetic basis vectors — no API keys, no production brain, so the live-serve
lock is moot. Semantic-embedding regressions remain the keyed eval suites'
job. Wired into `bun run verify` as check:eval-canary.
Verified-bug status at W0 ship: cycle-lock refresh + fencing (TODO-OPS-2
closed), stall-death parent unblock, started_at ×4, modality carry, import
+1 -1
View File
@@ -32,7 +32,7 @@ pure win. See the per-verb latency table in
calls `context_pack` / `delta` over MCP (they are on `--surface verbs`) or the
CLI (`gbrain context-pack`, `gbrain delta`) at the boundary and injects the
returned `text` (or renders the structured arms). This is the portable path —
no hooks required. It is the primary path for Codex (which has no hooks) and
no hooks required. It is the primary path for Codex and opencode (no wired hooks) and
for Postgres brains (which have no local IPC socket).
- **Push (PGLite + Claude Code):** the bundled hook framework fires
automatically at `SessionStart` (injects a warm pack — including the
+20 -4
View File
@@ -1,7 +1,7 @@
# GBrain Bootstrap — your harness as your agent
`gbrain bootstrap` turns a Claude Code or Codex session into a persistent personal
agent: identity files rendered from your own answers, a local PGLite brain,
`gbrain bootstrap` turns a Claude Code, Codex, or opencode session into a
persistent personal agent: identity files rendered from your own answers, a local PGLite brain,
per-turn context, session-triggered schedules, and a private GitHub repo as the
agent's durable, portable body. This guide is the full contract — what gets
installed, what runs when, what it can and cannot do, and how to undo all of it.
@@ -19,7 +19,7 @@ follows is `BOOTSTRAP_FOR_AGENTS.md` at the repo root, fetched at the
| Identity files (SOUL/USER/MEMORY/AGENTS/CLAUDE/HEARTBEAT/ACCESS_POLICY/GITHUB) | your workspace folder | loaded at session start |
| `agent.json` manifest + `brain/`, `memory/`, `skills/`, `state/` | workspace | — |
| Local brain (PGLite) | `~/.gbrain/` (never in the repo) | while a session's MCP serve is open |
| MCP registration (`gbrain serve`) | Claude Code: project scope by default; Codex: user-global (no scope flag) | spawned by your harness per session |
| MCP registration (`gbrain serve`) | Claude Code: project scope by default; Codex: user-global (no scope flag); opencode: user-global by default (project scope is an explicit opt-in — see the degradation matrix) | spawned by your harness per session |
| Hooks (Claude Code, ON by default) | local installs: `.claude/settings.local.json` (gitignored); cloud sandboxes: the COMMITTED `.claude/settings.json` (PATH-resolved, fail-open commands) | each prompt; fail-open; `--no-hooks` opts out at install, `GBRAIN_HOOKS=0` disables at runtime |
| Per-turn persistence | Stop hook → debounced, detached scan-gated push (per workspace; 5 min default, every turn in cloud sandboxes) | after each assistant turn; `GBRAIN_STOP_PUSH=0` disables; `GBRAIN_STOP_PUSH_DEBOUNCE_MIN` / config `hooks.stop_push_debounce_min` tune it |
| Session persistence | SessionEnd hook → scan-gated commit+push | at session end (note: the harness never fires SessionEnd on `/exit` — the per-turn push is what covers that) |
@@ -155,6 +155,7 @@ you'd apply to any journal: write what you'd be comfortable persisting.
| GitHub / `gh` | full local agent | off-machine durability (repo re-runnable later) |
| Hooks (Claude Code) | pull protocol via AGENTS.md gates | automatic per-turn context + session-end persistence |
| Codex (no wired hooks, no MCP scope flag) | pull protocol + MCP tools | per-turn push (stated plainly; not oversold — codex 0.147+ ships a hook system, but gbrain does not wire it yet) + the ability to confine MCP reach to one folder (`codex mcp add` is always user-global) |
| opencode (no wired hooks; scope INVERTED: user-global by default) | pull protocol (opencode reads AGENTS.md natively) + MCP tools; project scope available as an explicit opt-in | per-turn push (opencode ships a plugin/event system, but gbrain does not wire it yet). The project-scope default is deliberately NOT offered: opencode spawns project-config servers with no trust prompt, so a committed entry would auto-execute on every collaborator machine |
| Second simultaneous session | first session unaffected | second session's brain tools fail politely (one live serve per brain — v1 contract) |
| Postgres brain (incl. harness mode) | MCP tools every session + pull protocol | per-turn hook injection (`no_pglite_path`: the hook IPC socket is PGLite-only today; hooks stay pre-wired and light up when the engine-uniform listener lands) |
@@ -188,6 +189,15 @@ mode wires them in one command, with no `agent.json` and no interview:
INLINE in the codex config (0600) — framework-spawned codex inherits no
shell profile, so the env-var lane the `connect` path uses would never
reach it.
- opencode: one managed `mcp.gbrain` remote entry with the bearer header
INLINE in the user-global JSONC config (0600), written by the same
comment-preserving editor the workspace lane uses — the `{env:…}`
interpolation the `connect` path prefers would resolve empty under a
framework-spawned opencode for the same no-shell-profile reason.
Note: downgrading gbrain below the release that introduced opencode support
after wiring it leaves the opencode entry in place for manual removal —
edit the opencode config by hand, or re-upgrade and run
`gbrain bootstrap harness --remove`.
- Honesty on Postgres brains: per-turn injection is degraded (the matrix row
above); MCP is the active seam and the summary says so.
- `--status [--json]` probes the live truth (serve health, token validity via
@@ -274,6 +284,11 @@ that changed shape, a harness that stopped calling our MCP server):
a seeded, brain-only fact (falling back to a shell `gbrain query` if headless
stdio-MCP is unavailable).
opencode's real-binary door lives in
`test/e2e/install-real-opencode.serial.test.ts` (its writer-parity leg
handshakes gbrain's direct JSONC registration through the actual binary);
`docs/TESTING.md` carries the full door inventory and cadence policy.
These pay real API cost and take 30s2min per turn, so they are NOT in the PR
shard. Everything is hermetic (temp `HOME` / `CODEX_HOME` / `CLAUDE_CONFIG_DIR` /
`GBRAIN_HOME` per test — the operator's real `~/.claude`, `~/.gbrain`, `~/.codex`
@@ -294,7 +309,7 @@ bun test test/e2e/bootstrap-real-codex.serial.test.ts
## DX exploration harness (developer instrument, not a test)
The door tests prove the install WORKS; they say nothing about how it FEELS.
`test/helpers/tty-harness.ts` spawns any CLI (gbrain, `claude`, `codex`, `grok`) under a
`test/helpers/tty-harness.ts` spawns any CLI (gbrain, `claude`, `codex`, `grok`, `opencode`) under a
real pseudo-terminal (Bun's `terminal:` spawn option) and records every output
burst with a millisecond timestamp, so unnecessary pauses become a measurable
artifact (`computeStalls``stalls.md`) instead of a vibe. Same hermetic env as
@@ -313,6 +328,7 @@ bun run scripts/dx-explore.ts help # comprehension surfaces (no key
bun run scripts/dx-explore.ts init [--keyless] # interactive init, naive-user autopilot
bun run scripts/dx-explore.ts claude-install # REAL claude running the paste-in bootstrap
bun run scripts/dx-explore.ts codex-install # REAL codex, same
bun run scripts/dx-explore.ts opencode-install # REAL opencode running the paste-in bootstrap
bun run scripts/dx-explore.ts grok-install # REAL grok, brain-only GROK.md install (no bootstrap path)
bun run scripts/dx-explore.ts drive -- gbrain init # manual: steer a live TUI via a file channel
```
+24 -6
View File
@@ -361,16 +361,34 @@ claimable work waits. The escalation commands and thresholds live in the
[queue operations runbook](queue-operations-runbook.md) — that's the
canonical home for wedge recovery.
What can still bite: a *brief* blip during a long-running job can make
lock renewal miss, and the stall detector dead-letters the job after
`max_stalled` misses (schema column default 5; lock duration and stall
check interval are both 30 s).
What can still bite is now narrow. Lock renewal is verify-before-evict:
a thrown or timed-out renewal is never treated as loss — at the deadline
the worker asks the database the authoritative question (one fenced
re-check), so a starved-but-healthy job recovers its lease and keeps
working. Eviction happens only on a fenced miss (the row was genuinely
reclaimed — requeued with no attempt burned) or after a hard backstop
(default 2× the lease) during a total outage. Long LLM handlers also get
a 300 s lock lease by default (`HANDLER_DEFAULT_LOCK_DURATION_MS`)
instead of the worker-global 30 s, and the stall sweep grants a 15 s
reclaim grace so a just-recovered worker's renewal beats the sweep.
The remaining exposure: a genuinely dead worker's long-lease job waits
up to lease + grace + one sweep interval before requeue, and the stall
detector still dead-letters after `max_stalled` genuine misses (schema
column default 5).
Mixed-version fleets degrade gracefully: an old worker ignores the
`lock_duration_ms` column and runs the legacy 30 s behavior; new workers
honor old rows via the claim-time default. No drain or ordered restart
is required.
**Tune per-job.** `gbrain jobs submit` accepts `--max-stalled N`,
`--backoff-type fixed|exponential`, `--backoff-delay <ms>`,
`--backoff-jitter 0..1`, and `--timeout-ms N` as first-class flags.
`--timeout-ms N`, `--lock-duration-ms N` (lock lease, clamped to
[5 s, 1 h]), and `--backoff-jitter 0..1` as first-class flags.
These write onto the job row at submit time — which is what
`handleStalled()` reads — so per-job tuning is the real knob.
`handleStalled()` and the renewal timer read — so per-job tuning is the
real knob. The lock-renewal env knobs (incident escape hatches) are
documented in the [queue operations runbook](queue-operations-runbook.md).
### DO NOT pass `maxStalledCount` to `MinionWorker`
+2 -2
View File
@@ -12,7 +12,7 @@ The push channels share one zero-LLM core (`src/core/context/volunteer.ts`):
| `reflex` | automatic, inside the context engine | default-on for plugin hosts; nothing to call |
| `op` | `gbrain volunteer-context` / MCP `volunteer_context` | agents without the plugin; one call per turn |
| `watch` | `gbrain watch` | stream a transcript in, volunteered pages stream out |
| `claude-code` / `codex` | `gbrain hook user-prompt` (registered by `gbrain bootstrap`) | per-prompt injection inside a harness; see "Harness hooks" below |
| `claude-code` / `codex` / `opencode` | `gbrain hook user-prompt` (registered by `gbrain bootstrap`) | per-prompt injection inside a harness; see "Harness hooks" below |
## How it decides
@@ -74,7 +74,7 @@ this channel production-grade rather than spammy-and-invisible:
- **The feedback loop.** The serve logs each DELIVERED block's volunteered
pages and pointers to `context_volunteer_events` under the hook's channel
(`claude-code` by default; a codex hook registration passes
`--harness codex`). `gbrain volunteer-context --stats` then shows
`--harness codex` / `--harness opencode`). `gbrain volunteer-context --stats` then shows
per-harness precision, and `gbrain doctor`'s `volunteer_channels` check
shows which channels actually fire, with guidance for the two quiet cases:
"hook installed but never registered (restart the session)" and "registered
+51
View File
@@ -107,6 +107,57 @@ gbrain jobs smoke --wedge-rescue
`queue.add()` call. If you want a taller pile, raise the threshold via
`GBRAIN_QUEUE_WAITING_THRESHOLD=50 gbrain doctor`.
## Lock-renewal: reading an eviction, and the knobs
Since v0.46 lock renewal is **verify-before-evict**: a thrown or timed-out
renewal is never treated as loss. At the deadline the worker runs one fenced
re-check against the database — the only CERTAIN loss signal is that fenced
miss. Every renewal fault also writes a JSONL audit event
(`~/.gbrain/audit/lock-renewal-*.jsonl`) carrying the fields that answer the
first incident question — *was the database down, or was the worker starved?*
How to read a `gave_up` / eviction line:
| Field | Reading |
|---|---|
| `cause` | `call-timeout` = our own timer fired (starved loop, slow pool, or slow DB); `refused` = the driver threw (SQLSTATE in `error_code`); `fenced-lost` = certain reclaim, not an infrastructure fault. |
| `lateness_ms` | How late the renewal tick fired vs its own cadence. Tens of seconds = the WORKER was starved (the #4145 shape); ~0 with `refused` = the database was actually unreachable. |
| `load1` / `cores` | Raw loadavg at event time, with core count for normalization. |
| `overlap_skips` | Ticks skipped because a prior renewal call was still in flight. |
| `deadline_deferred` | The soft deadline passed but the fenced verify was unreachable — the job was KEPT and retried (the fence is the backstop). |
| `event_loop_delay …` (log line) | p99/max event-loop delay since the last successful renewal — the direct starvation measurement. |
A `Job N did not exit within 30s of abort` line after an infrastructure
abort is NOT an orphan leak: the handler is cooperatively cancelling. The
line carries the same cause/lateness/load fields. Caveat: eviction is
cooperative — an abort-IGNORING handler keeps running past every bound and
can duplicate external side effects until it exits; the worker only frees
the slot.
Env knobs (incident escape hatches; all validated, warn-once on bad values;
defaults derive from the per-job lease):
| Env var | Default | What it does |
|---|---|---|
| `GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS` | `min(lease/3, 15s)` | Per-call budget for each renewal attempt (raced + best-effort cancelled). |
| `GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS` | `min(lease/6, 30s)` | Headroom before lease expiry; the fenced verify fires when the NEXT tick would land past `lease - margin`. |
| `GBRAIN_LOCK_RENEWAL_HARD_EVICT_MS` | `2 × lease` | Hard local backstop when even the verify is unreachable (total outage). Floored to the soft deadline. Setting it TO the soft deadline approximates the legacy abort-at-deadline behavior. |
| `GBRAIN_LOCK_RENEWAL_MAX_FAILURES` | 3 | Audit-event labeling only — never gates eviction. |
| `GBRAIN_MINION_STALL_RECLAIM_GRACE_MS` | 15000 | Stall-sweep reclaim grace: a lease that lapsed within this window is not reclaimed (starved-owner head start). `0` restores the legacy `lock_until < now()` predicate. Capped at 600000 (10 min, warn-once + clamp) — an oversized value would otherwise disable stalled-job recovery fleet-wide. |
Cross-knob invariants are enforced with warn-once clamps (margin < lease/2,
call timeout ≤ renewal cadence, hard evict ≥ soft deadline) — a
misconfigured knob can degrade cadence but cannot silently re-break the
deadline math.
Per-job lease: `gbrain jobs submit --lock-duration-ms N` (clamped
[5 s, 1 h] — enforced at submit, re-applied to the resolved lease at
claim, and backed by a database range constraint; `--dry-run` echoes the
clamped value that will actually be stored); long LLM handlers default to 300 s via
`HANDLER_DEFAULT_LOCK_DURATION_MS` in `src/core/minions/handler-timeouts.ts`.
Renewal cadence is `min(lease/2, 60 s)`. Trade-off: a genuinely dead
worker's long-lease job requeues after lease + grace + one sweep interval.
## Self-check: is a worker even running?
```bash
+5 -2
View File
@@ -12,7 +12,10 @@ file, the workflow pins, and the affected assertions together.
version stamp; CI installs the RELEASE TAG `v2026.8.3` = commit `3c27eb62` — the two
differ by post-release main commits, same declared version. If a CI door run ever
diverges from these notes, re-observe against the tag checkout.)
- Installer sha256: `c118ff31618dc70339049ce71061b8f1351a1c70d9c2a236ed50d8a2550c550d`
- Installer sha256: `868ed3a91e0fabbff6d7418b3ede82bf4833652ec4e77196a42852fb35a9e5b9`
(refreshed 2026-08-15: upstream installer drifted past the prior pin —
reviewed; the `--commit` payload-pin path the door depends on is intact,
and the payload pins (tag+commit) are unchanged)
(download https://hermes-agent.nousresearch.com/install.sh to a file first; verify; then run)
- Installer flags used: `--skip-setup --non-interactive`; binary lands at `~/.local/bin/hermes`
- Python 3.11.15 via uv
@@ -93,7 +96,7 @@ non-interactive. `hermes cron tick` = run due jobs once and exit. `hermes cron l
`git -C ~/.hermes/hermes-agent rev-parse HEAD` and loud-fails on any mismatch, so an
installer that silently ignores unknown flags (or a moved checkout layout) can never
run unpinned upstream code on a runner that later holds secrets.
- `HERMES_INSTALL_SHA256: "c118ff31618dc70339049ce71061b8f1351a1c70d9c2a236ed50d8a2550c550d"`
- `HERMES_INSTALL_SHA256: "868ed3a91e0fabbff6d7418b3ede82bf4833652ec4e77196a42852fb35a9e5b9"`
- Door test asserts `hermes --version` output contains `v$HERMES_VERSION` when the env var is set.
- `hermes --version` output shape: `Hermes Agent v0.20.0 (2026.8.3)` + install dir + python lines.
+252
View File
@@ -0,0 +1,252 @@
# opencode CLI pin — observed behavior notes (v1.18.18)
Dev-facing companion to [OPENCODE.md](OPENCODE.md): every fact below was OBSERVED
against a real hermetic install (2026-08-15, macOS arm64), not researched from docs.
The claw-test OpencodeRunner, the install door e2e, and the heavy-tests
opencode-door CI job assert exactly these shapes — when opencode releases change
them, update this file, the workflow pins, and the affected assertions together
(`scripts/check-opencode-pin.sh` in `bun run verify` enforces the workflow-side
match). Where an observation CONTRADICTS opencode's docs, the observation wins and
the contradiction is called out inline.
Naming note: **opencode** (SST, opencode.ai, npm `opencode-ai`) is not **OpenClaw**
(the agent platform gbrain ships a runner for) and not the original `opencode` CLI
that was renamed Crush — see Troubleshooting in OPENCODE.md for the binary-name
collision.
<!-- opencode-pin: distribution_kind=npm -->
<!-- opencode-pin: npm_package=opencode-ai -->
<!-- opencode-pin: npm_version=1.18.18 -->
<!-- opencode-pin: npm_integrity=sha512-J+5HFq8tf+wPBBpBpMPSNjSytF2/EkNWYfFZh4si1d9auFbQriqDyqZv+vFUsLWERfdMU32Eajwuiq3rKBvZLQ== -->
<!-- opencode-pin: npm_linux_x64_integrity=sha512-WmeUnhljYJ252wywKTiW4bNDzsas2njpjPUEh0jM6HKNI4vFxJtREtzaWViY4AKEAcOkLWT8Ll17ixvcHz3AnA== -->
<!-- opencode-pin: npm_linux_arm64_integrity=sha512-e8D3g0qJEIzawEg2+ygW3vkZjAYL2ssyAx4GbihjwXwZFvlZZy5zRWWzdz5KLBoHSTl0FB73vNtnNeXONyHpVQ== -->
<!-- opencode-pin: opencode_version=1.18.18 -->
<!-- opencode-pin: observed_date=2026-08-15 -->
## Pin
- **opencode v1.18.18**, `opencode --version` output shape: bare `1.18.18`
version only, NO binary-name prefix, NO build hash (unlike grok's
`grok 1.0.4 (hash)`). The door's T1 shape assert is `/^\d+\.\d+\.\d+$/` on the
trimmed output; SST identity is discriminated by the `mcp`+`debug` subcommands
existing (`opencode debug paths` exits 0 and prints the path table below —
the renamed-to-Crush ancestor and other claimants have neither).
- **Provisioning (CI + local): pinned npm, pack-verify-install**
`opencode-ai@1.18.18`, registry integrity `sha512-J+5HFq…`. The CI job
`npm pack`s the wrapper AND the runner's platform payload first (pack
reports the integrity of the bytes it actually downloaded — closing the
view-then-install TOCTOU), asserts both against the stamps above, then
installs FROM the verified local wrapper tarball; the install-time platform
sub-package fetch is validated by npm against the same packument integrity
the pack step just byte-confirmed. The wrapper fans out to per-platform
payloads (`opencode-{darwin,linux,windows}-{arm64,x64}[-baseline|-musl]`) as
optionalDependencies at the same version; the LINUX payload integrities are
pinned separately because the wrapper's integrity covers only the wrapper
tarball. Darwin arm64 payload observed at
`sha512-VkG+bz8u8Xqg9NzPK+2/71nEd4DKKlo2NLZurQ1eLAzDnmb1CMYZif/o6Shl8YFuTuYU/30k6yufl4Zr0Ij64g==`
(informational — the CI runners are linux). Same npm version-immutability
assumption as the grok pin, stated explicitly.
- A curl installer (`https://opencode.ai/install`) exists but is NOT the pinned
lane; npm is.
## Pin-refresh cadence (this CLI ships near-continuously)
opencode releases far faster than grok (patch releases near-daily). The pinned
lane is the deterministic gate; the **canary leg** in `opencode-door` (schedule-
scoped, `continue-on-error`, installs `opencode-ai@latest`) exists to surface
drift BEFORE it strands the pin. Policy: when the canary leg reds or the pin is
>6 weeks old, run the re-observation checklist (bottom) against latest, bump the
stamps + workflow env pins together, and note behavior deltas in this file.
Do not chase every patch release; refresh on canary signal or the 6-week clock.
## Path seams — XDG honored; OPENCODE_CONFIG* env vars are INERT (verified)
`opencode debug paths` is the authoritative dump. Observed under
`HOME=<tmp> XDG_CONFIG_HOME=<tmp>/.config XDG_DATA_HOME=<tmp>/.local/share`:
```
config <XDG_CONFIG_HOME>/opencode (opencode.json + opencode.jsonc)
data <XDG_DATA_HOME>/opencode (auth.json, opencode.db*, log/, repos/)
state <tmp>/.local/state/opencode (locks/)
cache <tmp>/.cache/opencode (bin/)
tmp /tmp/opencode
```
- **HOME + XDG_CONFIG_HOME/XDG_DATA_HOME redirection works fully on macOS**
(nothing was written outside the hermetic home across the whole observation
run). The door uses HOME + both XDG vars, belt-and-suspenders.
- **DOCS-CONTRADICTION: `OPENCODE_CONFIG`, `OPENCODE_CONFIG_DIR`, and
`OPENCODE_CONFIG_CONTENT` had NO observable effect on config resolution in
1.18.18** — probes registered via each were absent from `mcp list`, while the
XDG-resolved global config was still read. gbrain's path helpers therefore
resolve via XDG only and deliberately do NOT honor `OPENCODE_CONFIG*`;
re-observe on version bump (if a future release activates them, the helpers
and this section change together). Hermetic child envs still DELETE all three
(defense against a future release activating them).
- Volatile paths (tripwire exclusions): `opencode.db`, `opencode.db-shm`,
`opencode.db-wal`, `log/`, `repos/` under data; `locks/` under state; `bin/`
under cache. The tripwire hashes only `opencode.json(c)` + `auth.json`.
- Vendor quirk: opencode writes a `.gitignore` (node_modules, package.json, …)
into the CONFIG dir on first touch.
## Config format — JSONC everywhere, both filenames merge (verified)
- `~/.config/opencode/opencode.jsonc` AND `~/.config/opencode/opencode.json`
are BOTH read when both exist (servers from each appeared simultaneously in
`mcp list`) — merge, not first-wins. opencode's own `mcp add` writes the
`.jsonc` name.
- **Comments parse in `.json`-named files too** (a `// comment` inside project
`opencode.json` did not break resolution). JSONC is the effective grammar for
every config file regardless of extension → gbrain's writer treats all
opencode configs as JSONC (jsonc-parser surgical edits; comments survive).
- Project config: `opencode.json` in the project root is read (lookup traverses
up); a project-scope entry appears alongside global entries.
- Unknown keys inside an `mcp.<name>` entry are TOLERATED in 1.18.18 (an
`_gbrain` probe key neither errored nor hid the server). gbrain still does
NOT write marker keys — ownership is judged by structural fingerprint — so a
future strict-schema flip cannot brick a user's opencode.
- `opencode debug config` prints the resolved merge (rendering has a doubled-
line quirk; treat it as a debug view, not a parse surface).
## `opencode mcp add` — observed facts
- Shape: `opencode mcp add <name> [--env KEY=VALUE]... -- <command> [args...]`
(local) or `opencode mcp add <name> --url <URL> [--header KEY=VALUE]...`
(remote). The `-- command` form is real but UNDOCUMENTED in `--help` (the
help lists only `--url/--env/--header`; the error copy for a bare add says
`Provide either --url <url> or a command after --`).
- **Always writes the GLOBAL `opencode.jsonc`** — even when a project
`opencode.json` with an `mcp` table exists in the cwd. There is NO scope
flag. Project-scope registration requires writing the file directly (gbrain's
writer does).
- **Add is lazy**: exit 0, no spawn, no prompt — for unreachable URLs and
nonexistent commands alike. Never treat add's exit code as a handshake.
- **Rewrites preserve comments and foreign keys** (a seeded `// comment` and a
`theme` key survived a subsequent add) — opencode uses a JSONC-preserving
editor internally; gbrain's writer matches that bar.
- `--header` values are stored verbatim, including `{env:VAR}` interpolation
syntax (`Authorization=Bearer {env:GBRAIN_REMOTE_TOKEN}` round-trips).
## Saved config schema (verbatim, from real adds)
```jsonc
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"gbrain": {
"type": "local",
"command": ["gbrain", "serve", "--surface", "verbs"],
"environment": { "GBRAIN_SOURCE": "workspace", "GBRAIN_HOME": "/tmp/<brain-home>" }
},
"gbrain-remote": {
"type": "remote",
"url": "https://brain.example/mcp",
"headers": { "Authorization": "Bearer {env:GBRAIN_REMOTE_TOKEN}" }
}
}
}
```
`enabled` is optional (absent = enabled). `oauth` was not written by the CLI and
is omitted by gbrain's writer (no OAuth interference with bearer headers was
observed). Local commands: an absolute `command[0]` works; PATH-resolved bare
`gbrain` resolves via the SPAWNING process's PATH (the door verifies the staged
bin-dir prepend).
## Probes — `mcp list` is the honest discriminator; `mcp debug` is NOT
- **`opencode mcp list` SPAWNS every configured local server and connects every
remote one**, then prints per-server status: `✓ <name> connected` or
`✗ <name> failed` with a reason line (`Executable not found in $PATH:
"gbrain"`, `SSE error: …`). THE door's keyless handshake proof. Caveats:
**exit code is 0 even when servers fail** (parse the text, assert
`✓ gbrain connected`), output is clack-style UI with ANSI codes, and there is
no `--json`.
- **`mcp list` is also a code-execution surface**: it spawned a PROJECT-defined
`type:local` command from a fresh checkout with NO prompt and NO trust gate
(verified with a touch-file probe). Two consequences: (1) gbrain's
bootstrap default scope for opencode is USER-GLOBAL — a committed project
entry would auto-spawn on every collaborator's machine; (2) any gbrain-run
probe uses `--pure` (kills external plugin autoload) + `OPENCODE_DISABLE_AUTOUPDATE=1`.
- `opencode mcp debug <name>` is OAUTH debugging only — on a local server it
prints `MCP server <name> is not a remote server` and exits 0. Not a
discriminator.
- No tool-count line exists in `mcp list` (grok's `7 tools discovered` has no
analog); tool discovery is proven by the SMOKE turn's `tool_use` events
instead.
## One-shot (`opencode run`) — KEYLESS WORKS (anonymous free tier)
- `opencode run "<msg>"` prints the ANSWER TEXT ALONE on stdout; the session
banner (`> build · <model>`) and UI go to stderr. Exit 0 on success; exit 1
with a structured JSON error (`"ref": "err_…"`) on failure (e.g. bogus
model).
- **Keyless runs WORK**: with zero credentials and no auth.json, `run` answers
via opencode's anonymous free tier (default model observed:
`opencode/big-pickle`; `opencode models` lists 8 keyless `opencode/*` models,
most `-free` suffixed; `opencode stats` reports $0.00). There is no
`Not signed in` wall in headless run mode.
- **MCP tools fire in keyless run mode WITHOUT `--auto`** (verified: the free
model called `gbrain_recall` and returned a seeded per-run nonce with
`--auto` absent). `--auto` exists (`auto-approve permissions that are not
explicitly denied (dangerous!)`) but the door does not need or use it.
- MCP tool naming: `<server>_<tool>` (observed `gbrain_recall`).
- `--format json` emits NDJSON events, every event
`{type, timestamp, sessionID, part}`; types observed: `step_start`,
`tool_use`, `text`, `step_finish`. Tool events carry
`part: {type:"tool", tool:"gbrain_recall", callID, state:{status:"completed",
input:{…}, output:"<stringified JSON>"}}` — `parseOpencodeJsonl` pins this.
- Model flag: `-m/--model <provider/model>` (`opencode/big-pickle` confirmed;
paid ids follow models.dev convention — see Pending auth).
- Keyless SMOKE end-to-end (proven 2026-08-15): pinned opencode + free model +
real `gbrain serve --surface verbs` (7 verbs banner) recalled a per-run nonce
through MCP with zero credentials, keyless PGLite brain.
## Environment — detectHarness + child-env facts (verified)
- Inside `run`'s bash tool, opencode sets **`OPENCODE=1`** and `OPENCODE_PID`
in child processes → `gbrain bootstrap`'s `detectHarness()` probes
`OPENCODE`.
- Auto-update kill: `OPENCODE_DISABLE_AUTOUPDATE=1` env + `"autoupdate": false`
config — the door seeds BOTH; version stayed pinned across every observed
run. `opencode upgrade` is the manual updater.
- Rules files: project `AGENTS.md` is loaded; a sibling `CLAUDE.md` is NOT
double-loaded (nonce test: only the AGENTS.md nonce surfaced) — AGENTS.md
wins per level, exactly as documented. gbrain's rendered pull-protocol
contract works unchanged.
- `.well-known/opencode` remote config: never observed to fire in any CLI run
(docs list it atop the lookup order). No kill needed today; re-observe on
version bump.
## Auth (only needed for PAID providers)
- Anonymous free tier needs nothing on disk; `auth.json` is only created by
`opencode auth login` at `<XDG_DATA_HOME>/opencode/auth.json`
(`opencode providers`, alias `auth`, prints the path).
- The optional paid door leg gates on `ANTHROPIC_API_KEY` (env-only) and
self-validates the model id against the authed `opencode models` output
before spending.
## When the door goes red (triage)
| Failure class | Signature | Remediation |
|---|---|---|
| npm pin drift | install step: version/integrity mismatch | Re-pin deliberately: bump `npm_version`+`npm_integrity` (+ platform stamps), run the re-observation checklist, update workflow env pins (check-opencode-pin.sh enforces the pair) |
| canary leg red, pinned leg green | latest-version leg fails install/asserts | Upstream changed shape — schedule a pin refresh; pinned lane still gates |
| version drift mid-run | `opencode --version` re-check ≠ pinned | Auto-update engaged — verify BOTH kills (env + config seed); re-pin if deliberate |
| `✗ gbrain failed` in `mcp list` | `Executable not found in $PATH` / spawn error | Staged bin dir missing from PATH, or abs path wrong — registration bug, not opencode drift |
| free-tier drift | keyless SMOKE stops answering / new auth wall | Re-observe keyless posture; if the free tier is gated, flip the SMOKE to the ANTHROPIC leg and re-pin this section |
| paid leg: model id unknown | models-gate assert fails before any spend | Update the pinned anthropic model id from the authed `opencode models` output |
| tripwire fired | manifest mismatch on `opencode.json(c)`/`auth.json` only | True isolation breach — stop and inspect; volatile-path drift alone must NOT fire |
| real door regression | handshake or nonce assert fails, pins intact | Bisect against the pinned version; file upstream if opencode-side |
Re-observation checklist on a version bump: npm pin captures (§Pin), help-surface
diff (`--help`, `run --help`, `mcp --help`, `mcp add --help`), the
add → saved-config → `mcp list` sequence (§add/§Probes), the keyless `run`
posture (§One-shot — free tier presence, stdout purity, MCP-without---auto),
`debug paths`, and the `OPENCODE_CONFIG*` inertness probe (§Path seams). The
spawn-gate probe (§Probes) re-runs whenever release notes mention MCP trust or
permissions.
## Pending auth (requires ANTHROPIC_API_KEY; the core door does NOT)
Authed `opencode models` list + exact `anthropic/<model>` id confirmation,
one paid one-shot smoke + per-turn cost note, `auth.json` verbatim shape after
`opencode auth login` (feeds evidence exclusions + TTY secretPaths), and
whether the authed TUI first-run differs from the keyless one pinned in the
dx scenario. The opencode-door paid leg self-validates the model id before
spending, so these pins harden the door but do not block it.
## Supported-version policy
gbrain's opencode integration is verified against **opencode v1.18.18** (this
pin). The canary CI leg tracks latest (continue-on-error); the pinned lane is
the deterministic gate. Keyless free-tier behavior is a LOAD-BEARING
observation (the SMOKE rides it) — treat free-tier changes as pin-refresh
triggers, not flakes.
+175
View File
@@ -0,0 +1,175 @@
# Connect GBrain to opencode
> This page is the MCP-registration reference for **opencode** — the SST
> terminal coding agent (opencode.ai, npm `opencode-ai`; not OpenClaw, and not
> the original `opencode` CLI that was renamed Crush — see Troubleshooting).
> For the full brain install — CLI, engine, skills, dream cycle — follow
> [INSTALL_FOR_AGENTS.md](../../INSTALL_FOR_AGENTS.md) first; this page wires
> the finished brain into opencode over stdio MCP. opencode is a
> **bootstrap-supported harness**: `gbrain bootstrap hooks --harness opencode`
> registers the brain for you (and `gbrain connect --agent opencode` handles
> remote brains — see below) — the commands on this page are the standalone
> manual recipe. Bootstrap's own registration additionally pins the workspace
> source (`GBRAIN_SOURCE`) and the full op surface, so the two are not
> byte-identical.
opencode spawns `gbrain serve` as a local stdio subprocess. No server, no
tunnel, no token needed. Works with both PGLite and Supabase engines — and
because opencode natively reads `AGENTS.md`, a gbrain workspace's rendered
brain contract loads with zero extra configuration.
## Register (recommended)
```bash
opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs
```
`--surface verbs` exposes the seven-verb memory protocol (`recall`,
`remember`, `entity`, `synthesize`, `forget`, `context_pack`, `delta`
[MEMORY_VERBS v1](../protocol/MEMORY_VERBS_v1.md)) instead of the full
100+-op catalog — the recommended starting surface for coding agents.
Three facts about `opencode mcp add`, all observed:
- **The local-command form is `-- <command> [args...]` after the flags**
it's real but missing from `--help` (which shows only `--url/--env/--header`).
`--env` is repeatable, one `KEY=VALUE` per flag.
- **Registration is lazy.** The add writes config and exits 0 without
connecting — even for a nonexistent command. Verify with `opencode mcp list`
(below), never with the add's exit code.
- **It always writes the USER-GLOBAL config**
(`~/.config/opencode/opencode.jsonc`) — there is no scope flag. For a
project-scoped entry, write the project `opencode.json` directly (next
section) — but read the sharing warning first.
## Direct config (equally supported)
Global (`~/.config/opencode/opencode.jsonc`) or project (`opencode.json` in
the repo root — opencode's lookup traverses up to the git root):
```jsonc
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"gbrain": {
"type": "local",
"command": ["gbrain", "serve", "--surface", "verbs"],
"environment": { "GBRAIN_HOME": "/home/alice-example" },
"enabled": true
}
}
}
```
Comments are fine — opencode parses JSONC in both `.json` and `.jsonc` files,
and both filenames are read (merged) when both exist. To remove gbrain,
delete the entry, or set `"enabled": false` to disable without losing it.
**Sharing warning for project config:** opencode spawns project-defined local
MCP servers with **no trust prompt** — a committed `opencode.json` carrying a
gbrain entry executes on every collaborator's machine. Teammates without
gbrain get a failing spawn each session; teammates WITH gbrain attach their
own `host` brain to your repo's context. Prefer the user-global config (the
gbrain bootstrap default); if you do commit a project entry, use the
PATH-resolved `"gbrain"` command form (never an absolute path) and tell
collaborators `"enabled": false` is the opt-out.
## Verify
```bash
opencode mcp list # the real probe: SPAWNS the server
```
`opencode mcp list` performs the actual spawn + handshake for every
configured server — expect `✓ gbrain connected`. A broken registration shows
`✗ gbrain failed` with the reason (e.g. `Executable not found in $PATH`).
Because it spawns everything — including any project `opencode.json` entries
in your cwd, with no trust prompt — run it from a directory you trust
(gbrain's own bootstrap verification probe runs from an empty temp directory
for exactly this reason, and skips the live probe entirely for project-scoped
registrations).
Two caveats: the exit code is 0 even when servers fail (read the output, not
`$?`), and `opencode mcp debug` is OAuth-only diagnostics — it is NOT a
handshake probe for local servers. Then one real round-trip:
```bash
opencode run "use the gbrain recall tool to answer: what did I import most recently?"
```
`opencode run` (headless one-shot) prints the final answer alone on stdout
(UI goes to stderr). MCP tools work in run mode without any permission flags.
## Remote brains (`gbrain connect`)
For a brain served over HTTP on another machine:
```bash
gbrain connect https://your-host/mcp --token gbrain_xxx --agent opencode [--install]
```
Without `--install` it prints the config block to add; with `--install` it
writes the entry directly into the user-global config (no opencode binary
required — the JSONC write IS the registration) and smoke-tests the token.
Either way the config stores only the `{env:GBRAIN_REMOTE_TOKEN}`
interpolation — opencode resolves the env var at read time, so the token
never lands in the file. Export `GBRAIN_REMOTE_TOKEN` in your shell profile.
`--force` replaces a gbrain-managed entry whose endpoint moved (a rotated
serve); an entry gbrain didn't write is never replaced — pick another
`--name`. (Framework-spawned opencode inherits no shell profile;
`gbrain bootstrap harness --harness opencode` covers that case with an
inline-bearer entry written 0600.)
## Auth + model pin
- **Keyless works.** opencode ships an anonymous free tier (default model
`opencode/big-pickle` at observation time) — headless runs and MCP tool
calls work with zero credentials. For paid providers, export the provider
key (e.g. `ANTHROPIC_API_KEY`) or run `opencode auth login` (credentials
land in `~/.local/share/opencode/auth.json`).
- **Model pin:** pass `-m <provider/model>` per call, or set `"model"` in the
config. `opencode models` lists what your credentials can reach.
- **Updates:** opencode self-updates by default. For pinned/reproducible
environments, set BOTH `"autoupdate": false` in config AND
`OPENCODE_DISABLE_AUTOUPDATE=1` in the environment.
## Pair with cron
opencode has no built-in cron; schedule headless one-shots with your system
scheduler:
```bash
# crontab: brain maintenance every 4 hours
0 */4 * * * opencode run "Run gbrain sync and report anything unusual"
```
See [docs/guides/cron-schedule.md](../guides/cron-schedule.md) for the full
brain maintenance protocol (sync, embed, dream cycle).
## Troubleshooting
- **Wrong `opencode` on PATH** — the name has prior claimants (the original
`opencode` project was renamed Crush). The SST CLI answers
`opencode --version` with a bare semver (`1.18.18`) and has `opencode mcp`
+ `opencode debug paths` subcommands. Install it via
`npm install -g opencode-ai` or `curl -fsSL https://opencode.ai/install | bash`.
- **opencode ≠ OpenClaw** — opencode (opencode.ai / SST) is the terminal
agent this page covers; OpenClaw is the agent platform with its own gbrain
runner and docs ([OPENCLAW.md](OPENCLAW.md)).
- **`✗ gbrain failed — Executable not found in $PATH`** — the registered
command was the bare `"gbrain"` name and opencode's PATH doesn't carry it.
Use the absolute binary path in the user-global config, or fix PATH.
- **Registered but nothing changed mid-session** — opencode reads config at
session start; restart opencode (or start a new session) after registering.
- **`OPENCODE_CONFIG` seems ignored** — observed inert in v1.18.18: only
`HOME`/`XDG_CONFIG_HOME` move the config location. Don't rely on it.
- **Which config won?**`opencode debug config` prints the resolved merge;
`opencode debug paths` prints every directory opencode uses.
- **Rules files** — opencode loads the project `AGENTS.md` (a sibling
`CLAUDE.md` is NOT double-loaded; AGENTS.md wins). gbrain's rendered
workspace contract rides this natively.
---
Verified against **opencode v1.18.18** (fast-moving project — the pin is
enforced in CI, with a latest-version canary leg watching for drift).
Dev-facing observed-behavior notes (exact flag semantics, exit-code caveats,
config schema, CI pin values) live in [OPENCODE-CLI-PIN.md](OPENCODE-CLI-PIN.md).
+5
View File
@@ -69,6 +69,11 @@ codex mcp add gbrain -- gbrain serve --surface verbs
grok mcp add gbrain -e "GBRAIN_HOME=$HOME" -- gbrain serve --surface verbs
```
**opencode** (verify with `opencode mcp list` — the add is lazy, and list SPAWNS the server)
```bash
opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs
```
**OpenClaw / any stdio MCP host** — register the server command
`gbrain serve --surface verbs`. Remote brains: `gbrain serve --http` on the
host, then `gbrain connect https://host/mcp --token gbrain_xxx --install` on
+24 -3
View File
@@ -1259,10 +1259,25 @@ grok mcp add gbrain -e "GBRAIN_HOME=$HOME" -- gbrain serve --surface verbs
The add is lazy (exit 0 without connecting) — verify with
`grok mcp doctor gbrain`, which spawns the server and must report
`7 tools discovered`. This is the brain-only install; the `gbrain bootstrap`
personal-agent path does not support Grok yet (Claude Code/Codex only).
personal-agent path does not support Grok yet (Claude Code, Codex, and opencode only).
Verified against Grok Build v1.0.4. Full reference:
[docs/mcp/GROK.md](docs/mcp/GROK.md).
**If you are opencode** (the SST terminal agent, opencode.ai — not OpenClaw):
you are a bootstrap-supported harness — for the full persistent-personal-agent
install, follow `BOOTSTRAP_FOR_AGENTS.md` instead of this page. For the
brain-only MCP registration:
```bash
opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs
```
The add is lazy (exit 0 without connecting) — verify with `opencode mcp list`,
which spawns the server and must show `✓ gbrain connected` (the exit code is 0
even on failure; read the output). Restart opencode afterwards — it reads
config at session start. Verified against opencode v1.18.18. Full reference:
[docs/mcp/OPENCODE.md](docs/mcp/OPENCODE.md).
Whether you scaffolded or not, read `skills/RESOLVER.md` (in your workspace, or the
bundled copy at `~/gbrain/skills/RESOLVER.md` when running from the cloned repo). It's
the skill dispatcher — tells you which skill to read for any task. Save this to your
@@ -1787,6 +1802,7 @@ GBrain exposes nearly all of its 100+ operations as MCP tools (stdio and HTTP; a
- **[Cursor / Windsurf / any stdio MCP client](docs/mcp/CLAUDE_CODE.md)** — same shape, add `{"command": "gbrain", "args": ["serve"]}` to your MCP config.
- **[Hermes](docs/mcp/HERMES.md)** — `printf 'Y\n' | hermes mcp add gbrain --env GBRAIN_HOME=$HOME --connect-timeout 60 --command $(which gbrain) --args serve`. Keep `--args` last, and verify with `hermes mcp test gbrain` (the add exits 0 even on failure).
- **[Grok Build](docs/mcp/GROK.md)** — `grok mcp add gbrain -e "GBRAIN_HOME=$HOME" -- gbrain serve --surface verbs`. The add is lazy (exit 0 without connecting) — verify with `grok mcp doctor gbrain`, which spawns the server and reports `7 tools discovered`. Verified against Grok Build v1.0.4.
- **[opencode](docs/mcp/OPENCODE.md)** (opencode.ai / SST — not OpenClaw) — `opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs`, or let `gbrain bootstrap hooks --harness opencode` write the config for you (opencode is a bootstrap-supported harness — it reads AGENTS.md natively). The add is lazy — verify with `opencode mcp list`, which spawns the server (`✓ gbrain connected`). Remote: `gbrain connect https://your-host/mcp --token gbrain_xxx --agent opencode [--install]` — the config stores only the `{env:GBRAIN_REMOTE_TOKEN}` interpolation. Verified against opencode v1.18.18.
- **[OpenClaw](docs/mcp/OPENCLAW.md)** — the ClawHub bundle plugin registers gbrain automatically (`openclaw.plugin.json` ships in this repo), or add `{"command": "gbrain", "args": ["serve"]}` to `~/.openclaw/config.json`'s `mcpServers`.
- **[Claude Desktop (Cowork)](docs/mcp/CLAUDE_DESKTOP.md)** — Settings → Integrations → add the URL of your HTTP server. Remote only; the local `claude_desktop_config.json` does not work for remote servers.
- **[Claude Cowork (team plan)](docs/mcp/CLAUDE_COWORK.md)** — org Owner adds the connector under Organization Settings → Connectors.
@@ -3946,7 +3962,7 @@ The push channels share one zero-LLM core (`src/core/context/volunteer.ts`):
| `reflex` | automatic, inside the context engine | default-on for plugin hosts; nothing to call |
| `op` | `gbrain volunteer-context` / MCP `volunteer_context` | agents without the plugin; one call per turn |
| `watch` | `gbrain watch` | stream a transcript in, volunteered pages stream out |
| `claude-code` / `codex` | `gbrain hook user-prompt` (registered by `gbrain bootstrap`) | per-prompt injection inside a harness; see "Harness hooks" below |
| `claude-code` / `codex` / `opencode` | `gbrain hook user-prompt` (registered by `gbrain bootstrap`) | per-prompt injection inside a harness; see "Harness hooks" below |
## How it decides
@@ -4008,7 +4024,7 @@ this channel production-grade rather than spammy-and-invisible:
- **The feedback loop.** The serve logs each DELIVERED block's volunteered
pages and pointers to `context_volunteer_events` under the hook's channel
(`claude-code` by default; a codex hook registration passes
`--harness codex`). `gbrain volunteer-context --stats` then shows
`--harness codex` / `--harness opencode`). `gbrain volunteer-context --stats` then shows
per-harness precision, and `gbrain doctor`'s `volunteer_channels` check
shows which channels actually fire, with guidance for the two quiet cases:
"hook installed but never registered (restart the session)" and "registered
@@ -4500,6 +4516,11 @@ codex mcp add gbrain -- gbrain serve --surface verbs
grok mcp add gbrain -e "GBRAIN_HOME=$HOME" -- gbrain serve --surface verbs
```
**opencode** (verify with `opencode mcp list` — the add is lazy, and list SPAWNS the server)
```bash
opencode mcp add gbrain --env GBRAIN_HOME=$HOME -- gbrain serve --surface verbs
```
**OpenClaw / any stdio MCP host** — register the server command
`gbrain serve --surface verbs`. Remote brains: `gbrain serve --http` on the
host, then `gbrain connect https://host/mcp --token gbrain_xxx --install` on
+1 -1
View File
@@ -1,7 +1,7 @@
{
"id": "gbrain-context-engine",
"name": "gbrain",
"version": "0.46.3.0",
"version": "0.46.6.0",
"description": "Personal knowledge brain with Postgres + pgvector hybrid search",
"family": "bundle-plugin",
"configSchema": {
+8 -1
View File
@@ -51,6 +51,8 @@
"check:cli-exec": "bash scripts/check-cli-executable.sh",
"check:engine-dynamic-import": "bash scripts/check-engine-dynamic-import.sh",
"check:grok-pin": "bash scripts/check-grok-pin.sh",
"check:opencode-pin": "bash scripts/check-opencode-pin.sh",
"check:pin-doc-privacy": "bash scripts/check-pin-doc-privacy.sh",
"check:gateway-routed": "bash scripts/check-gateway-routed-no-direct-anthropic.sh",
"check:worker-pool-atomicity": "bash scripts/check-worker-pool-atomicity.sh",
"check:doc-history": "bash scripts/check-key-files-current-state.sh",
@@ -92,6 +94,10 @@
"check:operations-filter-bypass": "bash scripts/check-operations-filter-bypass.sh",
"check:fixture-privacy": "bash scripts/check-fixture-privacy.sh",
"check:conversation-parser": "bun src/cli.ts eval conversation-parser test/fixtures/conversation-formats/all.jsonl --no-llm",
"check:eval-chronicle": "bun src/cli.ts eval chronicle",
"check:eval-canary": "bun run scripts/run-eval-canary.ts",
"check:pagetype-exhaustive": "bash scripts/check-pagetype-exhaustive.sh",
"check:pg-url-redaction": "bash scripts/check-pg-url-redaction.sh",
"check:source-scope-onboard": "bash scripts/check-source-scope-onboard.sh",
"postinstall": "bun run scripts/postinstall.ts",
"prepublish:clawhub": "bun run build:all",
@@ -132,6 +138,7 @@
"gray-matter": "^4.0.3",
"heic-decode": "^2.1.0",
"js-yaml": "^3.15.1",
"jsonc-parser": "^3.3.1",
"marked": "^18.0.2",
"openai": "^4.0.0",
"pgvector": "^0.2.0",
@@ -157,7 +164,7 @@
"bun": ">=1.3.10"
},
"license": "MIT",
"version": "0.46.3.0",
"version": "0.46.6.0",
"overrides": {
"@hono/node-server": "^2.0.5",
"fast-uri": "^3.1.5",
+26 -15
View File
@@ -25,14 +25,16 @@
# (d) Phase-list check [D5]: every `Phase: <name>` in BOOTSTRAP_FOR_AGENTS.md
# must appear in src/core/bootstrap/status.ts (the TS phase list is the
# single source; the runbook defers to it). Skips while either is absent.
# (e) Harness-scoping counter-signal pins: the MCP-scope consent is Claude
# Code only (Codex has no scope flag — `codex mcp add` is user-global).
# Tripwires against accidental deletion of the load-bearing prose, not
# proofs of placement: the runbook must carry the Codex bullet's
# "Do NOT offer an MCP scope choice" and the phase-3 "Claude Code only"
# scoping; questions.json's MCP_SCOPE.question must START WITH
# "(Claude Code only". Intentional rewording updates these pins in the
# same commit. Skips while the runbook/bank are absent.
# (e) Harness-scoping counter-signal pins: the MCP-scope consent applies on
# Claude Code and opencode (Codex has no scope flag — `codex mcp add` is
# user-global; opencode DEFAULTS to user-global — no trust gate on
# project-config servers). Tripwires against accidental deletion of the
# load-bearing prose, not proofs of placement: the runbook must carry the
# Codex bullet's "Do NOT offer an MCP scope choice" and the phase-3
# "Claude Code and opencode" scoping; questions.json's MCP_SCOPE.question
# must START WITH "(Claude Code and opencode". Intentional rewording
# updates these pins in the same commit. Skips while the runbook/bank are
# absent.
#
# BSD/GNU grep portable (no \t escapes). Uses `bun` for JSON parsing — the
# check runs via `bun run verify`, so bun is always present.
@@ -195,7 +197,7 @@ else
echo "SKIP: phase-list check (runbook or src/core/bootstrap/status.ts absent)"
fi
# ── (e) harness-scoping counter-signal pins (MCP scope is Claude Code only) ─
# ── (e) harness-scoping counter-signal pins (scope = Claude Code + opencode) ─
if [ -f "$RUNBOOK" ]; then
if ! grep -qF 'Do NOT offer an MCP scope choice' "$RUNBOOK"; then
fail=1
@@ -204,11 +206,20 @@ if [ -f "$RUNBOOK" ]; then
echo " without this line, Codex-door agents re-ask a dead question." >&2
echo " Rewording intentionally? Update this pin in the same commit." >&2
fi
if ! grep -qF 'Claude Code only' "$RUNBOOK"; then
if ! grep -qF 'Claude Code and opencode' "$RUNBOOK"; then
fail=1
echo "FAIL: BOOTSTRAP_FOR_AGENTS.md lost the 'Claude Code only' scoping on the" >&2
echo " MCP-scope consent (phase 3). Without it the consent reads as" >&2
echo " harness-blind and Codex-door agents ask it." >&2
echo "FAIL: BOOTSTRAP_FOR_AGENTS.md lost the 'Claude Code and opencode' scoping" >&2
echo " on the MCP-scope consent (phase 3). Without it the consent reads as" >&2
echo " harness-blind: Codex-door agents ask a dead question and opencode" >&2
echo " agents miss the inverted (user-global) default." >&2
echo " Rewording intentionally? Update this pin in the same commit." >&2
fi
if ! grep -qF 'NO trust prompt' "$RUNBOOK"; then
fail=1
echo "FAIL: BOOTSTRAP_FOR_AGENTS.md lost the opencode spawn-gate rationale" >&2
echo " ('NO trust prompt'). Without it agents recommend the Claude-style" >&2
echo " project default on opencode — where a committed project entry" >&2
echo " auto-executes on every collaborator machine." >&2
echo " Rewording intentionally? Update this pin in the same commit." >&2
fi
else
@@ -216,9 +227,9 @@ else
fi
if [ -f "$QUESTIONS" ] && command -v bun >/dev/null 2>&1; then
if ! GBRAIN_QJSON="$QUESTIONS" bun -e \
'const fs=require("fs");let b;try{b=JSON.parse(fs.readFileSync(process.env.GBRAIN_QJSON,"utf8"));}catch(e){process.exit(1);}if(!b.questions){process.exit(1);}const e=b.questions.MCP_SCOPE;const q=(e&&e.question)||"";process.exit(q.startsWith("(Claude Code only")&&e.phase==="interview"?0:1);'; then
'const fs=require("fs");let b;try{b=JSON.parse(fs.readFileSync(process.env.GBRAIN_QJSON,"utf8"));}catch(e){process.exit(1);}if(!b.questions){process.exit(1);}const e=b.questions.MCP_SCOPE;const q=(e&&e.question)||"";process.exit(q.startsWith("(Claude Code and opencode")&&e.phase==="interview"?0:1);'; then
fail=1
echo "FAIL: questions.json MCP_SCOPE.question must start with '(Claude Code only'" >&2
echo "FAIL: questions.json MCP_SCOPE.question must start with '(Claude Code and opencode'" >&2
echo " AND MCP_SCOPE.phase must be 'interview' (the consent is recorded" >&2
echo " pre-confirm during the interview; a 'wire' phase re-creates the" >&2
echo " bank-vs-runbook contradiction). Also fails when the questions" >&2
+149
View File
@@ -0,0 +1,149 @@
#!/usr/bin/env bash
# scripts/check-opencode-pin.sh — opencode pin consistency guard.
#
# OPENCODE-CLI-PIN.md is the single observed-behavior source for the opencode
# integration; its pins fan out to the heavy-tests opencode-door job env, the
# OpencodeRunner argv, and the door e2e assertions. The prose rule is "update
# together" — this guard turns the workflow half of that rule into CI:
#
# 1. docs/mcp/OPENCODE-CLI-PIN.md carries a machine-stable stamp block
# (`<!-- opencode-pin: key=value -->`, one per line) including
# distribution_kind (npm | installer).
# 2. The opencode-door job env in .github/workflows/heavy-tests.yml must carry
# EXACTLY the pin set for the chosen distribution_kind:
# npm: OPENCODE_VERSION==opencode_version, OPENCODE_NPM_PACKAGE==npm_package,
# OPENCODE_NPM_INTEGRITY==npm_integrity; no OPENCODE_INSTALL_SHA256.
# installer: OPENCODE_VERSION==opencode_version,
# OPENCODE_INSTALL_SHA256==installer_sha256; no OPENCODE_NPM_INTEGRITY.
# (The pin DOC may document both — the fallback path stays written down;
# exclusivity is about which pins the WORKFLOW actually enforces.)
#
# Greps are anchored to the opencode-door job block so a future canary matrix leg
# (or a second door job) cannot satisfy the check by accident.
#
# SKIP-GRACEFUL: missing pin doc, missing workflow, or no opencode-door job yet →
# SKIP (exit 0), matching scripts/check-bootstrap-tag.sh. Test override:
# GBRAIN_OPENCODE_PIN_GUARD_ROOT points file resolution at a fixture tree.
# BSD/GNU portable (no \t escapes, no GNU-only flags).
set -euo pipefail
ROOT="${GBRAIN_OPENCODE_PIN_GUARD_ROOT:-$(cd "$(dirname "$0")/.." && pwd)}"
PIN_FILE="$ROOT/docs/mcp/OPENCODE-CLI-PIN.md"
WORKFLOW="$ROOT/.github/workflows/heavy-tests.yml"
if [ ! -f "$WORKFLOW" ]; then
echo "check-opencode-pin: SKIP (no $WORKFLOW)"
exit 0
fi
if ! grep -q '^ opencode-door:' "$WORKFLOW"; then
echo "check-opencode-pin: SKIP (no opencode-door job in heavy-tests.yml yet)"
exit 0
fi
# Once the opencode-door job EXISTS, a missing pin doc is a FAILURE, not a skip —
# deleting/renaming the doc must not silently disable the supply-chain gate.
if [ ! -f "$PIN_FILE" ]; then
echo "check-opencode-pin: FAIL — opencode-door job exists but $PIN_FILE is missing (the pin doc is the gate's source of truth)" >&2
exit 1
fi
fail() {
echo "check-opencode-pin: FAIL — $1" >&2
exit 1
}
# --- 1. Parse the stamp block ------------------------------------------------
stamp() {
# First occurrence wins; a missing stamp yields the empty string (callers
# decide whether that is a failure) — the `|| true` keeps set -e/pipefail
# from treating grep's no-match exit as a script error.
{ grep -E "^<!-- opencode-pin: $1=" "$PIN_FILE" || true; } | head -1 \
| sed -e 's/^<!-- opencode-pin: [a-z0-9_]*=//' -e 's/ -->$//'
}
# Duplicate stamps are drift bait (two values, which one is real?).
dupes=$({ grep -E '^<!-- opencode-pin: ' "$PIN_FILE" || true; } | sed -e 's/^<!-- opencode-pin: //' -e 's/=.*$//' | sort | uniq -d)
[ -n "$dupes" ] && fail "duplicate opencode-pin stamp(s) in OPENCODE-CLI-PIN.md: $dupes"
DIST_KIND=$(stamp distribution_kind)
OPENCODE_VERSION_PIN=$(stamp opencode_version)
[ -n "$DIST_KIND" ] || fail "OPENCODE-CLI-PIN.md is missing the distribution_kind stamp"
[ -n "$OPENCODE_VERSION_PIN" ] || fail "OPENCODE-CLI-PIN.md is missing the opencode_version stamp"
case "$DIST_KIND" in
npm|installer) ;;
*) fail "distribution_kind stamp must be npm or installer; got '$DIST_KIND'" ;;
esac
# --- 2. Extract the opencode-door job block --------------------------------------
# Jobs sit at 2-space indent; the block ends at the next 2-space-indented key.
job_block=$(awk '
/^ opencode-door:/ { f = 1; print; next }
f && /^ [A-Za-z0-9_-]+:/ { exit }
f { print }
' "$WORKFLOW")
[ -n "$job_block" ] || fail "could not extract the opencode-door job block"
wf_env() {
# Strip either quote style: a YAML-formatter pass flipping double to single
# quotes must not read as pin drift.
{ printf '%s\n' "$job_block" | grep -E "^ $1:" || true; } | head -1 \
| sed -e "s/^ $1:[[:space:]]*//" -e 's/^"//' -e 's/"$//' -e "s/^'//" -e "s/'\$//"
}
WF_VERSION=$(wf_env OPENCODE_VERSION)
WF_NPM_PACKAGE=$(wf_env OPENCODE_NPM_PACKAGE)
WF_NPM_INTEGRITY=$(wf_env OPENCODE_NPM_INTEGRITY)
WF_INSTALL_SHA=$(wf_env OPENCODE_INSTALL_SHA256)
[ -n "$WF_VERSION" ] || fail "opencode-door job env is missing OPENCODE_VERSION"
[ "$WF_VERSION" = "$OPENCODE_VERSION_PIN" ] || fail "OPENCODE_VERSION drift — workflow '$WF_VERSION' vs pin-doc stamp '$OPENCODE_VERSION_PIN' (update together; see the pin doc's re-observation checklist)"
# EVERY OPENCODE_VERSION: env line in the WHOLE workflow (the real-agent-e2e
# door job carries a second copy) must equal the stamp — bumping the door job
# alone must never pass green. Env keys sit at line start after indentation,
# so comments mentioning the name never match.
all_wf_versions=$({ grep -E '^[[:space:]]*OPENCODE_VERSION:' "$WORKFLOW" || true; } \
| sed -e 's/^[[:space:]]*OPENCODE_VERSION:[[:space:]]*//' -e 's/^"//' -e 's/"$//' -e "s/^'//" -e "s/'\$//")
for v in $all_wf_versions; do
[ "$v" = "$OPENCODE_VERSION_PIN" ] || fail "an OPENCODE_VERSION occurrence elsewhere in heavy-tests.yml ('$v') disagrees with the pin-doc stamp '$OPENCODE_VERSION_PIN' — every copy in the workflow moves with the stamp"
done
if [ "$DIST_KIND" = "npm" ]; then
NPM_PACKAGE_PIN=$(stamp npm_package)
NPM_INTEGRITY_PIN=$(stamp npm_integrity)
[ -n "$NPM_PACKAGE_PIN" ] || fail "distribution_kind=npm but OPENCODE-CLI-PIN.md is missing the npm_package stamp"
[ -n "$NPM_INTEGRITY_PIN" ] || fail "distribution_kind=npm but OPENCODE-CLI-PIN.md is missing the npm_integrity stamp"
[ -n "$WF_NPM_PACKAGE" ] || fail "distribution_kind=npm but the opencode-door job env is missing OPENCODE_NPM_PACKAGE"
[ -n "$WF_NPM_INTEGRITY" ] || fail "distribution_kind=npm but the opencode-door job env is missing OPENCODE_NPM_INTEGRITY"
[ "$WF_NPM_PACKAGE" = "$NPM_PACKAGE_PIN" ] || fail "OPENCODE_NPM_PACKAGE drift — workflow '$WF_NPM_PACKAGE' vs stamp '$NPM_PACKAGE_PIN'"
[ "$WF_NPM_INTEGRITY" = "$NPM_INTEGRITY_PIN" ] || fail "OPENCODE_NPM_INTEGRITY drift — workflow vs stamp mismatch"
# npm_version is a documented near-duplicate of opencode_version — assert they
# agree so bumping one alone can never pass green.
NPM_VERSION_PIN=$(stamp npm_version)
if [ -n "$NPM_VERSION_PIN" ] && [ "$NPM_VERSION_PIN" != "$OPENCODE_VERSION_PIN" ]; then
fail "npm_version stamp ($NPM_VERSION_PIN) disagrees with opencode_version stamp ($OPENCODE_VERSION_PIN) — update together"
fi
# Platform-payload integrity stamps (the door job byte-pins the linux
# sub-packages too): when the pin doc carries them, the job env must match.
X64_PIN=$(stamp npm_linux_x64_integrity)
if [ -n "$X64_PIN" ]; then
WF_X64=$(wf_env OPENCODE_NPM_LINUX_X64_INTEGRITY)
[ -n "$WF_X64" ] || fail "pin doc stamps npm_linux_x64_integrity but the opencode-door job env is missing OPENCODE_NPM_LINUX_X64_INTEGRITY"
[ "$WF_X64" = "$X64_PIN" ] || fail "OPENCODE_NPM_LINUX_X64_INTEGRITY drift — workflow vs stamp mismatch"
fi
ARM64_PIN=$(stamp npm_linux_arm64_integrity)
if [ -n "$ARM64_PIN" ]; then
WF_ARM64=$(wf_env OPENCODE_NPM_LINUX_ARM64_INTEGRITY)
[ -n "$WF_ARM64" ] || fail "pin doc stamps npm_linux_arm64_integrity but the opencode-door job env is missing OPENCODE_NPM_LINUX_ARM64_INTEGRITY"
[ "$WF_ARM64" = "$ARM64_PIN" ] || fail "OPENCODE_NPM_LINUX_ARM64_INTEGRITY drift — workflow vs stamp mismatch"
fi
[ -z "$WF_INSTALL_SHA" ] || fail "distribution_kind=npm but the opencode-door job also pins OPENCODE_INSTALL_SHA256 — one provisioning mode only (mode exclusivity)"
else
INSTALL_SHA_PIN=$(stamp installer_sha256)
[ -n "$INSTALL_SHA_PIN" ] || fail "distribution_kind=installer but OPENCODE-CLI-PIN.md is missing the installer_sha256 stamp"
[ -n "$WF_INSTALL_SHA" ] || fail "distribution_kind=installer but the opencode-door job env is missing OPENCODE_INSTALL_SHA256"
[ "$WF_INSTALL_SHA" = "$INSTALL_SHA_PIN" ] || fail "OPENCODE_INSTALL_SHA256 drift — workflow vs stamp mismatch"
[ -z "$WF_NPM_INTEGRITY" ] || fail "distribution_kind=installer but the opencode-door job also pins OPENCODE_NPM_INTEGRITY — one provisioning mode only (mode exclusivity)"
fi
echo "check-opencode-pin: ok ($DIST_KIND mode, opencode $OPENCODE_VERSION_PIN)"
+5 -2
View File
@@ -21,7 +21,10 @@ ROOT=$(cd "$(dirname "$0")/.." && pwd)
# - The redactor itself: src/core/url-redact.ts
# - Test fixtures that build redacted strings from full URLs
# - Documentation comments referring to the pattern
ALLOW_REGEX='url-redact\.ts|test/url-redact\.test\.ts|/\* allow-pg-url-literal \*/'
# The marker text is the exemption; its comment wrapper is not load-bearing
# (inside a /** block comment a literal `*/` would terminate the comment, so
# block-comment examples carry the bare marker).
ALLOW_REGEX='url-redact\.ts|test/url-redact\.test\.ts|allow-pg-url-literal'
# The pattern matches an unredacted Postgres URL appearing in a string
# literal, NOT preceded by `redactPgUrl(` or `***@`. We also match any
@@ -47,6 +50,6 @@ echo "ERROR: unredacted postgres:// URL found in source. Use redactPgUrl() befor
echo ""
echo "$FILTERED"
echo ""
echo "Allowed exemption: append \"/* allow-pg-url-literal */\" comment on the line"
echo "Allowed exemption: append an allow-pg-url-literal comment marker on the line"
echo "(only for fixtures and the redactor itself)."
exit 1
+70
View File
@@ -0,0 +1,70 @@
#!/usr/bin/env bash
# scripts/check-pin-doc-privacy.sh — PIN-doc privacy guard.
#
# The docs/mcp/*-CLI-PIN.md files carry VERBATIM observation transcripts from
# real installs (help output, saved configs, error copy). That verbatim
# discipline is the point — but it is also exactly how an operator path
# (/Users/<name>/…), a key fragment, or an account id ends up committed and
# shipped with every release. This guard asserts the placeholder discipline:
#
# 1. No operator home paths: /Users/<name>/ or /home/<name>/ must appear as
# placeholders (<tmp>, $HOME, ~/) — never as a real username path.
# Bare `~/.grok`-style spellings are fine (that IS the placeholder).
# 2. No key material: long high-entropy tokens with known prefixes
# (sk-…, xai-…, gbrain_<64+hex-ish>, ANTHROPIC/OPENAI/XAI key shapes).
# npm `sha512-…` integrity pins are EXPECTED content — excluded.
# 3. No obvious account ids: emails outside example.com/invalid domains.
#
# SKIP-GRACEFUL: no pin docs yet → SKIP (exit 0). Test override:
# GBRAIN_PIN_PRIVACY_GUARD_ROOT points file resolution at a fixture tree.
# BSD/GNU portable.
set -uo pipefail
ROOT="${GBRAIN_PIN_PRIVACY_GUARD_ROOT:-$(cd "$(dirname "$0")/.." && pwd)}"
shopt -s nullglob
PIN_DOCS=("$ROOT"/docs/mcp/*-CLI-PIN.md)
shopt -u nullglob
if [ "${#PIN_DOCS[@]}" -eq 0 ]; then
echo "check-pin-doc-privacy: SKIP (no docs/mcp/*-CLI-PIN.md yet)"
exit 0
fi
fail=0
for doc in "${PIN_DOCS[@]}"; do
rel="${doc#"$ROOT"/}"
# 1. Operator home paths (a real username after /Users/ or /home/).
hits=$(grep -nE '(/Users|/home)/[A-Za-z][A-Za-z0-9._-]+/' "$doc" || true)
if [ -n "$hits" ]; then
fail=1
echo "FAIL: $rel carries operator home path(s) — replace with <tmp>/\$HOME/~ placeholders:" >&2
printf '%s\n' "$hits" | sed 's/^/ /' >&2
fi
# 2. Key material. sha512- npm integrity pins are expected; exclude lines
# carrying them before scanning for long secret-shaped runs.
hits=$(grep -v 'sha512-' "$doc" | grep -nE '(sk-[A-Za-z0-9_-]{20,}|xai-[A-Za-z0-9_-]{20,}|gbrain_[A-Za-z0-9]{32,}|AKIA[0-9A-Z]{16})' || true)
if [ -n "$hits" ]; then
fail=1
echo "FAIL: $rel carries key-shaped material — redact before committing:" >&2
printf '%s\n' "$hits" | sed 's/^/ /' >&2
fi
# 3. Emails outside the documentation-safe domains.
hits=$(grep -nE '[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}' "$doc" \
| grep -vE '@(example\.(com|org|net)|[A-Za-z0-9.-]*invalid)' || true)
if [ -n "$hits" ]; then
fail=1
echo "FAIL: $rel carries a non-placeholder email address:" >&2
printf '%s\n' "$hits" | sed 's/^/ /' >&2
fi
done
if [ "$fail" -ne 0 ]; then
echo "check-pin-doc-privacy: FAIL (pin docs ship with every release — placeholder discipline is the privacy IRON RULE)" >&2
exit 1
fi
echo "check-pin-doc-privacy: ok (${#PIN_DOCS[@]} pin doc(s))"
+9 -1
View File
@@ -41,6 +41,14 @@ ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
cd "$ROOT"
TARGET_DIR="${1:-test}"
# When scanning the default root, also lint evals/**/*.test.ts — those files
# are collected into the CI matrix (scripts/test-shard.sh) and must obey the
# same isolation rules as everything else CI executes. An explicit TARGET_DIR
# argument (guard self-test fixtures) scans only itself.
EXTRA_DIRS=""
if [ "$TARGET_DIR" = "test" ] && [ -d evals ]; then
EXTRA_DIRS="evals"
fi
ALLOWLIST_FILE="$ROOT/scripts/check-test-isolation.allowlist"
# Read allowlist (one filename per line, # comments allowed). Empty file
@@ -72,7 +80,7 @@ is_allowlisted() {
# Find non-serial unit test files (excluding test/e2e). Portable across
# bash 3.2 (macOS default) and bash 4+; no mapfile.
FILE_LIST="$(find "$TARGET_DIR" -name '*.test.ts' \
FILE_LIST="$(find "$TARGET_DIR" $EXTRA_DIRS -name '*.test.ts' \
-not -name '*.serial.test.ts' \
-not -path "*/e2e/*" \
-type f 2>/dev/null | sort)"
+4 -2
View File
@@ -101,10 +101,12 @@ done
IFS='|' eval 'PATTERN="${PATTERN_PARTS[*]}"'
# Find tool.
# evals/ joins the scan: its *.test.ts files are collected into the CI
# matrix (scripts/test-shard.sh) and carry the same privacy bar.
if command -v rg >/dev/null 2>&1; then
matches="$(rg -niH --no-heading -t ts "$PATTERN" test 2>/dev/null || true)"
matches="$(rg -niH --no-heading -t ts "$PATTERN" test evals 2>/dev/null || true)"
elif command -v grep >/dev/null 2>&1; then
matches="$(grep -rniE --include='*.test.ts' "$PATTERN" test 2>/dev/null || true)"
matches="$(grep -rniE --include='*.test.ts' "$PATTERN" test evals 2>/dev/null || true)"
else
echo "check-test-real-names: ERROR: neither rg nor grep available." >&2
exit 2
+2 -2
View File
@@ -67,13 +67,13 @@ if grep -Eq 'lockTimer[[:space:]]*=[[:space:]]*setInterval\([[:space:]]*async' "
echo " routes through src/core/minions/lock-renewal-tick.ts:"
echo
echo " setInterval(() => {"
echo " if (tickInFlight) return;"
echo " if (tickInFlight) { state.overlapSkips += 1; return; }"
echo " tickInFlight = true;"
echo " void runLockRenewalTick(deps, state)"
echo " .then(handleResult)"
echo " .catch(handlePostError)"
echo " .finally(() => { tickInFlight = false; });"
echo " }, lockDurationMs / 2);"
echo " }, renewalIntervalMs); // min(lease/2, 60s)"
exit 1
fi
+4 -4
View File
@@ -11,12 +11,12 @@
# bash scripts/ci-local.sh --clean # nuke named volumes for cold debug
# bash scripts/ci-local.sh --no-shard # debug: run E2E sequentially against postgres-1 only
#
# 4-way E2E sharding: 4 pgvector services on host ports 5434-5437. The 36 E2E
# files split N/4 per shard; shards run in parallel. Within a shard, files run
# 4-way E2E sharding: 4 pgvector services on host ports 5434-5437. The test/e2e/ file set splits
# roughly N/4 per shard; shards run in parallel. Within a shard, files run
# sequentially (TRUNCATE CASCADE no-race property documented in run-e2e.sh).
# Wall-time on a 16-core host: ~6 min sequential -> ~1.5-2 min sharded.
#
# Stronger than PR CI: PR CI runs only Tier 1's 2 files; this runs all 36.
# Stronger than PR CI: PR CI runs a handful of named files across its tiers; this runs every test/e2e file.
set -euo pipefail
@@ -231,7 +231,7 @@ else
echo "$SELECTED" | tr " " "\n" | grep -v "^$" > /tmp/e2e-selected.txt
fi'
else
# Empty file -> run-e2e.sh uses default glob (all 36 E2E files).
# Empty file -> run-e2e.sh uses default glob (every test/e2e file).
DIFF_E2E_PREP='> /tmp/e2e-selected.txt'
fi
RUN_PHASES_CMD="echo \"[runner] guards + typecheck (run once before sharding)\"
+57
View File
@@ -28,6 +28,8 @@
* so the interview completes unattended. Pays real API cost;
* takes 10-25 min. Run in background and watch session/screen.txt.
* codex-install Same for REAL `codex` (interactive TUI).
* opencode-install Same for REAL `opencode` (bootstrap-supported; the keyless
* run rides the anonymous free tier and should COMPLETE).
* drive -- <cmd> Manual mode: spawn ANY command under the PTY and steer it
* across separate shell calls via a file control channel:
* watch: cat <dir>/session/screen.txt
@@ -45,6 +47,7 @@
* bun run scripts/dx-explore.ts init
* bun run scripts/dx-explore.ts claude-install
* bun run scripts/dx-explore.ts codex-install
* bun run scripts/dx-explore.ts opencode-install [--keyless]
* bun run scripts/dx-explore.ts drive [--no-hermetic-home] -- gbrain init
* Options: --dir <out> transcript dir (default .context/dx-runs/<scenario>-<ts>)
* --gbrain <bin> use an existing gbrain binary (default: compile+cache)
@@ -833,6 +836,59 @@ async function scenarioGrokInstall(ctx: ScenarioCtx, args: CliArgs): Promise<voi
});
}
// ── scenario: opencode-install ───────────────────────────────────────────────
async function scenarioOpencodeInstall(ctx: ScenarioCtx, args: CliArgs): Promise<void> {
const home = tmp(ctx, 'gb-dx-home-');
const gbHome = tmp(ctx, 'gb-dx-gbhome-');
const ws = tmp(ctx, 'gb-dx-ws-');
const binDir = stageBinDir(ctx);
// Hermetic HOME + BOTH XDG dirs (config/auth/data all move — observed
// v1.18.18, OPENCODE-CLI-PIN.md §Path seams), seeded with the config half
// of the double autoupdate kill; the env half rides the session env below.
const xdgConfig = path.join(home, '.config');
const ocCfgDir = path.join(xdgConfig, 'opencode');
fs.mkdirSync(ocCfgDir, { recursive: true });
fs.writeFileSync(
path.join(ocCfgDir, 'opencode.json'),
JSON.stringify({ $schema: 'https://opencode.ai/config.json', autoupdate: false }, null, 2) + '\n',
);
// Auth travels env-only for the anthropic leg; a login flow would persist
// auth.json — pre-register the known candidate for the scrub (rm of a file
// that never appears is a no-op).
ctx.secretPaths.push(path.join(home, '.local', 'share', 'opencode', 'auth.json'));
spawnSync('git', ['init', '-q', ws]);
spawnSync('git', ['-C', ws, 'config', 'user.email', 'dx@example.com']);
spawnSync('git', ['-C', ws, 'config', 'user.name', 'DX Explore']);
log('REAL interactive opencode running the paste-in bootstrap (opencode is a bootstrap-supported harness)');
log('keyless runs ride the anonymous free tier (observed) — the flow should COMPLETE keyless; a sign-in wall here is itself a pin-refresh signal');
await runInstallSession(ctx, {
argv: ['opencode'],
cwd: ws,
env: {
HOME: home,
XDG_CONFIG_HOME: xdgConfig,
XDG_DATA_HOME: path.join(home, '.local', 'share'),
OPENCODE_DISABLE_AUTOUPDATE: '1',
GBRAIN_HOME: gbHome,
PATH: `${binDir}:${process.env.PATH ?? ''}`,
// Never let a first-run bounce the OPERATOR's browser for sign-in.
BROWSER: '/usr/bin/false',
},
extraAllow: ['ANTHROPIC_API_KEY'],
// --keyless drops provider keys AFTER extraAllow re-admission — on
// opencode that measures the FREE-TIER path, not a wall (observed).
dropEnv: args.keyless ? PROVIDER_KEY_NAMES : undefined,
prompt: installPrompt(),
meta: {
promptDeviation: 'unattended persona appendix + local runbook path + preinstalled binary',
runbook: 'BOOTSTRAP_FOR_AGENTS.md (local — opencode is bootstrap-supported)',
},
});
}
// ── scenario: drive (manual control channel) ─────────────────────────────────
async function scenarioDrive(ctx: ScenarioCtx, args: CliArgs): Promise<void> {
@@ -919,6 +975,7 @@ const SCENARIOS: Record<string, { needsGbrain: boolean; run: (ctx: ScenarioCtx,
'claude-install': { needsGbrain: true, run: scenarioClaudeInstall },
'codex-install': { needsGbrain: true, run: scenarioCodexInstall },
'grok-install': { needsGbrain: true, run: scenarioGrokInstall },
'opencode-install': { needsGbrain: true, run: scenarioOpencodeInstall },
drive: { needsGbrain: true, run: scenarioDrive },
};
+8 -6
View File
@@ -17,15 +17,15 @@
# guard class selftest notes
check-no-double-retry.sh scanner yes regex hole fixed in W0 (could not match `() =>`); perl multi-line pass replaces never-installed pcregrep
check-jsonb-pattern.sh scanner yes nested-paren hole fixed in W0; safe ::text::jsonb spelling stays unflagged
check-jsonb-params.mjs scanner yes positional $N::jsonb AST-lite scanner; argv/env root override
check-jsonb-params.mjs scanner yes positional $N::jsonb AST-lite scanner; argv/env root override; not in verify CHECKS: exercised by its unit test + self-test fixtures
check-batch-audit-site.sh scanner todo
check-bun-test-timeout.sh scanner todo
check-bun-test-timeout.sh scanner todo not in verify CHECKS: runs directly as a test.yml verify-job step
check-fixture-privacy.sh scanner todo
check-no-legacy-getconnection.sh scanner todo was reachable from neither verify nor CI pre-W0 (check:all only)
check-no-pii-in-agent-voice.sh scanner todo
check-operations-filter-bypass.sh scanner todo
check-pagetype-exhaustive.sh scanner todo
check-pg-url-redaction.sh scanner todo
check-pagetype-exhaustive.sh scanner todo wired into verify CHECKS (v0.45.x test/eval/CI pass; was registered-but-never-executed)
check-pg-url-redaction.sh scanner todo wired into verify CHECKS (v0.45.x test/eval/CI pass; was registered-but-never-executed)
check-privacy.sh scanner todo
check-progress-to-stdout.sh scanner todo
check-proposal-pii.sh scanner todo
@@ -47,12 +47,12 @@ check-exports-count.sh scanner todo was reachable from neither verify nor CI pre
check-trailing-newline.sh scanner todo was reachable from neither verify nor CI pre-W0 (check:all only)
check-test-isolation.sh scanner todo allowlist data file: check-test-isolation.allowlist
check-admin-build.sh buildfresh exempt runs the admin build; the build is the test
check-admin-embedded.sh buildfresh exempt embed freshness diff
check-admin-embedded.sh buildfresh exempt embed freshness diff; not in verify CHECKS: duplicates check:admin-build's build
check-admin-scope-drift.sh buildfresh exempt regenerates + diffs
check-bootstrap-templates.sh buildfresh exempt regenerates template tree + diffs
check-eval-glossary-fresh.sh buildfresh exempt regenerates + diffs
check-fuzz-purity.sh buildfresh exempt executes fuzz corpus
check-image-decoders-embedded.sh buildfresh exempt binary embed check
check-image-decoders-embedded.sh buildfresh exempt binary embed check; not in verify CHECKS: own bun build --compile too heavy per-verify
check-pglite-embedded.sh buildfresh exempt binary embed check
check-skills-manifest-fresh.sh buildfresh exempt regenerates + diffs
check-tool-catalog-fresh.sh buildfresh exempt regenerates + diffs
@@ -61,3 +61,5 @@ check-bootstrap-tag.sh repostate exempt VERSION stamp drift check
check-cli-executable.sh repostate exempt file-mode check
check-no-tracked-symlinks.sh repostate exempt git index state check
check-grok-pin.sh repostate exempt pin-stamp drift check (GROK-CLI-PIN.md stamps vs heavy-tests grok-door env); own bun guard tests in test/check-bootstrap-guards.test.ts
check-opencode-pin.sh repostate exempt pin-stamp drift check (OPENCODE-CLI-PIN.md stamps vs heavy-tests opencode-door env); own bun guard tests in test/check-bootstrap-guards.test.ts
check-pin-doc-privacy.sh repostate exempt PIN-doc placeholder discipline (no operator paths/key material/emails in docs/mcp/*-CLI-PIN.md); own bun guard tests in test/check-bootstrap-guards.test.ts
1 # CI guard registry (W0 fix-wave, Tier-1 #11 / D5.14).
17 # guard
18 check-no-double-retry.sh
19 check-jsonb-pattern.sh
20 check-jsonb-params.mjs
21 check-batch-audit-site.sh
22 check-bun-test-timeout.sh
23 check-fixture-privacy.sh
24 check-no-legacy-getconnection.sh
25 check-no-pii-in-agent-voice.sh
26 check-operations-filter-bypass.sh
27 check-pagetype-exhaustive.sh
28 check-pg-url-redaction.sh
29 check-privacy.sh
30 check-progress-to-stdout.sh
31 check-proposal-pii.sh
47 check-trailing-newline.sh
48 check-test-isolation.sh
49 check-admin-build.sh
50 check-admin-embedded.sh
51 check-admin-scope-drift.sh
52 check-bootstrap-templates.sh
53 check-eval-glossary-fresh.sh
54 check-fuzz-purity.sh
55 check-image-decoders-embedded.sh
56 check-pglite-embedded.sh
57 check-skills-manifest-fresh.sh
58 check-tool-catalog-fresh.sh
61 check-cli-executable.sh
62 check-no-tracked-symlinks.sh
63 check-grok-pin.sh
64 check-opencode-pin.sh
65 check-pin-doc-privacy.sh
+79
View File
@@ -0,0 +1,79 @@
# scripts/lib/test-env.sh — shared helpers for the test-runner family
# (test-shard.sh, run-serial-tests.sh, run-slow-tests.sh, run-unit-parallel.sh,
# run-verify-parallel.sh). Source AFTER cd'ing to the repo root:
#
# . scripts/lib/test-env.sh
#
# bash 3.2 compatible (macOS system bash): no mapfile, no wait -n, no ${var^^}.
# Every helper degrades gracefully inside the script-sandbox tests
# (test/scripts/run-unit-parallel.test.ts symlinks a minimal PATH with no
# sysctl/nproc/vm_stat/timeout and no package.json).
# ──────────────────────────────────────────────────────────────────────────
# CPU detection: Apple Silicon perf cores → Mac total physical → nproc → 4.
# Returns a single positive integer.
# ──────────────────────────────────────────────────────────────────────────
detect_cpus() {
local n=""
n=$(sysctl -n hw.perflevel0.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
n=$(sysctl -n hw.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
n=$(nproc 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
echo 4
}
# ──────────────────────────────────────────────────────────────────────────
# Available-memory detection (MB). macOS: vm_stat free + inactive +
# speculative + purgeable pages (inactive/purgeable are reclaimable on
# pressure, which is exactly the scenario we size for). Linux: MemAvailable.
# Unknown platform → 0, and the caller skips adaptation entirely.
# ──────────────────────────────────────────────────────────────────────────
detect_available_mem_mb() {
if command -v vm_stat >/dev/null 2>&1; then
vm_stat 2>/dev/null | awk '
/page size of/ { psize = $8 }
/Pages free/ { free = $NF }
/Pages inactive/ { inactive = $NF }
/Pages speculative/ { spec = $NF }
/Pages purgeable/ { purge = $NF }
END {
gsub(/\./, "", free); gsub(/\./, "", inactive)
gsub(/\./, "", spec); gsub(/\./, "", purge)
if (psize == 0) psize = 16384
printf "%d\n", (free + inactive + spec + purge) * psize / 1048576
}'
return
fi
if [ -r /proc/meminfo ]; then
awk '/MemAvailable/ { printf "%d\n", $2 / 1024; found = 1 } END { if (!found) print 0 }' /proc/meminfo
return
fi
echo 0
}
# ──────────────────────────────────────────────────────────────────────────
# PGLite schema snapshot: build (idempotent, ~40ms when fresh; mkdir-lock
# concurrency-safe; hash folds handler-migration source) and export
# GBRAIN_PGLITE_SNAPSHOT for child bun processes. 500+ test files each
# cold-boot PGLite + replay every migration without it (~3.5x per booting
# file — see docs/TESTING.md).
#
# No-op when GBRAIN_NO_SNAPSHOT=1 or when a parent runner already exported
# the path (double-building is harmless but noisy). Non-fatal on build
# failure — tests fall back to cold init. The one-line "active" echo makes
# a silent fall-back-to-cold-init regression visible in CI logs.
# $1: label for log lines (defaults to test-env).
# ──────────────────────────────────────────────────────────────────────────
ensure_pglite_snapshot() {
local label="${1:-test-env}"
[ "${GBRAIN_NO_SNAPSHOT:-0}" = "1" ] && return 0
if [ -n "${GBRAIN_PGLITE_SNAPSHOT:-}" ]; then
echo "[$label] PGLite snapshot active (inherited): $GBRAIN_PGLITE_SNAPSHOT" >&2
return 0
fi
if bun run build:pglite-snapshot >/dev/null 2>&1; then
export GBRAIN_PGLITE_SNAPSHOT=test/fixtures/pglite-snapshot.tar
echo "[$label] PGLite snapshot active: $GBRAIN_PGLITE_SNAPSHOT" >&2
else
echo "[$label] snapshot build failed (non-fatal) — tests run with cold init" >&2
fi
}
+30 -11
View File
@@ -87,13 +87,16 @@ mkdir -p "$E2E_TMP_HOME/.gbrain"
# internals survive untouched. We keep GBRAIN_HOME (just set above for HOME
# isolation); everything else GBRAIN_* is an operator override the suite must
# not inherit — which also scrubs GBRAIN_REAL_HERMES_E2E and
# GBRAIN_REAL_GROK_E2E, so the paid hermes/grok door suites structurally
# GBRAIN_REAL_GROK_E2E / GBRAIN_REAL_OPENCODE_E2E, so the real-agent door
# suites structurally
# cannot fire under this runner (their venue is heavy-tests.yml's direct bun
# test). GROK_ also drops an operator's GROK_BIN/GROK_HOME. Adapts GStack's
# test). GROK_ also drops an operator's GROK_BIN/GROK_HOME; OPENCODE_ drops
# OPENCODE_BIN and the OPENCODE_CONFIG* trio. Adapts GStack's
# buildHermeticEnv() allowlist to gbrain's shell E2E runner.
for _e2e_var in $(env | grep -oE '^(CONDUCTOR_|MCP_|OPENCLAW_|HERMES_|GROK_|GBRAIN_)[A-Za-z0-9_]*' | sort -u); do
for _e2e_var in $(env | grep -oE '^(CONDUCTOR_|MCP_|OPENCLAW_|HERMES_|GROK_|OPENCODE_|GBRAIN_)[A-Za-z0-9_]*' | sort -u); do
case "$_e2e_var" in
GBRAIN_HOME) ;; # required for HOME isolation (set above) — keep
GBRAIN_PGLITE_SNAPSHOT) ;; # snapshot fast-path fixture (exported by ci-local.sh / runners) — keep
GBRAIN_TEST_ALLOW_DATABASE_URL) ;; # #3485 preload opt-in (set above) — keep
GBRAIN_E2E_ALLOW_DB) ;; # #3485 name-floor opt-in — the guard's own error
# message tells operators to set it; stripping it
@@ -183,16 +186,32 @@ for f in "${files[@]}"; do
if [ -n "${DATABASE_URL:-}" ]; then
psql "$DATABASE_URL" -At -c "SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE pid != pg_backend_pid() AND datname = current_database()" >/dev/null 2>&1 || true
fi
# Hard outer timeout (180s per file). bun's --timeout covers tests AND
# hooks (measured on 1.3.14), but it's timer-based: a PGLite WASM call
# that blocks the event loop synchronously never lets the timer fire and
# the file wedges indefinitely. gtimeout/timeout SIGKILLs the file so the
# suite advances. gtimeout (macOS via coreutils) preferred; timeout (Linux)
# fallback; bare bun (no outer cap) if neither is installed.
# Hard outer timeout (default 180s per file; GBRAIN_E2E_FILE_TIMEOUT
# overrides). bun's --timeout covers tests AND hooks (measured on 1.3.14),
# but it's timer-based: a PGLite WASM call that blocks the event loop
# synchronously never lets the timer fire and the file wedges indefinitely.
# gtimeout/timeout SIGKILLs the file so the suite advances. gtimeout (macOS
# via coreutils) preferred; timeout (Linux) fallback; bare bun (no outer
# cap) if neither is installed.
#
# LLM-bound Tier-2 files (real provider round-trips when .env.testing
# carries keys) legitimately run past 180s — the ingest skill alone has
# been observed at ~131s — and were being SIGKILLed mid-run with no
# assertion output, which reads like a mystery failure. CI runs those
# files in their own job WITHOUT this wrapper (see .github/workflows/
# e2e.yml tier2), so the cap only ever bit local runs: give them 4x.
file_timeout="${GBRAIN_E2E_FILE_TIMEOUT:-180}"
# Digits-only validation (same strict positive-int posture as the TS env
# knobs): a malformed value falls back to the default instead of
# word-splitting into extra gtimeout arguments or breaking the 4x math.
case "$file_timeout" in ''|*[!0-9]*) file_timeout=180 ;; esac
case "$f" in
*/skills.test.ts|*/zeroentropy-live.test.ts) file_timeout=$((file_timeout * 4)) ;;
esac
if command -v gtimeout >/dev/null 2>&1; then
TIMEOUT_CMD="gtimeout 180"
TIMEOUT_CMD="gtimeout $file_timeout"
elif command -v timeout >/dev/null 2>&1; then
TIMEOUT_CMD="timeout 180"
TIMEOUT_CMD="timeout $file_timeout"
else
TIMEOUT_CMD=""
fi
+320
View File
@@ -0,0 +1,320 @@
/**
* scripts/run-eval-canary.ts hermetic CLI retrieval-quality canary.
*
* Boots a throwaway PGLite brain under a temp GBRAIN_HOME, seeds the qrels
* fixture corpus, then spawns the REAL gbrain CLI to run the qrels
* correctness gate with the deterministic embedder (basis-vector query
* embeddings). No API keys, no network, no writes to the personal brain,
* no writes to tracked files in check mode.
*
* Seeding is the "V2" shape (feasibility-spike finding, BINDING): the
* expected-top1 page carries its query text in the `timeline` column too,
* because page-grain FTS (`pages.search_vector`) indexes title(A) +
* timeline(C) ONLY `compiled_truth` is deliberately unindexed. With
* V1-style seeding (query text in compiled_truth only) the title arm votes
* only for the sibling page and expected_top1 is structurally 0.0.
*
* Modes:
* default check mode (CI): assert exit 0 + floors, print a one-line
* summary, clean up. Writes nothing to tracked files.
* record mode (pass the record flag) everything above PLUS append one
* EvalRunRecord-shaped JSONL line to
* <repo>/.gbrain-evals/eval-results.jsonl (the eval ledger).
*
* Honest scope: this gates the hybrid ranking pipeline (keyword/title/alias
* arms + RRF against gold qrels) with synthetic vectors. Semantic-embedding
* regressions remain the keyed eval suites' job.
*
* Budget: 60s under a saturated pool (two PGLite boots). If it breaches
* ~100s under contention, move the verify entry into the serial-tests CI
* job instead (same fallback as the chronicle eval).
*/
import { appendFileSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { dirname, join, resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
import { execSync, spawnSync } from 'node:child_process';
import type { BrainEngine } from '../src/core/engine.ts';
import type { ChunkInput } from '../src/core/types.ts';
import { basisEmbedding, parseLegacyQrels } from '../src/eval/deterministic-embed.ts';
import type { LegacyQrelsQuery } from '../src/eval/deterministic-embed.ts';
const ROOT = resolve(dirname(fileURLToPath(import.meta.url)), '..');
const QRELS_PATH = join(ROOT, 'test', 'fixtures', 'eval-baselines', 'qrels-search.json');
// The embedding space the throwaway brain is pinned to. The gateway must be
// configured with this BEFORE initSchema (the schema's vector(dims) columns
// derive from gateway config at initSchema time — config read only, no key
// needed), and the brain's GBRAIN_HOME config.json pins the same model+dims
// so the CLI subprocess resolves 1536 too.
const EMBEDDING_MODEL = 'openai:text-embedding-3-large';
const EMBEDDING_DIMENSIONS = 1536;
// Env the child must NOT inherit: engine reroutes (a stray DATABASE_URL
// flips the engine to postgres; a brain id reroutes to a mount), embedding
// overrides (would fight the pinned 1536 space), and provider keys (the
// canary must behave identically keyed and keyless — determinism by
// construction, not by the parent's shell profile).
const CHILD_ENV_STRIP = [
'DATABASE_URL',
'GBRAIN_DATABASE_URL',
'GBRAIN_BRAIN_ID',
'GBRAIN_SOURCE',
'GBRAIN_EMBEDDING_MODEL',
'GBRAIN_EMBEDDING_DIMENSIONS',
'OPENAI_API_KEY',
'ANTHROPIC_API_KEY',
'ZEROENTROPY_API_KEY',
'VOYAGE_API_KEY',
'OPENROUTER_API_KEY',
'DASHSCOPE_API_KEY',
'GOOGLE_GENERATIVE_AI_API_KEY',
'GEMINI_API_KEY',
];
// The legacy qrels parser lives with the embedder builder — one parser for
// the shape (re-exported here for the test that drives this runner).
export { parseLegacyQrels };
export type { LegacyQrelsQuery };
/**
* Seed the V2 canary corpus. For each query's relevant slugs:
* - putPage typed by prefix (person/company/note), title = slug tail;
* the expected-top1 page carries the primary text in BOTH
* compiled_truth and timeline (the V2 amendment timeline is what
* page-grain FTS indexes); siblings carry "Mentioned in context of
* <query>" in timeline.
* - upsertChunks with the same fixture text, basisEmbedding at the
* query's dim, token_count 10, chunk_source compiled_truth/timeline.
*/
export async function seedCanaryCorpus(engine: BrainEngine, queries: LegacyQrelsQuery[]): Promise<void> {
const seenSlugs = new Set<string>();
for (const q of queries) {
for (const slug of q.relevant_slugs) {
if (seenSlugs.has(slug)) continue;
seenSlugs.add(slug);
const isExpected = slug === q.first_relevant_slug;
const primaryText = `Primary content about ${q.query}`;
const mentionText = `Mentioned in context of ${q.query}`;
const type = slug.startsWith('people/')
? 'person'
: slug.startsWith('companies/')
? 'company'
: 'note';
await engine.putPage(slug, {
type,
title: slug.split('/').pop() ?? slug,
compiled_truth: isExpected ? primaryText : '',
// V2: the expected page's query text goes in timeline too — that is
// the page-grain-FTS-indexed column (title A + timeline C).
timeline: isExpected ? primaryText : mentionText,
});
const chunk: ChunkInput = {
chunk_index: 0,
chunk_text: isExpected ? primaryText : mentionText,
chunk_source: isExpected ? 'compiled_truth' : 'timeline',
embedding: basisEmbedding(q.embedding_dim, EMBEDDING_DIMENSIONS),
token_count: 10,
};
await engine.upsertChunks(slug, [chunk]);
}
}
}
interface GateJson {
verdict: 'pass' | 'fail';
correctness_gate: {
ran: boolean;
summary?: {
k: number;
queries_total: number;
queries_run: number;
queries_errored: number;
mean_recall_at_k: number;
first_relevant_hit_rate: number;
expected_top1_hit_rate: number;
expected_top1_denominator: number;
};
thresholds?: {
recall_at_k: number;
first_relevant_hit: number;
expected_top1: number;
};
breaches?: Array<Record<string, unknown>>;
};
}
function shortSha(): string {
try {
return execSync('git rev-parse --short HEAD', { cwd: ROOT, encoding: 'utf-8' }).trim();
} catch {
return 'unknown';
}
}
async function main(): Promise<number> {
const recordMode = process.argv.includes('--record');
const startedAt = Date.now();
const tmpHome = mkdtempSync(join(tmpdir(), 'gbrain-eval-canary-'));
// Keep the RUNNER's own gbrain home inside the sandbox too, so nothing in
// the seeding path can read or write the operator's real ~/.gbrain.
process.env.GBRAIN_HOME = tmpHome;
try {
const gbrainDir = join(tmpHome, '.gbrain');
mkdirSync(gbrainDir, { recursive: true });
const dbPath = join(gbrainDir, 'brain.pglite');
writeFileSync(
join(gbrainDir, 'config.json'),
JSON.stringify(
{
engine: 'pglite',
database_path: dbPath,
embedding_model: EMBEDDING_MODEL,
embedding_dimensions: EMBEDDING_DIMENSIONS,
},
null,
2,
) + '\n',
);
// Gateway config BEFORE initSchema (dims gotcha above). Empty env
// snapshot: no key is consulted, and none is needed for schema sizing.
const { configureGateway } = await import('../src/core/ai/gateway.ts');
configureGateway({
embedding_model: EMBEDDING_MODEL,
embedding_dimensions: EMBEDDING_DIMENSIONS,
env: {},
});
const { PGLiteEngine } = await import('../src/core/pglite-engine.ts');
const engine = new PGLiteEngine();
await engine.connect({ engine: 'pglite', database_path: dbPath });
await engine.initSchema();
const qrelsRaw = readFileSync(QRELS_PATH, 'utf-8');
await seedCanaryCorpus(engine, parseLegacyQrels(qrelsRaw));
// PGLite is single-writer: release the brain before the CLI child opens it.
await engine.disconnect();
const childEnv: Record<string, string | undefined> = { ...process.env };
for (const k of CHILD_ENV_STRIP) delete childEnv[k];
childEnv.GBRAIN_HOME = tmpHome;
// Spawn the REAL CLI. cwd is the temp home (not the repo) so repo-local
// dotfiles and Bun-auto-loaded .env files can't reroute the brain.
const child = spawnSync(
process.execPath,
[join(ROOT, 'src', 'cli.ts'), 'eval', 'gate', '--qrels', QRELS_PATH, '--embedder', 'deterministic', '--json'],
{
cwd: tmpHome,
env: childEnv as NodeJS.ProcessEnv,
encoding: 'utf-8',
timeout: 110_000,
maxBuffer: 32 * 1024 * 1024,
},
);
if (child.error) {
process.stderr.write(`[eval-canary] FAIL: could not spawn the CLI: ${child.error.message}\n`);
return 1;
}
if (child.status !== 0) {
process.stderr.write(`[eval-canary] FAIL: gate exit=${child.status ?? 'null(timeout/signal)'}\n`);
process.stderr.write(`[eval-canary] gate stdout tail:\n${(child.stdout ?? '').slice(-2000)}\n`);
process.stderr.write(`[eval-canary] gate stderr tail:\n${(child.stderr ?? '').slice(-2000)}\n`);
return 1;
}
// Parse the gate's JSON envelope (stdout carries only the envelope; any
// engine warnings go to stderr).
const stdout = child.stdout ?? '';
const jsonStart = stdout.indexOf('{');
if (jsonStart < 0) {
process.stderr.write(`[eval-canary] FAIL: no JSON found on gate stdout:\n${stdout.slice(-2000)}\n`);
return 1;
}
const gate = JSON.parse(stdout.slice(jsonStart)) as GateJson;
const summary = gate.correctness_gate.summary;
const floors = gate.correctness_gate.thresholds;
if (gate.verdict !== 'pass' || !summary || !floors) {
process.stderr.write(`[eval-canary] FAIL: verdict=${gate.verdict} summary=${JSON.stringify(summary)}\n`);
return 1;
}
// Exit 0 already implies floors held; assert explicitly anyway so a
// future exit-code regression in the gate can't silently pass the canary.
const breaches: string[] = [];
if (summary.queries_errored > 0) breaches.push(`queries_errored=${summary.queries_errored}`);
if (summary.mean_recall_at_k < floors.recall_at_k) breaches.push(`recall ${summary.mean_recall_at_k} < ${floors.recall_at_k}`);
if (summary.first_relevant_hit_rate < floors.first_relevant_hit) breaches.push(`first_relevant ${summary.first_relevant_hit_rate} < ${floors.first_relevant_hit}`);
if (summary.expected_top1_denominator > 0 && summary.expected_top1_hit_rate < floors.expected_top1) {
breaches.push(`expected_top1 ${summary.expected_top1_hit_rate} < ${floors.expected_top1}`);
}
if (breaches.length > 0) {
process.stderr.write(`[eval-canary] FAIL: ${breaches.join('; ')}\n`);
return 1;
}
const commit = shortSha();
const durationMs = Date.now() - startedAt;
process.stdout.write(
`[eval-canary] PASS commit=${commit}` +
` mean_recall_at_k=${summary.mean_recall_at_k.toFixed(4)}` +
` first_relevant_hit_rate=${summary.first_relevant_hit_rate.toFixed(4)}` +
` expected_top1_hit_rate=${summary.expected_top1_hit_rate.toFixed(4)}` +
` floors=${floors.recall_at_k}/${floors.first_relevant_hit}/${floors.expected_top1}` +
` k=${summary.k} queries=${summary.queries_run}/${summary.queries_total}` +
` duration_ms=${durationMs}\n`,
);
if (recordMode) {
// EvalRunRecord-shaped ledger line (matches src/commands/eval-run-all.ts;
// suite widened to the canary's own name, mode 'n/a' per the
// search-mode-independent convention).
const record = {
schema_version: 3,
run_id: `${commit}-retrieval-canary-na-0`,
ran_at: new Date().toISOString(),
suite: 'retrieval-canary',
mode: 'n/a',
commit,
seed: 0,
params: {
qrels: 'test/fixtures/eval-baselines/qrels-search.json',
embedder: 'deterministic',
k: summary.k,
metrics: {
mean_recall_at_k: summary.mean_recall_at_k,
first_relevant_hit_rate: summary.first_relevant_hit_rate,
expected_top1_hit_rate: summary.expected_top1_hit_rate,
expected_top1_denominator: summary.expected_top1_denominator,
queries_run: summary.queries_run,
queries_total: summary.queries_total,
},
floors: {
recall_at_k: floors.recall_at_k,
first_relevant_hit: floors.first_relevant_hit,
expected_top1: floors.expected_top1,
},
},
status: 'completed',
duration_ms: durationMs,
};
const ledgerDir = join(ROOT, '.gbrain-evals');
mkdirSync(ledgerDir, { recursive: true });
const ledgerPath = join(ledgerDir, 'eval-results.jsonl');
appendFileSync(ledgerPath, JSON.stringify(record) + '\n', 'utf-8');
process.stdout.write(`[eval-canary] recorded → ${ledgerPath}\n`);
}
return 0;
} catch (err) {
process.stderr.write(`[eval-canary] FAIL: ${(err as Error).stack ?? (err as Error).message}\n`);
return 1;
} finally {
rmSync(tmpHome, { recursive: true, force: true });
}
}
if (import.meta.main) {
process.exit(await main());
}
+273 -16
View File
@@ -1,18 +1,81 @@
#!/usr/bin/env bash
# scripts/run-serial-tests.sh — run *.serial.test.ts files with --max-concurrency=1.
# scripts/run-serial-tests.sh — run *.serial.test.ts files, one bun process per
# file, POOLED across files.
#
# Serial files are tests that share file-wide state (top-level mock.module,
# module-level singletons that intentionally cross test cases) and would race
# under intra-file concurrency. Discovered via filename suffix; no annotation
# inside the file is needed.
#
# Each file gets its OWN bun process. `--max-concurrency=1` alone was not
# enough: files in the same process share the module registry, so a top-level
# `mock.module(...)` in one file leaks into the next file's imports. Per-file
# processes give true isolation — and that isolation is per-PROCESS, not
# per-machine, so separate processes run CONCURRENTLY through the pool below.
# (The previous runner executed the ~140 processes strictly one-at-a-time:
# an 8.5-minute CI job whose serialization was never required by the
# quarantine contract.)
#
# Excluded by run-unit-shard.sh and run-unit-parallel.sh's parallel pass.
# Invoked separately by run-unit-parallel.sh after the parallel pass succeeds.
#
# Knobs:
# GBRAIN_SERIAL_POOL=N pool width (default min(detect_cpus, 4),
# then memory-adapted; 1 restores the old
# fully-sequential behavior)
# GBRAIN_SERIAL_FILE_TIMEOUT=S wall-clock kill per file (default 300;
# needs timeout/gtimeout on PATH, else no wrap)
# GBRAIN_TEST_MEM_PER_FILE_MB per-process memory budget (default 1536)
# GBRAIN_TEST_NO_MEM_ADAPT=1 skip the memory clamp
set -euo pipefail
# #3485: serial tests need no database — strip ambient DB URLs at this
# wrapper boundary (same four-layer guard as run-slow-tests.sh / the
# parallel runner) so the bunfig preload guard passes and nothing can
# reach a real brain.
unset DATABASE_URL GBRAIN_DATABASE_URL
cd "$(dirname "$0")/.."
. scripts/lib/test-env.sh
# ──────────────────────────────────────────────────────────────────────────
# EXCLUSIVE_FILES: files that must never run concurrently with anything else
# (machine-global state or contention-critical timing). They run sequentially
# AFTER the pool drains, without the wall-clock kill (a SIGKILL
# mid-registration could strand a real scheduled job). Growth guard:
# test/scripts/serial-files.test.ts fails when this list grows past 3
# entries — every addition needs a justification comment like the entries
# below.
# ──────────────────────────────────────────────────────────────────────────
EXCLUSIVE_FILES=(
# launchd/cron lifecycle arc: install → self-disable → reinstall →
# uninstall ordering against (PATH-shimmed) launchctl; the arc asserts
# machine-level sequencing and is the flake-class canary.
"test/autopilot-launchd-lifecycle.serial.test.ts"
# hardenBrainRepo({installCron:true}) executes REAL launchctl/crontab
# (src/core/brain-repo-durability.ts) — a concurrent or killed run could
# strand a real scheduled job on the machine.
"test/brain-durability-hook.serial.test.ts"
# hardenBrainRepo's own scaffolding commit fires the just-installed
# post-commit hook (background push) which races the synchronous
# push-probe on the same bare remote ("cannot lock ref" →
# needs_attention non-empty). The race is intra-call; pooled CPU
# contention widens the window past what the assertions tolerate
# (observed on a 4-vCPU CI runner, never locally). Sequential lane
# restores master-era timing until the probe learns to retry ref-lock
# contention.
"test/brain-repo-durability.serial.test.ts"
)
is_exclusive() {
local f="$1" e
for e in "${EXCLUSIVE_FILES[@]}"; do
[ "$f" = "$e" ] && return 0
done
return 1
}
# Use while-read for portability to macOS bash 3.2 (no mapfile).
files=()
while IFS= read -r f; do
@@ -24,35 +87,229 @@ if [ "${#files[@]}" -eq 0 ]; then
exit 0
fi
# --dry-run-list mirrors run-unit-shard.sh for inline checks/tests.
# --dry-run-list mirrors run-unit-shard.sh for inline checks/tests. Lists
# ALL discovered files, pooled and exclusive alike.
if [ "${1:-}" = "--dry-run-list" ]; then
printf '%s\n' "${files[@]}"
exit 0
fi
echo "[serial-tests] running ${#files[@]} file(s), one bun process per file"
ensure_pglite_snapshot "serial-tests"
# Each serial file gets its OWN bun process. `--max-concurrency=1` was not
# enough: files in the same process share the module registry, so a top-level
# `mock.module(...)` in one file leaks into the next file's imports
# (eval-takes-quality-runner mocks gateway.ts and the next file fails on
# `import { resetGateway }` because the mock factory didn't export it).
# Per-file processes give true isolation; cost is ~100ms startup × N files.
fail_count=0
failed_files=()
# Partition into pooled vs exclusive (exclusive entries missing from the
# discovered set are simply ignored — the list names repo files, and a
# sandbox copy of this script won't have them).
pool_files=()
exclusive_present=()
for f in "${files[@]}"; do
if ! bun test --max-concurrency=1 --timeout=60000 "$f"; then
fail_count=$((fail_count + 1))
failed_files+=("$f")
if is_exclusive "$f"; then
exclusive_present+=("$f")
else
pool_files+=("$f")
fi
done
# ──────────────────────────────────────────────────────────────────────────
# Pool sizing: min(detect_cpus, 4) — each pooled bun process can hold a
# PGLite WASM instance (~1.5GB) — then clamped by available memory (same
# layer-1 doctrine as run-unit-parallel.sh, 4GB OS reserve).
# ──────────────────────────────────────────────────────────────────────────
POOL="${GBRAIN_SERIAL_POOL:-}"
if [ -z "$POOL" ]; then
POOL=$(detect_cpus)
[ "$POOL" -gt 4 ] && POOL=4
if [ "${GBRAIN_TEST_NO_MEM_ADAPT:-0}" != "1" ]; then
MEM_PER_FILE_MB="${GBRAIN_TEST_MEM_PER_FILE_MB:-1536}"
AVAIL_MB=$(detect_available_mem_mb)
if [ "${AVAIL_MB:-0}" -gt 0 ] 2>/dev/null; then
BUDGET_MB=$((AVAIL_MB - 4096))
[ "$BUDGET_MB" -lt "$MEM_PER_FILE_MB" ] && BUDGET_MB="$MEM_PER_FILE_MB"
MAX_POOL=$((BUDGET_MB / MEM_PER_FILE_MB))
[ "$MAX_POOL" -lt 1 ] && MAX_POOL=1
[ "$POOL" -gt "$MAX_POOL" ] && POOL="$MAX_POOL"
fi
fi
fi
if ! printf '%s' "$POOL" | grep -qE '^[0-9]+$' || [ "$POOL" -lt 1 ]; then
echo "[serial-tests] ERROR: invalid pool size: $POOL" >&2
exit 2
fi
# Wall-clock kill per pooled file: contains the exit-hang class (a bun
# process that finishes its tests but never exits). SIGTERM first, SIGKILL
# after a grace period (`timeout -k`). macOS without coreutils has neither
# binary — run unwrapped there (CI is Linux and always wraps).
PER_FILE_TIMEOUT="${GBRAIN_SERIAL_FILE_TIMEOUT:-300}"
TIMEOUT_BIN=""
command -v timeout >/dev/null 2>&1 && TIMEOUT_BIN="timeout"
[ -z "$TIMEOUT_BIN" ] && command -v gtimeout >/dev/null 2>&1 && TIMEOUT_BIN="gtimeout"
LOG_DIR=$(mktemp -d "${TMPDIR:-/tmp}/gbrain-serial.XXXXXX")
trap 'rm -rf "$LOG_DIR"' EXIT
if [ -n "$TIMEOUT_BIN" ]; then
TIMEOUT_DESC="${PER_FILE_TIMEOUT}s via $TIMEOUT_BIN"
else
TIMEOUT_DESC="none (no timeout/gtimeout on PATH)"
fi
echo "[serial-tests] ${#files[@]} file(s): pool=$POOL (${#exclusive_present[@]} exclusive), per-file timeout=$TIMEOUT_DESC"
# Per-test timeout is 120s (not the fast-loop 60s): pooled contention can
# push a 30-50s file past 60s — the same flake class the slow lane hardened
# against in v0.40.10. The literal `bun test --max-concurrency=1` below is
# contract-pinned by test/scripts/serial-files.test.ts.
run_one_file() {
# $1 file, $2 log path, $3 exit-sentinel path, $4 wrap ("wrap"|"nowrap")
local f="$1" log="$2" exitf="$3" wrap="$4" rc=0
if [ "$wrap" = "wrap" ] && [ -n "$TIMEOUT_BIN" ]; then
"$TIMEOUT_BIN" -k 15 "$PER_FILE_TIMEOUT" \
bun test --max-concurrency=1 --timeout=120000 "$f" > "$log" 2>&1 || rc=$?
else
bun test --max-concurrency=1 --timeout=120000 "$f" > "$log" 2>&1 || rc=$?
fi
echo "$rc" > "$exitf"
}
start_epoch=$(date +%s)
idx=0
if [ "${#pool_files[@]}" -gt 0 ]; then
for f in "${pool_files[@]}"; do
while [ "$(jobs -rp | wc -l | tr -d ' ')" -ge "$POOL" ]; do
sleep 0.2
done
(
s=$(date +%s)
run_one_file "$f" "$LOG_DIR/$idx.log" "$LOG_DIR/$idx.exit" "wrap"
e=$(date +%s)
echo "$((e - s))" > "$LOG_DIR/$idx.dur"
) &
idx=$((idx + 1))
done
wait
fi
# Exclusive lane: sequential, unwrapped (see EXCLUSIVE_FILES comment).
if [ "${#exclusive_present[@]}" -gt 0 ]; then
for f in "${exclusive_present[@]}"; do
s=$(date +%s)
run_one_file "$f" "$LOG_DIR/$idx.log" "$LOG_DIR/$idx.exit" "nowrap"
e=$(date +%s)
echo "$((e - s))" > "$LOG_DIR/$idx.dur"
idx=$((idx + 1))
done
fi
# ──────────────────────────────────────────────────────────────────────────
# Aggregate from the exit sentinels.
# exit 0 → pass
# exit 124 → killed by the per-file wall-clock timeout (real
# failure: the exit-hang class this cap exists for)
# exit 143 / 137, or a → EXTERNAL-KILL class: a stray SIGTERM/SIGKILL from
# missing sentinel outside this runner (sibling workspaces' process
# cleanup, memory jetsam — the same class
# run-unit-parallel.sh rescues). Queued for ONE
# sequential rescue re-run below; a rescue that
# fails again is a real failure. Never a silent pass.
# anything else → real failure
# ──────────────────────────────────────────────────────────────────────────
ordered_files=()
if [ "${#pool_files[@]}" -gt 0 ]; then ordered_files+=("${pool_files[@]}"); fi
if [ "${#exclusive_present[@]}" -gt 0 ]; then ordered_files+=("${exclusive_present[@]}"); fi
fail_count=0
failed_files=()
rescue_files=()
# Aggregate pass count across pooled files, re-emitted below in bun's own
# " N pass" summary format so run-unit-parallel.sh's headline counter
# (bun_summary_count) still sees the serial suite's tests. Failing files'
# logs are cat'ed raw (their " N pass/fail" lines land in the stream
# directly), so only PASSING files accumulate here — no double counting.
pass_total=0
i=0
for f in "${ordered_files[@]}"; do
dur="?"
[ -f "$LOG_DIR/$i.dur" ] && dur=$(cat "$LOG_DIR/$i.dur")
if [ ! -f "$LOG_DIR/$i.exit" ]; then
echo "[serial-tests] KILLED ${dur}s $f — missing exit sentinel (external kill/OOM) — queued for serial rescue" >&2
rescue_files+=("$f")
else
rc=$(cat "$LOG_DIR/$i.exit")
if [ "$rc" = "0" ]; then
summary=$(grep -E '^ *[0-9]+ pass' "$LOG_DIR/$i.log" | tail -1 | tr -d ' ' || true)
n=$(printf '%s' "$summary" | grep -oE '^[0-9]+' || echo 0)
pass_total=$((pass_total + n))
echo "[serial-tests] PASS ${dur}s $f ${summary:+($summary)}"
elif [ "$rc" = "137" ] && [ "$dur" != "?" ] && [ "$dur" -ge "$PER_FILE_TIMEOUT" ] 2>/dev/null; then
# 137 with full duration = OUR timeout's SIGKILL escalation (a hang
# that ignored SIGTERM), not an external kill — a real failure; a
# rescue re-run would just re-hang for another ~315s.
echo "[serial-tests] FAIL ${dur}s $f — exit 137 (hang survived SIGTERM; killed by ${PER_FILE_TIMEOUT}s per-file timeout)" >&2
cat "$LOG_DIR/$i.log" >&2
fail_count=$((fail_count + 1))
failed_files+=("$f")
elif [ "$rc" = "143" ] || [ "$rc" = "137" ]; then
echo "[serial-tests] KILLED ${dur}s $f — exit $rc (external SIGTERM/SIGKILL) — queued for serial rescue" >&2
rescue_files+=("$f")
else
note=""
[ "$rc" = "124" ] && note=" (killed by ${PER_FILE_TIMEOUT}s per-file timeout)"
echo "[serial-tests] FAIL ${dur}s $f — exit $rc$note" >&2
cat "$LOG_DIR/$i.log" >&2
fail_count=$((fail_count + 1))
failed_files+=("$f")
fi
fi
i=$((i + 1))
done
# Rescue pass: one sequential, unpooled re-run per externally-killed file.
# Phantoms pass here and the run stays green (with a rescue note); real
# failures fail again and go red. Mirrors run-unit-parallel.sh's doctrine.
if [ "${#rescue_files[@]}" -gt 0 ]; then
echo "[serial-tests] rescue pass: ${#rescue_files[@]} externally-killed file(s), re-running serially" >&2
for f in "${rescue_files[@]}"; do
# Exclusive-lane files keep their no-kill contract on rescue too (the
# lane exists because a SIGKILL mid-registration strands real state).
wrap_mode="wrap"
is_exclusive "$f" && wrap_mode="nowrap"
s=$(date +%s)
run_one_file "$f" "$LOG_DIR/$i.log" "$LOG_DIR/$i.exit" "$wrap_mode"
e=$(date +%s)
rc=$(cat "$LOG_DIR/$i.exit" 2>/dev/null || echo 1)
if [ "$rc" = "0" ]; then
summary=$(grep -E '^ *[0-9]+ pass' "$LOG_DIR/$i.log" | tail -1 | tr -d ' ' || true)
n=$(printf '%s' "$summary" | grep -oE '^[0-9]+' || echo 0)
pass_total=$((pass_total + n))
echo "[serial-tests] PASS $((e - s))s $f ${summary:+($summary)} (rescued: external-kill phantom)"
else
echo "[serial-tests] FAIL $((e - s))s $f — exit $rc on rescue re-run" >&2
cat "$LOG_DIR/$i.log" >&2
fail_count=$((fail_count + 1))
failed_files+=("$f")
fi
i=$((i + 1))
done
fi
# Slowest-file table: feeds flake triage + future weight mining.
echo "[serial-tests] slowest files:"
i=0
for f in "${ordered_files[@]}"; do
[ -f "$LOG_DIR/$i.dur" ] && echo "$(cat "$LOG_DIR/$i.dur") $f"
i=$((i + 1))
done | sort -rn | head -10 | sed 's/^/ /'
total_epoch=$(( $(date +%s) - start_epoch ))
if [ "$fail_count" -gt 0 ]; then
echo "" >&2
echo "[serial-tests] $fail_count file(s) failed:" >&2
echo "[serial-tests] $fail_count file(s) failed (${total_epoch}s total):" >&2
for f in "${failed_files[@]}"; do
echo " - $f" >&2
done
exit 1
fi
echo "[serial-tests] all ${#files[@]} file(s) passed"
# bun-summary-format aggregate: run-unit-parallel.sh's headline counter
# (bun_summary_count awk: $1 numeric, $2 == "pass") reads this line — without
# it the serial suite's tests vanish from `bun run test`'s pass=N banner.
echo " $pass_total pass"
echo "[serial-tests] all ${#ordered_files[@]} file(s) passed in ${total_epoch}s (pool=$POOL)"
+3
View File
@@ -11,6 +11,9 @@ set -euo pipefail
unset DATABASE_URL GBRAIN_DATABASE_URL
cd "$(dirname "$0")/.."
. scripts/lib/test-env.sh
ensure_pglite_snapshot "run-slow-tests"
slow_files=()
while IFS= read -r f; do
slow_files+=("$f")
+5 -48
View File
@@ -62,54 +62,11 @@ cd "$(dirname "$0")/.."
# fan-out — shards inherit a finished fixture. Opt out: GBRAIN_NO_SNAPSHOT=1
# (the migration-replay canary tests clear the env themselves regardless).
# ──────────────────────────────────────────────────────────────────────────
if [ "${GBRAIN_NO_SNAPSHOT:-0}" != "1" ]; then
if bun run build:pglite-snapshot >/dev/null 2>&1; then
export GBRAIN_PGLITE_SNAPSHOT=test/fixtures/pglite-snapshot.tar
else
echo "[run-unit-parallel] snapshot build failed (non-fatal) — tests run with cold init" >&2
fi
fi
# ──────────────────────────────────────────────────────────────────────────
# CPU detection: Apple Silicon perf cores → Mac total physical → nproc → 4.
# Returns a single positive integer.
# ──────────────────────────────────────────────────────────────────────────
detect_cpus() {
local n=""
n=$(sysctl -n hw.perflevel0.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
n=$(sysctl -n hw.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
n=$(nproc 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
echo 4
}
# ──────────────────────────────────────────────────────────────────────────
# Available-memory detection (MB). macOS: vm_stat free + inactive +
# speculative + purgeable pages (inactive/purgeable are reclaimable on
# pressure, which is exactly the scenario we size for). Linux: MemAvailable.
# Unknown platform → 0, and the caller skips adaptation entirely.
# ──────────────────────────────────────────────────────────────────────────
detect_available_mem_mb() {
if command -v vm_stat >/dev/null 2>&1; then
vm_stat 2>/dev/null | awk '
/page size of/ { psize = $8 }
/Pages free/ { free = $NF }
/Pages inactive/ { inactive = $NF }
/Pages speculative/ { spec = $NF }
/Pages purgeable/ { purge = $NF }
END {
gsub(/\./, "", free); gsub(/\./, "", inactive)
gsub(/\./, "", spec); gsub(/\./, "", purge)
if (psize == 0) psize = 16384
printf "%d\n", (free + inactive + spec + purge) * psize / 1048576
}'
return
fi
if [ -r /proc/meminfo ]; then
awk '/MemAvailable/ { printf "%d\n", $2 / 1024; found = 1 } END { if (!found) print 0 }' /proc/meminfo
return
fi
echo 0
}
# detect_cpus / detect_available_mem_mb / ensure_pglite_snapshot live in the
# shared lib (also sourced by test-shard.sh, run-serial-tests.sh,
# run-slow-tests.sh) — one implementation, no copy drift.
. scripts/lib/test-env.sh
ensure_pglite_snapshot "run-unit-parallel"
# ──────────────────────────────────────────────────────────────────────────
# Argument parsing. --shards N override wins over $SHARDS; both are clamped.
+67 -21
View File
@@ -25,28 +25,61 @@ set -uo pipefail
cd "$(dirname "$0")/.."
# detect_cpus + ensure_pglite_snapshot (the PGLite-booting eval checks use
# the snapshot fast-path when the shape matches).
. scripts/lib/test-env.sh
# ──────────────────────────────────────────────────────────────────────────
# Checks to run. Order is irrelevant (parallel), but keep stable for log
# determinism + grep-ability. Each entry is a bun-script name (the
# `package.json` "scripts" key), invoked as `bun run <name>`.
# Checks to run. Each entry is a bun-script name (the `package.json`
# "scripts" key), invoked as `bun run <name>`.
#
# To add a check: append to this array. To skip in CI temporarily, comment
# the line — the parallel runner doesn't care about count.
# ORDER MATTERS for wallclock: the spawn loop below is capped at
# GBRAIN_VERIFY_MAX_PARALLEL workers, so the heaviest checks go FIRST
# (LPT-style — makespan ≈ max(longest check, total/POOL)). The heavy block:
# typecheck (tsc), two `cp -R src` + `bun build --compile` binary builds,
# the admin vite+tsc build, the fuzz bundles, guard self-tests, the
# PGLite-booting eval checks, and the whole-tree greps. Everything after is
# sub-second; that tail keeps its historical order for grep-ability.
#
# To add a check: append to the right block. To skip in CI temporarily,
# comment the line — the runner doesn't care about count.
# ──────────────────────────────────────────────────────────────────────────
CHECKS=(
# ── heavy (longest-first) ──
"typecheck"
"check:admin-build"
"check:wasm"
"check:pglite-embedded"
"check:fuzz-purity"
# W0 fix-wave (Tier-1 #11): guard self-tests — every scanner guard proves it
# can fail (bad fixture → exit 1) before it counts as coverage. Registry:
# scripts/guards-manifest.tsv (package.json's stale `check:all` copy deleted).
"check:guard-self-test"
# Chronicle eval: $0, deterministic, exit-0-only-on-perfect (6 gold tasks).
# Boots its own PGLite — budget ≤60s under a saturated pool; if it breaches
# ~100s under contention, move it into the serial-tests CI job instead.
"check:eval-chronicle"
# Retrieval canary: $0, hermetic, deterministic-embedder CLI run of the
# qrels correctness gate. Boots two PGLite processes (seed + real CLI) —
# budget ≤60s under a saturated pool; if it breaches ~100s under
# contention, move it into the serial-tests CI job instead (same fallback
# as eval-chronicle above).
"check:eval-canary"
"check:bootstrap-templates"
"check:skill-brain-first"
"check:conversation-parser"
"check:resolver"
"check:privacy"
"check:proposal-pii"
"check:test-names"
"check:test-isolation"
# ── light tail (sub-second greps; historical order) ──
"check:proposal-pii"
"check:jsonb"
"check:search-path"
"check:source-id-projection"
"check:source-config-leak"
"check:progress"
"check:no-tracked-symlinks"
"check:test-isolation"
"check:wasm"
"check:pglite-embedded"
"check:admin-build"
"check:admin-scope-drift"
"check:cli-exec"
"check:system-of-record"
@@ -55,33 +88,28 @@ CHECKS=(
"check:skills-manifest"
"check:no-pii-agent-voice"
"check:synthetic-corpus-privacy"
"check:skill-brain-first"
"check:fuzz-purity"
"check:operations-filter-bypass"
"check:gateway-routed"
"check:worker-pool-atomicity"
"check:doc-history"
"check:fixture-privacy"
"check:conversation-parser"
"check:resolver"
"check:source-scope-onboard"
"check:no-double-retry"
"check:batch-audit-site"
"check:engine-dynamic-import"
"check:grok-pin"
"check:opencode-pin"
"check:pin-doc-privacy"
"check:worker-lock-renewal-shape"
"check:bootstrap-tag"
"check:bootstrap-templates"
"check:skill-refs"
# W0 fix-wave (Tier-1 #11): guard self-tests — every scanner guard proves it
# can fail (bad fixture → exit 1) before it counts as coverage. Registry:
# scripts/guards-manifest.tsv (package.json's stale `check:all` copy deleted).
"check:guard-self-test"
# Previously reachable ONLY from the deleted check:all (i.e. never run):
"check:newlines"
"check:exports-count"
"check:no-legacy-getconnection"
"typecheck"
# Revived registered-but-never-executed guards (this pass):
"check:pagetype-exhaustive"
"check:pg-url-redaction"
)
if [ "${#CHECKS[@]}" -eq 0 ]; then
@@ -122,8 +150,22 @@ if command -v gtimeout >/dev/null 2>&1; then TIMEOUT_BIN="gtimeout"
elif command -v timeout >/dev/null 2>&1; then TIMEOUT_BIN="timeout"
fi
# Bounded worker pool. Unbounded fan-out ran two `cp -R src` +
# `bun build --compile` builds, the admin vite build, tsc, and ~40 greps
# simultaneously on a 4-vCPU CI runner — pushing slow checks into the
# 120s per-check timeout (the documented flake class on slower hosts).
# Default = detect_cpus so a many-core dev machine keeps its wide fan-out;
# escape hatch: GBRAIN_VERIFY_MAX_PARALLEL=999.
MAX_PAR="${GBRAIN_VERIFY_MAX_PARALLEL:-$(detect_cpus)}"
if ! printf '%s' "$MAX_PAR" | grep -qE '^[0-9]+$' || [ "$MAX_PAR" -lt 1 ]; then
echo "ERROR: invalid GBRAIN_VERIFY_MAX_PARALLEL: $MAX_PAR" >&2
exit 2
fi
ensure_pglite_snapshot "verify-parallel"
START_TS=$(date +%s)
echo "[verify-parallel] running ${#CHECKS[@]} checks in parallel (timeout=${TIMEOUT}s, logs=$LOG_DIR)" >&2
echo "[verify-parallel] running ${#CHECKS[@]} checks (pool=$MAX_PAR, timeout=${TIMEOUT}s, logs=$LOG_DIR)" >&2
# ──────────────────────────────────────────────────────────────────────────
# Spawn one background process per check. Each child captures its own exit
@@ -136,6 +178,10 @@ echo "[verify-parallel] running ${#CHECKS[@]} checks in parallel (timeout=${TIME
PIDS=()
SAFE_NAMES=()
for c in "${CHECKS[@]}"; do
# Throttle to the worker pool (bash 3.2 — no wait -n; jobs -rp reaps).
while [ "$(jobs -rp | wc -l | tr -d ' ')" -ge "$MAX_PAR" ]; do
sleep 0.1
done
safe="${c//:/_}"
SAFE_NAMES+=("$safe")
LOG_FILE="$LOG_DIR/$safe.log"
+21 -3
View File
@@ -96,6 +96,17 @@ export function computeMedian(values: number[]): number {
: sorted[mid]!;
}
/**
* Quantile (nearest-rank) of a list of numbers. Empty input returns 0.
* q in [0, 1]; q=0.75 is the missing-file fallback weight (see partition).
*/
export function computeQuantile(values: number[], q: number): number {
if (values.length === 0) return 0;
const sorted = [...values].sort((a, b) => a - b);
const idx = Math.min(sorted.length - 1, Math.max(0, Math.ceil(q * sorted.length) - 1));
return sorted[idx]!;
}
export interface PartitionOpts {
/**
* Weight to assign files that are absent from the weights map. Defaults
@@ -133,8 +144,15 @@ export function partition(
const shards: string[][] = Array.from({ length: n }, () => []);
if (files.length === 0) return shards;
// Compute fallback weight from the median of present weights, unless
// the caller supplied an explicit override.
// Compute fallback weight from the p75 of present weights, unless the
// caller supplied an explicit override. p75, not median: the weight
// distribution is extremely right-skewed (median ~27ms, mean ~800ms —
// most files are trivial greps, the tail boots PGLite), and files
// MISSING from the map skew heavy (new integration tests land unweighted
// more often than new pure-unit tests). A median fallback modeled 45% of
// the corpus at ~30ms and let one shard silently carry the unweighted
// heavies; p75 over-weights small new files slightly (harmless — LPT
// self-corrects on the next mine) instead of under-weighting big ones.
let fallback: number;
if (opts.fallbackWeight !== undefined) {
if (!Number.isFinite(opts.fallbackWeight) || opts.fallbackWeight < 0) {
@@ -144,7 +162,7 @@ export function partition(
}
fallback = opts.fallbackWeight;
} else {
fallback = computeMedian(Array.from(weights.values()));
fallback = computeQuantile(Array.from(weights.values()), 0.75);
}
// Cold-start guard: if the weights map is empty AND no explicit
// fallback was supplied, every effective weight would be 0 and LPT
+19 -2
View File
@@ -56,6 +56,8 @@ fi
cd "$(dirname "$0")/.."
. scripts/lib/test-env.sh
# Collect non-E2E, non-serial unit test files. Slow files INCLUDED — see
# header comment. Local run-unit-shard.sh excludes slow files (different
# policy by design).
@@ -72,7 +74,13 @@ cd "$(dirname "$0")/.."
# total bounded. With 10 matrix shards the per-shard total drops to ~272s.
# Dedicated jobs run in parallel so total CI wallclock = max(matrix ~4.5min,
# slow-eval ~3.3min, slow-entity-resolve-perf ~2.6min) ≈ 4.5min.
ALL_FILES=$(find test -name '*.test.ts' \
# evals/ is included: its *.test.ts files (eval-harness unit tests) were
# previously collected by NO runner — 45+ real tests never executed anywhere.
# Every collected evals file must be KEYLESS (no API keys, no network) —
# enforced by the allowlist guard in test/scripts/evals-collection.test.ts.
# The local fast loop (run-unit-shard.sh) stays test-only by design (see
# docs/TESTING.md "CI vs local: intentionally divergent file sets").
ALL_FILES=$(find test evals -name '*.test.ts' \
-not -name '*.serial.test.ts' \
-not -name 'eval-longmemeval-e2e.slow.test.ts' \
-not -name 'entity-resolve-perf.slow.test.ts' \
@@ -94,6 +102,12 @@ if [ "$DRY_RUN_LIST" = "1" ]; then
exit 0
fi
# Snapshot fast-path (after the dry-run exit so list-only calls stay
# instant): ~370 PGLite-booting matrix files pay ~3.1s cold init each
# without it. The echo inside makes silent cold-init regressions visible
# in CI logs.
ensure_pglite_snapshot "test-shard"
ALL_COUNT=$(printf '%s\n' "$ALL_FILES" | grep -c '^' || true)
SHARD_COUNT=$(printf '%s\n' "$SHARD_FILES" | grep -c '^' || true)
# grep -c on empty input returns 0 even with trailing newline edge cases
@@ -108,4 +122,7 @@ fi
# Convert newline-separated file list to argv. xargs handles the
# whitespace correctly without word-splitting on spaces in paths.
printf '%s\n' "$SHARD_FILES" | xargs bun test --timeout=60000
# --max-concurrency mirrors the local runner: unbounded intra-process
# concurrency under parallel PGLite boots produced real shard deaths (the
# 22-minute matrix timeout in test.yml records 13 of them).
printf '%s\n' "$SHARD_FILES" | xargs bun test --timeout=60000 --max-concurrency="${GBRAIN_TEST_MAX_CONCURRENCY:-4}"
+1201 -717
View File
File diff suppressed because it is too large Load Diff
+467 -40
View File
@@ -27,8 +27,9 @@
* B5 relay instruction), never a stack trace.
*/
import { mkdirSync, readdirSync } from 'node:fs';
import { basename, isAbsolute, join, resolve } from 'node:path';
import { existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { basename, dirname, isAbsolute, join, resolve } from 'node:path';
import { VERSION } from '../version.ts';
import { loadConfig, loadConfigFileOnly, toEngineConfig } from '../core/config.ts';
@@ -77,7 +78,16 @@ import {
statusHarness,
type HarnessDeps,
} from '../core/bootstrap/harness.ts';
import { codexConfigPath } from '../core/bootstrap/host-specs.ts';
import { codexConfigPath, opencodeConfigDir, opencodeGlobalConfigPath, opencodeProjectConfigPath } from '../core/bootstrap/host-specs.ts';
import {
opencodeEntryKind,
opencodeEntrySnippet,
opencodeRemoteEntryExists,
parseOpencodeConfig,
reconcileOpencodeSiblingGlobal,
removeOpencodeMcpEntry,
writeOpencodeMcpEntry,
} from '../core/bootstrap/opencode-json.ts';
import { promptLine } from '../core/cli-util.ts';
import {
appendInstallLog,
@@ -88,7 +98,7 @@ import {
} from '../core/bootstrap/status.ts';
import { verifyWorkspace, deriveWorkspaceSourceId } from '../core/bootstrap/verify.ts';
export const BOOTSTRAP_HELP = `gbrain bootstrap — paste-in agent install (Claude Code / Codex)
export const BOOTSTRAP_HELP = `gbrain bootstrap — paste-in agent install (Claude Code / Codex / opencode)
Usage: gbrain bootstrap <subcommand> [flags]
@@ -104,25 +114,29 @@ Subcommands (run \`gbrain bootstrap status\` first — it is the resume entrypoi
render [--force] [--only F] [--minimal]
Render identity files from the confirmed answers.
Never clobbers; --force backs up first.
hooks [--harness claude-code|codex] [--repair] [--no-hooks] [--gbrain-bin <path>]
hooks [--harness claude-code|codex|opencode] [--repair] [--no-hooks] [--gbrain-bin <path>]
Register MCP (+ per-turn hooks on Claude Code,
ON by default; --no-hooks opts out, GBRAIN_HOOKS=0
disables at runtime).
disables at runtime). opencode registrations are
written directly into its JSONC config (user-global
by default; MCP_SCOPE=project is an explicit opt-in
with a sharing warning).
repo Create the dedicated PRIVATE GitHub repo (or adopt
an EMPTY private repo you created under your own
account), verify the privacy bit via the API, push.
verify [--json] The whole install contract (round-trip, graph floor,
magic moment, scans, hooks smoke). Exit 0 or not done.
attach [--harness H] Machine two: adopt a cloned agent workspace.
harness [--harness claude-code|codex|all] [--url U | --port N] [--source ID]
harness [--harness claude-code|codex|opencode|all] [--url U | --port N] [--source ID]
[--token-name NAME | --token TOK] [--name MCPNAME] [--project DIR]...
[--no-hooks] [--no-capture] [--force] [--status] [--remove] [--yes] [--json]
Wire framework-spawned Claude Code / Codex sessions to a
RUNNING \`gbrain serve --http\` on this box (#4043): scoped
bearer token, user-scope MCP + headless pre-approval,
lifecycle hooks (user scope, or per --project dir), codex
config block. No agent.json needed. Idempotent; --remove
tears it down. (--local is an accepted no-op alias.)
Wire framework-spawned Claude Code / Codex / opencode
sessions to a RUNNING \`gbrain serve --http\` on this box
(#4043): scoped bearer token, user-scope MCP + headless
pre-approval, lifecycle hooks (user scope, or per --project
dir), codex config block, opencode config entry. No
agent.json needed. Idempotent; --remove tears it down.
(--local is an accepted no-op alias.)
cloud-setup-script Print the paste-ready cloud environment setup
script (installs the gbrain binary into the
environment snapshot; npm-based bun fetching
@@ -158,7 +172,7 @@ const SUBCOMMAND_HELP: Record<string, string> = {
' Create the dedicated PRIVATE GitHub repo (or adopt an EMPTY private repo you created\n' +
' under your own account), verify the privacy bit via the API, push.',
hooks:
'gbrain bootstrap hooks [--harness claude-code|codex] [--repair] [--no-hooks] [--gbrain-bin <path>]\n' +
'gbrain bootstrap hooks [--harness claude-code|codex|opencode] [--repair] [--no-hooks] [--gbrain-bin <path>]\n' +
' Register MCP (+ per-turn hooks on Claude Code, ON by default; --no-hooks opts out).',
verify:
'gbrain bootstrap verify [--json]\n' +
@@ -244,12 +258,26 @@ function shellQuoteForDisplay(arg: string): string {
// ── Shared plumbing ─────────────────────────────────────────────────────────
type Harness = 'claude-code' | 'codex';
type Harness = 'claude-code' | 'codex' | 'opencode';
/** Best-effort harness auto-detect; the --harness flag always wins. */
/** Every workspace-lane harness exhaustive-switch anchors key off this so
* a future member is a COMPILE error at each dispatch site, not a silent
* fall-through into another harness's branch (the union-widening trap: a
* `harness === 'claude-code' ? A : B` ternary routes every new member down
* B). */
const HARNESSES = ['claude-code', 'codex', 'opencode'] as const satisfies readonly Harness[];
function isHarness(v: string | undefined): v is Harness {
return (HARNESSES as readonly string[]).includes(v ?? '');
}
/** Best-effort harness auto-detect; the --harness flag always wins.
* opencode sets OPENCODE=1 (+OPENCODE_PID) in its bash-tool children
* verified against opencode 1.18.18 (OPENCODE-CLI-PIN.md §Environment). */
export function detectHarness(env: Record<string, string | undefined> = process.env): Harness | null {
if (env.CLAUDECODE || env.CLAUDE_CODE_ENTRYPOINT) return 'claude-code';
if (env.CODEX_HOME || env.CODEX_SANDBOX || env.CODEX_CI) return 'codex';
if (env.OPENCODE || env.OPENCODE_PID) return 'opencode';
return null;
}
@@ -285,7 +313,16 @@ async function verifyMcpTargetsWorkspace(
gbrainBin: string,
sourceId: string,
): Promise<'match' | 'mismatch' | 'unknown'> {
const bin = harness === 'claude-code' ? 'claude' : 'codex';
// Exec-lane harnesses only. opencode registrations go through the direct
// JSONC writer whose 4-state fingerprint IS the [FIX7] check (structural,
// no exec) — it never routes here; 'unknown' keeps a stray call honest.
const EXEC_HARNESS_BIN = {
'claude-code': 'claude',
codex: 'codex',
opencode: null,
} as const satisfies Record<Harness, string | null>;
const bin = EXEC_HARNESS_BIN[harness];
if (bin === null) return 'unknown';
let res;
try {
res = await runner([bin, 'mcp', 'get', name]);
@@ -300,6 +337,121 @@ async function verifyMcpTargetsWorkspace(
return hasBin && hasSource ? 'match' : 'mismatch';
}
/** Wall-clock cap on the best-effort `opencode mcp list` probe: `mcp list`
* SPAWNS every configured server, and a hung spawn must not hang the install
* on timeout the probe child is actually TERMINATED (SIGTERM, then SIGKILL
* ~2s later) and the result degrades to the could-not-confirm branch (code
* 124, repo-visibility's raced-runner convention). */
const OPENCODE_PROBE_TIMEOUT_MS = 20_000;
/** Injectable probe-spawn seam (the door serial tests capture argv + cwd +
* env and fake the child). The default holds the REAL process handle via
* Bun.spawn a Promise.race that merely abandons a hung `opencode mcp list`
* leaves its spawned MCP servers running (including the just-registered
* `gbrain serve`, which then squats the PGLite single-writer lock) and keeps
* the CLI's event loop alive past flushThenExit. */
export interface OpencodeProbeHandle {
exited: Promise<number>;
kill(force?: boolean): void;
stdout: Promise<string>;
stderr: Promise<string>;
/** Detach the child + its pipes from the event loop (called when the probe
* gives up on a hung child/grandchild so the CLI can still exit). */
unref?: () => void;
}
export type OpencodeProbeSpawn = (
argv: string[],
opts: { cwd: string; env: Record<string, string | undefined> },
) => OpencodeProbeHandle;
function defaultOpencodeProbeSpawn(
argv: string[],
opts: { cwd: string; env: Record<string, string | undefined> },
): OpencodeProbeHandle {
const proc = Bun.spawn(argv, {
cwd: opts.cwd,
env: opts.env as Record<string, string>,
stdin: 'ignore',
stdout: 'pipe',
stderr: 'pipe',
});
return {
exited: proc.exited,
kill: (force?: boolean) => {
try {
proc.kill(force ? 9 : undefined);
} catch {
/* already dead */
}
},
stdout: new Response(proc.stdout).text().catch(() => ''),
stderr: new Response(proc.stderr).text().catch(() => ''),
unref: () => {
try {
proc.unref();
} catch {
/* best-effort */
}
},
};
}
/** Run the opencode registration probe with OPENCODE_DISABLE_AUTOUPDATE=1 on
* the spawn env (OPENCODE-CLI-PIN.md §Probes: the auto-updater must never
* fire mid-probe) from an explicit `cwd` callers pass a fresh EMPTY temp
* dir, never the invoking cwd, because opencode merges a project
* opencode.json from cwd and spawns its local servers with NO trust prompt
* (a cloned malicious repo must not get code execution out of an install
* probe). On timeout the child is killed (SIGTERM SIGKILL) and the pipes
* are drained BOUNDED (a spawned MCP-server grandchild can inherit the pipe
* fds and hold them open past the direct child's death). Exported for the
* timeout-kill unit test. */
export async function runOpencodeProbe(
argv: string[],
opts: { cwd: string; spawn?: OpencodeProbeSpawn; timeoutMs?: number },
): Promise<{ code: number; stdout: string; stderr: string }> {
const spawnFn = opts.spawn ?? defaultOpencodeProbeSpawn;
const timeoutMs = opts.timeoutMs ?? OPENCODE_PROBE_TIMEOUT_MS;
const env: Record<string, string | undefined> = { ...process.env, OPENCODE_DISABLE_AUTOUPDATE: '1' };
let handle: OpencodeProbeHandle;
try {
handle = spawnFn(argv, { cwd: opts.cwd, env });
} catch (e) {
// Bun.spawn throws synchronously when the binary is absent — map to the
// shell's 127 convention so the caller's not-on-PATH branch fires.
return { code: 127, stdout: '', stderr: e instanceof Error ? e.message : String(e) };
}
// Bounded race helper that never leaves a live timer holding the loop.
const raceMs = async <T>(p: Promise<T>, ms: number, fallback: T): Promise<T> => {
let timer: ReturnType<typeof setTimeout> | undefined;
try {
return await Promise.race([p, new Promise<T>((res) => { timer = setTimeout(() => res(fallback), ms); })]);
} finally {
clearTimeout(timer);
}
};
let code = await raceMs<number | null>(handle.exited, timeoutMs, null);
const timedOut = code === null;
if (code === null) {
handle.kill(); // graceful first — opencode tears its servers down on TERM
code = await raceMs<number | null>(handle.exited, 2_000, null);
if (code === null) {
handle.kill(true); // SIGKILL is not refusable; the wait below is paranoia-bounded
code = await raceMs<number | null>(handle.exited, 2_000, null);
}
}
const drainCap = timedOut ? 2_000 : 5_000;
const [stdout, stderr] = await Promise.all([
raceMs(handle.stdout, drainCap, ''),
raceMs(handle.stderr, drainCap, ''),
]);
if (timedOut || code === null) {
handle.unref?.(); // a grandchild may still hold the pipes — never hold the CLI's exit
return { code: 124, stdout, stderr: stderr || `timeout after ${timeoutMs}ms` };
}
return { code, stdout, stderr };
}
async function withLock<T>(ws: string, fn: () => Promise<T>): Promise<T> {
const handle = await acquireBootstrapLock(ws);
try {
@@ -768,11 +920,21 @@ async function runRepo(ws: string, rest: string[], home: string, runner: ExecRun
});
}
async function runHooks(ws: string, rest: string[], home: string, runner: ExecRunner): Promise<number> {
const harnessFlag = flagValue(rest, '--harness') as Harness | undefined;
const harness = harnessFlag ?? detectHarness();
if (!harness || (harness !== 'claude-code' && harness !== 'codex')) {
console.error('cannot auto-detect the harness — pass --harness claude-code or --harness codex');
async function runHooks(
ws: string,
rest: string[],
home: string,
runner: ExecRunner,
probeSpawn?: OpencodeProbeSpawn,
): Promise<number> {
const harnessFlag = flagValue(rest, '--harness');
const harness = isHarness(harnessFlag) ? harnessFlag : harnessFlag ? null : detectHarness();
if (!harness) {
console.error(
harnessFlag
? `unknown --harness '${harnessFlag}' — pass --harness claude-code, codex, or opencode`
: 'cannot auto-detect the harness — pass --harness claude-code, codex, or opencode',
);
return 2;
}
// --repair is an idempotent-run alias: the same registration/write path as a
@@ -804,18 +966,44 @@ async function runHooks(ws: string, rest: string[], home: string, runner: ExecRu
return 2;
}
const mcpScope = ((consentAnswer(ws, 'MCP_SCOPE') ?? 'project').toLowerCase() === 'user' ? 'user' : 'project') as 'project' | 'user';
// Raw (unbanked) MCP_SCOPE answer — several harness branches need to know
// whether a human EXPLICITLY chose a scope vs the bank default filling in.
// typeof guard: readInterviewState validates `answers` is an object but not
// per-answer shapes — a hand-edited value of 3 must not throw.
const rawScopeAnswer = (() => {
const read = readInterviewState(ws);
const raw = read.ok ? read.state.answers['MCP_SCOPE'] : undefined;
// .trim(): a hand-edited or sloppily-recorded ' project' must not
// silently resolve to the user-global default (scope answers are
// security-relevant on opencode).
return raw?.skipped !== true && typeof raw?.value === 'string' ? raw.value.trim().toLowerCase() : undefined;
})();
// Scope resolution is per-harness (exhaustive switch — see HARNESSES):
// - claude-code: consent answer, bank default 'project' (the privacy-safe
// default: any other repo you open cannot read the brain).
// - codex: no scope flag exists; the value is ignored (note below).
// - opencode: default 'user' — OPPOSITE of claude-code, because opencode
// spawns project-config-defined servers with NO trust gate (verified,
// OPENCODE-CLI-PIN.md §Probes): a committed project entry would auto-spawn
// on every collaborator's machine. 'project' only via an EXPLICIT answer
// (the sharing warning prints at write time).
const mcpScope = ((): 'project' | 'user' => {
switch (harness) {
case 'claude-code':
return (consentAnswer(ws, 'MCP_SCOPE') ?? 'project').toLowerCase() === 'user' ? 'user' : 'project';
case 'codex':
return 'project'; // ignored — codex registrations are user-global (no scope flag)
case 'opencode':
return rawScopeAnswer === 'project' ? 'project' : 'user';
}
})();
// A persisted 'project' answer is meaningless on Codex (`codex mcp add` has no
// scope flag) — reachable via attach from a Claude Code machine or a pre-fix
// install. Fires on each hooks/repair run while the stale answer persists.
// Raw read, NOT consentAnswer: the bank default is 'project', so the resolved
// value would fire this note on every Codex install where no one was asked.
if (harness === 'codex') {
const read = readInterviewState(ws);
const raw = read.ok ? read.state.answers['MCP_SCOPE'] : undefined;
// typeof guard: readInterviewState validates `answers` is an object but not
// per-answer shapes — a hand-edited value of 3 must not throw.
if (raw?.skipped !== true && typeof raw?.value === 'string' && raw.value.toLowerCase() === 'project') {
if (rawScopeAnswer === 'project') {
console.error(
"note: the recorded MCP_SCOPE answer 'project' has no effect on Codex — " +
'`codex mcp add` has no scope flag; the registration is user-global (any repo ' +
@@ -845,6 +1033,27 @@ async function runHooks(ws: string, rest: string[], home: string, runner: ExecRu
);
return 0;
}
// Same ownership rule, opencode spelling: a REMOTE-type mcp.gbrain in the
// user-global config is either the harness lane's (inline bearer) or
// foreign — the stdio lane must not fight it in either case. BOTH global
// filenames are checked: opencode merges opencode.json AND opencode.jsonc
// when both exist, so a remote entry in EITHER file owns the name even
// when the path resolver would pick the other for writing.
if (
harness === 'opencode' &&
mcpScope === 'user' &&
[join(opencodeConfigDir(), 'opencode.jsonc'), join(opencodeConfigDir(), 'opencode.json')].some((p) =>
opencodeRemoteEntryExists(p, 'gbrain'),
)
) {
console.log(
"the 'gbrain' opencode MCP entry in the user-global config is a remote server (managed by " +
'`gbrain bootstrap harness`, or foreign) — skipping the stdio registration. Run ' +
'`gbrain bootstrap harness --remove` first (or remove the entry) if you want this ' +
'workspace-lane stdio registration instead.',
);
return 0;
}
return withLock(ws, async () => {
// 0. source_id visibility seam: `hooks` is the last ENGINE-FREE phase
@@ -889,6 +1098,157 @@ async function runHooks(ws: string, rest: string[], home: string, runner: ExecRu
// binary. The old early-return silently dropped hooks while the copy said
// only "MCP registration skipped".
let mcpSkipped = false;
if (harness === 'opencode') {
// Direct-writer lane (no exec): registrations land via the JSONC
// writer whose 4-state fingerprint is the [FIX7] check. Scope resolves
// to a FILE here — user → global config (absolute binary path),
// project → committed-candidate opencode.json (PATH-resolved command;
// no absolute machine paths in a file that travels, and no fail-open
// analog exists — the sharing warning below is the mitigation).
const configPath = mcpScope === 'project' ? opencodeProjectConfigPath(ws) : opencodeGlobalConfigPath();
const command =
mcpScope === 'project'
? ['gbrain', 'serve', '--surface', 'full']
: [gbrainBin, 'serve', '--surface', 'full'];
const entry = {
kind: 'local' as const,
name: 'gbrain',
command,
environment: { GBRAIN_SOURCE: sourceId, ...(gbrainHome ? { GBRAIN_HOME: gbrainHome } : {}) },
};
try {
// [X11] config-dir lock parity with the harness lane: the user-global
// config is shared across workspaces AND homes, so gbrain writers
// serialize on ITS directory. The project-scope file lives in the
// workspace root, which withLock(ws) already holds — the lock is
// non-reentrant, so the same-dir case skips the nested acquire.
const ocCfgDir = dirname(configPath);
let ocLock: { release(): void } | null = null;
if (resolve(ocCfgDir) !== resolve(ws)) {
mkdirSync(ocCfgDir, { recursive: true }); // the lock needs the dir; the writer mkdirs later anyway
ocLock = await acquireBootstrapLock(ocCfgDir);
}
let w: ReturnType<typeof writeOpencodeMcpEntry>;
try {
// [FIX7] parity: an existing entry pointing at a DIFFERENT workspace
// is warned about and replaced (same behavior as the exec lanes'
// mismatch path); a FOREIGN entry refuses inside the writer. The
// pre-check parse carries the same paste-by-hand snippet the writer
// uses so a corrupt config never strands the user.
const existingText = existsSync(configPath) ? readFileSync(configPath, 'utf8') : '';
const existingKind = opencodeEntryKind(
parseOpencodeConfig(existingText, configPath, opencodeEntrySnippet(entry)),
'gbrain',
{ sourceId },
);
if (existingKind === 'ours-other-source') {
console.error(`existing 'gbrain' opencode entry targets a DIFFERENT workspace — replacing it.`);
}
// Two-filename merge blind spot: opencode merges BOTH user-global
// filenames, so a same-name gbrain entry in the SIBLING file would
// survive this write as a shadow registration. Reconcile it under
// the same config-dir lock (ours → removed with a note; foreign →
// refuse loudly naming both files). User scope only — the project
// file has no observed sibling semantics.
if (mcpScope === 'user') {
const sib = reconcileOpencodeSiblingGlobal(configPath, 'gbrain', { sourceId });
for (const note of sib.notes) console.error(note);
}
w = writeOpencodeMcpEntry(configPath, entry, {
expect: { sourceId },
allowReplaceOtherSource: true,
});
} finally {
ocLock?.release();
}
console.log(
`MCP registered with opencode (scope: ${mcpScope === 'project' ? 'project (explicit opt-in)' : 'user-global'}) — ` +
`wrote ${w.configPath}${w.replacedPrior ? ' (replaced prior gbrain entry)' : ''}; ` +
'restart opencode (config is read at session start).',
);
for (const note of w.notes) console.error(note);
if (mcpScope === 'project') {
console.error(
'SHARING WARNING: opencode spawns project-config-defined MCP servers with NO trust prompt — ' +
'if this opencode.json is committed, every collaborator machine will spawn gbrain (teammates ' +
'without gbrain see a failing spawn each session; teammates WITH gbrain attach THEIR host ' +
'brain to this repo). The command is PATH-resolved ("gbrain" — requires gbrain on PATH); ' +
'the teammate opt-out is `"enabled": false` on the entry. The user-global default avoids all of this.' +
(gbrainHome
? ` Also: the entry embeds this machine's GBRAIN_HOME path (${gbrainHome}) — it won't be portable to other machines.`
: ''),
);
} else if (rawScopeAnswer === undefined) {
console.log(
"scope defaulted to user-global — opencode spawns project-defined servers with no trust gate, " +
'so the committed-file scope is explicit-opt-in only (record MCP_SCOPE=project to choose it).',
);
}
} catch (e) {
console.error((e as Error).message);
return 1;
}
// Registration smoke: the writer's post-render validation already
// proved the config parses and carries exactly our entry (that is the
// authoritative check). Best-effort live probe when the binary is on
// PATH: `opencode mcp list` SPAWNS servers (the honest discriminator)
// — run it with --pure (no external plugin autoload; `mcp list` is a
// code-execution surface otherwise) and skip it entirely when a
// plugin-bearing config is present (OPENCODE-CLI-PIN.md §Probes).
try {
const parsedCfg = parseOpencodeConfig(
existsSync(configPath) ? readFileSync(configPath, 'utf8') : '',
configPath,
);
if (mcpScope === 'project') {
// SECURITY: opencode merges the project opencode.json from the
// probe's cwd and spawns its local servers with NO trust prompt —
// running `mcp list` inside this workspace would execute whatever
// the (possibly just-cloned) repo's config names. Parse-back stays
// the authoritative check; the human runs the live probe.
console.log(
'live `opencode mcp list` probe skipped for project scope — config parse-back is authoritative; ' +
'run `opencode mcp list` yourself in this workspace to confirm.',
);
} else if (parsedCfg.plugin !== undefined) {
console.log('live `opencode mcp list` probe skipped (plugin-bearing config) — config parse-back is the verification.');
} else {
// SECURITY: the probe spawns from a fresh EMPTY temp dir, never the
// invoking cwd — no project opencode.json can load there (the same
// no-trust-prompt spawn surface as the project-scope skip above).
const probeCwd = mkdtempSync(join(tmpdir(), 'gbrain-opencode-probe-'));
let probe: { code: number; stdout: string; stderr: string };
try {
probe = await runOpencodeProbe(['opencode', 'mcp', 'list', '--pure'], {
cwd: probeCwd,
...(probeSpawn ? { spawn: probeSpawn } : {}),
});
} finally {
rmSync(probeCwd, { recursive: true, force: true });
}
// `mcp list` colorizes when a TTY-ish env leaks through — strip ANSI
// escapes before matching, and anchor the name on whitespace/EOL so
// a `gbrain-remote` entry can never satisfy a bare \bgbrain\b (\b
// matches before the hyphen).
const plain = probe.stdout.replace(/\u001b\[[0-9;]*m/g, '');
if (probe.code === 127) {
console.log('`opencode` is not on PATH — registration written; the config activates when opencode next starts here.');
} else if (probe.code === 0 && /✓\s+gbrain(\s|$)/.test(plain)) {
console.log('`opencode mcp list` handshake: ✓ gbrain connected.');
} else if (probe.code === 0 && /✗\s+gbrain(\s|$)/.test(plain)) {
console.error(
'WARNING: `opencode mcp list` reports ✗ gbrain failed — the spawn did not handshake ' +
'(is the gbrain binary path valid on this machine?). The exit code of `mcp list` is 0 even ' +
'on failure; this warning is from parsing its output.',
);
} else {
console.log('MCP registration written; could not confirm via `opencode mcp list` (best-effort probe).');
}
}
} catch {
/* smoke is best-effort */
}
} else {
const argvs =
harness === 'claude-code'
? registerClaudeMcp({ gbrainBin, scope: mcpScope, sourceId, ...(gbrainHome ? { gbrainHome } : {}) })
@@ -997,6 +1357,7 @@ async function runHooks(ws: string, rest: string[], home: string, runner: ExecRu
} catch {
/* smoke is best-effort */
}
} // end exec-lane registration (claude-code / codex)
// 3. Hooks (Claude Code only, consent-gated).
let hooksWritten = false;
@@ -1044,16 +1405,31 @@ async function runHooks(ws: string, rest: string[], home: string, runner: ExecRu
: 'hooks declined (HOOKS_CONSENT set to no) — the AGENTS.md pull protocol covers per-turn context instead; re-enable with `gbrain bootstrap hooks --harness claude-code`.',
);
}
} else {
} else if (harness === 'codex') {
console.log('gbrain does not wire Codex hooks yet — per-turn context is the AGENTS.md pull protocol (stated plainly; the codex hook lane is a filed follow-up).');
} else {
console.log(
'gbrain does not wire opencode\'s plugin/event system yet — per-turn context is the AGENTS.md ' +
'pull protocol, which opencode loads natively (the opencode plugin lane is a filed follow-up).',
);
}
// 4. Receipt registration record [CX2-12]. Detail records what actually
// landed; nothing landed at all (127 + no hooks) → no receipt entry.
if (!mcpSkipped || hooksWritten) {
const receiptScope = ((): string => {
switch (harness) {
case 'claude-code':
return mcpScope;
case 'codex':
return 'user'; // codex registrations are always user-global
case 'opencode':
return mcpScope; // user default; project only via explicit opt-in
}
})();
appendReceiptRegistration(home, ws, {
host: harness,
scope: harness === 'claude-code' ? mcpScope : 'user',
scope: receiptScope,
detail: hooksWritten ? (mcpSkipped ? 'hooks' : 'mcp+hooks') : 'mcp',
});
}
@@ -1252,15 +1628,63 @@ async function runUninstall(ws: string, rest: string[], home: string, runner: Ex
// Execute the structured host-registration removals the module returned.
for (const reg of result.registration_removals) {
if (reg.host === 'claude-code') {
const r = removeClaudeHooks(ws);
if (r.removed > 0) console.log(`removed ${r.removed} gbrain hook entr${r.removed === 1 ? 'y' : 'ies'} from ${r.settingsPath}`);
for (const note of r.notes) console.error(note);
const rm = await runner(['claude', 'mcp', 'remove', 'gbrain']);
if (rm.code !== 0) console.error('note: `claude mcp remove gbrain` did not succeed — remove it by hand if it lingers.');
} else {
const rm = await runner(['codex', 'mcp', 'remove', 'gbrain']);
if (rm.code !== 0) console.error('note: `codex mcp remove gbrain` did not succeed — remove it by hand if it lingers.');
switch (reg.host) {
case 'claude-code': {
const r = removeClaudeHooks(ws);
if (r.removed > 0) console.log(`removed ${r.removed} gbrain hook entr${r.removed === 1 ? 'y' : 'ies'} from ${r.settingsPath}`);
for (const note of r.notes) console.error(note);
const rm = await runner(['claude', 'mcp', 'remove', 'gbrain']);
if (rm.code !== 0) console.error('note: `claude mcp remove gbrain` did not succeed — remove it by hand if it lingers.');
break;
}
case 'codex': {
const rm = await runner(['codex', 'mcp', 'remove', 'gbrain']);
if (rm.code !== 0) console.error('note: `codex mcp remove gbrain` did not succeed — remove it by hand if it lingers.');
break;
}
case 'opencode': {
// Direct-writer removal (fingerprint-keyed; foreign entries refuse
// inside the module). Every candidate file best-effort — the
// receipt's scope names where the registration landed, but a stale
// entry in another file costs nothing to sweep. BOTH global
// filenames are swept: opencode merges opencode.json AND
// opencode.jsonc when both exist, so sweeping only the resolver's
// pick would strand a gbrain entry in the other file. The removal
// is expectation-keyed on THIS workspace's source id — a gbrain
// entry from a DIFFERENT workspace is skipped with a note, never
// silently deleted (it is not this uninstall's to remove).
const sweep = (p: string): void => {
try {
const r = removeOpencodeMcpEntry(p, 'gbrain', { sourceId: durabilitySourceId }, { skipOtherSource: true });
if (r.removed) console.log(`removed the gbrain opencode MCP entry from ${p}`);
for (const note of r.notes) console.error(note);
} catch (e) {
console.error(`note: could not remove the gbrain opencode entry from ${p}: ${(e as Error).message}`);
}
};
// Global files run under the config-dir bootstrap lock (the writer
// contract; harness.ts [X11] parity). Only when the dir exists — no
// dir means no config, and uninstall must not create one just to
// lock it.
const ocDir = opencodeConfigDir();
const globals = [join(ocDir, 'opencode.jsonc'), join(ocDir, 'opencode.json')].filter((p) => existsSync(p));
if (globals.length > 0) {
try {
const ocLock = await acquireBootstrapLock(ocDir);
try {
for (const p of globals) sweep(p);
} finally {
ocLock.release();
}
} catch (e) {
console.error(`note: could not lock the opencode config dir (${(e as Error).message}) — entries left for a re-run.`);
}
}
// The project file's dir IS the workspace, which withLock(ws)
// already holds — the lock is non-reentrant, so no nested acquire.
sweep(opencodeProjectConfigPath(ws));
break;
}
}
}
@@ -1321,6 +1745,9 @@ async function runUninstall(ws: string, rest: string[], home: string, runner: Ex
export interface RunBootstrapOpts {
/** Exec seam for gh/claude/codex subprocesses (tests inject a recorder). */
runner?: ExecRunner;
/** Spawn seam for the opencode `mcp list` probe (tests capture argv, cwd,
* and env; the default holds a real Bun.spawn handle so timeouts kill). */
probeSpawn?: OpencodeProbeSpawn;
}
/** Dispatch. Returns the process exit code (cli.ts passes it to setCliExitVerdict). */
@@ -1386,7 +1813,7 @@ export async function runBootstrap(args: string[], opts: RunBootstrapOpts = {}):
code = await runRepo(ws, rest, home, runner);
break;
case 'hooks':
code = await runHooks(ws, rest, home, runner);
code = await runHooks(ws, rest, home, runner, opts.probeSpawn);
break;
case 'verify':
code = await runVerify(ws, rest, home);
+6
View File
@@ -29,12 +29,14 @@ import { resolveAgentRunner, listRegisteredAgents, registerAgentRunner, validate
import { OpenClawRunner } from '../core/claw-test/runners/openclaw.ts';
import { HermesRunner } from '../core/claw-test/runners/hermes.ts';
import { GrokRunner } from '../core/claw-test/runners/grok.ts';
import { OpencodeRunner } from '../core/claw-test/runners/opencode.ts';
import { createTranscriptSink } from '../core/claw-test/transcript-capture.ts';
// Ensure built-in runners are registered.
registerAgentRunner('openclaw', () => new OpenClawRunner());
registerAgentRunner('hermes', () => new HermesRunner());
registerAgentRunner('grok', () => new GrokRunner());
registerAgentRunner('opencode', () => new OpencodeRunner());
interface HarnessOpts {
scenario: string;
@@ -417,6 +419,9 @@ const AGENT_INSTALL_HINTS: Record<string, string> = {
// Official xAI CLI only — the community superagent-ai grok-cli ships a
// colliding `grok` binary (docs/mcp/GROK-CLI-PIN.md).
grok: 'install grok (npm: @xai-official/grok, or https://x.ai/cli/install.sh) or set GROK_BIN',
// SST terminal agent — not OpenClaw, and not the renamed-to-Crush ancestor
// that shares the binary name (docs/mcp/OPENCODE-CLI-PIN.md).
opencode: 'install opencode (npm: opencode-ai, or https://opencode.ai/install) or set OPENCODE_BIN',
};
/**
@@ -1003,6 +1008,7 @@ Examples:
gbrain claw-test --scenario fresh-install
gbrain claw-test --scenario upgrade-from-v0.18 --keep-tempdir
gbrain claw-test --live --agent openclaw
gbrain claw-test --live --agent opencode
gbrain claw-test --live --agent hermes
gbrain claw-test --live --agent grok`);
}
+142 -9
View File
@@ -9,7 +9,7 @@
* needed for the connection.
*
* gbrain connect <mcp-url> [--token <bearer>] [--name gbrain]
* [--agent claude-code|codex|perplexity|generic]
* [--agent claude-code|codex|opencode|perplexity|generic]
* [--oauth [--register | --client-id ID --client-secret SECRET] [--scopes "read write"]]
* [--install] [--yes] [--json] [--show-token] [--force]
* [--timeout-ms N]
@@ -27,20 +27,34 @@
* only; --install runs it).
* - codex: `codex mcp add <name> --url <url> --bearer-token-env-var
* GBRAIN_REMOTE_TOKEN` (bearer via env var; --install runs it).
* - opencode: `opencode mcp add <name> --url <url> --header
* "Authorization=Bearer {env:GBRAIN_REMOTE_TOKEN}"` (the interpolation is
* stored literally; --install writes the entry directly via
* opencode-json.ts no binary needed).
* - perplexity: GUI connector (Settings Connectors). Supports bearer or
* OAuth; no --install.
* - generic: prints the connector fields for any other MCP client.
*/
import { execFileSync } from 'child_process';
import { mkdirSync } from 'node:fs';
import { dirname } from 'node:path';
import type { ConnectProbeResult } from '../core/connect-probe.ts';
import { probeBrainIdentity, DEFAULT_PROBE_TIMEOUT_MS } from '../core/connect-probe.ts';
import { opencodeGlobalConfigPath } from '../core/bootstrap/host-specs.ts';
import { acquireBootstrapLock } from '../core/bootstrap/lock.ts';
import {
GBRAIN_REMOTE_TOKEN_ENV,
reconcileOpencodeSiblingGlobal,
writeOpencodeMcpEntry,
} from '../core/bootstrap/opencode-json.ts';
import { promptLine } from '../core/cli-util.ts';
import {
NAME_RE,
REDACTED,
buildClaudeMcpAddArgv,
buildCodexMcpAddArgv,
buildOpencodeMcpAddArgv,
cmdString,
isValidName,
issuerFromMcpUrl,
@@ -58,6 +72,7 @@ export {
REDACTED,
buildClaudeMcpAddArgv,
buildCodexMcpAddArgv,
buildOpencodeMcpAddArgv,
cmdString,
isLinkLocalOrMetadata,
issuerFromMcpUrl,
@@ -69,7 +84,9 @@ export {
type UrlResult,
} from '../core/mcp-registration.ts';
export const ENV_VAR = 'GBRAIN_REMOTE_TOKEN';
// Defined from the writer's exported constant so the printed interpolation and
// the ownership fingerprint literal ({env:GBRAIN_REMOTE_TOKEN}) cannot drift.
export const ENV_VAR = GBRAIN_REMOTE_TOKEN_ENV;
export const PLACEHOLDER_TOKEN = '<paste-your-token>';
export const PLACEHOLDER_SECRET = '<paste-your-client-secret>';
export const DEFAULT_NAME = 'gbrain';
@@ -77,12 +94,12 @@ export const DEFAULT_SCOPES = 'read write';
// Single source of truth shared with the probe (was a duplicated 15_000 literal).
const DEFAULT_TIMEOUT_MS = DEFAULT_PROBE_TIMEOUT_MS;
export type AgentId = 'claude-code' | 'codex' | 'perplexity' | 'generic';
export type AgentId = 'claude-code' | 'codex' | 'opencode' | 'perplexity' | 'generic';
interface AgentSpec {
id: AgentId;
label: string; // human label for messages
binary?: string; // CLI binary backing --install ('claude' | 'codex')
binary?: string; // CLI binary backing --install ('claude' | 'codex'; opencode installs via the direct JSONC writer)
installable: boolean;
supportsOAuth: boolean; // accepts OAuth client-credentials connector fields
}
@@ -90,11 +107,14 @@ interface AgentSpec {
export const AGENT_SPECS: Record<AgentId, AgentSpec> = {
'claude-code': { id: 'claude-code', label: 'Claude Code', binary: 'claude', installable: true, supportsOAuth: false },
codex: { id: 'codex', label: 'Codex', binary: 'codex', installable: true, supportsOAuth: false },
// No `binary`: the opencode --install lane never execs a CLI (direct JSONC
// write), and it branches before the exec lane's `spec.binary` read.
opencode: { id: 'opencode', label: 'opencode', installable: true, supportsOAuth: false },
perplexity: { id: 'perplexity', label: 'Perplexity Computer', installable: false, supportsOAuth: true },
generic: { id: 'generic', label: 'your agent', installable: false, supportsOAuth: true },
};
export const AGENT_IDS: AgentId[] = ['claude-code', 'codex', 'perplexity', 'generic'];
export const AGENT_IDS: AgentId[] = ['claude-code', 'codex', 'opencode', 'perplexity', 'generic'];
// The named tools MUST be real MCP-exposed ops (verified by the round-trip
// E2E). `capture` is intentionally absent: it's a CLI-only convenience wrapper,
@@ -127,7 +147,7 @@ Usage:
gbrain connect <mcp-url> [--token <bearer>] [flags]
Prints a copy-paste setup block for your agent, or wires it up directly with
--install (claude-code + codex only). The MCP URL is your remote
--install (claude-code, codex + opencode). The MCP URL is your remote
'gbrain serve --http' endpoint; a bare host is rejected pass an explicit
https:// URL.
@@ -140,14 +160,15 @@ Auth:
Flags:
--token <bearer> Bearer token (else $${ENV_VAR}; from 'gbrain auth create')
--name <id> MCP server name in the agent (default: ${DEFAULT_NAME})
--agent <kind> claude-code (default) | codex | perplexity | generic
--agent <kind> claude-code (default) | codex | opencode | perplexity | generic
--oauth Use OAuth client credentials instead of a bearer token
--register With --oauth: mint a client on the host (gbrain auth register-client)
--client-id <id> With --oauth: use an existing OAuth client id
--client-secret <s> With --oauth: use an existing OAuth client secret
--scopes "<s>" With --oauth --register: client scopes (default: "${DEFAULT_SCOPES}")
--install Run the agent's MCP-add command, then smoke-test the token
(claude-code + codex only)
(claude-code + codex + opencode; opencode installs via a direct
config write no binary needed, token stays out of the file)
--yes Skip the install confirmation prompt
--force On --install, replace an existing server of the same name
--json Emit machine-readable JSON (secret redacted)
@@ -158,6 +179,7 @@ Examples:
gbrain connect https://brain.example.com/mcp --token gbrain_xxx
gbrain connect https://brain.example.com:3131 --install --yes
gbrain connect https://brain.example.com/mcp --token gbrain_xxx --agent codex
gbrain connect https://brain.example.com/mcp --token gbrain_xxx --agent opencode --install
gbrain connect https://brain.example.com/mcp --agent perplexity --oauth --register
gbrain connect https://brain.example.com/mcp --agent perplexity --oauth \\
--client-id gbrain_cl_xxx --client-secret gbrain_cs_xxx
@@ -225,6 +247,31 @@ function codexBlock(p: { name: string; url: string; token: string | null }): str
return lines.join('\n');
}
function opencodeBlock(p: { name: string; url: string; token: string | null }): string {
const tokenValue = p.token ?? PLACEHOLDER_TOKEN;
const cmd = cmdString('opencode', buildOpencodeMcpAddArgv({ name: p.name, url: p.url, envVar: ENV_VAR }));
const lines = [
'# Paste into opencode:',
'',
'Connect my knowledge brain, then learn what it can do:',
'',
` export ${ENV_VAR}=${shellQuote(tokenValue)}`,
` ${cmd}`,
'',
];
if (!p.token) lines.push(`Replace ${PLACEHOLDER_TOKEN} with a token from \`gbrain auth create "opencode"\` on the host.`, '');
lines.push(
`The config stores the literal \`{env:${ENV_VAR}}\` interpolation — opencode resolves it at read time, ` +
`so keep that variable exported in your shell profile; the token never lands in the config file. ` +
`Restart opencode after registering (config is read at session start).`,
'',
LEARN_INSTRUCTION,
'',
SECRET_NOTE,
);
return lines.join('\n');
}
function perplexityBearerBlock(p: { url: string; token: string | null }): string {
const tokenValue = p.token ?? PLACEHOLDER_TOKEN;
return [
@@ -296,6 +343,7 @@ export function buildConnectBlock(p: { agent: AgentId; name: string; url: string
switch (p.agent) {
case 'claude-code': return claudeBlock(p);
case 'codex': return codexBlock(p);
case 'opencode': return opencodeBlock(p);
case 'perplexity': return perplexityBearerBlock(p);
case 'generic': return genericBearerBlock(p);
}
@@ -330,6 +378,10 @@ export function buildJson(p: { url: string; name: string; agent: AgentId; token:
// Codex command carries no token (env-var name only), so it's safe verbatim.
command_argv = buildCodexMcpAddArgv({ name: p.name, url: p.url, envVar: ENV_VAR });
command = cmdString('codex', command_argv);
} else if (p.agent === 'opencode') {
// The literal {env:VAR} interpolation, not a token — safe verbatim.
command_argv = buildOpencodeMcpAddArgv({ name: p.name, url: p.url, envVar: ENV_VAR });
command = cmdString('opencode', command_argv);
}
return {
schema_version: 1,
@@ -363,6 +415,18 @@ export interface ConnectDeps {
probe(url: string, token: string, timeoutMs: number): Promise<ConnectProbeResult>;
env(name: string): string | undefined;
registerOAuthClient(name: string, scopes: string): RegisterResult;
/** opencode --install lane: direct JSONC write of a remote entry carrying
* the literal `{env:GBRAIN_REMOTE_TOKEN}` interpolation (no binary execed,
* no token on disk). Throws on a foreign same-name entry; an OURS entry at
* a different url refuses unless `allowReplaceOtherSource` (connect maps
* --force onto it). May be async: the default impl serializes on the
* config-dir bootstrap lock (the writer contract); sync test fakes remain
* assignable. */
writeOpencodeRemoteEntry(
name: string,
url: string,
opts?: { allowReplaceOtherSource?: boolean },
): { configPath: string; replacedPrior: boolean } | Promise<{ configPath: string; replacedPrior: boolean }>;
}
async function defaultPromptYesNo(question: string): Promise<boolean> {
@@ -420,6 +484,31 @@ const defaultDeps: ConnectDeps = {
probe: (url, token, timeoutMs) => probeBrainIdentity(url, token, { timeoutMs }),
env: (name) => process.env[name],
registerOAuthClient: defaultRegisterOAuthClient,
writeOpencodeRemoteEntry: async (name, url, opts) => {
// The writer's contract: callers hold acquireBootstrapLock on the config
// dir (harness.ts [X11] parity) — the user-global file is shared across
// workspaces and homes, so concurrent gbrain writers serialize here.
const configPath = opencodeGlobalConfigPath();
const cfgDir = dirname(configPath);
mkdirSync(cfgDir, { recursive: true }); // the lock needs the dir; the writer mkdirs later anyway
const lock = await acquireBootstrapLock(cfgDir);
try {
// Two-filename merge blind spot: opencode merges BOTH user-global
// filenames, so a same-name gbrain entry in the SIBLING file would
// survive this write as a shadow registration (ours → removed with a
// note; foreign → refuse loudly naming both files).
const sib = reconcileOpencodeSiblingGlobal(configPath, name, { url });
for (const note of sib.notes) console.error(note);
const r = writeOpencodeMcpEntry(
configPath,
{ kind: 'remote', name, url, tokenMode: 'env' },
{ expect: { url }, ...(opts?.allowReplaceOtherSource ? { allowReplaceOtherSource: true } : {}) },
);
return { configPath: r.configPath, replacedPrior: r.replacedPrior };
} finally {
lock.release();
}
},
};
// ---------------------------------------------------------------------------
@@ -595,8 +684,52 @@ export async function runConnect(args: string[], deps: ConnectDeps = defaultDeps
// --install path. token is guaranteed literal here (install mode resolveToken).
const realToken = token as string;
if (!spec.installable) {
fail(`--install supports claude-code and codex. ${spec.label} is set up through its own UI — drop --install to print the setup steps.`);
fail(`--install supports claude-code, codex, and opencode. ${spec.label} is set up through its own UI — drop --install to print the setup steps.`);
}
if (f.agent === 'opencode') {
// Direct-writer lane: no opencode binary required (the JSONC write IS the
// registration), and the config carries only the {env:VAR} interpolation
// — the writer's fingerprint handles idempotent re-runs and refuses a
// foreign same-name entry (--force cannot override THAT; pick --name).
// --force maps to the writer's allowReplaceOtherSource so an OURS entry
// at an old url (a rotated serve) is replaceable, mirroring the exec
// lanes' documented --force semantics.
if (!f.yes) {
if (!deps.isTTY()) {
fail('--install in a non-interactive shell requires --yes (refusing to register a credential-bearing MCP server without confirmation).');
}
const ok = await deps.promptYesNo(`Add MCP entry '${f.name}' -> ${url} to the opencode user-global config?`);
if (!ok) fail('Aborted.');
}
let w: { configPath: string; replacedPrior: boolean };
try {
w = await deps.writeOpencodeRemoteEntry(f.name, url, { allowReplaceOtherSource: f.force });
} catch (e) {
fail(redactToken((e as Error).message, realToken));
}
console.error(
`Added MCP entry '${f.name}' -> ${url} in ${w.configPath}` +
`${w.replacedPrior ? ' (replaced the prior gbrain entry)' : ''}. Restart opencode (config is read at session start).`,
);
if (deps.env(ENV_VAR) !== realToken) {
console.error(`opencode resolves {env:${ENV_VAR}} at read time. Add this to your shell profile so sessions can reach the brain:`);
console.error(` export ${ENV_VAR}=<your-token>`);
}
const ocProbe = await deps.probe(url, realToken, f.timeoutMs);
if (ocProbe.ok) {
console.error(`Verified: ${ocProbe.identity || 'brain reachable'}`);
console.error('');
console.error(LEARN_INSTRUCTION);
return;
}
console.error(
`Warning: registered '${f.name}', but the smoke-test did not verify (${ocProbe.reason}): ${redactToken(ocProbe.message, realToken)}`,
);
console.error('The agent will likely hit 401/errors until the token or URL is fixed.');
process.exit(1);
}
const binary = spec.binary as string; // 'claude' | 'codex'
if (!deps.hasBinary(binary)) {
fail(`${spec.label} CLI ('${binary}') not found on PATH. Install ${spec.label}, or drop --install to print the command to run manually.`);
+74 -3
View File
@@ -40,7 +40,7 @@ import {
parseQrelsFile,
type QrelsFile,
} from '../core/bench/qrels-file.ts';
import { runCorrectnessGate, type CorrectnessResult } from '../core/bench/correctness-gate.ts';
import { runCorrectnessGate, type CorrectnessGateOpts, type CorrectnessResult } from '../core/bench/correctness-gate.ts';
import { replayCore, type ReplaySummary } from './eval-replay.ts';
interface GateOpts {
@@ -55,6 +55,18 @@ interface GateOpts {
thresholdRecallAtK?: number;
thresholdFirstRelevantHit?: number;
thresholdExpectedTop1?: number;
/**
* Hermetic embedder selector. The only accepted value is 'deterministic':
* query embeddings come from the qrels fixture's basis-vector dims
* (src/eval/deterministic-embed.ts) instead of the gateway, so the
* correctness gate runs with no API keys. Correctness-gate-only; rejected
* when combined with the baseline regression gate (replay re-embeds
* captured queries via the gateway). Cache safety: this path drives bare
* `hybridSearch`, which never reads or writes the semantic query cache
* (both live in `hybridSearchCached`), so deterministic runs cannot
* poison cached production results by construction.
*/
embedder?: string;
}
interface Breach {
@@ -139,6 +151,10 @@ function parseArgs(args: string[]): GateOpts {
opts.thresholdExpectedTop1 = Number(next);
i++;
break;
case '--embedder':
opts.embedder = next;
i++;
break;
default:
break;
}
@@ -169,6 +185,13 @@ Thresholds (override baseline metadata; CLI > embedded > defaults):
--threshold-expected-top1 FLOAT Correctness: expected_top1-hit-rate floor (default ${DEFAULT_QRELS_THRESHOLDS.expected_top1})
-k, --k N Top-K for recall@K (default ${DEFAULT_QRELS_THRESHOLDS.k})
Hermetic mode (correctness gate only):
--embedder deterministic Embed queries as the qrels fixture's basis
vectors instead of calling the gateway
no API keys, fully reproducible (eval
canaries/CI). Rejected together with the
baseline regression gate.
Output:
--json Print JSON envelope to stdout
-h, --help Show this help
@@ -286,6 +309,7 @@ function runCorrectnessGateDispatch(
qrelsPath: string,
k: number,
cliOverrides: Pick<GateOpts, 'thresholdRecallAtK' | 'thresholdFirstRelevantHit' | 'thresholdExpectedTop1'>,
searchFn?: CorrectnessGateOpts['searchFn'],
): Promise<GateResult['correctness_gate']> {
return (async () => {
let qrelsFile: QrelsFile;
@@ -312,7 +336,7 @@ function runCorrectnessGateDispatch(
let result: CorrectnessResult;
try {
result = await runCorrectnessGate(engine, qrelsFile, { k });
result = await runCorrectnessGate(engine, qrelsFile, { k, ...(searchFn ? { searchFn } : {}) });
} catch (err) {
return {
ran: true,
@@ -448,6 +472,29 @@ export async function runEvalGate(engine: BrainEngine, args: string[]): Promise<
process.exit(2);
}
// Hermetic embedder validation. Only 'deterministic' is supported; the
// regression gate is out of scope (replay re-embeds captured queries via
// the gateway, which needs a provider key — defeating the hermetic point).
if (opts.embedder !== undefined) {
if (opts.embedder !== 'deterministic') {
console.error(
`Error: unsupported embedder "${opts.embedder}" — the only supported value is "deterministic".`,
);
process.exit(2);
}
if (opts.baseline) {
console.error(
'Error: the deterministic embedder cannot be combined with the baseline regression gate ' +
'(replay re-embeds captured queries via the gateway). Use it with the qrels correctness gate only.',
);
process.exit(2);
}
if (!opts.qrels) {
console.error('Error: the deterministic embedder requires a qrels file.');
process.exit(2);
}
}
const result: GateResult = {
schema_version: 1,
verdict: 'pass',
@@ -468,11 +515,35 @@ export async function runEvalGate(engine: BrainEngine, args: string[]): Promise<
if (opts.qrels) {
const k = opts.k ?? DEFAULT_QRELS_THRESHOLDS.k;
// Deterministic embedder: build a searchFn that threads basis-vector
// query embeddings (derived from the qrels fixture itself) into bare
// hybridSearch via the queryEmbedFn seam. The rest of the pipeline
// (keyword/title/alias arms, RRF, boosts) runs exactly as production.
let deterministicSearchFn: CorrectnessGateOpts['searchFn'] | undefined;
if (opts.embedder === 'deterministic') {
let queryEmbedFn: (text: string) => Float32Array;
try {
const { buildQrelsQueryEmbedFn } = await import('../eval/deterministic-embed.ts');
queryEmbedFn = buildQrelsQueryEmbedFn(readFileSync(opts.qrels, 'utf-8'));
} catch (err) {
console.error(
`Error: could not build the deterministic embedder from ${opts.qrels}: ${(err as Error).message}`,
);
process.exit(2);
}
const { hybridSearch } = await import('../core/search/hybrid.ts');
deterministicSearchFn = async (e, q, o) => {
const results = await hybridSearch(e, q, { limit: o.limit, queryEmbedFn });
return results.map(r => ({ source_id: r.source_id, slug: r.slug }));
};
}
result.correctness_gate = await runCorrectnessGateDispatch(engine, opts.qrels, k, {
thresholdRecallAtK: opts.thresholdRecallAtK,
thresholdFirstRelevantHit: opts.thresholdFirstRelevantHit,
thresholdExpectedTop1: opts.thresholdExpectedTop1,
});
}, deterministicSearchFn);
if (result.correctness_gate.breaches && result.correctness_gate.breaches.length > 0) {
result.verdict = 'fail';
}
+9 -9
View File
@@ -159,11 +159,11 @@ export interface HookIo {
/** TEST SEAM: user-prompt deadline override (wall-clock flake control). */
userPromptDeadlineMs?: number;
/**
* Feedback-loop attribution channel (`--harness <claude-code|codex>`).
* Feedback-loop attribution channel (`--harness <claude-code|codex|opencode>`).
* Default 'claude-code' the only harness bootstrap registers hooks for
* today; a codex hook registration passes the flag explicitly.
* today; a codex/opencode hook registration passes the flag explicitly.
*/
harness?: 'claude-code' | 'codex';
harness?: 'claude-code' | 'codex' | 'opencode';
}
// ── Entry point ─────────────────────────────────────────────────────────────
@@ -175,8 +175,8 @@ Events (wired into .claude/settings.local.json by gbrain bootstrap):
push status, hook health) to stdout
user-prompt read hook JSON on stdin, request per-turn context from a
running 'gbrain serve' over IPC, print additionalContext JSON
(--harness <claude-code|codex> sets the feedback-loop channel;
default claude-code, unknown values fall back to the default)
(--harness <claude-code|codex|opencode> sets the feedback-loop
channel; default claude-code, unknown values fall back to the default)
stop append to the per-session live buffer
session-end ingest the session transcript into the dream corpus
(secret-scanned), prune old corpus files, push the workspace
@@ -195,13 +195,13 @@ export async function runHook(args: string[], io: HookIo = {}): Promise<number>
write(io, USAGE + '\n');
return 0;
}
// `--harness <claude-code|codex>` — feedback-loop channel attribution for
// user-prompt. Unknown values fall back to the default (fail-open: a bad
// registration must never break the hook contract).
// `--harness <claude-code|codex|opencode>` — feedback-loop channel
// attribution for user-prompt. Unknown values fall back to the default
// (fail-open: a bad registration must never break the hook contract).
const harnessIdx = args.indexOf('--harness');
if (harnessIdx >= 0 && !io.harness) {
const v = args[harnessIdx + 1];
if (v === 'claude-code' || v === 'codex') io = { ...io, harness: v };
if (v === 'claude-code' || v === 'codex' || v === 'opencode') io = { ...io, harness: v };
}
if (!event || !['session-start', 'user-prompt', 'stop', 'session-end', 'compact'].includes(event)) {
process.stderr.write(USAGE + '\n');
+34 -2
View File
@@ -17,7 +17,7 @@ import type { PaceKeyOverrides } from '../core/pace-mode.ts';
import { loadConfig, isThinClient } from '../core/config.ts';
import { callRemoteTool, unpackToolResult } from '../core/mcp-client.ts';
import { parseNiceValue, applyNiceness, getEffectiveNiceness, formatNice } from '../core/minions/niceness.ts';
import { defaultTimeoutMsFor } from '../core/minions/handler-timeouts.ts';
import { defaultTimeoutMsFor, defaultLockDurationMsFor, clampLockDurationMs } from '../core/minions/handler-timeouts.ts';
function parseFlag(args: string[], flag: string): string | undefined {
const idx = args.indexOf(flag);
@@ -242,7 +242,17 @@ function formatTimeoutLines(job: MinionJob): string[] {
if (d != null) {
lines.push(` Timeout: (unset) — handler default ${d}ms stamps at claim`);
} else {
lines.push(` Timeout: (unset) — null-default wall-clock sweep applies (2 x lock-duration x max_stalled, ~5m at defaults)`);
lines.push(` Timeout: (unset) — null-default wall-clock sweep applies (2 x lock lease x max_stalled, ~5m at 30s-lease defaults)`);
}
}
// #4145: the lock lease line mirrors the timeout line — row value when
// stamped, otherwise the handler-map default that WILL stamp at claim.
if (job.lock_duration_ms != null) {
lines.push(` Lock lease: ${job.lock_duration_ms}ms (renewed at min(lease/2, 60s) cadence)`);
} else {
const lease = defaultLockDurationMsFor(job.name);
if (lease != null) {
lines.push(` Lock lease: (unset) — handler default ${lease}ms stamps at claim`);
}
}
return lines;
@@ -286,6 +296,7 @@ USAGE
[--max-waiting N]
[--backoff-type fixed|exponential] [--backoff-delay Nms]
[--backoff-jitter 0..1] [--timeout-ms Nms]
[--lock-duration-ms Nms]
[--idempotency-key K] [--queue Q] [--dry-run]
[--redact-secrets] (shell only; scrubs inherit
values from stdout/stderr)
@@ -455,6 +466,7 @@ USAGE
[--max-waiting N]
[--backoff-type fixed|exponential] [--backoff-delay Nms]
[--backoff-jitter 0..1] [--timeout-ms Nms]
[--lock-duration-ms Nms]
[--idempotency-key K] [--queue Q] [--dry-run]
[--redact-secrets]
@@ -471,6 +483,10 @@ OPTIONS
source before coalescing new submissions ([1,100])
--timeout-ms Nms Per-job wall-clock budget. Long-lane handlers get a
default from HANDLER_DEFAULT_TIMEOUT_MS when omitted.
--lock-duration-ms N Per-job lock lease (#4145). Clamped to [5s, 1h].
Long-lane handlers default to 300s via
HANDLER_DEFAULT_LOCK_DURATION_MS; others use the
worker default (30s).
--idempotency-key K At-most-one row per key (dead/cancelled free the key)
--queue Q Target queue (default: default)
--dry-run Print what would be submitted, submit nothing
@@ -580,6 +596,15 @@ export async function runJobs(engineOrNull: BrainEngine | null, args: string[]):
console.error('Error: --timeout-ms must be a positive integer (milliseconds)');
process.exit(1);
}
// #4145: per-job lock lease. Clamped to [5s,1h] in queue.add via
// clampLockDurationMs (shared with the MCP op); NULL falls to the
// handler map, then the worker default.
const lockDurationMsRaw = parseFlag(args, '--lock-duration-ms');
const lockDurationMs = lockDurationMsRaw !== undefined ? parseInt(lockDurationMsRaw, 10) : undefined;
if (lockDurationMsRaw !== undefined && (isNaN(lockDurationMs!) || lockDurationMs! <= 0)) {
console.error('Error: --lock-duration-ms must be a positive integer (milliseconds)');
process.exit(1);
}
const idempotencyKey = parseFlag(args, '--idempotency-key');
const queueName = parseFlag(args, '--queue') ?? 'default';
const dryRun = hasFlag(args, '--dry-run');
@@ -603,6 +628,12 @@ export async function runJobs(engineOrNull: BrainEngine | null, args: string[]):
if (backoffDelay !== undefined) console.log(` Backoff delay: ${backoffDelay}ms`);
if (backoffJitter !== undefined) console.log(` Backoff jitter: ${backoffJitter}`);
if (timeoutMs !== undefined) console.log(` Timeout: ${timeoutMs}ms`);
if (lockDurationMs !== undefined) {
// Echo what will actually be STORED (queue.add clamps to [5s,1h]);
// a dry-run that prints the raw out-of-range input lies.
const stored = clampLockDurationMs(lockDurationMs);
console.log(` Lock lease: ${stored}ms${stored !== lockDurationMs ? ` (clamped from ${lockDurationMs}ms)` : ''}`);
}
if (idempotencyKey) console.log(` Idempotency key: ${idempotencyKey}`);
if (delay > 0) console.log(` Delay: ${delay}ms`);
console.log(` Data: ${JSON.stringify(data)}`);
@@ -648,6 +679,7 @@ export async function runJobs(engineOrNull: BrainEngine | null, args: string[]):
backoff_delay: backoffDelay,
backoff_jitter: backoffJitter,
timeout_ms: timeoutMs,
lock_duration_ms: lockDurationMs,
idempotency_key: idempotencyKey,
queue: queueName,
}, trusted);
+46 -6
View File
@@ -47,6 +47,7 @@ import * as fs from 'fs';
import * as path from 'path';
import { createAuditWriter, resolveAuditDir, computeIsoWeekFilename } from './audit-writer.ts';
import { redactConnectionInfo } from './redact-connection-info.ts';
import type { LockRenewalTelemetryCtx } from '../minions/lock-renewal-tick.ts';
export type LockRenewalOutcome =
| 'failure'
@@ -77,6 +78,23 @@ export interface LockRenewalAuditEvent {
error_message_summary?: string;
/** Postgres SQLSTATE if present (e.g. '08006' for connection failure). */
error_code?: string;
// -- v0.46 additive telemetry (issue #4145 request 3). All optional: --
// -- pre-upgrade JSONL lines parse unchanged; the 4-outcome contract --
// -- is untouched. --
/** Failure-cause classification: call-timeout | refused | fenced-lost. */
cause?: string;
/** How late the renewal tick fired vs its own cadence (ms) — the local-starvation signal. */
lateness_ms?: number;
/** Interval callbacks skipped by the tickInFlight re-entrancy guard (overlap, NOT missed intervals). */
overlap_skips?: number;
/** os.loadavg()[0] at event time (raw, not normalized; 0 on Windows). */
load1?: number;
/** Core count paired with load1 so operators can normalize. */
cores?: number;
/** For success_after_failure: which path recovered — plain renewal or the at-deadline verify. */
via?: string;
/** True when the deadline was reached but eviction was deferred (verify threw pre-hard-deadline). */
deadline_deferred?: boolean;
}
const FEATURE_NAME = 'lock-renewal';
@@ -93,14 +111,33 @@ const writer = createAuditWriter<LockRenewalAuditEvent>({
* without writing to disk.
*/
export interface LockRenewalAuditSink {
logFailure(jobId: number, jobName: string, attempt: number, err: unknown): void;
logSuccessAfterFailure(jobId: number, jobName: string, recoveredAfterAttempts: number): void;
logGaveUp(jobId: number, jobName: string, totalFailures: number, err: unknown): void;
// Trailing ctx params are optional + additive (ENG-E2): the tick's
// structural SinkLike and pre-existing test fakes stay assignable.
logFailure(jobId: number, jobName: string, attempt: number, err: unknown, ctx?: LockRenewalTelemetryCtx): void;
logSuccessAfterFailure(jobId: number, jobName: string, recoveredAfterAttempts: number, ctx?: LockRenewalTelemetryCtx): void;
logGaveUp(jobId: number, jobName: string, totalFailures: number, err: unknown, ctx?: LockRenewalTelemetryCtx): void;
logExecuteJobRejected(jobId: number, jobName: string, err: unknown): void;
}
/**
* Copy only DEFINED ctx fields onto the event so absent telemetry stays
* absent from the JSONL (no `"load1":undefined` noise, stable byte size).
*/
function compactCtx(ctx?: LockRenewalTelemetryCtx): Partial<LockRenewalAuditEvent> {
if (!ctx) return {};
const out: Partial<LockRenewalAuditEvent> = {};
if (ctx.cause !== undefined) out.cause = ctx.cause;
if (ctx.lateness_ms !== undefined) out.lateness_ms = ctx.lateness_ms;
if (ctx.overlap_skips !== undefined) out.overlap_skips = ctx.overlap_skips;
if (ctx.load1 !== undefined) out.load1 = ctx.load1;
if (ctx.cores !== undefined) out.cores = ctx.cores;
if (ctx.via !== undefined) out.via = ctx.via;
if (ctx.deadline_deferred !== undefined) out.deadline_deferred = ctx.deadline_deferred;
return out;
}
export const lockRenewalAudit: LockRenewalAuditSink = {
logFailure(jobId, jobName, attempt, err) {
logFailure(jobId, jobName, attempt, err, ctx) {
writer.log({
job_id: jobId,
job_name: jobName,
@@ -108,18 +145,20 @@ export const lockRenewalAudit: LockRenewalAuditSink = {
outcome: 'failure',
error_message_summary: summarizeError(err),
error_code: extractErrorCode(err),
...compactCtx(ctx),
});
},
logSuccessAfterFailure(jobId, jobName, recoveredAfterAttempts) {
logSuccessAfterFailure(jobId, jobName, recoveredAfterAttempts, ctx) {
writer.log({
job_id: jobId,
job_name: jobName,
attempt: recoveredAfterAttempts,
outcome: 'success_after_failure',
// No error_message_summary or error_code: recovery has no error.
...compactCtx(ctx),
});
},
logGaveUp(jobId, jobName, totalFailures, err) {
logGaveUp(jobId, jobName, totalFailures, err, ctx) {
writer.log({
job_id: jobId,
job_name: jobName,
@@ -127,6 +166,7 @@ export const lockRenewalAudit: LockRenewalAuditSink = {
outcome: 'gave_up',
error_message_summary: summarizeError(err),
error_code: extractErrorCode(err),
...compactCtx(ctx),
});
},
logExecuteJobRejected(jobId, jobName, err) {
+2 -2
View File
@@ -8,7 +8,7 @@
* connection string into the error message:
* - `connection to server at "db.example.supabase.com" (1.2.3.4), port 5432 failed: ...`
* - `FATAL: password authentication failed for user "postgres"`
* - `could not connect to server: postgresql://user:pass@host:5432/db`
* - `could not connect to server: postgresql://user:pass@host:5432/db` (allow-pg-url-literal)
*
* If an operator pastes a JSONL audit dump into a GitHub issue or Slack,
* those errors leak credentials. The project's audit-as-debug-tool
@@ -34,7 +34,7 @@ interface RedactPattern {
* occurrences in a single string get redacted.
*/
const PATTERNS: ReadonlyArray<RedactPattern> = [
// postgres:// and postgresql:// URLs. Includes user:pass@host:port/db
// postgres:// and postgresql:// URLs. Includes user:pass@host:port/db /* allow-pg-url-literal */
// shapes plus query-string variants. Terminator is whitespace or
// common JSON/markdown delimiters.
{ kind: 'pg_url', re: /postgres(?:ql)?:\/\/[^\s"'>)]+/gi },
+103
View File
@@ -0,0 +1,103 @@
/**
* atomic-write.ts the ONE atomic config-file writer for bootstrap host
* surfaces (rule-of-three extraction: hooks.ts settings JSON, codex-toml.ts
* TOML text, opencode-json.ts JSONC text all swap through here).
*
* Semantics, hardened for shared user-scope targets [C10 / X11]:
* - The SYMLINK TARGET is resolved first so a dotfile-manager-linked config
* survives as a link (a bare rename would replace the link with a regular
* file). DANGLING links are resolved too (readlink, hop by hop): the write
* creates the missing target and the link survives.
* - tmp file uses a random suffix and inherits the EXISTING file's mode; a
* fresh file takes `freshMode` (caller's convention secret-bearing
* targets pass 0o600). `forceMode` overrides both (codex-toml forces 0600
* because the file carries a bearer token regardless of its prior mode).
* - chmod after write because writeFileSync's mode applies only on create.
*
* EOL and serialization stay caller-side: hooks.ts stringifies JSON,
* codex-toml converts to CRLF when the original was CRLF, opencode-json
* preserves EOLs naturally via jsonc-parser text splicing.
*/
import { randomBytes } from 'node:crypto';
import {
chmodSync,
existsSync,
lstatSync,
mkdirSync,
readlinkSync,
realpathSync,
renameSync,
statSync,
unlinkSync,
writeFileSync,
} from 'node:fs';
import { dirname, isAbsolute, resolve } from 'node:path';
/** Resolve the write target through symlinks, INCLUDING dangling ones.
* existsSync follows symlinks, so a DANGLING link reads "absent" and a bare
* rename would replace the link itself with a regular file instead the
* link text is resolved hop by hop (relative to each link's dir, bounded
* against loops) and the write lands at the final target, preserving the
* link the same way the live-symlink realpath branch does. */
function resolveWriteTarget(path: string): string {
if (existsSync(path)) {
// TOCTOU guard: the file can vanish between existsSync and realpathSync
// (a concurrent unlink), which would throw a raw ENOENT out of a writer
// that is perfectly able to proceed — fall through and treat the path as
// fresh/dangling instead.
try {
return realpathSync(path); // live file / live symlink chain
} catch {
/* raced away — resolve below */
}
}
let target = path;
for (let hops = 0; hops < 40; hops++) {
let st;
try {
st = lstatSync(target);
} catch {
return target; // truly absent — fresh-file target
}
if (!st.isSymbolicLink()) return target;
const linkText = readlinkSync(target);
target = isAbsolute(linkText) ? linkText : resolve(dirname(target), linkText);
}
return target; // pathological loop — bounded, last hop wins
}
export function atomicWriteTextFile(
path: string,
text: string,
opts?: { freshMode?: number; forceMode?: number },
): void {
const target = resolveWriteTarget(path);
mkdirSync(dirname(target), { recursive: true });
let mode: number | undefined;
if (opts?.forceMode !== undefined) {
mode = opts.forceMode;
} else {
try {
mode = statSync(target).mode & 0o777;
} catch {
mode = opts?.freshMode;
}
}
const tmp = `${target}.tmp-${randomBytes(6).toString('hex')}`;
// Failure hygiene: a throwing write/chmod/rename (ENOSPC, EACCES, target
// turned into a directory, …) must not leak the tmp file next to the
// user's config — unlink it best-effort and rethrow the original error.
try {
writeFileSync(tmp, text, { encoding: 'utf8', ...(mode !== undefined ? { mode } : {}) });
if (mode !== undefined) chmodSync(tmp, mode);
renameSync(tmp, target);
} catch (e) {
try {
unlinkSync(tmp);
} catch {
/* best-effort — the original error is the one that matters */
}
throw e;
}
}
+1 -1
View File
@@ -51,7 +51,7 @@ export interface AttachWorkspaceOptions {
/** The gbrain home receiving the install receipt (default: configDir()). */
gbrainHomeDir?: string;
/** Target harness for the hooks/MCP steps' descriptions. */
harness?: 'claude-code' | 'codex';
harness?: 'claude-code' | 'codex' | 'opencode';
/** Recorded as the receipt's created_by (the attaching binary's version). */
createdBy?: string;
}
+5 -20
View File
@@ -32,19 +32,8 @@
* brick).
*/
import { randomBytes } from 'node:crypto';
import {
chmodSync,
copyFileSync,
existsSync,
mkdirSync,
readFileSync,
realpathSync,
renameSync,
statSync,
writeFileSync,
} from 'node:fs';
import { dirname } from 'node:path';
import { chmodSync, copyFileSync, existsSync, readFileSync, statSync } from 'node:fs';
import { atomicWriteTextFile } from './atomic-write.ts';
import { CODEX_TOML_BLOCK_BEGIN, CODEX_TOML_BLOCK_END } from './host-specs.ts';
export interface CodexHttpServerBlock {
@@ -158,15 +147,11 @@ function renderBlock(block: CodexHttpServerBlock): string[] {
];
}
/** Atomic 0600 write preserving symlinks and the file's dominant EOL. */
/** Atomic 0600 write preserving symlinks and the file's dominant EOL
* (forceMode: the file carries a bearer token regardless of prior mode). */
function atomicWriteToml(configPath: string, unixText: string, crlf: boolean): void {
const target = existsSync(configPath) ? realpathSync(configPath) : configPath;
mkdirSync(dirname(target), { recursive: true });
const tmp = `${target}.tmp-${randomBytes(6).toString('hex')}`;
const out = crlf ? unixText.replace(/\n/g, '\r\n') : unixText;
writeFileSync(tmp, out, { encoding: 'utf8', mode: 0o600 });
chmodSync(tmp, 0o600);
renameSync(tmp, target);
atomicWriteTextFile(configPath, out, { forceMode: 0o600 });
}
/**
+2 -2
View File
@@ -130,7 +130,7 @@ export interface InstallReceipt {
* corpus dir, ). Uninstall removes exactly these, nothing else. */
created_paths: string[];
/** Host registrations bootstrap performed (for marker-keyed removal). */
registrations: Array<{ host: 'claude-code' | 'codex'; scope: string; detail?: string }>;
registrations: Array<{ host: 'claude-code' | 'codex' | 'opencode'; scope: string; detail?: string }>;
}
export function receiptPath(gbrainHomeDir: string): string {
@@ -223,7 +223,7 @@ export type HarnessTargetKind = 'mcp' | 'permission' | 'hooks';
export type HarnessTargetState = 'pending' | 'confirmed' | 'failed';
export interface HarnessTarget {
host: 'claude-code' | 'codex';
host: 'claude-code' | 'codex' | 'opencode';
kind: HarnessTargetKind;
state: HarnessTargetState;
/** user scope or a --project dir (hooks); user for mcp/permission. */
+274 -19
View File
@@ -1,8 +1,8 @@
/**
* harness.ts `gbrain bootstrap harness` (#4043): default brain wiring for
* agent-framework-driven coding (a downstream framework spawning Claude Code
* `claude -p` / codex exec on a box that already hosts a brain + a running
* `gbrain serve --http`).
* `claude -p` / codex exec / opencode run on a box that already hosts a brain
* + a running `gbrain serve --http`).
*
* What it wires, per harness:
* - Claude Code: user-scope HTTP MCP registration (`claude mcp add --scope
@@ -12,6 +12,10 @@
* required anywhere.
* - Codex: one managed `[mcp_servers.<name>]` TOML block with the inline
* bearer token (codex-toml.ts `codex mcp add` cannot express it).
* - opencode: one managed `mcp.<name>` remote entry with the inline bearer
* header in the user-global JSONC config (opencode-json.ts
* framework-spawned opencode inherits no shell env, so the `{env:…}`
* interpolation the connect lane uses would resolve empty here).
*
* Contracts folded from the CEO review + outside voice (letters reference the
* plan file):
@@ -36,7 +40,7 @@
* revoke defers with a typed message under a live PGLite serve.
*/
import { copyFileSync, existsSync, mkdirSync, readFileSync, statSync } from 'node:fs';
import { existsSync, mkdirSync, readFileSync, rmSync, statSync } from 'node:fs';
import { dirname, join, resolve } from 'node:path';
import { VERSION } from '../../version.ts';
@@ -66,6 +70,7 @@ import {
type HarnessReceipt,
type HarnessTarget,
} from './format.ts';
import { atomicWriteTextFile } from './atomic-write.ts';
import {
removeCodexHttpServerBlock,
writeCodexHttpServerBlock,
@@ -87,12 +92,22 @@ import {
claudeUserSettingsPath,
codexConfigPath,
mcpPermissionEntry,
opencodeConfigDir,
opencodeGlobalConfigPath,
type ClaudeHookEvent,
} from './host-specs.ts';
import {
opencodeEntryKind,
parseOpencodeConfig,
parseOpencodeEntryBearer,
reconcileOpencodeSiblingGlobal,
removeOpencodeMcpEntry,
writeOpencodeMcpEntry,
} from './opencode-json.ts';
// ── Flags ───────────────────────────────────────────────────────────────────
export type HarnessSelector = 'claude-code' | 'codex' | 'all';
export type HarnessSelector = 'claude-code' | 'codex' | 'opencode' | 'all';
export interface HarnessFlags {
harness: HarnessSelector;
@@ -143,8 +158,8 @@ export function parseHarnessArgs(rest: string[]): HarnessFlags {
};
const h = value('--harness');
if (h !== undefined) {
if (h !== 'claude-code' && h !== 'codex' && h !== 'all') {
out.error = `unknown --harness '${h}' — pass claude-code, codex, or all`;
if (h !== 'claude-code' && h !== 'codex' && h !== 'opencode' && h !== 'all') {
out.error = `unknown --harness '${h}' — pass claude-code, codex, opencode, or all`;
return out;
}
out.harness = h;
@@ -230,6 +245,8 @@ export interface HarnessDeps {
userSettingsPath?: string;
/** Resolved codex config path (tests point at a temp CODEX_HOME). */
codexConfig?: string;
/** Resolved opencode config path (tests point at a temp XDG_CONFIG_HOME). */
opencodeConfig?: string;
/** Engine-backed mint; tests inject a fake. */
mint?: (opts: {
name: string;
@@ -241,6 +258,7 @@ export interface HarnessDeps {
pgliteLiveServe?: () => boolean;
detectClaude?: () => boolean;
detectCodex?: () => boolean;
detectOpencode?: () => boolean;
gbrainBin?: string | null;
log?: (line: string) => void;
logError?: (line: string) => void;
@@ -256,6 +274,7 @@ function resolveDeps(deps: HarnessDeps): Required<Omit<HarnessDeps, 'gbrainBin'>
probeIdentity: deps.probeIdentity ?? ((url, token) => probeBrainIdentity(url, token)),
userSettingsPath: deps.userSettingsPath ?? claudeUserSettingsPath(),
codexConfig: deps.codexConfig ?? codexConfigPath(),
opencodeConfig: deps.opencodeConfig ?? opencodeGlobalConfigPath(),
mint: deps.mint ?? defaultMint,
revokeById: deps.revokeById ?? defaultRevokeById,
pgliteLiveServe: deps.pgliteLiveServe ?? defaultPgliteLiveServe,
@@ -263,6 +282,9 @@ function resolveDeps(deps: HarnessDeps): Required<Omit<HarnessDeps, 'gbrainBin'>
detectCodex:
deps.detectCodex ??
(() => whichSafe('codex') !== null || existsSync(deps.codexConfig ?? codexConfigPath())),
detectOpencode:
deps.detectOpencode ??
(() => whichSafe('opencode') !== null || existsSync(opencodeConfigDir())),
gbrainBin: deps.gbrainBin !== undefined ? deps.gbrainBin : null,
log: deps.log ?? ((l) => console.log(l)),
logError: deps.logError ?? ((l) => console.error(l)),
@@ -357,12 +379,14 @@ export function buildConsentBlock(p: {
url: string;
wireClaude: boolean;
wireCodex: boolean;
wireOpencode: boolean;
hooks: boolean;
capture: boolean;
hookScope: string;
name: string;
userSettingsPath: string;
codexConfig: string;
opencodeConfig: string;
}): string {
const lines: string[] = [
'gbrain bootstrap harness — wire framework-spawned coding sessions to this brain',
@@ -402,14 +426,23 @@ export function buildConsentBlock(p: {
`${p.codexConfig} (0600) — framework-spawned codex inherits no shell env, so an env-var token would not reach it.`,
);
}
if (p.wireOpencode) {
lines.push(
` ${n++}. opencode (user-global): write the mcp.${p.name} remote entry with the bearer token INLINE into ` +
`${p.opencodeConfig} (0600) — framework-spawned opencode inherits no shell env, so the {env:…} ` +
`interpolation would resolve empty.`,
);
}
// [X7] The reach statement matches what is ACTUALLY being wired — it must
// never claim a host or a hook lane this invocation does not touch.
const hosts =
p.wireClaude && p.wireCodex
? 'EVERY Claude Code and Codex session'
: p.wireClaude
? 'EVERY Claude Code session'
: 'EVERY Codex session';
// never claim a host or a hook lane this invocation does not touch. Joined
// list, not a ternary tree: a fourth harness must be a compile-time nudge
// here, not a silent mislabel.
const hostNames = [
...(p.wireClaude ? ['Claude Code'] : []),
...(p.wireCodex ? ['Codex'] : []),
...(p.wireOpencode ? ['opencode'] : []),
];
const hosts = `EVERY ${hostNames.join(' and ')} session`;
const hookLine = !p.wireClaude || !p.hooks
? 'No hooks are wired by this invocation.'
: p.capture
@@ -421,6 +454,7 @@ export function buildConsentBlock(p: {
'`gbrain auth revoke --id <id>` (see auth list)',
...(p.wireClaude ? [`\`claude mcp remove ${p.name} --scope user\``] : []),
...(p.wireCodex ? ['edit the codex config'] : []),
...(p.wireOpencode ? ['edit the opencode config'] : []),
];
lines.push(
'',
@@ -498,8 +532,42 @@ async function cleanupStalePriorTargets(
);
}
} else if (pt.host === 'codex' && pt.kind === 'mcp') {
const r = removeCodexHttpServerBlock(pt.path ?? d.codexConfig, pt.name ?? 'gbrain');
if (r.removed) d.log(`stale codex managed block removed from ${pt.path ?? d.codexConfig} (no longer planned).`);
// [X11] The caller holds only the claude config-dir lock; the codex
// config is a DIFFERENT shared file, and its read-modify-write must
// serialize on ITS directory's bootstrap lock too — a concurrent
// codex-dir-locked writer interleaving here would have one rename
// discard the other. Ordering matches apply/remove (claude/config
// dir first, then codex dir), so no lock-order inversion; the lock
// is non-reentrant, so the same-dir case skips the nested acquire.
const codexPath = pt.path ?? d.codexConfig;
const codexDir = dirname(codexPath);
mkdirSync(codexDir, { recursive: true });
const heldDir = resolve(dirname(d.userSettingsPath));
const lk = resolve(codexDir) === heldDir ? null : await acquireBootstrapLock(codexDir);
let r: ReturnType<typeof removeCodexHttpServerBlock>;
try {
r = removeCodexHttpServerBlock(codexPath, pt.name ?? 'gbrain');
} finally {
lk?.release();
}
if (r.removed) d.log(`stale codex managed block removed from ${codexPath} (no longer planned).`);
} else if (pt.host === 'opencode' && pt.kind === 'mcp') {
// Fingerprint-keyed against the PRIOR receipt's url — a foreign or
// rotated-away entry refuses inside the module (never guess). Same
// [X11] nested-lock discipline as the codex branch above (claude/
// config dir held by the caller, then the opencode dir here).
const ocPath = pt.path ?? d.opencodeConfig;
const ocDir = dirname(ocPath);
mkdirSync(ocDir, { recursive: true });
const heldDir = resolve(dirname(d.userSettingsPath));
const lk = resolve(ocDir) === heldDir ? null : await acquireBootstrapLock(ocDir);
let r: ReturnType<typeof removeOpencodeMcpEntry>;
try {
r = removeOpencodeMcpEntry(ocPath, pt.name ?? 'gbrain', { url: prior.url });
} finally {
lk?.release();
}
if (r.removed) d.log(`stale opencode entry removed from ${ocPath} (no longer planned).`);
}
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
@@ -539,14 +607,17 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
// explicit (we exec `claude mcp add`; it owns ~/.claude.json).
const wireClaude = (flags.harness === 'all' || flags.harness === 'claude-code') && d.detectClaude();
const wireCodex = flags.harness === 'codex' || (flags.harness === 'all' && d.detectCodex());
// opencode mirrors codex: an explicit --harness opencode FORCES wiring (the
// JSONC writer needs no opencode CLI and creates the config file itself).
const wireOpencode = flags.harness === 'opencode' || (flags.harness === 'all' && d.detectOpencode());
if (flags.harness === 'claude-code' && !wireClaude) {
d.logError('claude CLI not found on PATH — the user-scope MCP registration needs it (it owns ~/.claude.json).');
return 2;
}
if (!wireClaude && !wireCodex) {
if (!wireClaude && !wireCodex && !wireOpencode) {
d.logError(
'no harness detected on this box (claude CLI not on PATH; no codex install) — ' +
'pass --harness claude-code|codex explicitly if detection is wrong.',
'no harness detected on this box (claude CLI not on PATH; no codex install; no opencode install) — ' +
'pass --harness claude-code|codex|opencode explicitly if detection is wrong.',
);
return 2;
}
@@ -597,12 +668,14 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
url,
wireClaude,
wireCodex,
wireOpencode,
hooks: wireHooks,
capture: !flags.noCapture,
hookScope,
name: flags.name,
userSettingsPath: d.userSettingsPath,
codexConfig: d.codexConfig,
opencodeConfig: d.opencodeConfig,
});
d.log(consent);
if (!flags.yes) {
@@ -719,6 +792,17 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
mechanism: 'toml-block',
});
}
if (wireOpencode) {
targets.push({
host: 'opencode',
kind: 'mcp',
state: 'pending',
scope: 'user',
path: d.opencodeConfig,
name: flags.name,
mechanism: 'jsonc-entry',
});
}
// [X4] EVERY unrevoked prior minted id is carried — on the --token lane
// too. A failed rotation must never forget the token before last.
const carriedPreviousIds = [
@@ -805,6 +889,15 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
let oldClaudeReg: { url: string; token: string } | null = null;
let claudeReplaced = false;
let codexRollback: { path: string; backupPath: string | null; replacedPrior: boolean } | null = null;
let opencodeRollback: {
path: string;
backupPath: string | null;
replacedPrior: boolean;
/** Exact text this run's writer landed the rollback compares the LIVE
* file against it before restoring (the lock is released before the
* smoke, so a newer registration may have landed since). */
writtenText: string;
} | null = null;
let cfgLock: Awaited<ReturnType<typeof acquireBootstrapLock>> | null = null;
if (wireClaude) {
const cfgDir = dirname(d.userSettingsPath);
@@ -957,6 +1050,68 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
}
}
// 7b. opencode wiring — the managed JSONC entry is the single write
// mechanism (opencode-json.ts; fingerprint-owned, comment-preserving).
// Same [X11] lock discipline; same forced-wire posture as codex.
for (const t of targets) {
if (t.host !== 'opencode') continue;
try {
const ocDir = dirname(t.path!);
mkdirSync(ocDir, { recursive: true });
const ocLock = await acquireBootstrapLock(ocDir);
let r: ReturnType<typeof writeOpencodeMcpEntry>;
try {
// Ownership [C8]: an entry at OUR new url (idempotent re-run) or at
// the PRIOR receipt's url (rotation across a port change) is ours;
// anything else under the name refuses inside the writer. The prior
// url is offered as the expectation only when the current one does
// not classify the entry as ours.
let expectUrl = url;
if (prior && prior.url !== url && existsSync(t.path!)) {
try {
const parsedExisting = parseOpencodeConfig(readFileSync(t.path!, 'utf8'), t.path!);
if (
opencodeEntryKind(parsedExisting, flags.name, { url }) === 'foreign' &&
opencodeEntryKind(parsedExisting, flags.name, { url: prior.url }) === 'ours-same-source'
) {
expectUrl = prior.url;
}
} catch {
/* the writer's own read path raises the real error below */
}
}
// Two-filename merge blind spot: opencode merges BOTH user-global
// filenames, so a same-name gbrain entry in the SIBLING file would
// survive this write as a shadow registration. Reconcile under the
// same config-dir lock (ours → removed with a note; foreign →
// refuse loudly naming both files).
const sib = reconcileOpencodeSiblingGlobal(t.path!, flags.name, { url: expectUrl });
for (const note of sib.notes) d.log(note);
r = writeOpencodeMcpEntry(
t.path!,
{ kind: 'remote', name: flags.name, url, tokenMode: 'inline', bearerToken: token },
{ expect: { url: expectUrl }, allowReplaceOtherSource: true },
);
} finally {
ocLock.release();
}
for (const note of r.notes) d.logError(note);
opencodeRollback = { path: t.path!, backupPath: r.backupPath, replacedPrior: r.replacedPrior, writtenText: r.writtenText };
confirm(t);
d.log(
`opencode wired: mcp.${flags.name} remote entry with inline bearer header in ${t.path} (0600). ` +
'opencode ships a plugin/event system, but gbrain does not wire it yet — per-turn context on ' +
'opencode is MCP tools + the pull protocol (AGENTS.md loads natively). Restart opencode: it ' +
'reads config at session start.',
);
} catch (e) {
// Redaction parity with the claude lane: the writer's refusal messages
// can embed a paste-by-hand snippet, and the receipt + stderr must
// never carry the live bearer under any error shape.
failTarget(t, redactToken(e instanceof Error ? e.message : String(e), token));
}
}
// 8. Smoke [C3-enriched message; X10 verbs-surface honesty]. An
// unknown-tool tool_error means initialize + auth ALREADY succeeded — a
// serve running a narrowed --surface (e.g. verbs) is verified, not broken.
@@ -1010,7 +1165,9 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
const rbLock = await acquireBootstrapLock(dirname(codexRollback.path)); // [X11] parity
try {
if (codexRollback.backupPath && existsSync(codexRollback.backupPath)) {
copyFileSync(codexRollback.backupPath, codexRollback.path);
// Atomic restore: a crash mid-copy must never leave a torn config
// (the backup carries the previous bearer — 0600 stays forced).
atomicWriteTextFile(codexRollback.path, readFileSync(codexRollback.backupPath, 'utf8'), { forceMode: 0o600 });
} else if (!codexRollback.replacedPrior) {
removeCodexHttpServerBlock(codexRollback.path, flags.name);
}
@@ -1023,6 +1180,42 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
d.logError(`codex rollback failed: ${e instanceof Error ? e.message : String(e)} — re-run to converge.`);
}
}
if (opencodeRollback) {
try {
let failNote = 'rolled back to the previous opencode config after the failed smoke';
const rbLock = await acquireBootstrapLock(dirname(opencodeRollback.path)); // [X11] parity
try {
// Restore-guard: the config-dir lock was released before the smoke,
// so a NEWER registration (another run's) may have replaced ours —
// restoring this run's snapshot over it would clobber that newer
// wiring. Only restore when the live file still carries the EXACT
// text this run wrote; either way the fresh mint is revoked below.
const current = existsSync(opencodeRollback.path)
? readFileSync(opencodeRollback.path, 'utf8')
: '';
if (current !== opencodeRollback.writtenText) {
failNote =
'smoke failed; opencode rollback SKIPPED — the config changed after this run wrote it ' +
'(a newer registration exists); this run\'s fresh mint is still revoked';
d.log(failNote + '.');
} else if (opencodeRollback.backupPath && existsSync(opencodeRollback.backupPath)) {
// Atomic restore (codex-lane parity): never a torn config mid-crash.
atomicWriteTextFile(opencodeRollback.path, readFileSync(opencodeRollback.backupPath, 'utf8'), { forceMode: 0o600 });
// Consumed — the unique backup carries the previous bearer and
// has no consumer once restored.
try { rmSync(opencodeRollback.backupPath, { force: true }); } catch { /* best-effort */ }
} else if (!opencodeRollback.replacedPrior) {
removeOpencodeMcpEntry(opencodeRollback.path, flags.name, { url });
}
} finally {
rbLock.release();
}
const ot = targets.find((t) => t.host === 'opencode' && t.kind === 'mcp');
if (ot) failTarget(ot, failNote);
} catch (e) {
d.logError(`opencode rollback failed: ${e instanceof Error ? e.message : String(e)} — re-run to converge.`);
}
}
const mt = targets.find((t) => t.host === 'claude-code' && t.kind === 'mcp');
if (claudeReplaced && oldClaudeReg) {
await d.runner(['claude', 'mcp', 'remove', flags.name, '--scope', 'user']);
@@ -1080,6 +1273,18 @@ export async function applyHarness(flags: HarnessFlags, rawDeps: HarnessDeps): P
}
}
// The unique opencode backup carries the PREVIOUS bearer; once the new
// wiring is verified it has no consumer — unlink it so re-runs never
// accumulate token-bearing snapshots (failed runs consume it via the
// restore above; skipped restores leave it 0600 for manual recovery).
if (smokeOk && opencodeRollback?.backupPath) {
try {
rmSync(opencodeRollback.backupPath, { force: true });
} catch {
/* best-effort — it is 0600 either way */
}
}
// 9. [X3] Convergence cleanup — AFTER the smoke, so prior working wiring is
// never unwired on a run that failed to establish its replacement. Cleanup
// failures append as failed targets (blocking the rotation gate below) so
@@ -1286,6 +1491,47 @@ export async function removeHarness(flags: HarnessFlags, rawDeps: HarnessDeps):
? `Codex managed block removed from ${codexPath}.`
: `no managed block in ${codexPath} — counted as removed.`,
);
} else if (t.host === 'opencode') {
const ocPath = t.path ?? d.opencodeConfig;
// [C8] Ownership before removal: an entry now at a DIFFERENT url is
// another install's — skip with a note, cleared from the receipt
// (mirror of the claude not-ours branch; the module's foreign check
// would THROW, which reads as a failure rather than a skip).
let kind: ReturnType<typeof opencodeEntryKind> = 'absent';
if (existsSync(ocPath)) {
try {
kind = opencodeEntryKind(
parseOpencodeConfig(readFileSync(ocPath, 'utf8'), ocPath),
t.name ?? 'gbrain',
{ url: receipt.url },
);
} catch (e) {
throw new Error(`opencode config unreadable: ${e instanceof Error ? e.message : String(e)}`);
}
}
if (kind === 'absent') {
d.log(`opencode entry '${t.name}' already gone — counted as removed.`); // [F2]
} else if (kind !== 'ours-same-source') {
d.log(
`opencode entry '${t.name}' does not match this receipt's url (${receipt.url}) — owned by another ` +
'install; skipping, cleared from the receipt.',
);
} else {
const ocDir = dirname(ocPath);
mkdirSync(ocDir, { recursive: true });
const ocLock = resolve(ocDir) === resolve(rmCfgDir) ? null : await acquireBootstrapLock(ocDir);
let r: ReturnType<typeof removeOpencodeMcpEntry>;
try {
r = removeOpencodeMcpEntry(ocPath, t.name ?? 'gbrain', { url: receipt.url });
} finally {
ocLock?.release();
}
d.log(
r.removed
? `opencode managed entry removed from ${ocPath}.`
: `no managed entry in ${ocPath} — counted as removed.`,
);
}
}
} catch (e) {
anyFailed = true;
@@ -1443,6 +1689,15 @@ export async function statusHarness(flags: HarnessFlags, rawDeps: HarnessDeps):
}
}
}
if (!token) {
const ocMcp = receipt.targets.find((t) => t.host === 'opencode' && t.kind === 'mcp');
if (ocMcp?.path) {
// [C8] url-matched inside the helper: a foreign/rotated entry's bearer
// is never recovered (it was not issued for receipt.url).
token = parseOpencodeEntryBearer(ocMcp.path, ocMcp.name ?? 'gbrain', receipt.url);
if (token) tokenSource = 'opencode config entry';
}
}
let tokenLine: string;
let tokenVerified: boolean | 'unavailable' = 'unavailable';
+6 -25
View File
@@ -24,19 +24,9 @@
* (and GBRAIN_HOME when isolated) ride the registration itself.
*/
import { randomBytes } from 'node:crypto';
import {
chmodSync,
copyFileSync,
existsSync,
mkdirSync,
readFileSync,
realpathSync,
renameSync,
statSync,
writeFileSync,
} from 'node:fs';
import { dirname, isAbsolute, join } from 'node:path';
import { copyFileSync, existsSync, readFileSync } from 'node:fs';
import { isAbsolute, join } from 'node:path';
import { atomicWriteTextFile } from './atomic-write.ts';
import {
CLAUDE_COMMITTED_SETTINGS_FILE_RELPATH,
CLAUDE_HOOK_DEFAULT_TIMEOUT_SECS,
@@ -297,18 +287,9 @@ function stripOurEntries(groups: unknown[], marker: string = GBRAIN_HOOK_MARKER_
* but not for user-global config).
*/
function atomicWriteJson(path: string, value: unknown, freshMode?: number): void {
const target = existsSync(path) ? realpathSync(path) : path;
mkdirSync(dirname(target), { recursive: true });
let mode: number | undefined;
try {
mode = statSync(target).mode & 0o777;
} catch {
mode = freshMode; // fresh file: caller's convention (user-scope → 0600) [X11]
}
const tmp = `${target}.tmp-${randomBytes(6).toString('hex')}`;
writeFileSync(tmp, `${JSON.stringify(value, null, 2)}\n`, { encoding: 'utf8', ...(mode !== undefined ? { mode } : {}) });
if (mode !== undefined) chmodSync(tmp, mode); // writeFileSync mode applies only on create
renameSync(tmp, target);
// Shared bootstrap atomic writer (symlink-resolving, mode-inheriting) —
// fresh files take the caller's convention (user-scope → 0600) [X11].
atomicWriteTextFile(path, `${JSON.stringify(value, null, 2)}\n`, { freshMode });
}
/** Pre-write backup path per strategy; timestamped avoids the shared-slot loss. */
+111 -1
View File
@@ -19,8 +19,9 @@
* gbrain code.
*/
import { existsSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { basename, dirname, join } from 'node:path';
// ── Spec-target registry [ENG-7] ────────────────────────────────────────────
@@ -38,6 +39,7 @@ export interface HostSpecTarget {
export const CLAUDE_CODE_SPEC_ID = 'claude-code-2026-08';
export const CODEX_SPEC_ID = 'codex-2026-08';
export const OPENCODE_SPEC_ID = 'opencode-2026-08';
export const TARGETS: Record<string, HostSpecTarget> = {
[CLAUDE_CODE_SPEC_ID]: {
@@ -103,6 +105,43 @@ export const TARGETS: Record<string, HostSpecTarget> = {
'has no hooks". Some codex builds gate HTTP MCP servers behind ' +
'`experimental_use_rmcp_client = true` — probe at wiring time.',
},
[OPENCODE_SPEC_ID]: {
id: OPENCODE_SPEC_ID,
status: 'verified',
verifiedAt: '2026-08-15',
references: [
'docs/mcp/OPENCODE-CLI-PIN.md',
'https://opencode.ai/docs/mcp-servers/',
'opencode-ai 1.18.18 (hermetic observation run, macOS arm64, 2026-08-15)',
],
note:
'opencode (SST, opencode.ai — not OpenClaw). Config is JSONC everywhere: ' +
'comments parse in .json-named files, and global opencode.json AND ' +
'opencode.jsonc are BOTH read (merged) when both exist; `opencode mcp ' +
'add` writes the user-global opencode.jsonc via a comment-preserving ' +
'editor, so gbrain writes match that bar (jsonc-parser surgical edits, ' +
'opencode-json.ts). MCP entries: {type:"local", command[], environment, ' +
'enabled?} / {type:"remote", url, headers} — header values keep ' +
'`{env:VAR}` interpolation verbatim; unknown keys tolerated in 1.18.18 ' +
'but gbrain writes NO marker key (ownership is a structural ' +
'fingerprint — a future strict-schema flip must not brick the host). ' +
'`mcp add` has no scope flag (always user-global); project opencode.json ' +
'is read but a project-defined LOCAL server spawns with NO trust gate ' +
'(verified) — so gbrain defaults registration to USER scope and treats ' +
'project scope as explicit opt-in with a sharing warning. `mcp list` is ' +
'the honest discriminator (spawns servers; ✓/✗ text; exit 0 regardless); ' +
'`mcp debug` is OAuth-only. Keyless anonymous free tier answers headless ' +
'`run` AND drives MCP tool calls without --auto (load-bearing for the ' +
'door SMOKE). DOCS-CONTRADICTION pinned: OPENCODE_CONFIG / _CONFIG_DIR / ' +
'_CONFIG_CONTENT observed INERT in 1.18.18 — only HOME/XDG_CONFIG_HOME ' +
'move the config; path helpers resolve via XDG only. opencode sets ' +
'OPENCODE=1 (+OPENCODE_PID) in bash-tool children — detectHarness ' +
'probes OPENCODE. AGENTS.md loads natively; CLAUDE.md is NOT ' +
'double-loaded. opencode ships a JS plugin/event system — ' +
'OPENCODE_HAS_HOOKS=false means "gbrain does not wire it yet" (follow-up ' +
'filed), NOT "opencode has no hooks"; probes run with --pure + ' +
'OPENCODE_DISABLE_AUTOUPDATE=1 because mcp list autoloads plugins.',
},
};
// ── Claude Code shapes the writers consume ──────────────────────────────────
@@ -240,3 +279,74 @@ export const CODEX_HAS_HOOKS = false;
export const CODEX_TOML_BLOCK_BEGIN =
`# gbrain:${GBRAIN_HARNESS_MARKER_VALUE} begin - managed by \`gbrain bootstrap harness\`; do not edit inside`;
export const CODEX_TOML_BLOCK_END = `# gbrain:${GBRAIN_HARNESS_MARKER_VALUE} end`;
// ── opencode shapes ─────────────────────────────────────────────────────────
/**
* opencode config DIRECTORY (user-global). Resolution mirrors what the real
* binary was OBSERVED to do (OPENCODE-CLI-PIN.md §Path seams): XDG_CONFIG_HOME
* else $HOME/.config, then /opencode. The OPENCODE_CONFIG / OPENCODE_CONFIG_DIR
* / OPENCODE_CONFIG_CONTENT env vars are deliberately NOT honored here
* observed INERT in opencode 1.18.18 (probes registered through each were
* invisible to `mcp list` while the XDG-resolved config was still read), so
* honoring them would write registrations into a file opencode never reads (a
* silent no-op install). HOME is read from the env explicitly because Bun's
* homedir() reads the password database, not the HOME env var (the
* claudeUserSettingsPath lesson).
*/
export function opencodeConfigDir(): string {
const xdg = process.env.XDG_CONFIG_HOME?.trim();
if (xdg) return join(xdg, 'opencode');
const home = process.env.HOME?.trim();
return join(home || homedir(), '.config', 'opencode');
}
/**
* User-global opencode config FILE. Both `opencode.json` and `opencode.jsonc`
* are read (merged) by the host when both exist; gbrain edits the file that
* already carries content, preferring `.jsonc` (the name `opencode mcp add`
* itself writes) when both or neither exist one-owner-per-file keeps the
* merge unambiguous for `mcp.gbrain`.
*/
export function opencodeGlobalConfigPath(): string {
const dir = opencodeConfigDir();
const jsonc = join(dir, 'opencode.jsonc');
const json = join(dir, 'opencode.json');
if (existsSync(jsonc)) return jsonc;
if (existsSync(json)) return json;
return jsonc;
}
/**
* The OTHER member of the global filename pair for a given config path
* (`opencode.json` `opencode.jsonc` in the same dir), or null when the
* basename is not a pair member. opencode MERGES both files when both exist,
* so global WRITERS must reconcile `mcp.<name>` across the pair a same-name
* entry left in the sibling survives as a shadow registration whose merge
* winner is ambiguous (and a later removal of the primary "reveals" it).
* Callers apply this to the USER-GLOBAL pair only; project-scope sibling
* semantics are unobserved.
*/
export function opencodeGlobalSiblingPath(configPath: string): string | null {
const dir = dirname(configPath);
const base = basename(configPath);
if (base === 'opencode.json') return join(dir, 'opencode.jsonc');
if (base === 'opencode.jsonc') return join(dir, 'opencode.json');
return null;
}
/** Project-scope opencode config (docs-canonical name; opencode's lookup
* traverses up to the git root). Committed-file candidate the writer's
* PATH-resolved command + sharing-warning rules apply (OPENCODE.md). */
export function opencodeProjectConfigPath(workspaceDir: string): string {
return join(workspaceDir, 'opencode.json');
}
/**
* Whether gbrain WIRES opencode's hook/plugin system. False = not yet:
* opencode ships a JS plugin/event system (and `--pure` to suppress it), but
* gbrain's opencode plugin lane is a filed follow-up; per-turn context on
* opencode rides the pull-protocol AGENTS.md gates, which opencode loads
* natively (verified and CLAUDE.md is NOT double-loaded alongside it).
*/
export const OPENCODE_HAS_HOOKS = false;
+572
View File
@@ -0,0 +1,572 @@
/**
* opencode-json.ts managed `mcp.<name>` entry writer for opencode's JSONC
* configs (see TARGETS['opencode-2026-08'] in host-specs.ts and
* docs/mcp/OPENCODE-CLI-PIN.md for the verified format assumptions).
*
* Why a direct writer exists: `opencode mcp add` always targets the
* user-global opencode.jsonc (no scope flag), cannot set file modes (the
* harness lane's inline bearer needs 0600), and requires the binary on the
* box the writer covers project scope, secret hygiene, and offline/
* pre-install registration with one code path.
*
* Safety invariants (codex-toml.ts analog, adapted for JSONC):
* - ALL edits go through jsonc-parser `modify`/`applyEdits` text splicing
* that preserves comments, formatting, and EOLs byte-for-byte outside the
* edited range. opencode's own `mcp add` preserves comments (observed);
* gbrain matches that bar. JSON.parse is never used on config text.
* - Ownership is a STRUCTURAL FINGERPRINT, not a marker key (unknown keys
* are tolerated by opencode 1.18.18, but a future strict-schema flip must
* not brick the user's opencode): a local entry is ours when command[0] is
* gbrain-shaped AND environment.GBRAIN_SOURCE exists; source EQUALITY
* (not mere presence) splits `ours-same-source` from `ours-other-source`
* ([FIX7] parity with verifyMcpTargetsWorkspace) callers warn before
* overwriting another workspace's registration. A remote entry is ours
* when its url matches the caller's receipt, or when its Authorization
* header carries the `{env:GBRAIN_REMOTE_TOKEN}` interpolation (only the
* connect lane writes that). Anything else under our name is FOREIGN
* refuse, never guess.
* - Read-failure classes are distinct: ENOENT fresh file; empty/whitespace
* treated as `{}`; unreadable (EACCES etc.) refuse loudly (never
* clobber what cannot be read). A file that fails even JSONC parsing
* refuse with a paste-by-hand snippet.
* - Post-render validation before rename: the rendered text is re-parsed,
* our entry deep-asserted, and every OTHER top-level key asserted to
* survive; on any failure the original file is untouched.
* - Secrets hygiene: when the entry carries an inline bearer the target is
* forced 0600. Backups are UNIQUE per operation (`<config>.bak-<hex>`,
* returned in the result) so two overlapping runs can never clobber each
* other's snapshot, and a backup is chmod'd 0600 whenever the COPIED
* content carries an inline bearer (write AND remove paths on re-runs
* the backup carries the PREVIOUS token). Token-free entries inherit the
* file's existing mode.
* - Concurrency: callers hold acquireBootstrapLock (config-dir
* opencode-dir ordering, mirroring the codex lanes in harness.ts) the
* writer itself is lock-free like codex-toml.ts.
*/
import { randomBytes } from 'node:crypto';
import { chmodSync, copyFileSync, existsSync, readFileSync } from 'node:fs';
import { applyEdits, modify, parse as parseJsonc, printParseErrorCode, type ParseError } from 'jsonc-parser';
import { atomicWriteTextFile } from './atomic-write.ts';
import { opencodeGlobalSiblingPath } from './host-specs.ts';
export const GBRAIN_REMOTE_TOKEN_ENV = 'GBRAIN_REMOTE_TOKEN';
const ENV_INTERPOLATION = `{env:${GBRAIN_REMOTE_TOKEN_ENV}}`;
// ── Entry shapes ────────────────────────────────────────────────────────────
export interface OpencodeLocalEntry {
kind: 'local';
name: string;
/** argv command[0] is PATH-resolved "gbrain" (project scope, committed-
* file candidate) or an absolute binary path (user scope). */
command: string[];
environment: Record<string, string>;
}
export interface OpencodeRemoteEntry {
kind: 'remote';
name: string;
url: string;
/** 'inline' writes `Bearer <token>` (harness lane framework-spawned
* opencode inherits no shell profile; file forced 0600). 'env' writes the
* `{env:GBRAIN_REMOTE_TOKEN}` interpolation (connect lane token never
* enters the file). */
tokenMode: 'inline' | 'env';
bearerToken?: string;
}
export type OpencodeMcpEntry = OpencodeLocalEntry | OpencodeRemoteEntry;
export type OpencodeEntryKind =
| 'absent'
| 'ours-same-source'
| 'ours-other-source'
| 'foreign';
export interface OpencodeEntryExpectation {
/** GBRAIN_SOURCE the caller is registering (local entries). */
sourceId?: string;
/** Serve url from the caller's receipt (remote entries). */
url?: string;
}
export interface WriteOpencodeEntryResult {
configPath: string;
/** True when a prior gbrain-owned entry was replaced (idempotent re-run). */
replacedPrior: boolean;
/** Kind of the pre-existing entry (what was there before this write). */
priorKind: OpencodeEntryKind;
/** Unique per-write backup (`<config>.bak-<hex>`) of the prior file, or
* null on a fresh file. Callers that roll back restore from THIS path. */
backupPath: string | null;
/** The EXACT text this write landed rollback callers compare the current
* file content against it before restoring (a mismatch means a newer
* registration exists and a restore would clobber it). */
writtenText: string;
notes: string[];
}
export interface RemoveOpencodeEntryResult {
configPath: string;
removed: boolean;
backupPath: string | null;
notes: string[];
}
// ── Read + parse (failure classes are distinct) ─────────────────────────────
interface RawConfig {
text: string;
existed: boolean;
}
function readConfigRaw(configPath: string): RawConfig {
if (!existsSync(configPath)) return { text: '', existed: false };
let text: string;
try {
text = readFileSync(configPath, 'utf8');
} catch (e) {
throw new Error(
`${configPath} exists but cannot be read (${(e as Error).message}) — ` +
`refusing to touch a config that cannot be read back. Fix permissions and re-run.`,
);
}
return { text, existed: true };
}
/**
* Parse config text as JSONC (opencode's effective grammar for BOTH .json
* and .jsonc files OPENCODE-CLI-PIN.md §Config format). Empty/whitespace
* text parses as `{}`. Text that fails even JSONC parsing throws with a
* paste-by-hand snippet so the user is never stranded.
*/
export function parseOpencodeConfig(text: string, configPath: string, snippet?: string): Record<string, unknown> {
if (text.trim() === '') return {};
const errors: ParseError[] = [];
const parsed = parseJsonc(text, errors, { allowTrailingComma: true }) as unknown;
if (errors.length > 0) {
const first = errors[0];
throw new Error(
`${configPath} does not parse as JSONC (${printParseErrorCode(first.error)} at offset ${first.offset}) — ` +
`opencode itself cannot read it either. Fix the file, or add the entry by hand:\n${snippet ?? ''}`,
);
}
if (typeof parsed !== 'object' || parsed === null || Array.isArray(parsed)) {
throw new Error(`${configPath} is valid JSONC but not an object — fix the file and re-run.`);
}
return parsed as Record<string, unknown>;
}
// ── Ownership fingerprint ───────────────────────────────────────────────────
function isGbrainShapedCommand(command: unknown): boolean {
if (!Array.isArray(command) || command.length === 0) return false;
const head = command[0];
if (typeof head !== 'string') return false;
if (head === 'gbrain') return true; // PATH-resolved (project scope)
if (/[\\/]gbrain$/.test(head)) return true; // absolute binary path
// bun-run wrapper shim lane: `bun run <...>/gbrain/src/cli.ts` etc. The
// arg match is ANCHORED like the head-path lane: some arg must carry an
// exact `gbrain` path segment (or a hyphen-suffixed `gbrain-*` one) — a
// loose substring scan classified `bun run /opt/gbrainy-fork/src/cli.ts`
// as ours (both the `gbrain` substring and a bare `src/cli.ts$` matched).
// A gbrain-less `bun run /repo/src/cli.ts` is now NOT ours (fail-closed:
// gbrain refuses to touch what it cannot prove it owns).
if (head === 'bun' || head.endsWith('/bun')) {
return command.some((a) => typeof a === 'string' && /(?:^|[\\/])gbrain(?:[\\/-]|$)/.test(a));
}
// staged shim named gbrain-<suffix> (e.g. gbrain-shim from stageBinDir) —
// hyphen-anchored so a foreign /opt/bin/gbrainy is NOT ours.
return /[\\/]gbrain-[^\\/]*$/.test(head);
}
/**
* Classify the `mcp.<name>` entry in parsed config. The arbiter every lane
* consults before writing or removing (codexBlockOwnsName analog).
*/
export function opencodeEntryKind(
parsed: Record<string, unknown>,
name: string,
expect: OpencodeEntryExpectation = {},
): OpencodeEntryKind {
const mcp = parsed.mcp;
if (typeof mcp !== 'object' || mcp === null) return 'absent';
const entry = (mcp as Record<string, unknown>)[name];
if (entry === undefined) return 'absent';
if (typeof entry !== 'object' || entry === null) return 'foreign';
const e = entry as Record<string, unknown>;
if (e.type === 'local') {
if (!isGbrainShapedCommand(e.command)) return 'foreign';
const env = e.environment;
const src =
typeof env === 'object' && env !== null
? (env as Record<string, unknown>).GBRAIN_SOURCE
: undefined;
if (typeof src !== 'string' || src === '') return 'foreign';
// Kind mismatch (red-team): a caller expecting a REMOTE entry (harness /
// connect lanes pass expect.url) that finds a LOCAL gbrain entry is
// looking at ANOTHER lane's registration (the workspace stdio lane's) —
// never `ours-same-source`, or a silent replace (and a later --remove)
// would eat it. `ours-other-source` fires the refuse/confirm machinery.
if (expect.url !== undefined) return 'ours-other-source';
if (expect.sourceId === undefined) return 'ours-same-source';
return src === expect.sourceId ? 'ours-same-source' : 'ours-other-source';
}
if (e.type === 'remote') {
if (expect.url !== undefined && e.url === expect.url) return 'ours-same-source';
const headers = e.headers;
const auth =
typeof headers === 'object' && headers !== null
? (headers as Record<string, unknown>).Authorization
: undefined;
if (typeof auth === 'string' && auth.includes(ENV_INTERPOLATION)) {
// Only the gbrain connect lane writes the {env:GBRAIN_REMOTE_TOKEN}
// interpolation — unambiguously ours even without a receipt url. But a
// url mismatch (another serve) OR a LOCAL expectation (expect.sourceId
// — the workspace stdio lane; the kind-mismatch mirror of the local
// branch above) is another lane's wiring: ours-other-source.
return expect.url === undefined && expect.sourceId === undefined
? 'ours-same-source'
: 'ours-other-source';
}
return 'foreign';
}
return 'foreign';
}
// ── Rendering ───────────────────────────────────────────────────────────────
function entryValue(entry: OpencodeMcpEntry): Record<string, unknown> {
if (entry.kind === 'local') {
return {
type: 'local',
command: entry.command,
environment: entry.environment,
enabled: true,
};
}
const token =
entry.tokenMode === 'inline'
? `Bearer ${entry.bearerToken ?? ''}`
: `Bearer ${ENV_INTERPOLATION}`;
return {
type: 'remote',
url: entry.url,
headers: { Authorization: token },
enabled: true,
};
}
/** Copy-pasteable snippet for the refusal paths (the user is never stranded).
* SECURITY: an inline bearer is substituted with a literal placeholder the
* snippet rides thrown error messages (parse refusal, foreign refusal), and an
* error path must never embed the real secret in text that lands in logs,
* receipts, or stderr. Only the human-facing snippet changes; the write path
* still renders the real token. */
export function opencodeEntrySnippet(entry: OpencodeMcpEntry): string {
const safe: OpencodeMcpEntry =
entry.kind === 'remote' && entry.tokenMode === 'inline'
? { ...entry, bearerToken: '<paste-token-here>' }
: entry;
return JSON.stringify({ mcp: { [safe.name]: entryValue(safe) } }, null, 2);
}
function assertEntryName(name: string): void {
if (!/^[A-Za-z0-9_-]+$/.test(name)) {
throw new Error(
`MCP server name "${name}" is not a simple key ([A-Za-z0-9_-]+) — pick a simpler --name`,
);
}
}
function entryCarriesSecret(entry: OpencodeMcpEntry): boolean {
return entry.kind === 'remote' && entry.tokenMode === 'inline';
}
/** True when config text carries an INLINE bearer credential (any
* `Bearer <value>` that is not the `{env:…}` interpolation) the rule that
* decides whether a backup copy must be tightened to 0600. */
export function textCarriesInlineBearer(text: string): boolean {
return /Bearer\s+(?!\{env:)\S/.test(text);
}
/** Unique-suffix backup (`<config>.bak-<hex>`): two overlapping runs can
* never clobber each other's snapshot. The config-dir lock covers the WRITE,
* but a backup must survive until the caller's post-write verification (the
* harness network smoke) which runs AFTER the lock is released, so a fixed
* `.bak` name would let run B's writer overwrite run A's snapshot and a
* failed run A would then restore (and revoke against) run B's state.
* chmod 0600 whenever the copied content carries an inline bearer
* (copyFileSync onto a fresh path takes the source mode, but a hand-loosened
* source must not propagate a loose mode to a token-bearing backup). */
function createUniqueBackup(configPath: string, priorText: string): string {
const backupPath = `${configPath}.bak-${randomBytes(6).toString('hex')}`;
copyFileSync(configPath, backupPath);
if (textCarriesInlineBearer(priorText)) chmodSync(backupPath, 0o600);
return backupPath;
}
// No explicit eol: jsonc-parser detects and preserves the file's own EOLs
// (verified: a CRLF config keeps CRLF through modify/applyEdits).
const FORMATTING = { formattingOptions: { insertSpaces: true, tabSize: 2 } };
// ── Write ───────────────────────────────────────────────────────────────────
/**
* Idempotently write the managed `mcp.<name>` entry via a comment-preserving
* surgical edit. Refuses foreign entries; replaces ours-same-source silently;
* replaces ours-other-source only when `allowReplaceOtherSource` (callers
* warn first). Validates the render before the atomic swap.
*/
export function writeOpencodeMcpEntry(
configPath: string,
entry: OpencodeMcpEntry,
opts: { expect?: OpencodeEntryExpectation; allowReplaceOtherSource?: boolean } = {},
): WriteOpencodeEntryResult {
assertEntryName(entry.name);
if (entry.kind === 'remote' && entry.tokenMode === 'inline' && !entry.bearerToken) {
throw new Error('inline token mode requires a bearerToken');
}
const notes: string[] = [];
const snippet = opencodeEntrySnippet(entry);
const { text, existed } = readConfigRaw(configPath);
const parsed = parseOpencodeConfig(text, configPath, snippet);
const priorKind = opencodeEntryKind(parsed, entry.name, opts.expect);
if (priorKind === 'foreign') {
throw new Error(
`mcp.${entry.name} in ${configPath} is not a gbrain-managed entry — refusing to overwrite it. ` +
`Remove it (or pick another --name) and re-run.`,
);
}
if (priorKind === 'ours-other-source' && !opts.allowReplaceOtherSource) {
// Caller-appropriate refusal text: on the REMOTE path (expect.url — the
// harness/connect lanes) no GBRAIN_SOURCE is involved, and connect's
// documented escape hatch is --force; the GBRAIN_SOURCE wording belongs
// to the local/workspace lane only.
throw new Error(
opts.expect?.url !== undefined
? `mcp.${entry.name} in ${configPath} is a gbrain registration that does not match this endpoint ` +
`(${opts.expect.url}) — another install or lane owns it; pass --force to replace it, or pick another --name.`
: `mcp.${entry.name} in ${configPath} belongs to a DIFFERENT gbrain workspace ` +
`(GBRAIN_SOURCE mismatch) — re-run with the overwrite confirmation to reroute it, or pick another --name.`,
);
}
if (priorKind === 'ours-other-source') {
notes.push(
opts.expect?.url !== undefined
? `replaced a gbrain registration that did not match this endpoint (url/lane mismatch).`
: `replaced a gbrain registration that pointed at a different workspace (source mismatch).`,
);
}
const baseText = text.trim() === '' ? '{\n "$schema": "https://opencode.ai/config.json"\n}\n' : text;
const edits = modify(baseText, ['mcp', entry.name], entryValue(entry), FORMATTING);
const nextText = applyEdits(baseText, edits);
// Post-render validation: parse + deep-assert our entry + assert every
// OTHER top-level key survives. Any failure leaves the original untouched.
const rendered = parseOpencodeConfig(nextText, configPath, snippet);
const renderedMcp = rendered.mcp as Record<string, unknown> | undefined;
const ours = renderedMcp?.[entry.name];
if (JSON.stringify(ours) !== JSON.stringify(entryValue(entry))) {
throw new Error(
`post-render validation failed: mcp.${entry.name} did not round-trip — original file left untouched.`,
);
}
for (const key of Object.keys(parsed)) {
if (key === 'mcp') continue;
if (JSON.stringify(rendered[key]) !== JSON.stringify(parsed[key])) {
throw new Error(
`post-render validation failed: top-level key "${key}" changed — original file left untouched.`,
);
}
}
if (typeof parsed.mcp === 'object' && parsed.mcp !== null) {
for (const key of Object.keys(parsed.mcp as Record<string, unknown>)) {
if (key === entry.name) continue;
const before = (parsed.mcp as Record<string, unknown>)[key];
const after = renderedMcp?.[key];
if (JSON.stringify(after) !== JSON.stringify(before)) {
throw new Error(
`post-render validation failed: mcp.${key} (not ours) changed — original file left untouched.`,
);
}
}
}
const secret = entryCarriesSecret(entry);
let backupPath: string | null = null;
if (existed) {
backupPath = createUniqueBackup(configPath, text);
if (secret) chmodSync(backupPath, 0o600); // re-runs: the backup carries the previous token
}
atomicWriteTextFile(configPath, nextText, secret ? { forceMode: 0o600 } : { freshMode: 0o644 });
if (secret && existed) {
notes.push(`${configPath} tightened to 0600 — it now carries a bearer token.`);
}
return {
configPath,
replacedPrior: priorKind !== 'absent',
priorKind,
backupPath,
writtenText: nextText,
notes,
};
}
// ── Remove ──────────────────────────────────────────────────────────────────
/**
* Remove the managed entry (fingerprint-keyed; everything else survives
* byte-for-byte). Absent file / absent entry are calm no-ops. Foreign
* entries refuse removal never deletes what gbrain does not own.
* `skipOtherSource` turns an `ours-other-source` match into a calm skip-with-
* note instead of a removal (the uninstall sweep passes it: a gbrain entry
* from a DIFFERENT workspace is not this uninstall's to delete).
*/
export function removeOpencodeMcpEntry(
configPath: string,
name: string,
expect: OpencodeEntryExpectation = {},
opts: { skipOtherSource?: boolean } = {},
): RemoveOpencodeEntryResult {
assertEntryName(name);
const notes: string[] = [];
if (!existsSync(configPath)) {
return { configPath, removed: false, backupPath: null, notes: ['no opencode config — nothing to remove'] };
}
const { text } = readConfigRaw(configPath);
const parsed = parseOpencodeConfig(text, configPath);
const kind = opencodeEntryKind(parsed, name, expect);
if (kind === 'absent') {
return { configPath, removed: false, backupPath: null, notes: ['no gbrain-managed entry — nothing to remove'] };
}
if (kind === 'foreign') {
throw new Error(
`mcp.${name} in ${configPath} is not a gbrain-managed entry — refusing to remove it.`,
);
}
if (kind === 'ours-other-source' && opts.skipOtherSource) {
return {
configPath,
removed: false,
backupPath: null,
notes: [
`mcp.${name} in ${configPath} belongs to a DIFFERENT gbrain workspace (source mismatch) — left in place.`,
],
};
}
if (kind === 'ours-other-source') {
notes.push('removed a gbrain registration that pointed at a different workspace (source mismatch).');
}
const edits = modify(text, ['mcp', name], undefined, FORMATTING);
const nextText = applyEdits(text, edits);
parseOpencodeConfig(nextText, configPath); // never leave opencode unreadable
// Unique backup, 0600 when the copied content carries an inline bearer —
// the removed entry may BE the token-bearing one, and a fixed-name copy
// onto a pre-existing loose-mode backup would keep the loose mode.
const backupPath = createUniqueBackup(configPath, text);
atomicWriteTextFile(configPath, nextText);
return { configPath, removed: true, backupPath, notes };
}
// ── Sibling-global reconcile (the two-filename merge blind spot) ────────────
/**
* opencode merges the user-global `opencode.json` AND `opencode.jsonc` when
* both exist. Before writing `mcp.<name>` into one of them, reconcile the
* SIBLING file: an ours-classified entry there is removed (one owner per
* name left in place it survives as a shadow registration whose merge
* winner is ambiguous, and a later removal of the primary "reveals" it); a
* FOREIGN entry refuses loudly naming BOTH files (same refusal posture as
* the primary-file foreign case the merge winner is not ours to fight
* over). No-op when the path is not a global-pair member or the sibling is
* absent/entry-less. Callers hold the opencode config-dir bootstrap lock
* (both files share the dir one lock covers both) and call this ONLY for
* user-global writes (project-scope sibling semantics are unobserved).
*/
export function reconcileOpencodeSiblingGlobal(
configPath: string,
name: string,
expect: OpencodeEntryExpectation = {},
): { siblingPath: string | null; removed: boolean; notes: string[] } {
const siblingPath = opencodeGlobalSiblingPath(configPath);
if (!siblingPath || !existsSync(siblingPath)) return { siblingPath, removed: false, notes: [] };
const { text } = readConfigRaw(siblingPath);
const parsed = parseOpencodeConfig(text, siblingPath);
const kind = opencodeEntryKind(parsed, name, expect);
if (kind === 'absent') return { siblingPath, removed: false, notes: [] };
if (kind === 'foreign') {
throw new Error(
`mcp.${name} in ${siblingPath} is not a gbrain-managed entry — opencode merges ${siblingPath} AND ` +
`${configPath} when both exist, so writing mcp.${name} into ${configPath} would fight it with an ` +
`ambiguous merge winner. Remove it (or pick another --name) and re-run.`,
);
}
const r = removeOpencodeMcpEntry(siblingPath, name, expect);
const notes = [
`removed the gbrain mcp.${name} entry from ${siblingPath} — opencode merges both global filenames, and the ` +
`registration being written lands in ${configPath} (one owner per name).`,
...r.notes,
];
return { siblingPath, removed: r.removed, notes };
}
/**
* True when `mcp.<name>` exists as a REMOTE-type entry (regardless of
* ownership). The workspace stdio lane consults this before writing a local
* entry into the user-global config: a remote entry under our name is either
* the harness lane's (bootstrap harness) or foreign either way the stdio
* lane must not fight it (the codexBlockOwnsName analog, #4043 ownership
* rule). Best-effort: unreadable/unparseable configs return false (the write
* path re-checks with full refusal semantics).
*/
export function opencodeRemoteEntryExists(configPath: string, name: string): boolean {
try {
const { text, existed } = readConfigRaw(configPath);
if (!existed) return false;
const parsed = parseOpencodeConfig(text, configPath);
const mcp = parsed.mcp;
if (typeof mcp !== 'object' || mcp === null) return false;
const entry = (mcp as Record<string, unknown>)[name];
return typeof entry === 'object' && entry !== null && (entry as Record<string, unknown>).type === 'remote';
} catch {
return false;
}
}
// ── Status/recovery helpers ─────────────────────────────────────────────────
/**
* Recover the inline bearer from OUR remote entry (harness `--status` token
* liveness the receipt never stores the token). Returns null when the file
* or entry is absent, foreign, env-mode, or unreadable as JSONC.
*/
export function parseOpencodeEntryBearer(configPath: string, name: string, expectUrl?: string): string | null {
try {
const { text, existed } = readConfigRaw(configPath);
if (!existed) return null;
const parsed = parseOpencodeConfig(text, configPath);
const kind = opencodeEntryKind(parsed, name, { url: expectUrl });
if (kind !== 'ours-same-source') return null;
const entry = (parsed.mcp as Record<string, unknown>)[name] as Record<string, unknown>;
if (entry.type !== 'remote') return null;
const auth = (entry.headers as Record<string, unknown> | undefined)?.Authorization;
if (typeof auth !== 'string' || !auth.startsWith('Bearer ')) return null;
const token = auth.slice('Bearer '.length);
if (token.includes('{env:')) return null; // env-interpolated — no inline token to recover
return token || null;
} catch {
return null;
}
}
+5 -3
View File
@@ -209,7 +209,7 @@ export const PHASES: PhaseSpec[] = [
title: 'Identity interview (confirmed read-back)',
resume_hint:
'gbrain bootstrap interview --init, then --set each answer, then --confirm <hash>. ' +
'Claude Code only: also record the MCP scope consent (--set MCP_SCOPE <project|user>) BEFORE --confirm',
'Claude Code and opencode: also record the MCP scope consent (--set MCP_SCOPE <project|user>) BEFORE --confirm',
detect: (ws) => {
const exists = existsSync(interviewStatePath(ws));
const st = interviewStatus(ws);
@@ -265,8 +265,10 @@ export const PHASES: PhaseSpec[] = [
// outside the harness being wired). Advisory prose; the grep pins in
// scripts/check-bootstrap-templates.sh §(e) are the enforcement.
resume_hint:
'gbrain bootstrap hooks --harness <claude-code|codex> — MCP scope consent is ' +
'Claude Code only (recorded during the interview, pre-confirm); Codex registrations are always user-global (no scope flag)',
'gbrain bootstrap hooks --harness <claude-code|codex|opencode> — MCP scope consent applies on ' +
'Claude Code and opencode (recorded during the interview, pre-confirm; opencode defaults to ' +
'user-global — the sharing-safe choice, since it spawns project-config servers with no trust gate); ' +
'Codex registrations are always user-global (no scope flag)',
detect: (ws, ctx) => {
const regs = ctx.receipt?.registrations ?? [];
if (regs.length > 0) {
+1 -1
View File
@@ -98,7 +98,7 @@ export function resolveBrainDataDir(gbrainHomeDir: string): string {
// ---------------------------------------------------------------------------
export interface RegistrationRemovalRequest {
host: 'claude-code' | 'codex';
host: 'claude-code' | 'codex' | 'opencode';
scope: string;
detail?: string;
}
+7 -2
View File
@@ -570,7 +570,6 @@ async function runRoundtrip(
const putPage = findOp('put_page');
const getPage = findOp('get_page');
const queryOp = findOp('query');
const deletePage = findOp('delete_page');
await sweepProbeLeftovers(engine, ws, sourceId);
@@ -675,10 +674,16 @@ async function runRoundtrip(
}
// 6. Delete the probes [G13] — failure is a WARNING, never a verify fail.
// HARD delete via the engine primitive (same as sweepProbeLeftovers): the
// probe is not user content and verify is a trusted local caller. The
// delete_page OP is a v0.26.5 SOFT delete (sets deleted_at, row stays in
// pages until the 72h purge) — using it here left two probe tombstones in
// the user's brain after every verify run, visible to include_deleted
// readers and pinned as residue by the Postgres e2e cleanup assertion.
const deleteWarnings: string[] = [];
for (const slug of [VERIFY_PROBE_SLUG, VERIFY_PROBE_ENTITY_SLUG]) {
try {
await deletePage.handler(ctx, { slug });
await engine.deletePage(slug, { sourceId });
} catch (e) {
deleteWarnings.push(`${slug}: ${(e as Error).message}`);
}
+162
View File
@@ -0,0 +1,162 @@
/**
* opencode runner invokes the real `opencode` binary (SST terminal agent,
* opencode.ai) in a tempdir with a BRIEF.md prompt. Live mode only.
*
* Invocation pattern (verified against a pinned hermetic install, v1.18.18
* see docs/mcp/OPENCODE-CLI-PIN.md):
* opencode run "<brief>" --format default
*
* `run` is opencode's headless one-shot: prompt in, final answer text ALONE
* on stdout (banner/UI on stderr), exit 0. `--format default` is passed
* explicitly so an upstream default flip cannot silently change the
* transcript shape. NO `--auto`: MCP tool calls fire in run mode without any
* permission flag (verified the keyless SMOKE recalled a nonce through
* gbrain_recall with the flag absent). NO `-m`: live mode runs the
* OPERATOR's configured opencode, whose model pin (or the anonymous free
* tier) is the point of the measurement.
*
* Naming: opencode (SST, opencode.ai, npm `opencode-ai`) is not OpenClaw
* (the platform with its own runner) and not the original `opencode` CLI
* that became Crush the version preamble below makes a mis-bound claimant
* diagnosable (the SST CLI answers `--version` with a BARE semver).
*
* Hermeticity posture (deliberate): live mode runs the OPERATOR's configured
* opencode the real XDG config/data dirs are inherited unless
* XDG_CONFIG_HOME/XDG_DATA_HOME point elsewhere against a hermetic BRAIN.
* opencode-specific contamination channel (observed): the user-global
* opencode config is read for EVERY run, and a project opencode.json in the
* cwd spawns its local MCP servers with NO trust gate. Live-mode workspaces
* are harness-created tempdirs (no project config in reach), but a global
* mcp.gbrain entry would bind the operator's REAL brain while the oracle
* probes the hermetic one `invoke()` logs a loud warning for that case.
* The fully hermetic lane is the door e2e
* (install-real-opencode.serial.test.ts): fresh HOME + XDG dirs.
* OPENCODE_CONFIG_CONTENT (inline whole-config env) is deliberately NOT
* forwarded it is a config-shadowing channel, and it was observed inert in
* 1.18.18 anyway (OPENCODE-CLI-PIN.md §Path seams).
*
* Binary resolution: $OPENCODE_BIN > `which opencode` > unavailable.
*/
import { execFileSync } from 'child_process';
import { existsSync, readFileSync } from 'fs';
import { join } from 'path';
import {
BASE_ENV_ALLOWLIST,
detectBinary,
filterAllowlistEnv,
type AgentRunner,
type DetectResult,
type InvokeOpts,
type InvokeResult,
} from '../agent-runner.ts';
import { spawnWithCapture } from '../transcript-capture.ts';
/**
* Allow-list for env propagation when spawning opencode. The base list
* carries ONLY Anthropic + OpenAI provider keys opencode's headline
* feature is multi-provider, so the delta NAMES the additional provider
* keys a live-lane operator may be running on (xAI, Google, OpenRouter);
* anything not named here silently strips and reads as a misleading
* agent-auth failure. Plus the opencode seams: XDG dirs (config/auth
* relocation for hermetic callers), OPENCODE_CONFIG(_DIR) (observed inert
* in 1.18.18 but forwarded so a future release that activates them behaves
* the way the caller intended), and OPENCODE_DISABLE_AUTOUPDATE (the env
* half of the double autoupdate kill).
*/
const ENV_ALLOWLIST = [
...BASE_ENV_ALLOWLIST,
'XAI_API_KEY',
'GOOGLE_GENERATIVE_AI_API_KEY',
'GEMINI_API_KEY',
'OPENROUTER_API_KEY',
'XDG_CONFIG_HOME',
'XDG_DATA_HOME',
'OPENCODE_CONFIG',
'OPENCODE_CONFIG_DIR',
'OPENCODE_DISABLE_AUTOUPDATE',
];
export class OpencodeRunner implements AgentRunner {
readonly name = 'opencode';
async detect(): Promise<DetectResult> {
return detectBinary('OPENCODE_BIN', 'opencode');
}
async invoke(opts: InvokeOpts): Promise<InvokeResult> {
const detected = await this.detect();
if (!detected.available || !detected.binPath) {
throw new Error(`opencode runner unavailable: ${detected.reason ?? 'unknown'}`);
}
const args = ['run', opts.brief, '--format', 'default'];
const env = filterAllowlistEnv(ENV_ALLOWLIST, opts.env);
this.warnOnGlobalGbrainEntry(env);
// Version preamble: recorded as a plain stdout transcript event. The SST
// CLI answers with a BARE semver (`1.18.18` — no name, no build hash);
// any other shape means a colliding `opencode` claimant is bound.
// execFileSync (no shell) with the SAME filtered env as the turn itself.
try {
const version = execFileSync(detected.binPath, ['--version'], {
encoding: 'utf-8',
stdio: ['ignore', 'pipe', 'ignore'],
timeout: 15_000,
env: env as NodeJS.ProcessEnv,
cwd: opts.cwd,
}).trim();
opts.transcriptSink.write({
ts: Date.now(),
channel: 'stdout',
bytes: Buffer.from(`[opencode-runner preamble] version: ${version}\n`, 'utf-8'),
});
} catch {
// Preamble is diagnostic only — never fail the run for it.
}
const result = await spawnWithCapture(detected.binPath, args, {
cwd: opts.cwd,
env,
timeoutMs: opts.timeoutMs,
transcriptSink: opts.transcriptSink,
});
return { exitCode: result.exitCode, durationMs: result.durationMs };
}
/**
* Loud tripwire for the global-config contamination channel: when the
* config dir opencode will resolve carries an mcp.gbrain entry, a live
* turn routes gbrain tool calls at the operator's REAL brain while the
* oracle probes the hermetic one. Warning only (live mode deliberately
* runs the operator's agent); the hermetic lane is the door e2e. Checks
* BOTH filenames opencode merges opencode.json AND opencode.jsonc.
*/
private warnOnGlobalGbrainEntry(env: Record<string, string>): void {
try {
const xdg = env.XDG_CONFIG_HOME ?? process.env.XDG_CONFIG_HOME;
const home = env.HOME ?? process.env.HOME;
const cfgDir = xdg ? join(xdg, 'opencode') : home ? join(home, '.config', 'opencode') : null;
if (!cfgDir) return;
for (const file of ['opencode.jsonc', 'opencode.json']) {
const p = join(cfgDir, file);
if (!existsSync(p)) continue;
// Loose containment probe, not a parse: the global config is JSONC
// (comments legal), and a substring hit is enough for a warning.
const text = readFileSync(p, 'utf-8');
if (/"gbrain"\s*:/.test(text) && /"mcp"\s*:/.test(text)) {
console.warn(
`[opencode-runner] WARNING: ${p} carries an mcp.gbrain entry. opencode reads the ` +
"user-global config on every run, so this live turn may bind the OPERATOR'S REAL " +
'brain instead of the hermetic one. Use a scratch HOME/XDG_CONFIG_HOME for clean ' +
'measurements (docs/mcp/OPENCODE-CLI-PIN.md).',
);
return;
}
}
} catch {
// Best-effort tripwire — unreadable/invalid config is not an error here.
}
}
}
+6 -6
View File
@@ -23,7 +23,7 @@ export const CLI_FLAG_REGISTRY: Record<string, readonly string[]> = {
'backfill': ['--aliases', '--all', '--batch-size', '--brain', '--concurrency', '--dry-run', '--fresh', '--help', '--include-null-signature', '--json', '--keep-index', '--list', '--max-errors', '--max-rows', '--no-extract', '--pattern', '--pending', '--reset', '--resolve', '--resume', '--source', '--stale', '--supersessions', '--thin'],
'bench': ['--baseline', '--brain', '--explain', '--force', '--from', '--help', '--json', '--label', '--lang', '--limit', '--markdown', '--multimodal', '--near-symbol', '--restore-only', '--source', '--stale', '--symbol-kind', '--thin', '--threshold-jaccard', '--threshold-latency-multiplier', '--threshold-top1', '--to', '--tool'],
'book-mirror': ['--aliases', '--all', '--allow-empty', '--apply', '--asof', '--author', '--auto', '--background', '--bound-max-concurrent', '--bound-slug-prefixes', '--bound-source', '--bound-tools', '--brain', '--brain-wide-max-cost-usd', '--budget-usd-per-day', '--by-mention', '--chapters-dir', '--content', '--context-file', '--date', '--days', '--dry-run', '--entities', '--explain', '--fast', '--federated', '--file', '--follow', '--force', '--from-pages', '--help', '--http', '--image', '--include-null-signature', '--json', '--kind', '--limit', '--max-turns', '--max-usd', '--mode', '--model', '--multimodal', '--no-confirm', '--no-embedding', '--no-extract', '--no-follow', '--offset', '--path', '--pattern', '--pending', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--reset', '--resolve', '--save', '--session', '--session-id', '--since', '--slug', '--slugs', '--source', '--stale', '--stats', '--supersessions', '--surface', '--thin', '--timeout', '--timeout-ms', '--title', '--token-ttl', '--trusted-extraction', '--url', '--with-db', '--yes'],
'bootstrap': ['--abbrev-ref', '--abort', '--accept-visibility-change-consequences', '--active', '--all', '--allow-unverified-remote', '--brain', '--branch', '--cached', '--compile', '--confirm', '--count', '--delete-brain', '--diff-filter', '--env', '--error-unmatch', '--exclude-standard', '--fast', '--file', '--flag', '--force', '--from-pages', '--full', '--gbrain-bin', '--get', '--git-dir', '--git-path', '--harness', '--heads', '--help', '--home', '--hostname', '--http', '--id', '--init', '--install', '--is-inside-work-tree', '--isolated', '--jq', '--json', '--local', '--minimal', '--name', '--name-only', '--no-capture', '--no-cron', '--no-embedding', '--no-hooks', '--no-verify', '--once', '--only', '--others', '--pat-file', '--path', '--pglite', '--porcelain', '--port', '--private', '--project', '--push', '--push-only', '--quiet', '--rebase', '--remove', '--repair', '--scope', '--scopes', '--set', '--short', '--show', '--show-toplevel', '--skip', '--source', '--status', '--surface', '--token', '--token-name', '--token-ttl', '--unset-all', '--url', '--user-hooks', '--verify', '--version', '--visibility', '--workspace', '--yes'],
'bootstrap': ['--abbrev-ref', '--abort', '--accept-visibility-change-consequences', '--active', '--all', '--allow-unverified-remote', '--auto', '--brain', '--branch', '--cached', '--compile', '--confirm', '--count', '--delete-brain', '--diff-filter', '--env', '--error-unmatch', '--exclude-standard', '--fast', '--file', '--flag', '--force', '--from-pages', '--full', '--gbrain-bin', '--get', '--git-dir', '--git-path', '--harness', '--heads', '--help', '--home', '--hostname', '--http', '--id', '--init', '--install', '--is-inside-work-tree', '--isolated', '--jq', '--json', '--local', '--minimal', '--name', '--name-only', '--no-capture', '--no-cron', '--no-embedding', '--no-hooks', '--no-verify', '--once', '--only', '--others', '--pat-file', '--path', '--pglite', '--porcelain', '--port', '--private', '--project', '--pure', '--push', '--push-only', '--quiet', '--rebase', '--remove', '--repair', '--scope', '--scopes', '--set', '--short', '--show', '--show-toplevel', '--skip', '--source', '--status', '--surface', '--token', '--token-name', '--token-ttl', '--unset-all', '--url', '--user-hooks', '--verify', '--version', '--visibility', '--workspace', '--yes'],
'brainstorm': ['--aliases', '--all', '--brain', '--chunker-debug', '--code', '--compile', '--fast', '--file', '--fix', '--force', '--force-rechunk', '--force-resume', '--from-pages', '--full', '--help', '--http', '--include-null-signature', '--json', '--judge-model', '--lang', '--limit', '--list-runs', '--markdown', '--max-cost', '--max-far-set', '--max-ideas-per-judge-call', '--model', '--no-embed', '--no-embedding', '--no-extract', '--no-save', '--pattern', '--pending', '--reset', '--resolve', '--resume', '--retry-failed', '--retry-judge', '--save', '--source', '--stale', '--strict-budget', '--supersessions', '--surface', '--thin', '--timeout', '--token-ttl', '--yes'],
'cache': ['--brain', '--fast', '--force', '--from-pages', '--help', '--http', '--json', '--no-embedding', '--source', '--surface', '--token-ttl', '--yes'],
'calibration': ['--ab', '--aliases', '--all', '--allow-empty', '--apply', '--asof', '--auto', '--bound-max-concurrent', '--bound-slug-prefixes', '--bound-source', '--bound-tools', '--brain', '--budget-usd-per-day', '--by-mention', '--content', '--date', '--days', '--dry-run', '--entities', '--explain', '--fast', '--federated', '--file', '--follow', '--force', '--from-pages', '--help', '--holder', '--http', '--image', '--include-null-signature', '--json', '--key-prefix', '--kind', '--lang', '--limit', '--markdown', '--max-usd', '--mode', '--multimodal', '--near-symbol', '--no-embedding', '--no-extract', '--no-federated', '--offset', '--path', '--pattern', '--pending', '--phase', '--progress-interval', '--progress-json', '--quiet', '--regenerate', '--repo', '--reset', '--resolve', '--restore-only', '--save', '--scrub-gstack', '--session', '--session-id', '--since', '--slug', '--slugs', '--source', '--stale', '--stats', '--supersessions', '--surface', '--symbol-kind', '--thin', '--token-ttl', '--trusted-extraction', '--undo-wave', '--url', '--with-calibration', '--with-db', '--yes'],
@@ -32,20 +32,20 @@ export const CLI_FLAG_REGISTRY: Record<string, readonly string[]> = {
'check-backlinks': ['--background', '--brain', '--brain-wide-max-cost-usd', '--dir', '--dry-run', '--explain', '--follow', '--help', '--include-frontmatter', '--json', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--source', '--stale', '--timeout', '--type'],
'check-resolvable': ['--brain', '--dry-run', '--fix', '--help', '--json', '--skills-dir', '--source', '--strict', '--verbose'],
'check-update': ['--all', '--brain', '--check', '--dim', '--ff-only', '--help', '--json', '--markdown', '--migrate-only', '--non-interactive', '--refresh-cache', '--source', '--swap-only', '--to', '--version', '--yes'],
'claw-test': ['--ab', '--agent', '--all', '--auto-update', '--brain', '--break-lock', '--build-index', '--by-mention', '--compile', '--days', '--dir', '--exclusive', '--force', '--force-retry', '--force-schema', '--from-meetings', '--help', '--history', '--http', '--json', '--keep-tempdir', '--lang', '--list-agents', '--live', '--local', '--locks', '--markdown', '--max-age', '--message', '--multimodal', '--no-embed', '--no-embedding', '--no-extract', '--output-format', '--path', '--pglite', '--phase', '--priority', '--progress-json', '--refresh-unqualified', '--remediate', '--rollback', '--run-id', '--scenario', '--skip-verify', '--source', '--stale', '--surface', '--transcripts', '--undo-wave', '--use-captured-snapshot', '--version', '--with-calibration', '--yes'],
'claw-test': ['--ab', '--agent', '--all', '--auto', '--auto-update', '--brain', '--break-lock', '--build-index', '--by-mention', '--compile', '--days', '--dir', '--exclusive', '--force', '--force-retry', '--force-schema', '--format', '--from-meetings', '--help', '--history', '--http', '--json', '--keep-tempdir', '--lang', '--list-agents', '--live', '--local', '--locks', '--markdown', '--max-age', '--message', '--multimodal', '--no-embed', '--no-embedding', '--no-extract', '--output-format', '--path', '--pglite', '--phase', '--priority', '--progress-json', '--refresh-unqualified', '--remediate', '--rollback', '--run-id', '--scenario', '--skip-verify', '--source', '--stale', '--surface', '--transcripts', '--undo-wave', '--use-captured-snapshot', '--version', '--with-calibration', '--yes'],
'code-callees': ['--aliases', '--all', '--all-sources', '--brain', '--chunker-debug', '--clone-dir', '--confirm-destructive', '--federated', '--force', '--help', '--include-null-signature', '--json', '--limit', '--no-extract', '--no-federated', '--no-json', '--path', '--pattern', '--pending', '--repo', '--reset', '--resolve', '--restore-only', '--source', '--stale', '--supersessions', '--thin', '--url', '--url-managed', '--yes'],
'code-callers': ['--aliases', '--all', '--all-sources', '--brain', '--chunker-debug', '--clone-dir', '--confirm-destructive', '--federated', '--force', '--help', '--include-null-signature', '--json', '--limit', '--no-extract', '--no-federated', '--no-json', '--path', '--pattern', '--pending', '--repo', '--reset', '--resolve', '--restore-only', '--source', '--stale', '--supersessions', '--thin', '--url', '--url-managed', '--yes'],
'code-def': ['--aliases', '--all', '--brain', '--chunker-debug', '--help', '--include-null-signature', '--json', '--lang', '--limit', '--no-extract', '--no-json', '--pattern', '--pending', '--pretty', '--reset', '--resolve', '--source', '--stale', '--supersessions', '--thin', '--yes'],
'code-refs': ['--aliases', '--all', '--brain', '--chunker-debug', '--help', '--include-null-signature', '--json', '--lang', '--limit', '--no-extract', '--no-json', '--pattern', '--pending', '--reset', '--resolve', '--source', '--stale', '--supersessions', '--thin', '--yes'],
'config': ['--aliases', '--all', '--brain', '--column', '--coverage-override', '--detail', '--embedding-dimensions', '--embedding-model', '--fast', '--federated-read', '--follow', '--force', '--from-pages', '--help', '--http', '--include-null-signature', '--json', '--markdown', '--model', '--multimodal', '--no-embedding', '--no-extract', '--no-federated', '--pattern', '--pending', '--pglite', '--reset', '--resolve', '--source', '--stale', '--supersessions', '--surface', '--thin', '--token-ttl', '--yes'],
'connect': ['--agent', '--bearer-token-env-var', '--bind', '--brain', '--client-id', '--client-secret', '--force', '--grant-types', '--help', '--http', '--install', '--json', '--name', '--oauth', '--public-url', '--register', '--scope', '--scopes', '--show-token', '--source', '--timeout-ms', '--token', '--token-endpoint-auth-method', '--url', '--version', '--yes'],
'connect': ['--agent', '--auto', '--bearer-token-env-var', '--bind', '--brain', '--client-id', '--client-secret', '--delete-brain', '--env', '--force', '--grant-types', '--header', '--help', '--http', '--install', '--json', '--name', '--oauth', '--public-url', '--pure', '--register', '--remove', '--scope', '--scopes', '--show-token', '--source', '--status', '--timeout-ms', '--token', '--token-endpoint-auth-method', '--url', '--version', '--yes'],
'conversation-parser': ['--aliases', '--all', '--brain', '--help', '--include-null-signature', '--json', '--no-extract', '--pattern', '--pending', '--reset', '--resolve', '--source', '--stale', '--supersessions', '--thin'],
'doctor': ['--ab', '--abbrev-ref', '--abi', '--abort', '--aliases', '--all', '--allow-shell-jobs', '--allow-unverified-remote', '--auto', '--auto-fix', '--auto-update', '--background', '--batch', '--brain', '--brain-wide-max-cost-usd', '--branch', '--break-lock', '--build-index', '--by-mention', '--by-type', '--cached', '--check', '--column', '--compile', '--concurrency', '--confidence', '--confirm', '--content-audit', '--count', '--days', '--delete-brain', '--detach', '--detail', '--diff-filter', '--dim', '--dir', '--drain', '--dry-run', '--embedding-dimensions', '--embedding-model', '--exclude-standard', '--exclusive', '--explain', '--fast', '--file', '--fix', '--follow', '--force', '--force-break-lock', '--force-retry', '--force-schema', '--format', '--fresh', '--from-meetings', '--from-pages', '--full', '--get', '--git-dir', '--git-path', '--grant-types', '--harness', '--health-interval', '--help', '--history', '--home', '--http', '--include-flagged', '--include-frontmatter', '--include-null-signature', '--include-pseudo', '--index-audit', '--init', '--input', '--is-inside-work-tree', '--job-isolation', '--jq', '--json', '--lang', '--limit', '--local', '--locks', '--markdown', '--max-age', '--max-cost', '--max-cost-usd', '--max-crashes', '--max-jobs', '--max-rss', '--max-usd', '--mcp-only', '--migrate-only', '--model', '--multimodal', '--name-only', '--name-status', '--near-symbol', '--nice', '--no', '--no-cron', '--no-embed', '--no-embedding', '--no-extract', '--no-federated', '--no-mutate', '--no-verify', '--oauth-client-secret', '--older-than', '--once', '--others', '--overwrite', '--parallel', '--params', '--pat-file', '--path', '--pattern', '--pending', '--pglite', '--phase', '--pid-file', '--porcelain', '--priority', '--probe-pglite', '--progress-interval', '--progress-json', '--project', '--push-only', '--query', '--queue', '--quiet', '--rebase', '--rebuild-rollup', '--refresh', '--refresh-unqualified', '--regenerate', '--remediate', '--remediation-plan', '--remove', '--repo', '--reset', '--resolve', '--restore-only', '--resume', '--review-lower', '--rollback', '--scope', '--scopes', '--set', '--short', '--show-current', '--show-toplevel', '--since', '--skills-dir', '--skip-bare-tweet', '--skip-failed', '--skip-urls', '--skip-verify', '--slugs', '--source', '--source-id', '--stale', '--stats', '--status', '--strategy', '--strict', '--supabase', '--supersessions', '--surface', '--symbol-kind', '--target', '--target-score', '--thin', '--timeout', '--to', '--token', '--token-ttl', '--top-k', '--type', '--undo-wave', '--unsafe-bypass-dream-guard', '--unset-all', '--untracked-files', '--url', '--use-captured-snapshot', '--verbose', '--verify', '--version', '--window', '--with-calibration', '--workers', '--yes'],
'dream': ['--against', '--aliases', '--all', '--allow-regression', '--anchor', '--asof', '--audit-rejects', '--background', '--batch', '--brain', '--brain-wide-max-cost-usd', '--break-lock', '--budget-usd', '--budget-usd-answer', '--budget-usd-retrieval', '--by-type', '--by-type-floor', '--cancel-unmatched', '--code', '--committed-baseline', '--compare', '--compile', '--concurrent', '--ctx-size', '--cycles', '--date', '--detail', '--dimensions', '--dir', '--drain', '--dry-run', '--embedding-dimensions', '--embedding-model', '--embeddings', '--expansion', '--explain', '--fast', '--federated', '--fix', '--fixtures', '--follow', '--force', '--force-break-lock', '--force-rechunk', '--force-retry', '--format', '--from', '--from-db', '--from-pages', '--gold', '--harness', '--help', '--http', '--include-holdout', '--include-null-signature', '--input', '--install', '--json', '--judge-model', '--justification', '--keyword-only', '--lang', '--limit', '--llm', '--markdown', '--max-age', '--max-cost', '--max-cost-usd', '--max-runtime', '--max-tokens', '--max-usd', '--mcp-only', '--min-recall', '--mode', '--model', '--models', '--modes', '--multimodal', '--name', '--name-only', '--near-symbol', '--no', '--no-embed', '--no-embedding', '--no-extract', '--no-federated', '--no-llm', '--no-mutate', '--no-trajectory', '--once', '--out', '--output', '--output-dir', '--parallel', '--path', '--pattern', '--pending', '--pglite', '--phase', '--priority', '--progress-interval', '--progress-json', '--pull', '--quiet', '--receipt-dir', '--reconcile-queue', '--remediate', '--repo', '--reranking', '--reset', '--resolve', '--restore-only', '--resume-from', '--retrieval-only', '--rounds', '--rubric-version', '--save', '--seed', '--short', '--show-toplevel', '--since', '--skip-replay', '--slot-a-model', '--slot-b-model', '--slot-c-model', '--slug', '--slug-prefix', '--source', '--source-id', '--stale', '--suite', '--suites', '--supabase', '--supersessions', '--surface', '--symbol-kind', '--take', '--task', '--thin', '--threshold', '--timeout', '--to', '--token-ttl', '--top-k', '--undo', '--unsafe-bypass-dream-guard', '--update-baseline', '--verify', '--version', '--window', '--yes'],
'edges-backfill': ['--aliases', '--all', '--all-sources', '--brain', '--concurrency', '--federated', '--help', '--include-null-signature', '--json', '--max-age', '--max-chunks', '--max-cost-usd', '--no-extract', '--no-federated', '--older-than', '--path', '--pattern', '--pending', '--repo', '--reset', '--resolve', '--restore-only', '--source', '--stale', '--supersessions', '--thin', '--timeout', '--workers'],
'embed': ['--aliases', '--all', '--background', '--batch-size', '--brain', '--brain-wide-max-cost-usd', '--break-lock', '--catch-up', '--dry-run', '--embedding-dimensions', '--embedding-model', '--explain', '--fast', '--fix', '--follow', '--force', '--force-break-lock', '--from-pages', '--help', '--http', '--include-null-signature', '--json', '--lang', '--markdown', '--max-age', '--max-cost-usd', '--model', '--multimodal', '--name', '--near-symbol', '--no', '--no-embed', '--no-embedding', '--no-extract', '--pace', '--pace-max-concurrency', '--parallel', '--path', '--pattern', '--pending', '--pglite', '--prefix', '--priority', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--reset', '--resolve', '--restore-only', '--serial', '--slugs', '--source', '--stale', '--supabase', '--supersessions', '--surface', '--symbol-kind', '--thin', '--timeout', '--to', '--token-ttl', '--version'],
'enrich': ['--aliases', '--all', '--all-sources', '--allow-empty', '--apply', '--asof', '--auto', '--background', '--bound-max-concurrent', '--bound-slug-prefixes', '--bound-source', '--bound-tools', '--brain', '--brain-wide-max-cost-usd', '--break-lock', '--budget-usd-per-day', '--by-mention', '--clone-dir', '--code', '--concurrency', '--confirm-destructive', '--content', '--date', '--days', '--detail', '--dry-run', '--embedding-dimensions', '--embedding-model', '--entities', '--explain', '--fast', '--federated', '--file', '--fix', '--follow', '--force', '--force-break-lock', '--from-pages', '--help', '--http', '--image', '--include-null-signature', '--json', '--judge-model', '--kind', '--lang', '--limit', '--markdown', '--max-age', '--max-cost', '--max-cost-usd', '--max-runtime', '--max-usd', '--min-context', '--mode', '--model', '--multimodal', '--near-symbol', '--no', '--no-embed', '--no-embedding', '--no-extract', '--offset', '--older-than', '--order', '--path', '--pattern', '--pending', '--progress-interval', '--progress-json', '--quiet', '--reenrich-after', '--remediate', '--reset', '--resolve', '--restore-only', '--resume', '--save', '--session', '--session-id', '--since', '--slug', '--slugs', '--source', '--source-id', '--stale', '--stats', '--supersessions', '--surface', '--symbol-kind', '--thin', '--thin-threshold', '--timeout', '--to', '--token-ttl', '--trusted-extraction', '--types', '--url', '--url-managed', '--version', '--with-db', '--workers', '--yes'],
'eval': ['--ab-relational', '--against', '--aliases', '--all', '--allow-regression', '--background', '--baseline', '--batch', '--brain', '--brain-wide-max-cost-usd', '--budget-usd', '--budget-usd-answer', '--budget-usd-retrieval', '--committed-baseline', '--compare', '--compare-limit', '--concurrent', '--config-a', '--config-b', '--corpus', '--cycles', '--days', '--dedup-cosine', '--dedup-max-per-page', '--dedup-type-ratio', '--dimensions', '--distance-min', '--embedding-dimensions', '--embedding-model', '--expand', '--explain', '--fast', '--fixtures', '--follow', '--force', '--from-capture', '--from-db', '--from-pages', '--gold', '--grounding-min', '--harness', '--help', '--http', '--include-holdout', '--include-null-signature', '--input', '--json', '--judge', '--justification', '--k', '--limit', '--llm', '--max-pair-chars', '--max-tokens', '--max-usd', '--md', '--metric', '--min-recall', '--mode', '--model', '--models', '--modes', '--multimodal', '--name', '--no', '--no-cache', '--no-embed', '--no-embedding', '--no-expand', '--no-extract', '--no-llm', '--older-than', '--out', '--output', '--output-dir', '--parallel', '--pattern', '--pending', '--progress-interval', '--progress-json', '--qrels', '--queries-file', '--query', '--questions', '--quiet', '--receipt-dir', '--refresh-cache', '--remediate', '--reset', '--resolve', '--rrf-k', '--rubric-version', '--runs', '--sampling', '--save', '--seed', '--severity', '--short', '--show-toplevel', '--since', '--skip-replay', '--slot-a-model', '--slot-b-model', '--slot-c-model', '--slug', '--slug-prefix', '--source', '--stale', '--strategy', '--strict', '--suite', '--suites', '--supersessions', '--surface', '--task', '--thin', '--threshold', '--threshold-expected-top1', '--threshold-first-relevant-hit', '--threshold-jaccard', '--threshold-latency-multiplier', '--threshold-recall-at-k', '--threshold-top1', '--timeout', '--to', '--token-ttl', '--tool', '--top-k', '--top-regressions', '--until', '--update-baseline', '--usefulness-min', '--verbose', '--version', '--with-code-intel', '--yes'],
'eval': ['--ab-relational', '--against', '--aliases', '--all', '--allow-regression', '--background', '--baseline', '--batch', '--brain', '--brain-wide-max-cost-usd', '--budget-usd', '--budget-usd-answer', '--budget-usd-retrieval', '--committed-baseline', '--compare', '--compare-limit', '--concurrent', '--config-a', '--config-b', '--corpus', '--cycles', '--days', '--dedup-cosine', '--dedup-max-per-page', '--dedup-type-ratio', '--dimensions', '--distance-min', '--embedder', '--embedding-dimensions', '--embedding-model', '--expand', '--explain', '--fast', '--fixtures', '--follow', '--force', '--from-capture', '--from-db', '--from-pages', '--gold', '--grounding-min', '--harness', '--help', '--http', '--include-holdout', '--include-null-signature', '--input', '--json', '--judge', '--justification', '--k', '--limit', '--llm', '--max-pair-chars', '--max-tokens', '--max-usd', '--md', '--metric', '--min-recall', '--mode', '--model', '--models', '--modes', '--multimodal', '--name', '--no', '--no-cache', '--no-embed', '--no-embedding', '--no-expand', '--no-extract', '--no-llm', '--older-than', '--out', '--output', '--output-dir', '--parallel', '--pattern', '--pending', '--progress-interval', '--progress-json', '--qrels', '--queries-file', '--query', '--questions', '--quiet', '--receipt-dir', '--refresh-cache', '--remediate', '--reset', '--resolve', '--rrf-k', '--rubric-version', '--runs', '--sampling', '--save', '--seed', '--severity', '--short', '--show-toplevel', '--since', '--skip-replay', '--slot-a-model', '--slot-b-model', '--slot-c-model', '--slug', '--slug-prefix', '--source', '--stale', '--strategy', '--strict', '--suite', '--suites', '--supersessions', '--surface', '--task', '--thin', '--threshold', '--threshold-expected-top1', '--threshold-first-relevant-hit', '--threshold-jaccard', '--threshold-latency-multiplier', '--threshold-recall-at-k', '--threshold-top1', '--timeout', '--to', '--token-ttl', '--tool', '--top-k', '--top-regressions', '--until', '--update-baseline', '--usefulness-min', '--verbose', '--version', '--with-code-intel', '--yes'],
'export': ['--aliases', '--all', '--background', '--brain', '--brain-wide-max-cost-usd', '--dir', '--explain', '--federated', '--fix', '--follow', '--help', '--include-null-signature', '--json', '--lang', '--markdown', '--multimodal', '--near-symbol', '--no-extract', '--no-federated', '--path', '--pattern', '--pending', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--repo', '--reset', '--resolve', '--restore-only', '--slug-prefix', '--source', '--stale', '--supersessions', '--symbol-kind', '--thin', '--timeout', '--type'],
'extract': ['--aliases', '--all', '--background', '--brain', '--brain-wide-max-cost-usd', '--by-mention', '--catch-up', '--code', '--concurrency', '--dir', '--dry-run', '--explain', '--federated', '--follow', '--from-meetings', '--help', '--include-frontmatter', '--include-null-signature', '--infer-dates', '--json', '--kind', '--lang', '--markdown', '--max-age', '--max-cost-usd', '--multimodal', '--name-status', '--near-symbol', '--ner', '--no-extract', '--no-federated', '--older-than', '--pack', '--path', '--pattern', '--pending', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--repo', '--reset', '--resolve', '--restore-only', '--run-id', '--since', '--slug', '--source', '--source-id', '--stale', '--strategy', '--supersessions', '--symbol-kind', '--thin', '--timeout', '--type', '--verbose', '--workers', '--yes'],
'extract-conversation-facts': ['--aliases', '--all', '--all-sources', '--background', '--brain', '--brain-wide-max-cost-usd', '--break-lock', '--by-mention', '--clone-dir', '--code', '--concurrency', '--confirm-destructive', '--dry-run', '--embedding-dimensions', '--embedding-model', '--explain', '--fix', '--follow', '--force', '--force-break-lock', '--help', '--include-null-signature', '--json', '--judge-model', '--lang', '--limit', '--markdown', '--max-age', '--max-cost', '--max-cost-usd', '--max-runtime', '--model', '--multimodal', '--near-symbol', '--no', '--no-embed', '--no-embedding', '--no-extract', '--older-than', '--override-disabled', '--path', '--pattern', '--pending', '--pglite', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--reset', '--resolve', '--restore-only', '--segment-limit', '--session', '--since', '--sleep', '--slug', '--source', '--source-id', '--stale', '--supabase', '--supersessions', '--symbol-kind', '--thin', '--timeout', '--to', '--types', '--url', '--url-managed', '--version', '--workers', '--yes'],
@@ -56,12 +56,12 @@ export const CLI_FLAG_REGISTRY: Record<string, readonly string[]> = {
'friction': ['--agent', '--base', '--brain', '--compare', '--help', '--hint', '--json', '--kind', '--message', '--no-redact', '--phase', '--redact', '--run-id', '--severity', '--source', '--transcript-path', '--transcripts'],
'frontmatter': ['--aliases', '--all', '--allow-catch-all', '--brain', '--cached', '--diff-filter', '--dry-run', '--exclude-standard', '--fast', '--fix', '--force', '--from-pages', '--get', '--help', '--http', '--include-catch-all', '--include-null-signature', '--json', '--name-only', '--name-status', '--no-embedding', '--no-extract', '--no-verify', '--others', '--pattern', '--pending', '--reset', '--resolve', '--source', '--stale', '--strategy', '--supersessions', '--surface', '--thin', '--timeout', '--token-ttl', '--uninstall', '--write-back'],
'graph-query': ['--aliases', '--all', '--brain', '--depth', '--direction', '--explain', '--fast', '--force', '--from-pages', '--help', '--http', '--include-foreign', '--include-null-signature', '--json', '--lang', '--markdown', '--mcp-only', '--multimodal', '--near-symbol', '--no-embedding', '--no-extract', '--pattern', '--pending', '--reset', '--resolve', '--restore-only', '--source', '--stale', '--supersessions', '--surface', '--symbol-kind', '--thin', '--timeout', '--token-ttl', '--type'],
'hook': ['--aliases', '--all', '--allow-unverified-remote', '--batch-limit', '--brain', '--budget-ms', '--cached', '--count', '--delete-brain', '--detach', '--diff-filter', '--end-of-options', '--env', '--exclude-standard', '--fast', '--force', '--from-pages', '--get', '--harness', '--help', '--http', '--include-null-signature', '--jq', '--json', '--name-only', '--no-embedding', '--no-extract', '--once', '--others', '--path', '--pattern', '--pending', '--porcelain', '--project', '--quiet', '--remove', '--reset', '--resolve', '--show-current', '--show-toplevel', '--source', '--stale', '--stats', '--status', '--supersessions', '--surface', '--thin', '--timeout', '--token', '--token-ttl'],
'hook': ['--aliases', '--all', '--allow-unverified-remote', '--auto', '--batch-limit', '--brain', '--budget-ms', '--cached', '--count', '--delete-brain', '--detach', '--diff-filter', '--end-of-options', '--env', '--exclude-standard', '--fast', '--force', '--from-pages', '--get', '--harness', '--help', '--http', '--include-null-signature', '--jq', '--json', '--name-only', '--no-embedding', '--no-extract', '--once', '--others', '--path', '--pattern', '--pending', '--porcelain', '--project', '--pure', '--quiet', '--remove', '--reset', '--resolve', '--show-current', '--show-toplevel', '--source', '--stale', '--stats', '--status', '--supersessions', '--surface', '--thin', '--timeout', '--token', '--token-ttl'],
'import': ['--aliases', '--all', '--asof', '--background', '--brain', '--brain-wide-max-cost-usd', '--by-mention', '--cached', '--code', '--compile', '--concurrency', '--embedding-dimensions', '--embedding-model', '--exclude', '--exclude-standard', '--explain', '--fast', '--federated', '--fix', '--follow', '--force', '--force-rechunk', '--fresh', '--from-pages', '--full', '--help', '--http', '--include-gitignored', '--include-null-signature', '--json', '--lang', '--markdown', '--max-age', '--multimodal', '--name-status', '--no-embed', '--no-embedding', '--no-extract', '--no-federated', '--older-than', '--others', '--path', '--pattern', '--pending', '--pglite', '--priority', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--repo', '--reset', '--resolve', '--respect-gitignore', '--restore-only', '--since', '--skip-failed', '--source', '--source-id', '--stale', '--strategy', '--supabase', '--supersessions', '--surface', '--thin', '--timeout', '--token-ttl', '--url', '--workers'],
'init': ['--all', '--brain', '--chat-model', '--check', '--ctx-size', '--embedding-dimensions', '--embedding-model', '--embeddings', '--entity', '--expansion-model', '--fast', '--flag', '--force', '--from-pages', '--grant-types', '--help', '--http', '--issuer-url', '--json', '--judge-model', '--key', '--mcp-only', '--mcp-url', '--migrate-only', '--model', '--multimodal', '--no', '--no-embed', '--no-embedding', '--non-interactive', '--oauth-client-id', '--oauth-client-secret', '--path', '--pglite', '--provenance', '--reranking', '--schema-pack', '--scopes', '--skip-embed-check', '--source', '--stale', '--supabase', '--surface', '--to', '--token-ttl', '--touchpoint', '--url', '--version'],
'integrations': ['--auto', '--brain', '--dry-run', '--embeddings', '--fast', '--force', '--from-pages', '--help', '--http', '--json', '--no-embedding', '--overwrite', '--refresh', '--reranking', '--source', '--surface', '--target', '--token-ttl'],
'integrity': ['--aliases', '--all', '--auto', '--backend', '--background', '--brain', '--brain-wide-max-cost-usd', '--check', '--confidence', '--cost', '--dry-run', '--explain', '--fast', '--follow', '--force', '--fresh', '--from-pages', '--help', '--http', '--include-null-signature', '--json', '--limit', '--no-embedding', '--no-extract', '--pattern', '--pending', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--reset', '--resolve', '--review-lower', '--skip-bare-tweet', '--skip-urls', '--source', '--stale', '--supabase', '--supersessions', '--surface', '--thin', '--timeout', '--token-ttl', '--type', '--url'],
'jobs': ['--abbrev-ref', '--aliases', '--all', '--allow-empty', '--allow-protected', '--allow-shell-jobs', '--apply', '--asof', '--auto', '--auto-fix', '--auto-with-prompt', '--background', '--backoff-delay', '--backoff-jitter', '--backoff-type', '--batch', '--batch-size', '--bound-max-concurrent', '--bound-slug-prefixes', '--bound-source', '--bound-tools', '--brain', '--break-lock', '--budget-usd', '--budget-usd-per-day', '--by-mention', '--by-type', '--cached', '--catch-up', '--check', '--cli-path', '--cluster', '--cluster-errors', '--code', '--concurrency', '--confidence', '--confirm-destructive', '--content', '--date', '--days', '--delay', '--detach', '--diff-filter', '--dim', '--dimensions', '--dir', '--drain', '--dry-run', '--embedding-dimensions', '--embedding-model', '--empty', '--entities', '--exclude', '--exclude-standard', '--explain', '--fast', '--federated', '--federated-read', '--ff-only', '--file', '--fix', '--follow', '--force', '--force-break-lock', '--force-retry', '--format', '--fresh', '--from-meetings', '--from-pages', '--full', '--hard-deadline', '--health-interval', '--held-out', '--help', '--http', '--idempotency-key', '--image', '--include-frontmatter', '--include-gitignored', '--include-null-signature', '--infer-dates', '--inject-bootstrap', '--inline', '--input', '--install', '--interval', '--is-ancestor', '--job-id', '--job-isolation', '--json', '--kind', '--lang', '--limit', '--lock', '--markdown', '--max-age', '--max-attempts', '--max-cost-usd', '--max-crashes', '--max-rss', '--max-runtime-min', '--max-sources', '--max-stalled', '--max-usd', '--max-waiting', '--mcp-only', '--migrate-only', '--min-context', '--missing-path', '--mode', '--model', '--multimodal', '--name-only', '--name-status', '--near-symbol', '--ner', '--nice', '--no', '--no-auto-embed', '--no-embed', '--no-embedding', '--no-extract', '--no-federate', '--no-federated', '--no-gpg-sign', '--no-hard-deadline', '--no-inject', '--no-mutate', '--no-pull', '--no-renames', '--no-schema-pack', '--no-verify', '--no-worker', '--non-interactive', '--now', '--offset', '--older-than', '--once', '--order', '--orphan', '--others', '--output', '--override-disabled', '--pace', '--pace-max-concurrency', '--pack', '--parallel', '--params', '--path', '--pattern', '--pending', '--phase', '--pid-file', '--priority', '--progress-interval', '--progress-json', '--queue', '--quiet', '--redact-secrets', '--reenrich-after', '--refresh-cache', '--refresh-ms', '--remediate', '--remediation-plan', '--repo', '--reset', '--resolve', '--respect-gitignore', '--restore-only', '--resume', '--retry-failed', '--review-lower', '--run-id', '--save', '--segment-limit', '--serial', '--session', '--session-id', '--short', '--show-toplevel', '--sigkill-rescue', '--since', '--skip-bare-tweet', '--skip-failed', '--skip-urls', '--sleep', '--slug', '--slugs', '--source', '--source-id', '--src-subpath', '--stale', '--stats', '--status', '--strategy', '--supersessions', '--surface', '--swap-only', '--symbol-kind', '--target', '--target-score', '--thin', '--thin-threshold', '--timeout', '--timeout-ms', '--to', '--token-ttl', '--trusted-extraction', '--type', '--types', '--uninstall', '--unsafe-bypass-dream-guard', '--url', '--user', '--verbose', '--verify', '--version', '--watch', '--wedge-rescue', '--with-db', '--workers', '--yes'],
'jobs': ['--abbrev-ref', '--aliases', '--all', '--allow-empty', '--allow-protected', '--allow-shell-jobs', '--apply', '--asof', '--auto', '--auto-fix', '--auto-with-prompt', '--background', '--backoff-delay', '--backoff-jitter', '--backoff-type', '--batch', '--batch-size', '--bound-max-concurrent', '--bound-slug-prefixes', '--bound-source', '--bound-tools', '--brain', '--break-lock', '--budget-usd', '--budget-usd-per-day', '--by-mention', '--by-type', '--cached', '--catch-up', '--check', '--cli-path', '--cluster', '--cluster-errors', '--code', '--concurrency', '--confidence', '--confirm-destructive', '--content', '--date', '--days', '--delay', '--detach', '--diff-filter', '--dim', '--dimensions', '--dir', '--drain', '--dry-run', '--embedding-dimensions', '--embedding-model', '--empty', '--entities', '--exclude', '--exclude-standard', '--explain', '--fast', '--federated', '--federated-read', '--ff-only', '--file', '--fix', '--follow', '--force', '--force-break-lock', '--force-retry', '--format', '--fresh', '--from-meetings', '--from-pages', '--full', '--hard-deadline', '--health-interval', '--held-out', '--help', '--http', '--idempotency-key', '--image', '--include-frontmatter', '--include-gitignored', '--include-null-signature', '--infer-dates', '--inject-bootstrap', '--inline', '--input', '--install', '--interval', '--is-ancestor', '--job-id', '--job-isolation', '--json', '--kind', '--lang', '--limit', '--lock', '--lock-duration-ms', '--markdown', '--max-age', '--max-attempts', '--max-cost-usd', '--max-crashes', '--max-rss', '--max-runtime-min', '--max-sources', '--max-stalled', '--max-usd', '--max-waiting', '--mcp-only', '--migrate-only', '--min-context', '--missing-path', '--mode', '--model', '--multimodal', '--name-only', '--name-status', '--near-symbol', '--ner', '--nice', '--no', '--no-auto-embed', '--no-embed', '--no-embedding', '--no-extract', '--no-federate', '--no-federated', '--no-gpg-sign', '--no-hard-deadline', '--no-inject', '--no-mutate', '--no-pull', '--no-renames', '--no-schema-pack', '--no-verify', '--no-worker', '--non-interactive', '--now', '--offset', '--older-than', '--once', '--order', '--orphan', '--others', '--output', '--override-disabled', '--pace', '--pace-max-concurrency', '--pack', '--parallel', '--params', '--path', '--pattern', '--pending', '--phase', '--pid-file', '--priority', '--progress-interval', '--progress-json', '--queue', '--quiet', '--redact-secrets', '--reenrich-after', '--refresh-cache', '--refresh-ms', '--remediate', '--remediation-plan', '--repo', '--reset', '--resolve', '--respect-gitignore', '--restore-only', '--resume', '--retry-failed', '--review-lower', '--run-id', '--save', '--segment-limit', '--serial', '--session', '--session-id', '--short', '--show-toplevel', '--sigkill-rescue', '--since', '--skip-bare-tweet', '--skip-failed', '--skip-urls', '--sleep', '--slug', '--slugs', '--source', '--source-id', '--src-subpath', '--stale', '--stats', '--status', '--strategy', '--supersessions', '--surface', '--swap-only', '--symbol-kind', '--target', '--target-score', '--thin', '--thin-threshold', '--timeout', '--timeout-ms', '--to', '--token-ttl', '--trusted-extraction', '--type', '--types', '--uninstall', '--unsafe-bypass-dream-guard', '--url', '--user', '--verbose', '--verify', '--version', '--watch', '--wedge-rescue', '--with-db', '--workers', '--yes'],
'lint': ['--aliases', '--all', '--background', '--brain', '--brain-wide-max-cost-usd', '--dry-run', '--exclude', '--explain', '--fast', '--fix', '--follow', '--force', '--from-pages', '--help', '--http', '--include-null-signature', '--json', '--no-embedding', '--no-extract', '--pattern', '--pending', '--progress-interval', '--progress-json', '--quiet', '--remediate', '--reset', '--resolve', '--source', '--stale', '--supersessions', '--surface', '--thin', '--timeout', '--token-ttl'],
'lsd': ['--brain', '--force-resume', '--help', '--json', '--judge-model', '--limit', '--list-runs', '--max-cost', '--max-far-set', '--max-ideas-per-judge-call', '--no-save', '--resume', '--retry-judge', '--save', '--source', '--strict-budget', '--yes'],
'maintain': ['--aliases', '--all', '--background', '--brain', '--break-lock', '--by-mention', '--catch-up', '--column', '--concurrency', '--content-audit', '--count', '--detach', '--dim', '--dir', '--drain', '--dry-run', '--embedding-dimensions', '--embedding-model', '--explain', '--fast', '--fix', '--force', '--force-retry', '--force-schema', '--from-meetings', '--full', '--help', '--include-flagged', '--include-frontmatter', '--include-null-signature', '--index-audit', '--infer-dates', '--input', '--json', '--kind', '--lang', '--locks', '--markdown', '--max-cost', '--max-cost-usd', '--max-jobs', '--max-rss', '--max-usd', '--migrate-only', '--multimodal', '--near-symbol', '--ner', '--nice', '--no-extract', '--no-mutate', '--older-than', '--once', '--pack', '--parallel', '--params', '--path', '--pattern', '--pending', '--pglite', '--phase', '--pid-file', '--porcelain', '--probe-pglite', '--progress-json', '--query', '--queue', '--quiet', '--rebuild-rollup', '--regenerate', '--remediate', '--remediation-plan', '--reset', '--resolve', '--restore-only', '--resume', '--run-id', '--safe', '--scope', '--since', '--skills-dir', '--skip-failed', '--slugs', '--source', '--source-id', '--stale', '--status', '--supabase', '--supersessions', '--symbol-kind', '--target', '--target-score', '--thin', '--to', '--top-k', '--type', '--unsafe-bypass-dream-guard', '--url', '--verbose', '--window', '--workers', '--yes'],
+2 -2
View File
@@ -25,11 +25,11 @@ import { reflexPointerRationale } from './retrieval-reflex.ts';
export const VOLUNTEER_EVENTS_TTL_DAYS = 90;
/** Single source of truth for channel values — type + guards derive from it. */
export const VOLUNTEER_CHANNELS = ['op', 'reflex', 'watch', 'claude-code', 'codex'] as const;
export const VOLUNTEER_CHANNELS = ['op', 'reflex', 'watch', 'claude-code', 'codex', 'opencode'] as const;
export type VolunteerChannel = (typeof VOLUNTEER_CHANNELS)[number];
/** The harness subset — the ONLY channels a wire caller may claim. */
export const HARNESS_CHANNELS = ['claude-code', 'codex'] as const;
export const HARNESS_CHANNELS = ['claude-code', 'codex', 'opencode'] as const;
export type HarnessChannel = (typeof HARNESS_CHANNELS)[number];
/** Wire fallback: the only harness bootstrap registers hooks for today. */
+10
View File
@@ -141,6 +141,16 @@ export function buildCodexMcpAddArgv(p: { name: string; url: string; envVar: str
return ['mcp', 'add', p.name, '--url', p.url, '--bearer-token-env-var', p.envVar];
}
/**
* `opencode mcp add` argv for a remote HTTP server. The header value carries
* opencode's `{env:VAR}` interpolation LITERALLY opencode resolves it at
* read time, so the token never enters argv or the config file (verified
* against opencode 1.18.18; OPENCODE-CLI-PIN.md §mcp add).
*/
export function buildOpencodeMcpAddArgv(p: { name: string; url: string; envVar: string }): string[] {
return ['mcp', 'add', p.name, '--url', p.url, '--header', `Authorization=Bearer {env:${p.envVar}}`];
}
/**
* POSIX single-quote any arg that isn't already shell-safe, so `$()`, backticks,
* etc. in a token are inert literals when the block is pasted into a shell
+24
View File
@@ -5829,6 +5829,30 @@ export const MIGRATIONS: Migration[] = [
ALTER TABLE dream_verdicts ADD COLUMN IF NOT EXISTS triage_version INT;
`,
},
{
version: 130,
name: 'minion_jobs_lock_duration_ms',
// #4145: per-job lock lease. A single worker-global 30s lockDuration
// cannot serve both 2s shell jobs and 173s-average LLM subagent jobs —
// under host CPU saturation the renewal window was missed and healthy
// long-running handlers were force-evicted. The column carries an
// optional per-job lease; NULL means "worker default" (the pre-#4145
// behavior), so NO backfill is needed — the claim-time COALESCE against
// HANDLER_DEFAULT_LOCK_DURATION_MS (handler-timeouts.ts) owns all
// defaulting from here on. CHECK added via the idempotent drop-then-add
// pattern (v7 precedent) so migrated brains carry the same DB bound as
// fresh installs; the operative [5s,1h] range clamp lives app-side in
// clampLockDurationMs. No index: only per-row reads on already-indexed
// access paths (bootstrap-coverage: column-only, no probe needed).
// (Authored as v129; renumbered to v130 when #4152's
// dream_verdicts_triage_v1_columns landed the number first.)
idempotent: true,
sql: `
ALTER TABLE minion_jobs ADD COLUMN IF NOT EXISTS lock_duration_ms INTEGER;
ALTER TABLE minion_jobs DROP CONSTRAINT IF EXISTS chk_lock_duration_positive;
ALTER TABLE minion_jobs ADD CONSTRAINT chk_lock_duration_positive CHECK (lock_duration_ms IS NULL OR (lock_duration_ms >= 5000 AND lock_duration_ms <= 3600000));
`,
},
];
export const LATEST_VERSION = MIGRATIONS.length > 0
+75 -20
View File
@@ -1,28 +1,44 @@
/**
* handler-timeouts.ts per-handler default wall-clock budgets (#1737).
* handler-timeouts.ts per-handler-type defaults for the TWO time knobs a
* Minion job carries: the wall-clock budget (`timeout_ms`, #1737) and the
* lock lease (`lock_duration_ms`, #4145). They are DIFFERENT quantities
* the budget bounds total runtime (subagent: 30min), the lease bounds how
* long a dead worker's claim survives before the stall sweep reclaims it
* (subagent: 300s) and deriving one from the other would couple two
* unrelated tunables behind an implicit ratio policy. Both maps live here,
* under one header, precisely so they can't drift apart unseen.
*
* Short jobs (shell, lint, backlinks) want the tight default wall-clock
* (`2 * lockDuration * max_stalled`, computed in `handleWallClockTimeouts`
* when `timeout_ms IS NULL`). Long jobs do not: a 30-min LLM loop or a
* 10-15 min embed backfill submitted WITHOUT an explicit `timeout_ms` would
* inherit that short null-default and get wall-clock-killed mid-progress
* one half of #1737's thrash.
* WALL-CLOCK BUDGETS: short jobs (shell, lint, backlinks) want the tight
* default wall-clock (`2 * lock * max_stalled`, computed in
* `handleWallClockTimeouts` when `timeout_ms IS NULL`). Long jobs do not: a
* 30-min LLM loop submitted WITHOUT an explicit `timeout_ms` would inherit
* that short null-default and get wall-clock-killed mid-progress.
*
* Three layers apply the default (an explicit `opts.timeout_ms` always wins):
* LOCK LEASES: the worker-global default lease is 30s right for 2s shell
* jobs (fast dead-worker reclaim), structurally wrong for 173s-average LLM
* jobs, which had to renew ~12 consecutive times per run; under host CPU
* saturation a missed window force-evicted healthy handlers (the #4145
* incident). Long handlers get a 300s lease (5 renewal chances at the
* clamped 60s cadence); `shell` deliberately stays NULL worker default
* (child processes don't starve the loop, and verify-before-evict protects
* them anyway). Trade: dead-worker reclaim for mapped types slows to
* lease + reclaim grace + sweep cadence.
*
* Both defaults apply through the same layers (an explicit opts value
* always wins):
*
* 1. SUBMIT `MinionQueue.add` stamps the default onto the row. The value
* lives in `minion_jobs.timeout_ms`, not worker memory, so wall-clock
* behavior is stable across worker restart.
* 2. CLAIM `MinionQueue.claim` COALESCEs a NULL `timeout_ms` from this
* map (and derives `timeout_at` from the coalesced value). This is the
* durable invariant: it covers rows inserted before layer 1 existed and
* any writer that bypasses add(). Persisted by the claim UPDATE, so the
* restart-stability property holds here too.
* 3. ONE-SHOT migration v128 backfilled `timeout_ms` for non-terminal
* rows that predate both layers (they would otherwise never re-claim
* or die at the short null-default first). v128's values are a
* deliberate authoring-time SNAPSHOT of this map do NOT sync v128
* when editing the map below; layer 2 owns all future drift.
* lives on minion_jobs, not worker memory, so behavior is stable
* across worker restart.
* 2. CLAIM `MinionQueue.claim` COALESCEs a NULL column from the map
* (and derives `timeout_at` / `lock_until` from the coalesced value).
* This is the durable invariant: it covers rows inserted before
* layer 1 existed and any writer that bypasses add().
* 3. ONE-SHOT migration v128 backfilled `timeout_ms` for pre-#1737
* rows. v128's values are a deliberate authoring-time SNAPSHOT do
* NOT sync v128 when editing the maps; layer 2 owns all future drift.
* (`lock_duration_ms` needs no backfill: NULL simply means "worker
* default", which is the pre-#4145 behavior.)
*
* The 30-min anchor matches the explicit value cycle/patterns.ts already
* passes for subagent jobs, so this generalizes an existing convention
@@ -67,3 +83,42 @@ export const HANDLER_DEFAULT_TIMEOUT_MS: Readonly<Record<string, number>> = {
export function defaultTimeoutMsFor(jobName: string): number | null {
return HANDLER_DEFAULT_TIMEOUT_MS[jobName] ?? null;
}
const FIVE_MIN_MS = 5 * 60 * 1000;
const TWO_MIN_MS = 2 * 60 * 1000;
/**
* Default lock lease (ms) for long-running handler types (#4145). A handler
* not in this map returns `null` the worker-global `lockDuration` default
* (30s) applies see the header for why `shell` is deliberately absent.
* 300s = the issue's proposed lease for 173s-average LLM jobs (5 renewal
* chances at the 60s clamped cadence); 120s covers the single-LLM-call
* handlers whose p99 comfortably exceeds a 30s lease.
*/
export const HANDLER_DEFAULT_LOCK_DURATION_MS: Readonly<Record<string, number>> = {
subagent: FIVE_MIN_MS,
subagent_aggregator: FIVE_MIN_MS,
'embed-backfill': FIVE_MIN_MS,
'autopilot-cycle': FIVE_MIN_MS,
'autopilot-global-maintenance': FIVE_MIN_MS,
contextual_reindex_per_chunk: FIVE_MIN_MS,
chronicle_extract: TWO_MIN_MS,
'facts-absorb': TWO_MIN_MS,
};
export function defaultLockDurationMsFor(jobName: string): number | null {
return HANDLER_DEFAULT_LOCK_DURATION_MS[jobName] ?? null;
}
/** Clamp bounds for explicit per-job lease input [5s, 1h]. The floor
* kills the 1ms-lock foot-gun (instant stall-reclaim thrash); the ceiling
* kills the immortal-lock foot-gun (a crashed worker's claim surviving an
* hour). Shared by queue.add, the CLI flag, and the MCP op (ENG-E4) so
* three hand-copied bounds can't drift. The DB CHECK only enforces > 0;
* this clamp is the operative bound (precedent: max_stalled [1,100]). */
export const LOCK_DURATION_MS_MIN = 5_000;
export const LOCK_DURATION_MS_MAX = 3_600_000;
export function clampLockDurationMs(raw: number): number {
return Math.max(LOCK_DURATION_MS_MIN, Math.min(LOCK_DURATION_MS_MAX, Math.floor(raw)));
}
+429 -76
View File
@@ -18,14 +18,25 @@
* `Promise.race(call, timeoutPromise)` with `callTimeoutMs` so the
* call cannot pend longer than the configured budget.
*
* - **Threshold math**: with `lockDuration=30s` and `interval=15s`,
* a 3-strike count-based abort fires at t=45s but the lock has
* been reclaimable since t=30s a 15s window where another
* worker can claim the same job. `runLockRenewalTick` aborts based
* on `Date.now() - lastSuccessfulRenewalAt >= lockDuration -
* safetyMargin` (time-based), so the worker voluntarily releases
* BEFORE the stall detector can reclaim. The failure counter is
* kept for audit-event labeling only.
* - **Verify-before-evict (#4145, replaces the v0.41.22.2 abort-at-
* deadline doctrine)**: a thrown/timed-out renewal is NOT evidence of
* loss under event-loop starvation the UPDATE may even have landed
* server-side while the local race timeout won. When the NEXT tick
* would land past the soft deadline (`lockDuration - safetyMargin`,
* cadence-aware per CDX-4), the tick runs ONE bounded VERIFY renewal:
* fenced-true starved-but-ours, lease re-extended, keep working;
* fenced-false CERTAIN loss, abort (stall detector requeues, no
* attempt burned); verify unreachable defer and retry next tick,
* aborting only past the `hardEvictMs` backstop (a LOCAL decision
* under uncertainty that bounds blind external side effects during a
* total outage). The local clock only schedules WHEN to verify
* the DB fence decides WHETHER to evict; production binds `deps.now`
* to a monotonic source so wall-clock jumps can't distort the math.
* Accepted timing behavior (ENG-E1): at the 30s default a failed
* renewal (10s) + verify (10s) can span 20s > the 15s cadence, so
* tickInFlight skips one tick benign (a verify-success re-extends
* the lease and resets the baseline); do not "fix" it into an
* overlap. The failure counter is kept for audit-event labeling only.
*
* - **Cancelled-tick race**: if the job ends while a renewLock call
* is mid-flight, the IIFE in worker.ts must bail without writing
@@ -47,6 +58,43 @@
* below; tests + workers both consume it.
*/
/**
* Named error for the per-call timeout race so cause classification is
* name-based (`call-timeout` vs `refused`), never message-sniffing
* (issue #4145 request 3).
*/
export class RenewalCallTimeoutError extends Error {
constructor(what: string, ms: number) {
super(`${what} timed out after ${ms}ms`);
this.name = 'RenewalCallTimeoutError';
}
}
/**
* Why a renewal attempt failed, as far as the tick can tell:
* - `call-timeout` our own race timer fired (starved loop, slow pool,
* or slow DB the tick can't distinguish; lateness
* + load telemetry does).
* - `refused` the driver threw (SQLSTATE rides in the audit event).
* - `fenced-lost` the fenced UPDATE matched 0 rows: CERTAIN loss.
*/
export type RenewalFailureCause = 'call-timeout' | 'refused' | 'fenced-lost';
/**
* Optional telemetry threaded to the audit sink alongside each event.
* All fields additive the 4-outcome audit contract is unchanged and
* pre-upgrade JSONL lines parse fine without them.
*/
export interface LockRenewalTelemetryCtx {
cause?: RenewalFailureCause;
lateness_ms?: number;
overlap_skips?: number;
load1?: number;
cores?: number;
via?: 'renewal' | 'verify';
deadline_deferred?: boolean;
}
export interface LockRenewalKnobs {
/**
* Failure counter cap used ONLY for audit-event labeling.
@@ -56,17 +104,36 @@ export interface LockRenewalKnobs {
maxFailuresForAudit: number;
/**
* Per-renewLock-call timeout enforced via `Promise.race`.
* Env: `GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS`. Default: `lockDuration / 3`.
* Env: `GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS`. Default: `min(lockDuration / 3, 15s)`
* (the cap keeps a long per-job lease from inheriting a call budget that
* wedges tickInFlight across cadence windows).
* Bounds the "hung renewLock wedges the re-entrancy guard forever" vector.
*/
callTimeoutMs: number;
/**
* Time-based abort fires when `now - lastSuccessfulRenewalAt >=
* lockDuration - safetyMarginMs`. Default safety margin gives ~5s
* of headroom before another worker could reclaim the lock.
* Env: `GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS`. Default: `lockDuration / 6`.
* The soft deadline is `lockDuration - safetyMarginMs`: once the NEXT
* tick would land past it, the tick runs the at-deadline VERIFY renewal
* (fenced re-check) instead of retrying blindly. The margin is the
* headroom the verify has to complete before the lease actually lapses.
* Env: `GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS`. Default: `min(lockDuration / 6, 30s)`.
*/
safetyMarginMs: number;
/**
* The hard local backstop (#4145): when even the VERIFY renewal is
* unreachable (throws/times out) for this long since the last success,
* abort anyway. This is an uncertainty BOUND, not a certainty claim
* fenced-false is the only CERTAIN loss signal; the hard evict merely
* caps how long a handler with non-fenced EXTERNAL side effects may run
* blind during a total DB outage. APPROXIMATE by design: checked after
* a primary call + verify (overshoot up to ~2×callTimeoutMs + timer
* lateness), and eviction stays cooperative (an abort-ignoring handler
* outlives it the kill/reap follow-up TODO is the true bound).
* Setting this to the soft deadline approximates the legacy
* abort-at-deadline behavior (the verify still runs once).
* Env: `GBRAIN_LOCK_RENEWAL_HARD_EVICT_MS`. Default: `2 × lockDuration`,
* floored to the soft deadline (warn-once on floor).
*/
hardEvictMs: number;
}
/**
@@ -90,15 +157,36 @@ export function _resetKnobWarningsForTests(): void {
* operator who sets `GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS=abc` gets a
* loud-but-not-fatal nudge AND a working worker.
*/
/**
* The renewal cadence cap: a long lease renews every 60s (multiple chances
* per window, matching the cycle refresher's philosophy) instead of the
* bare lease/2. Leases 120s keep the legacy /2 exactly. ONE home for the
* formula the worker's timer and resolveLockRenewalKnobs' default both
* derive from it, so the relational validation always runs against the
* cadence production actually uses.
*/
export const RENEWAL_INTERVAL_CAP_MS = 60_000;
export function renewalIntervalFor(lockDurationMs: number): number {
return Math.max(1, Math.min(Math.floor(lockDurationMs / 2), RENEWAL_INTERVAL_CAP_MS));
}
export function resolveLockRenewalKnobs(
env: Record<string, string | undefined>,
lockDurationMs: number,
intervalMs: number = renewalIntervalFor(lockDurationMs),
): LockRenewalKnobs {
const defaultMaxFailures = 3;
const defaultCallTimeout = Math.max(1, Math.floor(lockDurationMs / 3));
const defaultSafetyMargin = Math.max(1, Math.floor(lockDurationMs / 6));
// #4145: the derived defaults CAP at 15s/30s so a long per-job lease
// (300s) doesn't inherit a 100s call timeout (which would wedge
// tickInFlight across cadence windows) or a 50s margin. The 30s worker
// default keeps today's 10s/5s exactly. Env overrides are still applied
// first and then pass the relational validation below.
const defaultCallTimeout = Math.min(Math.max(1, Math.floor(lockDurationMs / 3)), 15_000);
const defaultSafetyMargin = Math.min(Math.max(1, Math.floor(lockDurationMs / 6)), 30_000);
const defaultHardEvict = lockDurationMs * 2;
return {
const knobs: LockRenewalKnobs = {
maxFailuresForAudit: parsePositiveInt(
env.GBRAIN_LOCK_RENEWAL_MAX_FAILURES,
defaultMaxFailures,
@@ -114,7 +202,54 @@ export function resolveLockRenewalKnobs(
defaultSafetyMargin,
'GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS',
),
hardEvictMs: parsePositiveInt(
env.GBRAIN_LOCK_RENEWAL_HARD_EVICT_MS,
defaultHardEvict,
'GBRAIN_LOCK_RENEWAL_HARD_EVICT_MS',
),
};
// [CDX-10] Relational validation: positive-integer parsing alone lets a
// margin exceed the lease or a call timeout exceed the cadence, which
// silently re-breaks the deadline math. Clamp the offending knob (to a
// relationally-valid derivation) and warn once per process per knob.
// Runs per job launch against the EFFECTIVE per-job lease.
if (knobs.safetyMarginMs >= lockDurationMs / 2) {
warnRelationalClamp(
'GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS',
`safetyMarginMs (${knobs.safetyMarginMs}) must be < lockDuration/2 (${lockDurationMs / 2})`,
defaultSafetyMargin,
);
knobs.safetyMarginMs = defaultSafetyMargin;
}
if (knobs.callTimeoutMs > intervalMs) {
warnRelationalClamp(
'GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS',
`callTimeoutMs (${knobs.callTimeoutMs}) must be <= the renewal cadence (${intervalMs}) or a slow call wedges tickInFlight across intervals`,
intervalMs,
);
knobs.callTimeoutMs = intervalMs;
}
const softDeadline = lockDurationMs - knobs.safetyMarginMs;
if (knobs.hardEvictMs < softDeadline) {
warnRelationalClamp(
'GBRAIN_LOCK_RENEWAL_HARD_EVICT_MS',
`hardEvictMs (${knobs.hardEvictMs}) must be >= the soft deadline (${softDeadline})`,
softDeadline,
);
knobs.hardEvictMs = softDeadline;
}
return knobs;
}
function warnRelationalClamp(name: string, violation: string, clampedTo: number): void {
const key = `relational:${name}`;
if (!_warnedKnobs.has(key)) {
_warnedKnobs.add(key);
process.stderr.write(
`[lock-renewal] ${violation}; clamping to ${clampedTo}\n`,
);
}
}
function parsePositiveInt(raw: string | undefined, fallback: number, name: string): number {
@@ -148,11 +283,15 @@ function warnAndFallback(name: string, raw: string, fallback: number): number {
*/
export interface LockRenewalDeps {
/**
* The optional `opts.signal` is aborted when this call loses the tick's
* timeout race, so the underlying UPDATE is CANCELLED (postgres.js
* `.cancel()` via executeRawDirect) instead of orphaned on a checked-out
* pool slot for its full server-side duration the #6 starvation class.
* Optional-param widening keeps the legacy 3-arg test mocks compiling.
* The fenced renewal UPDATE. The optional `opts.signal` is aborted when
* this call loses the tick's timeout race (both attempts: primary AND the
* #4145 at-deadline verify), so the underlying UPDATE is CANCELLED
* (postgres.js `.cancel()` via executeRawDirect) instead of orphaned on a
* checked-out pool slot for its full server-side duration the #6
* starvation class. Cancellation is BEST-EFFORT (pool acquisition and PG
* protocol cancel are async; PGLite ignores the signal); the token FENCE
* is the correctness authority. Optional-param widening keeps the legacy
* 3-arg test mocks compiling.
*/
renewLock: (
jobId: number,
@@ -166,8 +305,9 @@ export interface LockRenewalDeps {
/**
* Injectable for hermetic Promise.race tests. Production:
* `globalThis.setTimeout`. The function must return a value that
* `clearTimeout` accepts, but this seam doesn't expose clearTimeout
* because the timeout race fires-and-forgets. The losing renewLock is no
* `clearTimeout` accepts: on the win path the race clears the losing
* timer (via global clearTimeout) so it doesn't later fire a stray
* no-op abort against a settled query. The losing renewLock is no
* longer merely abandoned: the timeout callback also aborts the per-call
* signal so the query releases its pool slot.
*/
@@ -190,6 +330,21 @@ export interface LockRenewalDeps {
* (optional-param) signature stays back-compatible with no-arg test mocks.
*/
reconnect?: (ctx?: { error?: unknown }) => Promise<void>;
/**
* OPTIONAL host-load probe for eviction telemetry (issue #4145 req 3).
* The worker binds `os.loadavg()[0]` + a cached core count. Every call
* is wrapped in try/catch here (CEO-F2) a throwing telemetry hook
* must never re-open the unhandledRejection class this module exists
* to close; on throw the event simply logs without load fields.
*/
loadSnapshot?: () => { load1: number; cores: number };
/**
* OPTIONAL success hook. The worker resets its event-loop-delay
* histogram here (R2-9) so an eviction-time sample attributes to the
* window since the LAST SUCCESSFUL renewal, not process lifetime.
* Best-effort: called inside try/catch.
*/
onRenewalSuccess?: () => void;
}
/**
@@ -198,9 +353,11 @@ export interface LockRenewalDeps {
* import the full audit module and inflate the test surface.
*/
export interface LockRenewalAuditSinkLike {
logFailure(jobId: number, jobName: string, attempt: number, err: unknown): void;
logSuccessAfterFailure(jobId: number, jobName: string, recoveredAfterAttempts: number): void;
logGaveUp(jobId: number, jobName: string, totalFailures: number, err: unknown): void;
// The trailing `ctx` is optional and additive (ENG-E2): pre-existing
// test fakes with the old arity stay structurally assignable.
logFailure(jobId: number, jobName: string, attempt: number, err: unknown, ctx?: LockRenewalTelemetryCtx): void;
logSuccessAfterFailure(jobId: number, jobName: string, recoveredAfterAttempts: number, ctx?: LockRenewalTelemetryCtx): void;
logGaveUp(jobId: number, jobName: string, totalFailures: number, err: unknown, ctx?: LockRenewalTelemetryCtx): void;
}
export interface LockRenewalState {
@@ -221,13 +378,45 @@ export interface LockRenewalState {
* cancellation event AND the post-await branch decisions.
*/
cancelled: () => boolean;
/**
* The renewal timer's cadence (ms) the SAME value the worker gave
* setInterval. Used for tick-lateness telemetry (below) and the
* cadence-aware verify trigger.
*/
intervalMs: number;
/**
* Timestamp of the previous tick's entry, on the SAME clock as
* `deps.now` (production binds a monotonic source; R2-4). Seeded to
* launch time. Lateness = max(0, now - lastTickFiredAt - intervalMs)
* is the PRIMARY local-starvation signal: interval callbacks COALESCE
* under a blocked event loop (one late callback, not N), so a
* missed-tick counter cannot measure starvation lateness can.
*/
lastTickFiredAt: number;
/**
* Count of interval callbacks skipped by the worker's tickInFlight
* re-entrancy guard. These are OVERLAP skips (a prior tick still in
* flight), NOT missed intervals see lastTickFiredAt. Incremented by
* the worker's interval closure; read here for telemetry.
*/
overlapSkips: number;
}
export type TickResult =
| { kind: 'ok' }
| { kind: 'cancelled' }
| { kind: 'lock_lost' }
| { kind: 'should_abort'; reason: 'lock-renewal-failed' };
| { kind: 'lock_lost'; cause: 'fenced-lost'; via: 'renewal' | 'verify' }
| {
kind: 'should_abort';
reason: 'lock-renewal-failed';
/** What the FINAL failed attempt looked like (R2-7 telemetry). */
cause: RenewalFailureCause;
latenessMs: number;
sinceLastSuccessMs: number;
overlapSkips: number;
load1?: number;
cores?: number;
};
/**
* Execute one renewal tick. Returns a tagged result the worker switches
@@ -238,69 +427,145 @@ export async function runLockRenewalTick(
deps: LockRenewalDeps,
state: LockRenewalState,
): Promise<TickResult> {
// Tick-lateness telemetry (CDX-13): how late did this callback fire vs
// its own cadence? Computed FIRST — even a cancelled tick advances the
// baseline so the next measurement stays honest.
const tickEnteredAt = deps.now();
const latenessMs = Math.max(0, tickEnteredAt - state.lastTickFiredAt - state.intervalMs);
state.lastTickFiredAt = tickEnteredAt;
if (state.cancelled()) return { kind: 'cancelled' };
let renewed: boolean;
// Per-call cancellation: when the timeout wins the race, abort the signal
// so the losing UPDATE releases its pool slot instead of holding it until
// the server finishes (issue #6 — an abandoned racer under a saturated
// pooler pinned a checked-out connection for minutes). On the win path the
// late-firing timer aborts an already-settled query, which runUnsafe
// ignores (abort listener removed in its .finally).
const callAbort = new AbortController();
try {
renewed = await Promise.race([
deps.renewLock(state.jobId, state.lockToken, state.lockDurationMs, { signal: callAbort.signal }),
new Promise<never>((_, reject) => {
deps.setTimeout(() => {
callAbort.abort();
reject(new Error(`renewLock timed out after ${state.knobs.callTimeoutMs}ms`));
}, state.knobs.callTimeoutMs);
}),
]);
renewed = await racedRenewLock(deps, state);
} catch (err) {
if (state.cancelled()) return { kind: 'cancelled' };
state.consecutiveFailures += 1;
const cause = classifyFailure(err);
const load = safeLoadSnapshot(deps);
const telemetry: LockRenewalTelemetryCtx = {
cause,
lateness_ms: latenessMs,
overlap_skips: state.overlapSkips,
...load,
};
// Defense-in-depth (codex C4): audit must never escape this catch.
try {
deps.audit.logFailure(state.jobId, state.jobName, state.consecutiveFailures, err);
deps.audit.logFailure(state.jobId, state.jobName, state.consecutiveFailures, err, telemetry);
} catch { /* audit best-effort */ }
const sinceLastSuccess = deps.now() - state.lastSuccessfulRenewalAt;
const deadline = state.lockDurationMs - state.knobs.safetyMarginMs;
if (sinceLastSuccess >= deadline) {
try {
deps.audit.logGaveUp(state.jobId, state.jobName, state.consecutiveFailures, err);
} catch { /* audit best-effort */ }
return { kind: 'should_abort', reason: 'lock-renewal-failed' };
}
// issue #1678 (Codex #2): not yet at the deadline, so we'll retry on the
// next tick. If the engine can rebuild its pool, do it ONCE now (bounded
// by callTimeoutMs) so the next renewLock sees a live connection instead
// of throwing the same reaped-socket error until the deadline. Best-effort:
// a reconnect throw/timeout is swallowed (next tick retries) and must NEVER
// escape this catch — that would re-introduce the unhandledRejection class
// this module was built to close.
if (deps.reconnect) {
const reconnect = deps.reconnect;
try {
await Promise.race([
// Thread the triggering renewLock error (CODEX impl review #2) so the
// engine can classify a CONNECTION_ENDED pooler reap as `reap_detected`.
reconnect({ error: err }),
new Promise<never>((_, reject) => {
deps.setTimeout(
() => reject(new Error(`reconnect timed out after ${state.knobs.callTimeoutMs}ms`)),
state.knobs.callTimeoutMs,
);
}),
]);
} catch { /* reconnect best-effort; next tick retries against a fresh attempt */ }
// [CDX-4] Cadence-aware trigger: run the at-deadline VERIFY when the
// NEXT tick would land past the soft deadline — `>= deadline` alone is
// unreachable under cadence quantization (a 300s lease with a 60s
// cadence has no tick between 240s and expiry; the first eligible tick
// would already be past the lease).
if (sinceLastSuccess + state.intervalMs < deadline) {
// Not near the deadline: retry on the next tick.
// issue #1678 (Codex #2): if the engine can rebuild its pool, do it
// ONCE now (bounded) so the next renewLock sees a live connection.
await attemptReconnectOnce(deps, state, err);
if (state.cancelled()) return { kind: 'cancelled' };
return { kind: 'ok' }; // counter incremented; not yet at deadline
}
return { kind: 'ok' }; // counter incremented; not yet at deadline
// ── At the deadline: VERIFY before evicting (#4145) ──────────────────
// The throw above is NOT evidence of loss (db-lock.ts doctrine:
// fenced-false = certain, throw = transient). Under event-loop
// starvation the renewal may even have LANDED server-side while our
// local race timeout won. renewLock is fenced on lock_token and does
// NOT check lock_until, so an expired-but-unstolen lease renews fine —
// ask the DB the authoritative question.
//
// Reconciling with queue.ts's "renewLock is deliberately not
// retry-wrapped" rationale (two-holders risk): that comment forbids
// BACKGROUND retries that outlive this tick's own timeout race. This
// verify is a synchronous, cancelled()-guarded, callTimeoutMs-bounded
// call INSIDE the tick's flow, and both UPDATEs are same-token
// idempotent lease extensions — a fenced row cannot gain two holders.
let verified: boolean | null = null; // null = the verify itself threw
let verifyErr: unknown = null;
try {
verified = await racedRenewLock(deps, state);
} catch (vErr) {
verifyErr = vErr;
}
if (state.cancelled()) return { kind: 'cancelled' };
if (verified === true) {
// Starved-but-ours: the lease is re-extended; the incident-saving path.
try {
deps.audit.logSuccessAfterFailure(
state.jobId, state.jobName, state.consecutiveFailures, { via: 'verify' },
);
} catch { /* audit best-effort */ }
state.consecutiveFailures = 0;
state.lastSuccessfulRenewalAt = deps.now();
try { deps.onRenewalSuccess?.(); } catch { /* telemetry best-effort */ }
return { kind: 'ok' };
}
if (verified === false) {
// Fenced miss: CERTAIN loss (stall sweep reclaimed / pauseJob).
return { kind: 'lock_lost', cause: 'fenced-lost', via: 'verify' };
}
// The verify was unreachable too — that is a second failed attempt.
// NOTE for audit readers: at the deadline `attempt` advances by 2 per
// tick (primary + verify) — the counter counts ATTEMPTS, not ticks.
state.consecutiveFailures += 1;
const verifyCause = classifyFailure(verifyErr);
// Re-sample load: the bounded verify can consume up to callTimeoutMs,
// and the verify-failure event should carry the load at ITS failure,
// not a snapshot stale by the verify's whole duration.
const verifyLoad = safeLoadSnapshot(deps);
const verifyTelemetry: LockRenewalTelemetryCtx = {
cause: verifyCause,
lateness_ms: latenessMs,
overlap_skips: state.overlapSkips,
...verifyLoad,
};
// Recompute elapsed AFTER the verify: the bounded verify itself can
// consume up to callTimeoutMs, and comparing the stale pre-verify value
// would defer one extra cadence past the advertised backstop.
const sinceLastSuccessAfterVerify = deps.now() - state.lastSuccessfulRenewalAt;
if (sinceLastSuccessAfterVerify >= state.knobs.hardEvictMs) {
// Hard backstop: a LOCAL decision under uncertainty, by design —
// bounds how long non-fenced external side effects run blind during
// a total outage. DB writes stay split-brain-safe via the fence.
try {
deps.audit.logGaveUp(
state.jobId, state.jobName, state.consecutiveFailures, verifyErr ?? err, verifyTelemetry,
);
} catch { /* audit best-effort */ }
return {
kind: 'should_abort',
reason: 'lock-renewal-failed',
cause: verifyCause,
latenessMs,
sinceLastSuccessMs: sinceLastSuccessAfterVerify,
overlapSkips: state.overlapSkips,
...verifyLoad,
};
}
// Deferred past the soft deadline: keep the job and retry next tick —
// the fence is the correctness backstop (a reclaim surfaces as a
// fenced-false on the next attempt; no attempt is burned). CEO-F3:
// give the next tick a live pool.
try {
deps.audit.logFailure(
state.jobId, state.jobName, state.consecutiveFailures, verifyErr ?? err,
{ ...verifyTelemetry, deadline_deferred: true },
);
} catch { /* audit best-effort */ }
await attemptReconnectOnce(deps, state, verifyErr ?? err);
if (state.cancelled()) return { kind: 'cancelled' };
return { kind: 'ok' };
}
if (state.cancelled()) return { kind: 'cancelled' };
@@ -308,8 +573,9 @@ export async function runLockRenewalTick(
// Token-fence failure: another worker reclaimed the row, or pauseJob
// cleared the token. NOT an infrastructure fault — no audit event
// (audit channel is for infrastructure faults only). The worker
// observes `lock_lost` and stderr-warns + aborts.
return { kind: 'lock_lost' };
// observes `lock_lost` and stderr-warns + aborts. This is the ONLY
// CERTAIN loss signal (exactly 0 rows matched the fence).
return { kind: 'lock_lost', cause: 'fenced-lost', via: 'renewal' };
}
if (state.consecutiveFailures > 0) {
@@ -318,10 +584,97 @@ export async function runLockRenewalTick(
state.jobId,
state.jobName,
state.consecutiveFailures,
{ via: 'renewal' },
);
} catch { /* audit best-effort */ }
state.consecutiveFailures = 0;
}
state.lastSuccessfulRenewalAt = deps.now();
// R2-9: let the worker reset its event-loop-delay histogram so the next
// eviction-time sample attributes to the window since THIS success.
try { deps.onRenewalSuccess?.(); } catch { /* telemetry best-effort */ }
return { kind: 'ok' };
}
/**
* CEO-F2: telemetry must never throw into the tick's control flow a
* throwing loadavg probe re-opening the unhandledRejection class would
* be a bitter irony. On throw, the event logs without load fields.
*/
function safeLoadSnapshot(deps: LockRenewalDeps): { load1?: number; cores?: number } {
try {
const s = deps.loadSnapshot?.();
return s ? { load1: s.load1, cores: s.cores } : {};
} catch {
return {};
}
}
/**
* One bounded renewal attempt: renewLock raced against callTimeoutMs.
* When the timeout wins the race it also ABORTS the per-call signal so the
* losing UPDATE releases its pool slot instead of holding it until the
* server finishes (issue #6 an abandoned racer under a saturated pooler
* pinned a checked-out connection for minutes). On the win path the
* late-firing timer aborts an already-settled query, which runUnsafe
* ignores (abort listener removed in its .finally). Cancellation is
* BEST-EFFORT (CDX-2/R2-2 the fence is the correctness authority; PGLite
* ignores the signal and resolves via the race alone). The race attaches
* handlers to both contenders, so a late loser settlement is absorbed,
* never an unhandledRejection.
*/
function racedRenewLock(deps: LockRenewalDeps, state: LockRenewalState): Promise<boolean> {
const controller = new AbortController();
let timer: unknown;
return Promise.race([
deps.renewLock(state.jobId, state.lockToken, state.lockDurationMs, { signal: controller.signal }),
new Promise<never>((_, reject) => {
timer = deps.setTimeout(() => {
try { controller.abort(); } catch { /* best-effort */ }
reject(new RenewalCallTimeoutError('renewLock', state.knobs.callTimeoutMs));
}, state.knobs.callTimeoutMs);
}),
]).finally(() => {
// Win-path hygiene: clear the losing timer so it can't fire a stray
// late abort at a settled query (absorbed by runUnsafe, but noisy).
// Test fakes may return null from the seam; clearTimeout(null) is a
// harmless no-op.
try { clearTimeout(timer as ReturnType<typeof setTimeout>); } catch { /* best-effort */ }
}) as Promise<boolean>;
}
function classifyFailure(err: unknown): RenewalFailureCause {
return err instanceof Error && err.name === 'RenewalCallTimeoutError' ? 'call-timeout' : 'refused';
}
/**
* issue #1678 (Codex #2): bounded, best-effort, ONCE-per-tick pool rebuild
* so the NEXT renewal attempt sees a live connection instead of throwing
* the same reaped-socket error. Threads the triggering error so the engine
* can classify a CONNECTION_ENDED pooler reap as `reap_detected`. A
* reconnect throw/timeout is swallowed it must NEVER escape into the
* tick's catch (the unhandledRejection class this module exists to close).
*/
async function attemptReconnectOnce(
deps: LockRenewalDeps,
state: LockRenewalState,
err: unknown,
): Promise<void> {
if (!deps.reconnect) return;
const reconnect = deps.reconnect;
let timer: unknown;
try {
await Promise.race([
reconnect({ error: err }),
new Promise<never>((_, reject) => {
timer = deps.setTimeout(
() => reject(new Error(`reconnect timed out after ${state.knobs.callTimeoutMs}ms`)),
state.knobs.callTimeoutMs,
);
}),
]);
} catch { /* reconnect best-effort; next tick retries against a fresh attempt */
} finally {
try { clearTimeout(timer as ReturnType<typeof setTimeout>); } catch { /* best-effort */ }
}
}
+111 -14
View File
@@ -16,7 +16,10 @@ import type {
import { rowToMinionJob, rowToInboxMessage, rowToAttachment } from './types.ts';
import { validateAttachment } from './attachments.ts';
import { isProtectedJobName } from './protected-names.ts';
import { defaultTimeoutMsFor, HANDLER_DEFAULT_TIMEOUT_MS } from './handler-timeouts.ts';
import {
defaultTimeoutMsFor, HANDLER_DEFAULT_TIMEOUT_MS,
defaultLockDurationMsFor, HANDLER_DEFAULT_LOCK_DURATION_MS, clampLockDurationMs,
} from './handler-timeouts.ts';
import {
withRetry, BULK_RETRY_OPTS, resolveBulkRetryOpts, computeNextDelay,
isRetryableConnError,
@@ -38,6 +41,62 @@ export interface TrustedSubmitOpts {
const MIGRATION_VERSION = 7;
const DEFAULT_MAX_SPAWN_DEPTH = 5;
/**
* Stall-sweep reclaim grace (#4145, CDX-7): don't reclaim a row whose
* `lock_until` lapsed within the last N ms. When a CPU-starved worker's
* event loop unblocks, its coalesced renewal tick and the stall sweep
* fire in the same burst if the sweep's UPDATE lands first it steals
* the OWNER'S live job. The grace is a HEAD-START for the owner's
* recovery renewal, not a guarantee: it only covers starvation bursts
* shorter than the grace, and a healthy second worker's sweep still
* wins beyond it. Minion analog of `GBRAIN_LOCK_STEAL_GRACE_SECONDS`
* (db-lock.ts), adapted because minion_jobs has no last_refreshed_at.
*
* Cost: dead-worker recovery becomes lock_until + grace + up to
* stalledInterval. Env `GBRAIN_MINION_STALL_RECLAIM_GRACE_MS` (0 allowed
* restores the exact legacy reclaim predicate).
*/
export const DEFAULT_STALL_RECLAIM_GRACE_MS = 15_000;
const _warnedGraceEnv = new Set<string>();
export function _resetStallGraceWarningsForTests(): void {
_warnedGraceEnv.clear();
}
export function resolveStallReclaimGraceMs(
env: Record<string, string | undefined> = process.env,
): number {
const raw = env.GBRAIN_MINION_STALL_RECLAIM_GRACE_MS;
if (raw === undefined || raw.trim() === '') return DEFAULT_STALL_RECLAIM_GRACE_MS;
// Unlike the lock-renewal knobs, 0 is a VALID value here (legacy reclaim).
if (!/^\d+$/.test(raw.trim())) {
if (!_warnedGraceEnv.has(raw)) {
_warnedGraceEnv.add(raw);
process.stderr.write(
`[minions] env GBRAIN_MINION_STALL_RECLAIM_GRACE_MS=${JSON.stringify(raw)} is not a non-negative integer; ` +
`falling back to default ${DEFAULT_STALL_RECLAIM_GRACE_MS}\n`,
);
}
return DEFAULT_STALL_RECLAIM_GRACE_MS;
}
const n = Number(raw.trim());
// Cap at 10 minutes: an absurd digit string (Number → huge/Infinity)
// would otherwise push the sweep cutoff to -infinity and silently
// disable stalled-job recovery altogether.
const MAX_GRACE_MS = 600_000;
if (n > MAX_GRACE_MS) {
if (!_warnedGraceEnv.has(raw)) {
_warnedGraceEnv.add(raw);
process.stderr.write(
`[minions] env GBRAIN_MINION_STALL_RECLAIM_GRACE_MS=${JSON.stringify(raw)} exceeds the ${MAX_GRACE_MS}ms cap; clamping\n`,
);
}
return MAX_GRACE_MS;
}
return n;
}
const DEFAULT_MAX_ATTACHMENT_BYTES = 5 * 1024 * 1024; // 5 MiB
const TERMINAL_STATUSES = ['completed', 'failed', 'dead', 'cancelled'] as const;
@@ -354,11 +413,11 @@ export class MinionQueue {
const baseCols = `name, queue, status, priority, data, max_attempts, backoff_type,
backoff_delay, backoff_jitter, delay_until, parent_job_id, on_child_fail,
depth, max_children, timeout_ms, remove_on_complete, remove_on_fail, idempotency_key,
depth, max_children, timeout_ms, lock_duration_ms, remove_on_complete, remove_on_fail, idempotency_key,
quiet_hours, stagger_key`;
const baseVals = `$1, $2, $3, $4, $5::jsonb, $6, $7, $8, $9, $10, $11, $12, $13, $14, $15, $16, $17, $18, $19::jsonb, $20`;
const baseVals = `$1, $2, $3, $4, $5::jsonb, $6, $7, $8, $9, $10, $11, $12, $13, $14, $15, $16, $17, $18, $19, $20::jsonb, $21`;
const cols = hasMaxStalled ? `${baseCols}, max_stalled` : baseCols;
const vals = hasMaxStalled ? `${baseVals}, $21` : baseVals;
const vals = hasMaxStalled ? `${baseVals}, $22` : baseVals;
const insertSql = opts?.idempotency_key
? `INSERT INTO minion_jobs (${cols})
@@ -388,6 +447,14 @@ export class MinionQueue {
// sane long wall-clock default stamped at submit when the caller didn't
// pass one, so they aren't killed mid-progress by the short null-default.
opts?.timeout_ms ?? defaultTimeoutMsFor(jobName),
// #4145: same three-layer pattern for the lock lease. Explicit input
// is clamped to [5s,1h]; absent → handler map default; NULL row =
// worker-global lockDuration at claim. INSERT-only (see the
// max_stalled footgun note above): an idempotency-key re-submit
// never mutates the first submitter's lease.
opts?.lock_duration_ms != null
? clampLockDurationMs(opts.lock_duration_ms)
: defaultLockDurationMsFor(jobName),
opts?.remove_on_complete ?? false,
opts?.remove_on_fail ?? false,
opts?.idempotency_key ?? null,
@@ -772,11 +839,26 @@ export class MinionQueue {
// Direct (session-mode) pool: claim opens the lock that renewLock then
// heartbeats. Both must live on a connection the transaction-mode pooler
// won't recycle mid-hold, or the lock orphans and the worker wedges.
//
// #4145: lock_duration_ms resolves row → handler map ($6, RAW object —
// same double-encode rule as $5) → worker default ($2), is STAMPED onto
// the row (durable, like timeout_ms), and lock_until derives from the
// same COALESCE (OLD-row semantics: repeat the expression, don't
// reference the assigned column). Both the stamp and lock_until are
// CASE-clamped to the [5s,1h] bound IN SQL (row/map resolution only —
// the worker-default fallback $2 is operator-configured, not row data,
// and tests/short-lived workers legitimately use sub-5s leases): the exposed
// submit surfaces clamp already, but a bypass-written row (direct SQL
// repair, foreign tooling) must not grant a ~24-day lease to a worker
// that crashes before its first renewal (or a 1ms one that thrashes).
const rows = await this.engine.executeRawDirect<Record<string, unknown>>(
`UPDATE minion_jobs SET
status = 'active',
lock_token = $1,
lock_until = now() + ($2::double precision * interval '1 millisecond'),
lock_until = now() + ((CASE WHEN COALESCE(lock_duration_ms, ($6::jsonb ->> name)::int) IS NULL THEN $2
ELSE LEAST(GREATEST(COALESCE(lock_duration_ms, ($6::jsonb ->> name)::int), 5000), 3600000) END)::double precision * interval '1 millisecond'),
lock_duration_ms = CASE WHEN COALESCE(lock_duration_ms, ($6::jsonb ->> name)::int) IS NULL THEN NULL
ELSE LEAST(GREATEST(COALESCE(lock_duration_ms, ($6::jsonb ->> name)::int), 5000), 3600000) END,
timeout_ms = COALESCE(timeout_ms, ($5::jsonb ->> name)::int),
timeout_at = CASE WHEN COALESCE(timeout_ms, ($5::jsonb ->> name)::int) IS NOT NULL
THEN now() + (COALESCE(timeout_ms, ($5::jsonb ->> name)::int)::double precision * interval '1 millisecond')
@@ -792,7 +874,7 @@ export class MinionQueue {
LIMIT 1
)
RETURNING *`,
[lockToken, lockDurationMs, queue, registeredNames, HANDLER_DEFAULT_TIMEOUT_MS]
[lockToken, lockDurationMs, queue, registeredNames, HANDLER_DEFAULT_TIMEOUT_MS, HANDLER_DEFAULT_LOCK_DURATION_MS]
);
return rows.length > 0 ? rowToMinionJob(rows[0]) : null;
}
@@ -959,7 +1041,7 @@ export class MinionQueue {
AND EXTRACT(EPOCH FROM (now() - started_at)) * 1000 >
CASE
WHEN timeout_ms IS NOT NULL THEN timeout_ms * 2
ELSE $1::double precision * 2 * GREATEST(max_stalled, 1)
ELSE COALESCE(lock_duration_ms, $1)::double precision * 2 * GREATEST(max_stalled, 1)
END`,
[lockDurationMs]
);
@@ -982,7 +1064,7 @@ export class MinionQueue {
AND EXTRACT(EPOCH FROM (now() - started_at)) * 1000 >
CASE
WHEN timeout_ms IS NOT NULL THEN timeout_ms * 2
ELSE $1::double precision * 2 * GREATEST(max_stalled, 1)
ELSE COALESCE(lock_duration_ms, $1)::double precision * 2 * GREATEST(max_stalled, 1)
END
FOR UPDATE SKIP LOCKED
)
@@ -1297,9 +1379,15 @@ export class MinionQueue {
/**
* Renew lock (token-fenced). Returns false if token mismatch (job was reclaimed).
*
* `opts.signal` cancels the in-flight UPDATE (postgres.js `.cancel()`) when the
* caller's timeout race gives up on it otherwise the abandoned query holds a
* checked-out pool slot for its full server-side duration (issue #6).
* Cancellation is BEST-EFFORT (#4145 CDX-2/R2-2): pool acquisition and PG
* protocol cancel are asynchronous, and PGLite ignores the signal so
* correctness never rests on it. A late-landing renewal UPDATE is fenced on
* OUR token, meaning it can only extend a lock nobody else has claimed;
* worst case is a stall-requeue delayed by one lease.
*/
async renewLock(
id: number,
@@ -1370,18 +1458,25 @@ export class MinionQueue {
}
/** Detect and handle stalled jobs. Single CTE, no off-by-one. Returns affected jobs. */
async handleStalled(): Promise<{ requeued: MinionJob[]; dead: MinionJob[] }> {
async handleStalled(graceMsOverride?: number): Promise<{ requeued: MinionJob[]; dead: MinionJob[] }> {
// W0 fix-wave (Tier-1 #4): the dead-letter branch previously emitted NO
// child_done and never unblocked aggregator parents — a child that died
// via max-stall stranded its parent in 'waiting-children' forever (the
// exact hang the v0.15 comment says was fixed for timeouts; there was no
// compensating sweep anywhere). Restructured into the parents-first
// discover/lock/kill shape (D5.12) with the shared killJobs() tail.
//
// #4145 (CDX-7): the reclaim predicate carries a grace — see
// resolveStallReclaimGraceMs. Callers (tests) may pass an explicit
// override; the worker sweep resolves from env/default.
const graceMs = graceMsOverride ?? resolveStallReclaimGraceMs();
return this.engine.transaction(async (tx) => {
const candidates = await tx.executeRaw<{ id: number; parent_job_id: number | null; stalled_counter: number; max_stalled: number }>(
`SELECT id, parent_job_id, stalled_counter, max_stalled
FROM minion_jobs
WHERE status = 'active' AND lock_until < now()`
WHERE status = 'active'
AND lock_until < now() - ($1::double precision * interval '1 millisecond')`,
[graceMs]
);
if (candidates.length === 0) return { requeued: [], dead: [] };
const ids = candidates.map(c => c.id);
@@ -1399,12 +1494,13 @@ export class MinionQueue {
WHERE id IN (
SELECT id FROM minion_jobs
WHERE id = ANY($1::bigint[])
AND status = 'active' AND lock_until < now()
AND status = 'active'
AND lock_until < now() - ($2::double precision * interval '1 millisecond')
AND stalled_counter + 1 < max_stalled
FOR UPDATE SKIP LOCKED
)
RETURNING *`,
[ids]
[ids, graceMs]
);
const deadRows = await tx.executeRaw<Record<string, unknown>>(
`UPDATE minion_jobs SET
@@ -1415,12 +1511,13 @@ export class MinionQueue {
WHERE id IN (
SELECT id FROM minion_jobs
WHERE id = ANY($1::bigint[])
AND status = 'active' AND lock_until < now()
AND status = 'active'
AND lock_until < now() - ($2::double precision * interval '1 millisecond')
AND stalled_counter + 1 >= max_stalled
FOR UPDATE SKIP LOCKED
)
RETURNING *`,
[ids]
[ids, graceMs]
);
// THE FIX: stall-death now notifies + unblocks parents like every
// other terminal kill. Outcome 'dead' (not 'timeout') so consumers can
+5
View File
@@ -71,6 +71,8 @@ export interface MinionJob {
max_children: number | null;
timeout_ms: number | null;
timeout_at: Date | null;
/** Per-job lock lease (ms, #4145). NULL = worker-global lockDuration default. */
lock_duration_ms: number | null;
remove_on_complete: boolean;
remove_on_fail: boolean;
idempotency_key: string | null;
@@ -128,6 +130,8 @@ export interface MinionJobInput {
max_children?: number;
/** Wall-clock per-job deadline in ms. Set on claim → timeout_at. Terminal on expire (no retry). */
timeout_ms?: number;
/** Per-job lock lease in ms (#4145). Clamped to [5s,1h]; NULL/undefined → handler map, then worker default. INSERT-only: an idempotency-key re-submit never mutates the first submitter's lease. */
lock_duration_ms?: number;
/** DELETE row on successful completion (after token rollup + child_done insert). */
remove_on_complete?: boolean;
/** DELETE row on terminal failure (after parent failure hook). */
@@ -435,6 +439,7 @@ export function rowToMinionJob(row: Record<string, unknown>): MinionJob {
depth: (row.depth as number) ?? 0,
max_children: (row.max_children as number) ?? null,
timeout_ms: (row.timeout_ms as number) ?? null,
lock_duration_ms: (row.lock_duration_ms as number) ?? null,
timeout_at: row.timeout_at ? new Date(row.timeout_at as string) : null,
remove_on_complete: row.remove_on_complete === true,
remove_on_fail: row.remove_on_fail === true,
+164 -16
View File
@@ -30,9 +30,12 @@ import { logLeasePressure } from './lease-pressure-audit.ts';
import {
runLockRenewalTick,
resolveLockRenewalKnobs,
renewalIntervalFor,
type LockRenewalDeps,
type LockRenewalState,
type TickResult,
} from './lock-renewal-tick.ts';
import { clampLockDurationMs } from './handler-timeouts.ts';
import {
runDbProbe,
getConnectionRouting,
@@ -48,6 +51,8 @@ import {
ChildNotClaimedError,
} from './child-job-runner.ts';
import { lockRenewalAudit } from '../audit/lock-renewal-audit.ts';
import { loadavg, cpus } from 'os';
import { monitorEventLoopDelay } from 'perf_hooks';
import { isRetryableConnError } from '../retry-matcher.ts';
import { reconnectAfterConnectionError as reconnectEngineAfterConnError } from './reconnect.ts';
@@ -230,6 +235,23 @@ export class MinionWorker extends EventEmitter {
private opts: Required<MinionWorkerOpts>;
/**
* Event-loop-delay histogram (CDX-1/R2-9, issue #4145): the DIRECT
* measurement of local starvation, sampled at eviction time and RESET
* on every successful lock renewal so a sample attributes to the
* window since the last success not process lifetime. The histogram
* is deliberately WORKER-global, not per-job: the event loop is one
* shared resource, and ANY job's successful renewal proves the loop was
* healthy enough to process a round-trip at that moment a legitimate
* truncation of the starvation window even for a sibling job that
* evicts moments later (its per-job discriminator is tick lateness,
* which IS per-job). Null when the runtime doesn't ship
* `monitorEventLoopDelay` (fail-open: eviction logs omit eld fields).
*/
private eldHistogram: ReturnType<typeof monitorEventLoopDelay> | null = null;
/** Core count cached once — pairs with raw loadavg in eviction telemetry. */
private readonly cpuCores: number;
constructor(
private engine: BrainEngine,
opts?: MinionWorkerOpts & MinionQueueOpts,
@@ -239,6 +261,17 @@ export class MinionWorker extends EventEmitter {
maxSpawnDepth: opts?.maxSpawnDepth,
maxAttachmentBytes: opts?.maxAttachmentBytes,
});
let cores = 0;
try { cores = cpus().length; } catch { /* telemetry best-effort */ }
this.cpuCores = cores;
try {
if (typeof monitorEventLoopDelay === 'function') {
// Created DISABLED: start() enables and stop() disables, so a
// constructed-but-never-started worker (setup failures, probe
// instances) never holds a native sampling timer.
this.eldHistogram = monitorEventLoopDelay({ resolution: 20 });
}
} catch { this.eldHistogram = null; /* fail-open */ }
this.opts = {
queue: opts?.queue ?? 'default',
concurrency: opts?.concurrency ?? 1,
@@ -346,6 +379,10 @@ export class MinionWorker extends EventEmitter {
await this.queue.ensureSchema();
this.running = true;
// R2-9 lifecycle: (re-)enable the event-loop-delay histogram for this
// run; stop() disables it so embedding hosts / test suites that cycle
// start()/stop() don't leak a ~50Hz native sampling timer per instance.
try { this.eldHistogram?.enable(); } catch { /* fail-open */ }
// Graceful shutdown. Fires shutdownAbort so handlers subscribed to
// `ctx.shutdownSignal` (currently: shell handler) can run their own cleanup
@@ -805,6 +842,7 @@ export class MinionWorker extends EventEmitter {
/** Stop the worker gracefully. */
stop(): void {
this.running = false;
try { this.eldHistogram?.disable(); } catch { /* fail-open */ }
}
/**
@@ -923,6 +961,40 @@ export class MinionWorker extends EventEmitter {
* `failJob` throwing during the same DB outage) can't propagate to
* the process-level handler and crash the daemon.
*/
/**
* One formatter for the classified abort telemetry (R2-7) so the
* should_abort warn and the 30s-later grace-evict line can never drift
* apart (they briefly did: `load1:` vs `load1_at_abort:`).
*/
private formatAbortMeta(meta: Extract<TickResult, { kind: 'lock_lost' | 'should_abort' }>): string {
if (meta.kind === 'lock_lost') {
return `cause: ${meta.cause}, via: ${meta.via}`;
}
return `cause: ${meta.cause}, since_last_success_ms: ${Math.round(meta.sinceLastSuccessMs)}, ` +
`tick_lateness_ms: ${Math.round(meta.latenessMs)}, overlap_skips: ${meta.overlapSkips}` +
`${meta.load1 !== undefined ? `, load1: ${meta.load1.toFixed(2)}/${meta.cores} cores` : ''}`;
}
/**
* Event-loop-delay sample for eviction log lines (CDX-1). The histogram
* resets on every successful renewal (R2-9 via deps.onRenewalSuccess),
* so these numbers attribute to the window since the last success
* i.e. exactly the window in which renewal was failing. nsms. Empty
* string when the runtime lacks the histogram or sampling throws
* (fail-open: never let telemetry break the eviction path).
*/
private formatEvictionTelemetry(): string {
const h = this.eldHistogram;
if (h === null) return '';
try {
const p99Ms = Math.round(h.percentile(99) / 1e6);
const maxMs = Math.round(h.max / 1e6);
return ` [event_loop_delay since last renewal: p99 ${p99Ms}ms, max ${maxMs}ms]`;
} catch {
return '';
}
}
private launchJob(job: MinionJob, lockToken: string): void {
const abort = new AbortController();
@@ -930,18 +1002,46 @@ export class MinionWorker extends EventEmitter {
let cancelled = false;
// --- re-entrancy guard for overlapping ticks during PgBouncer stalls ---
let tickInFlight = false;
// --- R2-7: the tick's final result, stashed at abort time so the ---
// --- grace-evict log (which fires 30s LATER) can report cause/ ---
// --- lateness/load instead of just the Error string. ---
let abortMeta: Extract<TickResult, { kind: 'lock_lost' | 'should_abort' }> | null = null;
// R2-4: ALL elapsed-time arithmetic in the renewal state machine runs
// on a monotonic clock — a wall-clock jump (NTP step, DST bug) must
// never evict or indefinitely defer. Date.now stays only in log/audit
// timestamps (the audit writer stamps its own `ts`).
const monotonicNow = () => performance.now();
// --- D3: pure-function lock renewal ---
const knobs = resolveLockRenewalKnobs(process.env, this.opts.lockDuration);
// #4145: the EFFECTIVE lease is per-job (claim stamped it from the
// handler map / explicit submit; NULL = worker default). The cadence
// clamps to 60s so a 300s lease renews 5x per window (matching the
// cycle refresher's multiple-chances philosophy) instead of the bare
// lease/2 = every 150s; leases ≤120s keep the legacy /2 exactly.
// Defense-in-depth: re-clamp the row value at consumption. The exposed
// submit surfaces already clamp, but the claim COALESCE trusts the row
// and the DB CHECK only enforces > 0 — a writer that bypasses add()
// (direct SQL repair, foreign tooling) could otherwise stamp a 1ms
// lease (renewal-storm setInterval) or a ~25-day one (weeks-long
// dead-worker pin).
const effectiveLockMs = job.lock_duration_ms != null
? clampLockDurationMs(job.lock_duration_ms)
: this.opts.lockDuration;
const renewalIntervalMs = renewalIntervalFor(effectiveLockMs);
const knobs = resolveLockRenewalKnobs(process.env, effectiveLockMs, renewalIntervalMs);
const renewalState: LockRenewalState = {
jobId: job.id,
jobName: job.name,
lockToken,
lockDurationMs: this.opts.lockDuration,
lockDurationMs: effectiveLockMs,
knobs,
lastSuccessfulRenewalAt: Date.now(),
lastSuccessfulRenewalAt: monotonicNow(),
consecutiveFailures: 0,
cancelled: () => cancelled,
intervalMs: renewalIntervalMs,
lastTickFiredAt: monotonicNow(),
overlapSkips: 0,
};
// issue #1678 (Codex #2): hand the tick a bounded reconnect-once hook when
// the engine owns a pool that a transaction-mode pooler can reap. Postgres
@@ -951,15 +1051,30 @@ export class MinionWorker extends EventEmitter {
const renewalDeps: LockRenewalDeps = {
renewLock: (id, tok, dur, opts) => this.queue.renewLock(id, tok, dur, opts),
audit: lockRenewalAudit,
now: Date.now,
// R2-4: monotonic — see monotonicNow above.
now: monotonicNow,
setTimeout: (cb, ms) => globalThis.setTimeout(cb, ms),
// Forward the tick's classified error (CODEX impl review #2) so a pooler
// reap during lock renewal is audited as reap_detected, not reconnect_other.
...(engineReconnect ? { reconnect: (ctx?: { error?: unknown }) => engineReconnect.call(this.engine, ctx) } : {}),
// Issue #4145 telemetry: raw loadavg + cached cores (CEO-F2: the tick
// try/catches every call), and the R2-9 histogram reset-on-success.
loadSnapshot: () => ({ load1: loadavg()[0], cores: this.cpuCores }),
onRenewalSuccess: () => { try { this.eldHistogram?.reset(); } catch { /* fail-open */ } },
};
const lockTimer = setInterval(() => {
if (tickInFlight) return;
if (tickInFlight) {
// Overlap skip (CDX-13): a prior tick is still awaiting its
// renewal call. Count it — this is NOT a missed interval (those
// coalesce and show up as tick LATENESS instead) — and ADVANCE the
// lateness baseline: this callback fired on schedule, so the next
// executed tick must not book the skipped window as event-loop
// lateness (that would misclassify a slow DB call as starvation).
renewalState.overlapSkips += 1;
renewalState.lastTickFiredAt = monotonicNow();
return;
}
tickInFlight = true;
void runLockRenewalTick(renewalDeps, renewalState)
.then((result) => {
@@ -970,13 +1085,24 @@ export class MinionWorker extends EventEmitter {
return;
case 'lock_lost':
if (!abort.signal.aborted) {
console.warn(`Lock lost for job ${job.id}, aborting execution`);
abortMeta = result;
console.warn(
`Lock lost for job ${job.id}, aborting execution ` +
`(${this.formatAbortMeta(result)}${this.formatEvictionTelemetry()})`,
);
clearInterval(lockTimer);
abort.abort(new Error('lock-lost'));
}
return;
case 'should_abort':
if (!abort.signal.aborted) {
abortMeta = result;
// Issue #4145 request 3: the one line that saves the 8h of
// forensics — WHY renewal failed + was the loop starved.
console.warn(
`Lock renewal failed for job ${job.id} (${job.name}); aborting ` +
`(${this.formatAbortMeta(result)}${this.formatEvictionTelemetry()})`,
);
clearInterval(lockTimer);
abort.abort(new Error(result.reason));
}
@@ -996,7 +1122,7 @@ export class MinionWorker extends EventEmitter {
.finally(() => {
tickInFlight = false;
});
}, this.opts.lockDuration / 2);
}, renewalIntervalMs);
// --- D8b: universal grace-eviction timer ---
// Fires for ANY abort reason (not just job.timeout_ms). Without
@@ -1012,18 +1138,35 @@ export class MinionWorker extends EventEmitter {
const reason = abort.signal.reason instanceof Error
? abort.signal.reason.message
: String(abort.signal.reason);
// R2-7: abortMeta carries the tick's classified cause + starvation
// telemetry captured AT abort time — the Error string alone would
// make this line read like an orphan leak (the #4145 forensics trap).
const meta = abortMeta === null ? '' : ` (${this.formatAbortMeta(abortMeta)})`;
console.warn(
`Job ${job.id} (${job.name}) did not exit within 30s of abort (reason: ${reason}). ` +
`Job ${job.id} (${job.name}) did not exit within 30s of abort (reason: ${reason}).${meta} ` +
`Force-evicting from inFlight to unblock worker. ` +
`The handler is still running but the worker will claim new jobs.`
`The handler is still running but the worker will claim new jobs.` +
this.formatEvictionTelemetry()
);
clearInterval(lockTimer);
this.inFlight.delete(job.id);
// D8a: don't failJob on infrastructure aborts (stall detector
// reclaims after lock expiry). Isolation mode: also skip — the
// group SIGKILL already fired and executeJob's own recording
// follows; a competing evict failJob('dead') could dead-letter a
// job with attempts remaining (adversarial-review P3).
// R2-1 generation-safety: delete ONLY our own execution's entry.
// After a force-evict, this job id can be requeued and re-claimed
// by THIS worker while the old handler is still alive — an
// unconditional delete-by-id would then remove the NEW
// execution's entry (concurrency undercount, lost tracking).
// The lockToken is minted per claim, so it is the generation.
if (this.inFlight.get(job.id)?.lockToken === lockToken) {
this.inFlight.delete(job.id);
}
// D8a: don't failJob if the abort was infrastructure. The stall
// detector will reclaim the row cleanly: a lock-renewal abort
// now fires only after the at-deadline VERIFY either returned
// fenced-false (row already reclaimed) or stayed unreachable
// past hardEvictMs (lease long expired) — see #4145
// verify-before-evict in lock-renewal-tick.ts. Isolation mode:
// also skip — the group SIGKILL already fired and executeJob's
// own recording follows; a competing evict failJob('dead') could
// dead-letter a job with attempts remaining (adversarial-review P3).
if (!INFRASTRUCTURE_ABORT_REASONS.has(reason) && this.opts.jobIsolation !== 'process') {
this.queue.failJob(
job.id,
@@ -1066,7 +1209,12 @@ export class MinionWorker extends EventEmitter {
clearInterval(lockTimer);
if (timeoutTimer) clearTimeout(timeoutTimer);
if (graceTimer) clearTimeout(graceTimer);
this.inFlight.delete(job.id);
// R2-1 generation-safety: a force-evicted execution's finally can
// fire long after the same job id was re-claimed by this worker.
// Only delete the entry if it is still OURS (token = generation).
if (this.inFlight.get(job.id)?.lockToken === lockToken) {
this.inFlight.delete(job.id);
}
this.jobsCompleted += 1;
this.checkMemoryLimit('post-job');
})
+5
View File
@@ -3588,6 +3588,7 @@ const submit_job: Operation = {
max_attempts: { type: 'number', description: 'Max retry attempts (default: 3)' },
delay: { type: 'number', description: 'Delay in ms before eligible' },
timeout_ms: { type: 'number', description: 'Per-job wall-clock timeout in ms; aborted job goes to dead' },
lock_duration_ms: { type: 'number', description: 'Per-job lock lease in ms (#4145). Out-of-range values are clamped to [5000, 3600000] — remote writers cannot pin an immortal lock. Omit to use the handler-type default (300s for long LLM handlers) or the worker default (30s).' },
},
mutating: true,
scope: 'admin',
@@ -3635,6 +3636,10 @@ const submit_job: Operation = {
max_attempts: (p.max_attempts as number) || 3,
delay: (p.delay as number) || undefined,
timeout_ms: (p.timeout_ms as number) || undefined,
// #4145 [CEO-F7/R2-6]: range enforcement lives in queue.add's
// clampLockDurationMs (ParamDef has no min/max support; wrong TYPE is
// rejected by the shared number validation upstream of this handler).
lock_duration_ms: (p.lock_duration_ms as number) || undefined,
}, trusted);
// v0.35.8.0: submit_job audit-log parity with the CLI path (codex F-CDX-4).
+97 -34
View File
@@ -117,35 +117,85 @@ const PGLITE_EDGE_BATCH_MAX_BIND_PARAMS = 30_000;
// silently fall through to a normal initSchema (snapshot is just an
// optimization, never authoritative).
let _snapshotWarnLogged = false;
// Per-process memo. MIGRATIONS + PGLITE_SCHEMA_SQL are static for the life of
// the process, so the schema hash is too; the version file and the ~42MB tar
// are read once per (path, process) instead of once per engine construction
// (a full suite constructs 600+ engines — the un-memoized loader re-read the
// tar and re-hashed 131 migration handler sources every time, ~84MB of
// transient allocation per call). A null entry means the path is terminally
// unusable this process (missing/stale/torn) — no retry per construction.
// The dims/model shape gate is deliberately NOT memoized: tests reconfigure
// the gateway mid-process (zembed/1280) and a mismatched engine must fall
// back to cold init even when an earlier engine loaded this same snapshot.
// Accepted limitation: a snapshot file rewritten mid-process is not observed;
// the only writer (build-pglite-snapshot.ts) runs before test fan-out.
let _snapshotSchemaHashMemo: string | null = null;
// blob stays null until the FIRST caller whose shape gate passes — a process
// whose gateway shape never matches the snapshot (the zembed/1280 test
// files) never pays the 42MB tar read at all.
const _snapshotFileMemo = new Map<string, { versionLines: string[]; blob: Blob | null } | null>();
let _snapshotTarReads = 0;
export function __snapshotMemoStatsForTests(): { tarReads: number; memoEntries: number } {
return { tarReads: _snapshotTarReads, memoEntries: _snapshotFileMemo.size };
}
export function __resetSnapshotMemoForTests(): void {
_snapshotSchemaHashMemo = null;
_snapshotFileMemo.clear();
_snapshotTarReads = 0;
_snapshotWarnLogged = false;
}
export function tryLoadSnapshot(snapshotPath: string): Blob | null {
try {
// Lazy require so production builds without these imports don't crash.
// eslint-disable-next-line @typescript-eslint/no-require-imports
const fs = require('node:fs') as typeof import('node:fs'); // engine-dynamic-import-ok
const crypto = require('node:crypto') as typeof import('node:crypto'); // engine-dynamic-import-ok
const { MIGRATIONS } = require('./migrate.ts') as typeof import('./migrate.ts'); // engine-dynamic-import-ok
const { PGLITE_SCHEMA_SQL } = require('./pglite-schema.ts') as typeof import('./pglite-schema.ts'); // engine-dynamic-import-ok
let entry = _snapshotFileMemo.get(snapshotPath);
if (entry === null) return null; // terminally unusable this process
if (entry === undefined) {
// First touch of this path in this process — do the file work once.
// Lazy require so production builds without these imports don't crash.
// eslint-disable-next-line @typescript-eslint/no-require-imports
const fs = require('node:fs') as typeof import('node:fs'); // engine-dynamic-import-ok
const crypto = require('node:crypto') as typeof import('node:crypto'); // engine-dynamic-import-ok
const { MIGRATIONS } = require('./migrate.ts') as typeof import('./migrate.ts'); // engine-dynamic-import-ok
const { PGLITE_SCHEMA_SQL } = require('./pglite-schema.ts') as typeof import('./pglite-schema.ts'); // engine-dynamic-import-ok
if (!fs.existsSync(snapshotPath)) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] GBRAIN_PGLITE_SNAPSHOT set but file missing: ${snapshotPath} — using normal init.`);
_snapshotWarnLogged = true;
if (!fs.existsSync(snapshotPath)) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] GBRAIN_PGLITE_SNAPSHOT set but file missing: ${snapshotPath} — using normal init.`);
_snapshotWarnLogged = true;
}
_snapshotFileMemo.set(snapshotPath, null);
return null;
}
return null;
}
const versionPath = snapshotPath.replace(/\.tar(?:\.gz)?$/, '.version');
if (!fs.existsSync(versionPath)) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] snapshot version file missing: ${versionPath} — using normal init.`);
_snapshotWarnLogged = true;
const versionPath = snapshotPath.replace(/\.tar(?:\.gz)?$/, '.version');
if (!fs.existsSync(versionPath)) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] snapshot version file missing: ${versionPath} — using normal init.`);
_snapshotWarnLogged = true;
}
_snapshotFileMemo.set(snapshotPath, null);
return null;
}
return null;
if (_snapshotSchemaHashMemo === null) {
_snapshotSchemaHashMemo = computeSnapshotSchemaHash(MIGRATIONS, PGLITE_SCHEMA_SQL, crypto);
}
const versionLines = fs.readFileSync(versionPath, 'utf8').trim().split('\n');
if (_snapshotSchemaHashMemo !== (versionLines[0] ?? '')) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] snapshot stale (schema hash mismatch) — using normal init. Rebuild with: bun run build:pglite-snapshot`);
_snapshotWarnLogged = true;
}
_snapshotFileMemo.set(snapshotPath, null);
return null;
}
entry = { versionLines, blob: null };
_snapshotFileMemo.set(snapshotPath, entry);
}
const expectedHash = computeSnapshotSchemaHash(MIGRATIONS, PGLITE_SCHEMA_SQL, crypto);
const versionLines = fs.readFileSync(versionPath, 'utf8').trim().split('\n');
const actualHash = versionLines[0] ?? '';
// W0 fix-wave: the version file's dims=/model= lines record the embedding
// shape the snapshot was BAKED with. A snapshot whose vector(dims) columns
@@ -154,6 +204,8 @@ export function tryLoadSnapshot(snapshotPath: string): Blob | null {
// fixture went default-on). Resolve our would-be shape through the same
// gateway-with-default fallback initSchema uses and refuse a mismatch.
// Version files without the shape lines (pre-W0) are treated as stale.
// Re-evaluated on EVERY call against the CURRENT gateway config — never
// memoized (see memo comment above).
let wantDims: number | string = DEFAULT_EMBEDDING_DIMENSIONS;
let wantModel: string = DEFAULT_EMBEDDING_MODEL;
try {
@@ -161,25 +213,30 @@ export function tryLoadSnapshot(snapshotPath: string): Blob | null {
wantDims = gw.getEmbeddingDimensions();
wantModel = gw.getEmbeddingModel();
} catch { /* gateway not configured — defaults, same as initSchema */ }
const shapeOk = versionLines[1] === `dims=${wantDims}` && versionLines[2] === `model=${wantModel}`;
const shapeOk = entry.versionLines[1] === `dims=${wantDims}` && entry.versionLines[2] === `model=${wantModel}`;
if (!shapeOk) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] snapshot embedding shape mismatch (want dims=${wantDims} model=${wantModel}, have ${versionLines[1] ?? 'none'} ${versionLines[2] ?? ''}) — using normal init. Rebuild with: bun run build:pglite-snapshot`);
console.warn(`[pglite] snapshot embedding shape mismatch (want dims=${wantDims} model=${wantModel}, have ${entry.versionLines[1] ?? 'none'} ${entry.versionLines[2] ?? ''}) — using normal init. Rebuild with: bun run build:pglite-snapshot`);
_snapshotWarnLogged = true;
}
return null;
}
if (expectedHash !== actualHash) {
if (!_snapshotWarnLogged) {
// eslint-disable-next-line no-console
console.warn(`[pglite] snapshot stale (schema hash mismatch) — using normal init. Rebuild with: bun run build:pglite-snapshot`);
_snapshotWarnLogged = true;
if (entry.blob === null) {
// Tar read deferred until the first shape-matching caller (see memo
// comment above). A torn/unreadable tar is terminal for the process.
try {
// eslint-disable-next-line @typescript-eslint/no-require-imports
const fs = require('node:fs') as typeof import('node:fs'); // engine-dynamic-import-ok
const buf = fs.readFileSync(snapshotPath);
_snapshotTarReads += 1;
entry.blob = new Blob([new Uint8Array(buf.buffer as ArrayBuffer, buf.byteOffset, buf.byteLength)]);
} catch {
_snapshotFileMemo.set(snapshotPath, null);
return null;
}
return null;
}
const buf = fs.readFileSync(snapshotPath);
return new Blob([buf]);
return entry.blob;
} catch {
// Any failure -> fall through to normal init. Never block tests.
return null;
@@ -3856,8 +3913,14 @@ export class PGLiteEngine implements BrainEngine {
COUNT(DISTINCT n.last_link_type) AS edge_count,
array_agg(DISTINCT n.last_link_type)
FILTER (WHERE n.last_link_type IS NOT NULL) AS via_link_types,
-- Final path tie-break (lexicographic) makes the pick deterministic
-- when a node is reachable at the same depth from multiple seeds;
-- without it the winner is plan/heap-order dependent and the two
-- engines (or two runs) can disagree. Relational retrieval is
-- documented deterministic; keep in lockstep with postgres-engine.ts.
(array_agg(array_to_string(n.path, chr(9))
ORDER BY n.depth ASC, array_length(n.path, 1) ASC))[1] AS path_str,
ORDER BY n.depth ASC, array_length(n.path, 1) ASC,
array_to_string(n.path, chr(9)) ASC))[1] AS path_str,
(SELECT cc.id FROM content_chunks cc
WHERE cc.page_id = n.id ORDER BY cc.chunk_index ASC LIMIT 1) AS canonical_chunk_id
FROM walk n
+3 -1
View File
@@ -462,6 +462,7 @@ CREATE TABLE IF NOT EXISTS minion_jobs (
depth INTEGER NOT NULL DEFAULT 0,
max_children INTEGER,
timeout_ms INTEGER,
lock_duration_ms INTEGER,
timeout_at TIMESTAMPTZ,
remove_on_complete BOOLEAN NOT NULL DEFAULT FALSE,
remove_on_fail BOOLEAN NOT NULL DEFAULT FALSE,
@@ -482,7 +483,8 @@ CREATE TABLE IF NOT EXISTS minion_jobs (
CONSTRAINT chk_nonnegative CHECK (attempts_made >= 0 AND attempts_started >= 0 AND stalled_counter >= 0 AND max_attempts >= 1 AND max_stalled >= 0),
CONSTRAINT chk_depth_nonnegative CHECK (depth >= 0),
CONSTRAINT chk_max_children_positive CHECK (max_children IS NULL OR max_children > 0),
CONSTRAINT chk_timeout_positive CHECK (timeout_ms IS NULL OR timeout_ms > 0)
CONSTRAINT chk_timeout_positive CHECK (timeout_ms IS NULL OR timeout_ms > 0),
CONSTRAINT chk_lock_duration_positive CHECK (lock_duration_ms IS NULL OR (lock_duration_ms >= 5000 AND lock_duration_ms <= 3600000))
);
CREATE INDEX IF NOT EXISTS idx_minion_jobs_claim ON minion_jobs (queue, priority ASC, created_at ASC) WHERE status = 'waiting';
+23 -11
View File
@@ -3785,8 +3785,14 @@ export class PostgresEngine implements BrainEngine {
COUNT(DISTINCT n.last_link_type) AS edge_count,
array_agg(DISTINCT n.last_link_type)
FILTER (WHERE n.last_link_type IS NOT NULL) AS via_link_types,
-- Final path tie-break (lexicographic) makes the pick deterministic
-- when a node is reachable at the same depth from multiple seeds;
-- without it the winner is plan/heap-order dependent and the two
-- engines (or two runs) can disagree. Relational retrieval is
-- documented deterministic; keep in lockstep with pglite-engine.ts.
(array_agg(array_to_string(n.path, chr(9))
ORDER BY n.depth ASC, array_length(n.path, 1) ASC))[1] AS path_str,
ORDER BY n.depth ASC, array_length(n.path, 1) ASC,
array_to_string(n.path, chr(9)) ASC))[1] AS path_str,
(SELECT cc.id FROM content_chunks cc
WHERE cc.page_id = n.id ORDER BY cc.chunk_index ASC LIMIT 1) AS canonical_chunk_id
FROM walk n
@@ -6242,18 +6248,17 @@ export class PostgresEngine implements BrainEngine {
params?: unknown[],
opts?: { signal?: AbortSignal },
): Promise<T[]> {
// #4145 R2-2 preflight: an ALREADY-aborted signal must short-circuit
// BEFORE the query is dispatched — the previous order created the
// pending query first and cancelled it after, which still burned a
// round-trip (and on a saturated pool, a slot). Cancellation remains
// BEST-EFFORT overall (PG protocol cancel is async); callers that need
// correctness must rely on their own fencing, not this signal.
if (opts?.signal?.aborted) {
throw new DOMException('aborted', 'AbortError');
}
const pending = conn.unsafe(sql, params as Parameters<typeof conn.unsafe>[1]);
if (opts?.signal) {
if (opts.signal.aborted) {
// .cancel() is fire-and-forget; the awaited query rejects with the
// postgres "query was cancelled" error which the caller catches.
try {
(pending as unknown as { cancel?: () => void }).cancel?.();
} catch {
// best-effort
}
throw new DOMException('aborted', 'AbortError');
}
const onAbort = () => {
try {
(pending as unknown as { cancel?: () => void }).cancel?.();
@@ -6312,6 +6317,13 @@ export class PostgresEngine implements BrainEngine {
params?: unknown[],
opts?: { signal?: AbortSignal },
): Promise<T[]> {
// #4145 R2-2: observe the signal BEFORE (potentially slow) direct-pool
// acquisition — a caller whose timeout already fired must not queue for
// a pool slot just to be cancelled afterwards. runUnsafe re-checks after
// acquisition.
if (opts?.signal?.aborted) {
throw new DOMException('aborted', 'AbortError');
}
// Inside an open transaction, _sql is the reserved tx connection (set via
// defineProperty in transaction()); never reroute off it.
const inTransaction = this._sql !== null && this.connectionManager?.peekReadPool() !== this._sql;
+3 -1
View File
@@ -926,6 +926,7 @@ CREATE TABLE IF NOT EXISTS minion_jobs (
depth INTEGER NOT NULL DEFAULT 0,
max_children INTEGER,
timeout_ms INTEGER,
lock_duration_ms INTEGER,
timeout_at TIMESTAMPTZ,
remove_on_complete BOOLEAN NOT NULL DEFAULT FALSE,
remove_on_fail BOOLEAN NOT NULL DEFAULT FALSE,
@@ -942,7 +943,8 @@ CREATE TABLE IF NOT EXISTS minion_jobs (
CONSTRAINT chk_nonnegative CHECK (attempts_made >= 0 AND attempts_started >= 0 AND stalled_counter >= 0 AND max_attempts >= 1 AND max_stalled >= 0),
CONSTRAINT chk_depth_nonnegative CHECK (depth >= 0),
CONSTRAINT chk_max_children_positive CHECK (max_children IS NULL OR max_children > 0),
CONSTRAINT chk_timeout_positive CHECK (timeout_ms IS NULL OR timeout_ms > 0)
CONSTRAINT chk_timeout_positive CHECK (timeout_ms IS NULL OR timeout_ms > 0),
CONSTRAINT chk_lock_duration_positive CHECK (lock_duration_ms IS NULL OR (lock_duration_ms >= 5000 AND lock_duration_ms <= 3600000))
);
CREATE INDEX IF NOT EXISTS idx_minion_jobs_claim ON minion_jobs (queue, priority ASC, created_at ASC) WHERE status = 'waiting';
+29 -3
View File
@@ -812,6 +812,21 @@ export interface HybridSearchOpts extends SearchOpts {
*/
_queryEmbedDeadline?: QueryEmbedDeadline;
/**
* Hermetic eval canaries/CI non-semantic embeddings. When set, the query
* embedding for the TEXT vector arm comes from this function (e.g. qrels
* basis vectors) INSTEAD of the gateway's query-embed path, and the
* no-embedding-provider keyword-only short-circuit is bypassed so the
* vector arm runs with no provider key configured at all. Never set on
* production paths; when absent, behavior is byte-for-byte unchanged.
*
* Cache note: bare `hybridSearch` neither reads nor writes the semantic
* query cache by construction both the lookup and the store live only in
* `hybridSearchCached` so a deterministic-embedding eval run through this
* seam cannot poison `query_cache` for production queries.
*/
queryEmbedFn?: (text: string) => Float32Array | Promise<Float32Array>;
/**
* INTERNAL cache-consult outcome threaded from `hybridSearchCached` into
* the inner `hybridSearch` so the ONE telemetry record per search (emitted
@@ -1264,7 +1279,10 @@ export async function hybridSearch(
earlyModality === 'both' ||
mayEscalateToMultimodal) &&
isAvailable('embedding', multimodalProviderProbe);
if (!isAvailable('embedding', providerProbe) && !willTryMultimodal) {
// Hermetic eval canaries/CI: a caller-supplied queryEmbedFn produces the
// vector-arm query embedding without the gateway, so provider
// availability is irrelevant — skip the keyword-only short-circuit.
if (!opts?.queryEmbedFn && !isAvailable('embedding', providerProbe) && !willTryMultimodal) {
// v0.43 — fuse the relational arm with keyword so typed-edge answers
// survive on the no-embedding-provider path (the relational win is most
// valuable exactly when vector is unavailable). The title arm fuses here
@@ -1502,12 +1520,20 @@ export async function hybridSearch(
// share one ~6s budget); direct callers get a fresh deadline. On timeout
// the embed rejects → salvage below (or keyword-only when all reject).
const embedDl = opts?._queryEmbedDeadline ?? makeQueryEmbedDeadline();
// Hermetic eval canaries/CI: queryEmbedFn (non-semantic deterministic
// embeddings) replaces the gateway query-embed for the text vector arm.
// No deadline needed — it's a synchronous-ish local computation with no
// network. Absent queryEmbedFn, the bounded gateway path is unchanged.
const embedOneQuery = (q: string): Promise<Float32Array> =>
opts?.queryEmbedFn
? Promise.resolve(opts.queryEmbedFn(q))
: embedQueryBounded(q, embedOpts, embedDl);
if (!searchSalvageEnabled()) {
// ENG-7 kill switch (GBRAIN_SEARCH_SALVAGE=off): pre-wave
// all-or-nothing fan-outs — one variant's failure abandons every
// embedding and falls back to keyword-only.
try {
const embeddings = await Promise.all(queries.map(q => embedQueryBounded(q, embedOpts, embedDl)));
const embeddings = await Promise.all(queries.map(q => embedOneQuery(q)));
queryEmbedding = embeddings[0];
const textLists = await Promise.all(
embeddings.map(emb => engine.searchVector(emb, searchOpts)),
@@ -1537,7 +1563,7 @@ export async function hybridSearch(
// WP2/T3 (ENG-15) salvage fan-outs: allSettled on BOTH the embed
// fan-out and the searchVector fan-out so one variant's failure no
// longer abandons the survivors (the query-vs-search asymmetry fix).
const settled = await Promise.allSettled(queries.map(q => embedQueryBounded(q, embedOpts, embedDl)));
const settled = await Promise.allSettled(queries.map(q => embedOneQuery(q)));
const okEmbeds: Float32Array[] = [];
const embedFailures: unknown[] = [];
for (const s of settled) {
+83
View File
@@ -0,0 +1,83 @@
/**
* Deterministic basis-vector embeddings for hermetic eval canaries/CI.
*
* NON-SEMANTIC embeddings: each query embeds as a unit basis vector at a
* fixed dimension, so retrieval through the full hybrid pipeline (vector +
* keyword/title/alias arms + RRF) is exactly reproducible with no API keys,
* no network, and no provider drift. Used by the qrels correctness gate's
* deterministic embedder path and the retrieval canary runner
* (scripts/run-eval-canary.ts). Mirrors the basis-vector convention in
* test/eval-replay-gate.test.ts and test/fixtures/eval-baselines/
* qrels-search.json (each fixture query carries an `embedding_dim`).
*/
/** Unit basis vector with 1.0 at `idx % dim` and 0.0 elsewhere. */
export function basisEmbedding(idx: number, dim = 1536): Float32Array {
const emb = new Float32Array(dim);
emb[idx % dim] = 1.0;
return emb;
}
/**
* 32-bit FNV-1a hash. Deterministic, dependency-free; used only to derive a
* stable fallback basis dimension for query texts not present in the qrels
* fixture.
*/
export function fnv1a(text: string): number {
let h = 0x811c9dc5;
for (let i = 0; i < text.length; i++) {
h ^= text.charCodeAt(i);
h = Math.imul(h, 0x01000193);
}
return h | 0;
}
export interface LegacyQrelsQuery {
query_id: string;
query: string;
embedding_dim: number;
relevant_slugs: string[];
first_relevant_slug: string;
}
/**
* Parse the legacy qrels fixture shape
* ({queries: [{query, embedding_dim, relevant_slugs, first_relevant_slug}]}).
* THE single parser for this shape the canary runner and the embedder
* builder both consume it. Throws on malformed JSON or a missing `queries`
* array (callers surface a usage error; the gate's own qrels parser reports
* shape problems in detail).
*/
export function parseLegacyQrels(raw: string): LegacyQrelsQuery[] {
const parsed = JSON.parse(raw) as { queries?: unknown };
if (!Array.isArray(parsed.queries)) {
throw new Error('qrels fixture missing "queries" array');
}
return parsed.queries as LegacyQrelsQuery[];
}
/**
* Build a query-embed function from a raw qrels fixture (the legacy shape:
* `{queries: [{query, embedding_dim, ...}]}`). Known query texts map to
* `basisEmbedding(embedding_dim)`; unknown texts fall back to a
* deterministic FNV-1a-derived basis dimension in [100, 1099] outside the
* fixture's low dims, so an unknown query can never accidentally vote for a
* fixture page's basis direction (fixture dims are small integers).
*
* Throws on malformed JSON or a missing `queries` array (caller surfaces a
* usage error; the gate's own qrels parser reports shape problems in detail).
*/
export function buildQrelsQueryEmbedFn(qrelsRaw: string): (text: string) => Float32Array {
const queries = parseLegacyQrels(qrelsRaw);
const dimByQuery = new Map<string, number>();
for (const q of queries) {
if (typeof (q as { query?: unknown })?.query === 'string' && typeof (q as { embedding_dim?: unknown })?.embedding_dim === 'number') {
dimByQuery.set(q.query, q.embedding_dim);
}
}
return (text: string): Float32Array => {
const dim = dimByQuery.get(text);
if (dim !== undefined) return basisEmbedding(dim);
return basisEmbedding(100 + (Math.abs(fnv1a(text)) % 1000));
};
}
+3 -1
View File
@@ -922,6 +922,7 @@ CREATE TABLE IF NOT EXISTS minion_jobs (
depth INTEGER NOT NULL DEFAULT 0,
max_children INTEGER,
timeout_ms INTEGER,
lock_duration_ms INTEGER,
timeout_at TIMESTAMPTZ,
remove_on_complete BOOLEAN NOT NULL DEFAULT FALSE,
remove_on_fail BOOLEAN NOT NULL DEFAULT FALSE,
@@ -938,7 +939,8 @@ CREATE TABLE IF NOT EXISTS minion_jobs (
CONSTRAINT chk_nonnegative CHECK (attempts_made >= 0 AND attempts_started >= 0 AND stalled_counter >= 0 AND max_attempts >= 1 AND max_stalled >= 0),
CONSTRAINT chk_depth_nonnegative CHECK (depth >= 0),
CONSTRAINT chk_max_children_positive CHECK (max_children IS NULL OR max_children > 0),
CONSTRAINT chk_timeout_positive CHECK (timeout_ms IS NULL OR timeout_ms > 0)
CONSTRAINT chk_timeout_positive CHECK (timeout_ms IS NULL OR timeout_ms > 0),
CONSTRAINT chk_lock_duration_positive CHECK (lock_duration_ms IS NULL OR (lock_duration_ms >= 5000 AND lock_duration_ms <= 3600000))
);
CREATE INDEX IF NOT EXISTS idx_minion_jobs_claim ON minion_jobs (queue, priority ASC, created_at ASC) WHERE status = 'waiting';
@@ -48,6 +48,14 @@ registration is always user-global and the tradeoff above is the standing
state. Off-ramps: `codex mcp remove gbrain` removes just the registration;
`gbrain bootstrap uninstall` is the full teardown.
opencode: the scope logic is INVERTED from Claude Code. opencode spawns
servers from a project `opencode.json` with NO trust prompt, so a
project-scoped registration in a repo you share means anyone who checks the
repo out gets the entry executed on open. gbrain therefore defaults to a
user-global registration; project scope is an explicit opt-in that prints a
sharing warning. Off-ramps: `gbrain bootstrap uninstall` removes the entry
from both scope files; deleting the `mcp.gbrain` key by hand also works.
## The transcript corpus
Session transcripts are retained locally (outside this repo, mode 0700, pruned
+1 -1
View File
@@ -72,7 +72,7 @@ during long work reads as broken.
**Gate 2 — Recover missed context.** Scan the conversation for earlier messages that
never got processed. Before sending the final reply, rescan for anything that
arrived mid-turn. On a harness WITHOUT hooks (Codex — pull protocol), also run
arrived mid-turn. On a harness WITHOUT hooks (Codex / opencode — pull protocol), also run
`gbrain bootstrap status` once at the start of a conversation: it surfaces a
failing or stale workspace push that hook-carrying harnesses would have shown
automatically. If it reports the push FAILING, tell {{PRINCIPAL_NAME}} plainly —
+2 -2
View File
@@ -10,8 +10,8 @@ remain private.
anything matching the deny list in `.gitignore`.
- **How it syncs:** `gbrain sources push` — a secret-scan-gated commit + push that
refuses public remotes. On Claude Code it runs automatically per turn
(debounced) and at session end via hooks; on Codex (no hook system) run it at
natural stopping points (the AGENTS.md gate reminds you). If background
(debounced) and at session end via hooks; on Codex or opencode (no wired hook
system) run it at natural stopping points (the AGENTS.md gate reminds you). If background
persistence is enabled, a git post-commit hook auto-pushes each commit and a
30-minute pull job keeps multi-machine checkouts fresh. Run it by hand after
meaningful changes on any harness.
+2 -2
View File
@@ -118,7 +118,7 @@
"MCP_SCOPE": {
"consent": true,
"phase": "interview",
"question": "(Claude Code only. Codex has no scope flag — its registrations are always user-global; on Codex, state that plainly instead of asking.) Register the brain for THIS folder only (recommended — any other repo you open cannot read it), or for every session on this machine (your agent everywhere, but any repo you open can query your brain, and two open sessions will contend for the local database)?",
"question": "(Claude Code and opencode. Codex has no scope flag — its registrations are always user-global; on Codex, state that plainly instead of asking. On opencode the DEFAULT is user-global — the sharing-safe choice, because opencode spawns project-config-defined servers with no trust gate; offer 'project' only as a deliberate opt-in and state the committed-file consequence.) Register the brain for THIS folder only (recommended on Claude Code — any other repo you open cannot read it), or for every session on this machine (your agent everywhere, but any repo you open can query your brain, and two open sessions will contend for the local database)?",
"default": "project",
"allowed": ["project", "user"],
"maxLength": 8
@@ -163,7 +163,7 @@
"VOICE_BANNED": { "default": "- Opening with filler (\"Great question\", \"I'd be happy to help\", \"Absolutely\")\n- Hedging when a take exists\n- \"It's not X, it's Y\" constructions\n- Narrating the writing process inside a document", "maxLength": 1024, "shape": "list" },
"SAFETY_RED_LINES": { "default": "- Never send money or make purchases without explicit per-instance approval\n- Never message third parties as the principal without sign-off on the exact text\n- Never delete data that cannot be restored\n- Never share the principal's private information with anyone but the principal", "maxLength": 2048, "shape": "list" },
"QUIET_HOURS": { "default": "23:00-08:00 local", "maxLength": 64 },
"SURFACE_PRIMARY": { "default": "this workspace (Claude Code / Codex)", "maxLength": 128 },
"SURFACE_PRIMARY": { "default": "this workspace (Claude Code / Codex / opencode)", "maxLength": 128 },
"SURFACE_MULTIUSER": { "default": "single-principal", "allowed": ["single-principal", "shared"], "maxLength": 32 },
"MEMORY_WHAT_MATTERS": { "default": "- Corrections the principal makes (these become standing rules)\n- Commitments made in either direction, with dates\n- Preferences stated once that should never need restating\n- Facts about people and projects the principal works with", "maxLength": 2048, "shape": "list" },
"PRINCIPAL_PROJECTS": { "default": "*(none recorded yet — add as they come up)*", "maxLength": 4096, "shape": "list" },
@@ -48,6 +48,14 @@ registration is always user-global and the tradeoff above is the standing
state. Off-ramps: `codex mcp remove gbrain` removes just the registration;
`gbrain bootstrap uninstall` is the full teardown.
opencode: the scope logic is INVERTED from Claude Code. opencode spawns
servers from a project `opencode.json` with NO trust prompt, so a
project-scoped registration in a repo you share means anyone who checks the
repo out gets the entry executed on open. gbrain therefore defaults to a
user-global registration; project scope is an explicit opt-in that prints a
sharing warning. Off-ramps: `gbrain bootstrap uninstall` removes the entry
from both scope files; deleting the `mcp.gbrain` key by hand also works.
## The transcript corpus
Session transcripts are retained locally (outside this repo, mode 0700, pruned
+1 -1
View File
@@ -76,7 +76,7 @@ during long work reads as broken.
**Gate 2 — Recover missed context.** Scan the conversation for earlier messages that
never got processed. Before sending the final reply, rescan for anything that
arrived mid-turn. On a harness WITHOUT hooks (Codex — pull protocol), also run
arrived mid-turn. On a harness WITHOUT hooks (Codex / opencode — pull protocol), also run
`gbrain bootstrap status` once at the start of a conversation: it surfaces a
failing or stale workspace push that hook-carrying harnesses would have shown
automatically. If it reports the push FAILING, tell {{PRINCIPAL_NAME}} plainly —
+2 -2
View File
@@ -10,8 +10,8 @@ remain private.
anything matching the deny list in `.gitignore`.
- **How it syncs:** `gbrain sources push` — a secret-scan-gated commit + push that
refuses public remotes. On Claude Code it runs automatically per turn
(debounced) and at session end via hooks; on Codex (no hook system) run it at
natural stopping points (the AGENTS.md gate reminds you). If background
(debounced) and at session end via hooks; on Codex or opencode (no wired hook
system) run it at natural stopping points (the AGENTS.md gate reminds you). If background
persistence is enabled, a git post-commit hook auto-pushes each commit and a
30-minute pull job keeps multi-machine checkouts fresh. Run it by hand after
meaningful changes on any harness.
+1 -1
View File
@@ -1,6 +1,6 @@
# gbrain agent workspace — template
<!-- gbrain-template-stamp: 0.46.3.0 -->
<!-- gbrain-template-stamp: 0.46.6.0 -->
This repository is the **"Use this template"** distribution artifact for a
[gbrain](https://github.com/garrytan/gbrain) personal-agent workspace — the same
+1 -1
View File
@@ -7,7 +7,7 @@ valence of what they said, stop and re-derive from their words.
- **Name:** {{PRINCIPAL_NAME}}
- **Timezone:** America/Los_Angeles
- **Primary surface:** this workspace (Claude Code / Codex)
- **Primary surface:** this workspace (Claude Code / Codex / opencode)
## Context
+63
View File
@@ -85,6 +85,69 @@ describe('lockRenewalAudit: 4-outcome contract', () => {
});
});
describe('lockRenewalAudit: v0.46 (#4145) additive telemetry fields', () => {
test('case 9a — ctx fields round-trip through the JSONL (cause/lateness/overlap/load/via)', async () => {
await withEnv({ GBRAIN_AUDIT_DIR: tmpDir }, async () => {
lockRenewalAudit.logFailure(7, 'subagent', 2, new Error('x'), {
cause: 'call-timeout',
lateness_ms: 40_000,
overlap_skips: 1,
load1: 28.12,
cores: 32,
deadline_deferred: true,
});
lockRenewalAudit.logSuccessAfterFailure(7, 'subagent', 2, { via: 'verify' });
const result = readRecentLockRenewalEvents(24);
expect(result.events).toHaveLength(2);
expect(result.events[0]).toMatchObject({
outcome: 'failure',
cause: 'call-timeout',
lateness_ms: 40_000,
overlap_skips: 1,
load1: 28.12,
cores: 32,
deadline_deferred: true,
});
expect(result.events[1]).toMatchObject({ outcome: 'success_after_failure', via: 'verify' });
});
});
test('case 9b — omitted ctx leaves the new keys OUT of the JSONL entirely (no undefined noise)', async () => {
await withEnv({ GBRAIN_AUDIT_DIR: tmpDir }, async () => {
lockRenewalAudit.logGaveUp(8, 'sync', 3, new Error('y'));
const result = readRecentLockRenewalEvents(24);
expect(result.events).toHaveLength(1);
const raw = result.events[0] as unknown as Record<string, unknown>;
for (const key of ['cause', 'lateness_ms', 'overlap_skips', 'load1', 'cores', 'via', 'deadline_deferred']) {
expect(key in raw).toBe(false);
}
});
});
test('case 9c — pre-upgrade JSONL lines (no telemetry fields) still parse in readback', async () => {
await withEnv({ GBRAIN_AUDIT_DIR: tmpDir }, async () => {
// A line exactly as a pre-v0.46 build would have written it.
const legacy = JSON.stringify({
ts: new Date().toISOString(),
job_id: 5,
job_name: 'embed',
attempt: 1,
outcome: 'failure',
error_message_summary: 'Connection terminated',
error_code: '08006',
});
const file = path.join(tmpDir, computeIsoWeekFilename(LOCK_RENEWAL_FEATURE_NAME, new Date()));
fs.mkdirSync(path.dirname(file), { recursive: true });
fs.writeFileSync(file, `${legacy}\n`);
const result = readRecentLockRenewalEvents(24);
expect(result.corrupted_lines).toBe(0);
expect(result.events).toHaveLength(1);
expect(result.events[0].outcome).toBe('failure');
expect(result.events[0].cause).toBeUndefined();
});
});
});
describe('lockRenewalAudit: privacy via redactor (D9)', () => {
test('case 5a — logFailure with PG connection-failure error: no DSN/IP in JSONL', async () => {
await withEnv({ GBRAIN_AUDIT_DIR: tmpDir }, async () => {
+141
View File
@@ -0,0 +1,141 @@
/**
* atomic-write.ts the ONE atomic config writer for bootstrap host surfaces.
* Pins the symlink-preservation contract (live AND dangling links survive as
* links; the write lands at the resolved target) plus the mode ladder:
* forceMode > existing-file mode > freshMode.
*
* The dangling case is the red-team finding: existsSync FOLLOWS symlinks, so
* a dangling link reads "absent" and a naive rename would replace the link
* itself with a regular file a dotfile-manager layout whose target was
* cleaned up would silently stop being managed.
*/
import { afterEach, beforeEach, describe, expect, test } from 'bun:test';
import {
chmodSync,
existsSync,
lstatSync,
mkdirSync,
mkdtempSync,
readdirSync,
readFileSync,
rmSync,
statSync,
symlinkSync,
writeFileSync,
} from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { atomicWriteTextFile } from '../src/core/bootstrap/atomic-write.ts';
let dir: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), 'gbrain-atomic-write-'));
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
describe('atomicWriteTextFile — symlink preservation', () => {
test('live symlink: write lands in the TARGET, the link survives as a link', () => {
const target = join(dir, 'dotfiles', 'config.jsonc');
mkdirSync(join(dir, 'dotfiles'), { recursive: true });
writeFileSync(target, '{"old":true}');
const link = join(dir, 'config.jsonc');
symlinkSync(target, link);
atomicWriteTextFile(link, '{"new":true}');
expect(lstatSync(link).isSymbolicLink()).toBe(true);
expect(readFileSync(target, 'utf8')).toBe('{"new":true}');
expect(readFileSync(link, 'utf8')).toBe('{"new":true}');
});
test('DANGLING symlink: the missing target is created (parent dir too) and the link survives', () => {
const target = join(dir, 'dotfiles', 'nested', 'config.jsonc'); // dir does not exist either
const link = join(dir, 'config.jsonc');
symlinkSync(target, link);
expect(existsSync(link)).toBe(false); // existsSync follows the link — the trap
atomicWriteTextFile(link, '{"created":true}', { freshMode: 0o600 });
expect(lstatSync(link).isSymbolicLink()).toBe(true); // NOT replaced by a regular file
expect(readFileSync(target, 'utf8')).toBe('{"created":true}');
expect(readFileSync(link, 'utf8')).toBe('{"created":true}');
expect(statSync(target).mode & 0o777).toBe(0o600); // fresh target takes freshMode
});
test('DANGLING symlink with RELATIVE link text resolves against the link dir', () => {
const link = join(dir, 'config.jsonc');
symlinkSync(join('sub', 'real.jsonc'), link); // relative, target absent
atomicWriteTextFile(link, 'relative-ok');
expect(lstatSync(link).isSymbolicLink()).toBe(true);
expect(readFileSync(join(dir, 'sub', 'real.jsonc'), 'utf8')).toBe('relative-ok');
expect(readFileSync(link, 'utf8')).toBe('relative-ok');
});
});
describe('atomicWriteTextFile — mode ladder', () => {
test('fresh file (ENOENT) takes freshMode', () => {
const p = join(dir, 'fresh.json');
atomicWriteTextFile(p, '{}', { freshMode: 0o600 });
expect(statSync(p).mode & 0o777).toBe(0o600);
});
test('fresh file without freshMode gets the platform default (no chmod)', () => {
const p = join(dir, 'fresh-default.json');
atomicWriteTextFile(p, '{}');
expect(existsSync(p)).toBe(true); // mode is umask-dependent; existence is the pin
});
test('forceMode overrides an existing file\'s looser mode', () => {
const p = join(dir, 'secret.toml');
writeFileSync(p, 'old');
chmodSync(p, 0o644); // a known loose mode first
atomicWriteTextFile(p, 'new', { forceMode: 0o600 });
expect(statSync(p).mode & 0o777).toBe(0o600);
expect(readFileSync(p, 'utf8')).toBe('new');
});
test('existing file\'s own mode is inherited when neither force nor fresh applies', () => {
const p = join(dir, 'keep-mode.json');
writeFileSync(p, 'old');
chmodSync(p, 0o640);
atomicWriteTextFile(p, 'new', { freshMode: 0o600 }); // freshMode must NOT apply — the file exists
expect(statSync(p).mode & 0o777).toBe(0o640);
expect(readFileSync(p, 'utf8')).toBe('new');
});
});
describe('atomicWriteTextFile — failure hygiene (no tmp litter)', () => {
test('a failing rename does not leak the .tmp- file (target is a directory → rename throws)', () => {
// The resolved target being a DIRECTORY makes writeFileSync of the tmp
// succeed but renameSync(tmp, target) throw — the exact mid-sequence
// failure shape (ENOSPC/EACCES class) that used to strand tmp litter
// next to the user's config.
const target = join(dir, 'config.jsonc');
mkdirSync(target, { recursive: true });
expect(() => atomicWriteTextFile(target, '{"x":1}')).toThrow();
const litter = readdirSync(dir).filter((n) => n.includes('.tmp-'));
expect(litter).toEqual([]);
expect(lstatSync(target).isDirectory()).toBe(true); // target untouched
});
test('a failing write in a read-only dir throws without leaving litter behind', () => {
if (process.getuid?.() === 0) return; // root ignores modes
const ro = join(dir, 'ro');
mkdirSync(ro, { recursive: true });
writeFileSync(join(ro, 'config.jsonc'), 'old');
chmodSync(ro, 0o500);
try {
expect(() => atomicWriteTextFile(join(ro, 'config.jsonc'), 'new')).toThrow();
} finally {
chmodSync(ro, 0o700);
}
const litter = readdirSync(ro).filter((n) => n.includes('.tmp-'));
expect(litter).toEqual([]);
expect(readFileSync(join(ro, 'config.jsonc'), 'utf8')).toBe('old'); // original intact
});
});

Some files were not shown because too many files have changed in this diff Show More