mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 08:53:22 +00:00
* docs(designs): agent-bootstrap plan + design docs (normative, review-absorbed) The scrubbed, in-repo sources of truth for the gbrain bootstrap wave: AGENT_BOOTSTRAP_DESIGN.md (product scope/sequencing) and AGENT_BOOTSTRAP_PLAN.md (implementation; all review-finding IDs inlined). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): format spec, question bank, identity templates, bundled assets agent.json manifest (format_version 1, initialized sentinel) + machine-local install receipt [CX2-1, CX2-12]; 12-question/6-required interview bank with consent keys and a persist:false sink for the optional provider key [CX2-13]; ten {{TOKEN}} identity templates (generic, adapted to gbrain ops — gates call recall/query/put_page, write-through-ops rule, keyless agent-authored facts, silence contract); assets embedded compiled-binary-safe via file-type imports [ENG-6]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): runbook, README paste block, bootstrap guide, TODOS entries BOOTSTRAP_FOR_AGENTS.md (agent-driven install runbook: CLI phase list is the source of truth, never-invent rules, Codex approvals preflight, keyless posture, failure-modes table, version stamp for the skew check); README gains the full-agent paste block pinned to latest-stable inside the Claude Code/Codex quick start (memory-only tier stays); docs/guides/bootstrap.md carries the full install/security/consent/degradation/uninstall contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(designs): spike instrument for the bootstrap wave (build order 0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): interview + render engines Interview gate with read-back confirm-hash (any later answer change clears the confirmation — the hostile single-batch case is structurally impossible), per-answer provenance, caps + escaping at set time, config-sink routing for the provider key; renderer with hard-fail token sweep, subordinate fencing of principal input, never-clobber + backups, deterministic minimal mode for the template repo, scaled byte floors. 58 unit tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): private-repo lifecycle — repo create, attach, uninstall, run lock gh-gated private repo creation with API-verified privacy (rate-limit distinct from public), refuse-foreign-origin with attach as the sanctioned path, atomic bootstrap mutex (pid liveness + age + token), receipt-keyed uninstall that never wholesale-deletes the gbrain home and only offers --delete-brain for a brain it created; read-only PGLite lock probe (never opens the engine). 54 unit tests, injectable exec seam throughout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(release): latest-stable ref, template-repo publish job, bootstrap CI guards release.yml advances the latest-stable tag only after assets publish (the paste block's permanent ref — copies in the wild never rot) and gains a PAT-gated publish-template job verified against the vendored tree; two skip-graceful guards (sanctioned-ref + runbook stamp; template/token bijection + placeholder assertion + generator byte-diff) wired into verify; README + runbook re-admitted to the CI cache hash; vendored deterministic template tree generated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(context): IPC v2 turn_context + 8KB assembly + visibility resolver + session identity Discriminated-union IPC with handler map, protocol echo (stale-serve detection), shared-secret gate, server-side source binding, per-kind budgets; turn-context assembly (reflex pointers + volunteered pages + world-only hot facts) under a data-not-instructions envelope trimmed to the harness's 10KB hook-output cap; facts.default_visibility resolved through one helper at all four sites (explicit caller wins, typos fail closed); typed sessionId threads _meta.session_id into the hot-memory cache key. 50 new tests; 180 adjacent tests confirmed green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(persistence): secret-scan, gbrain sources push, durability unification Pattern secret scanner (own runtime allowlist, redacted previews, corpus-write redaction mode); sources push runs the whole scan→stage→commit→pull→push sequence under one cross-platform lock (mkdir-atomic, pid+age+token) with a deny-glob backstop, commit-first divergence-safe pull, refuse-unverifiable visibility, and push-status telemetry; gbrain-home choke point unifies GBRAIN_HOME semantics with config (0700); durability is parent-repo-aware and rotates its push log at 0600. 35 new tests; 200 existing green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sources): harden/pull gates accept sources inside a parent git repo The bootstrap workspace registers brain/ (a subdirectory) as the source; the durability core already resolves the repo root, so the command gates now check inside-a-repo rather than .git-right-here [CX2-3]. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(serve): resident maintenance sweep + keyless capability probe The lock-owning serve process now closes the persistence loop: startup (3s post-connect, best-effort, unref'd) and idle (10-min quiet intervals through the injectable timer seam) sweeps run facts-fence reconciliation, deterministic link/timeline extraction over recent workspace pages, and spend-gated corpus ingest (skipped keyless — agent-authored fences cover it). gbrain sweep --once is the trusted CLI seam bootstrap verify uses. Capability probe renders the honest keyless/keyed report. Full reuse of the cycle extractor + extract cores; 26 new tests, neighbors green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(hooks): engine-free gbrain hook command, settings writers, transcript parser Four hook events (session-start digest + crashed-session recovery push, user-prompt turn-context injection under an 800ms deadline and the 10KB cap, stop buffers, session-end corpus write with redaction/retention/dedup + best-effort push); structural JSON settings merger keyed by a _gbrain marker (foreign hooks and permissions survive); dated host-spec registry; Claude Code .jsonl parser as a spec-target with a scrubbed 7-shape fixture. Heartbeat is counters-only by construction. 59 tests; zero engine modules in the import graph. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(bootstrap): cross-link the full-agent path from the connection docs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): dispatcher, verify, status — the command assembled gbrain bootstrap {status,interview,render,repo,hooks,verify,uninstall,attach}: engine-free except verify (owns its engine, in-process sweep — no live-serve conflict); phase list is the TS source of truth with install.jsonl telemetry and the support blob; verify's fail-soft check suite covers the real write path (put_page → write-through file → sweep → graph floor → recall), passes keyless, persists snapshots, and ends with the first-run tour. cli.ts wired per the three-touchpoint rule; doctor gains the bootstrap check group (silent on machines with no bootstrap state). 28 new tests; 353 adjacent green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: KEY_FILES bootstrap cluster + CLAUDE.md dispatcher row (+ build:llms) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): e2e pins — hook-under-live-serve, attach, degraded modes, compiled binary, Docker harness The permanent pins: a real serve holds the PGLite lock while the engine-free hook completes (and a direct engine open provably throws LiveServeLockError); stale-socket fail-open; machine-2 attach with marker-keyed hook repair; decline-everything installs verify green with every degradation named; the compiled binary renders bundled templates in an empty cwd. Offline Docker harness (networkless, read-only) gated into heavy-tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bootstrap): register doctor check categories + system-of-record allow comments The six bootstrap doctor checks join OPS_CHECK_NAMES; the sweep's batch link/ timeline inserts carry the explicit extract-path allow comments (the sweep IS the extraction path for workspace pages). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard wedge cap tracks suite growth (1500s -> 1800s) At ~9000 tests a healthy shard finished at 1466s and two progressing shards were false-killed at the old cap; 1800s restores ~25% headroom over the slowest observed healthy shard. Real hangs still hit it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ci): cache-hash policy — README + runbook edits must invalidate [C2] The old deny-list assertion predates the paste block; README.md and BOOTSTRAP_FOR_AGENTS.md are policy-doc re-admissions now, so their edits must change the hash (a paste-block edit shipping under a cached green was the C2 hole). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): classify post-suite exit-hangs as warn-pass; file the leak forensics A shard killed by the wedge watchdog with every assigned file started and zero fail markers did all its work and leaked a handle at exit — pre-existing and master-reproducible (P1 TODO carries the full bisect forensics). Bun's per-test timeout turns a hung test into a (fail), so the classifier cannot mask one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard cap 2400s — the count-balanced heavy shard needs it under contention Observed: the heavy shard still progressing 22s before an 1800s kill while siblings finish at 1150-1550s (split balances file count, not weight). Filed the load-sensitive WAL-repair flake (pre-existing, master's v0.42.75.0 wave). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): quarantine env-mutating suites to the serial lane check-test-isolation R1: six new files mutate GBRAIN_HOME/env at module scope — the serial lane (one process per file) is the guard's prescribed home for them. All 114 tests pass post-rename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): per-harness install sections — Codex, Claude Code, then OpenClaw/Hermes Each harness gets its own complete paste-block section (desktop app first, terminal noted — Claude Code CLI is the identical harness; Codex CLI works pull-based today); the OpenClaw/Hermes platform path keeps equal weight with its one-click deploys and INSTALL_FOR_AGENTS block intact; memory-only and remote-connect tiers consolidated under 'Lighter ways in'. Supersedes the review's D5 ordering by user direction; stale heading references updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): Codex as the recommended first step; OpenClaw/Hermes framed as-intended, high-cost The install section now routes newcomers explicitly: Codex first (subscription-priced, nothing to deploy), OpenClaw/Hermes as GBrain used the way it was designed — always on, at real server + API cost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: doc-audit code chasers — broken recovery hints, stale op description, auto_link key Four small code fixes surfaced by the markdown accuracy audit: - doctor's auto-RLS recovery hint pointed at `apply-migrations --force-retry 35`, which cannot work (--force-retry targets the vX.Y.Z orchestrator registry, not the numeric schema MIGRATIONS array). Hint now points at the recreate SQL in docs/guides/rls-and-you.md; test pins against regression. - v0_11_0 migration printed the same broken-mechanism class of hint (`config set minion_mode` writes DB config nothing reads); now names `apply-migrations --mode` + preferences.json, the real setter. - submit_job's op description hardcoded a stale handler list; now points at registerBuiltinHandlers as the source plus the --follow discovery trick. - `auto_link` added to KNOWN_CONFIG_KEYS: read by link-extraction, reconcile-links, and sweep, and documented as the off-switch in brain-ops/maintain, but the allowlist rejected `config set auto_link false`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: repo-wide accuracy + MECE reform from the 9-bucket markdown audit A code-grounded audit of every markdown file (root, architecture, guides, mcp, tutorials, docs-root, operations/eval/designs, skills, recipes) followed by a fix wave with per-bucket ownership. Four classes of change: Accuracy — every documented command/flag verified against src/ before writing: dead commands replaced with working ones (pages purge-deleted, jobs watch --follow, gbrain restore, import-based Obsidian flow, space-separated --scopes, real thin-client recipes, working isolation verification, real supervisor restart procedure, curl-based ngrok health check, real minion_mode setter); count drift fixed with rot-proof phrasing (100+ ops, 50+ bundled skills via skills/manifest.json, 140+ engine methods, KNOBS_HASH_VERSION pointer instead of hardcoded versions); stale claims corrected (search-mode defaults, RETRIEVAL pipeline order incl. autocut, sentinel rules, refusal-list mechanism, engine snapshot, shard cap 2400s + EXIT-HANG classifier in TESTING.md, latest-stable + publish-template documented in RELEASING.md as release.yml promises). MECE — one home per concept, pointers elsewhere: test isolation → TESTING.md; OAuth registration + --bind/--public-url lore → DEPLOY.md; mode bundles → guides/search-modes.md (the home the CLAUDE.md dispatcher always promised); merge contract → schema-packs.md; WAL ladder → ENGINES.md; quiet-hours → quiet-hours.md; capture taxonomy → entity-detection.md; person-page taxonomy → compiled-truth.md; brain-first protocol → brain-first-lookup.md; refresh semantics → refresh-algorithm.md; KEY_FILES.md deduplicated (58 extension entries merged, one entry per file); infra-layer.md rewritten as a pointer page. Privacy — placeholder sweep across guides, docs, skills, and recipes per the iron rule; per-release narration stripped from reference docs (current-state prose only). Bootstrap coverage — AGENTS.md pointer, RESOLVER routing row, INSTALL.md path, tutorial cross-links, keyless-mode sections in spend-controls/headless-install. skills.lock.json regenerated; llms.txt/llms-full.txt rebuilt. Gates: verify 36/36, typecheck clean, doctor 96/96, skills-integrity + resolver + build-llms + config-set + migrations all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme): refresh production-brain stats to current brain-repo counts 155,795 pages / 24,589 people / 5,340 companies, counted from the brain repo's current HEAD; the "100K-page brain" framing moves to 150K to match. Cron-fleet count unchanged (its store lives on the deployment host, not in the repos available for verification). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scan,push): modern OpenAI/Voyage key patterns, scan staged blobs not disk - secret-scan matches sk-proj-/sk-svcacct-/sk-None- and pa- Voyage keys (the bare sk- pattern missed every current OpenAI key format). - workspacePush stages first, then scans the staged index blobs via git cat-file, closing the scan-then-stage TOCTOU where a file changed between snapshot and commit shipped unscanned. - shared binary-sniff helper, memoized glob regexes, atomic push-status write, and tests for pull_conflict + gitignored deny-match paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): regenerate flag registry for new commands, harden shard classifier + release token - cli-flag-registry.generated.ts regenerated: bootstrap/hook/sweep and sources push --message/--allow-unverified-remote were missing, so the strict #2185 validator rejected real invocations and skipped the new commands entirely. - EXIT-HANG shard classifier now requires every assigned file to have started before warn-passing a watchdog kill (was fail-open). - release.yml passes TEMPLATE_REPO_PAT via http.extraheader, off the argv. - compiled-binary e2e fails loud in CI instead of a silent permanent skip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(hook,sweep,ipc): non-blocking hook pushes, bounded sweep + cache, source-bound resolve - session-start/session-end no longer run synchronous git + inline push inside their self-deadline; a detached child does the push and the hook returns immediately (blocked Claude Code startup for minutes on a dirty tree before). - serve sweep drops the unbounded listAllPageRefs, resolves only candidate targets, claims corpus files atomically (no double-LLM-spend race), and caps the fence LIKE scan; heartbeat writes are O_APPEND with rare compaction. - hot-memory cache evicts expired entries and bounds entry count (the key is caller-controlled via _meta.session_id). - v1 resolve IPC honors boundSourceId like turn_context; turn-context runs its arms concurrently. doctor reads push/heartbeat thresholds from hook.ts. - new tests: doctor bootstrap checks, hook push-gate + deadline, concurrent sweep claims, cache eviction, bound-source resolve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bootstrap): origin-ownership gate, world visibility, collision-safe source id, consent + templates - repo adoption requires an exact receipt repo_url match or authed-owner check (undefined repo_url was a wildcard); create verifies privacy BEFORE the first push. - verify sets facts.default_visibility=world if unset, so agent-authored facts surface in per-turn context (they defaulted private before). - source_id derives a path-hash suffix when 'workspace' is taken by another checkout; every consumer reads manifest.source_id. - skipped HOOKS_CONSENT now declines (was falling through to default yes); --minimal refuses on an initialized manifest; tilde fences escaped. - MCP registration pins --surface full; status hard-fails a public origin (template door); receipt writers guard against newer/corrupt receipts; uninstall only claims brain-deleted after a real rm. - templates ship jobs disabled + provider-consent + support-relay lines; soul-audit re-runs over the shared interview bank. TODOS: 11 follow-ups. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): close adversarial-review findings — scan fails closed, whole-PEM redaction, bound repo push Cross-model adversarial pass (Claude + Codex) on the bootstrap wave: - secret scan fails CLOSED: an unreadable, oversized, or binary staged blob now blocks the push (blocked_unscannable, exit 5) instead of committing unscanned; only a confirmed staged deletion is skipped. This was the headline "block secrets before they leave the machine" property failing open. - private-key redaction spans the whole PEM block (header+body+footer), not just the header line — the base64 body no longer survives into the corpus the sweep sends to an extraction provider. - bootstrap repo commits the workspace (secret-scan-gated) before the first push and verifies the remote actually received it, so a push-fail retry can't adopt an empty remote as success. - privacy verify is re-bound to origin immediately before push (a concurrent origin rewrite between verify and push is refused). - session-end corpus write is atomic and clears the stale ingested/in-progress sidecars so a resumed session's appended transcript is re-ingested. - public-origin refusal enforced at render (not only status); MCP "already registered" is verified to target this workspace, not blessed blindly; verify probe cleanup scopes deletes to its own slugs, not a token substring; allowlist fingerprint floor raised 8→16 hex. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code Turns a Codex or Claude Code session into a persistent personal agent: interview-rendered identity files, a local PGLite brain, per-turn context via serve IPC (Claude Code hooks / Codex pull protocol), session-triggered persistence, and a private GitHub repo as the agent's portable body. Keyless- first (the harness model is the LLM; one optional key adds embeddings + extraction). New `gbrain bootstrap` command family + `gbrain hook` + `gbrain sweep`; doctor bootstrap health checks; latest-stable distribution ref + template-repo publish job. Opt-in, additive — existing installs untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: regenerate flag registry for security-fix flags; drop fabricated gbrain capabilities doc ref CI caught two real failures under the merged state: - the flag registry lagged the blocked_unscannable/exit-5 flags the security round added, tripping the #2185 freshness guard. - headless-install.md described the keyless capability report as a `gbrain capabilities` command, which the #3502 doc-command resolver rejects — reworded to prose (the real surface is bootstrap verify's report). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: sync KEY_FILES + bootstrap plan to security-fix behavior Cross-referenced the security-fix round against the reference docs and corrected the drift those commits introduced: - workspace-push.ts entry: stage-FIRST-then-scan order (the TOCTOU fix), fail-closed blocked_unscannable, and the sources-push status -> exit-code map. - hooks.ts entry: MCP registration pins `serve --surface full`. - hook.ts entry: session-start/session-end pushes run in a detached child (non-blocking); atomic corpus write clears stale sidecars. - bootstrap.ts entry: render hard-refuses a public origin (template door). - verify.ts entry: source_id collision resolution (workspace-<path-hash>). - AGENT_BOOTSTRAP_PLAN as-shipped delta note for the scan/stage reorder. llms bundle unchanged (KEY_FILES is link-only); build:llms and test/build-llms.test.ts green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): silence SC2016 on the intentional askpass literal in release.yml The one-shot GIT_ASKPASS script must contain literal $1 and $TEMPLATE_REPO_PAT so they expand when /bin/sh runs it at git's credential prompt, not when the outer shell writes the file — single quotes are correct. Add a scoped shellcheck disable so actionlint passes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(skills): declare 'bootstrap my data' trigger in cold-start frontmatter The doc reform added 'bootstrap my data' to cold-start's RESOLVER.md row (to disambiguate data-bootstrap from agent-bootstrap) but not to the skill's own frontmatter triggers, tripping the RESOLVER↔frontmatter round-trip contract (resolver.test.ts). Declare it; regenerate skills.lock.json. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bootstrap): default per-turn hooks + search mode ON without a prompt Installing gbrain for your coding agent IS the consent for the behaviors that make it work, so stop re-litigating them with install-time questions whose "no" defeats the product: - Per-turn hooks (Claude Code) install ON by default — no prompt. Off-ramps: `--no-hooks` at install, `GBRAIN_HOOKS=0` at runtime, `bootstrap uninstall`. The "hooks installed" line now surfaces the kill switch so default-on is never silent. A persisted HOOKS_CONSENT=no (interview --skip) still declines. - Search mode defaults to `balanced` silently (nobody knows the modes at install; `gbrain search modes` changes it any time). - MCP scope stays the ONE deliberate prompt — project vs user is a real cross-repo privacy choice, not friction. Marks the two consents `silent: true` in the question bank (new QuestionSpec field), rewrites the runbook phases so the agent no longer asks them, adds the `--no-hooks` flag (+ registry regen), and adds default-on / opt-out tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): real end-to-end coverage — cross-session recall, per-turn content, Codex door, realistic corpus Closes the seven e2e gaps a coverage audit surfaced: the plumbing was well-unit-tested but the product claims ("Codex works, context shows up every turn with real content, it remembers across restarts, machine two recovers, Postgres works") were unproven end to end. Test-only wave — zero src changes. - Hermetic synthetic corpus (test/fixtures/bootstrap-corpus/ + a loader helper): 12 interlinked pages (52 edges, timelines), 12 world/private beliefs, 8 gold queries — curated from the gbrain-evals synthetic corpora, 100% placeholder names, so recall is asserted on a real multi-entity brain instead of a 2-node self-planted probe. - GAP1 magic moment: author a fact via the real write path, disconnect the engine, reopen against the same DB, recall it — a real session boundary, not verify.ts's same-connection SQL read-back. Plus a source-isolation assertion. - GAP2 per-turn content: hook-under-serve Pin 1 now seeds a known fact and asserts its text lands in the injected block AND private beliefs never do (was: empty brain, empty_block accepted as a pass). - GAP3 Codex door: assert the rendered AGENTS.md carries the Gate-3 brain-first pull protocol; make the fake codex shim implement `mcp get` so the [FIX7] target-verification can actually fail; the Docker cold-machine harness now exercises the hooks/MCP registration step instead of skipping it. - GAP4 corpus recall: turn-context + verify graph-floor/qrels run on the real multi-entity brain with real edges. - GAP5 attach: machine-two now re-ingests the cloned brain/ into a fresh DB and recalls a fact authored only on machine one — the multi-device payoff. - GAP6 keyed + Postgres (env-gated): real embeddings prove semantic recall a paraphrase query can reach but keyless BM25 cannot; bootstrap verify drives a real Postgres engine (skipIf DATABASE_URL/keys absent). - GAP7 persistence: session-end runs the REAL push (not the mocked seam) to a local bare remote and the remote receives the content; a planted secret is blocked at the gate; the 15-min cron installs and fires a scan-gated push. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(bootstrap): real-agent e2e — drive the actual claude + codex binaries end to end Closes the audit's biggest gap ("no real harness ever drives a turn"). Adapts gstack's PTY/headless agent harness to prove the bootstrap install + smoke work against the REAL binaries, not PATH shims. Test/CI/docs only — zero src changes. - test/helpers/agent-harness.ts: hermetic clean-room child env (ported from gstack; drops CONDUCTOR_/CLAUDE_/GSTACK_/MCP_/GBRAIN_, promotes GSTACK_ANTHROPIC_API_KEY→ANTHROPIC_API_KEY), real-binary resolvers + auth probes, headless `claude -p --output-format stream-json` and `codex exec --json` turn runners, a gbrain stdio MCP-config writer, and a keyless brain seeder. + a fixture-parse unit test (no binary needed). - test/e2e/bootstrap-real-claude.serial.test.ts: real `gbrain bootstrap` install → REAL `claude mcp add` (verified via `claude mcp get`) → verify exit 0 → a real `claude -p --mcp-config --strict-mcp-config` turn that invokes mcp__gbrain__search and answers from the brain (proven: toolCalls include mcp__gbrain__search, final text carries the seeded fact). - test/e2e/bootstrap-real-codex.serial.test.ts: same install with REAL `codex mcp add` into a real ~/.codex/config.toml + Gate-3 pull-protocol assertion, then a real `codex exec --json` turn surfacing the fact (MCP or the pull- protocol shell path). Bounded retry absorbs codex's occasional MCP-call cancellation without softening the fact-requiring assertion. - Everything hermetic (temp HOME/CLAUDE_CONFIG_DIR/CODEX_HOME/GBRAIN_HOME; real ~/.codex auth copied read-only) and skipIf-gated so it self-skips cleanly where the binaries/auth are absent. - heavy-tests.yml: gated `real-agent-e2e` job (nightly/label, never the PR shard; no-op on a runner without authed binaries). - TODOS: compiled `gbrain` binary can't serve a PGLite brain (bun compile omits the WASM/extension payloads); harness falls back to `bun run` serve. Verified against live claude 4.6 + codex 0.147.0: 15 pass / 0 fail; verify 36/36; typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): real-agent-e2e job — bash array + --timeout (actionlint SC2086 + bun-test-timeout guard) The real-agent-e2e job's file loop used an unquoted $FILES (SC2086) and ran `bun test` without --timeout (check-bun-test-timeout guard). Switch to a bash array and add --timeout=600000 (real-agent turns are slow; the tests self-skip without authed binaries so it's a no-op elsewhere). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pglite): embed WASM + extension assets so the compiled binary can serve A `bun build --compile` gbrain binary could not `serve` a PGLite brain: the compile bundles JS but not PGLite's runtime payload (pglite.wasm, initdb.wasm, pglite.data, vector/pg_trgm tarballs), so `serve` on PGLite died with a bunfs/ENOENT. Now the assets ride inside the binary. - src/core/pglite-embedded-assets.ts: embeds the five assets via `import … with { type: 'file' }` (the ENG-6 idiom) and exposes getEmbeddedPgliteOptions() → { pgliteWasmModule, initdbWasmModule, fsBundle, extensions:{vector,pg_trgm} }. WASM/fsBundle are consumed as bytes; the two extension tarballs are materialized to a content-addressed temp file (atomic, size-verified reuse) because PGLite reads them via fs.createReadStream, which cannot read a /$bunfs path. Unconditional (works in bun-run and compiled), so no fragile mode branch. - src/core/pglite-engine.ts: static-import getEmbeddedPgliteOptions (engine path stays static per the engine-dynamic-import invariant); spread into both PGlite.create sites (initial + WAL-repair retry). The bunfs classifier stays as a backstop but no longer fires for a correct binary. - scripts/check-pglite-embedded.sh (+ smoketest): compiles a focused binary and asserts it boots PGLite, CREATE EXTENSION vector/pg_trgm, and round-trips a page — wired into `bun run verify` (now 37 checks), check:all, and check:pglite-embedded. Fail-soft only when compile is unavailable. - agent-harness.ts probeCompiledPglite now passes → the real-agent e2e uses the fast compiled MCP server. TODOS: the P2 "can't serve PGLite" item is closed. Verified: fresh compiled binary ran `search`/`query` against a PGLite brain and returned the seeded row (no bunfs/ENOENT); verify 37/37; pglite-engine 120/0 source-mode; typecheck clean; engine-dynamic-import + parity guards pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
515 lines
27 KiB
Markdown
515 lines
27 KiB
Markdown
# Releasing & contributing (gbrain)
|
|
|
|
The full release + contributor process. CLAUDE.md keeps the ship-critical IRON RULES
|
|
inline (the Version-locations table, branch=workspace, post-ship `/document-release`,
|
|
the Privacy + Responsible-disclosure rules, PR-title-version-first, never-hand-roll-ship)
|
|
and points here for everything else. **Before any ship, read this in full. Use `/ship` —
|
|
never hand-roll a release.**
|
|
|
|
## Pre-ship requirements
|
|
|
|
Before shipping (/ship) or reviewing (/review), always run the full test suite.
|
|
Two equivalent paths:
|
|
|
|
**Path A — local CI gate (recommended, v0.23.1+):**
|
|
- `bun run ci:local` runs the entire stack inside Docker: gitleaks (host),
|
|
guards + typecheck, then 4-shard parallel unit + E2E against four pgvector
|
|
containers plus a transaction-mode PgBouncer service (unit phase keeps
|
|
`DATABASE_URL` unset; `--no-shard` for the legacy sequential flow). Stronger
|
|
than PR CI's 2-file Tier 1 set; closer to what nightly Tier 1 catches. Spins
|
|
up + tears down postgres automatically via `docker-compose.ci.yml`. Override
|
|
the host port with `GBRAIN_CI_PG_PORT=5435 bun run ci:local` if 5434 collides.
|
|
- `bun run ci:local:diff` runs only the E2E files matched by the diff selector
|
|
(`scripts/select-e2e.ts`), falling back to ALL E2E files on unmapped src/
|
|
paths or schema/skills/package.json changes. Fast iteration during a focused
|
|
branch.
|
|
|
|
**Path B — manual lifecycle (still supported):**
|
|
- `bun test` — unit tests (no database required)
|
|
- Follow the "E2E test DB lifecycle" steps in
|
|
[docs/TESTING.md](TESTING.md) to spin up the test DB, run
|
|
`bun run test:e2e`, then tear it down.
|
|
|
|
Both must pass. Do not ship with failing E2E tests. Do not skip E2E tests.
|
|
|
|
**Always run typecheck before pushing.** Neither `bun test` (the bun runner)
|
|
nor `bun run test` gates on types — `bun run test` is just
|
|
`bash scripts/run-unit-parallel.sh` (the sharded unit runner; no typecheck,
|
|
no shell pre-checks — see the test-tier table in [docs/TESTING.md](TESTING.md)).
|
|
Three ways to actually gate on types:
|
|
|
|
1. `bun run verify` — runs the shell guard checks (privacy, jsonb, source-id,
|
|
progress-to-stdout, …) plus `bun run typecheck` in parallel
|
|
(`scripts/run-verify-parallel.sh`). Use this mid-branch.
|
|
2. `bun run typecheck` — `tsc --noEmit` standalone. Fast (~5s on this repo).
|
|
3. `bun run ci:local` — the full local CI gate from Path A.
|
|
|
|
The trap is: writing a new test, running `bun test test/foo.test.ts`,
|
|
seeing it pass, pushing — and CI's separate typecheck stage rejects an
|
|
invalid type literal that the runner accepted. Caught one of these
|
|
shipping the v0.23.2 round-trip E2E (`type: 'reflection'` is not a
|
|
member of `PageType`). Run `bun run typecheck` once before push, even
|
|
when only test files changed.
|
|
|
|
|
|
## CHANGELOG + VERSION are branch-scoped
|
|
|
|
**VERSION and CHANGELOG describe what THIS branch adds vs master, not how we got
|
|
here.** Every feature branch that ships gets its own version bump and CHANGELOG
|
|
entry. The entry is product release notes for users; it is not a log of internal
|
|
decisions, review rounds, or codex findings.
|
|
|
|
**Write the CHANGELOG entry at /ship time, not during development.** Mid-branch
|
|
iterations, review rounds (CEO/Eng/Codex/DX), and implementation detours belong
|
|
in the plan file at `~/.claude/plans/`, not in the CHANGELOG. One unified entry
|
|
per branch, covering what the branch added vs the base branch.
|
|
|
|
**Never edit a CHANGELOG entry that already landed on master.** If master has
|
|
v0.18.2 and your branch adds features, bump to the next version (v0.19.0, not
|
|
editing master's v0.18.2). When merging master into your branch, master may
|
|
bring new CHANGELOG entries above yours — push your entry above master's
|
|
latest and verify:
|
|
|
|
- Does CHANGELOG have your branch's own entry separate from master's entries?
|
|
- Is VERSION higher than master's VERSION?
|
|
- Is your entry the topmost `## [X.Y.Z]` entry?
|
|
- `grep "^## \[" CHANGELOG.md` shows a contiguous version sequence?
|
|
|
|
If any answer is no, fix it before continuing.
|
|
|
|
**CHANGELOG is for users, not contributors.** Write like product release notes:
|
|
|
|
- Lead with what the user can now **do** that they couldn't before. Sell the capability.
|
|
- Plain language, not implementation details. "You can now..." not "Refactored the..."
|
|
- **Never mention internal artifacts**: plan file IDs, decision tags (D-CX-#, F-ENG-#),
|
|
review rounds, codex findings, subcontractor credits. These are invisible to users.
|
|
- Put contributor-facing changes in a separate `### For contributors` section at the bottom.
|
|
- Every entry should make someone think "oh nice, I want to try that."
|
|
|
|
**What to omit:**
|
|
- "Codex caught X that the CEO review missed" — private process detail.
|
|
- "D-CX-3 split errors/warnings" — tag is meaningless to users; name the feature instead.
|
|
- "Fix-wave PR #N supersedes #M" — supersede chains belong in PR bodies, not release notes.
|
|
- "215 new cases, 3 decisions applied, 7 reviews cleared" — these are planning-mode metrics.
|
|
|
|
**What to keep:**
|
|
- The user-facing change: what commands exist now, what flag was added, what behavior fixed.
|
|
- Numbers that mean something to the user: TTHW, commands that timed out before, detection counts.
|
|
- Upgrade instructions: `gbrain upgrade` + any manual step if needed.
|
|
- Credit to external contributors when a community PR was incorporated.
|
|
|
|
## CHANGELOG voice + release-summary format
|
|
|
|
**IRON RULE: the CHANGELOG describes what the user gets, not how the work
|
|
happened.** Nobody reading release notes cares that codex caught a bug, that
|
|
the plan went through CEO + eng review, that the migration was originally
|
|
numbered v68 and renumbered to v79 during master merge, or that two
|
|
review rounds caught architectural mistakes. The reader cares what
|
|
`gbrain brainstorm` does and how to use it. If a fact only exists because
|
|
of the development process, it does NOT belong in the CHANGELOG.
|
|
|
|
**Specifically forbidden in CHANGELOG entries:**
|
|
|
|
- Any mention of review processes (CEO review, eng review, codex review,
|
|
plan-eng-review, outside voice, adversarial review, autoplan, /review).
|
|
- "What we caught and fixed before merging" sections. Bugs found pre-merge
|
|
are not changes — they're things that didn't ship.
|
|
- Plan file references, plan IDs, plan decision tags (D1, D14, D-CDX-3).
|
|
- Migration version drama ("originally v68", "renumbered to v77", "claimed
|
|
by parallel waves") — just say "Migration v79 adds X." If the user
|
|
cares about migration ordering, they read the diff.
|
|
- Round counts, finding counts, decision counts ("25 findings across 2
|
|
rounds", "8 architectural decisions", "5/6 expansions accepted").
|
|
- Names of internal collaborators ("codex caught", "the reviewer flagged",
|
|
"Claude noticed").
|
|
- "Plan + reviews" summary bullets. The plan lives in `~/.claude/plans/`;
|
|
if a future reader wants the backstory they can grep there.
|
|
- Any wording that frames a shipped feature as a *recovery* from a planning
|
|
mistake ("the first plan was wrong", "we corrected the approach", "the
|
|
shipped version supersedes the original design").
|
|
|
|
**Smell test:** read the entry as a stranger who has never touched gbrain.
|
|
If any sentence makes them think "why are you telling me this?", cut it.
|
|
Every sentence in the release-summary AND in the itemized changes must
|
|
answer one of three questions: *What can I now do? How do I use it? What
|
|
should I watch for after I upgrade?*
|
|
|
|
Every version entry in `CHANGELOG.md` MUST start with a release-summary section in
|
|
the GStack/Garry voice — one viewport's worth of prose + tables that lands like a
|
|
verdict, not marketing. The itemized changelog (subsections, bullets, files) goes
|
|
BELOW that summary, separated by a `### Itemized changes` header.
|
|
|
|
The release-summary section gets read by humans, by the auto-update agent, and by
|
|
anyone deciding whether to upgrade. The itemized list is for agents that need to
|
|
know exactly what changed.
|
|
|
|
### Release-summary template
|
|
|
|
**Iron rule: lead ELI10, get precise after.** The first ~150 words of every entry
|
|
must be readable by someone who does NOT know gbrain's internals. No file paths,
|
|
no function names, no internal constants, no acronyms (no "RRF", no "knobsHash",
|
|
no "MODE_BUNDLES", no "CDX-4"), no jargon that requires reading the codebase to
|
|
parse. Lead with the user-visible behavior change, in everyday English, like
|
|
you're explaining it to a smart engineer who has never opened the repo.
|
|
|
|
THEN, once the reader knows what shipped and why they'd care, drill into the
|
|
precise details: real file paths, real function names, real config keys, real
|
|
numbers. The precision part is required (the entry is also the technical record
|
|
of what changed), but it lives AFTER the plain-English lead, never before it.
|
|
|
|
The shape:
|
|
|
|
1. **One-line bold headline.** What changed for the user, in human English. No
|
|
jargon. No internal terms. Example good: "Your search stops boosting weak
|
|
pages just because they have a lot of links pointing at them." Example bad:
|
|
"PostFusionOpts gains floorRatio; KNOBS_HASH_VERSION bumped 2→3."
|
|
2. **Plain-English opener** (~3-5 sentences). Describe the problem this fixes in
|
|
everyday terms. Pretend the reader has a brain full of meeting notes and
|
|
people pages and wants to know if this release helps them. Concrete example
|
|
beats abstract description.
|
|
3. **A "How to turn it on" or "How to use it" section** with paste-ready
|
|
commands. Real flags, real config keys. This is where precision starts.
|
|
4. **A "What you'd see in a concrete example" or "The X numbers that matter"
|
|
section** with a table. Use everyday-language column headers ("Page",
|
|
"Match quality", "Has many backlinks?") even when the underlying mechanism
|
|
is technical. The table teaches what the feature does without requiring the
|
|
reader to understand how.
|
|
5. **A "What's safe to know about" or "Things to watch" section** for caveats,
|
|
side effects, cache invalidation, mid-deploy notes. Still in plain language.
|
|
6. **A "What we caught and fixed before merging" section** if the work went
|
|
through review (CEO/eng/codex/outside-voice). Translate review findings into
|
|
plain English. "We caught a stale-cache bug" beats "knobsHash() did not
|
|
include floorRatio in the v=2 hash input."
|
|
7. **`### Itemized changes`** (precision lives here). File paths, function
|
|
names, types, constants, line numbers. This section is for engineers who
|
|
need to know exactly what moved.
|
|
|
|
Voice rules (apply throughout):
|
|
- No em dashes (use commas, periods, "...").
|
|
- No AI vocabulary (delve, robust, comprehensive, nuanced, fundamental, etc.) or
|
|
banned phrases ("here's the kicker", "the bottom line", etc.).
|
|
- Real numbers, real file names, real commands AFTER the ELI10 lead. Not "fast"
|
|
but "~30s on 30K pages." In the ELI10 lead, "fast enough that you won't
|
|
notice" or "~30 seconds even on a big brain."
|
|
- Short paragraphs, mix one-sentence punches with 2-3 sentence runs.
|
|
- Connect to user outcomes: "the agent does ~3x less reading" beats "improved
|
|
precision."
|
|
- Be direct about quality. "Well-designed" or "this is a mess." No dancing.
|
|
|
|
**The smell test:** if someone who has never opened gbrain reads the first 150
|
|
words and walks away knowing what shipped and whether they care, the entry
|
|
passes. If they need to grep the codebase to follow along, rewrite the lead.
|
|
|
|
**Canonical examples in this CHANGELOG:** v0.35.6.0 (floor-ratio gate, written
|
|
ELI10-lead-first), v0.34.4.0 (embed stale fix wave). Use those shapes when in
|
|
doubt. Avoid the shape of entries that lead with internal constants or release
|
|
mechanics; those exist in older history but should not be the model for new
|
|
work.
|
|
|
|
Source material to pull from:
|
|
- CHANGELOG.md previous entry for prior context
|
|
- Latest `gbrain-evals/docs/benchmarks/[latest].md` for headline numbers (sibling repo)
|
|
- Recent commits (`git log <prev-version>..HEAD --oneline`) for what shipped
|
|
- Don't make up numbers. If a metric isn't in a benchmark or production data, don't
|
|
include it. Say "no measurement yet" if asked.
|
|
|
|
Target length: ~250-350 words for the summary. Should render as one viewport.
|
|
|
|
### "To take advantage of v[version]" block (required, v0.13+)
|
|
|
|
After the release-summary and BEFORE `### Itemized changes`, every `## [X.Y.Z]`
|
|
entry MUST include a human-readable self-repair block under the heading
|
|
`## To take advantage of v[version]`.
|
|
|
|
Why: `gbrain upgrade` runs `gbrain post-upgrade` which runs `gbrain apply-migrations`.
|
|
This chain has a known weak link — `upgrade.ts` catches post-upgrade failures as
|
|
best-effort (so the binary still works). When that chain silently fails, users end
|
|
up with half-upgraded brains. The self-repair block gives them a paste-ready
|
|
recovery path; the v0.13+ `~/.gbrain/upgrade-errors.jsonl` trail + `gbrain doctor`
|
|
integration close the loop.
|
|
|
|
Template (adapt the verify commands per release):
|
|
|
|
```markdown
|
|
## To take advantage of v[version]
|
|
|
|
`gbrain upgrade` should do this automatically. If it didn't, or if `gbrain doctor`
|
|
warns about a partial migration:
|
|
|
|
1. **Run the orchestrator manually:**
|
|
```bash
|
|
gbrain apply-migrations --yes
|
|
```
|
|
2. **Your agent reads `skills/migrations/v[version].md` the next time you interact with it.**
|
|
[One sentence on whether headless agents need manual action, or whether the
|
|
orchestrator already handled the mechanical side.]
|
|
3. **Verify the outcome:**
|
|
```bash
|
|
[release-specific verify commands, e.g. `gbrain graph ... --depth 2`]
|
|
gbrain stats
|
|
```
|
|
4. **If any step fails or the numbers look wrong,** please file an issue:
|
|
https://github.com/garrytan/gbrain/issues with:
|
|
- output of `gbrain doctor`
|
|
- contents of `~/.gbrain/upgrade-errors.jsonl` if it exists
|
|
- which step broke
|
|
|
|
This feedback loop is how the gbrain maintainers find fragile upgrade paths. Thank you.
|
|
```
|
|
|
|
**Skip this block** for patches that are pure bug fixes with zero user-facing action
|
|
(rare). If the release has a schema migration, data backfill, or new feature the
|
|
user needs to verify, the block is required.
|
|
|
|
The v0.13.0 entry in CHANGELOG.md is the canonical example.
|
|
|
|
### Itemized changes (the existing rules)
|
|
|
|
Below the release summary, write `### Itemized changes` and continue with the
|
|
detailed subsections (Knowledge Graph Layer, Schema migrations, Security hardening,
|
|
Tests, etc.). Same rules as before:
|
|
|
|
- Lead with what the user can now DO that they couldn't before
|
|
- Frame as benefits and capabilities, not files changed or code written
|
|
- Make the user think "hell yeah, I want that"
|
|
- Bad: "Added GBRAIN_VERIFY.md installation verification runbook"
|
|
- Good: "Your agent now verifies the entire GBrain installation end-to-end, catching
|
|
silent sync failures and stale embeddings before they bite you"
|
|
- Bad: "Setup skill Phase H and Phase I added"
|
|
- Good: "New installs automatically set up live sync so your brain never falls behind"
|
|
- **Always credit community contributions.** When a CHANGELOG entry includes work from
|
|
a community PR, name the contributor with `Contributed by @username`. Contributors
|
|
did real work. Thank them publicly every time, no exceptions.
|
|
|
|
### Reference: v0.12.0 entry as canonical example
|
|
|
|
The v0.12.0 entry in CHANGELOG.md is the canonical example of the format. Match its
|
|
structure for every future version: bold headline, lead paragraph, "numbers that
|
|
matter" with BrainBench-style before/after table, "what this means" closer, then
|
|
`### Itemized changes` with the detailed sections below.
|
|
|
|
## Version migrations
|
|
|
|
Create a migration file at `skills/migrations/v[version].md` when a release
|
|
includes changes that existing users need to act on. The auto-update agent
|
|
reads these files post-upgrade (see `docs/guides/upgrades-auto-update.md`)
|
|
and executes them.
|
|
|
|
**You need a migration file when:**
|
|
- New setup step that existing installs don't have (e.g., v0.5.0 added live sync,
|
|
existing users need to set it up, not just new installs)
|
|
- New SKILLPACK section with a MUST ADD setup requirement
|
|
- Schema changes that require `gbrain init` or manual SQL
|
|
- Changed defaults that affect existing behavior
|
|
- Deprecated commands or flags that need replacement
|
|
- New verification steps that should run on existing installs
|
|
- New cron jobs or background processes that should be registered
|
|
|
|
**You do NOT need a migration file when:**
|
|
- Bug fixes with no behavior changes
|
|
- Documentation-only improvements (the agent re-reads docs automatically)
|
|
- New optional features that don't affect existing setups
|
|
- Performance improvements that are transparent
|
|
|
|
**The key test:** if an existing user upgrades and does nothing else, will their
|
|
brain work worse than before? If yes, migration file. If no, skip it.
|
|
|
|
Write migration files as agent instructions, not technical notes. Tell the agent
|
|
what to do, step by step, with exact commands. See `skills/migrations/v0.5.0.md`
|
|
for the pattern.
|
|
|
|
## Migration is canonical, not advisory
|
|
|
|
GBrain's job is to deliver a canonical, working setup to every user on upgrade.
|
|
Anything that looks like a "host-repo change" — AGENTS.md, cron manifests,
|
|
launchctl units, config files outside `~/.gbrain/` — is a GBrain migration
|
|
step, not a nudge we leave for the host-repo maintainer. Migrations edit host
|
|
files (with backups) to make the canonical setup real. Exceptions: changes
|
|
that require human judgment (content edits, renames that break semantics,
|
|
host-specific handler registration where shell-exec would be an RCE surface).
|
|
Everything mechanical ships in the migration.
|
|
|
|
**Test:** if shipping a feature requires a sentence that starts with "in
|
|
your AGENTS.md, add…" or "in your cron/jobs.json, rewrite…", the migration
|
|
orchestrator should be doing that edit, not the user.
|
|
|
|
**The exception is host-specific code.** For custom Minion handlers
|
|
(host-specific integrations like inbox sweeps or third-party API scanners), shipping them as a
|
|
data file the worker would exec is an RCE surface. Those get registered in
|
|
the host's own repo via the plugin contract (`docs/guides/plugin-handlers.md`);
|
|
the migration orchestrator emits a structured TODO to
|
|
`~/.gbrain/migrations/pending-host-work.jsonl` + the host agent walks the
|
|
TODOs using `skills/migrations/v0.11.0.md` — stays host-agnostic, still
|
|
canonical.
|
|
|
|
|
|
## Schema state tracking
|
|
|
|
`~/.gbrain/update-state.json` tracks which recommended schema directories the user
|
|
adopted, declined, or added custom. The auto-update agent
|
|
(`docs/guides/upgrades-auto-update.md`) reads this during upgrades to suggest new schema additions without re-suggesting
|
|
things the user already declined. The setup skill writes the initial state during
|
|
Phase C/E. Never modify a user's custom directories or re-suggest declined ones.
|
|
|
|
## GitHub Actions SHA maintenance
|
|
|
|
All GitHub Actions in `.github/workflows/` are pinned to commit SHAs. Before shipping
|
|
(`/ship`) or reviewing (`/review`), check for stale pins and update them:
|
|
|
|
```bash
|
|
for action in actions/checkout oven-sh/setup-bun actions/upload-artifact actions/download-artifact softprops/action-gh-release gitleaks/gitleaks-action; do
|
|
tag=$(grep -r "$action@" .github/workflows/ | head -1 | grep -o '#.*' | tr -d '# ')
|
|
[ -n "$tag" ] && echo "$action@$tag: $(gh api repos/$action/git/ref/tags/$tag --jq .object.sha 2>/dev/null)"
|
|
done
|
|
```
|
|
|
|
If any SHA differs from what's in the workflow files, update the pin and version comment.
|
|
|
|
## GitHub releases (binary assets + self-update) — #3521
|
|
|
|
`.github/workflows/release.yml` publishes a GitHub release automatically for
|
|
**every VERSION bump that lands on master** (trigger: push to master touching
|
|
`VERSION`, plus `workflow_dispatch` for a manual first run or repair). No
|
|
manual tag push is part of the ship flow — the workflow reads `VERSION` (the
|
|
single source of truth), mints tag `v<VERSION>` at the pushed commit, titles
|
|
the release the same, uses that version's `CHANGELOG.md` entry as the notes
|
|
(`scripts/changelog-entry.sh`; falls back to a CHANGELOG link if the entry is
|
|
missing), and attaches the compiled binaries.
|
|
|
|
### The `latest-stable` tag
|
|
|
|
The **final step of the release job** force-advances the `latest-stable` tag to
|
|
the release commit (`git push origin "+${GITHUB_SHA}:refs/tags/latest-stable"`).
|
|
`latest-stable` is the single sanctioned distribution ref: the README paste
|
|
block, the `BOOTSTRAP_FOR_AGENTS.md` fetch URL, and
|
|
`bun install -g github:garrytan/gbrain#latest-stable` all reference it
|
|
permanently, so paste blocks copied into the wild never rot and there is no 404
|
|
window between VERSION landing and assets publishing.
|
|
`scripts/check-bootstrap-tag.sh` keeps the entry docs pinned to this ref.
|
|
|
|
Because it moves ONLY after binaries + provenance attestation have fully
|
|
published, a half-built release never advances it. If the tag-advance step
|
|
alone fails, re-advance by hand (a full workflow re-run would skip — the
|
|
release already exists with all assets):
|
|
|
|
```bash
|
|
git push origin "+refs/tags/v<VERSION>^{commit}:refs/tags/latest-stable"
|
|
```
|
|
|
|
### The `publish-template` job
|
|
|
|
After the release job, a `publish-template` job force-pushes the rendered
|
|
agent-workspace template repo (the GitHub "Use this template" door,
|
|
`vars.TEMPLATE_REPO`, default `garrytan/gbrain-agent-template`) from CI only —
|
|
no human pushes it by hand, so what adopters clone is exactly what this repo
|
|
reviewed. It is guarded three ways: the release above fully published; the
|
|
vendored tree `templates/bootstrap/template-repo/` exists (skip, never fail,
|
|
if not); and the `TEMPLATE_REPO_PAT` secret is configured (skip if not).
|
|
Before pushing, it regenerates the template tree
|
|
(`bun run scripts/generate-template-repo.ts`) and byte-diffs it against the
|
|
vendored copy — a mismatch fails the job; regenerate + commit the vendored
|
|
tree (`scripts/check-bootstrap-templates.sh` runs the same diff offline in
|
|
`bun run verify`).
|
|
|
|
**`TEMPLATE_REPO_PAT` scope:** a fine-grained PAT with `contents: write` on
|
|
the template repository ONLY — no other repositories, no other permissions.
|
|
Configure it as a repo secret; when absent, template publishing is disabled
|
|
and the job skips cleanly.
|
|
|
|
Why every bump, not selective: `gbrain check-update` resolves the latest
|
|
version from `VERSION` on master, while binary self-update
|
|
(`src/core/binary-self-update.ts`) downloads assets from `releases/latest`.
|
|
Any release that lags `VERSION` tells binary installs an upgrade exists that
|
|
self-update cannot apply. `releases/latest` must track `VERSION`.
|
|
|
|
Invariants:
|
|
|
|
- **Asset names are a contract.** The build matrix's `artifact:` names must
|
|
equal what `expectedAssetName()` in `src/core/binary-self-update.ts`
|
|
returns (`gbrain-darwin-arm64`, `gbrain-linux-x64` today). Adding a
|
|
platform means updating BOTH plus the version job's completeness check;
|
|
`test/release-workflow.test.ts` pins all of it.
|
|
- **Idempotent + self-repairing.** The version job skips when a release for
|
|
`v<VERSION>` already exists with all expected assets; a partial release
|
|
(tag but no release, or missing assets) is completed on re-run. Racing
|
|
master pushes queue via the `release` concurrency group — a skipped
|
|
intermediate version is fine, latest is what matters.
|
|
- **Historical tags are never rewritten.** Old 3-segment versions keep their
|
|
history; every new 4-segment `VERSION` mints a fresh tag.
|
|
- **Permissions stay scoped.** `contents: write` lives on the release job
|
|
only; everything else runs read-only.
|
|
- **Never advance `latest-stable` on a partial release.** The tag moves only
|
|
as the final release-job step, after every asset has published. Manual
|
|
re-advances must point at a fully published `v<VERSION>` release.
|
|
|
|
## PR descriptions cover the whole branch
|
|
|
|
Pull request titles and bodies must describe **everything in the PR diff against the
|
|
base branch**, not just the most recent commit you made. When you open or update a
|
|
PR, walk the full commit range with `git log --oneline <base>..<head>` and write the
|
|
body to cover all of it. Group by feature area (schema, code, tests, docs) — not
|
|
chronologically by commit.
|
|
|
|
This matters because reviewers read the PR body to understand what's shipping. If
|
|
the body only covers your last commit, they miss everything else and can't review
|
|
properly. A 7-commit PR with a body that describes commit 7 is worse than no body
|
|
at all — it actively misleads.
|
|
|
|
When in doubt, run `gh pr view <N> --json commits --jq '[.commits[].messageHeadline]'`
|
|
to see what's actually in the PR before writing the body.
|
|
|
|
## Community PR wave process
|
|
|
|
Never merge external PRs directly into master. Instead, use the "fix wave" workflow:
|
|
|
|
1. **Categorize** — group PRs by theme (bug fixes, features, infra, docs)
|
|
2. **Deduplicate** — if two PRs fix the same thing, pick the one that changes fewer
|
|
lines. Close the other with a note pointing to the winner.
|
|
3. **Collector branch** — create a feature branch (e.g. `garrytan/fix-wave-N`), cherry-pick
|
|
or manually re-implement the best fixes from each PR. Do NOT merge PR branches directly —
|
|
read the diff, understand the fix, and write it yourself if needed.
|
|
4. **Test the wave** — verify with `bun test && bun run test:e2e` (full E2E lifecycle).
|
|
Every fix in the wave must have test coverage.
|
|
5. **Close with context** — every closed PR gets a comment explaining why and what (if
|
|
anything) supersedes it. Contributors did real work; respect that with clear communication
|
|
and thank them.
|
|
6. **Ship as one PR** — single PR to master with all attributions preserved via
|
|
`Co-Authored-By:` trailers. Include a summary of what merged and what closed.
|
|
|
|
**Community PR guardrails:**
|
|
- Always AskUserQuestion before accepting commits that touch voice, tone, or
|
|
promotional material (README intro, CHANGELOG voice, skill templates).
|
|
- Never auto-merge PRs that remove YC references or "neutralize" the founder perspective.
|
|
- Preserve contributor attribution in commit messages.
|
|
|
|
## Checking out PRs from garrytan-agents
|
|
|
|
`garrytan-agents` is the AI-authored PR account and is NOT a collaborator on
|
|
this repo. Its PRs live in a fork, so GitHub Actions triggered by
|
|
`pull_request` events on those PRs do not receive base-repo secrets. Any CI
|
|
job that needs `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or similar will fail
|
|
with empty-env auth errors, regardless of what's set on the base repo. This
|
|
is a GitHub security default, not a config bug.
|
|
|
|
When the user says "check out <PR link>" and the PR is from `garrytan-agents`
|
|
(or any other non-collaborator fork), move the branch into the base repo
|
|
before running CI:
|
|
|
|
1. `gh pr checkout <N>` — pull down the fork's branch. Note the PR number and
|
|
head branch name (`gh pr view <N> --json headRefName --jq .headRefName`).
|
|
2. `git push origin HEAD:<branch-name>` — push the same branch to the base
|
|
repo (origin points at `garrytan/gbrain`, not the fork). This is the move
|
|
that gives CI access to secrets.
|
|
3. `gh pr close <N> --comment "moving to base-repo branch for secret access"`
|
|
— close the fork PR so the queue stays clean.
|
|
4. `gh pr create --base master --head <branch-name>` — open the replacement
|
|
PR from the base-repo branch. **Preserve the original PR's title and body
|
|
verbatim** (`gh pr view <N> --json title,body`); contributor attribution
|
|
moves to a `Co-Authored-By:` trailer if needed.
|
|
|
|
Why this over alternatives: adding `garrytan-agents` as a collaborator, or
|
|
flipping the repo-wide "send secrets to fork PRs" toggle, both broaden
|
|
secret distribution to every fork PR from that account or any fork. Moving
|
|
the branch keeps secret scope tight to just the one PR being shipped.
|
|
|