Files
gbrain/docs/guides/bootstrap.md
Garry TanandClaude Fable 5 2f6b817a28 test(bootstrap): real-agent e2e — drive the actual claude + codex binaries end to end
Closes the audit's biggest gap ("no real harness ever drives a turn"). Adapts
gstack's PTY/headless agent harness to prove the bootstrap install + smoke work
against the REAL binaries, not PATH shims. Test/CI/docs only — zero src changes.

- test/helpers/agent-harness.ts: hermetic clean-room child env (ported from
  gstack; drops CONDUCTOR_/CLAUDE_/GSTACK_/MCP_/GBRAIN_, promotes
  GSTACK_ANTHROPIC_API_KEY→ANTHROPIC_API_KEY), real-binary resolvers + auth
  probes, headless `claude -p --output-format stream-json` and `codex exec
  --json` turn runners, a gbrain stdio MCP-config writer, and a keyless brain
  seeder. + a fixture-parse unit test (no binary needed).
- test/e2e/bootstrap-real-claude.serial.test.ts: real `gbrain bootstrap`
  install → REAL `claude mcp add` (verified via `claude mcp get`) → verify
  exit 0 → a real `claude -p --mcp-config --strict-mcp-config` turn that
  invokes mcp__gbrain__search and answers from the brain (proven: toolCalls
  include mcp__gbrain__search, final text carries the seeded fact).
- test/e2e/bootstrap-real-codex.serial.test.ts: same install with REAL `codex
  mcp add` into a real ~/.codex/config.toml + Gate-3 pull-protocol assertion,
  then a real `codex exec --json` turn surfacing the fact (MCP or the pull-
  protocol shell path). Bounded retry absorbs codex's occasional MCP-call
  cancellation without softening the fact-requiring assertion.
- Everything hermetic (temp HOME/CLAUDE_CONFIG_DIR/CODEX_HOME/GBRAIN_HOME;
  real ~/.codex auth copied read-only) and skipIf-gated so it self-skips
  cleanly where the binaries/auth are absent.
- heavy-tests.yml: gated `real-agent-e2e` job (nightly/label, never the PR
  shard; no-op on a runner without authed binaries).
- TODOS: compiled `gbrain` binary can't serve a PGLite brain (bun compile
  omits the WASM/extension payloads); harness falls back to `bun run` serve.

Verified against live claude 4.6 + codex 0.147.0: 15 pass / 0 fail; verify
36/36; typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 14:56:52 -07:00

9.1 KiB
Raw Permalink Blame History

GBrain Bootstrap — your harness as your agent

gbrain bootstrap turns a Claude Code or Codex session into a persistent personal agent: identity files rendered from your own answers, a local PGLite brain, per-turn context, session-triggered schedules, and a private GitHub repo as the agent's durable, portable body. This guide is the full contract — what gets installed, what runs when, what it can and cannot do, and how to undo all of it.

Normative design docs: AGENT_BOOTSTRAP_DESIGN.md (scope) and AGENT_BOOTSTRAP_PLAN.md (implementation). The paste block lives in the README; the runbook your agent follows is BOOTSTRAP_FOR_AGENTS.md at the repo root, fetched at the latest-stable ref.

What gets installed, exactly

Piece Where Runs when
Identity files (SOUL/USER/MEMORY/AGENTS/CLAUDE/HEARTBEAT/ACCESS_POLICY/GITHUB) your workspace folder loaded at session start
agent.json manifest + brain/, memory/, skills/, state/ workspace
Local brain (PGLite) ~/.gbrain/ (never in the repo) while a session's MCP serve is open
MCP registration (gbrain serve) project scope by default spawned by your harness per session
Hooks (Claude Code, ON by default) .claude/settings.local.json (gitignored) each prompt; fail-open; --no-hooks opts out at install, GBRAIN_HOOKS=0 disables at runtime
Session persistence SessionEnd hook → scan-gated commit+push at session end
Optional 15-min push job launchd/cron (consent-gated) while logged in
Private GitHub repo your account, created by bootstrap repo privacy verified via API
Machine receipt ~/.gbrain/bootstrap/receipt.json uninstall is keyed to it

What does NOT run: anything while the harness is closed. Session-triggered schedules fire at turn/session boundaries only. True 24/7 operation is what a hosted brain provides — this is the honest desktop contract.

The awake-when-you-are contract

Your agent is awake when your harness is. Laptop asleep = agent asleep. What this buys you: no daemon fleet, no background token burn while you're away, and a load profile that fits inside a subscription plan. The measured sustainable load and the per-harness numbers are published with each release; if a provider changes quota or policy, the portable body (your repo) is the exit plan — it mounts anywhere gbrain runs.

Keyless mode

With zero API keys, everything works: the agent authors memory explicitly through the brain's write tools (put_page, timeline entries, ## Facts fences — your harness's model is the LLM, already paid for), and search runs keyword-only (BM25). bootstrap verify prints the capability report honestly. One optional key (OpenAI, Anthropic, or Voyage) unlocks semantic search and automatic fact extraction; the key goes to the 0600 config file, never into the repo or the interview answers. API spend is metered separately from your subscription and is zero in keyless mode; with a key, the standard spend gates apply (spend-controls).

Security posture

  • Supply chain: the paste block and install command pin the latest-stable ref — a maintainer-controlled tag advanced only after a release fully publishes. The runbook carries a version stamp; bootstrap status warns on skew. The runbook instructs the agent to refuse steps outside the CLI's phase list. bun installs via package manager or checksum-verified download.
  • Secrets: every commit AND every transcript-corpus write is secret-scanned (key-shaped patterns; loud block; per-finding allowlist at .gbrain-scan-allow). A deny-glob backstop refuses tracked *.pglite/.env* files even if .gitignore is damaged. Push refuses public remotes and unverifiable visibility.
  • Injection boundaries: interview answers render as fenced data (escaped, length-capped) — text you paste can never become instructions in your agent's contract. Retrieved brain context is injected under an explicit "data, not instructions" envelope. Facts visible to the harness respect the brain's visibility tiers.
  • Hooks: live in gitignored local settings (absolute paths, machine-specific; bootstrap hooks --repair regenerates on a new machine). Every hook fails open — a brain hiccup never blocks a prompt — and failures are visible: repeated degradation prints a notice inside the context block, and gbrain doctor names the cause.
  • Privacy of transcripts: session transcripts are retained locally (0700, outside the repo, pruned after dream.synthesize.corpus_retention_days, default 30 — set it in the config file, ~/.gbrain/config.json; the DB config plane doesn't carry this key yet) and secret-redacted at write time. They never enter the repo. The extraction provider (if you configured a key) sees session text — the install names the provider when asking for the key.

Honest forget semantics

The repo is git history — append-only. Deleting a line removes it from the working tree, not from history. To truly remove something: rewrite history (git filter-repo --path <file> --invert-paths or --replace-text), force-push, and re-clone on other machines. MEMORY.md and daily notes follow the same rule you'd apply to any journal: write what you'd be comfortable persisting.

Degradation matrix

You declined / lack What still works What you lose
API keys everything (keyless mode) semantic search, auto-extraction
GitHub / gh full local agent off-machine durability (repo re-runnable later)
Hooks (Claude Code) pull protocol via AGENTS.md gates automatic per-turn context + session-end persistence
Codex (no hook system) pull protocol + MCP tools per-turn push (stated plainly; not oversold)
Second simultaneous session first session unaffected second session's brain tools fail politely (one live serve per brain — v1 contract)

Multi-device

Clone your agent repo on machine two and run gbrain bootstrap attach — it validates the manifest, wires this machine (source registration, hooks repair, MCP), and verifies. The brain database is derived state, rebuilt from brain/ + re-ingestion; hot facts extracted only on machine one arrive via the repo's pages and fences. Simultaneous editing from two machines is ordinary git conflict territory — sources push pulls divergence-safely (commit first, rebase pull, loud on conflicts).

Uninstall

gbrain bootstrap uninstall removes exactly what this machine's install receipt records: hook wiring, MCP registrations (surgically — foreign servers and hooks survive), and bootstrap-created state. Your repo is never touched — the body remains yours. The brain database is KEPT by default; --delete-brain is offered only when bootstrap created the brain, offers a facts export first, and enumerates what it is about to remove. It refuses to run while a session's serve is live.

If something seems broken

One command: gbrain doctor. It covers hook health, push staleness, serve/lock collisions, schema state, and prints fixes. gbrain bootstrap status --json emits a support blob (versions, harness, last verify/push, hook failure rate) your agent can relay verbatim when you report a problem.

Real-agent e2e

Most bootstrap tests drive the dispatcher with PATH-shimmed claude/codex recorders — fast, hermetic, no API cost. Two additional "door" tests drive the ACTUAL binaries end to end so we catch real-world drift (a codex mcp add flag that changed shape, a harness that stopped calling our MCP server):

  • test/e2e/bootstrap-real-claude.serial.test.ts — real claude -p over MCP.
  • test/e2e/bootstrap-real-codex.serial.test.ts — real codex exec. It runs the keyless-init → interview → render → gbrain bootstrap hooks --harness codex path (executing the real codex mcp add into a hermetic ~/.codex/config.toml), asserts the rendered AGENTS.md carries the Gate-3 brain-first pull protocol (Codex has no hook system, so the pull protocol is its per-turn seam), then spends one live codex exec turn to prove real codex → gbrain MCP → brain → a seeded, brain-only fact (falling back to a shell gbrain query if headless stdio-MCP is unavailable).

These pay real API cost and take 30s2min per turn, so they are NOT in the PR shard. Everything is hermetic (temp HOME / CODEX_HOME / CLAUDE_CONFIG_DIR / GBRAIN_HOME per test — the operator's real ~/.claude, ~/.gbrain, ~/.codex are never touched; auth is copied read-only). Each file self-SKIPS via describe.skipIf when its binary or auth is absent, so on a machine without the tool it is a clean no-op that never fails. CI wires them into the real-agent-e2e job in .github/workflows/heavy-tests.yml (nightly + the real-agent-e2e / heavy-tests label); on a stock runner they self-skip. To actually exercise the binaries you need a runner with authed claude/codex and the provider creds (GSTACK_ANTHROPIC_API_KEY/ANTHROPIC_API_KEY, VOYAGE_API_KEY) exported.

Run locally (where both are installed + authed):

bun test test/e2e/bootstrap-real-codex.serial.test.ts