Files
gbrain/recipes/twilio-voice-brain.md
T
Garry TanandClaude Fable 5 d35c9c9e44 v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code (#3975)
* docs(designs): agent-bootstrap plan + design docs (normative, review-absorbed)

The scrubbed, in-repo sources of truth for the gbrain bootstrap wave:
AGENT_BOOTSTRAP_DESIGN.md (product scope/sequencing) and
AGENT_BOOTSTRAP_PLAN.md (implementation; all review-finding IDs inlined).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): format spec, question bank, identity templates, bundled assets

agent.json manifest (format_version 1, initialized sentinel) + machine-local
install receipt [CX2-1, CX2-12]; 12-question/6-required interview bank with
consent keys and a persist:false sink for the optional provider key [CX2-13];
ten {{TOKEN}} identity templates (generic, adapted to gbrain ops — gates call
recall/query/put_page, write-through-ops rule, keyless agent-authored facts,
silence contract); assets embedded compiled-binary-safe via file-type imports
[ENG-6].

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): runbook, README paste block, bootstrap guide, TODOS entries

BOOTSTRAP_FOR_AGENTS.md (agent-driven install runbook: CLI phase list is the
source of truth, never-invent rules, Codex approvals preflight, keyless posture,
failure-modes table, version stamp for the skew check); README gains the
full-agent paste block pinned to latest-stable inside the Claude Code/Codex
quick start (memory-only tier stays); docs/guides/bootstrap.md carries the full
install/security/consent/degradation/uninstall contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(designs): spike instrument for the bootstrap wave (build order 0)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): interview + render engines

Interview gate with read-back confirm-hash (any later answer change clears the
confirmation — the hostile single-batch case is structurally impossible),
per-answer provenance, caps + escaping at set time, config-sink routing for the
provider key; renderer with hard-fail token sweep, subordinate fencing of
principal input, never-clobber + backups, deterministic minimal mode for the
template repo, scaled byte floors. 58 unit tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): private-repo lifecycle — repo create, attach, uninstall, run lock

gh-gated private repo creation with API-verified privacy (rate-limit distinct
from public), refuse-foreign-origin with attach as the sanctioned path, atomic
bootstrap mutex (pid liveness + age + token), receipt-keyed uninstall that
never wholesale-deletes the gbrain home and only offers --delete-brain for a
brain it created; read-only PGLite lock probe (never opens the engine).
54 unit tests, injectable exec seam throughout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(release): latest-stable ref, template-repo publish job, bootstrap CI guards

release.yml advances the latest-stable tag only after assets publish (the paste
block's permanent ref — copies in the wild never rot) and gains a PAT-gated
publish-template job verified against the vendored tree; two skip-graceful
guards (sanctioned-ref + runbook stamp; template/token bijection + placeholder
assertion + generator byte-diff) wired into verify; README + runbook re-admitted
to the CI cache hash; vendored deterministic template tree generated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(context): IPC v2 turn_context + 8KB assembly + visibility resolver + session identity

Discriminated-union IPC with handler map, protocol echo (stale-serve detection),
shared-secret gate, server-side source binding, per-kind budgets; turn-context
assembly (reflex pointers + volunteered pages + world-only hot facts) under a
data-not-instructions envelope trimmed to the harness's 10KB hook-output cap;
facts.default_visibility resolved through one helper at all four sites (explicit
caller wins, typos fail closed); typed sessionId threads _meta.session_id into
the hot-memory cache key. 50 new tests; 180 adjacent tests confirmed green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(persistence): secret-scan, gbrain sources push, durability unification

Pattern secret scanner (own runtime allowlist, redacted previews, corpus-write
redaction mode); sources push runs the whole scan→stage→commit→pull→push
sequence under one cross-platform lock (mkdir-atomic, pid+age+token) with a
deny-glob backstop, commit-first divergence-safe pull, refuse-unverifiable
visibility, and push-status telemetry; gbrain-home choke point unifies
GBRAIN_HOME semantics with config (0700); durability is parent-repo-aware and
rotates its push log at 0600. 35 new tests; 200 existing green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(sources): harden/pull gates accept sources inside a parent git repo

The bootstrap workspace registers brain/ (a subdirectory) as the source; the
durability core already resolves the repo root, so the command gates now check
inside-a-repo rather than .git-right-here [CX2-3].

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(serve): resident maintenance sweep + keyless capability probe

The lock-owning serve process now closes the persistence loop: startup (3s
post-connect, best-effort, unref'd) and idle (10-min quiet intervals through
the injectable timer seam) sweeps run facts-fence reconciliation, deterministic
link/timeline extraction over recent workspace pages, and spend-gated corpus
ingest (skipped keyless — agent-authored fences cover it). gbrain sweep --once
is the trusted CLI seam bootstrap verify uses. Capability probe renders the
honest keyless/keyed report. Full reuse of the cycle extractor + extract cores;
26 new tests, neighbors green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(hooks): engine-free gbrain hook command, settings writers, transcript parser

Four hook events (session-start digest + crashed-session recovery push,
user-prompt turn-context injection under an 800ms deadline and the 10KB cap,
stop buffers, session-end corpus write with redaction/retention/dedup +
best-effort push); structural JSON settings merger keyed by a _gbrain marker
(foreign hooks and permissions survive); dated host-spec registry; Claude Code
.jsonl parser as a spec-target with a scrubbed 7-shape fixture. Heartbeat is
counters-only by construction. 59 tests; zero engine modules in the import
graph.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(bootstrap): cross-link the full-agent path from the connection docs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): dispatcher, verify, status — the command assembled

gbrain bootstrap {status,interview,render,repo,hooks,verify,uninstall,attach}:
engine-free except verify (owns its engine, in-process sweep — no live-serve
conflict); phase list is the TS source of truth with install.jsonl telemetry
and the support blob; verify's fail-soft check suite covers the real write path
(put_page → write-through file → sweep → graph floor → recall), passes keyless,
persists snapshots, and ends with the first-run tour. cli.ts wired per the
three-touchpoint rule; doctor gains the bootstrap check group (silent on
machines with no bootstrap state). 28 new tests; 353 adjacent green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: KEY_FILES bootstrap cluster + CLAUDE.md dispatcher row (+ build:llms)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): e2e pins — hook-under-live-serve, attach, degraded modes, compiled binary, Docker harness

The permanent pins: a real serve holds the PGLite lock while the engine-free
hook completes (and a direct engine open provably throws LiveServeLockError);
stale-socket fail-open; machine-2 attach with marker-keyed hook repair;
decline-everything installs verify green with every degradation named; the
compiled binary renders bundled templates in an empty cwd. Offline Docker
harness (networkless, read-only) gated into heavy-tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): register doctor check categories + system-of-record allow comments

The six bootstrap doctor checks join OPS_CHECK_NAMES; the sweep's batch link/
timeline inserts carry the explicit extract-path allow comments (the sweep IS
the extraction path for workspace pages).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): shard wedge cap tracks suite growth (1500s -> 1800s)

At ~9000 tests a healthy shard finished at 1466s and two progressing shards
were false-killed at the old cap; 1800s restores ~25% headroom over the
slowest observed healthy shard. Real hangs still hit it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): cache-hash policy — README + runbook edits must invalidate [C2]

The old deny-list assertion predates the paste block; README.md and
BOOTSTRAP_FOR_AGENTS.md are policy-doc re-admissions now, so their edits must
change the hash (a paste-block edit shipping under a cached green was the C2
hole).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): classify post-suite exit-hangs as warn-pass; file the leak forensics

A shard killed by the wedge watchdog with every assigned file started and zero
fail markers did all its work and leaked a handle at exit — pre-existing and
master-reproducible (P1 TODO carries the full bisect forensics). Bun's per-test
timeout turns a hung test into a (fail), so the classifier cannot mask one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test): shard cap 2400s — the count-balanced heavy shard needs it under contention

Observed: the heavy shard still progressing 22s before an 1800s kill while
siblings finish at 1150-1550s (split balances file count, not weight). Filed
the load-sensitive WAL-repair flake (pre-existing, master's v0.42.75.0 wave).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): quarantine env-mutating suites to the serial lane

check-test-isolation R1: six new files mutate GBRAIN_HOME/env at module scope —
the serial lane (one process per file) is the guard's prescribed home for them.
All 114 tests pass post-rename.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): per-harness install sections — Codex, Claude Code, then OpenClaw/Hermes

Each harness gets its own complete paste-block section (desktop app first,
terminal noted — Claude Code CLI is the identical harness; Codex CLI works
pull-based today); the OpenClaw/Hermes platform path keeps equal weight with
its one-click deploys and INSTALL_FOR_AGENTS block intact; memory-only and
remote-connect tiers consolidated under 'Lighter ways in'. Supersedes the
review's D5 ordering by user direction; stale heading references updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): Codex as the recommended first step; OpenClaw/Hermes framed as-intended, high-cost

The install section now routes newcomers explicitly: Codex first
(subscription-priced, nothing to deploy), OpenClaw/Hermes as GBrain used the
way it was designed — always on, at real server + API cost.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: doc-audit code chasers — broken recovery hints, stale op description, auto_link key

Four small code fixes surfaced by the markdown accuracy audit:
- doctor's auto-RLS recovery hint pointed at `apply-migrations --force-retry 35`,
  which cannot work (--force-retry targets the vX.Y.Z orchestrator registry, not
  the numeric schema MIGRATIONS array). Hint now points at the recreate SQL in
  docs/guides/rls-and-you.md; test pins against regression.
- v0_11_0 migration printed the same broken-mechanism class of hint
  (`config set minion_mode` writes DB config nothing reads); now names
  `apply-migrations --mode` + preferences.json, the real setter.
- submit_job's op description hardcoded a stale handler list; now points at
  registerBuiltinHandlers as the source plus the --follow discovery trick.
- `auto_link` added to KNOWN_CONFIG_KEYS: read by link-extraction, reconcile-links,
  and sweep, and documented as the off-switch in brain-ops/maintain, but the
  allowlist rejected `config set auto_link false`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: repo-wide accuracy + MECE reform from the 9-bucket markdown audit

A code-grounded audit of every markdown file (root, architecture, guides,
mcp, tutorials, docs-root, operations/eval/designs, skills, recipes) followed
by a fix wave with per-bucket ownership. Four classes of change:

Accuracy — every documented command/flag verified against src/ before writing:
dead commands replaced with working ones (pages purge-deleted, jobs watch
--follow, gbrain restore, import-based Obsidian flow, space-separated --scopes,
real thin-client recipes, working isolation verification, real supervisor
restart procedure, curl-based ngrok health check, real minion_mode setter);
count drift fixed with rot-proof phrasing (100+ ops, 50+ bundled skills via
skills/manifest.json, 140+ engine methods, KNOBS_HASH_VERSION pointer instead
of hardcoded versions); stale claims corrected (search-mode defaults, RETRIEVAL
pipeline order incl. autocut, sentinel rules, refusal-list mechanism, engine
snapshot, shard cap 2400s + EXIT-HANG classifier in TESTING.md, latest-stable +
publish-template documented in RELEASING.md as release.yml promises).

MECE — one home per concept, pointers elsewhere: test isolation → TESTING.md;
OAuth registration + --bind/--public-url lore → DEPLOY.md; mode bundles →
guides/search-modes.md (the home the CLAUDE.md dispatcher always promised);
merge contract → schema-packs.md; WAL ladder → ENGINES.md; quiet-hours →
quiet-hours.md; capture taxonomy → entity-detection.md; person-page taxonomy →
compiled-truth.md; brain-first protocol → brain-first-lookup.md; refresh
semantics → refresh-algorithm.md; KEY_FILES.md deduplicated (58 extension
entries merged, one entry per file); infra-layer.md rewritten as a pointer page.

Privacy — placeholder sweep across guides, docs, skills, and recipes per the
iron rule; per-release narration stripped from reference docs (current-state
prose only).

Bootstrap coverage — AGENTS.md pointer, RESOLVER routing row, INSTALL.md path,
tutorial cross-links, keyless-mode sections in spend-controls/headless-install.

skills.lock.json regenerated; llms.txt/llms-full.txt rebuilt. Gates: verify
36/36, typecheck clean, doctor 96/96, skills-integrity + resolver + build-llms
+ config-set + migrations all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme): refresh production-brain stats to current brain-repo counts

155,795 pages / 24,589 people / 5,340 companies, counted from the brain
repo's current HEAD; the "100K-page brain" framing moves to 150K to match.
Cron-fleet count unchanged (its store lives on the deployment host, not in
the repos available for verification).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scan,push): modern OpenAI/Voyage key patterns, scan staged blobs not disk

- secret-scan matches sk-proj-/sk-svcacct-/sk-None- and pa- Voyage keys
  (the bare sk- pattern missed every current OpenAI key format).
- workspacePush stages first, then scans the staged index blobs via
  git cat-file, closing the scan-then-stage TOCTOU where a file changed
  between snapshot and commit shipped unscanned.
- shared binary-sniff helper, memoized glob regexes, atomic push-status
  write, and tests for pull_conflict + gitignored deny-match paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): regenerate flag registry for new commands, harden shard classifier + release token

- cli-flag-registry.generated.ts regenerated: bootstrap/hook/sweep and
  sources push --message/--allow-unverified-remote were missing, so the
  strict #2185 validator rejected real invocations and skipped the new
  commands entirely.
- EXIT-HANG shard classifier now requires every assigned file to have
  started before warn-passing a watchdog kill (was fail-open).
- release.yml passes TEMPLATE_REPO_PAT via http.extraheader, off the argv.
- compiled-binary e2e fails loud in CI instead of a silent permanent skip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(hook,sweep,ipc): non-blocking hook pushes, bounded sweep + cache, source-bound resolve

- session-start/session-end no longer run synchronous git + inline push
  inside their self-deadline; a detached child does the push and the hook
  returns immediately (blocked Claude Code startup for minutes on a dirty
  tree before).
- serve sweep drops the unbounded listAllPageRefs, resolves only candidate
  targets, claims corpus files atomically (no double-LLM-spend race), and
  caps the fence LIKE scan; heartbeat writes are O_APPEND with rare compaction.
- hot-memory cache evicts expired entries and bounds entry count (the key is
  caller-controlled via _meta.session_id).
- v1 resolve IPC honors boundSourceId like turn_context; turn-context runs
  its arms concurrently. doctor reads push/heartbeat thresholds from hook.ts.
- new tests: doctor bootstrap checks, hook push-gate + deadline, concurrent
  sweep claims, cache eviction, bound-source resolve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): origin-ownership gate, world visibility, collision-safe source id, consent + templates

- repo adoption requires an exact receipt repo_url match or authed-owner
  check (undefined repo_url was a wildcard); create verifies privacy BEFORE
  the first push.
- verify sets facts.default_visibility=world if unset, so agent-authored
  facts surface in per-turn context (they defaulted private before).
- source_id derives a path-hash suffix when 'workspace' is taken by another
  checkout; every consumer reads manifest.source_id.
- skipped HOOKS_CONSENT now declines (was falling through to default yes);
  --minimal refuses on an initialized manifest; tilde fences escaped.
- MCP registration pins --surface full; status hard-fails a public origin
  (template door); receipt writers guard against newer/corrupt receipts;
  uninstall only claims brain-deleted after a real rm.
- templates ship jobs disabled + provider-consent + support-relay lines;
  soul-audit re-runs over the shared interview bank. TODOS: 11 follow-ups.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(security): close adversarial-review findings — scan fails closed, whole-PEM redaction, bound repo push

Cross-model adversarial pass (Claude + Codex) on the bootstrap wave:

- secret scan fails CLOSED: an unreadable, oversized, or binary staged blob
  now blocks the push (blocked_unscannable, exit 5) instead of committing
  unscanned; only a confirmed staged deletion is skipped. This was the
  headline "block secrets before they leave the machine" property failing open.
- private-key redaction spans the whole PEM block (header+body+footer), not
  just the header line — the base64 body no longer survives into the corpus
  the sweep sends to an extraction provider.
- bootstrap repo commits the workspace (secret-scan-gated) before the first
  push and verifies the remote actually received it, so a push-fail retry
  can't adopt an empty remote as success.
- privacy verify is re-bound to origin immediately before push (a concurrent
  origin rewrite between verify and push is refused).
- session-end corpus write is atomic and clears the stale ingested/in-progress
  sidecars so a resumed session's appended transcript is re-ingested.
- public-origin refusal enforced at render (not only status); MCP "already
  registered" is verified to target this workspace, not blessed blindly;
  verify probe cleanup scopes deletes to its own slugs, not a token substring;
  allowlist fingerprint floor raised 8→16 hex.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.45.0.0 feat(bootstrap): paste-in personal-agent install for Codex + Claude Code

Turns a Codex or Claude Code session into a persistent personal agent:
interview-rendered identity files, a local PGLite brain, per-turn context via
serve IPC (Claude Code hooks / Codex pull protocol), session-triggered
persistence, and a private GitHub repo as the agent's portable body. Keyless-
first (the harness model is the LLM; one optional key adds embeddings +
extraction). New `gbrain bootstrap` command family + `gbrain hook` + `gbrain
sweep`; doctor bootstrap health checks; latest-stable distribution ref +
template-repo publish job. Opt-in, additive — existing installs untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: regenerate flag registry for security-fix flags; drop fabricated gbrain capabilities doc ref

CI caught two real failures under the merged state:
- the flag registry lagged the blocked_unscannable/exit-5 flags the security
  round added, tripping the #2185 freshness guard.
- headless-install.md described the keyless capability report as a
  `gbrain capabilities` command, which the #3502 doc-command resolver
  rejects — reworded to prose (the real surface is bootstrap verify's report).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: sync KEY_FILES + bootstrap plan to security-fix behavior

Cross-referenced the security-fix round against the reference docs and
corrected the drift those commits introduced:

- workspace-push.ts entry: stage-FIRST-then-scan order (the TOCTOU fix),
  fail-closed blocked_unscannable, and the sources-push status -> exit-code map.
- hooks.ts entry: MCP registration pins `serve --surface full`.
- hook.ts entry: session-start/session-end pushes run in a detached child
  (non-blocking); atomic corpus write clears stale sidecars.
- bootstrap.ts entry: render hard-refuses a public origin (template door).
- verify.ts entry: source_id collision resolution (workspace-<path-hash>).
- AGENT_BOOTSTRAP_PLAN as-shipped delta note for the scan/stage reorder.

llms bundle unchanged (KEY_FILES is link-only); build:llms and
test/build-llms.test.ts green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): silence SC2016 on the intentional askpass literal in release.yml

The one-shot GIT_ASKPASS script must contain literal $1 and
$TEMPLATE_REPO_PAT so they expand when /bin/sh runs it at git's credential
prompt, not when the outer shell writes the file — single quotes are correct.
Add a scoped shellcheck disable so actionlint passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(skills): declare 'bootstrap my data' trigger in cold-start frontmatter

The doc reform added 'bootstrap my data' to cold-start's RESOLVER.md row (to
disambiguate data-bootstrap from agent-bootstrap) but not to the skill's own
frontmatter triggers, tripping the RESOLVER↔frontmatter round-trip contract
(resolver.test.ts). Declare it; regenerate skills.lock.json.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): default per-turn hooks + search mode ON without a prompt

Installing gbrain for your coding agent IS the consent for the behaviors that
make it work, so stop re-litigating them with install-time questions whose
"no" defeats the product:

- Per-turn hooks (Claude Code) install ON by default — no prompt. Off-ramps:
  `--no-hooks` at install, `GBRAIN_HOOKS=0` at runtime, `bootstrap uninstall`.
  The "hooks installed" line now surfaces the kill switch so default-on is
  never silent. A persisted HOOKS_CONSENT=no (interview --skip) still declines.
- Search mode defaults to `balanced` silently (nobody knows the modes at
  install; `gbrain search modes` changes it any time).
- MCP scope stays the ONE deliberate prompt — project vs user is a real
  cross-repo privacy choice, not friction.

Marks the two consents `silent: true` in the question bank (new QuestionSpec
field), rewrites the runbook phases so the agent no longer asks them, adds the
`--no-hooks` flag (+ registry regen), and adds default-on / opt-out tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): real end-to-end coverage — cross-session recall, per-turn content, Codex door, realistic corpus

Closes the seven e2e gaps a coverage audit surfaced: the plumbing was
well-unit-tested but the product claims ("Codex works, context shows up every
turn with real content, it remembers across restarts, machine two recovers,
Postgres works") were unproven end to end. Test-only wave — zero src changes.

- Hermetic synthetic corpus (test/fixtures/bootstrap-corpus/ + a loader helper):
  12 interlinked pages (52 edges, timelines), 12 world/private beliefs, 8 gold
  queries — curated from the gbrain-evals synthetic corpora, 100% placeholder
  names, so recall is asserted on a real multi-entity brain instead of a
  2-node self-planted probe.
- GAP1 magic moment: author a fact via the real write path, disconnect the
  engine, reopen against the same DB, recall it — a real session boundary, not
  verify.ts's same-connection SQL read-back. Plus a source-isolation assertion.
- GAP2 per-turn content: hook-under-serve Pin 1 now seeds a known fact and
  asserts its text lands in the injected block AND private beliefs never do
  (was: empty brain, empty_block accepted as a pass).
- GAP3 Codex door: assert the rendered AGENTS.md carries the Gate-3 brain-first
  pull protocol; make the fake codex shim implement `mcp get` so the [FIX7]
  target-verification can actually fail; the Docker cold-machine harness now
  exercises the hooks/MCP registration step instead of skipping it.
- GAP4 corpus recall: turn-context + verify graph-floor/qrels run on the real
  multi-entity brain with real edges.
- GAP5 attach: machine-two now re-ingests the cloned brain/ into a fresh DB and
  recalls a fact authored only on machine one — the multi-device payoff.
- GAP6 keyed + Postgres (env-gated): real embeddings prove semantic recall a
  paraphrase query can reach but keyless BM25 cannot; bootstrap verify drives a
  real Postgres engine (skipIf DATABASE_URL/keys absent).
- GAP7 persistence: session-end runs the REAL push (not the mocked seam) to a
  local bare remote and the remote receives the content; a planted secret is
  blocked at the gate; the 15-min cron installs and fires a scan-gated push.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(bootstrap): real-agent e2e — drive the actual claude + codex binaries end to end

Closes the audit's biggest gap ("no real harness ever drives a turn"). Adapts
gstack's PTY/headless agent harness to prove the bootstrap install + smoke work
against the REAL binaries, not PATH shims. Test/CI/docs only — zero src changes.

- test/helpers/agent-harness.ts: hermetic clean-room child env (ported from
  gstack; drops CONDUCTOR_/CLAUDE_/GSTACK_/MCP_/GBRAIN_, promotes
  GSTACK_ANTHROPIC_API_KEY→ANTHROPIC_API_KEY), real-binary resolvers + auth
  probes, headless `claude -p --output-format stream-json` and `codex exec
  --json` turn runners, a gbrain stdio MCP-config writer, and a keyless brain
  seeder. + a fixture-parse unit test (no binary needed).
- test/e2e/bootstrap-real-claude.serial.test.ts: real `gbrain bootstrap`
  install → REAL `claude mcp add` (verified via `claude mcp get`) → verify
  exit 0 → a real `claude -p --mcp-config --strict-mcp-config` turn that
  invokes mcp__gbrain__search and answers from the brain (proven: toolCalls
  include mcp__gbrain__search, final text carries the seeded fact).
- test/e2e/bootstrap-real-codex.serial.test.ts: same install with REAL `codex
  mcp add` into a real ~/.codex/config.toml + Gate-3 pull-protocol assertion,
  then a real `codex exec --json` turn surfacing the fact (MCP or the pull-
  protocol shell path). Bounded retry absorbs codex's occasional MCP-call
  cancellation without softening the fact-requiring assertion.
- Everything hermetic (temp HOME/CLAUDE_CONFIG_DIR/CODEX_HOME/GBRAIN_HOME;
  real ~/.codex auth copied read-only) and skipIf-gated so it self-skips
  cleanly where the binaries/auth are absent.
- heavy-tests.yml: gated `real-agent-e2e` job (nightly/label, never the PR
  shard; no-op on a runner without authed binaries).
- TODOS: compiled `gbrain` binary can't serve a PGLite brain (bun compile
  omits the WASM/extension payloads); harness falls back to `bun run` serve.

Verified against live claude 4.6 + codex 0.147.0: 15 pass / 0 fail; verify
36/36; typecheck clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): real-agent-e2e job — bash array + --timeout (actionlint SC2086 + bun-test-timeout guard)

The real-agent-e2e job's file loop used an unquoted $FILES (SC2086) and ran
`bun test` without --timeout (check-bun-test-timeout guard). Switch to a bash
array and add --timeout=600000 (real-agent turns are slow; the tests self-skip
without authed binaries so it's a no-op elsewhere).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pglite): embed WASM + extension assets so the compiled binary can serve

A `bun build --compile` gbrain binary could not `serve` a PGLite brain: the
compile bundles JS but not PGLite's runtime payload (pglite.wasm, initdb.wasm,
pglite.data, vector/pg_trgm tarballs), so `serve` on PGLite died with a
bunfs/ENOENT. Now the assets ride inside the binary.

- src/core/pglite-embedded-assets.ts: embeds the five assets via
  `import … with { type: 'file' }` (the ENG-6 idiom) and exposes
  getEmbeddedPgliteOptions() → { pgliteWasmModule, initdbWasmModule, fsBundle,
  extensions:{vector,pg_trgm} }. WASM/fsBundle are consumed as bytes; the two
  extension tarballs are materialized to a content-addressed temp file (atomic,
  size-verified reuse) because PGLite reads them via fs.createReadStream, which
  cannot read a /$bunfs path. Unconditional (works in bun-run and compiled),
  so no fragile mode branch.
- src/core/pglite-engine.ts: static-import getEmbeddedPgliteOptions (engine
  path stays static per the engine-dynamic-import invariant); spread into both
  PGlite.create sites (initial + WAL-repair retry). The bunfs classifier stays
  as a backstop but no longer fires for a correct binary.
- scripts/check-pglite-embedded.sh (+ smoketest): compiles a focused binary and
  asserts it boots PGLite, CREATE EXTENSION vector/pg_trgm, and round-trips a
  page — wired into `bun run verify` (now 37 checks), check:all, and
  check:pglite-embedded. Fail-soft only when compile is unavailable.
- agent-harness.ts probeCompiledPglite now passes → the real-agent e2e uses the
  fast compiled MCP server. TODOS: the P2 "can't serve PGLite" item is closed.

Verified: fresh compiled binary ran `search`/`query` against a PGLite brain and
returned the seeded row (no bunfs/ENOENT); verify 37/37; pglite-engine 120/0
source-mode; typecheck clean; engine-dynamic-import + parity guards pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 16:42:11 -07:00

35 KiB

id, name, version, description, category, requires, secrets, health_checks, setup_time, cost_estimate
id name version description category requires secrets health_checks setup_time cost_estimate
twilio-voice-brain Voice-to-Brain (DEPRECATED — see agent-voice) 0.8.2 DEPRECATED. New installs use `gbrain integrations install agent-voice` — the copy-into-host-repo paradigm with WebRTC-first browser client + Mars/Venus personas + read-only tool router. This recipe stays as a redirect for existing Twilio installs; it is frozen (no longer updated) and will be removed in a future release. sense
ngrok-tunnel
name description where
TWILIO_ACCOUNT_SID Twilio account SID (starts with AC) https://www.twilio.com/console — visible on the main dashboard after login
name description where
TWILIO_AUTH_TOKEN Twilio auth token (click "Show" next to the SID on the dashboard) https://www.twilio.com/console — click "Show" under Auth Token on the main dashboard
name description where
OPENAI_API_KEY OpenAI API key (needs Realtime API access enabled on your account) https://platform.openai.com/api-keys — click "+ Create new secret key", copy immediately (you can't see it again)
type url auth auth_user auth_pass label
http https://api.twilio.com/2010-04-01/Accounts/$TWILIO_ACCOUNT_SID.json basic $TWILIO_ACCOUNT_SID $TWILIO_AUTH_TOKEN Twilio API
type url auth auth_token label
http https://api.openai.com/v1/models bearer $OPENAI_API_KEY OpenAI API
30 min $15-25/mo (Twilio number $1-2 + voice $0.01/min, OpenAI Realtime $0.06/min input + $0.24/min output)

Voice-to-Brain: Phone Calls That Create Brain Pages

⚠️ DEPRECATED. New installs should use the agent-voice recipe — a WebRTC-first voice agent with Mars + Venus personas, copy-into-host-repo install paradigm, and read-only tool router. This recipe stays as a redirect for operators with existing Twilio installs. It is frozen: no longer updated, and will be removed in a future release once existing installs have migrated.

Migration: gbrain integrations install agent-voice --target <your-repo> copies a working reference into your host agent repo where you own the edits. The new recipe includes a Twilio bridge in code/lib/twilio-bridge.mjs for operators who still want phone inbound, but the WebRTC /call?test=1 flow is the headline experience.

Where this recipe's content now lives (canonical homes):

  • ngrok tunnel setup + watchdog → ngrok-tunnel.md
  • post-call transcript pipeline → the voice-post-call skill in the agent-voice bundle
  • voice production patterns (unicode sanitize, PII scrub, identity-first prompt, conversation timing) → shipped as code in the agent-voice bundle

The copies below are kept only so existing installs have a self-contained runbook; they may drift from current gbrain behavior. Treat the homes above as authoritative.

Call a phone number. Talk. A structured brain page appears with entity detection, cross-references, and a summary posted to your messaging app.

IMPORTANT: Instructions for the Agent

You are the installer. This recipe is written for YOU (the AI agent) to execute on behalf of the user. Follow these instructions precisely.

Why sequential execution matters: Each step depends on the previous one:

  • Step 1 validates prerequisites. If GBrain isn't configured, nothing else works.
  • Step 2 collects credentials. If a credential is wrong, Steps 5-7 will silently fail.
  • Step 3 creates the ngrok tunnel. Step 5 needs the ngrok URL for the Twilio webhook.
  • Step 5 configures Twilio. Step 7 (smoke test) needs Twilio configured to reach your server.

Do not skip steps. Do not reorder steps. Do not batch multiple steps.

Stop points (MUST pause and verify before continuing):

  • After Step 1: all prerequisites pass? If not, fix before proceeding.
  • After each credential in Step 2: validation passes? If not, help the user fix it.
  • After Step 6: health check passes? If not, debug before smoke test.
  • After Step 7: brain page created? If not, troubleshoot before declaring success.

When something fails: Tell the user EXACTLY what failed, what it means, and what to try. Never say "something went wrong." Say "Twilio returned a 401, which means the auth token is incorrect. Let's re-enter it."

Architecture

Two pipeline options:

Option A: OpenAI Realtime (turnkey, simpler)

Caller (phone)
  ↓ Twilio (WebSocket, g711_ulaw audio — no transcoding)
Voice Server (Node.js, your machine or cloud)
  ↓↑ OpenAI Realtime API (STT + LLM + TTS in one pipeline)
  ↓ Function calls during conversation
GBrain MCP (semantic search, page reads, page writes)
  ↓ Post-call
Brain page created (meetings/YYYY-MM-DD-call-{caller}.md)
Summary posted to messaging app (Telegram/Slack/Discord)

Option B: DIY STT+LLM+TTS (full control, production-grade)

Caller (phone or WebRTC browser)
  ↓ Twilio WebSocket OR WebRTC
Voice Server (Node.js)
  ↓ Deepgram STT (streaming speech-to-text, speaker diarization)
  ↓ Claude API (streaming SSE, sentence-boundary dispatch)
  ↓ Cartesia / OpenAI TTS (text-to-speech, low latency)
  ↓ Function calls during conversation
GBrain MCP (semantic search, page reads, page writes)
  ↓ Post-call
Brain page + audio upload + transcript storage

Why v2 (Option B)? OpenAI Realtime is a black box — you can't control STT quality, swap LLMs, or debug audio issues. The DIY stack gives you transparent Deepgram+Claude+TTS with full control over each stage. Trade-off: more integration work, but you own the pipeline.

Production-tested v2 architecture (pipeline.mjs, ~250 lines):

  • Streaming SSE from Claude with sentence-boundary TTS dispatch
  • 20-turn conversation history cap (prevents context bloat)
  • Reconnect logic with exponential backoff on STT/TTS disconnects
  • Periodic keepalives to prevent WebSocket timeout
  • Audio endpointing for natural turn-taking
  • Smart VAD (Silero) as default with push-to-talk fallback

Opinionated Defaults

These are production-tested defaults from a real deployment. Customize after setup.

Caller routing (prompt-based, enforced server-side):

  • Owner: OTP challenge via secure channel, then full access (read + write + gateway)
  • Trusted contacts: callback verification, scoped write access
  • Known contacts (brain score >= 4): warm greeting by name, offer to transfer
  • Unknown callers: screen, ask name + reason, take message

Security:

  • Twilio signature validation on /voice endpoint (X-Twilio-Signature header)
  • Unauthenticated callers never see write tools
  • Caller ID is NOT trusted for auth (OTP or callback required)

Setup Flow

Step 1: Check Prerequisites

STOP if any check fails. Fix before proceeding.

Run these checks and report results to the user:

# 1. Verify GBrain is configured
gbrain doctor --json

If this fails: "GBrain isn't set up yet. Let's run gbrain init --supabase first."

# 2. Verify Node.js 18+
node --version

If missing or < 18: "Node.js 18+ is required. Install it: https://nodejs.org/en/download"

# 3. Check if ngrok is installed
which ngrok

If missing:

Tell the user: "All prerequisites checked. [N/3 passed]. [List any that failed and how to fix.]"

Step 2: Collect and Validate Credentials

Ask for each credential ONE AT A TIME. Validate IMMEDIATELY. Do not proceed to the next credential until the current one validates.

Credential 1: Twilio Account SID + Auth Token

Tell the user: "I need your Twilio Account SID and Auth Token. Here's exactly where to find them:

  1. Go to https://www.twilio.com/console (sign up free if you don't have an account)
  2. After logging in, you'll see your Account SID right on the main dashboard (it starts with 'AC' followed by 32 characters)
  3. Below it you'll see Auth Token — click 'Show' to reveal it
  4. Copy both values and paste them to me"

After the user provides them, validate immediately:

curl -s -u "$TWILIO_ACCOUNT_SID:$TWILIO_AUTH_TOKEN" \
  "https://api.twilio.com/2010-04-01/Accounts/$TWILIO_ACCOUNT_SID.json" \
  | grep -q '"status"' \
  && echo "PASS: Twilio credentials valid" \
  || echo "FAIL: Twilio credentials invalid — double-check the SID starts with AC and the auth token is correct"

If validation fails: "That didn't work. Common issues: (1) the SID should start with 'AC', (2) make sure you clicked 'Show' to reveal the auth token and copied the full value, (3) if you just created the account, wait 30 seconds and try again."

STOP HERE until Twilio validates.

Credential 2: OpenAI API Key

Tell the user: "I need your OpenAI API key. Here's exactly where to get one:

  1. Go to https://platform.openai.com/api-keys
  2. Click '+ Create new secret key' (top right)
  3. Name it something like 'gbrain-voice'
  4. Click 'Create secret key'
  5. Copy the key immediately — you won't be able to see it again after closing the dialog
  6. Paste it to me

Note: your OpenAI account needs Realtime API access. Most accounts have it by default."

After the user provides it, validate immediately:

curl -sf -H "Authorization: Bearer $OPENAI_API_KEY" \
  https://api.openai.com/v1/models > /dev/null \
  && echo "PASS: OpenAI key valid" \
  || echo "FAIL: OpenAI key invalid — make sure you copied the full key (starts with sk-)"

If validation fails: "That didn't work. Common issues: (1) the key starts with 'sk-', (2) make sure you copied the entire key (it's long), (3) if you just created it, it's active immediately — no delay needed."

STOP HERE until OpenAI validates.

Credential 3: ngrok Account (Hobby tier recommended)

Tell the user: "I need your ngrok auth token. I strongly recommend the Hobby tier ($8/mo) because it gives you a fixed domain that never changes. With the free tier, your URL changes every time ngrok restarts, breaking Twilio and Claude Desktop.

  1. Go to https://dashboard.ngrok.com/signup (sign up)
  2. Recommended: Go to https://dashboard.ngrok.com/billing and upgrade to Hobby ($8/mo). This gives you a fixed domain.
  3. If you upgraded: go to https://dashboard.ngrok.com/domains and click '+ New Domain'. Choose a name (e.g., your-brain-voice.ngrok.app).
  4. Go to https://dashboard.ngrok.com/get-started/your-authtoken
  5. Copy your Authtoken and paste it to me
  6. Also tell me your fixed domain name (if you created one)"
ngrok config add-authtoken $NGROK_TOKEN \
  && echo "PASS: ngrok configured" \
  || echo "FAIL: ngrok auth token rejected"

If user has a fixed domain, use --url flag (Step 3 below). If user stayed on free tier, URLs will change on restart (the watchdog handles this).

Credential 4: Messaging Platform (for call summaries)

Ask the user: "Where should I send call summaries? Options: Telegram, Slack, or Discord."

Based on their choice:

  • Telegram: "Create a bot via @BotFather on Telegram, copy the bot token, and tell me which chat/group to send summaries to." Validate: curl -sf "https://api.telegram.org/bot$TOKEN/getMe" | grep -q '"ok":true'
  • Slack: "Create an Incoming Webhook at https://api.slack.com/apps → your app → Incoming Webhooks → Add New. Copy the webhook URL." Validate: curl -sf -X POST -d '{"text":"GBrain voice test"}' $WEBHOOK_URL
  • Discord: "Go to your server → channel settings → Integrations → Webhooks → New Webhook. Copy the webhook URL." Validate: curl -sf -X POST -H "Content-Type: application/json" -d '{"content":"GBrain voice test"}' $WEBHOOK_URL

Tell the user: "All credentials validated. Moving to server setup."

Step 3: Start ngrok Tunnel

# With fixed domain (Hobby tier — recommended):
ngrok http 8765 --url your-brain-voice.ngrok.app

# Without fixed domain (free tier — URL changes on restart):
ngrok http 8765

If using a fixed domain, the URL is always https://your-brain-voice.ngrok.app. If using free tier, copy the URL from the ngrok output (changes every restart).

Note: ngrok runs in the foreground. Run it in a background process or new terminal tab.

The same ngrok account can also serve your GBrain MCP server (see ngrok-tunnel recipe for the full multi-service pattern).

Step 4: Create Voice Server

Create the voice server directory and install dependencies:

mkdir -p voice-agent && cd voice-agent
npm init -y
npm install ws express

The voice server needs these components in server.mjs:

  1. HTTP server on port 8765 with:

    • POST /voice — returns TwiML that opens a WebSocket media stream to /ws
    • GET /health — returns { ok: true }
    • Twilio signature validation (X-Twilio-Signature header) on /voice
  2. WebSocket handler at /ws that:

    • Accepts Twilio media stream (g711_ulaw audio)
    • Opens a second WebSocket to wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview
    • Bridges audio bidirectionally (no transcoding — both sides use g711_ulaw)
    • Handles response.function_call_arguments.done events from OpenAI (tool execution)
    • Sends tool results back via conversation.item.create with type function_call_output
  3. System prompt builder that takes caller phone number and returns:

    • Appropriate greeting based on caller routing rules
    • Available tools (read-only for unauthenticated, full for authenticated)
    • Instructions: "You are a voice assistant. Search the brain before answering questions. Take messages from unknown callers. Never hang up first."
  4. Tool executor that:

    • Spawns GBrain MCP client (gbrain serve as stdio child process)
    • Routes function calls: search_braingbrain query, lookup_persongbrain search + gbrain get
    • Gates write tools behind authentication
  5. Post-call handler that:

    • Saves transcript to brain/meetings/YYYY-MM-DD-call-{caller}.md
    • Posts summary to the user's messaging platform
    • Runs gbrain sync --no-pull --no-embed to index the new page
  6. WebRTC endpoint (optional, for browser-based calling):

    • POST /session — accepts SDP offer, forwards to OpenAI Realtime /v1/realtime/calls as multipart form-data, returns SDP answer
    • GET /call — serves a web client HTML page with:
      • WebRTC connection to OpenAI Realtime API
      • RNNoise WASM noise suppression (AudioWorklet)
      • Push-to-talk AND auto-VAD mode switching
      • Pipeline: Microphone → RNNoise denoise → MediaStream → WebRTC → OpenAI
    • POST /tool — receives tool calls from the WebRTC data channel, executes them, returns results
    • This lets users call the voice agent from a browser tab instead of a phone

    WebRTC session creation pseudocode:

    POST /session:
      sdp = request.body  // caller's SDP offer
    
      sessionConfig = JSON.stringify({
        type: 'realtime',
        model: 'gpt-4o-realtime-preview',
        audio: { output: { voice: VOICE } },
        instructions: buildPrompt(null),
        tools: TOOL_SETS.unauthenticated,
      })
    
      // Use native FormData (Node 18+) — NOT manual multipart
      fd = new FormData()
      fd.set('sdp', sdp)
      fd.set('session', sessionConfig)
    
      response = POST 'https://api.openai.com/v1/realtime/calls'
        Authorization: Bearer OPENAI_API_KEY
        body: fd   // fetch() sets Content-Type automatically
    
      return response.text()  // SDP answer
    

    Important WebRTC gotchas:

    • voice goes under audio.output.voice, not top-level
    • Do NOT send turn_detection in session config (not accepted by /v1/realtime/calls)
    • Do NOT send session.update on connect (server already configured it)
    • All session.update calls must include type: 'realtime' to avoid session.type errors
    • input_audio_transcription is NOT supported over WebRTC data channel — use Whisper post-call on recorded audio instead
    • Trigger greeting via data channel after WebRTC connects

Reference implementation: The architecture above and the OpenAI Realtime API docs (https://platform.openai.com/docs/guides/realtime) provide the building blocks.

Step 5: Configure Twilio Phone Number

Tell the user: "Now I need to set up your Twilio phone number. Here's what to do:

  1. Go to https://www.twilio.com/console/phone-numbers/search
  2. Search for a number (pick your area code or any available number)
  3. Click 'Buy' next to the number you want (costs $1-2/month)
  4. After purchase, go to https://www.twilio.com/console/phone-numbers/incoming
  5. Click on your new number
  6. Scroll to 'Voice Configuration'
  7. Under 'A call comes in', select 'Webhook'
  8. Enter: https://YOUR-NGROK-URL.ngrok-free.app/voice
  9. Method: HTTP POST
  10. Click 'Save configuration'
  11. Tell me the phone number you purchased"

Or if the user prefers CLI:

# Buy a number (US local)
twilio phone-numbers:buy:local --area-code 415

# Configure webhook
twilio phone-numbers:update PHONE_SID \
  --voice-url https://YOUR-NGROK-URL.ngrok-free.app/voice \
  --voice-method POST

Step 6: Start Voice Server and Verify

cd voice-agent && node server.mjs

STOP and verify:

curl -sf http://localhost:8765/health && echo "Voice server: running" || echo "Voice server: NOT running"

If not running: check the server logs for errors. Common issues:

  • Port 8765 already in use: lsof -i :8765 to find what's using it
  • Missing environment variables: make sure OPENAI_API_KEY is set
  • Module not found: run npm install again

Step 7: Smoke Test (Outbound Call)

This is the magical moment. The agent calls the USER to prove the system works.

Tell the user: "Your phone is about to ring. Pick up and talk for about 30 seconds. Say something like 'Hey, I'm testing my new voice-to-brain system. Remind me to check the quarterly numbers tomorrow.' When you're done, hang up."

curl -X POST "https://api.twilio.com/2010-04-01/Accounts/$TWILIO_ACCOUNT_SID/Calls.json" \
  --data-urlencode "To=USER_PHONE_NUMBER" \
  --data-urlencode "From=TWILIO_PHONE_NUMBER" \
  --data-urlencode "Url=https://YOUR-NGROK-URL.ngrok-free.app/voice" \
  -u "$TWILIO_ACCOUNT_SID:$TWILIO_AUTH_TOKEN"

After the call ends, verify ALL of these:

  1. Messaging notification arrived with call summary
  2. Brain page exists:
    gbrain search "call" --limit 1
    
  3. The brain page has: transcript, entity mentions, action items

If the smoke test fails:

  • No ring: check Twilio console for error logs at https://www.twilio.com/console/debugger
  • Ring but no voice: check ngrok tunnel is up, check OpenAI key is valid
  • Voice works but no brain page: check post-call handler logs, run gbrain sync manually
  • Brain page but no messaging: check messaging bot token is valid

STOP HERE until the smoke test passes. Do not declare success until the user confirms they received the messaging notification AND the brain page exists.

Step 8: Set Up Inbound Calling

Tell the user: "The smoke test passed — voice-to-brain is live! Your number is [TWILIO_NUMBER]. Now let's set up inbound calling."

  1. Twilio webhook is already configured from Step 5
  2. Ask: "Do you want calls to your existing phone to forward to this number after a few rings? That way you answer if you can, and the voice agent picks up if you don't."
  3. Configure caller routing rules in the system prompt
  4. Add the user's phone number as the "owner" number for full access

Step 9: Watchdog (Auto-restart)

# Cron watchdog (every 2 minutes) — add to crontab
*/2 * * * * curl -sf http://localhost:8765/health > /dev/null || (cd /path/to/voice-agent && node server.mjs >> /tmp/voice-agent.log 2>&1 &)

If using ngrok, also set up URL monitoring (free ngrok URLs change on restart):

# Check if ngrok URL changed, update Twilio if so
NGROK_URL=$(curl -s http://localhost:4040/api/tunnels 2>/dev/null | grep -o '"public_url":"https://[^"]*' | grep -o 'https://.*')
if [ -n "$NGROK_URL" ]; then
  twilio phone-numbers:update PHONE_SID --voice-url "$NGROK_URL/voice"
fi

Step 10: Log Setup Completion

mkdir -p ~/.gbrain/integrations/twilio-voice-brain
echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","event":"setup_complete","source_version":"0.8.2","status":"ok","details":{"phone":"TWILIO_NUMBER","deployment":"local+ngrok"}}' >> ~/.gbrain/integrations/twilio-voice-brain/heartbeat.jsonl

Tell the user: "Voice-to-brain is fully set up. Your number is [NUMBER]. Here's what happens now: anyone who calls gets screened by the voice agent. Known contacts get a warm greeting. Unknown callers leave a message. Every call creates a brain page with the full transcript, and you get a summary on [their messaging platform]. The watchdog restarts the server if it crashes."

Cost Estimate

Component Monthly Cost Source
Twilio phone number $1-2/mo Twilio pricing
Twilio voice minutes (100 min) $1-2/mo $0.0085-0.015/min depending on direction
OpenAI Realtime input (100 min) $6/mo $0.06/min
OpenAI Realtime output (50 min) $12/mo $0.24/min
ngrok (free tier) $0 Static domain: $8/mo
Total estimate $20-22/mo For ~100 min of calls

Troubleshooting

Calls don't connect:

Voice agent doesn't respond:

  • Check OpenAI key: the validation command from Step 2 should still pass
  • Check server logs for WebSocket errors (look for "connection refused" or "401")
  • Verify Realtime API access: not all OpenAI accounts have it. Check https://platform.openai.com/docs/guides/realtime

Brain pages not created after call:

  • Run gbrain doctor — if it fails, the database connection is broken
  • Check if the post-call handler ran (look in server logs for "transcript saved")
  • Run gbrain sync manually to force indexing
  • Check file permissions on the brain repo directory

ngrok URL keeps changing:

  • Free ngrok URLs change every time ngrok restarts
  • The watchdog (Step 9) handles this automatically
  • For a permanent URL: upgrade to ngrok paid ($8/mo) for a static domain, or deploy to Fly.io/Railway instead

Note on Option B credentials: If using the DIY pipeline (Option B), you will also need API keys for your chosen STT provider (e.g., Deepgram) and TTS provider (e.g., Cartesia, OpenAI TTS). Collect and validate these during Step 2 alongside the Twilio and OpenAI credentials listed above.

Critical Production Fixes (v0.8.1)

These are NOT optional. They prevent real production failures discovered in a deployment handling daily calls.

Unicode Crash Fix (CRITICAL)

Problem: Em dashes (--), arrows (->), and other non-ASCII characters in the prompt context cause broken surrogate pairs that crash the Twilio WebSocket connection. Phone calls drop silently.

Fix: Replace ALL non-ASCII characters with ASCII equivalents throughout the entire prompt file before sending to Twilio. This is invisible in development (browsers handle unicode fine) and catastrophic in production.

function sanitizeForTwilio(text) {
  return text
    .replace(/[\u2014\u2013]/g, '--')   // em/en dash
    .replace(/[\u2018\u2019]/g, "'")     // smart quotes
    .replace(/[\u201C\u201D]/g, '"')     // smart double quotes
    .replace(/\u2192/g, '->')              // right arrow
    .replace(/\u2190/g, '<-')              // left arrow
    .replace(/[\u2026]/g, '...')         // ellipsis
    .replace(/[^\x00-\x7F]/g, '')        // strip remaining non-ASCII
}

PII Scrub from Voice Context (CRITICAL)

Problem: Brain context loaded into the voice prompt may contain phone numbers, email addresses, and other PII. The voice agent reads these aloud to callers.

Fix: Regex-strip PII from all voice context before injecting into the prompt:

  • Phone numbers: /\+?\d[\d\s\-().]{7,}\d/g
  • Email addresses: /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}/g
  • URLs with auth tokens or API keys
  • Any string matching common credential patterns

Identity-First Prompt (IMPORTANT)

Problem: Voice agents lose their identity mid-conversation. Saying "You are NOT Claude" doesn't stick. The model reverts to its base persona.

Fix: Put identity FIRST in the system prompt, before any context or rules:

# You ARE [Agent Name]
You are [Name], a voice assistant who works with [Brain Name].
You are NOT Claude. You are NOT a general AI assistant.
[Name] has their own personality: [traits].

# Context
[... brain context, calendar, tasks ...]

# Rules
[... behavioral rules ...]

Positioning identity before context ensures the model sees it first and maintains it throughout the conversation.

Problem: If post-call processing fails, the call audio is lost forever.

Fix: Auto-upload ALL call audio immediately on call end:

  • Twilio calls: download the MP3 recording URL from Twilio
  • WebRTC calls: capture via MediaRecorder (webm/opus format)
  • Upload via gbrain files upload-raw <audio-file> --page meetings/call-slug --type call-recording
  • GBrain auto-routes: small files stay in git, large files go to cloud storage with .redirect.yaml pointer. Files >= 100 MB use TUS resumable upload.
  • Generate signed URLs for playback: gbrain files signed-url <storage-path>
  • This ensures every call has a recoverable audio source regardless of whether the transcript or brain page was created successfully

Smart VAD as Default

Problem: Push-to-talk is unnatural on phone calls. Server-side VAD has variable quality.

Fix: Default to Smart VAD (Silero VAD) for voice activity detection:

  • Better endpointing than server-side VAD
  • Fewer false triggers in noisy environments
  • PTT available as fallback (UI toggle for WebRTC clients)
  • Presets: quiet (0.7 threshold), normal (0.85), noisy (0.95), very_noisy (0.98)

These patterns come from a production voice deployment handling real calls daily. They are NOT required for basic setup. Implement them AFTER the smoke test passes. Each pattern is self-contained and optional.

Agent Identity & Engagement

Identity Separation

Problem: A voice agent pretending to be the full AI system creates uncanny valley. Pattern: The voice agent picks its own name and personality, distinct from the main AI brain. "I work with [Brain], [Owner]'s AI." Lighter, more playful, more curious.

Pre-Computed Bid System

Problem: Dead air kills engagement. Voice agents wait passively. Pattern: At call start, scan live context and pre-compute up to 10 engagement bids. Two types: informative (tasks, calendar, social monitoring) and relational (curiosity templates). Bids go INTO the prompt so the agent picks from a list. Use bids #1 and #2 for greeting, cycle the rest during conversation. Never ask "anything else?" — bring up the next bid.

Context-First Prompt

Problem: Voice agent greets generically because it doesn't know what's happening today. Pattern: Load live context at call start: tasks, calendar, location, social monitoring, morning briefing. Position context FIRST in the prompt (before rules) so the model sees it immediately and uses it in the greeting. Try/catch per section. Cap 500-1000 chars each.

Proactive Advisor Mode

Problem: Voice agents are reactive task machines. Pattern: The agent drives the conversation. Anticipate decisions on stale tasks. Suggest capitalizing on trending items. Connect upcoming events with brain context. "Dead air is your enemy" — fill every pause. Never wait passively.

Conversation Timing (the #1 fix)

Problem: Voice agents interrupt mid-thought AND go silent when the caller is done. Both feel terrible. Early "fill every pause" instructions cause the agent to talk over the caller while they're thinking. Pattern: Replace blanket "never be silent" with nuanced timing rules:

  • Caller talking or thinking: SHUT UP. Even 3-5 second pauses mid-thought, wait. Incomplete sentence or mid-story = still thinking. Do not interrupt.
  • Caller done (complete thought + 2-3 seconds silence): NOW respond. Use a bid, ask a follow-up, or pivot to the next topic.
  • Detection heuristic: Incomplete sentence = still thinking. Complete statement + silence = done. Question directed at you = respond immediately.
  • Hard rule: Never let silence go past 5 seconds after a COMPLETE thought.

Add this as a labeled section in the system prompt (e.g., # CRITICAL: Conversation Timing) positioned prominently so the model sees it early. This came from real usage feedback and is the single highest-impact voice quality improvement.

No Repetition Rule

Problem: Voice agent cycles back to the same bid multiple times in a call. Pattern: Add to the system prompt: "Do NOT repeat yourself. If you already said something, move to the NEXT bid. Vary your responses." Simple but addresses a real annoyance that compounds over longer calls.

Prompt Engineering

Radical Prompt Compression

Problem: Long system prompts increase latency and cost on every turn. Pattern: Compress aggressively. Production went 13K to 4.7K tokens (65% cut). Bullets over prose, cut repetition, behavior-first. Every token costs latency + money.

OpenAI Realtime Prompting Guide Structure

Problem: Prose paragraphs parse slowly for the model. Pattern: Use labeled markdown sections: # Role & Objective, # Personality & Tone, # Rules, # Conversation Flow with state machine substates (## State 1: VERIFY, ## State 2: GREETING, ## State 3: CONVERSATION), # Trust.

Auth-Before-Speech

Problem: Auth flow adds dead air at call start. Pattern: Call the auth tool BEFORE speaking any greeting. Then speak "Hey, code's on its way." Shaves seconds off the round-trip.

Brain Escalation

Problem: Voice agent can't answer complex questions that need the full brain. Pattern: If caller says "talk to [Brain]" or asks a deep question, immediately route to main AI via gateway tool with verbal bridge: "one sec, checking with [Brain]."

Call Reliability

Stuck Watchdog

Problem: Calls go silent when VAD stalls or tool execution hangs. Pattern: 20-second timer. If no audio out: clear input buffer, inject "you still there?" system message, force response.create.

Never Hang Up

Problem: AI agents try to end calls. Pattern: Hard prompt rule: only the caller decides when the call ends. Never say goodbye, "I'll let you go," or wrap-up language. If silence, ask "you still there?"

Thinking Sound

Problem: Dead air during slow tool execution. Pattern: Pre-generate g711_ulaw audio chunks in a JSON array. Loop at 20ms intervals during slow tools (brain search, web lookup). Stop when tool result returns.

Fallback TwiML

Problem: Voice agent crashes, callers get silence. Pattern: /fallback endpoint returns TwiML forwarding to owner's cell. Configure as Twilio fallback URL.

Authentication & Authorization

Tool Set Architecture

Problem: Unauthenticated callers accessing write operations. Pattern: Four sets: READ_TOOLS (all callers), WRITE_TOOLS (owner), SCOPED_WRITE_TOOLS (trusted users), GATEWAY_TOOLS (authenticated). LLM doesn't see write tools until auth succeeds. Upgrade via session.update with new tools array. All session.update calls must include type: 'realtime'.

Trusted User Auth with Callback

Problem: People other than the owner need authenticated access. Pattern: Phone registry + callback verification. Each user gets a scope: full, household, content, operational. Scope determines which tools they access.

Caller Routing

Problem: Different callers need different experiences. Pattern: buildPrompt(callerPhone) returns different system prompts: owner (OTP), trusted (callback), inner circle (warm greeting + transfer), known (greeting, message), unknown (screen + message).

Voice Quality

Dynamic VAD / Noise Mode

Problem: Background noise causes false triggers or missed speech. Pattern: set_noise_mode tool adjusts VAD threshold mid-call. Presets: quiet (0.7), normal (0.85), noisy (0.95), very_noisy (0.98). Agent calls proactively on noise.

On-Screen Debug UI

Problem: console.log is useless when testing from a phone. Pattern: WebRTC client displays tool calls, results, errors, and key events inline.

Real-Time Awareness

Live Moment Capture

Problem: Important things said during a call are lost if the call drops or the post-call summary tool doesn't fire. Pattern: When the caller shares something important (feedback, ideas, personal stories, decisions), log it in real-time using a log_voice_request tool. Don't wait until the call ends. Tell the caller: "Got that, sending it to [Brain] now." Also stream key moments to [messaging platform] during the call so the main agent has awareness before the call is over.

Belt-and-Suspenders Post-Call

Problem: Post-call processing depends on the voice agent remembering to call the post_call_summary tool. If the call drops or the agent forgets, the call is lost. Pattern: Both the tool-based AND the automatic call-end handler should post structured signals. The call-end handler (fires on WebSocket close or /call-end) should post to [messaging platform] with:

  • Audio file path
  • Transcript file path (or warning if missing)
  • Tools used during the call
  • Explicit instruction: "[Brain]: Read the call, summarize, take action."

This ensures every call gets processed regardless of whether the voice agent remembered to call the summary tool. Belt and suspenders.

Post-Call Processing

Mandatory 3-Step Post-Call

Problem: Main agent doesn't know a call happened. Pattern: Every call ends with three steps:

  1. Messaging notification — summary to [messaging platform]
  2. Transcript to brainbrain/meetings/YYYY-MM-DD-call-{caller}.md
  3. Audio to storage — Twilio MP3 or WebRTC webm/opus, uploaded to cloud storage

WebRTC Audio + Transcript Parity

Problem: WebRTC calls don't go through Twilio, no automatic logging. Pattern: Client captures audio (MediaRecorder, webm/opus) and transcript (per-turn POST to /transcript). On call end, POST to /call-end saves JSON log. Both channels produce identical output formats. Note: input_audio_transcription is NOT supported over WebRTC data channel — use Whisper post-call instead.

Dual API Event Handling

Problem: OpenAI Realtime API changed event names. Pattern: Handle both response.audio.delta (old) and response.output_audio.delta (new). Same for .done events. Future-proofs against API changes.

Brain Query Optimization

Report-Aware Query Routing

Problem: Voice queries about specific topics trigger slow vector searches. Pattern: Check the question against a keyword map BEFORE full brain search:

Keyword Report Loaded
email, inbox, mail inbox sweep report
social, twitter, mentions social engagement report
briefing, morning morning briefing
meeting meeting sync report
slack slack scan report
content, ideas content ideas report

Load up to 2,500 chars of matching report. Break after first match. Fall back to full brain search if no keyword matches.