Compare commits

...
Author SHA1 Message Date
Wintermute acdee0905a feat: code indexing + multi-repo support
Tree-sitter-based code chunker for TS/JS/Python/Ruby/Go.
Splits code at semantic boundaries (functions, classes, types, exports).
Each chunk includes structured header for embedding context.

Multi-repo config: `gbrain repos add/list/remove`, `gbrain sync --all`.
Strategy-aware sync: markdown (default), code, or auto.
New PageType 'code' for code file pages.

Backward compatible: no config changes = existing behavior preserved.
All 37 sync tests pass, typecheck clean.
2026-04-22 15:34:15 +00:00
55ca4984b2 feat: v0.17.0 — gbrain dream + runCycle primitive (one cycle, two CLIs) (#321)
* fix(sync): honor --dry-run in full-sync path + expose embedded count

Precondition for v0.17 brain maintenance cycle (runCycle primitive).

The full-sync path (performFullSync) previously called runImport() even
when opts.dryRun was true, silently writing to the DB and advancing
sync.last_commit. `gbrain sync --dry-run` on a fresh brain (or with
--full) would mutate state without warning.

Fix:
  - performFullSync now early-returns a `dry_run` SyncResult when
    opts.dryRun is set. Walks the repo via collectMarkdownFiles +
    isSyncable to count what WOULD be imported. No writes, no git
    state advance.
  - SyncResult gains an `embedded: number` field (required). Tracks
    pages re-embedded during the sync's auto-embed step. Existing
    return sites set 0; the synced + first_sync paths set real counts
    (best-estimate until commit 2 sharpens runEmbedCore's return type).
  - first_sync path now returns real added + chunksCreated counts
    from runImport instead of hardcoded zeros.
  - printSyncResult shows embedded count in human output.

Tests (test/sync.test.ts, new `performSync dry-run never writes`
block, PGLite + temp git repo, no DATABASE_URL required):
  - first-sync --dry-run: no pages, no sync.last_commit
  - incremental --dry-run after real sync: bookmark unchanged
  - --full --dry-run: no reimport, bookmark unchanged
  - SyncResult.embedded is a number

Codex outside-voice caught this. Would have shipped silent DB writes
on dry-run for anyone using `gbrain sync --dry-run --full`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(embed): add dry-run mode + return EmbedResult with counts

Precondition for v0.17 brain maintenance cycle (runCycle primitive).

runEmbedCore previously returned Promise<void> and had no dry-run mode.
That made it impossible for runCycle to (a) report accurate embedded
counts or (b) honor --dry-run without also skipping the entire embed
phase (which would have required runCycle to know embed's internal
semantics — a layering violation).

Changes:
  - EmbedOpts gains `dryRun?: boolean`. When set, embedPage and
    embedAll enumerate stale chunks (or would-be-created chunks for
    unchunked pages, via local chunkText without engine.upsertChunks)
    but never call embedBatch and never write to the engine.
  - runEmbedCore: Promise<void> -> Promise<EmbedResult>. Result shape:
    { embedded, skipped, would_embed, total_chunks, pages_processed,
      dryRun }.
    embedded = chunks newly embedded (0 in dryRun).
    would_embed = chunks that WOULD be embedded (0 in non-dryRun).
    skipped = chunks with pre-existing embeddings.
  - runEmbed CLI wrapper honors --dry-run flag and returns the result
    through. `gbrain embed --stale --dry-run` is now a safe preview.
  - Callers ignoring the return value (sync auto-embed, autopilot
    inline fallback, jobs.ts handlers, CLI) keep compiling — the new
    return type is additive for `await` callers.

Tests (test/embed.test.ts, new `runEmbedCore --dry-run` block, uses
the existing mock.module embedBatch pattern, no API key required):
  - dry-run --all: zero embedBatch calls, zero upsertChunks calls,
    would_embed matches stale chunk total
  - dry-run --stale correctly splits stale vs already-embedded counts
  - dry-run --slugs on a single page tallies per-chunk counts
  - non-dry-run regression guard: embedded count matches across
    concurrent workers

Codex outside-voice flagged the Promise<void> return as a blocker for
accurate CycleReport.totals.pages_embedded.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(orphans): engine-injected queries, drop db.getConnection() global

Precondition for v0.17 brain maintenance cycle (runCycle primitive).

findOrphans + queryOrphanPages previously reached into the postgres-js
singleton via db.getConnection(), which (a) didn't compose with
runCycle's explicit-engine contract and (b) was wrong for PGLite test
fixtures and for any caller not using the default global connection.
Codex outside-voice flagged this as a blocker.

Changes:
  - BrainEngine interface gains findOrphanPages() — returns pages with
    no inbound links via the same NOT EXISTS anti-join. Implemented on
    both postgres-engine (sql tag) and pglite-engine (db.query).
  - findOrphans signature: findOrphans(engine, { includePseudo }).
    Engine is required. Uses engine.findOrphanPages() and
    engine.getStats().page_count instead of raw SQL + global counts.
  - queryOrphanPages signature: queryOrphanPages(engine). Delegates to
    engine.findOrphanPages().
  - src/commands/orphans.ts drops the `import * as db` — no more
    global-state coupling.
  - Callers updated: src/core/operations.ts find_orphans handler now
    passes ctx.engine through; runOrphans CLI entry uses its engine arg.
  - No signature change needed in cli.ts (it was already passing engine
    via CLI_ONLY dispatch).

Tests (test/orphans.test.ts, new `findOrphans (engine-injected)`
describe block, PGLite in-memory, no DATABASE_URL required):
  - links correctly scope orphans (alice links to bob -> bob not
    an orphan; alice is)
  - includePseudo:true surfaces _atlas-style pages
  - queryOrphanPages delegates to passed engine
  - empty brain returns {orphans: [], total_pages: 0} without crashing

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(cycle): add runCycle primitive in src/core/cycle.ts

The brain maintenance cycle as a single function. Six phases in
semantically-driven order (fix files → sync → extract → embed →
report orphans). Pure composition of existing library calls — no
execSync, no subprocess anti-patterns, no regex-parsed output.

    ┌───────────────────────────────────────────────────┐
    │ runCycle(engine, opts) → CycleReport              │
    │   Phase 1: lint --fix         (fs writes)         │
    │   Phase 2: backlinks --fix    (fs writes)         │
    │   Phase 3: sync               (DB picks up 1+2)   │
    │   Phase 4: extract            (DB picks up links) │
    │   Phase 5: embed --stale      (DB writes)         │
    │   Phase 6: orphans            (DB read, report)   │
    └───────────────────────────────────────────────────┘

Why the commit-4 primitive:

  - CEO + Eng + Codex reviews all converged on "extract one cycle
    function, wire both dream and autopilot through it." Two CLIs,
    one definition of what the brain does overnight.
  - Phase order was wrong in PR #309's original dream.ts (sync
    before lint+backlinks lost the "fix files, then index them"
    semantic).
  - This commit is the bisectable foundation; commit 5 (dream)
    and commit 6 (autopilot+jobs) just call into it.

Coordination — the codex-flagged blocker:

Session-scoped pg_try_advisory_lock does not survive PgBouncer
transaction pooling (the v0.15.4 fix made pooled connections the
default). Replaced with a DB lock table (gbrain_cycle_locks) that
works through every pooler:

  - Acquire: INSERT ... ON CONFLICT DO UPDATE ... WHERE ttl < NOW()
  - Refresh: UPDATE ttl_expires_at between phases via hook
  - Release: DELETE in finally{}
  - TTL: 30 min; crashed holders auto-release

PGLite / engine=null path uses a file lock at ~/.gbrain/cycle.lock
with PID liveness check. kill(pid, 0) with EPERM treated as alive
(so init/launchd-pid holders aren't mis-classified as stale).

Lock-skip: only phases that mutate state (lint, backlinks, sync,
extract, embed) trigger lock acquisition. orphans is read-only.
Single-phase --phase orphans runs never block on a held lock.

Engine-null mode preserved: filesystem phases run, DB phases skip
with {status:'skipped', reason:'no_database'}. Matches current
dream's capability that would have been lost if runCycle required
a connected engine.

Contract details:

  - CycleReport has schema_version:"1" (stable, additive) so agents
    consuming --json can rely on the shape
  - status: 'ok' | 'clean' | 'partial' | 'skipped' | 'failed'.
    'clean' = ran successfully with zero activity; agents trivially
    detect a healthy brain.
  - PhaseResult.error: { class, code, message, hint?, docs_url? }
    (Stripe-API-tier structured failure info) when status='fail'
  - yieldBetweenPhases hook: awaited between EVERY phase and before
    return, runs even after phase failure, exceptions logged but
    non-fatal. Required so the Minions autopilot-cycle handler can
    renew its job lock between phases (prevents the v0.14 stall-death
    regression codex flagged).
  - git pull explicit: opts.pull defaults to false (cron-safe).
    Autopilot daemon callers opt in if user configured it.
  - extract phase doesn't have a dry-run mode in the underlying
    library function, so runCycle honestly skips extract when
    dryRun=true (status:'skipped', reason:'no_dry_run_support').

Schema migration v16: gbrain_cycle_locks table + idx_cycle_locks_ttl.
Also appended to src/schema.sql and src/core/pglite-schema.ts for
fresh installs. schema-embedded.ts regenerated via build:schema.

Tests (test/core/cycle.test.ts, PGLite in-memory + mocked library
functions, no DATABASE_URL required):

  - dryRun × phases matrix: dryRun:true reaches lint/backlinks/sync/
    embed; extract is honestly skipped
  - Phase selection: default runs all 6 in order; --phase lint runs
    only lint; --phase orphans runs only orphans
  - Lock semantics: acquire + release on mutating phases, skip
    entirely for read-only selections
  - cycle_already_running: seeded live-holder lock → status:skipped,
    zero phase runs; TTL-expired holder → auto-claimed
  - Engine null: filesystem phases run, DB phases skip
  - File lock (engine=null) blocks when PID 1 holds lock with fresh
    mtime — exercises the PID liveness branch including EPERM
  - Status derivation: 'ok' vs 'clean' vs 'partial' vs 'skipped'
  - yieldBetweenPhases called N times, hook exceptions non-fatal

Next: commit 5 rewrites dream.ts as a thin CLI alias over runCycle,
commit 6 migrates autopilot daemon + jobs.ts handler to delegate to
runCycle too.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(dream): add gbrain dream CLI as a thin alias over runCycle

`gbrain dream` is the README brand-promise command: "the agent runs
while I sleep, the dream cycle ... I wake up and the brain is smarter."
Cron-friendly, JSON-reportable, phase-selectable. Same maintenance
cycle as `gbrain autopilot`, just scheduled differently — both
converge on runCycle (added in commit 4) so there's one source of
truth for what happens overnight.

Contract:
  gbrain dream                       # full 6-phase cycle
  gbrain dream --dry-run             # preview, no writes
  gbrain dream --json                # CycleReport JSON (agent-readable)
  gbrain dream --phase <name>        # single-phase run
  gbrain dream --pull                # git pull before syncing
  gbrain dream --dir /path/to/brain  # explicit brain location

Cron: 0 2 * * * gbrain dream --json >> /var/log/gbrain-dream.log

Behavior details:
  - Brain-dir resolution: requires explicit --dir OR sync.repo_path
    in engine config. No more walk-up-cwd-for-.git footgun that
    PR #309's original dream.ts had (would lint unrelated git repos).
  - engine=null mode preserved via cli.ts's try/catch around
    connectEngine — filesystem phases (lint, backlinks) still run
    without a DB, DB phases report skipped/no_database in the output.
  - status=clean prints "Brain is healthy. N phase(s) checked in Ns."
    status=skipped prints the reason (cycle_already_running, etc.).
    Partial/failed prints the phase-by-phase detail.
  - Exit code 1 when status=failed (cron spots real problems).
    'partial' is not a failure — warnings shouldn't page you.
  - --help text cross-references `autopilot --install` for users
    who want continuous maintenance as a daemon.

CLI registration (src/cli.ts):
  - 'dream' added to CLI_ONLY
  - handleCliOnly has a pre-engine branch mirroring doctor's pattern:
    try connectEngine() → ok path; catch → runDream(null, args) so
    filesystem phases still run when DB is down
  - Help text updated with one-line dream entry and autopilot cross-ref

Tests (test/dream.test.ts, real PGLite + real library calls, no mocks
to avoid `mock.module` leakage across test files):
  - brainDir resolution: explicit --dir wins, engine config fallback,
    missing + nonexistent errors
  - phase selection: --phase lint|orphans produces single-phase report
  - phase validation: --phase garbage exits 1
  - output: --json parses as CycleReport with schema_version:"1"
  - human output mentions "Brain is healthy" on clean status
  - dry-run: cycle runs but DB stays untouched
  - exit code: clean/ok/partial do not call process.exit

Also (test/core/cycle.test.ts): refactored to use beforeAll/afterAll
with one shared PGLite engine per describe + truncateCycleLocks
between tests. Cuts test time from ~11s to ~4s; avoids the 15-migration
penalty per test that was causing parallel-suite timeout flakes.

Co-Authored-By: Wintermute <wintermute@garrytan.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: v0.17.0 — autopilot + jobs delegate to runCycle (unifies the cycle)

Autopilot daemon (`--inline` path) and Minions `autopilot-cycle`
handler both now delegate to `runCycle` (introduced in commit 4).
Three callers, one cycle definition:

  1. `gbrain dream`                        — one-shot cron cycle
  2. `gbrain autopilot` daemon inline path — scheduled cycles
  3. `autopilot-cycle` Minions handler     — durable queue with retry

All three share:
  - Same 6 phases in same order (lint → backlinks → sync → extract →
    embed → orphans)
  - Same DB lock table coordination (`gbrain_cycle_locks`)
  - Same yieldBetweenPhases discipline (prevents v0.14 stall-death)
  - Same structured CycleReport output

Autopilot inline path gains lint + orphan sweep that the old path
skipped. Minions autopilot-cycle handler also gains lint + orphans.
Users who run `gbrain autopilot --install` see 6-phase reports in
`gbrain jobs get <id>` starting on next interval. No config change
required.

Changes:
  - `src/commands/autopilot.ts`: inline fallback path (~20 lines)
    replaces the ~22-line sync+extract+embed sequence with a single
    runCycle call. Uses pull:true (matches pre-v0.17 autopilot
    behavior). Uses setImmediate yield hook. Status/failure reporting
    derives from CycleReport.status. `--help` cross-references `gbrain
    dream` for one-shot use.
  - `src/commands/jobs.ts:579` (`autopilot-cycle` handler): replaces
    the 4-step try/catch sequence with a runCycle call. Returns
    `{ partial, status, report }` so `gbrain jobs get <id>` shows the
    full structured CycleReport. Preserves partial-failure semantic
    (one phase failing does NOT throw; next cycle still runs).
    yieldBetweenPhases yields the event loop between phases for the
    worker's lock-renewal timer.

Release scaffolding:
  - VERSION: 0.16.0 → 0.17.0
  - CHANGELOG.md: v0.17.0 entry in GStack voice — headline, numbers
    table, "what this means" paragraph, "To take advantage" block
    per CLAUDE.md post-ship rules. Itemized changes below the fold.
    Credit to @Wintermute for the original PR #309 thesis.
  - skills/migrations/v0.17.0.md: documents what changed for
    upgrading users. No mechanical action required — schema migration
    v16 (cycle locks table) + handler delegation both apply
    automatically. Includes opt-out paths for users who don't want
    their daemon modifying files (use `dream --phase orphans` in cron
    and skip autopilot-install, or other explicit configs).
  - CLAUDE.md: new entries for `src/core/cycle.ts` and
    `src/commands/dream.ts` with contract details.

Tests: no new test file needed for this commit — the cycle primitive
is extensively tested in test/core/cycle.test.ts (18 cases), dream
in test/dream.test.ts (11), and autopilot's delegation is mechanical
(calls runCycle with specific opts). The handler contract is covered
implicitly: if runCycle returns a CycleReport, the handler wraps it
in `{ partial, status, report }` — nothing else to assert.

Verified:
  - `bun test test/autopilot-install.test.ts test/autopilot-resolve-cli.test.ts test/core/cycle.test.ts test/dream.test.ts` → 37 pass, 0 fail

Completes the v0.17.0 feature: 6 bisectable commits on one branch
(garrytan/v0.17-dream-cycle), ready to push as one PR.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): add runCycle + dream E2E coverage against real Postgres

Gap from the v0.17 commit series: PR #321 shipped unit-level tests
for runCycle (test/core/cycle.test.ts) and dream (test/dream.test.ts)
but no E2E coverage that exercises the real Postgres paths. Filling
that in before merge.

  test/e2e/cycle.test.ts (6 cases):
    - schema migration v16 created gbrain_cycle_locks + index
    - dry-run full cycle: zero DB writes + lock table empty after
    - live cycle: pages + chunks materialize, sync.last_commit set
    - concurrent cycle blocked by lock → status:'skipped'
    - TTL-expired lock auto-claimed (crashed-holder recovery)
    - --phase orphans skips lock entirely (read-only optimization)

  test/e2e/dream.test.ts (3 cases):
    - dream --dry-run --json emits valid CycleReport + DB stays empty
    - dream (no --dry-run) syncs pages into real DB
    - dream --phase orphans doesn't touch the cycle-lock table

Both files mock embedBatch via mock.module so the embed phase never
calls OpenAI even when the full 6-phase cycle runs (zero API cost,
zero flakiness from network calls).

Verified locally:
  - `docker run pgvector/pgvector:pg16` on port 5434
  - `DATABASE_URL=... bun test test/e2e/cycle.test.ts test/e2e/dream.test.ts` → 9 pass, 0 fail
  - Full E2E suite (`bun run test:e2e`): 16 files, 150 tests, 0 fail
  - Container torn down after: `docker stop + rm gbrain-test-pg`

Per CLAUDE.md E2E test DB lifecycle. These tests skip gracefully when
DATABASE_URL isn't set (via hasDatabase() helper + describe.skip).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Wintermute <wintermute@garrytan.com>
2026-04-22 08:23:24 -07:00
Wintermute 35967645f3 fix: doctor --fix — 7 DRY violations resolved (inline Iron Law → convention reference) 2026-04-22 09:11:32 +00:00
Garry TanandClaude Opus 4.7 dcd13dd638 feat: v0.16.4 — gbrain check-resolvable CLI + skillify-check wiring (#325)
* Merge origin/master into garrytan/check-resolvable-v1

Resolves CHANGELOG.md conflict: preserved v0.16.1/v0.16.2/v0.16.3 upstream
entries and added v0.16.4 (check-resolvable ship) above them.

* refactor: extract findRepoRoot to src/core/repo-root.ts

Moves findRepoRoot() from private in doctor.ts to a zero-dependency shared
module with a parameterized startDir for test hermeticity. Doctor imports
the shared version; no behavior change (default arg matches prior semantics).

The new gbrain check-resolvable CLI needs findRepoRoot too; importing from
doctor.ts would drag in DB/progress dependencies.

* feat: gbrain check-resolvable CLI wrapper

Standalone CLI gate over checkResolvable(). Exits 1 on any issue (warnings
or errors) per the README:259 contract, stricter than doctor's resolver_health
which ignores warnings. Doctor has 15 other checks to lean on; the standalone
command has nowhere to hide.

- Stable JSON envelope: {ok, skillsDir, report, autoFix, deferred, error, message}
- --fix auto-applies DRY fixes via autoFixDryViolations before re-checking
- --dry-run with --fix previews without writing; autoFix.fixed shows diff
- --verbose prints the deferred-checks note (Checks 5 + 6)
- --skills-dir PATH for hermetic test runs
- Permissive on unknown flags, matching lint/orphans/publish convention

Checks 5 (trigger routing eval) and 6 (brain filing) are tracked as separate
GitHub issues and surfaced via the deferred[] field in --json output.

Covered by 17 new test cases (flag parsing, JSON envelope shape, exit-code
regression gates, --fix wiring, --verbose output).

* chore: bump version and changelog (v0.16.4)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore: track check-resolvable issue-URL swap in TODOS

Defers the filing of GitHub tracking issues for Checks 5 (trigger routing
eval) and 6 (brain filing) plus the TBD-check-5/TBD-check-6 URL replacement
in src/commands/check-resolvable.ts. Unblocks merging PR #325.

* test: fix repo-root CI failure — assert parity, not path contents

The 'default arg uses process.cwd()' test asserted the returned path
matched /honolulu/, which is the local workspace name but not the CI
runner's checkout path (/home/runner/work/gbrain/gbrain). The test's
real purpose is behavioral parity: findRepoRoot() === findRepoRoot(cwd).
Assert that directly instead of pattern-matching paths.

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-22 02:07:00 -07:00
96178d726e fix(subagent): v0.16.3 — bind Anthropic SDK correctly + enable tsc in CI (#318)
* fix(subagent): bind Anthropic SDK messages.create() correctly

The makeSubagentHandler was casting `new Anthropic()` directly to
MessagesClient, but MessagesClient.create() maps to sdk.messages.create(),
not sdk.create(). Every subagent job immediately died with:

  client.create is not a function

Fix: wrap the SDK instance so .create() delegates to .messages.create()
with proper `this` binding via .bind(sdk.messages).

Discovered on first production run of gbrain agent against Supabase.

Co-Authored-By: Wintermute <wintermute@openclaw.ai>

* chore(ci): add typescript typecheck to test pipeline + clean up baseline errors

Root cause infra gap that let the v0.16.0 subagent bug ship: CI ran
only `bun test`, which transpiles types without checking them. Type
errors only surfaced at runtime, in production.

Changes:
- Add `typescript` devDep and a `typecheck` npm script (`tsc --noEmit`).
- Chain `bun run typecheck` into `bun run test` so developers get the
  same pipeline locally that CI runs.
- Flip `.github/workflows/test.yml` to invoke `bun run test` (the npm
  script, including typecheck) instead of `bun test` (runner only).
- Clean up 100+ pre-existing type errors across 30+ files so the first
  run of `tsc --noEmit` is green. Root causes were:
  - `databaseUrl` → `database_url` rename drift in test fixtures (9 files)
  - `PageType` union missing `'meeting'` / `'note'` entries that are
    already used in both src and tests (link-extraction.ts comments
    acknowledged the gap)
  - `GBrainConfig.storage` field never declared despite being read in
    files.ts and operations.ts
  - `ErrorCode` union missing `'permission_denied'`
  - `OrchestratorOpts` shape changed; test callers not updated
  - Dead-code comparisons in migration orchestrators against narrowed
    status types
  - postgres.js `Row`-callback type drift on several `.map()` calls
  - Buffer-as-BodyInit assignment in supabase.ts (real but non-fatal
    runtime bug; Uint8Array slice works and is type-correct)
  - Various `as X` single-step casts that now need `as unknown as X`
    per TS's stricter structural-conversion rules
- Bump `beforeAll` hook timeout to 30s on four PGLite-heavy tests that
  were flaky under parallel test execution: wait-for-completion,
  extract-fs, e2e/search-quality, e2e/graph-quality. All pass in
  isolation; timeouts only happened when dozens of PGLite instances
  init'd simultaneously.

The new CI pipeline now fails on any type error across src/ or test/,
giving us the compile-time regression guard the subagent fix depends on.

* fix(subagent): bind Anthropic SDK messages.create() correctly

Shipped bug: v0.16.0 cast `new Anthropic()` to `MessagesClient`, but
`.create()` lives at `sdk.messages.create`, not on the top-level client.
Every subagent job in production died on first LLM call with
`client.create is not a function`. Discovered on the first `gbrain agent
run` against Supabase.

Fix: assign `sdk.messages` directly to the `MessagesClient` slot.
`sdk.messages` IS the object with a callable `.create()`; the original
bug was picking the wrong entry point on the SDK. No helper, no
wrapper, no `.bind()` — JS method-call semantics preserve `this` at
the call site because `subagent.ts:336` invokes `client.create(...)`
with `client === sdk.messages`.

The one-line assignment also typechecks cleanly against the existing
`MessagesClient` interface (SDK's first `create` overload:
`(MessageCreateParamsNonStreaming, Core.RequestOptions?) =>
APIPromise<Message>` is assignable structurally). This gives us
compile-time regression protection: anyone reverting to
`new Anthropic()` would fail tsc because `Anthropic` has no top-level
`.create`. (The companion chore commit puts `tsc --noEmit` in CI so
this guard is enforced.)

Also adds a `makeAnthropic?: () => Anthropic` dep-injection seam so
the factory default construction branch is testable without real API
calls. Regression test drives one handler turn through a fake SDK,
asserting `sdk.messages.create` is actually called. If someone later
reverts to `new Anthropic()`, both guards fire: tsc fails AND the test
fails.

Co-Authored-By: Wintermute <wintermute@garrytan.com>

* chore(tests): add bunfig.toml + 60s hook timeouts to stabilize PGLite-heavy suites

After turning on tsc in CI (previous commit), running the full `bun run test`
suite in one shot triggered flaky `beforeEach/afterEach hook timed out`
failures on 8+ test files. Every failure traced to PGLite WASM init
contention when many test files spin up fresh PGLite instances in parallel;
each one alone passes in isolation.

- `bunfig.toml` sets the global test hook timeout to 60s (default is 5s),
  covering every test file without per-file edits.
- Individual `beforeAll(fn, 60_000)` / `beforeEach(fn, 15_000)` calls on
  the 8 tests that flaked most stay in place as explicit safety nets so
  a future bunfig config change doesn't silently re-introduce the flake.

Result: 1997 pass, 0 fail on `bun run test` (117 tests added since the
prior baseline by picking up typecheck-gated passes). No infrastructure
flake tolerated in CI.

* chore: bump version and changelog (v0.16.3)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Wintermute <wintermute@garrytan.com>
Co-authored-by: Wintermute <wintermute@openclaw.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 01:34:22 -07:00
Garry TanandClaude Opus 4.7 418d955fd3 docs: v0.16.1 — minions worker deployment guide (from #287) (#317)
* docs: v0.16.1 — minions worker deployment guide (from #287)

New docs/guides/minions-deployment.md covering persistent worker deploy
patterns (watchdog cron, inline --follow for cron-only workloads) plus
the sharp edges of running gbrain jobs work against Supabase in
production.

Addresses a real gap: existing minions docs (minions-fix.md,
minions-shell-jobs.md) cover schema repair and shell-job security,
not deploy patterns. With v0.16.0's durable agent runtime, the
persistent worker is now load-bearing for subagent + subagent_aggregator
handlers too, so a supervised deploy story matters.

Pre-landing accuracy pass corrected five factual bugs against current
source:
- max_stalled column default (5, not 1 or 3)
- stalled-jobs smoke-test query (active, not waiting)
- watchdog SIGTERM-to-SIGKILL grace (10s minimum, not 2s)
- cron env pattern (crontab env lines, not source ~/.bashrc)
- --follow exit semantics (blocks until submitted job is terminal,
  not until queue is empty)

Docs-only. No code changed. Zero migration required.

Contributed by a downstream agent fork via #287.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: credit Wintermute correctly in v0.16.1 CHANGELOG

Wintermute is gbrain's own OpenClaw instance running in production, not a
community contributor. The original CHANGELOG framing ("community contributor
@wintermute") understated the funnier truth: the agent built on top of the
project wrote the deploy guide for the project after hitting its sharp edges
in production. Dogfooding with extra steps.

Co-Authored-By: Wintermute (OpenClaw) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: rewrite minions deployment guide for agent line-by-line execution

Fixes 12 findings from reading v0.16.1 guide as-an-agent would:

Real bugs:
- Crontab syntax wrong for user crontabs (6-field format dumped into
  `crontab -e` got "bad minute" or parsed `user` as the command). Now two
  labeled blocks: 5-field for `crontab -e`, 6-field for `/etc/crontab`.
- Watchdog restart loop (old shutdown lines in unrotated log re-matched
  every 5 min forever). New `minion-watchdog.sh` writes 2-line PID file
  (PID + restart epoch) and only considers log lines newer than the
  epoch. Regex rewritten explicit (mawk rejects `{n}` intervals).
- Credentials in world-readable /etc/crontab. Secrets move to
  /etc/gbrain.env (mode 600), referenced via BASH_ENV in crontab.

Structural:
- Preconditions block (5 fail-fast checks).
- "Which option?" decision tree.
- Template variable table (6 vars documented).
- Upgrade section (v0.13.x -> v0.16.2 checklist).
- Option 3: systemd.service + Procfile + fly.toml.partial snippets.
- Uninstall section.
- `--follow` example uses `gbrain embed --stale` (a real command) instead
  of the fictional `gbrain enrich`.
- Dead-end "Proposed CLI flags (not yet implemented)" replaced with a
  "Tune per-job today" callout pointing at flags that exist.
- Known Issues rewritten as imperatives.

Also wires `docs/guides/minions-deployment.md` into `scripts/llms-config.ts`
under the Configuration section so remote agents fetching llms.txt /
llms-full.txt see the guide by name.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.16.2)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: sync v0.16.2 CHANGELOG with the actual --follow example in the guide

The shipped docs/guides/minions-deployment.md uses `gbrain embed --stale`
(a real command) but the v0.16.2 CHANGELOG entry still referenced
`gbrain enrich --brain $GBRAIN_WORKSPACE` (the older draft). Bring the
CHANGELOG in line with what actually shipped.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 00:01:08 -07:00
Garry TanandClaude Opus 4.7 0e9f8814a5 feat: v0.16.0 — durable agent runtime (gbrain agent + subagent handler + plugin loader) (#258)
* refactor(mcp): extract buildToolDefs helper for subagent tool registry reuse

The inline operations.map(...) block in src/mcp/server.ts became the only
source of truth for agent-facing tool definitions. Extract into a reusable
exported helper so the v0.15 subagent tool registry can call it with a
filtered OPERATIONS subset instead of duplicating the shape.

Byte-for-byte equivalence regression pinned in test/mcp-tool-defs.test.ts —
legacy inline mapping kept verbatim inside the test so any future drift
between the new helper and the pre-extraction MCP schema fails loudly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(operations): subagent-aware OperationContext + put_page namespace

Adds three optional fields to OperationContext:
  - jobId?: number       — the currently running Minion job id
  - subagentId?: number  — the owning subagent job id for tool-dispatched calls
  - viaSubagent?: boolean — FAIL-CLOSED flag for agent-path gating

put_page now enforces a namespace rule when invoked on the subagent tool
dispatch path (viaSubagent=true): writes MUST target
`wiki/agents/<subagentId>/...`. Anchored, slash-boundary enforced so a
collision like `wiki/agents/12evil/...` can't impersonate subagent 12.

The check runs BEFORE the dry-run short-circuit so preview calls surface
the same rejection. Fail-closed: a missing subagentId with viaSubagent=true
rejects every slug rather than letting a dispatcher bug open a hole.

Existing callers unaffected — all three fields are optional and the legacy
put_page behavior is unchanged when viaSubagent is undefined/false.

12 regression + namespace tests pin:
  - local CLI writes (viaSubagent unset) accept arbitrary slugs
  - MCP writes (remote=true, viaSubagent unset) accept arbitrary slugs
  - subagent-path: anchored prefix accepted, wrong id rejected, prefix-
    collision defeated, leading-slash rejected, bare-prefix rejected,
    fail-closed on missing/NaN subagentId, permission_denied code emitted

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(schema): v0.15.0 subagent runtime tables + migration orchestrator

Adds three new tables for the durable LLM agent runtime:

  subagent_messages         — Anthropic message-block persistence.
                              Parallel tool_use blocks in one assistant
                              message live in content_blocks JSONB, not
                              across rows (fixes the (job_id, turn_idx, role)
                              misdesign codex caught in v0.13 drafting).

  subagent_tool_executions  — Two-phase tool ledger. INSERT pending before
                              execute, UPDATE complete/failed after. Replay
                              re-runs pending rows only if the tool is
                              idempotent (v1 ships only idempotent tools so
                              this is preventive).

  subagent_rate_leases      — Lease-based concurrency cap for outbound
                              providers (e.g. anthropic:messages). Stale
                              leases auto-prune on next acquire so crashed
                              workers can't strand capacity.

All DDL uses CREATE TABLE/INDEX IF NOT EXISTS — order-independent vs
PR #244's initSchema() reorder, and idempotent across fresh-install +
upgrade paths. Shipped in both src/schema.sql (Postgres) and
src/core/pglite-schema.ts (PGLite); schema-embedded.ts regenerated.

Migration orchestrator v0_15_0.ts (phases: schema → verify → record).
v0_14_0.ts is a no-op stub so the registry's version sequence stays
gapless (v0.14.0 shipped shell-jobs — code change, no DB migration).

10 unit tests for registry wiring, ordering, dry-run phase behavior, and
schema-embedded table presence. test/apply-migrations.test.ts updated for
the two new registry entries.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): emit child_done on every terminal + max_stalled per-job + terminal set fix

Three correctness fixes the v0.15 subagent aggregator spine depends on:

1. child_done emission on ALL terminal transitions, not just success.
   - completeJob already emitted on success — now also tags outcome='complete'.
   - failJob newly emits on terminal 'failed' or 'dead' (outcome='failed'|'dead',
     error=<text>), BEFORE the parent-terminal UPDATE so the EXISTS guard on
     the inbox INSERT doesn't skip it on fail_parent paths (codex catch).
   - cancelJob now emits outcome='cancelled' per descendant with a parent.
   - handleTimeouts now emits outcome='timeout' per timed-out child.
   ChildDoneMessage gains optional { outcome, error } — backwards compatible
   (legacy writers omitted them; consumers treat absent outcome as 'complete').

2. Parent-resolution terminal set now includes 'failed'.
   Pre-v0.15 the `NOT EXISTS (... status NOT IN ('completed','dead','cancelled'))`
   guard treated a failed child as still-pending, stranding aggregator parents
   that chose on_child_fail='continue' or 'ignore' in waiting-children forever.
   Expanded to {completed, failed, dead, cancelled} everywhere parent resolution
   reads child status (completeJob inline, failJob remove_dep + continue,
   cancelJob sweep, handleTimeouts sweep, and the resolveParent method itself).

3. MinionJobInput.max_stalled threads through MinionQueue.add() on INSERT.
   Column exists with default 1 — that is "first stall → dead", which defeats
   crash recovery for long-running handlers. Subagent children will set
   max_stalled: 3 to survive mid-run worker kills. Second-submitter under an
   idempotency-key hit does NOT mutate the existing row (codex-flagged
   footgun — first-submit options are load-bearing state).

13 unit tests pin: emission on each of completeJob/failJob/cancelJob/
handleTimeouts, insertion order on fail_parent, terminal-set expansion with
continue policy, max_stalled default + override + idempotency behavior.

E2E tier 1 (Postgres) passes 141 tests unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): rate-leases + waitForCompletion infra for v0.15 subagent

Two infrastructure modules the subagent handler spine depends on:

rate-leases.ts — lease-based concurrency cap for outbound providers
(anthropic:messages, openai:*, etc.). Counter-based limiters leak capacity
on worker crash; leases are owner-tagged rows with expires_at that
auto-prune on the next acquire. Two-phase: txn-scoped pg_advisory_xact_lock
guards the check-then-insert so concurrent acquires can't both win the
"last slot". renewLeaseWithBackoff retries 3x (250/500/1000ms) for mid-
call DB blips — on persistent failure the LLM-loop caller aborts with a
renewable error so the worker re-claims and the rate invariant is
preserved. Owner FK cascades clean up leases on job deletion.

wait-for-completion.ts — poll-until-terminal helper for CLI callers.
Minions' NOTIFY is worker-side only; `gbrain agent run --follow` polls
getJob() until status is {completed, failed, dead, cancelled}. TimeoutError
carries jobId + elapsedMs and does NOT cancel the job — the user can
inspect via `gbrain jobs get <id>` later. Supports AbortSignal for Ctrl-C
without throwing. Default pollMs is 1000 on Postgres, 250 on PGLite (inline
CLI has no network RTT).

21 unit tests cover: single/multi acquire under cap, rejection past cap,
release frees slot, different keys are independent, stale prune, cascade
on owner delete, renew bumps expires_at, renew on missing is false,
backoff path success + pruned short-circuit. waitForCompletion: fast-path
terminal, transitions mid-wait (completed/failed/cancelled), TimeoutError
shape, abort-signal early exit, non-existent job error.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): subagent ToolDef types + brain-tool registry (v0.15)

Types first so the handler has a stable contract:
  - SubagentHandlerData / AggregatorHandlerData — the two job.data shapes
  - ToolCtx (engine, jobId, remote, signal) + ToolDef (name, description,
    input_schema, idempotent, execute) — Anthropic-envelope, distinct from
    the MCP McpToolDef extraction landed earlier
  - ContentBlock discriminated union for subagent_messages.content_blocks
  - SubagentStopReason + SubagentResult emitted on terminal completion

brain-allowlist.ts derives one ToolDef per allow-listed OPERATION. Reuses
the ParamDef → JSONSchema shape from the MCP extraction in a local helper
(Anthropic's input_schema field diverges from MCP's inputSchema by a
character). The 11-name allow-list is read-safe + put_page — every
destructive / filesystem / identity-mutating op stays off by default.

put_page gets a namespace-wrapped tool schema: `slug` pattern = anchored
`^wiki/agents/<subagentId>/.+`. The server-side check in put_page op
(shipped in prior commit) is still the authoritative gate — the schema
just helps the model write correct slugs first-try. `subagentId` is
plumbed into the ToolCtx so the viaSubagent=true fail-closed path lights
up on every tool-dispatched put_page.

filterAllowedTools narrows a registry by subagent_def's allowed_tools
frontmatter field. Rejects unknown names at load time (no silent drop —
typos in a skills/subagents/*.md would otherwise ship to prod with a
tool silently missing).

18 tests pin: every allowlist name exists in OPERATIONS (catches upstream
rename), Anthropic name regex, put_page namespace pattern per-subagent,
execute() routes through the op handler with viaSubagent=true, out-of-
namespace put_page throws permission_denied, filter passes prefixed +
unprefixed names, rejects unknowns, deduplicates.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): subagent-audit JSONL + transcript renderer

Two small plumbing pieces the v0.15 subagent handler + `gbrain agent logs`
depend on:

subagent-audit.ts — JSONL-rotated audit log mirroring the shell-audit
pattern. Two event flavors: submission (one line per job submit) and
heartbeat (one line per turn boundary — llm_call_started / completed /
tool_called / tool_result / tool_failed). Heartbeats fix the "--follow on
a long Anthropic call shows nothing for 30 seconds" problem codex flagged.
Never logs prompts or tool inputs (PII risk — subagent input_vars may
carry user-supplied free text); DOES log tokens, ms_elapsed, tool_name,
first 200 chars of error text. Rotates weekly via ISO week. `readSubagent
AuditForJob` is the readback path for `gbrain agent logs` — scans the
current + prior week file so job boundaries across weeks still resolve.
`GBRAIN_AUDIT_DIR` overrides the default ~/.gbrain/audit/ for container
deploys.

transcript.ts — renders subagent_messages + subagent_tool_executions to
markdown. Message order is authoritative; tool rows splice under their
owning assistant tool_use by tool_use_id. Handles text, tool_use (with
pending / complete / failed execution rows), tool_result (skipped if
we already rendered the owning tool_use — avoids double-printing), and
unknown block types (fenced JSON dump for diagnostics). Output is
UTF-8-safe truncated at maxOutputBytes.

21 unit tests: ISO week filename rotation (incl. 2027-01-01 → W53-2026
boundary), submission + heartbeat write shapes, 200-char error cap, best-
effort write failure doesn't throw, readback filters by job_id and
sinceIso. Transcript: empty input, ordering, token line, tool_use +
complete/failed/pending execution rendering, truncation, unknown-block
diagnostic dump.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): subagent LLM-loop handler with crash-resumable replay

The main event: runs one Anthropic Messages API conversation with tool
use, persists every turn + tool execution, and resumes cleanly after a
worker kill anywhere in the loop.

Design points that carry the v0.15 guarantees:

  1. Two-phase tool persistence. INSERT status='pending' before dispatch,
     UPDATE to 'complete' or 'failed' after. subagent_messages rows are
     the canonical conversation; subagent_tool_executions rows are the
     canonical "did this tool run + what did it return". Either DB commit
     is atomic, so replay has a single source of truth.

  2. Replay reconciliation. If the last persisted message is an assistant
     with tool_use blocks AND no following synthesized user message, we
     crashed mid-dispatch. On resume, finish those tools first (respecting
     idempotent flag for 'pending' rows), synthesize the user turn, and
     THEN call the LLM again. Non-idempotent pending rows abort the job
     with a clear error — v0.15 ships only idempotent tools so this is
     preventive.

  3. Rate lease around every LLM call. acquireLease before, releaseLease
     after (both success and error paths). acquired=false throws
     RateLeaseUnavailableError — the worker treats it as a renewable
     error and re-claims later, so a temporary capacity cap doesn't fail
     the job terminally.

  4. Anthropic prompt caching. system block gets cache_control=ephemeral;
     the LAST tool def gets it too (Anthropic caches everything up to and
     including the marked block). ~10x cost reduction on multi-turn
     agents per the plan.

  5. Dual-signal abort. AbortSignal.any merges ctx.signal (timeout / lock
     loss / cancel) with ctx.shutdownSignal (worker SIGTERM). Both feed
     the Anthropic call's AbortSignal; mid-turn abort bails before the
     next LLM call with whatever turns are already persisted. Node ≥ 20
     has AbortSignal.any; older runtimes get a manual-merge polyfill.

  6. Injectable Anthropic client. The real SDK implements MessagesClient
     structurally; tests inject a FakeMessagesClient that scripts
     responses.

12 unit tests pin: no-tool happy path, single tool_use complete, tool
throws → failed row + loop continues, unknown tool name rejection,
max_turns cap, crash-then-resume with partial state, replay skips already-
complete tool execs without re-invoking execute, non-idempotent pending
rejects on resume, lease acquire + release roundtrip, RateLeaseUnavailable
under cap-full, missing prompt validation, allowed_tools unknown-name.

NOT in v0.15: refusal detection (stop_reason + content shape), stop_reason
=max_tokens partial recovery, mid-call lease renewal with backoff loop.
All three are documented as P2 items in the plan file.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): subagent_aggregator handler with mixed-outcome rendering

Claims AFTER all subagent children resolve — by then Lane 1B's queue
changes have posted one child_done message per terminal transition into
this job's inbox (complete / failed / dead / cancelled / timeout). The
aggregator reads those, builds a deterministic markdown summary, and
returns it as the handler result.

Not an LLM call in v0.15 — output is reproducible concatenation so
fan-out runs stay comparable. v0.16+ can add an LLM synthesis pass
behind an opt-in flag.

Contract:
  - empty children_ids → `(no children)` marker
  - missing child_done (shouldn't happen under v0.15 invariants but
    possible if a terminal-state path slipped past Lane 1B) → counted as
    failed with "no child_done message observed" error
  - non-complete outcomes: result is null in the output so no payload
    leaks alongside a failure label
  - children appear in the order children_ids was supplied
  - custom aggregate_prompt_template replaces the markdown header

13 unit tests cover: empty input, all-success, mixed outcomes, result
suppression on failure, missing child_done handling, order preservation,
custom template, progress + log emission, stringified JSONB payload
parsing, non-child_done inbox filtering, legacy-writer outcome fallback,
and internal helper edges.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): GBRAIN_PLUGIN_PATH loader + plugin-authors guide (v0.15)

Plumbing that makes Wintermute (and future downstream agents) day-1
usable on v0.15. Host repos drop a `gbrain.plugin.json` + `subagents/`
directory somewhere, set GBRAIN_PLUGIN_PATH (colon-separated like \$PATH),
and their custom subagent defs load at worker startup.

Path policy is strict: absolute paths only. Relative, ~-prefixed, and
URL-style (https://, file://) all rejected with warnings — the user
controls where plugins live. Non-existent paths and files (not dirs) are
warned and skipped so a typo doesn't crash worker startup.

Collision policy: left-wins. If two plugins ship a subagent with the same
name, the first one in GBRAIN_PLUGIN_PATH keeps it and the other gets a
warning naming both sources. Deterministic + debuggable.

Trust policy: plugins ship subagent defs ONLY. Cannot declare new tools,
cannot extend the brain allow-list, cannot override safety flags. The
subagent def's `allowed_tools:` frontmatter MUST subset the derived
registry — validation happens at load time (worker startup), not at
dispatch time, so a typo in a skill gives a loud startup error instead
of silently "tool never fires at 3am."

Manifest `plugin_version: "gbrain-plugin-v1"` locks the contract. Unknown
versions rejected. `subagents` field escape attempts (`../../../etc` etc)
rejected. gray-matter handles the markdown frontmatter parse — subagent
defs don't conform to the page schema, so we don't use parseMarkdown.

docs/guides/plugin-authors.md is the Wintermute-facing walkthrough.
Covers the minimum viable plugin shape, the three policies, the
frontmatter fields, known caveats (audit JSONL is local-only, tool calls
always run remote=true, put_page is namespace-scoped).

22 unit tests pin path rejection, missing/invalid manifest, unsupported
version, escape-attempt, basename fallback for missing frontmatter.name,
allowed_tools round-trip, unknown-tool rejection with validAgentToolNames,
empty env, multi-path, collision warning with left-wins, trimmed paths,
manifest-rejection as warning.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(cli): gbrain agent run + logs + worker registration (v0.15 Lane 4H)

Three integration seams wired:

src/commands/agent.ts — \`gbrain agent run\`. Submits subagent jobs (or a
fan-out of N + aggregator) under the trusted-submit flag so the
PROTECTED_JOB_NAMES guard doesn't reject. Fan-out path creates the
aggregator first (so children can reference its id as parent), submits
each child with on_child_fail='continue' (required by Lane 1B's terminal-
set + child_done machinery), then jsonb_set's the aggregator's
children_ids. Short-circuits a 1-entry manifest to a single subagent
with no aggregator. Follow mode runs agent-logs streaming + waitFor
Completion in parallel and exits on terminal status; detach prints the
job id and exits. Ctrl-C is handled as detach, not cancel — the job
keeps running, consistent with durability invariants.

src/commands/agent-logs.ts — \`gbrain agent logs\`. Merges ~/.gbrain/audit/
subagent-jobs-*.jsonl (heartbeats + submissions) with subagent_messages
(persisted conversation) in one chronological stream. --follow polls at
1s and exits when the job hits terminal. --since accepts ISO-8601 OR
relative shorthand (5m / 1h / 2d). Writes transcript tail (full message
+ tool tree) only for terminal jobs, so mid-run --follow doesn't spam a
half-rendered transcript.

src/commands/jobs.ts registerBuiltinHandlers — matches the shell-handler
opt-in shape. GBRAIN_ALLOW_LLM_JOBS=1 registers the subagent +
subagent_aggregator handlers, then loads plugins from GBRAIN_PLUGIN_PATH
with validAgentToolNames pulled from BRAIN_TOOL_ALLOWLIST. Every plugin
warning + loaded-plugin line prints to stderr, mirroring the openclaw-
seam startup convention.

src/core/minions/protected-names.ts — subagent + subagent_aggregator
join the protected set. MCP submit_job returns permission_denied; only
trusted-CLI callers (with allowProtectedSubmit) can insert these rows.

src/cli.ts — adds 'agent' to CLI_ONLY + dispatches it like 'jobs'.

Test fallout: subagent-handler.test.ts + subagent-transcript.test.ts
helpers now submit under allowProtectedSubmit (they insert rows named
'subagent' directly against the queue). 23 new tests in agent-cli.test.ts
cover: flag parsing (including --detach implies !follow, --tools comma
split, -- terminator, unknown flag throw), --since parse (ISO, relative
5m/2h/1d, unparseable error), protected-name guard for all three names,
trusted-submit gate, and a fan-out integration check that verifies the
aggregator + children shape after --fanout-manifest.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): rename max_children test's spawned jobs off the protected 'subagent' name

The spawn-storm test submitted 50 literal-string 'subagent' children to
exercise the max_children row-lock serialization. In v0.15 'subagent' is
a PROTECTED_JOB_NAME (CLI-only; trusted submit required), so the old
literal submission now throws before reaching the row-lock check.

The test is about max_children semantics, not the v0.15 subagent runtime
specifically — rename the child name to 'child_worker' so the test
exercises the exact same queue.add path without tripping the new guard.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(ship): v0.15.0 — VERSION, CHANGELOG, README, upgrading-agents, CLAUDE.md

Bumps VERSION → 0.15.0 and package.json → 0.15.0 (resolves the pre-existing
drift — on master, VERSION=0.14.0 but package.json=0.13.1; src/version.ts
reads package.json, so this is what the binary prints now).

CHANGELOG lands the release-summary entry in the GStack voice + the full
itemized change list (11 new modules, 3 new tables, queue correctness
fixes, trust-model additions, 159 new unit tests). Voice rules respected
— no em dashes, no AI vocabulary, real file names + real numbers.

README gets a "Durable agents: `gbrain agent` (v0.15)" section next to
the Minions block, with the three canonical CLI shapes (single run,
fanout-manifest, logs --follow) and a pointer to plugin-authors.md.

docs/UPGRADING_DOWNSTREAM_AGENTS.md gets a full v0.15.0 section covering
the four adoption steps downstream agents (Wintermute and similar) need:
(1) worker opt-in via GBRAIN_ALLOW_LLM_JOBS, (2) moving custom subagent
defs to a plugin repo, (3) replacing ephemeral subagent runs with durable
`gbrain agent run`, (4) the put_page namespace rule for agent-driven writes.

CLAUDE.md updated with concise per-file descriptions for every new module:
the handler, aggregator, audit, rate-leases, wait-for-completion,
transcript, plugin-loader, brain-allowlist, tool-defs extraction, agent
CLI + logs CLI, and the registerBuiltinHandlers wiring for subagent
handlers + plugin-loader.

Verified: binary builds (940 modules, 89ms compile), prints `gbrain 0.15.0`,
`gbrain agent --help` shows the new subcommand shape. 170 new tests pass
(full v0.15 surface). Full unit suite passes bar one parallel-load
flake on a pre-existing E2E (graph-quality, passes in isolation).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(minions): drop GBRAIN_ALLOW_LLM_JOBS flag — subagent handlers always-on

The env flag was ceremony. Shell jobs need the flag because they execute
arbitrary CLI commands (RCE surface). Subagent jobs don't — they call the
Anthropic API with whatever ANTHROPIC_API_KEY is in env, so the key is
already the cost gate (no key → SDK fails on the first turn). And
who-can-submit is already protected by PROTECTED_JOB_NAMES +
TrustedSubmitOpts: MCP callers get permission_denied; only `gbrain agent
run` with allowProtectedSubmit can insert subagent / subagent_aggregator
rows. The flag added nothing the existing guards didn't already give us.

registerBuiltinHandlers now always registers subagent + subagent_aggregator
and loads GBRAIN_PLUGIN_PATH plugins. Worker startup prints:

  [minion worker] subagent handlers enabled

instead of the conditional enabled/disabled pair. Plugin discovery runs
unconditionally — empty PATH is a no-op.

README, CHANGELOG, docs/UPGRADING_DOWNSTREAM_AGENTS, CLAUDE.md, agent CLI
help text, and subagent handler docstring all updated to drop the flag
reference. Shell handler's GBRAIN_ALLOW_SHELL_JOBS gate is untouched —
separate concern (RCE, not billing).

Full suite: 1859 pass, 0 fail.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: scrub private agent-fork name from all public artifacts

Enforces the rule added to CLAUDE.md (privacy section): never say
`Wintermute` in any CHANGELOG, README, doc, PR, or commit message.
Reader-facing copy says `your OpenClaw` (the term covers every
downstream OpenClaw deployment — Wintermute, Hermes, AlphaClaw — in
one umbrella the reader already recognizes). First-person /
origin-story copy says `Garry's OpenClaw` (honest that this is the
production deployment driving the feature, without exposing the
private agent's name).

Swept across:
  CHANGELOG.md (v0.15 entry + 4 historical mentions)
  README.md
  TODOS.md
  docs/UPGRADING_DOWNSTREAM_AGENTS.md
  docs/guides/plugin-authors.md (including example plugin names)
  docs/guides/plugin-handlers.md
  docs/guides/minions-fix.md
  docs/designs/KNOWLEDGE_RUNTIME.md (27 refs, mostly analytical)
  docs/benchmarks/2026-04-18-minions-vs-openclaw-production.md
  skills/migrations/v0.11.0.md
  skills/skillpack-check/SKILL.md
  scripts/skillify-check.ts
  src/commands/doctor.ts
  src/commands/migrations/v0_15_0.ts
  src/commands/skillpack-check.ts
  src/core/enrichment/completeness.ts
  src/core/minions/plugin-loader.ts
  src/core/operations.ts
  src/core/output/scaffold.ts

Intentionally kept (these mentions define/test the rule itself):
  CLAUDE.md — the privacy rule section necessarily uses the literal
  name to define the restriction and examples
  test/plugin-loader.test.ts — fixture name in a plugin-loading test;
  renaming risks breaking assertion logic
  test/integrations.test.ts — the word appears in a privacy-regex
  test that explicitly enforces name redaction
  test/doctor-minions-check.test.ts — a comment referencing the rule
  CEO plan artifact at ~/.gstack/projects/… — private, not distributed

Binary builds (941 modules), 198/198 relevant tests pass, `gbrain --version`
prints `0.15.0`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: gitignore bun --compile artifacts with a glob, not specific hashes

Each `bun build --compile` emits a fresh hash-named `.*-*.bun-build` file
in cwd. The prior entries listed two specific hashes that were already
stale, so every build after those created a new untracked file requiring
manual cleanup.

Replace the two stale entries with `*.bun-build` so any current or future
compile artifact is ignored automatically.

Verified: ran `bun build --compile`, got two new `.*-*.bun-build` files,
`git status` stays clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(ship): rename v0.15.0 → v0.16.0

gbrain master is at 0.14.2. Other 0.15.x PRs may land before/after
this one — we bump the minor (new capability) and lock to 0.16.0 so
ordering with concurrent work doesn't matter.

Touches:
- VERSION: 0.15.0 → 0.16.0
- package.json: 0.15.0 → 0.16.0
- Rename src/commands/migrations/v0_15_0.ts → v0_16_0.ts (+ all
  version strings inside + import in index.ts registry)
- Rename test/migrations-v0_15_0.test.ts → migrations-v0_16_0.test.ts
- test/apply-migrations.test.ts: skippedFuture lists now reference
  '0.16.0'
- test/put-page-namespace.test.ts + test/mcp-tool-defs.test.ts: Lane
  comment refs updated
- src/schema.sql + src/core/pglite-schema.ts: "v0.15.0" section
  comment updated; src/core/schema-embedded.ts regenerated
- CHANGELOG.md: top entry renamed to [0.16.0]; inline v0_15_0 /
  v0.15.0 refs swept
- docs/UPGRADING_DOWNSTREAM_AGENTS.md: section heading v0.15.0 → v0.16.0

Verified: `gbrain --version` prints 0.16.0, migration registry /
buildPlan / put_page / mcp-tool-defs / handlers tests all green
(49/49).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: reframe v0.16 durability headline around OpenClaw crashes

"Laptop closed mid-run" framing implied a consumer workflow. Real pain is
OpenClaw subagents dying daily on worker kill, memory blip, or timeout.
Headline + README copy match the body now.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: regenerate llms-full.txt after README copy change

Regen drift guard caught the README edit from 83beec4.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 21:14:17 -07:00
133 changed files with 12214 additions and 404 deletions
+1 -1
View File
@@ -28,4 +28,4 @@ jobs:
with:
bun-version: latest
- run: bun install
- run: bun test
- run: bun run test
+3 -2
View File
@@ -5,8 +5,9 @@ bin/
.env
.env.*
!.env.*.example
.18a49dfd730ff378-00000000.bun-build
.18a49f9dfb996f70-00000000.bun-build
# Bun --compile temp artifacts. Each build emits a new hash-named .bun-build
# file in cwd; glob catches all of them.
*.bun-build
.gstack/
supabase/.temp/
.claude/skills/
+426 -5
View File
@@ -2,6 +2,427 @@
All notable changes to GBrain will be documented in this file.
## [0.17.0] - 2026-04-22
## **`gbrain dream`. Run the brain maintenance cycle while you sleep.**
## **One primitive, two CLIs. Autopilot gains lint + orphan sweep automatically.**
The README has promised "the dream cycle" for a year. v0.17 makes it real as a first-class command. `gbrain dream` runs one maintenance cycle and exits, designed for cron. Same six phases as `gbrain autopilot` — they both delegate to the new `runCycle` primitive in `src/core/cycle.ts`. One source of truth for what your brain does overnight.
Phase order is semantically driven: **fix files first, then index them**. Lint and backlinks write to disk. Sync picks them up into the DB. Extract links the graph. Embed refreshes vectors. Orphan sweep reports the gaps. If your autopilot daemon was doing sync-before-lint (which PR #309's original dream.ts also got wrong), your fixes landed the next cycle instead of the current one. Fixed.
Autopilot users upgrading get lint + orphan sweep for free. No config change. `gbrain jobs list` shows the full 6-phase report now. If you don't want the daemon modifying files, `gbrain dream --phase orphans` in cron keeps autopilot for embed+sync and gives you manual control over the writes.
### The numbers that matter
Measured against a v0.16 baseline. Lines-of-code delta is net-small: runCycle adds ~500 lines, but the new dream.ts is 80 lines (vs the 446-line original in PR #309), and autopilot's two-path branching collapses to one delegated call.
| Metric | BEFORE v0.17 | AFTER v0.17 | Δ |
|--------|--------------|-------------|---|
| `gbrain dream --dry-run` mutates DB | Yes (full-sync + embed silently wrote) | No (every phase honors dry-run) | correctness |
| Sources of truth for "the cycle" | 3-4 (dream inline, dream shell-outs, autopilot inline, Minions handler) | 1 (`runCycle`) | DRY win |
| Phase order: fix-then-index | No (sync before lint) | Yes (lint → backlinks → sync → extract → embed → orphans) | semantics |
| Coordination across daemon + cron + Minions worker | Lockfile heuristic with 6 known holes | DB lock table + PID-liveness file lock | primitive upgrade |
| Works under PgBouncer transaction pooling | No (session-scoped `pg_try_advisory_lock`) | Yes (TTL row, refreshed between phases) | Supabase-safe |
| `findRepoRoot` walks into wrong git repo | Yes (10 levels of cwd) | No (explicit --dir OR configured sync.repo_path) | footgun fixed |
| Autopilot daemon phase count | 4 (sync+extract+embed+backlinks in Minions mode; no backlinks inline) | 6 (+lint +orphans) | feature parity |
| CycleReport shape stability for agents | N/A | `schema_version: "1"` (stable, additive only) | API contract |
### What this means for your workflow
Cron users: one line. `0 2 * * * gbrain dream --json >> /var/log/gbrain-dream.log`. You get a structured `CycleReport` every morning with per-phase timing, counts, and any errors tagged with `{class, code, message, hint, docs_url}`.
Autopilot users: nothing to do. Your daemon picks up the new phases on next cycle. If you want to see them: `gbrain jobs get <autopilot-cycle-id>` shows the full report.
Reviewers/codex caught three plan-breakers during multi-round review that would have shipped silent DB writes on dry-run: (1) `performSync`'s full-sync path was ignoring `opts.dryRun`, (2) `runEmbedCore` had no dry-run mode and returned void, (3) `findOrphans` used `db.getConnection()` global and didn't compose with a passed engine. All three are fixed as preconditions (commits 1-3 of the 6-commit bisectable series).
Credit: @Wintermute for the original `gbrain dream` thesis (PR #309). The brand-promise framing survived; the implementation got redesigned from scratch around the runCycle primitive after CEO + Eng + Codex + DX review found structural issues.
## To take advantage of v0.17.0
`gbrain upgrade` should do this automatically. If it didn't, or if `gbrain doctor` warns about a partial migration:
1. **Run the migration orchestrator manually:**
```bash
gbrain apply-migrations --yes
```
2. **Your agent reads `skills/migrations/v0.17.0.md` the next time you interact with it.** No mechanical host-repo action required; the schema migration (v16 cycle-lock table) and the behavior shift in autopilot's inline path both apply automatically.
3. **Verify the outcome:**
```bash
gbrain dream --help # new command exists
gbrain dream --dry-run --json # safe preview
gbrain doctor # should show no pending migrations
```
Autopilot users: `gbrain jobs list --status complete | head -5` and inspect an `autopilot-cycle` job with `gbrain jobs get <id>` — the report now includes 6 phases.
4. **If any step fails or the numbers look wrong,** please file an issue: https://github.com/garrytan/gbrain/issues with:
- output of `gbrain doctor`
- contents of `~/.gbrain/upgrade-errors.jsonl` if it exists
- which step broke
This feedback loop is how the gbrain maintainers find fragile upgrade paths. Thank you.
### Itemized changes
**New CLI command: `gbrain dream`**
- One-shot maintenance cycle for cron. Exits when done. Flags: `--dry-run`, `--json`, `--phase <name>`, `--pull`, `--dir <path>`, `--help`.
- `--help` shows cron example + cross-reference to `autopilot --install` for continuous daemon.
- Empty-state output is intentionally satisfying: `Brain is healthy. 6 phase(s) checked in 2.3s.` Agents detect it via `status: "clean"`.
- Exit code 1 on `status: "failed"`. Warnings (`status: "partial"`) are not failures — don't page someone.
- `--dir` OR `sync.repo_path` config required. No more walk-up-cwd-for-.git footgun.
**New primitive: `src/core/cycle.ts`**
- `runCycle(engine: BrainEngine | null, opts: CycleOpts): Promise<CycleReport>`.
- Six phases in order: lint → backlinks → sync → extract → embed → orphans.
- `CycleReport` has `schema_version: "1"` (stable, additive). `status: 'ok' | 'clean' | 'partial' | 'skipped' | 'failed'` with `reason` field on skipped.
- `PhaseResult.error: { class, code, message, hint?, docs_url? }` on fail. Stripe-API-tier structured errors.
- `yieldBetweenPhases` hook awaited between every phase + before return. Required for Minions worker lock renewal. Exceptions non-fatal.
- Engine nullable — filesystem phases run without DB; DB phases skip with `reason: "no_database"`.
- Lock-skip: read-only phase selections (`--phase orphans`) skip lock acquisition.
**New schema: `gbrain_cycle_locks` (migration v16)**
- DB lock table with TTL (30 min), replaces session-scoped `pg_try_advisory_lock` which the v0.15.4 PgBouncer-transaction-pooler fix silently broke.
- Refreshed between phases via the yield hook. Crashed holders auto-release on TTL expiry.
- PGLite + engine=null use a file-based fallback at `~/.gbrain/cycle.lock` with PID-liveness check (EPERM treated as alive so PID 1 holders aren't mis-classified).
**Autopilot + Minions integration**
- Autopilot's inline fallback path (`--inline` flag + PGLite mode) now delegates to `runCycle`. Gains lint + orphan phases it didn't run before. Uses `pull: true` by default (preserves pre-v0.17 pull semantics).
- Minions `autopilot-cycle` handler (in `src/commands/jobs.ts`) also delegates to `runCycle`. Returns `{ partial, status, report }` so `gbrain jobs get <id>` surfaces the full structured report.
- `gbrain autopilot --install` install/uninstall/launchd/systemd/crontab machinery untouched.
- `gbrain autopilot --help` now cross-references `gbrain dream`.
**Precondition fixes (required for the runCycle primitive to compose cleanly)**
- `src/commands/sync.ts`: `performFullSync` honors `opts.dryRun` in first-sync + `--full` paths. Was silently calling `runImport` regardless. `SyncResult.embedded: number` field added; `first_sync` path now returns real counts from `runImport` (was hardcoded to 0).
- `src/commands/embed.ts`: `runEmbedCore` adds `dryRun?: boolean` opt and returns `EmbedResult { embedded, skipped, would_embed, total_chunks, pages_processed, dryRun }` instead of `void`. `gbrain embed --stale --dry-run` is now a safe preview.
- `src/commands/orphans.ts`: `findOrphans(engine, opts)` takes a `BrainEngine` parameter. Added `findOrphanPages()` method to `BrainEngine` interface + implementations on both `postgres-engine` and `pglite-engine`. Drops `db.getConnection()` global — findOrphans now composes with test-injected engines and works on PGLite.
**Tests (all run in CI, no DATABASE_URL or API keys required)**
- `test/sync.test.ts`: 4 new cases. First-sync dry-run, incremental dry-run, `--full` dry-run, SyncResult.embedded shape. PGLite + temp git repo.
- `test/embed.test.ts`: 4 new cases. Dry-run with stale chunks, dry-run stale-vs-fresh split, dry-run --slugs, non-dry-run regression guard. Mocked `embedBatch`.
- `test/orphans.test.ts`: 4 new cases. Engine-injected findOrphans, includePseudo flag, queryOrphanPages delegation, empty-brain edge. PGLite.
- `test/core/cycle.test.ts` (new): 18 cases covering dryRun × phases × lock_held × engine-null. Shared PGLite engine per describe via beforeAll + truncateCycleLocks (cuts test time ~3x vs per-test init).
- `test/dream.test.ts` (rewritten, 11 cases): brainDir resolution, phase selection, phase validation, JSON output shape, dry-run propagation, exit-code semantics. Real PGLite + real library calls (no `mock.module` to avoid leakage).
**Docs**
- `skills/migrations/v0.17.0.md`: new. Informational, no mechanical action required.
- `CHANGELOG.md` + `CLAUDE.md`: updated.
**PR #309 disposition**
- Closed with credit to @Wintermute. Their thesis ("`gbrain dream` as first-class CLI verb") was right; the implementation got redesigned around the runCycle primitive after deep review surfaced structural issues in the fold approach.
- `Co-Authored-By: Wintermute` preserved on commit 5 (the dream.ts rewrite).
---
## [0.16.4] - 2026-04-22
## **`gbrain check-resolvable` ships. The command the README promised for weeks.**
## **Agents and CI finally have a one-shot skill-tree gate that actually exits non-zero when anything is off.**
The `resolver_health` logic has lived inside `gbrain doctor` since v0.11. The README claimed a standalone `gbrain check-resolvable` shipped too ... it didn't. Scripts referenced it. Skillify's 10-item checklist referenced it. The binary just shrugged. Fixed.
`gbrain check-resolvable` runs the same four checks doctor runs (reachability, MECE overlap, MECE gap, DRY violations) but with a stricter contract: **exits 1 on any issue, errors AND warnings**. Doctor's resolver_health block still exits 0 on warnings-only because doctor has 15 other checks to lean on. The standalone command has nowhere to hide. CI can finally gate on a single command instead of parsing `gbrain doctor --json`.
The JSON output is a stable envelope, one shape for success and error: `{ok, skillsDir, report, autoFix, deferred, error, message}`. No more "did it succeed? let me see which keys are present." The `deferred` array names the two checks still pending (trigger routing eval, brain filing) with links to their tracking issues, so agents reading the JSON know the current coverage boundary.
`scripts/skillify-check.ts` is now machine-gated. Item #8 on the skillify 10-item checklist used to print "run: gbrain check-resolvable" and pass unconditionally. Now it subprocess-calls the real command and asserts on the exit code. Binary-missing fails loud instead of silently passing ... the kind of silent false-pass that used to put broken skills on the shelf.
## To take advantage of v0.16.4
No migration needed. `gbrain upgrade` brings the binary; nothing to apply. Try it:
```bash
gbrain check-resolvable # human output, like doctor's resolver section
gbrain check-resolvable --json | jq .ok # machine-readable gate for CI
gbrain check-resolvable --fix --dry-run # preview DRY auto-fixes without writing
```
Wire it into your CI:
```bash
gbrain check-resolvable || exit 1 # fails the build on any warning/error
```
### Itemized changes
**New command**
- `gbrain check-resolvable [--json] [--fix] [--dry-run] [--verbose] [--skills-dir PATH] [--help]` — standalone skill-tree gate. Covers reachability, MECE overlap, MECE gap, DRY violations. Exits 1 on any issue.
- Stable JSON envelope (`ok`, `skillsDir`, `report`, `autoFix`, `deferred`, `error`, `message`) — one shape for both success and error paths.
- `--fix` auto-applies DRY fixes via `autoFixDryViolations` before re-checking (same ordering as `doctor --fix`).
- `--dry-run` with `--fix` previews without writing; the JSON `autoFix.fixed` array shows what would change.
- `--verbose` prints the Deferred checks note with issue URLs so nobody forgets Checks 5 and 6 are still tracked.
**Deferred to separate issues**
- Check 5: trigger routing eval — verify every skill's own frontmatter trigger routes to itself in RESOLVER.md. Surfaced via the CLI's `deferred[]` output block.
- Check 6: brain filing validation — verify mutating skills register the brain directories they write to. Same surface.
**Shared refactor**
- `src/core/repo-root.ts` — extracted `findRepoRoot()` from `doctor.ts` to a zero-dependency shared module with a parameterized `startDir` for test hermeticity. Doctor imports the shared version; no behavior change (default arg matches prior semantics).
- `src/commands/doctor.ts` — updated to import the shared `findRepoRoot`.
**Skillify integration**
- `scripts/skillify-check.ts` — item #8 ("check-resolvable gate") now subprocess-calls `gbrain check-resolvable --json` and gates on the exit code. Result is cached per process so iterating many skills only runs the subprocess once. Binary-missing fails loud via explicit `spawn` error handling ... no silent false-pass.
**Tests (22 new cases)**
- `test/repo-root.test.ts` — 4 cases for the extracted `findRepoRoot()` (first-iter hit, walks up, returns null, default arg behavioral parity).
- `test/check-resolvable-cli.test.ts` — 17 cases split between direct unit tests (flag parsing, resolveSkillsDir, DEFERRED constants) and subprocess integration tests (help, JSON envelope shape, exit-code regression gates for warnings AND errors, `--fix --dry-run` wiring, `--verbose` output).
- `test/skillify-check.test.ts` — 2 new cases for the check-resolvable wiring: loud failure when binary is missing (no silent pass), happy path when a synthetic gbrain returns `ok: true`.
**Contract note for CI users**
- `gbrain check-resolvable` exits 1 on warnings AND errors. `gbrain doctor`'s resolver_health block still exits 0 on warnings-only. If you scripted against doctor's looser gate, `check-resolvable` will bite harder ... on purpose. This honors the README:259 contract: "Exits non-zero if anything is off."
---
## [0.16.3] - 2026-04-22
## **`gbrain agent run` actually runs now. The subagent SDK wiring that shipped broken in v0.16.0 is fixed.**
## **Every `.ts` file in the repo typechecks on every `bun run test`. Silent regressions end here.**
v0.16.0 shipped with the headline feature, `gbrain agent run`, unable to make a single LLM call. `makeSubagentHandler` cast `new Anthropic()` straight to `MessagesClient`, but the SDK exposes `.create()` at `sdk.messages.create`, not on the top-level client. Every subagent job in production died on the first call with `client.create is not a function`. The type system would have caught it. Nothing was running the type system.
The root cause isn't the casting bug. It's that `bun test` transpiles TypeScript without type-checking it, and `bun test` was the entire CI pipeline. Invalid types ran until they hit runtime. This release fixes the symptom (one-line change, `deps.client ?? new Anthropic().messages`, which typechecks cleanly against `MessagesClient` because `sdk.messages` IS the right object) and closes the hole that let it ship (`tsc --noEmit` now runs on every `bun run test`, and the CI workflow runs `bun run test` not `bun test`). Two independent guards: anyone reverting to `new Anthropic()` fails the type check; a new regression test drives one handler turn through an injected fake SDK and fails loudly if the factory default branch breaks.
Closing the CI gap surfaced 100+ pre-existing type errors across 30+ files: `databaseUrl` → `database_url` rename drift, missing `"meeting"` / `"note"` entries in the `PageType` union that both src and tests already used, a Buffer-as-BodyInit assignment in the Supabase uploader, dead-code comparisons against narrowed status types in the migration orchestrators, and several `as X` casts that TS 5.6 requires be spelled `as unknown as X`. All cleaned up. The first tsc run is green.
### The numbers that matter
From the merged branch after both the fix and the infra cleanup landed locally against master.
| Metric | Before | After | Δ |
|---|---|---|---|
| `bun run typecheck` errors | 104 | 0 | -104 |
| `gbrain agent run` in prod | 100% failure on first LLM call | Works | ✅ |
| Test file count | ~75 | ~75 (+1 regression test block) | +1 |
| `bun run test` pass rate | 1962 pass / 4 fail (PGLite flake under parallel load) | 1997 pass / 0 fail | +35 pass, -4 fail |
| CI test-gate steps | `bun test` (no type check) | `bun run test` (jsonb guard + progress-to-stdout guard + `tsc --noEmit` + `bun test`) | 1→4 |
| Regression guards on this bug class | 0 | 2 (compile-time via `tsc`, runtime via `makeAnthropic` injection test) | +2 |
The 104 → 0 isn't a refactor. Every error was a real correctness signal TS had been trying to send that nobody was listening for. Most were trivial to fix (`as unknown as X`, one missing union member, one rename propagation). The Buffer/BodyInit one in Supabase upload is a live bug — `fetch(url, {body: buf})` works today in Node/Bun but has no type guarantee; the fix copies `data.buffer, data.byteOffset, data.byteLength` into a `Uint8Array` slice that is genuinely assignable to `BodyInit`.
### What this means for operators
`gbrain agent run "say hello"` against a Supabase brain completes end-to-end after this upgrade. No stuck subagent jobs, no `client.create is not a function` traceback. v0.16.0 users should upgrade immediately — the feature that release was named for did not work.
### Itemized changes
#### `gbrain agent run` now works against the real Anthropic SDK
- `src/core/minions/handlers/subagent.ts` — factory default construction replaced with `const client: MessagesClient = deps.client ?? makeAnthropic().messages`. The SDK's `Messages` resource is already the right object; no helper, no wrapper, no `.bind()` needed (method-call semantics preserve `this`). `const makeAnthropic = deps.makeAnthropic ?? (() => new Anthropic())` adds a dependency-injection seam so tests can exercise the default branch without a real API key or network call.
- `test/subagent-handler.test.ts` — new `describe('makeSubagentHandler default client construction')` block drives a full handler turn through a fake SDK injected via `makeAnthropic`. If anyone reverts `.messages` or reintroduces a `new Anthropic()` top-level cast, this test fails loudly.
#### CI type-checking is now real
- `package.json` — added `typescript@^5.6.0` as devDep; added `"typecheck": "tsc --noEmit"` script; chained `bun run typecheck` into `"test"` so local `bun run test` and CI run identical pipelines (grep guards + typecheck + bun test).
- `.github/workflows/test.yml` — CI now runs `bun run test` (the npm script) instead of `bun test` (the runner). One line. Biggest-leverage change in the release.
#### 100+ pre-existing type errors cleaned up
So `tsc --noEmit` actually stays green. All mechanical, zero behavior change. Groups:
- **`databaseUrl` → `database_url` rename drift** in 9 test fixtures (test/agent-cli, test/brain-allowlist, test/minions-shell, test/minions, test/queue-child-done, test/rate-leases, test/subagent-handler, test/subagent-transcript, test/wait-for-completion).
- **`PageType` union** in `src/core/types.ts` gained `'meeting'` and `'note'` entries. Both were already used in src (`link-extraction.ts` had a code comment acknowledging the gap) and across 6 test files. The union was just out of date.
- **`GBrainConfig.storage`** field declared in `src/core/config.ts` — the code at `src/commands/files.ts` and `src/core/operations.ts` was reading `config.storage` with 18 inferred-type errors.
- **`ErrorCode`** union in `src/core/operations.ts` gained `'permission_denied'`; the code was throwing this exact string but the union disagreed.
- **Dead-code comparisons** removed from `src/commands/migrations/v0_12_0.ts`, `v0_12_2.ts`, `v0_13_0.ts`, `v0_16_0.ts` — each orchestrator had an early-return on `a.status === 'failed'` followed later by a redundant check against a then-narrowed type. TS correctly flagged the later check as always-false.
- **postgres.js `Row` callback typing** on `src/core/postgres-engine.ts` — 6 `.map((r: { slug: string }) => r.slug)` callbacks rewritten as `.map((r) => r.slug as string)` to match postgres.js's `Row` generic. Same behavior, correct signature.
- **Buffer → BodyInit** in `src/core/storage/supabase.ts:58,129` — `body: data` (Buffer) replaced with `body: new Uint8Array(data.buffer, data.byteOffset, data.byteLength) as BodyInit`. Zero-copy view of the same bytes, structurally assignable to `BodyInit`, no runtime change.
- **Various `as X` casts** upgraded to `as unknown as X` where TS 5.6's stricter structural-conversion rules rejected the single-step cast. Affected: `src/core/file-resolver.ts` (3), `src/core/minions/handlers/subagent-aggregator.ts`, `src/core/minions/worker.ts`, `src/commands/orphans.ts`, `src/commands/repair-jsonb.ts`, `src/core/postgres-engine.ts` (2 RowList → array conversions).
#### Test suite stability
- `bunfig.toml` — new file. Sets `[test].timeout = 60_000` globally. PGLite WASM init is slow enough that the default 5-second hook timeout flakes when many test files spin up PGLite instances in parallel on a loaded machine.
- 8 test files (`test/wait-for-completion`, `test/extract-fs`, `test/subagent-handler`, `test/minions-shell`, `test/minions-quiet-hours`, `test/integrity`, `test/e2e/graph-quality`, `test/e2e/search-quality`) additionally declare `beforeAll(fn, 60_000)` / `beforeEach(fn, 15_000)` as explicit safety nets — redundant with `bunfig.toml` today, but stays as belt-and-suspenders if the bunfig schema ever changes.
## To take advantage of v0.16.3
`gbrain upgrade` should do this automatically. If it didn't, or if `gbrain doctor` warns about anything:
1. **Verify your brain still runs:**
```bash
gbrain doctor
```
2. **Verify the agent runtime works:**
```bash
gbrain agent run "say hello"
```
Should complete end-to-end. If it fails with `client.create is not a function`, the upgrade didn't land — run `gbrain upgrade` again.
3. **No migrations required.** No schema changes in this release. Fix is in the handler code, not the DB.
4. **If any step fails,** please file an issue: https://github.com/garrytan/gbrain/issues with:
- output of `gbrain doctor`
- output of `gbrain agent run "say hello"`
- contents of `~/.gbrain/upgrade-errors.jsonl` if it exists
### Itemized changes
---
## [0.16.2] - 2026-04-22
## **The deployment guide now reads like a runbook an agent can execute line-by-line.**
## **Three real bugs from v0.16.1 fixed, nine DX gaps closed.**
v0.16.1 shipped the Minions worker deployment guide. Re-reading it as the agent it was written for, top-to-bottom, copy-pasting every block, surfaced twelve issues a human skim-reader would not catch. Three are real bugs that break a first-time deploy. Nine are structural gaps that force the agent to invent values.
The bugs: the crontab example used `*/5 * * * * user bash /path/...` which is `/etc/crontab` format only, so an agent running `crontab -e` and pasting it got "bad minute" or parsed `user` as the command. The watchdog script grepped `tail -20` of an unrotated log for shutdown markers, so every 5-minute tick after the first restart re-matched the old shutdown line forever and killed the healthy worker on loop. And `DATABASE_URL=postgresql://user:pass@...` lived directly in `/etc/crontab`, which is mode 644 (world-readable).
The gaps: no preconditions block, no "which option should I pick" selector, hardcoded `/path/to/...` and `/my/workspace` throughout with no template-variable legend, no upgrade section (so an agent coming from v0.13.x had no idea `GBRAIN_ALLOW_SHELL_JOBS=1` is now required or that `max_stalled` flipped from 1 to 5), no alternative to bare cron for Fly/Render/systemd deployments, a "Proposed CLI flags (not yet implemented)" block that an agent would copy and get `unrecognized flag`, and a `MinionWorker.maxStalledCount` note that did not tell the agent what to do.
### What this means for operators
The guide is now copy-pasteable without invention. Every `$VAR` is documented in a table at the top. Every code block runs as-is on the target it claims. The watchdog writes a two-line PID file (PID + restart epoch) and the shutdown check only considers log lines newer than the epoch, which is the actual fix for the restart loop. Secrets live in `/etc/gbrain.env` (mode 600), referenced via `BASH_ENV=/etc/gbrain.env` in crontab. A new Option 3 ships a systemd unit, a Procfile, and a fly.toml fragment so Fly/Render/Railway/systemd users skip cron entirely. The upgrade section walks the v0.13.x → v0.16.2 checklist (stop worker, apply migrations, add `GBRAIN_ALLOW_SHELL_JOBS`, swap the watchdog).
The shipped watchdog was verified against an abbreviated end-to-end test (3 ticks in ~30 seconds inside an Ubuntu 22.04 container): tick 1 starts the worker and writes the 2-line PID file; tick 2 sees a shutdown line with a 1-hour-old timestamp and correctly does nothing; tick 3 sees a fresh shutdown line and correctly restarts. The regex was caught and fixed during the test when mawk rejected `{n}` interval quantifiers. The systemd unit was smoked in a privileged container with `Restart=always` firing a second banner after a 10-second `RestartSec` window, confirming crash-recovery works before any host ever boots the unit.
## To take advantage of v0.16.2
`gbrain upgrade` pulls the new guide. If you deployed under v0.16.1 with the original watchdog, swap it:
1. **Re-read the guide:**
```bash
less docs/guides/minions-deployment.md
```
2. **Swap the watchdog script.** The v0.16.1 version has the restart-loop bug:
```bash
sudo install -m 755 docs/guides/minions-deployment-snippets/minion-watchdog.sh \
/usr/local/bin/minion-watchdog.sh
```
3. **Move secrets out of crontab.** Put `DATABASE_URL` and `GBRAIN_ALLOW_SHELL_JOBS=1` into `/etc/gbrain.env` (mode 600), reference it from crontab via `BASH_ENV=/etc/gbrain.env`.
4. **Fix the cron form.** If you pasted the v0.16.1 `*/5 * * * * user bash ...` into `crontab -e`, drop the `user` column and the explicit `bash` prefix.
5. **If you have shell access to a long-running box,** consider Option 3 (systemd) instead of Option 1 (watchdog). systemd replaces the watchdog entirely and is the cleanest path.
No schema change. No data migration. Docs + snippets only.
### Itemized changes
**Fixed**
- **Crontab syntax now matches the target.** Two labeled blocks: 5-field for `crontab -e`, 6-field with user column for `/etc/crontab`. An agent no longer hits "bad minute" or has `user` parsed as the command.
- **Watchdog restart loop killed.** The shipped `minion-watchdog.sh` writes a two-line PID file (PID on line 1, restart epoch on line 2) and only considers log lines whose ISO-8601 timestamp is newer than the epoch. Stale shutdown lines from earlier restarts no longer re-match every 5 minutes forever. Regex rewritten to use explicit `[0-9][0-9][0-9][0-9]` instead of `{4}` intervals because mawk (Debian/Ubuntu's default awk) rejects interval quantifiers. Verified end-to-end in a 3-tick abbreviated test inside Ubuntu 22.04.
- **Credentials off the world-readable filesystem.** Secrets move to `/etc/gbrain.env` (mode 600, owned by the worker user), referenced via `BASH_ENV=/etc/gbrain.env` in crontab. `/etc/crontab` is mode 644 and user crontabs under `/var/spool/cron/` are readable by root. A new `gbrain.env.example` ships in-repo with the full env surface.
**Added**
- **Preconditions block.** Five checks at the top of the guide: `gbrain` on PATH, DB connectivity, schema version, crontab write access, and the `GBRAIN_ALLOW_SHELL_JOBS=1` requirement for shell-job workers. Agent fails fast on setup, not content.
- **Decision tree.** "Which option?" selector at the top of the deployment section. Subagent workloads and long jobs take Option 1. Scheduled scripts take Option 2. No shell access take Option 3. Replaces the previous "recommended for X" prose that forced re-reading.
- **Template variable table.** Six variables (`$GBRAIN_BIN`, `$GBRAIN_WORKER_USER`, `$GBRAIN_WORKER_PID_FILE`, `$GBRAIN_WORKER_LOG_FILE`, `$GBRAIN_WORKSPACE`, `$GBRAIN_ENV_FILE`) with meaning and typical value. Agent substitutes once, everything downstream lands correctly.
- **Upgrade section.** v0.13.x → v0.16.2 checklist: stop the worker, run migrations, add `GBRAIN_ALLOW_SHELL_JOBS=1` for shell jobs, handle the `max_stalled` default flip from 1 to 5, swap the v0.16.1 watchdog for the current one.
- **Option 3: service manager.** New `systemd.service`, `Procfile`, and `fly.toml.partial` ship under `docs/guides/minions-deployment-snippets/`. systemd replaces the watchdog entirely with `Restart=always` + `RestartSec=10s` and runs the worker as an unprivileged user with `PrivateTmp`, `ProtectSystem=strict`, and `ReadWritePaths`. Smoked end-to-end in a privileged container: banner fired twice across a 10-second restart cycle, `Restart=always` honored, unit enabled for boot persistence.
- **Uninstall section.** One-paragraph rollback for each option.
- **`docs/guides/minions-deployment.md` listed in `scripts/llms-config.ts`.** Remote agents fetching `llms.txt` or `llms-full.txt` now see the deployment guide without having to guess its path.
**Changed**
- **`--follow` example uses a gbrain subcommand, not `node my-script.mjs`.** The new example submits `gbrain embed --stale` as a shell job on a dedicated queue with `--timeout-ms 600000`. Maps directly onto how an OpenClaw-style agent actually schedules brain maintenance.
- **"Proposed CLI flags (not yet implemented)" dead-end removed.** Replaced with a "Tune per-job today" callout pointing at the `gbrain jobs submit` flags that exist in source (`--max-stalled`, `--backoff-type`, `--backoff-delay`, `--backoff-jitter`, `--timeout-ms`, `--idempotency-key` — all first-class since v0.13.1).
- **Known Issues rewritten as imperatives.** "DO NOT pass `maxStalledCount` to `MinionWorker`" leads the paragraph, followed by the reason and the correct knob (`gbrain jobs submit --max-stalled N`). Zombie-shell-children section leads with the 10s / 30s numbers and the action.
Contributed by garrytan (issue report), fixes verified by an abbreviated end-to-end test suite (render-check + watchdog 3-tick + systemd container smoke + `bun test` + full E2E DB lifecycle).
## [0.16.1] - 2026-04-22
## **Minions worker deployment, finally documented.**
## **If you run `gbrain jobs work` in production, there's now a guide for the sharp edges.**
Garry's OpenClaw (gbrain's own instance, out there actually running `gbrain jobs work` in production) wrote a real deployment guide for the Minions worker, the piece of gbrain most operators hit next after getting sync running. Agents dogfooding the project they live on is a weird, good feedback loop. Two patterns: a watchdog cron for persistent workers, and an inline `--follow` for cron-only workloads. It covers the connection-drop, stall-detector, and zombie-child traps that show up once your brain is actually working for you. Every command and every default in the guide is checked against current source (`max_stalled = 5`, not 1 or 3; `--follow` exits on submitted-job-terminal, not queue-empty; stalled jobs show up as `active`, not `waiting`). Nothing about this was obvious, and nothing about it was in the docs before.
With v0.16.0's durable agent runtime now shipping, the persistent worker is load-bearing for a lot more (`subagent` + `subagent_aggregator` handlers run there too). A supervised deployment story is the sharp end of the stick.
### What this means for operators
If you have been running the Minions worker under `nohup` with no restart story, this guide is the missing manual. Copy the watchdog script, paste the crontab env lines (`SHELL=/bin/bash`, `PATH`, `DATABASE_URL`, `GBRAIN_ALLOW_SHELL_JOBS=1`), and wire the cron to run every 5 minutes. You get a restart loop that handles the three silent-death modes: DB connection blip, lock-renewal stall, event loop wedge.
If you are running scheduled shell jobs only, skip the persistent worker and use `--follow`. 2-3 seconds of startup overhead is trivial when your job runs for a minute.
Docs-only release. No code changed. Zero migration required.
## To take advantage of v0.16.1
`gbrain upgrade` pulls the new guide. Read it:
1. **Open the guide:**
```bash
less docs/guides/minions-deployment.md
```
Or browse it on GitHub.
2. **Persistent worker:** copy `minion-watchdog.sh`, set crontab env lines, wire a `*/5 * * * *` cron.
3. **Scheduled shell jobs only:** rewrite your cron as `gbrain jobs submit shell ... --follow --timeout-ms N` and drop the persistent worker entirely.
4. **The "Proposed CLI flags" section** (`--lock-duration` / `--max-stalled` / `--stall-interval` on `gbrain jobs work`): those are on the roadmap. Per-job `--max-stalled` on `gbrain jobs submit` is already real and writes to the row's column directly.
### Itemized changes
**Added**
- **Minions worker deployment guide** — new `docs/guides/minions-deployment.md` covering watchdog cron patterns, inline `--follow` for cron-only workloads, and the sharp edges of running `gbrain jobs work` against Supabase in production. Addresses a real gap: existing Minions docs (`minions-fix.md`, `minions-shell-jobs.md`) cover schema repair and shell-job security, not deploy patterns. Contributed by your OpenClaw via #287. Pre-landing accuracy pass corrected five factual bugs against current source: the `max_stalled` column default (5, not 1 or 3), the stalled-jobs smoke-test query (`active`, not `waiting`), the SIGTERM-to-SIGKILL grace window (10s minimum, not 2s), the cron env pattern (crontab env lines, not `source ~/.bashrc`), and the `--follow` exit semantics (blocks until submitted job is terminal, not until queue is empty).
## [0.16.0] - 2026-04-20
## **Durable agents land. Your LLM loops survive crashes, timeouts, and worker restarts now.**
## **OpenClaw died mid-run? Come back, resume from the last committed turn.**
Your OpenClaw crashes daily. Not "sometimes." Daily. An 8-turn OpenClaw subagent fires a tool call, the worker dies on a memory blip, all eight turns of context are gone, and there's nothing to do but start over from turn zero. This release kills that. `gbrain agent run` submits an Anthropic Messages API conversation as a first-class Minion job: every turn persists to `subagent_messages`, every tool call is a two-phase ledger row (`pending` → `complete | failed`), and replay on worker restart picks up from exactly the last committed turn. Crash-safe by construction, not by hope.
Fan-out works the same way. `--fanout-manifest` splits N prompts across N subagent children plus one aggregator. Children run `on_child_fail: 'continue'` so one failing run doesn't cascade, and the aggregator claims after all children reach ANY terminal state (complete, failed, dead, cancelled, timeout) and writes a mixed-outcome summary. No polling loop, no dead parents stranded in `waiting-children`.
Plugins work. Host repos drop a `gbrain.plugin.json` + `subagents/*.md` dir somewhere on `GBRAIN_PLUGIN_PATH`, and their custom subagent defs load at worker startup. Your OpenClaw ships its meeting-ingestion, signal-detector, and daily-task-prep subagents in its own repo now; gbrain discovers them day one. Collision rule is deterministic (left-wins with a loud warning). Trust boundary is strict on purpose: plugins ship DEFS, not tools. Tool allow-list stays here.
### The numbers that matter
Measured on the v0.15 branch against real Postgres via `bun run test:e2e`, plus the 159 new unit tests across 10 new test files. Coverage: 12 new runtime modules, 53+ code paths + user flows traced, 3 critical regression tests for the shell-jobs queue surface.
| Metric | BEFORE v0.15 | AFTER v0.15 | Δ |
|----------------------------------------------------------|------------------------------------|---------------------------------------------|--------------------------------------|
| Your OpenClaw run survives worker kill mid-tool-call | No (start over) | Yes (resume from last committed turn) | crash-recovery unlocked |
| Fan-out run with 1 failed child out of N | Aggregator fails | Aggregator still claims + summarizes | mixed-outcome aggregation works |
| `gbrain agent logs --follow` during long Anthropic call | Silent (looks frozen) | Heartbeat line per turn boundary | visible progress |
| Tool-use replay on resume | N/A (no resume) | Idempotent re-run, non-idempotent aborts | two-phase protocol |
| `put_page` exposure to agent-driven writes | Full write surface | Namespace-scoped `wiki/agents/<id>/…` | fail-closed, server-enforced |
| Plugin subagent defs for downstream hosts | Not supported | `GBRAIN_PLUGIN_PATH` + validated at startup | OpenClaw day-1 usable |
| Rate-lease capacity leaks on worker crash | Counter-based (leaks) | Lease-based (auto-prune on next acquire) | no starvation after SIGKILL |
| Anthropic prompt cache on 40-turn agent | Per-turn cold | `cache_control: ephemeral` on system + tools | ~10x cost reduction (best-case) |
### What this means for your OpenClaw
You stop rerunning from zero. A crash at 3am that used to lose two hours of turns now costs you whatever fraction of one turn was in-flight when the worker died. The rest of the conversation is rows in `subagent_messages` and `subagent_tool_executions`, and the next worker claim replays from there. `gbrain agent logs <job>` shows you where it died, which tool it was running, and what came back from the last successful call. Real debugging, not guessing.
Credit: shell-jobs (v0.14) established every pattern v0.15 reuses — handler signature, dual-signal abort, ctx.updateTokens, protected-names, trusted-submit, JSONL audit log, timeout_ms. Codex caught the Mode A "transparent Agent() interception" impossibility during plan review and saved the shape of this work. The v0.15 handler is what survives on the other side of that review.
### Itemized changes
**New capability: `gbrain agent` CLI**
- `gbrain agent run <prompt> [--subagent-def|--model|--max-turns|--tools|--timeout-ms|--fanout-manifest|--follow|--detach]` — submits a subagent job (or fan-out of N subagents + aggregator) under the trusted-submit flag. Follow mode tails status + logs until terminal; detach prints the job id and exits. Ctrl-C detaches (job keeps running), does not cancel.
- `gbrain agent logs <job_id> [--follow] [--since ISO-or-relative]` — merges the JSONL heartbeat audit with persisted `subagent_messages` into one chronological timeline. `--since 5m` / `1h` / `2d` shorthand supported. Transcript tail renders the full message + tool tree only after the job is terminal.
- Always registered on the worker (no separate env flag). `ANTHROPIC_API_KEY` is the natural cost gate — no key, the SDK call fails immediately. Who-can-submit is already gated by `PROTECTED_JOB_NAMES` + `TrustedSubmitOpts` so only the trusted-CLI path can insert `subagent` / `subagent_aggregator` rows.
**New durability primitives**
- `src/core/minions/handlers/subagent.ts` — the LLM-loop handler. Two-phase tool persistence, replay reconciliation for mid-dispatch crashes, dual-signal abort (`ctx.signal` + `ctx.shutdownSignal`), Anthropic prompt caching on system + tool defs, injectable `MessagesClient` for mocking.
- `src/core/minions/handlers/subagent-aggregator.ts` — claims AFTER all children resolve (Lane 1B's queue changes guarantee each terminal child posts a `child_done` inbox message), produces deterministic mixed-outcome markdown summary.
- `src/core/minions/rate-leases.ts` — lease-based concurrency cap for outbound providers. Owner-tagged rows with `expires_at` auto-prune on acquire, so a crashed worker can't strand capacity. `pg_advisory_xact_lock` guards the check-then-insert.
- `src/core/minions/wait-for-completion.ts` — poll-until-terminal helper for CLI callers. `TimeoutError` does NOT cancel the job; AbortSignal exits cleanly. Default `pollMs`: 1000 on Postgres, 250 on PGLite inline.
- `src/core/minions/handlers/subagent-audit.ts` — JSONL audit + heartbeat writer. Rotates weekly via ISO week. `readSubagentAuditForJob` is the readback path for `gbrain agent logs`.
- `src/core/minions/transcript.ts` — messages + tool executions → markdown renderer. UTF-8-safe truncation; unknown block types fall through to JSON for diagnostics.
- `src/core/minions/tools/brain-allowlist.ts` — derives the subagent tool registry from `src/core/operations.ts`. 11-name allow-list (read-only + deterministic `put_page`). `put_page` schema is namespace-wrapped per subagent so the model writes correct slugs first-try; the server-side check in `put_page` is the authoritative gate.
- `src/core/minions/plugin-loader.ts` — `GBRAIN_PLUGIN_PATH` (colon-separated absolute paths like `PATH`) + `gbrain.plugin.json` manifest + `subagents/*.md` defs. Strict path policy, left-wins collision, plugins ship DEFS only (no new tools), `allowed_tools:` validated at load time.
- `src/mcp/tool-defs.ts` — extracted from an inline `operations.map(...)` block in the MCP server so subagent + MCP use the same source of truth. Byte-for-byte equivalence pinned by regression test.
**Schema (3 new tables + OperationContext fields + migration orchestrator)**
- `subagent_messages` — Anthropic message-block persistence. `(job_id, message_idx)` UNIQUE; `content_blocks JSONB` holds parallel tool_use blocks in one assistant message.
- `subagent_tool_executions` — two-phase ledger. `(job_id, tool_use_id)` UNIQUE; status: `pending | complete | failed`.
- `subagent_rate_leases` — lease-based concurrency control. CASCADE deletes on owning job removal so no leaked rows.
- `OperationContext` gains `jobId?`, `subagentId?`, and `viaSubagent?` (fail-closed signal for agent-path gating). Added to `src/core/operations.ts`.
- `src/commands/migrations/v0_15_0.ts` — post-upgrade orchestrator (phases: schema → verify → record). `v0_14_0.ts` noop stub keeps the registry version sequence gapless.
**Queue correctness fixes**
- `failJob`, `cancelJob`, and `handleTimeouts` all emit `child_done` inbox messages with `outcome: 'complete' | 'failed' | 'dead' | 'cancelled' | 'timeout'`. Pre-v0.15 only `completeJob` emitted; failed/cancelled/timed-out children silently stranded aggregator-style parents.
- Parent-resolution terminal set expanded from `{completed, dead, cancelled}` to include `'failed'` everywhere parent-state is checked. A failed child with `on_child_fail: 'continue'` now correctly unblocks the parent.
- `failJob` emits `child_done` BEFORE the parent-terminal UPDATE. Without insertion ordering, the EXISTS guard on the inbox INSERT would skip the row on `fail_parent` paths (caught by codex iteration 3).
- `MinionJobInput.max_stalled` threads through `MinionQueue.add()` as INSERT param (not UPDATE on idempotency replay — that would mutate first-submitter state).
**Trust model**
- `subagent` and `subagent_aggregator` join `PROTECTED_JOB_NAMES`. MCP `submit_job` returns `permission_denied`; only `gbrain agent run` (with `allowProtectedSubmit`) can insert these rows.
- `put_page` gains a server-side fail-closed namespace check: when `ctx.viaSubagent === true`, `slug` MUST match `^wiki/agents/<subagentId>/.+` — even if `subagentId` is undefined (dispatcher bug must not open a hole).
**Docs**
- `docs/guides/plugin-authors.md` — downstream-OpenClaw-facing walkthrough (minimum viable plugin, path + collision + trust policies, frontmatter fields, caveats).
- 12 bisectable commits on `garrytan/minions-seam`, each PR-worthy on its own; the full series lands v0.15.0 end-to-end.
**Tests**
- 159 new unit tests across 10 new files: `mcp-tool-defs`, `put-page-namespace`, `migrations-v0_15_0`, `queue-child-done`, `rate-leases`, `wait-for-completion`, `brain-allowlist`, `subagent-audit`, `subagent-transcript`, `subagent-handler`, `subagent-aggregator`, `plugin-loader`, `agent-cli`.
- 3 critical regression tests pin the shell-jobs queue surface: `failJob` child_done behavior, `put_page` namespace path for non-subagent callers, MCP `buildToolDefs` byte-equivalence.
- E2E `minions-resilience.test.ts` updated: the max_children test renames its spawned children off the now-protected `subagent` name.
## [0.15.4] - 2026-04-21
## **PgBouncer transaction-mode prepared statements, fixed at the pool.**
@@ -565,7 +986,7 @@ Three new migrations, all idempotent, apply automatically on `gbrain init` / upg
- **Strict-mode default flip.** BrainWriter ships with `strict_mode=lint`. The flip to strict requires a 7-day soak + BrainBench regression ≤1pt + zero false-positive count.
- **Sandboxed user plugins.** v0.13 ships builtins only. User-provided TS modules deferred pending a real isolation story (worker_threads or vm2) in a follow-on release.
- **`openai_embedding` refactor.** Deferred to PR 1.5 post-flip; embedding is a hot path.
- **Wintermute `claw-bridge`.** Adoption path is documentation-only this release.
- **OpenClaw `claw-bridge`.** Adoption path is documentation-only this release.
### Tests
@@ -585,7 +1006,7 @@ Three new migrations, all idempotent, apply automatically on `gbrain init` / upg
Four subcommands: `check` (read-only report with `--json`, `--type`, `--limit`), `auto` (three-bucket repair with `--confidence`, `--review-lower`, `--dry-run`, `--fresh`, `--limit`), `review` (prints queue path + count), `reset-progress`. Nine bare-tweet phrase regexes. External-link extraction for optional dead-link probing. Repairs route through `BrainWriter.transaction`.
#### BudgetLedger + CompletenessScorer (`src/core/enrichment/`)
`BudgetLedger.reserve` returns `{kind:'held'}` or `{kind:'exhausted'}`. FOR UPDATE serializes concurrent reserves. `commit`, `rollback`, `cleanupExpired`. Midnight rollover via `Intl.DateTimeFormat` en-CA in configured IANA tz. Seven per-type rubrics + default (weights sum to 1.0). Person rubric's `non_redundancy` and `recency_score` kill Wintermute's length-only heuristic + 30-day-re-enrich-forever pathologies.
`BudgetLedger.reserve` returns `{kind:'held'}` or `{kind:'exhausted'}`. FOR UPDATE serializes concurrent reserves. `commit`, `rollback`, `cleanupExpired`. Midnight rollover via `Intl.DateTimeFormat` en-CA in configured IANA tz. Seven per-type rubrics + default (weights sum to 1.0). Person rubric's `non_redundancy` and `recency_score` kill Garry's OpenClaw's length-only heuristic + 30-day-re-enrich-forever pathologies.
#### Minions scheduler polish (`src/core/minions/`)
`quiet-hours.ts` — pure `evaluateQuietHours(cfg, now?)`. Wrap-around windows. Unknown tz fails open. `stagger.ts` — FNV-1a → 059 deterministic across runtimes. `worker.ts` integrated: post-claim evaluation, defer → `delayed/+15m`, skip → `cancelled`.
@@ -952,7 +1373,7 @@ Your brain now wires itself. Every page write automatically extracts entity refe
- **Auto-link on every page write.** When you `gbrain put` a page that mentions `[Alice](people/alice)` or `[Acme](companies/acme)`, those links land in the graph automatically. Stale links (refs no longer in the page text) are removed in the same call. Run a quick `gbrain put` and the brain knows who's connected to whom. To opt out: `gbrain config set auto_link false`.
- **Typed relationships.** Inferred from context using deterministic regex (zero LLM calls): `attended` (meeting -> person), `works_at` (CEO of, VP at, joined as), `invested_in` (invested in, backed by), `founded` (founded, co-founded), `advises` (advises, board member), `source` (frontmatter), `mentions` (default). On a 80-page benchmark brain: 94% type accuracy.
- **`gbrain extract --source db`.** New mode for the existing `gbrain extract <links|timeline|all>` command that walks pages from the engine instead of from disk. Works for live brains backed by Postgres or PGLite without a local markdown checkout — exactly what an MCP-driven Wintermute or OpenClaw setup needs. Filesystem mode (`--source fs`) is unchanged and still the default.
- **`gbrain extract --source db`.** New mode for the existing `gbrain extract <links|timeline|all>` command that walks pages from the engine instead of from disk. Works for live brains backed by Postgres or PGLite without a local markdown checkout — exactly what an MCP-driven OpenClaw setup needs. Filesystem mode (`--source fs`) is unchanged and still the default.
- **`gbrain graph-query <slug>` for relationship traversal.** "Who works at Acme?" → `gbrain graph-query companies/acme --type works_at --direction in`. "Who attended meetings with Alice?" → `gbrain graph-query people/alice --type attended --depth 2`. Returns typed edges with depth, not just nodes. Backed by a new `traversePaths()` engine method on both PGLite and Postgres with cycle prevention (no exponential blowup on cyclic subgraphs).
- **Graph-powered search ranking.** Hybrid search now applies a small backlink boost after cosine re-scoring (`score *= 1 + 0.05 * log(1 + backlink_count)`). Well-connected entities surface higher in results. Works in both keyword-only and full hybrid paths. Tested on the new `test/benchmark-graph-quality.ts` (80 pages, 35 queries, A/B/C comparison) — relational query recall jumps from ~30% (search alone) to 100% (graph traversal).
- **Graph health metrics in `gbrain health`.** New `link_coverage` and `timeline_coverage` percentages on entity pages (person/company), plus `most_connected` top-5 list. The `dead_links` field is dropped (always 0 under ON DELETE CASCADE — was a phantom metric). The `brain_score` composite formula stays but now reflects a sharper graph signal.
@@ -1015,7 +1436,7 @@ CLI wrappers (`runExtract`, `runEmbed`, etc.) stay as thin arg-parsers that catc
### Added — skillify ships as a first-class gbrain skill
Ported from Wintermute, proven in production. Paired with `gbrain check-resolvable` gives a user-controllable equivalent of Hermes' auto-skill-creation — you decide when and what, the tooling keeps the 10-item checklist honest.
Ported from Garry's OpenClaw, proven in production. Paired with `gbrain check-resolvable` gives a user-controllable equivalent of Hermes' auto-skill-creation — you decide when and what, the tooling keeps the 10-item checklist honest.
- `skills/skillify/SKILL.md` — the meta skill. Triggers: "skillify this", "is this a skill?", "make this proper".
- `scripts/skillify-check.ts` — machine-readable audit. `--json` for CI, `--recent` to check files modified in the last 7 days.
@@ -1183,7 +1604,7 @@ Wave 3 fixes were contributed by **@garagon** (PRs #105-#109) and **@Hybirdss**
| **cron-scheduler** | Schedule staggering (5-min offsets), quiet hours (timezone-aware with wake-up override), thin job prompts. | 21 cron jobs at :00 is a thundering herd. Staggering prevents it. Quiet hours mean no 3 AM notifications. Wake-up override releases the backlog. |
| **reports** | Timestamped reports with keyword routing. "What's the latest briefing?" maps to the right report directory. | Cheap replacement for vector search on frequent queries. Don't embed. Load the file. |
| **testing** | Validates every skill has SKILL.md with frontmatter, manifest coverage, resolver coverage. The CI for your skill system. | 3 skills and you need validation. 24 skills and you need it yesterday. Catches dead references, missing sections, MECE violations. |
| **soul-audit** | 6-phase interview that generates SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md. Your agent's identity, built from your answers. | What makes Wintermute feel like Wintermute. Without personality and access control, every agent feels the same. |
| **soul-audit** | 6-phase interview that generates SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md. Your agent's identity, built from your answers. | What makes your OpenClaw feel like yours. Without personality and access control, every agent feels the same. |
| **webhook-transforms** | External events (SMS, meetings, social mentions) converted into brain pages with entity extraction. Dead-letter queue for failures. | Your brain ingests signals from everywhere. Not just conversations, but every webhook, every notification, every external event. |
### Infrastructure (new in v0.10.0)
+31 -1
View File
@@ -43,6 +43,8 @@ strict behavior when unset.
- `src/commands/eval.ts``gbrain eval` command: single-run table + A/B config comparison
- `src/core/embedding.ts` — OpenAI text-embedding-3-large, batch, retry, backoff
- `src/core/check-resolvable.ts` — Resolver validation: reachability, MECE overlap, DRY checks, structured fix objects. v0.14.1: `CROSS_CUTTING_PATTERNS.conventions` is an array (notability gate accepts both `conventions/quality.md` and `_brain-filing-rules.md`). New `extractDelegationTargets()` parses `> **Convention:**`, `> **Filing rule:**`, and inline backtick references. DRY suppression is proximity-based via `DRY_PROXIMITY_LINES = 40`.
- `src/core/repo-root.ts` — Shared `findRepoRoot(startDir?)` (v0.16.4): walks up from `startDir` (default `process.cwd()`) looking for `skills/RESOLVER.md`. Zero-dependency module imported by both `doctor.ts` and `check-resolvable.ts`. Parameterized `startDir` makes tests hermetic.
- `src/commands/check-resolvable.ts` — Standalone CLI wrapper (v0.16.4) over `checkResolvable()`. Exports `parseFlags`, `resolveSkillsDir`, `DEFERRED`, `runCheckResolvable`. Exit rule: **1 on any issue (warnings OR errors)**, stricter than doctor's `ok` flag — honors README:259. Stable JSON envelope `{ok, skillsDir, report, autoFix, deferred, error, message}` — same shape on success and error paths. `--fix` path runs `autoFixDryViolations` BEFORE `checkResolvable` (same ordering as doctor). `deferred[]` array surfaces pending Checks 5 (trigger routing eval) and 6 (brain filing) with issue URLs. `scripts/skillify-check.ts` subprocess-calls `gbrain check-resolvable --json` (cached per process) and fails loud on binary-missing — no silent false-pass.
- `src/core/dry-fix.ts``gbrain doctor --fix` engine. `autoFixDryViolations(fixes, {dryRun})` rewrites inlined rules to `> **Convention:** see [path](path).` callouts via three shape-aware expanders (bullet / blockquote / paragraph). Five guards: working-tree-dirty (`getWorkingTreeStatus()` returns 3-state `'clean' | 'dirty' | 'not_a_repo'`), no-git-backup, inside-code-fence, already-delegated (40-line proximity, consistent with detector), ambiguous-multi-match, block-is-callout. `execFileSync` array args (no shell — no injection surface). EOF newline preserved.
- `src/core/backoff.ts` — Adaptive load-aware throttling: CPU/memory checks, exponential backoff, active hours multiplier
- `src/core/fail-improve.ts` — Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generation
@@ -59,8 +61,19 @@ strict behavior when unset.
- `src/core/minions/protected-names.ts` — side-effect-free constant module exporting `PROTECTED_JOB_NAMES` + `isProtectedJobName()`. Kept pure so queue core can import without loading handler modules.
- `src/core/minions/handlers/shell.ts``shell` job handler. Spawns `/bin/sh -c cmd` (absolute path, PATH-override-safe) or `argv[0] argv[1..]` (no shell). Env allowlist: `PATH, HOME, USER, LANG, TZ, NODE_ENV` + caller `env:` overrides. UTF-8-safe stdout/stderr tail via `string_decoder.StringDecoder`. Abort (either `ctx.signal` or `ctx.shutdownSignal`) fires SIGTERM → 5s grace → SIGKILL on child. Requires `GBRAIN_ALLOW_SHELL_JOBS=1` on worker (gated by `registerBuiltinHandlers`).
- `src/core/minions/handlers/shell-audit.ts` — per-submission JSONL audit trail at `~/.gbrain/audit/shell-jobs-YYYY-Www.jsonl` (ISO-week rotation; override via `GBRAIN_AUDIT_DIR`). Best-effort: `mkdirSync(recursive)` + `appendFileSync`; failures logged to stderr, submission not blocked. Logs cmd (first 80 chars) or argv (JSON array). Never logs env values.
- `src/core/minions/handlers/subagent.ts` (v0.15) — LLM-loop handler. Two-phase tool persistence (pending → complete/failed), replay reconciliation for mid-dispatch crashes, dual-signal abort (`ctx.signal` + `ctx.shutdownSignal`), Anthropic prompt caching on system + tool defs. `makeSubagentHandler({engine, client?, ...})` factory; `MessagesClient` is an injectable interface the real SDK implements structurally. Throws `RateLeaseUnavailableError` (renewable) when rate-lease capacity is full.
- `src/core/minions/handlers/subagent-aggregator.ts` (v0.15) — `subagent_aggregator` handler. Claims AFTER all children resolve (queue changes guarantee every terminal child posts a `child_done` inbox message with outcome). Reads inbox via `ctx.readInbox()`, builds deterministic mixed-outcome markdown summary. No LLM call in v0.15.
- `src/core/minions/handlers/subagent-audit.ts` (v0.15) — JSONL audit + heartbeat writer at `~/.gbrain/audit/subagent-jobs-YYYY-Www.jsonl`. Events: `submission` (one line per submit) + `heartbeat` (per turn boundary: `llm_call_started | llm_call_completed | tool_called | tool_result | tool_failed`). Never logs prompts or tool inputs. `readSubagentAuditForJob(jobId, {sinceIso})` is the readback path for `gbrain agent logs`.
- `src/core/minions/rate-leases.ts` (v0.15) — lease-based concurrency cap for outbound providers (default key `anthropic:messages`, max via `GBRAIN_ANTHROPIC_MAX_INFLIGHT`). Owner-tagged rows with `expires_at` auto-prune on acquire; `pg_advisory_xact_lock` guards check-then-insert; CASCADE on owning job deletion. `renewLeaseWithBackoff` retries 3x (250/500/1000ms).
- `src/core/minions/wait-for-completion.ts` (v0.15) — poll-until-terminal helper for CLI callers. `TimeoutError` does NOT cancel the job; `AbortSignal` exits without throwing. Default `pollMs`: 1000 on Postgres, 250 on PGLite inline.
- `src/core/minions/transcript.ts` (v0.15) — renders `subagent_messages` + `subagent_tool_executions` to markdown. Tool rows splice under their owning assistant `tool_use` by `tool_use_id`. UTF-8-safe truncation; unknown block types fall through to fenced JSON.
- `src/core/minions/plugin-loader.ts` (v0.15) — `GBRAIN_PLUGIN_PATH` discovery. Absolute paths only, left-wins collision, `gbrain.plugin.json` with `plugin_version: "gbrain-plugin-v1"`, plugins ship DEFS only (no new tools), `allowed_tools:` validated at load time against the derived registry.
- `src/core/minions/tools/brain-allowlist.ts` (v0.15) — derives subagent tool registry from `src/core/operations.ts`. 11-name allow-list: `query`, `search`, `get_page`, `list_pages`, `file_list`, `file_url`, `get_backlinks`, `traverse_graph`, `resolve_slugs`, `get_ingest_log`, `put_page`. `put_page` schema is namespace-wrapped per subagent (`^wiki/agents/<subagentId>/.+`); the `put_page` op's server-side check is the authoritative gate via `ctx.viaSubagent` fail-closed.
- `src/mcp/tool-defs.ts` (v0.15) — extracted `buildToolDefs(ops)` helper. MCP server + subagent tool registry both call it; byte-for-byte equivalence pinned by `test/mcp-tool-defs.test.ts`.
- `src/core/minions/attachments.ts` — Attachment validation (path traversal, null byte, oversize, base64, duplicate detection)
- `src/commands/jobs.ts``gbrain jobs` CLI subcommands + `gbrain jobs work` daemon. v0.13.1 surfaces the full `MinionJobInput` retry/backoff/timeout/idempotency surface as first-class CLI flags on `jobs submit`: `--max-stalled`, `--backoff-type fixed|exponential`, `--backoff-delay`, `--backoff-jitter`, `--timeout-ms`, `--idempotency-key`. `jobs smoke --sigkill-rescue` is the opt-in regression guard for #219.
- `src/commands/agent.ts` (v0.16)`gbrain agent run <prompt> [flags]` CLI. Submits `subagent` (or N children + 1 aggregator) under `{allowProtectedSubmit: true}`. Single-entry `--fanout-manifest` short-circuits. Children get `on_child_fail: 'continue'` + `max_stalled: 3`. `--follow` is the default on TTY; streams logs + polls `waitForCompletion` in parallel. Ctrl-C detaches, does not cancel.
- `src/commands/agent-logs.ts` (v0.16) — `gbrain agent logs <job> [--follow] [--since]`. Merges JSONL heartbeat audit + `subagent_messages` into a chronological timeline. `parseSince` accepts ISO-8601 or relative (`5m`, `1h`, `2d`). Transcript tail renders only for terminal jobs.
- `src/commands/jobs.ts``gbrain jobs` CLI subcommands + `gbrain jobs work` daemon. v0.13.1 surfaces the full `MinionJobInput` retry/backoff/timeout/idempotency surface as first-class CLI flags on `jobs submit`: `--max-stalled`, `--backoff-type fixed|exponential`, `--backoff-delay`, `--backoff-jitter`, `--timeout-ms`, `--idempotency-key`. `jobs smoke --sigkill-rescue` is the opt-in regression guard for #219. v0.16 wires `registerBuiltinHandlers` to always register `subagent` + `subagent_aggregator` (no env flag — `ANTHROPIC_API_KEY` is the natural cost gate, trust is via `PROTECTED_JOB_NAMES`) and loads `GBRAIN_PLUGIN_PATH` plugins at worker startup with a loud startup-line per plugin. `shell` handler still gated by `GBRAIN_ALLOW_SHELL_JOBS=1` (RCE surface, separate concern).
- `src/commands/features.ts``gbrain features --json --auto-fix`: usage scan + feature adoption salesman
- `src/commands/autopilot.ts``gbrain autopilot --install`: self-maintaining brain daemon (sync+extract+embed)
- `src/mcp/server.ts` — MCP stdio server (generated from operations)
@@ -73,6 +86,8 @@ strict behavior when unset.
- `src/core/migrate.ts` — schema-migration runner. Owns the `MIGRATIONS` array (source of truth for schema DDL). v0.14.2 extended the `Migration` interface with `sqlFor?: { postgres?, pglite? }` (engine-specific SQL overrides `sql`) and `transaction?: boolean` (set to false for `CREATE INDEX CONCURRENTLY`, which Postgres refuses inside a transaction; ignored on PGLite since it has no concurrent writers). Migration v14 (fix wave) uses a handler branching on `engine.kind` to run CONCURRENTLY on Postgres (with a pre-drop of any invalid remnant via `pg_index.indisvalid`) and plain `CREATE INDEX` on PGLite. v15 bumps `minion_jobs.max_stalled` default 1→5 and backfills existing non-terminal rows.
- `src/core/progress.ts` — Shared bulk-action progress reporter. Writes to stderr. Modes: `auto` (TTY: `\r`-rewriting; non-TTY: plain lines), `human`, `json` (JSONL), `quiet`. Rate-gated by `minIntervalMs` and `minItems`. `startHeartbeat(reporter, note)` helper for single long queries. `child()` composes phase paths. Singleton SIGINT/SIGTERM coordinator emits `abort` events for every live phase. EPIPE defense on both sync throws and stream `'error'` events. Zero dependencies. Introduced in v0.15.2.
- `src/core/cli-options.ts` — Global CLI flag parser. `parseGlobalFlags(argv)` returns `{cliOpts, rest}` with `--quiet` / `--progress-json` / `--progress-interval=<ms>` stripped. `getCliOptions()` / `setCliOptions()` expose a module-level singleton so commands reach the resolved flags without parameter threading. `cliOptsToProgressOptions()` maps to reporter options. `childGlobalFlags()` returns the flag suffix to append to `execSync('gbrain ...')` calls in migration orchestrators. `OperationContext.cliOpts` extends shared-op dispatch for MCP callers.
- `src/core/cycle.ts` — v0.17 brain maintenance cycle primitive. `runCycle(engine: BrainEngine | null, opts: CycleOpts): Promise<CycleReport>` composes 6 phases in semantically-driven order (lint → backlinks → sync → extract → embed → orphans). Three callers: `gbrain dream` CLI, `gbrain autopilot` daemon's inline path, and the Minions `autopilot-cycle` handler (`src/commands/jobs.ts`). One source of truth for what the brain does overnight. Coordination via `gbrain_cycle_locks` DB table (TTL-based; works through PgBouncer transaction pooling, unlike session-scoped `pg_try_advisory_lock`) + `~/.gbrain/cycle.lock` file lock with PID-liveness for PGLite / engine=null mode. `CycleReport.schema_version: "1"` is the stable agent-consumable shape. `PhaseResult.error: { class, code, message, hint?, docs_url? }` is Stripe-API-tier structured failure info. `yieldBetweenPhases` hook awaited between every phase — Minions handler uses this to renew its job lock and prevent v0.14 stall-death regression. Engine nullable: filesystem phases (lint, backlinks) run without DB; DB phases skip with `status: "skipped", reason: "no_database"`. Lock-skip: read-only phase selections (`--phase orphans`) bypass the cycle lock.
- `src/commands/dream.ts` — v0.17 `gbrain dream` CLI. ~80-line thin alias over `runCycle`. brainDir resolution requires explicit `--dir` OR `sync.repo_path` config (no more walk-up-cwd-for-.git footgun). Flags: `--dry-run`, `--json`, `--phase <name>`, `--pull`, `--dir <path>`. Exit code 1 on status=failed (partial/warn not fatal — don't page on warnings).
- `scripts/check-progress-to-stdout.sh` — CI guard against regressing to `\r`-on-stdout progress. Wired into `bun run test` via `scripts/check-progress-to-stdout.sh && bun test` in package.json.
- `docs/progress-events.md` — Canonical JSON event schema reference. Stable from v0.15.2, additive only.
- `src/core/markdown.ts` — Frontmatter parsing + body splitter. `splitBody` requires an explicit timeline sentinel (`<!-- timeline -->`, `--- timeline ---`, or `---` immediately before `## Timeline`/`## History`). Plain `---` in body text is a markdown horizontal rule, not a separator. `inferType` auto-types `/wiki/analysis/` → analysis, `/wiki/guides/` → guide, `/wiki/hardware/` → hardware, `/wiki/architecture/` → architecture, `/writing/` → writing (plus the existing people/companies/deals/etc heuristics).
@@ -557,6 +572,21 @@ search engines, surfaced in cross-references, and distributed with every release
- Example meeting → `meetings/2026-04-03` (generic date is fine)
- Example user → `you` or `the user`, never a proper name
**Specific rule: never say `Wintermute` in any CHANGELOG, README, doc, PR, or
commit message.** When the temptation is to illustrate with the real fork name:
- Reader-facing copy → `your OpenClaw` (covers Wintermute, Hermes, AlphaClaw,
and any other downstream OpenClaw deployment in one term the reader already
recognizes).
- First-person / origin-story copy → `Garry's OpenClaw` (honest that this is
the production deployment driving the feature, without exposing the private
agent's name).
`Wintermute` may appear in private artifacts (scratch plans under
`~/.gstack/projects/…`, memory files, conversation transcripts, CEO-review
plans) — those aren't distributed. Anything checked into this repo or shipped
in a release must use the OpenClaw phrasing above. Sweeping a stale reference
is a small clean-up PR, not a debate.
**When in doubt, ask yourself:** "Would this query reveal private information
about the user's contacts, investments, or portfolio if it were read by a
stranger?" If yes, replace with generic placeholders.
+19
View File
@@ -230,6 +230,25 @@ If anything's off, `actions[]` tells you the exact command to run. For deeper tr
Moving gateway crons to Minions (deterministic scripts, zero LLM tokens per fire): [`docs/guides/minions-shell-jobs.md`](docs/guides/minions-shell-jobs.md).
## Durable agents: `gbrain agent` (v0.15)
Your subagent runs survive crashes now. OpenClaw died mid-run? The worker re-claims on restart and replays from the last committed turn. Fan-out across 50 shards, one shard crashes — the aggregator still claims after every child reaches a terminal state and writes a mixed-outcome summary. Tool calls persist as a two-phase ledger (`pending``complete | failed`) so replay is safe by construction, not by hope.
```bash
# Submit a single-subagent run
gbrain agent run "summarize my last 10 journal pages"
# Fan out N prompts across N subagent children + 1 aggregator
gbrain agent run "analyze every page" \
--fanout-manifest manifests/pages.json \
--subagent-def analyzer
# Tail a running job (heartbeat per turn + full transcript on completion)
gbrain agent logs 1247 --follow --since 5m
```
Durability is the point: every Anthropic turn commits to `subagent_messages`, every tool call to `subagent_tool_executions`. Worker kills, OpenClaw crashes, timeouts — all resumable. Host repos (your OpenClaw, etc.) ship their own subagent definitions via `GBRAIN_PLUGIN_PATH` + a `gbrain.plugin.json` manifest: see [`docs/guides/plugin-authors.md`](docs/guides/plugin-authors.md). Requires `ANTHROPIC_API_KEY` on the worker.
## Skillify: your skills tree stops being a black box
Hermes and similar agent frameworks auto-create skills as a background behavior. Fine until you don't know what the agent shipped. Checklists decay. Tests drift. Resolver entries get stale. Six months later you've got an opaque pile of "skills" that nobody has read, nobody has tested, and nobody is sure still work.
+17 -1
View File
@@ -1,5 +1,21 @@
# TODOS
## check-resolvable
### File tracking issues for Checks 5 + 6 (deferred in PR #325)
**Priority:** P2
**What:** `src/commands/check-resolvable.ts` currently points `DEFERRED[].issue` at GitHub issue search URLs (`?q=TBD-check-5`, `?q=TBD-check-6`). File real tracking issues and grep-replace both placeholders with the real URLs.
**Why:** v0.16.4 shipped `gbrain check-resolvable` with 4 of the 6 checks from the original spec. Checks 5 (trigger routing eval) and 6 (brain filing) were explicitly deferred during plan-ceo-review because they each need new detection logic. The CLI's `deferred[]` JSON field is meant to surface these to agents so they know the coverage boundary — the TBD placeholders do the right thing mechanically but aren't clickable.
**How:**
1. `gh issue create -t "check-resolvable Check 5: trigger routing eval" -b "..."` — detection: every skill's own frontmatter trigger should match the RESOLVER.md entry pointing at that skill. Needs new issue type (e.g. `mis_route`).
2. `gh issue create -t "check-resolvable Check 6: brain filing validation" -b "..."` — detection: scan SKILL.md body for brain paths (e.g., `brain/people/`, `brain/companies/`), cross-reference with `skills/_brain-filing-rules.md`. Flag mutating skills missing entries.
3. Replace `TBD-check-5` and `TBD-check-6` in `src/commands/check-resolvable.ts` with the real issue URLs.
**Effort:** ~15 min mechanical (issue filing + grep-replace). Implementation of the checks themselves is a separate, larger piece of work — the TODO here is just the issue filing + URL swap.
## P1 (BrainBench v1.1 — categories deferred from PR #188)
### BrainBench Cat 5: Source Attribution / Provenance
@@ -173,7 +189,7 @@ board" — likely an advisor-role page prior plus verb-pattern combinations.
**Cons:** Requires adding `sender_id` or `access_tier` to `OperationContext`. Each mutating operation needs a permission check. Medium implementation effort.
**Context:** From CEO review + Codex outside voice (2026-04-13). Prompt-layer access control works in practice (same model as Wintermute) but is not sufficient for remote MCP where direct tool calls bypass the agent's prompt.
**Context:** From CEO review + Codex outside voice (2026-04-13). Prompt-layer access control works in practice (same model as Garry's OpenClaw) but is not sufficient for remote MCP where direct tool calls bypass the agent's prompt.
**Depends on:** v0.10.0 GStackBrain skill layer (shipped).
+1 -1
View File
@@ -1 +1 @@
0.15.4
0.17.0
+9
View File
@@ -14,9 +14,12 @@
"openai": "^4.0.0",
"pgvector": "^0.2.0",
"postgres": "^3.4.0",
"tree-sitter-wasms": "0.1.13",
"web-tree-sitter": "0.22.6",
},
"devDependencies": {
"@types/bun": "latest",
"typescript": "^5.6.0",
},
},
},
@@ -452,10 +455,14 @@
"tr46": ["tr46@0.0.3", "", {}, "sha512-N3WMsuqV66lT30CrXNbEjx4GEwlow3v6rr4mCcv6prnfwhS01rkgyFdjPNBYd9br7LpXV1+Emh01fHnq2Gdgrw=="],
"tree-sitter-wasms": ["tree-sitter-wasms@0.1.13", "", { "dependencies": { "tree-sitter-wasms": "^0.1.11" } }, "sha512-wT+cR6DwaIz80/vho3AvSF0N4txuNx/5bcRKoXouOfClpxh/qqrF4URNLQXbbt8MaAxeksZcZd1j8gcGjc+QxQ=="],
"tslib": ["tslib@2.8.1", "", {}, "sha512-oJFu94HQb+KVduSUQL7wnpmqnfmLsOA/nAh6b6EH0wCEoK0/mPeXU6c3wKDV83MkOuHPRHtSXKKU99IBazS/2w=="],
"type-is": ["type-is@2.0.1", "", { "dependencies": { "content-type": "^1.0.5", "media-typer": "^1.1.0", "mime-types": "^3.0.0" } }, "sha512-OZs6gsjF4vMp32qrCbiVSkrFmXtG/AZhY3t0iAMrMBiAZyV9oALtXO8hsrHbMXF9x6L3grlFuwW2oAz7cav+Gw=="],
"typescript": ["typescript@5.9.3", "", { "bin": { "tsc": "bin/tsc", "tsserver": "bin/tsserver" } }, "sha512-jl1vZzPDinLr9eUt3J/t7V6FgNEw9QjvBPdysz9KfQDD41fQrC2Y4vKQdiaUpFT4bXlb1RHhLpp8wtm6M5TgSw=="],
"undici-types": ["undici-types@5.26.5", "", {}, "sha512-JlCMO+ehdEIKqlFxk6IfVoAUVmgz7cU7zD/h9XZ0qzeosSHmUJVOzSQvvYSYWXkFXC+IfLKSIffhv0sVZup6pA=="],
"unpipe": ["unpipe@1.0.0", "", {}, "sha512-pjy2bYhSsufwWlKwPc+l3cN7+wuJlK6uz0YdJEOlQDbl6jo/YlPi4mb8agUkVC8BF7V8NuzeyPNqRksA3hztKQ=="],
@@ -464,6 +471,8 @@
"web-streams-polyfill": ["web-streams-polyfill@4.0.0-beta.3", "", {}, "sha512-QW95TCTaHmsYfHDybGMwO5IJIM93I/6vTRk+daHTWFPhwh+C8Cg7j7XyKrwrj8Ib6vYXe0ocYNrmzY4xAAN6ug=="],
"web-tree-sitter": ["web-tree-sitter@0.22.6", "", {}, "sha512-hS87TH71Zd6mGAmYCvlgxeGDjqd9GTeqXNqTT+u0Gs51uIozNIaaq/kUAbV/Zf56jb2ZOyG8BxZs2GG9wbLi6Q=="],
"webidl-conversions": ["webidl-conversions@3.0.1", "", {}, "sha512-2JAn3z8AR6rjK8Sm8orRC0h/bcl/DqL7tRPdGZ4I1CjdF+EaMLmYxBHyXuKL849eucPFhvBoxMsflfOb8kxaeQ=="],
"whatwg-url": ["whatwg-url@5.0.0", "", { "dependencies": { "tr46": "~0.0.3", "webidl-conversions": "^3.0.0" } }, "sha512-saE57nupxk6v3HY35+jzBwYa0rKSy0XR8JSxZPwgLr7ys0IBzhGviA1/TUGJLmSVqs8pb9AnvICXEuOHLprYTw=="],
+6
View File
@@ -0,0 +1,6 @@
[test]
# PGLite initialization can be slow under parallel test execution.
# Default 5s is too short when many test files boot PGLite instances at once.
# 60s is the empirical ceiling we observed before the first file's beforeAll
# completed on a loaded machine.
timeout = 60_000
+100
View File
@@ -358,6 +358,106 @@ upcoming `gbrain crontab-to-minions <file>` helper is P1 in TODOS.
---
## v0.16.0: durable agent runtime
v0.15 ships `gbrain agent run` / `gbrain agent logs`, a new `subagent` handler
type in Minions, and a plugin contract for host-repo subagent defs. None of the
existing skills need surgery. The question for downstream agents is *how* to
adopt the new runtime, not how to patch around a breaking change.
### 1. Run a worker with an Anthropic key
The subagent handlers (`subagent` and `subagent_aggregator`) are always
registered on the worker. No separate opt-in flag — `ANTHROPIC_API_KEY` is
the natural cost gate (no key, the SDK call fails on the first turn), and
who-can-submit is already protected (`PROTECTED_JOB_NAMES` + trusted-submit:
MCP callers get `permission_denied`; only `gbrain agent run` can insert
these rows).
```bash
ANTHROPIC_API_KEY=sk-ant-... gbrain jobs work
```
Worker startup prints:
```
[minion worker] subagent handlers enabled
```
### 2. Ship your subagents as a plugin (OpenClaw + similar)
Move your custom subagent definitions out of your gbrain fork and into your own
repo as a plugin. Concretely:
```
~/<your-agent>/gbrain-plugin/
├── gbrain.plugin.json
└── subagents/
├── meeting-ingestion.md
├── signal-detector.md
└── daily-task-prep.md
```
`gbrain.plugin.json`:
```json
{
"name": "your-openclaw",
"version": "2026.4.20",
"plugin_version": "gbrain-plugin-v1"
}
```
Each `subagents/*.md` is a plain-text agent definition — YAML frontmatter +
body-as-system-prompt. Recognized frontmatter fields: `name`, `model`,
`max_turns`, `allowed_tools` (must subset the derived brain-tool registry).
Turn it on:
```bash
export GBRAIN_PLUGIN_PATH="$HOME/<your-agent>/gbrain-plugin"
```
Worker startup prints `[plugin-loader] loaded '<name>' v<ver> (N subagents)`
per plugin; any rejection (bad manifest, unknown tool in `allowed_tools`,
version mismatch) shows up as a loud warning at startup, not a silent dispatch-
time failure. See `docs/guides/plugin-authors.md` for the full contract.
### 3. Replace ephemeral subagent runs with durable ones
If your agent currently spawns ephemeral subagents (OpenClaw `Agent()`, ad-hoc
Anthropic API calls, etc.) for work that should survive crashes, sleeps, or
worker restarts, migrate those to `gbrain agent run`. The durability is free:
```bash
gbrain agent run "analyze my last 50 journal pages for recurring themes" \
--subagent-def analyzer --fanout-manifest manifests/journal-pages.json
```
Every turn persists to `subagent_messages`, every tool call is a two-phase
ledger, and `gbrain agent logs <job>` shows where it died + what the last
successful call returned. No more "re-run from scratch because the session
context evaporated."
### 4. `put_page` from subagents writes under an agent namespace
If you adopted the v0.15 subagent runtime, note that `put_page` calls
originating from a subagent's tool dispatch MUST target
`wiki/agents/<subagent_id>/...`. The schema shown to the model enforces this
on first try; a server-side fail-closed check rejects anything else. This
does NOT affect your skill files, CLI put_page calls, or MCP put_page —
only tool-dispatched writes from inside an LLM loop.
Aggregation output (the final "here's what all N children found" brain page)
goes via a separate trusted CLI path, not through a subagent tool call, so
it can write anywhere you want.
Iron rule: **never grant an agent write access beyond its namespace**. The
server-side check exists because dispatcher bugs happen; treat it as defense
in depth, not the primary boundary.
---
## Future versions
When gbrain ships a new version, this doc will be updated with the diffs for that
@@ -1,7 +1,7 @@
# Production Benchmark: Minions vs OpenClaw Sub-agents (Real Deployment)
**Date:** 2026-04-18
**Environment:** Wintermute on Render (ephemeral container, Supabase Postgres)
**Environment:** Garry's OpenClaw on Render (ephemeral container, Supabase Postgres)
**GBrain:** v0.11.0 (minions-jobs branch)
**OpenClaw:** 2026.4.10
**Brain:** 45,798 pages, 98K chunks, 25K links, 79K timeline entries
+27 -27
View File
@@ -8,9 +8,9 @@
## 0. Context
During a CEO review of a narrow two-feature plan (bare-tweet citation repair + completeness score, borrowed from Feynman), the scope was reframed. The narrow plan duplicated work Wintermute already does and missed the real leverage point: **the bespoke abstractions hiding inside Wintermute — resolvers, enrichment orchestration, scheduling, deterministic output — should live in GBrain as first-class primitives.**
During a CEO review of a narrow two-feature plan (bare-tweet citation repair + completeness score, borrowed from Feynman), the scope was reframed. The narrow plan duplicated work Garry's OpenClaw already does and missed the real leverage point: **the bespoke abstractions hiding inside OpenClaw — resolvers, enrichment orchestration, scheduling, deterministic output — should live in GBrain as first-class primitives.**
North star: *"When Wintermute's Claw upgrades to this version of GBrain, it should immediately recognize brilliance and completeness and say 'It's time to switch to these abstractions.'"*
North star: *"When Garry's OpenClaw's Claw upgrades to this version of GBrain, it should immediately recognize brilliance and completeness and say 'It's time to switch to these abstractions.'"*
That is the test this document is designed against. Everything else is downstream.
@@ -67,7 +67,7 @@ An earlier implementation could ship L1 + L4 first (the two "purest" layers) and
### 3.1 What's broken today
Wintermute has **69 distinct external-lookup patterns** across X API (14 shapes), Perplexity, Mistral OCR, Gmail, Calendar, Slack, GitHub, YouTube, Diarize.io, YC tools, OSINT collectors, and brain-local lookups. Each one is a bespoke script under `scripts/` with its own error handling, retry logic, and output shape. GBrain has 3 ad-hoc wrappers (`embedding.ts`, `transcription.ts`, `enrichment-service.ts`) that don't share an interface.
Garry's OpenClaw has **69 distinct external-lookup patterns** across X API (14 shapes), Perplexity, Mistral OCR, Gmail, Calendar, Slack, GitHub, YouTube, Diarize.io, YC tools, OSINT collectors, and brain-local lookups. Each one is a bespoke script under `scripts/` with its own error handling, retry logic, and output shape. GBrain has 3 ad-hoc wrappers (`embedding.ts`, `transcription.ts`, `enrichment-service.ts`) that don't share an interface.
Common consequences:
- No uniform retry/backoff strategy (some scripts retry, most don't)
@@ -187,7 +187,7 @@ Existing `src/core/fail-improve.ts` is the deterministic-first/LLM-fallback patt
### 3.7 Reference implementations to ship
The Wintermute survey inventoried 69 resolver shapes. Shipping all of them is wrong (over-scoped); shipping zero is under-scoped. The dogfood set:
The OpenClaw survey inventoried 69 resolver shapes. Shipping all of them is wrong (over-scoped); shipping zero is under-scoped. The dogfood set:
| # | Resolver | Purpose | Used by |
|---|---|---|---|
@@ -198,7 +198,7 @@ The Wintermute survey inventoried 69 resolver shapes. Shipping all of them is wr
| 5 | `perplexity_query` | Query → synthesis + citations | Enrichment Orchestrator |
| 6 | `text_to_entities` | LLM entity extraction (structured JSON) | Enrichment Orchestrator |
The remaining 63 Wintermute patterns port incrementally, driven by user need. Each port is a new YAML + module under `recipes/` or `~/.gbrain/resolvers/` with no framework changes.
The remaining 63 OpenClaw patterns port incrementally, driven by user need. Each port is a new YAML + module under `recipes/` or `~/.gbrain/resolvers/` with no framework changes.
---
@@ -206,7 +206,7 @@ The remaining 63 Wintermute patterns port incrementally, driven by user need. Ea
### 4.1 What's broken today
Wintermute's enrichment is **polished at the data layer, hacky at the control layer**:
Garry's OpenClaw's enrichment is **polished at the data layer, hacky at the control layer**:
- **Completeness = "length > 500 chars + no `needs-enrichment` tag"** (`lib/enrich.mjs:351-355`). Naïve. A rich page of repetitive Perplexity summaries (see `brain/people/0interestrates.md` — 38 repeating blocks) passes this check.
- **30-day auto-re-enrichment** runs forever. No "done" state. A person met once in 2023 still gets re-researched monthly.
@@ -342,9 +342,9 @@ await writer.transaction(async (tx) => {
### 5.1 What's broken today
Wintermute's cron is **externally-driven JSON** (`cron/jobs.json`) with ~30 jobs manually stagger-offset at different minutes. GBrain has **zero native scheduling**`src/commands/autopilot.ts` is a single daemon loop, and `docs/guides/cron-schedule.md` is architectural guidance, not code.
Garry's OpenClaw's cron is **externally-driven JSON** (`cron/jobs.json`) with ~30 jobs manually stagger-offset at different minutes. GBrain has **zero native scheduling**`src/commands/autopilot.ts` is a single daemon loop, and `docs/guides/cron-schedule.md` is architectural guidance, not code.
Failures observed in Wintermute's actual state:
Failures observed in Garry's OpenClaw's actual state:
- `X OAuth2 Token Refresh`: 11 consecutive timeouts (critical-path silent failure)
- `flight-tracker daily scan`: 5 consecutive timeouts
- `morning-briefing`: 4 consecutive timeouts
@@ -378,9 +378,9 @@ export interface ScheduledResolver extends Resolver<void, ScheduledResult> {
}
```
### 5.3 Enforcement vs convention (the key delta from Wintermute)
### 5.3 Enforcement vs convention (the key delta from Garry's OpenClaw)
| Concern | Wintermute today | Knowledge Runtime |
| Concern | Garry's OpenClaw today | Knowledge Runtime |
|---|---|---|
| Quiet hours | Checked inside each skill (trust-based) | Enforced at scheduler, skill cannot override |
| Staggering | Manual minute-offset in `jobs.json` | Scheduler assigns slots via hashed staggerKey |
@@ -405,7 +405,7 @@ Every scheduled run emits structured events: `started`, `skipped-quiet-hours`, `
- `engine.logIngest` (audit trail in brain DB)
- Optional webhook (Slack/Telegram for the user)
`gbrain doctor` reads the event log and reports: current circuit-breaker state, any resolver with > 3 consecutive failures, any resolver that hasn't fired within 3× its interval (freshness SLA like Wintermute's `freshness-check.mjs` but built-in).
`gbrain doctor` reads the event log and reports: current circuit-breaker state, any resolver with > 3 consecutive failures, any resolver that hasn't fired within 3× its interval (freshness SLA like Garry's OpenClaw's `freshness-check.mjs` but built-in).
---
@@ -415,9 +415,9 @@ Every scheduled run emits structured events: `started`, `skipped-quiet-hours`, `
**Iron Law: LLM picks WHAT. Code guarantees WHERE and HOW.**
Wintermute's existing `lib/enrich.mjs:buildTweetEntry` is close to this — tweet URLs are built from `tweet.id` returned by the X API, never from LLM memory. But:
Garry's OpenClaw's existing `lib/enrich.mjs:buildTweetEntry` is close to this — tweet URLs are built from `tweet.id` returned by the X API, never from LLM memory. But:
- A past incident: *"Sub-agent test #2 FAILED — hallucinated 'Philip Leung' entity links across all daily files. LLM rewriting of daily files is too error-prone."* (Wintermute memory log, 2026-04-13.)
- A past incident: *"Sub-agent test #2 FAILED — hallucinated 'Philip Leung' entity links across all daily files. LLM rewriting of daily files is too error-prone."* (Garry's OpenClaw memory log, 2026-04-13.)
- Back-links depend on `appendTimeline` being called everywhere; skips are silent.
- Slug collisions are unchecked (no conflict detection on `slugify`).
- Citation format is post-hoc linted weekly, not pre-write enforced.
@@ -461,7 +461,7 @@ export class Scaffolder {
// "[Source: [X/garrytan, 2026-04-18](https://x.com/garrytan/status/123456)]"
}
emailCitation(account: string, messageId: string, subject: string): string {
// deterministic Gmail URL per Wintermute pattern
// deterministic Gmail URL per OpenClaw pattern
}
sourceCitation(resolverResult: ResolverResult<unknown>): string {
// pulls .source, .fetchedAt, .raw from the result
@@ -563,7 +563,7 @@ Each phase ships independently, passes full E2E, is feature-flagged, and is reve
- L4 core: `BrainWriter.transaction`, `Scaffolder`, `SlugRegistry` with conflict detection.
- Pre-write validators: citation, link, back-link, triple-HR.
- Migrate `src/commands/publish.ts` + `src/commands/backlinks.ts` to route through BrainWriter.
- **Now** Wintermute's "Philip Leung" hallucination is structurally impossible — LLM output passes through JSON-Schema validator before reaching Scaffolder.
- **Now** Garry's OpenClaw's "Philip Leung" hallucination is structurally impossible — LLM output passes through JSON-Schema validator before reaching Scaffolder.
### Phase 3 — `gbrain integrity` command (human: ~0.5 wk / CC: ~2 h)
- Ship the originally-scoped user-facing feature on top of the new foundation.
@@ -582,14 +582,14 @@ Each phase ships independently, passes full E2E, is feature-flagged, and is reve
- Migrate `src/commands/autopilot.ts` to a ScheduledResolver set.
- Ship `gbrain schedule list|run|pause|tail` CLI for observability.
### Phase 6 — Port 58 Wintermute resolvers (human: ~1.5 wk / CC: ~6 h)
### Phase 6 — Port 58 OpenClaw resolvers (human: ~1.5 wk / CC: ~6 h)
- `perplexity_query`, `text_to_entities`, `mistral_ocr_pdf`, `x_search_all`, `x_user_to_tweets`, `gmail_query_to_threads`, `calendar_date_to_events`.
- Each ships as YAML + TS module under `resolvers/builtin/` — **proof of the plugin format.**
### Phase 7 — Wintermute Claw Adoption Integration (human: ~1 wk / CC: ~4 h)
- Write `docs/wintermute/ADOPTION.md` showing Wintermute how to replace its 69 bespoke scripts with calls to `gbrain registry.resolve(...)`.
- Ship a `gbrain claw-bridge` subcommand that proxies Wintermute's current script invocations to the resolver registry — zero-edit adoption path.
- **This is the test of the north star.** If Wintermute can stand up a 1-line shim and drop `scripts/x-api-client.mjs`, the abstraction succeeded.
### Phase 7 — OpenClaw Adoption Integration (human: ~1 wk / CC: ~4 h)
- Write `docs/openclaw/ADOPTION.md` showing your OpenClaw how to replace its 69 bespoke scripts with calls to `gbrain registry.resolve(...)`.
- Ship a `gbrain claw-bridge` subcommand that proxies Garry's OpenClaw's current script invocations to the resolver registry — zero-edit adoption path.
- **This is the test of the north star.** If your OpenClaw can stand up a 1-line shim and drop `scripts/x-api-client.mjs`, the abstraction succeeded.
Total: human: ~10 weeks / CC: ~42 hours / calendar with single implementer: ~34 weeks.
@@ -649,7 +649,7 @@ src/commands/
integrity.ts # ships in Phase 3, replaces Feynman Phase A/B
schedule.ts # gbrain schedule list|run|pause|tail (Phase 5)
docs/wintermute/
docs/openclaw/
ADOPTION.md # written in Phase 7
```
@@ -685,19 +685,19 @@ Every Resolver implementation tested against the interface spec. Table-driven: r
- Simulate API timeout mid-transaction; transaction must roll back completely.
- Corrupted state file; scheduler must escalate, not silently skip.
### Regression tests vs. Wintermute behavior
For each Wintermute pattern we port (e.g. X-handle → tweet URL), a regression test proves the new resolver produces the same answer on real-world inputs from the brain audit. This is the "Wintermute would adopt" proof.
### Regression tests vs. Garry's OpenClaw behavior
For each OpenClaw pattern we port (e.g. X-handle → tweet URL), a regression test proves the new resolver produces the same answer on real-world inputs from the brain audit. This is the "your OpenClaw would adopt" proof.
---
## 11. Open Questions (flagged for CEO re-review)
1. **Scope shape.** Is this the right four-layer decomposition, or are some layers better left to Wintermute (e.g. Scheduling lives above GBrain, not in it)?
1. **Scope shape.** Is this the right four-layer decomposition, or are some layers better left to OpenClaw (e.g. Scheduling lives above GBrain, not in it)?
2. **Phase 3 user-value break.** Does Phase 3 (user-visible `gbrain integrity`) ship early enough, or do we need an even smaller MVP?
3. **LLM-as-resolver.** Should `text_to_entities` be a Resolver, or does that blur the "code vs LLM" line the invariant relies on?
4. **Plugin format.** YAML + TS module (§3.5) vs. pure TS module with decorator-style metadata. Latter is more type-safe; former is more discoverable.
5. **Cross-resolver transactions.** Do we support "atomic fetch-from-Perplexity + write-to-brain" at the L2 layer? Current design says yes; implementation is tricky (Perplexity call isn't rollbackable).
6. **Wintermute bridge scope.** Phase 7 `gbrain claw-bridge` — is that worth a phase of its own, or should adoption be documentation-only?
6. **OpenClaw bridge scope.** Phase 7 `gbrain claw-bridge` — is that worth a phase of its own, or should adoption be documentation-only?
7. **Completeness rubric coverage.** Do we define rubrics for all 9 PageTypes upfront, or ship people/company/meeting first and extend incrementally?
8. **Budget config UX.** Hard daily cap is strict; should we also expose a soft-cap warning mode, and how is the cap set (env var? config file? prompt on first use?)
9. **Backwards compat.** `src/commands/publish.ts` and `src/commands/backlinks.ts` have been running cleanly for weeks. Refactoring through BrainWriter carries migration risk. Acceptable?
@@ -705,12 +705,12 @@ For each Wintermute pattern we port (e.g. X-handle → tweet URL), a regression
---
## 12. Verification (the "Wintermute would adopt" test)
## 12. Verification (the "your OpenClaw would adopt" test)
The design succeeds iff:
- [ ] A user can add a new resolver by dropping a YAML + TS module in `~/.gbrain/resolvers/` without editing GBrain source.
- [ ] Wintermute can delete `scripts/x-api-client.mjs` and replace all callers with 1-line `await registry.resolve('x_handle_to_tweet', ...)`.
- [ ] Your OpenClaw can delete `scripts/x-api-client.mjs` and replace all callers with 1-line `await registry.resolve('x_handle_to_tweet', ...)`.
- [ ] No brain page can be written with a bare tweet reference, a missing back-link, or an unverified URL (validators catch it pre-commit).
- [ ] Running `gbrain integrity --auto --confidence 0.8` over a real brain fixes ≥1,000 of the 1,424 known bare-tweet citations without human review.
- [ ] Full E2E test suite passes on both PGLite + Postgres engines.
@@ -0,0 +1,10 @@
# Procfile — Render / Railway / Heroku.
#
# Fly.io users: see fly.toml.partial instead.
#
# Set secrets via the platform's env UI or CLI (e.g. `heroku config:set`,
# `render env:set`, `railway variables set`). At minimum:
# DATABASE_URL=postgresql://...
# GBRAIN_ALLOW_SHELL_JOBS=1 # only if submitting shell jobs
worker: gbrain jobs work --concurrency 2
@@ -0,0 +1,22 @@
# fly.toml — partial. Merge into your existing fly.toml.
#
# Set secrets once (never commit them):
# fly secrets set DATABASE_URL='postgresql://user:pass@host:6543/db?prepare=false'
# fly secrets set GBRAIN_ALLOW_SHELL_JOBS=1 # only if submitting shell jobs
# fly secrets set ANTHROPIC_API_KEY=... # optional
#
# Fly.io auto-restarts the process on crash — no watchdog needed.
[processes]
worker = "gbrain jobs work --concurrency 2"
# Scale the worker process to 1 machine (job queue serializes work; more
# machines means higher concurrency but also more Postgres connections).
# fly scale count worker=1
# If you want the worker in its own VM size:
# [[vm]]
# processes = ["worker"]
# memory = "512mb"
# cpu_kind = "shared"
# cpus = 1
@@ -0,0 +1,35 @@
# /etc/gbrain.env — secrets + env for the gbrain worker.
#
# Install:
# sudo install -m 600 -o $GBRAIN_WORKER_USER -g $GBRAIN_WORKER_USER \
# gbrain.env.example /etc/gbrain.env
# sudoedit /etc/gbrain.env # fill in real values
#
# Referenced from crontab via BASH_ENV=/etc/gbrain.env, or from systemd
# via EnvironmentFile=/etc/gbrain.env. Never commit real secrets.
# --- Required ---------------------------------------------------------------
# Postgres connection string. For Supabase transaction pooler, include
# prepare=false (see CLAUDE.md #284/#286).
DATABASE_URL=postgresql://user:pass@host:6543/db?prepare=false
# --- Required if you submit `shell` jobs ------------------------------------
# Only the worker process needs this. Submitters do not.
GBRAIN_ALLOW_SHELL_JOBS=1
# --- Optional ---------------------------------------------------------------
# LLM provider keys (needed for `subagent` handler, transcription, enrichment).
# ANTHROPIC_API_KEY=
# OPENAI_API_KEY=
# Custom handler plugins (see docs/guides/plugin-handlers.md).
# GBRAIN_PLUGIN_PATH=/etc/gbrain/plugins
# Pool size tuning for Supabase transaction pooler (default 10; drop to 2
# if you hit MaxClients during upgrade subprocess spawns).
# GBRAIN_POOL_SIZE=2
# Connection-level concurrency cap for Anthropic Messages API.
# GBRAIN_ANTHROPIC_MAX_INFLIGHT=4
+68
View File
@@ -0,0 +1,68 @@
#!/bin/bash
# minion-watchdog.sh — restart gbrain jobs work if the process is dead or
# has logged a shutdown marker since its last start.
#
# Fixes the v0.16.1 restart-loop bug: old shutdown lines from previous
# restarts stayed in the unrotated log and every tick re-matched them
# forever. This version writes a restart epoch to line 2 of the PID file
# and only considers log lines newer than that epoch.
#
# Run every 5 minutes from crontab. See docs/guides/minions-deployment.md.
set -u
PID_FILE="${GBRAIN_WORKER_PID_FILE:-/tmp/gbrain-worker.pid}"
LOG_FILE="${GBRAIN_WORKER_LOG_FILE:-/tmp/gbrain-worker.log}"
GBRAIN="${GBRAIN_BIN:-/usr/local/bin/gbrain}"
CONCURRENCY="${GBRAIN_WORKER_CONCURRENCY:-2}"
start_worker() {
# stderr merged so banner lines ("[minion worker] shell handler enabled",
# "worker shutting down") all land in $LOG_FILE.
nohup "$GBRAIN" jobs work --concurrency "$CONCURRENCY" \
> "$LOG_FILE" 2>&1 &
local pid=$!
# Line 1: PID. Line 2: restart epoch (seconds since 1970).
# Readers that want just PID use `head -n1 "$PID_FILE"`.
printf '%s\n%s\n' "$pid" "$(date +%s)" > "$PID_FILE"
}
shutdown_since_restart() {
# Only match shutdown lines logged AFTER the most recent restart epoch.
# Worker log lines start with ISO-8601 UTC timestamps ("2026-04-21T19:05:12Z ...").
local restart_epoch
restart_epoch=$(sed -n '2p' "$PID_FILE" 2>/dev/null || echo 0)
[ -z "$restart_epoch" ] && restart_epoch=0
# POSIX-portable regex (no {n} intervals — mawk on Debian/Ubuntu rejects them).
awk -v since="$restart_epoch" '
match($0, /^[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9]T[0-9:.+Z-]+/) {
ts_str = substr($0, RSTART, RLENGTH)
cmd = "date -d \"" ts_str "\" +%s 2>/dev/null"
cmd | getline ts
close(cmd)
if (ts + 0 > since + 0) print
}
' "$LOG_FILE" 2>/dev/null | grep -q "worker stopped\|worker shutting down"
}
if [ -f "$PID_FILE" ]; then
PID=$(head -n1 "$PID_FILE")
if [ -n "$PID" ] && kill -0 "$PID" 2>/dev/null; then
# Process alive — check whether the worker logged an internal shutdown
# AFTER the last start. If yes, worker is dead-inside; restart.
if shutdown_since_restart; then
kill "$PID" 2>/dev/null
# 10s grace: covers shell handler's 5s child SIGTERM→SIGKILL window
# and leaves room for in-flight jobs to flush. Bump to 30 if your
# jobs run > 10s.
sleep 10
kill -9 "$PID" 2>/dev/null
start_worker
fi
else
# PID file exists but process is gone (crash / kill -9 / reboot).
start_worker
fi
else
start_worker
fi
@@ -0,0 +1,44 @@
[Unit]
Description=gbrain minion worker
Documentation=https://github.com/garrytan/gbrain/blob/master/docs/guides/minions-deployment.md
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
# Runs as an unprivileged user that owns the brain repo and any shell-job cwds.
# Create with: sudo useradd --system --home /srv/gbrain --shell /usr/sbin/nologin gbrain
User=gbrain
Group=gbrain
WorkingDirectory=/srv/gbrain
# Env file is mode 600, owned by User=. Do not put secrets in this unit.
EnvironmentFile=/etc/gbrain.env
ExecStart=/usr/local/bin/gbrain jobs work --concurrency 2
# Replaces the cron watchdog. systemd restarts on any non-zero exit.
Restart=always
RestartSec=10s
# Graceful shutdown: SIGTERM → wait → SIGKILL. 30s matches worker grace
# for in-flight jobs and the shell handler's 5s child SIGTERM window.
KillSignal=SIGTERM
TimeoutStopSec=30s
StandardOutput=journal
StandardError=journal
SyslogIdentifier=gbrain-worker
# Default 1024 is tight for Bun + Postgres pool + concurrent subagent LLM calls.
LimitNOFILE=65535
# Hardening (optional — remove if they break your deployment).
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=read-only
ReadWritePaths=/srv/gbrain
[Install]
WantedBy=multi-user.target
+323
View File
@@ -0,0 +1,323 @@
# Minions Worker Deployment Guide
Deploy `gbrain jobs work` so it stays running across crashes, reboots, and
Postgres connection blips. Written for agents to execute line-by-line.
## The problem
The persistent worker can die silently from:
- Database connection drops (Supabase/Postgres maintenance or network blips).
- Lock-renewal failures → the stall detector eventually dead-letters jobs.
- Bun process crashes with no automatic restart.
- Internal event-loop death (PID alive, worker loop stopped).
When the worker dies, submitted jobs sit in `waiting` forever. Nothing in
gbrain core auto-restarts the worker — that's what this guide wires up.
## Variables used in this guide
Substitute these once before copy-pasting any snippet.
| Variable | Meaning | Typical value |
|---|---|---|
| `$GBRAIN_BIN` | Absolute path to the `gbrain` binary | `$(command -v gbrain)` — often `/usr/local/bin/gbrain` or `~/.bun/bin/gbrain` |
| `$GBRAIN_WORKER_USER` | OS user that owns the worker process | the same user that ran `gbrain init`; never `root` |
| `$GBRAIN_WORKER_PID_FILE` | Worker PID + restart-epoch file | `/tmp/gbrain-worker.pid` (or `/var/run/gbrain/worker.pid` for systemd) |
| `$GBRAIN_WORKER_LOG_FILE` | Worker log sink (stdout + stderr merged) | `/tmp/gbrain-worker.log` (or `/var/log/gbrain/worker.log`) |
| `$GBRAIN_WORKSPACE` | `cwd` for shell jobs submitted by this deployment | absolute path, e.g. `/srv/my-brain` |
| `$GBRAIN_ENV_FILE` | Secrets file sourced by crontab / systemd | `/etc/gbrain.env` (mode 600) |
## Preconditions
Run these before Step 1 of any option. Fail fast if something is wrong.
```bash
# 1. gbrain is on PATH and resolves to an absolute location.
command -v gbrain || { echo "gbrain not on PATH. Install, then retry."; exit 1; }
# 2. DATABASE_URL points at reachable Postgres (or PGLite path exists).
gbrain doctor --fast --json | jq '.checks[] | select(.name=="db_connectivity")'
# 3. Schema is up to date. If version=0 or status=="fail", fix it first:
# gbrain apply-migrations --yes
gbrain doctor --fast --json | jq '.checks[] | select(.name=="schema_version")'
# 4. You have write access to at least one crontab mechanism.
crontab -l >/dev/null 2>&1 && echo "user crontab OK"
[ -w /etc/crontab ] && echo "/etc/crontab OK"
# 5. If you plan to submit `shell` jobs, the WORKER process needs
# GBRAIN_ALLOW_SHELL_JOBS=1 (submitters do not). The handler is gated
# in registerBuiltinHandlers(); without the flag the worker startup
# line reads "shell handler disabled (...)".
```
## Which option?
- Your workload runs LLM subagents (`gbrain agent run`) or jobs that take
> 30 s → **Option 1** (watchdog cron + persistent worker).
- Your workload is short deterministic scripts on a fixed schedule (every
3 h, daily, weekly) → **Option 2** (inline `--follow`).
- You don't have shell access to a long-running box (Fly/Render/Railway,
or any systemd host) → **Option 3** (service manager — replaces cron).
## Option 1: watchdog cron + persistent worker
A 5-minute cron checks whether the worker process is alive **and** whether
it has logged an internal shutdown since its last start. Restarts if either
condition fails.
### 1a. Install the env file (secrets stay out of crontab)
Never paste `DATABASE_URL` or API keys into crontab. `/etc/crontab` is
mode 644 (world-readable); user crontabs under `/var/spool/cron/` are
readable by `root`. Use the shipped env-file template:
```bash
sudo install -m 600 -o $GBRAIN_WORKER_USER -g $GBRAIN_WORKER_USER \
docs/guides/minions-deployment-snippets/gbrain.env.example /etc/gbrain.env
sudoedit /etc/gbrain.env
```
Fill in the connection string and `GBRAIN_ALLOW_SHELL_JOBS=1` (if
applicable). See
[`gbrain.env.example`](./minions-deployment-snippets/gbrain.env.example)
for the full list.
### 1b. Install the watchdog script
The [`minion-watchdog.sh`](./minions-deployment-snippets/minion-watchdog.sh)
ships in-repo and writes a two-line PID file (PID on line 1, restart epoch
on line 2). The restart-epoch marker is how the watchdog distinguishes
stale shutdown lines in the log from current ones — without it, every tick
after the first restart would match an old `worker shutting down` line and
loop forever.
Requires GNU coreutils (Linux default). On macOS/BSD install via
`brew install coreutils` and alias `date` to `gdate` in the cron env if you
want to test the watchdog locally; production Linux boxes work as-is.
```bash
sudo install -m 755 -o $GBRAIN_WORKER_USER -g $GBRAIN_WORKER_USER \
docs/guides/minions-deployment-snippets/minion-watchdog.sh \
/usr/local/bin/minion-watchdog.sh
```
### 1c. Wire into cron
Pick the form that matches the crontab you're editing.
**If you ran `crontab -e`** (user crontab — 5-field, no user column):
```
SHELL=/bin/bash
PATH=/usr/local/bin:/usr/bin:/bin
BASH_ENV=/etc/gbrain.env
*/5 * * * * /usr/local/bin/minion-watchdog.sh
```
**If you edited `/etc/crontab` directly** (system crontab — 6-field, with
user column):
```
SHELL=/bin/bash
PATH=/usr/local/bin:/usr/bin:/bin
BASH_ENV=/etc/gbrain.env
*/5 * * * * gbrain /usr/local/bin/minion-watchdog.sh
```
In both forms, `BASH_ENV=/etc/gbrain.env` tells non-interactive bash to
source the env file before running the watchdog — that's how the
connection string and `GBRAIN_ALLOW_SHELL_JOBS` reach the worker without
landing in the world-readable crontab itself.
### 1d. Log rotation
The watchdog appends to the worker log across restarts. If you expect the
file to grow unbounded, rotate it externally with `logrotate`:
```
# /etc/logrotate.d/gbrain-worker
/tmp/gbrain-worker.log {
daily
rotate 7
missingok
notifempty
copytruncate
}
```
`copytruncate` is important — the watchdog's restart-epoch check survives
it (the epoch is compared against in-log timestamps, not file inode).
## Option 2: inline `--follow` (no persistent worker)
Each cron run brings its own temporary worker. `--follow` starts one on
the queue and blocks until the just-submitted job reaches a terminal state
(`completed` / `failed` / `dead` / `cancelled`). 2-3 s startup overhead
per job; negligible vs job duration for scheduled work.
Example: nightly brain enrichment as a shell job.
```bash
GBRAIN_ALLOW_SHELL_JOBS=1 gbrain jobs submit shell \
--queue nightly-enrich \
--params "{\"cmd\":\"$GBRAIN_BIN embed --stale\",\"cwd\":\"$GBRAIN_WORKSPACE\"}" \
--follow \
--timeout-ms 600000
```
Replace `gbrain embed --stale` with whichever gbrain subcommand you're
scheduling (`sync`, `extract`, `orphans`, `doctor`, `check-backlinks`,
`lint`, `autopilot`). If you're shelling out to a non-gbrain binary,
keep its absolute path in the `cmd`.
**Shared-queue gotcha.** If other jobs are already waiting on the same
queue with higher priority or earlier `created_at`, the temporary worker
processes those first before reaching yours. `--follow` still exits only
when YOUR job finishes. For strict single-job semantics on shared queues,
use a dedicated queue name like `nightly-enrich` above.
## Option 3: service manager (systemd / Fly / Render / Railway)
Replaces the watchdog entirely. No cron, no PID file, no restart-loop.
The service manager owns liveness.
### systemd (Linux hosts with shell access)
```bash
# Create the worker user if it doesn't exist.
sudo useradd --system --home "$GBRAIN_WORKSPACE" --shell /usr/sbin/nologin gbrain \
2>/dev/null || true
sudo mkdir -p "$GBRAIN_WORKSPACE" && sudo chown gbrain:gbrain "$GBRAIN_WORKSPACE"
# Install the unit file, substituting /srv/gbrain → your workspace path.
sudo install -m 644 docs/guides/minions-deployment-snippets/systemd.service \
/etc/systemd/system/gbrain-worker.service
sudo sed -i "s|/srv/gbrain|$GBRAIN_WORKSPACE|g" \
/etc/systemd/system/gbrain-worker.service
# See 1a above for /etc/gbrain.env install.
sudo systemctl daemon-reload
sudo systemctl enable --now gbrain-worker
sudo systemctl status gbrain-worker
journalctl -u gbrain-worker -n 50
```
`Restart=always` + `RestartSec=10s` give you crash-loop recovery. The unit
runs as an unprivileged `gbrain` user with `PrivateTmp`, `ProtectSystem=strict`,
and `ReadWritePaths=$GBRAIN_WORKSPACE`. `LimitNOFILE=65535` in the shipped
unit covers Bun + Postgres pool + concurrent LLM subagent calls without
hitting the default 1024 cap.
### Fly.io
Merge the `[processes]` block from
[`fly.toml.partial`](./minions-deployment-snippets/fly.toml.partial) into
your existing `fly.toml`. Set secrets with `fly secrets set`
Fly auto-restarts the process on crash.
### Render / Railway / Heroku
Drop [`Procfile`](./minions-deployment-snippets/Procfile) at the repo root.
Set the connection string and `GBRAIN_ALLOW_SHELL_JOBS=1` via the
platform's env UI or CLI.
## Upgrading an existing deployment
If you deployed on v0.13.x or earlier, walk this checklist:
1. **Stop the worker before upgrading.**
`kill $(head -n1 /tmp/gbrain-worker.pid)` and wait for the process to
exit. Skipping this risks an in-flight job landing partial schema.
2. **Run `gbrain upgrade`**. Then `gbrain apply-migrations --yes` if
`gbrain doctor` reports any migration as `partial` or `pending`.
3. **If you run shell jobs:** from v0.14 onward, the worker requires
`GBRAIN_ALLOW_SHELL_JOBS=1` to register the `shell` handler. Add it to
`/etc/gbrain.env`. Submitters don't need the flag; only the worker does.
4. **If you tuned your watchdog for `max_stalled=1`:** v0.14.3 migration
v15 raised the schema default to 5 and backfilled existing non-terminal
rows. A watchdog tuned around 1-strike dead-lettering will now
over-restart because it takes 5 misses to dead-letter. Switch to the
shipped watchdog (which keys on log markers, not job state).
5. **If your v0.16.1 watchdog is still running:** it has a restart-loop
bug (old shutdown lines in the unrotated log re-match every 5 min
forever). Install the current `minion-watchdog.sh` from this guide's
snippets — it writes a restart epoch into the PID file and only
considers log lines newer than that epoch.
6. **Verify.** `gbrain doctor` should report zero `pending` or `partial`
migrations. `gbrain jobs stats` should show no unexplained growth in
`dead` between pre- and post-upgrade.
## Known issues
### Supabase connection drops
The worker uses a single Postgres connection. If Supabase drops it
(maintenance, connection limits, network blip), lock renewal fails
silently. The stall detector then dead-letters the job after
`max_stalled` misses.
**Current defaults that make this worse:**
- `lockDuration: 30000` (30 s) — too short for long jobs during connection blips.
- `max_stalled: 5` (schema column default on master — see `src/schema.sql`
and `src/core/pglite-schema.ts`). Five missed heartbeats before dead-letter.
- `stalledInterval: 30000` (30 s) — checks too aggressively.
**Tune per-job today.** `gbrain jobs submit` accepts `--max-stalled N`,
`--backoff-type fixed|exponential`, `--backoff-delay <ms>`,
`--backoff-jitter 0..1`, and `--timeout-ms N` as first-class flags
(since v0.13.1). These write onto the job row at submit time — which is
what `handleStalled()` reads — so per-job tuning is the real knob today.
Worker-level `--lock-duration` / `--stall-interval` are on the roadmap;
until they land, rely on per-job `--max-stalled` plus the watchdog (or
systemd) for worker health.
### DO NOT pass `maxStalledCount` to `MinionWorker`
It's a no-op. The stall detector reads the row's `max_stalled` column
(set at submit time), not the worker opt in `src/core/minions/worker.ts:74`.
Use `gbrain jobs submit --max-stalled N` per-job instead.
### Zombie shell children
When the Bun worker crashes hard, child processes from shell jobs can
become zombies. The watchdog's 10 s `SIGTERM → SIGKILL` window covers the
shell handler's 5 s child-kill grace (`KILL_GRACE_MS`). For long-running
shell jobs, bump the watchdog's `sleep 10` to `sleep 30` so the worker
has time to flush in-flight jobs before the kill.
## Smoke test
```bash
# Worker alive?
kill -0 $(head -n1 /tmp/gbrain-worker.pid) 2>/dev/null && echo ALIVE || echo DEAD
# Aggregate queue health.
gbrain jobs stats
# Jobs currently stalled (still `active` with expired lock_until, pre-requeue).
gbrain jobs list --status active --limit 10
# Dead-lettered jobs.
gbrain jobs list --status dead --limit 10
# Shell handler registered? (stderr banner merged into log via 2>&1.)
grep "shell handler enabled" /tmp/gbrain-worker.log
```
## Uninstall
- **Option 1 (watchdog cron):** `crontab -e`, delete the watchdog line.
`kill $(head -n1 /tmp/gbrain-worker.pid) && rm /tmp/gbrain-worker.pid`.
Optionally `sudo rm /etc/gbrain.env /usr/local/bin/minion-watchdog.sh`.
- **Option 2 (inline `--follow`):** remove the cron entry. Nothing else to
clean up — temporary workers exit with their jobs.
- **Option 3 (systemd):** `sudo systemctl disable --now gbrain-worker`,
then `sudo rm /etc/systemd/system/gbrain-worker.service /etc/gbrain.env`,
then `sudo systemctl daemon-reload`.
- **Option 3 (Fly/Render/Railway):** delete the `worker` process from
`fly.toml` / `Procfile` and redeploy. Secrets set via `fly secrets`
persist until `fly secrets unset`.
+1 -1
View File
@@ -73,7 +73,7 @@ F. Install gbrain autopilot --install (env-aware)
G. Record append completed.jsonl status:"complete"
```
If Phase E emits TODOs for host-specific handlers (e.g. Wintermute's
If Phase E emits TODOs for host-specific handlers (e.g. your OpenClaw's
~29 non-gbrain crons), the migration finishes with `status: "partial"`.
Your host agent walks the TODOs using `skills/migrations/v0.11.0.md` +
`docs/guides/plugin-handlers.md`, ships handler registrations in the
+163
View File
@@ -0,0 +1,163 @@
# Plugin authors guide (v0.15)
`gbrain` discovers subagent definitions from outside this repo via
`GBRAIN_PLUGIN_PATH`. If you maintain a downstream agent (your OpenClaw
deployment, a workflow host, a private tool) and want to ship custom
subagents alongside it, drop a plugin directory on that env path.
This guide is for plugin authors. The CLI user doesn't need to read it.
## Minimum viable plugin
```
/path/to/my-plugin/
├── gbrain.plugin.json
└── subagents/
└── my-summarizer.md
```
`gbrain.plugin.json`:
```json
{
"name": "my-plugin",
"version": "1.0.0",
"plugin_version": "gbrain-plugin-v1"
}
```
`subagents/my-summarizer.md`:
```markdown
---
name: my-summarizer
model: claude-sonnet-4-6
allowed_tools:
- brain_search
- brain_get_page
---
You are a brain page summarizer. Given a slug, fetch the page and produce
a 3-sentence summary.
```
## Turning it on
```bash
export GBRAIN_PLUGIN_PATH="/path/to/my-plugin"
gbrain jobs work # worker startup prints the plugin load line
gbrain agent run "summarize meetings/2026-04-20" --subagent-def my-summarizer
```
Multiple plugins: colon-separated, just like `$PATH`.
```bash
export GBRAIN_PLUGIN_PATH="/path/to/plugin-a:/path/to/plugin-b"
```
## Rules (strict by design)
**Path policy.** Absolute paths only. Relative paths, `~`-prefixed paths,
and URL-style paths (`https://`, `file://`) are rejected with a warning.
You control where your plugin lives on disk; `gbrain` doesn't guess.
**Collision policy.** If two plugins ship a subagent with the same `name`,
the one listed FIRST in `GBRAIN_PLUGIN_PATH` wins. The other is dropped
with a warning naming both sources.
**Trust policy.** Plugins ship subagent definitions ONLY in v0.15:
- You **cannot** declare new tools.
- You **cannot** extend the brain tool allow-list.
- You **cannot** override any `agentSafe` or similar flag.
- Your `allowed_tools:` frontmatter field MUST subset the derived brain
tool registry. Names not in the registry are rejected at plugin load
time (worker startup), NOT at subagent dispatch time — so a typo in
your plugin gives you a loud startup error, not a silent "tool never
fires" at 3am.
v0.16+ may open up plugin-declared tools with a separate contract. Don't
expect it.
## `gbrain.plugin.json`
| field | type | required | notes |
|------------------|--------|----------|--------------------------------------------------------------------|
| `name` | string | yes | Human-readable plugin id. Shows up in warnings and collision logs. |
| `version` | string | yes | Your plugin's semver. Informational. |
| `plugin_version` | string | yes | Contract lock. Must equal `"gbrain-plugin-v1"` for v0.15. |
| `subagents` | string | no | Subdir name (default `subagents`). Escape-attempts are rejected. |
| `description` | string | no | Shown in future `gbrain plugin list`. |
## Subagent definition files
Plain markdown with YAML frontmatter. The body is the system prompt. The
frontmatter controls runtime behavior.
Recognized frontmatter fields:
| field | type | required | notes |
|-----------------|----------|----------|-----------------------------------------------------------------------------------------|
| `name` | string | no | Subagent identifier used as `--subagent-def`. Defaults to the file basename. |
| `model` | string | no | Anthropic model id. Defaults to the handler default (sonnet). |
| `max_turns` | number | no | Cap on assistant turns. Defaults to 20. |
| `allowed_tools` | string[] | no | Whitelist of tool names. Must subset the derived brain registry. Rejected on mismatch. |
Unknown frontmatter fields are preserved but ignored by the handler. v0.16
may consume more of them.
## Caveats that will bite you
1. **Plugin definitions can't change during a run.** The loader reads the
disk once at worker startup. Editing a subagent def doesn't re-take
effect until you restart the worker. This is deliberate — live
reloads would break crash-resumable replay.
2. **`~/.gbrain/audit/subagent-jobs-*.jsonl` is local only.** If your
worker runs on a different host than the `gbrain agent logs` caller,
the CLI won't see heartbeats from that worker. v0.16 will unify this;
for now assume worker + CLI share a filesystem.
3. **Tool calls always run with `ctx.remote = true`.** Even on local CLI
invocation. Tools that gate on `remote=true` (file_upload's strict
confinement, put_page's namespace check) will apply. Good default; a
subagent definition that wants local-filesystem reach beyond the brain
can't have it.
4. **`put_page` writes are namespace-scoped.** A subagent with id 42 can
only write under `wiki/agents/42/...`. This is enforced both in the
tool schema (the slug pattern shown to the model) AND server-side in
the `put_page` operation (fail-closed if `viaSubagent=true`). Don't
try to route around it; you'll get `permission_denied`.
## Example: a downstream-OpenClaw plugin
```
~/your-openclaw/
└── gbrain-plugin/
├── gbrain.plugin.json
└── subagents/
├── meeting-ingestion.md
├── signal-detector.md
└── daily-task-prep.md
```
`~/your-openclaw/gbrain-plugin/gbrain.plugin.json`:
```json
{
"name": "your-openclaw",
"version": "2026.4.20",
"plugin_version": "gbrain-plugin-v1",
"description": "Your OpenClaw's personal-brain subagents"
}
```
Environment:
```bash
export GBRAIN_PLUGIN_PATH="$HOME/your-openclaw/gbrain-plugin"
```
Then your OpenClaw calls `gbrain agent run --subagent-def meeting-ingestion
--fanout-by transcript ...` and its definitions load automatically.
+3 -3
View File
@@ -4,8 +4,8 @@ GBrain's Minion worker ships with seven built-in handlers: `sync`,
`embed`, `lint`, `import`, `extract`, `backlinks`, `autopilot-cycle`.
These cover every background operation the gbrain CLI itself performs.
Host platforms (Wintermute, other OpenClaw deployments, future hosts)
register their own handlers via a plugin bootstrap that imports
Host platforms (OpenClaw deployments, future hosts) register their own
handlers via a plugin bootstrap that imports
`gbrain/minions`. No `handlers.json`-style data file — handlers are
code, loaded by the worker, with the same trust model as any other
code in the host's repo.
@@ -58,7 +58,7 @@ async function main() {
main().catch(err => { console.error(err); process.exit(1); });
```
Ship this as a separate binary in the host repo (e.g. `wintermute-worker`)
Ship this as a separate binary in the host repo (e.g. `your-openclaw-worker`)
or as a side-effect module that the stock `gbrain jobs work` command
auto-loads on startup (configurable via a host-provided entry point).
+481 -2
View File
@@ -122,6 +122,8 @@ strict behavior when unset.
- `src/commands/eval.ts` — `gbrain eval` command: single-run table + A/B config comparison
- `src/core/embedding.ts` — OpenAI text-embedding-3-large, batch, retry, backoff
- `src/core/check-resolvable.ts` — Resolver validation: reachability, MECE overlap, DRY checks, structured fix objects. v0.14.1: `CROSS_CUTTING_PATTERNS.conventions` is an array (notability gate accepts both `conventions/quality.md` and `_brain-filing-rules.md`). New `extractDelegationTargets()` parses `> **Convention:**`, `> **Filing rule:**`, and inline backtick references. DRY suppression is proximity-based via `DRY_PROXIMITY_LINES = 40`.
- `src/core/repo-root.ts` — Shared `findRepoRoot(startDir?)` (v0.16.4): walks up from `startDir` (default `process.cwd()`) looking for `skills/RESOLVER.md`. Zero-dependency module imported by both `doctor.ts` and `check-resolvable.ts`. Parameterized `startDir` makes tests hermetic.
- `src/commands/check-resolvable.ts` — Standalone CLI wrapper (v0.16.4) over `checkResolvable()`. Exports `parseFlags`, `resolveSkillsDir`, `DEFERRED`, `runCheckResolvable`. Exit rule: **1 on any issue (warnings OR errors)**, stricter than doctor's `ok` flag — honors README:259. Stable JSON envelope `{ok, skillsDir, report, autoFix, deferred, error, message}` — same shape on success and error paths. `--fix` path runs `autoFixDryViolations` BEFORE `checkResolvable` (same ordering as doctor). `deferred[]` array surfaces pending Checks 5 (trigger routing eval) and 6 (brain filing) with issue URLs. `scripts/skillify-check.ts` subprocess-calls `gbrain check-resolvable --json` (cached per process) and fails loud on binary-missing — no silent false-pass.
- `src/core/dry-fix.ts` — `gbrain doctor --fix` engine. `autoFixDryViolations(fixes, {dryRun})` rewrites inlined rules to `> **Convention:** see [path](path).` callouts via three shape-aware expanders (bullet / blockquote / paragraph). Five guards: working-tree-dirty (`getWorkingTreeStatus()` returns 3-state `'clean' | 'dirty' | 'not_a_repo'`), no-git-backup, inside-code-fence, already-delegated (40-line proximity, consistent with detector), ambiguous-multi-match, block-is-callout. `execFileSync` array args (no shell — no injection surface). EOF newline preserved.
- `src/core/backoff.ts` — Adaptive load-aware throttling: CPU/memory checks, exponential backoff, active hours multiplier
- `src/core/fail-improve.ts` — Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generation
@@ -138,8 +140,19 @@ strict behavior when unset.
- `src/core/minions/protected-names.ts` — side-effect-free constant module exporting `PROTECTED_JOB_NAMES` + `isProtectedJobName()`. Kept pure so queue core can import without loading handler modules.
- `src/core/minions/handlers/shell.ts` — `shell` job handler. Spawns `/bin/sh -c cmd` (absolute path, PATH-override-safe) or `argv[0] argv[1..]` (no shell). Env allowlist: `PATH, HOME, USER, LANG, TZ, NODE_ENV` + caller `env:` overrides. UTF-8-safe stdout/stderr tail via `string_decoder.StringDecoder`. Abort (either `ctx.signal` or `ctx.shutdownSignal`) fires SIGTERM → 5s grace → SIGKILL on child. Requires `GBRAIN_ALLOW_SHELL_JOBS=1` on worker (gated by `registerBuiltinHandlers`).
- `src/core/minions/handlers/shell-audit.ts` — per-submission JSONL audit trail at `~/.gbrain/audit/shell-jobs-YYYY-Www.jsonl` (ISO-week rotation; override via `GBRAIN_AUDIT_DIR`). Best-effort: `mkdirSync(recursive)` + `appendFileSync`; failures logged to stderr, submission not blocked. Logs cmd (first 80 chars) or argv (JSON array). Never logs env values.
- `src/core/minions/handlers/subagent.ts` (v0.15) — LLM-loop handler. Two-phase tool persistence (pending → complete/failed), replay reconciliation for mid-dispatch crashes, dual-signal abort (`ctx.signal` + `ctx.shutdownSignal`), Anthropic prompt caching on system + tool defs. `makeSubagentHandler({engine, client?, ...})` factory; `MessagesClient` is an injectable interface the real SDK implements structurally. Throws `RateLeaseUnavailableError` (renewable) when rate-lease capacity is full.
- `src/core/minions/handlers/subagent-aggregator.ts` (v0.15) — `subagent_aggregator` handler. Claims AFTER all children resolve (queue changes guarantee every terminal child posts a `child_done` inbox message with outcome). Reads inbox via `ctx.readInbox()`, builds deterministic mixed-outcome markdown summary. No LLM call in v0.15.
- `src/core/minions/handlers/subagent-audit.ts` (v0.15) — JSONL audit + heartbeat writer at `~/.gbrain/audit/subagent-jobs-YYYY-Www.jsonl`. Events: `submission` (one line per submit) + `heartbeat` (per turn boundary: `llm_call_started | llm_call_completed | tool_called | tool_result | tool_failed`). Never logs prompts or tool inputs. `readSubagentAuditForJob(jobId, {sinceIso})` is the readback path for `gbrain agent logs`.
- `src/core/minions/rate-leases.ts` (v0.15) — lease-based concurrency cap for outbound providers (default key `anthropic:messages`, max via `GBRAIN_ANTHROPIC_MAX_INFLIGHT`). Owner-tagged rows with `expires_at` auto-prune on acquire; `pg_advisory_xact_lock` guards check-then-insert; CASCADE on owning job deletion. `renewLeaseWithBackoff` retries 3x (250/500/1000ms).
- `src/core/minions/wait-for-completion.ts` (v0.15) — poll-until-terminal helper for CLI callers. `TimeoutError` does NOT cancel the job; `AbortSignal` exits without throwing. Default `pollMs`: 1000 on Postgres, 250 on PGLite inline.
- `src/core/minions/transcript.ts` (v0.15) — renders `subagent_messages` + `subagent_tool_executions` to markdown. Tool rows splice under their owning assistant `tool_use` by `tool_use_id`. UTF-8-safe truncation; unknown block types fall through to fenced JSON.
- `src/core/minions/plugin-loader.ts` (v0.15) — `GBRAIN_PLUGIN_PATH` discovery. Absolute paths only, left-wins collision, `gbrain.plugin.json` with `plugin_version: "gbrain-plugin-v1"`, plugins ship DEFS only (no new tools), `allowed_tools:` validated at load time against the derived registry.
- `src/core/minions/tools/brain-allowlist.ts` (v0.15) — derives subagent tool registry from `src/core/operations.ts`. 11-name allow-list: `query`, `search`, `get_page`, `list_pages`, `file_list`, `file_url`, `get_backlinks`, `traverse_graph`, `resolve_slugs`, `get_ingest_log`, `put_page`. `put_page` schema is namespace-wrapped per subagent (`^wiki/agents/<subagentId>/.+`); the `put_page` op's server-side check is the authoritative gate via `ctx.viaSubagent` fail-closed.
- `src/mcp/tool-defs.ts` (v0.15) — extracted `buildToolDefs(ops)` helper. MCP server + subagent tool registry both call it; byte-for-byte equivalence pinned by `test/mcp-tool-defs.test.ts`.
- `src/core/minions/attachments.ts` — Attachment validation (path traversal, null byte, oversize, base64, duplicate detection)
- `src/commands/jobs.ts` — `gbrain jobs` CLI subcommands + `gbrain jobs work` daemon. v0.13.1 surfaces the full `MinionJobInput` retry/backoff/timeout/idempotency surface as first-class CLI flags on `jobs submit`: `--max-stalled`, `--backoff-type fixed|exponential`, `--backoff-delay`, `--backoff-jitter`, `--timeout-ms`, `--idempotency-key`. `jobs smoke --sigkill-rescue` is the opt-in regression guard for #219.
- `src/commands/agent.ts` (v0.16) — `gbrain agent run <prompt> [flags]` CLI. Submits `subagent` (or N children + 1 aggregator) under `{allowProtectedSubmit: true}`. Single-entry `--fanout-manifest` short-circuits. Children get `on_child_fail: 'continue'` + `max_stalled: 3`. `--follow` is the default on TTY; streams logs + polls `waitForCompletion` in parallel. Ctrl-C detaches, does not cancel.
- `src/commands/agent-logs.ts` (v0.16) — `gbrain agent logs <job> [--follow] [--since]`. Merges JSONL heartbeat audit + `subagent_messages` into a chronological timeline. `parseSince` accepts ISO-8601 or relative (`5m`, `1h`, `2d`). Transcript tail renders only for terminal jobs.
- `src/commands/jobs.ts` — `gbrain jobs` CLI subcommands + `gbrain jobs work` daemon. v0.13.1 surfaces the full `MinionJobInput` retry/backoff/timeout/idempotency surface as first-class CLI flags on `jobs submit`: `--max-stalled`, `--backoff-type fixed|exponential`, `--backoff-delay`, `--backoff-jitter`, `--timeout-ms`, `--idempotency-key`. `jobs smoke --sigkill-rescue` is the opt-in regression guard for #219. v0.16 wires `registerBuiltinHandlers` to always register `subagent` + `subagent_aggregator` (no env flag — `ANTHROPIC_API_KEY` is the natural cost gate, trust is via `PROTECTED_JOB_NAMES`) and loads `GBRAIN_PLUGIN_PATH` plugins at worker startup with a loud startup-line per plugin. `shell` handler still gated by `GBRAIN_ALLOW_SHELL_JOBS=1` (RCE surface, separate concern).
- `src/commands/features.ts` — `gbrain features --json --auto-fix`: usage scan + feature adoption salesman
- `src/commands/autopilot.ts` — `gbrain autopilot --install`: self-maintaining brain daemon (sync+extract+embed)
- `src/mcp/server.ts` — MCP stdio server (generated from operations)
@@ -152,6 +165,8 @@ strict behavior when unset.
- `src/core/migrate.ts` — schema-migration runner. Owns the `MIGRATIONS` array (source of truth for schema DDL). v0.14.2 extended the `Migration` interface with `sqlFor?: { postgres?, pglite? }` (engine-specific SQL overrides `sql`) and `transaction?: boolean` (set to false for `CREATE INDEX CONCURRENTLY`, which Postgres refuses inside a transaction; ignored on PGLite since it has no concurrent writers). Migration v14 (fix wave) uses a handler branching on `engine.kind` to run CONCURRENTLY on Postgres (with a pre-drop of any invalid remnant via `pg_index.indisvalid`) and plain `CREATE INDEX` on PGLite. v15 bumps `minion_jobs.max_stalled` default 1→5 and backfills existing non-terminal rows.
- `src/core/progress.ts` — Shared bulk-action progress reporter. Writes to stderr. Modes: `auto` (TTY: `\r`-rewriting; non-TTY: plain lines), `human`, `json` (JSONL), `quiet`. Rate-gated by `minIntervalMs` and `minItems`. `startHeartbeat(reporter, note)` helper for single long queries. `child()` composes phase paths. Singleton SIGINT/SIGTERM coordinator emits `abort` events for every live phase. EPIPE defense on both sync throws and stream `'error'` events. Zero dependencies. Introduced in v0.15.2.
- `src/core/cli-options.ts` — Global CLI flag parser. `parseGlobalFlags(argv)` returns `{cliOpts, rest}` with `--quiet` / `--progress-json` / `--progress-interval=<ms>` stripped. `getCliOptions()` / `setCliOptions()` expose a module-level singleton so commands reach the resolved flags without parameter threading. `cliOptsToProgressOptions()` maps to reporter options. `childGlobalFlags()` returns the flag suffix to append to `execSync('gbrain ...')` calls in migration orchestrators. `OperationContext.cliOpts` extends shared-op dispatch for MCP callers.
- `src/core/cycle.ts` — v0.17 brain maintenance cycle primitive. `runCycle(engine: BrainEngine | null, opts: CycleOpts): Promise<CycleReport>` composes 6 phases in semantically-driven order (lint → backlinks → sync → extract → embed → orphans). Three callers: `gbrain dream` CLI, `gbrain autopilot` daemon's inline path, and the Minions `autopilot-cycle` handler (`src/commands/jobs.ts`). One source of truth for what the brain does overnight. Coordination via `gbrain_cycle_locks` DB table (TTL-based; works through PgBouncer transaction pooling, unlike session-scoped `pg_try_advisory_lock`) + `~/.gbrain/cycle.lock` file lock with PID-liveness for PGLite / engine=null mode. `CycleReport.schema_version: "1"` is the stable agent-consumable shape. `PhaseResult.error: { class, code, message, hint?, docs_url? }` is Stripe-API-tier structured failure info. `yieldBetweenPhases` hook awaited between every phase — Minions handler uses this to renew its job lock and prevent v0.14 stall-death regression. Engine nullable: filesystem phases (lint, backlinks) run without DB; DB phases skip with `status: "skipped", reason: "no_database"`. Lock-skip: read-only phase selections (`--phase orphans`) bypass the cycle lock.
- `src/commands/dream.ts` — v0.17 `gbrain dream` CLI. ~80-line thin alias over `runCycle`. brainDir resolution requires explicit `--dir` OR `sync.repo_path` config (no more walk-up-cwd-for-.git footgun). Flags: `--dry-run`, `--json`, `--phase <name>`, `--pull`, `--dir <path>`. Exit code 1 on status=failed (partial/warn not fatal — don't page on warnings).
- `scripts/check-progress-to-stdout.sh` — CI guard against regressing to `\r`-on-stdout progress. Wired into `bun run test` via `scripts/check-progress-to-stdout.sh && bun test` in package.json.
- `docs/progress-events.md` — Canonical JSON event schema reference. Stable from v0.15.2, additive only.
- `src/core/markdown.ts` — Frontmatter parsing + body splitter. `splitBody` requires an explicit timeline sentinel (`<!-- timeline -->`, `--- timeline ---`, or `---` immediately before `## Timeline`/`## History`). Plain `---` in body text is a markdown horizontal rule, not a separator. `inferType` auto-types `/wiki/analysis/` → analysis, `/wiki/guides/` → guide, `/wiki/hardware/` → hardware, `/wiki/architecture/` → architecture, `/writing/` → writing (plus the existing people/companies/deals/etc heuristics).
@@ -636,6 +651,21 @@ search engines, surfaced in cross-references, and distributed with every release
- Example meeting → `meetings/2026-04-03` (generic date is fine)
- Example user → `you` or `the user`, never a proper name
**Specific rule: never say `Wintermute` in any CHANGELOG, README, doc, PR, or
commit message.** When the temptation is to illustrate with the real fork name:
- Reader-facing copy → `your OpenClaw` (covers Wintermute, Hermes, AlphaClaw,
and any other downstream OpenClaw deployment in one term the reader already
recognizes).
- First-person / origin-story copy → `Garry's OpenClaw` (honest that this is
the production deployment driving the feature, without exposing the private
agent's name).
`Wintermute` may appear in private artifacts (scratch plans under
`~/.gstack/projects/…`, memory files, conversation transcripts, CEO-review
plans) — those aren't distributed. Anything checked into this repo or shipped
in a release must use the OpenClaw phrasing above. Sweeping a stale reference
is a small clean-up PR, not a debate.
**When in doubt, ask yourself:** "Would this query reveal private information
about the user's contacts, investments, or portfolio if it were read by a
stranger?" If yes, replace with generic placeholders.
@@ -1258,6 +1288,25 @@ If anything's off, `actions[]` tells you the exact command to run. For deeper tr
Moving gateway crons to Minions (deterministic scripts, zero LLM tokens per fire): [`docs/guides/minions-shell-jobs.md`](docs/guides/minions-shell-jobs.md).
## Durable agents: `gbrain agent` (v0.15)
Your subagent runs survive crashes now. OpenClaw died mid-run? The worker re-claims on restart and replays from the last committed turn. Fan-out across 50 shards, one shard crashes — the aggregator still claims after every child reaches a terminal state and writes a mixed-outcome summary. Tool calls persist as a two-phase ledger (`pending` → `complete | failed`) so replay is safe by construction, not by hope.
```bash
# Submit a single-subagent run
gbrain agent run "summarize my last 10 journal pages"
# Fan out N prompts across N subagent children + 1 aggregator
gbrain agent run "analyze every page" \
--fanout-manifest manifests/pages.json \
--subagent-def analyzer
# Tail a running job (heartbeat per turn + full transcript on completion)
gbrain agent logs 1247 --follow --since 5m
```
Durability is the point: every Anthropic turn commits to `subagent_messages`, every tool call to `subagent_tool_executions`. Worker kills, OpenClaw crashes, timeouts — all resumable. Host repos (your OpenClaw, etc.) ship their own subagent definitions via `GBRAIN_PLUGIN_PATH` + a `gbrain.plugin.json` manifest: see [`docs/guides/plugin-authors.md`](docs/guides/plugin-authors.md). Requires `ANTHROPIC_API_KEY` on the worker.
## Skillify: your skills tree stops being a black box
Hermes and similar agent frameworks auto-create skills as a background behavior. Fine until you don't know what the agent shipped. Checklists decay. Tests drift. Resolver entries get stale. Six months later you've got an opaque pile of "skills" that nobody has read, nobody has tested, and nobody is sure still work.
@@ -3236,6 +3285,336 @@ echo "Dream cycle complete at $(date)"
---
## docs/guides/minions-deployment.md
Source: https://raw.githubusercontent.com/garrytan/gbrain/master/docs/guides/minions-deployment.md
# Minions Worker Deployment Guide
Deploy `gbrain jobs work` so it stays running across crashes, reboots, and
Postgres connection blips. Written for agents to execute line-by-line.
## The problem
The persistent worker can die silently from:
- Database connection drops (Supabase/Postgres maintenance or network blips).
- Lock-renewal failures → the stall detector eventually dead-letters jobs.
- Bun process crashes with no automatic restart.
- Internal event-loop death (PID alive, worker loop stopped).
When the worker dies, submitted jobs sit in `waiting` forever. Nothing in
gbrain core auto-restarts the worker — that's what this guide wires up.
## Variables used in this guide
Substitute these once before copy-pasting any snippet.
| Variable | Meaning | Typical value |
|---|---|---|
| `$GBRAIN_BIN` | Absolute path to the `gbrain` binary | `$(command -v gbrain)` — often `/usr/local/bin/gbrain` or `~/.bun/bin/gbrain` |
| `$GBRAIN_WORKER_USER` | OS user that owns the worker process | the same user that ran `gbrain init`; never `root` |
| `$GBRAIN_WORKER_PID_FILE` | Worker PID + restart-epoch file | `/tmp/gbrain-worker.pid` (or `/var/run/gbrain/worker.pid` for systemd) |
| `$GBRAIN_WORKER_LOG_FILE` | Worker log sink (stdout + stderr merged) | `/tmp/gbrain-worker.log` (or `/var/log/gbrain/worker.log`) |
| `$GBRAIN_WORKSPACE` | `cwd` for shell jobs submitted by this deployment | absolute path, e.g. `/srv/my-brain` |
| `$GBRAIN_ENV_FILE` | Secrets file sourced by crontab / systemd | `/etc/gbrain.env` (mode 600) |
## Preconditions
Run these before Step 1 of any option. Fail fast if something is wrong.
```bash
# 1. gbrain is on PATH and resolves to an absolute location.
command -v gbrain || { echo "gbrain not on PATH. Install, then retry."; exit 1; }
# 2. DATABASE_URL points at reachable Postgres (or PGLite path exists).
gbrain doctor --fast --json | jq '.checks[] | select(.name=="db_connectivity")'
# 3. Schema is up to date. If version=0 or status=="fail", fix it first:
# gbrain apply-migrations --yes
gbrain doctor --fast --json | jq '.checks[] | select(.name=="schema_version")'
# 4. You have write access to at least one crontab mechanism.
crontab -l >/dev/null 2>&1 && echo "user crontab OK"
[ -w /etc/crontab ] && echo "/etc/crontab OK"
# 5. If you plan to submit `shell` jobs, the WORKER process needs
# GBRAIN_ALLOW_SHELL_JOBS=1 (submitters do not). The handler is gated
# in registerBuiltinHandlers(); without the flag the worker startup
# line reads "shell handler disabled (...)".
```
## Which option?
- Your workload runs LLM subagents (`gbrain agent run`) or jobs that take
> 30 s → **Option 1** (watchdog cron + persistent worker).
- Your workload is short deterministic scripts on a fixed schedule (every
3 h, daily, weekly) → **Option 2** (inline `--follow`).
- You don't have shell access to a long-running box (Fly/Render/Railway,
or any systemd host) → **Option 3** (service manager — replaces cron).
## Option 1: watchdog cron + persistent worker
A 5-minute cron checks whether the worker process is alive **and** whether
it has logged an internal shutdown since its last start. Restarts if either
condition fails.
### 1a. Install the env file (secrets stay out of crontab)
Never paste `DATABASE_URL` or API keys into crontab. `/etc/crontab` is
mode 644 (world-readable); user crontabs under `/var/spool/cron/` are
readable by `root`. Use the shipped env-file template:
```bash
sudo install -m 600 -o $GBRAIN_WORKER_USER -g $GBRAIN_WORKER_USER \
docs/guides/minions-deployment-snippets/gbrain.env.example /etc/gbrain.env
sudoedit /etc/gbrain.env
```
Fill in the connection string and `GBRAIN_ALLOW_SHELL_JOBS=1` (if
applicable). See
[`gbrain.env.example`](./minions-deployment-snippets/gbrain.env.example)
for the full list.
### 1b. Install the watchdog script
The [`minion-watchdog.sh`](./minions-deployment-snippets/minion-watchdog.sh)
ships in-repo and writes a two-line PID file (PID on line 1, restart epoch
on line 2). The restart-epoch marker is how the watchdog distinguishes
stale shutdown lines in the log from current ones — without it, every tick
after the first restart would match an old `worker shutting down` line and
loop forever.
Requires GNU coreutils (Linux default). On macOS/BSD install via
`brew install coreutils` and alias `date` to `gdate` in the cron env if you
want to test the watchdog locally; production Linux boxes work as-is.
```bash
sudo install -m 755 -o $GBRAIN_WORKER_USER -g $GBRAIN_WORKER_USER \
docs/guides/minions-deployment-snippets/minion-watchdog.sh \
/usr/local/bin/minion-watchdog.sh
```
### 1c. Wire into cron
Pick the form that matches the crontab you're editing.
**If you ran `crontab -e`** (user crontab — 5-field, no user column):
```
SHELL=/bin/bash
PATH=/usr/local/bin:/usr/bin:/bin
BASH_ENV=/etc/gbrain.env
*/5 * * * * /usr/local/bin/minion-watchdog.sh
```
**If you edited `/etc/crontab` directly** (system crontab — 6-field, with
user column):
```
SHELL=/bin/bash
PATH=/usr/local/bin:/usr/bin:/bin
BASH_ENV=/etc/gbrain.env
*/5 * * * * gbrain /usr/local/bin/minion-watchdog.sh
```
In both forms, `BASH_ENV=/etc/gbrain.env` tells non-interactive bash to
source the env file before running the watchdog — that's how the
connection string and `GBRAIN_ALLOW_SHELL_JOBS` reach the worker without
landing in the world-readable crontab itself.
### 1d. Log rotation
The watchdog appends to the worker log across restarts. If you expect the
file to grow unbounded, rotate it externally with `logrotate`:
```
# /etc/logrotate.d/gbrain-worker
/tmp/gbrain-worker.log {
daily
rotate 7
missingok
notifempty
copytruncate
}
```
`copytruncate` is important — the watchdog's restart-epoch check survives
it (the epoch is compared against in-log timestamps, not file inode).
## Option 2: inline `--follow` (no persistent worker)
Each cron run brings its own temporary worker. `--follow` starts one on
the queue and blocks until the just-submitted job reaches a terminal state
(`completed` / `failed` / `dead` / `cancelled`). 2-3 s startup overhead
per job; negligible vs job duration for scheduled work.
Example: nightly brain enrichment as a shell job.
```bash
GBRAIN_ALLOW_SHELL_JOBS=1 gbrain jobs submit shell \
--queue nightly-enrich \
--params "{\"cmd\":\"$GBRAIN_BIN embed --stale\",\"cwd\":\"$GBRAIN_WORKSPACE\"}" \
--follow \
--timeout-ms 600000
```
Replace `gbrain embed --stale` with whichever gbrain subcommand you're
scheduling (`sync`, `extract`, `orphans`, `doctor`, `check-backlinks`,
`lint`, `autopilot`). If you're shelling out to a non-gbrain binary,
keep its absolute path in the `cmd`.
**Shared-queue gotcha.** If other jobs are already waiting on the same
queue with higher priority or earlier `created_at`, the temporary worker
processes those first before reaching yours. `--follow` still exits only
when YOUR job finishes. For strict single-job semantics on shared queues,
use a dedicated queue name like `nightly-enrich` above.
## Option 3: service manager (systemd / Fly / Render / Railway)
Replaces the watchdog entirely. No cron, no PID file, no restart-loop.
The service manager owns liveness.
### systemd (Linux hosts with shell access)
```bash
# Create the worker user if it doesn't exist.
sudo useradd --system --home "$GBRAIN_WORKSPACE" --shell /usr/sbin/nologin gbrain \
2>/dev/null || true
sudo mkdir -p "$GBRAIN_WORKSPACE" && sudo chown gbrain:gbrain "$GBRAIN_WORKSPACE"
# Install the unit file, substituting /srv/gbrain → your workspace path.
sudo install -m 644 docs/guides/minions-deployment-snippets/systemd.service \
/etc/systemd/system/gbrain-worker.service
sudo sed -i "s|/srv/gbrain|$GBRAIN_WORKSPACE|g" \
/etc/systemd/system/gbrain-worker.service
# See 1a above for /etc/gbrain.env install.
sudo systemctl daemon-reload
sudo systemctl enable --now gbrain-worker
sudo systemctl status gbrain-worker
journalctl -u gbrain-worker -n 50
```
`Restart=always` + `RestartSec=10s` give you crash-loop recovery. The unit
runs as an unprivileged `gbrain` user with `PrivateTmp`, `ProtectSystem=strict`,
and `ReadWritePaths=$GBRAIN_WORKSPACE`. `LimitNOFILE=65535` in the shipped
unit covers Bun + Postgres pool + concurrent LLM subagent calls without
hitting the default 1024 cap.
### Fly.io
Merge the `[processes]` block from
[`fly.toml.partial`](./minions-deployment-snippets/fly.toml.partial) into
your existing `fly.toml`. Set secrets with `fly secrets set` —
Fly auto-restarts the process on crash.
### Render / Railway / Heroku
Drop [`Procfile`](./minions-deployment-snippets/Procfile) at the repo root.
Set the connection string and `GBRAIN_ALLOW_SHELL_JOBS=1` via the
platform's env UI or CLI.
## Upgrading an existing deployment
If you deployed on v0.13.x or earlier, walk this checklist:
1. **Stop the worker before upgrading.**
`kill $(head -n1 /tmp/gbrain-worker.pid)` and wait for the process to
exit. Skipping this risks an in-flight job landing partial schema.
2. **Run `gbrain upgrade`**. Then `gbrain apply-migrations --yes` if
`gbrain doctor` reports any migration as `partial` or `pending`.
3. **If you run shell jobs:** from v0.14 onward, the worker requires
`GBRAIN_ALLOW_SHELL_JOBS=1` to register the `shell` handler. Add it to
`/etc/gbrain.env`. Submitters don't need the flag; only the worker does.
4. **If you tuned your watchdog for `max_stalled=1`:** v0.14.3 migration
v15 raised the schema default to 5 and backfilled existing non-terminal
rows. A watchdog tuned around 1-strike dead-lettering will now
over-restart because it takes 5 misses to dead-letter. Switch to the
shipped watchdog (which keys on log markers, not job state).
5. **If your v0.16.1 watchdog is still running:** it has a restart-loop
bug (old shutdown lines in the unrotated log re-match every 5 min
forever). Install the current `minion-watchdog.sh` from this guide's
snippets — it writes a restart epoch into the PID file and only
considers log lines newer than that epoch.
6. **Verify.** `gbrain doctor` should report zero `pending` or `partial`
migrations. `gbrain jobs stats` should show no unexplained growth in
`dead` between pre- and post-upgrade.
## Known issues
### Supabase connection drops
The worker uses a single Postgres connection. If Supabase drops it
(maintenance, connection limits, network blip), lock renewal fails
silently. The stall detector then dead-letters the job after
`max_stalled` misses.
**Current defaults that make this worse:**
- `lockDuration: 30000` (30 s) — too short for long jobs during connection blips.
- `max_stalled: 5` (schema column default on master — see `src/schema.sql`
and `src/core/pglite-schema.ts`). Five missed heartbeats before dead-letter.
- `stalledInterval: 30000` (30 s) — checks too aggressively.
**Tune per-job today.** `gbrain jobs submit` accepts `--max-stalled N`,
`--backoff-type fixed|exponential`, `--backoff-delay <ms>`,
`--backoff-jitter 0..1`, and `--timeout-ms N` as first-class flags
(since v0.13.1). These write onto the job row at submit time — which is
what `handleStalled()` reads — so per-job tuning is the real knob today.
Worker-level `--lock-duration` / `--stall-interval` are on the roadmap;
until they land, rely on per-job `--max-stalled` plus the watchdog (or
systemd) for worker health.
### DO NOT pass `maxStalledCount` to `MinionWorker`
It's a no-op. The stall detector reads the row's `max_stalled` column
(set at submit time), not the worker opt in `src/core/minions/worker.ts:74`.
Use `gbrain jobs submit --max-stalled N` per-job instead.
### Zombie shell children
When the Bun worker crashes hard, child processes from shell jobs can
become zombies. The watchdog's 10 s `SIGTERM → SIGKILL` window covers the
shell handler's 5 s child-kill grace (`KILL_GRACE_MS`). For long-running
shell jobs, bump the watchdog's `sleep 10` to `sleep 30` so the worker
has time to flush in-flight jobs before the kill.
## Smoke test
```bash
# Worker alive?
kill -0 $(head -n1 /tmp/gbrain-worker.pid) 2>/dev/null && echo ALIVE || echo DEAD
# Aggregate queue health.
gbrain jobs stats
# Jobs currently stalled (still `active` with expired lock_until, pre-requeue).
gbrain jobs list --status active --limit 10
# Dead-lettered jobs.
gbrain jobs list --status dead --limit 10
# Shell handler registered? (stderr banner merged into log via 2>&1.)
grep "shell handler enabled" /tmp/gbrain-worker.log
```
## Uninstall
- **Option 1 (watchdog cron):** `crontab -e`, delete the watchdog line.
`kill $(head -n1 /tmp/gbrain-worker.pid) && rm /tmp/gbrain-worker.pid`.
Optionally `sudo rm /etc/gbrain.env /usr/local/bin/minion-watchdog.sh`.
- **Option 2 (inline `--follow`):** remove the cron entry. Nothing else to
clean up — temporary workers exit with their jobs.
- **Option 3 (systemd):** `sudo systemctl disable --now gbrain-worker`,
then `sudo rm /etc/systemd/system/gbrain-worker.service /etc/gbrain.env`,
then `sudo systemctl daemon-reload`.
- **Option 3 (Fly/Render/Railway):** delete the `worker` process from
`fly.toml` / `Procfile` and redeploy. Secrets set via `fly secrets`
persist until `fly secrets unset`.
---
## docs/guides/quiet-hours.md
Source: https://raw.githubusercontent.com/garrytan/gbrain/master/docs/guides/quiet-hours.md
@@ -3918,7 +4297,7 @@ F. Install gbrain autopilot --install (env-aware)
G. Record append completed.jsonl status:"complete"
```
If Phase E emits TODOs for host-specific handlers (e.g. Wintermute's
If Phase E emits TODOs for host-specific handlers (e.g. your OpenClaw's
~29 non-gbrain crons), the migration finishes with `status: "partial"`.
Your host agent walks the TODOs using `skills/migrations/v0.11.0.md` +
`docs/guides/plugin-handlers.md`, ships handler registrations in the
@@ -4444,6 +4823,106 @@ upcoming `gbrain crontab-to-minions <file>` helper is P1 in TODOS.
---
## v0.16.0: durable agent runtime
v0.15 ships `gbrain agent run` / `gbrain agent logs`, a new `subagent` handler
type in Minions, and a plugin contract for host-repo subagent defs. None of the
existing skills need surgery. The question for downstream agents is *how* to
adopt the new runtime, not how to patch around a breaking change.
### 1. Run a worker with an Anthropic key
The subagent handlers (`subagent` and `subagent_aggregator`) are always
registered on the worker. No separate opt-in flag — `ANTHROPIC_API_KEY` is
the natural cost gate (no key, the SDK call fails on the first turn), and
who-can-submit is already protected (`PROTECTED_JOB_NAMES` + trusted-submit:
MCP callers get `permission_denied`; only `gbrain agent run` can insert
these rows).
```bash
ANTHROPIC_API_KEY=sk-ant-... gbrain jobs work
```
Worker startup prints:
```
[minion worker] subagent handlers enabled
```
### 2. Ship your subagents as a plugin (OpenClaw + similar)
Move your custom subagent definitions out of your gbrain fork and into your own
repo as a plugin. Concretely:
```
~/<your-agent>/gbrain-plugin/
├── gbrain.plugin.json
└── subagents/
├── meeting-ingestion.md
├── signal-detector.md
└── daily-task-prep.md
```
`gbrain.plugin.json`:
```json
{
"name": "your-openclaw",
"version": "2026.4.20",
"plugin_version": "gbrain-plugin-v1"
}
```
Each `subagents/*.md` is a plain-text agent definition — YAML frontmatter +
body-as-system-prompt. Recognized frontmatter fields: `name`, `model`,
`max_turns`, `allowed_tools` (must subset the derived brain-tool registry).
Turn it on:
```bash
export GBRAIN_PLUGIN_PATH="$HOME/<your-agent>/gbrain-plugin"
```
Worker startup prints `[plugin-loader] loaded '<name>' v<ver> (N subagents)`
per plugin; any rejection (bad manifest, unknown tool in `allowed_tools`,
version mismatch) shows up as a loud warning at startup, not a silent dispatch-
time failure. See `docs/guides/plugin-authors.md` for the full contract.
### 3. Replace ephemeral subagent runs with durable ones
If your agent currently spawns ephemeral subagents (OpenClaw `Agent()`, ad-hoc
Anthropic API calls, etc.) for work that should survive crashes, sleeps, or
worker restarts, migrate those to `gbrain agent run`. The durability is free:
```bash
gbrain agent run "analyze my last 50 journal pages for recurring themes" \
--subagent-def analyzer --fanout-manifest manifests/journal-pages.json
```
Every turn persists to `subagent_messages`, every tool call is a two-phase
ledger, and `gbrain agent logs <job>` shows where it died + what the last
successful call returned. No more "re-run from scratch because the session
context evaporated."
### 4. `put_page` from subagents writes under an agent namespace
If you adopted the v0.15 subagent runtime, note that `put_page` calls
originating from a subagent's tool dispatch MUST target
`wiki/agents/<subagent_id>/...`. The schema shown to the model enforces this
on first try; a server-side fail-closed check rejects anything else. This
does NOT affect your skill files, CLI put_page calls, or MCP put_page —
only tool-dispatched writes from inside an LLM loop.
Aggregation output (the final "here's what all N children found" brain page)
goes via a separate trusted CLI path, not through a subagent tool call, so
it can write anywhere you want.
Iron rule: **never grant an agent write access beyond its namespace**. The
server-side check exists because dispatcher bugs happen; treat it as defense
in depth, not the primary boundary.
---
## Future versions
When gbrain ships a new version, this doc will be updated with the diffs for that
+1
View File
@@ -18,6 +18,7 @@ Repo: https://github.com/garrytan/gbrain
- [docs/GBRAIN_RECOMMENDED_SCHEMA.md](https://raw.githubusercontent.com/garrytan/gbrain/master/docs/GBRAIN_RECOMMENDED_SCHEMA.md): MECE directory structure (people/, companies/, concepts/).
- [docs/guides/live-sync.md](https://raw.githubusercontent.com/garrytan/gbrain/master/docs/guides/live-sync.md): Incremental markdown sync setup.
- [docs/guides/cron-schedule.md](https://raw.githubusercontent.com/garrytan/gbrain/master/docs/guides/cron-schedule.md): Recurring job scheduling.
- [docs/guides/minions-deployment.md](https://raw.githubusercontent.com/garrytan/gbrain/master/docs/guides/minions-deployment.md): Deploying the gbrain jobs worker: crontab + watchdog, inline --follow, systemd/Procfile/fly.toml, upgrade checklist.
- [docs/guides/quiet-hours.md](https://raw.githubusercontent.com/garrytan/gbrain/master/docs/guides/quiet-hours.md): Notification hold + timezone-aware delivery.
- [docs/mcp/DEPLOY.md](https://raw.githubusercontent.com/garrytan/gbrain/master/docs/mcp/DEPLOY.md): MCP server deployment.
+8 -4
View File
@@ -1,6 +1,6 @@
{
"name": "gbrain",
"version": "0.15.4",
"version": "0.16.4",
"description": "Postgres-native personal knowledge brain with hybrid RAG search",
"type": "module",
"main": "src/core/index.ts",
@@ -21,8 +21,9 @@
"build:all": "bun build --compile --target=bun-darwin-arm64 --outfile bin/gbrain-darwin-arm64 src/cli.ts && bun build --compile --target=bun-linux-x64 --outfile bin/gbrain-linux-x64 src/cli.ts",
"build:schema": "bash scripts/build-schema.sh",
"build:llms": "bun run scripts/build-llms.ts",
"test": "scripts/check-jsonb-pattern.sh && scripts/check-progress-to-stdout.sh && bun test",
"test": "scripts/check-jsonb-pattern.sh && scripts/check-progress-to-stdout.sh && bun run typecheck && bun test",
"test:e2e": "bash scripts/run-e2e.sh",
"typecheck": "tsc --noEmit",
"check:jsonb": "scripts/check-jsonb-pattern.sh",
"check:progress": "scripts/check-progress-to-stdout.sh",
"postinstall": "command -v gbrain >/dev/null 2>&1 && gbrain apply-migrations --yes --non-interactive || echo '[gbrain] postinstall skipped. If installed via bun install -g github:...: run `gbrain doctor` and `gbrain apply-migrations --yes` manually. See https://github.com/garrytan/gbrain/issues/218' 1>&2",
@@ -43,10 +44,13 @@
"marked": "^18.0.0",
"openai": "^4.0.0",
"pgvector": "^0.2.0",
"postgres": "^3.4.0"
"postgres": "^3.4.0",
"tree-sitter-wasms": "0.1.13",
"web-tree-sitter": "0.22.6"
},
"devDependencies": {
"@types/bun": "latest"
"@types/bun": "latest",
"typescript": "^5.6.0"
},
"trustedDependencies": [
"@electric-sql/pglite"
+6
View File
@@ -92,6 +92,12 @@ export const SECTIONS: DocSection[] = [
description: "Recurring job scheduling.",
path: "docs/guides/cron-schedule.md",
},
{
title: "docs/guides/minions-deployment.md",
description:
"Deploying the gbrain jobs worker: crontab + watchdog, inline --follow, systemd/Procfile/fly.toml, upgrade checklist.",
path: "docs/guides/minions-deployment.md",
},
{
title: "docs/guides/quiet-hours.md",
description: "Notification hold + timezone-aware delivery.",
+48 -7
View File
@@ -15,7 +15,7 @@
* Returns JSON when --json is passed: { path, score, total, items,
* recommendation }. Exit code is 0 when score == total, 1 otherwise.
*
* Ported from ~/git/wintermute/workspace/scripts/skillify-check.mjs
* Ported from ~/git/your-openclaw/workspace/scripts/skillify-check.mjs
* (genericized: paths computed from $PROJECT_ROOT + runtime test-dir
* detection; replaces the manual `grep AGENTS.md` check with a reference
* to `gbrain check-resolvable` which validates the resolver better).
@@ -23,6 +23,7 @@
import { existsSync, readFileSync, readdirSync, statSync } from 'fs';
import { join, basename, dirname, resolve } from 'path';
import { spawnSync } from 'child_process';
function projectRoot(): string {
// Walk up from cwd until we find a package.json — that's the repo root.
@@ -64,6 +65,45 @@ function checkOptional(name: string, passed: boolean, detail?: string): CheckIte
return { name, passed, required: false, detail };
}
/**
* Invoke `gbrain check-resolvable --json` once and cache the result for the
* process lifetime. Binary-missing surfaces a loud error instead of silently
* passing this is the critical guard the failure-mode audit flagged.
*/
interface ResolverResult {
ok: boolean;
detail: string;
}
let _resolverCache: ResolverResult | null = null;
function runCheckResolvableCached(): ResolverResult {
if (_resolverCache) return _resolverCache;
try {
const res = spawnSync('gbrain', ['check-resolvable', '--json'], {
encoding: 'utf-8',
maxBuffer: 10 * 1024 * 1024,
});
if (res.error || res.status === null) {
const reason = res.error?.message ?? 'spawn returned null status';
console.error(`[skillify] gbrain check-resolvable not runnable: ${reason}`);
_resolverCache = { ok: false, detail: `check-resolvable unavailable: ${reason}` };
return _resolverCache;
}
const payload = JSON.parse(res.stdout);
if (payload.ok === true) {
_resolverCache = { ok: true, detail: 'all skill-tree checks pass' };
} else {
const count = payload.report?.issues?.length ?? 0;
const err = payload.error ? ` (${payload.error})` : '';
_resolverCache = { ok: false, detail: `${count} issue(s)${err} — run: gbrain check-resolvable` };
}
return _resolverCache;
} catch (err) {
console.error(`[skillify] check-resolvable parse failed: ${err}`);
_resolverCache = { ok: false, detail: `check-resolvable parse error: ${err}` };
return _resolverCache;
}
}
/**
* Guess the skill-directory name from a script path.
* scripts/frameio-scraper.ts frameio-scraper
@@ -184,12 +224,13 @@ function runCheck(target: string): {
}
items.push(checkOptional('Resolver trigger eval', hasTriggerEval));
// 8. check-resolvable — we don't run it here (side effects + cost); we
// report whether the SKILL.md exists at all, which is the ground-truth
// input check-resolvable would consume.
items.push(checkOptional('check-resolvable input present',
existsSync(skillMd) && existsSync(RESOLVER_MD),
'run: gbrain check-resolvable'));
// 8. check-resolvable — invoke the real gate. Cached per process so
// iterating many skills only runs the subprocess once. Binary-missing
// is surfaced loudly so a silent false-pass can't happen.
const resolverResult = runCheckResolvableCached();
items.push(checkOptional('check-resolvable gate',
resolverResult.ok,
resolverResult.detail));
// 9. E2E — same as item 4 but required.
items.push(check('E2E test (either under e2e/ or integration test)', hasE2E, 'try /qa or test/e2e/'));
+1 -1
View File
@@ -37,7 +37,7 @@ This skill guarantees:
> **Filing rule:** Read `skills/_brain-filing-rules.md` before creating any new page.
## Iron Law: Back-Linking (MANDATORY)
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Every mention of a person or company with a brain page MUST create a back-link
FROM that entity's page TO the page mentioning them. An unlinked mention is a
+1 -1
View File
@@ -36,7 +36,7 @@ This skill guarantees:
- Every fact has an inline `[Source: ...]` citation
- Filing follows primary subject rules (not format-based)
## Iron Law: Back-Linking (MANDATORY)
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Every mention of a person or company with a brain page MUST create a back-link.
Format: `- **YYYY-MM-DD** | Referenced in [page title](path) — brief context`
+1 -1
View File
@@ -29,7 +29,7 @@ Ingest meetings, articles, media, documents, and conversations into the brain.
- State sections are rewritten with current best understanding, never appended to.
- Entity detection fires on every inbound message; notable entities get pages or updates.
## Iron Law: Back-Linking (MANDATORY)
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Every mention of a person or company with a brain page MUST create a back-link
FROM that entity's page TO the page mentioning them. An unlinked mention is a
+1 -1
View File
@@ -39,7 +39,7 @@ This skill guarantees:
- Raw source files preserved via `gbrain files upload-raw`
- Filing by primary subject, not by media format
## Iron Law: Back-Linking (MANDATORY)
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Every mention of a person or company with a brain page MUST create a back-link.
+1 -1
View File
@@ -34,7 +34,7 @@ This skill guarantees:
- Meeting is NOT fully ingested until enrich runs for every entity
- Back-links created bidirectionally
## Iron Law: Back-Linking (MANDATORY)
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Every attendee and company mentioned MUST get a back-link from their page to
the meeting page. An unlinked mention is a broken brain.
+3 -3
View File
@@ -9,7 +9,7 @@ feature_pitch:
# v0.11.0 Migration: Minions — host-agent instruction manual
**Audience: host agents (Wintermute, other OpenClaw deployments, future
**Audience: host agents (OpenClaw deployments, future
hosts) reading this AFTER `gbrain apply-migrations` has run its
mechanical phases.** The orchestrator in
`src/commands/migrations/v0_11_0.ts` is the runtime source of truth for
@@ -32,7 +32,7 @@ Non-empty? Each line is a TODO. Each `type` routes to a section below.
Gbrain rewrites cron entries whose handler name matches a gbrain
builtin (`sync`, `embed`, `lint`, `import`, `extract`, `backlinks`,
`autopilot-cycle`). For host-specific handlers (e.g. `ea-inbox-sweep`,
`frameio-scan`, `x-dm-triage`, `calendar-sync` on Wintermute), gbrain
`frameio-scan`, `x-dm-triage`, `calendar-sync` on your OpenClaw), gbrain
leaves the manifest alone and emits a TODO with shape:
```json
@@ -69,7 +69,7 @@ await worker.start();
### (b) Ship the bootstrap in your host repo
Autopilot already spawns `gbrain jobs work` as a child. Configure it to
spawn your custom worker binary (e.g. `wintermute-worker`) instead, or
spawn your custom worker binary (e.g. `your-openclaw-worker`) instead, or
register handlers as a side-effect module that the stock worker loads on
startup. Either path is documented in `plugin-handlers.md`.
+167
View File
@@ -0,0 +1,167 @@
---
version: 0.17.0
feature_pitch:
headline: "One brain maintenance cycle, two CLIs. `gbrain dream` delivers the README promise."
description: |
The README has said "the agent runs while I sleep, the dream cycle
scans every conversation, enriches missing entities, fixes broken
citations, consolidates memory" for a year. v0.17 makes that real
as a first-class command (`gbrain dream`) backed by one shared
primitive (`runCycle`). Autopilot users get lint + orphan sweep
added to their nightly cycle automatically — no config change.
Cron users get a single legible verb: `0 2 * * * gbrain dream`.
Both converge on the same phase order (lint → backlinks → sync →
extract → embed → orphans) so file fixes land in the DB the same
night, not the next.
recipe: null
tiers: null
---
# v0.17.0 Migration: `gbrain dream` + unified maintenance cycle
**Audience: agents + humans upgrading from v0.16.x. There is no
mechanical migration step required — the schema migration (v16
cycle-lock table) and behavior changes all apply automatically on
upgrade. This file documents what changed, how to verify it, and
the one opt-out users may care about.**
## What changed
### New command: `gbrain dream`
The brand-promise one-liner. Runs one brain maintenance cycle and
exits. Designed for cron.
```
gbrain dream # full 6-phase cycle
gbrain dream --dry-run --json # preview, agent-readable
gbrain dream --phase lint # single-phase (fast, targeted)
gbrain dream --pull # git pull before syncing
0 2 * * * gbrain dream --json # nightly cron
```
See `gbrain dream --help` for the full flag reference.
### Autopilot now runs lint + orphan sweep
`gbrain autopilot --install` users: on upgrade, your daemon's cycle
gains two phases it didn't run before:
- **lint --fix** — auto-fixes LLM artifacts, placeholder dates, bad
citations across the brain. Modifies files on disk.
- **orphan sweep** — reports (read-only) pages with no inbound
wikilinks. Visible in `gbrain jobs list` output for each
`autopilot-cycle` job.
No action required. The new phases run on the daemon's existing
interval.
### Shared primitive: `src/core/cycle.ts`
Three callers (dream CLI, autopilot inline path, autopilot-cycle
Minions handler) now all delegate to `runCycle(engine, opts)`. One
source of truth for what happens overnight.
### Cycle coordination via a DB lock table
`gbrain_cycle_locks` (new table, migration v16) replaces
session-scoped `pg_try_advisory_lock` which the v0.15.4
PgBouncer-transaction-pooler fix silently broke. The table has a
TTL (30 min), refreshed between phases, so crashed holders
auto-release.
## Verify after upgrade
```bash
# 1. Dream command exists:
gbrain dream --help
# 2. Run a dry cycle (safe, no writes):
gbrain dream --dry-run --json
# 3. If you run autopilot --install:
gbrain jobs list --status complete | head -5
# Each `autopilot-cycle` entry now has 6 phases in its report,
# not 4. Check a recent one with `gbrain jobs get <id>`.
# 4. Schema migration landed:
gbrain doctor # should show no pending migrations
```
Expected `gbrain dream --dry-run` output on a healthy brain:
```
Brain is healthy. 6 phase(s) checked in 1.3s.
```
Or with `--json`:
```json
{
"schema_version": "1",
"status": "clean",
"phases": [...],
"totals": { "lint_fixes": 0, "backlinks_added": 0, ... }
}
```
## Opt-outs for autopilot-installed users
If you explicitly do NOT want autopilot's daemon modifying files
(lint + backlinks phases write to disk):
**Option 1: disable those phases in cron-dream but keep autopilot
running.** Since dream is separate, you can run just the phases
you want from cron without touching autopilot:
```bash
# e.g. only re-embed and orphan-sweep nightly, skip file mutations:
0 2 * * * gbrain dream --phase orphans
```
**Option 2: uninstall autopilot and use cron-dream only.**
```bash
gbrain autopilot --uninstall
# Then add to your crontab:
0 2 * * * gbrain dream --pull
```
**Option 3: accept the default.** The new phases are conservative:
lint only fixes known-safe artifacts (em dashes, placeholder dates),
never destructive. Back-link fills are additive. If something does
go wrong, `gbrain dream --dry-run` always tells you what WOULD
change before you run it for real.
## Troubleshooting
**"cycle_already_running" in dream output:**
Another cycle (probably autopilot's daemon) is holding the lock.
Expected behavior — dream skipped to avoid racing the daemon. The
daemon's next interval will pick up the work.
**`gbrain dream --dry-run` reports changes when you expected none:**
Check `gbrain doctor` for drift: lint issues, stale embeddings,
missing back-links. Dream's dry-run is the honest preview of what
autopilot's daemon will do on its next cycle.
**Minion `autopilot-cycle` jobs failing after upgrade:**
Open a GitHub issue with the output of `gbrain jobs get <id>` for a
failing job. The new runCycle-backed handler preserves the
partial-failure semantic (one phase failing doesn't block future
cycles), but specific phases may surface new error classes.
## What did NOT change
- `gbrain autopilot --install` machinery (launchd / systemd /
crontab generators). Existing installs keep working.
- `~/.gbrain/autopilot.lock` daemon-singleton lockfile. Separate
concern from the new per-cycle lock.
- `gbrain jobs` interface. `gbrain jobs get <id>` now shows a
richer report structure (schema_version:"1"), but the surface
API is stable.
---
*This migration file is informational only. No mechanical step is
required — all changes apply automatically on `gbrain upgrade`.*
+1 -1
View File
@@ -275,7 +275,7 @@ Inject the key patterns into the agent's system context or AGENTS.md:
1. **Brain-agent loop** (Section 2): read before responding, write after learning
2. **Entity detection** (Section 3): spawn on every message, capture people/companies/ideas
3. **Source attribution** (Section 7): every fact needs `[Source: ...]`
4. **Iron law back-linking** (Section 15.4): every mention links back to the entity page
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Tell the user: "The production agent guide is at docs/GBRAIN_SKILLPACK.md. It covers
the brain-agent loop, entity detection, enrichment, meeting ingestion, and cron
+1 -1
View File
@@ -39,7 +39,7 @@ This skill guarantees:
- Back-links all entity mentions (Iron Law)
- Citations on every fact written
## Iron Law: Back-Linking (MANDATORY)
> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
Every time this skill creates or updates a brain page that mentions a person or company:
1. Check if that person/company has a brain page
+2 -2
View File
@@ -4,7 +4,7 @@ version: 1.0.0
description: |
Run `gbrain skillpack-check` to produce an agent-readable JSON health report
for the gbrain install. Wraps `gbrain doctor` + `gbrain apply-migrations
--list` so a host agent (Wintermute's morning-briefing, any OpenClaw cron)
--list` so a host agent (your OpenClaw's morning-briefing, any OpenClaw cron)
can see at a glance whether the skillpack needs attention.
Use when the user asks "is gbrain healthy?", when a cron fires a morning
@@ -40,7 +40,7 @@ Exit code:
## When to run
- **Daily cron** (e.g. Wintermute's `morning-briefing`): `gbrain skillpack-check --quiet`.
- **Daily cron** (e.g. your OpenClaw's `morning-briefing`): `gbrain skillpack-check --quiet`.
Exit code alone tells you if anything is wrong; surface a one-liner in the
briefing only when exit != 0. No JSON noise in happy-path briefings.
- **On demand**: `gbrain skillpack-check` for the full JSON when debugging.
+44 -1
View File
@@ -19,7 +19,7 @@ for (const op of operations) {
}
// CLI-only commands that bypass the operation layer
const CLI_ONLY = new Set(['init', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'sync', 'extract', 'features', 'autopilot', 'graph-query', 'jobs', 'apply-migrations', 'skillpack-check', 'resolvers', 'integrity', 'repair-jsonb', 'orphans']);
const CLI_ONLY = new Set(['init', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'sync', 'extract', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'dream', 'check-resolvable', 'repos']);
async function main() {
// Parse global flags (--quiet / --progress-json / --progress-interval)
@@ -310,6 +310,16 @@ async function handleCliOnly(command: string, args: string[]) {
await runLint(args);
return;
}
if (command === 'check-resolvable') {
const { runCheckResolvable } = await import('./commands/check-resolvable.ts');
await runCheckResolvable(args);
return;
}
if (command === 'repos') {
const { handleRepos } = await import('./commands/repos.ts');
await handleRepos(args);
return;
}
if (command === 'report') {
const { runReport } = await import('./commands/report.ts');
await runReport(args);
@@ -358,6 +368,25 @@ async function handleCliOnly(command: string, args: string[]) {
return;
}
if (command === 'dream') {
// Dream mirrors doctor's pattern: filesystem phases run without a DB,
// so an engine connection failure is non-fatal. runCycle honestly
// reports DB phases as skipped when engine is null.
const { runDream } = await import('./commands/dream.ts');
let eng: BrainEngine | null = null;
try {
eng = await connectEngine();
} catch {
// DB unavailable — lint + backlinks still run against the brain dir.
}
try {
await runDream(eng, args);
} finally {
if (eng) await eng.disconnect();
}
return;
}
// All remaining CLI-only commands need a DB connection
const engine = await connectEngine();
try {
@@ -413,6 +442,11 @@ async function handleCliOnly(command: string, args: string[]) {
await runJobs(engine, args);
break;
}
case 'agent': {
const { runAgent } = await import('./commands/agent.ts');
await runAgent(engine, args);
break;
}
case 'sync': {
const { runSync } = await import('./commands/sync.ts');
await runSync(engine, args);
@@ -552,8 +586,17 @@ TOOLS
check-backlinks <check|fix> [dir] Find/fix missing back-links across brain
lint <dir|file> [--fix] Catch LLM artifacts, placeholder dates, bad frontmatter
orphans [--json] [--count] Find pages with no inbound wikilinks
dream [--dry-run] [--json] Run the overnight maintenance cycle once (cron-friendly).
See also: autopilot --install (continuous daemon).
check-resolvable [--json] [--fix] Validate skill tree (reachability/MECE/DRY)
report --type <name> --content ... Save timestamped report to brain/reports/
MULTI-REPO
repos list Show configured repos
repos add <path> [--name N] Add a repo [--strategy markdown|code|auto]
repos remove <name> Remove a repo
sync --all Sync all configured repos
JOBS (Minions)
jobs submit <name> [--params JSON] Submit background job [--follow] [--dry-run]
jobs list [--status S] [--limit N] List jobs
+185
View File
@@ -0,0 +1,185 @@
/**
* `gbrain agent logs <job_id> [--follow] [--since <spec>]`
*
* Reads two sources and merges them chronologically:
* - ~/.gbrain/audit/subagent-jobs-*.jsonl (heartbeat + submission events
* lives on the WORKER's filesystem, so this CLI's effectiveness is
* host-local today; see docs/guides/plugin-authors.md caveat #2)
* - subagent_messages (DB rows, authoritative for persisted conversation)
*
* No new DB tables; all the infrastructure landed in prior Lane commits.
*/
import type { BrainEngine } from '../core/engine.ts';
import { readSubagentAuditForJob } from '../core/minions/handlers/subagent-audit.ts';
import type { SubagentAuditEvent } from '../core/minions/handlers/subagent-audit.ts';
import { loadTranscriptRows, renderTranscript } from '../core/minions/transcript.ts';
import type { SubagentMessageRow } from '../core/minions/transcript.ts';
export interface AgentLogsOpts {
follow?: boolean;
/** ISO-8601 timestamp OR relative like "5m" / "1h" / "2d". */
since?: string;
/** Override poll interval for --follow. Default 1000ms. */
pollMs?: number;
/** Injectable writer for testing; default process.stdout.write. */
write?: (s: string) => void;
/** Abort to cut off a --follow loop cleanly (tests + Ctrl-C). */
signal?: AbortSignal;
}
const TERMINAL_STATUSES = new Set(['completed', 'failed', 'dead', 'cancelled']);
export async function runAgentLogs(
engine: BrainEngine,
jobId: number,
opts: AgentLogsOpts = {},
): Promise<void> {
const write = opts.write ?? ((s: string) => { process.stdout.write(s); });
const sinceIso = parseSince(opts.since);
// Seeded render: dump everything we have right now.
let lastTs: string | undefined = sinceIso;
lastTs = await dumpSince(engine, jobId, lastTs, write);
if (!opts.follow) return;
const pollMs = opts.pollMs ?? 1000;
while (!opts.signal?.aborted) {
await sleep(pollMs, opts.signal);
lastTs = await dumpSince(engine, jobId, lastTs, write);
// Break on terminal job status so --follow exits once the run is done.
const status = await readJobStatus(engine, jobId);
if (status && TERMINAL_STATUSES.has(status)) {
write(`\n[gbrain agent] job ${jobId} reached terminal state: ${status}\n`);
return;
}
}
}
/**
* Dump events with ts >= sinceIso. Returns the max ts seen so the next
* poll round filters cleanly. When `sinceIso` is undefined on first call,
* everything is dumped.
*/
async function dumpSince(
engine: BrainEngine,
jobId: number,
sinceIso: string | undefined,
write: (s: string) => void,
): Promise<string | undefined> {
const audit = readSubagentAuditForJob(jobId, sinceIso ? { sinceIso } : {});
const { messages, tools } = await loadTranscriptRows(engine, jobId);
// Merge audit events + message rows into one timeline ordered by ts.
const merged: Array<{ ts: string; line: string }> = [];
for (const e of audit) {
if (sinceIso && e.ts <= sinceIso) continue;
merged.push({ ts: e.ts, line: formatAudit(e) });
}
for (const m of messages) {
const ts = m.ended_at.toISOString();
if (sinceIso && ts <= sinceIso) continue;
merged.push({ ts, line: formatMessage(m) });
}
merged.sort((a, b) => a.ts.localeCompare(b.ts));
let maxTs = sinceIso;
for (const item of merged) {
write(`${item.ts} ${item.line}\n`);
if (!maxTs || item.ts > maxTs) maxTs = item.ts;
}
// Transcript tail (renders the full message/tool tree) only if we
// actually have messages and the job is in a terminal state. This
// avoids spamming a half-rendered transcript mid-run.
if (messages.length > 0 && !sinceIso) {
const status = await readJobStatus(engine, jobId);
if (status && TERMINAL_STATUSES.has(status)) {
write('\n');
write(renderTranscript(messages, tools));
write('\n');
}
}
return maxTs;
}
function formatAudit(e: SubagentAuditEvent): string {
if (e.type === 'submission') {
return `[submission] ${e.caller} model=${e.model ?? '?'} tools=${e.tools_count ?? 0}`;
}
// heartbeat
const parts = [`[${e.event}]`, `turn=${e.turn_idx}`];
if (e.tool_name) parts.push(`tool=${e.tool_name}`);
if (e.ms_elapsed != null) parts.push(`${e.ms_elapsed}ms`);
if (e.tokens) {
const t = e.tokens;
const tokStr = [
t.in ? `in=${t.in}` : null,
t.out ? `out=${t.out}` : null,
t.cache_read ? `cache_read=${t.cache_read}` : null,
t.cache_create ? `cache_create=${t.cache_create}` : null,
].filter(Boolean).join(' ');
if (tokStr) parts.push(`tokens(${tokStr})`);
}
if (e.error) parts.push(`error="${e.error.slice(0, 100)}"`);
return parts.join(' ');
}
function formatMessage(m: SubagentMessageRow): string {
const blockTypes = m.content_blocks.map(b => b.type).join(',');
return `[message #${m.message_idx} ${m.role}] blocks=${blockTypes || '(empty)'}`;
}
async function readJobStatus(engine: BrainEngine, jobId: number): Promise<string | null> {
const rows = await engine.executeRaw<{ status: string }>(
`SELECT status FROM minion_jobs WHERE id = $1`,
[jobId],
);
return rows[0]?.status ?? null;
}
const RELATIVE_RE = /^(\d+)\s*(s|m|h|d)$/i;
/** Parse `--since`. Accepts ISO-8601 or relative ("5m", "1h", "2d"). */
export function parseSince(input: string | undefined): string | undefined {
if (!input) return undefined;
const trimmed = input.trim();
if (!trimmed) return undefined;
const rel = RELATIVE_RE.exec(trimmed);
if (rel) {
const [, nStr, unitRaw] = rel;
const unit = unitRaw!.toLowerCase();
const n = parseInt(nStr!, 10);
const mult = unit === 's' ? 1000
: unit === 'm' ? 60_000
: unit === 'h' ? 3_600_000
: 86_400_000; // 'd'
return new Date(Date.now() - n * mult).toISOString();
}
// Assume ISO. `new Date(input).toISOString()` both validates and
// normalizes; invalid ISO throws.
const d = new Date(trimmed);
if (isNaN(d.getTime())) {
throw new Error(`--since: could not parse "${input}" as ISO-8601 or relative (e.g. "5m", "1h")`);
}
return d.toISOString();
}
function sleep(ms: number, signal?: AbortSignal): Promise<void> {
return new Promise((resolve) => {
const t = setTimeout(() => { signal?.removeEventListener('abort', onAbort); resolve(); }, ms);
const onAbort = () => { clearTimeout(t); resolve(); };
signal?.addEventListener('abort', onAbort, { once: true });
});
}
export const __testing = {
parseSince,
formatAudit,
formatMessage,
dumpSince,
};
+333
View File
@@ -0,0 +1,333 @@
/**
* `gbrain agent` CLI: the user-facing entry point for the v0.15 subagent
* runtime.
*
* gbrain agent run <prompt> [flags]
* gbrain agent logs <job_id> [--follow] [--since <spec>]
*
* `run` submits a subagent job (or fan-out of N subagents + aggregator)
* under the trusted-submit flag so the PROTECTED_JOB_NAMES guard doesn't
* reject. It does NOT execute the loop here the handler runs in a
* `gbrain jobs work` process. `--follow` tails status until terminal;
* without `--follow` (or with `--detach`) the CLI prints the job id and
* exits, leaving the user to check back with `gbrain agent logs`.
*/
import * as fs from 'node:fs';
import type { BrainEngine } from '../core/engine.ts';
import { MinionQueue } from '../core/minions/queue.ts';
import { waitForCompletion, TimeoutError } from '../core/minions/wait-for-completion.ts';
import type { MinionJobInput, SubagentHandlerData, AggregatorHandlerData } from '../core/minions/types.ts';
import { runAgentLogs } from './agent-logs.ts';
// ── arg parsing helpers ────────────────────────────────────
function parseFlag(args: string[], flag: string): string | undefined {
const idx = args.indexOf(flag);
return idx >= 0 && idx + 1 < args.length ? args[idx + 1] : undefined;
}
function hasFlag(args: string[], flag: string): boolean { return args.includes(flag); }
/** Keep CLI args that look like flags from being eaten as the prompt. */
function isKnownFlag(s: string): boolean {
return s.startsWith('--');
}
// ── command dispatcher ────────────────────────────────────
export async function runAgent(engine: BrainEngine, args: string[]): Promise<void> {
const sub = args[0];
if (!sub || sub === '--help' || sub === '-h') {
printHelp();
return;
}
switch (sub) {
case 'run':
await runAgentRun(engine, args.slice(1));
return;
case 'logs':
await runAgentLogsCmd(engine, args.slice(1));
return;
default:
console.error(`gbrain agent: unknown subcommand "${sub}"`);
printHelp();
process.exit(2);
}
}
function printHelp(): void {
console.log(`gbrain agent — durable LLM agent runs (v0.15)
USAGE
gbrain agent run <prompt> [flags]
gbrain agent logs <job_id> [--follow] [--since <spec>]
SUBMITTING
gbrain agent run <prompt>
--subagent-def <name> Named plugin subagent (from GBRAIN_PLUGIN_PATH)
--model <id> Anthropic model id (defaults to sonnet)
--max-turns <n> Max assistant turns (default 20)
--tools a,b,c Subset of registered tool names (comma list)
--timeout-ms <n> Per-job wall-clock timeout
--fanout-manifest <path> JSON array of {prompt, input_vars?} one child each
--follow Tail status until terminal (default on TTY)
--detach Submit + print job id, exit immediately
Flags after \`run\` up to the first unrecognized token are parsed; the
remainder is the prompt. Use \`--\` to explicitly terminate flag parsing.
VIEWING
gbrain agent logs <job_id>
--follow Keep polling until the job reaches terminal
--since <spec> ISO-8601 timestamp OR relative ("5m","1h","2d")
NOTES
Submitting subagent jobs is trusted-only; MCP submitters receive
permission_denied. The worker needs ANTHROPIC_API_KEY set, or the
first LLM turn of a claimed job fails.
`);
}
// ── `gbrain agent run` ────────────────────────────────────
interface RunFlags {
subagentDef?: string;
model?: string;
maxTurns?: number;
tools?: string[];
timeoutMs?: number;
fanoutManifest?: string;
follow: boolean;
detach: boolean;
}
function parseRunFlags(args: string[]): { flags: RunFlags; rest: string[] } {
const flags: RunFlags = {
follow: process.stdout.isTTY === true,
detach: false,
};
let i = 0;
while (i < args.length) {
const a = args[i];
if (a === '--') { i++; break; }
if (!isKnownFlag(a!)) break;
switch (a) {
case '--subagent-def': flags.subagentDef = args[++i]; i++; break;
case '--model': flags.model = args[++i]; i++; break;
case '--max-turns': flags.maxTurns = parseInt(args[++i] ?? '', 10); i++; break;
case '--tools': flags.tools = (args[++i] ?? '').split(',').map(s => s.trim()).filter(Boolean); i++; break;
case '--timeout-ms': flags.timeoutMs = parseInt(args[++i] ?? '', 10); i++; break;
case '--fanout-manifest': flags.fanoutManifest = args[++i]; i++; break;
case '--follow': flags.follow = true; i++; break;
case '--no-follow': flags.follow = false; i++; break;
case '--detach': flags.detach = true; flags.follow = false; i++; break;
default:
throw new Error(`unknown flag: ${a}. Run \`gbrain agent run --help\` for usage.`);
}
}
return { flags, rest: args.slice(i) };
}
export async function runAgentRun(engine: BrainEngine, args: string[]): Promise<void> {
const { flags, rest } = parseRunFlags(args);
const queue = new MinionQueue(engine);
// Fan-out path: --fanout-manifest supplies explicit child inputs. The
// aggregator submits first (so its id is available as parent for each
// child); children submit with on_child_fail='continue' so mixed
// outcomes don't cascade; aggregator waits in waiting-children until
// Lane 1B's terminal-set check unblocks it.
if (flags.fanoutManifest) {
await runFanout(engine, queue, flags, rest.join(' '));
return;
}
const prompt = rest.join(' ').trim();
if (!prompt) {
console.error('gbrain agent run: prompt is required');
process.exit(2);
}
const data: SubagentHandlerData = { prompt };
if (flags.subagentDef) data.subagent_def = flags.subagentDef;
if (flags.model) data.model = flags.model;
if (flags.maxTurns) data.max_turns = flags.maxTurns;
if (flags.tools && flags.tools.length > 0) data.allowed_tools = flags.tools;
const submitOpts: Partial<MinionJobInput> = { max_stalled: 3 };
if (flags.timeoutMs) submitOpts.timeout_ms = flags.timeoutMs;
const job = await queue.add('subagent', data as unknown as Record<string, unknown>, submitOpts, {
allowProtectedSubmit: true,
});
process.stderr.write(`submitted: job ${job.id} (subagent)\n`);
if (flags.detach || !flags.follow) {
process.stdout.write(String(job.id) + '\n');
return;
}
await followJob(engine, queue, job.id, flags.timeoutMs);
}
// ── fan-out ───────────────────────────────────────────────
async function runFanout(engine: BrainEngine, queue: MinionQueue, flags: RunFlags, promptTemplate: string): Promise<void> {
const manifestPath = flags.fanoutManifest!;
let manifest: Array<{ prompt?: string; input_vars?: Record<string, unknown> }>;
try {
const raw = fs.readFileSync(manifestPath, 'utf8');
const parsed = JSON.parse(raw);
if (!Array.isArray(parsed)) throw new Error('manifest must be a JSON array');
manifest = parsed as typeof manifest;
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
console.error(`gbrain agent run: invalid --fanout-manifest ${manifestPath}: ${msg}`);
process.exit(2);
}
if (manifest.length === 0) {
console.error('gbrain agent run: --fanout-manifest is empty; nothing to run');
process.exit(2);
}
// Short-circuit: 1 entry → single subagent, no aggregator.
if (manifest.length === 1) {
const entry = manifest[0]!;
const data: SubagentHandlerData = {
prompt: entry.prompt ?? promptTemplate,
...(entry.input_vars ? { input_vars: entry.input_vars } : {}),
...(flags.subagentDef ? { subagent_def: flags.subagentDef } : {}),
...(flags.model ? { model: flags.model } : {}),
...(flags.maxTurns ? { max_turns: flags.maxTurns } : {}),
...(flags.tools && flags.tools.length > 0 ? { allowed_tools: flags.tools } : {}),
};
const submitOpts: Partial<MinionJobInput> = { max_stalled: 3 };
if (flags.timeoutMs) submitOpts.timeout_ms = flags.timeoutMs;
const job = await queue.add('subagent', data as unknown as Record<string, unknown>, submitOpts, {
allowProtectedSubmit: true,
});
process.stderr.write(`submitted: job ${job.id} (single-entry manifest short-circuit)\n`);
if (flags.detach || !flags.follow) { process.stdout.write(`${job.id}\n`); return; }
await followJob(engine, queue, job.id, flags.timeoutMs);
return;
}
// N-entry fan-out: aggregator first (so we have its id as parent), then
// N children, then flip the aggregator's children_ids to include them.
const aggregatorSeed: AggregatorHandlerData = { children_ids: [] };
const aggregator = await queue.add(
'subagent_aggregator',
aggregatorSeed as unknown as Record<string, unknown>,
{ max_stalled: 3 },
{ allowProtectedSubmit: true },
);
const childIds: number[] = [];
for (const entry of manifest) {
const data: SubagentHandlerData = {
prompt: entry.prompt ?? promptTemplate,
...(entry.input_vars ? { input_vars: entry.input_vars } : {}),
...(flags.subagentDef ? { subagent_def: flags.subagentDef } : {}),
...(flags.model ? { model: flags.model } : {}),
...(flags.maxTurns ? { max_turns: flags.maxTurns } : {}),
...(flags.tools && flags.tools.length > 0 ? { allowed_tools: flags.tools } : {}),
};
const submitOpts: Partial<MinionJobInput> = {
parent_job_id: aggregator.id,
on_child_fail: 'continue', // mixed-outcome aggregation
max_stalled: 3,
};
if (flags.timeoutMs) submitOpts.timeout_ms = flags.timeoutMs;
const child = await queue.add('subagent', data as unknown as Record<string, unknown>, submitOpts, {
allowProtectedSubmit: true,
});
childIds.push(child.id);
}
// Update the aggregator's data with the final children_ids. We have to
// do this after submission because each add() returns the committed
// row's id; the aggregator's seed started with an empty array.
await engine.executeRaw(
`UPDATE minion_jobs SET data = jsonb_set(data, '{children_ids}', $1::jsonb) WHERE id = $2`,
[JSON.stringify(childIds), aggregator.id],
);
process.stderr.write(
`submitted: aggregator job ${aggregator.id} + ${childIds.length} subagent children ` +
`(${childIds[0]}..${childIds[childIds.length - 1]})\n`,
);
if (flags.detach || !flags.follow) {
process.stdout.write(`${aggregator.id}\n`);
return;
}
await followJob(engine, queue, aggregator.id, flags.timeoutMs);
}
// ── follow ────────────────────────────────────────────────
async function followJob(engine: BrainEngine, queue: MinionQueue, jobId: number, timeoutMs?: number): Promise<void> {
process.stderr.write(`[gbrain agent] following job ${jobId} (Ctrl-C to detach)...\n`);
const ac = new AbortController();
const onSigint = () => ac.abort();
process.once('SIGINT', onSigint);
try {
// Streaming logs happen in the background; we poll the terminal state
// in parallel so the function returns as soon as the job completes.
const logsP = runAgentLogs(engine, jobId, { follow: true, signal: ac.signal, pollMs: 1000 });
try {
const job = await waitForCompletion(queue, jobId, {
timeoutMs: timeoutMs ?? 24 * 60 * 60 * 1000,
pollMs: 1000,
signal: ac.signal,
});
ac.abort();
await logsP.catch(() => {});
process.stderr.write(`[gbrain agent] job ${jobId} terminal: ${job.status}\n`);
if (job.result != null) process.stdout.write(JSON.stringify(job.result, null, 2) + '\n');
if (job.status !== 'completed') process.exit(1);
} catch (e) {
if (e instanceof TimeoutError) {
process.stderr.write(`[gbrain agent] timeout after ${e.elapsedMs}ms — job is still running. Check with: gbrain jobs get ${jobId}\n`);
process.exit(3);
}
throw e;
}
} finally {
process.removeListener('SIGINT', onSigint);
}
}
// ── `gbrain agent logs` ────────────────────────────────────
async function runAgentLogsCmd(engine: BrainEngine, args: string[]): Promise<void> {
const jobIdStr = args.find(a => !isKnownFlag(a));
if (!jobIdStr) {
console.error('gbrain agent logs: <job_id> is required');
process.exit(2);
}
const jobId = parseInt(jobIdStr, 10);
if (!Number.isFinite(jobId) || jobId <= 0) {
console.error(`gbrain agent logs: "${jobIdStr}" is not a valid job id`);
process.exit(2);
}
const follow = hasFlag(args, '--follow');
const since = parseFlag(args, '--since');
const ac = new AbortController();
const onSigint = () => ac.abort();
process.once('SIGINT', onSigint);
try {
await runAgentLogs(engine, jobId, { follow, since, signal: ac.signal });
} finally {
process.removeListener('SIGINT', onSigint);
}
}
// Expose for tests.
export const __testing = {
parseRunFlags,
};
+33 -20
View File
@@ -77,7 +77,15 @@ export function resolveGbrainCliPath(): string {
export async function runAutopilot(engine: BrainEngine, args: string[]) {
if (args.includes('--help') || args.includes('-h')) {
console.log('Usage: gbrain autopilot [--repo <path>] [--interval N] [--json]\n gbrain autopilot --install [--repo <path>]\n gbrain autopilot --uninstall\n gbrain autopilot --status [--json]\n\nSelf-maintaining brain daemon. Runs sync + extract + embed + backlinks in a loop.');
console.log(
'Usage: gbrain autopilot [--repo <path>] [--interval N] [--json]\n' +
' gbrain autopilot --install [--repo <path>]\n' +
' gbrain autopilot --uninstall\n' +
' gbrain autopilot --status [--json]\n\n' +
'Self-maintaining brain daemon. Runs the full maintenance cycle\n' +
'(lint + backlinks + sync + extract + embed + orphans) on an interval.\n\n' +
'For a one-shot cron-triggered cycle, see `gbrain dream`.',
);
return;
}
@@ -233,27 +241,32 @@ export async function runAutopilot(engine: BrainEngine, args: string[]) {
}
} catch (e) { logError('dispatch', e); cycleOk = false; }
} else {
// Inline fallback — same as pre-v0.11.1 behavior.
// 1. Sync
// Inline fallback — delegate to runCycle so lint + backlinks +
// orphan sweep run too (previously this path only did sync +
// extract + embed, which didn't match the Minions-dispatch
// path's phase set). Now both converge on the same primitive.
try {
const { performSync } = await import('./sync.ts');
const result = await performSync(engine, { repoPath, noEmbed: true });
if (result.status === 'synced') {
console.log(`[sync] +${result.added} ~${result.modified} -${result.deleted}`);
const { runCycle } = await import('../core/cycle.ts');
const report = await runCycle(engine, {
brainDir: repoPath,
// Autopilot daemon path: pulls by default (matches
// pre-v0.17 autopilot behavior). CLI dream defaults false
// for cron safety; that choice is scoped to dream only.
pull: true,
yieldBetweenPhases: async () => {
await new Promise(r => setImmediate(r));
},
});
if (report.status === 'failed' || report.status === 'partial') {
cycleOk = false;
}
} catch (e) { logError('sync', e); cycleOk = false; }
// 2. Extract (full brain, incremental dedup handles repeats)
try {
const { runExtractCore } = await import('./extract.ts');
await runExtractCore(engine, { mode: 'all', dir: repoPath });
} catch (e) { logError('extract', e); cycleOk = false; }
// 3. Embed stale
try {
const { runEmbedCore } = await import('./embed.ts');
await runEmbedCore(engine, { stale: true });
} catch (e) { logError('embed', e); cycleOk = false; }
if (jsonMode) {
process.stderr.write(JSON.stringify({ event: 'cycle-inline', status: report.status, duration_ms: report.duration_ms, totals: report.totals }) + '\n');
} else {
const t = report.totals;
console.log(`[cycle-inline ${report.status}] lint=${t.lint_fixes} backlinks=${t.backlinks_added} synced=${t.pages_synced} extracted=${t.pages_extracted} embedded=${t.pages_embedded} orphans=${t.orphans_found}`);
}
} catch (e) { logError('cycle-inline', e); cycleOk = false; }
}
// 4. Health check + adaptive interval (same for both paths)
+260
View File
@@ -0,0 +1,260 @@
/**
* gbrain check-resolvable Standalone CLI gate for skill-tree integrity.
*
* Thin wrapper over `src/core/check-resolvable.ts`. Exit-code rule is stricter
* than `gbrain doctor`'s resolver_health: this command exits 1 on ANY issue
* (errors OR warnings) so CI can gate on a single command. Honors the README
* contract: "Exits non-zero if anything is off."
*
* Currently covers 4 of 6 checks from the original design: reachability,
* MECE overlap, MECE gap, DRY violations. Checks 5 (trigger routing eval)
* and 6 (brain filing) are tracked as separate GitHub issues and surfaced
* via the `deferred` field in --json output.
*/
import { resolve as resolvePath, isAbsolute } from 'path';
import {
checkResolvable,
autoFixDryViolations,
type ResolvableReport,
type ResolvableIssue,
type AutoFixReport,
} from '../core/check-resolvable.ts';
import { findRepoRoot } from '../core/repo-root.ts';
// ---------------------------------------------------------------------------
// Types
// ---------------------------------------------------------------------------
export interface DeferredCheck {
check: number;
name: string;
issue: string;
}
export interface Envelope {
ok: boolean;
skillsDir: string | null;
report: ResolvableReport | null;
autoFix: AutoFixReport | null;
deferred: DeferredCheck[];
error: 'no_skills_dir' | null;
message: string | null;
}
export interface Flags {
help: boolean;
json: boolean;
fix: boolean;
dryRun: boolean;
verbose: boolean;
skillsDir: string | null;
}
// TBD: fill these issue URLs after filing the GitHub issues pre-PR.
// grep for 'TBD-check-5' / 'TBD-check-6' before shipping.
export const DEFERRED: DeferredCheck[] = [
{
check: 5,
name: 'trigger_routing_eval',
issue: 'https://github.com/garrytan/gbrain/issues?q=TBD-check-5',
},
{
check: 6,
name: 'brain_filing',
issue: 'https://github.com/garrytan/gbrain/issues?q=TBD-check-6',
},
];
const HELP_TEXT = `gbrain check-resolvable [options]
Validate the skill tree: reachability, MECE overlap, DRY violations, and
gap detection. Exits non-zero if any issues are found (errors OR warnings).
Options:
--json Machine-readable JSON (stable envelope)
--fix Apply DRY auto-fixes before checking
--dry-run With --fix, preview only; no writes
--verbose Show passing checks and the deferred-check note
--skills-dir PATH Override the auto-detected skills/ directory
--help Show this message
Deferred to separate issues (see --json .deferred[]):
- Check 5: trigger routing eval
- Check 6: brain filing
`;
// ---------------------------------------------------------------------------
// Flag parsing — permissive on unknown flags, matching lint/orphans/publish.
// ---------------------------------------------------------------------------
export function parseFlags(argv: string[]): Flags {
const flags: Flags = {
help: false,
json: false,
fix: false,
dryRun: false,
verbose: false,
skillsDir: null,
};
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === '--help' || a === '-h') flags.help = true;
else if (a === '--json') flags.json = true;
else if (a === '--fix') flags.fix = true;
else if (a === '--dry-run') flags.dryRun = true;
else if (a === '--verbose') flags.verbose = true;
else if (a === '--skills-dir') {
flags.skillsDir = argv[i + 1] ?? null;
i++;
} else if (a?.startsWith('--skills-dir=')) {
flags.skillsDir = a.slice('--skills-dir='.length) || null;
}
// unknown flags silently ignored
}
return flags;
}
// ---------------------------------------------------------------------------
// Skills-dir resolution
// ---------------------------------------------------------------------------
export function resolveSkillsDir(flags: Flags): { dir: string | null; error: Envelope['error']; message: string | null } {
if (flags.skillsDir) {
const dir = isAbsolute(flags.skillsDir)
? flags.skillsDir
: resolvePath(process.cwd(), flags.skillsDir);
return { dir, error: null, message: null };
}
const repoRoot = findRepoRoot();
if (!repoRoot) {
return {
dir: null,
error: 'no_skills_dir',
message:
'Could not locate skills/RESOLVER.md from cwd. Pass --skills-dir <path> or run from inside a gbrain repo.',
};
}
return { dir: resolvePath(repoRoot, 'skills'), error: null, message: null };
}
// ---------------------------------------------------------------------------
// Human output (mirrors doctor's resolver_health formatting)
// ---------------------------------------------------------------------------
function renderHuman(env: Envelope, flags: Flags): void {
if (env.error === 'no_skills_dir') {
console.error(env.message);
return;
}
const report = env.report!;
if (flags.fix && env.autoFix) {
printAutoFixHuman(env.autoFix, flags.dryRun);
}
if (report.ok && report.issues.length === 0) {
console.log(`resolver_health: OK — ${report.summary.total_skills} skills, all reachable`);
} else {
const errors = report.issues.filter(i => i.severity === 'error');
const warnings = report.issues.filter(i => i.severity === 'warning');
const status = errors.length > 0 ? 'FAIL' : 'WARN';
console.log(
`resolver_health: ${status}${report.issues.length} issue(s): ${errors.length} error(s), ${warnings.length} warning(s)`,
);
for (const iss of report.issues) {
console.log(formatIssueLine(iss));
}
}
if (flags.verbose) {
const urls = DEFERRED.map(d => `${d.name} (${d.issue})`).join(', ');
console.log(`Deferred: ${urls}`);
}
}
function formatIssueLine(iss: ResolvableIssue): string {
const type = iss.type.padEnd(18);
const skill = iss.skill.padEnd(24);
return `${type} ${skill} ${iss.action}`;
}
function printAutoFixHuman(autoFix: AutoFixReport, dryRun: boolean): void {
const verb = dryRun ? 'PROPOSED' : 'APPLIED';
for (const outcome of autoFix.fixed) {
console.log(`[${verb}] ${outcome.skillPath} (${outcome.patternLabel})`);
}
const n = autoFix.fixed.length;
const s = autoFix.skipped.length;
if (n === 0 && s === 0) {
console.log('check-resolvable --fix: no DRY violations to repair.');
return;
}
const label = dryRun ? 'fixes proposed' : 'fixes applied';
console.log(`${n} ${label}${s > 0 ? `, ${s} skipped:` : '.'}`);
for (const sk of autoFix.skipped) {
const hint = sk.reason === 'working_tree_dirty' ? ' (run `git stash` first)' : '';
console.log(` - ${sk.skillPath}: ${sk.reason}${hint}`);
}
if (dryRun && n > 0) console.log('Run without --dry-run to apply.\n');
}
// ---------------------------------------------------------------------------
// Entry point
// ---------------------------------------------------------------------------
export async function runCheckResolvable(args: string[]): Promise<void> {
const flags = parseFlags(args);
if (flags.help) {
console.log(HELP_TEXT);
process.exit(0);
}
const { dir, error, message } = resolveSkillsDir(flags);
if (error === 'no_skills_dir') {
const env: Envelope = {
ok: false,
skillsDir: null,
report: null,
autoFix: null,
deferred: DEFERRED,
error,
message,
};
if (flags.json) {
console.log(JSON.stringify(env, null, 2));
} else {
renderHuman(env, flags);
}
process.exit(1);
}
const skillsDir = dir!;
let autoFix: AutoFixReport | null = null;
if (flags.fix) {
autoFix = autoFixDryViolations(skillsDir, { dryRun: flags.dryRun });
}
const report = checkResolvable(skillsDir);
const env: Envelope = {
ok: report.issues.length === 0,
skillsDir,
report,
autoFix,
deferred: DEFERRED,
error: null,
message: null,
};
if (flags.json) {
console.log(JSON.stringify(env, null, 2));
} else {
renderHuman(env, flags);
}
process.exit(env.ok ? 0 : 1);
}
+2 -12
View File
@@ -3,6 +3,7 @@ import * as db from '../core/db.ts';
import { LATEST_VERSION } from '../core/migrate.ts';
import { checkResolvable } from '../core/check-resolvable.ts';
import { autoFixDryViolations, type AutoFixReport, type FixOutcome } from '../core/dry-fix.ts';
import { findRepoRoot } from '../core/repo-root.ts';
import { loadCompletedMigrations } from '../core/preferences.ts';
import { createProgress, startHeartbeat, type ProgressReporter } from '../core/progress.ts';
import { getCliOptions, cliOptsToProgressOptions } from '../core/cli-options.ts';
@@ -96,7 +97,7 @@ export async function runDoctor(engine: BrainEngine | null, args: string[], dbSo
// status:"complete" for the same version, the install is mid-migration.
// Typical cause: v0.11.0 stopgap wrote a partial record but nobody ran
// `gbrain apply-migrations --yes` afterward. This check fires on every
// `gbrain doctor` invocation so Wintermute's health skill catches it.
// `gbrain doctor` invocation so your OpenClaw's health skill catches it.
try {
const completed = loadCompletedMigrations();
const byVersion = new Map<string, { complete: boolean; partial: boolean }>();
@@ -602,17 +603,6 @@ function printAutoFixReport(report: AutoFixReport, dryRun: boolean, jsonOutput:
if (dryRun && n > 0) console.log('\nRun without --dry-run to apply.');
}
/** Find the GBrain repo root by walking up from cwd looking for skills/RESOLVER.md */
function findRepoRoot(): string | null {
let dir = process.cwd();
for (let i = 0; i < 10; i++) {
if (existsSync(join(dir, 'skills', 'RESOLVER.md'))) return dir;
const parent = join(dir, '..');
if (parent === dir) break;
dir = parent;
}
return null;
}
/** Quick skill conformance check — frontmatter + required sections */
function checkSkillConformance(skillsDir: string): Check {
+209
View File
@@ -0,0 +1,209 @@
/**
* gbrain dream run one brain maintenance cycle.
*
* The README brand promise: "the agent runs while I sleep, the dream
* cycle ... I wake up and the brain is smarter." Cron-friendly, JSON
* report, phase-selectable.
*
* Thin alias over runCycle (src/core/cycle.ts). Both this command and
* `gbrain autopilot` converge on the same primitive so there's one
* source of truth for what "overnight maintenance" means.
*
* Usage:
* gbrain dream # full 6-phase cycle
* gbrain dream --dry-run # preview, no writes
* gbrain dream --json # CycleReport JSON (for agents)
* gbrain dream --phase lint # run a single phase
* gbrain dream --pull # also git pull the brain repo
* gbrain dream --dir /path/to/brain # explicit brain location
*
* Cron: 0 2 * * * gbrain dream --json >> /var/log/gbrain-dream.log
*
* Related: `gbrain autopilot --install` for continuous daemonized
* maintenance. dream is the one-shot, autopilot is the scheduler.
*/
import type { BrainEngine } from '../core/engine.ts';
import {
runCycle,
ALL_PHASES,
type CyclePhase,
type CycleReport,
} from '../core/cycle.ts';
import { existsSync } from 'fs';
interface DreamArgs {
json: boolean;
dryRun: boolean;
pull: boolean;
phase: CyclePhase | null;
dir: string | null;
help: boolean;
}
function parseArgs(args: string[]): DreamArgs {
const phaseIdx = args.indexOf('--phase');
const rawPhase = phaseIdx !== -1 ? args[phaseIdx + 1] : null;
const phase = rawPhase && (ALL_PHASES as string[]).includes(rawPhase)
? (rawPhase as CyclePhase)
: null;
if (rawPhase && !phase) {
console.error(`Unknown phase "${rawPhase}". Valid: ${ALL_PHASES.join(', ')}`);
process.exit(1);
}
const dirIdx = args.indexOf('--dir');
const dir = dirIdx !== -1 ? args[dirIdx + 1] : null;
return {
json: args.includes('--json'),
dryRun: args.includes('--dry-run'),
pull: args.includes('--pull'),
phase,
dir,
help: args.includes('--help') || args.includes('-h'),
};
}
/**
* Resolve the brain directory without the `findRepoRoot` footgun.
*
* Prior dream.ts walked up 10 levels of cwd looking for `.git` and would
* happily run lint + sync against an unrelated git repo the user happened
* to be cd'd into. This resolver only trusts two sources:
* 1. An explicit --dir argument.
* 2. The `sync.repo_path` config key set by `gbrain init` (engine-backed).
*
* If neither is available, we error out instead of guessing.
*/
async function resolveBrainDir(
engine: BrainEngine | null,
explicit: string | null,
): Promise<string> {
if (explicit) {
if (!existsSync(explicit)) {
console.error(`--dir path does not exist: ${explicit}`);
process.exit(1);
}
return explicit;
}
if (engine) {
const configured = await engine.getConfig('sync.repo_path');
if (configured && existsSync(configured)) {
return configured;
}
}
console.error(
'No brain directory found. Pass --dir <path> or configure one via `gbrain init`.',
);
process.exit(1);
}
function printHelp() {
console.log(`Usage: gbrain dream [options]
Run one brain maintenance cycle: lint, backlinks, orphan sweep, sync,
extract, and embed. Designed for cron (exits when done).
Options:
--dry-run Preview all fixes without writing (fs or DB)
--json Emit the CycleReport as JSON (agent-readable)
--phase <name> Run a single phase: ${ALL_PHASES.join(' | ')}
--pull git pull the brain repo before syncing (default: no pull)
--dir <path> Brain directory (default: configured brain)
--help, -h Show this help
Examples:
gbrain dream
gbrain dream --dry-run --json
gbrain dream --phase lint
0 2 * * * gbrain dream --json # nightly via cron
Related:
gbrain autopilot --install # continuous maintenance as a daemon
gbrain autopilot # same maintenance cycle, scheduled
`);
}
// ─── Human-friendly report printing ────────────────────────────────
function printHuman(report: CycleReport) {
if (report.status === 'skipped') {
if (report.reason === 'cycle_already_running') {
console.log(`Skipped: another cycle is already running. (locked)`);
} else if (report.reason === 'no_database') {
console.log(`Skipped: no database available.`);
} else {
console.log(`Skipped: ${report.reason ?? 'unknown reason'}.`);
}
return;
}
if (report.status === 'clean') {
console.log(
`Brain is healthy. ${report.phases.length} phase(s) checked in ${(report.duration_ms / 1000).toFixed(1)}s.`,
);
return;
}
console.log(`Dream cycle (${report.status}) in ${(report.duration_ms / 1000).toFixed(1)}s:`);
for (const p of report.phases) {
const icon =
p.status === 'ok' ? '✓' :
p.status === 'warn' ? '!' :
p.status === 'skipped' ? '-' : '✗';
const line = ` ${icon} ${p.phase.padEnd(10)} ${p.summary}`;
console.log(line);
if (p.error) {
const hint = p.error.hint ? ` (${p.error.hint})` : '';
console.log(` [${p.error.class}/${p.error.code}] ${p.error.message}${hint}`);
}
}
const t = report.totals;
const hasTotals =
t.lint_fixes > 0 || t.backlinks_added > 0 || t.pages_synced > 0 ||
t.pages_extracted > 0 || t.pages_embedded > 0 || t.orphans_found > 0;
if (hasTotals) {
console.log(
` totals: lint=${t.lint_fixes} backlinks=${t.backlinks_added} synced=${t.pages_synced} extracted=${t.pages_extracted} embedded=${t.pages_embedded} orphans=${t.orphans_found}`,
);
}
}
// ─── CLI entry ─────────────────────────────────────────────────────
export async function runDream(engine: BrainEngine | null, args: string[]): Promise<CycleReport | void> {
const opts = parseArgs(args);
if (opts.help) {
printHelp();
return;
}
const brainDir = await resolveBrainDir(engine, opts.dir);
const phases: CyclePhase[] | undefined = opts.phase ? [opts.phase] : undefined;
const report = await runCycle(engine, {
brainDir,
dryRun: opts.dryRun,
pull: opts.pull,
phases,
});
if (opts.json) {
console.log(JSON.stringify(report, null, 2));
} else {
printHuman(report);
}
// Exit non-zero when the cycle failed overall (helps cron spot real problems).
// 'partial' is not a failure — it means some phase warned but the cycle ran.
if (report.status === 'failed') {
process.exit(1);
}
return report;
}
+111 -23
View File
@@ -14,6 +14,12 @@ export interface EmbedOpts {
slugs?: string[];
/** Embed a single page. */
slug?: string;
/**
* Dry run: enumerate what WOULD be embedded (stale chunk counts)
* without calling the embedding model or writing to the engine.
* Safe to call with no API key. Used by runCycle's dryRun propagation.
*/
dryRun?: boolean;
/**
* Optional progress callback. Called after each page. CLI wrappers
* supply a reporter.tick()-backed implementation; Minion handlers
@@ -23,48 +29,86 @@ export interface EmbedOpts {
onProgress?: (done: number, total: number, embedded: number) => void;
}
/**
* Structured result from a library-level embed run.
*
* In dryRun mode, `embedded = 0` and `would_embed` holds the count of
* stale chunks that WOULD have been sent to the embedding model. In
* non-dryRun mode, `embedded` holds the real count and `would_embed = 0`.
* `skipped` counts chunks that already had embeddings (nothing to do).
*/
export interface EmbedResult {
/** Chunks newly embedded in this run (0 in dryRun). */
embedded: number;
/** Chunks with pre-existing embeddings, skipped. */
skipped: number;
/** Chunks that would be embedded if not for dryRun (0 in non-dryRun). */
would_embed: number;
/** Total chunks considered across all processed pages. */
total_chunks: number;
/** Number of pages processed (whether or not they had stale chunks). */
pages_processed: number;
/** True if this run was a dry-run. */
dryRun: boolean;
}
/**
* Library-level embed. Throws on validation errors; per-page embed failures
* are logged to stderr but do not throw (matches the existing CLI semantics
* for batch runs). Safe to call from Minions handlers no process.exit.
*
* Returns EmbedResult with accurate counts so callers (runCycle, sync
* auto-embed step) can report embeddings in their own structured output.
*/
export async function runEmbedCore(engine: BrainEngine, opts: EmbedOpts): Promise<void> {
export async function runEmbedCore(engine: BrainEngine, opts: EmbedOpts): Promise<EmbedResult> {
const result: EmbedResult = {
embedded: 0,
skipped: 0,
would_embed: 0,
total_chunks: 0,
pages_processed: 0,
dryRun: !!opts.dryRun,
};
if (opts.slugs && opts.slugs.length > 0) {
for (const s of opts.slugs) {
try { await embedPage(engine, s); } catch (e: unknown) {
try {
await embedPage(engine, s, !!opts.dryRun, result);
} catch (e: unknown) {
console.error(` Error embedding ${s}: ${e instanceof Error ? e.message : e}`);
}
}
return;
return result;
}
if (opts.all || opts.stale) {
await embedAll(engine, !!opts.stale, opts.onProgress);
return;
await embedAll(engine, !!opts.stale, !!opts.dryRun, result, opts.onProgress);
return result;
}
if (opts.slug) {
await embedPage(engine, opts.slug);
return;
await embedPage(engine, opts.slug, !!opts.dryRun, result);
return result;
}
throw new Error('No embed target specified. Pass { slug }, { slugs }, { all }, or { stale }.');
}
export async function runEmbed(engine: BrainEngine, args: string[]) {
export async function runEmbed(engine: BrainEngine, args: string[]): Promise<EmbedResult | undefined> {
const slugsIdx = args.indexOf('--slugs');
const all = args.includes('--all');
const stale = args.includes('--stale');
const dryRun = args.includes('--dry-run');
let opts: EmbedOpts;
if (slugsIdx >= 0) {
opts = { slugs: args.slice(slugsIdx + 1).filter(a => !a.startsWith('--')) };
opts = { slugs: args.slice(slugsIdx + 1).filter(a => !a.startsWith('--')), dryRun };
} else if (all || stale) {
opts = { all, stale };
opts = { all, stale, dryRun };
} else {
const slug = args.find(a => !a.startsWith('--'));
if (!slug) {
console.error('Usage: gbrain embed [<slug>|--all|--stale|--slugs s1 s2 ...]');
console.error('Usage: gbrain embed [<slug>|--all|--stale|--slugs s1 s2 ...] [--dry-run]');
process.exit(1);
}
opts = { slug };
opts = { slug, dryRun };
}
// CLI path: wire a reporter so --progress-json / --quiet / TTY rendering
@@ -81,8 +125,9 @@ export async function runEmbed(engine: BrainEngine, args: string[]) {
};
try {
await runEmbedCore(engine, opts);
const result = await runEmbedCore(engine, opts);
if (progressStarted) progress.finish();
return result;
} catch (e) {
if (progressStarted) progress.finish();
console.error(e instanceof Error ? e.message : String(e));
@@ -90,16 +135,22 @@ export async function runEmbed(engine: BrainEngine, args: string[]) {
}
}
async function embedPage(engine: BrainEngine, slug: string) {
async function embedPage(
engine: BrainEngine,
slug: string,
dryRun: boolean,
result: EmbedResult,
) {
const page = await engine.getPage(slug);
if (!page) {
throw new Error(`Page not found: ${slug}`);
}
// Get existing chunks or create new ones
// Get existing chunks or create new ones.
// In dryRun, we still chunk the text locally to count what WOULD be
// embedded — but we never write chunks or call the embedding model.
let chunks = await engine.getChunks(slug);
if (chunks.length === 0) {
// Create chunks first
const inputs: ChunkInput[] = [];
if (page.compiled_truth.trim()) {
for (const c of chunkText(page.compiled_truth)) {
@@ -111,6 +162,15 @@ async function embedPage(engine: BrainEngine, slug: string) {
inputs.push({ chunk_index: inputs.length, chunk_text: c.text, chunk_source: 'timeline' });
}
}
if (dryRun) {
// Count what chunking WOULD produce, without writing.
result.total_chunks += inputs.length;
result.would_embed += inputs.length;
result.pages_processed++;
return;
}
if (inputs.length > 0) {
await engine.upsertChunks(slug, inputs);
chunks = await engine.getChunks(slug);
@@ -119,8 +179,18 @@ async function embedPage(engine: BrainEngine, slug: string) {
// Embed chunks without embeddings
const toEmbed = chunks.filter(c => !c.embedded_at);
result.total_chunks += chunks.length;
result.skipped += chunks.length - toEmbed.length;
if (toEmbed.length === 0) {
console.log(`${slug}: all ${chunks.length} chunks already embedded`);
result.pages_processed++;
return;
}
if (dryRun) {
result.would_embed += toEmbed.length;
result.pages_processed++;
return;
}
@@ -138,17 +208,19 @@ async function embedPage(engine: BrainEngine, slug: string) {
}));
await engine.upsertChunks(slug, updated);
result.embedded += toEmbed.length;
result.pages_processed++;
console.log(`${slug}: embedded ${toEmbed.length} chunks`);
}
async function embedAll(
engine: BrainEngine,
staleOnly: boolean,
dryRun: boolean,
result: EmbedResult,
onProgress?: (done: number, total: number, embedded: number) => void,
) {
const pages = await engine.listPages({ limit: 100000 });
let total = 0;
let embedded = 0;
let processed = 0;
// Concurrency limit for parallel page embedding.
@@ -167,9 +239,21 @@ async function embedAll(
? chunks.filter(c => !c.embedded_at)
: chunks;
result.total_chunks += chunks.length;
result.skipped += chunks.length - toEmbed.length;
if (toEmbed.length === 0) {
processed++;
onProgress?.(processed, pages.length, embedded);
result.pages_processed++;
onProgress?.(processed, pages.length, result.embedded);
return;
}
if (dryRun) {
result.would_embed += toEmbed.length;
processed++;
result.pages_processed++;
onProgress?.(processed, pages.length, result.embedded);
return;
}
@@ -189,14 +273,14 @@ async function embedAll(
token_count: c.token_count || Math.ceil(c.chunk_text.length / 4),
}));
await engine.upsertChunks(page.slug, updated);
embedded += toEmbed.length;
result.embedded += toEmbed.length;
} catch (e: unknown) {
console.error(`\n Error embedding ${page.slug}: ${e instanceof Error ? e.message : e}`);
}
total += toEmbed.length;
processed++;
onProgress?.(processed, pages.length, embedded);
result.pages_processed++;
onProgress?.(processed, pages.length, result.embedded);
}
// Sliding worker pool: N workers share a queue and each pulls the
@@ -216,5 +300,9 @@ async function embedAll(
await Promise.all(Array.from({ length: numWorkers }, () => worker()));
// Stdout summary preserved for scripts/tests that grep for counts.
console.log(`Embedded ${embedded} chunks across ${pages.length} pages`);
if (dryRun) {
console.log(`[dry-run] Would embed ${result.would_embed} chunks across ${pages.length} pages`);
} else {
console.log(`Embedded ${result.embedded} chunks across ${pages.length} pages`);
}
}
+1 -1
View File
@@ -421,7 +421,7 @@ async function mirrorFiles(args: string[]) {
// Write .supabase marker
const marker = stringify({
synced_at: new Date().toISOString(),
bucket: config.storage.bucket || 'brain-files',
bucket: (config.storage as { bucket?: string })?.bucket || 'brain-files',
prefix: basename(dir) + '/',
file_count: uploaded,
});
+3 -2
View File
@@ -38,12 +38,13 @@ export async function runImport(engine: BrainEngine, args: string[], opts: { com
// Find dir: first non-flag arg that isn't a value for --workers
const flagValues = new Set<number>();
if (workersIdx !== -1) flagValues.add(workersIdx + 1);
const dir = args.find((a, i) => !a.startsWith('--') && !flagValues.has(i));
const dirArg = args.find((a, i) => !a.startsWith('--') && !flagValues.has(i));
if (!dir) {
if (!dirArg) {
console.error('Usage: gbrain import <dir> [--no-embed] [--workers N] [--fresh] [--json]');
process.exit(1);
}
const dir: string = dirArg; // narrowed; survives closure capture
// Collect all .md files
const allFiles = collectMarkdownFiles(dir);
+60 -46
View File
@@ -569,58 +569,39 @@ export async function registerBuiltinHandlers(worker: MinionWorker, engine: Brai
return await runBacklinksCore({ action, dir, dryRun: !!job.data.dryRun });
});
// The killer handler. Autopilot submits ONE `autopilot-cycle` per cycle
// (idempotency_key on cycle slot) instead of a 4-job parent-child DAG,
// because Minions' parent/child is NOT a depends_on primitive (Codex
// H3/H4). Each step is wrapped in its own try/catch; the handler returns
// `{ partial: true, failed_steps: [...] }` when any step fails. It does
// NOT throw on partial failure — that would cause the Minion to retry,
// and an intermittent extract bug would block every future cycle.
// Autopilot-cycle handler: delegates to runCycle. Shares the exact same
// phase set and ordering as `gbrain dream` and autopilot's inline path —
// one source of truth for what the brain does overnight.
//
// Yields the event loop between phases so the worker's lock-renewal
// timer (src/core/minions/worker.ts) can fire. Without this the v0.14
// stall-death regression returns: long CPU-bound phases starve the
// renewal callback and the stalled-sweeper kills the job.
//
// Phase failures surface as report.status='partial' (via runCycle's
// derivation); the handler returns { partial, status, report } so
// `gbrain jobs get <id>` shows the full structured report. Does NOT
// throw on partial: a flaky phase must not block every future cycle.
worker.register('autopilot-cycle', async (job) => {
const { performSync } = await import('./sync.ts');
const { runExtractCore } = await import('./extract.ts');
const { runEmbedCore } = await import('./embed.ts');
const { runBacklinksCore } = await import('./backlinks.ts');
const { runCycle } = await import('../core/cycle.ts');
const repoPath = typeof job.data.repoPath === 'string'
? job.data.repoPath
: (await engine.getConfig('sync.repo_path')) ?? '.';
const steps: Record<string, unknown> = {};
const failed: string[] = [];
const report = await runCycle(engine, {
brainDir: repoPath,
pull: true, // autopilot daemon opts into git pull
yieldBetweenPhases: async () => {
// Yield to the event loop so worker lock-renewal can fire.
await new Promise<void>(r => setImmediate(r));
},
});
// Bug 8 — Between phases, yield to the event loop. The worker's lock
// renewal runs on a timer (src/core/minions/worker.ts); without a
// periodic yield, long CPU-bound phases starve the renewal callback
// and the job gets killed by the stalled-sweeper. A single
// `await new Promise(r => setImmediate(r))` gives the timer a chance
// to fire. The per-phase body is async+await already, so each phase
// internally yields on its own I/O boundaries — this is a belt for
// the gap between phases.
//
// Follow-up (deferred to v0.15): thread ctx.signal / ctx.shutdownSignal
// through each core fn so mid-phase cancellation works on huge brains.
const yieldToLoop = () => new Promise<void>(r => setImmediate(r));
try { steps.sync = await performSync(engine, { repoPath, noEmbed: true }); }
catch (e) { steps.sync = { error: e instanceof Error ? e.message : String(e) }; failed.push('sync'); }
await yieldToLoop();
try { steps.extract = await runExtractCore(engine, { mode: 'all', dir: repoPath }); }
catch (e) { steps.extract = { error: e instanceof Error ? e.message : String(e) }; failed.push('extract'); }
await yieldToLoop();
try { await runEmbedCore(engine, { stale: true }); steps.embed = { embedded: true }; }
catch (e) { steps.embed = { error: e instanceof Error ? e.message : String(e) }; failed.push('embed'); }
await yieldToLoop();
try { steps.backlinks = await runBacklinksCore({ action: 'fix', dir: repoPath }); }
catch (e) { steps.backlinks = { error: e instanceof Error ? e.message : String(e) }; failed.push('backlinks'); }
if (failed.length > 0) {
return { partial: true, failed_steps: failed, steps };
}
return { partial: false, steps };
return {
partial: report.status === 'partial' || report.status === 'failed',
status: report.status,
report,
};
});
// Shell handler: registered ONLY when GBRAIN_ALLOW_SHELL_JOBS=1 is set on the
@@ -634,4 +615,37 @@ export async function registerBuiltinHandlers(worker: MinionWorker, engine: Brai
} else {
process.stderr.write('[minion worker] shell handler disabled (set GBRAIN_ALLOW_SHELL_JOBS=1 to enable)\n');
}
// v0.15 subagent handlers: always-on. Unlike shell (which needs an env
// flag because of RCE surface), subagent only calls the Anthropic API
// with the operator's own ANTHROPIC_API_KEY — no key, the SDK call
// fails immediately. Who-can-submit is already gated by
// PROTECTED_JOB_NAMES + TrustedSubmitOpts (MCP can't submit subagent
// jobs; only the CLI path with allowProtectedSubmit can). No separate
// cost-ceremony env flag needed.
const { makeSubagentHandler } = await import('../core/minions/handlers/subagent.ts');
const { subagentAggregatorHandler } = await import('../core/minions/handlers/subagent-aggregator.ts');
worker.register('subagent', makeSubagentHandler({ engine }));
worker.register('subagent_aggregator', subagentAggregatorHandler);
process.stderr.write('[minion worker] subagent handlers enabled\n');
// Plugin discovery — one line per discovered plugin (mirrors the
// openclaw-seam startup line convention from v0.11+). Loaded
// unconditionally; empty GBRAIN_PLUGIN_PATH is a no-op.
try {
const { loadPluginsFromEnv } = await import('../core/minions/plugin-loader.ts');
const { BRAIN_TOOL_ALLOWLIST } = await import('../core/minions/tools/brain-allowlist.ts');
const validNames = new Set<string>();
for (const n of BRAIN_TOOL_ALLOWLIST) validNames.add(`brain_${n}`);
const loaded = loadPluginsFromEnv({ validAgentToolNames: validNames });
for (const w of loaded.warnings) process.stderr.write(w + '\n');
for (const p of loaded.plugins) {
process.stderr.write(
`[plugin-loader] loaded '${p.manifest.name}' v${p.manifest.version} (${p.subagents.length} subagents)\n`,
);
}
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
process.stderr.write(`[plugin-loader] discovery failed: ${msg}\n`);
}
}
+2
View File
@@ -17,6 +17,7 @@ import { v0_12_2 } from './v0_12_2.ts';
import { v0_13_0 } from './v0_13_0.ts';
import { v0_13_1 } from './v0_13_1.ts';
import { v0_14_0 } from './v0_14_0.ts';
import { v0_16_0 } from './v0_16_0.ts';
export const migrations: Migration[] = [
v0_11_0,
@@ -25,6 +26,7 @@ export const migrations: Migration[] = [
v0_13_0,
v0_13_1,
v0_14_0,
v0_16_0,
];
/** Look up a migration by exact version string. */
+2 -3
View File
@@ -217,10 +217,9 @@ async function orchestrator(opts: OrchestratorOpts): Promise<OrchestratorResult>
phases.push(e);
// F. Record
// a.status was narrowed to 'skipped' | 'complete' by the early return above.
const overallStatus: 'complete' | 'partial' | 'failed' =
a.status === 'failed' ? 'failed' :
phases.some(p => p.status === 'failed') ? 'partial' :
'complete';
phases.some(p => p.status === 'failed') ? 'partial' : 'complete';
return finalizeResult(phases, overallStatus);
}
+2 -3
View File
@@ -105,10 +105,9 @@ async function orchestrator(opts: OrchestratorOpts): Promise<OrchestratorResult>
const c = phaseCVerify(opts);
phases.push(c);
// a.status and b.status were narrowed to 'skipped' | 'complete' by early returns above.
const overallStatus: 'complete' | 'partial' | 'failed' =
a.status === 'failed' || b.status === 'failed' ? 'failed' :
c.status === 'failed' ? 'partial' :
'complete';
c.status === 'failed' ? 'partial' : 'complete';
return finalizeResult(phases, overallStatus);
}
+2 -3
View File
@@ -129,10 +129,9 @@ async function orchestrator(opts: OrchestratorOpts): Promise<OrchestratorResult>
const c = phaseCVerify(opts);
phases.push(c);
// a.status and b.status were narrowed to 'skipped' | 'complete' by early returns above.
const overallStatus: 'complete' | 'partial' | 'failed' =
a.status === 'failed' || b.status === 'failed' ? 'failed' :
c.status === 'failed' ? 'partial' :
'complete';
c.status === 'failed' ? 'partial' : 'complete';
return finalizeResult(phases, overallStatus);
}
+140
View File
@@ -0,0 +1,140 @@
/**
* v0.16.0 migration orchestrator Subagent runtime schema.
*
* Adds three tables for durable LLM agent loops:
* - subagent_messages Anthropic message-block persistence
* - subagent_tool_executions Two-phase tool ledger (pending/complete/failed)
* - subagent_rate_leases Lease-based concurrency cap
*
* All DDL is `CREATE TABLE IF NOT EXISTS` and ships in src/schema.sql +
* src/core/pglite-schema.ts (both Postgres and PGLite fresh-install paths).
* This orchestrator's job is therefore only to VERIFY the tables exist after
* `gbrain init --migrate-only` has run, so an upgrade that somehow skipped
* the schema step fails loudly instead of silently.
*
* Phases (all idempotent):
* A. Schema gbrain init --migrate-only (creates tables via SCHEMA_SQL).
* B. Verify confirm all three tables exist.
* C. Record append completed.jsonl.
*/
import { execSync } from 'child_process';
import type { Migration, OrchestratorOpts, OrchestratorResult, OrchestratorPhaseResult } from './types.ts';
import { appendCompletedMigration } from '../../core/preferences.ts';
import { loadConfig, toEngineConfig } from '../../core/config.ts';
import { createEngine } from '../../core/engine-factory.ts';
const REQUIRED_TABLES = ['subagent_messages', 'subagent_tool_executions', 'subagent_rate_leases'] as const;
// ── Phase A — Schema ────────────────────────────────────────
function phaseASchema(opts: OrchestratorOpts): OrchestratorPhaseResult {
if (opts.dryRun) return { name: 'schema', status: 'skipped', detail: 'dry-run' };
try {
execSync('gbrain init --migrate-only', { stdio: 'inherit', timeout: 60_000, env: process.env });
return { name: 'schema', status: 'complete' };
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
return { name: 'schema', status: 'failed', detail: msg };
}
}
// ── Phase B — Verify tables exist ───────────────────────────
async function phaseBVerify(opts: OrchestratorOpts): Promise<OrchestratorPhaseResult> {
if (opts.dryRun) return { name: 'verify', status: 'skipped', detail: 'dry-run' };
try {
const config = loadConfig();
if (!config) {
return { name: 'verify', status: 'skipped', detail: 'no brain configured' };
}
const engine = await createEngine(toEngineConfig(config));
await engine.connect(toEngineConfig(config));
try {
const rows = await engine.executeRaw<{ table_name: string }>(
`SELECT table_name FROM information_schema.tables
WHERE table_schema = current_schema()
AND table_name IN ('subagent_messages','subagent_tool_executions','subagent_rate_leases')`,
);
const found = new Set(rows.map(r => r.table_name));
const missing = REQUIRED_TABLES.filter(t => !found.has(t));
if (missing.length > 0) {
return {
name: 'verify',
status: 'failed',
detail: `missing tables: ${missing.join(', ')}`,
};
}
return { name: 'verify', status: 'complete', detail: `${REQUIRED_TABLES.length} tables present` };
} finally {
try { await engine.disconnect(); } catch {}
}
} catch (e) {
return {
name: 'verify',
status: 'failed',
detail: e instanceof Error ? e.message : String(e),
};
}
}
// ── Orchestrator ────────────────────────────────────────────
async function orchestrator(opts: OrchestratorOpts): Promise<OrchestratorResult> {
console.log('');
console.log('=== v0.16.0 — Subagent runtime schema ===');
if (opts.dryRun) console.log(' (dry-run; no side effects)');
console.log('');
const phases: OrchestratorPhaseResult[] = [];
const a = phaseASchema(opts);
phases.push(a);
if (a.status === 'failed') return finalize(phases, 'failed');
const b = await phaseBVerify(opts);
phases.push(b);
// a.status was narrowed to 'skipped' | 'complete' by the early return above.
const status: 'complete' | 'partial' | 'failed' =
b.status === 'failed' ? 'partial' : 'complete';
return finalize(phases, status);
}
function finalize(phases: OrchestratorPhaseResult[], status: 'complete' | 'partial' | 'failed'): OrchestratorResult {
if (status !== 'failed') {
try {
appendCompletedMigration({
version: '0.16.0',
completed_at: new Date().toISOString(),
status: status as 'complete' | 'partial',
phases: phases.map(p => ({ name: p.name, status: p.status })),
});
} catch {
// Recording is best-effort.
}
}
return { version: '0.16.0', status, phases };
}
export const v0_16_0: Migration = {
version: '0.16.0',
featurePitch: {
headline: 'Durable LLM agents land in the brain — survive crashes, sleeps, and worker restarts.',
description:
'v0.16.0 adds the subagent runtime: run long-running, fan-out Anthropic LLM loops ' +
'as first-class Minion jobs. Crash-resumable turn persistence, two-phase tool ledger, ' +
'lease-based rate limit, parent-child fan-out with aggregation. Entry points: `gbrain ' +
'agent run` and `gbrain agent logs`. See docs/guides/plugin-authors.md for shipping ' +
'custom subagent defs from a host repo (your OpenClaw etc.).',
},
orchestrator,
};
/** Exported for unit tests. */
export const __testing = {
phaseASchema,
phaseBVerify,
REQUIRED_TABLES,
};
+22 -25
View File
@@ -13,7 +13,6 @@
*/
import type { BrainEngine } from '../core/engine.ts';
import * as db from '../core/db.ts';
import { createProgress, startHeartbeat } from '../core/progress.ts';
import { getCliOptions, cliOptsToProgressOptions } from '../core/cli-options.ts';
@@ -99,30 +98,30 @@ export function deriveDomain(frontmatterDomain: string | null | undefined, slug:
// --- Core query ---
/**
* Find pages with no inbound links.
* Returns raw rows from the DB (all pages regardless of filter).
* Find pages with no inbound links via the engine's built-in helper.
* Returns raw rows (all pages regardless of filter).
*
* As of v0.17: takes an engine argument. Composes with runCycle which
* passes an explicit engine. No more db.getConnection() global fixes
* the PGLite-vs-Postgres + test-fixture coupling codex flagged.
*/
export async function queryOrphanPages(): Promise<{ slug: string; title: string; domain: string | null }[]> {
const sql = db.getConnection();
const rows = await sql`
SELECT
p.slug,
COALESCE(p.title, p.slug) AS title,
p.frontmatter->>'domain' AS domain
FROM pages p
WHERE NOT EXISTS (
SELECT 1 FROM links l WHERE l.to_page_id = p.id
)
ORDER BY p.slug
`;
return rows as { slug: string; title: string; domain: string | null }[];
export async function queryOrphanPages(
engine: BrainEngine,
): Promise<{ slug: string; title: string; domain: string | null }[]> {
return engine.findOrphanPages();
}
/**
* Find orphan pages, with optional pseudo-page filtering.
* Returns structured OrphanResult with totals.
*
* As of v0.17: `engine` is required. See queryOrphanPages for rationale.
*/
export async function findOrphans(includePseudo: boolean = false): Promise<OrphanResult> {
export async function findOrphans(
engine: BrainEngine,
opts: { includePseudo?: boolean } = {},
): Promise<OrphanResult> {
const includePseudo = !!opts.includePseudo;
// The NOT EXISTS anti-join over pages × links can take seconds on 50K-page
// brains. Heartbeat every second so agents see the scan is alive. Keyset
// pagination was considered and rejected: without an index on
@@ -134,12 +133,10 @@ export async function findOrphans(includePseudo: boolean = false): Promise<Orpha
let allOrphans: { slug: string; title: string; domain: string | null }[];
let total: number;
try {
allOrphans = await queryOrphanPages();
allOrphans = await engine.findOrphanPages();
// Count total pages in DB for the summary line
const sql = db.getConnection();
const [{ count: totalPagesCount }] = await sql`SELECT count(*)::int AS count FROM pages`;
total = Number(totalPagesCount);
const stats = await engine.getStats();
total = stats.page_count;
} finally {
stopHb();
progress.finish();
@@ -206,7 +203,7 @@ export function formatOrphansText(result: OrphanResult): string {
// --- CLI entry point ---
export async function runOrphans(_engine: BrainEngine, args: string[]) {
export async function runOrphans(engine: BrainEngine, args: string[]) {
const json = args.includes('--json');
const count = args.includes('--count');
const includePseudo = args.includes('--include-pseudo');
@@ -228,7 +225,7 @@ Summary line: N orphans out of M linkable pages (K total; K-M excluded)
return;
}
const result = await findOrphans(includePseudo);
const result = await findOrphans(engine, { includePseudo });
if (count) {
console.log(String(result.total_orphans));
+1 -1
View File
@@ -117,7 +117,7 @@ export async function repairJsonb(opts: RepairOpts = { dryRun: false }): Promise
const rows = await sql.unsafe(
`SELECT count(*)::int AS n FROM ${t.table} WHERE jsonb_typeof(${t.column}) = 'string'`,
);
repaired = (rows[0] as { n: number }).n;
repaired = (rows[0] as unknown as { n: number }).n;
} else {
const rows = await sql.unsafe(
`UPDATE ${t.table}
+121
View File
@@ -0,0 +1,121 @@
/**
* CLI: gbrain repos list|add|remove
* Multi-repo management for code + knowledge indexing.
*/
import { resolve } from 'path';
import { existsSync } from 'fs';
import {
loadRepoConfigs,
addRepoConfig,
removeRepoConfig,
normalizeRepoName,
type RepoStrategy,
} from '../core/multi-repo.ts';
export async function handleRepos(args: string[]): Promise<void> {
const sub = args[0];
if (!sub || sub === 'list') {
return reposList();
}
if (sub === 'add') {
return reposAdd(args.slice(1));
}
if (sub === 'remove' || sub === 'rm') {
return reposRemove(args.slice(1));
}
console.error(`Unknown repos subcommand: ${sub}`);
console.error('Usage: gbrain repos [list|add|remove]');
process.exit(1);
}
function reposList(): void {
const repos = loadRepoConfigs();
if (repos.length === 0) {
console.log('No repos configured. Use `gbrain repos add <path>` to add one.');
return;
}
console.log(`${repos.length} repo(s) configured:\n`);
for (const repo of repos) {
const enabled = repo.syncEnabled !== false ? '✓' : '✗';
const includes = repo.include?.length ? ` include=[${repo.include.join(',')}]` : '';
const excludes = repo.exclude?.length ? ` exclude=[${repo.exclude.join(',')}]` : '';
console.log(` ${enabled} ${repo.name} (${repo.strategy}) → ${repo.path}${includes}${excludes}`);
}
}
function reposAdd(args: string[]): void {
if (args.length === 0) {
console.error('Usage: gbrain repos add <path> [--name <name>] [--strategy markdown|code|auto]');
process.exit(1);
}
const repoPath = resolve(args[0]);
if (!existsSync(repoPath)) {
console.error(`Path does not exist: ${repoPath}`);
process.exit(1);
}
let name = normalizeRepoName(repoPath);
let strategy: RepoStrategy = 'auto';
const include: string[] = [];
const exclude: string[] = [];
for (let i = 1; i < args.length; i++) {
if (args[i] === '--name' && args[i + 1]) {
name = args[++i];
} else if (args[i] === '--strategy' && args[i + 1]) {
const s = args[++i];
if (s === 'markdown' || s === 'code' || s === 'auto') {
strategy = s;
} else {
console.error(`Invalid strategy: ${s}. Must be markdown, code, or auto.`);
process.exit(1);
}
} else if (args[i] === '--include' && args[i + 1]) {
include.push(args[++i]);
} else if (args[i] === '--exclude' && args[i + 1]) {
exclude.push(args[++i]);
}
}
try {
const repos = addRepoConfig({
path: repoPath,
name,
strategy,
include: include.length > 0 ? include : undefined,
exclude: exclude.length > 0 ? exclude : undefined,
syncEnabled: true,
});
console.log(`Added repo "${name}" (${strategy}) → ${repoPath}`);
console.log(`${repos.length} repo(s) total. Run \`gbrain sync --all\` to index.`);
} catch (e: unknown) {
console.error(e instanceof Error ? e.message : String(e));
process.exit(1);
}
}
function reposRemove(args: string[]): void {
if (args.length === 0) {
console.error('Usage: gbrain repos remove <name>');
process.exit(1);
}
const name = args[0];
try {
const repos = removeRepoConfig(name);
console.log(`Removed repo "${name}". ${repos.length} repo(s) remaining.`);
} catch (e: unknown) {
console.error(e instanceof Error ? e.message : String(e));
process.exit(1);
}
}
+5 -2
View File
@@ -25,9 +25,12 @@ import { xHandleToTweetResolver } from '../core/resolvers/builtin/x-api/handle-t
* to call from multiple entry points.
*/
export function registerBuiltinResolvers(registry = getDefaultRegistry()): void {
const builtins = [urlReachableResolver, xHandleToTweetResolver] as const;
// Cast each element to the widest shape the registry accepts. The tuple
// element types diverge (different Input/Output generics) so the union
// type would not satisfy registry.register's single-signature parameter.
const builtins = [urlReachableResolver, xHandleToTweetResolver];
for (const r of builtins) {
if (!registry.has(r.id)) registry.register(r);
if (!registry.has(r.id)) registry.register(r as Parameters<typeof registry.register>[0]);
}
}
+1 -1
View File
@@ -2,7 +2,7 @@
* `gbrain skillpack-check` agent-readable health report.
*
* Wraps `gbrain doctor --json` + `gbrain apply-migrations --list` into a
* single JSON blob a host agent (Wintermute's morning-briefing, any
* single JSON blob a host agent (your OpenClaw's morning-briefing, any
* OpenClaw cron) can consume without parsing two subcommands.
*
* Usage:
+94 -11
View File
@@ -24,6 +24,8 @@ export interface SyncResult {
deleted: number;
renamed: number;
chunksCreated: number;
/** Pages re-embedded during this sync's auto-embed step. 0 if --no-embed or skipped. */
embedded: number;
pagesAffected: string[];
failedFiles?: number; // count of parse failures (Bug 9)
}
@@ -39,6 +41,8 @@ export interface SyncOpts {
skipFailed?: boolean;
/** Bug 9 — re-attempt unacknowledged failures explicitly (CLI --retry-failed). */
retryFailed?: boolean;
/** Multi-repo: sync strategy override (markdown, code, auto). */
strategy?: 'markdown' | 'code' | 'auto';
}
function git(repoPath: string, ...args: string[]): string {
@@ -116,6 +120,7 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
toCommit: headCommit,
added: 0, modified: 0, deleted: 0, renamed: 0,
chunksCreated: 0,
embedded: 0,
pagesAffected: [],
};
}
@@ -124,16 +129,17 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
const diffOutput = git(repoPath, 'diff', '--name-status', '-M', `${lastCommit}..${headCommit}`);
const manifest = buildSyncManifest(diffOutput);
// Filter to syncable files
// Filter to syncable files (strategy-aware)
const syncOpts = opts.strategy ? { strategy: opts.strategy } : undefined;
const filtered: SyncManifest = {
added: manifest.added.filter(p => isSyncable(p)),
modified: manifest.modified.filter(p => isSyncable(p)),
deleted: manifest.deleted.filter(p => isSyncable(p)),
renamed: manifest.renamed.filter(r => isSyncable(r.to)),
added: manifest.added.filter(p => isSyncable(p, syncOpts)),
modified: manifest.modified.filter(p => isSyncable(p, syncOpts)),
deleted: manifest.deleted.filter(p => isSyncable(p, syncOpts)),
renamed: manifest.renamed.filter(r => isSyncable(r.to, syncOpts)),
};
// Delete pages that became un-syncable (modified but filtered out)
const unsyncableModified = manifest.modified.filter(p => !isSyncable(p));
const unsyncableModified = manifest.modified.filter(p => !isSyncable(p, syncOpts));
for (const path of unsyncableModified) {
const slug = pathToSlug(path);
try {
@@ -165,6 +171,7 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
deleted: filtered.deleted.length,
renamed: filtered.renamed.length,
chunksCreated: 0,
embedded: 0,
pagesAffected: [],
};
}
@@ -179,6 +186,7 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
toCommit: headCommit,
added: 0, modified: 0, deleted: 0, renamed: 0,
chunksCreated: 0,
embedded: 0,
pagesAffected: [],
};
}
@@ -301,6 +309,7 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
deleted: filtered.deleted.length,
renamed: filtered.renamed.length,
chunksCreated,
embedded: 0,
pagesAffected,
failedFiles: failedFiles.length,
};
@@ -338,10 +347,15 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
}
// Auto-embed (skip for large syncs — embedding calls OpenAI)
let embedded = 0;
if (!noEmbed && pagesAffected.length > 0 && pagesAffected.length <= 100) {
try {
const { runEmbed } = await import('./embed.ts');
await runEmbed(engine, ['--slugs', ...pagesAffected]);
// Before commit 2 lands: runEmbed is void. Best estimate is pagesAffected,
// since runEmbed re-embeds every requested slug. Commit 2 sharpens this
// with EmbedResult.embedded.
embedded = pagesAffected.length;
} catch { /* embedding is best-effort */ }
} else if (noEmbed || totalChanges > 100) {
console.log(`Text imported. Run 'gbrain embed --stale' to generate embeddings.`);
@@ -356,6 +370,7 @@ export async function performSync(engine: BrainEngine, opts: SyncOpts): Promise<
deleted: filtered.deleted.length,
renamed: filtered.renamed.length,
chunksCreated,
embedded,
pagesAffected,
};
}
@@ -366,6 +381,33 @@ async function performFullSync(
headCommit: string,
opts: SyncOpts,
): Promise<SyncResult> {
// Dry-run: walk the repo, count syncable files, return without writing.
// Fixes the silent-write-on-dry-run bug where performFullSync called
// runImport unconditionally regardless of opts.dryRun.
if (opts.dryRun) {
const { collectMarkdownFiles } = await import('./import.ts');
const allFiles = collectMarkdownFiles(repoPath);
const syncableRelPaths = allFiles
.map(abs => relative(repoPath, abs))
.filter(rel => isSyncable(rel));
console.log(
`Full-sync dry run: ${syncableRelPaths.length} file(s) would be imported ` +
`from ${repoPath} @ ${headCommit.slice(0, 8)}.`,
);
return {
status: 'dry_run',
fromCommit: null,
toCommit: headCommit,
added: syncableRelPaths.length,
modified: 0,
deleted: 0,
renamed: 0,
chunksCreated: 0,
embedded: 0,
pagesAffected: [],
};
}
console.log(`Running full import of ${repoPath}...`);
const { runImport } = await import('./import.ts');
const importArgs = [repoPath];
@@ -391,6 +433,7 @@ async function performFullSync(
toCommit: headCommit,
added: 0, modified: 0, deleted: 0, renamed: 0,
chunksCreated: result.chunksCreated,
embedded: 0,
pagesAffected: [],
failedFiles: result.failures.length,
};
@@ -404,11 +447,15 @@ async function performFullSync(
await engine.setConfig('sync.last_run', new Date().toISOString());
await engine.setConfig('sync.repo_path', repoPath);
// Full sync doesn't track pagesAffected, so fall back to embed --stale
// Full sync doesn't track pagesAffected, so fall back to embed --stale.
// Before commit 2: runEmbed is void; use result.imported as best estimate of
// pages touched. Commit 2 sharpens this with real EmbedResult counts.
let embedded = 0;
if (!opts.noEmbed) {
try {
const { runEmbed } = await import('./embed.ts');
await runEmbed(engine, ['--stale']);
embedded = result.imported;
} catch { /* embedding is best-effort */ }
}
@@ -416,8 +463,12 @@ async function performFullSync(
status: 'first_sync',
fromCommit: null,
toCommit: headCommit,
added: 0, modified: 0, deleted: 0, renamed: 0,
chunksCreated: 0,
added: result.imported,
modified: 0,
deleted: 0,
renamed: 0,
chunksCreated: result.chunksCreated,
embedded,
pagesAffected: [],
};
}
@@ -433,8 +484,39 @@ export async function runSync(engine: BrainEngine, args: string[]) {
const noEmbed = args.includes('--no-embed');
const skipFailed = args.includes('--skip-failed');
const retryFailed = args.includes('--retry-failed');
const syncAll = args.includes('--all');
const strategyArg = args.find((a, i) => args[i - 1] === '--strategy') as SyncOpts['strategy'] | undefined;
const opts: SyncOpts = { repoPath, dryRun, full, noPull, noEmbed, skipFailed, retryFailed };
// Multi-repo: --all syncs all configured repos
if (syncAll) {
const { loadRepoConfigs } = await import('../core/multi-repo.ts');
const repos = loadRepoConfigs();
if (repos.length === 0) {
console.log('No repos configured. Use `gbrain repos add <path>` first.');
return;
}
for (const repo of repos) {
if (repo.syncEnabled === false) {
console.log(`Skipping disabled repo: ${repo.name}`);
continue;
}
console.log(`\n--- Syncing repo: ${repo.name} (${repo.strategy}) ---`);
const repoOpts: SyncOpts = {
repoPath: repo.path,
dryRun, full, noPull, noEmbed, skipFailed, retryFailed,
strategy: repo.strategy,
};
try {
const result = await performSync(engine, repoOpts);
printSyncResult(result);
} catch (e: unknown) {
console.error(`Error syncing ${repo.name}: ${e instanceof Error ? e.message : String(e)}`);
}
}
return;
}
const opts: SyncOpts = { repoPath, dryRun, full, noPull, noEmbed, skipFailed, retryFailed, strategy: strategyArg };
// Bug 9 — --retry-failed: before running normal sync, clear acknowledgment
// flags so the sync picks them up as fresh work. The actual re-attempt
@@ -489,10 +571,11 @@ function printSyncResult(result: SyncResult) {
case 'synced':
console.log(`Synced ${result.fromCommit?.slice(0, 8)}..${result.toCommit.slice(0, 8)}:`);
console.log(` +${result.added} added, ~${result.modified} modified, -${result.deleted} deleted, R${result.renamed} renamed`);
console.log(` ${result.chunksCreated} chunks created`);
console.log(` ${result.chunksCreated} chunks created${result.embedded > 0 ? `, ${result.embedded} pages embedded` : ''}`);
break;
case 'first_sync':
console.log(`First sync complete. Checkpoint: ${result.toCommit.slice(0, 8)}`);
console.log(` ${result.added} file(s) imported, ${result.chunksCreated} chunks${result.embedded > 0 ? `, ${result.embedded} pages embedded` : ''}`);
break;
case 'dry_run':
break; // already printed in performSync
+408
View File
@@ -0,0 +1,408 @@
/**
* Code Chunker Tree-Sitter-Based Semantic Code Splitting
*
* Uses web-tree-sitter (WASM) to parse code files into AST, then extracts
* semantic units (functions, classes, types, exports) as chunks.
*
* Each chunk includes a structured header with language, file path, line range,
* and symbol name so embeddings capture both context and code content.
*
* Supports: TypeScript, TSX, JavaScript, Python, Ruby, Go.
* Falls back to recursive text chunker for unsupported languages.
*/
import { chunkText as recursiveChunk } from './recursive.ts';
// Lazy-loaded tree-sitter module (v0.22.x API: Parser is default export)
let Parser: typeof import('web-tree-sitter') | null = null;
async function getParser(): Promise<typeof import('web-tree-sitter')> {
if (!Parser) {
Parser = (await import('web-tree-sitter')).default || await import('web-tree-sitter');
}
return Parser;
}
export type SupportedCodeLanguage = 'typescript' | 'tsx' | 'javascript' | 'python' | 'ruby' | 'go';
export interface CodeChunkMetadata {
symbolName: string | null;
symbolType: string;
filePath: string;
language: SupportedCodeLanguage;
startLine: number;
endLine: number;
}
export interface CodeChunk {
text: string;
index: number;
metadata: CodeChunkMetadata;
}
export interface CodeChunkOptions {
chunkSizeTokens?: number;
largeChunkThresholdTokens?: number;
fallbackChunkSizeWords?: number;
fallbackOverlapWords?: number;
}
const GRAMMAR_FILES: Record<SupportedCodeLanguage, string> = {
typescript: 'tree-sitter-typescript.wasm',
tsx: 'tree-sitter-tsx.wasm',
javascript: 'tree-sitter-javascript.wasm',
python: 'tree-sitter-python.wasm',
ruby: 'tree-sitter-ruby.wasm',
go: 'tree-sitter-go.wasm',
};
const TOP_LEVEL_TYPES: Record<SupportedCodeLanguage, Set<string>> = {
typescript: new Set([
'function_declaration',
'class_declaration',
'abstract_class_declaration',
'interface_declaration',
'type_alias_declaration',
'enum_declaration',
'lexical_declaration',
'variable_declaration',
'export_statement',
]),
tsx: new Set([
'function_declaration',
'class_declaration',
'interface_declaration',
'type_alias_declaration',
'enum_declaration',
'lexical_declaration',
'variable_declaration',
'export_statement',
]),
javascript: new Set([
'function_declaration',
'class_declaration',
'lexical_declaration',
'variable_declaration',
'export_statement',
]),
python: new Set([
'function_definition',
'class_definition',
'import_statement',
'import_from_statement',
'assignment',
]),
ruby: new Set([
'class',
'module',
'method',
'singleton_method',
'assignment',
]),
go: new Set([
'function_declaration',
'method_declaration',
'type_declaration',
'const_declaration',
'var_declaration',
'import_declaration',
]),
};
const BODY_NODE_TYPES = new Set([
'statement_block',
'block',
'class_body',
'module_body',
'body_statement',
'body',
]);
let initDone = false;
let initPromise: Promise<void> | null = null;
const languageCache = new Map<SupportedCodeLanguage, any>();
// ---------- Public API ----------
export function detectCodeLanguage(filePath: string): SupportedCodeLanguage | null {
const lower = filePath.toLowerCase();
if (lower.endsWith('.tsx')) return 'tsx';
if (lower.endsWith('.ts')) return 'typescript';
if (lower.endsWith('.js') || lower.endsWith('.jsx') || lower.endsWith('.mjs') || lower.endsWith('.cjs')) return 'javascript';
if (lower.endsWith('.py')) return 'python';
if (lower.endsWith('.rb')) return 'ruby';
if (lower.endsWith('.go')) return 'go';
return null;
}
export async function chunkCodeText(
source: string,
filePath: string,
opts: CodeChunkOptions = {},
): Promise<CodeChunk[]> {
const language = detectCodeLanguage(filePath);
if (!language) {
return fallbackChunks(source, filePath, 'javascript', opts);
}
if (!source.trim()) return [];
const largeThreshold = opts.largeChunkThresholdTokens ?? 1000;
const chunkTarget = opts.chunkSizeTokens ?? 300;
try {
await ensureInit();
const P = await getParser();
const parser = new (P as any)();
const grammar = await loadLanguage(language);
parser.setLanguage(grammar);
const tree = parser.parse(source);
if (!tree) {
parser.delete();
return fallbackChunks(source, filePath, language, opts);
}
const root = tree.rootNode;
const topLevelTypes = TOP_LEVEL_TYPES[language];
const semanticNodes = root.namedChildren.filter((n: any) => topLevelTypes.has(n.type));
if (semanticNodes.length === 0) {
tree.delete();
parser.delete();
return fallbackChunks(source, filePath, language, opts);
}
const chunks: CodeChunk[] = [];
for (const node of semanticNodes) {
const symbolName = extractSymbolName(node);
const symbolType = normalizeSymbolType(node.type);
const nodeText = source.slice(node.startIndex, node.endIndex).trim();
if (!nodeText) continue;
if (estimateTokens(nodeText) <= largeThreshold) {
chunks.push(buildChunk({
body: nodeText, filePath, language, symbolName, symbolType,
startLine: node.startPosition.row + 1,
endLine: node.endPosition.row + 1,
index: chunks.length,
}));
continue;
}
// Split very large nodes at nested block boundaries
const subRanges = splitLargeNode(node, source, chunkTarget);
if (subRanges.length === 0) {
chunks.push(buildChunk({
body: nodeText, filePath, language, symbolName, symbolType,
startLine: node.startPosition.row + 1,
endLine: node.endPosition.row + 1,
index: chunks.length,
}));
continue;
}
for (const range of subRanges) {
const body = source.slice(range.startIndex, range.endIndex).trim();
if (!body) continue;
chunks.push(buildChunk({
body, filePath, language, symbolName, symbolType,
startLine: range.startLine, endLine: range.endLine,
index: chunks.length,
}));
}
}
tree.delete();
parser.delete();
return chunks.length > 0 ? chunks : fallbackChunks(source, filePath, language, opts);
} catch {
return fallbackChunks(source, filePath, language, opts);
}
}
// ---------- Internals ----------
function fallbackChunks(
source: string,
filePath: string,
language: SupportedCodeLanguage,
opts: CodeChunkOptions,
): CodeChunk[] {
const size = opts.fallbackChunkSizeWords ?? 300;
const overlap = opts.fallbackOverlapWords ?? 50;
return recursiveChunk(source, { chunkSize: size, chunkOverlap: overlap }).map((chunk, index) =>
buildChunk({
body: chunk.text, filePath, language,
symbolName: null, symbolType: 'module',
startLine: 1, endLine: countLines(chunk.text),
index,
}),
);
}
function buildChunk(input: {
body: string;
filePath: string;
language: SupportedCodeLanguage;
symbolName: string | null;
symbolType: string;
startLine: number;
endLine: number;
index: number;
}): CodeChunk {
const symbol = input.symbolName ? `${input.symbolType} ${input.symbolName}` : input.symbolType;
const header = `[${displayLang(input.language)}] ${input.filePath}:${input.startLine}-${input.endLine} ${symbol}`;
return {
index: input.index,
text: `${header}\n\n${input.body}`,
metadata: {
symbolName: input.symbolName,
symbolType: input.symbolType,
filePath: input.filePath,
language: input.language,
startLine: input.startLine,
endLine: input.endLine,
},
};
}
interface SplitRange {
startIndex: number;
endIndex: number;
startLine: number;
endLine: number;
}
function splitLargeNode(node: any, source: string, chunkTarget: number): SplitRange[] {
const body =
node.childForFieldName('body') ||
node.namedChildren.find((c: any) => BODY_NODE_TYPES.has(c.type)) ||
null;
if (!body || body.namedChildren.length < 2) return [];
const children = body.namedChildren.filter((c: any) => !c.isExtra);
if (children.length < 2) return [];
const ranges: SplitRange[] = [];
let curStart = children[0].startIndex;
let curStartLine = children[0].startPosition.row + 1;
let curEnd = children[0].endIndex;
let curEndLine = children[0].endPosition.row + 1;
let curTokens = estimateTokens(source.slice(curStart, curEnd));
for (let i = 1; i < children.length; i++) {
const child = children[i];
const childTokens = estimateTokens(source.slice(child.startIndex, child.endIndex));
if (curTokens + childTokens > Math.ceil(chunkTarget * 1.5)) {
ranges.push({ startIndex: curStart, endIndex: curEnd, startLine: curStartLine, endLine: curEndLine });
curStart = child.startIndex;
curStartLine = child.startPosition.row + 1;
curEnd = child.endIndex;
curEndLine = child.endPosition.row + 1;
curTokens = childTokens;
} else {
curEnd = child.endIndex;
curEndLine = child.endPosition.row + 1;
curTokens += childTokens;
}
}
ranges.push({ startIndex: curStart, endIndex: curEnd, startLine: curStartLine, endLine: curEndLine });
return ranges;
}
function extractSymbolName(node: any): string | null {
const directName = node.childForFieldName('name');
if (directName?.text?.trim()) return sanitize(directName.text);
const declaration = node.childForFieldName('declaration');
if (declaration) {
const nested = extractSymbolName(declaration);
if (nested) return nested;
}
for (const child of node.namedChildren) {
if (child.type.endsWith('identifier') || child.type === 'constant') {
const v = sanitize(child.text);
if (v) return v;
}
}
return null;
}
function normalizeSymbolType(type: string): string {
if (type.includes('function') || type === 'method' || type === 'singleton_method') return 'function';
if (type.includes('class')) return 'class';
if (type.includes('interface')) return 'interface';
if (type.includes('type_alias')) return 'type';
if (type.includes('enum')) return 'enum';
if (type.includes('module')) return 'module';
if (type.includes('import')) return 'import';
return type.replace(/_/g, ' ');
}
function sanitize(name: string): string {
return name.replace(/[\n\r\t]+/g, ' ').replace(/\s+/g, ' ').trim();
}
function estimateTokens(text: string): number {
return Math.max(1, Math.ceil(text.length / 4));
}
function displayLang(lang: SupportedCodeLanguage): string {
const map: Record<SupportedCodeLanguage, string> = {
typescript: 'TypeScript', tsx: 'TSX', javascript: 'JavaScript',
python: 'Python', ruby: 'Ruby', go: 'Go',
};
return map[lang];
}
function countLines(text: string): number {
return text ? text.split('\n').length : 0;
}
// ---------- Tree-sitter init ----------
async function ensureInit(): Promise<void> {
if (initDone) return;
if (!initPromise) {
initPromise = (async () => {
const P = await getParser();
// v0.22.x: init takes locateFile for the WASM module
const wasmPath = new URL('../../../node_modules/web-tree-sitter/tree-sitter.wasm', import.meta.url);
let resolved: string;
try {
const { fileURLToPath } = await import('url');
resolved = fileURLToPath(wasmPath);
} catch {
resolved = wasmPath.pathname;
}
await (P as any).init({ locateFile: () => resolved });
initDone = true;
})();
}
await initPromise;
}
async function loadLanguage(language: SupportedCodeLanguage): Promise<any> {
if (languageCache.has(language)) return languageCache.get(language);
const P = await getParser();
const grammarUrl = new URL(
`../../../node_modules/tree-sitter-wasms/out/${GRAMMAR_FILES[language]}`,
import.meta.url,
);
let resolved: string;
try {
const { fileURLToPath } = await import('url');
resolved = fileURLToPath(grammarUrl);
} catch {
resolved = grammarUrl.pathname;
}
const lang = await (P as any).Language.load(resolved);
languageCache.set(language, lang);
return lang;
}
+15
View File
@@ -27,8 +27,23 @@ export interface GBrainConfig {
engine: 'postgres' | 'pglite';
database_url?: string;
database_path?: string;
repos?: Array<{
path: string;
name: string;
strategy: 'markdown' | 'code' | 'auto';
include?: string[];
exclude?: string[];
syncEnabled?: boolean;
}>;
openai_api_key?: string;
anthropic_api_key?: string;
/**
* Optional storage backend config (S3/Supabase/local). Shape matches
* `StorageConfig` in `./storage.ts`. Typed as `unknown` here to avoid
* a cyclic import; callers pass this through `createStorage()` which
* validates the shape at runtime.
*/
storage?: unknown;
}
/**
+817
View File
@@ -0,0 +1,817 @@
/**
* src/core/cycle.ts The brain maintenance cycle primitive.
*
* Composes lint, backlinks, sync, extract, embed, and orphans into
* one honest unit of work. Called from:
* - `gbrain dream` (CLI alias; one-shot cron-triggered cycle)
* - `gbrain autopilot` (daemon; scheduled on an interval)
* - Minions `autopilot-cycle` handler (durable queue; retry + observability)
*
* All three converge on runCycle() so there's one source of truth for
* what "overnight maintenance" means.
*
* PHASE ORDER (semantically driven fix files first, then index):
*
*
* Phase 1: lint --fix (filesystem writes, no DB)
* Phase 2: backlinks --fix (filesystem writes, no DB)
* Phase 3: sync (DB picks up phases 1+2)
* Phase 4: extract (DB picks up links from sync)
* Phase 5: embed --stale (DB writes)
* Phase 6: orphans (DB read, report only)
*
*
* COORDINATION:
*
* Postgres: a row in gbrain_cycle_locks with a TTL (30 min). Refreshed
* between phases via yieldBetweenPhases. Works through PgBouncer
* transaction pooling (session-scoped pg_try_advisory_lock does not).
*
* PGLite / engine=null: a file lock at ~/.gbrain/cycle.lock holding
* the PID + mtime. Same 30-min TTL semantics.
*
* LOCK-SKIP:
*
* Filesystem-only or read-only phase selections (lint, backlinks,
* orphans) skip the lock. Only DB-write phases (sync, extract, embed)
* trigger lock acquisition.
*/
import { existsSync, readFileSync, writeFileSync, unlinkSync, mkdirSync, statSync } from 'fs';
import { join } from 'path';
import { homedir, hostname } from 'os';
import type { BrainEngine } from './engine.ts';
import { createProgress, type ProgressReporter } from './progress.ts';
import { getCliOptions, cliOptsToProgressOptions } from './cli-options.ts';
// ─── Types ─────────────────────────────────────────────────────────
export type CyclePhase = 'lint' | 'backlinks' | 'sync' | 'extract' | 'embed' | 'orphans';
export const ALL_PHASES: CyclePhase[] = [
'lint',
'backlinks',
'sync',
'extract',
'embed',
'orphans',
];
/**
* Phases that mutate state (filesystem or DB) and therefore should
* coordinate via the cycle lock. Only orphans is truly read-only
* and skips the lock.
*/
const NEEDS_LOCK_PHASES: ReadonlySet<CyclePhase> = new Set([
'lint',
'backlinks',
'sync',
'extract',
'embed',
]);
export type PhaseStatus = 'ok' | 'warn' | 'fail' | 'skipped';
export interface PhaseError {
/** Error class for machine branching — e.g., 'DatabaseConnection', 'Timeout', 'LLMError', 'FilesystemError', 'InternalError'. */
class: string;
/** System error code or short identifier, e.g., 'ECONNREFUSED', 'ETIMEDOUT', 'UNKNOWN'. */
code: string;
/** Human-readable single-line message. */
message: string;
/** Optional suggestion of what to try next. */
hint?: string;
/** Optional link to a troubleshooting doc. */
docs_url?: string;
}
export interface PhaseResult {
phase: CyclePhase;
status: PhaseStatus;
duration_ms: number;
summary: string;
details: Record<string, unknown>;
error?: PhaseError;
}
export type CycleStatus = 'ok' | 'clean' | 'partial' | 'skipped' | 'failed';
export interface CycleReport {
/** Additive schema. Bumped on breaking changes. */
schema_version: '1';
timestamp: string;
duration_ms: number;
/**
* Overall status derived from phase results:
* - 'clean' : ran successfully, zero fixes/writes across every phase
* - 'ok' : ran successfully, some work was done
* - 'partial' : at least one phase warned or failed, others ran
* - 'skipped' : cycle did not run (lock held by another holder)
* - 'failed' : lock acquired but all attempted phases failed
*/
status: CycleStatus;
/** Present when status = 'skipped'. E.g., 'cycle_already_running' or 'no_database'. */
reason?: string;
brain_dir: string | null;
phases: PhaseResult[];
totals: {
lint_fixes: number;
backlinks_added: number;
pages_synced: number;
pages_extracted: number;
pages_embedded: number;
orphans_found: number;
};
}
export interface CycleOpts {
/** If true, no writes to filesystem or DB. All phases honor this. */
dryRun?: boolean;
/** Defaults to ALL_PHASES. Pass a subset for --phase lint etc. */
phases?: CyclePhase[];
/** Brain directory (git repo). Required for filesystem phases. */
brainDir: string;
/** Whether sync should run `git pull`. Default false (cron-safe). */
pull?: boolean;
/**
* Called between phases AND before runCycle returns. Awaited even
* after phase failure. Hook exceptions are logged, never fatal.
* Minions handlers pass a function that yields + renews the job lock
* + refreshes the cycle-lock-table TTL.
*/
yieldBetweenPhases?: () => Promise<void>;
}
// ─── Lock primitives ───────────────────────────────────────────────
const CYCLE_LOCK_ID = 'gbrain-cycle';
const LOCK_TTL_MS = 30 * 60 * 1000; // 30 minutes
const LOCK_FILE_PATH_DEFAULT = join(homedir(), '.gbrain', 'cycle.lock');
interface LockHandle {
release: () => Promise<void>;
refresh: () => Promise<void>;
}
/**
* Acquire the Postgres-backed cycle lock.
* Returns a LockHandle on success, or null if another live holder has it.
*
* Uses INSERT ... ON CONFLICT (id) DO UPDATE ... WHERE ttl_expires_at < NOW()
* RETURNING *. An empty RETURNING means the existing row is still live.
* Crashed holders auto-release: when their TTL expires, the next
* acquirer's UPDATE branch fires and takes over.
*/
async function acquirePostgresLock(engine: BrainEngine): Promise<LockHandle | null> {
const pid = process.pid;
const host = hostname();
// Engine-agnostic: BrainEngine exposes findOrphanPages etc., but not raw SQL.
// We reach through the engine's internal connection for this lock operation.
// Both engines expose `sql` (postgres-js tag) or `db.query` (PGLite).
const maybePG = engine as unknown as { sql?: (...args: unknown[]) => Promise<unknown> };
const maybePGLite = engine as unknown as { db?: { query: (sql: string, params?: unknown[]) => Promise<{ rows: unknown[] }> } };
if (engine.kind === 'postgres' && maybePG.sql) {
const sql = maybePG.sql as any;
const rows: Array<{ id: string }> = await sql`
INSERT INTO gbrain_cycle_locks (id, holder_pid, holder_host, acquired_at, ttl_expires_at)
VALUES (${CYCLE_LOCK_ID}, ${pid}, ${host}, NOW(), NOW() + INTERVAL '30 minutes')
ON CONFLICT (id) DO UPDATE
SET holder_pid = ${pid},
holder_host = ${host},
acquired_at = NOW(),
ttl_expires_at = NOW() + INTERVAL '30 minutes'
WHERE gbrain_cycle_locks.ttl_expires_at < NOW()
RETURNING id
`;
if (rows.length === 0) return null; // live holder
return {
refresh: async () => {
await sql`
UPDATE gbrain_cycle_locks
SET ttl_expires_at = NOW() + INTERVAL '30 minutes'
WHERE id = ${CYCLE_LOCK_ID} AND holder_pid = ${pid}
`;
},
release: async () => {
await sql`
DELETE FROM gbrain_cycle_locks
WHERE id = ${CYCLE_LOCK_ID} AND holder_pid = ${pid}
`;
},
};
}
if (engine.kind === 'pglite' && maybePGLite.db) {
// PGLite is single-writer; the DB row is belt-and-braces on top of the
// file lock. Callers always hold the file lock first, so this UPSERT
// is race-free against other processes.
const db = maybePGLite.db;
const { rows } = await db.query(
`INSERT INTO gbrain_cycle_locks (id, holder_pid, holder_host, acquired_at, ttl_expires_at)
VALUES ($1, $2, $3, NOW(), NOW() + INTERVAL '30 minutes')
ON CONFLICT (id) DO UPDATE
SET holder_pid = $2,
holder_host = $3,
acquired_at = NOW(),
ttl_expires_at = NOW() + INTERVAL '30 minutes'
WHERE gbrain_cycle_locks.ttl_expires_at < NOW()
RETURNING id`,
[CYCLE_LOCK_ID, pid, host],
);
if (rows.length === 0) return null;
return {
refresh: async () => {
await db.query(
`UPDATE gbrain_cycle_locks
SET ttl_expires_at = NOW() + INTERVAL '30 minutes'
WHERE id = $1 AND holder_pid = $2`,
[CYCLE_LOCK_ID, pid],
);
},
release: async () => {
await db.query(
`DELETE FROM gbrain_cycle_locks WHERE id = $1 AND holder_pid = $2`,
[CYCLE_LOCK_ID, pid],
);
},
};
}
throw new Error(`Unknown engine kind: ${engine.kind}`);
}
/**
* Acquire the file-based cycle lock (used when engine === null).
* Returns a LockHandle on success, or null if a live holder has it.
*
* The file contains `{pid}\n{iso-timestamp}`. Staleness = mtime older
* than LOCK_TTL_MS OR the PID is no longer alive on this host.
*/
function acquireFileLock(lockPath = LOCK_FILE_PATH_DEFAULT): LockHandle | null {
mkdirSync(join(lockPath, '..'), { recursive: true });
const pid = process.pid;
if (existsSync(lockPath)) {
// Check TTL.
try {
const st = statSync(lockPath);
const ageMs = Date.now() - st.mtimeMs;
const existingContent = readFileSync(lockPath, 'utf-8').trim();
const existingPid = parseInt(existingContent.split('\n')[0] || '0', 10);
// PID liveness check (same host only). kill(pid, 0) distinguishes:
// - success → process exists, caller can signal it
// - error ESRCH → no such process (truly dead)
// - error EPERM → process exists but caller can't signal it
// (e.g., PID 1/init on unix) → still alive
// Any error code OTHER than ESRCH means the PID is alive.
let pidAlive = false;
if (existingPid > 0 && existingPid !== pid) {
try {
process.kill(existingPid, 0);
pidAlive = true;
} catch (e) {
const code = (e as NodeJS.ErrnoException).code;
pidAlive = code !== 'ESRCH';
}
} else if (existingPid === pid) {
// Our own stale lock (same pid, previous run) — treat as stale.
pidAlive = false;
}
if (pidAlive && ageMs < LOCK_TTL_MS) {
return null; // live holder
}
// Stale lock — fall through to overwrite.
} catch {
// Any read/stat error: treat as stale.
}
}
writeFileSync(lockPath, `${pid}\n${new Date().toISOString()}\n`);
return {
refresh: async () => {
try {
writeFileSync(lockPath, `${pid}\n${new Date().toISOString()}\n`);
} catch {
/* non-fatal — a next-run stale check will notice */
}
},
release: async () => {
try {
const content = readFileSync(lockPath, 'utf-8').trim();
const heldPid = parseInt(content.split('\n')[0] || '0', 10);
if (heldPid === pid) unlinkSync(lockPath);
} catch {
/* already gone */
}
},
};
}
// ─── Helpers ───────────────────────────────────────────────────────
function makeErrorFromException(e: unknown, fallbackClass = 'InternalError'): PhaseError {
const err = e instanceof Error ? e : new Error(String(e));
// Node errors often have .code (e.g., 'ECONNREFUSED').
const code = (err as NodeJS.ErrnoException).code || 'UNKNOWN';
let className = fallbackClass;
if (code === 'ECONNREFUSED' || code === 'ENOTFOUND') className = 'DatabaseConnection';
if (code === 'ETIMEDOUT') className = 'Timeout';
if (/OpenAI|embed/i.test(err.message)) className = 'LLMError';
if (/ENOENT|EACCES|EISDIR|ENOTDIR/.test(code)) className = 'FilesystemError';
return {
class: className,
code,
message: err.message.slice(0, 200),
};
}
async function timePhase<T>(fn: () => Promise<T>): Promise<{ result: T; duration_ms: number }> {
const start = performance.now();
const result = await fn();
return { result, duration_ms: Math.round(performance.now() - start) };
}
async function safeYield(hook?: () => Promise<void>) {
if (!hook) return;
try {
await hook();
} catch (e) {
console.warn(`[cycle] yieldBetweenPhases hook error (non-fatal): ${e instanceof Error ? e.message : String(e)}`);
}
}
// ─── Phase runners ─────────────────────────────────────────────────
async function runPhaseLint(brainDir: string, dryRun: boolean): Promise<PhaseResult> {
try {
const { runLintCore } = await import('../commands/lint.ts');
const result = await runLintCore({ target: brainDir, fix: true, dryRun });
const issues = result.total_issues ?? 0;
const fixed = result.total_fixed ?? 0;
const remaining = Math.max(0, issues - fixed);
// 'ok' when nothing noteworthy remains:
// - no issues at all, or
// - non-dry-run and everything fixable was fixed.
// 'warn' when issues remain after the run.
const status: PhaseStatus =
issues === 0 || (!dryRun && remaining === 0) ? 'ok' : 'warn';
return {
phase: 'lint',
status,
duration_ms: 0, // set by caller
summary: dryRun
? `${issues} issue(s) found (dry-run, no writes)`
: `${fixed} fix(es) applied, ${remaining} remaining`,
details: { issues, fixed, pages_scanned: result.pages_scanned, dryRun },
};
} catch (e) {
return {
phase: 'lint',
status: 'fail',
duration_ms: 0,
summary: 'lint phase failed',
details: {},
error: makeErrorFromException(e),
};
}
}
async function runPhaseBacklinks(brainDir: string, dryRun: boolean): Promise<PhaseResult> {
try {
// Library function path — the v0.15 backlinks.ts exports
// runBacklinksCore when --fix is requested.
const { runBacklinksCore } = await import('../commands/backlinks.ts');
const result = await runBacklinksCore({
action: 'fix',
dir: brainDir,
dryRun,
});
const gaps = result.gaps_found ?? 0;
const added = result.fixed ?? 0;
const remaining = Math.max(0, gaps - added);
const status: PhaseStatus =
gaps === 0 || (!dryRun && remaining === 0) ? 'ok' : 'warn';
return {
phase: 'backlinks',
status,
duration_ms: 0,
summary: dryRun
? `${gaps} missing back-link(s) (dry-run)`
: `${added} back-link(s) added, ${remaining} remaining`,
details: { gaps, added, pages_affected: result.pages_affected, dryRun },
};
} catch (e) {
return {
phase: 'backlinks',
status: 'fail',
duration_ms: 0,
summary: 'backlinks phase failed',
details: {},
error: makeErrorFromException(e),
};
}
}
async function runPhaseSync(
engine: BrainEngine,
brainDir: string,
dryRun: boolean,
pull: boolean,
): Promise<PhaseResult> {
try {
const { performSync } = await import('../commands/sync.ts');
const result = await performSync(engine, {
repoPath: brainDir,
dryRun,
noPull: !pull,
noEmbed: true, // embed is a separate phase
});
const syncedCount = result.added + result.modified;
return {
phase: 'sync',
status: result.status === 'blocked_by_failures' ? 'warn' : 'ok',
duration_ms: 0,
summary: dryRun
? `${syncedCount} page(s) would sync, ${result.deleted} would delete`
: `+${result.added} added, ~${result.modified} modified, -${result.deleted} deleted`,
details: {
added: result.added,
modified: result.modified,
deleted: result.deleted,
renamed: result.renamed,
chunksCreated: result.chunksCreated,
failedFiles: result.failedFiles ?? 0,
syncStatus: result.status,
dryRun,
},
};
} catch (e) {
return {
phase: 'sync',
status: 'fail',
duration_ms: 0,
summary: 'sync phase failed',
details: {},
error: makeErrorFromException(e),
};
}
}
async function runPhaseExtract(
engine: BrainEngine,
brainDir: string,
dryRun: boolean,
): Promise<PhaseResult> {
try {
const { runExtractCore } = await import('../commands/extract.ts');
// Extract is read-mostly against the filesystem + write to links table.
// Honor dryRun by skipping with a 'skipped' entry: extract doesn't have
// a clean dry-run mode today and runCycle should be honest about it.
if (dryRun) {
return {
phase: 'extract',
status: 'skipped',
duration_ms: 0,
summary: 'dry-run: extract phase skipped (no dry-run mode yet)',
details: { dryRun: true, reason: 'no_dry_run_support' },
};
}
const result = await runExtractCore(engine, { mode: 'all', dir: brainDir });
const linksCreated = result?.links_created ?? 0;
const timelineCreated = result?.timeline_entries_created ?? 0;
return {
phase: 'extract',
status: 'ok',
duration_ms: 0,
summary: `${linksCreated} link(s), ${timelineCreated} timeline entries`,
details: { linksCreated, timelineCreated, pages_processed: result?.pages_processed ?? 0 },
};
} catch (e) {
return {
phase: 'extract',
status: 'fail',
duration_ms: 0,
summary: 'extract phase failed',
details: {},
error: makeErrorFromException(e),
};
}
}
async function runPhaseEmbed(engine: BrainEngine, dryRun: boolean): Promise<PhaseResult> {
try {
const { runEmbedCore } = await import('../commands/embed.ts');
const result = await runEmbedCore(engine, { stale: true, dryRun });
const embeddedCount = dryRun ? result.would_embed : result.embedded;
return {
phase: 'embed',
status: 'ok',
duration_ms: 0,
summary: dryRun
? `${result.would_embed} chunk(s) would be embedded (dry-run)`
: `${result.embedded} chunk(s) newly embedded (${result.skipped} already had embeddings)`,
details: {
embedded: result.embedded,
skipped: result.skipped,
would_embed: result.would_embed,
total_chunks: result.total_chunks,
pages_processed: result.pages_processed,
dryRun,
// Convenience field used by CycleReport.totals.pages_embedded.
// In dry-run, this counts pages with stale chunks that would
// have been processed (same semantic as a real run).
pages_embedded_count: dryRun ? result.pages_processed : embeddedCount > 0 ? result.pages_processed : 0,
},
};
} catch (e) {
return {
phase: 'embed',
status: 'fail',
duration_ms: 0,
summary: 'embed phase failed',
details: {},
error: makeErrorFromException(e),
};
}
}
async function runPhaseOrphans(engine: BrainEngine): Promise<PhaseResult> {
try {
const { findOrphans } = await import('../commands/orphans.ts');
const result = await findOrphans(engine);
const count = result.total_orphans;
return {
phase: 'orphans',
status: count > 20 ? 'warn' : 'ok',
duration_ms: 0,
summary: `${count} orphan page(s) out of ${result.total_pages} total`,
details: {
total_orphans: count,
total_pages: result.total_pages,
excluded: result.excluded,
},
};
} catch (e) {
return {
phase: 'orphans',
status: 'fail',
duration_ms: 0,
summary: 'orphans phase failed',
details: {},
error: makeErrorFromException(e),
};
}
}
// ─── Main ──────────────────────────────────────────────────────────
/**
* Run the brain maintenance cycle.
*
* Engine may be null: filesystem phases (lint, backlinks) still run;
* DB-dependent phases skip with status='skipped', reason='no_database'.
*
* Acquires the cycle lock for any DB-write phase selection. Non-DB-write
* selections (e.g., --phase lint) skip the lock as an optimization so
* single-phase runs are always responsive even if another cycle is live.
*/
export async function runCycle(
engine: BrainEngine | null,
opts: CycleOpts,
): Promise<CycleReport> {
const start = performance.now();
const phases = opts.phases ?? ALL_PHASES;
const dryRun = !!opts.dryRun;
const pull = !!opts.pull;
const timestamp = new Date().toISOString();
const phaseResults: PhaseResult[] = [];
const progress = createProgress(cliOptsToProgressOptions(getCliOptions()));
// Decide if we need the cycle lock: any state-mutating phase in the selection.
const needsLock = phases.some(p => NEEDS_LOCK_PHASES.has(p));
let lock: LockHandle | null = null;
if (needsLock) {
if (engine) {
try {
lock = await acquirePostgresLock(engine);
} catch (e) {
// Lock acquisition failed catastrophically (e.g., migration missing).
// Return a failed report rather than silently running without a lock.
return {
schema_version: '1',
timestamp,
duration_ms: Math.round(performance.now() - start),
status: 'failed',
reason: 'lock_acquisition_error',
brain_dir: opts.brainDir,
phases: [
{
phase: 'sync',
status: 'fail',
duration_ms: 0,
summary: 'could not acquire cycle lock',
details: {},
error: makeErrorFromException(e, 'DatabaseConnection'),
},
],
totals: emptyTotals(),
};
}
} else {
lock = acquireFileLock();
}
if (lock === null) {
return {
schema_version: '1',
timestamp,
duration_ms: Math.round(performance.now() - start),
status: 'skipped',
reason: 'cycle_already_running',
brain_dir: opts.brainDir,
phases: [],
totals: emptyTotals(),
};
}
}
try {
// ── Phase 1: lint ────────────────────────────────────────────
if (phases.includes('lint')) {
progress.start('cycle.lint');
const { result, duration_ms } = await timePhase(() => runPhaseLint(opts.brainDir, dryRun));
result.duration_ms = duration_ms;
phaseResults.push(result);
progress.finish();
await safeYield(opts.yieldBetweenPhases);
}
// ── Phase 2: backlinks ──────────────────────────────────────
if (phases.includes('backlinks')) {
progress.start('cycle.backlinks');
const { result, duration_ms } = await timePhase(() => runPhaseBacklinks(opts.brainDir, dryRun));
result.duration_ms = duration_ms;
phaseResults.push(result);
progress.finish();
await safeYield(opts.yieldBetweenPhases);
}
// ── Phase 3: sync ───────────────────────────────────────────
if (phases.includes('sync')) {
if (!engine) {
phaseResults.push({
phase: 'sync',
status: 'skipped',
duration_ms: 0,
summary: 'no database connected',
details: { reason: 'no_database' },
});
} else {
progress.start('cycle.sync');
const { result, duration_ms } = await timePhase(() => runPhaseSync(engine, opts.brainDir, dryRun, pull));
result.duration_ms = duration_ms;
phaseResults.push(result);
progress.finish();
}
await safeYield(opts.yieldBetweenPhases);
}
// ── Phase 4: extract ────────────────────────────────────────
if (phases.includes('extract')) {
if (!engine) {
phaseResults.push({
phase: 'extract',
status: 'skipped',
duration_ms: 0,
summary: 'no database connected',
details: { reason: 'no_database' },
});
} else {
progress.start('cycle.extract');
const { result, duration_ms } = await timePhase(() => runPhaseExtract(engine, opts.brainDir, dryRun));
result.duration_ms = duration_ms;
phaseResults.push(result);
progress.finish();
}
await safeYield(opts.yieldBetweenPhases);
}
// ── Phase 5: embed ──────────────────────────────────────────
if (phases.includes('embed')) {
if (!engine) {
phaseResults.push({
phase: 'embed',
status: 'skipped',
duration_ms: 0,
summary: 'no database connected',
details: { reason: 'no_database' },
});
} else {
progress.start('cycle.embed');
const { result, duration_ms } = await timePhase(() => runPhaseEmbed(engine, dryRun));
result.duration_ms = duration_ms;
phaseResults.push(result);
progress.finish();
}
await safeYield(opts.yieldBetweenPhases);
}
// ── Phase 6: orphans ────────────────────────────────────────
if (phases.includes('orphans')) {
if (!engine) {
phaseResults.push({
phase: 'orphans',
status: 'skipped',
duration_ms: 0,
summary: 'no database connected',
details: { reason: 'no_database' },
});
} else {
progress.start('cycle.orphans');
const { result, duration_ms } = await timePhase(() => runPhaseOrphans(engine));
result.duration_ms = duration_ms;
phaseResults.push(result);
progress.finish();
}
await safeYield(opts.yieldBetweenPhases);
}
} finally {
if (lock) {
try { await lock.release(); } catch { /* best-effort */ }
}
}
const duration_ms = Math.round(performance.now() - start);
const totals = extractTotals(phaseResults);
const status = deriveStatus(phaseResults, totals);
return {
schema_version: '1',
timestamp,
duration_ms,
status,
brain_dir: opts.brainDir,
phases: phaseResults,
totals,
};
}
// ─── Totals + status derivation ────────────────────────────────────
function emptyTotals(): CycleReport['totals'] {
return {
lint_fixes: 0,
backlinks_added: 0,
pages_synced: 0,
pages_extracted: 0,
pages_embedded: 0,
orphans_found: 0,
};
}
function extractTotals(phases: PhaseResult[]): CycleReport['totals'] {
const t = emptyTotals();
for (const p of phases) {
if (p.phase === 'lint' && p.details) {
t.lint_fixes = Number(p.details.fixed ?? 0);
} else if (p.phase === 'backlinks' && p.details) {
t.backlinks_added = Number(p.details.added ?? 0);
} else if (p.phase === 'sync' && p.details) {
t.pages_synced = Number(p.details.added ?? 0) + Number(p.details.modified ?? 0);
} else if (p.phase === 'extract' && p.details) {
t.pages_extracted = Number(p.details.linksCreated ?? 0);
} else if (p.phase === 'embed' && p.details) {
// In dry-run, use would_embed as the "activity" measure; else embedded.
const dryRun = p.details.dryRun === true;
t.pages_embedded = dryRun
? Number(p.details.would_embed ?? 0)
: Number(p.details.embedded ?? 0);
} else if (p.phase === 'orphans' && p.details) {
t.orphans_found = Number(p.details.total_orphans ?? 0);
}
}
return t;
}
function deriveStatus(phases: PhaseResult[], totals: CycleReport['totals']): CycleStatus {
if (phases.length === 0) return 'failed';
const anyFailed = phases.some(p => p.status === 'fail');
const allFailed = phases.every(p => p.status === 'fail');
const anyWarn = phases.some(p => p.status === 'warn');
if (allFailed) return 'failed';
if (anyFailed || anyWarn) return 'partial';
// All phases 'ok' or 'skipped'. Distinguish clean (no activity) from ok (work done).
const anyWork =
totals.lint_fixes > 0 ||
totals.backlinks_added > 0 ||
totals.pages_synced > 0 ||
totals.pages_extracted > 0 ||
totals.pages_embedded > 0;
return anyWork ? 'ok' : 'clean';
}
+1 -1
View File
@@ -159,5 +159,5 @@ export async function withTransaction<T>(fn: (tx: ReturnType<typeof postgres>) =
const conn = getConnection();
return conn.begin(async (tx) => {
return fn(tx as unknown as ReturnType<typeof postgres>);
});
}) as Promise<T>;
}
+7
View File
@@ -152,6 +152,13 @@ export interface BrainEngine {
* Slugs with zero inbound links are present in the map with value 0.
*/
getBacklinkCounts(slugs: string[]): Promise<Map<string, number>>;
/**
* Return every page with no inbound links (from any source).
* Domain comes from the frontmatter `domain` field (null if unset).
* The caller filters pseudo-pages + derives display domain.
* Used by `gbrain orphans` and `runCycle`'s orphan sweep phase.
*/
findOrphanPages(): Promise<Array<{ slug: string; title: string; domain: string | null }>>;
// Tags
addTag(slug: string, tag: string): Promise<void>;
+3 -3
View File
@@ -113,8 +113,8 @@ export async function enrichEntity(
let timelineAdded = false;
try {
await engine.addTimelineEntry(slug, {
date: new Date().toISOString().split('T')[0],
content: `Referenced in [${request.sourceSlug}](${request.sourceSlug}) — ${request.context}`,
date: new Date().toISOString().split('T')[0] ?? '',
summary: `Referenced in [${request.sourceSlug}](${request.sourceSlug}) — ${request.context}`,
source: request.sourceSlug,
});
timelineAdded = true;
@@ -161,7 +161,7 @@ export async function enrichEntities(
}
const result = await enrichEntity(engine, req);
results.push(result);
config?.onProgress?.(results.length, requests.length, req.name);
config?.onProgress?.(results.length, requests.length, req.entityName);
}
return results;
}
+1 -1
View File
@@ -1,7 +1,7 @@
/**
* CompletenessScorer per-entity-type rubrics, 0.01.0 score per page.
*
* Replaces Wintermute's length-based heuristic ("compiled_truth > 500 chars")
* Replaces Garry's OpenClaw's length-based heuristic ("compiled_truth > 500 chars")
* with a weighted rubric that actually reflects whether a page would be
* useful to answer a query. Runs on demand; BrainWriter invokes it on
* write to cache the score in frontmatter.
+3 -3
View File
@@ -119,18 +119,18 @@ export async function resolveFile(
/** Parse v0.9+ .redirect.yaml pointer */
export function parseRedirectYaml(path: string): RedirectYaml {
const content = readFileSync(path, 'utf-8');
return parseYaml(content) as RedirectYaml;
return parseYaml(content) as unknown as RedirectYaml;
}
/** Parse legacy v0.8 .redirect breadcrumb */
export function parseRedirect(path: string): RedirectInfo {
const content = readFileSync(path, 'utf-8');
return parseYaml(content) as RedirectInfo;
return parseYaml(content) as unknown as RedirectInfo;
}
export function parseMarker(path: string): MarkerInfo {
const content = readFileSync(path, 'utf-8');
return parseYaml(content) as MarkerInfo;
return parseYaml(content) as unknown as MarkerInfo;
}
/** Human-readable file size */
+86 -1
View File
@@ -1,10 +1,12 @@
import { readFileSync, statSync, lstatSync } from 'fs';
import { basename } from 'path';
import { createHash } from 'crypto';
import type { BrainEngine } from './engine.ts';
import { parseMarkdown } from './markdown.ts';
import { chunkText } from './chunkers/recursive.ts';
import { chunkCodeText, detectCodeLanguage } from './chunkers/code.ts';
import { embedBatch } from './embedding.ts';
import { slugifyPath } from './sync.ts';
import { slugifyPath, slugifyCodePath, isCodeFilePath } from './sync.ts';
import type { ChunkInput, PageType } from './types.ts';
/**
@@ -184,6 +186,12 @@ export async function importFromFile(
}
const content = readFileSync(filePath, 'utf-8');
// Route code files through the code import path
if (isCodeFilePath(relativePath)) {
return importCodeFile(engine, relativePath, content, opts);
}
const parsed = parseMarkdown(content, relativePath);
// Enforce path-authoritative slug. parseMarkdown prefers frontmatter.slug over
@@ -206,6 +214,83 @@ export async function importFromFile(
return importFromContent(engine, expectedSlug, content, opts);
}
/**
* Import a code file. Bypasses markdown parsing entirely.
* Uses tree-sitter code chunker for semantic splitting.
* Page type is 'code', slug includes file extension.
*/
export async function importCodeFile(
engine: BrainEngine,
relativePath: string,
content: string,
opts: { noEmbed?: boolean } = {},
): Promise<ImportResult> {
const slug = slugifyCodePath(relativePath);
const lang = detectCodeLanguage(relativePath) || 'unknown';
const title = `${relativePath} (${lang})`;
const byteLength = Buffer.byteLength(content, 'utf-8');
if (byteLength > MAX_FILE_SIZE) {
return { slug, status: 'skipped', chunks: 0, error: `Code file too large (${byteLength} bytes)` };
}
// Hash for idempotency
const hash = createHash('sha256')
.update(JSON.stringify({ title, type: 'code', content, lang }))
.digest('hex');
const existing = await engine.getPage(slug);
if (existing?.content_hash === hash) {
return { slug, status: 'skipped', chunks: 0 };
}
// Chunk via tree-sitter code chunker
const codeChunks = await chunkCodeText(content, relativePath);
const chunks: ChunkInput[] = codeChunks.map((c, i) => ({
chunk_index: i,
chunk_text: c.text,
chunk_source: 'compiled_truth' as const,
}));
// Embed
if (!opts.noEmbed && chunks.length > 0) {
try {
const embeddings = await embedBatch(chunks.map(c => c.chunk_text));
for (let i = 0; i < chunks.length; i++) {
chunks[i].embedding = embeddings[i];
chunks[i].token_count = Math.ceil(chunks[i].chunk_text.length / 4);
}
} catch (e: unknown) {
console.warn(`[gbrain] embedding failed for code file ${slug}: ${e instanceof Error ? e.message : String(e)}`);
}
}
// Store
await engine.transaction(async (tx) => {
if (existing) await tx.createVersion(slug);
await tx.putPage(slug, {
type: 'code' as PageType,
title,
compiled_truth: content,
timeline: '',
frontmatter: { language: lang, file: relativePath },
content_hash: hash,
});
await tx.addTag(slug, 'code');
await tx.addTag(slug, lang);
if (chunks.length > 0) {
await tx.upsertChunks(slug, chunks);
} else {
await tx.deleteChunks(slug);
}
});
return { slug, status: 'imported', chunks: chunks.length };
}
// Backward compat
export const importFile = importFromFile;
export type ImportFileResult = ImportResult;
+26
View File
@@ -466,6 +466,32 @@ export const MIGRATIONS: Migration[] = [
AND max_stalled < 5;
`,
},
{
version: 16,
name: 'cycle_locks_table',
// v0.17 brain maintenance cycle (runCycle primitive).
// PgBouncer transaction pooling strips session-scoped advisory locks
// (pg_try_advisory_lock) across connection checkouts, so we can't use
// them as the cycle-coordination primitive. A row with a TTL works
// through every pooler: any backend can SELECT/UPDATE/DELETE it, no
// session state required.
//
// Acquire: INSERT ... ON CONFLICT (id) DO UPDATE ... WHERE ttl_expires_at < NOW()
// returning ... — empty RETURNING = lock held by live holder.
// Refresh: UPDATE ... SET ttl_expires_at = NOW() + interval '30 min'
// WHERE id = 'gbrain-cycle' AND holder_pid = <my pid> — between phases.
// Release: DELETE WHERE id = 'gbrain-cycle' AND holder_pid = <my pid>.
sql: `
CREATE TABLE IF NOT EXISTS gbrain_cycle_locks (
id TEXT PRIMARY KEY,
holder_pid INT NOT NULL,
holder_host TEXT,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
ttl_expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_cycle_locks_ttl ON gbrain_cycle_locks(ttl_expires_at);
`,
},
];
export const LATEST_VERSION = MIGRATIONS.length > 0
+1 -1
View File
@@ -269,7 +269,7 @@ export async function shellHandler(ctx: MinionJobContext): Promise<ShellJobResul
if (ctx.signal.aborted) sigAbort();
if (ctx.shutdownSignal.aborted) shutdownAbort();
const exitCode: number = await new Promise((resolve, reject) => {
const exitCode: number = await new Promise<number>((resolve, reject) => {
proc.on('error', (err) => {
reject(err);
});
@@ -0,0 +1,169 @@
/**
* subagent_aggregator handler (v0.15).
*
* This is the job that CLAIMS after all subagent children resolve and
* produces the final aggregated output. Not a polling parent Lane 1B's
* queue changes make every terminal child transition (complete/failed/
* dead/cancelled/timeout) emit a child_done message into this job's
* inbox, AND flip this job out of waiting-children once all kids are
* terminal. When we claim, all N child_done messages are already in
* minion_inbox.
*
* The aggregator does NOT re-call Anthropic in v0.15. It reads child
* results from child_done messages, builds a markdown summary, and
* returns it as the handler result. If children produced brain pages
* under wiki/agents/<child_id>/..., those are referenced by slug not
* re-embedded into the summary blob.
*
* v0.16+ will add an LLM synthesis pass for richer summaries. The v0.15
* output is deterministic string concatenation so fan-out runs stay
* reproducible.
*/
import type { MinionJobContext, ChildDoneMessage, ChildOutcome } from '../types.ts';
import type { AggregatorHandlerData } from '../types.ts';
export interface AggregatorResult {
/** Per-child record in the order children_ids was supplied. */
children: Array<{
child_id: number;
job_name: string;
outcome: ChildOutcome;
error: string | null;
/** JSON-parsed result payload for successful children. null on failure/cancel/timeout. */
result: unknown;
}>;
/** Counts by outcome — quick shape for logs + tests. */
summary: Record<ChildOutcome, number>;
/** Rendered markdown, suitable for attaching to the job row or writing as a brain page. */
markdown: string;
}
/** v0.15 aggregator: synchronous read from inbox, no LLM call. */
export async function subagentAggregatorHandler(ctx: MinionJobContext): Promise<AggregatorResult> {
const data = (ctx.data ?? {}) as unknown as AggregatorHandlerData;
const expectedIds = Array.isArray(data.children_ids) ? data.children_ids : [];
if (expectedIds.length === 0) {
return {
children: [],
summary: emptySummary(),
markdown: '# Aggregated subagent results\n\n_(no children)_',
};
}
// Read every child_done inbox message addressed to this job. By the time
// we're claimed, the queue layer has posted one per child terminal
// transition. The `readInbox` method marks messages as read so future
// claims don't re-process them.
const messages = await ctx.readInbox();
const childDoneByChildId = new Map<number, ChildDoneMessage>();
for (const m of messages) {
const payload = parseChildDone(m.payload);
if (!payload) continue;
childDoneByChildId.set(payload.child_id, payload);
}
const summary = emptySummary();
const children: AggregatorResult['children'] = expectedIds.map(childId => {
const msg = childDoneByChildId.get(childId);
if (!msg) {
// Missing — shouldn't happen under the v0.15 invariants (every
// terminal path emits child_done). Surface as a failure row so the
// aggregator is honest about what it knows.
summary.failed = (summary.failed ?? 0) + 1;
return {
child_id: childId,
job_name: '',
outcome: 'failed',
error: 'no child_done message observed in inbox',
result: null,
};
}
const outcome: ChildOutcome = msg.outcome ?? 'complete';
summary[outcome] = (summary[outcome] ?? 0) + 1;
return {
child_id: childId,
job_name: msg.job_name,
outcome,
error: msg.error ?? null,
result: outcome === 'complete' ? msg.result : null,
};
});
const markdown = renderMarkdown(children, summary, data.aggregate_prompt_template);
await ctx.updateProgress({ total: expectedIds.length, summary });
await ctx.log(`aggregated ${expectedIds.length} children — ${formatSummary(summary)}`);
return { children, summary, markdown };
}
// ── internal ────────────────────────────────────────────────
function emptySummary(): Record<ChildOutcome, number> {
return { complete: 0, failed: 0, dead: 0, cancelled: 0, timeout: 0 };
}
function formatSummary(s: Record<ChildOutcome, number>): string {
return Object.entries(s)
.filter(([, n]) => n > 0)
.map(([k, n]) => `${k}=${n}`)
.join(', ');
}
function parseChildDone(payload: unknown): ChildDoneMessage | null {
const obj = typeof payload === 'string' ? safeParse(payload) : payload;
if (!obj || typeof obj !== 'object') return null;
const rec = obj as Record<string, unknown>;
if (rec.type !== 'child_done' || typeof rec.child_id !== 'number') return null;
return {
type: 'child_done',
child_id: rec.child_id,
job_name: typeof rec.job_name === 'string' ? rec.job_name : '',
result: rec.result,
outcome: typeof rec.outcome === 'string' ? rec.outcome as ChildOutcome : undefined,
error: typeof rec.error === 'string' ? rec.error : null,
};
}
function safeParse(raw: string): unknown {
try { return JSON.parse(raw); } catch { return null; }
}
function renderMarkdown(
children: AggregatorResult['children'],
summary: Record<ChildOutcome, number>,
template?: string,
): string {
const header = template && template.trim().length > 0
? template
: '# Aggregated subagent results';
const parts: string[] = [header, ''];
parts.push(`- total: ${children.length}`);
for (const [outcome, n] of Object.entries(summary)) {
if (n > 0) parts.push(`- ${outcome}: ${n}`);
}
parts.push('');
for (const c of children) {
parts.push(`## child ${c.child_id} (${c.job_name || 'unknown'}) — ${c.outcome}`);
if (c.error) parts.push(`> error: ${c.error}`);
if (c.outcome === 'complete' && c.result !== undefined) {
parts.push('```json', JSON.stringify(c.result, null, 2), '```');
}
parts.push('');
}
return parts.join('\n').replace(/\n{3,}/g, '\n\n');
}
// ── Testing surface ─────────────────────────────────────────
export const __testing = {
emptySummary,
formatSummary,
parseChildDone,
renderMarkdown,
};
+137
View File
@@ -0,0 +1,137 @@
/**
* Subagent audit + heartbeat log. JSONL, file-rotated weekly, best-effort.
*
* Two event flavors:
* - submission: one line per subagent job submit (mirrors shell-audit).
* - heartbeat: one line per LLM turn boundary (started / completed) so
* `gbrain agent logs <job> --follow` has fresh content to
* show during long Anthropic calls. Without these, a
* 30-second model call produces zero output between turns
* and --follow looks frozen.
*
* Never logs prompts, tool inputs, or full tool outputs (PII risk input
* vars may contain emails, free text from the user, etc.). DO log
* non-identifying operational fields: tokens, duration, model, tool_name.
*
* `GBRAIN_AUDIT_DIR` overrides the default ~/.gbrain/audit/ path useful
* for container deploys with a read-only $HOME.
*/
import * as fs from 'node:fs';
import * as path from 'node:path';
import { resolveAuditDir } from './shell-audit.ts';
export interface SubagentSubmissionEvent {
ts: string;
type: 'submission';
caller: 'cli' | 'mcp' | 'worker';
remote: boolean;
job_id: number;
parent_job_id?: number | null;
model?: string;
tools_count?: number;
allowed_tools?: string[];
}
export interface SubagentHeartbeatEvent {
ts: string;
type: 'heartbeat';
job_id: number;
event: 'llm_call_started' | 'llm_call_completed' | 'tool_called' | 'tool_result' | 'tool_failed';
turn_idx: number;
/** Tool name for tool_* events. Never the input — that may contain secrets. */
tool_name?: string;
/** ms elapsed for *_completed / tool_result / tool_failed. */
ms_elapsed?: number;
/** Token rollup for llm_call_completed. Per-turn, not cumulative. */
tokens?: { in?: number; out?: number; cache_read?: number; cache_create?: number };
/** Short error text for tool_failed. First 200 chars. */
error?: string;
}
export type SubagentAuditEvent = SubagentSubmissionEvent | SubagentHeartbeatEvent;
/** File name, rotated by ISO week. `subagent-jobs-YYYY-Www.jsonl`. */
export function computeSubagentAuditFilename(now: Date = new Date()): string {
const d = new Date(Date.UTC(now.getUTCFullYear(), now.getUTCMonth(), now.getUTCDate()));
const dayNum = (d.getUTCDay() + 6) % 7;
d.setUTCDate(d.getUTCDate() - dayNum + 3);
const isoYear = d.getUTCFullYear();
const firstThursday = new Date(Date.UTC(isoYear, 0, 4));
const firstThursdayDayNum = (firstThursday.getUTCDay() + 6) % 7;
firstThursday.setUTCDate(firstThursday.getUTCDate() - firstThursdayDayNum + 3);
const weekNum = Math.round((d.getTime() - firstThursday.getTime()) / (7 * 86400000)) + 1;
const ww = String(weekNum).padStart(2, '0');
return `subagent-jobs-${isoYear}-W${ww}.jsonl`;
}
/** Low-level append. Best-effort; write failure goes to stderr + keep running. */
function append(event: SubagentAuditEvent): void {
const dir = resolveAuditDir();
const file = path.join(dir, computeSubagentAuditFilename());
const line = JSON.stringify(event) + '\n';
try {
fs.mkdirSync(dir, { recursive: true });
fs.appendFileSync(file, line, { encoding: 'utf8' });
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
process.stderr.write(`[subagent-audit] write failed (${msg}); job continues\n`);
}
}
export function logSubagentSubmission(event: Omit<SubagentSubmissionEvent, 'ts' | 'type'>): void {
append({ ...event, ts: new Date().toISOString(), type: 'submission' });
}
export function logSubagentHeartbeat(event: Omit<SubagentHeartbeatEvent, 'ts' | 'type'>): void {
// Defensive: trim error text to avoid accidentally writing huge stack traces.
const trimmed = event.error ? { ...event, error: event.error.slice(0, 200) } : event;
append({ ...trimmed, ts: new Date().toISOString(), type: 'heartbeat' });
}
/**
* Read back all audit events for a job id from the current + prior week
* files. Used by `gbrain agent logs <job>`. Returns chronological order.
*
* `sinceIso` (if present) filters to events with ts >= sinceIso.
*/
export function readSubagentAuditForJob(jobId: number, opts: { sinceIso?: string } = {}): SubagentAuditEvent[] {
const dir = resolveAuditDir();
if (!fs.existsSync(dir)) return [];
const now = new Date();
const thisWeek = computeSubagentAuditFilename(now);
const weekAgo = computeSubagentAuditFilename(new Date(now.getTime() - 7 * 86400000));
const candidates = [...new Set([weekAgo, thisWeek])];
const out: SubagentAuditEvent[] = [];
for (const name of candidates) {
const file = path.join(dir, name);
if (!fs.existsSync(file)) continue;
let raw: string;
try {
raw = fs.readFileSync(file, 'utf8');
} catch {
continue;
}
for (const line of raw.split('\n')) {
if (!line) continue;
let ev: SubagentAuditEvent;
try {
ev = JSON.parse(line) as SubagentAuditEvent;
} catch {
continue;
}
// Submission events have job_id at top level; heartbeats too. Both safe.
if ((ev as { job_id?: number }).job_id !== jobId) continue;
if (opts.sinceIso && ev.ts < opts.sinceIso) continue;
out.push(ev);
}
}
return out.sort((a, b) => a.ts.localeCompare(b.ts));
}
/** Exported for unit tests. */
export const __testing = {
append,
};
+710
View File
@@ -0,0 +1,710 @@
/**
* Subagent LLM-loop handler (v0.15).
*
* Runs one Anthropic Messages API conversation with tool use. The loop is
* crash-resumable: subagent_messages + subagent_tool_executions together
* are the single source of truth about where the conversation is. On
* resume after a worker kill, we load all committed rows, trust any tool
* execution marked 'complete' or 'failed', and re-run 'pending' ones only
* for idempotent tools.
*
* Safety rails:
* - rate leases around every LLM call (acquire call release). Mid-
* call renewal with backoff. Persistent renewal failure aborts as a
* renewable error so the worker re-claims.
* - dual-signal abort wiring (ctx.signal + ctx.shutdownSignal) drains
* the in-flight call and commits whatever turns are already persisted.
* - Anthropic prompt cache markers on system + tools blocks.
* - token rollup via ctx.updateTokens per turn.
*
* NOT in v0.15: refusal detection, stop_reason=max_tokens partial
* recovery, parallel tool-use dispatch (runs tools sequentially; the
* Messages API allows parallel tool_use blocks and the replay tolerates
* them, but v1 dispatches serially for simplicity). All three are tracked
* as P2 items in the plan file.
*/
import Anthropic from '@anthropic-ai/sdk';
import type { MinionJobContext, MinionJob } from '../types.ts';
import type {
ContentBlock,
SubagentHandlerData,
SubagentResult,
SubagentStopReason,
ToolDef,
} from '../types.ts';
import type { BrainEngine } from '../../engine.ts';
import type { GBrainConfig } from '../../config.ts';
import { loadConfig } from '../../config.ts';
import { buildBrainTools, filterAllowedTools } from '../tools/brain-allowlist.ts';
import {
acquireLease,
releaseLease,
renewLeaseWithBackoff,
} from '../rate-leases.ts';
import {
logSubagentSubmission,
logSubagentHeartbeat,
} from './subagent-audit.ts';
// ── Defaults ────────────────────────────────────────────────
const DEFAULT_MODEL = 'claude-sonnet-4-6';
const DEFAULT_MAX_TURNS = 20;
const DEFAULT_RATE_KEY = 'anthropic:messages';
const DEFAULT_MAX_CONCURRENT = Number(process.env.GBRAIN_ANTHROPIC_MAX_INFLIGHT ?? '8');
const DEFAULT_LEASE_TTL_MS = 120_000;
const DEFAULT_SYSTEM = 'You are a helpful assistant running as a gbrain subagent.';
// ── Injectable surfaces (for tests) ─────────────────────────
/**
* Anthropic Messages client. The real Anthropic SDK implements this
* structurally; tests can substitute a mock without the SDK import.
*/
export interface MessagesClient {
create(params: Anthropic.MessageCreateParamsNonStreaming, opts?: { signal?: AbortSignal }): Promise<Anthropic.Message>;
}
export interface SubagentDeps {
/** Engine for DB-backed ops (tools + message persistence + rate leases). */
engine: BrainEngine;
/** Anthropic client. Defaults to the SDK-constructed client. */
client?: MessagesClient;
/**
* Anthropic SDK constructor. Defaults to `() => new Anthropic()`.
* Overridable in tests so the factory default-client branch is
* exercisable without an ANTHROPIC_API_KEY or a real API call.
* When `deps.client` is provided, this is unused.
*/
makeAnthropic?: () => Anthropic;
/** Config (MCP, brain, etc.). Defaults to loadConfig(). */
config?: GBrainConfig;
/** Rate-lease key. Defaults to `anthropic:messages`. */
rateLeaseKey?: string;
/** Max concurrent inflight calls on that key. Defaults to GBRAIN_ANTHROPIC_MAX_INFLIGHT or 8. */
maxConcurrent?: number;
/** Lease TTL. Defaults to 120s. */
leaseTtlMs?: number;
/**
* Override tool registry. When omitted, buildBrainTools is called with
* the caller's subagentId at dispatch time.
*/
toolRegistry?: ToolDef[];
}
// ── Types for internal state ────────────────────────────────
interface PersistedMessage {
message_idx: number;
role: 'user' | 'assistant';
content_blocks: ContentBlock[];
tokens_in: number | null;
tokens_out: number | null;
tokens_cache_read: number | null;
tokens_cache_create: number | null;
model: string | null;
}
interface PersistedToolExec {
message_idx: number;
tool_use_id: string;
tool_name: string;
input: unknown;
status: 'pending' | 'complete' | 'failed';
output: unknown;
error: string | null;
}
// ── Public handler factory ──────────────────────────────────
/**
* Build a subagent handler bound to a specific engine. `registerBuiltin
* Handlers` wires this up as `worker.register('subagent', handler)` at
* worker startup. Always registered `ANTHROPIC_API_KEY` is the natural
* cost gate and `PROTECTED_JOB_NAMES` gates submission.
*/
export function makeSubagentHandler(deps: SubagentDeps) {
const engine = deps.engine;
// sdk.messages IS the MessagesClient-shaped object. The v0.16.0 bug was
// casting new Anthropic() (top level) to MessagesClient, but .create()
// lives at sdk.messages.create. Assigning sdk.messages directly gets the
// right object; JS method-call semantics preserve `this` at the call
// site (subagent.ts invokes client.create(...) with client === sdk.messages).
const makeAnthropic = deps.makeAnthropic ?? (() => new Anthropic());
const client: MessagesClient = deps.client ?? makeAnthropic().messages;
const config = deps.config ?? loadConfig() ?? ({ engine: 'postgres' } as GBrainConfig);
const rateLeaseKey = deps.rateLeaseKey ?? DEFAULT_RATE_KEY;
const maxConcurrent = deps.maxConcurrent ?? DEFAULT_MAX_CONCURRENT;
const leaseTtlMs = deps.leaseTtlMs ?? DEFAULT_LEASE_TTL_MS;
return async function subagentHandler(ctx: MinionJobContext): Promise<SubagentResult> {
const data = (ctx.data ?? {}) as unknown as SubagentHandlerData;
if (!data.prompt || typeof data.prompt !== 'string') {
throw new Error('subagent job data.prompt is required (string)');
}
const model = data.model ?? DEFAULT_MODEL;
const maxTurns = data.max_turns ?? DEFAULT_MAX_TURNS;
const systemPrompt = data.system ?? DEFAULT_SYSTEM;
// Build the tool registry bound to THIS job as the owning subagent.
const registry = deps.toolRegistry ?? buildBrainTools({
subagentId: ctx.id,
engine,
config,
});
const toolDefs = data.allowed_tools && data.allowed_tools.length > 0
? filterAllowedTools(registry, data.allowed_tools)
: registry;
logSubagentSubmission({
caller: 'worker',
remote: true,
job_id: ctx.id,
model,
tools_count: toolDefs.length,
allowed_tools: toolDefs.map(t => t.name),
});
// ── Load prior state (replay) ───────────────────────────
const priorMessages = await loadPriorMessages(engine, ctx.id);
const priorTools = await loadPriorTools(engine, ctx.id);
const priorToolByUseId = new Map(priorTools.map(t => [t.tool_use_id, t]));
// Rebuild the Anthropic messages array from persisted rows.
const anthroMessages: Anthropic.MessageParam[] = priorMessages.length > 0
? priorMessages.map(m => ({ role: m.role, content: m.content_blocks as any }))
: [{ role: 'user', content: data.prompt }];
// If we had no prior messages, persist the seed user message.
let nextMessageIdx = priorMessages.length;
if (priorMessages.length === 0) {
await persistMessage(engine, ctx.id, {
message_idx: 0,
role: 'user',
content_blocks: [{ type: 'text', text: data.prompt }],
tokens_in: null,
tokens_out: null,
tokens_cache_read: null,
tokens_cache_create: null,
model: null,
});
nextMessageIdx = 1;
}
// Token rollup.
const tokenTotals = { in: 0, out: 0, cache_read: 0, cache_create: 0 };
for (const m of priorMessages) {
if (m.tokens_in) tokenTotals.in += m.tokens_in;
if (m.tokens_out) tokenTotals.out += m.tokens_out;
if (m.tokens_cache_read) tokenTotals.cache_read += m.tokens_cache_read;
if (m.tokens_cache_create) tokenTotals.cache_create += m.tokens_cache_create;
}
// Count assistant messages already persisted toward max_turns.
let assistantTurns = priorMessages.filter(m => m.role === 'assistant').length;
// ── Replay reconciliation ───────────────────────────────
//
// If the last persisted message is an assistant with tool_use blocks
// AND no subsequent user message has been synthesized yet, we crashed
// mid-tool-dispatch. Finish those tools now so the next LLM call sees
// a consistent conversation.
const last = priorMessages[priorMessages.length - 1];
if (last && last.role === 'assistant') {
const pendingToolUses = last.content_blocks.filter(
(b): b is { type: 'tool_use'; id: string; name: string; input: unknown } & Record<string, unknown> =>
b.type === 'tool_use',
);
if (pendingToolUses.length > 0) {
const synthesizedResults: ContentBlock[] = [];
for (const use of pendingToolUses) {
const prior = priorToolByUseId.get(use.id);
if (prior?.status === 'complete') {
synthesizedResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: asStringIfNotObject(prior.output),
} as ContentBlock);
continue;
}
if (prior?.status === 'failed') {
synthesizedResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: prior.error ?? 'tool failed',
is_error: true,
} as ContentBlock);
continue;
}
// pending or no row yet — try to dispatch.
const toolDef = toolDefs.find(t => t.name === use.name);
if (!toolDef) {
await persistToolExecFailed(
engine, ctx.id, last.message_idx, use.id, use.name, use.input,
`tool "${use.name}" is not in the registry for this subagent`,
);
synthesizedResults.push({
type: 'tool_result', tool_use_id: use.id,
content: `tool "${use.name}" is not available`, is_error: true,
} as ContentBlock);
continue;
}
if (prior?.status === 'pending' && !toolDef.idempotent) {
throw new Error(`non-idempotent tool "${use.name}" pending on resume; cannot safely re-run`);
}
await persistToolExecPending(engine, ctx.id, last.message_idx, use.id, use.name, use.input);
try {
const output = await toolDef.execute(use.input, {
engine, jobId: ctx.id, remote: true, signal: ctx.signal,
});
await persistToolExecComplete(engine, ctx.id, use.id, output);
synthesizedResults.push({
type: 'tool_result', tool_use_id: use.id,
content: asStringIfNotObject(output),
} as ContentBlock);
} catch (e) {
const errText = e instanceof Error ? (e.stack ?? e.message) : String(e);
await persistToolExecFailed(engine, ctx.id, last.message_idx, use.id, use.name, use.input, errText);
synthesizedResults.push({
type: 'tool_result', tool_use_id: use.id,
content: errText, is_error: true,
} as ContentBlock);
}
}
// Persist the synthesized user turn so next-resume picks up here.
const userIdx = nextMessageIdx++;
await persistMessage(engine, ctx.id, {
message_idx: userIdx,
role: 'user',
content_blocks: synthesizedResults,
tokens_in: null, tokens_out: null, tokens_cache_read: null, tokens_cache_create: null, model: null,
});
anthroMessages.push({ role: 'user', content: synthesizedResults as any });
}
}
// ── Main loop ───────────────────────────────────────────
let stopReason: SubagentStopReason = 'error';
let finalText = '';
while (true) {
if (assistantTurns >= maxTurns) {
stopReason = 'max_turns';
break;
}
if (ctx.signal.aborted || ctx.shutdownSignal.aborted) {
stopReason = 'error';
throw new Error('subagent aborted before turn');
}
// 1. Acquire rate lease for the outbound call.
const lease = await acquireLease(engine, rateLeaseKey, ctx.id, maxConcurrent, { ttlMs: leaseTtlMs });
if (!lease.acquired) {
// No slots — treat as a renewable error so the worker re-claims
// the job later. Don't fail terminally.
throw new RateLeaseUnavailableError(rateLeaseKey, lease.activeCount, lease.maxConcurrent);
}
let assistantMsg: Anthropic.Message;
const turnIdx = assistantTurns;
const t0 = Date.now();
logSubagentHeartbeat({ job_id: ctx.id, event: 'llm_call_started', turn_idx: turnIdx });
// Renewal is short-lived; for single-call turns the initial TTL
// covers the whole request. A mid-call renewal loop would add
// complexity; for v0.15 we lean on the 120s TTL + abort-on-signal.
try {
const params: Anthropic.MessageCreateParamsNonStreaming = {
model,
max_tokens: 4096,
system: [
{ type: 'text', text: systemPrompt, cache_control: { type: 'ephemeral' } },
] as any,
messages: anthroMessages,
...(toolDefs.length > 0
? {
tools: toolDefs.map((t, i) => {
const def: any = {
name: t.name,
description: t.description,
input_schema: t.input_schema,
};
// Cache only the last tool def — Anthropic treats cache_control
// as "cache everything up to and including this block".
if (i === toolDefs.length - 1) def.cache_control = { type: 'ephemeral' };
return def;
}),
}
: {}),
};
const combinedSignal = mergeSignals(ctx.signal, ctx.shutdownSignal);
assistantMsg = await client.create(params, { signal: combinedSignal });
} catch (err) {
// Release lease eagerly on error so we don't starve capacity.
await releaseLease(engine, lease.leaseId!).catch(() => {});
throw err;
}
// 2. Release lease as soon as the call returns. Tool execution runs
// outside the lease — tool calls use their own capacity.
await releaseLease(engine, lease.leaseId!).catch(() => {});
const ms = Date.now() - t0;
const inTokens = assistantMsg.usage?.input_tokens ?? 0;
const outTokens = assistantMsg.usage?.output_tokens ?? 0;
const cacheRead = (assistantMsg.usage as any)?.cache_read_input_tokens ?? 0;
const cacheCreate = (assistantMsg.usage as any)?.cache_creation_input_tokens ?? 0;
tokenTotals.in += inTokens;
tokenTotals.out += outTokens;
tokenTotals.cache_read += cacheRead;
tokenTotals.cache_create += cacheCreate;
logSubagentHeartbeat({
job_id: ctx.id,
event: 'llm_call_completed',
turn_idx: turnIdx,
ms_elapsed: ms,
tokens: { in: inTokens, out: outTokens, cache_read: cacheRead, cache_create: cacheCreate },
});
// Update job-level token rollup (best-effort; may throw if lock lost).
await ctx.updateTokens({
input: inTokens,
output: outTokens,
cache_read: cacheRead,
});
const blocks = assistantMsg.content as ContentBlock[];
// 3. Persist the assistant message BEFORE tool dispatch so replay
// sees a consistent state.
const assistantIdx = nextMessageIdx++;
await persistMessage(engine, ctx.id, {
message_idx: assistantIdx,
role: 'assistant',
content_blocks: blocks,
tokens_in: inTokens,
tokens_out: outTokens,
tokens_cache_read: cacheRead,
tokens_cache_create: cacheCreate,
model,
});
anthroMessages.push({ role: 'assistant', content: blocks as any });
assistantTurns++;
// 4. Collect tool_use blocks. If none, we're done.
const toolUses = blocks.filter(
(b): b is { type: 'tool_use'; id: string; name: string; input: unknown } & Record<string, unknown> =>
b.type === 'tool_use',
);
if (toolUses.length === 0) {
stopReason = 'end_turn';
// Concatenate text blocks as the final answer.
finalText = blocks
.filter(b => b.type === 'text' && typeof b.text === 'string')
.map(b => b.text as string)
.join('\n');
break;
}
// 5. Dispatch each tool_use. Two-phase persist (pending → complete/failed).
const toolResults: ContentBlock[] = [];
for (const use of toolUses) {
if (ctx.signal.aborted || ctx.shutdownSignal.aborted) {
throw new Error('subagent aborted during tool dispatch');
}
const toolName = use.name;
const toolDef = toolDefs.find(t => t.name === toolName);
if (!toolDef) {
// Model called a tool we didn't expose. Mark execution failed
// with a clear error and feed the error back in the next turn.
await persistToolExecFailed(
engine, ctx.id, assistantIdx, use.id, toolName, use.input,
`tool "${toolName}" is not in the registry for this subagent`,
);
toolResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: `tool "${toolName}" is not available`,
is_error: true,
} as ContentBlock);
logSubagentHeartbeat({
job_id: ctx.id,
event: 'tool_failed',
turn_idx: turnIdx,
tool_name: toolName,
error: 'not in registry',
});
continue;
}
// Replay: if we already have a row for this tool_use_id, trust it
// unless status='pending' and the tool is idempotent (re-run).
const prior = priorToolByUseId.get(use.id);
if (prior && prior.status === 'complete') {
toolResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: asStringIfNotObject(prior.output),
} as ContentBlock);
continue;
}
if (prior && prior.status === 'failed') {
toolResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: prior.error ?? 'tool failed',
is_error: true,
} as ContentBlock);
continue;
}
if (prior && prior.status === 'pending' && !toolDef.idempotent) {
// Non-idempotent and we don't know the outcome — fail the job.
throw new Error(`non-idempotent tool "${toolName}" pending on resume; cannot safely re-run`);
}
// Fresh or idempotent-replay dispatch.
await persistToolExecPending(engine, ctx.id, assistantIdx, use.id, toolName, use.input);
logSubagentHeartbeat({ job_id: ctx.id, event: 'tool_called', turn_idx: turnIdx, tool_name: toolName });
const toolStart = Date.now();
try {
const output = await toolDef.execute(use.input, {
engine,
jobId: ctx.id,
remote: true,
signal: ctx.signal,
});
await persistToolExecComplete(engine, ctx.id, use.id, output);
logSubagentHeartbeat({
job_id: ctx.id,
event: 'tool_result',
turn_idx: turnIdx,
tool_name: toolName,
ms_elapsed: Date.now() - toolStart,
});
toolResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: asStringIfNotObject(output),
} as ContentBlock);
} catch (e) {
const errText = e instanceof Error
? (e.stack ?? e.message)
: String(e);
await persistToolExecFailed(engine, ctx.id, assistantIdx, use.id, toolName, use.input, errText);
logSubagentHeartbeat({
job_id: ctx.id,
event: 'tool_failed',
turn_idx: turnIdx,
tool_name: toolName,
ms_elapsed: Date.now() - toolStart,
error: errText,
});
toolResults.push({
type: 'tool_result',
tool_use_id: use.id,
content: errText,
is_error: true,
} as ContentBlock);
}
}
// 6. Append the synthesized user turn (tool_result wrappers) to the
// conversation and persist it so replay picks it up.
const userIdx = nextMessageIdx++;
await persistMessage(engine, ctx.id, {
message_idx: userIdx,
role: 'user',
content_blocks: toolResults,
tokens_in: null,
tokens_out: null,
tokens_cache_read: null,
tokens_cache_create: null,
model: null,
});
anthroMessages.push({ role: 'user', content: toolResults as any });
}
return {
result: finalText,
turns_count: assistantTurns,
stop_reason: stopReason,
tokens: tokenTotals,
};
};
}
// ── Internal: persistence ───────────────────────────────────
async function loadPriorMessages(engine: BrainEngine, jobId: number): Promise<PersistedMessage[]> {
const rows = await engine.executeRaw<Record<string, unknown>>(
`SELECT message_idx, role, content_blocks, tokens_in, tokens_out,
tokens_cache_read, tokens_cache_create, model
FROM subagent_messages
WHERE job_id = $1
ORDER BY message_idx ASC`,
[jobId],
);
return rows.map(r => ({
message_idx: r.message_idx as number,
role: r.role as 'user' | 'assistant',
content_blocks: (typeof r.content_blocks === 'string'
? JSON.parse(r.content_blocks as string)
: r.content_blocks) as ContentBlock[],
tokens_in: (r.tokens_in as number) ?? null,
tokens_out: (r.tokens_out as number) ?? null,
tokens_cache_read: (r.tokens_cache_read as number) ?? null,
tokens_cache_create: (r.tokens_cache_create as number) ?? null,
model: (r.model as string) ?? null,
}));
}
async function loadPriorTools(engine: BrainEngine, jobId: number): Promise<PersistedToolExec[]> {
const rows = await engine.executeRaw<Record<string, unknown>>(
`SELECT message_idx, tool_use_id, tool_name, input, status, output, error
FROM subagent_tool_executions
WHERE job_id = $1`,
[jobId],
);
return rows.map(r => ({
message_idx: r.message_idx as number,
tool_use_id: r.tool_use_id as string,
tool_name: r.tool_name as string,
input: typeof r.input === 'string' ? JSON.parse(r.input) : r.input,
status: r.status as 'pending' | 'complete' | 'failed',
output: r.output == null
? null
: (typeof r.output === 'string' ? JSON.parse(r.output) : r.output),
error: (r.error as string) ?? null,
}));
}
async function persistMessage(engine: BrainEngine, jobId: number, msg: PersistedMessage): Promise<void> {
await engine.executeRaw(
`INSERT INTO subagent_messages (job_id, message_idx, role, content_blocks,
tokens_in, tokens_out, tokens_cache_read, tokens_cache_create, model)
VALUES ($1, $2, $3, $4::jsonb, $5, $6, $7, $8, $9)
ON CONFLICT (job_id, message_idx) DO NOTHING`,
[
jobId,
msg.message_idx,
msg.role,
JSON.stringify(msg.content_blocks),
msg.tokens_in,
msg.tokens_out,
msg.tokens_cache_read,
msg.tokens_cache_create,
msg.model,
],
);
}
async function persistToolExecPending(
engine: BrainEngine,
jobId: number,
messageIdx: number,
toolUseId: string,
toolName: string,
input: unknown,
): Promise<void> {
await engine.executeRaw(
`INSERT INTO subagent_tool_executions (job_id, message_idx, tool_use_id, tool_name, input, status)
VALUES ($1, $2, $3, $4, $5::jsonb, 'pending')
ON CONFLICT (job_id, tool_use_id) DO NOTHING`,
[jobId, messageIdx, toolUseId, toolName, JSON.stringify(input)],
);
}
async function persistToolExecComplete(
engine: BrainEngine,
jobId: number,
toolUseId: string,
output: unknown,
): Promise<void> {
await engine.executeRaw(
`UPDATE subagent_tool_executions
SET status = 'complete', output = $3::jsonb, ended_at = now()
WHERE job_id = $1 AND tool_use_id = $2`,
[jobId, toolUseId, JSON.stringify(output)],
);
}
async function persistToolExecFailed(
engine: BrainEngine,
jobId: number,
messageIdx: number,
toolUseId: string,
toolName: string,
input: unknown,
error: string,
): Promise<void> {
// INSERT-or-UPDATE to failed — covers both "no pending row yet" (tool
// rejected upfront) and "pending row exists" (tool threw mid-execute).
await engine.executeRaw(
`INSERT INTO subagent_tool_executions (job_id, message_idx, tool_use_id, tool_name, input, status, error, ended_at)
VALUES ($1, $2, $3, $4, $5::jsonb, 'failed', $6, now())
ON CONFLICT (job_id, tool_use_id) DO UPDATE
SET status = 'failed', error = EXCLUDED.error, ended_at = now()`,
[jobId, messageIdx, toolUseId, toolName, JSON.stringify(input), error],
);
}
// ── Internal: helpers ───────────────────────────────────────
function asStringIfNotObject(value: unknown): string {
if (typeof value === 'string') return value;
try {
return JSON.stringify(value);
} catch {
return String(value);
}
}
/**
* Merge two AbortSignals into one. Fires when either source aborts. No-op
* polyfill when AbortSignal.any isn't available yet (Node 20 has it).
*/
function mergeSignals(a: AbortSignal, b: AbortSignal): AbortSignal {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const anyFn = (AbortSignal as any).any;
if (typeof anyFn === 'function') return anyFn([a, b]) as AbortSignal;
// Manual merge.
const ac = new AbortController();
if (a.aborted || b.aborted) ac.abort();
else {
a.addEventListener('abort', () => ac.abort(), { once: true });
b.addEventListener('abort', () => ac.abort(), { once: true });
}
return ac.signal;
}
/**
* Error thrown when acquireLease returns acquired=false. The worker
* treats this as a renewable error job goes back to waiting with
* backoff, no terminal fail.
*/
export class RateLeaseUnavailableError extends Error {
constructor(public key: string, public active: number, public max: number) {
super(`rate lease "${key}" full (${active}/${max})`);
this.name = 'RateLeaseUnavailableError';
}
}
// ── Testing surface ─────────────────────────────────────────
export const __testing = {
loadPriorMessages,
loadPriorTools,
persistMessage,
persistToolExecPending,
persistToolExecComplete,
persistToolExecFailed,
asStringIfNotObject,
DEFAULT_MODEL,
};
+235
View File
@@ -0,0 +1,235 @@
/**
* GBRAIN_PLUGIN_PATH loader for host-repo subagent definitions (v0.15).
*
* Your OpenClaw (and future downstream agents) ship custom subagent defs
* from their own repos. gbrain discovers them at worker startup via
* GBRAIN_PLUGIN_PATH = colon-separated absolute paths (like $PATH). Each
* path must contain a gbrain.plugin.json manifest describing the plugin
* and a subagents/ subdirectory holding `*.md` definition files.
*
* Path policy is strict on purpose:
* - ABSOLUTE paths only. Relative paths and `~` prefixes are rejected
* (no implicit cwd or home expansion too easy to pick up a tampered
* sibling directory).
* - Remote URLs (http://, https://, file://) rejected. Plugin loading
* must go through the filesystem so the user controls what's there.
* - Non-existent paths logged and skipped (do not fail worker startup).
*
* Collision policy: left-to-right wins. A warning goes to stderr naming
* both sides of the collision.
*
* Trust policy: plugins ship subagent *defs* only. They cannot declare
* new tools, cannot extend the brain-allowlist, cannot override
* agent-safe flags. The `allowed_tools:` frontmatter field of a subagent
* def must subset the derived registry validation happens at plugin
* load time, NOT at subagent dispatch time, so a typo in a plugin skill
* fails loudly at worker startup instead of silently disabling a tool.
*
* Manifest version (`plugin_version`) locks the contract shape. Unknown
* versions are rejected so the authoritative definition is whatever this
* version of gbrain understands.
*/
import * as fs from 'node:fs';
import * as path from 'node:path';
import matter from 'gray-matter';
export const SUPPORTED_PLUGIN_VERSION = 'gbrain-plugin-v1';
export interface PluginManifest {
name: string;
version: string;
plugin_version: string;
subagents?: string;
description?: string;
}
export interface SubagentDefinition {
/** The plugin that shipped this def. */
plugin_name: string;
/** Stable agent name used as `subagent_def` by CLI callers. */
name: string;
/** Full path to the .md file on disk, for debug surfaces. */
source_path: string;
frontmatter: Record<string, unknown>;
/** Markdown body (system prompt content). */
body: string;
/** Optional allowed_tools list (frontmatter). Subset of registry. */
allowed_tools?: string[];
}
export interface PluginLoadResult {
/** Successfully loaded plugins with their subagents. */
plugins: Array<{ manifest: PluginManifest; rootDir: string; subagents: SubagentDefinition[] }>;
/** Per-path warnings (rejected, missing, malformed) collected during load. */
warnings: string[];
}
export interface LoadOpts {
/**
* Registry names the plugin's subagent `allowed_tools` must subset. When
* present, any frontmatter entry not in this set fails the plugin load.
* Pass `undefined` to skip validation (early worker startup before the
* registry is built but production callers should always pass it).
*/
validAgentToolNames?: ReadonlySet<string>;
/** Override the PATH env (for tests). */
envPath?: string;
}
/** Public entry point: load every plugin directory from GBRAIN_PLUGIN_PATH. */
export function loadPluginsFromEnv(opts: LoadOpts = {}): PluginLoadResult {
const raw = opts.envPath ?? process.env.GBRAIN_PLUGIN_PATH ?? '';
const paths = raw.split(':').map(s => s.trim()).filter(Boolean);
const result: PluginLoadResult = { plugins: [], warnings: [] };
// Left-wins collision tracking.
const subagentByName = new Map<string, { pluginName: string; pathLeft: string }>();
for (const p of paths) {
const rejection = rejectIfNotAbsolute(p);
if (rejection) { result.warnings.push(rejection); continue; }
if (!fs.existsSync(p)) {
result.warnings.push(`[plugin-loader] path does not exist, skipping: ${p}`);
continue;
}
if (!fs.statSync(p).isDirectory()) {
result.warnings.push(`[plugin-loader] not a directory, skipping: ${p}`);
continue;
}
try {
const loaded = loadSinglePlugin(p, opts);
if ('error' in loaded) {
result.warnings.push(`[plugin-loader] rejected ${p}: ${loaded.error}`);
continue;
}
const accepted: SubagentDefinition[] = [];
for (const sa of loaded.subagents) {
const prior = subagentByName.get(sa.name);
if (prior) {
result.warnings.push(
`[plugin-loader] collision: subagent '${sa.name}' from '${loaded.manifest.name}' at ${p} ` +
`shadowed by earlier '${prior.pluginName}' at ${prior.pathLeft} (first wins)`,
);
continue;
}
subagentByName.set(sa.name, { pluginName: loaded.manifest.name, pathLeft: p });
accepted.push(sa);
}
result.plugins.push({ manifest: loaded.manifest, rootDir: p, subagents: accepted });
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
result.warnings.push(`[plugin-loader] unexpected error loading ${p}: ${msg}`);
}
}
return result;
}
function rejectIfNotAbsolute(p: string): string | null {
if (/^[a-z][a-z0-9+.-]*:\/\//i.test(p)) {
return `[plugin-loader] remote URL rejected: ${p}`;
}
if (p.startsWith('~')) {
return `[plugin-loader] ~-prefixed path rejected (expand explicitly): ${p}`;
}
if (!path.isAbsolute(p)) {
return `[plugin-loader] relative path rejected: ${p}`;
}
return null;
}
export interface LoadedPlugin {
manifest: PluginManifest;
subagents: SubagentDefinition[];
}
/**
* Load one plugin directory. Returns a union so callers can differentiate
* rejection (loud but non-fatal) from an empty plugin (fatal-ish the
* manifest parsed but contributes nothing).
*/
export function loadSinglePlugin(
rootDir: string,
opts: LoadOpts = {},
): LoadedPlugin | { error: string } {
const manifestPath = path.join(rootDir, 'gbrain.plugin.json');
if (!fs.existsSync(manifestPath)) {
return { error: 'missing gbrain.plugin.json' };
}
let manifest: PluginManifest;
try {
const raw = fs.readFileSync(manifestPath, 'utf8');
manifest = JSON.parse(raw) as PluginManifest;
} catch (e) {
return { error: `invalid manifest JSON: ${e instanceof Error ? e.message : String(e)}` };
}
if (typeof manifest.name !== 'string' || manifest.name.length === 0) {
return { error: 'manifest missing required "name" field' };
}
if (manifest.plugin_version !== SUPPORTED_PLUGIN_VERSION) {
return {
error: `unsupported plugin_version "${manifest.plugin_version}" (gbrain supports "${SUPPORTED_PLUGIN_VERSION}")`,
};
}
const subagentsDirRel = manifest.subagents ?? 'subagents';
const subagentsDir = path.resolve(rootDir, subagentsDirRel);
// Prevent `../` escape via the manifest's `subagents` field.
if (!subagentsDir.startsWith(rootDir + path.sep) && subagentsDir !== rootDir) {
return { error: `subagents path escapes plugin root: ${subagentsDirRel}` };
}
const subagents: SubagentDefinition[] = [];
if (fs.existsSync(subagentsDir) && fs.statSync(subagentsDir).isDirectory()) {
for (const entry of fs.readdirSync(subagentsDir)) {
if (!entry.endsWith('.md')) continue;
const sourcePath = path.join(subagentsDir, entry);
try {
const raw = fs.readFileSync(sourcePath, 'utf8');
const parsed = matter(raw);
const frontmatter = (parsed.data ?? {}) as Record<string, unknown>;
const body = parsed.content ?? '';
const name = typeof frontmatter.name === 'string'
? frontmatter.name
: entry.replace(/\.md$/, '');
const allowed = Array.isArray(frontmatter.allowed_tools)
? (frontmatter.allowed_tools as unknown[]).filter(x => typeof x === 'string') as string[]
: undefined;
if (allowed && opts.validAgentToolNames) {
const missing = allowed.filter(t => !opts.validAgentToolNames!.has(t));
if (missing.length > 0) {
return {
error: `subagent '${name}' allowed_tools references unknown tools: ${missing.join(', ')}`,
};
}
}
subagents.push({
plugin_name: manifest.name,
name,
source_path: sourcePath,
frontmatter,
body,
allowed_tools: allowed,
});
} catch (e) {
return { error: `could not parse ${sourcePath}: ${e instanceof Error ? e.message : String(e)}` };
}
}
}
return { manifest, subagents };
}
/** Testing surface. */
export const __testing = {
rejectIfNotAbsolute,
SUPPORTED_PLUGIN_VERSION,
};
+9 -1
View File
@@ -12,7 +12,15 @@
* pay them at module load.
*/
export const PROTECTED_JOB_NAMES: ReadonlySet<string> = new Set(['shell']);
export const PROTECTED_JOB_NAMES: ReadonlySet<string> = new Set([
'shell',
// v0.15: subagent + aggregator are protected because they call the
// Anthropic API. MCP callers can't submit them directly; only the
// `gbrain agent run` CLI path (which sets allowProtectedSubmit) or a
// trusted local `submit_job` (ctx.remote=false) can insert these rows.
'subagent',
'subagent_aggregator',
]);
/** Check a job name against the protected set. Normalizes whitespace first. */
export function isProtectedJobName(name: string): boolean {
+207 -50
View File
@@ -134,11 +134,17 @@ export class MinionQueue {
// 3. Insert child. Use ON CONFLICT for idempotency; if a concurrent submit
// raced past the fast-path SELECT, the unique index catches it here.
// v13 quiet_hours + stagger_key always present (null fallback; schema
// stores NULL). v15 max_stalled is conditional: provided values get
// clamped to [1, 100] and included in the INSERT; omitted values
// skip the column so the schema DEFAULT (5 as of v0.14.1) kicks in.
// Keeps the app layer from hardcoding the schema default constant.
// quiet_hours + stagger_key always present (null fallback; schema
// stores NULL). max_stalled is conditional: provided values get
// clamped to [1, 100] and included in the INSERT; omitted values
// skip the column so the schema DEFAULT (5 as of v0.14.1) kicks in.
// Keeps the app layer from hardcoding the schema default constant.
//
// Footgun note (codex iter 3): threading max_stalled on INSERT only is
// deliberate. An idempotency-key hit returns the EXISTING row via the
// fast-path SELECT above — we do NOT UPDATE max_stalled on a re-submit,
// because letting a second submitter mutate the first submitter's
// durability semantics is a nasty surprise.
const hasMaxStalled = opts?.max_stalled !== undefined && opts.max_stalled !== null;
const clampedMaxStalled = hasMaxStalled
? Math.max(1, Math.min(100, Math.floor(opts!.max_stalled as number)))
@@ -286,29 +292,83 @@ export class MinionQueue {
* Returns the *root* (the job matching id), not an arbitrary descendant.
*/
async cancelJob(id: number): Promise<MinionJob | null> {
const rows = await this.engine.executeRaw<Record<string, unknown>>(
`WITH RECURSIVE descendants AS (
SELECT id, 0 AS d FROM minion_jobs WHERE id = $1
UNION ALL
SELECT m.id, descendants.d + 1
FROM minion_jobs m
JOIN descendants ON m.parent_job_id = descendants.id
WHERE descendants.d < 100
)
UPDATE minion_jobs SET
status = 'cancelled',
lock_token = NULL,
lock_until = NULL,
finished_at = now(),
updated_at = now()
WHERE id IN (SELECT id FROM descendants)
AND status IN ('waiting','active','delayed','waiting-children','paused')
RETURNING *`,
[id]
);
if (rows.length === 0) return null;
const root = rows.find(r => (r.id as number) === id);
return root ? rowToMinionJob(root) : null;
return this.engine.transaction(async (tx) => {
const rows = await tx.executeRaw<Record<string, unknown>>(
`WITH RECURSIVE descendants AS (
SELECT id, 0 AS d FROM minion_jobs WHERE id = $1
UNION ALL
SELECT m.id, descendants.d + 1
FROM minion_jobs m
JOIN descendants ON m.parent_job_id = descendants.id
WHERE descendants.d < 100
)
UPDATE minion_jobs SET
status = 'cancelled',
lock_token = NULL,
lock_until = NULL,
finished_at = now(),
updated_at = now()
WHERE id IN (SELECT id FROM descendants)
AND status IN ('waiting','active','delayed','waiting-children','paused')
RETURNING *`,
[id]
);
if (rows.length === 0) return null;
// v0.15: emit child_done(outcome='cancelled') for every cancelled row
// that had a parent. Without this, an aggregator waiting for N
// child_done messages hangs forever when a child is cancelled (codex
// iteration 3). Also unblock any aggregator parents whose last
// non-terminal child we just cancelled.
const parentIds = new Set<number>();
for (const r of rows) {
const childId = r.id as number;
const parentJobId = r.parent_job_id as number | null;
const name = r.name as string;
// Skip the root if it's the caller's cancel target AND has no parent.
// Descendants whose parent got cancelled in the same sweep still
// benefit from the inbox message — their parent exits waiting-children
// via the resolve sweep below even though the parent is itself
// cancelled (EXISTS guard on inbox INSERT handles it).
if (parentJobId == null) continue;
parentIds.add(parentJobId);
const childDone: ChildDoneMessage = {
type: 'child_done',
child_id: childId,
job_name: name,
result: null,
outcome: 'cancelled',
error: 'cancelled',
};
await tx.executeRaw(
`INSERT INTO minion_inbox (job_id, sender, payload)
SELECT $1, 'minions', $2::jsonb
WHERE EXISTS (
SELECT 1 FROM minion_jobs
WHERE id = $1 AND status NOT IN ('completed','failed','dead','cancelled')
)`,
[parentJobId, childDone]
);
}
// Resolve any non-cancelled aggregator parents sitting on
// waiting-children whose last open child we just cancelled.
for (const parentId of parentIds) {
await tx.executeRaw(
`UPDATE minion_jobs SET status = 'waiting', updated_at = now()
WHERE id = $1 AND status = 'waiting-children'
AND NOT EXISTS (
SELECT 1 FROM minion_jobs
WHERE parent_job_id = $1
AND status NOT IN ('completed', 'failed', 'dead', 'cancelled')
)`,
[parentId]
);
}
const root = rows.find(r => (r.id as number) === id);
return root ? rowToMinionJob(root) : null;
});
}
/** Re-queue a failed or dead job for retry. */
@@ -440,21 +500,67 @@ export class MinionQueue {
* but will be caught the next one (after re-claim). Never double-handled.
*/
async handleTimeouts(): Promise<MinionJob[]> {
const rows = await this.engine.executeRaw<Record<string, unknown>>(
`UPDATE minion_jobs SET
status = 'dead',
error_text = 'timeout exceeded',
lock_token = NULL,
lock_until = NULL,
finished_at = now(),
updated_at = now()
WHERE status = 'active'
AND timeout_at IS NOT NULL
AND timeout_at < now()
AND lock_until > now()
RETURNING *`
);
return rows.map(rowToMinionJob);
return this.engine.transaction(async (tx) => {
const rows = await tx.executeRaw<Record<string, unknown>>(
`UPDATE minion_jobs SET
status = 'dead',
error_text = 'timeout exceeded',
lock_token = NULL,
lock_until = NULL,
finished_at = now(),
updated_at = now()
WHERE status = 'active'
AND timeout_at IS NOT NULL
AND timeout_at < now()
AND lock_until > now()
RETURNING *`
);
// v0.15: emit child_done(outcome='timeout') for every timed-out job that
// had a parent. Without this, an aggregator waiting for N child_done
// messages hangs forever when a child times out (codex iteration 3).
// Outcome 'timeout' is distinct from 'dead' so consumers can distinguish
// "timed out during run" from "died via max-stall".
const parentIds = new Set<number>();
for (const r of rows) {
const parentJobId = r.parent_job_id as number | null;
if (parentJobId == null) continue;
parentIds.add(parentJobId);
const childDone: ChildDoneMessage = {
type: 'child_done',
child_id: r.id as number,
job_name: r.name as string,
result: null,
outcome: 'timeout',
error: 'timeout exceeded',
};
await tx.executeRaw(
`INSERT INTO minion_inbox (job_id, sender, payload)
SELECT $1, 'minions', $2::jsonb
WHERE EXISTS (
SELECT 1 FROM minion_jobs
WHERE id = $1 AND status NOT IN ('completed','failed','dead','cancelled')
)`,
[parentJobId, childDone]
);
}
// Unblock any aggregator parents whose last open child we just killed.
for (const parentId of parentIds) {
await tx.executeRaw(
`UPDATE minion_jobs SET status = 'waiting', updated_at = now()
WHERE id = $1 AND status = 'waiting-children'
AND NOT EXISTS (
SELECT 1 FROM minion_jobs
WHERE parent_job_id = $1
AND status NOT IN ('completed', 'failed', 'dead', 'cancelled')
)`,
[parentId]
);
}
return rows.map(rowToMinionJob);
});
}
/**
@@ -524,6 +630,7 @@ export class MinionQueue {
child_id: completed.id,
job_name: completed.name,
result: result ?? null,
outcome: 'complete',
};
await tx.executeRaw(
`INSERT INTO minion_inbox (job_id, sender, payload)
@@ -535,14 +642,17 @@ export class MinionQueue {
[completed.parent_job_id, childDone]
);
// Fold-in resolveParent: flip parent to waiting once all children done.
// Fold-in resolveParent: flip parent to waiting once all children are
// in ANY terminal state. Terminal set includes 'failed' so a failed
// child with on_child_fail='continue'/'ignore' doesn't strand the
// parent in waiting-children forever (v0.15 aggregator fix).
await tx.executeRaw(
`UPDATE minion_jobs SET status = 'waiting', updated_at = now()
WHERE id = $1 AND status = 'waiting-children'
AND NOT EXISTS (
SELECT 1 FROM minion_jobs
WHERE parent_job_id = $1
AND status NOT IN ('completed', 'dead', 'cancelled')
AND status NOT IN ('completed', 'failed', 'dead', 'cancelled')
)`,
[completed.parent_job_id]
);
@@ -616,6 +726,29 @@ export class MinionQueue {
// Parent hook on terminal failure.
if (terminal && failed.parent_job_id) {
// v0.15: emit child_done(outcome='failed') BEFORE any parent-terminal
// update. Insertion order matters because `completeJob`'s inbox-write
// EXISTS guard skips writes once the parent is 'failed' — if we let
// the fail_parent UPDATE run first, this inbox row would be dropped
// for aggregator-style parents that still want to count it (codex).
const childDone: ChildDoneMessage = {
type: 'child_done',
child_id: failed.id,
job_name: failed.name,
result: null,
outcome: newStatus === 'dead' ? 'dead' : 'failed',
error: errorText,
};
await tx.executeRaw(
`INSERT INTO minion_inbox (job_id, sender, payload)
SELECT $1, 'minions', $2::jsonb
WHERE EXISTS (
SELECT 1 FROM minion_jobs
WHERE id = $1 AND status NOT IN ('completed','failed','dead','cancelled')
)`,
[failed.parent_job_id, childDone]
);
if (failed.on_child_fail === 'fail_parent') {
await tx.executeRaw(
`UPDATE minion_jobs SET status = 'failed',
@@ -628,19 +761,37 @@ export class MinionQueue {
`UPDATE minion_jobs SET parent_job_id = NULL, updated_at = now() WHERE id = $1`,
[failed.id]
);
// After dropping the dep, try to resolve the parent if all OTHER kids are done.
// After dropping the dep, try to resolve the parent if all OTHER
// kids are terminal. Terminal set includes 'failed' (v0.15).
await tx.executeRaw(
`UPDATE minion_jobs SET status = 'waiting', updated_at = now()
WHERE id = $1 AND status = 'waiting-children'
AND NOT EXISTS (
SELECT 1 FROM minion_jobs
WHERE parent_job_id = $1
AND status NOT IN ('completed', 'dead', 'cancelled')
AND status NOT IN ('completed', 'failed', 'dead', 'cancelled')
)`,
[failed.parent_job_id]
);
} else {
// 'ignore' / 'continue': parent stays in waiting-children waiting on
// siblings. With v0.15 terminal-set expansion + child_done emission
// above, an aggregator sibling-count model now works: all N children
// reach terminal → completeJob on a sibling (or the LAST terminal
// transition here) flips parent → waiting once no non-terminal kids
// remain. Run the resolve check here so the last child transitioning
// via THIS code path still unblocks the parent.
await tx.executeRaw(
`UPDATE minion_jobs SET status = 'waiting', updated_at = now()
WHERE id = $1 AND status = 'waiting-children'
AND NOT EXISTS (
SELECT 1 FROM minion_jobs
WHERE parent_job_id = $1
AND status NOT IN ('completed', 'failed', 'dead', 'cancelled')
)`,
[failed.parent_job_id]
);
}
// 'ignore' / 'continue' → parent stays in waiting-children waiting on siblings
}
// remove_on_fail cleanup AFTER parent hook.
@@ -725,7 +876,13 @@ export class MinionQueue {
return { requeued, dead };
}
/** Check if all children of a parent are done. If so, unblock parent. */
/**
* Check if all children of a parent are in ANY terminal state. If so,
* unblock parent (flip waiting-children waiting).
*
* v0.15: terminal set includes 'failed' so a child failing with
* on_child_fail='continue'/'ignore' doesn't strand the parent.
*/
async resolveParent(parentId: number): Promise<MinionJob | null> {
const rows = await this.engine.executeRaw<Record<string, unknown>>(
`UPDATE minion_jobs SET status = 'waiting', updated_at = now()
@@ -733,7 +890,7 @@ export class MinionQueue {
AND NOT EXISTS (
SELECT 1 FROM minion_jobs
WHERE parent_job_id = $1
AND status NOT IN ('completed', 'dead', 'cancelled')
AND status NOT IN ('completed', 'failed', 'dead', 'cancelled')
)
RETURNING *`,
[parentId]
+152
View File
@@ -0,0 +1,152 @@
/**
* Lease-based rate limiter for outbound providers (e.g. anthropic:messages).
*
* Counter-based limiters leak capacity when a worker crashes mid-call
* (counter never decrements). Leases are owner-tagged rows with an expires_at
* timestamp crash recovery is free: any row past expires_at is considered
* dead on the next acquire and pruned before the active-count check.
*
* Two-phase acquire:
* 1. Pre-prune: DELETE expired leases for this key (same txn).
* 2. Check-then-insert under a txn-scoped advisory lock so two concurrent
* acquires can't both see "one slot left".
*
* The owner is always a Minion job id; the lease is CASCADE-tied to
* minion_jobs so an out-of-band row DELETE (prune, cancel) doesn't leave
* stale leases. Mid-call renewal bumps expires_at in-place.
*/
import type { BrainEngine } from '../engine.ts';
/**
* Acquisition result. If `acquired=false`, the caller should back off and
* retry we don't queue, we just reject.
*/
export interface LeaseAcquireResult {
acquired: boolean;
/** The lease row id, present only when acquired=true. */
leaseId?: number;
/** Active count seen at acquire time (for diagnostics). */
activeCount: number;
/** max_concurrent that was checked against. */
maxConcurrent: number;
}
/**
* Convert a key string to a stable int64 for pg_advisory_xact_lock. Simple
* FNV-1a is fine the lock space is per-transaction and we only need
* different keys to (usually) hash to different locks.
*/
function hashKey(key: string): bigint {
// FNV-1a 64-bit
let h = 0xcbf29ce484222325n;
const prime = 0x100000001b3n;
for (let i = 0; i < key.length; i++) {
h ^= BigInt(key.charCodeAt(i));
h = (h * prime) & 0xffffffffffffffffn;
}
// Fit into signed int64 for PG bigint. The high bit gets clipped in the
// arithmetic above already, but be explicit.
const signBit = 0x8000000000000000n;
return h & (signBit - 1n);
}
const DEFAULT_TTL_MS = 120_000;
export interface AcquireOpts {
ttlMs?: number;
}
/**
* Attempt to acquire a lease on `key`. Returns `{acquired: false}` when the
* active count (after pre-pruning stale rows) would exceed maxConcurrent.
*
* The call MUST run inside a transaction for the advisory lock + insert to
* be atomic. Pass in the engine the helper wraps the txn internally.
*/
export async function acquireLease(
engine: BrainEngine,
key: string,
ownerJobId: number,
maxConcurrent: number,
opts: AcquireOpts = {},
): Promise<LeaseAcquireResult> {
const ttlMs = opts.ttlMs ?? DEFAULT_TTL_MS;
const lockKey = hashKey(key);
return engine.transaction(async (tx) => {
// txn-scoped advisory lock keyed on the rate-lease key name. Released
// automatically when the txn commits/rolls back.
await tx.executeRaw(`SELECT pg_advisory_xact_lock($1::bigint)`, [lockKey.toString()]);
// Pre-prune stale leases for this key.
await tx.executeRaw(
`DELETE FROM subagent_rate_leases WHERE key = $1 AND expires_at <= now()`,
[key],
);
const countRows = await tx.executeRaw<{ count: string | number }>(
`SELECT count(*)::text AS count FROM subagent_rate_leases WHERE key = $1`,
[key],
);
const activeCount = parseInt(String(countRows[0]?.count ?? '0'), 10);
if (activeCount >= maxConcurrent) {
return { acquired: false, activeCount, maxConcurrent };
}
const rows = await tx.executeRaw<{ id: number }>(
`INSERT INTO subagent_rate_leases (key, owner_job_id, expires_at)
VALUES ($1, $2, now() + ($3::double precision * interval '1 millisecond'))
RETURNING id`,
[key, ownerJobId, ttlMs],
);
const leaseId = rows[0]!.id;
return { acquired: true, leaseId, activeCount: activeCount + 1, maxConcurrent };
});
}
/**
* Renew a lease's expires_at (mid-call). Returns true if the lease still
* exists (was renewed), false if it was pruned (caller must re-acquire or
* abort).
*/
export async function renewLease(engine: BrainEngine, leaseId: number, ttlMs = DEFAULT_TTL_MS): Promise<boolean> {
const rows = await engine.executeRaw<{ id: number }>(
`UPDATE subagent_rate_leases
SET expires_at = now() + ($2::double precision * interval '1 millisecond')
WHERE id = $1
RETURNING id`,
[leaseId, ttlMs],
);
return rows.length > 0;
}
/**
* Release a lease explicitly. Idempotent a missing lease returns silently
* (it was pruned or the owning job row cascade-deleted it).
*/
export async function releaseLease(engine: BrainEngine, leaseId: number): Promise<void> {
await engine.executeRaw(`DELETE FROM subagent_rate_leases WHERE id = $1`, [leaseId]);
}
/**
* Attempt to renew with 3x exponential backoff (250ms / 500ms / 1s). Used
* mid-LLM-call when the first renewal attempt hits a DB blip. On all-three
* failure the caller must abort with a renewable error so the worker
* re-claims the job.
*/
export async function renewLeaseWithBackoff(engine: BrainEngine, leaseId: number, ttlMs = DEFAULT_TTL_MS): Promise<boolean> {
const delays = [0, 250, 500, 1000]; // first attempt immediate, then 250/500/1000
for (const delay of delays) {
if (delay > 0) await new Promise(r => setTimeout(r, delay));
try {
if (await renewLease(engine, leaseId, ttlMs)) return true;
// Lease is gone (pruned). No point retrying — caller must abort.
return false;
} catch {
// DB blip; fall through to next delay.
}
}
return false;
}
+225
View File
@@ -0,0 +1,225 @@
/**
* Derive the subagent brain-tool registry from src/core/operations.ts.
*
* Single source of truth: the MCP server already maps OPERATIONS tool defs.
* We reuse the same ParamDef-shape JSONSchema conversion (lives in
* buildToolDefs for MCP) and wrap each allowed op with an execute() that
* invokes its handler under a subagent-tagged OperationContext.
*
* Filtering is NAME-based (not by OperationContext.remote, which is a
* call-time flag, not operation metadata codex catch). The allow-list
* below is reviewed manually; adding a new op here is an explicit security
* decision.
*
* put_page: allowed, but the subagent tool-schema wraps its `slug` with a
* per-subagent namespace regex so the model can only write under
* `wiki/agents/<subagentId>/...`. The put_page operation also has a server-
* side fail-closed check (see src/core/operations.ts) that catches any
* dispatcher bug where viaSubagent=true but subagentId is missing.
*
* In v0.15 every allow-list op is treated as idempotent for the two-phase
* replay path. put_page with a deterministic slug is idempotent at the row
* level; repeats re-derive the same embedding over identical content.
*/
import type { BrainEngine } from '../../engine.ts';
import type { GBrainConfig } from '../../config.ts';
import { operations } from '../../operations.ts';
import type { Operation, OperationContext } from '../../operations.ts';
import type { ToolCtx, ToolDef } from '../types.ts';
/**
* v0.15 brain-tool allow-list. Review carefully when extending. Op names
* verified against origin/master:src/core/operations.ts (post shell-jobs +
* Knowledge Runtime).
*
* Read-only (all safe):
* query, search, get_page, list_pages, file_list, file_url,
* get_backlinks, traverse_graph, resolve_slugs, get_ingest_log
*
* Conditional write:
* put_page (namespace-enforced by the tool schema + server-side check)
*
* Every name below MUST exist in src/core/operations.ts OPERATIONS; the
* brain-allowlist test pins this invariant so an upstream rename fails CI
* instead of silently dropping a tool.
*/
export const BRAIN_TOOL_ALLOWLIST: ReadonlySet<string> = new Set([
'query',
'search',
'get_page',
'list_pages',
'file_list',
'file_url',
'get_backlinks',
'traverse_graph',
'resolve_slugs',
'get_ingest_log',
'put_page',
]);
/** Matches Anthropic's tool-name constraint. No dots. */
const ANTHROPIC_NAME_RE = /^[a-zA-Z0-9_-]{1,64}$/;
function sanitizeToolName(opName: string): string {
// Prefix with brain_ and replace any non-conforming char. For the v0.15
// allow-list, every op name is already a valid simple identifier, so this
// is defense-in-depth.
const prefixed = `brain_${opName}`.replace(/[^a-zA-Z0-9_-]/g, '_');
return prefixed.slice(0, 64);
}
/**
* Convert an Operation.params (ParamDef) map to an Anthropic-compatible
* JSONSchema.input_schema. Same shape MCP uses inline ParamDef.type
* narrows to a subset of JSONSchema types.
*/
function paramsToInputSchema(op: Operation): Record<string, unknown> {
return {
type: 'object' as const,
properties: Object.fromEntries(
Object.entries(op.params).map(([k, v]) => [k, {
type: v.type === 'array' ? 'array' : v.type,
...(v.description ? { description: v.description } : {}),
...(v.enum ? { enum: v.enum } : {}),
...(v.items ? { items: { type: v.items.type } } : {}),
}]),
),
required: Object.entries(op.params).filter(([, v]) => v.required).map(([k]) => k),
};
}
/**
* For put_page specifically, the tool schema shown to the model constrains
* `slug` to `wiki/agents/<subagentId>/...`. The server-side check in
* operations.ts is the authoritative gate; this just helps the model write
* correct slugs on the first try.
*/
function namespacedPutPageSchema(op: Operation, subagentId: number): Record<string, unknown> {
const base = paramsToInputSchema(op);
const props = (base.properties as Record<string, Record<string, unknown>>) ?? {};
if (props.slug) {
props.slug = {
...props.slug,
description: `Page slug. MUST start with "wiki/agents/${subagentId}/" (agents can only write under their own namespace).`,
pattern: `^wiki/agents/${subagentId}/.+`,
};
}
return { ...base, properties: props };
}
/** Args required to build the registry for a given subagent job. */
export interface BuildBrainToolsOpts {
subagentId: number;
engine: BrainEngine;
config: GBrainConfig;
/** Optional filter: only include names in this set. */
allowedNames?: ReadonlySet<string>;
}
interface OpContextDeps {
engine: BrainEngine;
config: GBrainConfig;
subagentId: number;
jobId: number;
signal?: AbortSignal;
}
function buildOpContext(deps: OpContextDeps): OperationContext {
return {
engine: deps.engine,
config: deps.config,
logger: {
info: (msg: string) => process.stderr.write(`[subagent-tool:${deps.jobId}] ${msg}\n`),
warn: (msg: string) => process.stderr.write(`[subagent-tool:${deps.jobId}] WARN: ${msg}\n`),
error: (msg: string) => process.stderr.write(`[subagent-tool:${deps.jobId}] ERROR: ${msg}\n`),
},
dryRun: false,
remote: true, // match MCP trust boundary
jobId: deps.jobId,
subagentId: deps.subagentId,
viaSubagent: true, // FAIL-CLOSED: put_page etc. enforce namespace
};
}
/**
* Build the subagent brain-tool registry. One ToolDef per allow-listed op,
* with a namespace-wrapped schema for put_page.
*
* Call this once per subagent-job claim; the registry is keyed to the job's
* subagentId + engine handle, so it's not shareable across jobs.
*/
export function buildBrainTools(opts: BuildBrainToolsOpts): ToolDef[] {
const filter = opts.allowedNames ?? BRAIN_TOOL_ALLOWLIST;
const picked: Operation[] = operations.filter(
op => BRAIN_TOOL_ALLOWLIST.has(op.name) && filter.has(op.name),
);
return picked.map<ToolDef>(op => {
const schema = op.name === 'put_page'
? namespacedPutPageSchema(op, opts.subagentId)
: paramsToInputSchema(op);
const toolName = sanitizeToolName(op.name);
if (!ANTHROPIC_NAME_RE.test(toolName)) {
throw new Error(`brain tool name ${toolName} does not match Anthropic constraint`);
}
return {
name: toolName,
description: op.description,
input_schema: schema,
// v0.15 ships only idempotent brain tools (every allow-listed op is
// deterministic over its input; put_page re-writes the same slug).
idempotent: true,
async execute(input: unknown, ctx: ToolCtx): Promise<unknown> {
const opCtx = buildOpContext({
engine: ctx.engine,
config: opts.config,
subagentId: opts.subagentId,
jobId: ctx.jobId,
signal: ctx.signal,
});
const params = (input && typeof input === 'object') ? input as Record<string, unknown> : {};
return op.handler(opCtx, params);
},
};
});
}
/**
* Apply the caller's `allowed_tools` subset to a registry. Unknown tool
* names throw a clear error at load time (NOT silently ignored) so
* subagent defs with a typo don't ship to prod wondering why a tool
* never fires.
*/
export function filterAllowedTools(registry: ToolDef[], allowedToolNames: string[]): ToolDef[] {
const indexByName = new Map(registry.map(t => [t.name, t]));
// Also index by the un-prefixed op name (for friendlier allowed_tools entries).
const indexByShort = new Map(
registry.map(t => [t.name.replace(/^brain_/, ''), t]),
);
const seen = new Set<string>();
const picked: ToolDef[] = [];
for (const requested of allowedToolNames) {
const match = indexByName.get(requested) ?? indexByShort.get(requested);
if (!match) {
throw new Error(
`subagent allowed_tools references unknown tool "${requested}". ` +
`Known: ${[...indexByName.keys()].join(', ')}`,
);
}
if (seen.has(match.name)) continue;
seen.add(match.name);
picked.push(match);
}
return picked;
}
/** Exported for unit tests (stable surface). */
export const __testing = {
sanitizeToolName,
paramsToInputSchema,
namespacedPutPageSchema,
ANTHROPIC_NAME_RE,
};
+229
View File
@@ -0,0 +1,229 @@
/**
* Render a subagent conversation to markdown.
*
* Two inputs:
* - subagent_messages rows (persisted Anthropic message-block arrays)
* - subagent_tool_executions rows (two-phase tool ledger used to show
* tool outputs alongside the model's tool_use calls)
*
* The output is suitable for:
* - an attachment on the completed subagent job row
* - inline display in `gbrain agent logs <job>` after the heartbeat stream
* - committing as a brain page under wiki/agents/<subagentId>/transcript-*
*
* Does NOT redact anything the caller writes to a location they control.
* For PII-sensitive deployments, pass through a sanitizer before persisting.
*/
import type { BrainEngine } from '../engine.ts';
import type { ContentBlock } from './types.ts';
export interface SubagentMessageRow {
id: number;
job_id: number;
message_idx: number;
role: 'user' | 'assistant';
content_blocks: ContentBlock[];
tokens_in: number | null;
tokens_out: number | null;
tokens_cache_read: number | null;
tokens_cache_create: number | null;
model: string | null;
ended_at: Date;
}
export interface SubagentToolExecRow {
id: number;
job_id: number;
message_idx: number;
tool_use_id: string;
tool_name: string;
input: unknown;
status: 'pending' | 'complete' | 'failed';
output: unknown;
error: string | null;
}
/** Fetch both row sets for a job in one shot. */
export async function loadTranscriptRows(
engine: BrainEngine,
jobId: number,
): Promise<{ messages: SubagentMessageRow[]; tools: SubagentToolExecRow[] }> {
const msgRows = await engine.executeRaw<Record<string, unknown>>(
`SELECT id, job_id, message_idx, role, content_blocks, tokens_in, tokens_out,
tokens_cache_read, tokens_cache_create, model, ended_at
FROM subagent_messages
WHERE job_id = $1
ORDER BY message_idx ASC`,
[jobId],
);
const toolRows = await engine.executeRaw<Record<string, unknown>>(
`SELECT id, job_id, message_idx, tool_use_id, tool_name, input, status, output, error
FROM subagent_tool_executions
WHERE job_id = $1
ORDER BY id ASC`,
[jobId],
);
return {
messages: msgRows.map(normalizeMessage),
tools: toolRows.map(normalizeTool),
};
}
function normalizeMessage(row: Record<string, unknown>): SubagentMessageRow {
const blocks = row.content_blocks;
const parsedBlocks: ContentBlock[] = typeof blocks === 'string'
? (JSON.parse(blocks) as ContentBlock[])
: (blocks as ContentBlock[]) ?? [];
return {
id: row.id as number,
job_id: row.job_id as number,
message_idx: row.message_idx as number,
role: row.role as 'user' | 'assistant',
content_blocks: parsedBlocks,
tokens_in: (row.tokens_in as number) ?? null,
tokens_out: (row.tokens_out as number) ?? null,
tokens_cache_read: (row.tokens_cache_read as number) ?? null,
tokens_cache_create: (row.tokens_cache_create as number) ?? null,
model: (row.model as string) ?? null,
ended_at: new Date(row.ended_at as string),
};
}
function normalizeTool(row: Record<string, unknown>): SubagentToolExecRow {
const input = typeof row.input === 'string' ? JSON.parse(row.input) : row.input;
const output = row.output == null
? null
: (typeof row.output === 'string' ? JSON.parse(row.output) : row.output);
return {
id: row.id as number,
job_id: row.job_id as number,
message_idx: row.message_idx as number,
tool_use_id: row.tool_use_id as string,
tool_name: row.tool_name as string,
input,
status: row.status as 'pending' | 'complete' | 'failed',
output,
error: (row.error as string) ?? null,
};
}
export interface RenderTranscriptOpts {
/** Trim long tool outputs in the markdown. Default: 4 KiB per output. */
maxOutputBytes?: number;
}
/**
* Render messages + tool executions to markdown. Message order is
* authoritative; tool rows are spliced under their owning assistant message
* by tool_use_id.
*/
export function renderTranscript(
messages: SubagentMessageRow[],
tools: SubagentToolExecRow[],
opts: RenderTranscriptOpts = {},
): string {
const maxOut = opts.maxOutputBytes ?? 4096;
const toolById = new Map<string, SubagentToolExecRow>(
tools.map(t => [t.tool_use_id, t]),
);
const out: string[] = [];
out.push('# Subagent transcript', '');
if (messages.length === 0) {
out.push('_(no messages)_');
return out.join('\n');
}
const first = messages[0]!;
out.push(`- job_id: ${first.job_id}`);
out.push(`- messages: ${messages.length}`);
if (first.model) out.push(`- model: ${first.model}`);
out.push('');
for (const msg of messages) {
out.push(`## Message ${msg.message_idx}${msg.role}`);
if (msg.tokens_in != null || msg.tokens_out != null) {
const parts: string[] = [];
if (msg.tokens_in) parts.push(`in=${msg.tokens_in}`);
if (msg.tokens_out) parts.push(`out=${msg.tokens_out}`);
if (msg.tokens_cache_read) parts.push(`cache_read=${msg.tokens_cache_read}`);
if (msg.tokens_cache_create) parts.push(`cache_create=${msg.tokens_cache_create}`);
if (parts.length > 0) out.push(`> tokens: ${parts.join(' ')}`);
}
out.push('');
for (const block of msg.content_blocks) {
renderBlock(block, toolById, maxOut, out);
}
out.push('');
}
return out.join('\n').replace(/\n{3,}/g, '\n\n');
}
function renderBlock(
block: ContentBlock,
toolById: Map<string, SubagentToolExecRow>,
maxOutputBytes: number,
out: string[],
): void {
if (block.type === 'text' && typeof block.text === 'string') {
out.push(block.text);
out.push('');
return;
}
if (block.type === 'tool_use') {
const name = typeof block.name === 'string' ? block.name : '<unknown>';
const inputStr = safeJson(block.input, 2);
out.push(`**tool_use** \`${name}\` (id=\`${block.id ?? '?'}\`)`);
out.push('```json', inputStr, '```');
const toolRow = block.id && typeof block.id === 'string' ? toolById.get(block.id) : undefined;
if (toolRow) {
out.push(`→ status: **${toolRow.status}**`);
if (toolRow.status === 'complete') {
out.push('```json', truncate(safeJson(toolRow.output, 2), maxOutputBytes), '```');
} else if (toolRow.status === 'failed') {
out.push(`> error: ${toolRow.error ?? '(no error text)'}`);
} else if (toolRow.status === 'pending') {
out.push('> pending (no resolution recorded yet)');
}
}
out.push('');
return;
}
if (block.type === 'tool_result') {
// Most tool_result blocks live inside user messages echoing back the
// assistant's tool_use. We skip them here because the owning tool_use
// block already rendered the execution row. If the user message carries
// a raw tool_result with no matching tool_use (rare), dump it raw.
if (!block.tool_use_id || !toolById.has(block.tool_use_id as string)) {
out.push('**tool_result** (no matching tool_use in this transcript)');
out.push('```json', truncate(safeJson(block.content, 2), maxOutputBytes), '```');
out.push('');
}
return;
}
// Unknown block type — dump as a fenced JSON block for diagnostics.
out.push(`**${block.type}**`);
out.push('```json', truncate(safeJson(block, 2), maxOutputBytes), '```');
out.push('');
}
function safeJson(value: unknown, indent = 0): string {
try {
return JSON.stringify(value, null, indent);
} catch {
return String(value);
}
}
function truncate(s: string, maxBytes: number): string {
if (Buffer.byteLength(s, 'utf8') <= maxBytes) return s;
// Slice bytewise via Buffer so we don't split a multibyte char awkwardly.
const buf = Buffer.from(s, 'utf8').slice(0, maxBytes);
return buf.toString('utf8') + `\n... [truncated at ${maxBytes} bytes]`;
}
+145 -5
View File
@@ -104,9 +104,11 @@ export interface MinionJobInput {
backoff_delay?: number;
backoff_jitter?: number;
/**
* Max number of stall windows before dead-letter. Default is the schema
* default (5 as of v0.13.1). Clamped to [1, 100] on insert values
* outside that range are silently coerced. See migration v13.
* Per-job override for how many stall windows are tolerated before the
* queue dead-letters the job. When omitted, the schema column DEFAULT
* applies (bumped 1 3 in v0.14, now 5 as of v0.13.1's audit). Clamped
* to [1, 100] on insert. For long-running handlers (LLM loops etc.) that
* should survive a worker kill mid-run, set max_stalled: 3+.
*/
max_stalled?: number;
delay?: number; // ms delay before eligible
@@ -208,14 +210,37 @@ export function rowToInboxMessage(row: Record<string, unknown>): InboxMessage {
};
}
// --- Child-done inbox message (auto-posted on completeJob) ---
// --- Child-done inbox message (auto-posted on every terminal transition) ---
/**
* Posted into the parent's inbox when a child reaches a terminal state.
*
* Pre-v0.15: only success paths (completeJob) emitted this. Failed/dead/
* cancelled children produced no payload, which stranded aggregator-style
* parents that needed to wait for N children regardless of outcome.
*
* v0.15: failJob, cancelJob, and handleTimeouts also emit child_done with
* the appropriate `outcome`, so the aggregator handler can count "N children
* resolved" without worrying about which rail each one took.
*
* Backwards compatible: old ChildDoneMessage consumers only read child_id,
* job_name, and result (non-null on success). Outcome and error are additive.
*/
export type ChildOutcome = 'complete' | 'failed' | 'dead' | 'cancelled' | 'timeout';
/** Posted into the parent's inbox when a child completes successfully. */
export interface ChildDoneMessage {
type: 'child_done';
child_id: number;
job_name: string;
result: unknown;
/**
* Terminal outcome. When absent (from a pre-v0.15 writer that didn't set
* it), consumers should treat the message as 'complete' the legacy writer
* only emitted on success paths.
*/
outcome?: ChildOutcome;
/** Set when outcome !== 'complete'. Mirrors minion_jobs.error_text. */
error?: string | null;
}
// --- Attachments (v7) ---
@@ -336,3 +361,118 @@ export function rowToMinionJob(row: Record<string, unknown>): MinionJob {
updated_at: new Date(row.updated_at as string),
};
}
// ---------------------------------------------------------------------------
// Subagent runtime (v0.15+)
// ---------------------------------------------------------------------------
/**
* Input payload for the 'subagent' handler. Shape is intentionally narrow
* tool registry and provider config resolve via handler-side defaults + env,
* not per-job data, so restart/replay uses the same behavior.
*/
export interface SubagentHandlerData {
/** Top-level user turn kicking off the loop. */
prompt: string;
/** Optional subagent definition path (skills/subagents/*.md or plugin). */
subagent_def?: string;
/** Anthropic model id. Defaults to sonnet at handler resolution time. */
model?: string;
/** Max assistant turns before the loop fails with stop_reason='max_turns'. */
max_turns?: number;
/**
* Whitelist of tool names the agent may call. MUST be a subset of the
* derived registry names invalid entries are rejected at tool-dispatch
* time, not silently ignored. Empty array = no tools.
*/
allowed_tools?: string[];
/** System prompt override. When omitted, the handler builds one. */
system?: string;
/** Template variables for subagent_def. Arbitrary JSON-serializable. */
input_vars?: Record<string, unknown>;
}
/**
* Input for the 'subagent_aggregator' handler. Claims AFTER all children
* resolve and aggregates their results into a brain page.
*/
export interface AggregatorHandlerData {
/** The subagent child job ids this aggregator is waiting on. */
children_ids: number[];
/**
* Optional template for the synthesis prompt. When omitted, the handler
* uses a generic "summarize these N results" prompt.
*/
aggregate_prompt_template?: string;
/**
* Target slug for the aggregated brain page. When present, a trusted-CLI
* put_page (viaSubagent=false) writes the final aggregation there.
*/
output_slug?: string;
}
/** Tool execution context passed to every ToolDef.execute. */
export interface ToolCtx {
/** Engine for DB-backed tools (brain_query, put_page, etc.). */
engine: import('../engine.ts').BrainEngine;
/** The subagent job id (used for audit + put_page namespace enforcement). */
jobId: number;
/** Always true for LLM-invoked tools — matches MCP trust boundary. */
remote: true;
/** Fired on cooperative abort (timeout, lock loss, cancel, SIGTERM). */
signal?: AbortSignal;
}
/**
* A tool the subagent can call. Names match Anthropic's constraint
* `^[a-zA-Z0-9_-]{1,64}$` no dots. The input_schema is the JSONSchema
* shipped to the Anthropic Messages API verbatim; ToolDef is the single
* Anthropic-compatible envelope, not an MCP McpToolDef (those have a
* different shape ".inputSchema" vs ".input_schema").
*
* `idempotent: true` is required for the two-phase replay path: on resume,
* a 'pending' row can be re-executed. Non-idempotent tools need a separate
* resume policy and are not supported in v0.15.
*/
export interface ToolDef {
name: string;
description: string;
input_schema: Record<string, unknown>;
idempotent: boolean;
execute(input: unknown, ctx: ToolCtx): Promise<unknown>;
}
/**
* Anthropic content-block subset we persist in subagent_messages.content_blocks.
* This is structural we don't gatekeep on unknown block types (future SDK
* additions pass through). Use the string-literal discriminant on 'type'.
*/
export type ContentBlock =
| { type: 'text'; text: string; [k: string]: unknown }
| { type: 'tool_use'; id: string; name: string; input: unknown; [k: string]: unknown }
| { type: 'tool_result'; tool_use_id: string; content: unknown; is_error?: boolean; [k: string]: unknown }
| { type: string; [k: string]: unknown };
/** Stop reason reported to the caller when the subagent loop terminates. */
export type SubagentStopReason =
| 'end_turn' // Anthropic says end_turn and last message has no tool_use
| 'max_turns' // hit max_turns budget before end_turn
| 'refusal' // detected via stop_reason + content shape
| 'error'; // unrecoverable (empty response retry exhausted, etc.)
/** Terminal result payload emitted by the subagent handler. */
export interface SubagentResult {
/** Concatenated text from the final assistant message. */
result: string;
/** Number of assistant turns consumed. */
turns_count: number;
/** Why the loop stopped. */
stop_reason: SubagentStopReason;
/** Rollup of tokens across all turns. */
tokens: {
in: number;
out: number;
cache_read: number;
cache_create: number;
};
}
+94
View File
@@ -0,0 +1,94 @@
/**
* Poll-until-terminal helper for CLI callers. Minions doesn't ship a
* notification stream for arbitrary callers (the NOTIFY trigger is worker-
* side), so `gbrain agent run --follow` on the CLI side polls getJob() until
* the job reaches a terminal state.
*
* On timeout, the job is NOT cancelled the user can `gbrain jobs get <id>`
* later to check. Explicit cancellation is the user's call via `gbrain jobs
* cancel <id>`.
*/
import type { MinionQueue } from './queue.ts';
import type { MinionJob, MinionJobStatus } from './types.ts';
export class TimeoutError extends Error {
constructor(public readonly jobId: number, public readonly elapsedMs: number) {
super(`timeout after ${elapsedMs}ms waiting for job ${jobId}`);
this.name = 'TimeoutError';
}
}
const TERMINAL_STATES: readonly MinionJobStatus[] = ['completed', 'failed', 'dead', 'cancelled'] as const;
const TERMINAL_SET = new Set<MinionJobStatus>(TERMINAL_STATES);
export interface WaitOpts {
/** Abort after this many ms. Default: 24h (long enough for most durable runs). */
timeoutMs?: number;
/**
* Poll interval. Defaults:
* - 1000ms on Postgres (lighter load, concurrent followers scale)
* - 250ms when the caller knows it's on PGLite inline (single process,
* no network RTT)
* Callers pass the appropriate value explicitly this module doesn't
* introspect the engine.
*/
pollMs?: number;
/** Optional AbortSignal — on abort, the poll loop exits early (no TimeoutError). */
signal?: AbortSignal;
}
export async function waitForCompletion(
queue: MinionQueue,
jobId: number,
opts: WaitOpts = {},
): Promise<MinionJob> {
const timeoutMs = opts.timeoutMs ?? 24 * 60 * 60 * 1000;
const pollMs = opts.pollMs ?? 1000;
const started = Date.now();
// Fast-path first read (don't wait pollMs just to learn it's already done).
let job = await queue.getJob(jobId);
if (!job) throw new Error(`job ${jobId} not found`);
if (TERMINAL_SET.has(job.status)) return job;
while (true) {
if (opts.signal?.aborted) {
// Caller aborted. Return the last-seen snapshot rather than throwing —
// the job itself is still alive queue-side, and the caller knows they
// aborted.
return job;
}
const elapsed = Date.now() - started;
if (elapsed >= timeoutMs) {
throw new TimeoutError(jobId, elapsed);
}
const remaining = timeoutMs - elapsed;
const sleep = Math.min(pollMs, remaining);
await delay(sleep, opts.signal);
job = await queue.getJob(jobId);
if (!job) throw new Error(`job ${jobId} disappeared mid-wait`);
if (TERMINAL_SET.has(job.status)) return job;
}
}
function delay(ms: number, signal?: AbortSignal): Promise<void> {
if (ms <= 0) return Promise.resolve();
return new Promise((resolve) => {
const t = setTimeout(() => {
signal?.removeEventListener('abort', onAbort);
resolve();
}, ms);
const onAbort = () => {
clearTimeout(t);
resolve();
};
signal?.addEventListener('abort', onAbort, { once: true });
});
}
// Exported for unit tests.
export const __testing = {
TERMINAL_STATES,
};
+1 -1
View File
@@ -30,7 +30,7 @@ import { evaluateQuietHours, type QuietHoursConfig } from './quiet-hours.ts';
function readQuietHoursConfig(job: MinionJob): QuietHoursConfig | null {
const cfg = (job as MinionJob & { quiet_hours?: unknown }).quiet_hours;
if (!cfg || typeof cfg !== 'object') return null;
return cfg as QuietHoursConfig;
return cfg as unknown as QuietHoursConfig;
}
/** Per-job in-flight state (isolated per job, not shared on the worker). */
+120
View File
@@ -0,0 +1,120 @@
import { chmodSync, existsSync, mkdirSync, readFileSync, writeFileSync } from 'fs';
import { resolve } from 'path';
import { configDir, configPath } from './config.ts';
export type RepoStrategy = 'markdown' | 'code' | 'auto';
export interface RepoConfig {
path: string;
name: string;
strategy: RepoStrategy;
include?: string[];
exclude?: string[];
syncEnabled?: boolean;
}
interface RawConfig {
[key: string]: unknown;
repos?: unknown;
}
function readRawConfig(): RawConfig {
try {
const raw = readFileSync(configPath(), 'utf-8');
return JSON.parse(raw) as RawConfig;
} catch {
return {};
}
}
function writeRawConfig(config: RawConfig): void {
mkdirSync(configDir(), { recursive: true });
writeFileSync(configPath(), JSON.stringify(config, null, 2) + '\n', { mode: 0o600 });
try {
chmodSync(configPath(), 0o600);
} catch {
// chmod may fail on some platforms
}
}
export function normalizeRepoName(repoPath: string): string {
const normalized = repoPath.replace(/\\/g, '/').replace(/\/+$/g, '');
const parts = normalized.split('/').filter(Boolean);
return parts[parts.length - 1] || 'repo';
}
export function loadRepoConfigs(): RepoConfig[] {
if (!existsSync(configPath())) return [];
const parsed = readRawConfig();
if (!Array.isArray(parsed.repos)) return [];
const repos: RepoConfig[] = [];
for (const candidate of parsed.repos) {
if (!candidate || typeof candidate !== 'object') continue;
const row = candidate as Record<string, unknown>;
if (typeof row.path !== 'string' || typeof row.name !== 'string') continue;
const strategy = row.strategy;
const normalizedStrategy: RepoStrategy = strategy === 'code' || strategy === 'auto' ? strategy : 'markdown';
repos.push({
path: resolve(row.path),
name: row.name,
strategy: normalizedStrategy,
include: Array.isArray(row.include) ? row.include.map(String) : undefined,
exclude: Array.isArray(row.exclude) ? row.exclude.map(String) : undefined,
syncEnabled: typeof row.syncEnabled === 'boolean' ? row.syncEnabled : true,
});
}
return repos;
}
export function saveRepoConfigs(repos: RepoConfig[]): void {
const parsed = readRawConfig();
parsed.repos = repos.map((repo) => ({
path: resolve(repo.path),
name: repo.name,
strategy: repo.strategy,
...(repo.include && repo.include.length > 0 ? { include: repo.include } : {}),
...(repo.exclude && repo.exclude.length > 0 ? { exclude: repo.exclude } : {}),
...(repo.syncEnabled === false ? { syncEnabled: false } : {}),
}));
writeRawConfig(parsed);
}
export function addRepoConfig(repo: RepoConfig): RepoConfig[] {
const repos = loadRepoConfigs();
const normalized: RepoConfig = {
...repo,
path: resolve(repo.path),
name: repo.name || normalizeRepoName(repo.path),
strategy: repo.strategy || 'auto',
syncEnabled: repo.syncEnabled ?? true,
};
const nameTaken = repos.find((r) => r.name === normalized.name);
if (nameTaken) {
throw new Error(`Repo name already exists: ${normalized.name}`);
}
const pathTaken = repos.find((r) => resolve(r.path) === normalized.path);
if (pathTaken) {
throw new Error(`Repo path already configured: ${normalized.path}`);
}
const updated = [...repos, normalized];
saveRepoConfigs(updated);
return updated;
}
export function removeRepoConfig(name: string): RepoConfig[] {
const repos = loadRepoConfigs();
const next = repos.filter((r) => r.name !== name);
if (next.length === repos.length) {
throw new Error(`Repo not found: ${name}`);
}
saveRepoConfigs(next);
return next;
}
+39 -5
View File
@@ -24,7 +24,8 @@ export type ErrorCode =
| 'embedding_failed'
| 'storage_error'
| 'bucket_not_found'
| 'database_error';
| 'database_error'
| 'permission_denied';
export class OperationError extends Error {
constructor(
@@ -167,6 +168,19 @@ export interface OperationContext {
* When unset, operations MUST default to the stricter (remote=true) behavior.
*/
remote?: boolean;
/**
* Subagent runtime context (v0.16+). Set by the subagent tool dispatcher when
* dispatching an op as a tool call from an LLM loop. Used to enforce per-op
* agent policy (e.g. put_page namespace rule).
*
* `viaSubagent` is the FAIL-CLOSED flag: when true, agent-facing policy MUST
* be enforced even if `subagentId` happens to be undefined (a bug in the
* dispatcher must not bypass the guard). `subagentId` is the owning subagent
* job id; `jobId` is the current Minion job id (aggregator or subagent).
*/
jobId?: number;
subagentId?: number;
viaSubagent?: boolean;
/**
* Resolved global CLI options (--quiet / --progress-json / --progress-interval).
* CLI callers populate this from `getCliOptions()`. MCP / library callers
@@ -235,8 +249,28 @@ const put_page: Operation = {
},
mutating: true,
handler: async (ctx, p) => {
if (ctx.dryRun) return { dry_run: true, action: 'put_page', slug: p.slug };
const slug = p.slug as string;
// Subagent namespace enforcement (v0.15+). Runs BEFORE the dry-run
// short-circuit so preview calls surface the same rejection. Confines
// LLM-driven writes to wiki/agents/<subagentId>/... — no leading slash
// (slug grammar rejects that), anchored, slash-boundary to defeat prefix
// collisions like `wiki/agents/12evil/*` impersonating subagent 12.
//
// FAIL-CLOSED: `viaSubagent=true` enforces the check even if the
// dispatcher forgot to populate `subagentId`. Agent-originated writes
// without an owning subagent id are rejected outright.
if (ctx.viaSubagent === true) {
if (typeof ctx.subagentId !== 'number' || Number.isNaN(ctx.subagentId)) {
throw new OperationError('permission_denied', 'put_page via subagent requires ctx.subagentId');
}
const prefix = `wiki/agents/${ctx.subagentId}/`;
if (!slug.startsWith(prefix) || slug.length === prefix.length) {
throw new OperationError('permission_denied', `put_page via subagent must write under '${prefix}...'`);
}
}
if (ctx.dryRun) return { dry_run: true, action: 'put_page', slug: p.slug };
// Skip embedding when no OpenAI key is configured. importFromContent's existing
// try/catch around embed only catches; without a key the OpenAI client would
// attempt 5 retries with exponential backoff (up to ~2 minutes total) before
@@ -675,7 +709,7 @@ const get_backlinks: Operation = {
* grows a `visited` array per path; in `direction=both` the join is `OR`-based and
* fans out exponentially. Without a cap, a remote MCP caller can pass depth=1e6
* and burn memory/CPU on the database. 10 hops is well beyond any realistic
* relationship query (Wintermute's "people who attended meetings with Alice"
* relationship query (your OpenClaw's "people who attended meetings with Alice"
* is 2 hops; the deepest meaningful chain in our test data is 4).
*/
const TRAVERSE_DEPTH_CAP = 10;
@@ -1256,9 +1290,9 @@ const find_orphans: Operation = {
description: 'Include auto-generated and pseudo pages (default: false)',
},
},
handler: async (_ctx, p) => {
handler: async (ctx, p) => {
const { findOrphans } = await import('../commands/orphans.ts');
return findOrphans((p.include_pseudo as boolean) || false);
return findOrphans(ctx.engine, { includePseudo: (p.include_pseudo as boolean) || false });
},
cliHints: { name: 'orphans', hidden: true },
};
+2 -2
View File
@@ -5,7 +5,7 @@
* WHERE and HOW. Every user-visible URL, every citation, every wikilink is
* assembled from resolver outputs or structured IDs never from LLM text.
*
* Example (from the Wintermute memory log, 2026-04-13): an agent was asked
* Example (from Garry's OpenClaw memory log, 2026-04-13): an agent was asked
* to rewrite daily files and it invented a "Philip Leung" entity that didn't
* exist. With the Scaffolder, the LLM writes "the attendee was mentioned
* again" and code writes the actual `[Philip Leung](people/philip-leung.md)`
@@ -65,7 +65,7 @@ export interface EmailCitationInput {
* Canonical email citation with a deep link that opens the actual thread:
* [Source: email "Subject line", 2026-04-18](https://mail.google.com/mail/u/?authuser=...#inbox/...)
*
* URL shape matches the pattern Wintermute's ingest pipeline builds from API
* URL shape matches the pattern Garry's OpenClaw's ingest pipeline builds from API
* responses, so brain-page links and agent-generated links use the same
* format (cross-tool consistency).
*/
+16 -1
View File
@@ -113,7 +113,7 @@ export class PGLiteEngine implements BrainEngine {
async putPage(slug: string, page: PageInput): Promise<Page> {
slug = validateSlug(slug);
const hash = page.content_hash || contentHash(page.compiled_truth, page.timeline || '');
const hash = page.content_hash || contentHash(page);
const frontmatter = page.frontmatter || {};
const { rows } = await this.db.query(
@@ -655,6 +655,21 @@ export class PGLiteEngine implements BrainEngine {
return result;
}
async findOrphanPages(): Promise<Array<{ slug: string; title: string; domain: string | null }>> {
const { rows } = await this.db.query(
`SELECT
p.slug,
COALESCE(p.title, p.slug) AS title,
p.frontmatter->>'domain' AS domain
FROM pages p
WHERE NOT EXISTS (
SELECT 1 FROM links l WHERE l.to_page_id = p.id
)
ORDER BY p.slug`
);
return rows as Array<{ slug: string; title: string; domain: string | null }>;
}
// Tags
async addTag(slug: string, tag: string): Promise<void> {
await this.db.query(
+3 -1
View File
@@ -59,7 +59,9 @@ export async function acquireLock(dataDir: string | undefined, opts?: { timeoutM
return { lockDir: '', acquired: true };
}
mkdirSync(dataDir, { recursive: true });
// `lockDir` being set implies `dataDir` is set (see getLockDir), but TS
// can't derive that across helper boundaries.
mkdirSync(dataDir as string, { recursive: true });
const timeoutMs = opts?.timeoutMs ?? 30_000; // 30 second default timeout
const startTime = Date.now();
+62
View File
@@ -264,6 +264,68 @@ CREATE INDEX IF NOT EXISTS idx_minion_attachments_job ON minion_attachments (job
-- NOTE: SET STORAGE EXTERNAL is omitted on PGLite; it's a Postgres TOAST optimization
-- and PGLite may not support it. Postgres path applies it via migration v7.
-- ============================================================
-- Subagent runtime (v0.16.0) durable LLM loops
-- ============================================================
CREATE TABLE IF NOT EXISTS subagent_messages (
id BIGSERIAL PRIMARY KEY,
job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
message_idx INTEGER NOT NULL,
role TEXT NOT NULL,
content_blocks JSONB NOT NULL,
tokens_in INTEGER,
tokens_out INTEGER,
tokens_cache_read INTEGER,
tokens_cache_create INTEGER,
model TEXT,
ended_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uniq_subagent_messages_idx UNIQUE (job_id, message_idx),
CONSTRAINT chk_subagent_messages_role CHECK (role IN ('user','assistant'))
);
CREATE INDEX IF NOT EXISTS idx_subagent_messages_job ON subagent_messages (job_id, message_idx);
CREATE TABLE IF NOT EXISTS subagent_tool_executions (
id BIGSERIAL PRIMARY KEY,
job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
message_idx INTEGER NOT NULL,
tool_use_id TEXT NOT NULL,
tool_name TEXT NOT NULL,
input JSONB NOT NULL,
status TEXT NOT NULL,
output JSONB,
error TEXT,
started_at TIMESTAMPTZ NOT NULL DEFAULT now(),
ended_at TIMESTAMPTZ,
CONSTRAINT uniq_subagent_tools_use_id UNIQUE (job_id, tool_use_id),
CONSTRAINT chk_subagent_tools_status CHECK (status IN ('pending','complete','failed'))
);
CREATE INDEX IF NOT EXISTS idx_subagent_tools_job ON subagent_tool_executions (job_id, status);
CREATE TABLE IF NOT EXISTS subagent_rate_leases (
id BIGSERIAL PRIMARY KEY,
key TEXT NOT NULL,
owner_job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT now(),
expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_rate_leases_key_expires ON subagent_rate_leases (key, expires_at);
-- ============================================================
-- Cycle coordination lock v0.17 runCycle primitive
-- ============================================================
-- See src/schema.sql for full rationale. One row per active cycle.
-- PGLite is single-writer, so the lock doubly protects: the DB-level
-- row + the file lock at ~/.gbrain/cycle.lock prevent concurrent
-- CLI invocations from racing.
CREATE TABLE IF NOT EXISTS gbrain_cycle_locks (
id TEXT PRIMARY KEY,
holder_pid INT NOT NULL,
holder_host TEXT,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
ttl_expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_cycle_locks_ttl ON gbrain_cycle_locks(ttl_expires_at);
-- ============================================================
-- Trigger-based search_vector (spans pages + timeline_entries)
-- ============================================================
+29 -13
View File
@@ -4,7 +4,7 @@ import { MAX_SEARCH_LIMIT, clampSearchLimit } from './engine.ts';
import { runMigrations } from './migrate.ts';
import { SCHEMA_SQL } from './schema-embedded.ts';
import type {
Page, PageInput, PageFilters,
Page, PageInput, PageFilters, PageType,
Chunk, ChunkInput,
SearchResult, SearchOpts,
Link, GraphNode, GraphPath,
@@ -95,7 +95,7 @@ export class PostgresEngine implements BrainEngine {
Object.defineProperty(txEngine, 'sql', { get: () => tx });
Object.defineProperty(txEngine, '_sql', { value: tx as unknown as ReturnType<typeof postgres>, writable: false });
return fn(txEngine);
});
}) as Promise<T>;
}
// Pages CRUD
@@ -117,7 +117,7 @@ export class PostgresEngine implements BrainEngine {
const rows = await sql`
INSERT INTO pages (slug, type, title, compiled_truth, timeline, frontmatter, content_hash, updated_at)
VALUES (${slug}, ${page.type}, ${page.title}, ${page.compiled_truth}, ${page.timeline || ''}, ${sql.json(frontmatter)}, ${hash}, now())
VALUES (${slug}, ${page.type}, ${page.title}, ${page.compiled_truth}, ${page.timeline || ''}, ${sql.json(frontmatter as Parameters<typeof sql.json>[0])}, ${hash}, now())
ON CONFLICT (slug) DO UPDATE SET
type = EXCLUDED.type,
title = EXCLUDED.title,
@@ -165,7 +165,7 @@ export class PostgresEngine implements BrainEngine {
async getAllSlugs(): Promise<Set<string>> {
const sql = this.sql;
const rows = await sql`SELECT slug FROM pages`;
return new Set(rows.map((r: { slug: string }) => r.slug));
return new Set(rows.map((r) => r.slug as string));
}
async resolveSlugs(partial: string): Promise<string[]> {
@@ -183,7 +183,7 @@ export class PostgresEngine implements BrainEngine {
ORDER BY sim DESC
LIMIT 5
`;
return fuzzy.map((r: { slug: string }) => r.slug);
return fuzzy.map((r) => r.slug as string);
}
// Search
@@ -344,7 +344,7 @@ export class PostgresEngine implements BrainEngine {
model = COALESCE(EXCLUDED.model, content_chunks.model),
token_count = EXCLUDED.token_count,
embedded_at = COALESCE(EXCLUDED.embedded_at, content_chunks.embedded_at)`,
params,
params as Parameters<typeof sql.unsafe>[1],
);
}
@@ -356,7 +356,7 @@ export class PostgresEngine implements BrainEngine {
WHERE p.slug = ${slug}
ORDER BY cc.chunk_index
`;
return rows.map(rowToChunk);
return rows.map((r) => rowToChunk(r as Record<string, unknown>));
}
async deleteChunks(slug: string): Promise<void> {
@@ -693,12 +693,28 @@ export class PostgresEngine implements BrainEngine {
WHERE p.slug = ANY(${slugs}::text[])
GROUP BY p.slug
`;
for (const r of rows as { slug: string; cnt: number }[]) {
for (const r of rows as unknown as { slug: string; cnt: number }[]) {
result.set(r.slug, Number(r.cnt));
}
return result;
}
async findOrphanPages(): Promise<Array<{ slug: string; title: string; domain: string | null }>> {
const sql = this.sql;
const rows = await sql`
SELECT
p.slug,
COALESCE(p.title, p.slug) AS title,
p.frontmatter->>'domain' AS domain
FROM pages p
WHERE NOT EXISTS (
SELECT 1 FROM links l WHERE l.to_page_id = p.id
)
ORDER BY p.slug
`;
return rows as unknown as Array<{ slug: string; title: string; domain: string | null }>;
}
// Tags
async addTag(slug: string, tag: string): Promise<void> {
const sql = this.sql;
@@ -729,7 +745,7 @@ export class PostgresEngine implements BrainEngine {
WHERE page_id = (SELECT id FROM pages WHERE slug = ${slug})
ORDER BY tag
`;
return rows.map((r: { tag: string }) => r.tag);
return rows.map((r) => r.tag as string);
}
// Timeline
@@ -814,7 +830,7 @@ export class PostgresEngine implements BrainEngine {
const sql = this.sql;
const result = await sql`
INSERT INTO raw_data (page_id, source, data)
SELECT id, ${source}, ${sql.json(data as Record<string, unknown>)}
SELECT id, ${source}, ${sql.json(data as Parameters<typeof sql.json>[0])}
FROM pages WHERE slug = ${slug}
ON CONFLICT (page_id, source) DO UPDATE SET
data = EXCLUDED.data,
@@ -987,7 +1003,7 @@ export class PostgresEngine implements BrainEngine {
dead_links: deadLinks,
link_coverage: Number(h.link_coverage),
timeline_coverage: Number(h.timeline_coverage),
most_connected: (connected as { slug: string; link_count: number }[]).map(c => ({
most_connected: (connected as unknown as { slug: string; link_count: number }[]).map(c => ({
slug: c.slug,
link_count: Number(c.link_count),
})),
@@ -1060,11 +1076,11 @@ export class PostgresEngine implements BrainEngine {
WHERE p.slug = ${slug}
ORDER BY cc.chunk_index
`;
return rows.map((r: Record<string, unknown>) => rowToChunk(r, true));
return rows.map((r) => rowToChunk(r as Record<string, unknown>, true));
}
async executeRaw<T = Record<string, unknown>>(sql: string, params?: unknown[]): Promise<T[]> {
const conn = this.sql;
return conn.unsafe(sql, params) as unknown as T[];
return conn.unsafe(sql, params as Parameters<typeof conn.unsafe>[1]) as unknown as T[];
}
}
+21
View File
@@ -0,0 +1,21 @@
import { existsSync } from 'fs';
import { join } from 'path';
/**
* Walk up from `startDir` looking for `skills/RESOLVER.md` the marker of a
* gbrain repo root. Returns the absolute directory containing `skills/` or
* null if no such directory is found within 10 levels.
*
* `startDir` is parameterized so tests can run hermetically against fixtures.
* Default matches the prior `doctor.ts`-private implementation.
*/
export function findRepoRoot(startDir: string = process.cwd()): string | null {
let dir = startDir;
for (let i = 0; i < 10; i++) {
if (existsSync(join(dir, 'skills', 'RESOLVER.md'))) return dir;
const parent = join(dir, '..');
if (parent === dir) break;
dir = parent;
}
return null;
}
+74
View File
@@ -356,6 +356,80 @@ CREATE TABLE IF NOT EXISTS minion_attachments (
CREATE INDEX IF NOT EXISTS idx_minion_attachments_job ON minion_attachments (job_id);
ALTER TABLE minion_attachments ALTER COLUMN content SET STORAGE EXTERNAL;
-- ============================================================
-- Subagent runtime (v0.16.0) durable LLM loops
-- ============================================================
-- Anthropic-native message blocks, one row per Messages API message. Parallel
-- tool_use blocks in one assistant message live in content_blocks JSONB,
-- not across rows.
CREATE TABLE IF NOT EXISTS subagent_messages (
id BIGSERIAL PRIMARY KEY,
job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
message_idx INTEGER NOT NULL,
role TEXT NOT NULL,
content_blocks JSONB NOT NULL,
tokens_in INTEGER,
tokens_out INTEGER,
tokens_cache_read INTEGER,
tokens_cache_create INTEGER,
model TEXT,
ended_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uniq_subagent_messages_idx UNIQUE (job_id, message_idx),
CONSTRAINT chk_subagent_messages_role CHECK (role IN ('user','assistant'))
);
CREATE INDEX IF NOT EXISTS idx_subagent_messages_job ON subagent_messages (job_id, message_idx);
-- Two-phase tool execution ledger. Before tool call: INSERT status='pending'.
-- After success: UPDATE to 'complete' + output. On failure: 'failed' + error.
-- Replay re-runs 'pending' rows only if the tool is idempotent.
CREATE TABLE IF NOT EXISTS subagent_tool_executions (
id BIGSERIAL PRIMARY KEY,
job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
message_idx INTEGER NOT NULL,
tool_use_id TEXT NOT NULL,
tool_name TEXT NOT NULL,
input JSONB NOT NULL,
status TEXT NOT NULL,
output JSONB,
error TEXT,
started_at TIMESTAMPTZ NOT NULL DEFAULT now(),
ended_at TIMESTAMPTZ,
CONSTRAINT uniq_subagent_tools_use_id UNIQUE (job_id, tool_use_id),
CONSTRAINT chk_subagent_tools_status CHECK (status IN ('pending','complete','failed'))
);
CREATE INDEX IF NOT EXISTS idx_subagent_tools_job ON subagent_tool_executions (job_id, status);
-- Rate-lease table concurrency cap on outbound providers (e.g.
-- anthropic:messages). Acquire: INSERT if active < max_concurrent under
-- advisory lock. Release: DELETE. Stale leases (expires_at past) auto-prune
-- on next acquire so crashed workers can't strand capacity.
CREATE TABLE IF NOT EXISTS subagent_rate_leases (
id BIGSERIAL PRIMARY KEY,
key TEXT NOT NULL,
owner_job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT now(),
expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_rate_leases_key_expires ON subagent_rate_leases (key, expires_at);
-- ============================================================
-- Cycle coordination lock v0.17 runCycle primitive
-- ============================================================
-- One row per active cycle. Any caller (autopilot daemon, Minions
-- autopilot-cycle handler, gbrain dream CLI) tries to acquire this
-- row before running a DB-write phase. Holders refresh ttl_expires_at
-- between phases; crashed holders auto-release once TTL expires.
-- Works through PgBouncer transaction pooling, unlike session-scoped
-- pg_try_advisory_lock.
CREATE TABLE IF NOT EXISTS gbrain_cycle_locks (
id TEXT PRIMARY KEY,
holder_pid INT NOT NULL,
holder_host TEXT,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
ttl_expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_cycle_locks_ttl ON gbrain_cycle_locks(ttl_expires_at);
-- NOTIFY trigger for real-time job events (Postgres only, not PGLite)
CREATE OR REPLACE FUNCTION notify_minion_job_change() RETURNS trigger AS \$\$
BEGIN
+2 -2
View File
@@ -55,7 +55,7 @@ export class SupabaseStorage implements StorageBackend {
'Content-Type': mime || 'application/octet-stream',
'x-upsert': 'true',
},
body: data,
body: new Uint8Array(data.buffer, data.byteOffset, data.byteLength) as BodyInit,
});
if (!res.ok) {
const body = await res.text();
@@ -126,7 +126,7 @@ export class SupabaseStorage implements StorageBackend {
'Content-Type': 'application/offset+octet-stream',
'Content-Length': String(chunk.length),
},
body: chunk,
body: new Uint8Array(chunk.buffer, chunk.byteOffset, chunk.byteLength) as BodyInit,
});
if (!patchRes.ok) {
+85 -5
View File
@@ -24,6 +24,16 @@ export interface RawManifestEntry {
oldPath?: string;
}
export type SyncStrategy = 'markdown' | 'code' | 'auto';
interface SyncableOptions {
strategy?: SyncStrategy;
include?: string[];
exclude?: string[];
}
const CODE_EXTENSIONS = new Set(['.ts', '.tsx', '.js', '.jsx', '.mjs', '.cjs', '.py', '.rb', '.go']);
/**
* Parse the output of `git diff --name-status -M LAST..HEAD` into structured entries.
*
@@ -72,12 +82,60 @@ export function buildSyncManifest(gitDiffOutput: string): SyncManifest {
return manifest;
}
export function isCodeFilePath(path: string): boolean {
const lower = path.toLowerCase();
for (const ext of CODE_EXTENSIONS) {
if (lower.endsWith(ext)) return true;
}
return false;
}
function isMarkdownFilePath(path: string): boolean {
return path.endsWith('.md') || path.endsWith('.mdx');
}
function isAllowedByStrategy(path: string, strategy: SyncStrategy): boolean {
if (strategy === 'markdown') return isMarkdownFilePath(path);
if (strategy === 'code') return isCodeFilePath(path);
return isMarkdownFilePath(path) || isCodeFilePath(path);
}
function globToRegex(pattern: string): RegExp {
let regex = '^';
for (let i = 0; i < pattern.length; i++) {
const ch = pattern[i];
if (ch === '*') {
const next = pattern[i + 1];
if (next === '*') {
regex += '.*';
i++;
} else {
regex += '[^/]*';
}
continue;
}
if (ch === '?') { regex += '.'; continue; }
if ('\\.[]{}()+-^$|'.includes(ch)) { regex += `\\${ch}`; continue; }
regex += ch;
}
regex += '$';
return new RegExp(regex);
}
function matchesAnyGlob(path: string, patterns?: string[]): boolean {
if (!patterns || patterns.length === 0) return false;
const normalized = path.replace(/\\/g, '/');
return patterns.some((pattern) => globToRegex(pattern).test(normalized));
}
/**
* Filter a file path to determine if it should be synced to GBrain.
* Strategy-aware: 'markdown' (default) = .md/.mdx only, 'code' = code files only, 'auto' = both.
*/
export function isSyncable(path: string): boolean {
// Must be .md or .mdx
if (!path.endsWith('.md') && !path.endsWith('.mdx')) return false;
export function isSyncable(path: string, opts: SyncableOptions = {}): boolean {
const strategy = opts.strategy || 'markdown';
if (!isAllowedByStrategy(path, strategy)) return false;
// Skip hidden directories
if (path.split('/').some(p => p.startsWith('.'))) return false;
@@ -93,6 +151,9 @@ export function isSyncable(path: string): boolean {
// Skip ops/ directory
if (path.startsWith('ops/')) return false;
if (opts.include && opts.include.length > 0 && !matchesAnyGlob(path, opts.include)) return false;
if (opts.exclude && opts.exclude.length > 0 && matchesAnyGlob(path, opts.exclude)) return false;
return true;
}
@@ -125,11 +186,30 @@ export function slugifyPath(filePath: string): string {
return path.split('/').map(slugifySegment).filter(Boolean).join('/');
}
/**
* Slugify a code file path: flatten into a single slug segment with dots hyphens.
* e.g. 'src/core/chunkers/code.ts' 'src-core-chunkers-code-ts'
*/
export function slugifyCodePath(filePath: string): string {
let path = filePath.replace(/\\/g, '/');
path = path.replace(/^\.?\//, '');
return path
.split('/')
.map(segment => slugifySegment(segment.replace(/\./g, '-')))
.filter(Boolean)
.join('-');
}
/**
* Convert a repo-relative file path to a GBrain page slug.
*/
export function pathToSlug(filePath: string, repoPrefix?: string): string {
let slug = slugifyPath(filePath);
export function pathToSlug(
filePath: string,
repoPrefix?: string,
options: { pageKind?: 'markdown' | 'code' } = {},
): string {
const pageKind = options.pageKind || 'markdown';
let slug = pageKind === 'code' ? slugifyCodePath(filePath) : slugifyPath(filePath);
if (repoPrefix) slug = `${repoPrefix}/${slug}`;
return slug.toLowerCase();
}
+1 -1
View File
@@ -1,5 +1,5 @@
// Page types
export type PageType = 'person' | 'company' | 'deal' | 'yc' | 'civic' | 'project' | 'concept' | 'source' | 'media' | 'writing' | 'analysis' | 'guide' | 'hardware' | 'architecture';
export type PageType = 'person' | 'company' | 'deal' | 'yc' | 'civic' | 'project' | 'concept' | 'source' | 'media' | 'writing' | 'analysis' | 'guide' | 'hardware' | 'architecture' | 'meeting' | 'note' | 'code';
export interface Page {
id: number;
+5 -19
View File
@@ -6,6 +6,7 @@ import { operations, OperationError } from '../core/operations.ts';
import type { Operation, OperationContext } from '../core/operations.ts';
import { loadConfig } from '../core/config.ts';
import { VERSION } from '../version.ts';
import { buildToolDefs } from './tool-defs.ts';
/** Validate required params exist and have the expected type */
function validateParams(op: Operation, params: Record<string, unknown>): string | null {
@@ -32,26 +33,11 @@ export async function startMcpServer(engine: BrainEngine) {
{ capabilities: { tools: {} } },
);
// Generate tool definitions from operations
// Generate tool definitions from operations. Extracted to buildToolDefs so
// the subagent tool registry (v0.15+) can call the same mapper against a
// filtered OPERATIONS subset instead of duplicating this shape.
server.setRequestHandler(ListToolsRequestSchema, async () => ({
tools: operations.map(op => ({
name: op.name,
description: op.description,
inputSchema: {
type: 'object' as const,
properties: Object.fromEntries(
Object.entries(op.params).map(([k, v]) => [k, {
type: v.type === 'array' ? 'array' : v.type,
...(v.description ? { description: v.description } : {}),
...(v.enum ? { enum: v.enum } : {}),
...(v.items ? { items: { type: v.items.type } } : {}),
}]),
),
required: Object.entries(op.params)
.filter(([, v]) => v.required)
.map(([k]) => k),
},
})),
tools: buildToolDefs(operations),
}));
// Dispatch tool calls to operation handlers
+32
View File
@@ -0,0 +1,32 @@
import type { Operation } from '../core/operations.ts';
export interface McpToolDef {
name: string;
description: string;
inputSchema: {
type: 'object';
properties: Record<string, unknown>;
required: string[];
};
}
export function buildToolDefs(ops: Operation[]): McpToolDef[] {
return ops.map(op => ({
name: op.name,
description: op.description,
inputSchema: {
type: 'object' as const,
properties: Object.fromEntries(
Object.entries(op.params).map(([k, v]) => [k, {
type: v.type === 'array' ? 'array' : v.type,
...(v.description ? { description: v.description } : {}),
...(v.enum ? { enum: v.enum } : {}),
...(v.items ? { items: { type: v.items.type } } : {}),
}]),
),
required: Object.entries(op.params)
.filter(([, v]) => v.required)
.map(([k]) => k),
},
}));
}
+74
View File
@@ -352,6 +352,80 @@ CREATE TABLE IF NOT EXISTS minion_attachments (
CREATE INDEX IF NOT EXISTS idx_minion_attachments_job ON minion_attachments (job_id);
ALTER TABLE minion_attachments ALTER COLUMN content SET STORAGE EXTERNAL;
-- ============================================================
-- Subagent runtime (v0.16.0) — durable LLM loops
-- ============================================================
-- Anthropic-native message blocks, one row per Messages API message. Parallel
-- tool_use blocks in one assistant message live in content_blocks JSONB,
-- not across rows.
CREATE TABLE IF NOT EXISTS subagent_messages (
id BIGSERIAL PRIMARY KEY,
job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
message_idx INTEGER NOT NULL,
role TEXT NOT NULL,
content_blocks JSONB NOT NULL,
tokens_in INTEGER,
tokens_out INTEGER,
tokens_cache_read INTEGER,
tokens_cache_create INTEGER,
model TEXT,
ended_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uniq_subagent_messages_idx UNIQUE (job_id, message_idx),
CONSTRAINT chk_subagent_messages_role CHECK (role IN ('user','assistant'))
);
CREATE INDEX IF NOT EXISTS idx_subagent_messages_job ON subagent_messages (job_id, message_idx);
-- Two-phase tool execution ledger. Before tool call: INSERT status='pending'.
-- After success: UPDATE to 'complete' + output. On failure: 'failed' + error.
-- Replay re-runs 'pending' rows only if the tool is idempotent.
CREATE TABLE IF NOT EXISTS subagent_tool_executions (
id BIGSERIAL PRIMARY KEY,
job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
message_idx INTEGER NOT NULL,
tool_use_id TEXT NOT NULL,
tool_name TEXT NOT NULL,
input JSONB NOT NULL,
status TEXT NOT NULL,
output JSONB,
error TEXT,
started_at TIMESTAMPTZ NOT NULL DEFAULT now(),
ended_at TIMESTAMPTZ,
CONSTRAINT uniq_subagent_tools_use_id UNIQUE (job_id, tool_use_id),
CONSTRAINT chk_subagent_tools_status CHECK (status IN ('pending','complete','failed'))
);
CREATE INDEX IF NOT EXISTS idx_subagent_tools_job ON subagent_tool_executions (job_id, status);
-- Rate-lease table — concurrency cap on outbound providers (e.g.
-- anthropic:messages). Acquire: INSERT if active < max_concurrent under
-- advisory lock. Release: DELETE. Stale leases (expires_at past) auto-prune
-- on next acquire so crashed workers can't strand capacity.
CREATE TABLE IF NOT EXISTS subagent_rate_leases (
id BIGSERIAL PRIMARY KEY,
key TEXT NOT NULL,
owner_job_id BIGINT NOT NULL REFERENCES minion_jobs(id) ON DELETE CASCADE,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT now(),
expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_rate_leases_key_expires ON subagent_rate_leases (key, expires_at);
-- ============================================================
-- Cycle coordination lock — v0.17 runCycle primitive
-- ============================================================
-- One row per active cycle. Any caller (autopilot daemon, Minions
-- autopilot-cycle handler, gbrain dream CLI) tries to acquire this
-- row before running a DB-write phase. Holders refresh ttl_expires_at
-- between phases; crashed holders auto-release once TTL expires.
-- Works through PgBouncer transaction pooling, unlike session-scoped
-- pg_try_advisory_lock.
CREATE TABLE IF NOT EXISTS gbrain_cycle_locks (
id TEXT PRIMARY KEY,
holder_pid INT NOT NULL,
holder_host TEXT,
acquired_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
ttl_expires_at TIMESTAMPTZ NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_cycle_locks_ttl ON gbrain_cycle_locks(ttl_expires_at);
-- NOTIFY trigger for real-time job events (Postgres only, not PGLite)
CREATE OR REPLACE FUNCTION notify_minion_job_change() RETURNS trigger AS $$
BEGIN
+228
View File
@@ -0,0 +1,228 @@
/**
* `gbrain agent` CLI tests. Covers arg parsing, --since parser, and the
* submit path end-to-end against PGLite so we verify trusted submission,
* protected-name guard, and fan-out wiring.
*
* The full handler-run loop is NOT exercised here (tested in subagent-
* handler.test.ts). This file checks the CLI's submission + orchestration
* glue.
*/
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
import * as fs from 'node:fs';
import * as path from 'node:path';
import * as os from 'node:os';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { MinionQueue } from '../src/core/minions/queue.ts';
import { __testing as agentTesting } from '../src/commands/agent.ts';
import { parseSince } from '../src/commands/agent-logs.ts';
import { isProtectedJobName, PROTECTED_JOB_NAMES } from '../src/core/minions/protected-names.ts';
let engine: PGLiteEngine;
let queue: MinionQueue;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({ database_url: '' });
await engine.initSchema();
queue = new MinionQueue(engine);
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await engine.executeRaw('DELETE FROM minion_jobs');
});
describe('parseRunFlags', () => {
test('follow defaults off when stdout is non-TTY (test env)', () => {
const { flags, rest } = agentTesting.parseRunFlags(['hello', 'world']);
expect(flags.follow).toBe(process.stdout.isTTY === true);
expect(rest).toEqual(['hello', 'world']);
});
test('flags before prompt are parsed, unknown token ends flag parsing', () => {
const { flags, rest } = agentTesting.parseRunFlags([
'--model', 'claude-opus-4-7', '--max-turns', '30', 'summarize', 'everything',
]);
expect(flags.model).toBe('claude-opus-4-7');
expect(flags.maxTurns).toBe(30);
expect(rest).toEqual(['summarize', 'everything']);
});
test('--tools comma-split', () => {
const { flags } = agentTesting.parseRunFlags(['--tools', 'brain_search, brain_get_page', 'prompt']);
expect(flags.tools).toEqual(['brain_search', 'brain_get_page']);
});
test('--detach implies !follow', () => {
const { flags } = agentTesting.parseRunFlags(['--detach', 'x']);
expect(flags.detach).toBe(true);
expect(flags.follow).toBe(false);
});
test('double-dash ends flag parsing explicitly', () => {
const { flags, rest } = agentTesting.parseRunFlags(['--model', 'm', '--', '--not-a-flag']);
expect(flags.model).toBe('m');
expect(rest).toEqual(['--not-a-flag']);
});
test('unknown flag throws', () => {
expect(() => agentTesting.parseRunFlags(['--what', 'x'])).toThrow(/unknown flag/);
});
test('--subagent-def + --timeout-ms parsed', () => {
const { flags } = agentTesting.parseRunFlags([
'--subagent-def', 'researcher', '--timeout-ms', '60000', 'hello',
]);
expect(flags.subagentDef).toBe('researcher');
expect(flags.timeoutMs).toBe(60000);
});
test('--fanout-manifest parsed', () => {
const { flags } = agentTesting.parseRunFlags(['--fanout-manifest', '/tmp/m.json']);
expect(flags.fanoutManifest).toBe('/tmp/m.json');
});
});
describe('parseSince', () => {
test('returns undefined on empty input', () => {
expect(parseSince(undefined)).toBeUndefined();
expect(parseSince('')).toBeUndefined();
});
test('parses ISO-8601 timestamps', () => {
const iso = '2026-04-20T12:00:00.000Z';
expect(parseSince(iso)).toBe(iso);
});
test('parses relative 5m', () => {
const out = parseSince('5m')!;
const parsed = new Date(out).getTime();
const now = Date.now();
expect(now - parsed).toBeGreaterThanOrEqual(5 * 60 * 1000 - 1000);
expect(now - parsed).toBeLessThan(5 * 60 * 1000 + 1000);
});
test('parses relative 2h', () => {
const out = parseSince('2h')!;
const delta = Date.now() - new Date(out).getTime();
expect(delta).toBeGreaterThanOrEqual(2 * 3600 * 1000 - 1000);
});
test('parses relative 1d', () => {
const out = parseSince('1d')!;
const delta = Date.now() - new Date(out).getTime();
expect(delta).toBeGreaterThanOrEqual(86_400_000 - 1000);
});
test('throws on unparseable input', () => {
expect(() => parseSince('not-a-date')).toThrow(/could not parse/);
});
});
describe('protected-name guard includes subagent + aggregator', () => {
test('shell stays protected', () => {
expect(isProtectedJobName('shell')).toBe(true);
expect(PROTECTED_JOB_NAMES.has('shell')).toBe(true);
});
test('subagent is protected (v0.15)', () => {
expect(isProtectedJobName('subagent')).toBe(true);
});
test('subagent_aggregator is protected (v0.15)', () => {
expect(isProtectedJobName('subagent_aggregator')).toBe(true);
});
test('a random non-protected name is not protected', () => {
expect(isProtectedJobName('sync')).toBe(false);
});
test('trim normalization still blocks " subagent "', () => {
expect(isProtectedJobName(' subagent ')).toBe(true);
});
});
describe('queue.add trusted-submit gate for subagent', () => {
test('subagent without allowProtectedSubmit throws', async () => {
await expect(queue.add('subagent', { prompt: 'hi' })).rejects.toThrow();
});
test('subagent with allowProtectedSubmit succeeds', async () => {
const job = await queue.add('subagent', { prompt: 'hi' }, {}, { allowProtectedSubmit: true });
expect(job.name).toBe('subagent');
expect(job.status).toBe('waiting');
});
test('subagent_aggregator gated the same way', async () => {
await expect(queue.add('subagent_aggregator', { children_ids: [] })).rejects.toThrow();
const ok = await queue.add('subagent_aggregator', { children_ids: [1] }, {}, {
allowProtectedSubmit: true,
});
expect(ok.name).toBe('subagent_aggregator');
});
});
describe('fan-out manifest shape (integration)', () => {
test('fanout-manifest with 3 entries creates 3 subagent children + 1 aggregator', async () => {
// Manually replicate what runAgentRun does for --fanout-manifest > 1.
// We don't invoke runAgentRun (it calls process.exit on error) — we
// assert that the plumbing works via direct queue calls with the
// same flags it uses.
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'fanout-'));
try {
const manifestPath = path.join(tmp, 'm.json');
fs.writeFileSync(manifestPath, JSON.stringify([
{ prompt: 'chunk 1' }, { prompt: 'chunk 2' }, { prompt: 'chunk 3' },
]));
// Aggregator first.
const agg = await queue.add(
'subagent_aggregator',
{ children_ids: [] },
{ max_stalled: 3 },
{ allowProtectedSubmit: true },
);
const kids: number[] = [];
for (const p of ['chunk 1', 'chunk 2', 'chunk 3']) {
const c = await queue.add(
'subagent',
{ prompt: p },
{ parent_job_id: agg.id, on_child_fail: 'continue', max_stalled: 3 },
{ allowProtectedSubmit: true },
);
kids.push(c.id);
}
await engine.executeRaw(
`UPDATE minion_jobs SET data = jsonb_set(data, '{children_ids}', $1::jsonb) WHERE id = $2`,
[JSON.stringify(kids), agg.id],
);
// Aggregator should be in waiting-children since kids were submitted
// with parent_job_id = agg.id (Lane 1B behavior).
const aggNow = await queue.getJob(agg.id);
expect(aggNow?.status).toBe('waiting-children');
// Aggregator's data.children_ids reflects the spawned children.
const dataRow = await engine.executeRaw<{ data: unknown }>(
`SELECT data FROM minion_jobs WHERE id = $1`, [agg.id],
);
const data = typeof dataRow[0]!.data === 'string'
? JSON.parse(dataRow[0]!.data as string)
: dataRow[0]!.data as Record<string, unknown>;
expect(data.children_ids).toEqual(kids);
// Each child should have on_child_fail = 'continue'.
const childRows = await engine.executeRaw<{ on_child_fail: string }>(
`SELECT on_child_fail FROM minion_jobs WHERE parent_job_id = $1`, [agg.id],
);
expect(childRows.length).toBe(3);
expect(childRows.every(r => r.on_child_fail === 'continue')).toBe(true);
} finally {
fs.rmSync(tmp, { recursive: true, force: true });
}
});
});
+9 -8
View File
@@ -103,10 +103,10 @@ describe('buildPlan — diff against completed + installed VERSION', () => {
expect(plan.partial).toEqual([]);
expect(plan.pending.map(m => m.version)).toContain('0.11.0');
// Future migrations (registered but newer than installed VERSION) land in
// skippedFuture until the binary catches up. v0.13.0 = frontmatter graph
// (master), v0.13.1 = Knowledge Runtime grandfather, v0.14.0 = shell
// jobs + autopilot cooperative.
expect(plan.skippedFuture.map(m => m.version)).toEqual(['0.12.0', '0.12.2', '0.13.0', '0.13.1', '0.14.0']);
// skippedFuture until the binary catches up. v0.13.0 = frontmatter graph,
// v0.13.1 = Knowledge Runtime grandfather, v0.14.0 = shell jobs +
// autopilot cooperative, v0.16.0 = subagent runtime (this branch).
expect(plan.skippedFuture.map(m => m.version)).toEqual(['0.12.0', '0.12.2', '0.13.0', '0.13.1', '0.14.0', '0.16.0']);
});
test('already applied → v0.11.0 lands in `applied` bucket, not pending', () => {
@@ -142,10 +142,11 @@ describe('buildPlan — diff against completed + installed VERSION', () => {
const idx = indexCompleted([]);
const plan = buildPlan(idx, '0.12.0');
expect(plan.pending.map(m => m.version)).toContain('0.11.0');
// v0.12.2, v0.13.0, v0.13.1, and v0.14.0 were added later; installed=0.12.0
// means they belong in skippedFuture, not pending. v0.11.0 and v0.12.0
// stay pending despite being ≤ installed — that is the H9 invariant.
expect(plan.skippedFuture.map(m => m.version)).toEqual(['0.12.2', '0.13.0', '0.13.1', '0.14.0']);
// v0.12.2, v0.13.0, v0.13.1, v0.14.0, and v0.16.0 were added later;
// installed=0.12.0 means they belong in skippedFuture, not pending. v0.11.0
// and v0.12.0 stay pending despite being ≤ installed — that is the H9
// invariant.
expect(plan.skippedFuture.map(m => m.version)).toEqual(['0.12.2', '0.13.0', '0.13.1', '0.14.0', '0.16.0']);
});
test('--migration filter narrows to one version', () => {
+1 -1
View File
@@ -425,7 +425,7 @@ async function measureBaselineRelational(
}
const ENTITY_REF_RE = /\[[^\]]+\]\(([^)]+)\)|\b((?:people|companies|meetings|concepts)\/[a-z0-9-]+)\b/gi;
const perQuery: Array<{ question: string; expected: number; found: number }> = [];
const perQuery: Array<{ question: string; expected: number; found: number; returned: number }> = [];
let totalExpected = 0, totalFound = 0;
let totalReturned = 0, totalValid = 0;
+177
View File
@@ -0,0 +1,177 @@
/**
* Subagent brain-tool registry tests. Covers:
* - every allow-list name exists in OPERATIONS (catches renames upstream)
* - Anthropic tool-name constraint enforced
* - put_page schema is namespace-wrapped per subagent
* - execute() invokes the op handler with viaSubagent=true + subagentId
* - filterAllowedTools narrows registry + rejects unknown names
* - denied ops (file_upload etc.) do NOT appear in the registry
*/
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { operations, OperationError } from '../src/core/operations.ts';
import {
BRAIN_TOOL_ALLOWLIST,
buildBrainTools,
filterAllowedTools,
__testing,
} from '../src/core/minions/tools/brain-allowlist.ts';
import type { GBrainConfig } from '../src/core/config.ts';
import type { ToolCtx } from '../src/core/minions/types.ts';
let engine: PGLiteEngine;
const config: GBrainConfig = { engine: 'pglite' } as GBrainConfig;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({ database_url: '' });
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await engine.executeRaw('DELETE FROM pages');
});
describe('BRAIN_TOOL_ALLOWLIST', () => {
test('every name exists in src/core/operations.ts OPERATIONS', () => {
const opNames = new Set(operations.map(o => o.name));
const missing = [...BRAIN_TOOL_ALLOWLIST].filter(n => !opNames.has(n));
expect(missing).toEqual([]);
});
test('contains the read-only 10 + put_page', () => {
expect(BRAIN_TOOL_ALLOWLIST.size).toBe(11);
expect(BRAIN_TOOL_ALLOWLIST.has('query')).toBe(true);
expect(BRAIN_TOOL_ALLOWLIST.has('search')).toBe(true);
expect(BRAIN_TOOL_ALLOWLIST.has('get_page')).toBe(true);
expect(BRAIN_TOOL_ALLOWLIST.has('list_pages')).toBe(true);
expect(BRAIN_TOOL_ALLOWLIST.has('put_page')).toBe(true);
});
test('does NOT contain destructive ops', () => {
expect(BRAIN_TOOL_ALLOWLIST.has('file_upload')).toBe(false);
expect(BRAIN_TOOL_ALLOWLIST.has('delete_page')).toBe(false);
expect(BRAIN_TOOL_ALLOWLIST.has('delete_file')).toBe(false);
expect(BRAIN_TOOL_ALLOWLIST.has('sync')).toBe(false);
});
});
describe('buildBrainTools', () => {
test('produces one ToolDef per allow-listed op that exists in operations.ts', () => {
const tools = buildBrainTools({ subagentId: 42, engine, config });
const opNames = new Set(operations.map(o => o.name));
const expected = [...BRAIN_TOOL_ALLOWLIST].filter(n => opNames.has(n)).length;
expect(tools.length).toBe(expected);
});
test('tool names are brain_<op> and match Anthropic constraint', () => {
const tools = buildBrainTools({ subagentId: 7, engine, config });
for (const t of tools) {
expect(t.name).toMatch(__testing.ANTHROPIC_NAME_RE);
expect(t.name.startsWith('brain_')).toBe(true);
}
});
test('tools are flagged idempotent in v0.15', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
expect(tools.every(t => t.idempotent === true)).toBe(true);
});
test('tools carry the op description verbatim', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
const getPage = tools.find(t => t.name === 'brain_get_page');
const op = operations.find(o => o.name === 'get_page');
expect(getPage?.description).toBe(op!.description);
});
test('put_page schema is namespace-wrapped per subagent', () => {
const tools42 = buildBrainTools({ subagentId: 42, engine, config });
const putPage42 = tools42.find(t => t.name === 'brain_put_page');
const slug42 = ((putPage42!.input_schema as any).properties as any).slug;
expect(slug42.pattern).toBe('^wiki/agents/42/.+');
expect(slug42.description).toContain('wiki/agents/42/');
const tools7 = buildBrainTools({ subagentId: 7, engine, config });
const putPage7 = tools7.find(t => t.name === 'brain_put_page');
const slug7 = ((putPage7!.input_schema as any).properties as any).slug;
expect(slug7.pattern).toBe('^wiki/agents/7/.+');
});
test('non-put_page tools do NOT get a pattern on slug', () => {
const tools = buildBrainTools({ subagentId: 42, engine, config });
const getPage = tools.find(t => t.name === 'brain_get_page');
const slug = ((getPage!.input_schema as any).properties as any).slug;
expect(slug).toBeDefined();
expect(slug.pattern).toBeUndefined();
});
test('execute() on put_page with valid namespace slug succeeds', async () => {
const tools = buildBrainTools({ subagentId: 42, engine, config });
const putPage = tools.find(t => t.name === 'brain_put_page');
const ctx: ToolCtx = { engine, jobId: 1, remote: true };
const res = await putPage!.execute(
{ slug: 'wiki/agents/42/notes', content: '---\ntitle: Notes\n---\nbody' },
ctx,
);
expect(res).toBeTruthy();
});
test('execute() on put_page with out-of-namespace slug throws permission_denied', async () => {
const tools = buildBrainTools({ subagentId: 42, engine, config });
const putPage = tools.find(t => t.name === 'brain_put_page');
const ctx: ToolCtx = { engine, jobId: 1, remote: true };
await expect(
putPage!.execute(
{ slug: 'wiki/analysis/stomp', content: '---\ntitle: x\n---\nb' },
ctx,
),
).rejects.toBeInstanceOf(OperationError);
});
});
describe('filterAllowedTools', () => {
test('passes prefixed names through', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
const filtered = filterAllowedTools(tools, ['brain_get_page', 'brain_search']);
expect(filtered.map(t => t.name)).toEqual(['brain_get_page', 'brain_search']);
});
test('accepts un-prefixed names as a convenience', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
const filtered = filterAllowedTools(tools, ['get_page', 'search']);
expect(filtered.map(t => t.name)).toEqual(['brain_get_page', 'brain_search']);
});
test('rejects unknown tool names (no silent ignore)', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
expect(() => filterAllowedTools(tools, ['brain_typo_nope'])).toThrow(/unknown tool/);
});
test('deduplicates when both prefixed + unprefixed given', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
const filtered = filterAllowedTools(tools, ['brain_get_page', 'get_page']);
expect(filtered.length).toBe(1);
});
test('empty array yields empty registry', () => {
const tools = buildBrainTools({ subagentId: 1, engine, config });
expect(filterAllowedTools(tools, [])).toEqual([]);
});
});
describe('sanitizeToolName', () => {
test('returns within 64 chars', () => {
// Synthetic: simulate an op name long enough to need slicing.
const long = 'a'.repeat(100);
expect(__testing.sanitizeToolName(long).length).toBeLessThanOrEqual(64);
});
test('replaces non-conforming chars with _', () => {
expect(__testing.sanitizeToolName('foo.bar')).toBe('brain_foo_bar');
});
});

Some files were not shown because too many files have changed in this diff Show More