Files
gbrain/test/bootstrap-verify.serial.test.ts
T
Garry TanandClaude Fable 5 6411150071 v0.45.11.0 feat(bootstrap): OOBE hand-off (own the brain + cold-start) + DX polish waves (#4047)
* feat(bootstrap): TTY DX exploration harness + Krug onboarding fix wave

Add a real-PTY exploration harness and land 16 verified "Don't Make Me
Think" fixes on the paste-in install experience for Claude Code and Codex.

Harness:
- test/helpers/tty-harness.ts — spawns any CLI (gbrain/claude/codex) under a
  real pseudo-terminal (Bun terminal: spawn), timestamps every output burst,
  and turns silence windows into a measurable stall report. Hermetic; pure
  helpers unit-tested in test/tty-harness.test.ts.
- scripts/dx-explore.ts — drives the fresh-user funnel (help / init / real
  claude-install / real codex-install / manual drive mode), writing
  transcripts to .context/dx-runs/ (gitignored).

Fixes (all adversarially verified against the code first):
- Keyless bare `gbrain init` completes in keyless mode instead of exit 1;
  multi-key non-TTY auto-picks the canonical default; typo stays fail-loud.
- Provider picker probe-gates ollama (daemon-up != model-pulled) and offers
  an explicit "continue keyless" option that is the bare-Enter default.
- Fresh-brain init prints one schema-setup line instead of ~240 migration
  names (GBRAIN_MIGRATE_VERBOSE=1 restores detail).
- Init epilogue: memory-verbs funnel is last-on-screen; skills advisory
  compacted for init; Mod Status trimmed.
- PGLite live-serve lock error names the fix (close the agent session).
- Mode-picker banner interpolates the applied mode; expansion-key gate is
  Anthropic/OpenAI/Google, not OpenAI-only.
- Missing `claude` binary skips MCP but still installs hooks; honest copy.
- Foreign MCP-registration removal targets the conflicting scope and fails
  loud if it does not land.
- Upgrade marker compares the running binary to latest and self-spawns via
  execPath, so a current/newer binary no longer nags from a stale cache.
- interview --set/--skip after --confirm warns it voided the confirmation.
- init --help matches behavior; init --supabase fails loud on non-TTY.
- Provider capabilities attributed per provider across README / runbook /
  questions bank / bootstrap.md.
- First-run tour: restart-first, prompt 3 true on day one, withheld on FAIL;
  README gives Codex the same scripted magic moment.
- Empty-brain "0 takes" onboard nudge suppressed.
- Broken settings.local.json aborts the hooks write fail-closed instead of
  silently dropping the user's permissions.

Regenerated cli-flag-registry.generated.ts and llms-full.txt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): second DX polish wave — clean the success screen + honest copy (F17-F21)

Follow-up to the DX fix wave, closing the top-5 remaining gaps the scorecard
flagged (all human-facing polish, not survival):

F17 — machine markers no longer leak to humans:
- verify report drops the `[D3.6]` plan-tag from the first_run_tour detail.
- the raw `UPGRADE_AVAILABLE <cur> <latest>` marker line prints ONLY on a
  non-TTY stderr (parsers still get it); an interactive human sees just the
  "gbrain X -> Y available" sentence.
- per-migration "what changed" notices (v123/v124, incl. the #2704 ref) are
  suppressed on a FRESH-install replay via a module quiet flag; upgrades still
  narrate. (GBRAIN_MIGRATE_VERBOSE=1 restores them.)

F18 — one obvious next action on the init success screen: the memory-verbs
demo is the single "→ Do this next" hero, last on screen; import/migrate/doctor
collapse into one terse "More:" footer; the graph block only shows for a
non-empty brain.

F19 — README "moment it clicks" is now the genuine cross-session brain
round-trip (remember → restart → recall), explicitly distinguished from the
identity-file recall, on both the Codex and Claude Code paths.

F20 — the compact init skills advisory is human-voiced (no `[AGENT]`
stage-direction on the human-facing success screen; the mode-picker's
agent-directed block stays gated to the non-TTY channel).

F21 — time promise reconciled: headline is ~15 min (personal-agent path) /
~30 min (always-on OpenClaw/Hermes); the runbook's search-mode line no longer
claims "balanced" when keyless applies "conservative". README hooks copy says
"on by default, with an opt-out" to match the runbook.

Regenerated llms-full.txt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bootstrap): address two-model adversarial review of the DX wave

Fixes the regressions the 5-specialist + red-team + Claude/Codex adversarial
pass found in the F1–F21 changes, each with a test:

- Keyless upgrade hint pointed at `config set embedding_model`, which config.ts
  hard-refuses as a schema-sizing no-op — now names the working re-init recipe
  (`gbrain init --force --pglite --embedding-model <id>`), zero-key AND multi-key
  paths.
- Multi-key TTY picker offered "continue keyless" but the caller aborted on it —
  now honors keyless like the zero-key path.
- Detached update-refresh spawn used a `/gbrain$/` basename check that misfires
  for a renamed/official-named compiled binary (`gbrain-darwin-arm64`) and
  prepends the /$bunfs entrypoint — now detects dev-vs-compiled by the runtime
  basename (bun|node) so the refresh always runs.
- `bootstrap status` reported the wire phase "done" on a hooks-only receipt
  (host CLI missing at wire time) — now "partial" with a re-run hint, so a
  resuming agent doesn't trust a false complete.
- Post-repair MCP mismatch re-verifies and aborts instead of blessing a
  registration a racing writer may have re-claimed.
- probeOpenAICompat's abort timer now spans the body read (was cleared before
  it), so a stalled `/v1/models` body can't hang init past the 1s cap.
- Centralized the 4-copy stale-cache upgrade predicate into
  `pendingUpgradeVersion`; UPGRADE_AVAILABLE gains a GBRAIN_FORCE_UPGRADE_MARKER
  override for PTY-based agent harnesses.
- Mode picker's expansion-key gate adds GEMINI_API_KEY; picker prompt is
  article-aware ("an embedding" / "a chat"); dead `!brainEmpty` clause removed;
  migrate.ts try/finally widened + stamp failures named in quiet mode.
- DX harness: credential copies scrubbed even on SIGINT/interrupt (+chmod 600),
  child process TREE reaped on teardown, advisory made fail-open, KEY_MAP typed
  as a literal union.

New tests: migrate quiet-replay, self-upgrade pending predicate + negative
cache cases, bootstrap 127/scoped-remove/broken-settings dispatch, interview
invalidation flag, verify tour-withheld-on-FAIL, init keyless/supabase/multi-key,
init-nudge branches, ai-probes model parsing. Regenerated flag registry +
template-repo.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.45.8.0 fix(bootstrap): onboarding DX polish wave (F17-F21) + review fixes

DX fix wave on the paste-in install/first-run experience for Claude Code and
Codex, driven by a new real-PTY exploration harness. Keyless init completes
instead of erroring, the migration wall collapses to one line, the success
screen leads with one action, and the "magic moment" copy points at the genuine
cross-session round-trip. Full detail in CHANGELOG.

Version trio + openclaw manifest + runbook stamp bumped to 0.45.8.0; CHANGELOG
release entry; TODOS onboarding-DX follow-ups filed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(bootstrap): OOBE hand-off — you own the brain, cold-start is skill #1

A working install now ends by making the two facts that matter actually land:

- `gbrain bootstrap verify` prints (and returns as `handoff` in --json) an
  ownership block — the actual private-repo URL with what owning it means
  (read it, `gbrain bootstrap attach` on machine two, delete it and the brain
  is gone), or the local-only variant pointing at `gbrain bootstrap repo` —
  followed by the ONE next action: run the cold-start skill (Gmail/calendar/
  contacts via ClawVisor, an OAuth vault so the agent never holds raw tokens;
  or offline archives), one consented phase at a time. Withheld on FAIL like
  the tour; shape stays unconditional for machine consumers.
- cold-start ships in the downstream bundle (61 skills): its plugin exclusion
  ("host onboarding flow") predated the v0.45 personal-agent bootstrap and is
  deliberately reversed — the paste-in audience is exactly who day-one
  onboarding is for. It now LEADS the recommended set (ahead of book-mirror:
  every flagship skill only becomes magical once the brain holds the user's
  real life).
- New drift guard: every recommended slug must be scaffoldable from the
  plugin bundle — recommended-but-unscaffoldable is a dead-end CTA and now
  fails the suite.
- Runbook Hand off rewritten around the two must-land facts + the on-the-spot
  cold-start offer; README's Codex and Claude Code paths carry the same two
  follow-ups.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* v0.45.10.0 feat(bootstrap): the OOBE hand-off release

Version trio + runbook stamp + template tree to 0.45.10.0; CHANGELOG entry;
llms bundles regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): stop memory-verbs-conformance leaking a fake-keyed gateway into shard-mates

The deterministic-embedder helper configures the MODULE-GLOBAL gateway with a
fake OpenAI key; the file's afterAll never reset it. The bunfig preload's
per-test restore only fires when the gateway is UNCONFIGURED, so the fake-keyed
config persisted for every later file in the shard process — turn-context's
corpus writes then embedded against real OpenAI and 401'd (CI shard-8 failure;
shard re-binning from this branch's new test files exposed it).

Fix both sides: conformance's afterAll now resetGateway()s back to the preload
baseline and nulls both test transports; turn-context's beforeAll does the same
defensively so it stays hermetic regardless of shard composition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 15:35:05 -07:00

476 lines
20 KiB
TypeScript

/**
* bootstrap verify — the D8 check union against a hermetic in-memory PGLite
* brain (the resolve-ipc-v2 pattern): keyless roundtrip [G1/CX-P0.5],
* graph floor [CX2-5], magic moment via the ## Facts fence, probe cleanup
* [G13], failure-shape checks (token sweep / byte floors / secret scan), and
* snapshot retention [B2].
*/
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { addSource } from '../src/core/sources-ops.ts';
import { readManifest, writeManifest } from '../src/core/bootstrap/format.ts';
import { initState, setAnswer, confirm, readBackHash } from '../src/core/bootstrap/interview.ts';
import { loadQuestionBank } from '../src/core/bootstrap/assets.ts';
import { byteFloors } from '../src/core/bootstrap/render.ts';
import {
verifyWorkspace,
VERIFY_PROBE_SLUG,
VERIFY_PROBE_ENTITY_SLUG,
VERIFY_MAGIC_TOKEN,
VERIFY_SNAPSHOTS_KEPT,
FIRST_RUN_TOUR,
} from '../src/core/bootstrap/verify.ts';
import { listVerifyRuns } from '../src/core/bootstrap/status.ts';
import type { CapabilityReport } from '../src/core/capability.ts';
import { operations, type OperationContext } from '../src/core/operations.ts';
import { loadCorpusPages, loadCorpusQueries } from './helpers/bootstrap-corpus.ts';
const KEYLESS: CapabilityReport = {
embeddings: { available: false },
extraction: { available: false },
search: 'keyword-only',
mode: 'keyless',
};
let engine: PGLiteEngine;
let tmpParent: string;
let home: string;
let ws: string;
let prevHome: string | undefined;
function pad(text: string, bytes: number): string {
let out = text;
while (Buffer.byteLength(out, 'utf8') < bytes) {
out += '\nSubstantive identity prose rendered from real interview answers, not filler headers.';
}
return out;
}
beforeAll(async () => {
tmpParent = mkdtempSync(join(tmpdir(), 'gb-verify-'));
home = join(tmpParent, '.gbrain');
mkdirSync(join(home, 'bootstrap'), { recursive: true });
ws = mkdtempSync(join(tmpdir(), 'gb-verify-ws-'));
prevHome = process.env.GBRAIN_HOME;
process.env.GBRAIN_HOME = tmpParent;
// Workspace shape: initialized manifest + interview answers (6 required,
// confirmed) + identity files above the 6-answer byte floors + brain/.
writeManifest(ws, {
format_version: 1,
initialized: true,
agent_name: 'Verify Test Agent',
created_by: 'test',
created_at: new Date().toISOString(),
source_id: 'workspace',
});
const bank = loadQuestionBank();
const required = bank.interviewKeys.filter((k) => bank.questions[k]?.required === true);
expect(initState(ws).ok).toBe(true);
const answers: Record<string, string> = {
AGENT_NAME: 'Testa',
PRINCIPAL_NAME: 'Pat Example',
AGENT_PURPOSE: 'Maintain the research corpus and draft the weekly memo without re-briefing.',
AGENT_TOP_JOBS: '- corpus upkeep\n- weekly memo\n- meeting prep',
PRINCIPAL_CONTEXT: 'Runs a small research group; builds internal tooling; cares about signal over noise.',
VOICE_REGISTER: 'Direct: three options, the second one wins.',
};
for (const key of required) {
const r = setAnswer(ws, key, answers[key] ?? `a real answer for ${key}`);
if (!r.ok) throw new Error(`setAnswer(${key}) failed: ${r.message}`);
}
const h = readBackHash(ws);
if (!h.ok) throw new Error(h.message);
expect(confirm(ws, h.hash).ok).toBe(true);
const floors = byteFloors(required.length);
writeFileSync(join(ws, 'SOUL.md'), pad('# Soul\n\nIdentity rendered from answers.\n', floors['SOUL.md'] + 200));
writeFileSync(join(ws, 'USER.md'), pad('# User\n\nTheir literal words are ground truth.\n', floors['USER.md'] + 100));
mkdirSync(join(ws, 'brain'), { recursive: true });
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
// The workspace source: brain/ registered with its own local_path so the
// put_page write-through materializes committed files under brain/ [G1].
await addSource(engine, { id: 'workspace', localPath: join(ws, 'brain'), force: true });
}, 240_000);
afterAll(async () => {
try {
await engine.disconnect();
} catch {
/* noop */
}
if (prevHome === undefined) delete process.env.GBRAIN_HOME;
else process.env.GBRAIN_HOME = prevHome;
rmSync(tmpParent, { recursive: true, force: true });
rmSync(ws, { recursive: true, force: true });
});
function check(checks: Array<{ id: string; ok: boolean; warn?: boolean; detail: string }>, id: string) {
const found = checks.filter((c) => c.id === id);
expect(found.length).toBeGreaterThan(0);
return found;
}
describe('verifyWorkspace — keyless pass', () => {
test('roundtrip + graph floor + magic moment pass with ZERO api keys; probes cleaned up', async () => {
// Snapshot-retention setup [B2]: pre-seed old snapshots so the prune path runs.
for (let i = 0; i < VERIFY_SNAPSHOTS_KEPT; i++) {
writeFileSync(
join(home, 'bootstrap', `verify-2020-01-0${i + 1}T00-00-00-000Z.json`),
JSON.stringify({ ts: `2020-01-0${i + 1}T00:00:00.000Z`, ok: true, checks: [] }),
);
}
const res = await verifyWorkspace(engine, ws, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
});
// Individual checks.
expect(check(res.checks, 'doctor_green')[0].ok).toBe(true);
expect(check(res.checks, 'token_sweep')[0].ok).toBe(true);
expect(check(res.checks, 'byte_floors')[0].ok).toBe(true);
expect(check(res.checks, 'secret_scan')[0].ok).toBe(true);
expect(check(res.checks, 'deny_globs')[0].ok).toBe(true);
expect(check(res.checks, 'repo_privacy')[0].ok).toBe(true); // local-only
// execution_env is informational and NEVER gates [D-cloud].
expect(check(res.checks, 'execution_env')[0].ok).toBe(true);
for (const c of check(res.checks, 'roundtrip')) expect(c.ok).toBe(true);
expect(check(res.checks, 'graph_floor')[0].ok).toBe(true);
expect(check(res.checks, 'magic_moment')[0].ok).toBe(true);
expect(check(res.checks, 'capability_report')[0].detail).toContain('keyless');
expect(check(res.checks, 'hooks_smoke')[0].ok).toBe(true); // not installed → not applicable
expect(check(res.checks, 'first_run_tour')[0].ok).toBe(true);
expect(res.ok).toBe(true);
// The report embeds the capability block + the three scripted prompts [D3.6/A4].
expect(res.report).toContain('keyless mode');
for (const prompt of FIRST_RUN_TOUR) {
expect(res.report).toContain(prompt);
}
expect(res.tour).toEqual([...FIRST_RUN_TOUR]);
// The OOBE hand-off block prints after the tour on PASS: ownership (this
// ws has no origin remote → the local-only variant with the repo upgrade
// path) and the ONE next action (the cold-start skill via ClawVisor).
expect(res.report).toContain('What you own');
expect(res.report).toContain('gbrain bootstrap repo');
expect(res.report).toContain('cold-start');
expect(res.report).toContain('ClawVisor');
expect(res.handoff.length).toBeGreaterThan(0);
// Probe cleanup [G13]: pages, files, and the reconciled fact are gone.
expect(existsSync(join(ws, 'brain', `${VERIFY_PROBE_SLUG}.md`))).toBe(false);
expect(existsSync(join(ws, 'brain', `${VERIFY_PROBE_ENTITY_SLUG}.md`))).toBe(false);
const facts = await engine.executeRaw<{ fact: string }>(
`SELECT fact FROM facts WHERE source_id = $1 AND fact LIKE $2`,
['workspace', `%${VERIFY_MAGIC_TOKEN}%`],
);
expect(facts.length).toBe(0);
// Snapshot persisted + retention holds at VERIFY_SNAPSHOTS_KEPT [B2].
const runs = listVerifyRuns(home);
expect(runs.length).toBeLessThanOrEqual(VERIFY_SNAPSHOTS_KEPT);
expect(runs[0].ok).toBe(true); // newest = this run
}, 240_000);
test('a user fact that merely CONTAINS the magic-token substring survives verify [8b]', async () => {
// A real user fact from a different page, plus a legacy NULL-slug fact,
// both mentioning the token STRING. The probe cleanup must scope to the
// probe page's source_markdown_slug, never a `fact LIKE %token%` match, so
// neither of these is swept.
await engine.executeRaw(
`INSERT INTO facts (source_id, fact, source, source_markdown_slug, visibility)
VALUES ($1, $2, 'user', $3, 'world')`,
['workspace', `a real note that mentions ${VERIFY_MAGIC_TOKEN} in passing`, 'wiki/real-user-note'],
);
await engine.executeRaw(
`INSERT INTO facts (source_id, fact, source, visibility)
VALUES ($1, $2, 'user', 'world')`,
['workspace', `a legacy NULL-slug fact mentioning ${VERIFY_MAGIC_TOKEN}`],
);
const res = await verifyWorkspace(engine, ws, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
});
expect(res.ok).toBe(true);
// The probe fact (source_markdown_slug = probe slug) IS cleaned up …
const probeFacts = await engine.executeRaw<{ id: number }>(
`SELECT id FROM facts WHERE source_id = $1 AND source_markdown_slug = $2`,
['workspace', VERIFY_PROBE_SLUG],
);
expect(probeFacts.length).toBe(0);
// … but BOTH user facts that merely mention the token string survive.
const survivors = await engine.executeRaw<{ fact: string }>(
`SELECT fact FROM facts WHERE source_id = $1 AND fact LIKE $2`,
['workspace', `%${VERIFY_MAGIC_TOKEN}%`],
);
expect(survivors.length).toBe(2);
// Cleanup so later tests start from a clean facts table.
await engine.executeRaw(`DELETE FROM facts WHERE source_id = $1 AND fact LIKE $2`, [
'workspace', `%${VERIFY_MAGIC_TOKEN}%`,
]);
}, 240_000);
test('failure shapes: unresolved token + under-floor USER.md + planted secret all surface', async () => {
const soulPath = join(ws, 'SOUL.md');
const userPath = join(ws, 'USER.md');
const githubPath = join(ws, 'GITHUB.md');
const soulOriginal = readFileSync(soulPath, 'utf8');
const userOriginal = readFileSync(userPath, 'utf8');
try {
writeFileSync(githubPath, '# GitHub\n\nRepo: {{GITHUB_REPO_URL}}\n');
writeFileSync(userPath, '# tiny\n');
writeFileSync(soulPath, soulOriginal + '\napi dump: sk-AAAAAAAAAAAAAAAAAAAAAAAA\n');
const res = await verifyWorkspace(engine, ws, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
});
const token = check(res.checks, 'token_sweep')[0];
expect(token.ok).toBe(false);
expect(token.detail).toContain('GITHUB.md');
expect(token.detail).toContain('GITHUB_REPO_URL');
const floors = check(res.checks, 'byte_floors')[0];
expect(floors.ok).toBe(false);
expect(floors.detail).toContain('USER.md');
const scan = check(res.checks, 'secret_scan')[0];
expect(scan.ok).toBe(false);
expect(scan.detail).toContain('openai');
// Redaction discipline: the finding detail NEVER carries the secret value.
expect(scan.detail).not.toContain('sk-AAAAAAAAAAAAAAAAAAAAAAAA');
expect(res.ok).toBe(false);
// Tour gating on FAIL: the report says fix-first and withholds the
// celebration prompts ("broken, but go enjoy it" is a mixed signal) …
expect(res.report).toContain('Fix the FAIL checks above');
expect(res.report).not.toContain('Who am I to you?');
// … the hand-off block is withheld with the tour (celebrating ownership
// of a FAILED install is the same mixed signal) …
expect(res.report).not.toContain('What you own');
expect(res.report).not.toContain('cold-start');
// … while the returned tour + handoff arrays stay unconditional so
// machine consumers (--json) keep a stable shape, and the check names
// the gate.
expect(res.tour).toEqual([...FIRST_RUN_TOUR]);
expect(res.handoff.length).toBeGreaterThan(0);
expect(check(res.checks, 'first_run_tour')[0].detail).toContain('withheld');
} finally {
rmSync(githubPath, { force: true });
writeFileSync(userPath, userOriginal);
writeFileSync(soulPath, soulOriginal);
}
}, 240_000);
});
describe('verifyWorkspace — engine-plane side effects', () => {
test('[CX-P1.1] facts.default_visibility: unset → set to world; explicit private → untouched', async () => {
const e2 = new PGLiteEngine();
await e2.connect({});
await e2.initSchema();
const ws2 = mkdtempSync(join(tmpdir(), 'gb-verify-vis-'));
try {
mkdirSync(join(ws2, 'brain'), { recursive: true });
expect(await e2.getConfig('facts.default_visibility')).toBeNull();
const res = await verifyWorkspace(e2, ws2, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
sweepBudgetMs: 5_000,
});
// Fresh engine: posture set to world through the engine config plane…
expect(await e2.getConfig('facts.default_visibility')).toBe('world');
// …and the report names the posture + where to flip it, in one line.
const check = res.checks.find((c) => c.id === 'facts_visibility')!;
expect(check.ok).toBe(true);
expect(check.detail).toContain('world');
expect(check.detail).toContain('facts.default_visibility');
expect(check.detail).toContain('gbrain config set');
// Pre-set explicit value survives verify (set-if-unset, never override).
await e2.setConfig('facts.default_visibility', 'private');
const res2 = await verifyWorkspace(e2, ws2, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
sweepBudgetMs: 5_000,
});
expect(await e2.getConfig('facts.default_visibility')).toBe('private');
const check2 = res2.checks.find((c) => c.id === 'facts_visibility')!;
expect(check2.detail).toContain('private');
expect(check2.detail).toContain('untouched');
} finally {
await e2.disconnect();
rmSync(ws2, { recursive: true, force: true });
}
}, 240_000);
test('mcp_surface=verbs in config → loud WARN naming the --surface full fix', async () => {
const cfgPath = join(home, 'config.json');
writeFileSync(cfgPath, JSON.stringify({ engine: 'pglite', mcp_surface: 'verbs' }), 'utf8');
try {
const res = await verifyWorkspace(engine, ws, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
});
const check = res.checks.find((c) => c.id === 'mcp_surface')!;
expect(check.ok).toBe(true);
expect(check.warn).toBe(true);
expect(check.detail).toContain('verbs');
expect(check.detail).toContain('--surface full');
} finally {
rmSync(cfgPath, { force: true });
}
}, 240_000);
});
describe('verifyWorkspace — source_id collision resolution', () => {
test('single workspace: id registered to THIS checkout stays unchanged (no collision check emitted)', async () => {
const res = await verifyWorkspace(engine, ws, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
});
const m = readManifest(ws);
expect(m.state).toBe('initialized');
if (m.state === 'initialized') expect(m.manifest.source_id).toBe('workspace');
expect(res.checks.find((c) => c.id === 'source_id')).toBeUndefined();
}, 240_000);
test("second workspace claiming a taken 'workspace' id derives workspace-<8char-hash> and persists it", async () => {
const e2 = new PGLiteEngine();
await e2.connect({});
await e2.initSchema();
const firstBrain = mkdtempSync(join(tmpdir(), 'gb-verify-first-brain-'));
const ws2 = mkdtempSync(join(tmpdir(), 'gb-verify-second-ws-'));
try {
// The brain already has 'workspace' registered to ANOTHER checkout.
await addSource(e2, { id: 'workspace', localPath: firstBrain, force: true });
mkdirSync(join(ws2, 'brain'), { recursive: true });
writeManifest(ws2, {
format_version: 1,
initialized: true,
agent_name: 'Second',
created_by: 'test',
created_at: new Date().toISOString(),
source_id: 'workspace',
});
const res = await verifyWorkspace(e2, ws2, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
sweepBudgetMs: 5_000,
});
const m = readManifest(ws2);
expect(m.state).toBe('initialized');
if (m.state !== 'initialized') throw new Error('unreachable');
const derived = m.manifest.source_id;
expect(derived).toMatch(/^workspace-[0-9a-f]{8}$/);
const check = res.checks.find((c) => c.id === 'source_id')!;
expect(check.warn).toBe(true);
expect(check.detail).toContain(derived);
expect(check.detail).toContain(`gbrain sources add ${derived}`);
expect(check.detail).toContain('hooks --repair');
// Re-run is stable: the derived id is not itself in collision.
const res2 = await verifyWorkspace(e2, ws2, {
sourceId: derived,
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
sweepBudgetMs: 5_000,
});
const m2 = readManifest(ws2);
if (m2.state !== 'initialized') throw new Error('unreachable');
expect(m2.manifest.source_id).toBe(derived);
expect(res2.checks.find((c) => c.id === 'source_id')).toBeUndefined();
} finally {
await e2.disconnect();
rmSync(firstBrain, { recursive: true, force: true });
rmSync(ws2, { recursive: true, force: true });
}
}, 240_000);
});
describe('verifyWorkspace — real corpus graph floor + qrels recall', () => {
test('auto-link builds real multi-entity edges; verify passes on the populated brain; gold queries recall the expected pages', async () => {
// Seed the synthetic world into the SAME workspace source verify checks,
// through the real put_page handler so auto-link builds REAL edges — a
// 12-page graph, not the 2-node self-planted probe pair verify writes.
const loaded = await loadCorpusPages(engine, { sourceId: 'workspace' });
expect(loaded).toContain('people/alice-example');
expect(loaded).toContain('companies/ridge-platform');
// Real edges exist from the wikilinks in the page bodies (not the probe):
// alice-example → ridge-platform, with multiple outbound edges on a real
// entity, and the reverse edge answers through the backlink table.
const aliceLinks = await engine.getLinks('people/alice-example', { sourceId: 'workspace' });
const aliceTargets = aliceLinks.map((l) => l.to_slug);
expect(aliceTargets).toContain('companies/ridge-platform');
expect(aliceTargets.length).toBeGreaterThanOrEqual(2);
const ridgeBacklinks = await engine.getBacklinks('companies/ridge-platform', { sourceId: 'workspace' });
expect(ridgeBacklinks.map((l) => l.from_slug)).toContain('people/alice-example');
// verify still runs green end-to-end on the now-populated brain, and its
// graph_floor / roundtrip checks pass alongside the real corpus graph.
const res = await verifyWorkspace(engine, ws, {
sourceId: 'workspace',
gbrainHomeDir: home,
capabilities: KEYLESS,
skipHooksSmoke: true,
});
if (!res.ok) console.error(res.report);
expect(res.ok).toBe(true);
expect(check(res.checks, 'graph_floor')[0].ok).toBe(true);
for (const c of check(res.checks, 'roundtrip')) expect(c.ok).toBe(true);
// qrels: each gold page-recall case returns its expected slug through the
// REAL keyword-search query op (keyless), over the multi-page corpus.
const queryOp = operations.find((o) => o.name === 'query')!;
const ctx: OperationContext = {
engine,
config: { engine: 'pglite' } as never,
logger: { info: () => {}, warn: () => {}, error: () => {} },
dryRun: false,
remote: false,
sourceId: 'workspace',
};
const goldPageCases = loadCorpusQueries().filter((q) => q.kind === 'page');
expect(goldPageCases.length).toBeGreaterThan(0);
const misses: string[] = [];
for (const q of goldPageCases) {
const result = await queryOp.handler(ctx, { query: q.query, limit: 10, expand: false });
if (!JSON.stringify(result).includes(q.expect_slug!)) {
misses.push(`${q.id}: "${q.query}" did not recall ${q.expect_slug}`);
}
}
expect(misses).toEqual([]);
}, 240_000);
});