mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 00:48:18 +00:00
* fix(schema): bootstrap timeline_entries.event_page_id forward reference — un-wedge pre-v121 upgrades (#2626 #2594 #2579 #2537 #2536) v0.42.56.0 (Chronicle, migration v121) added timeline_entries.event_page_id and two partial indexes in the embedded schema blobs without extending applyForwardReferenceBootstrap — any brain whose timeline_entries predates v121 wedged initSchema at blob replay ("column event_page_id does not exist") before runMigrations could apply v121, with no in-band recovery. - Add the timeline_entries.event_page_id probe + column-only ALTER to applyForwardReferenceBootstrap in BOTH engines; FK + partial indexes land via the idempotent v121 / blob replay afterwards. Stays in the always-run bootstrap (never a migration hook — those skip oddly-stamped brains). - REQUIRED_BOOTSTRAP_COVERAGE entry + strip blocks in both runtime tests. - e2e: pre-v121 rewind → full initSchema converges (indexes re-created); wedged-brain recovery — a brain that already FAILED the upgrade attempt converges on retry with full final shape (column + FK + both partial indexes) and no ledger residue. Absorbs PR #2548 (@chetan-guevara) and the e2e test from PR #2623 (@colinagent) — thank you both. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(schema): coverage guard cross-references migration-added columns — close the scan hole that shipped the v121 wedge The A2 static check treated a column as covered when the current CREATE TABLE body declared it. But a column that is BOTH in the blob's CREATE TABLE AND added by a migration is a forward reference by definition — on pre-existing tables CREATE TABLE IF NOT EXISTS no-ops and the blob's CREATE INDEX crashes initSchema before runMigrations can help. That mask is exactly how timeline_entries.event_page_id passed the guard while wedging every pre-v121 brain. - buildIndexRefCoveragePredicate: migration-added columns (from extractAddedColumnsFromMigrations over the MIGRATIONS array) require a bootstrap ALTER; CREATE TABLE presence no longer counts for them. - Unit test pins the v121 regression shape red/green with synthetic inputs; the A2 test pins the incident triple directly (migration-added + blob-indexed + bootstrap-covered). - The strengthened predicate immediately surfaced two more latent wedges of the same class: minion_jobs.timeout_at + minion_jobs.idempotency_key (migration v7, blob-indexed, unprobed). Added probes in both engines + coverage entries + runtime strip blocks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): serve --http hides the generated admin token from non-TTY output by default (#2624) Generated admin bootstrap tokens printed into container/log-aggregator stdout on every headless start. shouldSuppressBootstrapPrint now defaults to hidden unless stderr is an interactive TTY; env-sourced tokens are never printed; --print-admin-token is the explicit escape hatch for capturing the value on a trusted non-TTY start; --suppress-bootstrap-token still overrides everything. Unit-tested across all five postures. Absorbs PR #2625 (@irresi) — thank you. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): legacy bearer tokens honor permissions.takes_holders over serve --http (#2529) GBrainOAuthProvider.verifyAccessToken never returned takesHoldersAllowList, so the serve --http dispatch site always fell back to ['world'] — remote MCP callers with an operator-configured takes_holders grant saw only public takes. The legacy branch now extracts permissions.takes_holders exactly like src/mcp/http-transport.ts (fail-safe ['world'] default, non-string entries dropped, malformed permissions JSON fails closed without throwing), and AuthInfo carries the field as a typed contract. OAuth-registered clients have no takes_holders storage on oauth_clients; that lane is design work tracked in TODOS (column migration + DCR/ registration surface), not part of this hotfix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): get_chunks honors the federated source grant + stops shipping embedding vectors (#2555, getChunks half of #2544) The get_chunks op still used the pre-#2200 scalar pattern (ctx.sourceId ? {sourceId} : {}) and engine.getChunks had no sourceIds[] support — a federated client that could read a page via get_page got [] from get_chunks. The op now routes through sourceScopeOpts (canonical ladder: federated array > scalar floor > nothing) and both engines gain the getPage-style sourceIds[] precedence branch; the unset-opts 'default' floor is preserved for local callers (importCodeFile contract). While in the function: SELECT cc.* pulled every embedding vector over the wire per chunk only for rowToChunk to discard them — replaced with the explicit non-vector column list in both engines (the getChunks half of #2544; the per-put_page getAllSlugs half is tracked separately). getChunksWithEmbeddings stays scalar-only by design (engine-internal, zero remote-reachable callers — documented at the interface). Tests: op-level federated repro + isolation + default-floor bleed guard (PGLite), engine precedence + Chunk-shape pin, and a DATABASE_URL-gated engine-parity test covering all three scope shapes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(auth): /admin/api/register-client accepts source + federatedRead bindings (#2143 enabler) The HTTP register endpoint hardcoded source_id='default' and federated_read=undefined — only the CLI could mint a client bound to a non-default source, so HTTP-registered MCP clients wrote into 'default' regardless of intent. The endpoint now accepts optional source / federatedRead body fields, validated via assertValidSourceId with a structured 400 on bad input; omitting both preserves the historical default. The admin-UI form layer is a tracked follow-up. Absorbs PR #2016 — thank you. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scope,calibration): think reads over MCP; source-scoped takes reads; calibration CLI reachability + model resolution (#2078-class #2451) - think op → scope:'read' for OAuth/MCP clients: the handler already forces save/take off for remote callers before persistence, so a read-scoped token can think without a write grant; local CLI persistence unchanged. Scope-annotation test carries an explicit remote-gated allowlist. - takes_list / takes_search / takes_scorecard / takes_calibration route through sourceScopeOpts (the #2200 class on the takes read lane) with engine support in BOTH engines + tests. - 'calibration' added to CLI_ONLY (the command was registered but unreachable — dispatch-gap class) and calibration_profile/voice-gate resolve models through the canonical gateway tier resolver instead of bare ids that parseModelId rejects. - BigInt-safe local-op output normalization (bigintToStringReplacer, postgres.js wire parity) — first half of the #2450 fix; the formatResult default case lands with the cli-output commit. Absorbs PR #2598 (@colinagent) and PR #2452 (@spinsirr) — thank you both. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(cli): output correctness — BigInt-safe rendering, search --json, files bigint, and 5 unreachable commands (#2450 #2527 #2042 #2035-class) - normalizeLocalResult wraps the local-op output round-trip with the bigint→string replacer (postgres.js wire parity; a bare stringify THROWS on BIGSERIAL keys); formatResult's default renderer gets the same replacer so nothing upstream can crash it. - search/query --json: CLI-local formatter flag threaded through the shared formatter — stdout is a parseable result array, never human text on the --json path (the #2042 residual). - file_list normalizes size_bytes (Postgres BIGINT → Number) so MCP serialization and the CLI KB math survive; null preserved. - NEW dispatch-gap guard: every handleCliOnly top-level case label must be reachable via CLI_ONLY. It immediately caught FIVE live unreachable commands: pages, backfill, reconcile-links, notability-eval (added to CLI_ONLY), and the documented 'gbrain search modes|stats|tune' dashboards (pre-fix, 'search modes' silently keyword-searched the word "modes") — now routed via a pre-dispatch subcommand gate. 'whoknows' stays on its op-alias route (collision guard); tracked with PR #2509. Absorbs PR #2494 and PR #2531 (@javieraldape) and adapts PR #472 (@vinsew) — thank you. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(db): updateSourceConfig survives mixed-array config rows + repair/doctor cover the subagent jsonb columns (#2251) The array-coercion branch called jsonb_each(elem) bare — a mixed array (e.g. '["x", {"last_full_cycle_at": ...}]') threw 'cannot call jsonb_each on a non-object' DURING row production, permanently failing every subsequent updateSourceConfig (last_full_cycle_at could never be written again). Non-object elements are now neutralized inline via a CASE-guarded jsonb_each; the row self-heals to a flat object on the next write, object elements' keys recovered. Pinned ungated on PGLite (real Postgres semantics) and via a DATABASE_URL-gated e2e on the real engine. repair-jsonb + doctor's jsonb_integrity check extend from 5 to 8 columns (subagent_messages.content_blocks, subagent_tool_executions.input/output — historical damage rows from the pre-v0.42.53.0 positional double-encode; the write paths themselves were fixed in #2375) with a to_regclass skip for brains predating those tables. Adapts the repair/doctor extension from PR #597 (@vinsew) — thank you. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(cli): dry-run honesty — strict unknown-flag rejection CLI-wide + unify-types worker defaults to dry-run (#2185 #1575) #2185: 'gbrain init --migrate-only --dry-run' applied REAL migrations — flags are read ad hoc (args.includes) so anything a handler doesn't look for was silently ignored, including intent-bearing safety flags. The CLI now validates every flag pre-dispatch and pre-engine: - Op commands validate against the operation contract (op.params + CLI-local --json/--explain), mirroring parseOpArgs so flag values that begin with '--' are never misread. parseOpArgs also gains the --key=value inline form (previously parsed as a junk key that consumed the NEXT token). - CLI_ONLY commands validate against a GENERATED per-command registry (scripts/generate-flag-registry.ts scans each command's case block + imported modules + one level of relative imports; deliberately over-inclusive so a missed flag can't break a working invocation). Committed as src/core/cli-flag-registry.generated.ts; 'bun run build:flag-registry' regenerates. - Passthrough by construction: everything after '--', plus call / config / 'jobs submit' payloads (handler-defined params are their contract). - Guards: sweep test (every command × nonsense flag → error), acceptance tests (real flags, --no- negation, = form, -- passthrough), drift guard (every CLI_ONLY member has an entry), freshness guard (committed registry == fresh generator run), subprocess smokes incl. the literal #2185 repro failing loud with zero engine work. BREAKING: scripts passing stray flags now fail loud with "Unknown flag --x for 'gbrain <cmd>'" — that is the point. #1575: the unify-types worker registration passed apply ?? true while the handler documents 'Default false (dry-run)' — the canonical operator invocation destructively retyped 25K+ pages by default. Now ?? false with a structural test; explicit --params '{"apply":true}' is the only way to mutate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(todos): file fix-wave 1 follow-ups (OAuth takes_holders design, parseFlags end-state, whoknows routing, #2544 half, #1558 UI, #2536 diagnostics) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): memory-safe unit runner — adaptive concurrency + serial OOM rescue pass A default run (4 shards × 4 intra-shard files) holds up to 16 concurrent PGLite WASM instances (~1.5GB each). With sibling Conductor workspaces running their own suites, PGLite connect failed with 'Out of memory' across every shard at once — 369 phantom test failures on a healthy branch, indistinguishable from real breakage at a glance. Two default-on layers in scripts/run-unit-parallel.sh: 1. Memory-aware sizing: total concurrency is capped to available memory (vm_stat on macOS, MemAvailable on Linux) at GBRAIN_TEST_MEM_PER_FILE_MB (default 1536) per concurrent file, shedding shards before intra-shard width. Quiet machines are unaffected (banner: mem-ok); pressured ones degrade instead of OOMing (banner: mem-adapted AxB→CxD). 2. Serial OOM rescue: failures whose shard log carries the WASM out-of-memory signature are re-run at --max-concurrency 1 after the fan-out drains. Phantoms pass serially → run goes green with an oom_rescued note and the failure blocks marked superseded; real failures fail again and stay red. Plain assertion failures never match the signature and never enter the rescue lane (existing exit-code and failure-log contract tests unchanged). Escape hatches: GBRAIN_TEST_NO_MEM_ADAPT=1, GBRAIN_TEST_NO_OOM_FALLBACK=1. Tests: OOM-once fixture rescued to exit 0; kill-switch stays red; banner advertises the sizing verdict. Documented in docs/TESTING.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test,cli): merge-seam repairs — shard timeout for the tripled suite, init flag-error contract, think scope exemption Three post-merge repairs surfaced by the first full-suite run: - Shard timeout 1500s -> 3000s: the suite roughly tripled since the cap was sized (~3900 -> 11k+ tests; PGLite inits replay 120 migrations, was 92). Two shards were killed mid-progress at 1500s. - The #2185 pre-dispatch validator now emits the same error contract as init.ts's in-handler check it preempts: lowercase 'unknown flag' on stderr + structured {status:'error', reason:'invalid_flag'} on stdout for --json callers (pinned by test/init-migrate-only.test.ts). - test/operations-trust-boundary.test.ts gets the same documented remote-gated allowlist for think's read scope (#2598) that test/oauth.test.ts already carries. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): rescue lane also covers externally-killed shards (sibling-workspace pkill / memory jetsam) Second phantom class observed on a multi-workspace Conductor machine: 3 shards SIGTERM'd + 1 SIGKILL'd at ~700s under a 3000s cap, all mid-progress — an external killer, not a wedge. The dead shards then poisoned the serial pass (lock/state residue → 18 more phantoms), and every one of the 18 passed standalone. The runner now stamps per-shard start/end epochs; a shard dying on 143/137 before 80% of SHARD_TIMEOUT is classified externally-killed and its file list joins the serial rescue queue (real wedges die AT the cap and stay red). Serial-pass failures that occur while any shard was externally killed are treated as suspect residue and rescued too. Structural tests pin the detector, threshold, and routing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(auth): keep tokenEndpointAuthMethod terminal in the register-client destructure (PKCE structural contract) The #2016 absorb appended source/federatedRead after tokenEndpointAuthMethod; test/fix-wave-structural.test.ts pins tokenEndpointAuthMethod as the final destructured field (v0.36.1.x #1077 PKCE regression contract). The added fields move into the regex's optional-middle slot. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): re-apply the #2598 remote-gated scope exemption to master's oauth scope-annotation test Taking master's v0.42.74.0 oauth.test.ts (its #2529 implementation) dropped the think read-scope allowlist that PR #2598 carries; re-applied to match test/operations-trust-boundary.test.ts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): pre-landing review batch — CLI-local flag acceptance, json= coherence, re-landed getChunks trim, payload-aware jsonb repair, rescue-lane tightening Ship Step 9 findings (checklist pass + 4 specialists), all verified before fixing: - CRITICAL: the #2185 validator rejected --source/--dry-run on op commands (makeContext CLI-locals consumed outside the op contract) — 'gbrain search "x" --source y' exited 1. Exempted with parser-mirroring value consumption + unit/subprocess tests incl. global-flag acceptance. - CRITICAL: --json=<v> diverged between validator (accepted) and parseOpArgs (junk-key path consumed the NEXT token, corrupting positionals). Parser now handles --json=true|false; =-forms of bare-only CLI-locals reject loud. parseOpArgs inline-= suite added (regression rule). - CRITICAL: the master merge silently restored SELECT cc.* in both engines' getChunks while docs claimed the #2544 trim. Re-landed the explicit non-vector column list + a source-level structural pin so a merge can't silently undo it again. - repair-jsonb/doctor: the subagent columns legitimately hold jsonb string scalars (persistToolExec binds pre-serialized strings) — unconditional unwrap would abort the repair run or corrupt legit values. jsonPayloadOnly predicate (JSON-container content only) on those 3 targets, mirrored in doctor + parameterized to_regclass + behavioral test (damage flagged, legit string ignored, absent table skipped). - runner rescue-lane tightening: serial failures rescue-eligible only with their own OOM signature or after an external shard kill (residue), never because a sibling shard OOM'd — flaky serial tests stay red. Rescue passes no longer double-count into TOTAL_PASS; shard timeout scales when mem-adaptation sheds shards; negative-path tests (mixed run stays red, deterministic OOM-signature failure stays red). - fail-closed remote spelling at 2 forward sites (ctx.remote !== false per the CLAUDE.md invariant); registry generator drops template-literal flag prefixes; real-PG e2es for the v121 + minion_jobs wedge classes; stale comments corrected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate flag registry after master merge (#3864 added extract help flags) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): red-team batch — dispatch-order validation, real --dry-run boolean, rescue-lane timeout/isolation/cap, parseable-payload repair gate Red-team pass (post-specialist) findings, all verified before fixing: - CRITICAL: validateCommandFlags checked the op lane before CLI_ONLY while dispatch runs CLI_ONLY first — dual-lane commands (think/salience/ anomalies) were validated against the WRONG contract, rejecting documented invocations ('salience --kind entity'). Lane order now mirrors dispatch. - CRITICAL: --dry-run was blessed as legal on op commands but parseOpArgs never SET it (trailing → nothing → ctx.dryRun false → the REAL destructive action ran; leading → consumed the next token). Now a CLI-local boolean exactly like --json, with --dry-run=false support and regression tests. - CRITICAL: the rescue lane ran bun test WITHOUT --timeout=60000 (bun default 5s) — PGLite phantoms re-failed on timeout and were mislabeled 'confirmed real'. Both rescue invocations now mirror the shard flags; serial files re-run one process per file (run-serial-tests.sh isolation contract); rescue wallclock capped at 2x the shard timeout. - CRITICAL: the jsonPayloadOnly probe matched container-LOOKING invalid JSON ('[INFO] fetch complete') whose repair cast would throw and abort the run mid-loop. Predicate now gates on pg_input_is_valid (PG16+, same floor as the existing IS JSON usage) + per-target catch records and continues; doctor mirrors; behavioral test covers the lookalike row. - Registry generator bounded at handleCliOnly's closing brace (the LAST case block absorbed ~100 junk flags from the rest of cli.ts, neutering strict validation for it); uppercase flag typos reject loudly in both lanes (handlers are lowercase-sensitive); --json=true spelling gets the structured invalid_flag envelope; get_chunks __all__ narrowing filed as a Wave 3 TODO. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(engines): getChunks trimmed SELECT must carry modality (codex P1-B) The #2544 egress trim replaced SELECT cc.* with an explicit column list but omitted cc.modality — every rowToChunk field except the vector must survive the trim, or the embed round-trip (getChunks -> upsertChunks) rewrites image chunks as text. Both engines; the structural pin now iterates the full rowToChunk field list instead of spot-checking. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(cli): safety flags require consumption evidence in the flag registry (codex P1-A) upgrade.ts prints a help hint naming another command's --dry-run; that literal is depth-0 text for post-upgrade, so the generator allowlisted --dry-run there — recreating the exact #2185 repro this wave kills (post-upgrade --dry-run accepted, ignored, migrations run for real). Safety flags now need a tight-quoted standalone literal (an args read like has('--dry-run')) before the registry grants them; prose bleed embeds the flag inside a longer string and never qualifies. Regenerated registry drops --dry-run from post-upgrade, keeps genuine consumers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v0.42.76.0 fix(upgrade,security,cli): strict flag validation, bootstrap-wedge class kill, federated chunk scope Version bump + CHANGELOG for the fix wave: CLI-wide unknown-flag rejection with a generated per-command registry (#2185, #1575 class), minion_jobs bootstrap probes + migration-aware coverage guard (v121 wedge class), get_chunks federated scope + egress trim (#2555, half of #2544), think read-scope over MCP (#2598), register-client source bindings (#2016, #2143 enabler), repair-jsonb/doctor subagent columns with a parse-validated damage predicate, memory-safe unit-test runner. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v0.42.76.0 CONTRIBUTING.md: test-runner claims match the memory-safe 4-shard default (was 8-shard) and the CLI-only command recipe now includes the build:flag-registry regen step. KEY_FILES.md: repair-jsonb entry updated to the 8-column parse-gated current state; new entries for the strict flag-validation subsystem and src/core/source-id.ts; operations/engine/ serve-http entries updated for get_chunks scope ladder + trimmed SELECT and the register-client HTTP source bindings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply doc-review fixes for v0.42.76.0 Cross-model doc review caught stragglers: docs/TESTING.md still said 8-shard in the file taxonomy, carried a two-generations-stale shard timeout default (600s -> 3000s), and didn't name the new run-unit-parallel regression test or the remaining runner knobs; CONTRIBUTING.md's fast-loop file count predated the tripled suite (92+ -> 1000+); test-count claims unified at 3700+; the CLIENT_FENCED_WRITE_OPS comment in operations.ts still described think as scope write; KEY_FILES names the exported findUnknownFlag. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): rescue lane survives CI — strip ::group:: prefixes, satisfy the bun-test-timeout guard Two CI-only breaks from the master merge: (1) under GITHUB_ACTIONS the shard wraps file sections as ::group::path.test.ts, so the rescue pass extracted literal ::group:: non-paths that matched zero test files — failing_files_in_log now strips the prefix; (2) master's new check-bun-test-timeout guard greps for bare 'bun test' and tripped on run_rescue's comment text (the invocations themselves carry --timeout=60000) — comment reworded. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): shard-mechanics tests disable mem-adaptation — CI's 7GB runner collapsed explicit 2 shards to 1 The runner deliberately adapts even explicit --shards to available memory (GBRAIN_TEST_NO_MEM_ADAPT=1 is the escape hatch); on GitHub's ~7GB runners that collapsed the tests' 2-shard sandbox runs to 1 shard, breaking every 'shard 1/2:' expectation while passing locally. The tests pin shard MECHANICS with tiny synthetic files, so they now set the escape hatch; the one test that checks the mem banner overrides it back on. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
239 lines
11 KiB
TypeScript
239 lines
11 KiB
TypeScript
/**
|
|
* v0.28: integration test that proves the per-token takes-holder allow-list
|
|
* filters server-side through the dispatch layer (Codex P0 #3 fix
|
|
* verification). PGLite-only; no DATABASE_URL required.
|
|
*
|
|
* Threads:
|
|
* 1. Auth wires `permissions.takes_holders` from `access_tokens` → AuthResult
|
|
* 2. HTTP transport passes `auth.takesHoldersAllowList` to dispatchToolCall
|
|
* 3. dispatch.ts threads it into OperationContext.takesHoldersAllowList
|
|
* 4. takes_list / takes_search ops pass it to engine.listTakes / .searchTakes
|
|
* 5. engine SQL applies `AND holder = ANY($allowList)`
|
|
*
|
|
* This test exercises step 3-5 directly through dispatchToolCall.
|
|
*/
|
|
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
|
import { withoutAnthropicKey } from './helpers/no-anthropic-key.ts';
|
|
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
|
import { dispatchToolCall } from '../src/mcp/dispatch.ts';
|
|
import { TAKES_FENCE_BEGIN, TAKES_FENCE_END } from '../src/core/takes-fence.ts';
|
|
import { operationsByName } from '../src/core/operations.ts';
|
|
|
|
let engine: PGLiteEngine;
|
|
let alicePageId: number;
|
|
|
|
beforeAll(async () => {
|
|
engine = new PGLiteEngine();
|
|
await engine.connect({});
|
|
await engine.initSchema();
|
|
const alice = await engine.putPage('people/alice-example', {
|
|
title: 'Alice', type: 'person', compiled_truth: '## Takes\n',
|
|
});
|
|
alicePageId = alice.id;
|
|
// Seed three takes by three holders. Public fact, garry's bet, brain's hunch.
|
|
await engine.addTakesBatch([
|
|
{ page_id: alicePageId, row_num: 1, claim: 'CEO of Acme', kind: 'fact', holder: 'world', weight: 1.0 },
|
|
{ page_id: alicePageId, row_num: 2, claim: 'Strong technical founder', kind: 'take', holder: 'garry', weight: 0.85 },
|
|
{ page_id: alicePageId, row_num: 3, claim: 'Seemed burned out in last OH', kind: 'hunch', holder: 'brain', weight: 0.4 },
|
|
]);
|
|
});
|
|
|
|
afterAll(async () => {
|
|
await engine.disconnect();
|
|
});
|
|
|
|
function parseResult(result: { content: Array<{ text: string }>; isError?: boolean }): unknown {
|
|
expect(result.isError).toBeFalsy();
|
|
return JSON.parse(result.content[0].text);
|
|
}
|
|
|
|
describe('per-token takes-holder allow-list — takes_list', () => {
|
|
test('default (no allow-list, local CLI) returns all holders', async () => {
|
|
const result = await dispatchToolCall(engine, 'takes_list', { page_slug: 'people/alice-example' }, {
|
|
remote: false, // Local CLI: no allow-list applied.
|
|
});
|
|
const takes = parseResult(result) as Array<{ holder: string; claim: string }>;
|
|
const holders = takes.map(t => t.holder).sort();
|
|
expect(holders).toEqual(['brain', 'garry', 'world']);
|
|
});
|
|
|
|
test('allow-list ["world"] (default-deny token) returns ONLY world holders', async () => {
|
|
const result = await dispatchToolCall(engine, 'takes_list', { page_slug: 'people/alice-example' }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world'],
|
|
});
|
|
const takes = parseResult(result) as Array<{ holder: string; claim: string }>;
|
|
expect(takes).toHaveLength(1);
|
|
expect(takes[0].holder).toBe('world');
|
|
expect(takes[0].claim).toBe('CEO of Acme');
|
|
});
|
|
|
|
test('allow-list ["world", "garry"] returns world + garry, hides brain hunches', async () => {
|
|
const result = await dispatchToolCall(engine, 'takes_list', { page_slug: 'people/alice-example' }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world', 'garry'],
|
|
});
|
|
const takes = parseResult(result) as Array<{ holder: string }>;
|
|
const holders = takes.map(t => t.holder).sort();
|
|
expect(holders).toEqual(['garry', 'world']);
|
|
});
|
|
|
|
test('allow-list with no overlap returns empty (no fallback to default)', async () => {
|
|
const result = await dispatchToolCall(engine, 'takes_list', { page_slug: 'people/alice-example' }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['nonexistent-holder'],
|
|
});
|
|
const takes = parseResult(result) as unknown[];
|
|
expect(takes).toHaveLength(0);
|
|
});
|
|
});
|
|
|
|
describe('per-token takes-holder allow-list — takes_search', () => {
|
|
test('allow-list ["world"] filters search hits to public claims only', async () => {
|
|
const result = await dispatchToolCall(engine, 'takes_search', { query: 'founder' }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world'],
|
|
});
|
|
const hits = parseResult(result) as Array<{ holder: string; claim: string }>;
|
|
expect(hits.every(h => h.holder === 'world')).toBe(true);
|
|
});
|
|
|
|
test('no allow-list (local) sees all holders in search', async () => {
|
|
const result = await dispatchToolCall(engine, 'takes_search', { query: 'founder' }, {
|
|
remote: false,
|
|
});
|
|
const hits = parseResult(result) as Array<{ holder: string }>;
|
|
// 'Strong technical founder' (garry) should match
|
|
expect(hits.some(h => h.holder === 'garry')).toBe(true);
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Page-body channel: get_page / get_versions must respect the same allow-list.
|
|
// Take rows are stored in TWO places per the extract-takes contract: the
|
|
// `takes` table (filtered by the SQL `holder = ANY($allowList)` clause) and
|
|
// inline in `pages.compiled_truth` between TAKES_FENCE markers as a markdown
|
|
// table. Without a strip on the page-CRUD path, a `world`-only token reading
|
|
// `get_page <slug>` recovers every non-`world` claim verbatim from the body.
|
|
// ---------------------------------------------------------------------------
|
|
|
|
describe('per-token takes-holder allow-list — get_page body channel', () => {
|
|
const SLUG = 'people/bob-example';
|
|
const FENCE_BODY =
|
|
'## Takes\n\n' +
|
|
`${TAKES_FENCE_BEGIN}\n` +
|
|
'\n| # | claim | kind | who | weight | since | source |\n' +
|
|
'|---|---|---|---|---|---|---|\n' +
|
|
'| 1 | CEO of Widget | fact | world | 1.0 | 2017-01 | Crustdata |\n' +
|
|
'| 2 | Strong technical founder | take | garry | 0.85 | 2026-04-29 | OH |\n' +
|
|
'| 3 | Seemed burned out in last OH | hunch | brain | 0.4 | 2026-05-01 | private |\n\n' +
|
|
`${TAKES_FENCE_END}\n` +
|
|
'\nFooter content stays.\n';
|
|
|
|
beforeAll(async () => {
|
|
await engine.putPage(SLUG, { title: 'Bob', type: 'person', compiled_truth: FENCE_BODY });
|
|
});
|
|
|
|
test('remote token with allow-list strips fence from compiled_truth', async () => {
|
|
const result = await dispatchToolCall(engine, 'get_page', { slug: SLUG }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world'],
|
|
});
|
|
const page = parseResult(result) as { compiled_truth: string };
|
|
expect(page.compiled_truth).not.toContain(TAKES_FENCE_BEGIN);
|
|
expect(page.compiled_truth).not.toContain(TAKES_FENCE_END);
|
|
expect(page.compiled_truth).not.toContain('Strong technical founder');
|
|
expect(page.compiled_truth).not.toContain('Seemed burned out');
|
|
expect(page.compiled_truth).not.toContain('| garry |');
|
|
expect(page.compiled_truth).not.toContain('| brain |');
|
|
// Surrounding body kept intact.
|
|
expect(page.compiled_truth).toContain('Footer content stays.');
|
|
});
|
|
|
|
test('local CLI (no allow-list) preserves the fence — backwards compatibility', async () => {
|
|
const result = await dispatchToolCall(engine, 'get_page', { slug: SLUG }, {
|
|
remote: false,
|
|
});
|
|
const page = parseResult(result) as { compiled_truth: string };
|
|
expect(page.compiled_truth).toContain(TAKES_FENCE_BEGIN);
|
|
expect(page.compiled_truth).toContain('Seemed burned out');
|
|
});
|
|
|
|
test('fuzzy resolution path also strips for remote token', async () => {
|
|
const result = await dispatchToolCall(engine, 'get_page', { slug: 'people/bob-example', fuzzy: true }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world', 'garry'],
|
|
});
|
|
const page = parseResult(result) as { compiled_truth: string };
|
|
// Allow-list does not yet re-render filtered rows; whole fence is stripped.
|
|
// Pinned so future re-rendering work is an additive change, not a silent
|
|
// semantic flip.
|
|
expect(page.compiled_truth).not.toContain(TAKES_FENCE_BEGIN);
|
|
expect(page.compiled_truth).not.toContain('Strong technical founder');
|
|
});
|
|
});
|
|
|
|
describe('per-token takes-holder allow-list — get_versions body channel', () => {
|
|
const SLUG = 'people/carol-example';
|
|
const FENCE_BODY =
|
|
`${TAKES_FENCE_BEGIN}\n| # | claim | kind | who |\n|---|---|---|---|\n| 1 | private hunch | hunch | brain |\n${TAKES_FENCE_END}\n`;
|
|
|
|
beforeAll(async () => {
|
|
await engine.putPage(SLUG, { title: 'Carol', type: 'person', compiled_truth: FENCE_BODY });
|
|
await engine.createVersion(SLUG); // snapshot now has the fence
|
|
});
|
|
|
|
test('remote token with allow-list strips fence from every snapshot', async () => {
|
|
const result = await dispatchToolCall(engine, 'get_versions', { slug: SLUG }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world'],
|
|
});
|
|
const versions = parseResult(result) as Array<{ compiled_truth: string }>;
|
|
expect(versions.length).toBeGreaterThan(0);
|
|
for (const v of versions) {
|
|
expect(v.compiled_truth).not.toContain(TAKES_FENCE_BEGIN);
|
|
expect(v.compiled_truth).not.toContain('private hunch');
|
|
}
|
|
});
|
|
|
|
test('local CLI sees historical takes in snapshots', async () => {
|
|
const result = await dispatchToolCall(engine, 'get_versions', { slug: SLUG }, {
|
|
remote: false,
|
|
});
|
|
const versions = parseResult(result) as Array<{ compiled_truth: string }>;
|
|
expect(versions.some(v => v.compiled_truth.includes('private hunch'))).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('think op — read-only on remote callers (Lane D landed)', () => {
|
|
test('think is read-scoped for MCP while local persistence remains possible', () => {
|
|
expect(operationsByName.think.scope).toBe('read');
|
|
expect(operationsByName.think.mutating).toBe(true);
|
|
});
|
|
|
|
test('remote save/take is forced read-only via remote_persisted_blocked flag', async () => {
|
|
// Hermetic no-key: neutralize BOTH env var AND ~/.gbrain config key, else a
|
|
// configured machine fires a real LLM call and the warning flips to
|
|
// LLM_OUTPUT_NOT_JSON. runThink then returns gather-only + NO_ANTHROPIC_API_KEY.
|
|
const result = await withoutAnthropicKey(() => dispatchToolCall(engine, 'think', { question: 'q', save: true, take: true }, {
|
|
remote: true, sourceId: 'default', takesHoldersAllowList: ['world', 'garry', 'brain'],
|
|
}));
|
|
const env = parseResult(result) as {
|
|
remote_persisted_blocked: boolean;
|
|
saved_slug: string | null;
|
|
warnings: string[];
|
|
};
|
|
// Codex P1 #7: remote save/take is silently disabled.
|
|
expect(env.remote_persisted_blocked).toBe(true);
|
|
expect(env.saved_slug).toBeNull();
|
|
// Without API key, gather succeeds but synthesis is skipped.
|
|
expect(env.warnings).toContain('NO_ANTHROPIC_API_KEY');
|
|
});
|
|
|
|
test('local-CLI think runs full pipeline (gather-only without API key)', async () => {
|
|
const result = await withoutAnthropicKey(() => dispatchToolCall(engine, 'think', { question: 'q', save: true }, {
|
|
remote: false,
|
|
}));
|
|
const env = parseResult(result) as {
|
|
warnings: string[];
|
|
remote_persisted_blocked: boolean;
|
|
};
|
|
expect(env.remote_persisted_blocked).toBe(false);
|
|
// Without API key, returns gather-only + warning. With key, would actually synthesize.
|
|
expect(env.warnings).toContain('NO_ANTHROPIC_API_KEY');
|
|
});
|
|
});
|