Compare commits

..
Author SHA1 Message Date
Garry TanandClaude Fable 5 c275fa3fab fix(cli): pre-set GBRAIN_EVAL_CAPTURE/SCRUB_PII env wins over DB stash
The #1475 stash unconditionally overwrote the env vars from the merged DB
plane, silently clobbering an operator's exported value — inverting the
env-above-config precedence the same code comment advertises. Unlike
GBRAIN_EMBEDDING_MULTIMODAL, these keys have no loadConfig() env mapping,
so file/env-wins in loadConfigWithEngine never covered them. Guard the
stash writes: DB fills the gap only when the env var is unset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:06:59 -07:00
Garry TanandClaude Fable 5 ee095f1ca2 test(progress): assert net-zero live-reporter leak, not process-global zero
The signal-handler test asserted __liveReporterCountForTest() === 0 — a
process-global absolute that any earlier test file in the same bun shard
can poison (a production path that skips finish() on an error branch
leaves one live entry behind). Shard-5 LPT packing started co-locating
such a file before progress.test.ts, failing this test deterministically
on CI across unrelated PRs while master stayed green by packing luck.

Snapshot the count before the 50 lifecycles and assert no NET leak —
the same prior-state tolerance the handler assertion already applies
via installedBefore. The test still catches its own regression class
(any leak from these lifecycles shows up as delta >= 1).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 11:22:40 -07:00
Garry Tan 5f123c1404 Merge remote-tracking branch 'origin/master' into fix/backlog-c136
# Conflicts:
#	src/cli.ts
2026-07-22 11:00:24 -07:00
d9eb027bdd fix(openclaw): declare gbrain plugin manifest entry (takeover of #2551) (#3185)
Add the OpenClaw-required top-level id to openclaw.plugin.json, export a
direct register(api) entrypoint from src/openclaw-context-engine.ts, add a
manifest regression test, and document that skillpack harvest must preserve
OpenClaw-native manifest fields (id, configSchema, contracts).

llms bundles regenerated (bun run build:llms) — no content drift.

Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: Filip <FilipHarald@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 01:32:11 -07:00
62e009d192 fix(skillopt): emit proposed.md in no-mutate mode (#2635) (#3182)
Takeover of #2719 (fork head; rebased onto origin/master).

- writeProposed now writes both best.md (current-best pointer) and
  proposed.md (stable human-review artifact); returns the proposal path.
- Orchestrator reports the real proposed.md path for accepted --no-mutate runs.
- Tutorial updated; llms bundles regenerated (no content drift — tutorial
  is not inlined in the bundle).

Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: Ziyang Guo <121015044+RerankerGuo@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 00:23:10 -07:00
314fefa560 fix(readme): correct broken OpenClaw and Hermes project links (#1961) (#3179)
Point the OpenClaw and Hermes anchors in the "Have your agent install
it" section at their real upstream repos; the previous openclawagents
org URLs 404. Regenerated llms-full.txt to match.

Takeover of #1961 (fork branch) rebased onto current master.

Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: jessems <jessems@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 20:22:27 -07:00
7f841fae7f feat(maintain): safe maintenance automation + shared orphan-exclusion policy (#3015) (#3023)
Ports #3015 by @jdewoski-cmd onto current master:

- src/core/orphan-policy.ts centralizes the orphan-reporting exclusion
  convention so `gbrain orphans`, doctor's orphan_ratio, and both engines'
  getHealth orphan_pages can no longer drift.
- getHealth stale_pages now uses the link-extractor stale watermark
  (countStalePagesForExtraction) so health agrees with what `gbrain extract
  --stale` will actually process.
- New `gbrain maintain` command: dry-run by default, `--safe` applies only
  the conservative runbook actions (DB-backed stale extraction + source-scoped
  dream cycles for doctor cycle_freshness findings), `--json` for structured
  before/action/after reports. Frontmatter mutations, schema-pack upgrades,
  and semantic hub links stay review-only by design.

Changed from the original PR: the shared defaults carried slugs specific to
the contributor's own brain ('josa-secrets/', '*-ga4-property-id.md',
'*-josa-test', literal 'welcome'/'untitled' fixtures). Global defaults now
carry only GBrain-wide conventions; brain-specific exclusions move to a new
per-brain config plane the policy reads through loadOrphanPolicyOverrides:

    gbrain config set orphans.exclude_prefixes "my-private-folder/,archive/"
    gbrain config set orphans.exclude_slugs "some-one-off-page"

Both engines' getHealth and the orphans command thread the overrides;
tests cover the neutral defaults, the override plane, and health parity.
Also registered `maintain` in CLI_ONLY_SELF_HELP so `gbrain maintain --help`
reaches the command's own usage block.

Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: jdewoski-cmd <jdewoski@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 19:52:18 -07:00
1fabbb9849 fix(links): resolve path-qualified wikilinks outside DIR_PATTERN in the DB/put_page path (#2866)
The generic wikilink pass (issue #972) forwarded the raw literal to
resolveBasenameMatches, whose index is keyed by final path segments —
so [[notes/struktura]] (any dir outside DIR_PATTERN) silently produced
zero edges from `extract links --source db` and put_page auto-link,
while the FS extractor resolves the identical content (resolveSlugAll
strips the dirname before its basename lookup).

Query by the literal's final segment, then keep only matches whose slug
ends with the written path — [[notes/struktura]] can resolve to
vault/notes/struktura but never attach to wiki/struktura. Bare literals
are untouched. Flag-gated by link_resolution.global_basename as before.

Co-authored-by: YMYD <53603073+OJ-OnJourney@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Time Attakc <89218912+time-attack@users.noreply.github.com>
2026-07-21 18:45:23 -07:00
64920f83c9 fix(embed): preserve code-chunk metadata across re-embed (#769) (#1232)
Closes #769. Every re-embed pass clobbered code-chunk metadata
(language, symbol_name, symbol_type, start_line, end_line,
parent_symbol_path, doc_comment, symbol_name_qualified) to NULL,
disabling code-def queries across thousands of indexed chunks.

Two complementary fixes:

embed.ts — three re-upsert call sites (embedPage, embedAll
non-stale, embedAllStale autopilot path) build ChunkInputs from
loaded chunks; they were stripping the 8 metadata fields. New
preserveCodeMetadata helper threads those fields through
consistently. Integrated cleanly with v0.34.4.0's cursor-paginated
--stale hardening — the wrap sits inside the worker function
between embedBatchWithBackoff and engine.upsertChunks.

postgres-engine.ts + pglite-engine.ts — upsertChunks ON CONFLICT
clause OVERWROTE metadata columns from EXCLUDED. Asymmetric vs the
embedding/embedded_at columns which already used a chunk_text-gated
CASE pattern (re-chunk → trust EXCLUDED, re-embed → COALESCE
preserve). Applied the same pattern to all 8 metadata columns.

Three regression tests in test/embed.serial.test.ts cover --stale
(autopilot), --all, and --slugs paths. Each loads a chunk with
full metadata, runs runEmbed, and asserts engine.upsertChunks
receives the metadata round-tripped. Coexists with master's D5
embedBatchWithBackoff test block.

Backfill required after deploy: \`gbrain sync --strategy code
--force --source <id>\` per code source to re-populate metadata via
the chunker. Without backfill, existing NULL columns stay NULL —
re-embed alone never produces metadata, only the chunker does.

Originally landed as part of PR #768 (the wave that bundled #767 +
fix; this PR carries the #769 fix alone with no scope overlap.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Time Attakc <89218912+time-attack@users.noreply.github.com>
2026-07-21 18:13:45 -07:00
Benjamin D. SmithandTime Attakc e861b92da7 feat(synopsis): tail-truncate documentText for small-model chat handlers (#1427)
* feat(synopsis): tail-truncate documentText for small-model chat handlers

Small local chat models (Gemma 4 E2B, Qwen3 4B) get dramatically
slower on long contexts even at 131K declared windows. A 73K-char
page synopsis on Gemma 4 E2B takes 60-120s, exceeding the worker's
default 30s `lockDuration` and tripping `lock-lost` errors.

Add `SYNOPSIS_DOC_MAX_CHARS` env-overridable cap (default 32768
chars, ~8K tokens) applied in `buildUserPrompt`. Truncate the TAIL
so the head (title, frontmatter, intro) preserves the document-level
anchor the synopsis needs.

Anthropic Haiku is unaffected at this cap; bump via
`GBRAIN_SYNOPSIS_DOC_MAX_CHARS` for frontier models that want
richer document anchoring.

Belt-and-suspenders companion to commit 0aaff691 (--lock-duration
flag on the worker). Combined: bumping lock TTL gives the handler
more time, AND truncating doc cap makes the handler complete faster.
Either alone helps; both together get the synopsis backfill running
reliably on small local LLMs.

Verified: real 1383-chunk personal brain backfill at
GBRAIN_SYNOPSIS_MODEL=lmstudio:google/gemma-4-e2b +
GBRAIN_SYNOPSIS_DOC_MAX_CHARS=16384 +
`gbrain jobs work --concurrency 4 --lock-duration 300000`
transitions from "lock-lost on every transcript page" to "no
deaths, no stalls, steady throughput."

RECOVERY REBUILD 2026-05-26 of original ac213aa6.

* fix: fold SYNOPSIS_DOC_MAX_CHARS into corpus_generation hash

Codex review of #1427 flagged that changing GBRAIN_SYNOPSIS_DOC_MAX_CHARS
shifts the synopsis prompt + downstream embeddings for long documents
but was NOT folded into the computeCorpusGeneration hash. Pages
re-embedded with a different cap would retain the same
corpus_generation, defeating the v0.40.3.0 D27 P1-5 cache invalidation
contract.

Three changes:

1. Export SYNOPSIS_DOC_MAX_CHARS from src/core/page-summary.ts
2. computeCorpusGeneration accepts optional synopsisDocMaxChars param.
   When set, folded into hash via '|doc_cap=<N>'. Omitted for
   non-synopsis modes (title / none don't consult the cap) so existing
   pre-PR caches stay valid for those.
3. Service-layer call sites (2 in contextual-retrieval-service.ts)
   pass SYNOPSIS_DOC_MAX_CHARS when attemptMode/resolution.mode is
   per_chunk_synopsis, undefined otherwise.
4. import-file.ts inline path passes undefined (per_chunk_synopsis
   refused upstream there).

One-time effect: per_chunk_synopsis pages re-embedded post-PR get a
NEW corpus_generation including the cap. v0.40.3.0 query_cache.page_generations
contract auto-invalidates cached query results on first re-embed.
Future cap changes track correctly.

Addresses codex review P2 on PR #1427.

---------

Co-authored-by: Time Attakc <89218912+time-attack@users.noreply.github.com>
2026-07-21 16:50:43 -07:00
948ccc7b4f fix(cli): wire gbrain bench publish dispatcher + honor DB-plane eval.capture
Two backlog fixes:

1. Takeover of #1476 (closes #1474): 'bench' was missing from CLI_ONLY and
   runBenchPublish was imported nowhere, so the documented
   'gbrain bench publish' hit 'Unknown command'. Adds the 'bench' token to
   master's current CLI_ONLY set (the original PR rewrote the line from a
   stale v0.41.14 snapshot, deleting ~15 newer commands), routes bench
   through a no-DB bypass (bench publish is pure file I/O), and adds
   'bench' to CLI_ONLY_SELF_HELP so --help reaches bench-publish's own
   usage text.

2. Fixes #1475: 'gbrain config set eval.capture true' persisted to the DB
   plane but was never read — the capture gate reads the sync file-plane
   ctx.config. loadConfigWithEngine now sparse-merges eval.capture /
   eval.scrub_pii (file wins per key), connectEngine stashes the merged
   values on GBRAIN_EVAL_CAPTURE / GBRAIN_EVAL_SCRUB_PII (same pattern as
   the multimodal flags), and the gates consult the stash when the file
   plane is silent.

Co-authored-by: Mr-B-1 <Mr-B-1@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 14:30:23 -07:00
54 changed files with 1276 additions and 1621 deletions
+2 -2
View File
@@ -71,8 +71,8 @@ GBrain is designed to be installed and operated by an AI agent. The fastest path
If you don't already have an AI agent platform running, start with one of these. Both are designed to read GBrain's install protocol and execute it:
- **[OpenClaw](https://github.com/openclawagents/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
- **[Hermes](https://github.com/openclawagents/hermes)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
- **[OpenClaw](https://github.com/openclaw/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
- **[Hermes](https://github.com/NousResearch/hermes-agent)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
Then paste this into your agent:
-38
View File
@@ -115,7 +115,6 @@ Full subcommand reference:
```
gbrain sources add <id> --path <p> [--name <n>] [--federated|--no-federated] [--force]
[--include <glob>...] [--exclude <glob>...]
Register a source. id: [a-z0-9](?:[a-z0-9-]{0,30}[a-z0-9])?
--path must be a git repo (or a subdirectory of one) — see
"The git requirement for --path sources" below. --force
@@ -132,43 +131,6 @@ gbrain sources federate <id>
gbrain sources unfederate <id>
```
## Filtering what gets synced (--include / --exclude)
`--include` and `--exclude` on `gbrain sources add` accept repeatable glob
patterns and are honored by every subsequent sync AND lint of the source.
Common Obsidian vault setups need to exclude authoring scaffolding so it
doesn't pollute search:
```bash
# Skip Templates/, Drafts/, and the smart-env sidecar; everything else syncs.
gbrain sources add vault \
--path ~/Documents/vault --federated \
--exclude 'Templates/**' \
--exclude 'Drafts/**' \
--exclude '.smart-env/**'
# Or: only sync the people/ and companies/ subtrees of a CRM vault.
gbrain sources add crm \
--path ~/Documents/crm --no-federated \
--include 'people/**' \
--include 'companies/**'
```
Both persist into `sources.config.include_globs` / `exclude_globs` arrays.
The filter runs `include` first, then `exclude`, so a path inside
`people/**` is still rejected if it also matches `exclude_globs`. Globs use
the same matcher as the rest of gbrain's sync classifier (`matchesAnyGlob`
in `src/core/sync.ts`) and are matched against the source-root-relative
path. Exclusion is conservative: it never deletes previously-imported pages.
`gbrain sync --include <glob> --exclude <glob>` and
`gbrain lint <dir> --include <glob> --exclude <glob>` take the same
repeatable flags for one-off scope changes; for lint the persisted source
globs are auto-applied when the lint target matches a source's `local_path`.
Changing the persisted globs on an existing source triggers a full re-walk
on the next sync (the source's config fingerprint invalidates the
"already up to date" gate).
## The git requirement for --path sources
Every `--path` source must be a git repository (or live inside one — a
+3 -1
View File
@@ -131,7 +131,9 @@ into gbrain so other clients can scaffold it. Default behavior:
`~/.gbrain/harvest-private-patterns.txt` plus built-in defaults
(canonical private fork name, common email regex, Slack channel pattern). Any
match → rollback (delete the harvested files) and exit non-zero.
- `openclaw.plugin.json` updated with the new slug, sorted.
- `openclaw.plugin.json` updated with the new slug, sorted. Harvest must preserve
the top-level OpenClaw-native plugin fields (`id`, `configSchema`, `contracts`)
because OpenClaw validates those before it can install the package.
- `--no-lint` bypasses the linter (after a manual editorial scrub).
Use the `skillpack-harvest` skill (its companion editorial workflow)
@@ -233,13 +233,14 @@ keep it or `git checkout` to throw it away. Nothing is committed for you.
**For a skill that ships with gbrain** (anything under the gbrain repo's own
`skills/`): SkillOpt refuses to overwrite it by default and writes the winner to
`skills/<name>/skillopt/best.md` instead, so an optimization pass can never
silently mutate a skill other people depend on. Two ways to handle that:
`skills/<name>/skillopt/proposed.md` instead (while keeping `best.md` as the
optimizer's current-best pointer), so an optimization pass can never silently
mutate a skill other people depend on. Two ways to handle that:
```bash
# See the proposed improvement without touching SKILL.md (works for ANY skill):
gbrain skillopt meeting-prep --split 1:1:1 --no-mutate
# → writes skills/meeting-prep/skillopt/best.md (the proposed rewrite), prints its path. Copy what you want.
# → writes skills/meeting-prep/skillopt/proposed.md, updates best.md, and prints the proposal path.
# Actually rewrite a bundled skill (explicit opt-in + an independent held-out set):
gbrain skillopt brain-ops --split 1:1:1 --allow-mutate-bundled \
+2 -2
View File
@@ -1565,8 +1565,8 @@ GBrain is designed to be installed and operated by an AI agent. The fastest path
If you don't already have an AI agent platform running, start with one of these. Both are designed to read GBrain's install protocol and execute it:
- **[OpenClaw](https://github.com/openclawagents/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
- **[Hermes](https://github.com/openclawagents/hermes)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
- **[OpenClaw](https://github.com/openclaw/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
- **[Hermes](https://github.com/NousResearch/hermes-agent)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
Then paste this into your agent:
+1
View File
@@ -1,4 +1,5 @@
{
"id": "gbrain-context-engine",
"name": "gbrain",
"version": "0.32.3.0",
"description": "Personal knowledge brain with Postgres + pgvector hybrid search",
+2 -1
View File
@@ -266,4 +266,5 @@ editorial pass.
(e.g. `src/commands/<slug>.ts` if the host SKILL.md declares it
in frontmatter)
- gbrain's `openclaw.plugin.json` — adds the slug to `skills:`
array, sorted alphabetically
array, sorted alphabetically, without removing OpenClaw-native plugin fields
like `id`, `configSchema`, or `contracts`
+3 -1
View File
@@ -57,6 +57,8 @@ This mode guarantees:
- `skills/manifest.json` lists every skill directory
- `skills/RESOLVER.md` references every skill in the manifest
- `openclaw.plugin.json` `skills[]` round-trips with both
- `openclaw.plugin.json` keeps OpenClaw install-required native plugin fields
(`id`, object `configSchema`, and `contracts.contextEngines` when applicable)
- No MECE violations (duplicate triggers across skills)
### Phases
@@ -72,7 +74,7 @@ This mode guarantees:
### Automation
```bash
bun test test/skills-conformance.test.ts test/resolver.test.ts
bun test test/skills-conformance.test.ts test/resolver.test.ts test/openclaw-plugin-manifest.test.ts
```
The CI-gated check is the package.json `test` script.
+50 -1
View File
@@ -54,7 +54,7 @@ export function bigintToStringReplacer(_key: string, value: unknown): unknown {
}
// CLI-only commands that bypass the operation layer
export const CLI_ONLY = new Set(['init', 'reinit-pglite', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'sync', 'extract', 'extract-conversation-facts', 'enrich', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'skillpack', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'sources', 'mounts', 'dream', 'check-resolvable', 'routing-eval', 'skillify', 'smoke-test', 'providers', 'storage', 'repos', 'code-def', 'code-refs', 'reindex', 'reindex-code', 'reindex-frontmatter', 'code-callers', 'code-callees', 'reconcile-links', 'frontmatter', 'auth', 'friction', 'claw-test', 'book-mirror', 'takes', 'think', 'salience', 'anomalies', 'calibration', 'transcripts', 'models', 'remote', 'recall', 'forget', 'edges-backfill', 'cache', 'ze-switch', 'founder', 'brainstorm', 'lsd', 'schema', 'capture', 'onboard', 'conversation-parser', 'status', 'connect', 'skillopt', 'quarantine', 'self-upgrade', 'advisor', 'watch', 'reindex-search-vector']);
export const CLI_ONLY = new Set(['init', 'reinit-pglite', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'bench', 'sync', 'extract', 'extract-conversation-facts', 'enrich', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'skillpack', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'maintain', 'sources', 'mounts', 'dream', 'check-resolvable', 'routing-eval', 'skillify', 'smoke-test', 'providers', 'storage', 'repos', 'code-def', 'code-refs', 'reindex', 'reindex-code', 'reindex-frontmatter', 'code-callers', 'code-callees', 'reconcile-links', 'frontmatter', 'auth', 'friction', 'claw-test', 'book-mirror', 'takes', 'think', 'salience', 'anomalies', 'calibration', 'transcripts', 'models', 'remote', 'recall', 'forget', 'edges-backfill', 'cache', 'ze-switch', 'founder', 'brainstorm', 'lsd', 'schema', 'capture', 'onboard', 'conversation-parser', 'status', 'connect', 'skillopt', 'quarantine', 'self-upgrade', 'advisor', 'watch', 'reindex-search-vector']);
// CLI-only commands whose handlers print their own --help text. These are
// excluded from the generic short-circuit so detailed per-command and
// per-subcommand usage stays reachable.
@@ -78,6 +78,8 @@ const CLI_ONLY_SELF_HELP = new Set([
'capture',
// v0.42 self-upgrade ships its own usage (flags + the agent-skill story).
'self-upgrade',
// maintain (#3015) prints its own usage block (modes + not-auto-applied list).
'maintain',
// v0.43 (#2095): watch ships WATCH_HELP (flags + the stdin-turn protocol).
'watch',
// v0.37 fix wave (Lane D.4 + CDX2-12): sync's --no-embed flag was
@@ -104,6 +106,9 @@ const CLI_ONLY_SELF_HELP = new Set([
// `gbrain connect --help` prints its own usage (flags + examples) from
// runConnect; route around the generic one-line short-circuit.
'connect',
// #1474: bench-publish ships its own detailed HELP (flags, exit codes,
// the export → publish → gate loop). Route around the generic stub.
'bench',
]);
// v114 (#1941): alias -> operation lookup, kept separate from `cliOps` so
@@ -1427,6 +1432,29 @@ async function handleCliOnly(command: string, args: string[]) {
return;
}
// #1474: `gbrain bench publish` is pure file I/O (reads a captured
// eval-candidates NDJSON from `gbrain eval export`, writes a baseline
// NDJSON). No DB access; bypass connectEngine entirely so the documented
// export → publish → gate loop works on machines without a brain.
// The v0.41.1 wave shipped bench-publish.ts + docs/eval-bench.md but this
// dispatcher case was never added, so the command hit 'Unknown command'.
if (command === 'bench') {
if (args[0] === 'publish') {
const { runBenchPublish } = await import('./commands/bench-publish.ts');
await runBenchPublish(args.slice(1));
return;
}
if (args.length === 0 || args[0] === '--help' || args[0] === '-h') {
const { runBenchPublish } = await import('./commands/bench-publish.ts');
await runBenchPublish(['--help']);
return;
}
console.error(`Unknown bench subcommand: ${args[0]}`);
console.error('Usage: gbrain bench publish --from <captured.ndjson> --to <baseline.ndjson> [flags]');
console.error(' See docs/eval-bench.md for the full loop: eval export → bench publish → eval gate');
process.exit(2);
}
// v0.42.x (#2390): `gbrain eval chronicle` is deterministic — brings its own
// in-memory PGLite, no DB/gateway. CI fixture gate runs anywhere.
if (command === 'eval' && args[0] === 'chronicle') {
@@ -1757,6 +1785,11 @@ async function handleCliOnly(command: string, args: string[]) {
await runOrphans(engine, args);
break;
}
case 'maintain': {
const { runMaintain } = await import('./commands/maintain.ts');
await runMaintain(engine, args);
break;
}
// v0.32.7 CJK wave — post-upgrade markdown re-chunk sweep.
// v0.36 Phase 3 wave — `gbrain reindex --multimodal` re-embeds content_chunks
// into the unified Voyage multimodal-3 column.
@@ -2217,6 +2250,22 @@ async function connectEngine(opts?: { probeOnly?: boolean }): Promise<BrainEngin
if (merged.embedding_image_ocr_model !== undefined) {
process.env.GBRAIN_EMBEDDING_IMAGE_OCR_MODEL = merged.embedding_image_ocr_model;
}
// #1475: stash the merged eval.* flags the same way. The capture gate
// (isEvalCaptureEnabled / isEvalScrubEnabled) runs against ctx.config,
// which is built from the sync file-plane loadConfig() in both the CLI
// op path and MCP dispatch — it never sees the DB plane directly. The
// gates consult this stash when the file plane is silent, so
// `gbrain config set eval.capture true` actually turns capture on.
// A pre-set env value wins over the DB plane (env-above-config, the
// incident escape hatch) — unlike GBRAIN_EMBEDDING_MULTIMODAL these
// keys have no loadConfig() env mapping, so without this guard the
// DB stash would silently clobber an operator's export.
if (process.env.GBRAIN_EVAL_CAPTURE === undefined && merged.eval?.capture !== undefined) {
process.env.GBRAIN_EVAL_CAPTURE = String(merged.eval.capture);
}
if (process.env.GBRAIN_EVAL_SCRUB_PII === undefined && merged.eval?.scrub_pii !== undefined) {
process.env.GBRAIN_EVAL_SCRUB_PII = String(merged.eval.scrub_pii);
}
// Always re-configure with merged values when DB merge succeeded. The
// trigger used to be field-name-gated (only when embedding_multimodal_model
// was set); that coupled the gate to the field set and would silently
+34 -4
View File
@@ -581,7 +581,7 @@ async function embedPage(
for (let j = 0; j < toEmbed.length; j++) {
embeddingMap.set(toEmbed[j].chunk_index, embeddings[j]);
}
const updated: ChunkInput[] = chunks.map(c => ({
const updated: ChunkInput[] = chunks.map(c => preserveCodeMetadata(c, {
chunk_index: c.chunk_index,
chunk_text: c.chunk_text,
chunk_source: c.chunk_source,
@@ -605,6 +605,31 @@ async function embedPage(
slog(`${slug}: embedded ${toEmbed.length} chunks`);
}
/**
* Carry code-chunk metadata (language, symbol_name, symbol_type, line range,
* parent scope, doc comment, qualified name) from a loaded Chunk back into a
* ChunkInput destined for upsertChunks.
*
* Issue #769: every re-embed used to strip these fields, and upsertChunks
* overwrites (does not COALESCE) the metadata columns from EXCLUDED, so
* each pass clobbered code-def's primary index to NULL. Pulling the
* preservation into one helper keeps the three re-embed call sites
* (embedPage, embedAll non-stale, embedAllStale) in lock-step.
*/
function preserveCodeMetadata(loaded: any, base: ChunkInput): ChunkInput {
return {
...base,
language: loaded.language ?? undefined,
symbol_name: loaded.symbol_name ?? undefined,
symbol_type: loaded.symbol_type ?? undefined,
start_line: loaded.start_line ?? undefined,
end_line: loaded.end_line ?? undefined,
parent_symbol_path: loaded.parent_symbol_path ?? undefined,
doc_comment: loaded.doc_comment ?? undefined,
symbol_name_qualified: loaded.symbol_name_qualified ?? undefined,
};
}
async function embedAll(
engine: BrainEngine,
staleOnly: boolean,
@@ -717,8 +742,10 @@ async function embedAll(
for (let j = 0; j < toEmbed.length; j++) {
embeddingMap.set(toEmbed[j].chunk_index, embeddings[j]);
}
// Preserve ALL chunks, only update embeddings for stale ones
const updated: ChunkInput[] = chunks.map(c => ({
// Preserve ALL chunks, only update embeddings for stale ones.
// preserveCodeMetadata threads code-chunk metadata (#769) so re-embed
// doesn't clobber language/symbol_name/symbol_type to NULL.
const updated: ChunkInput[] = chunks.map(c => preserveCodeMetadata(c, {
chunk_index: c.chunk_index,
chunk_text: c.chunk_text,
chunk_source: c.chunk_source,
@@ -1012,7 +1039,10 @@ async function embedAllStale(
for (let j = 0; j < stale.length; j++) {
staleIdxToEmbedding.set(stale[j].chunk_index, embeddings[j]);
}
const merged: ChunkInput[] = existing.map(c => ({
// preserveCodeMetadata threads code-chunk metadata (#769) so the
// autopilot --stale path doesn't clobber language/symbol_name/etc
// to NULL on every cycle.
const merged: ChunkInput[] = existing.map(c => preserveCodeMetadata(c, {
chunk_index: c.chunk_index,
chunk_text: c.chunk_text,
chunk_source: c.chunk_source,
+1 -1
View File
@@ -1651,7 +1651,7 @@ async function extractTimelineFromDB(
* make re-extraction idempotent). EVERY processed page is stamped, including
* zero-link pages — they WERE processed.
*/
async function extractStaleFromDB(
export async function extractStaleFromDB(
engine: BrainEngine,
opts: {
dryRun: boolean;
-11
View File
@@ -53,13 +53,6 @@ export async function runImport(
strategy?: SyncStrategy;
sourceId?: string;
managedBookmark?: boolean;
/**
* #2156: allow-list glob patterns only dir-relative paths matching at
* least one pattern are imported. Applied BEFORE `exclude`. Threaded by
* performFullSync from `gbrain sync --include` / the source row's
* persisted `config.include_globs`.
*/
include?: string[];
/**
* #753/#774: glob patterns to exclude from the import (same semantics as
* `isSyncable`'s `exclude` matched against the dir-relative path).
@@ -222,10 +215,6 @@ export async function runImport(
);
const fileTypeLabel = strategy === 'code' ? 'code'
: strategy === 'auto' ? 'syncable' : 'markdown';
// #2156: apply --include allow-list globs first (threaded by performFullSync).
if (opts.include && opts.include.length > 0) {
allFiles = allFiles.filter(abs => matchesAnyGlob(relative(dir, abs), opts.include));
}
// #753/#774: apply --exclude glob patterns (threaded by performFullSync).
if (opts.exclude && opts.exclude.length > 0) {
const beforeExclude = allFiles.length;
+12 -172
View File
@@ -17,7 +17,7 @@
*/
import { readFileSync, writeFileSync, readdirSync, statSync, lstatSync, existsSync } from 'fs';
import { join, relative, resolve } from 'path';
import { join, relative } from 'path';
import { isAborted } from '../core/abort-check.ts';
import { parseMarkdown, type ParseValidationCode } from '../core/markdown.ts';
import {
@@ -26,9 +26,7 @@ import {
DEFAULT_BYTES_WARN,
} from '../core/content-sanity.ts';
import { loadOperatorLiterals } from '../core/content-sanity-literals.ts';
import { loadConfig, loadConfigWithEngine, toEngineConfig, gbrainPath } from '../core/config.ts';
import { matchesAnyGlob } from '../core/sync.ts';
import { parseGlobList } from './sync.ts';
import { loadConfig, loadConfigWithEngine, gbrainPath } from '../core/config.ts';
import type { BrainEngine } from '../core/engine.ts';
export interface LintIssue {
@@ -380,89 +378,21 @@ async function resolveLintContentSanity(
};
}
/** Collect markdown files from a directory.
*
* When `opts.include` or `opts.exclude` are set, each candidate `.md` path's
* POSIX-style relative path (relative to `dir`) is matched against the same
* glob semantics sync uses (`matchesAnyGlob`). `include` allow-lists;
* `exclude` deny-lists. Empty or undefined arrays leave the filter
* unengaged. Symmetric with `isSyncable` in `src/core/sync.ts` so a
* source-config `exclude_globs` honored by `gbrain sync` is also honored
* by `gbrain lint` against the same dir.
*/
function collectPages(
dir: string,
opts: { include?: string[]; exclude?: string[] } = {},
): string[] {
const { include, exclude } = opts;
const haveInclude = !!(include && include.length > 0);
const haveExclude = !!(exclude && exclude.length > 0);
/** Collect markdown files from a directory */
function collectPages(dir: string): string[] {
const pages: string[] = [];
function walk(d: string) {
for (const entry of readdirSync(d)) {
if (entry.startsWith('.') || entry.startsWith('_')) continue;
const full = join(d, entry);
if (lstatSync(full).isDirectory()) walk(full);
else if (entry.endsWith('.md')) {
if (haveInclude || haveExclude) {
// Match against the path RELATIVE to `dir` (the source root),
// normalized to POSIX separators by matchesAnyGlob. A
// source-config glob like `Resources/veriff/**` is anchored at
// the source root; matching against the absolute path would
// require the user to anchor on their `$HOME` or repo prefix,
// which is brittle.
const rel = relative(dir, full);
if (haveInclude && !matchesAnyGlob(rel, include)) continue;
if (haveExclude && matchesAnyGlob(rel, exclude)) continue;
}
pages.push(full);
}
else if (entry.endsWith('.md')) pages.push(full);
}
}
walk(dir);
return pages.sort();
}
/** Look up the source row whose `local_path` resolves to the same absolute
* directory as `target`, and return its persisted `include_globs` /
* `exclude_globs` as parsed string arrays. Returns an empty object when no
* matching source exists, when the row has no globs configured, or when the
* lookup throws (best-effort auto-resolution must never break standalone
* lint on brains without a sources table).
*
* Mirrors how `syncOneSource` lifts the same fields off `src.config` before
* threading them into `SyncOpts.include` / `SyncOpts.exclude`.
*/
async function resolveSourceGlobsForTarget(
engine: BrainEngine,
target: string,
): Promise<{ include?: string[]; exclude?: string[] }> {
try {
const absTarget = resolve(target);
const rows = await engine.executeRaw<{ config: unknown }>(
`SELECT config FROM sources
WHERE archived IS NOT TRUE
AND local_path IS NOT NULL
AND local_path = $1
LIMIT 1`,
[absTarget],
);
if (rows.length === 0) return {};
const cfg = (rows[0].config && typeof rows[0].config === 'object')
? rows[0].config as Record<string, unknown>
: {};
return {
include: parseGlobList(cfg.include_globs),
exclude: parseGlobList(cfg.exclude_globs),
};
} catch {
// Engine not connected, sources table missing on a fresh brain, RLS
// denial in an unusual scope — all best-effort. Lint proceeds without
// filtering rather than fail-closed.
return {};
}
}
export interface LintOpts {
target: string;
fix?: boolean;
@@ -484,22 +414,6 @@ export interface LintOpts {
* yields + checks this every 200 pages.
*/
signal?: AbortSignal;
/**
* Glob filters threaded into the file walker. When set, paths relative to
* `target` are matched against the patterns using the same semantics as
* `gbrain sync` (`matchesAnyGlob` in `src/core/sync.ts`). `include`
* allow-lists; `exclude` deny-lists; both unset == no filter.
*
* When BOTH are unset AND `engine` is provided, `runLintCore` attempts to
* auto-resolve them from the `sources` row whose `local_path` matches
* `target` symmetric with `syncOneSource`, so a user who has run
* `gbrain sources add --exclude 'Resources/veriff/**'` sees the same
* exclusion applied to `gbrain lint <same-dir>` and to the cycle.lint
* phase without restating it on every invocation. Explicit caller-supplied
* arrays always win over the source-row lift.
*/
include?: string[];
exclude?: string[];
}
export interface LintResult {
@@ -526,21 +440,7 @@ export async function runLintCore(opts: LintOpts): Promise<LintResult> {
}
const isSingleFile = statSync(opts.target).isFile();
// Resolve glob filters. Explicit caller-supplied include/exclude win;
// otherwise lift from `sources.config.{include,exclude}_globs` when an
// engine is available and the target matches a known source's local_path.
// Single-file lints skip the resolve entirely — globs are a directory
// walk concern.
let include = opts.include;
let exclude = opts.exclude;
const haveExplicit = (include && include.length > 0) || (exclude && exclude.length > 0);
if (!isSingleFile && !haveExplicit && opts.engine) {
const resolved = await resolveSourceGlobsForTarget(opts.engine, opts.target);
include = resolved.include;
exclude = resolved.exclude;
}
const pages = isSingleFile ? [opts.target] : collectPages(opts.target, { include, exclude });
const pages = isSingleFile ? [opts.target] : collectPages(opts.target);
// Resolve content-sanity config once for this lint run (D1: lift DB
// config when reachable). Caller can pre-pass via opts.contentSanity
@@ -591,27 +491,14 @@ export async function runLintCore(opts: LintOpts): Promise<LintResult> {
}
export async function runLint(args: string[]) {
const target = args.find(a => !a.startsWith('--') && !args[args.indexOf(a) - 1]?.match(/^--(include|exclude)$/));
const target = args.find(a => !a.startsWith('--'));
const doFix = args.includes('--fix');
const dryRun = args.includes('--dry-run');
// Parse repeatable `--include <glob>` and `--exclude <glob>` flags.
// Symmetric with `gbrain sources add --include / --exclude` from PR #2157;
// explicit flags here override the source-config lift performed below for
// dir-mode lints.
const cliInclude: string[] = [];
const cliExclude: string[] = [];
for (let i = 0; i < args.length; i++) {
if (args[i] === '--include' && i + 1 < args.length) cliInclude.push(args[++i]);
else if (args[i] === '--exclude' && i + 1 < args.length) cliExclude.push(args[++i]);
}
if (!target) {
console.error('Usage: gbrain lint <dir|file.md> [--fix] [--dry-run] [--include <glob>]... [--exclude <glob>]...');
console.error(' --fix Auto-fix fixable issues (LLM preambles, code fences)');
console.error(' --dry-run Preview fixes without writing');
console.error(' --include <glob> Repeatable; only lint paths matching at least one pattern');
console.error(' --exclude <glob> Repeatable; skip paths matching any pattern (applied after --include)');
console.error('Usage: gbrain lint <dir|file.md> [--fix] [--dry-run]');
console.error(' --fix Auto-fix fixable issues (LLM preambles, code fences)');
console.error(' --dry-run Preview fixes without writing');
process.exit(1);
}
@@ -623,44 +510,7 @@ export async function runLint(args: string[]) {
// Single file or directory — print human detail as we go, then rely on
// Core for the aggregate numbers at the end.
const isSingleFile = statSync(target).isFile();
// Resolve glob filters for directory lints. Explicit CLI flags win;
// otherwise lift from `sources.config.{include,exclude}_globs` matching
// `target`. Connect a transient engine for the lookup only when (a) no
// explicit flags were passed AND (b) file/env config suggests an engine is
// available — mirrors the connect-disconnect pattern in
// `resolveLintContentSanity` (issue #1678: standalone CLI never shares the
// db.ts singleton, so create + dispose here is safe).
let runInclude: string[] | undefined = cliInclude.length > 0 ? cliInclude : undefined;
let runExclude: string[] | undefined = cliExclude.length > 0 ? cliExclude : undefined;
if (!isSingleFile && runInclude === undefined && runExclude === undefined) {
const base = loadConfig();
if (base?.database_url || base?.database_path) {
try {
const { createEngine } = await import('../core/engine-factory.ts');
const { connectWithRetry } = await import('../core/db.ts');
const engineCfg = toEngineConfig(base);
const engine = await createEngine(engineCfg);
try {
// Use the same connect path the rest of the CLI uses
// (`connectEngine` in cli.ts). `engine.connect({})` with empty
// opts drops the URL — confirmed by direct probe. `noRetry: true`
// keeps the standalone lint snappy (no retry tax when the brain
// happens to be unreachable; auto-resolve degrades to no-filter).
await connectWithRetry(engine, engineCfg, { noRetry: true });
const lifted = await resolveSourceGlobsForTarget(engine, target);
runInclude = lifted.include;
runExclude = lifted.exclude;
} finally {
await engine.disconnect().catch(() => { /* best-effort */ });
}
} catch {
// best-effort; fall through to no-filter
}
}
}
const pages = isSingleFile ? [target] : collectPages(target, { include: runInclude, exclude: runExclude });
const pages = isSingleFile ? [target] : collectPages(target);
// Progress on stderr. Stdout keeps the per-issue human output it always had.
const { createProgress } = await import('../core/progress.ts');
@@ -707,17 +557,7 @@ export async function runLint(args: string[]) {
// produces canonical numbers for the summary line).
// Pass contentSanity through so runLintCore skips its own resolve
// (we already resolved once for the human-detail loop above).
// Pass include/exclude so the aggregate scope matches the human-detail
// walk above — otherwise the summary line reports the unfiltered count
// even though the per-page details were already filtered.
const result = await runLintCore({
target,
fix: doFix,
dryRun,
contentSanity,
include: runInclude,
exclude: runExclude,
});
const result = await runLintCore({ target, fix: doFix, dryRun, contentSanity });
console.log(`\n${result.pages_scanned} pages scanned. ${result.total_issues} issue(s) in ${result.pages_with_issues} page(s).`);
if (doFix) {
console.log(`${dryRun ? '(dry run) ' : ''}${result.total_fixed} auto-fixed.`);
+224
View File
@@ -0,0 +1,224 @@
/**
* gbrain maintain conservative self-healing maintenance.
*
* This command automates the safe parts of the operator runbook:
* - stale link/timeline extraction
* - stale per-source dream cycles when doctor reports cycle_freshness
*
* It deliberately does NOT mutate source files, apply schema-pack upgrades, or
* invent semantic hub links. Those need review or a separate command with an
* auditable proposal surface.
*/
import { existsSync } from 'fs';
import type { BrainEngine } from '../core/engine.ts';
import type { BrainHealth } from '../core/types.ts';
import { buildChecks, computeDoctorReport, type DoctorReport, type Check } from './doctor.ts';
import { extractStaleFromDB } from './extract.ts';
import { runCycle, type CycleReport } from '../core/cycle.ts';
type ActionStatus = 'ok' | 'would_apply' | 'applied' | 'blocked' | 'skipped';
export interface MaintenanceAction {
name: string;
status: ActionStatus;
message: string;
details?: Record<string, unknown>;
}
export interface MaintainOptions {
json: boolean;
safe: boolean;
dryRun: boolean;
help: boolean;
}
export interface MaintainReport {
mode: 'dry-run' | 'safe';
before: {
health: BrainHealth;
doctor: DoctorReport;
};
actions: MaintenanceAction[];
after: {
health: BrainHealth;
doctor: DoctorReport;
};
}
export function parseMaintainArgs(args: string[]): MaintainOptions {
const safe = args.includes('--safe');
return {
json: args.includes('--json'),
safe,
dryRun: args.includes('--dry-run') || !safe,
help: args.includes('--help') || args.includes('-h'),
};
}
export function extractCycleFreshnessSourceIds(checks: Check[]): string[] {
const ids = new Set<string>();
for (const check of checks) {
if (check.name !== 'cycle_freshness' || check.status === 'ok') continue;
const re = /Source '([^']+)' last cycled/g;
for (const match of check.message.matchAll(re)) {
const id = match[1]?.trim();
if (id) ids.add(id);
}
}
return [...ids].sort();
}
async function buildDoctorReport(engine: BrainEngine): Promise<DoctorReport> {
const checks = await buildChecks(engine, ['--json', '--scope=brain']);
return computeDoctorReport(checks);
}
async function runStaleExtraction(
engine: BrainEngine,
beforeHealth: BrainHealth,
dryRun: boolean,
): Promise<MaintenanceAction> {
if (beforeHealth.stale_pages <= 0) {
return { name: 'extract_stale', status: 'ok', message: 'No stale pages.' };
}
if (dryRun) {
return {
name: 'extract_stale',
status: 'would_apply',
message: `Would run DB-backed stale extraction for ${beforeHealth.stale_pages} page(s).`,
details: { stale_pages: beforeHealth.stale_pages },
};
}
const result = await extractStaleFromDB(engine, {
dryRun: false,
jsonMode: false,
includeFrontmatter: false,
catchUp: false,
});
return {
name: 'extract_stale',
status: 'applied',
message: `Processed ${result.pagesProcessed} stale page(s); ${result.staleRemaining} remain.`,
details: {
links_created: result.linksCreated,
timeline_created: result.timelineCreated,
pages_processed: result.pagesProcessed,
stale_remaining: result.staleRemaining,
},
};
}
async function runCycleFreshnessMaintenance(
engine: BrainEngine,
beforeDoctor: DoctorReport,
dryRun: boolean,
): Promise<MaintenanceAction[]> {
const sourceIds = extractCycleFreshnessSourceIds(beforeDoctor.checks);
if (sourceIds.length === 0) {
return [{ name: 'cycle_freshness', status: 'ok', message: 'All sources cycled recently.' }];
}
if (dryRun) {
return sourceIds.map((sourceId) => ({
name: 'cycle_freshness',
status: 'would_apply',
message: `Would run source-scoped dream cycle for ${sourceId}.`,
details: { source_id: sourceId },
}));
}
const sources = await engine.listAllSources();
const actions: MaintenanceAction[] = [];
for (const sourceId of sourceIds) {
const source = sources.find((s) => s.id === sourceId);
const localPath = source?.local_path ?? null;
const brainDir = localPath && existsSync(localPath) ? localPath : null;
const report: CycleReport = await runCycle(engine, {
brainDir,
dryRun: false,
pull: false,
sourceId,
});
actions.push({
name: 'cycle_freshness',
status: report.status === 'failed' ? 'blocked' : 'applied',
message: `Ran source-scoped dream cycle for ${sourceId}: ${report.status}.`,
details: {
source_id: sourceId,
brain_dir: brainDir,
cycle_status: report.status,
phases: report.phases.map((p) => ({ phase: p.phase, status: p.status })),
},
});
}
return actions;
}
export async function runMaintain(engine: BrainEngine, args: string[]): Promise<MaintainReport | void> {
const opts = parseMaintainArgs(args);
if (opts.help) {
console.log(`Usage: gbrain maintain [--safe] [--dry-run] [--json]
Conservative self-healing maintenance.
Modes:
--dry-run Preview safe actions without writes. Default when --safe is absent.
--safe Apply safe actions: stale extraction and source cycle freshness.
--json Emit a structured before/action/after report.
Not auto-applied:
source-file frontmatter fixes, schema-pack upgrades, atom-pack changes,
semantic hub-link guesses, and destructive cleanup.
`);
return;
}
const beforeHealth = await engine.getHealth();
const beforeDoctor = await buildDoctorReport(engine);
const actions: MaintenanceAction[] = [];
actions.push(await runStaleExtraction(engine, beforeHealth, opts.dryRun));
actions.push(...await runCycleFreshnessMaintenance(engine, beforeDoctor, opts.dryRun));
const afterHealth = await engine.getHealth();
const afterDoctor = await buildDoctorReport(engine);
const report: MaintainReport = {
mode: opts.dryRun ? 'dry-run' : 'safe',
before: { health: beforeHealth, doctor: beforeDoctor },
actions,
after: { health: afterHealth, doctor: afterDoctor },
};
if (opts.json) {
console.log(JSON.stringify(report, null, 2));
} else {
printMaintainReport(report);
}
return report;
}
function printMaintainReport(report: MaintainReport): void {
console.log(`GBrain maintain (${report.mode})`);
console.log(
`Before: brain_score=${Math.round(report.before.health.brain_score)}/100 ` +
`stale=${report.before.health.stale_pages} islands=${report.before.health.orphan_pages} ` +
`doctor=${report.before.doctor.status}`,
);
for (const action of report.actions) {
console.log(` ${action.status}: ${action.name}${action.message}`);
}
console.log(
`After: brain_score=${Math.round(report.after.health.brain_score)}/100 ` +
`stale=${report.after.health.stale_pages} islands=${report.after.health.orphan_pages} ` +
`doctor=${report.after.doctor.status}`,
);
if (report.mode === 'dry-run') {
console.log('Run `gbrain maintain --safe` to apply safe actions.');
}
}
+10 -55
View File
@@ -15,6 +15,11 @@
import type { BrainEngine } from '../core/engine.ts';
import { createProgress, startHeartbeat } from '../core/progress.ts';
import { getCliOptions, cliOptsToProgressOptions } from '../core/cli-options.ts';
import {
shouldExcludeFromOrphanReporting,
loadOrphanPolicyOverrides,
type OrphanPolicyOverrides,
} from '../core/orphan-policy.ts';
// --- Types ---
@@ -32,65 +37,14 @@ export interface OrphanResult {
excluded: number;
}
// --- Filter constants ---
/** Slug suffixes that are always auto-generated root files */
const AUTO_SUFFIX_PATTERNS = ['/_index', '/log'];
/** Page slugs that are pseudo-pages by convention */
const PSEUDO_SLUGS = new Set(['_atlas', '_index', '_stats', '_orphans', '_scratch', 'claude']);
/** Slug segment that marks raw sources */
const RAW_SEGMENT = '/raw/';
/** Slug prefixes where no inbound links is expected */
const DENY_PREFIXES = [
'output/',
'dashboards/',
'scripts/',
'templates/',
'openclaw/config/',
];
/** First slug segments where no inbound links is expected */
const FIRST_SEGMENT_EXCLUSIONS = new Set([
'scratch',
'thoughts',
'catalog',
'entities',
'raw',
'atoms',
'skills',
]);
// --- Filter logic ---
/**
* Returns true if a slug should be excluded from orphan reporting by default.
* These are pages where having no inbound links is expected / not a content problem.
*/
export function shouldExclude(slug: string): boolean {
// Pseudo-pages (exact match)
if (PSEUDO_SLUGS.has(slug)) return true;
// Auto-generated suffix patterns
for (const suffix of AUTO_SUFFIX_PATTERNS) {
if (slug.endsWith(suffix)) return true;
}
// Raw source slugs
if (slug.includes(RAW_SEGMENT)) return true;
// Deny-prefix slugs
for (const prefix of DENY_PREFIXES) {
if (slug.startsWith(prefix)) return true;
}
// First-segment exclusions
const firstSegment = slug.split('/')[0];
if (FIRST_SEGMENT_EXCLUSIONS.has(firstSegment)) return true;
return false;
export function shouldExclude(slug: string, overrides?: OrphanPolicyOverrides): boolean {
return shouldExcludeFromOrphanReporting(slug, overrides);
}
/**
@@ -156,6 +110,7 @@ export async function findOrphans(
let allOrphans: { slug: string; title: string; domain: string | null }[];
let total: number;
let excludedAll: number;
const overrides = includePseudo ? undefined : await loadOrphanPolicyOverrides(engine);
try {
allOrphans = await engine.findOrphanPages(
sourceIds ? { sourceIds } : sourceId ? { sourceId } : undefined,
@@ -184,7 +139,7 @@ export async function findOrphans(
total = liveRows.length;
excludedAll = includePseudo
? 0
: liveRows.reduce((n, r) => n + (shouldExclude(r.slug) ? 1 : 0), 0);
: liveRows.reduce((n, r) => n + (shouldExclude(r.slug, overrides) ? 1 : 0), 0);
} finally {
stopHb();
progress.finish();
@@ -192,7 +147,7 @@ export async function findOrphans(
const filtered = includePseudo
? allOrphans
: allOrphans.filter(row => !shouldExclude(row.slug));
: allOrphans.filter(row => !shouldExclude(row.slug, overrides));
const orphans: OrphanPage[] = filtered.map(row => ({
slug: row.slug,
+1 -34
View File
@@ -122,8 +122,7 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
if (!id) {
console.error(
'Usage: gbrain sources add <id> [--path <path> | --url <https-url>] ' +
'[--name <display>] [--federated|--no-federated] [--clone-dir <path>] [--force] ' +
'[--include <glob>...] [--exclude <glob>...]',
'[--name <display>] [--federated|--no-federated] [--clone-dir <path>] [--force]',
);
process.exit(2);
}
@@ -136,12 +135,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
let patFile: string | undefined;
let noHarden = false;
let force = false;
// Repeatable. `--include 'people/**' --include 'companies/**'` accumulates.
// Persisted into sources.config.include_globs / .exclude_globs and read at
// sync time by commands/sync.ts so `Templates/`, `.smart-env/`, `Drafts/`
// and other vault scaffolding can be skipped without renaming directories.
const includeGlobs: string[] = [];
const excludeGlobs: string[] = [];
for (let i = 1; i < args.length; i++) {
const a = args[i];
@@ -154,24 +147,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
if (a === '--pat-file') { patFile = args[++i]; continue; }
if (a === '--no-harden') { noHarden = true; continue; }
if (a === '--force') { force = true; continue; }
if (a === '--include') {
const v = args[++i];
if (!v || v.startsWith('--')) {
console.error('Error: --include requires a glob argument (e.g. --include "people/**")');
process.exit(2);
}
includeGlobs.push(v);
continue;
}
if (a === '--exclude') {
const v = args[++i];
if (!v || v.startsWith('--')) {
console.error('Error: --exclude requires a glob argument (e.g. --exclude "Templates/**")');
process.exit(2);
}
excludeGlobs.push(v);
continue;
}
console.error(`Unknown flag: ${a}`);
process.exit(2);
}
@@ -192,8 +167,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
federated,
cloneDir,
force,
includeGlobs: includeGlobs.length > 0 ? includeGlobs : undefined,
excludeGlobs: excludeGlobs.length > 0 ? excludeGlobs : undefined,
});
// Topology A discovery: if the just-added source carries a brain-resident
@@ -217,12 +190,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
console.log(
` federated: ${fed}${fed ? ' — appears in cross-source default search' : ' — only searched when explicitly named via --source'}`,
);
if (includeGlobs.length > 0) {
console.log(` include globs: ${includeGlobs.join(', ')}`);
}
if (excludeGlobs.length > 0) {
console.log(` exclude globs: ${excludeGlobs.join(', ')}`);
}
// v0.42.44 — auto-harden managed clones for git durability the moment a brain
// repo is added with a PAT. Best-effort: NEVER fail `add` if hardening fails.
+18 -256
View File
@@ -1,7 +1,6 @@
import { existsSync, readFileSync, writeFileSync, statSync, realpathSync } from 'fs';
import { execFileSync } from 'child_process';
import { join, relative } from 'path';
import { createHash } from 'crypto';
import type { BrainEngine } from '../core/engine.ts';
import { DELETE_BATCH_SIZE } from '../core/engine-constants.ts';
import { importFile } from '../core/import-file.ts';
@@ -757,23 +756,12 @@ export interface SyncOpts {
* are rejected before any git op runs.
*/
srcSubpath?: string;
/**
* #2156 glob patterns files must match to be synced (allow-list).
* Populated from the source row's persisted `config.include_globs`
* (set via `gbrain sources add --include <glob>`) or the repeatable
* `--include` CLI flag. Matched against the scope-relative path, same
* anchoring as `exclude`. `exclude` is applied after `include`: a path
* matching an include pattern is still rejected if it also matches an
* exclude pattern. Empty arrays are the same as undefined (no filter).
*/
include?: string[];
/**
* #753/#774 glob patterns for files to exclude from sync (repeatable
* `--exclude` on the CLI; #2156: also populated from the source row's
* persisted `config.exclude_globs`). Matched against the scope-relative
* path in both the full-sync and incremental paths. Excluded files are
* never imported; exclusion does NOT delete previously-imported pages
* (conservative, matching the #1433 metafile posture).
* `--exclude` on the CLI). Matched against the scope-relative path in both
* the full-sync and incremental paths. Excluded files are never imported;
* exclusion does NOT delete previously-imported pages (conservative,
* matching the #1433 metafile posture).
*/
exclude?: string[];
/**
@@ -1165,29 +1153,6 @@ function unique<T>(items: T[]): T[] {
// `src/core/sync-delta.ts` (re-imported below) so the inline cost estimator
// prices detached sources through the same code the executor imports them with.
/**
* Defensive parse for the JSONB-loaded `config.include_globs` / `config.exclude_globs`
* arrays read off the sources row. The column is a free-form JSONB and could
* contain anything coerce to a string-only array, drop empties, and return
* undefined when the result has no useful entries so the caller can decide
* not to engage glob-filtering at all.
*/
export function parseGlobList(value: unknown): string[] | undefined {
if (!Array.isArray(value)) return undefined;
const globs = value.filter((v): v is string => typeof v === 'string' && v.length > 0);
return globs.length > 0 ? globs : undefined;
}
/**
* Union of CLI-supplied glob patterns (one-off, this invocation) and the
* source row's persisted config globs (every sync). Deduped; undefined when
* neither side has entries so `SyncOpts` stays unset and no filter engages.
*/
export function mergeGlobs(cli: string[], persisted: string[] | undefined): string[] | undefined {
const merged = [...new Set([...cli, ...(persisted ?? [])])];
return merged.length > 0 ? merged : undefined;
}
// v0.18.0 Step 5: source-scoped sync state helpers. When opts.sourceId
// is set, read/write the per-source row instead of the global config
// keys. These wrappers centralize the branch so every read/write site
@@ -1350,125 +1315,6 @@ async function writeChunkerVersion(
);
}
/**
* #2157 follow-on: detect when sources.config has shifted in a way that
* affects which paths the walker will include this run. The "Already up
* to date" gate at performSync's git-HEAD equality check honored chunker
* version match but ignored config drift a user who changes
* `sources.config.exclude_globs` mid-life got "Already up to date" on
* the next sync because git HEAD was unchanged, with no observable
* effect until `gbrain sync --full`.
*
* Fingerprint covers exactly the walk-affecting fields that flow from
* `sources.config` into `SyncOpts` at the syncOneSource call site:
* `strategy`, `include_globs`, `exclude_globs`. CLI-supplied --include
* / --exclude overrides do NOT participate they are one-off scope
* changes, not source state, and shouldn't invalidate the row's
* checkpoint. (A user running `gbrain sync --exclude X` on a row whose
* stored config has no X is intentionally narrowing this one pass; on
* the next no-flags sync, the row config governs again.)
*
* Array order is normalized (alphabetical, post-defensive-parse) so
* `["a/**", "b/**"]` and `["b/**", "a/**"]` fingerprint identically.
* `parseGlobList` shares the same defensive coercion as the call site
* that builds SyncOpts, so hand-edited or pre-normalization rows
* fingerprint to the same shape the walker actually sees.
*/
export function computeSourceConfigFingerprint(rawConfig: unknown): string {
const cfg = (rawConfig || {}) as {
strategy?: unknown;
include_globs?: unknown;
exclude_globs?: unknown;
};
const canonical = JSON.stringify({
strategy: typeof cfg.strategy === 'string' ? cfg.strategy : null,
include_globs: (parseGlobList(cfg.include_globs) ?? []).slice().sort(),
exclude_globs: (parseGlobList(cfg.exclude_globs) ?? []).slice().sort(),
});
return createHash('sha256').update(canonical).digest('hex');
}
/**
* Read the per-source fingerprint stamp. NULL on pre-migration rows or
* sources that have never been synced the gate treats NULL as
* "fingerprint unknown" and skips the invalidation check so first-time
* post-upgrade syncs don't spuriously force-full.
*/
export async function readConfigFingerprint(
engine: BrainEngine,
sourceId: string | undefined,
): Promise<string | null> {
if (!sourceId) return null;
const rows = await engine.executeRaw<{ config_fingerprint: string | null }>(
`SELECT config_fingerprint FROM sources WHERE id = $1`,
[sourceId],
);
return rows[0]?.config_fingerprint ?? null;
}
export async function writeConfigFingerprint(
engine: BrainEngine,
sourceId: string | undefined,
fingerprint: string,
): Promise<void> {
if (!sourceId) return;
await engine.executeRaw(
`UPDATE sources SET config_fingerprint = $1 WHERE id = $2`,
[fingerprint, sourceId],
);
}
/**
* Read the raw `sources.config` value for the named source. Returns an
* empty object for missing or never-configured rows. The reader is
* defensive about legacy double-encoded JSONB rows (`{"federated":true}`
* stored as a JSON string scalar, the #2339 class) the engine's
* `r.config` may arrive as either a string or an object, and both are
* normalized to an object before the fingerprint computation walks the
* keys.
*/
async function readSourceConfig(
engine: BrainEngine,
sourceId: string | undefined,
): Promise<unknown> {
if (!sourceId) return {};
const rows = await engine.executeRaw<{ config: unknown }>(
`SELECT config FROM sources WHERE id = $1`,
[sourceId],
);
const raw = rows[0]?.config;
if (raw === null || raw === undefined) return {};
if (typeof raw === 'string') {
try { return JSON.parse(raw); } catch { return {}; }
}
return raw;
}
/**
* Read-hash-stamp wrapper for sync-completion sites outside the gate's
* scope (e.g. `performFullSync`'s `advanceFull` closure, which doesn't
* see `performSync`'s cached `currentConfigFp` because it's a separate
* function). Reads the row's current config and stamps a fresh
* fingerprint.
*
* Race note: if `sources.config` was mutated between the gate's read
* and this stamp, the freshly-read value wins. The walker still used
* the gate-time effective globs (already captured into `opts.include` /
* `opts.exclude` upstream), so the stamp can drift from what was
* actually walked. In practice mid-sync mutations are rare and the
* NEXT sync will re-evaluate against the latest config anyway, so the
* minor staleness is acceptable and avoids threading the gate-time
* fingerprint through every helper signature.
*/
async function stampSourceConfigFingerprint(
engine: BrainEngine,
sourceId: string | undefined,
): Promise<void> {
if (!sourceId) return;
const cfg = await readSourceConfig(engine, sourceId);
await writeConfigFingerprint(engine, sourceId, computeSourceConfigFingerprint(cfg));
}
/**
* v0.40 Federated Sync v2: `gbrain sync trigger --source <id> [--priority high|normal|low]`
*
@@ -2317,25 +2163,7 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
detachedWorkingTreeManifest.deleted.length > 0 ||
detachedWorkingTreeManifest.renamed.length > 0);
// #2157 follow-on: parallel gate for sources.config drift. Without
// this, changing `sources.config.exclude_globs` (or include_globs /
// strategy) on a synced source had no observable effect on the next
// sync because git HEAD was unchanged — the "Already up to date"
// branch below returned without re-walking. Mismatch path mirrors the
// chunker_version gate exactly so both kinds of drift route through
// the same `performFullSync` recovery.
//
// NULL stored fingerprint is "never stamped" (pre-v125 brain OR fresh
// source whose first sync hasn't completed yet). Treated as
// pass-through in the up-to-date check — first post-upgrade sync
// stamps the column quietly so subsequent passes have a baseline.
const storedConfigFp = await readConfigFingerprint(engine, opts.sourceId);
const currentSourceConfig = await readSourceConfig(engine, opts.sourceId);
const currentConfigFp = computeSourceConfigFingerprint(currentSourceConfig);
const configMismatch = storedConfigFp !== null && storedConfigFp !== currentConfigFp;
const configNeverStamped = storedConfigFp === null && opts.sourceId !== undefined;
if (lastCommit === headCommit && !versionMismatch && !versionNeverSet && !hasDetachedWorkingTreeChanges && !configMismatch) {
if (lastCommit === headCommit && !versionMismatch && !versionNeverSet && !hasDetachedWorkingTreeChanges) {
// v0.42.52.0 (PR #22xx): bump last_sync_at as a heartbeat on every successful
// 0-changes sync. D4 invariant ("never advance last_commit on partial") is
// preserved: last_sync_at is a monitoring signal (doctor sync_freshness
@@ -2348,14 +2176,6 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
[opts.sourceId],
);
}
// First post-upgrade sync on a pre-v125 brain lands here with
// configNeverStamped=true; stamp the fingerprint so the gate has a
// baseline for the NEXT pass. A spurious re-walk on the upgrade
// pass would surprise users; quietly establishing the baseline does
// not.
if (configNeverStamped) {
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
}
return {
status: 'up_to_date',
fromCommit: lastCommit,
@@ -2367,21 +2187,13 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
};
}
if ((versionMismatch || versionNeverSet || configMismatch) && lastCommit === headCommit) {
const reasons: string[] = [];
if (versionMismatch || versionNeverSet) {
reasons.push(`chunker_version=${storedVersion ?? 'unset'}${currentVersion}`);
}
if (configMismatch) {
reasons.push(`config_fingerprint=${storedConfigFp?.slice(0, 8)}${currentConfigFp.slice(0, 8)}`);
}
if ((versionMismatch || versionNeverSet) && lastCommit === headCommit) {
slog(
`[sync] full re-walk forced (${reasons.join(', ')}): ` +
`git HEAD unchanged but a walk-affecting setting advanced.`,
`[sync] chunker_version gate: stored=${storedVersion ?? 'unset'}, current=${currentVersion}. ` +
`Forcing full re-chunk pass (git HEAD unchanged but pipeline version advanced).`,
);
const result = await performFullSync(engine, fullSyncRoots, headCommit, opts);
await writeChunkerVersion(engine, opts.sourceId, currentVersion);
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
return result;
}
@@ -2425,16 +2237,8 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
scoped && p.startsWith(syncScopeRelPath + '/') ? p.slice(syncScopeRelPath.length + 1) : p;
const excluded = (p: string): boolean =>
opts.exclude !== undefined && opts.exclude.length > 0 && matchesAnyGlob(scopeRel(p), opts.exclude);
// #2156: include globs are an allow-list, same scope-relative anchoring as
// exclude. Populated from the source row's persisted config.include_globs
// (or CLI --include). Deliberately NOT threaded into syncOpts/isSyncable:
// the unsyncable-cleanup loop below deletes pages for non-metafile
// classifications, and glob filtering must stay conservative (never delete
// previously-imported pages — the documented #1433 posture for --exclude).
const included = (p: string): boolean =>
opts.include === undefined || opts.include.length === 0 || matchesAnyGlob(scopeRel(p), opts.include);
// Filter to syncable files (strategy-aware + scope-aware + glob-aware)
// Filter to syncable files (strategy-aware + scope-aware + exclude-aware)
const syncOpts = opts.strategy ? { strategy: opts.strategy } : undefined;
// #1970 (F-C): a rename whose DESTINATION is unsyncable drops out of BOTH
// `renamed` (only `r.to` is kept below) AND `deleted` (git emits it as `R`,
@@ -2448,13 +2252,13 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
!(inScope(r.to) && isSyncable(r.to, syncOpts)))
.map(r => r.from);
const filtered: SyncManifest = {
added: manifest.added.filter(p => inScope(p) && included(p) && !excluded(p) && isSyncable(p, syncOpts)),
modified: manifest.modified.filter(p => inScope(p) && included(p) && !excluded(p) && isSyncable(p, syncOpts)),
added: manifest.added.filter(p => inScope(p) && !excluded(p) && isSyncable(p, syncOpts)),
modified: manifest.modified.filter(p => inScope(p) && !excluded(p) && isSyncable(p, syncOpts)),
deleted: unique([
...manifest.deleted.filter(p => inScope(p) && isSyncable(p, syncOpts)),
...renamedToUnsyncable,
]),
renamed: manifest.renamed.filter(r => inScope(r.to) && included(r.to) && !excluded(r.to) && isSyncable(r.to, syncOpts)),
renamed: manifest.renamed.filter(r => inScope(r.to) && !excluded(r.to) && isSyncable(r.to, syncOpts)),
};
// NAV-4: warn when --exclude filtered out every candidate change — almost
@@ -2551,7 +2355,6 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
await writeSyncAnchor(engine, opts.sourceId, 'last_commit', pin, commitTimeMs(gitContextRoot, pin));
await engine.setConfig('sync.last_run', new Date().toISOString());
await writeChunkerVersion(engine, opts.sourceId, String(CHUNKER_VERSION));
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
await clearOpCheckpoint(engine, ckpt.paths);
await clearOpCheckpoint(engine, ckpt.target);
return {
@@ -3376,7 +3179,6 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
await engine.setConfig('sync.last_run', new Date().toISOString());
await writeSyncAnchor(engine, opts.sourceId, 'repo_path', anchorPath);
await writeChunkerVersion(engine, opts.sourceId, String(CHUNKER_VERSION));
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
await clearOpCheckpoint(engine, ckpt.paths);
await clearOpCheckpoint(engine, ckpt.target);
};
@@ -3628,9 +3430,6 @@ async function performFullSync(
// files were waiting.
if (opts.dryRun) {
let allFiles = collectSyncableFiles(syncScopeRoot, { strategy: opts.strategy ?? 'markdown' });
if (opts.include && opts.include.length > 0) {
allFiles = allFiles.filter(abs => matchesAnyGlob(relative(syncScopeRoot, abs), opts.include));
}
if (opts.exclude && opts.exclude.length > 0) {
allFiles = allFiles.filter(abs => !matchesAnyGlob(relative(syncScopeRoot, abs), opts.exclude));
}
@@ -3676,7 +3475,6 @@ async function performFullSync(
commit: headCommit,
strategy: opts.strategy,
sourceId: opts.sourceId,
include: opts.include,
exclude: opts.exclude,
slugRoot,
// issue #1939: performFullSync owns the failure ledger + bookmark via the
@@ -3706,7 +3504,6 @@ async function performFullSync(
await engine.setConfig('sync.last_run', new Date().toISOString());
await writeSyncAnchor(engine, opts.sourceId, 'repo_path', anchorPath);
await writeChunkerVersion(engine, opts.sourceId, String(CHUNKER_VERSION));
await stampSourceConfigFingerprint(engine, opts.sourceId);
};
const fullGate = await applySyncFailureGate({
@@ -4170,12 +3967,8 @@ Options:
run at the repo root; imports are scoped to the subdir
and slugs stay root-relative (wiki/page1). Passing the
subdirectory directly as --repo also works.
--include <glob> Only sync files matching at least one glob (repeatable;
matched against the scope-relative path). Merged with
the source's persisted config.include_globs.
--exclude <glob> Exclude files matching the glob from sync (repeatable;
matched against the scope-relative path; applied after
--include). Merged with config.exclude_globs.
matched against the scope-relative path).
--dry-run Show what would be synced without writing.
--skip-failed Acknowledge previously-recorded sync failures so
the bookmark can advance past unparseable files.
@@ -4336,17 +4129,14 @@ See also:
}
const strategyArg = args.find((a, i) => args[i - 1] === '--strategy') as SyncOpts['strategy'] | undefined;
// #753/#774: monorepo subdir-source flags. --exclude is repeatable.
// #2156: --include is the allow-list counterpart, same repeatable shape.
const srcSubpath = args.find((a, i) => args[i - 1] === '--src-subpath') || undefined;
const excludePatterns: string[] = [];
const includePatterns: string[] = [];
for (let i = 0; i < args.length; i++) {
if (args[i] === '--exclude' && i + 1 < args.length) excludePatterns.push(args[i + 1]);
if (args[i] === '--include' && i + 1 < args.length) includePatterns.push(args[i + 1]);
}
if (syncAll && (srcSubpath || excludePatterns.length > 0 || includePatterns.length > 0)) {
if (syncAll && (srcSubpath || excludePatterns.length > 0)) {
console.error(
`--src-subpath/--include/--exclude scope a single sync invocation; they cannot be combined with --all. ` +
`--src-subpath/--exclude scope a single sync invocation; they cannot be combined with --all. ` +
`For --all runs, register the subdirectory as the source's local_path instead ` +
`(gbrain sources add <id> --path <repo>/<subdir>).`,
);
@@ -4547,11 +4337,7 @@ See also:
const onAllSigint = () => { try { allInterrupt.abort(new Error('SIGINT')); } catch { /* */ } };
const runOne = async (src: typeof sources[number]): Promise<SyncResult> => {
const cfg = (src.config || {}) as {
strategy?: 'markdown' | 'code' | 'auto';
include_globs?: unknown;
exclude_globs?: unknown;
};
const cfg = (src.config || {}) as { strategy?: 'markdown' | 'code' | 'auto' };
// D18: parallel path defers embed; auto-enqueue embed-backfill after.
// v0.42.42.0 (#2139): `autoDeferEmbeds` (the inline gate tripped in a
// non-TTY session) ALSO forces deferral — global by design (the gate's
@@ -4589,8 +4375,6 @@ See also:
skipFailed, retryFailed, noSchemaPack,
sourceId: src.id,
strategy: cfg.strategy,
include: parseGlobList(cfg.include_globs),
exclude: parseGlobList(cfg.exclude_globs),
concurrency,
signal: composeAbortSignals(allInterrupt.signal, controller?.signal),
};
@@ -4802,27 +4586,11 @@ See also:
// lock released by its own finally) instead of a hard cut.
const singleSourceInterrupt = new AbortController();
const onSingleSourceSigint = () => { try { singleSourceInterrupt.abort(new Error('SIGINT')); } catch { /* */ } };
// Read persisted include/exclude globs from the source row, mirroring the
// --all fan-out's `runOne` closure above. Best-effort: a fetch failure
// falls through to "no glob filters", preserving pre-existing behavior.
// sourceId is always set here (resolveSourceWithTier ran above), so this
// path never silently runs without source-config awareness.
let sourceCfg: { include_globs?: unknown; exclude_globs?: unknown } = {};
try {
const { fetchSource } = await import('../core/sources-load.ts');
const src = await fetchSource(engine, sourceId);
if (src?.config && typeof src.config === 'object') {
sourceCfg = src.config as { include_globs?: unknown; exclude_globs?: unknown };
}
} catch { /* fall through to no filters */ }
const opts: SyncOpts = {
repoPath, dryRun, full, noPull, noEmbed, noExtract, skipFailed, retryFailed, noSchemaPack, sourceId,
strategy: strategyArg, concurrency,
srcSubpath,
// #2156: union of the repeatable CLI flags (one-off, this invocation
// only) and the source row's persisted config globs (every sync).
include: mergeGlobs(includePatterns, parseGlobList(sourceCfg.include_globs)),
exclude: mergeGlobs(excludePatterns, parseGlobList(sourceCfg.exclude_globs)),
exclude: excludePatterns.length > 0 ? excludePatterns : undefined,
signal: composeAbortSignals(singleSourceInterrupt.signal, singleSourceController?.signal),
};
@@ -5050,11 +4818,7 @@ export async function syncOneSource(
noExtract?: boolean;
},
): Promise<{ result: SyncResult; log: string }> {
const cfg = (src.config || {}) as {
strategy?: 'markdown' | 'code' | 'auto';
include_globs?: unknown;
exclude_globs?: unknown;
};
const cfg = (src.config || {}) as { strategy?: 'markdown' | 'code' | 'auto' };
const log = `\n--- Syncing source: ${src.name} ---\n`;
const repoOpts: SyncOpts = {
repoPath: src.local_path!,
@@ -5068,8 +4832,6 @@ export async function syncOneSource(
noSchemaPack: shared.noSchemaPack,
sourceId: src.id,
strategy: cfg.strategy,
include: parseGlobList(cfg.include_globs),
exclude: parseGlobList(cfg.exclude_globs),
concurrency: shared.concurrency,
// lockId defaults to `gbrain-sync:${src.id}` via the invariant in
// performSync (no explicit override needed — sourceId triggers it).
+19
View File
@@ -816,6 +816,25 @@ export async function loadConfigWithEngine(
merged.dream = mergedDream;
}
// #1475: eval.* DB-plane merge. `gbrain config set eval.capture true`
// writes the DB plane (both keys are in KNOWN_CONFIG_KEYS, so `set`
// accepts them silently), but the capture gate (isEvalCaptureEnabled)
// reads the merged config. Without this merge the DB value was written
// and never read — capture only fired via GBRAIN_CONTRIBUTOR_MODE=1.
// Sparse per-key merge: file/env wins per key, DB fills the gaps.
const dbEvalCapture = await dbBool('eval.capture');
const dbEvalScrub = await dbBool('eval.scrub_pii');
const mergedEval: NonNullable<GBrainConfig['eval']> = { ...(merged.eval ?? {}) };
if (mergedEval.capture === undefined && dbEvalCapture !== undefined) {
mergedEval.capture = dbEvalCapture;
}
if (mergedEval.scrub_pii === undefined && dbEvalScrub !== undefined) {
mergedEval.scrub_pii = dbEvalScrub;
}
if (Object.keys(mergedEval).length > 0) {
merged.eval = mergedEval;
}
return merged;
}
+22 -5
View File
@@ -54,6 +54,7 @@ import {
import {
generatePerChunkSynopsis,
SYNOPSIS_PROMPT_VERSION,
SYNOPSIS_DOC_MAX_CHARS,
type GeneratePerChunkSynopsisResult,
} from './page-summary.ts';
import {
@@ -103,8 +104,17 @@ function getEmbeddingModelTag(): string {
export function computeCorpusGeneration(args: {
crMode: CRMode;
haikuModel: string;
/**
* Resolved `SYNOPSIS_DOC_MAX_CHARS` for per_chunk_synopsis runs. When
* present, folded into the hash so changes to
* `GBRAIN_SYNOPSIS_DOC_MAX_CHARS` invalidate the prior cache cleanly.
* Omit for `crMode !== 'per_chunk_synopsis'` title / none modes
* don't consult the cap and the field stays out of the hash for
* back-compat with pre-cap embeddings.
*/
synopsisDocMaxChars?: number;
}): string {
return createHash('sha256')
const h = createHash('sha256')
.update(args.crMode)
.update('|')
.update(String(SYNOPSIS_PROMPT_VERSION))
@@ -113,9 +123,11 @@ export function computeCorpusGeneration(args: {
.update('|')
.update(String(TITLE_WRAPPER_VERSION))
.update('|')
.update(getEmbeddingModelTag())
.digest('hex')
.slice(0, 16);
.update(getEmbeddingModelTag());
if (args.synopsisDocMaxChars !== undefined) {
h.update('|doc_cap=').update(String(args.synopsisDocMaxChars));
}
return h.digest('hex').slice(0, 16);
}
/**
@@ -253,7 +265,11 @@ export async function reembedPageWithContextualRetrieval(
args.pageSlug,
args.sourceId,
resolution.mode,
computeCorpusGeneration({ crMode: resolution.mode, haikuModel: args.haikuModel ?? DEFAULT_HAIKU_MODEL }),
computeCorpusGeneration({
crMode: resolution.mode,
haikuModel: args.haikuModel ?? DEFAULT_HAIKU_MODEL,
synopsisDocMaxChars: resolution.mode === 'per_chunk_synopsis' ? SYNOPSIS_DOC_MAX_CHARS : undefined,
}),
);
return { kind: 'skipped', reason: 'no_chunks' };
}
@@ -282,6 +298,7 @@ export async function reembedPageWithContextualRetrieval(
const corpus_generation = computeCorpusGeneration({
crMode: attemptMode,
haikuModel,
synopsisDocMaxChars: attemptMode === 'per_chunk_synopsis' ? SYNOPSIS_DOC_MAX_CHARS : undefined,
});
// ── PHASE 2: single DB transaction ───────────────────────────
+12 -1
View File
@@ -251,6 +251,14 @@ registerBackgroundWorkDrainer({
export function isEvalCaptureEnabled(config: GBrainConfig | null | undefined): boolean {
if (config?.eval?.capture === true) return true;
if (config?.eval?.capture === false) return false;
// #1475: DB-plane stash. `gbrain config set eval.capture true` lands in the
// config table; connectEngine stamps the merged value here because
// ctx.config is the sync file-plane load and never sees the DB plane.
// Explicit per-key setting (file above, DB here) beats the broad
// CONTRIBUTOR_MODE flag, matching how file-plane `false` already wins.
// Doubles as a direct operator env knob.
if (process.env.GBRAIN_EVAL_CAPTURE === 'true') return true;
if (process.env.GBRAIN_EVAL_CAPTURE === 'false') return false;
return process.env.GBRAIN_CONTRIBUTOR_MODE === '1';
}
@@ -263,5 +271,8 @@ export function isEvalCaptureEnabled(config: GBrainConfig | null | undefined): b
* have explicit `capture: true`.
*/
export function isEvalScrubEnabled(config: GBrainConfig | null | undefined): boolean {
return config?.eval?.scrub_pii !== false;
if (config?.eval?.scrub_pii === false) return false;
if (config?.eval?.scrub_pii === true) return true;
// #1475: DB-plane stash — see isEvalCaptureEnabled. Default stays true.
return process.env.GBRAIN_EVAL_SCRUB_PII !== 'false';
}
+5
View File
@@ -733,6 +733,11 @@ export async function importFromContent(
: computeCorpusGeneration({
crMode: effectiveCRMode,
haikuModel: 'anthropic:claude-haiku-4-5-20251001',
// Inline import-file path never uses per_chunk_synopsis (refuses
// upstream); pass undefined so the doc-cap field stays out of
// the hash here. Per_chunk_synopsis runs through the Minion
// backfill handler which threads SYNOPSIS_DOC_MAX_CHARS through
// the service layer.
});
// Transaction wraps all DB writes. Every per-page tx call carries the
+16 -1
View File
@@ -489,7 +489,22 @@ export async function extractPageLinks(
// text inside `[[...]]` before any `|`), NOT the display alias
// (ref.name = match[2]). `[[struktura|the project]]` must resolve
// `struktura`, not "the project". The display text is for context only.
const matches = await resolver.resolveBasenameMatches(ref.slug);
//
// The literal may be path-qualified (`[[notes/struktura]]`). The FS
// path (resolveSlugAll) strips the dirname before its basename lookup,
// but this path passed the raw literal to an index keyed by final
// segments only — so every slash-containing wikilink outside
// DIR_PATTERN silently resolved to nothing. Query by the final
// segment, then use the written path as a disambiguation filter
// (the analogue of the FS ancestor walk honoring the written path):
// a match must end with the literal, so `[[notes/struktura]]` can
// resolve to `vault/notes/struktura` but never to `wiki/struktura`.
const slashIdx = ref.slug.lastIndexOf('/');
const basename = slashIdx === -1 ? ref.slug : ref.slug.slice(slashIdx + 1);
let matches = await resolver.resolveBasenameMatches(basename);
if (slashIdx !== -1) {
matches = matches.filter(m => m === ref.slug || m.endsWith(`/${ref.slug}`));
}
if (matches.length === 0) continue;
const idx = content.indexOf(ref.slug);
const context = idx >= 0 ? excerpt(content, idx, 240) : ref.name;
-26
View File
@@ -5671,32 +5671,6 @@ export const MIGRATIONS: Migration[] = [
`);
},
},
{
version: 125,
name: 'sources_config_fingerprint',
// #2157 follow-on: the "Already up to date" gate at sync.ts honors
// git-HEAD equality + chunker-version match but ignored source-config
// drift. A user who runs `gbrain sources add default --exclude
// 'Templates/**'` AFTER an initial sync got "Already up to date" on
// the next pass because git HEAD was unchanged — the new exclusion
// never reached the walk until `gbrain sync --full`.
//
// This column caches a SHA-256 fingerprint of the walk-affecting
// fields in `sources.config` (strategy + include_globs +
// exclude_globs); mismatches trigger a full re-walk via the same code
// path as a chunker_version bump.
//
// NULL on pre-migration rows is treated as "not yet stamped" by
// readConfigFingerprint, so the FIRST sync after upgrade is normal
// (no spurious force-full just because the column was added).
//
// Keep in sync with src/schema.sql and src/core/schema-embedded.ts.
idempotent: true,
sql: `
ALTER TABLE sources
ADD COLUMN IF NOT EXISTS config_fingerprint TEXT;
`,
},
];
export const LATEST_VERSION = MIGRATIONS.length > 0
+116
View File
@@ -0,0 +1,116 @@
/**
* Shared orphan-reporting exclusion policy.
*
* These are pages where "no inbound links" is expected and should not count
* against health. Keep this in core so the CLI orphan report and engine health
* dashboard cannot drift.
*
* Defaults are GBrain-wide conventions only. Brain-specific exclusions
* (private folder names, one-off fixture slugs) belong in the brain's own
* config, not here:
*
* gbrain config set orphans.exclude_prefixes "my-private-folder/,archive/"
* gbrain config set orphans.exclude_slugs "some-one-off-page"
*/
const AUTO_SUFFIX_PATTERNS = ['/_index', '/log'];
const PSEUDO_SLUGS = new Set(['_atlas', '_index', '_stats', '_orphans', '_scratch', 'claude']);
const RAW_SEGMENT = '/raw/';
const DENY_PREFIXES = [
'output/',
'dashboards/',
'scripts/',
'templates/',
'_templates/',
'openclaw/config/',
'extracts/',
];
const FIRST_SEGMENT_EXCLUSIONS = new Set([
'scratch',
'thoughts',
'catalog',
'entities',
'raw',
'atoms',
'skills',
'dreaming',
'daily',
]);
const ROOT_DATE_SLUG = /^\d{4}-\d{2}-\d{2}(?:-.+)?$/;
function isAgentWorkspaceConvention(slug: string): boolean {
if (!slug.startsWith('agents/')) return false;
if (slug.includes('/memory/dreaming/')) return true;
return /^agents\/[^/]+\/(?:agents|identity|soul|tools|user|heartbeat|dreams|dormant)$/.test(slug);
}
/** Per-brain additions to the convention defaults (from config). */
export interface OrphanPolicyOverrides {
excludePrefixes?: string[];
excludeSlugs?: string[];
}
/** Config keys for per-brain orphan exclusions (comma-separated values). */
export const ORPHAN_EXCLUDE_PREFIXES_KEY = 'orphans.exclude_prefixes';
export const ORPHAN_EXCLUDE_SLUGS_KEY = 'orphans.exclude_slugs';
function parseList(value: string | null): string[] {
if (!value) return [];
return value.split(',').map(s => s.trim()).filter(Boolean);
}
/**
* Load per-brain orphan exclusions from the brain config table. Callers with
* an engine in hand (getHealth, `gbrain orphans`) pass the result as the
* second argument to shouldExcludeFromOrphanReporting.
*/
export async function loadOrphanPolicyOverrides(
engine: { getConfig(key: string): Promise<string | null> },
): Promise<OrphanPolicyOverrides> {
const [prefixes, slugs] = await Promise.all([
engine.getConfig(ORPHAN_EXCLUDE_PREFIXES_KEY),
engine.getConfig(ORPHAN_EXCLUDE_SLUGS_KEY),
]);
return { excludePrefixes: parseList(prefixes), excludeSlugs: parseList(slugs) };
}
export function shouldExcludeFromOrphanReporting(
slug: string,
overrides?: OrphanPolicyOverrides,
): boolean {
if (PSEUDO_SLUGS.has(slug)) return true;
for (const suffix of AUTO_SUFFIX_PATTERNS) {
if (slug.endsWith(suffix)) return true;
}
if (slug.includes(RAW_SEGMENT)) return true;
if (slug.includes('/daily/')) return true;
for (const prefix of DENY_PREFIXES) {
if (slug.startsWith(prefix)) return true;
}
const firstSegment = slug.split('/')[0];
if (FIRST_SEGMENT_EXCLUSIONS.has(firstSegment)) return true;
if (ROOT_DATE_SLUG.test(slug)) return true;
if (slug.startsWith('_brain-')) return true;
if (isAgentWorkspaceConvention(slug)) return true;
if (overrides) {
if (overrides.excludeSlugs?.includes(slug)) return true;
for (const prefix of overrides.excludePrefixes ?? []) {
if (slug.startsWith(prefix)) return true;
}
}
return false;
}
+36 -1
View File
@@ -44,6 +44,33 @@ const HAIKU_MAX_TOKENS = 200;
/** Default model when caller doesn't override. Resolves through the gateway. */
const DEFAULT_SYNOPSIS_MODEL = 'anthropic:claude-haiku-4-5-20251001';
/**
* Hard cap on `documentText` length (chars) before send.
*
* 2026-05-25 fix wave: small local chat models (Gemma 4 E2B, Qwen3 4B) get
* dramatically slower on long contexts even with 131K-token windows declared.
* A 73K-char page synopsis on Gemma 4 E2B takes 60-120s, exceeding the
* worker's default 30s `lockDuration` and tripping `lock-lost` errors.
*
* Truncate to a budget that fits a small model's effective throughput while
* preserving enough document context for the synopsis to be useful. Truncates
* the TAIL because the head (title, frontmatter, intro) carries the
* document-level anchor the synopsis needs.
*
* Override per workload via `GBRAIN_SYNOPSIS_DOC_MAX_CHARS`. Default 32768
* (~8K tokens at 4 chars/tok) keeps small-model synopsis under ~30s.
* Anthropic Haiku is unaffected at this cap; bump higher when running
* frontier models if you want richer document anchoring.
*/
export const SYNOPSIS_DOC_MAX_CHARS = (() => {
const env = process.env.GBRAIN_SYNOPSIS_DOC_MAX_CHARS;
if (env && /^\d+$/.test(env)) {
const n = parseInt(env, 10);
if (n >= 512 && n <= 1_048_576) return n;
}
return 32768;
})();
/**
* Synopsis prompt version. Folded into corpus_generation so prompt edits
* invalidate prior embeddings via the v0.40.3.0 query_cache.page_generations
@@ -188,11 +215,19 @@ function buildUserPrompt(
documentText: string,
chunkText: string,
): string {
// Tail-truncate `documentText` to `SYNOPSIS_DOC_MAX_CHARS` so small local
// chat models don't stall on >100KB pages. Head preserved (title block,
// frontmatter, intro paragraphs carry the document-level anchor).
let trimmedDoc = documentText;
if (documentText.length > SYNOPSIS_DOC_MAX_CHARS) {
trimmedDoc = documentText.slice(0, SYNOPSIS_DOC_MAX_CHARS) +
`\n\n[... ${documentText.length - SYNOPSIS_DOC_MAX_CHARS} chars truncated for synopsis budget ...]`;
}
return [
`<page_title>${pageTitle}</page_title>`,
'',
'<full_document>',
documentText,
trimmedDoc,
'</full_document>',
'',
'<chunk>',
+30 -19
View File
@@ -57,6 +57,8 @@ import { finalizeLastSeen } from './chronicle/last-seen.ts';
import { computeAnomaliesFromBuckets } from './cycle/anomaly.ts';
import { resolveBoostMap, resolveHardExcludes } from './search/source-boost.ts';
import { buildSourceFactorCase, buildHardExcludeClause, buildVisibilityClause, buildRecencyComponentSql, buildBestPerPagePoolCte, buildOrFallbackWebsearchQuery } from './search/sql-ranking.ts';
import { shouldExcludeFromOrphanReporting, loadOrphanPolicyOverrides } from './orphan-policy.ts';
import { LINK_EXTRACTOR_VERSION_TS } from './link-extraction.ts';
import {
normalizeEngineColumn,
buildVectorCastFragment,
@@ -2322,6 +2324,10 @@ export class PGLiteEngine implements BrainEngine {
// v0.40.3.0 D24 NULL→non-NULL race fix mirrors postgres-engine.ts. Two writers
// racing on the same chunk previously raced last-write-wins; the fix lets the
// fresher `embedded_at` win in the text-unchanged branch.
//
// Code-chunk metadata columns follow the same chunk_text-gated CASE pattern as `embedding`
// (#769). Re-chunk trusts EXCLUDED outright; pure re-embed COALESCEs so a caller carrying
// only embedding-shaped fields doesn't clobber metadata to NULL.
await this.db.query(
`INSERT INTO content_chunks ${cols} VALUES ${rowParts.join(', ')}
ON CONFLICT (page_id, chunk_index) DO UPDATE SET
@@ -2345,14 +2351,14 @@ export class PGLiteEngine implements BrainEngine {
THEN EXCLUDED.embedded_at
ELSE content_chunks.embedded_at
END,
language = EXCLUDED.language,
symbol_name = EXCLUDED.symbol_name,
symbol_type = EXCLUDED.symbol_type,
start_line = EXCLUDED.start_line,
end_line = EXCLUDED.end_line,
parent_symbol_path = EXCLUDED.parent_symbol_path,
doc_comment = EXCLUDED.doc_comment,
symbol_name_qualified = EXCLUDED.symbol_name_qualified,
language = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.language ELSE COALESCE(EXCLUDED.language, content_chunks.language) END,
symbol_name = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name ELSE COALESCE(EXCLUDED.symbol_name, content_chunks.symbol_name) END,
symbol_type = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_type ELSE COALESCE(EXCLUDED.symbol_type, content_chunks.symbol_type) END,
start_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.start_line ELSE COALESCE(EXCLUDED.start_line, content_chunks.start_line) END,
end_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.end_line ELSE COALESCE(EXCLUDED.end_line, content_chunks.end_line) END,
parent_symbol_path = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.parent_symbol_path ELSE COALESCE(EXCLUDED.parent_symbol_path, content_chunks.parent_symbol_path) END,
doc_comment = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.doc_comment ELSE COALESCE(EXCLUDED.doc_comment, content_chunks.doc_comment) END,
symbol_name_qualified = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name_qualified ELSE COALESCE(EXCLUDED.symbol_name_qualified, content_chunks.symbol_name_qualified) END,
modality = EXCLUDED.modality,
embedding_image = COALESCE(EXCLUDED.embedding_image, content_chunks.embedding_image)`,
params
@@ -5207,15 +5213,10 @@ export class PGLiteEngine implements BrainEngine {
(SELECT count(*) FROM pages) as page_count,
(SELECT count(*) FROM content_chunks WHERE embedded_at IS NOT NULL)::float /
GREATEST((SELECT count(*) FROM content_chunks), 1)::float as embed_coverage,
(SELECT count(*) FROM pages p
WHERE p.updated_at < (SELECT MAX(te.created_at) FROM timeline_entries te WHERE te.page_id = p.id)
) as stale_pages,
-- Bug 11 orphan = islanded (no inbound AND no outbound).
-- See BrainHealth.orphan_pages docstring; docs updated to match this.
(SELECT count(*) FROM pages p
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
) as orphan_pages,
0 as stale_pages,
-- Bug 11 orphan = islanded (no inbound AND no outbound). The raw
-- list is filtered in TS using the shared orphan-reporting policy.
0 as orphan_pages,
(SELECT count(*) FROM links l
WHERE NOT EXISTS (SELECT 1 FROM pages p WHERE p.id = l.to_page_id)
) as dead_links,
@@ -5240,10 +5241,20 @@ export class PGLiteEngine implements BrainEngine {
LIMIT 5
`);
const { rows: islandedRows } = await this.db.query(`
SELECT p.slug
FROM pages p
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
`);
const r = h as Record<string, unknown>;
const pageCount = Number(r.page_count);
const embedCoverage = Number(r.embed_coverage);
const orphanPages = Number(r.orphan_pages);
const stalePages = await this.countStalePagesForExtraction({ versionTs: LINK_EXTRACTOR_VERSION_TS });
const orphanOverrides = await loadOrphanPolicyOverrides(this);
const orphanPages = (islandedRows as { slug: string }[])
.filter(row => !shouldExcludeFromOrphanReporting(row.slug, orphanOverrides)).length;
const deadLinks = Number(r.dead_links);
const linkCount = Number(r.link_count);
const pagesWithTimeline = Number(r.pages_with_timeline);
@@ -5271,7 +5282,7 @@ export class PGLiteEngine implements BrainEngine {
return {
page_count: pageCount,
embed_coverage: embedCoverage,
stale_pages: Number(r.stale_pages),
stale_pages: stalePages,
orphan_pages: orphanPages,
missing_embeddings: Number(r.missing_embeddings),
brain_score: brainScore,
+33 -22
View File
@@ -67,6 +67,8 @@ import { resolveBoostMap, resolveHardExcludes } from './search/source-boost.ts';
import { buildSourceFactorCase, buildHardExcludeClause, buildVisibilityClause, buildRecencyComponentSql, buildBestPerPagePoolCte, buildOrFallbackWebsearchQuery } from './search/sql-ranking.ts';
import { DEFAULT_EMBEDDING_MODEL, DEFAULT_EMBEDDING_DIMENSIONS } from './ai/defaults.ts';
import { DELETE_BATCH_SIZE } from './engine-constants.ts';
import { shouldExcludeFromOrphanReporting, loadOrphanPolicyOverrides } from './orphan-policy.ts';
import { LINK_EXTRACTOR_VERSION_TS } from './link-extraction.ts';
function escapeSqlStringLiteral(value: string): string {
return value.replace(/'/g, "''");
@@ -2473,6 +2475,13 @@ export class PostgresEngine implements BrainEngine {
// - new is fresher (embedded_at > existing.embedded_at) → take new
// - otherwise → keep existing (slower writer with stale embedding loses)
// Mirrored in pglite-engine.ts; pinned by test/e2e/concurrent-embed-race.test.ts.
//
// Code-chunk metadata columns (language / symbol_name / symbol_type / line range /
// parent_symbol_path / doc_comment / symbol_name_qualified) follow the SAME chunk_text-gated
// CASE pattern as `embedding` (#769). Re-chunk (chunk_text changed) trusts EXCLUDED outright;
// pure re-embed (chunk_text unchanged) COALESCEs so a caller that only carries embedding
// doesn't clobber metadata to NULL. Without this, every embed --stale pass nuked code-def's
// primary index for thousands of chunks at once.
await sql.unsafe(
`INSERT INTO content_chunks ${cols} VALUES ${rows.join(', ')}
ON CONFLICT (page_id, chunk_index) DO UPDATE SET
@@ -2496,14 +2505,14 @@ export class PostgresEngine implements BrainEngine {
THEN EXCLUDED.embedded_at
ELSE content_chunks.embedded_at
END,
language = EXCLUDED.language,
symbol_name = EXCLUDED.symbol_name,
symbol_type = EXCLUDED.symbol_type,
start_line = EXCLUDED.start_line,
end_line = EXCLUDED.end_line,
parent_symbol_path = EXCLUDED.parent_symbol_path,
doc_comment = EXCLUDED.doc_comment,
symbol_name_qualified = EXCLUDED.symbol_name_qualified,
language = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.language ELSE COALESCE(EXCLUDED.language, content_chunks.language) END,
symbol_name = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name ELSE COALESCE(EXCLUDED.symbol_name, content_chunks.symbol_name) END,
symbol_type = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_type ELSE COALESCE(EXCLUDED.symbol_type, content_chunks.symbol_type) END,
start_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.start_line ELSE COALESCE(EXCLUDED.start_line, content_chunks.start_line) END,
end_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.end_line ELSE COALESCE(EXCLUDED.end_line, content_chunks.end_line) END,
parent_symbol_path = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.parent_symbol_path ELSE COALESCE(EXCLUDED.parent_symbol_path, content_chunks.parent_symbol_path) END,
doc_comment = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.doc_comment ELSE COALESCE(EXCLUDED.doc_comment, content_chunks.doc_comment) END,
symbol_name_qualified = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name_qualified ELSE COALESCE(EXCLUDED.symbol_name_qualified, content_chunks.symbol_name_qualified) END,
modality = EXCLUDED.modality,
embedding_image = COALESCE(EXCLUDED.embedding_image, content_chunks.embedding_image)`,
params as Parameters<typeof sql.unsafe>[1],
@@ -5313,11 +5322,9 @@ export class PostgresEngine implements BrainEngine {
async getHealth(): Promise<BrainHealth> {
const sql = this.sql;
// Bug 11 doc-drift fix — orphan_pages means "islanded" (no inbound AND
// no outbound links), aligning both engines with the user-facing
// definition. The type comment previously said "no inbound" but the
// SQL required both — docs now match code so users can trust the
// number. A hub page that links out to many but has no back-references
// is working as intended, not an orphan.
// no outbound links). The raw islanded list is filtered through the same
// policy as `gbrain orphans` so convention pages do not count against
// dashboard health.
const [h] = await sql`
WITH entity_pages AS (
SELECT id, slug FROM pages WHERE type IN ('person', 'company')
@@ -5326,13 +5333,8 @@ export class PostgresEngine implements BrainEngine {
(SELECT count(*) FROM pages) as page_count,
(SELECT count(*) FROM content_chunks WHERE embedded_at IS NOT NULL)::float /
GREATEST((SELECT count(*) FROM content_chunks), 1)::float as embed_coverage,
(SELECT count(*) FROM pages p
WHERE p.updated_at < (SELECT MAX(te.created_at) FROM timeline_entries te WHERE te.page_id = p.id)
) as stale_pages,
(SELECT count(*) FROM pages p
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
) as orphan_pages,
0 as stale_pages,
0 as orphan_pages,
(SELECT count(*) FROM links l
WHERE NOT EXISTS (SELECT 1 FROM pages p WHERE p.id = l.to_page_id)
) as dead_links,
@@ -5356,9 +5358,18 @@ export class PostgresEngine implements BrainEngine {
LIMIT 5
`;
const islandedRows = await sql<{ slug: string }[]>`
SELECT p.slug
FROM pages p
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
`;
const pageCount = Number(h.page_count);
const embedCoverage = Number(h.embed_coverage);
const orphanPages = Number(h.orphan_pages);
const stalePages = await this.countStalePagesForExtraction({ versionTs: LINK_EXTRACTOR_VERSION_TS });
const orphanOverrides = await loadOrphanPolicyOverrides(this);
const orphanPages = islandedRows.filter(row => !shouldExcludeFromOrphanReporting(row.slug, orphanOverrides)).length;
const deadLinks = Number(h.dead_links);
const linkCount = Number(h.link_count);
const pagesWithTimeline = Number(h.pages_with_timeline);
@@ -5386,7 +5397,7 @@ export class PostgresEngine implements BrainEngine {
return {
page_count: pageCount,
embed_coverage: embedCoverage,
stale_pages: Number(h.stale_pages),
stale_pages: stalePages,
orphan_pages: orphanPages,
missing_embeddings: Number(h.missing_embeddings),
brain_score: brainScore,
-7
View File
@@ -39,13 +39,6 @@ CREATE TABLE IF NOT EXISTS sources (
-- bypassing the git-HEAD up_to_date early-return so CHUNKER_VERSION bumps
-- actually trigger re-chunking on upgrade.
chunker_version TEXT,
-- #2157 follow-on: SHA-256 fingerprint of the walk-affecting fields in
-- \`config\` (strategy + include_globs + exclude_globs). Mismatch forces a
-- full re-walk via the same code path as chunker_version, so a user who
-- changes \`sources.config.exclude_globs\` mid-life doesn't get "Already up
-- to date" on the next sync. NULL on pre-migration rows is treated as
-- "not yet stamped" and skips the gate (preserves first-run semantics).
config_fingerprint TEXT,
-- v0.26.5: soft-delete + recovery window. \`archive\` flips archived=true and
-- sets archive_expires_at = now() + 72h. The autopilot purge phase
-- hard-deletes rows where archive_expires_at <= now(). Promoted from a
+10 -4
View File
@@ -93,7 +93,13 @@ import { resolveLrSchedule } from './lr-schedule.ts';
import { preflight, formatPreflightReport } from './preflight.ts';
import { isRejected, loadRejectedBuffer, makeRejectedEntry, saveRejectedBuffer } from './rejected-buffer.ts';
import { runReflect, runOneShotRewrite, describeJudges } from './reflect.ts';
import { acceptCandidate, bestPath, revertAllPending, skillPath, writeProposed } from './version-store.ts';
import {
acceptCandidate,
proposedPath as proposedFilePath,
revertAllPending,
skillPath,
writeProposed,
} from './version-store.ts';
import { runValidationGate, scoreSkillOnTasks } from './validate-gate.ts';
import { ROLLOUT_SUCCESS_THRESHOLD } from './types.ts';
import type { SkillOptOpts, EditOp, RunReceipt, BenchmarkTask } from './types.ts';
@@ -702,9 +708,9 @@ async function runOptimizationLoop(
// to the catch's assignment values only (it can't prove the async callback ran).
const finalOutcome = outcome as 'accepted' | 'no_improvement' | 'aborted' | 'errored';
if (!mutateDecision.mutate && finalOutcome === 'accepted') {
// best.md was written by writeProposed() in the accept branch (no-mutate
// path); it doubles as proposed.md for human review. SKILL.md untouched.
proposedPath = bestPath(skillsDir, skillName);
// writeProposed() emitted both the best pointer and the stable review
// artifact in the accept branch. SKILL.md remains untouched.
proposedPath = proposedFilePath(skillsDir, skillName);
} else if (mutateDecision.mutate) {
mutatedSkillFile = finalOutcome === 'accepted';
}
+15 -9
View File
@@ -23,6 +23,7 @@
*
* history.json
* best.md
* proposed.md
* versions/
* v0001_e1_s1.md
* v0002_e1_s2.md
@@ -52,6 +53,10 @@ export function bestPath(skillsDir: string, skillName: string): string {
return path.join(skilloptDir(skillsDir, skillName), 'best.md');
}
export function proposedPath(skillsDir: string, skillName: string): string {
return path.join(skilloptDir(skillsDir, skillName), 'proposed.md');
}
export function skillPath(skillsDir: string, skillName: string): string {
return path.join(skillsDir, skillName, 'SKILL.md');
}
@@ -171,17 +176,18 @@ export function acceptCandidate(input: AcceptInput): AcceptResult {
}
/**
* Write the candidate to `best.md` (which doubles as `proposed.md`) WITHOUT
* touching SKILL.md or the history ledger. Used by the `--no-mutate` /
* bundled-without-allow paths: the optimizer found a better candidate but the
* caller opted out of in-place mutation, so we surface it for human review.
* Returns the path written. Atomic (.tmp + rename).
* Write the candidate to both `best.md` and `proposed.md` WITHOUT touching
* SKILL.md or the history ledger. `best.md` remains the optimizer's current
* best pointer; `proposed.md` is the stable human-review artifact promised by
* `--no-mutate`. Returns the proposal path. Each write is atomic (.tmp + rename).
*/
export function writeProposed(skillsDir: string, skillName: string, candidateText: string): string {
const p = bestPath(skillsDir, skillName);
fs.mkdirSync(path.dirname(p), { recursive: true });
atomicWrite(p, candidateText);
return p;
const best = bestPath(skillsDir, skillName);
const proposed = proposedPath(skillsDir, skillName);
fs.mkdirSync(path.dirname(best), { recursive: true });
atomicWrite(best, candidateText);
atomicWrite(proposed, candidateText);
return proposed;
}
/**
-25
View File
@@ -155,19 +155,6 @@ export interface AddSourceOpts {
* runs). Does NOT auto-`git init` anything see `addSource` docstring.
*/
force?: boolean;
/**
* Glob filters persisted into `sources.config.include_globs` /
* `sources.config.exclude_globs`. Read at sync time by
* `commands/sync.ts:syncOneSource` and the single-source path, threaded
* into `isSyncable` / `unsyncableReason` (their `SyncableOptions` shape
* has carried this contract since v0.41.13).
*
* Empty / unspecified arrays are not persisted at all (no `[]` written
* to the JSONB), which keeps the row identical to today for sources
* that don't use filtering.
*/
includeGlobs?: string[];
excludeGlobs?: string[];
}
export interface RemoveSourceOpts {
@@ -442,12 +429,6 @@ export async function addSource(
if (opts.federated !== null && opts.federated !== undefined) {
config.federated = opts.federated;
}
if (opts.includeGlobs && opts.includeGlobs.length > 0) {
config.include_globs = opts.includeGlobs;
}
if (opts.excludeGlobs && opts.excludeGlobs.length > 0) {
config.exclude_globs = opts.excludeGlobs;
}
const displayName = opts.name ?? opts.id;
try {
@@ -527,12 +508,6 @@ export async function addSource(
if (opts.federated !== null && opts.federated !== undefined) {
config.federated = opts.federated;
}
if (opts.includeGlobs && opts.includeGlobs.length > 0) {
config.include_globs = opts.includeGlobs;
}
if (opts.excludeGlobs && opts.excludeGlobs.length > 0) {
config.exclude_globs = opts.excludeGlobs;
}
const displayName = opts.name ?? opts.id;
await engine.executeRaw(
`INSERT INTO sources (id, name, local_path, config)
-7
View File
@@ -219,13 +219,6 @@ function globToRegex(pattern: string): RegExp {
return new RegExp(regex);
}
/**
* Test a normalized POSIX-style path against an array of glob patterns. Returns
* true if any pattern matches. Empty / undefined `patterns` returns false (no
* filter engaged). Exported so non-sync surfaces (lint walker, future ingest
* variants) can apply the same glob semantics as `isSyncable` without
* re-declaring `globToRegex`.
*/
export function matchesAnyGlob(path: string, patterns?: string[]): boolean {
if (!patterns || patterns.length === 0) return false;
const normalized = path.replace(/\\/g, '/');
+16 -15
View File
@@ -63,25 +63,26 @@ interface PluginCtx {
[key: string]: unknown;
}
export function register(api: PluginApi) {
api.registerContextEngine(ENGINE_ID, (ctx: PluginCtx) => {
const hostResolver =
typeof ctx.resolveEntities === 'function'
? ctx.resolveEntities
: typeof ctx.brainQuery === 'function'
? ctx.brainQuery
: undefined;
return createGBrainContextEngine({
workspaceDir: ctx.workspaceDir,
resolveEntities: hostResolver,
});
});
}
const entry: PluginEntry = {
id: 'gbrain-context-engine',
name: 'GBrain Context Engine',
description: 'Deterministic temporal/spatial context injection on every turn',
register(api: PluginApi) {
api.registerContextEngine(ENGINE_ID, (ctx: PluginCtx) => {
const hostResolver =
typeof ctx.resolveEntities === 'function'
? ctx.resolveEntities
: typeof ctx.brainQuery === 'function'
? ctx.brainQuery
: undefined;
return createGBrainContextEngine({
workspaceDir: ctx.workspaceDir,
resolveEntities: hostResolver,
});
});
},
register,
};
export default entry;
-7
View File
@@ -35,13 +35,6 @@ CREATE TABLE IF NOT EXISTS sources (
-- bypassing the git-HEAD up_to_date early-return so CHUNKER_VERSION bumps
-- actually trigger re-chunking on upgrade.
chunker_version TEXT,
-- #2157 follow-on: SHA-256 fingerprint of the walk-affecting fields in
-- `config` (strategy + include_globs + exclude_globs). Mismatch forces a
-- full re-walk via the same code path as chunker_version, so a user who
-- changes `sources.config.exclude_globs` mid-life doesn't get "Already up
-- to date" on the next sync. NULL on pre-migration rows is treated as
-- "not yet stamped" and skips the gate (preserves first-run semantics).
config_fingerprint TEXT,
-- v0.26.5: soft-delete + recovery window. `archive` flips archived=true and
-- sets archive_expires_at = now() + 72h. The autopilot purge phase
-- hard-deletes rows where archive_expires_at <= now(). Promoted from a
+67
View File
@@ -0,0 +1,67 @@
// #1474: the v0.41.1 wave shipped bench-publish.ts + docs/eval-bench.md
// advertising `gbrain bench publish`, but the cli.ts dispatcher case was never
// added — the documented command hit 'Unknown command'. These tests spawn the
// real CLI (no DB needed; bench publish is pure file I/O) and fail on any
// regression of the dispatcher wiring.
import { describe, expect, test } from 'bun:test';
import { mkdtempSync, writeFileSync, existsSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { spawnSync } from 'node:child_process';
function runCli(args: string[]): { stdout: string; stderr: string; code: number } {
const result = spawnSync(process.execPath, ['run', 'src/cli.ts', 'bench', ...args], {
encoding: 'utf8',
cwd: process.cwd(),
env: { ...process.env },
});
return { stdout: result.stdout ?? '', stderr: result.stderr ?? '', code: result.status ?? -1 };
}
describe('gbrain bench dispatcher (#1474)', () => {
test('bench --help reaches bench-publish help without a DB (was: Unknown command)', () => {
const { stdout, stderr, code } = runCli(['--help']);
expect(stderr).not.toContain('Unknown command');
expect(code).toBe(0);
expect(stdout).toContain('gbrain bench publish');
expect(stdout).toContain('--from');
});
test('unknown bench subcommand exits 2 with usage', () => {
const { stderr, code } = runCli(['bogus']);
expect(code).toBe(2);
expect(stderr).toContain('Unknown bench subcommand');
expect(stderr).toContain('bench publish');
});
test('bench publish roundtrip: captured NDJSON in, baseline file out', () => {
const tmp = mkdtempSync(join(tmpdir(), 'bench-cli-'));
try {
const row = {
tool_name: 'query',
query: 'hello world',
retrieved_slugs: ['slug-a'],
retrieved_chunk_ids: [1],
source_ids: ['default'],
expand_enabled: false,
detail: 'medium',
detail_resolved: 'medium',
vector_enabled: true,
expansion_applied: false,
latency_ms: 100,
remote: false,
job_id: null,
subagent_id: null,
};
const from = join(tmp, 'captured.ndjson');
const to = join(tmp, 'personal.baseline.ndjson');
writeFileSync(from, `${JSON.stringify(row)}\n`);
const { code, stderr } = runCli(['publish', '--from', from, '--to', to]);
expect(stderr).not.toContain('Unknown command');
expect(code).toBe(0);
expect(existsSync(to)).toBe(true);
} finally {
rmSync(tmp, { recursive: true, force: true });
}
});
});
+53
View File
@@ -172,6 +172,59 @@ describe('issue #972 — DB-source (gbrain extract links --source db)', () => {
expect(strk!.link_type).toBe('wikilink_basename');
});
test('flag ON → path-qualified wikilink outside DIR_PATTERN resolves via DB path', async () => {
// `[[notes/struktura]]` — `notes` is not in DIR_PATTERN, so the ref
// reaches the generic pass with its dirname intact. Regression: the DB
// path queried the basename index with the raw literal (which is keyed
// by final segments only), so path-qualified wikilinks outside
// DIR_PATTERN silently produced zero edges while the FS path resolved
// the identical content.
await engine.putPage('notes/struktura', {
type: 'concept' as any, title: 'Struktura Notes',
compiled_truth: '', timeline: '',
});
await engine.putPage('concepts/knowledge-graph', {
type: 'concept', title: 'Knowledge Graph',
compiled_truth: 'Background in [[notes/struktura]].', timeline: '',
});
await engine.setConfig('link_resolution.global_basename', 'true');
await runExtract(engine, ['links', '--source', 'db']);
const outLinks = await engine.getLinks('concepts/knowledge-graph');
const strk = outLinks.find(l => l.to_slug === 'notes/struktura');
expect(strk).toBeDefined();
expect(strk!.link_type).toBe('wikilink_basename');
expect(strk!.link_source).toBe('wikilink-resolved');
});
test('path-qualified wikilink never attaches to a basename-only sibling', async () => {
// Both notes/struktura and wiki/struktura exist. The author wrote
// `[[notes/struktura]]` — the written path must exclude wiki/struktura
// (a bare `[[struktura]]` would legitimately match both).
await engine.putPage('notes/struktura', {
type: 'concept' as any, title: 'Struktura Notes',
compiled_truth: '', timeline: '',
});
await engine.putPage('wiki/struktura', {
type: 'concept' as any, title: 'Struktura Wiki',
compiled_truth: '', timeline: '',
});
await engine.putPage('concepts/x', {
type: 'concept', title: 'X',
compiled_truth: 'See [[notes/struktura]].', timeline: '',
});
await engine.setConfig('link_resolution.global_basename', 'true');
await runExtract(engine, ['links', '--source', 'db']);
const outLinks = await engine.getLinks('concepts/x');
const basenameLinks = outLinks
.filter(l => l.link_type === 'wikilink_basename')
.map(l => l.to_slug);
expect(basenameLinks).toEqual(['notes/struktura']);
});
test('flag OFF → no basename edges via DB path (back-compat)', async () => {
await engine.putPage('projects/struktura', {
type: 'project', title: 'Struktura',
+4 -4
View File
@@ -39,6 +39,7 @@ import { runSkillOpt } from '../../src/core/skillopt/orchestrator.ts';
import {
bestPath,
loadHistory,
proposedPath,
skillPath,
} from '../../src/core/skillopt/version-store.ts';
import { loadRejectedBuffer } from '../../src/core/skillopt/rejected-buffer.ts';
@@ -741,7 +742,7 @@ describe('skillopt T3 — F11 held-out gate, ablation opts, no-DB-pollution', ()
} finally { fixture.cleanup(); }
});
test('--no-mutate writes proposed.md (best.md), leaves SKILL.md untouched', async () => {
test('--no-mutate writes proposed.md and best.md, leaves SKILL.md untouched', async () => {
const fixture = setupFixture(SKILL_PEOPLE_ONLY, CITATIONS_BENCHMARK);
try {
installStub({
@@ -753,10 +754,9 @@ describe('skillopt T3 — F11 held-out gate, ablation opts, no-DB-pollution', ()
const result = await runOnce(fixture, { noMutate: true });
expect(result.outcome).toBe('accepted');
expect(result.mutatedSkillFile).toBe(false);
expect(result.proposedPath).toBeDefined();
// proposed.md (best.md) exists and carries the improvement.
expect(fs.existsSync(result.proposedPath!)).toBe(true);
expect(result.proposedPath).toBe(proposedPath(fixture.skillsDir, SKILL));
expect(fs.readFileSync(result.proposedPath!, 'utf8')).toContain('## Citations');
expect(fs.readFileSync(bestPath(fixture.skillsDir, SKILL), 'utf8')).toContain('## Citations');
// SKILL.md on disk is UNCHANGED (still People-only).
const skill = fs.readFileSync(skillPath(fixture.skillsDir, SKILL), 'utf8');
expect(skill).not.toContain('## Citations');
+104
View File
@@ -803,3 +803,107 @@ describe('embedAllStale --source threading (D7)', () => {
expect((firstCallOpts as { sourceId?: string }).sourceId).toBe('media-corpus');
});
});
// ────────────────────────────────────────────────────────────────
// Code metadata preservation across re-embed (regression for #769)
// ────────────────────────────────────────────────────────────────
//
// gbrain v0.30.1 and earlier silently clobbered code-chunk metadata
// (language, symbol_name, symbol_type, start_line, end_line,
// parent_symbol_path, doc_comment, symbol_name_qualified) on every
// re-embed pass. The chunker populated those columns at import time,
// but embed.ts loaded chunks via getChunks then mapped them to a
// stripped ChunkInput carrying only 5 fields. upsertChunks then
// OVERWROTE (not COALESCEd) the metadata columns from EXCLUDED, so
// re-embed wiped them to NULL. End result on a real brain: 4875 code
// pages, 47866 chunks, all with NULL language/symbol_name/symbol_type;
// code-def returned 0 hits across every indexed repo.
//
// All three runEmbed paths (--stale autopilot, --all, --slugs) must
// thread metadata through the re-upsert. Tests below assert that the
// engine.upsertChunks call carries the same metadata it loaded.
describe('runEmbed preserves code-chunk metadata across re-embed (regression for #769)', () => {
const fullCodeChunk = {
chunk_index: 0,
chunk_text: '[Java] foo/Bar.java:10-20 method baz',
chunk_source: 'compiled_truth' as const,
embedded_at: null,
token_count: 12,
language: 'java',
symbol_name: 'baz',
symbol_type: 'function',
start_line: 10,
end_line: 20,
parent_symbol_path: ['Bar'],
doc_comment: 'does the thing',
symbol_name_qualified: 'Bar.baz',
};
function metadataOf(chunk: any) {
return {
language: chunk.language,
symbol_name: chunk.symbol_name,
symbol_type: chunk.symbol_type,
start_line: chunk.start_line,
end_line: chunk.end_line,
parent_symbol_path: chunk.parent_symbol_path,
doc_comment: chunk.doc_comment,
symbol_name_qualified: chunk.symbol_name_qualified,
};
}
test('--stale (autopilot path) carries code metadata into upsertChunks', async () => {
const stale = [{
slug: 'code-page',
chunk_index: 0,
chunk_text: fullCodeChunk.chunk_text,
chunk_source: 'compiled_truth',
model: null,
token_count: 12,
}];
let upsertChunkArgs: any[] | null = null;
const engine = mockEngine({
countStaleChunks: async () => 1,
listStaleChunks: async () => stale,
getChunks: async () => [fullCodeChunk],
upsertChunks: async (_slug: string, chunks: any[]) => { upsertChunkArgs = chunks; },
});
await runEmbed(engine, ['--stale']);
expect(upsertChunkArgs).not.toBeNull();
expect(upsertChunkArgs!).toHaveLength(1);
expect(metadataOf(upsertChunkArgs![0])).toEqual(metadataOf(fullCodeChunk));
});
test('--all (full re-embed) carries code metadata into upsertChunks', async () => {
let upsertChunkArgs: any[] | null = null;
const engine = mockEngine({
listPages: async () => [{ slug: 'code-page' }],
getChunks: async () => [fullCodeChunk],
upsertChunks: async (_slug: string, chunks: any[]) => { upsertChunkArgs = chunks; },
});
await runEmbed(engine, ['--all']);
expect(upsertChunkArgs).not.toBeNull();
expect(upsertChunkArgs!).toHaveLength(1);
expect(metadataOf(upsertChunkArgs![0])).toEqual(metadataOf(fullCodeChunk));
});
test('--slugs (per-page embed) carries code metadata into upsertChunks', async () => {
let upsertChunkArgs: any[] | null = null;
const engine = mockEngine({
getPage: async () => ({ slug: 'code-page', compiled_truth: 'x', timeline: '' }),
getChunks: async () => [fullCodeChunk],
upsertChunks: async (_slug: string, chunks: any[]) => { upsertChunkArgs = chunks; },
});
await runEmbed(engine, ['--slugs', 'code-page']);
expect(upsertChunkArgs).not.toBeNull();
expect(upsertChunkArgs!).toHaveLength(1);
expect(metadataOf(upsertChunkArgs![0])).toEqual(metadataOf(fullCodeChunk));
});
});
+59
View File
@@ -309,3 +309,62 @@ describe('isEvalCaptureEnabled / isEvalScrubEnabled (CONTRIBUTOR_MODE-gated)', (
} finally { restore(); }
});
});
describe('DB-plane stash (#1475): GBRAIN_EVAL_CAPTURE / GBRAIN_EVAL_SCRUB_PII', () => {
// connectEngine stamps `gbrain config set eval.capture` (DB plane) onto
// these env vars because ctx.config is the sync file-plane load. Without
// the stash check the DB value was written and never read.
const origCapture = process.env.GBRAIN_EVAL_CAPTURE;
const origScrub = process.env.GBRAIN_EVAL_SCRUB_PII;
const origMode = process.env.GBRAIN_CONTRIBUTOR_MODE;
const restore = () => {
if (origCapture === undefined) delete process.env.GBRAIN_EVAL_CAPTURE;
else process.env.GBRAIN_EVAL_CAPTURE = origCapture;
if (origScrub === undefined) delete process.env.GBRAIN_EVAL_SCRUB_PII;
else process.env.GBRAIN_EVAL_SCRUB_PII = origScrub;
if (origMode === undefined) delete process.env.GBRAIN_CONTRIBUTOR_MODE;
else process.env.GBRAIN_CONTRIBUTOR_MODE = origMode;
};
test('stash=true turns capture on when file plane is silent (the #1475 repro)', () => {
delete process.env.GBRAIN_CONTRIBUTOR_MODE;
process.env.GBRAIN_EVAL_CAPTURE = 'true';
try {
expect(isEvalCaptureEnabled(null)).toBe(true);
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const noEval: any = { engine: 'pglite' };
expect(isEvalCaptureEnabled(noEval)).toBe(true);
} finally { restore(); }
});
test('stash=false wins over CONTRIBUTOR_MODE=1 (explicit per-key beats broad flag)', () => {
process.env.GBRAIN_CONTRIBUTOR_MODE = '1';
process.env.GBRAIN_EVAL_CAPTURE = 'false';
try {
expect(isEvalCaptureEnabled(null)).toBe(false);
} finally { restore(); }
});
test('file-plane explicit value still wins over the stash', () => {
process.env.GBRAIN_EVAL_CAPTURE = 'true';
try {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const disabled: any = { engine: 'pglite', eval: { capture: false } };
expect(isEvalCaptureEnabled(disabled)).toBe(false);
} finally { restore(); }
});
test('scrub stash: false disables, file plane wins, default stays true', () => {
process.env.GBRAIN_EVAL_SCRUB_PII = 'false';
try {
expect(isEvalScrubEnabled(null)).toBe(false);
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const fileWins: any = { engine: 'pglite', eval: { scrub_pii: true } };
expect(isEvalScrubEnabled(fileWins)).toBe(true);
} finally { restore(); }
delete process.env.GBRAIN_EVAL_SCRUB_PII;
try {
expect(isEvalScrubEnabled(null)).toBe(true);
} finally { restore(); }
});
});
+71
View File
@@ -403,6 +403,77 @@ describe('extractPageLinks', () => {
expect(candidates).toEqual([]);
});
test('path-qualified wikilink outside DIR_PATTERN queries by final segment', async () => {
// `[[notes/struktura]]` (dir not in DIR_PATTERN) falls to the generic
// pass. The resolver's basename index is keyed by final path segments,
// so the lookup must strip the dirname — mirroring the FS path
// (resolveSlugAll). Regression: the raw literal was passed through,
// which never matched, so these links silently dropped.
const seen: string[] = [];
const resolver: SlugResolver = {
resolve: async () => null,
resolveBasenameMatches: async (name) => {
seen.push(name);
return name === 'struktura' ? ['notes/struktura'] : [];
},
};
const { candidates } = await extractPageLinks(
'concepts/x', 'See [[notes/struktura]].',
{}, 'concept', resolver, { globalBasename: true },
);
expect(seen).toContain('struktura');
expect(seen).not.toContain('notes/struktura');
expect(candidates.map(c => c.targetSlug)).toEqual(['notes/struktura']);
expect(candidates[0].linkType).toBe('wikilink_basename');
expect(candidates[0].linkSource).toBe('wikilink-resolved');
});
test('path-qualified wikilink keeps only matches ending with the written path', async () => {
// The written path disambiguates: `[[notes/struktura]]` must never
// attach to `wiki/struktura` even though both share the basename.
const resolver: SlugResolver = {
resolve: async () => null,
resolveBasenameMatches: async (name) =>
name === 'struktura' ? ['notes/struktura', 'wiki/struktura'] : [],
};
const { candidates } = await extractPageLinks(
'concepts/x', 'See [[notes/struktura]].',
{}, 'concept', resolver, { globalBasename: true },
);
expect(candidates.map(c => c.targetSlug)).toEqual(['notes/struktura']);
});
test('path-qualified wikilink matches a deeper real slug by path suffix', async () => {
// The page lives at vault/notes/struktura; the author wrote the shorter
// tail `[[notes/struktura]]`. Suffix matching connects them, while the
// basename-only sibling `wiki/struktura` stays excluded.
const resolver: SlugResolver = {
resolve: async () => null,
resolveBasenameMatches: async (name) =>
name === 'struktura' ? ['vault/notes/struktura', 'wiki/struktura'] : [],
};
const { candidates } = await extractPageLinks(
'concepts/x', 'See [[notes/struktura]].',
{}, 'concept', resolver, { globalBasename: true },
);
expect(candidates.map(c => c.targetSlug)).toEqual(['vault/notes/struktura']);
});
test('path-qualified self-link is dropped like the bare form', async () => {
// `[[notes/struktura]]` written on notes/struktura itself must not
// produce a self-loop (same guard as the bare `[[own-tail]]` case).
const resolver: SlugResolver = {
resolve: async () => null,
resolveBasenameMatches: async (name) =>
name === 'struktura' ? ['notes/struktura'] : [],
};
const { candidates } = await extractPageLinks(
'notes/struktura', 'See [[notes/struktura]].',
{}, 'concept', resolver, { globalBasename: true },
);
expect(candidates).toEqual([]);
});
test('bare wikilink resolution does not interfere with DIR_PATTERN wikilinks', async () => {
// 2b refs (people/alice) take the verb-inferred type;
// 2c refs (struktura) take wikilink_basename. Same call.
-153
View File
@@ -1,153 +0,0 @@
/**
* `gbrain lint` source-glob filter walker integration.
*
* PR #2157 (commit cf9a3b18, `feat/sync-source-glob-filters`) wired
* `sources.config.include_globs` / `exclude_globs` into `gbrain sync` so a
* user could exclude `Resources/veriff/**` and have every subsequent sync
* honor it. The lint command walked the same source dirs blind and emitted
* findings against paths the user had already declared out of scope a
* half-finished feature.
*
* This patch extends the same persisted glob contract to lint:
* - `gbrain lint` gains `--include / --exclude` flags (parallel to sync).
* - `runLintCore` lifts `sources.config.{include,exclude}_globs` for any
* target whose absolute path matches a source row's `local_path`, so the
* cycle.lint phase + Minion lint handlers honor the same filter without
* restating it.
* - The walker in `collectPages` applies the filter using the SAME
* `matchesAnyGlob` helper sync uses, anchored at the target dir (so a
* persisted `Resources/veriff/**` glob written against the source root
* works without rewriting it as an absolute path).
*
* These tests pin the walker contract. The engine-side lift
* (`resolveSourceGlobsForTarget`) is best-effort by design (returns `{}` on
* any error) and is exercised by the dream-cycle lint phase end-to-end.
*/
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
import { mkdtempSync, mkdirSync, writeFileSync, rmSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';
// runLintCore is the library entry — the same surface the cycle.lint phase
// and Minion handlers call. Exercising it covers the walker via its real
// callsite; testing `collectPages` directly would skip the wiring.
import { runLintCore } from '../src/commands/lint.ts';
// A self-contained content-sanity stub so the test never touches a real
// engine / config file. Empty operator-literal list keeps the content-sanity
// pass silent so the only findings come from the structural rules
// (no-frontmatter etc.).
const STUB_CS = {
fail_on_throw: false,
warn_on_throw: false,
bytes_warn: 1024 * 1024,
operator_literals: [],
};
describe('runLintCore — source-glob walker filter', () => {
let root: string;
beforeAll(() => {
root = mkdtempSync(join(tmpdir(), 'gbrain-lint-globs-'));
// Three subtrees with mixed structured / archive-style content.
// All pages have `# Title` headers but no frontmatter so each one
// emits at least one `no-frontmatter` issue under the default rule set.
mkdirSync(join(root, 'Notes'), { recursive: true });
mkdirSync(join(root, 'Resources', 'veriff'), { recursive: true });
mkdirSync(join(root, 'Resources', 'prior-art', 'archive-v1'), { recursive: true });
writeFileSync(join(root, 'Notes', 'a.md'), '# A\nbody\n');
writeFileSync(join(root, 'Notes', 'b.md'), '# B\nbody\n');
writeFileSync(join(root, 'Resources', 'veriff', 'spec-1.md'), '# Veriff spec 1\nbody\n');
writeFileSync(join(root, 'Resources', 'veriff', 'spec-2.md'), '# Veriff spec 2\nbody\n');
writeFileSync(join(root, 'Resources', 'prior-art', 'archive-v1', 'old.md'), '# Old\nbody\n');
});
afterAll(() => {
rmSync(root, { recursive: true, force: true });
});
test('no filter — walks every .md (regression guard for default behavior)', async () => {
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
});
expect(result.pages_scanned).toBe(5);
expect(result.pages_with_issues).toBeGreaterThan(0);
});
test('exclude glob skips matching paths (Resources/veriff/** off-limits)', async () => {
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
exclude: ['Resources/veriff/**'],
});
// 5 total minus 2 veriff specs = 3 pages walked.
expect(result.pages_scanned).toBe(3);
});
test('exclude with multiple patterns is union (veriff + prior-art both skipped)', async () => {
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
exclude: ['Resources/veriff/**', 'Resources/prior-art/**'],
});
// 5 total minus 3 (2 veriff + 1 archive-v1) = 2 pages walked.
expect(result.pages_scanned).toBe(2);
});
test('include glob narrows the walk to matching paths only', async () => {
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
include: ['Notes/**'],
});
expect(result.pages_scanned).toBe(2);
});
test('exclude runs AFTER include (same precedence as `gbrain sync`)', async () => {
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
include: ['**/*.md'],
exclude: ['Resources/**'],
});
// include lets everything through; exclude drops the 3 Resources/* files.
expect(result.pages_scanned).toBe(2);
});
test('empty include / exclude arrays do NOT engage the filter', async () => {
// Symmetric with `parseGlobList` returning undefined for empty input —
// an empty include would otherwise classify every path as a miss and
// silently zero out the lint scope. Pin the guard at the walker level.
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
include: [],
exclude: [],
});
expect(result.pages_scanned).toBe(5);
});
test('exclude semantics match sync — `**` matches across path segments', async () => {
const result = await runLintCore({
target: root,
contentSanity: STUB_CS,
exclude: ['**/spec-*.md'],
});
// Both Veriff specs match the deep glob; Notes + archive-v1 survive.
expect(result.pages_scanned).toBe(3);
});
test('single-file target bypasses the filter (file mode is not a walk)', async () => {
// A user lints one .md explicitly: filters are a directory-walk concern,
// so the file is processed even if its name would match an exclude.
const result = await runLintCore({
target: join(root, 'Resources', 'veriff', 'spec-1.md'),
contentSanity: STUB_CS,
exclude: ['Resources/veriff/**'],
});
expect(result.pages_scanned).toBe(1);
});
});
+25
View File
@@ -302,4 +302,29 @@ describe('loadConfigWithEngine (Phase 4 / F3)', () => {
expect(merged?.engine).toBe('pglite');
});
});
describe('eval.* DB-plane merge (#1475)', () => {
test('gbrain config set eval.capture true reaches the merged config', async () => {
// The #1475 repro: DB plane has eval.capture=true, file plane silent.
// Pre-fix the merge skipped eval.* entirely and capture never fired.
const base: GBrainConfig = { engine: 'pglite' };
const engine = makeEngine({ 'eval.capture': 'true', 'eval.scrub_pii': 'false' });
const merged = await loadConfigWithEngine(engine, base);
expect(merged?.eval?.capture).toBe(true);
expect(merged?.eval?.scrub_pii).toBe(false);
});
test('file plane wins per key; DB fills only the gaps', async () => {
const base: GBrainConfig = { engine: 'pglite', eval: { capture: false } };
const engine = makeEngine({ 'eval.capture': 'true', 'eval.scrub_pii': 'false' });
const merged = await loadConfigWithEngine(engine, base);
expect(merged?.eval?.capture).toBe(false); // file wins
expect(merged?.eval?.scrub_pii).toBe(false); // DB fills the gap
});
test('no eval keys anywhere leaves cfg.eval undefined', async () => {
const merged = await loadConfigWithEngine(makeEngine({}), { engine: 'pglite' });
expect(merged?.eval).toBeUndefined();
});
});
});
+62
View File
@@ -0,0 +1,62 @@
import { describe, expect, test } from 'bun:test';
import {
extractCycleFreshnessSourceIds,
parseMaintainArgs,
} from '../src/commands/maintain.ts';
import type { Check } from '../src/commands/doctor.ts';
describe('maintain args', () => {
test('defaults to dry-run unless --safe is explicit', () => {
expect(parseMaintainArgs([])).toMatchObject({
safe: false,
dryRun: true,
json: false,
});
});
test('--safe enables mutating safe mode', () => {
expect(parseMaintainArgs(['--safe', '--json'])).toMatchObject({
safe: true,
dryRun: false,
json: true,
});
});
test('--dry-run wins over --safe', () => {
expect(parseMaintainArgs(['--safe', '--dry-run'])).toMatchObject({
safe: true,
dryRun: true,
});
});
});
describe('cycle freshness source extraction', () => {
test('extracts stale source ids from doctor messages', () => {
const checks: Check[] = [
{
name: 'cycle_freshness',
status: 'fail',
message: "Source 'brain-sync-remote-teffur' last cycled 40h ago. Run `gbrain dream --source <id>`.",
},
{
name: 'cycle_freshness',
status: 'fail',
message: "Source 'wiki' last cycled 25h ago. Source 'wiki' last cycled 25h ago.",
},
];
expect(extractCycleFreshnessSourceIds(checks)).toEqual([
'brain-sync-remote-teffur',
'wiki',
]);
});
test('ignores ok and unrelated checks', () => {
const checks: Check[] = [
{ name: 'cycle_freshness', status: 'ok', message: "Source 'fresh' last cycled recently." },
{ name: 'frontmatter_integrity', status: 'warn', message: "Source 'wiki' has frontmatter issues." },
];
expect(extractCycleFreshnessSourceIds(checks)).toEqual([]);
});
});
+17
View File
@@ -0,0 +1,17 @@
import { describe, expect, it } from 'bun:test';
import { readFileSync } from 'fs';
import { join } from 'path';
describe('root OpenClaw plugin manifest', () => {
it('declares the id required by OpenClaw plugin installs', () => {
const manifest = JSON.parse(readFileSync(join(import.meta.dir, '..', 'openclaw.plugin.json'), 'utf8'));
const entrySource = readFileSync(join(import.meta.dir, '..', 'src', 'openclaw-context-engine.ts'), 'utf8');
const entryId = entrySource.match(/id:\s*'([^']+)'/)?.[1];
expect(manifest.id).toBe(entryId);
expect(manifest.configSchema).toBeDefined();
expect(typeof manifest.configSchema).toBe('object');
expect(manifest.contracts?.contextEngines).toContain('gbrain-context');
expect(entrySource).toContain('export function register');
});
});
+56
View File
@@ -186,11 +186,67 @@ describe('shouldExclude — orphan filter regression (preserve curation)', () =>
expect(shouldExclude('entities/anonymous')).toBe(true);
expect(shouldExclude('atoms/fact-123')).toBe(true);
expect(shouldExclude('skills/gbrain-operations')).toBe(true);
expect(shouldExclude('dreaming/light/2026-07-20')).toBe(true);
expect(shouldExclude('daily/2026-07-20')).toBe(true);
expect(shouldExclude('agent-openclaw/daily/2026-07-20')).toBe(true);
});
test('workspace convention slugs are excluded', () => {
expect(shouldExclude('_brain-conventions')).toBe(true);
expect(shouldExclude('_templates/decision')).toBe(true);
expect(shouldExclude('extracts/2026-06-30/takes.proposed/round-single')).toBe(true);
expect(shouldExclude('2026-07-20')).toBe(true);
expect(shouldExclude('2026-07-20-qa-sweep')).toBe(true);
expect(shouldExclude('agents/arya/identity')).toBe(true);
expect(shouldExclude('agents/arya/memory/dreaming/deep/2026-07-20')).toBe(true);
});
test('regular slugs are NOT excluded', () => {
expect(shouldExclude('people/alice')).toBe(false);
expect(shouldExclude('companies/acme')).toBe(false);
expect(shouldExclude('writing/post-1')).toBe(false);
expect(shouldExclude('agents/arya/qa-reports/launch-review')).toBe(false);
});
});
describe('getHealth orphan_pages uses shared exclusion policy', () => {
test('excluded convention islands do not count against health', async () => {
await engine.putPage('_templates/decision', {
type: 'template', title: 'Decision', compiled_truth: 'template', timeline: '', frontmatter: {},
});
await engine.putPage('skills/arya/source-check', {
type: 'concept', title: 'Skill', compiled_truth: 'skill', timeline: '', frontmatter: {},
});
await engine.putPage('agents/arya/identity', {
type: 'note', title: 'Identity', compiled_truth: 'identity', timeline: '', frontmatter: {},
});
await engine.putPage('people/alice', {
type: 'person', title: 'Alice', compiled_truth: 'real island', timeline: '', frontmatter: {},
});
const health = await engine.getHealth();
expect(health.orphan_pages).toBe(1);
});
test('per-brain config overrides (orphans.exclude_*) also apply to health', async () => {
await engine.putPage('my-private-folder/secret-ref', {
type: 'note', title: 'Ref', compiled_truth: 'ref', timeline: '', frontmatter: {},
});
await engine.putPage('one-off-fixture-page', {
type: 'note', title: 'Fixture', compiled_truth: 'fixture', timeline: '', frontmatter: {},
});
await engine.putPage('people/alice', {
type: 'person', title: 'Alice', compiled_truth: 'real island', timeline: '', frontmatter: {},
});
expect((await engine.getHealth()).orphan_pages).toBe(3);
await engine.setConfig('orphans.exclude_prefixes', 'my-private-folder/');
await engine.setConfig('orphans.exclude_slugs', 'one-off-fixture-page');
expect((await engine.getHealth()).orphan_pages).toBe(1);
await engine.unsetConfig('orphans.exclude_prefixes');
await engine.unsetConfig('orphans.exclude_slugs');
});
});
+38
View File
@@ -66,6 +66,10 @@ describe('shouldExclude', () => {
expect(shouldExclude('templates/meeting-note')).toBe(true);
});
test('excludes deny-prefix: _templates/', () => {
expect(shouldExclude('_templates/meeting-note')).toBe(true);
});
test('excludes deny-prefix: openclaw/config/', () => {
expect(shouldExclude('openclaw/config/agent')).toBe(true);
});
@@ -86,10 +90,44 @@ describe('shouldExclude', () => {
expect(shouldExclude('entities/product-hunt')).toBe(true);
});
test('excludes first-segment: skills, dreaming, and daily', () => {
expect(shouldExclude('skills/arya/source-check')).toBe(true);
expect(shouldExclude('dreaming/light/2026-07-20')).toBe(true);
expect(shouldExclude('daily/2026-07-20')).toBe(true);
expect(shouldExclude('agent-openclaw/daily/2026-07-20')).toBe(true);
});
test('excludes root date logs and agent workspace conventions', () => {
expect(shouldExclude('_brain-conventions')).toBe(true);
expect(shouldExclude('2026-07-20')).toBe(true);
expect(shouldExclude('2026-07-20-qa-sweep')).toBe(true);
expect(shouldExclude('agents/arya/identity')).toBe(true);
expect(shouldExclude('agents/arya/memory/dreaming/deep/2026-07-20')).toBe(true);
});
test('excludes generated extracts', () => {
expect(shouldExclude('extracts/2026-06-30/takes.proposed/round-single')).toBe(true);
});
test('brain-specific exclusions come from config overrides, not global defaults', () => {
// No baked-in defaults for these:
expect(shouldExclude('my-private-folder/some-secret-ref.md')).toBe(false);
expect(shouldExclude('one-off-fixture-page')).toBe(false);
// The per-brain config plane (orphans.exclude_prefixes / exclude_slugs):
const overrides = {
excludePrefixes: ['my-private-folder/'],
excludeSlugs: ['one-off-fixture-page'],
};
expect(shouldExclude('my-private-folder/some-secret-ref.md', overrides)).toBe(true);
expect(shouldExclude('one-off-fixture-page', overrides)).toBe(true);
expect(shouldExclude('people/jane-doe', overrides)).toBe(false);
});
test('does NOT exclude a normal content page', () => {
expect(shouldExclude('companies/acme')).toBe(false);
expect(shouldExclude('people/jane-doe')).toBe(false);
expect(shouldExclude('projects/gbrain')).toBe(false);
expect(shouldExclude('agents/arya/qa-reports/launch-review')).toBe(false);
});
test('does NOT exclude a page ending with log-like text that is not /log', () => {
+8 -2
View File
@@ -218,15 +218,21 @@ describe('progress reporter', () => {
test('only one process-level signal handler installed across many reporters', () => {
// Baseline: one handler already installed by prior tests in this file.
const installedBefore = __signalHandlerInstalledForTest();
// liveReporters is process-global: earlier test files in the same shard
// can leave a live entry behind (e.g. a production path that skips
// finish() on an error branch). Assert NET-zero leak from THIS test's
// lifecycles, not an absolute zero we don't control — same tolerance
// the handler assertion below already applies via `installedBefore`.
const liveBefore = __liveReporterCountForTest();
const { stream } = sink(false);
for (let i = 0; i < 50; i++) {
const p = createProgress({ mode: 'json', stream, minIntervalMs: 0, minItems: 1 });
p.start(`phase_${i}`, 1);
p.finish();
}
// After 50 reporter lifecycles, still exactly one handler and zero leaked live entries.
// After 50 reporter lifecycles, still exactly one handler and zero NEWLY leaked live entries.
expect(__signalHandlerInstalledForTest()).toBe(installedBefore || true);
expect(__liveReporterCountForTest()).toBe(0);
expect(__liveReporterCountForTest()).toBe(liveBefore);
});
test('startHeartbeat() fires heartbeats and stop() clears', async () => {
-6
View File
@@ -694,12 +694,6 @@ const COLUMN_EXEMPTIONS = new Set<string>([
'minion_jobs.quiet_hours',
'minion_jobs.stagger_key',
'sources.chunker_version',
// #2157 follow-on (migration v125). TEXT column read by performSync's
// `Already up to date` gate; not referenced by any CREATE INDEX. Same
// upgrade-path coverage as sources.chunker_version above: fresh installs
// get it via the CREATE TABLE in src/schema.sql + schema-embedded.ts;
// pre-existing brains get it via the idempotent ALTER TABLE in v125.
'sources.config_fingerprint',
'access_tokens.permissions',
'takes.resolved_quality',
'pages.emotional_weight_recomputed_at',
+15
View File
@@ -12,9 +12,11 @@ import {
bestPath,
historyPath,
loadHistory,
proposedPath,
revertAllPending,
skillPath,
versionsDir,
writeProposed,
} from '../../src/core/skillopt/version-store.ts';
let tmpDir: string;
@@ -79,6 +81,19 @@ describe('acceptCandidate (D8 two-phase commit)', () => {
});
});
describe('writeProposed', () => {
test('writes distinct best and proposed artifacts without mutating SKILL.md (#2635)', () => {
const candidate = '---\nname: test\n---\nproposed body\n';
const written = writeProposed(tmpDir, SKILL, candidate);
expect(written).toBe(proposedPath(tmpDir, SKILL));
expect(fs.readFileSync(bestPath(tmpDir, SKILL), 'utf8')).toBe(candidate);
expect(fs.readFileSync(proposedPath(tmpDir, SKILL), 'utf8')).toBe(candidate);
expect(fs.readFileSync(skillPath(tmpDir, SKILL), 'utf8')).toContain('baseline body');
});
});
describe('revertAllPending (D8 crash recovery)', () => {
test('no-op when no pending rows', () => {
const reverted = revertAllPending(tmpDir, SKILL);
-121
View File
@@ -159,127 +159,6 @@ describe('sources add', () => {
await expect(runSources(engine, ['add', 'plans', '--path', '/tmp/gstack/plans']))
.rejects.toThrow(/overlaps with existing source "gstack"/);
});
// Glob filters — TODO #3 from the brettdavies fork recon. Pre-fix, the
// `SyncableOptions` shape in `src/core/sync.ts` had been carrying
// `include` / `exclude` since v0.41.13, but commands/sync.ts:1454 never
// populated them and `sources add` had no flag to persist them — so users
// had no way to tell gbrain to skip `Templates/` in an Obsidian vault.
test('--exclude persists glob into sources.config.exclude_globs', async () => {
const { engine, calls } = makeStub({
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
id: 'vault',
name: 'vault',
local_path: '/tmp/vault',
last_commit: null,
last_sync_at: null,
config: '{"exclude_globs":["Templates/**"]}',
created_at: new Date(),
}],
});
await runSources(engine, ['add', 'vault', '--path', '/tmp/vault', '--exclude', 'Templates/**']);
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
expect(insert!.params[3]).toBe('{"exclude_globs":["Templates/**"]}');
});
test('--include persists glob into sources.config.include_globs', async () => {
const { engine, calls } = makeStub({
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
id: 'wiki',
name: 'wiki',
local_path: '/tmp/wiki',
last_commit: null,
last_sync_at: null,
config: '{"include_globs":["people/**"]}',
created_at: new Date(),
}],
});
await runSources(engine, ['add', 'wiki', '--path', '/tmp/wiki', '--include', 'people/**']);
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
expect(insert!.params[3]).toBe('{"include_globs":["people/**"]}');
});
test('--exclude is repeatable; preserves order', async () => {
const { engine, calls } = makeStub({
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
id: 'vault',
name: 'vault',
local_path: '/tmp/vault',
last_commit: null,
last_sync_at: null,
config: '{}',
created_at: new Date(),
}],
});
await runSources(engine, [
'add', 'vault', '--path', '/tmp/vault',
'--exclude', 'Templates/**',
'--exclude', '.smart-env/**',
'--exclude', 'Drafts/**',
]);
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
expect(insert!.params[3]).toBe('{"exclude_globs":["Templates/**",".smart-env/**","Drafts/**"]}');
});
test('--include and --exclude compose in one command (federated source with both filter axes)', async () => {
const { engine, calls } = makeStub({
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
id: 'vault',
name: 'vault',
local_path: '/tmp/vault',
last_commit: null,
last_sync_at: null,
config: '{"federated":true,"include_globs":["people/**"],"exclude_globs":["Templates/**"]}',
created_at: new Date(),
}],
});
await runSources(engine, [
'add', 'vault', '--path', '/tmp/vault', '--federated',
'--include', 'people/**',
'--exclude', 'Templates/**',
]);
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
expect(insert!.params[3]).toBe(
'{"federated":true,"include_globs":["people/**"],"exclude_globs":["Templates/**"]}',
);
});
test('omitted glob flags leave config untouched (no [] entries persisted)', async () => {
// Regression guard: empty glob arrays must NOT be written. Otherwise a
// brain that never opts into filtering grows {"include_globs": [],
// "exclude_globs": []} cruft in every source row, and the parseGlobList
// path would return undefined anyway (the cruft is purely noise).
const { engine, calls } = makeStub({
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
id: 'gstack',
name: 'gstack',
local_path: '/tmp/gstack',
last_commit: null,
last_sync_at: null,
config: '{}',
created_at: new Date(),
}],
});
await runSources(engine, ['add', 'gstack', '--path', '/tmp/gstack']);
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
expect(insert!.params[3]).toBe('{}');
});
test('--exclude requires a glob argument', async () => {
const { engine } = makeStub();
const code = await withExitCapture(() => runSources(engine, [
'add', 'vault', '--path', '/tmp/vault', '--exclude',
]));
expect(code).toBe(2);
});
test('--include rejects a flag-like value (--include --path looks like a typo)', async () => {
const { engine } = makeStub();
const code = await withExitCapture(() => runSources(engine, [
'add', 'vault', '--path', '/tmp/vault', '--include', '--federated',
]));
expect(code).toBe(2);
});
});
// ── add — #2707 git-repo validation (CLI wiring) ───────────────
-95
View File
@@ -1,95 +0,0 @@
/**
* #2157 follow-on (migration v125) end-to-end gate wiring.
*
* `test/sync-config-fingerprint.test.ts` pins the persistence + comparison
* primitives (compute/read/write). This file pins the WIRING inside
* `performSync`: with git HEAD unchanged, a drift in the walk-affecting
* `sources.config` fields must break out of the "Already up to date" early
* return and force a full re-walk and the re-stamped fingerprint must
* settle the gate back to `up_to_date` on the following pass. Deleting the
* `configMismatch` term from the gate condition fails this test; none of the
* primitive tests would catch that.
*/
import { test, expect, beforeAll, afterAll } from 'bun:test';
import { mkdtempSync, mkdirSync, writeFileSync, rmSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';
import { execFileSync } from 'child_process';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { performSync } from '../src/commands/sync.ts';
let engine: PGLiteEngine;
let repoPath: string;
function git(cwd: string, ...args: string[]) {
execFileSync('git', args, { cwd, stdio: 'pipe' });
}
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
repoPath = mkdtempSync(join(tmpdir(), 'gbrain-fp-gate-'));
mkdirSync(join(repoPath, 'wiki'));
mkdirSync(join(repoPath, 'memory'));
writeFileSync(join(repoPath, 'wiki', 'page1.md'), '# Page 1\n\nbody\n');
writeFileSync(join(repoPath, 'memory', 'note1.md'), '# Note 1\n\nbody\n');
git(repoPath, 'init');
git(repoPath, 'add', '-A');
git(repoPath, '-c', 'user.email=t@example.com', '-c', 'user.name=t', 'commit', '-m', 'init');
await engine.executeRaw(
`INSERT INTO sources (id, name, local_path, config) VALUES ($1, $2, $3, $4::text::jsonb)`,
['vault', 'vault', repoPath, JSON.stringify({ include_globs: ['wiki/**'] })],
);
}, 60_000);
afterAll(async () => {
await engine?.disconnect();
rmSync(repoPath, { recursive: true, force: true });
});
test('config-glob drift with unchanged git HEAD forces a re-walk, then settles', async () => {
// First sync: row config include_globs = ['wiki/**'], caller threads it
// (as syncOneSource / the single-source CLI path do). memory/* skipped.
const first = await performSync(engine, {
repoPath, sourceId: 'vault', include: ['wiki/**'],
noPull: true, noEmbed: true, full: true,
});
expect(first.status).toBe('first_sync');
expect(await engine.getPage('wiki/page1')).not.toBeNull();
expect(await engine.getPage('memory/note1')).toBeNull();
// No drift, HEAD unchanged: gate stays quiet.
const second = await performSync(engine, {
repoPath, sourceId: 'vault', include: ['wiki/**'],
noPull: true, noEmbed: true,
});
expect(second.status).toBe('up_to_date');
// User widens the persisted globs (what `gbrain sources add --include`
// writes). Git HEAD has NOT moved.
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[JSON.stringify({ include_globs: ['wiki/**', 'memory/**'] }), 'vault'],
);
// Pre-fix this returned `up_to_date` (HEAD unchanged) and memory/note1
// stayed missing until a manual `--full`. The fingerprint gate must force
// the full re-walk instead.
const third = await performSync(engine, {
repoPath, sourceId: 'vault', include: ['wiki/**', 'memory/**'],
noPull: true, noEmbed: true,
});
expect(third.status).not.toBe('up_to_date');
expect(await engine.getPage('memory/note1')).not.toBeNull();
// Re-stamped fingerprint matches the current row: gate settles.
const fourth = await performSync(engine, {
repoPath, sourceId: 'vault', include: ['wiki/**', 'memory/**'],
noPull: true, noEmbed: true,
});
expect(fourth.status).toBe('up_to_date');
});
-234
View File
@@ -1,234 +0,0 @@
/**
* #2157 follow-on (migration v125 `sources.config_fingerprint`).
*
* The "Already up to date" gate at performSync's git-HEAD equality check
* honored chunker_version match but ignored `sources.config` drift.
* Changing `sources.config.exclude_globs` (or include_globs / strategy)
* had no observable effect on the next sync because git HEAD was
* unchanged the gate returned early and the new walk scope never
* applied. This file exercises the persistence shape + drift detection
* end-to-end on PGLite, including:
*
* - Migration v125 actually adds the column (regression guard against
* a future re-numbering or accidental deletion).
* - read/write round-trips preserve the value.
* - The fingerprint differs across the three walk-affecting fields
* and is order-insensitive on the array fields.
* - NULL fingerprint on pre-v125 rows treats as "not stamped" so a
* first post-upgrade sync doesn't spuriously force-full.
* - A toggle-and-revert leaves the stored fingerprint matching the
* current row, so the gate stays quiet.
*
* The wired-up gate behavior (force-full triggered on mismatch) is
* exercised by the existing sync end-to-end tests; here we pin the
* persistence + comparison primitives the gate depends on.
*/
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
import {
computeSourceConfigFingerprint,
readConfigFingerprint,
writeConfigFingerprint,
} from '../src/commands/sync.ts';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
}, 60_000);
afterAll(async () => {
await engine?.disconnect();
});
/** Insert a fresh source row with the given config. Returns the id. */
async function makeSource(
id: string,
config: Record<string, unknown> = {},
): Promise<string> {
await engine.executeRaw(
`INSERT INTO sources (id, name, local_path, config) VALUES ($1, $2, $3, $4::text::jsonb)`,
[id, id, `/tmp/${id}`, JSON.stringify(config)],
);
return id;
}
describe('migration v125 — sources.config_fingerprint column', () => {
test('column exists on the sources table', async () => {
const rows = await engine.executeRaw<{ column_name: string }>(
`SELECT column_name FROM information_schema.columns
WHERE table_name = 'sources' AND column_name = 'config_fingerprint'`,
);
expect(rows).toHaveLength(1);
});
test('column is nullable (preserves pre-migration row semantics)', async () => {
const rows = await engine.executeRaw<{ is_nullable: string }>(
`SELECT is_nullable FROM information_schema.columns
WHERE table_name = 'sources' AND column_name = 'config_fingerprint'`,
);
expect(rows[0]?.is_nullable).toBe('YES');
});
});
describe('readConfigFingerprint / writeConfigFingerprint — persistence round-trip', () => {
test('round-trip: write then read returns the same value', async () => {
const id = await makeSource('rt-basic', { exclude_globs: ['Templates/**'] });
const fp = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
await writeConfigFingerprint(engine, id, fp);
const got = await readConfigFingerprint(engine, id);
expect(got).toBe(fp);
});
test('NULL on never-stamped row (pre-v125 semantics)', async () => {
const id = await makeSource('rt-never');
const got = await readConfigFingerprint(engine, id);
expect(got).toBeNull();
});
test('undefined sourceId returns null (legacy non-source-scoped sync)', async () => {
const got = await readConfigFingerprint(engine, undefined);
expect(got).toBeNull();
});
test('write with undefined sourceId is a no-op (does not throw)', async () => {
// The legacy global-sync code path hits this branch; the guard must
// be silent rather than fail the sync run.
await writeConfigFingerprint(engine, undefined, 'deadbeef'.repeat(8));
// No assertion beyond "did not throw"; the function returns void.
});
test('overwrite: a second write replaces the prior fingerprint', async () => {
const id = await makeSource('rt-overwrite');
await writeConfigFingerprint(engine, id, 'a'.repeat(64));
await writeConfigFingerprint(engine, id, 'b'.repeat(64));
const got = await readConfigFingerprint(engine, id);
expect(got).toBe('b'.repeat(64));
});
});
describe('end-to-end drift simulation — the gate semantics this column enables', () => {
test('first stamp matches computed fingerprint of the row config', async () => {
const cfg = { exclude_globs: ['Templates/**', 'Photos/**'], strategy: 'markdown' };
const id = await makeSource('e2e-first-stamp', cfg);
const computed = computeSourceConfigFingerprint(cfg);
await writeConfigFingerprint(engine, id, computed);
expect(await readConfigFingerprint(engine, id)).toBe(computed);
});
test('exclude_globs mutation makes stored != current (drift detected)', async () => {
const before = { exclude_globs: ['Templates/**'] };
const after = { exclude_globs: ['Templates/**', 'Photos/**'] };
const id = await makeSource('e2e-exclude-drift', before);
const beforeFp = computeSourceConfigFingerprint(before);
await writeConfigFingerprint(engine, id, beforeFp);
// Simulate the user mutating sources.config via `gbrain sources add
// --exclude`. The gate's next read of (stored, computed-from-current)
// detects the drift and forces a re-walk.
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[JSON.stringify(after), id],
);
const afterFp = computeSourceConfigFingerprint(after);
const stored = await readConfigFingerprint(engine, id);
expect(stored).toBe(beforeFp);
expect(stored).not.toBe(afterFp);
});
test('toggle-and-revert: add then remove same pattern leaves stored matching current', async () => {
const original = { exclude_globs: ['Templates/**'] };
const id = await makeSource('e2e-toggle', original);
const originalFp = computeSourceConfigFingerprint(original);
await writeConfigFingerprint(engine, id, originalFp);
// Add a pattern (drift) then remove it (revert).
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[JSON.stringify({ exclude_globs: ['Templates/**', 'Photos/**'] }), id],
);
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[JSON.stringify(original), id],
);
const revertedFp = computeSourceConfigFingerprint(original);
expect(revertedFp).toBe(originalFp);
expect(await readConfigFingerprint(engine, id)).toBe(originalFp);
// ⇒ Gate compares storedFp (==originalFp) to currentFp (==originalFp): no drift, no force-full.
});
test('include_globs drift detected independently', async () => {
const before = { include_globs: ['people/**'] };
const after = { include_globs: ['people/**', 'companies/**'] };
const id = await makeSource('e2e-include-drift', before);
await writeConfigFingerprint(engine, id, computeSourceConfigFingerprint(before));
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[JSON.stringify(after), id],
);
const stored = await readConfigFingerprint(engine, id);
const current = computeSourceConfigFingerprint(after);
expect(stored).not.toBe(current);
});
test('strategy drift detected', async () => {
const before = { strategy: 'markdown' };
const after = { strategy: 'code' };
const id = await makeSource('e2e-strategy-drift', before);
await writeConfigFingerprint(engine, id, computeSourceConfigFingerprint(before));
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[JSON.stringify(after), id],
);
expect(await readConfigFingerprint(engine, id))
.not.toBe(computeSourceConfigFingerprint(after));
});
test('mutating unrelated config field (federated) does NOT drift', async () => {
// The fingerprint hashes ONLY walk-affecting fields. Federation
// changes search visibility, not the walk set — must not invalidate
// the checkpoint.
const id = await makeSource('e2e-federated-toggle', {
federated: true,
exclude_globs: ['Templates/**'],
});
await writeConfigFingerprint(
engine,
id,
computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] }),
);
await engine.executeRaw(
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
[
JSON.stringify({ federated: false, exclude_globs: ['Templates/**'] }),
id,
],
);
const stored = await readConfigFingerprint(engine, id);
const current = computeSourceConfigFingerprint({
federated: false,
exclude_globs: ['Templates/**'],
});
expect(stored).toBe(current);
});
test('double-encoded JSONB config (the sources-add stringify bug) hashes equivalently to the parsed object', async () => {
// `gbrain sources add` writes `JSON.stringify(config)::jsonb`, which
// double-encodes the value into a JSON-string scalar (`"{\"x\":1}"`)
// rather than a proper JSONB object. The defensive reader in
// postgres-engine.ts:1274 + readSourceConfig parses the string back
// before the fingerprint sees it, so a double-encoded row and a
// properly-shaped row must fingerprint identically.
const cfg = { exclude_globs: ['Templates/**'], strategy: 'markdown' };
const direct = computeSourceConfigFingerprint(cfg);
// The pure compute fn handles a pre-parsed object; the persistence
// layer's job is to deliver a parsed object. We assert that the
// round-trip a real read would produce (parse the string scalar)
// hashes to the same value.
const parsed = JSON.parse(JSON.stringify(cfg));
expect(computeSourceConfigFingerprint(parsed)).toBe(direct);
});
});
-63
View File
@@ -355,69 +355,6 @@ describe('sync monorepo subdir-source support (#753/#774)', () => {
expect(await engine.getPage('wiki/draft-a')).toBeNull();
});
// ─────────────────────────────────────────────────────────────────────────
// --include: allow-list counterpart (#2156). Same scope-relative anchoring
// as --exclude; exclude applies after include.
// ─────────────────────────────────────────────────────────────────────────
test('--include: only matching files import on full sync', async () => {
const { performSync } = await import('../src/commands/sync.ts');
const result = await performSync(engine, {
repoPath,
include: ['wiki/**'],
noPull: true,
noEmbed: true,
full: true,
});
expect(result.status).toBe('first_sync');
expect(result.added).toBe(2); // wiki/page1 + wiki/page2; memory/* miss the allow-list
expect(await engine.getPage('wiki/page1')).not.toBeNull();
expect(await engine.getPage('memory/note1')).toBeNull();
});
test('--include applies to the incremental path too', async () => {
const { performSync } = await import('../src/commands/sync.ts');
const first = await performSync(engine, {
repoPath,
include: ['wiki/**'],
noPull: true,
noEmbed: true,
full: true,
});
expect(first.status).toBe('first_sync');
writeFileSync(join(repoPath, 'wiki', 'page3.md'), mdPage('Wiki Page 3'));
writeFileSync(join(repoPath, 'memory', 'note3.md'), mdPage('Memory Note 3'));
gitCommit(repoPath, 'more pages');
const second = await performSync(engine, {
repoPath,
include: ['wiki/**'],
noPull: true,
noEmbed: true,
});
expect(second.status).toBe('synced');
expect(second.added).toBe(1); // wiki/page3 only; memory/note3 misses the allow-list
expect(await engine.getPage('wiki/page3')).not.toBeNull();
expect(await engine.getPage('memory/note3')).toBeNull();
});
test('--exclude applies after --include (path in both is rejected)', async () => {
const { performSync } = await import('../src/commands/sync.ts');
const result = await performSync(engine, {
repoPath,
include: ['wiki/**'],
exclude: ['wiki/page2.md'],
noPull: true,
noEmbed: true,
full: true,
});
expect(result.status).toBe('first_sync');
expect(result.added).toBe(1); // page1 only: page2 included then excluded
expect(await engine.getPage('wiki/page1')).not.toBeNull();
expect(await engine.getPage('wiki/page2')).toBeNull();
});
// ─────────────────────────────────────────────────────────────────────────
// --exclude '**/*' emits warning (NAV-4)
// ─────────────────────────────────────────────────────────────────────────
-212
View File
@@ -1,212 +0,0 @@
/**
* TODO #3 `parseGlobList` defensive parse.
*
* `sources.config` is a JSONB column with no schema. The runtime can find
* anything in `config.include_globs` / `config.exclude_globs`:
* - A user `gbrain sources add` wrote `["people/**"]` (the happy path).
* - A stray hand-edit wrote `"people/**"` (string, not array).
* - A future migration's null default.
* - A test fixture that left the column at `{}`.
*
* The parse must produce `string[] | undefined` so the downstream
* `SyncOpts.include` / `SyncOpts.exclude` are either undefined (no filter)
* or a non-empty list of usable globs. Returning `[]` would make
* `commands/sync.ts:1454` engage the filter loop with an empty allow-list
* that classifies every path as `include-glob-miss`.
*/
import { describe, test, expect } from 'bun:test';
import { parseGlobList, mergeGlobs, computeSourceConfigFingerprint } from '../src/commands/sync.ts';
describe('parseGlobList — JSONB-safe coercion to string[] | undefined', () => {
test('happy path: array of strings round-trips identically', () => {
expect(parseGlobList(['people/**', 'companies/**'])).toEqual(['people/**', 'companies/**']);
});
test('single-element array returned as-is', () => {
expect(parseGlobList(['Templates/**'])).toEqual(['Templates/**']);
});
test('non-array values return undefined (string, object, number, null)', () => {
expect(parseGlobList('Templates/**')).toBeUndefined();
expect(parseGlobList({ globs: ['Templates/**'] })).toBeUndefined();
expect(parseGlobList(42)).toBeUndefined();
expect(parseGlobList(null)).toBeUndefined();
expect(parseGlobList(undefined)).toBeUndefined();
});
test('empty array returns undefined (no engagement of the filter loop)', () => {
// Critical: a literal `[]` must not slip through. Empty `include` in
// SyncableOptions silently passes everything (good), but empty
// `exclude` is fine too — the real motivation is to keep `SyncOpts`
// unset so callers can ignore the field entirely. Symmetric with the
// `omitted glob flags leave config untouched` regression guard in
// sources.test.ts.
expect(parseGlobList([])).toBeUndefined();
});
test('mixed array drops non-string entries and keeps the rest', () => {
expect(parseGlobList(['people/**', 42, null, 'companies/**'])).toEqual([
'people/**',
'companies/**',
]);
});
test('empty strings dropped (a `""` glob would match every path)', () => {
expect(parseGlobList(['', 'people/**', ''])).toEqual(['people/**']);
});
test('array of only empty strings collapses to undefined', () => {
expect(parseGlobList(['', '', ''])).toBeUndefined();
});
});
describe('mergeGlobs — CLI flags union with persisted source-config globs', () => {
test('both sides present: union, deduped, CLI first', () => {
expect(mergeGlobs(['a/**', 'b/**'], ['b/**', 'c/**'])).toEqual(['a/**', 'b/**', 'c/**']);
});
test('CLI only', () => {
expect(mergeGlobs(['a/**'], undefined)).toEqual(['a/**']);
});
test('persisted only', () => {
expect(mergeGlobs([], ['Templates/**'])).toEqual(['Templates/**']);
});
test('neither side: undefined so SyncOpts stays unset', () => {
expect(mergeGlobs([], undefined)).toBeUndefined();
});
});
/**
* #2157 follow-on (sources_config_fingerprint, migration v125).
*
* computeSourceConfigFingerprint hashes the walk-affecting fields of
* sources.config (strategy + include_globs + exclude_globs) so the
* "Already up to date" gate at performSync's git-HEAD equality check
* can detect drift and force a re-walk. These cases pin the contract
* the gate depends on:
*
* - Deterministic over equivalent inputs (order-insensitive,
* defensively-coerced via parseGlobList).
* - Sensitive to each walk-affecting field separately.
* - Insensitive to fields the walker doesn't read (federated,
* unrelated keys).
* - A toggle-and-revert is a no-op (returns to the original hash).
*
* Without the canonicalization the gate would fire spuriously on
* cosmetic changes (e.g. a user re-ordering their exclude list) and
* miss real drift (e.g. an add-then-remove that nets to a different
* effective set than the stored fingerprint).
*/
describe('computeSourceConfigFingerprint — walk-affecting config drift detector', () => {
test('empty config produces a stable hash', () => {
const a = computeSourceConfigFingerprint({});
const b = computeSourceConfigFingerprint({});
expect(a).toBe(b);
expect(a).toMatch(/^[a-f0-9]{64}$/);
});
test('null / undefined / missing config all hash the same', () => {
const empty = computeSourceConfigFingerprint({});
expect(computeSourceConfigFingerprint(null)).toBe(empty);
expect(computeSourceConfigFingerprint(undefined)).toBe(empty);
});
test('same config → same hash (deterministic)', () => {
const cfg = { strategy: 'markdown', exclude_globs: ['Templates/**', 'Photos/**'] };
expect(computeSourceConfigFingerprint(cfg)).toBe(computeSourceConfigFingerprint(cfg));
});
test('array order does not affect hash (canonical sort)', () => {
const a = computeSourceConfigFingerprint({ exclude_globs: ['a/**', 'b/**', 'c/**'] });
const b = computeSourceConfigFingerprint({ exclude_globs: ['c/**', 'a/**', 'b/**'] });
expect(a).toBe(b);
});
test('exclude_globs change → different hash', () => {
const a = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
const b = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**', 'Photos/**'] });
expect(a).not.toBe(b);
});
test('include_globs change → different hash', () => {
const a = computeSourceConfigFingerprint({ include_globs: ['people/**'] });
const b = computeSourceConfigFingerprint({ include_globs: ['people/**', 'companies/**'] });
expect(a).not.toBe(b);
});
test('strategy change → different hash', () => {
const a = computeSourceConfigFingerprint({ strategy: 'markdown' });
const b = computeSourceConfigFingerprint({ strategy: 'code' });
expect(a).not.toBe(b);
});
test('strategy unset vs set differ', () => {
const unset = computeSourceConfigFingerprint({});
const set = computeSourceConfigFingerprint({ strategy: 'markdown' });
expect(unset).not.toBe(set);
});
test('add-then-remove returns to original hash (toggle is a no-op)', () => {
const original = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
const added = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**', 'Photos/**'] });
const reverted = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
expect(added).not.toBe(original);
expect(reverted).toBe(original);
});
test('non-walk-affecting fields are ignored (federated, unrelated keys)', () => {
const a = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'], federated: true });
const b = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'], federated: false });
const c = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'], some_unrelated_key: 'value' });
expect(a).toBe(b);
expect(a).toBe(c);
});
test('non-string strategy coerced to null (defensive)', () => {
// A hand-edited row could leave `strategy: 42` or `strategy: {}` — both
// collapse to the same shape as `strategy: undefined` so the fingerprint
// doesn't reflect a value the walker can't honor anyway.
const empty = computeSourceConfigFingerprint({});
expect(computeSourceConfigFingerprint({ strategy: 42 })).toBe(empty);
expect(computeSourceConfigFingerprint({ strategy: {} })).toBe(empty);
expect(computeSourceConfigFingerprint({ strategy: null })).toBe(empty);
});
test('defensive parsing: mixed-type glob arrays hash same as cleaned arrays', () => {
// parseGlobList drops non-string + empty entries; the fingerprint must
// reflect what the walker actually uses, not what the raw row says.
const dirty = computeSourceConfigFingerprint({
exclude_globs: ['Templates/**', 42, null, '', 'Photos/**'],
});
const clean = computeSourceConfigFingerprint({
exclude_globs: ['Templates/**', 'Photos/**'],
});
expect(dirty).toBe(clean);
});
test('empty array and missing field hash identically', () => {
const missing = computeSourceConfigFingerprint({});
const emptyArray = computeSourceConfigFingerprint({ exclude_globs: [] });
const emptyAfterClean = computeSourceConfigFingerprint({ exclude_globs: ['', '', ''] });
expect(emptyArray).toBe(missing);
expect(emptyAfterClean).toBe(missing);
});
test('non-array exclude_globs (string, object) hash same as missing', () => {
const missing = computeSourceConfigFingerprint({});
expect(computeSourceConfigFingerprint({ exclude_globs: 'Templates/**' })).toBe(missing);
expect(computeSourceConfigFingerprint({ exclude_globs: { foo: 'bar' } })).toBe(missing);
});
test('SHA-256 output shape: 64 hex characters', () => {
const fp = computeSourceConfigFingerprint({
strategy: 'markdown',
include_globs: ['people/**'],
exclude_globs: ['Templates/**', '.git/**'],
});
expect(fp).toMatch(/^[a-f0-9]{64}$/);
});
});