mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-16 09:52:22 +00:00
Compare commits
11
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c275fa3fab | ||
|
|
ee095f1ca2 | ||
|
|
5f123c1404 | ||
|
|
d9eb027bdd | ||
|
|
62e009d192 | ||
|
|
314fefa560 | ||
|
|
7f841fae7f | ||
|
|
1fabbb9849 | ||
|
|
64920f83c9 | ||
|
|
e861b92da7 | ||
|
|
948ccc7b4f |
@@ -71,8 +71,8 @@ GBrain is designed to be installed and operated by an AI agent. The fastest path
|
||||
|
||||
If you don't already have an AI agent platform running, start with one of these. Both are designed to read GBrain's install protocol and execute it:
|
||||
|
||||
- **[OpenClaw](https://github.com/openclawagents/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
|
||||
- **[Hermes](https://github.com/openclawagents/hermes)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
|
||||
- **[OpenClaw](https://github.com/openclaw/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
|
||||
- **[Hermes](https://github.com/NousResearch/hermes-agent)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
|
||||
|
||||
Then paste this into your agent:
|
||||
|
||||
|
||||
@@ -115,7 +115,6 @@ Full subcommand reference:
|
||||
|
||||
```
|
||||
gbrain sources add <id> --path <p> [--name <n>] [--federated|--no-federated] [--force]
|
||||
[--include <glob>...] [--exclude <glob>...]
|
||||
Register a source. id: [a-z0-9](?:[a-z0-9-]{0,30}[a-z0-9])?
|
||||
--path must be a git repo (or a subdirectory of one) — see
|
||||
"The git requirement for --path sources" below. --force
|
||||
@@ -132,43 +131,6 @@ gbrain sources federate <id>
|
||||
gbrain sources unfederate <id>
|
||||
```
|
||||
|
||||
## Filtering what gets synced (--include / --exclude)
|
||||
|
||||
`--include` and `--exclude` on `gbrain sources add` accept repeatable glob
|
||||
patterns and are honored by every subsequent sync AND lint of the source.
|
||||
Common Obsidian vault setups need to exclude authoring scaffolding so it
|
||||
doesn't pollute search:
|
||||
|
||||
```bash
|
||||
# Skip Templates/, Drafts/, and the smart-env sidecar; everything else syncs.
|
||||
gbrain sources add vault \
|
||||
--path ~/Documents/vault --federated \
|
||||
--exclude 'Templates/**' \
|
||||
--exclude 'Drafts/**' \
|
||||
--exclude '.smart-env/**'
|
||||
|
||||
# Or: only sync the people/ and companies/ subtrees of a CRM vault.
|
||||
gbrain sources add crm \
|
||||
--path ~/Documents/crm --no-federated \
|
||||
--include 'people/**' \
|
||||
--include 'companies/**'
|
||||
```
|
||||
|
||||
Both persist into `sources.config.include_globs` / `exclude_globs` arrays.
|
||||
The filter runs `include` first, then `exclude`, so a path inside
|
||||
`people/**` is still rejected if it also matches `exclude_globs`. Globs use
|
||||
the same matcher as the rest of gbrain's sync classifier (`matchesAnyGlob`
|
||||
in `src/core/sync.ts`) and are matched against the source-root-relative
|
||||
path. Exclusion is conservative: it never deletes previously-imported pages.
|
||||
|
||||
`gbrain sync --include <glob> --exclude <glob>` and
|
||||
`gbrain lint <dir> --include <glob> --exclude <glob>` take the same
|
||||
repeatable flags for one-off scope changes; for lint the persisted source
|
||||
globs are auto-applied when the lint target matches a source's `local_path`.
|
||||
Changing the persisted globs on an existing source triggers a full re-walk
|
||||
on the next sync (the source's config fingerprint invalidates the
|
||||
"already up to date" gate).
|
||||
|
||||
## The git requirement for --path sources
|
||||
|
||||
Every `--path` source must be a git repository (or live inside one — a
|
||||
|
||||
@@ -131,7 +131,9 @@ into gbrain so other clients can scaffold it. Default behavior:
|
||||
`~/.gbrain/harvest-private-patterns.txt` plus built-in defaults
|
||||
(canonical private fork name, common email regex, Slack channel pattern). Any
|
||||
match → rollback (delete the harvested files) and exit non-zero.
|
||||
- `openclaw.plugin.json` updated with the new slug, sorted.
|
||||
- `openclaw.plugin.json` updated with the new slug, sorted. Harvest must preserve
|
||||
the top-level OpenClaw-native plugin fields (`id`, `configSchema`, `contracts`)
|
||||
because OpenClaw validates those before it can install the package.
|
||||
- `--no-lint` bypasses the linter (after a manual editorial scrub).
|
||||
|
||||
Use the `skillpack-harvest` skill (its companion editorial workflow)
|
||||
|
||||
@@ -233,13 +233,14 @@ keep it or `git checkout` to throw it away. Nothing is committed for you.
|
||||
|
||||
**For a skill that ships with gbrain** (anything under the gbrain repo's own
|
||||
`skills/`): SkillOpt refuses to overwrite it by default and writes the winner to
|
||||
`skills/<name>/skillopt/best.md` instead, so an optimization pass can never
|
||||
silently mutate a skill other people depend on. Two ways to handle that:
|
||||
`skills/<name>/skillopt/proposed.md` instead (while keeping `best.md` as the
|
||||
optimizer's current-best pointer), so an optimization pass can never silently
|
||||
mutate a skill other people depend on. Two ways to handle that:
|
||||
|
||||
```bash
|
||||
# See the proposed improvement without touching SKILL.md (works for ANY skill):
|
||||
gbrain skillopt meeting-prep --split 1:1:1 --no-mutate
|
||||
# → writes skills/meeting-prep/skillopt/best.md (the proposed rewrite), prints its path. Copy what you want.
|
||||
# → writes skills/meeting-prep/skillopt/proposed.md, updates best.md, and prints the proposal path.
|
||||
|
||||
# Actually rewrite a bundled skill (explicit opt-in + an independent held-out set):
|
||||
gbrain skillopt brain-ops --split 1:1:1 --allow-mutate-bundled \
|
||||
|
||||
+2
-2
@@ -1565,8 +1565,8 @@ GBrain is designed to be installed and operated by an AI agent. The fastest path
|
||||
|
||||
If you don't already have an AI agent platform running, start with one of these. Both are designed to read GBrain's install protocol and execute it:
|
||||
|
||||
- **[OpenClaw](https://github.com/openclawagents/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
|
||||
- **[Hermes](https://github.com/openclawagents/hermes)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
|
||||
- **[OpenClaw](https://github.com/openclaw/openclaw)** — deploy [AlphaClaw on Render](https://render.com/deploy?repo=https://github.com/chrysb/alphaclaw) (one click, 8GB+ RAM)
|
||||
- **[Hermes](https://github.com/NousResearch/hermes-agent)** — deploy on [Railway](https://github.com/praveen-ks-2001/hermes-agent-template) (one click)
|
||||
|
||||
Then paste this into your agent:
|
||||
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
{
|
||||
"id": "gbrain-context-engine",
|
||||
"name": "gbrain",
|
||||
"version": "0.32.3.0",
|
||||
"description": "Personal knowledge brain with Postgres + pgvector hybrid search",
|
||||
|
||||
@@ -266,4 +266,5 @@ editorial pass.
|
||||
(e.g. `src/commands/<slug>.ts` if the host SKILL.md declares it
|
||||
in frontmatter)
|
||||
- gbrain's `openclaw.plugin.json` — adds the slug to `skills:`
|
||||
array, sorted alphabetically
|
||||
array, sorted alphabetically, without removing OpenClaw-native plugin fields
|
||||
like `id`, `configSchema`, or `contracts`
|
||||
|
||||
@@ -57,6 +57,8 @@ This mode guarantees:
|
||||
- `skills/manifest.json` lists every skill directory
|
||||
- `skills/RESOLVER.md` references every skill in the manifest
|
||||
- `openclaw.plugin.json` `skills[]` round-trips with both
|
||||
- `openclaw.plugin.json` keeps OpenClaw install-required native plugin fields
|
||||
(`id`, object `configSchema`, and `contracts.contextEngines` when applicable)
|
||||
- No MECE violations (duplicate triggers across skills)
|
||||
|
||||
### Phases
|
||||
@@ -72,7 +74,7 @@ This mode guarantees:
|
||||
### Automation
|
||||
|
||||
```bash
|
||||
bun test test/skills-conformance.test.ts test/resolver.test.ts
|
||||
bun test test/skills-conformance.test.ts test/resolver.test.ts test/openclaw-plugin-manifest.test.ts
|
||||
```
|
||||
|
||||
The CI-gated check is the package.json `test` script.
|
||||
|
||||
+50
-1
@@ -54,7 +54,7 @@ export function bigintToStringReplacer(_key: string, value: unknown): unknown {
|
||||
}
|
||||
|
||||
// CLI-only commands that bypass the operation layer
|
||||
export const CLI_ONLY = new Set(['init', 'reinit-pglite', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'sync', 'extract', 'extract-conversation-facts', 'enrich', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'skillpack', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'sources', 'mounts', 'dream', 'check-resolvable', 'routing-eval', 'skillify', 'smoke-test', 'providers', 'storage', 'repos', 'code-def', 'code-refs', 'reindex', 'reindex-code', 'reindex-frontmatter', 'code-callers', 'code-callees', 'reconcile-links', 'frontmatter', 'auth', 'friction', 'claw-test', 'book-mirror', 'takes', 'think', 'salience', 'anomalies', 'calibration', 'transcripts', 'models', 'remote', 'recall', 'forget', 'edges-backfill', 'cache', 'ze-switch', 'founder', 'brainstorm', 'lsd', 'schema', 'capture', 'onboard', 'conversation-parser', 'status', 'connect', 'skillopt', 'quarantine', 'self-upgrade', 'advisor', 'watch', 'reindex-search-vector']);
|
||||
export const CLI_ONLY = new Set(['init', 'reinit-pglite', 'upgrade', 'post-upgrade', 'check-update', 'integrations', 'publish', 'check-backlinks', 'lint', 'report', 'import', 'export', 'files', 'embed', 'serve', 'call', 'config', 'doctor', 'migrate', 'eval', 'bench', 'sync', 'extract', 'extract-conversation-facts', 'enrich', 'features', 'autopilot', 'graph-query', 'jobs', 'agent', 'apply-migrations', 'skillpack-check', 'skillpack', 'resolvers', 'integrity', 'repair-jsonb', 'orphans', 'maintain', 'sources', 'mounts', 'dream', 'check-resolvable', 'routing-eval', 'skillify', 'smoke-test', 'providers', 'storage', 'repos', 'code-def', 'code-refs', 'reindex', 'reindex-code', 'reindex-frontmatter', 'code-callers', 'code-callees', 'reconcile-links', 'frontmatter', 'auth', 'friction', 'claw-test', 'book-mirror', 'takes', 'think', 'salience', 'anomalies', 'calibration', 'transcripts', 'models', 'remote', 'recall', 'forget', 'edges-backfill', 'cache', 'ze-switch', 'founder', 'brainstorm', 'lsd', 'schema', 'capture', 'onboard', 'conversation-parser', 'status', 'connect', 'skillopt', 'quarantine', 'self-upgrade', 'advisor', 'watch', 'reindex-search-vector']);
|
||||
// CLI-only commands whose handlers print their own --help text. These are
|
||||
// excluded from the generic short-circuit so detailed per-command and
|
||||
// per-subcommand usage stays reachable.
|
||||
@@ -78,6 +78,8 @@ const CLI_ONLY_SELF_HELP = new Set([
|
||||
'capture',
|
||||
// v0.42 self-upgrade ships its own usage (flags + the agent-skill story).
|
||||
'self-upgrade',
|
||||
// maintain (#3015) prints its own usage block (modes + not-auto-applied list).
|
||||
'maintain',
|
||||
// v0.43 (#2095): watch ships WATCH_HELP (flags + the stdin-turn protocol).
|
||||
'watch',
|
||||
// v0.37 fix wave (Lane D.4 + CDX2-12): sync's --no-embed flag was
|
||||
@@ -104,6 +106,9 @@ const CLI_ONLY_SELF_HELP = new Set([
|
||||
// `gbrain connect --help` prints its own usage (flags + examples) from
|
||||
// runConnect; route around the generic one-line short-circuit.
|
||||
'connect',
|
||||
// #1474: bench-publish ships its own detailed HELP (flags, exit codes,
|
||||
// the export → publish → gate loop). Route around the generic stub.
|
||||
'bench',
|
||||
]);
|
||||
|
||||
// v114 (#1941): alias -> operation lookup, kept separate from `cliOps` so
|
||||
@@ -1427,6 +1432,29 @@ async function handleCliOnly(command: string, args: string[]) {
|
||||
return;
|
||||
}
|
||||
|
||||
// #1474: `gbrain bench publish` is pure file I/O (reads a captured
|
||||
// eval-candidates NDJSON from `gbrain eval export`, writes a baseline
|
||||
// NDJSON). No DB access; bypass connectEngine entirely so the documented
|
||||
// export → publish → gate loop works on machines without a brain.
|
||||
// The v0.41.1 wave shipped bench-publish.ts + docs/eval-bench.md but this
|
||||
// dispatcher case was never added, so the command hit 'Unknown command'.
|
||||
if (command === 'bench') {
|
||||
if (args[0] === 'publish') {
|
||||
const { runBenchPublish } = await import('./commands/bench-publish.ts');
|
||||
await runBenchPublish(args.slice(1));
|
||||
return;
|
||||
}
|
||||
if (args.length === 0 || args[0] === '--help' || args[0] === '-h') {
|
||||
const { runBenchPublish } = await import('./commands/bench-publish.ts');
|
||||
await runBenchPublish(['--help']);
|
||||
return;
|
||||
}
|
||||
console.error(`Unknown bench subcommand: ${args[0]}`);
|
||||
console.error('Usage: gbrain bench publish --from <captured.ndjson> --to <baseline.ndjson> [flags]');
|
||||
console.error(' See docs/eval-bench.md for the full loop: eval export → bench publish → eval gate');
|
||||
process.exit(2);
|
||||
}
|
||||
|
||||
// v0.42.x (#2390): `gbrain eval chronicle` is deterministic — brings its own
|
||||
// in-memory PGLite, no DB/gateway. CI fixture gate runs anywhere.
|
||||
if (command === 'eval' && args[0] === 'chronicle') {
|
||||
@@ -1757,6 +1785,11 @@ async function handleCliOnly(command: string, args: string[]) {
|
||||
await runOrphans(engine, args);
|
||||
break;
|
||||
}
|
||||
case 'maintain': {
|
||||
const { runMaintain } = await import('./commands/maintain.ts');
|
||||
await runMaintain(engine, args);
|
||||
break;
|
||||
}
|
||||
// v0.32.7 CJK wave — post-upgrade markdown re-chunk sweep.
|
||||
// v0.36 Phase 3 wave — `gbrain reindex --multimodal` re-embeds content_chunks
|
||||
// into the unified Voyage multimodal-3 column.
|
||||
@@ -2217,6 +2250,22 @@ async function connectEngine(opts?: { probeOnly?: boolean }): Promise<BrainEngin
|
||||
if (merged.embedding_image_ocr_model !== undefined) {
|
||||
process.env.GBRAIN_EMBEDDING_IMAGE_OCR_MODEL = merged.embedding_image_ocr_model;
|
||||
}
|
||||
// #1475: stash the merged eval.* flags the same way. The capture gate
|
||||
// (isEvalCaptureEnabled / isEvalScrubEnabled) runs against ctx.config,
|
||||
// which is built from the sync file-plane loadConfig() in both the CLI
|
||||
// op path and MCP dispatch — it never sees the DB plane directly. The
|
||||
// gates consult this stash when the file plane is silent, so
|
||||
// `gbrain config set eval.capture true` actually turns capture on.
|
||||
// A pre-set env value wins over the DB plane (env-above-config, the
|
||||
// incident escape hatch) — unlike GBRAIN_EMBEDDING_MULTIMODAL these
|
||||
// keys have no loadConfig() env mapping, so without this guard the
|
||||
// DB stash would silently clobber an operator's export.
|
||||
if (process.env.GBRAIN_EVAL_CAPTURE === undefined && merged.eval?.capture !== undefined) {
|
||||
process.env.GBRAIN_EVAL_CAPTURE = String(merged.eval.capture);
|
||||
}
|
||||
if (process.env.GBRAIN_EVAL_SCRUB_PII === undefined && merged.eval?.scrub_pii !== undefined) {
|
||||
process.env.GBRAIN_EVAL_SCRUB_PII = String(merged.eval.scrub_pii);
|
||||
}
|
||||
// Always re-configure with merged values when DB merge succeeded. The
|
||||
// trigger used to be field-name-gated (only when embedding_multimodal_model
|
||||
// was set); that coupled the gate to the field set and would silently
|
||||
|
||||
+34
-4
@@ -581,7 +581,7 @@ async function embedPage(
|
||||
for (let j = 0; j < toEmbed.length; j++) {
|
||||
embeddingMap.set(toEmbed[j].chunk_index, embeddings[j]);
|
||||
}
|
||||
const updated: ChunkInput[] = chunks.map(c => ({
|
||||
const updated: ChunkInput[] = chunks.map(c => preserveCodeMetadata(c, {
|
||||
chunk_index: c.chunk_index,
|
||||
chunk_text: c.chunk_text,
|
||||
chunk_source: c.chunk_source,
|
||||
@@ -605,6 +605,31 @@ async function embedPage(
|
||||
slog(`${slug}: embedded ${toEmbed.length} chunks`);
|
||||
}
|
||||
|
||||
/**
|
||||
* Carry code-chunk metadata (language, symbol_name, symbol_type, line range,
|
||||
* parent scope, doc comment, qualified name) from a loaded Chunk back into a
|
||||
* ChunkInput destined for upsertChunks.
|
||||
*
|
||||
* Issue #769: every re-embed used to strip these fields, and upsertChunks
|
||||
* overwrites (does not COALESCE) the metadata columns from EXCLUDED, so
|
||||
* each pass clobbered code-def's primary index to NULL. Pulling the
|
||||
* preservation into one helper keeps the three re-embed call sites
|
||||
* (embedPage, embedAll non-stale, embedAllStale) in lock-step.
|
||||
*/
|
||||
function preserveCodeMetadata(loaded: any, base: ChunkInput): ChunkInput {
|
||||
return {
|
||||
...base,
|
||||
language: loaded.language ?? undefined,
|
||||
symbol_name: loaded.symbol_name ?? undefined,
|
||||
symbol_type: loaded.symbol_type ?? undefined,
|
||||
start_line: loaded.start_line ?? undefined,
|
||||
end_line: loaded.end_line ?? undefined,
|
||||
parent_symbol_path: loaded.parent_symbol_path ?? undefined,
|
||||
doc_comment: loaded.doc_comment ?? undefined,
|
||||
symbol_name_qualified: loaded.symbol_name_qualified ?? undefined,
|
||||
};
|
||||
}
|
||||
|
||||
async function embedAll(
|
||||
engine: BrainEngine,
|
||||
staleOnly: boolean,
|
||||
@@ -717,8 +742,10 @@ async function embedAll(
|
||||
for (let j = 0; j < toEmbed.length; j++) {
|
||||
embeddingMap.set(toEmbed[j].chunk_index, embeddings[j]);
|
||||
}
|
||||
// Preserve ALL chunks, only update embeddings for stale ones
|
||||
const updated: ChunkInput[] = chunks.map(c => ({
|
||||
// Preserve ALL chunks, only update embeddings for stale ones.
|
||||
// preserveCodeMetadata threads code-chunk metadata (#769) so re-embed
|
||||
// doesn't clobber language/symbol_name/symbol_type to NULL.
|
||||
const updated: ChunkInput[] = chunks.map(c => preserveCodeMetadata(c, {
|
||||
chunk_index: c.chunk_index,
|
||||
chunk_text: c.chunk_text,
|
||||
chunk_source: c.chunk_source,
|
||||
@@ -1012,7 +1039,10 @@ async function embedAllStale(
|
||||
for (let j = 0; j < stale.length; j++) {
|
||||
staleIdxToEmbedding.set(stale[j].chunk_index, embeddings[j]);
|
||||
}
|
||||
const merged: ChunkInput[] = existing.map(c => ({
|
||||
// preserveCodeMetadata threads code-chunk metadata (#769) so the
|
||||
// autopilot --stale path doesn't clobber language/symbol_name/etc
|
||||
// to NULL on every cycle.
|
||||
const merged: ChunkInput[] = existing.map(c => preserveCodeMetadata(c, {
|
||||
chunk_index: c.chunk_index,
|
||||
chunk_text: c.chunk_text,
|
||||
chunk_source: c.chunk_source,
|
||||
|
||||
@@ -1651,7 +1651,7 @@ async function extractTimelineFromDB(
|
||||
* make re-extraction idempotent). EVERY processed page is stamped, including
|
||||
* zero-link pages — they WERE processed.
|
||||
*/
|
||||
async function extractStaleFromDB(
|
||||
export async function extractStaleFromDB(
|
||||
engine: BrainEngine,
|
||||
opts: {
|
||||
dryRun: boolean;
|
||||
|
||||
@@ -53,13 +53,6 @@ export async function runImport(
|
||||
strategy?: SyncStrategy;
|
||||
sourceId?: string;
|
||||
managedBookmark?: boolean;
|
||||
/**
|
||||
* #2156: allow-list glob patterns — only dir-relative paths matching at
|
||||
* least one pattern are imported. Applied BEFORE `exclude`. Threaded by
|
||||
* performFullSync from `gbrain sync --include` / the source row's
|
||||
* persisted `config.include_globs`.
|
||||
*/
|
||||
include?: string[];
|
||||
/**
|
||||
* #753/#774: glob patterns to exclude from the import (same semantics as
|
||||
* `isSyncable`'s `exclude` — matched against the dir-relative path).
|
||||
@@ -222,10 +215,6 @@ export async function runImport(
|
||||
);
|
||||
const fileTypeLabel = strategy === 'code' ? 'code'
|
||||
: strategy === 'auto' ? 'syncable' : 'markdown';
|
||||
// #2156: apply --include allow-list globs first (threaded by performFullSync).
|
||||
if (opts.include && opts.include.length > 0) {
|
||||
allFiles = allFiles.filter(abs => matchesAnyGlob(relative(dir, abs), opts.include));
|
||||
}
|
||||
// #753/#774: apply --exclude glob patterns (threaded by performFullSync).
|
||||
if (opts.exclude && opts.exclude.length > 0) {
|
||||
const beforeExclude = allFiles.length;
|
||||
|
||||
+12
-172
@@ -17,7 +17,7 @@
|
||||
*/
|
||||
|
||||
import { readFileSync, writeFileSync, readdirSync, statSync, lstatSync, existsSync } from 'fs';
|
||||
import { join, relative, resolve } from 'path';
|
||||
import { join, relative } from 'path';
|
||||
import { isAborted } from '../core/abort-check.ts';
|
||||
import { parseMarkdown, type ParseValidationCode } from '../core/markdown.ts';
|
||||
import {
|
||||
@@ -26,9 +26,7 @@ import {
|
||||
DEFAULT_BYTES_WARN,
|
||||
} from '../core/content-sanity.ts';
|
||||
import { loadOperatorLiterals } from '../core/content-sanity-literals.ts';
|
||||
import { loadConfig, loadConfigWithEngine, toEngineConfig, gbrainPath } from '../core/config.ts';
|
||||
import { matchesAnyGlob } from '../core/sync.ts';
|
||||
import { parseGlobList } from './sync.ts';
|
||||
import { loadConfig, loadConfigWithEngine, gbrainPath } from '../core/config.ts';
|
||||
import type { BrainEngine } from '../core/engine.ts';
|
||||
|
||||
export interface LintIssue {
|
||||
@@ -380,89 +378,21 @@ async function resolveLintContentSanity(
|
||||
};
|
||||
}
|
||||
|
||||
/** Collect markdown files from a directory.
|
||||
*
|
||||
* When `opts.include` or `opts.exclude` are set, each candidate `.md` path's
|
||||
* POSIX-style relative path (relative to `dir`) is matched against the same
|
||||
* glob semantics sync uses (`matchesAnyGlob`). `include` allow-lists;
|
||||
* `exclude` deny-lists. Empty or undefined arrays leave the filter
|
||||
* unengaged. Symmetric with `isSyncable` in `src/core/sync.ts` so a
|
||||
* source-config `exclude_globs` honored by `gbrain sync` is also honored
|
||||
* by `gbrain lint` against the same dir.
|
||||
*/
|
||||
function collectPages(
|
||||
dir: string,
|
||||
opts: { include?: string[]; exclude?: string[] } = {},
|
||||
): string[] {
|
||||
const { include, exclude } = opts;
|
||||
const haveInclude = !!(include && include.length > 0);
|
||||
const haveExclude = !!(exclude && exclude.length > 0);
|
||||
/** Collect markdown files from a directory */
|
||||
function collectPages(dir: string): string[] {
|
||||
const pages: string[] = [];
|
||||
function walk(d: string) {
|
||||
for (const entry of readdirSync(d)) {
|
||||
if (entry.startsWith('.') || entry.startsWith('_')) continue;
|
||||
const full = join(d, entry);
|
||||
if (lstatSync(full).isDirectory()) walk(full);
|
||||
else if (entry.endsWith('.md')) {
|
||||
if (haveInclude || haveExclude) {
|
||||
// Match against the path RELATIVE to `dir` (the source root),
|
||||
// normalized to POSIX separators by matchesAnyGlob. A
|
||||
// source-config glob like `Resources/veriff/**` is anchored at
|
||||
// the source root; matching against the absolute path would
|
||||
// require the user to anchor on their `$HOME` or repo prefix,
|
||||
// which is brittle.
|
||||
const rel = relative(dir, full);
|
||||
if (haveInclude && !matchesAnyGlob(rel, include)) continue;
|
||||
if (haveExclude && matchesAnyGlob(rel, exclude)) continue;
|
||||
}
|
||||
pages.push(full);
|
||||
}
|
||||
else if (entry.endsWith('.md')) pages.push(full);
|
||||
}
|
||||
}
|
||||
walk(dir);
|
||||
return pages.sort();
|
||||
}
|
||||
|
||||
/** Look up the source row whose `local_path` resolves to the same absolute
|
||||
* directory as `target`, and return its persisted `include_globs` /
|
||||
* `exclude_globs` as parsed string arrays. Returns an empty object when no
|
||||
* matching source exists, when the row has no globs configured, or when the
|
||||
* lookup throws (best-effort — auto-resolution must never break standalone
|
||||
* lint on brains without a sources table).
|
||||
*
|
||||
* Mirrors how `syncOneSource` lifts the same fields off `src.config` before
|
||||
* threading them into `SyncOpts.include` / `SyncOpts.exclude`.
|
||||
*/
|
||||
async function resolveSourceGlobsForTarget(
|
||||
engine: BrainEngine,
|
||||
target: string,
|
||||
): Promise<{ include?: string[]; exclude?: string[] }> {
|
||||
try {
|
||||
const absTarget = resolve(target);
|
||||
const rows = await engine.executeRaw<{ config: unknown }>(
|
||||
`SELECT config FROM sources
|
||||
WHERE archived IS NOT TRUE
|
||||
AND local_path IS NOT NULL
|
||||
AND local_path = $1
|
||||
LIMIT 1`,
|
||||
[absTarget],
|
||||
);
|
||||
if (rows.length === 0) return {};
|
||||
const cfg = (rows[0].config && typeof rows[0].config === 'object')
|
||||
? rows[0].config as Record<string, unknown>
|
||||
: {};
|
||||
return {
|
||||
include: parseGlobList(cfg.include_globs),
|
||||
exclude: parseGlobList(cfg.exclude_globs),
|
||||
};
|
||||
} catch {
|
||||
// Engine not connected, sources table missing on a fresh brain, RLS
|
||||
// denial in an unusual scope — all best-effort. Lint proceeds without
|
||||
// filtering rather than fail-closed.
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
export interface LintOpts {
|
||||
target: string;
|
||||
fix?: boolean;
|
||||
@@ -484,22 +414,6 @@ export interface LintOpts {
|
||||
* yields + checks this every 200 pages.
|
||||
*/
|
||||
signal?: AbortSignal;
|
||||
/**
|
||||
* Glob filters threaded into the file walker. When set, paths relative to
|
||||
* `target` are matched against the patterns using the same semantics as
|
||||
* `gbrain sync` (`matchesAnyGlob` in `src/core/sync.ts`). `include`
|
||||
* allow-lists; `exclude` deny-lists; both unset == no filter.
|
||||
*
|
||||
* When BOTH are unset AND `engine` is provided, `runLintCore` attempts to
|
||||
* auto-resolve them from the `sources` row whose `local_path` matches
|
||||
* `target` — symmetric with `syncOneSource`, so a user who has run
|
||||
* `gbrain sources add --exclude 'Resources/veriff/**'` sees the same
|
||||
* exclusion applied to `gbrain lint <same-dir>` and to the cycle.lint
|
||||
* phase without restating it on every invocation. Explicit caller-supplied
|
||||
* arrays always win over the source-row lift.
|
||||
*/
|
||||
include?: string[];
|
||||
exclude?: string[];
|
||||
}
|
||||
|
||||
export interface LintResult {
|
||||
@@ -526,21 +440,7 @@ export async function runLintCore(opts: LintOpts): Promise<LintResult> {
|
||||
}
|
||||
|
||||
const isSingleFile = statSync(opts.target).isFile();
|
||||
|
||||
// Resolve glob filters. Explicit caller-supplied include/exclude win;
|
||||
// otherwise lift from `sources.config.{include,exclude}_globs` when an
|
||||
// engine is available and the target matches a known source's local_path.
|
||||
// Single-file lints skip the resolve entirely — globs are a directory
|
||||
// walk concern.
|
||||
let include = opts.include;
|
||||
let exclude = opts.exclude;
|
||||
const haveExplicit = (include && include.length > 0) || (exclude && exclude.length > 0);
|
||||
if (!isSingleFile && !haveExplicit && opts.engine) {
|
||||
const resolved = await resolveSourceGlobsForTarget(opts.engine, opts.target);
|
||||
include = resolved.include;
|
||||
exclude = resolved.exclude;
|
||||
}
|
||||
const pages = isSingleFile ? [opts.target] : collectPages(opts.target, { include, exclude });
|
||||
const pages = isSingleFile ? [opts.target] : collectPages(opts.target);
|
||||
|
||||
// Resolve content-sanity config once for this lint run (D1: lift DB
|
||||
// config when reachable). Caller can pre-pass via opts.contentSanity
|
||||
@@ -591,27 +491,14 @@ export async function runLintCore(opts: LintOpts): Promise<LintResult> {
|
||||
}
|
||||
|
||||
export async function runLint(args: string[]) {
|
||||
const target = args.find(a => !a.startsWith('--') && !args[args.indexOf(a) - 1]?.match(/^--(include|exclude)$/));
|
||||
const target = args.find(a => !a.startsWith('--'));
|
||||
const doFix = args.includes('--fix');
|
||||
const dryRun = args.includes('--dry-run');
|
||||
|
||||
// Parse repeatable `--include <glob>` and `--exclude <glob>` flags.
|
||||
// Symmetric with `gbrain sources add --include / --exclude` from PR #2157;
|
||||
// explicit flags here override the source-config lift performed below for
|
||||
// dir-mode lints.
|
||||
const cliInclude: string[] = [];
|
||||
const cliExclude: string[] = [];
|
||||
for (let i = 0; i < args.length; i++) {
|
||||
if (args[i] === '--include' && i + 1 < args.length) cliInclude.push(args[++i]);
|
||||
else if (args[i] === '--exclude' && i + 1 < args.length) cliExclude.push(args[++i]);
|
||||
}
|
||||
|
||||
if (!target) {
|
||||
console.error('Usage: gbrain lint <dir|file.md> [--fix] [--dry-run] [--include <glob>]... [--exclude <glob>]...');
|
||||
console.error(' --fix Auto-fix fixable issues (LLM preambles, code fences)');
|
||||
console.error(' --dry-run Preview fixes without writing');
|
||||
console.error(' --include <glob> Repeatable; only lint paths matching at least one pattern');
|
||||
console.error(' --exclude <glob> Repeatable; skip paths matching any pattern (applied after --include)');
|
||||
console.error('Usage: gbrain lint <dir|file.md> [--fix] [--dry-run]');
|
||||
console.error(' --fix Auto-fix fixable issues (LLM preambles, code fences)');
|
||||
console.error(' --dry-run Preview fixes without writing');
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
@@ -623,44 +510,7 @@ export async function runLint(args: string[]) {
|
||||
// Single file or directory — print human detail as we go, then rely on
|
||||
// Core for the aggregate numbers at the end.
|
||||
const isSingleFile = statSync(target).isFile();
|
||||
|
||||
// Resolve glob filters for directory lints. Explicit CLI flags win;
|
||||
// otherwise lift from `sources.config.{include,exclude}_globs` matching
|
||||
// `target`. Connect a transient engine for the lookup only when (a) no
|
||||
// explicit flags were passed AND (b) file/env config suggests an engine is
|
||||
// available — mirrors the connect-disconnect pattern in
|
||||
// `resolveLintContentSanity` (issue #1678: standalone CLI never shares the
|
||||
// db.ts singleton, so create + dispose here is safe).
|
||||
let runInclude: string[] | undefined = cliInclude.length > 0 ? cliInclude : undefined;
|
||||
let runExclude: string[] | undefined = cliExclude.length > 0 ? cliExclude : undefined;
|
||||
if (!isSingleFile && runInclude === undefined && runExclude === undefined) {
|
||||
const base = loadConfig();
|
||||
if (base?.database_url || base?.database_path) {
|
||||
try {
|
||||
const { createEngine } = await import('../core/engine-factory.ts');
|
||||
const { connectWithRetry } = await import('../core/db.ts');
|
||||
const engineCfg = toEngineConfig(base);
|
||||
const engine = await createEngine(engineCfg);
|
||||
try {
|
||||
// Use the same connect path the rest of the CLI uses
|
||||
// (`connectEngine` in cli.ts). `engine.connect({})` with empty
|
||||
// opts drops the URL — confirmed by direct probe. `noRetry: true`
|
||||
// keeps the standalone lint snappy (no retry tax when the brain
|
||||
// happens to be unreachable; auto-resolve degrades to no-filter).
|
||||
await connectWithRetry(engine, engineCfg, { noRetry: true });
|
||||
const lifted = await resolveSourceGlobsForTarget(engine, target);
|
||||
runInclude = lifted.include;
|
||||
runExclude = lifted.exclude;
|
||||
} finally {
|
||||
await engine.disconnect().catch(() => { /* best-effort */ });
|
||||
}
|
||||
} catch {
|
||||
// best-effort; fall through to no-filter
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
const pages = isSingleFile ? [target] : collectPages(target, { include: runInclude, exclude: runExclude });
|
||||
const pages = isSingleFile ? [target] : collectPages(target);
|
||||
|
||||
// Progress on stderr. Stdout keeps the per-issue human output it always had.
|
||||
const { createProgress } = await import('../core/progress.ts');
|
||||
@@ -707,17 +557,7 @@ export async function runLint(args: string[]) {
|
||||
// produces canonical numbers for the summary line).
|
||||
// Pass contentSanity through so runLintCore skips its own resolve
|
||||
// (we already resolved once for the human-detail loop above).
|
||||
// Pass include/exclude so the aggregate scope matches the human-detail
|
||||
// walk above — otherwise the summary line reports the unfiltered count
|
||||
// even though the per-page details were already filtered.
|
||||
const result = await runLintCore({
|
||||
target,
|
||||
fix: doFix,
|
||||
dryRun,
|
||||
contentSanity,
|
||||
include: runInclude,
|
||||
exclude: runExclude,
|
||||
});
|
||||
const result = await runLintCore({ target, fix: doFix, dryRun, contentSanity });
|
||||
console.log(`\n${result.pages_scanned} pages scanned. ${result.total_issues} issue(s) in ${result.pages_with_issues} page(s).`);
|
||||
if (doFix) {
|
||||
console.log(`${dryRun ? '(dry run) ' : ''}${result.total_fixed} auto-fixed.`);
|
||||
|
||||
@@ -0,0 +1,224 @@
|
||||
/**
|
||||
* gbrain maintain — conservative self-healing maintenance.
|
||||
*
|
||||
* This command automates the safe parts of the operator runbook:
|
||||
* - stale link/timeline extraction
|
||||
* - stale per-source dream cycles when doctor reports cycle_freshness
|
||||
*
|
||||
* It deliberately does NOT mutate source files, apply schema-pack upgrades, or
|
||||
* invent semantic hub links. Those need review or a separate command with an
|
||||
* auditable proposal surface.
|
||||
*/
|
||||
|
||||
import { existsSync } from 'fs';
|
||||
import type { BrainEngine } from '../core/engine.ts';
|
||||
import type { BrainHealth } from '../core/types.ts';
|
||||
import { buildChecks, computeDoctorReport, type DoctorReport, type Check } from './doctor.ts';
|
||||
import { extractStaleFromDB } from './extract.ts';
|
||||
import { runCycle, type CycleReport } from '../core/cycle.ts';
|
||||
|
||||
type ActionStatus = 'ok' | 'would_apply' | 'applied' | 'blocked' | 'skipped';
|
||||
|
||||
export interface MaintenanceAction {
|
||||
name: string;
|
||||
status: ActionStatus;
|
||||
message: string;
|
||||
details?: Record<string, unknown>;
|
||||
}
|
||||
|
||||
export interface MaintainOptions {
|
||||
json: boolean;
|
||||
safe: boolean;
|
||||
dryRun: boolean;
|
||||
help: boolean;
|
||||
}
|
||||
|
||||
export interface MaintainReport {
|
||||
mode: 'dry-run' | 'safe';
|
||||
before: {
|
||||
health: BrainHealth;
|
||||
doctor: DoctorReport;
|
||||
};
|
||||
actions: MaintenanceAction[];
|
||||
after: {
|
||||
health: BrainHealth;
|
||||
doctor: DoctorReport;
|
||||
};
|
||||
}
|
||||
|
||||
export function parseMaintainArgs(args: string[]): MaintainOptions {
|
||||
const safe = args.includes('--safe');
|
||||
return {
|
||||
json: args.includes('--json'),
|
||||
safe,
|
||||
dryRun: args.includes('--dry-run') || !safe,
|
||||
help: args.includes('--help') || args.includes('-h'),
|
||||
};
|
||||
}
|
||||
|
||||
export function extractCycleFreshnessSourceIds(checks: Check[]): string[] {
|
||||
const ids = new Set<string>();
|
||||
for (const check of checks) {
|
||||
if (check.name !== 'cycle_freshness' || check.status === 'ok') continue;
|
||||
const re = /Source '([^']+)' last cycled/g;
|
||||
for (const match of check.message.matchAll(re)) {
|
||||
const id = match[1]?.trim();
|
||||
if (id) ids.add(id);
|
||||
}
|
||||
}
|
||||
return [...ids].sort();
|
||||
}
|
||||
|
||||
async function buildDoctorReport(engine: BrainEngine): Promise<DoctorReport> {
|
||||
const checks = await buildChecks(engine, ['--json', '--scope=brain']);
|
||||
return computeDoctorReport(checks);
|
||||
}
|
||||
|
||||
async function runStaleExtraction(
|
||||
engine: BrainEngine,
|
||||
beforeHealth: BrainHealth,
|
||||
dryRun: boolean,
|
||||
): Promise<MaintenanceAction> {
|
||||
if (beforeHealth.stale_pages <= 0) {
|
||||
return { name: 'extract_stale', status: 'ok', message: 'No stale pages.' };
|
||||
}
|
||||
|
||||
if (dryRun) {
|
||||
return {
|
||||
name: 'extract_stale',
|
||||
status: 'would_apply',
|
||||
message: `Would run DB-backed stale extraction for ${beforeHealth.stale_pages} page(s).`,
|
||||
details: { stale_pages: beforeHealth.stale_pages },
|
||||
};
|
||||
}
|
||||
|
||||
const result = await extractStaleFromDB(engine, {
|
||||
dryRun: false,
|
||||
jsonMode: false,
|
||||
includeFrontmatter: false,
|
||||
catchUp: false,
|
||||
});
|
||||
|
||||
return {
|
||||
name: 'extract_stale',
|
||||
status: 'applied',
|
||||
message: `Processed ${result.pagesProcessed} stale page(s); ${result.staleRemaining} remain.`,
|
||||
details: {
|
||||
links_created: result.linksCreated,
|
||||
timeline_created: result.timelineCreated,
|
||||
pages_processed: result.pagesProcessed,
|
||||
stale_remaining: result.staleRemaining,
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
async function runCycleFreshnessMaintenance(
|
||||
engine: BrainEngine,
|
||||
beforeDoctor: DoctorReport,
|
||||
dryRun: boolean,
|
||||
): Promise<MaintenanceAction[]> {
|
||||
const sourceIds = extractCycleFreshnessSourceIds(beforeDoctor.checks);
|
||||
if (sourceIds.length === 0) {
|
||||
return [{ name: 'cycle_freshness', status: 'ok', message: 'All sources cycled recently.' }];
|
||||
}
|
||||
|
||||
if (dryRun) {
|
||||
return sourceIds.map((sourceId) => ({
|
||||
name: 'cycle_freshness',
|
||||
status: 'would_apply',
|
||||
message: `Would run source-scoped dream cycle for ${sourceId}.`,
|
||||
details: { source_id: sourceId },
|
||||
}));
|
||||
}
|
||||
|
||||
const sources = await engine.listAllSources();
|
||||
const actions: MaintenanceAction[] = [];
|
||||
|
||||
for (const sourceId of sourceIds) {
|
||||
const source = sources.find((s) => s.id === sourceId);
|
||||
const localPath = source?.local_path ?? null;
|
||||
const brainDir = localPath && existsSync(localPath) ? localPath : null;
|
||||
const report: CycleReport = await runCycle(engine, {
|
||||
brainDir,
|
||||
dryRun: false,
|
||||
pull: false,
|
||||
sourceId,
|
||||
});
|
||||
actions.push({
|
||||
name: 'cycle_freshness',
|
||||
status: report.status === 'failed' ? 'blocked' : 'applied',
|
||||
message: `Ran source-scoped dream cycle for ${sourceId}: ${report.status}.`,
|
||||
details: {
|
||||
source_id: sourceId,
|
||||
brain_dir: brainDir,
|
||||
cycle_status: report.status,
|
||||
phases: report.phases.map((p) => ({ phase: p.phase, status: p.status })),
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
return actions;
|
||||
}
|
||||
|
||||
export async function runMaintain(engine: BrainEngine, args: string[]): Promise<MaintainReport | void> {
|
||||
const opts = parseMaintainArgs(args);
|
||||
if (opts.help) {
|
||||
console.log(`Usage: gbrain maintain [--safe] [--dry-run] [--json]
|
||||
|
||||
Conservative self-healing maintenance.
|
||||
|
||||
Modes:
|
||||
--dry-run Preview safe actions without writes. Default when --safe is absent.
|
||||
--safe Apply safe actions: stale extraction and source cycle freshness.
|
||||
--json Emit a structured before/action/after report.
|
||||
|
||||
Not auto-applied:
|
||||
source-file frontmatter fixes, schema-pack upgrades, atom-pack changes,
|
||||
semantic hub-link guesses, and destructive cleanup.
|
||||
`);
|
||||
return;
|
||||
}
|
||||
|
||||
const beforeHealth = await engine.getHealth();
|
||||
const beforeDoctor = await buildDoctorReport(engine);
|
||||
const actions: MaintenanceAction[] = [];
|
||||
|
||||
actions.push(await runStaleExtraction(engine, beforeHealth, opts.dryRun));
|
||||
actions.push(...await runCycleFreshnessMaintenance(engine, beforeDoctor, opts.dryRun));
|
||||
|
||||
const afterHealth = await engine.getHealth();
|
||||
const afterDoctor = await buildDoctorReport(engine);
|
||||
const report: MaintainReport = {
|
||||
mode: opts.dryRun ? 'dry-run' : 'safe',
|
||||
before: { health: beforeHealth, doctor: beforeDoctor },
|
||||
actions,
|
||||
after: { health: afterHealth, doctor: afterDoctor },
|
||||
};
|
||||
|
||||
if (opts.json) {
|
||||
console.log(JSON.stringify(report, null, 2));
|
||||
} else {
|
||||
printMaintainReport(report);
|
||||
}
|
||||
return report;
|
||||
}
|
||||
|
||||
function printMaintainReport(report: MaintainReport): void {
|
||||
console.log(`GBrain maintain (${report.mode})`);
|
||||
console.log(
|
||||
`Before: brain_score=${Math.round(report.before.health.brain_score)}/100 ` +
|
||||
`stale=${report.before.health.stale_pages} islands=${report.before.health.orphan_pages} ` +
|
||||
`doctor=${report.before.doctor.status}`,
|
||||
);
|
||||
for (const action of report.actions) {
|
||||
console.log(` ${action.status}: ${action.name} — ${action.message}`);
|
||||
}
|
||||
console.log(
|
||||
`After: brain_score=${Math.round(report.after.health.brain_score)}/100 ` +
|
||||
`stale=${report.after.health.stale_pages} islands=${report.after.health.orphan_pages} ` +
|
||||
`doctor=${report.after.doctor.status}`,
|
||||
);
|
||||
if (report.mode === 'dry-run') {
|
||||
console.log('Run `gbrain maintain --safe` to apply safe actions.');
|
||||
}
|
||||
}
|
||||
+10
-55
@@ -15,6 +15,11 @@
|
||||
import type { BrainEngine } from '../core/engine.ts';
|
||||
import { createProgress, startHeartbeat } from '../core/progress.ts';
|
||||
import { getCliOptions, cliOptsToProgressOptions } from '../core/cli-options.ts';
|
||||
import {
|
||||
shouldExcludeFromOrphanReporting,
|
||||
loadOrphanPolicyOverrides,
|
||||
type OrphanPolicyOverrides,
|
||||
} from '../core/orphan-policy.ts';
|
||||
|
||||
// --- Types ---
|
||||
|
||||
@@ -32,65 +37,14 @@ export interface OrphanResult {
|
||||
excluded: number;
|
||||
}
|
||||
|
||||
// --- Filter constants ---
|
||||
|
||||
/** Slug suffixes that are always auto-generated root files */
|
||||
const AUTO_SUFFIX_PATTERNS = ['/_index', '/log'];
|
||||
|
||||
/** Page slugs that are pseudo-pages by convention */
|
||||
const PSEUDO_SLUGS = new Set(['_atlas', '_index', '_stats', '_orphans', '_scratch', 'claude']);
|
||||
|
||||
/** Slug segment that marks raw sources */
|
||||
const RAW_SEGMENT = '/raw/';
|
||||
|
||||
/** Slug prefixes where no inbound links is expected */
|
||||
const DENY_PREFIXES = [
|
||||
'output/',
|
||||
'dashboards/',
|
||||
'scripts/',
|
||||
'templates/',
|
||||
'openclaw/config/',
|
||||
];
|
||||
|
||||
/** First slug segments where no inbound links is expected */
|
||||
const FIRST_SEGMENT_EXCLUSIONS = new Set([
|
||||
'scratch',
|
||||
'thoughts',
|
||||
'catalog',
|
||||
'entities',
|
||||
'raw',
|
||||
'atoms',
|
||||
'skills',
|
||||
]);
|
||||
|
||||
// --- Filter logic ---
|
||||
|
||||
/**
|
||||
* Returns true if a slug should be excluded from orphan reporting by default.
|
||||
* These are pages where having no inbound links is expected / not a content problem.
|
||||
*/
|
||||
export function shouldExclude(slug: string): boolean {
|
||||
// Pseudo-pages (exact match)
|
||||
if (PSEUDO_SLUGS.has(slug)) return true;
|
||||
|
||||
// Auto-generated suffix patterns
|
||||
for (const suffix of AUTO_SUFFIX_PATTERNS) {
|
||||
if (slug.endsWith(suffix)) return true;
|
||||
}
|
||||
|
||||
// Raw source slugs
|
||||
if (slug.includes(RAW_SEGMENT)) return true;
|
||||
|
||||
// Deny-prefix slugs
|
||||
for (const prefix of DENY_PREFIXES) {
|
||||
if (slug.startsWith(prefix)) return true;
|
||||
}
|
||||
|
||||
// First-segment exclusions
|
||||
const firstSegment = slug.split('/')[0];
|
||||
if (FIRST_SEGMENT_EXCLUSIONS.has(firstSegment)) return true;
|
||||
|
||||
return false;
|
||||
export function shouldExclude(slug: string, overrides?: OrphanPolicyOverrides): boolean {
|
||||
return shouldExcludeFromOrphanReporting(slug, overrides);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -156,6 +110,7 @@ export async function findOrphans(
|
||||
let allOrphans: { slug: string; title: string; domain: string | null }[];
|
||||
let total: number;
|
||||
let excludedAll: number;
|
||||
const overrides = includePseudo ? undefined : await loadOrphanPolicyOverrides(engine);
|
||||
try {
|
||||
allOrphans = await engine.findOrphanPages(
|
||||
sourceIds ? { sourceIds } : sourceId ? { sourceId } : undefined,
|
||||
@@ -184,7 +139,7 @@ export async function findOrphans(
|
||||
total = liveRows.length;
|
||||
excludedAll = includePseudo
|
||||
? 0
|
||||
: liveRows.reduce((n, r) => n + (shouldExclude(r.slug) ? 1 : 0), 0);
|
||||
: liveRows.reduce((n, r) => n + (shouldExclude(r.slug, overrides) ? 1 : 0), 0);
|
||||
} finally {
|
||||
stopHb();
|
||||
progress.finish();
|
||||
@@ -192,7 +147,7 @@ export async function findOrphans(
|
||||
|
||||
const filtered = includePseudo
|
||||
? allOrphans
|
||||
: allOrphans.filter(row => !shouldExclude(row.slug));
|
||||
: allOrphans.filter(row => !shouldExclude(row.slug, overrides));
|
||||
|
||||
const orphans: OrphanPage[] = filtered.map(row => ({
|
||||
slug: row.slug,
|
||||
|
||||
+1
-34
@@ -122,8 +122,7 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
|
||||
if (!id) {
|
||||
console.error(
|
||||
'Usage: gbrain sources add <id> [--path <path> | --url <https-url>] ' +
|
||||
'[--name <display>] [--federated|--no-federated] [--clone-dir <path>] [--force] ' +
|
||||
'[--include <glob>...] [--exclude <glob>...]',
|
||||
'[--name <display>] [--federated|--no-federated] [--clone-dir <path>] [--force]',
|
||||
);
|
||||
process.exit(2);
|
||||
}
|
||||
@@ -136,12 +135,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
|
||||
let patFile: string | undefined;
|
||||
let noHarden = false;
|
||||
let force = false;
|
||||
// Repeatable. `--include 'people/**' --include 'companies/**'` accumulates.
|
||||
// Persisted into sources.config.include_globs / .exclude_globs and read at
|
||||
// sync time by commands/sync.ts so `Templates/`, `.smart-env/`, `Drafts/`
|
||||
// and other vault scaffolding can be skipped without renaming directories.
|
||||
const includeGlobs: string[] = [];
|
||||
const excludeGlobs: string[] = [];
|
||||
|
||||
for (let i = 1; i < args.length; i++) {
|
||||
const a = args[i];
|
||||
@@ -154,24 +147,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
|
||||
if (a === '--pat-file') { patFile = args[++i]; continue; }
|
||||
if (a === '--no-harden') { noHarden = true; continue; }
|
||||
if (a === '--force') { force = true; continue; }
|
||||
if (a === '--include') {
|
||||
const v = args[++i];
|
||||
if (!v || v.startsWith('--')) {
|
||||
console.error('Error: --include requires a glob argument (e.g. --include "people/**")');
|
||||
process.exit(2);
|
||||
}
|
||||
includeGlobs.push(v);
|
||||
continue;
|
||||
}
|
||||
if (a === '--exclude') {
|
||||
const v = args[++i];
|
||||
if (!v || v.startsWith('--')) {
|
||||
console.error('Error: --exclude requires a glob argument (e.g. --exclude "Templates/**")');
|
||||
process.exit(2);
|
||||
}
|
||||
excludeGlobs.push(v);
|
||||
continue;
|
||||
}
|
||||
console.error(`Unknown flag: ${a}`);
|
||||
process.exit(2);
|
||||
}
|
||||
@@ -192,8 +167,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
|
||||
federated,
|
||||
cloneDir,
|
||||
force,
|
||||
includeGlobs: includeGlobs.length > 0 ? includeGlobs : undefined,
|
||||
excludeGlobs: excludeGlobs.length > 0 ? excludeGlobs : undefined,
|
||||
});
|
||||
|
||||
// Topology A discovery: if the just-added source carries a brain-resident
|
||||
@@ -217,12 +190,6 @@ async function runAdd(engine: BrainEngine, args: string[]): Promise<void> {
|
||||
console.log(
|
||||
` federated: ${fed}${fed ? ' — appears in cross-source default search' : ' — only searched when explicitly named via --source'}`,
|
||||
);
|
||||
if (includeGlobs.length > 0) {
|
||||
console.log(` include globs: ${includeGlobs.join(', ')}`);
|
||||
}
|
||||
if (excludeGlobs.length > 0) {
|
||||
console.log(` exclude globs: ${excludeGlobs.join(', ')}`);
|
||||
}
|
||||
|
||||
// v0.42.44 — auto-harden managed clones for git durability the moment a brain
|
||||
// repo is added with a PAT. Best-effort: NEVER fail `add` if hardening fails.
|
||||
|
||||
+18
-256
@@ -1,7 +1,6 @@
|
||||
import { existsSync, readFileSync, writeFileSync, statSync, realpathSync } from 'fs';
|
||||
import { execFileSync } from 'child_process';
|
||||
import { join, relative } from 'path';
|
||||
import { createHash } from 'crypto';
|
||||
import type { BrainEngine } from '../core/engine.ts';
|
||||
import { DELETE_BATCH_SIZE } from '../core/engine-constants.ts';
|
||||
import { importFile } from '../core/import-file.ts';
|
||||
@@ -757,23 +756,12 @@ export interface SyncOpts {
|
||||
* are rejected before any git op runs.
|
||||
*/
|
||||
srcSubpath?: string;
|
||||
/**
|
||||
* #2156 — glob patterns files must match to be synced (allow-list).
|
||||
* Populated from the source row's persisted `config.include_globs`
|
||||
* (set via `gbrain sources add --include <glob>`) or the repeatable
|
||||
* `--include` CLI flag. Matched against the scope-relative path, same
|
||||
* anchoring as `exclude`. `exclude` is applied after `include`: a path
|
||||
* matching an include pattern is still rejected if it also matches an
|
||||
* exclude pattern. Empty arrays are the same as undefined (no filter).
|
||||
*/
|
||||
include?: string[];
|
||||
/**
|
||||
* #753/#774 — glob patterns for files to exclude from sync (repeatable
|
||||
* `--exclude` on the CLI; #2156: also populated from the source row's
|
||||
* persisted `config.exclude_globs`). Matched against the scope-relative
|
||||
* path in both the full-sync and incremental paths. Excluded files are
|
||||
* never imported; exclusion does NOT delete previously-imported pages
|
||||
* (conservative, matching the #1433 metafile posture).
|
||||
* `--exclude` on the CLI). Matched against the scope-relative path in both
|
||||
* the full-sync and incremental paths. Excluded files are never imported;
|
||||
* exclusion does NOT delete previously-imported pages (conservative,
|
||||
* matching the #1433 metafile posture).
|
||||
*/
|
||||
exclude?: string[];
|
||||
/**
|
||||
@@ -1165,29 +1153,6 @@ function unique<T>(items: T[]): T[] {
|
||||
// `src/core/sync-delta.ts` (re-imported below) so the inline cost estimator
|
||||
// prices detached sources through the same code the executor imports them with.
|
||||
|
||||
/**
|
||||
* Defensive parse for the JSONB-loaded `config.include_globs` / `config.exclude_globs`
|
||||
* arrays read off the sources row. The column is a free-form JSONB and could
|
||||
* contain anything — coerce to a string-only array, drop empties, and return
|
||||
* undefined when the result has no useful entries so the caller can decide
|
||||
* not to engage glob-filtering at all.
|
||||
*/
|
||||
export function parseGlobList(value: unknown): string[] | undefined {
|
||||
if (!Array.isArray(value)) return undefined;
|
||||
const globs = value.filter((v): v is string => typeof v === 'string' && v.length > 0);
|
||||
return globs.length > 0 ? globs : undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Union of CLI-supplied glob patterns (one-off, this invocation) and the
|
||||
* source row's persisted config globs (every sync). Deduped; undefined when
|
||||
* neither side has entries so `SyncOpts` stays unset and no filter engages.
|
||||
*/
|
||||
export function mergeGlobs(cli: string[], persisted: string[] | undefined): string[] | undefined {
|
||||
const merged = [...new Set([...cli, ...(persisted ?? [])])];
|
||||
return merged.length > 0 ? merged : undefined;
|
||||
}
|
||||
|
||||
// v0.18.0 Step 5: source-scoped sync state helpers. When opts.sourceId
|
||||
// is set, read/write the per-source row instead of the global config
|
||||
// keys. These wrappers centralize the branch so every read/write site
|
||||
@@ -1350,125 +1315,6 @@ async function writeChunkerVersion(
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* #2157 follow-on: detect when sources.config has shifted in a way that
|
||||
* affects which paths the walker will include this run. The "Already up
|
||||
* to date" gate at performSync's git-HEAD equality check honored chunker
|
||||
* version match but ignored config drift — a user who changes
|
||||
* `sources.config.exclude_globs` mid-life got "Already up to date" on
|
||||
* the next sync because git HEAD was unchanged, with no observable
|
||||
* effect until `gbrain sync --full`.
|
||||
*
|
||||
* Fingerprint covers exactly the walk-affecting fields that flow from
|
||||
* `sources.config` into `SyncOpts` at the syncOneSource call site:
|
||||
* `strategy`, `include_globs`, `exclude_globs`. CLI-supplied --include
|
||||
* / --exclude overrides do NOT participate — they are one-off scope
|
||||
* changes, not source state, and shouldn't invalidate the row's
|
||||
* checkpoint. (A user running `gbrain sync --exclude X` on a row whose
|
||||
* stored config has no X is intentionally narrowing this one pass; on
|
||||
* the next no-flags sync, the row config governs again.)
|
||||
*
|
||||
* Array order is normalized (alphabetical, post-defensive-parse) so
|
||||
* `["a/**", "b/**"]` and `["b/**", "a/**"]` fingerprint identically.
|
||||
* `parseGlobList` shares the same defensive coercion as the call site
|
||||
* that builds SyncOpts, so hand-edited or pre-normalization rows
|
||||
* fingerprint to the same shape the walker actually sees.
|
||||
*/
|
||||
export function computeSourceConfigFingerprint(rawConfig: unknown): string {
|
||||
const cfg = (rawConfig || {}) as {
|
||||
strategy?: unknown;
|
||||
include_globs?: unknown;
|
||||
exclude_globs?: unknown;
|
||||
};
|
||||
const canonical = JSON.stringify({
|
||||
strategy: typeof cfg.strategy === 'string' ? cfg.strategy : null,
|
||||
include_globs: (parseGlobList(cfg.include_globs) ?? []).slice().sort(),
|
||||
exclude_globs: (parseGlobList(cfg.exclude_globs) ?? []).slice().sort(),
|
||||
});
|
||||
return createHash('sha256').update(canonical).digest('hex');
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the per-source fingerprint stamp. NULL on pre-migration rows or
|
||||
* sources that have never been synced — the gate treats NULL as
|
||||
* "fingerprint unknown" and skips the invalidation check so first-time
|
||||
* post-upgrade syncs don't spuriously force-full.
|
||||
*/
|
||||
export async function readConfigFingerprint(
|
||||
engine: BrainEngine,
|
||||
sourceId: string | undefined,
|
||||
): Promise<string | null> {
|
||||
if (!sourceId) return null;
|
||||
const rows = await engine.executeRaw<{ config_fingerprint: string | null }>(
|
||||
`SELECT config_fingerprint FROM sources WHERE id = $1`,
|
||||
[sourceId],
|
||||
);
|
||||
return rows[0]?.config_fingerprint ?? null;
|
||||
}
|
||||
|
||||
export async function writeConfigFingerprint(
|
||||
engine: BrainEngine,
|
||||
sourceId: string | undefined,
|
||||
fingerprint: string,
|
||||
): Promise<void> {
|
||||
if (!sourceId) return;
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config_fingerprint = $1 WHERE id = $2`,
|
||||
[fingerprint, sourceId],
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the raw `sources.config` value for the named source. Returns an
|
||||
* empty object for missing or never-configured rows. The reader is
|
||||
* defensive about legacy double-encoded JSONB rows (`{"federated":true}`
|
||||
* stored as a JSON string scalar, the #2339 class) — the engine's
|
||||
* `r.config` may arrive as either a string or an object, and both are
|
||||
* normalized to an object before the fingerprint computation walks the
|
||||
* keys.
|
||||
*/
|
||||
async function readSourceConfig(
|
||||
engine: BrainEngine,
|
||||
sourceId: string | undefined,
|
||||
): Promise<unknown> {
|
||||
if (!sourceId) return {};
|
||||
const rows = await engine.executeRaw<{ config: unknown }>(
|
||||
`SELECT config FROM sources WHERE id = $1`,
|
||||
[sourceId],
|
||||
);
|
||||
const raw = rows[0]?.config;
|
||||
if (raw === null || raw === undefined) return {};
|
||||
if (typeof raw === 'string') {
|
||||
try { return JSON.parse(raw); } catch { return {}; }
|
||||
}
|
||||
return raw;
|
||||
}
|
||||
|
||||
/**
|
||||
* Read-hash-stamp wrapper for sync-completion sites outside the gate's
|
||||
* scope (e.g. `performFullSync`'s `advanceFull` closure, which doesn't
|
||||
* see `performSync`'s cached `currentConfigFp` because it's a separate
|
||||
* function). Reads the row's current config and stamps a fresh
|
||||
* fingerprint.
|
||||
*
|
||||
* Race note: if `sources.config` was mutated between the gate's read
|
||||
* and this stamp, the freshly-read value wins. The walker still used
|
||||
* the gate-time effective globs (already captured into `opts.include` /
|
||||
* `opts.exclude` upstream), so the stamp can drift from what was
|
||||
* actually walked. In practice mid-sync mutations are rare and the
|
||||
* NEXT sync will re-evaluate against the latest config anyway, so the
|
||||
* minor staleness is acceptable and avoids threading the gate-time
|
||||
* fingerprint through every helper signature.
|
||||
*/
|
||||
async function stampSourceConfigFingerprint(
|
||||
engine: BrainEngine,
|
||||
sourceId: string | undefined,
|
||||
): Promise<void> {
|
||||
if (!sourceId) return;
|
||||
const cfg = await readSourceConfig(engine, sourceId);
|
||||
await writeConfigFingerprint(engine, sourceId, computeSourceConfigFingerprint(cfg));
|
||||
}
|
||||
|
||||
/**
|
||||
* v0.40 Federated Sync v2: `gbrain sync trigger --source <id> [--priority high|normal|low]`
|
||||
*
|
||||
@@ -2317,25 +2163,7 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
detachedWorkingTreeManifest.deleted.length > 0 ||
|
||||
detachedWorkingTreeManifest.renamed.length > 0);
|
||||
|
||||
// #2157 follow-on: parallel gate for sources.config drift. Without
|
||||
// this, changing `sources.config.exclude_globs` (or include_globs /
|
||||
// strategy) on a synced source had no observable effect on the next
|
||||
// sync because git HEAD was unchanged — the "Already up to date"
|
||||
// branch below returned without re-walking. Mismatch path mirrors the
|
||||
// chunker_version gate exactly so both kinds of drift route through
|
||||
// the same `performFullSync` recovery.
|
||||
//
|
||||
// NULL stored fingerprint is "never stamped" (pre-v125 brain OR fresh
|
||||
// source whose first sync hasn't completed yet). Treated as
|
||||
// pass-through in the up-to-date check — first post-upgrade sync
|
||||
// stamps the column quietly so subsequent passes have a baseline.
|
||||
const storedConfigFp = await readConfigFingerprint(engine, opts.sourceId);
|
||||
const currentSourceConfig = await readSourceConfig(engine, opts.sourceId);
|
||||
const currentConfigFp = computeSourceConfigFingerprint(currentSourceConfig);
|
||||
const configMismatch = storedConfigFp !== null && storedConfigFp !== currentConfigFp;
|
||||
const configNeverStamped = storedConfigFp === null && opts.sourceId !== undefined;
|
||||
|
||||
if (lastCommit === headCommit && !versionMismatch && !versionNeverSet && !hasDetachedWorkingTreeChanges && !configMismatch) {
|
||||
if (lastCommit === headCommit && !versionMismatch && !versionNeverSet && !hasDetachedWorkingTreeChanges) {
|
||||
// v0.42.52.0 (PR #22xx): bump last_sync_at as a heartbeat on every successful
|
||||
// 0-changes sync. D4 invariant ("never advance last_commit on partial") is
|
||||
// preserved: last_sync_at is a monitoring signal (doctor sync_freshness
|
||||
@@ -2348,14 +2176,6 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
[opts.sourceId],
|
||||
);
|
||||
}
|
||||
// First post-upgrade sync on a pre-v125 brain lands here with
|
||||
// configNeverStamped=true; stamp the fingerprint so the gate has a
|
||||
// baseline for the NEXT pass. A spurious re-walk on the upgrade
|
||||
// pass would surprise users; quietly establishing the baseline does
|
||||
// not.
|
||||
if (configNeverStamped) {
|
||||
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
|
||||
}
|
||||
return {
|
||||
status: 'up_to_date',
|
||||
fromCommit: lastCommit,
|
||||
@@ -2367,21 +2187,13 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
};
|
||||
}
|
||||
|
||||
if ((versionMismatch || versionNeverSet || configMismatch) && lastCommit === headCommit) {
|
||||
const reasons: string[] = [];
|
||||
if (versionMismatch || versionNeverSet) {
|
||||
reasons.push(`chunker_version=${storedVersion ?? 'unset'}→${currentVersion}`);
|
||||
}
|
||||
if (configMismatch) {
|
||||
reasons.push(`config_fingerprint=${storedConfigFp?.slice(0, 8)}→${currentConfigFp.slice(0, 8)}`);
|
||||
}
|
||||
if ((versionMismatch || versionNeverSet) && lastCommit === headCommit) {
|
||||
slog(
|
||||
`[sync] full re-walk forced (${reasons.join(', ')}): ` +
|
||||
`git HEAD unchanged but a walk-affecting setting advanced.`,
|
||||
`[sync] chunker_version gate: stored=${storedVersion ?? 'unset'}, current=${currentVersion}. ` +
|
||||
`Forcing full re-chunk pass (git HEAD unchanged but pipeline version advanced).`,
|
||||
);
|
||||
const result = await performFullSync(engine, fullSyncRoots, headCommit, opts);
|
||||
await writeChunkerVersion(engine, opts.sourceId, currentVersion);
|
||||
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
|
||||
return result;
|
||||
}
|
||||
|
||||
@@ -2425,16 +2237,8 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
scoped && p.startsWith(syncScopeRelPath + '/') ? p.slice(syncScopeRelPath.length + 1) : p;
|
||||
const excluded = (p: string): boolean =>
|
||||
opts.exclude !== undefined && opts.exclude.length > 0 && matchesAnyGlob(scopeRel(p), opts.exclude);
|
||||
// #2156: include globs are an allow-list, same scope-relative anchoring as
|
||||
// exclude. Populated from the source row's persisted config.include_globs
|
||||
// (or CLI --include). Deliberately NOT threaded into syncOpts/isSyncable:
|
||||
// the unsyncable-cleanup loop below deletes pages for non-metafile
|
||||
// classifications, and glob filtering must stay conservative (never delete
|
||||
// previously-imported pages — the documented #1433 posture for --exclude).
|
||||
const included = (p: string): boolean =>
|
||||
opts.include === undefined || opts.include.length === 0 || matchesAnyGlob(scopeRel(p), opts.include);
|
||||
|
||||
// Filter to syncable files (strategy-aware + scope-aware + glob-aware)
|
||||
// Filter to syncable files (strategy-aware + scope-aware + exclude-aware)
|
||||
const syncOpts = opts.strategy ? { strategy: opts.strategy } : undefined;
|
||||
// #1970 (F-C): a rename whose DESTINATION is unsyncable drops out of BOTH
|
||||
// `renamed` (only `r.to` is kept below) AND `deleted` (git emits it as `R`,
|
||||
@@ -2448,13 +2252,13 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
!(inScope(r.to) && isSyncable(r.to, syncOpts)))
|
||||
.map(r => r.from);
|
||||
const filtered: SyncManifest = {
|
||||
added: manifest.added.filter(p => inScope(p) && included(p) && !excluded(p) && isSyncable(p, syncOpts)),
|
||||
modified: manifest.modified.filter(p => inScope(p) && included(p) && !excluded(p) && isSyncable(p, syncOpts)),
|
||||
added: manifest.added.filter(p => inScope(p) && !excluded(p) && isSyncable(p, syncOpts)),
|
||||
modified: manifest.modified.filter(p => inScope(p) && !excluded(p) && isSyncable(p, syncOpts)),
|
||||
deleted: unique([
|
||||
...manifest.deleted.filter(p => inScope(p) && isSyncable(p, syncOpts)),
|
||||
...renamedToUnsyncable,
|
||||
]),
|
||||
renamed: manifest.renamed.filter(r => inScope(r.to) && included(r.to) && !excluded(r.to) && isSyncable(r.to, syncOpts)),
|
||||
renamed: manifest.renamed.filter(r => inScope(r.to) && !excluded(r.to) && isSyncable(r.to, syncOpts)),
|
||||
};
|
||||
|
||||
// NAV-4: warn when --exclude filtered out every candidate change — almost
|
||||
@@ -2551,7 +2355,6 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
await writeSyncAnchor(engine, opts.sourceId, 'last_commit', pin, commitTimeMs(gitContextRoot, pin));
|
||||
await engine.setConfig('sync.last_run', new Date().toISOString());
|
||||
await writeChunkerVersion(engine, opts.sourceId, String(CHUNKER_VERSION));
|
||||
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
|
||||
await clearOpCheckpoint(engine, ckpt.paths);
|
||||
await clearOpCheckpoint(engine, ckpt.target);
|
||||
return {
|
||||
@@ -3376,7 +3179,6 @@ async function performSyncInner(engine: BrainEngine, opts: SyncOpts): Promise<Sy
|
||||
await engine.setConfig('sync.last_run', new Date().toISOString());
|
||||
await writeSyncAnchor(engine, opts.sourceId, 'repo_path', anchorPath);
|
||||
await writeChunkerVersion(engine, opts.sourceId, String(CHUNKER_VERSION));
|
||||
await writeConfigFingerprint(engine, opts.sourceId, currentConfigFp);
|
||||
await clearOpCheckpoint(engine, ckpt.paths);
|
||||
await clearOpCheckpoint(engine, ckpt.target);
|
||||
};
|
||||
@@ -3628,9 +3430,6 @@ async function performFullSync(
|
||||
// files were waiting.
|
||||
if (opts.dryRun) {
|
||||
let allFiles = collectSyncableFiles(syncScopeRoot, { strategy: opts.strategy ?? 'markdown' });
|
||||
if (opts.include && opts.include.length > 0) {
|
||||
allFiles = allFiles.filter(abs => matchesAnyGlob(relative(syncScopeRoot, abs), opts.include));
|
||||
}
|
||||
if (opts.exclude && opts.exclude.length > 0) {
|
||||
allFiles = allFiles.filter(abs => !matchesAnyGlob(relative(syncScopeRoot, abs), opts.exclude));
|
||||
}
|
||||
@@ -3676,7 +3475,6 @@ async function performFullSync(
|
||||
commit: headCommit,
|
||||
strategy: opts.strategy,
|
||||
sourceId: opts.sourceId,
|
||||
include: opts.include,
|
||||
exclude: opts.exclude,
|
||||
slugRoot,
|
||||
// issue #1939: performFullSync owns the failure ledger + bookmark via the
|
||||
@@ -3706,7 +3504,6 @@ async function performFullSync(
|
||||
await engine.setConfig('sync.last_run', new Date().toISOString());
|
||||
await writeSyncAnchor(engine, opts.sourceId, 'repo_path', anchorPath);
|
||||
await writeChunkerVersion(engine, opts.sourceId, String(CHUNKER_VERSION));
|
||||
await stampSourceConfigFingerprint(engine, opts.sourceId);
|
||||
};
|
||||
|
||||
const fullGate = await applySyncFailureGate({
|
||||
@@ -4170,12 +3967,8 @@ Options:
|
||||
run at the repo root; imports are scoped to the subdir
|
||||
and slugs stay root-relative (wiki/page1). Passing the
|
||||
subdirectory directly as --repo also works.
|
||||
--include <glob> Only sync files matching at least one glob (repeatable;
|
||||
matched against the scope-relative path). Merged with
|
||||
the source's persisted config.include_globs.
|
||||
--exclude <glob> Exclude files matching the glob from sync (repeatable;
|
||||
matched against the scope-relative path; applied after
|
||||
--include). Merged with config.exclude_globs.
|
||||
matched against the scope-relative path).
|
||||
--dry-run Show what would be synced without writing.
|
||||
--skip-failed Acknowledge previously-recorded sync failures so
|
||||
the bookmark can advance past unparseable files.
|
||||
@@ -4336,17 +4129,14 @@ See also:
|
||||
}
|
||||
const strategyArg = args.find((a, i) => args[i - 1] === '--strategy') as SyncOpts['strategy'] | undefined;
|
||||
// #753/#774: monorepo subdir-source flags. --exclude is repeatable.
|
||||
// #2156: --include is the allow-list counterpart, same repeatable shape.
|
||||
const srcSubpath = args.find((a, i) => args[i - 1] === '--src-subpath') || undefined;
|
||||
const excludePatterns: string[] = [];
|
||||
const includePatterns: string[] = [];
|
||||
for (let i = 0; i < args.length; i++) {
|
||||
if (args[i] === '--exclude' && i + 1 < args.length) excludePatterns.push(args[i + 1]);
|
||||
if (args[i] === '--include' && i + 1 < args.length) includePatterns.push(args[i + 1]);
|
||||
}
|
||||
if (syncAll && (srcSubpath || excludePatterns.length > 0 || includePatterns.length > 0)) {
|
||||
if (syncAll && (srcSubpath || excludePatterns.length > 0)) {
|
||||
console.error(
|
||||
`--src-subpath/--include/--exclude scope a single sync invocation; they cannot be combined with --all. ` +
|
||||
`--src-subpath/--exclude scope a single sync invocation; they cannot be combined with --all. ` +
|
||||
`For --all runs, register the subdirectory as the source's local_path instead ` +
|
||||
`(gbrain sources add <id> --path <repo>/<subdir>).`,
|
||||
);
|
||||
@@ -4547,11 +4337,7 @@ See also:
|
||||
const onAllSigint = () => { try { allInterrupt.abort(new Error('SIGINT')); } catch { /* */ } };
|
||||
|
||||
const runOne = async (src: typeof sources[number]): Promise<SyncResult> => {
|
||||
const cfg = (src.config || {}) as {
|
||||
strategy?: 'markdown' | 'code' | 'auto';
|
||||
include_globs?: unknown;
|
||||
exclude_globs?: unknown;
|
||||
};
|
||||
const cfg = (src.config || {}) as { strategy?: 'markdown' | 'code' | 'auto' };
|
||||
// D18: parallel path defers embed; auto-enqueue embed-backfill after.
|
||||
// v0.42.42.0 (#2139): `autoDeferEmbeds` (the inline gate tripped in a
|
||||
// non-TTY session) ALSO forces deferral — global by design (the gate's
|
||||
@@ -4589,8 +4375,6 @@ See also:
|
||||
skipFailed, retryFailed, noSchemaPack,
|
||||
sourceId: src.id,
|
||||
strategy: cfg.strategy,
|
||||
include: parseGlobList(cfg.include_globs),
|
||||
exclude: parseGlobList(cfg.exclude_globs),
|
||||
concurrency,
|
||||
signal: composeAbortSignals(allInterrupt.signal, controller?.signal),
|
||||
};
|
||||
@@ -4802,27 +4586,11 @@ See also:
|
||||
// lock released by its own finally) instead of a hard cut.
|
||||
const singleSourceInterrupt = new AbortController();
|
||||
const onSingleSourceSigint = () => { try { singleSourceInterrupt.abort(new Error('SIGINT')); } catch { /* */ } };
|
||||
// Read persisted include/exclude globs from the source row, mirroring the
|
||||
// --all fan-out's `runOne` closure above. Best-effort: a fetch failure
|
||||
// falls through to "no glob filters", preserving pre-existing behavior.
|
||||
// sourceId is always set here (resolveSourceWithTier ran above), so this
|
||||
// path never silently runs without source-config awareness.
|
||||
let sourceCfg: { include_globs?: unknown; exclude_globs?: unknown } = {};
|
||||
try {
|
||||
const { fetchSource } = await import('../core/sources-load.ts');
|
||||
const src = await fetchSource(engine, sourceId);
|
||||
if (src?.config && typeof src.config === 'object') {
|
||||
sourceCfg = src.config as { include_globs?: unknown; exclude_globs?: unknown };
|
||||
}
|
||||
} catch { /* fall through to no filters */ }
|
||||
const opts: SyncOpts = {
|
||||
repoPath, dryRun, full, noPull, noEmbed, noExtract, skipFailed, retryFailed, noSchemaPack, sourceId,
|
||||
strategy: strategyArg, concurrency,
|
||||
srcSubpath,
|
||||
// #2156: union of the repeatable CLI flags (one-off, this invocation
|
||||
// only) and the source row's persisted config globs (every sync).
|
||||
include: mergeGlobs(includePatterns, parseGlobList(sourceCfg.include_globs)),
|
||||
exclude: mergeGlobs(excludePatterns, parseGlobList(sourceCfg.exclude_globs)),
|
||||
exclude: excludePatterns.length > 0 ? excludePatterns : undefined,
|
||||
signal: composeAbortSignals(singleSourceInterrupt.signal, singleSourceController?.signal),
|
||||
};
|
||||
|
||||
@@ -5050,11 +4818,7 @@ export async function syncOneSource(
|
||||
noExtract?: boolean;
|
||||
},
|
||||
): Promise<{ result: SyncResult; log: string }> {
|
||||
const cfg = (src.config || {}) as {
|
||||
strategy?: 'markdown' | 'code' | 'auto';
|
||||
include_globs?: unknown;
|
||||
exclude_globs?: unknown;
|
||||
};
|
||||
const cfg = (src.config || {}) as { strategy?: 'markdown' | 'code' | 'auto' };
|
||||
const log = `\n--- Syncing source: ${src.name} ---\n`;
|
||||
const repoOpts: SyncOpts = {
|
||||
repoPath: src.local_path!,
|
||||
@@ -5068,8 +4832,6 @@ export async function syncOneSource(
|
||||
noSchemaPack: shared.noSchemaPack,
|
||||
sourceId: src.id,
|
||||
strategy: cfg.strategy,
|
||||
include: parseGlobList(cfg.include_globs),
|
||||
exclude: parseGlobList(cfg.exclude_globs),
|
||||
concurrency: shared.concurrency,
|
||||
// lockId defaults to `gbrain-sync:${src.id}` via the invariant in
|
||||
// performSync (no explicit override needed — sourceId triggers it).
|
||||
|
||||
@@ -816,6 +816,25 @@ export async function loadConfigWithEngine(
|
||||
merged.dream = mergedDream;
|
||||
}
|
||||
|
||||
// #1475: eval.* DB-plane merge. `gbrain config set eval.capture true`
|
||||
// writes the DB plane (both keys are in KNOWN_CONFIG_KEYS, so `set`
|
||||
// accepts them silently), but the capture gate (isEvalCaptureEnabled)
|
||||
// reads the merged config. Without this merge the DB value was written
|
||||
// and never read — capture only fired via GBRAIN_CONTRIBUTOR_MODE=1.
|
||||
// Sparse per-key merge: file/env wins per key, DB fills the gaps.
|
||||
const dbEvalCapture = await dbBool('eval.capture');
|
||||
const dbEvalScrub = await dbBool('eval.scrub_pii');
|
||||
const mergedEval: NonNullable<GBrainConfig['eval']> = { ...(merged.eval ?? {}) };
|
||||
if (mergedEval.capture === undefined && dbEvalCapture !== undefined) {
|
||||
mergedEval.capture = dbEvalCapture;
|
||||
}
|
||||
if (mergedEval.scrub_pii === undefined && dbEvalScrub !== undefined) {
|
||||
mergedEval.scrub_pii = dbEvalScrub;
|
||||
}
|
||||
if (Object.keys(mergedEval).length > 0) {
|
||||
merged.eval = mergedEval;
|
||||
}
|
||||
|
||||
return merged;
|
||||
}
|
||||
|
||||
|
||||
@@ -54,6 +54,7 @@ import {
|
||||
import {
|
||||
generatePerChunkSynopsis,
|
||||
SYNOPSIS_PROMPT_VERSION,
|
||||
SYNOPSIS_DOC_MAX_CHARS,
|
||||
type GeneratePerChunkSynopsisResult,
|
||||
} from './page-summary.ts';
|
||||
import {
|
||||
@@ -103,8 +104,17 @@ function getEmbeddingModelTag(): string {
|
||||
export function computeCorpusGeneration(args: {
|
||||
crMode: CRMode;
|
||||
haikuModel: string;
|
||||
/**
|
||||
* Resolved `SYNOPSIS_DOC_MAX_CHARS` for per_chunk_synopsis runs. When
|
||||
* present, folded into the hash so changes to
|
||||
* `GBRAIN_SYNOPSIS_DOC_MAX_CHARS` invalidate the prior cache cleanly.
|
||||
* Omit for `crMode !== 'per_chunk_synopsis'` — title / none modes
|
||||
* don't consult the cap and the field stays out of the hash for
|
||||
* back-compat with pre-cap embeddings.
|
||||
*/
|
||||
synopsisDocMaxChars?: number;
|
||||
}): string {
|
||||
return createHash('sha256')
|
||||
const h = createHash('sha256')
|
||||
.update(args.crMode)
|
||||
.update('|')
|
||||
.update(String(SYNOPSIS_PROMPT_VERSION))
|
||||
@@ -113,9 +123,11 @@ export function computeCorpusGeneration(args: {
|
||||
.update('|')
|
||||
.update(String(TITLE_WRAPPER_VERSION))
|
||||
.update('|')
|
||||
.update(getEmbeddingModelTag())
|
||||
.digest('hex')
|
||||
.slice(0, 16);
|
||||
.update(getEmbeddingModelTag());
|
||||
if (args.synopsisDocMaxChars !== undefined) {
|
||||
h.update('|doc_cap=').update(String(args.synopsisDocMaxChars));
|
||||
}
|
||||
return h.digest('hex').slice(0, 16);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -253,7 +265,11 @@ export async function reembedPageWithContextualRetrieval(
|
||||
args.pageSlug,
|
||||
args.sourceId,
|
||||
resolution.mode,
|
||||
computeCorpusGeneration({ crMode: resolution.mode, haikuModel: args.haikuModel ?? DEFAULT_HAIKU_MODEL }),
|
||||
computeCorpusGeneration({
|
||||
crMode: resolution.mode,
|
||||
haikuModel: args.haikuModel ?? DEFAULT_HAIKU_MODEL,
|
||||
synopsisDocMaxChars: resolution.mode === 'per_chunk_synopsis' ? SYNOPSIS_DOC_MAX_CHARS : undefined,
|
||||
}),
|
||||
);
|
||||
return { kind: 'skipped', reason: 'no_chunks' };
|
||||
}
|
||||
@@ -282,6 +298,7 @@ export async function reembedPageWithContextualRetrieval(
|
||||
const corpus_generation = computeCorpusGeneration({
|
||||
crMode: attemptMode,
|
||||
haikuModel,
|
||||
synopsisDocMaxChars: attemptMode === 'per_chunk_synopsis' ? SYNOPSIS_DOC_MAX_CHARS : undefined,
|
||||
});
|
||||
|
||||
// ── PHASE 2: single DB transaction ───────────────────────────
|
||||
|
||||
@@ -251,6 +251,14 @@ registerBackgroundWorkDrainer({
|
||||
export function isEvalCaptureEnabled(config: GBrainConfig | null | undefined): boolean {
|
||||
if (config?.eval?.capture === true) return true;
|
||||
if (config?.eval?.capture === false) return false;
|
||||
// #1475: DB-plane stash. `gbrain config set eval.capture true` lands in the
|
||||
// config table; connectEngine stamps the merged value here because
|
||||
// ctx.config is the sync file-plane load and never sees the DB plane.
|
||||
// Explicit per-key setting (file above, DB here) beats the broad
|
||||
// CONTRIBUTOR_MODE flag, matching how file-plane `false` already wins.
|
||||
// Doubles as a direct operator env knob.
|
||||
if (process.env.GBRAIN_EVAL_CAPTURE === 'true') return true;
|
||||
if (process.env.GBRAIN_EVAL_CAPTURE === 'false') return false;
|
||||
return process.env.GBRAIN_CONTRIBUTOR_MODE === '1';
|
||||
}
|
||||
|
||||
@@ -263,5 +271,8 @@ export function isEvalCaptureEnabled(config: GBrainConfig | null | undefined): b
|
||||
* have explicit `capture: true`.
|
||||
*/
|
||||
export function isEvalScrubEnabled(config: GBrainConfig | null | undefined): boolean {
|
||||
return config?.eval?.scrub_pii !== false;
|
||||
if (config?.eval?.scrub_pii === false) return false;
|
||||
if (config?.eval?.scrub_pii === true) return true;
|
||||
// #1475: DB-plane stash — see isEvalCaptureEnabled. Default stays true.
|
||||
return process.env.GBRAIN_EVAL_SCRUB_PII !== 'false';
|
||||
}
|
||||
|
||||
@@ -733,6 +733,11 @@ export async function importFromContent(
|
||||
: computeCorpusGeneration({
|
||||
crMode: effectiveCRMode,
|
||||
haikuModel: 'anthropic:claude-haiku-4-5-20251001',
|
||||
// Inline import-file path never uses per_chunk_synopsis (refuses
|
||||
// upstream); pass undefined so the doc-cap field stays out of
|
||||
// the hash here. Per_chunk_synopsis runs through the Minion
|
||||
// backfill handler which threads SYNOPSIS_DOC_MAX_CHARS through
|
||||
// the service layer.
|
||||
});
|
||||
|
||||
// Transaction wraps all DB writes. Every per-page tx call carries the
|
||||
|
||||
@@ -489,7 +489,22 @@ export async function extractPageLinks(
|
||||
// text inside `[[...]]` before any `|`), NOT the display alias
|
||||
// (ref.name = match[2]). `[[struktura|the project]]` must resolve
|
||||
// `struktura`, not "the project". The display text is for context only.
|
||||
const matches = await resolver.resolveBasenameMatches(ref.slug);
|
||||
//
|
||||
// The literal may be path-qualified (`[[notes/struktura]]`). The FS
|
||||
// path (resolveSlugAll) strips the dirname before its basename lookup,
|
||||
// but this path passed the raw literal to an index keyed by final
|
||||
// segments only — so every slash-containing wikilink outside
|
||||
// DIR_PATTERN silently resolved to nothing. Query by the final
|
||||
// segment, then use the written path as a disambiguation filter
|
||||
// (the analogue of the FS ancestor walk honoring the written path):
|
||||
// a match must end with the literal, so `[[notes/struktura]]` can
|
||||
// resolve to `vault/notes/struktura` but never to `wiki/struktura`.
|
||||
const slashIdx = ref.slug.lastIndexOf('/');
|
||||
const basename = slashIdx === -1 ? ref.slug : ref.slug.slice(slashIdx + 1);
|
||||
let matches = await resolver.resolveBasenameMatches(basename);
|
||||
if (slashIdx !== -1) {
|
||||
matches = matches.filter(m => m === ref.slug || m.endsWith(`/${ref.slug}`));
|
||||
}
|
||||
if (matches.length === 0) continue;
|
||||
const idx = content.indexOf(ref.slug);
|
||||
const context = idx >= 0 ? excerpt(content, idx, 240) : ref.name;
|
||||
|
||||
@@ -5671,32 +5671,6 @@ export const MIGRATIONS: Migration[] = [
|
||||
`);
|
||||
},
|
||||
},
|
||||
{
|
||||
version: 125,
|
||||
name: 'sources_config_fingerprint',
|
||||
// #2157 follow-on: the "Already up to date" gate at sync.ts honors
|
||||
// git-HEAD equality + chunker-version match but ignored source-config
|
||||
// drift. A user who runs `gbrain sources add default --exclude
|
||||
// 'Templates/**'` AFTER an initial sync got "Already up to date" on
|
||||
// the next pass because git HEAD was unchanged — the new exclusion
|
||||
// never reached the walk until `gbrain sync --full`.
|
||||
//
|
||||
// This column caches a SHA-256 fingerprint of the walk-affecting
|
||||
// fields in `sources.config` (strategy + include_globs +
|
||||
// exclude_globs); mismatches trigger a full re-walk via the same code
|
||||
// path as a chunker_version bump.
|
||||
//
|
||||
// NULL on pre-migration rows is treated as "not yet stamped" by
|
||||
// readConfigFingerprint, so the FIRST sync after upgrade is normal
|
||||
// (no spurious force-full just because the column was added).
|
||||
//
|
||||
// Keep in sync with src/schema.sql and src/core/schema-embedded.ts.
|
||||
idempotent: true,
|
||||
sql: `
|
||||
ALTER TABLE sources
|
||||
ADD COLUMN IF NOT EXISTS config_fingerprint TEXT;
|
||||
`,
|
||||
},
|
||||
];
|
||||
|
||||
export const LATEST_VERSION = MIGRATIONS.length > 0
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
/**
|
||||
* Shared orphan-reporting exclusion policy.
|
||||
*
|
||||
* These are pages where "no inbound links" is expected and should not count
|
||||
* against health. Keep this in core so the CLI orphan report and engine health
|
||||
* dashboard cannot drift.
|
||||
*
|
||||
* Defaults are GBrain-wide conventions only. Brain-specific exclusions
|
||||
* (private folder names, one-off fixture slugs) belong in the brain's own
|
||||
* config, not here:
|
||||
*
|
||||
* gbrain config set orphans.exclude_prefixes "my-private-folder/,archive/"
|
||||
* gbrain config set orphans.exclude_slugs "some-one-off-page"
|
||||
*/
|
||||
|
||||
const AUTO_SUFFIX_PATTERNS = ['/_index', '/log'];
|
||||
|
||||
const PSEUDO_SLUGS = new Set(['_atlas', '_index', '_stats', '_orphans', '_scratch', 'claude']);
|
||||
|
||||
const RAW_SEGMENT = '/raw/';
|
||||
|
||||
const DENY_PREFIXES = [
|
||||
'output/',
|
||||
'dashboards/',
|
||||
'scripts/',
|
||||
'templates/',
|
||||
'_templates/',
|
||||
'openclaw/config/',
|
||||
'extracts/',
|
||||
];
|
||||
|
||||
const FIRST_SEGMENT_EXCLUSIONS = new Set([
|
||||
'scratch',
|
||||
'thoughts',
|
||||
'catalog',
|
||||
'entities',
|
||||
'raw',
|
||||
'atoms',
|
||||
'skills',
|
||||
'dreaming',
|
||||
'daily',
|
||||
]);
|
||||
|
||||
const ROOT_DATE_SLUG = /^\d{4}-\d{2}-\d{2}(?:-.+)?$/;
|
||||
|
||||
function isAgentWorkspaceConvention(slug: string): boolean {
|
||||
if (!slug.startsWith('agents/')) return false;
|
||||
if (slug.includes('/memory/dreaming/')) return true;
|
||||
return /^agents\/[^/]+\/(?:agents|identity|soul|tools|user|heartbeat|dreams|dormant)$/.test(slug);
|
||||
}
|
||||
|
||||
/** Per-brain additions to the convention defaults (from config). */
|
||||
export interface OrphanPolicyOverrides {
|
||||
excludePrefixes?: string[];
|
||||
excludeSlugs?: string[];
|
||||
}
|
||||
|
||||
/** Config keys for per-brain orphan exclusions (comma-separated values). */
|
||||
export const ORPHAN_EXCLUDE_PREFIXES_KEY = 'orphans.exclude_prefixes';
|
||||
export const ORPHAN_EXCLUDE_SLUGS_KEY = 'orphans.exclude_slugs';
|
||||
|
||||
function parseList(value: string | null): string[] {
|
||||
if (!value) return [];
|
||||
return value.split(',').map(s => s.trim()).filter(Boolean);
|
||||
}
|
||||
|
||||
/**
|
||||
* Load per-brain orphan exclusions from the brain config table. Callers with
|
||||
* an engine in hand (getHealth, `gbrain orphans`) pass the result as the
|
||||
* second argument to shouldExcludeFromOrphanReporting.
|
||||
*/
|
||||
export async function loadOrphanPolicyOverrides(
|
||||
engine: { getConfig(key: string): Promise<string | null> },
|
||||
): Promise<OrphanPolicyOverrides> {
|
||||
const [prefixes, slugs] = await Promise.all([
|
||||
engine.getConfig(ORPHAN_EXCLUDE_PREFIXES_KEY),
|
||||
engine.getConfig(ORPHAN_EXCLUDE_SLUGS_KEY),
|
||||
]);
|
||||
return { excludePrefixes: parseList(prefixes), excludeSlugs: parseList(slugs) };
|
||||
}
|
||||
|
||||
export function shouldExcludeFromOrphanReporting(
|
||||
slug: string,
|
||||
overrides?: OrphanPolicyOverrides,
|
||||
): boolean {
|
||||
if (PSEUDO_SLUGS.has(slug)) return true;
|
||||
|
||||
for (const suffix of AUTO_SUFFIX_PATTERNS) {
|
||||
if (slug.endsWith(suffix)) return true;
|
||||
}
|
||||
|
||||
if (slug.includes(RAW_SEGMENT)) return true;
|
||||
if (slug.includes('/daily/')) return true;
|
||||
|
||||
for (const prefix of DENY_PREFIXES) {
|
||||
if (slug.startsWith(prefix)) return true;
|
||||
}
|
||||
|
||||
const firstSegment = slug.split('/')[0];
|
||||
if (FIRST_SEGMENT_EXCLUSIONS.has(firstSegment)) return true;
|
||||
|
||||
if (ROOT_DATE_SLUG.test(slug)) return true;
|
||||
|
||||
if (slug.startsWith('_brain-')) return true;
|
||||
|
||||
if (isAgentWorkspaceConvention(slug)) return true;
|
||||
|
||||
if (overrides) {
|
||||
if (overrides.excludeSlugs?.includes(slug)) return true;
|
||||
for (const prefix of overrides.excludePrefixes ?? []) {
|
||||
if (slug.startsWith(prefix)) return true;
|
||||
}
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
@@ -44,6 +44,33 @@ const HAIKU_MAX_TOKENS = 200;
|
||||
/** Default model when caller doesn't override. Resolves through the gateway. */
|
||||
const DEFAULT_SYNOPSIS_MODEL = 'anthropic:claude-haiku-4-5-20251001';
|
||||
|
||||
/**
|
||||
* Hard cap on `documentText` length (chars) before send.
|
||||
*
|
||||
* 2026-05-25 fix wave: small local chat models (Gemma 4 E2B, Qwen3 4B) get
|
||||
* dramatically slower on long contexts even with 131K-token windows declared.
|
||||
* A 73K-char page synopsis on Gemma 4 E2B takes 60-120s, exceeding the
|
||||
* worker's default 30s `lockDuration` and tripping `lock-lost` errors.
|
||||
*
|
||||
* Truncate to a budget that fits a small model's effective throughput while
|
||||
* preserving enough document context for the synopsis to be useful. Truncates
|
||||
* the TAIL because the head (title, frontmatter, intro) carries the
|
||||
* document-level anchor the synopsis needs.
|
||||
*
|
||||
* Override per workload via `GBRAIN_SYNOPSIS_DOC_MAX_CHARS`. Default 32768
|
||||
* (~8K tokens at 4 chars/tok) keeps small-model synopsis under ~30s.
|
||||
* Anthropic Haiku is unaffected at this cap; bump higher when running
|
||||
* frontier models if you want richer document anchoring.
|
||||
*/
|
||||
export const SYNOPSIS_DOC_MAX_CHARS = (() => {
|
||||
const env = process.env.GBRAIN_SYNOPSIS_DOC_MAX_CHARS;
|
||||
if (env && /^\d+$/.test(env)) {
|
||||
const n = parseInt(env, 10);
|
||||
if (n >= 512 && n <= 1_048_576) return n;
|
||||
}
|
||||
return 32768;
|
||||
})();
|
||||
|
||||
/**
|
||||
* Synopsis prompt version. Folded into corpus_generation so prompt edits
|
||||
* invalidate prior embeddings via the v0.40.3.0 query_cache.page_generations
|
||||
@@ -188,11 +215,19 @@ function buildUserPrompt(
|
||||
documentText: string,
|
||||
chunkText: string,
|
||||
): string {
|
||||
// Tail-truncate `documentText` to `SYNOPSIS_DOC_MAX_CHARS` so small local
|
||||
// chat models don't stall on >100KB pages. Head preserved (title block,
|
||||
// frontmatter, intro paragraphs carry the document-level anchor).
|
||||
let trimmedDoc = documentText;
|
||||
if (documentText.length > SYNOPSIS_DOC_MAX_CHARS) {
|
||||
trimmedDoc = documentText.slice(0, SYNOPSIS_DOC_MAX_CHARS) +
|
||||
`\n\n[... ${documentText.length - SYNOPSIS_DOC_MAX_CHARS} chars truncated for synopsis budget ...]`;
|
||||
}
|
||||
return [
|
||||
`<page_title>${pageTitle}</page_title>`,
|
||||
'',
|
||||
'<full_document>',
|
||||
documentText,
|
||||
trimmedDoc,
|
||||
'</full_document>',
|
||||
'',
|
||||
'<chunk>',
|
||||
|
||||
+30
-19
@@ -57,6 +57,8 @@ import { finalizeLastSeen } from './chronicle/last-seen.ts';
|
||||
import { computeAnomaliesFromBuckets } from './cycle/anomaly.ts';
|
||||
import { resolveBoostMap, resolveHardExcludes } from './search/source-boost.ts';
|
||||
import { buildSourceFactorCase, buildHardExcludeClause, buildVisibilityClause, buildRecencyComponentSql, buildBestPerPagePoolCte, buildOrFallbackWebsearchQuery } from './search/sql-ranking.ts';
|
||||
import { shouldExcludeFromOrphanReporting, loadOrphanPolicyOverrides } from './orphan-policy.ts';
|
||||
import { LINK_EXTRACTOR_VERSION_TS } from './link-extraction.ts';
|
||||
import {
|
||||
normalizeEngineColumn,
|
||||
buildVectorCastFragment,
|
||||
@@ -2322,6 +2324,10 @@ export class PGLiteEngine implements BrainEngine {
|
||||
// v0.40.3.0 D24 NULL→non-NULL race fix mirrors postgres-engine.ts. Two writers
|
||||
// racing on the same chunk previously raced last-write-wins; the fix lets the
|
||||
// fresher `embedded_at` win in the text-unchanged branch.
|
||||
//
|
||||
// Code-chunk metadata columns follow the same chunk_text-gated CASE pattern as `embedding`
|
||||
// (#769). Re-chunk trusts EXCLUDED outright; pure re-embed COALESCEs so a caller carrying
|
||||
// only embedding-shaped fields doesn't clobber metadata to NULL.
|
||||
await this.db.query(
|
||||
`INSERT INTO content_chunks ${cols} VALUES ${rowParts.join(', ')}
|
||||
ON CONFLICT (page_id, chunk_index) DO UPDATE SET
|
||||
@@ -2345,14 +2351,14 @@ export class PGLiteEngine implements BrainEngine {
|
||||
THEN EXCLUDED.embedded_at
|
||||
ELSE content_chunks.embedded_at
|
||||
END,
|
||||
language = EXCLUDED.language,
|
||||
symbol_name = EXCLUDED.symbol_name,
|
||||
symbol_type = EXCLUDED.symbol_type,
|
||||
start_line = EXCLUDED.start_line,
|
||||
end_line = EXCLUDED.end_line,
|
||||
parent_symbol_path = EXCLUDED.parent_symbol_path,
|
||||
doc_comment = EXCLUDED.doc_comment,
|
||||
symbol_name_qualified = EXCLUDED.symbol_name_qualified,
|
||||
language = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.language ELSE COALESCE(EXCLUDED.language, content_chunks.language) END,
|
||||
symbol_name = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name ELSE COALESCE(EXCLUDED.symbol_name, content_chunks.symbol_name) END,
|
||||
symbol_type = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_type ELSE COALESCE(EXCLUDED.symbol_type, content_chunks.symbol_type) END,
|
||||
start_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.start_line ELSE COALESCE(EXCLUDED.start_line, content_chunks.start_line) END,
|
||||
end_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.end_line ELSE COALESCE(EXCLUDED.end_line, content_chunks.end_line) END,
|
||||
parent_symbol_path = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.parent_symbol_path ELSE COALESCE(EXCLUDED.parent_symbol_path, content_chunks.parent_symbol_path) END,
|
||||
doc_comment = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.doc_comment ELSE COALESCE(EXCLUDED.doc_comment, content_chunks.doc_comment) END,
|
||||
symbol_name_qualified = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name_qualified ELSE COALESCE(EXCLUDED.symbol_name_qualified, content_chunks.symbol_name_qualified) END,
|
||||
modality = EXCLUDED.modality,
|
||||
embedding_image = COALESCE(EXCLUDED.embedding_image, content_chunks.embedding_image)`,
|
||||
params
|
||||
@@ -5207,15 +5213,10 @@ export class PGLiteEngine implements BrainEngine {
|
||||
(SELECT count(*) FROM pages) as page_count,
|
||||
(SELECT count(*) FROM content_chunks WHERE embedded_at IS NOT NULL)::float /
|
||||
GREATEST((SELECT count(*) FROM content_chunks), 1)::float as embed_coverage,
|
||||
(SELECT count(*) FROM pages p
|
||||
WHERE p.updated_at < (SELECT MAX(te.created_at) FROM timeline_entries te WHERE te.page_id = p.id)
|
||||
) as stale_pages,
|
||||
-- Bug 11 — orphan = islanded (no inbound AND no outbound).
|
||||
-- See BrainHealth.orphan_pages docstring; docs updated to match this.
|
||||
(SELECT count(*) FROM pages p
|
||||
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
|
||||
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
|
||||
) as orphan_pages,
|
||||
0 as stale_pages,
|
||||
-- Bug 11 — orphan = islanded (no inbound AND no outbound). The raw
|
||||
-- list is filtered in TS using the shared orphan-reporting policy.
|
||||
0 as orphan_pages,
|
||||
(SELECT count(*) FROM links l
|
||||
WHERE NOT EXISTS (SELECT 1 FROM pages p WHERE p.id = l.to_page_id)
|
||||
) as dead_links,
|
||||
@@ -5240,10 +5241,20 @@ export class PGLiteEngine implements BrainEngine {
|
||||
LIMIT 5
|
||||
`);
|
||||
|
||||
const { rows: islandedRows } = await this.db.query(`
|
||||
SELECT p.slug
|
||||
FROM pages p
|
||||
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
|
||||
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
|
||||
`);
|
||||
|
||||
const r = h as Record<string, unknown>;
|
||||
const pageCount = Number(r.page_count);
|
||||
const embedCoverage = Number(r.embed_coverage);
|
||||
const orphanPages = Number(r.orphan_pages);
|
||||
const stalePages = await this.countStalePagesForExtraction({ versionTs: LINK_EXTRACTOR_VERSION_TS });
|
||||
const orphanOverrides = await loadOrphanPolicyOverrides(this);
|
||||
const orphanPages = (islandedRows as { slug: string }[])
|
||||
.filter(row => !shouldExcludeFromOrphanReporting(row.slug, orphanOverrides)).length;
|
||||
const deadLinks = Number(r.dead_links);
|
||||
const linkCount = Number(r.link_count);
|
||||
const pagesWithTimeline = Number(r.pages_with_timeline);
|
||||
@@ -5271,7 +5282,7 @@ export class PGLiteEngine implements BrainEngine {
|
||||
return {
|
||||
page_count: pageCount,
|
||||
embed_coverage: embedCoverage,
|
||||
stale_pages: Number(r.stale_pages),
|
||||
stale_pages: stalePages,
|
||||
orphan_pages: orphanPages,
|
||||
missing_embeddings: Number(r.missing_embeddings),
|
||||
brain_score: brainScore,
|
||||
|
||||
+33
-22
@@ -67,6 +67,8 @@ import { resolveBoostMap, resolveHardExcludes } from './search/source-boost.ts';
|
||||
import { buildSourceFactorCase, buildHardExcludeClause, buildVisibilityClause, buildRecencyComponentSql, buildBestPerPagePoolCte, buildOrFallbackWebsearchQuery } from './search/sql-ranking.ts';
|
||||
import { DEFAULT_EMBEDDING_MODEL, DEFAULT_EMBEDDING_DIMENSIONS } from './ai/defaults.ts';
|
||||
import { DELETE_BATCH_SIZE } from './engine-constants.ts';
|
||||
import { shouldExcludeFromOrphanReporting, loadOrphanPolicyOverrides } from './orphan-policy.ts';
|
||||
import { LINK_EXTRACTOR_VERSION_TS } from './link-extraction.ts';
|
||||
|
||||
function escapeSqlStringLiteral(value: string): string {
|
||||
return value.replace(/'/g, "''");
|
||||
@@ -2473,6 +2475,13 @@ export class PostgresEngine implements BrainEngine {
|
||||
// - new is fresher (embedded_at > existing.embedded_at) → take new
|
||||
// - otherwise → keep existing (slower writer with stale embedding loses)
|
||||
// Mirrored in pglite-engine.ts; pinned by test/e2e/concurrent-embed-race.test.ts.
|
||||
//
|
||||
// Code-chunk metadata columns (language / symbol_name / symbol_type / line range /
|
||||
// parent_symbol_path / doc_comment / symbol_name_qualified) follow the SAME chunk_text-gated
|
||||
// CASE pattern as `embedding` (#769). Re-chunk (chunk_text changed) trusts EXCLUDED outright;
|
||||
// pure re-embed (chunk_text unchanged) COALESCEs so a caller that only carries embedding
|
||||
// doesn't clobber metadata to NULL. Without this, every embed --stale pass nuked code-def's
|
||||
// primary index for thousands of chunks at once.
|
||||
await sql.unsafe(
|
||||
`INSERT INTO content_chunks ${cols} VALUES ${rows.join(', ')}
|
||||
ON CONFLICT (page_id, chunk_index) DO UPDATE SET
|
||||
@@ -2496,14 +2505,14 @@ export class PostgresEngine implements BrainEngine {
|
||||
THEN EXCLUDED.embedded_at
|
||||
ELSE content_chunks.embedded_at
|
||||
END,
|
||||
language = EXCLUDED.language,
|
||||
symbol_name = EXCLUDED.symbol_name,
|
||||
symbol_type = EXCLUDED.symbol_type,
|
||||
start_line = EXCLUDED.start_line,
|
||||
end_line = EXCLUDED.end_line,
|
||||
parent_symbol_path = EXCLUDED.parent_symbol_path,
|
||||
doc_comment = EXCLUDED.doc_comment,
|
||||
symbol_name_qualified = EXCLUDED.symbol_name_qualified,
|
||||
language = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.language ELSE COALESCE(EXCLUDED.language, content_chunks.language) END,
|
||||
symbol_name = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name ELSE COALESCE(EXCLUDED.symbol_name, content_chunks.symbol_name) END,
|
||||
symbol_type = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_type ELSE COALESCE(EXCLUDED.symbol_type, content_chunks.symbol_type) END,
|
||||
start_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.start_line ELSE COALESCE(EXCLUDED.start_line, content_chunks.start_line) END,
|
||||
end_line = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.end_line ELSE COALESCE(EXCLUDED.end_line, content_chunks.end_line) END,
|
||||
parent_symbol_path = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.parent_symbol_path ELSE COALESCE(EXCLUDED.parent_symbol_path, content_chunks.parent_symbol_path) END,
|
||||
doc_comment = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.doc_comment ELSE COALESCE(EXCLUDED.doc_comment, content_chunks.doc_comment) END,
|
||||
symbol_name_qualified = CASE WHEN EXCLUDED.chunk_text != content_chunks.chunk_text THEN EXCLUDED.symbol_name_qualified ELSE COALESCE(EXCLUDED.symbol_name_qualified, content_chunks.symbol_name_qualified) END,
|
||||
modality = EXCLUDED.modality,
|
||||
embedding_image = COALESCE(EXCLUDED.embedding_image, content_chunks.embedding_image)`,
|
||||
params as Parameters<typeof sql.unsafe>[1],
|
||||
@@ -5313,11 +5322,9 @@ export class PostgresEngine implements BrainEngine {
|
||||
async getHealth(): Promise<BrainHealth> {
|
||||
const sql = this.sql;
|
||||
// Bug 11 doc-drift fix — orphan_pages means "islanded" (no inbound AND
|
||||
// no outbound links), aligning both engines with the user-facing
|
||||
// definition. The type comment previously said "no inbound" but the
|
||||
// SQL required both — docs now match code so users can trust the
|
||||
// number. A hub page that links out to many but has no back-references
|
||||
// is working as intended, not an orphan.
|
||||
// no outbound links). The raw islanded list is filtered through the same
|
||||
// policy as `gbrain orphans` so convention pages do not count against
|
||||
// dashboard health.
|
||||
const [h] = await sql`
|
||||
WITH entity_pages AS (
|
||||
SELECT id, slug FROM pages WHERE type IN ('person', 'company')
|
||||
@@ -5326,13 +5333,8 @@ export class PostgresEngine implements BrainEngine {
|
||||
(SELECT count(*) FROM pages) as page_count,
|
||||
(SELECT count(*) FROM content_chunks WHERE embedded_at IS NOT NULL)::float /
|
||||
GREATEST((SELECT count(*) FROM content_chunks), 1)::float as embed_coverage,
|
||||
(SELECT count(*) FROM pages p
|
||||
WHERE p.updated_at < (SELECT MAX(te.created_at) FROM timeline_entries te WHERE te.page_id = p.id)
|
||||
) as stale_pages,
|
||||
(SELECT count(*) FROM pages p
|
||||
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
|
||||
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
|
||||
) as orphan_pages,
|
||||
0 as stale_pages,
|
||||
0 as orphan_pages,
|
||||
(SELECT count(*) FROM links l
|
||||
WHERE NOT EXISTS (SELECT 1 FROM pages p WHERE p.id = l.to_page_id)
|
||||
) as dead_links,
|
||||
@@ -5356,9 +5358,18 @@ export class PostgresEngine implements BrainEngine {
|
||||
LIMIT 5
|
||||
`;
|
||||
|
||||
const islandedRows = await sql<{ slug: string }[]>`
|
||||
SELECT p.slug
|
||||
FROM pages p
|
||||
WHERE NOT EXISTS (SELECT 1 FROM links l WHERE l.to_page_id = p.id)
|
||||
AND NOT EXISTS (SELECT 1 FROM links l WHERE l.from_page_id = p.id)
|
||||
`;
|
||||
|
||||
const pageCount = Number(h.page_count);
|
||||
const embedCoverage = Number(h.embed_coverage);
|
||||
const orphanPages = Number(h.orphan_pages);
|
||||
const stalePages = await this.countStalePagesForExtraction({ versionTs: LINK_EXTRACTOR_VERSION_TS });
|
||||
const orphanOverrides = await loadOrphanPolicyOverrides(this);
|
||||
const orphanPages = islandedRows.filter(row => !shouldExcludeFromOrphanReporting(row.slug, orphanOverrides)).length;
|
||||
const deadLinks = Number(h.dead_links);
|
||||
const linkCount = Number(h.link_count);
|
||||
const pagesWithTimeline = Number(h.pages_with_timeline);
|
||||
@@ -5386,7 +5397,7 @@ export class PostgresEngine implements BrainEngine {
|
||||
return {
|
||||
page_count: pageCount,
|
||||
embed_coverage: embedCoverage,
|
||||
stale_pages: Number(h.stale_pages),
|
||||
stale_pages: stalePages,
|
||||
orphan_pages: orphanPages,
|
||||
missing_embeddings: Number(h.missing_embeddings),
|
||||
brain_score: brainScore,
|
||||
|
||||
@@ -39,13 +39,6 @@ CREATE TABLE IF NOT EXISTS sources (
|
||||
-- bypassing the git-HEAD up_to_date early-return so CHUNKER_VERSION bumps
|
||||
-- actually trigger re-chunking on upgrade.
|
||||
chunker_version TEXT,
|
||||
-- #2157 follow-on: SHA-256 fingerprint of the walk-affecting fields in
|
||||
-- \`config\` (strategy + include_globs + exclude_globs). Mismatch forces a
|
||||
-- full re-walk via the same code path as chunker_version, so a user who
|
||||
-- changes \`sources.config.exclude_globs\` mid-life doesn't get "Already up
|
||||
-- to date" on the next sync. NULL on pre-migration rows is treated as
|
||||
-- "not yet stamped" and skips the gate (preserves first-run semantics).
|
||||
config_fingerprint TEXT,
|
||||
-- v0.26.5: soft-delete + recovery window. \`archive\` flips archived=true and
|
||||
-- sets archive_expires_at = now() + 72h. The autopilot purge phase
|
||||
-- hard-deletes rows where archive_expires_at <= now(). Promoted from a
|
||||
|
||||
@@ -93,7 +93,13 @@ import { resolveLrSchedule } from './lr-schedule.ts';
|
||||
import { preflight, formatPreflightReport } from './preflight.ts';
|
||||
import { isRejected, loadRejectedBuffer, makeRejectedEntry, saveRejectedBuffer } from './rejected-buffer.ts';
|
||||
import { runReflect, runOneShotRewrite, describeJudges } from './reflect.ts';
|
||||
import { acceptCandidate, bestPath, revertAllPending, skillPath, writeProposed } from './version-store.ts';
|
||||
import {
|
||||
acceptCandidate,
|
||||
proposedPath as proposedFilePath,
|
||||
revertAllPending,
|
||||
skillPath,
|
||||
writeProposed,
|
||||
} from './version-store.ts';
|
||||
import { runValidationGate, scoreSkillOnTasks } from './validate-gate.ts';
|
||||
import { ROLLOUT_SUCCESS_THRESHOLD } from './types.ts';
|
||||
import type { SkillOptOpts, EditOp, RunReceipt, BenchmarkTask } from './types.ts';
|
||||
@@ -702,9 +708,9 @@ async function runOptimizationLoop(
|
||||
// to the catch's assignment values only (it can't prove the async callback ran).
|
||||
const finalOutcome = outcome as 'accepted' | 'no_improvement' | 'aborted' | 'errored';
|
||||
if (!mutateDecision.mutate && finalOutcome === 'accepted') {
|
||||
// best.md was written by writeProposed() in the accept branch (no-mutate
|
||||
// path); it doubles as proposed.md for human review. SKILL.md untouched.
|
||||
proposedPath = bestPath(skillsDir, skillName);
|
||||
// writeProposed() emitted both the best pointer and the stable review
|
||||
// artifact in the accept branch. SKILL.md remains untouched.
|
||||
proposedPath = proposedFilePath(skillsDir, skillName);
|
||||
} else if (mutateDecision.mutate) {
|
||||
mutatedSkillFile = finalOutcome === 'accepted';
|
||||
}
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
*
|
||||
* history.json
|
||||
* best.md
|
||||
* proposed.md
|
||||
* versions/
|
||||
* v0001_e1_s1.md
|
||||
* v0002_e1_s2.md
|
||||
@@ -52,6 +53,10 @@ export function bestPath(skillsDir: string, skillName: string): string {
|
||||
return path.join(skilloptDir(skillsDir, skillName), 'best.md');
|
||||
}
|
||||
|
||||
export function proposedPath(skillsDir: string, skillName: string): string {
|
||||
return path.join(skilloptDir(skillsDir, skillName), 'proposed.md');
|
||||
}
|
||||
|
||||
export function skillPath(skillsDir: string, skillName: string): string {
|
||||
return path.join(skillsDir, skillName, 'SKILL.md');
|
||||
}
|
||||
@@ -171,17 +176,18 @@ export function acceptCandidate(input: AcceptInput): AcceptResult {
|
||||
}
|
||||
|
||||
/**
|
||||
* Write the candidate to `best.md` (which doubles as `proposed.md`) WITHOUT
|
||||
* touching SKILL.md or the history ledger. Used by the `--no-mutate` /
|
||||
* bundled-without-allow paths: the optimizer found a better candidate but the
|
||||
* caller opted out of in-place mutation, so we surface it for human review.
|
||||
* Returns the path written. Atomic (.tmp + rename).
|
||||
* Write the candidate to both `best.md` and `proposed.md` WITHOUT touching
|
||||
* SKILL.md or the history ledger. `best.md` remains the optimizer's current
|
||||
* best pointer; `proposed.md` is the stable human-review artifact promised by
|
||||
* `--no-mutate`. Returns the proposal path. Each write is atomic (.tmp + rename).
|
||||
*/
|
||||
export function writeProposed(skillsDir: string, skillName: string, candidateText: string): string {
|
||||
const p = bestPath(skillsDir, skillName);
|
||||
fs.mkdirSync(path.dirname(p), { recursive: true });
|
||||
atomicWrite(p, candidateText);
|
||||
return p;
|
||||
const best = bestPath(skillsDir, skillName);
|
||||
const proposed = proposedPath(skillsDir, skillName);
|
||||
fs.mkdirSync(path.dirname(best), { recursive: true });
|
||||
atomicWrite(best, candidateText);
|
||||
atomicWrite(proposed, candidateText);
|
||||
return proposed;
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -155,19 +155,6 @@ export interface AddSourceOpts {
|
||||
* runs). Does NOT auto-`git init` anything — see `addSource` docstring.
|
||||
*/
|
||||
force?: boolean;
|
||||
/**
|
||||
* Glob filters persisted into `sources.config.include_globs` /
|
||||
* `sources.config.exclude_globs`. Read at sync time by
|
||||
* `commands/sync.ts:syncOneSource` and the single-source path, threaded
|
||||
* into `isSyncable` / `unsyncableReason` (their `SyncableOptions` shape
|
||||
* has carried this contract since v0.41.13).
|
||||
*
|
||||
* Empty / unspecified arrays are not persisted at all (no `[]` written
|
||||
* to the JSONB), which keeps the row identical to today for sources
|
||||
* that don't use filtering.
|
||||
*/
|
||||
includeGlobs?: string[];
|
||||
excludeGlobs?: string[];
|
||||
}
|
||||
|
||||
export interface RemoveSourceOpts {
|
||||
@@ -442,12 +429,6 @@ export async function addSource(
|
||||
if (opts.federated !== null && opts.federated !== undefined) {
|
||||
config.federated = opts.federated;
|
||||
}
|
||||
if (opts.includeGlobs && opts.includeGlobs.length > 0) {
|
||||
config.include_globs = opts.includeGlobs;
|
||||
}
|
||||
if (opts.excludeGlobs && opts.excludeGlobs.length > 0) {
|
||||
config.exclude_globs = opts.excludeGlobs;
|
||||
}
|
||||
const displayName = opts.name ?? opts.id;
|
||||
|
||||
try {
|
||||
@@ -527,12 +508,6 @@ export async function addSource(
|
||||
if (opts.federated !== null && opts.federated !== undefined) {
|
||||
config.federated = opts.federated;
|
||||
}
|
||||
if (opts.includeGlobs && opts.includeGlobs.length > 0) {
|
||||
config.include_globs = opts.includeGlobs;
|
||||
}
|
||||
if (opts.excludeGlobs && opts.excludeGlobs.length > 0) {
|
||||
config.exclude_globs = opts.excludeGlobs;
|
||||
}
|
||||
const displayName = opts.name ?? opts.id;
|
||||
await engine.executeRaw(
|
||||
`INSERT INTO sources (id, name, local_path, config)
|
||||
|
||||
@@ -219,13 +219,6 @@ function globToRegex(pattern: string): RegExp {
|
||||
return new RegExp(regex);
|
||||
}
|
||||
|
||||
/**
|
||||
* Test a normalized POSIX-style path against an array of glob patterns. Returns
|
||||
* true if any pattern matches. Empty / undefined `patterns` returns false (no
|
||||
* filter engaged). Exported so non-sync surfaces (lint walker, future ingest
|
||||
* variants) can apply the same glob semantics as `isSyncable` without
|
||||
* re-declaring `globToRegex`.
|
||||
*/
|
||||
export function matchesAnyGlob(path: string, patterns?: string[]): boolean {
|
||||
if (!patterns || patterns.length === 0) return false;
|
||||
const normalized = path.replace(/\\/g, '/');
|
||||
|
||||
@@ -63,25 +63,26 @@ interface PluginCtx {
|
||||
[key: string]: unknown;
|
||||
}
|
||||
|
||||
export function register(api: PluginApi) {
|
||||
api.registerContextEngine(ENGINE_ID, (ctx: PluginCtx) => {
|
||||
const hostResolver =
|
||||
typeof ctx.resolveEntities === 'function'
|
||||
? ctx.resolveEntities
|
||||
: typeof ctx.brainQuery === 'function'
|
||||
? ctx.brainQuery
|
||||
: undefined;
|
||||
return createGBrainContextEngine({
|
||||
workspaceDir: ctx.workspaceDir,
|
||||
resolveEntities: hostResolver,
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
const entry: PluginEntry = {
|
||||
id: 'gbrain-context-engine',
|
||||
name: 'GBrain Context Engine',
|
||||
description: 'Deterministic temporal/spatial context injection on every turn',
|
||||
|
||||
register(api: PluginApi) {
|
||||
api.registerContextEngine(ENGINE_ID, (ctx: PluginCtx) => {
|
||||
const hostResolver =
|
||||
typeof ctx.resolveEntities === 'function'
|
||||
? ctx.resolveEntities
|
||||
: typeof ctx.brainQuery === 'function'
|
||||
? ctx.brainQuery
|
||||
: undefined;
|
||||
return createGBrainContextEngine({
|
||||
workspaceDir: ctx.workspaceDir,
|
||||
resolveEntities: hostResolver,
|
||||
});
|
||||
});
|
||||
},
|
||||
register,
|
||||
};
|
||||
|
||||
export default entry;
|
||||
|
||||
@@ -35,13 +35,6 @@ CREATE TABLE IF NOT EXISTS sources (
|
||||
-- bypassing the git-HEAD up_to_date early-return so CHUNKER_VERSION bumps
|
||||
-- actually trigger re-chunking on upgrade.
|
||||
chunker_version TEXT,
|
||||
-- #2157 follow-on: SHA-256 fingerprint of the walk-affecting fields in
|
||||
-- `config` (strategy + include_globs + exclude_globs). Mismatch forces a
|
||||
-- full re-walk via the same code path as chunker_version, so a user who
|
||||
-- changes `sources.config.exclude_globs` mid-life doesn't get "Already up
|
||||
-- to date" on the next sync. NULL on pre-migration rows is treated as
|
||||
-- "not yet stamped" and skips the gate (preserves first-run semantics).
|
||||
config_fingerprint TEXT,
|
||||
-- v0.26.5: soft-delete + recovery window. `archive` flips archived=true and
|
||||
-- sets archive_expires_at = now() + 72h. The autopilot purge phase
|
||||
-- hard-deletes rows where archive_expires_at <= now(). Promoted from a
|
||||
|
||||
@@ -0,0 +1,67 @@
|
||||
// #1474: the v0.41.1 wave shipped bench-publish.ts + docs/eval-bench.md
|
||||
// advertising `gbrain bench publish`, but the cli.ts dispatcher case was never
|
||||
// added — the documented command hit 'Unknown command'. These tests spawn the
|
||||
// real CLI (no DB needed; bench publish is pure file I/O) and fail on any
|
||||
// regression of the dispatcher wiring.
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import { mkdtempSync, writeFileSync, existsSync, rmSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { spawnSync } from 'node:child_process';
|
||||
|
||||
function runCli(args: string[]): { stdout: string; stderr: string; code: number } {
|
||||
const result = spawnSync(process.execPath, ['run', 'src/cli.ts', 'bench', ...args], {
|
||||
encoding: 'utf8',
|
||||
cwd: process.cwd(),
|
||||
env: { ...process.env },
|
||||
});
|
||||
return { stdout: result.stdout ?? '', stderr: result.stderr ?? '', code: result.status ?? -1 };
|
||||
}
|
||||
|
||||
describe('gbrain bench dispatcher (#1474)', () => {
|
||||
test('bench --help reaches bench-publish help without a DB (was: Unknown command)', () => {
|
||||
const { stdout, stderr, code } = runCli(['--help']);
|
||||
expect(stderr).not.toContain('Unknown command');
|
||||
expect(code).toBe(0);
|
||||
expect(stdout).toContain('gbrain bench publish');
|
||||
expect(stdout).toContain('--from');
|
||||
});
|
||||
|
||||
test('unknown bench subcommand exits 2 with usage', () => {
|
||||
const { stderr, code } = runCli(['bogus']);
|
||||
expect(code).toBe(2);
|
||||
expect(stderr).toContain('Unknown bench subcommand');
|
||||
expect(stderr).toContain('bench publish');
|
||||
});
|
||||
|
||||
test('bench publish roundtrip: captured NDJSON in, baseline file out', () => {
|
||||
const tmp = mkdtempSync(join(tmpdir(), 'bench-cli-'));
|
||||
try {
|
||||
const row = {
|
||||
tool_name: 'query',
|
||||
query: 'hello world',
|
||||
retrieved_slugs: ['slug-a'],
|
||||
retrieved_chunk_ids: [1],
|
||||
source_ids: ['default'],
|
||||
expand_enabled: false,
|
||||
detail: 'medium',
|
||||
detail_resolved: 'medium',
|
||||
vector_enabled: true,
|
||||
expansion_applied: false,
|
||||
latency_ms: 100,
|
||||
remote: false,
|
||||
job_id: null,
|
||||
subagent_id: null,
|
||||
};
|
||||
const from = join(tmp, 'captured.ndjson');
|
||||
const to = join(tmp, 'personal.baseline.ndjson');
|
||||
writeFileSync(from, `${JSON.stringify(row)}\n`);
|
||||
const { code, stderr } = runCli(['publish', '--from', from, '--to', to]);
|
||||
expect(stderr).not.toContain('Unknown command');
|
||||
expect(code).toBe(0);
|
||||
expect(existsSync(to)).toBe(true);
|
||||
} finally {
|
||||
rmSync(tmp, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -172,6 +172,59 @@ describe('issue #972 — DB-source (gbrain extract links --source db)', () => {
|
||||
expect(strk!.link_type).toBe('wikilink_basename');
|
||||
});
|
||||
|
||||
test('flag ON → path-qualified wikilink outside DIR_PATTERN resolves via DB path', async () => {
|
||||
// `[[notes/struktura]]` — `notes` is not in DIR_PATTERN, so the ref
|
||||
// reaches the generic pass with its dirname intact. Regression: the DB
|
||||
// path queried the basename index with the raw literal (which is keyed
|
||||
// by final segments only), so path-qualified wikilinks outside
|
||||
// DIR_PATTERN silently produced zero edges while the FS path resolved
|
||||
// the identical content.
|
||||
await engine.putPage('notes/struktura', {
|
||||
type: 'concept' as any, title: 'Struktura Notes',
|
||||
compiled_truth: '', timeline: '',
|
||||
});
|
||||
await engine.putPage('concepts/knowledge-graph', {
|
||||
type: 'concept', title: 'Knowledge Graph',
|
||||
compiled_truth: 'Background in [[notes/struktura]].', timeline: '',
|
||||
});
|
||||
await engine.setConfig('link_resolution.global_basename', 'true');
|
||||
|
||||
await runExtract(engine, ['links', '--source', 'db']);
|
||||
|
||||
const outLinks = await engine.getLinks('concepts/knowledge-graph');
|
||||
const strk = outLinks.find(l => l.to_slug === 'notes/struktura');
|
||||
expect(strk).toBeDefined();
|
||||
expect(strk!.link_type).toBe('wikilink_basename');
|
||||
expect(strk!.link_source).toBe('wikilink-resolved');
|
||||
});
|
||||
|
||||
test('path-qualified wikilink never attaches to a basename-only sibling', async () => {
|
||||
// Both notes/struktura and wiki/struktura exist. The author wrote
|
||||
// `[[notes/struktura]]` — the written path must exclude wiki/struktura
|
||||
// (a bare `[[struktura]]` would legitimately match both).
|
||||
await engine.putPage('notes/struktura', {
|
||||
type: 'concept' as any, title: 'Struktura Notes',
|
||||
compiled_truth: '', timeline: '',
|
||||
});
|
||||
await engine.putPage('wiki/struktura', {
|
||||
type: 'concept' as any, title: 'Struktura Wiki',
|
||||
compiled_truth: '', timeline: '',
|
||||
});
|
||||
await engine.putPage('concepts/x', {
|
||||
type: 'concept', title: 'X',
|
||||
compiled_truth: 'See [[notes/struktura]].', timeline: '',
|
||||
});
|
||||
await engine.setConfig('link_resolution.global_basename', 'true');
|
||||
|
||||
await runExtract(engine, ['links', '--source', 'db']);
|
||||
|
||||
const outLinks = await engine.getLinks('concepts/x');
|
||||
const basenameLinks = outLinks
|
||||
.filter(l => l.link_type === 'wikilink_basename')
|
||||
.map(l => l.to_slug);
|
||||
expect(basenameLinks).toEqual(['notes/struktura']);
|
||||
});
|
||||
|
||||
test('flag OFF → no basename edges via DB path (back-compat)', async () => {
|
||||
await engine.putPage('projects/struktura', {
|
||||
type: 'project', title: 'Struktura',
|
||||
|
||||
@@ -39,6 +39,7 @@ import { runSkillOpt } from '../../src/core/skillopt/orchestrator.ts';
|
||||
import {
|
||||
bestPath,
|
||||
loadHistory,
|
||||
proposedPath,
|
||||
skillPath,
|
||||
} from '../../src/core/skillopt/version-store.ts';
|
||||
import { loadRejectedBuffer } from '../../src/core/skillopt/rejected-buffer.ts';
|
||||
@@ -741,7 +742,7 @@ describe('skillopt T3 — F11 held-out gate, ablation opts, no-DB-pollution', ()
|
||||
} finally { fixture.cleanup(); }
|
||||
});
|
||||
|
||||
test('--no-mutate writes proposed.md (best.md), leaves SKILL.md untouched', async () => {
|
||||
test('--no-mutate writes proposed.md and best.md, leaves SKILL.md untouched', async () => {
|
||||
const fixture = setupFixture(SKILL_PEOPLE_ONLY, CITATIONS_BENCHMARK);
|
||||
try {
|
||||
installStub({
|
||||
@@ -753,10 +754,9 @@ describe('skillopt T3 — F11 held-out gate, ablation opts, no-DB-pollution', ()
|
||||
const result = await runOnce(fixture, { noMutate: true });
|
||||
expect(result.outcome).toBe('accepted');
|
||||
expect(result.mutatedSkillFile).toBe(false);
|
||||
expect(result.proposedPath).toBeDefined();
|
||||
// proposed.md (best.md) exists and carries the improvement.
|
||||
expect(fs.existsSync(result.proposedPath!)).toBe(true);
|
||||
expect(result.proposedPath).toBe(proposedPath(fixture.skillsDir, SKILL));
|
||||
expect(fs.readFileSync(result.proposedPath!, 'utf8')).toContain('## Citations');
|
||||
expect(fs.readFileSync(bestPath(fixture.skillsDir, SKILL), 'utf8')).toContain('## Citations');
|
||||
// SKILL.md on disk is UNCHANGED (still People-only).
|
||||
const skill = fs.readFileSync(skillPath(fixture.skillsDir, SKILL), 'utf8');
|
||||
expect(skill).not.toContain('## Citations');
|
||||
|
||||
@@ -803,3 +803,107 @@ describe('embedAllStale --source threading (D7)', () => {
|
||||
expect((firstCallOpts as { sourceId?: string }).sourceId).toBe('media-corpus');
|
||||
});
|
||||
});
|
||||
|
||||
// ────────────────────────────────────────────────────────────────
|
||||
// Code metadata preservation across re-embed (regression for #769)
|
||||
// ────────────────────────────────────────────────────────────────
|
||||
//
|
||||
// gbrain v0.30.1 and earlier silently clobbered code-chunk metadata
|
||||
// (language, symbol_name, symbol_type, start_line, end_line,
|
||||
// parent_symbol_path, doc_comment, symbol_name_qualified) on every
|
||||
// re-embed pass. The chunker populated those columns at import time,
|
||||
// but embed.ts loaded chunks via getChunks then mapped them to a
|
||||
// stripped ChunkInput carrying only 5 fields. upsertChunks then
|
||||
// OVERWROTE (not COALESCEd) the metadata columns from EXCLUDED, so
|
||||
// re-embed wiped them to NULL. End result on a real brain: 4875 code
|
||||
// pages, 47866 chunks, all with NULL language/symbol_name/symbol_type;
|
||||
// code-def returned 0 hits across every indexed repo.
|
||||
//
|
||||
// All three runEmbed paths (--stale autopilot, --all, --slugs) must
|
||||
// thread metadata through the re-upsert. Tests below assert that the
|
||||
// engine.upsertChunks call carries the same metadata it loaded.
|
||||
|
||||
describe('runEmbed preserves code-chunk metadata across re-embed (regression for #769)', () => {
|
||||
const fullCodeChunk = {
|
||||
chunk_index: 0,
|
||||
chunk_text: '[Java] foo/Bar.java:10-20 method baz',
|
||||
chunk_source: 'compiled_truth' as const,
|
||||
embedded_at: null,
|
||||
token_count: 12,
|
||||
language: 'java',
|
||||
symbol_name: 'baz',
|
||||
symbol_type: 'function',
|
||||
start_line: 10,
|
||||
end_line: 20,
|
||||
parent_symbol_path: ['Bar'],
|
||||
doc_comment: 'does the thing',
|
||||
symbol_name_qualified: 'Bar.baz',
|
||||
};
|
||||
|
||||
function metadataOf(chunk: any) {
|
||||
return {
|
||||
language: chunk.language,
|
||||
symbol_name: chunk.symbol_name,
|
||||
symbol_type: chunk.symbol_type,
|
||||
start_line: chunk.start_line,
|
||||
end_line: chunk.end_line,
|
||||
parent_symbol_path: chunk.parent_symbol_path,
|
||||
doc_comment: chunk.doc_comment,
|
||||
symbol_name_qualified: chunk.symbol_name_qualified,
|
||||
};
|
||||
}
|
||||
|
||||
test('--stale (autopilot path) carries code metadata into upsertChunks', async () => {
|
||||
const stale = [{
|
||||
slug: 'code-page',
|
||||
chunk_index: 0,
|
||||
chunk_text: fullCodeChunk.chunk_text,
|
||||
chunk_source: 'compiled_truth',
|
||||
model: null,
|
||||
token_count: 12,
|
||||
}];
|
||||
let upsertChunkArgs: any[] | null = null;
|
||||
const engine = mockEngine({
|
||||
countStaleChunks: async () => 1,
|
||||
listStaleChunks: async () => stale,
|
||||
getChunks: async () => [fullCodeChunk],
|
||||
upsertChunks: async (_slug: string, chunks: any[]) => { upsertChunkArgs = chunks; },
|
||||
});
|
||||
|
||||
await runEmbed(engine, ['--stale']);
|
||||
|
||||
expect(upsertChunkArgs).not.toBeNull();
|
||||
expect(upsertChunkArgs!).toHaveLength(1);
|
||||
expect(metadataOf(upsertChunkArgs![0])).toEqual(metadataOf(fullCodeChunk));
|
||||
});
|
||||
|
||||
test('--all (full re-embed) carries code metadata into upsertChunks', async () => {
|
||||
let upsertChunkArgs: any[] | null = null;
|
||||
const engine = mockEngine({
|
||||
listPages: async () => [{ slug: 'code-page' }],
|
||||
getChunks: async () => [fullCodeChunk],
|
||||
upsertChunks: async (_slug: string, chunks: any[]) => { upsertChunkArgs = chunks; },
|
||||
});
|
||||
|
||||
await runEmbed(engine, ['--all']);
|
||||
|
||||
expect(upsertChunkArgs).not.toBeNull();
|
||||
expect(upsertChunkArgs!).toHaveLength(1);
|
||||
expect(metadataOf(upsertChunkArgs![0])).toEqual(metadataOf(fullCodeChunk));
|
||||
});
|
||||
|
||||
test('--slugs (per-page embed) carries code metadata into upsertChunks', async () => {
|
||||
let upsertChunkArgs: any[] | null = null;
|
||||
const engine = mockEngine({
|
||||
getPage: async () => ({ slug: 'code-page', compiled_truth: 'x', timeline: '' }),
|
||||
getChunks: async () => [fullCodeChunk],
|
||||
upsertChunks: async (_slug: string, chunks: any[]) => { upsertChunkArgs = chunks; },
|
||||
});
|
||||
|
||||
await runEmbed(engine, ['--slugs', 'code-page']);
|
||||
|
||||
expect(upsertChunkArgs).not.toBeNull();
|
||||
expect(upsertChunkArgs!).toHaveLength(1);
|
||||
expect(metadataOf(upsertChunkArgs![0])).toEqual(metadataOf(fullCodeChunk));
|
||||
});
|
||||
});
|
||||
|
||||
@@ -309,3 +309,62 @@ describe('isEvalCaptureEnabled / isEvalScrubEnabled (CONTRIBUTOR_MODE-gated)', (
|
||||
} finally { restore(); }
|
||||
});
|
||||
});
|
||||
|
||||
describe('DB-plane stash (#1475): GBRAIN_EVAL_CAPTURE / GBRAIN_EVAL_SCRUB_PII', () => {
|
||||
// connectEngine stamps `gbrain config set eval.capture` (DB plane) onto
|
||||
// these env vars because ctx.config is the sync file-plane load. Without
|
||||
// the stash check the DB value was written and never read.
|
||||
const origCapture = process.env.GBRAIN_EVAL_CAPTURE;
|
||||
const origScrub = process.env.GBRAIN_EVAL_SCRUB_PII;
|
||||
const origMode = process.env.GBRAIN_CONTRIBUTOR_MODE;
|
||||
const restore = () => {
|
||||
if (origCapture === undefined) delete process.env.GBRAIN_EVAL_CAPTURE;
|
||||
else process.env.GBRAIN_EVAL_CAPTURE = origCapture;
|
||||
if (origScrub === undefined) delete process.env.GBRAIN_EVAL_SCRUB_PII;
|
||||
else process.env.GBRAIN_EVAL_SCRUB_PII = origScrub;
|
||||
if (origMode === undefined) delete process.env.GBRAIN_CONTRIBUTOR_MODE;
|
||||
else process.env.GBRAIN_CONTRIBUTOR_MODE = origMode;
|
||||
};
|
||||
|
||||
test('stash=true turns capture on when file plane is silent (the #1475 repro)', () => {
|
||||
delete process.env.GBRAIN_CONTRIBUTOR_MODE;
|
||||
process.env.GBRAIN_EVAL_CAPTURE = 'true';
|
||||
try {
|
||||
expect(isEvalCaptureEnabled(null)).toBe(true);
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
const noEval: any = { engine: 'pglite' };
|
||||
expect(isEvalCaptureEnabled(noEval)).toBe(true);
|
||||
} finally { restore(); }
|
||||
});
|
||||
|
||||
test('stash=false wins over CONTRIBUTOR_MODE=1 (explicit per-key beats broad flag)', () => {
|
||||
process.env.GBRAIN_CONTRIBUTOR_MODE = '1';
|
||||
process.env.GBRAIN_EVAL_CAPTURE = 'false';
|
||||
try {
|
||||
expect(isEvalCaptureEnabled(null)).toBe(false);
|
||||
} finally { restore(); }
|
||||
});
|
||||
|
||||
test('file-plane explicit value still wins over the stash', () => {
|
||||
process.env.GBRAIN_EVAL_CAPTURE = 'true';
|
||||
try {
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
const disabled: any = { engine: 'pglite', eval: { capture: false } };
|
||||
expect(isEvalCaptureEnabled(disabled)).toBe(false);
|
||||
} finally { restore(); }
|
||||
});
|
||||
|
||||
test('scrub stash: false disables, file plane wins, default stays true', () => {
|
||||
process.env.GBRAIN_EVAL_SCRUB_PII = 'false';
|
||||
try {
|
||||
expect(isEvalScrubEnabled(null)).toBe(false);
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
const fileWins: any = { engine: 'pglite', eval: { scrub_pii: true } };
|
||||
expect(isEvalScrubEnabled(fileWins)).toBe(true);
|
||||
} finally { restore(); }
|
||||
delete process.env.GBRAIN_EVAL_SCRUB_PII;
|
||||
try {
|
||||
expect(isEvalScrubEnabled(null)).toBe(true);
|
||||
} finally { restore(); }
|
||||
});
|
||||
});
|
||||
|
||||
@@ -403,6 +403,77 @@ describe('extractPageLinks', () => {
|
||||
expect(candidates).toEqual([]);
|
||||
});
|
||||
|
||||
test('path-qualified wikilink outside DIR_PATTERN queries by final segment', async () => {
|
||||
// `[[notes/struktura]]` (dir not in DIR_PATTERN) falls to the generic
|
||||
// pass. The resolver's basename index is keyed by final path segments,
|
||||
// so the lookup must strip the dirname — mirroring the FS path
|
||||
// (resolveSlugAll). Regression: the raw literal was passed through,
|
||||
// which never matched, so these links silently dropped.
|
||||
const seen: string[] = [];
|
||||
const resolver: SlugResolver = {
|
||||
resolve: async () => null,
|
||||
resolveBasenameMatches: async (name) => {
|
||||
seen.push(name);
|
||||
return name === 'struktura' ? ['notes/struktura'] : [];
|
||||
},
|
||||
};
|
||||
const { candidates } = await extractPageLinks(
|
||||
'concepts/x', 'See [[notes/struktura]].',
|
||||
{}, 'concept', resolver, { globalBasename: true },
|
||||
);
|
||||
expect(seen).toContain('struktura');
|
||||
expect(seen).not.toContain('notes/struktura');
|
||||
expect(candidates.map(c => c.targetSlug)).toEqual(['notes/struktura']);
|
||||
expect(candidates[0].linkType).toBe('wikilink_basename');
|
||||
expect(candidates[0].linkSource).toBe('wikilink-resolved');
|
||||
});
|
||||
|
||||
test('path-qualified wikilink keeps only matches ending with the written path', async () => {
|
||||
// The written path disambiguates: `[[notes/struktura]]` must never
|
||||
// attach to `wiki/struktura` even though both share the basename.
|
||||
const resolver: SlugResolver = {
|
||||
resolve: async () => null,
|
||||
resolveBasenameMatches: async (name) =>
|
||||
name === 'struktura' ? ['notes/struktura', 'wiki/struktura'] : [],
|
||||
};
|
||||
const { candidates } = await extractPageLinks(
|
||||
'concepts/x', 'See [[notes/struktura]].',
|
||||
{}, 'concept', resolver, { globalBasename: true },
|
||||
);
|
||||
expect(candidates.map(c => c.targetSlug)).toEqual(['notes/struktura']);
|
||||
});
|
||||
|
||||
test('path-qualified wikilink matches a deeper real slug by path suffix', async () => {
|
||||
// The page lives at vault/notes/struktura; the author wrote the shorter
|
||||
// tail `[[notes/struktura]]`. Suffix matching connects them, while the
|
||||
// basename-only sibling `wiki/struktura` stays excluded.
|
||||
const resolver: SlugResolver = {
|
||||
resolve: async () => null,
|
||||
resolveBasenameMatches: async (name) =>
|
||||
name === 'struktura' ? ['vault/notes/struktura', 'wiki/struktura'] : [],
|
||||
};
|
||||
const { candidates } = await extractPageLinks(
|
||||
'concepts/x', 'See [[notes/struktura]].',
|
||||
{}, 'concept', resolver, { globalBasename: true },
|
||||
);
|
||||
expect(candidates.map(c => c.targetSlug)).toEqual(['vault/notes/struktura']);
|
||||
});
|
||||
|
||||
test('path-qualified self-link is dropped like the bare form', async () => {
|
||||
// `[[notes/struktura]]` written on notes/struktura itself must not
|
||||
// produce a self-loop (same guard as the bare `[[own-tail]]` case).
|
||||
const resolver: SlugResolver = {
|
||||
resolve: async () => null,
|
||||
resolveBasenameMatches: async (name) =>
|
||||
name === 'struktura' ? ['notes/struktura'] : [],
|
||||
};
|
||||
const { candidates } = await extractPageLinks(
|
||||
'notes/struktura', 'See [[notes/struktura]].',
|
||||
{}, 'concept', resolver, { globalBasename: true },
|
||||
);
|
||||
expect(candidates).toEqual([]);
|
||||
});
|
||||
|
||||
test('bare wikilink resolution does not interfere with DIR_PATTERN wikilinks', async () => {
|
||||
// 2b refs (people/alice) take the verb-inferred type;
|
||||
// 2c refs (struktura) take wikilink_basename. Same call.
|
||||
|
||||
@@ -1,153 +0,0 @@
|
||||
/**
|
||||
* `gbrain lint` source-glob filter — walker integration.
|
||||
*
|
||||
* PR #2157 (commit cf9a3b18, `feat/sync-source-glob-filters`) wired
|
||||
* `sources.config.include_globs` / `exclude_globs` into `gbrain sync` so a
|
||||
* user could exclude `Resources/veriff/**` and have every subsequent sync
|
||||
* honor it. The lint command walked the same source dirs blind and emitted
|
||||
* findings against paths the user had already declared out of scope — a
|
||||
* half-finished feature.
|
||||
*
|
||||
* This patch extends the same persisted glob contract to lint:
|
||||
* - `gbrain lint` gains `--include / --exclude` flags (parallel to sync).
|
||||
* - `runLintCore` lifts `sources.config.{include,exclude}_globs` for any
|
||||
* target whose absolute path matches a source row's `local_path`, so the
|
||||
* cycle.lint phase + Minion lint handlers honor the same filter without
|
||||
* restating it.
|
||||
* - The walker in `collectPages` applies the filter using the SAME
|
||||
* `matchesAnyGlob` helper sync uses, anchored at the target dir (so a
|
||||
* persisted `Resources/veriff/**` glob written against the source root
|
||||
* works without rewriting it as an absolute path).
|
||||
*
|
||||
* These tests pin the walker contract. The engine-side lift
|
||||
* (`resolveSourceGlobsForTarget`) is best-effort by design (returns `{}` on
|
||||
* any error) and is exercised by the dream-cycle lint phase end-to-end.
|
||||
*/
|
||||
|
||||
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, rmSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { tmpdir } from 'os';
|
||||
|
||||
// runLintCore is the library entry — the same surface the cycle.lint phase
|
||||
// and Minion handlers call. Exercising it covers the walker via its real
|
||||
// callsite; testing `collectPages` directly would skip the wiring.
|
||||
import { runLintCore } from '../src/commands/lint.ts';
|
||||
|
||||
// A self-contained content-sanity stub so the test never touches a real
|
||||
// engine / config file. Empty operator-literal list keeps the content-sanity
|
||||
// pass silent so the only findings come from the structural rules
|
||||
// (no-frontmatter etc.).
|
||||
const STUB_CS = {
|
||||
fail_on_throw: false,
|
||||
warn_on_throw: false,
|
||||
bytes_warn: 1024 * 1024,
|
||||
operator_literals: [],
|
||||
};
|
||||
|
||||
describe('runLintCore — source-glob walker filter', () => {
|
||||
let root: string;
|
||||
|
||||
beforeAll(() => {
|
||||
root = mkdtempSync(join(tmpdir(), 'gbrain-lint-globs-'));
|
||||
// Three subtrees with mixed structured / archive-style content.
|
||||
// All pages have `# Title` headers but no frontmatter so each one
|
||||
// emits at least one `no-frontmatter` issue under the default rule set.
|
||||
mkdirSync(join(root, 'Notes'), { recursive: true });
|
||||
mkdirSync(join(root, 'Resources', 'veriff'), { recursive: true });
|
||||
mkdirSync(join(root, 'Resources', 'prior-art', 'archive-v1'), { recursive: true });
|
||||
|
||||
writeFileSync(join(root, 'Notes', 'a.md'), '# A\nbody\n');
|
||||
writeFileSync(join(root, 'Notes', 'b.md'), '# B\nbody\n');
|
||||
writeFileSync(join(root, 'Resources', 'veriff', 'spec-1.md'), '# Veriff spec 1\nbody\n');
|
||||
writeFileSync(join(root, 'Resources', 'veriff', 'spec-2.md'), '# Veriff spec 2\nbody\n');
|
||||
writeFileSync(join(root, 'Resources', 'prior-art', 'archive-v1', 'old.md'), '# Old\nbody\n');
|
||||
});
|
||||
|
||||
afterAll(() => {
|
||||
rmSync(root, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
test('no filter — walks every .md (regression guard for default behavior)', async () => {
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
});
|
||||
expect(result.pages_scanned).toBe(5);
|
||||
expect(result.pages_with_issues).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
test('exclude glob skips matching paths (Resources/veriff/** off-limits)', async () => {
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
exclude: ['Resources/veriff/**'],
|
||||
});
|
||||
// 5 total minus 2 veriff specs = 3 pages walked.
|
||||
expect(result.pages_scanned).toBe(3);
|
||||
});
|
||||
|
||||
test('exclude with multiple patterns is union (veriff + prior-art both skipped)', async () => {
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
exclude: ['Resources/veriff/**', 'Resources/prior-art/**'],
|
||||
});
|
||||
// 5 total minus 3 (2 veriff + 1 archive-v1) = 2 pages walked.
|
||||
expect(result.pages_scanned).toBe(2);
|
||||
});
|
||||
|
||||
test('include glob narrows the walk to matching paths only', async () => {
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
include: ['Notes/**'],
|
||||
});
|
||||
expect(result.pages_scanned).toBe(2);
|
||||
});
|
||||
|
||||
test('exclude runs AFTER include (same precedence as `gbrain sync`)', async () => {
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
include: ['**/*.md'],
|
||||
exclude: ['Resources/**'],
|
||||
});
|
||||
// include lets everything through; exclude drops the 3 Resources/* files.
|
||||
expect(result.pages_scanned).toBe(2);
|
||||
});
|
||||
|
||||
test('empty include / exclude arrays do NOT engage the filter', async () => {
|
||||
// Symmetric with `parseGlobList` returning undefined for empty input —
|
||||
// an empty include would otherwise classify every path as a miss and
|
||||
// silently zero out the lint scope. Pin the guard at the walker level.
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
include: [],
|
||||
exclude: [],
|
||||
});
|
||||
expect(result.pages_scanned).toBe(5);
|
||||
});
|
||||
|
||||
test('exclude semantics match sync — `**` matches across path segments', async () => {
|
||||
const result = await runLintCore({
|
||||
target: root,
|
||||
contentSanity: STUB_CS,
|
||||
exclude: ['**/spec-*.md'],
|
||||
});
|
||||
// Both Veriff specs match the deep glob; Notes + archive-v1 survive.
|
||||
expect(result.pages_scanned).toBe(3);
|
||||
});
|
||||
|
||||
test('single-file target bypasses the filter (file mode is not a walk)', async () => {
|
||||
// A user lints one .md explicitly: filters are a directory-walk concern,
|
||||
// so the file is processed even if its name would match an exclude.
|
||||
const result = await runLintCore({
|
||||
target: join(root, 'Resources', 'veriff', 'spec-1.md'),
|
||||
contentSanity: STUB_CS,
|
||||
exclude: ['Resources/veriff/**'],
|
||||
});
|
||||
expect(result.pages_scanned).toBe(1);
|
||||
});
|
||||
});
|
||||
@@ -302,4 +302,29 @@ describe('loadConfigWithEngine (Phase 4 / F3)', () => {
|
||||
expect(merged?.engine).toBe('pglite');
|
||||
});
|
||||
});
|
||||
|
||||
describe('eval.* DB-plane merge (#1475)', () => {
|
||||
test('gbrain config set eval.capture true reaches the merged config', async () => {
|
||||
// The #1475 repro: DB plane has eval.capture=true, file plane silent.
|
||||
// Pre-fix the merge skipped eval.* entirely and capture never fired.
|
||||
const base: GBrainConfig = { engine: 'pglite' };
|
||||
const engine = makeEngine({ 'eval.capture': 'true', 'eval.scrub_pii': 'false' });
|
||||
const merged = await loadConfigWithEngine(engine, base);
|
||||
expect(merged?.eval?.capture).toBe(true);
|
||||
expect(merged?.eval?.scrub_pii).toBe(false);
|
||||
});
|
||||
|
||||
test('file plane wins per key; DB fills only the gaps', async () => {
|
||||
const base: GBrainConfig = { engine: 'pglite', eval: { capture: false } };
|
||||
const engine = makeEngine({ 'eval.capture': 'true', 'eval.scrub_pii': 'false' });
|
||||
const merged = await loadConfigWithEngine(engine, base);
|
||||
expect(merged?.eval?.capture).toBe(false); // file wins
|
||||
expect(merged?.eval?.scrub_pii).toBe(false); // DB fills the gap
|
||||
});
|
||||
|
||||
test('no eval keys anywhere leaves cfg.eval undefined', async () => {
|
||||
const merged = await loadConfigWithEngine(makeEngine({}), { engine: 'pglite' });
|
||||
expect(merged?.eval).toBeUndefined();
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
import { describe, expect, test } from 'bun:test';
|
||||
import {
|
||||
extractCycleFreshnessSourceIds,
|
||||
parseMaintainArgs,
|
||||
} from '../src/commands/maintain.ts';
|
||||
import type { Check } from '../src/commands/doctor.ts';
|
||||
|
||||
describe('maintain args', () => {
|
||||
test('defaults to dry-run unless --safe is explicit', () => {
|
||||
expect(parseMaintainArgs([])).toMatchObject({
|
||||
safe: false,
|
||||
dryRun: true,
|
||||
json: false,
|
||||
});
|
||||
});
|
||||
|
||||
test('--safe enables mutating safe mode', () => {
|
||||
expect(parseMaintainArgs(['--safe', '--json'])).toMatchObject({
|
||||
safe: true,
|
||||
dryRun: false,
|
||||
json: true,
|
||||
});
|
||||
});
|
||||
|
||||
test('--dry-run wins over --safe', () => {
|
||||
expect(parseMaintainArgs(['--safe', '--dry-run'])).toMatchObject({
|
||||
safe: true,
|
||||
dryRun: true,
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe('cycle freshness source extraction', () => {
|
||||
test('extracts stale source ids from doctor messages', () => {
|
||||
const checks: Check[] = [
|
||||
{
|
||||
name: 'cycle_freshness',
|
||||
status: 'fail',
|
||||
message: "Source 'brain-sync-remote-teffur' last cycled 40h ago. Run `gbrain dream --source <id>`.",
|
||||
},
|
||||
{
|
||||
name: 'cycle_freshness',
|
||||
status: 'fail',
|
||||
message: "Source 'wiki' last cycled 25h ago. Source 'wiki' last cycled 25h ago.",
|
||||
},
|
||||
];
|
||||
|
||||
expect(extractCycleFreshnessSourceIds(checks)).toEqual([
|
||||
'brain-sync-remote-teffur',
|
||||
'wiki',
|
||||
]);
|
||||
});
|
||||
|
||||
test('ignores ok and unrelated checks', () => {
|
||||
const checks: Check[] = [
|
||||
{ name: 'cycle_freshness', status: 'ok', message: "Source 'fresh' last cycled recently." },
|
||||
{ name: 'frontmatter_integrity', status: 'warn', message: "Source 'wiki' has frontmatter issues." },
|
||||
];
|
||||
|
||||
expect(extractCycleFreshnessSourceIds(checks)).toEqual([]);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,17 @@
|
||||
import { describe, expect, it } from 'bun:test';
|
||||
import { readFileSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
|
||||
describe('root OpenClaw plugin manifest', () => {
|
||||
it('declares the id required by OpenClaw plugin installs', () => {
|
||||
const manifest = JSON.parse(readFileSync(join(import.meta.dir, '..', 'openclaw.plugin.json'), 'utf8'));
|
||||
const entrySource = readFileSync(join(import.meta.dir, '..', 'src', 'openclaw-context-engine.ts'), 'utf8');
|
||||
const entryId = entrySource.match(/id:\s*'([^']+)'/)?.[1];
|
||||
|
||||
expect(manifest.id).toBe(entryId);
|
||||
expect(manifest.configSchema).toBeDefined();
|
||||
expect(typeof manifest.configSchema).toBe('object');
|
||||
expect(manifest.contracts?.contextEngines).toContain('gbrain-context');
|
||||
expect(entrySource).toContain('export function register');
|
||||
});
|
||||
});
|
||||
@@ -186,11 +186,67 @@ describe('shouldExclude — orphan filter regression (preserve curation)', () =>
|
||||
expect(shouldExclude('entities/anonymous')).toBe(true);
|
||||
expect(shouldExclude('atoms/fact-123')).toBe(true);
|
||||
expect(shouldExclude('skills/gbrain-operations')).toBe(true);
|
||||
expect(shouldExclude('dreaming/light/2026-07-20')).toBe(true);
|
||||
expect(shouldExclude('daily/2026-07-20')).toBe(true);
|
||||
expect(shouldExclude('agent-openclaw/daily/2026-07-20')).toBe(true);
|
||||
});
|
||||
|
||||
test('workspace convention slugs are excluded', () => {
|
||||
expect(shouldExclude('_brain-conventions')).toBe(true);
|
||||
expect(shouldExclude('_templates/decision')).toBe(true);
|
||||
expect(shouldExclude('extracts/2026-06-30/takes.proposed/round-single')).toBe(true);
|
||||
expect(shouldExclude('2026-07-20')).toBe(true);
|
||||
expect(shouldExclude('2026-07-20-qa-sweep')).toBe(true);
|
||||
expect(shouldExclude('agents/arya/identity')).toBe(true);
|
||||
expect(shouldExclude('agents/arya/memory/dreaming/deep/2026-07-20')).toBe(true);
|
||||
});
|
||||
|
||||
test('regular slugs are NOT excluded', () => {
|
||||
expect(shouldExclude('people/alice')).toBe(false);
|
||||
expect(shouldExclude('companies/acme')).toBe(false);
|
||||
expect(shouldExclude('writing/post-1')).toBe(false);
|
||||
expect(shouldExclude('agents/arya/qa-reports/launch-review')).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('getHealth orphan_pages uses shared exclusion policy', () => {
|
||||
test('excluded convention islands do not count against health', async () => {
|
||||
await engine.putPage('_templates/decision', {
|
||||
type: 'template', title: 'Decision', compiled_truth: 'template', timeline: '', frontmatter: {},
|
||||
});
|
||||
await engine.putPage('skills/arya/source-check', {
|
||||
type: 'concept', title: 'Skill', compiled_truth: 'skill', timeline: '', frontmatter: {},
|
||||
});
|
||||
await engine.putPage('agents/arya/identity', {
|
||||
type: 'note', title: 'Identity', compiled_truth: 'identity', timeline: '', frontmatter: {},
|
||||
});
|
||||
await engine.putPage('people/alice', {
|
||||
type: 'person', title: 'Alice', compiled_truth: 'real island', timeline: '', frontmatter: {},
|
||||
});
|
||||
|
||||
const health = await engine.getHealth();
|
||||
|
||||
expect(health.orphan_pages).toBe(1);
|
||||
});
|
||||
|
||||
test('per-brain config overrides (orphans.exclude_*) also apply to health', async () => {
|
||||
await engine.putPage('my-private-folder/secret-ref', {
|
||||
type: 'note', title: 'Ref', compiled_truth: 'ref', timeline: '', frontmatter: {},
|
||||
});
|
||||
await engine.putPage('one-off-fixture-page', {
|
||||
type: 'note', title: 'Fixture', compiled_truth: 'fixture', timeline: '', frontmatter: {},
|
||||
});
|
||||
await engine.putPage('people/alice', {
|
||||
type: 'person', title: 'Alice', compiled_truth: 'real island', timeline: '', frontmatter: {},
|
||||
});
|
||||
|
||||
expect((await engine.getHealth()).orphan_pages).toBe(3);
|
||||
|
||||
await engine.setConfig('orphans.exclude_prefixes', 'my-private-folder/');
|
||||
await engine.setConfig('orphans.exclude_slugs', 'one-off-fixture-page');
|
||||
expect((await engine.getHealth()).orphan_pages).toBe(1);
|
||||
|
||||
await engine.unsetConfig('orphans.exclude_prefixes');
|
||||
await engine.unsetConfig('orphans.exclude_slugs');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -66,6 +66,10 @@ describe('shouldExclude', () => {
|
||||
expect(shouldExclude('templates/meeting-note')).toBe(true);
|
||||
});
|
||||
|
||||
test('excludes deny-prefix: _templates/', () => {
|
||||
expect(shouldExclude('_templates/meeting-note')).toBe(true);
|
||||
});
|
||||
|
||||
test('excludes deny-prefix: openclaw/config/', () => {
|
||||
expect(shouldExclude('openclaw/config/agent')).toBe(true);
|
||||
});
|
||||
@@ -86,10 +90,44 @@ describe('shouldExclude', () => {
|
||||
expect(shouldExclude('entities/product-hunt')).toBe(true);
|
||||
});
|
||||
|
||||
test('excludes first-segment: skills, dreaming, and daily', () => {
|
||||
expect(shouldExclude('skills/arya/source-check')).toBe(true);
|
||||
expect(shouldExclude('dreaming/light/2026-07-20')).toBe(true);
|
||||
expect(shouldExclude('daily/2026-07-20')).toBe(true);
|
||||
expect(shouldExclude('agent-openclaw/daily/2026-07-20')).toBe(true);
|
||||
});
|
||||
|
||||
test('excludes root date logs and agent workspace conventions', () => {
|
||||
expect(shouldExclude('_brain-conventions')).toBe(true);
|
||||
expect(shouldExclude('2026-07-20')).toBe(true);
|
||||
expect(shouldExclude('2026-07-20-qa-sweep')).toBe(true);
|
||||
expect(shouldExclude('agents/arya/identity')).toBe(true);
|
||||
expect(shouldExclude('agents/arya/memory/dreaming/deep/2026-07-20')).toBe(true);
|
||||
});
|
||||
|
||||
test('excludes generated extracts', () => {
|
||||
expect(shouldExclude('extracts/2026-06-30/takes.proposed/round-single')).toBe(true);
|
||||
});
|
||||
|
||||
test('brain-specific exclusions come from config overrides, not global defaults', () => {
|
||||
// No baked-in defaults for these:
|
||||
expect(shouldExclude('my-private-folder/some-secret-ref.md')).toBe(false);
|
||||
expect(shouldExclude('one-off-fixture-page')).toBe(false);
|
||||
// The per-brain config plane (orphans.exclude_prefixes / exclude_slugs):
|
||||
const overrides = {
|
||||
excludePrefixes: ['my-private-folder/'],
|
||||
excludeSlugs: ['one-off-fixture-page'],
|
||||
};
|
||||
expect(shouldExclude('my-private-folder/some-secret-ref.md', overrides)).toBe(true);
|
||||
expect(shouldExclude('one-off-fixture-page', overrides)).toBe(true);
|
||||
expect(shouldExclude('people/jane-doe', overrides)).toBe(false);
|
||||
});
|
||||
|
||||
test('does NOT exclude a normal content page', () => {
|
||||
expect(shouldExclude('companies/acme')).toBe(false);
|
||||
expect(shouldExclude('people/jane-doe')).toBe(false);
|
||||
expect(shouldExclude('projects/gbrain')).toBe(false);
|
||||
expect(shouldExclude('agents/arya/qa-reports/launch-review')).toBe(false);
|
||||
});
|
||||
|
||||
test('does NOT exclude a page ending with log-like text that is not /log', () => {
|
||||
|
||||
@@ -218,15 +218,21 @@ describe('progress reporter', () => {
|
||||
test('only one process-level signal handler installed across many reporters', () => {
|
||||
// Baseline: one handler already installed by prior tests in this file.
|
||||
const installedBefore = __signalHandlerInstalledForTest();
|
||||
// liveReporters is process-global: earlier test files in the same shard
|
||||
// can leave a live entry behind (e.g. a production path that skips
|
||||
// finish() on an error branch). Assert NET-zero leak from THIS test's
|
||||
// lifecycles, not an absolute zero we don't control — same tolerance
|
||||
// the handler assertion below already applies via `installedBefore`.
|
||||
const liveBefore = __liveReporterCountForTest();
|
||||
const { stream } = sink(false);
|
||||
for (let i = 0; i < 50; i++) {
|
||||
const p = createProgress({ mode: 'json', stream, minIntervalMs: 0, minItems: 1 });
|
||||
p.start(`phase_${i}`, 1);
|
||||
p.finish();
|
||||
}
|
||||
// After 50 reporter lifecycles, still exactly one handler and zero leaked live entries.
|
||||
// After 50 reporter lifecycles, still exactly one handler and zero NEWLY leaked live entries.
|
||||
expect(__signalHandlerInstalledForTest()).toBe(installedBefore || true);
|
||||
expect(__liveReporterCountForTest()).toBe(0);
|
||||
expect(__liveReporterCountForTest()).toBe(liveBefore);
|
||||
});
|
||||
|
||||
test('startHeartbeat() fires heartbeats and stop() clears', async () => {
|
||||
|
||||
@@ -694,12 +694,6 @@ const COLUMN_EXEMPTIONS = new Set<string>([
|
||||
'minion_jobs.quiet_hours',
|
||||
'minion_jobs.stagger_key',
|
||||
'sources.chunker_version',
|
||||
// #2157 follow-on (migration v125). TEXT column read by performSync's
|
||||
// `Already up to date` gate; not referenced by any CREATE INDEX. Same
|
||||
// upgrade-path coverage as sources.chunker_version above: fresh installs
|
||||
// get it via the CREATE TABLE in src/schema.sql + schema-embedded.ts;
|
||||
// pre-existing brains get it via the idempotent ALTER TABLE in v125.
|
||||
'sources.config_fingerprint',
|
||||
'access_tokens.permissions',
|
||||
'takes.resolved_quality',
|
||||
'pages.emotional_weight_recomputed_at',
|
||||
|
||||
@@ -12,9 +12,11 @@ import {
|
||||
bestPath,
|
||||
historyPath,
|
||||
loadHistory,
|
||||
proposedPath,
|
||||
revertAllPending,
|
||||
skillPath,
|
||||
versionsDir,
|
||||
writeProposed,
|
||||
} from '../../src/core/skillopt/version-store.ts';
|
||||
|
||||
let tmpDir: string;
|
||||
@@ -79,6 +81,19 @@ describe('acceptCandidate (D8 two-phase commit)', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('writeProposed', () => {
|
||||
test('writes distinct best and proposed artifacts without mutating SKILL.md (#2635)', () => {
|
||||
const candidate = '---\nname: test\n---\nproposed body\n';
|
||||
|
||||
const written = writeProposed(tmpDir, SKILL, candidate);
|
||||
|
||||
expect(written).toBe(proposedPath(tmpDir, SKILL));
|
||||
expect(fs.readFileSync(bestPath(tmpDir, SKILL), 'utf8')).toBe(candidate);
|
||||
expect(fs.readFileSync(proposedPath(tmpDir, SKILL), 'utf8')).toBe(candidate);
|
||||
expect(fs.readFileSync(skillPath(tmpDir, SKILL), 'utf8')).toContain('baseline body');
|
||||
});
|
||||
});
|
||||
|
||||
describe('revertAllPending (D8 crash recovery)', () => {
|
||||
test('no-op when no pending rows', () => {
|
||||
const reverted = revertAllPending(tmpDir, SKILL);
|
||||
|
||||
@@ -159,127 +159,6 @@ describe('sources add', () => {
|
||||
await expect(runSources(engine, ['add', 'plans', '--path', '/tmp/gstack/plans']))
|
||||
.rejects.toThrow(/overlaps with existing source "gstack"/);
|
||||
});
|
||||
|
||||
// Glob filters — TODO #3 from the brettdavies fork recon. Pre-fix, the
|
||||
// `SyncableOptions` shape in `src/core/sync.ts` had been carrying
|
||||
// `include` / `exclude` since v0.41.13, but commands/sync.ts:1454 never
|
||||
// populated them and `sources add` had no flag to persist them — so users
|
||||
// had no way to tell gbrain to skip `Templates/` in an Obsidian vault.
|
||||
test('--exclude persists glob into sources.config.exclude_globs', async () => {
|
||||
const { engine, calls } = makeStub({
|
||||
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
|
||||
id: 'vault',
|
||||
name: 'vault',
|
||||
local_path: '/tmp/vault',
|
||||
last_commit: null,
|
||||
last_sync_at: null,
|
||||
config: '{"exclude_globs":["Templates/**"]}',
|
||||
created_at: new Date(),
|
||||
}],
|
||||
});
|
||||
await runSources(engine, ['add', 'vault', '--path', '/tmp/vault', '--exclude', 'Templates/**']);
|
||||
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
|
||||
expect(insert!.params[3]).toBe('{"exclude_globs":["Templates/**"]}');
|
||||
});
|
||||
|
||||
test('--include persists glob into sources.config.include_globs', async () => {
|
||||
const { engine, calls } = makeStub({
|
||||
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
|
||||
id: 'wiki',
|
||||
name: 'wiki',
|
||||
local_path: '/tmp/wiki',
|
||||
last_commit: null,
|
||||
last_sync_at: null,
|
||||
config: '{"include_globs":["people/**"]}',
|
||||
created_at: new Date(),
|
||||
}],
|
||||
});
|
||||
await runSources(engine, ['add', 'wiki', '--path', '/tmp/wiki', '--include', 'people/**']);
|
||||
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
|
||||
expect(insert!.params[3]).toBe('{"include_globs":["people/**"]}');
|
||||
});
|
||||
|
||||
test('--exclude is repeatable; preserves order', async () => {
|
||||
const { engine, calls } = makeStub({
|
||||
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
|
||||
id: 'vault',
|
||||
name: 'vault',
|
||||
local_path: '/tmp/vault',
|
||||
last_commit: null,
|
||||
last_sync_at: null,
|
||||
config: '{}',
|
||||
created_at: new Date(),
|
||||
}],
|
||||
});
|
||||
await runSources(engine, [
|
||||
'add', 'vault', '--path', '/tmp/vault',
|
||||
'--exclude', 'Templates/**',
|
||||
'--exclude', '.smart-env/**',
|
||||
'--exclude', 'Drafts/**',
|
||||
]);
|
||||
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
|
||||
expect(insert!.params[3]).toBe('{"exclude_globs":["Templates/**",".smart-env/**","Drafts/**"]}');
|
||||
});
|
||||
|
||||
test('--include and --exclude compose in one command (federated source with both filter axes)', async () => {
|
||||
const { engine, calls } = makeStub({
|
||||
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
|
||||
id: 'vault',
|
||||
name: 'vault',
|
||||
local_path: '/tmp/vault',
|
||||
last_commit: null,
|
||||
last_sync_at: null,
|
||||
config: '{"federated":true,"include_globs":["people/**"],"exclude_globs":["Templates/**"]}',
|
||||
created_at: new Date(),
|
||||
}],
|
||||
});
|
||||
await runSources(engine, [
|
||||
'add', 'vault', '--path', '/tmp/vault', '--federated',
|
||||
'--include', 'people/**',
|
||||
'--exclude', 'Templates/**',
|
||||
]);
|
||||
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
|
||||
expect(insert!.params[3]).toBe(
|
||||
'{"federated":true,"include_globs":["people/**"],"exclude_globs":["Templates/**"]}',
|
||||
);
|
||||
});
|
||||
|
||||
test('omitted glob flags leave config untouched (no [] entries persisted)', async () => {
|
||||
// Regression guard: empty glob arrays must NOT be written. Otherwise a
|
||||
// brain that never opts into filtering grows {"include_globs": [],
|
||||
// "exclude_globs": []} cruft in every source row, and the parseGlobList
|
||||
// path would return undefined anyway (the cruft is purely noise).
|
||||
const { engine, calls } = makeStub({
|
||||
'SELECT id, name, local_path, last_commit, last_sync_at, config, created_at': [{
|
||||
id: 'gstack',
|
||||
name: 'gstack',
|
||||
local_path: '/tmp/gstack',
|
||||
last_commit: null,
|
||||
last_sync_at: null,
|
||||
config: '{}',
|
||||
created_at: new Date(),
|
||||
}],
|
||||
});
|
||||
await runSources(engine, ['add', 'gstack', '--path', '/tmp/gstack']);
|
||||
const insert = calls.find(c => c.sql.includes('INSERT INTO sources'));
|
||||
expect(insert!.params[3]).toBe('{}');
|
||||
});
|
||||
|
||||
test('--exclude requires a glob argument', async () => {
|
||||
const { engine } = makeStub();
|
||||
const code = await withExitCapture(() => runSources(engine, [
|
||||
'add', 'vault', '--path', '/tmp/vault', '--exclude',
|
||||
]));
|
||||
expect(code).toBe(2);
|
||||
});
|
||||
|
||||
test('--include rejects a flag-like value (--include --path looks like a typo)', async () => {
|
||||
const { engine } = makeStub();
|
||||
const code = await withExitCapture(() => runSources(engine, [
|
||||
'add', 'vault', '--path', '/tmp/vault', '--include', '--federated',
|
||||
]));
|
||||
expect(code).toBe(2);
|
||||
});
|
||||
});
|
||||
|
||||
// ── add — #2707 git-repo validation (CLI wiring) ───────────────
|
||||
|
||||
@@ -1,95 +0,0 @@
|
||||
/**
|
||||
* #2157 follow-on (migration v125) — end-to-end gate wiring.
|
||||
*
|
||||
* `test/sync-config-fingerprint.test.ts` pins the persistence + comparison
|
||||
* primitives (compute/read/write). This file pins the WIRING inside
|
||||
* `performSync`: with git HEAD unchanged, a drift in the walk-affecting
|
||||
* `sources.config` fields must break out of the "Already up to date" early
|
||||
* return and force a full re-walk — and the re-stamped fingerprint must
|
||||
* settle the gate back to `up_to_date` on the following pass. Deleting the
|
||||
* `configMismatch` term from the gate condition fails this test; none of the
|
||||
* primitive tests would catch that.
|
||||
*/
|
||||
|
||||
import { test, expect, beforeAll, afterAll } from 'bun:test';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, rmSync } from 'fs';
|
||||
import { join } from 'path';
|
||||
import { tmpdir } from 'os';
|
||||
import { execFileSync } from 'child_process';
|
||||
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
||||
import { performSync } from '../src/commands/sync.ts';
|
||||
|
||||
let engine: PGLiteEngine;
|
||||
let repoPath: string;
|
||||
|
||||
function git(cwd: string, ...args: string[]) {
|
||||
execFileSync('git', args, { cwd, stdio: 'pipe' });
|
||||
}
|
||||
|
||||
beforeAll(async () => {
|
||||
engine = new PGLiteEngine();
|
||||
await engine.connect({});
|
||||
await engine.initSchema();
|
||||
|
||||
repoPath = mkdtempSync(join(tmpdir(), 'gbrain-fp-gate-'));
|
||||
mkdirSync(join(repoPath, 'wiki'));
|
||||
mkdirSync(join(repoPath, 'memory'));
|
||||
writeFileSync(join(repoPath, 'wiki', 'page1.md'), '# Page 1\n\nbody\n');
|
||||
writeFileSync(join(repoPath, 'memory', 'note1.md'), '# Note 1\n\nbody\n');
|
||||
git(repoPath, 'init');
|
||||
git(repoPath, 'add', '-A');
|
||||
git(repoPath, '-c', 'user.email=t@example.com', '-c', 'user.name=t', 'commit', '-m', 'init');
|
||||
|
||||
await engine.executeRaw(
|
||||
`INSERT INTO sources (id, name, local_path, config) VALUES ($1, $2, $3, $4::text::jsonb)`,
|
||||
['vault', 'vault', repoPath, JSON.stringify({ include_globs: ['wiki/**'] })],
|
||||
);
|
||||
}, 60_000);
|
||||
|
||||
afterAll(async () => {
|
||||
await engine?.disconnect();
|
||||
rmSync(repoPath, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
test('config-glob drift with unchanged git HEAD forces a re-walk, then settles', async () => {
|
||||
// First sync: row config include_globs = ['wiki/**'], caller threads it
|
||||
// (as syncOneSource / the single-source CLI path do). memory/* skipped.
|
||||
const first = await performSync(engine, {
|
||||
repoPath, sourceId: 'vault', include: ['wiki/**'],
|
||||
noPull: true, noEmbed: true, full: true,
|
||||
});
|
||||
expect(first.status).toBe('first_sync');
|
||||
expect(await engine.getPage('wiki/page1')).not.toBeNull();
|
||||
expect(await engine.getPage('memory/note1')).toBeNull();
|
||||
|
||||
// No drift, HEAD unchanged: gate stays quiet.
|
||||
const second = await performSync(engine, {
|
||||
repoPath, sourceId: 'vault', include: ['wiki/**'],
|
||||
noPull: true, noEmbed: true,
|
||||
});
|
||||
expect(second.status).toBe('up_to_date');
|
||||
|
||||
// User widens the persisted globs (what `gbrain sources add --include`
|
||||
// writes). Git HEAD has NOT moved.
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[JSON.stringify({ include_globs: ['wiki/**', 'memory/**'] }), 'vault'],
|
||||
);
|
||||
|
||||
// Pre-fix this returned `up_to_date` (HEAD unchanged) and memory/note1
|
||||
// stayed missing until a manual `--full`. The fingerprint gate must force
|
||||
// the full re-walk instead.
|
||||
const third = await performSync(engine, {
|
||||
repoPath, sourceId: 'vault', include: ['wiki/**', 'memory/**'],
|
||||
noPull: true, noEmbed: true,
|
||||
});
|
||||
expect(third.status).not.toBe('up_to_date');
|
||||
expect(await engine.getPage('memory/note1')).not.toBeNull();
|
||||
|
||||
// Re-stamped fingerprint matches the current row: gate settles.
|
||||
const fourth = await performSync(engine, {
|
||||
repoPath, sourceId: 'vault', include: ['wiki/**', 'memory/**'],
|
||||
noPull: true, noEmbed: true,
|
||||
});
|
||||
expect(fourth.status).toBe('up_to_date');
|
||||
});
|
||||
@@ -1,234 +0,0 @@
|
||||
/**
|
||||
* #2157 follow-on (migration v125 — `sources.config_fingerprint`).
|
||||
*
|
||||
* The "Already up to date" gate at performSync's git-HEAD equality check
|
||||
* honored chunker_version match but ignored `sources.config` drift.
|
||||
* Changing `sources.config.exclude_globs` (or include_globs / strategy)
|
||||
* had no observable effect on the next sync because git HEAD was
|
||||
* unchanged — the gate returned early and the new walk scope never
|
||||
* applied. This file exercises the persistence shape + drift detection
|
||||
* end-to-end on PGLite, including:
|
||||
*
|
||||
* - Migration v125 actually adds the column (regression guard against
|
||||
* a future re-numbering or accidental deletion).
|
||||
* - read/write round-trips preserve the value.
|
||||
* - The fingerprint differs across the three walk-affecting fields
|
||||
* and is order-insensitive on the array fields.
|
||||
* - NULL fingerprint on pre-v125 rows treats as "not stamped" so a
|
||||
* first post-upgrade sync doesn't spuriously force-full.
|
||||
* - A toggle-and-revert leaves the stored fingerprint matching the
|
||||
* current row, so the gate stays quiet.
|
||||
*
|
||||
* The wired-up gate behavior (force-full triggered on mismatch) is
|
||||
* exercised by the existing sync end-to-end tests; here we pin the
|
||||
* persistence + comparison primitives the gate depends on.
|
||||
*/
|
||||
|
||||
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
||||
import {
|
||||
computeSourceConfigFingerprint,
|
||||
readConfigFingerprint,
|
||||
writeConfigFingerprint,
|
||||
} from '../src/commands/sync.ts';
|
||||
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
||||
|
||||
let engine: PGLiteEngine;
|
||||
|
||||
beforeAll(async () => {
|
||||
engine = new PGLiteEngine();
|
||||
await engine.connect({});
|
||||
await engine.initSchema();
|
||||
}, 60_000);
|
||||
|
||||
afterAll(async () => {
|
||||
await engine?.disconnect();
|
||||
});
|
||||
|
||||
/** Insert a fresh source row with the given config. Returns the id. */
|
||||
async function makeSource(
|
||||
id: string,
|
||||
config: Record<string, unknown> = {},
|
||||
): Promise<string> {
|
||||
await engine.executeRaw(
|
||||
`INSERT INTO sources (id, name, local_path, config) VALUES ($1, $2, $3, $4::text::jsonb)`,
|
||||
[id, id, `/tmp/${id}`, JSON.stringify(config)],
|
||||
);
|
||||
return id;
|
||||
}
|
||||
|
||||
describe('migration v125 — sources.config_fingerprint column', () => {
|
||||
test('column exists on the sources table', async () => {
|
||||
const rows = await engine.executeRaw<{ column_name: string }>(
|
||||
`SELECT column_name FROM information_schema.columns
|
||||
WHERE table_name = 'sources' AND column_name = 'config_fingerprint'`,
|
||||
);
|
||||
expect(rows).toHaveLength(1);
|
||||
});
|
||||
|
||||
test('column is nullable (preserves pre-migration row semantics)', async () => {
|
||||
const rows = await engine.executeRaw<{ is_nullable: string }>(
|
||||
`SELECT is_nullable FROM information_schema.columns
|
||||
WHERE table_name = 'sources' AND column_name = 'config_fingerprint'`,
|
||||
);
|
||||
expect(rows[0]?.is_nullable).toBe('YES');
|
||||
});
|
||||
});
|
||||
|
||||
describe('readConfigFingerprint / writeConfigFingerprint — persistence round-trip', () => {
|
||||
test('round-trip: write then read returns the same value', async () => {
|
||||
const id = await makeSource('rt-basic', { exclude_globs: ['Templates/**'] });
|
||||
const fp = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
|
||||
await writeConfigFingerprint(engine, id, fp);
|
||||
const got = await readConfigFingerprint(engine, id);
|
||||
expect(got).toBe(fp);
|
||||
});
|
||||
|
||||
test('NULL on never-stamped row (pre-v125 semantics)', async () => {
|
||||
const id = await makeSource('rt-never');
|
||||
const got = await readConfigFingerprint(engine, id);
|
||||
expect(got).toBeNull();
|
||||
});
|
||||
|
||||
test('undefined sourceId returns null (legacy non-source-scoped sync)', async () => {
|
||||
const got = await readConfigFingerprint(engine, undefined);
|
||||
expect(got).toBeNull();
|
||||
});
|
||||
|
||||
test('write with undefined sourceId is a no-op (does not throw)', async () => {
|
||||
// The legacy global-sync code path hits this branch; the guard must
|
||||
// be silent rather than fail the sync run.
|
||||
await writeConfigFingerprint(engine, undefined, 'deadbeef'.repeat(8));
|
||||
// No assertion beyond "did not throw"; the function returns void.
|
||||
});
|
||||
|
||||
test('overwrite: a second write replaces the prior fingerprint', async () => {
|
||||
const id = await makeSource('rt-overwrite');
|
||||
await writeConfigFingerprint(engine, id, 'a'.repeat(64));
|
||||
await writeConfigFingerprint(engine, id, 'b'.repeat(64));
|
||||
const got = await readConfigFingerprint(engine, id);
|
||||
expect(got).toBe('b'.repeat(64));
|
||||
});
|
||||
});
|
||||
|
||||
describe('end-to-end drift simulation — the gate semantics this column enables', () => {
|
||||
test('first stamp matches computed fingerprint of the row config', async () => {
|
||||
const cfg = { exclude_globs: ['Templates/**', 'Photos/**'], strategy: 'markdown' };
|
||||
const id = await makeSource('e2e-first-stamp', cfg);
|
||||
const computed = computeSourceConfigFingerprint(cfg);
|
||||
await writeConfigFingerprint(engine, id, computed);
|
||||
expect(await readConfigFingerprint(engine, id)).toBe(computed);
|
||||
});
|
||||
|
||||
test('exclude_globs mutation makes stored != current (drift detected)', async () => {
|
||||
const before = { exclude_globs: ['Templates/**'] };
|
||||
const after = { exclude_globs: ['Templates/**', 'Photos/**'] };
|
||||
const id = await makeSource('e2e-exclude-drift', before);
|
||||
const beforeFp = computeSourceConfigFingerprint(before);
|
||||
await writeConfigFingerprint(engine, id, beforeFp);
|
||||
|
||||
// Simulate the user mutating sources.config via `gbrain sources add
|
||||
// --exclude`. The gate's next read of (stored, computed-from-current)
|
||||
// detects the drift and forces a re-walk.
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[JSON.stringify(after), id],
|
||||
);
|
||||
const afterFp = computeSourceConfigFingerprint(after);
|
||||
const stored = await readConfigFingerprint(engine, id);
|
||||
expect(stored).toBe(beforeFp);
|
||||
expect(stored).not.toBe(afterFp);
|
||||
});
|
||||
|
||||
test('toggle-and-revert: add then remove same pattern leaves stored matching current', async () => {
|
||||
const original = { exclude_globs: ['Templates/**'] };
|
||||
const id = await makeSource('e2e-toggle', original);
|
||||
const originalFp = computeSourceConfigFingerprint(original);
|
||||
await writeConfigFingerprint(engine, id, originalFp);
|
||||
|
||||
// Add a pattern (drift) then remove it (revert).
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[JSON.stringify({ exclude_globs: ['Templates/**', 'Photos/**'] }), id],
|
||||
);
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[JSON.stringify(original), id],
|
||||
);
|
||||
const revertedFp = computeSourceConfigFingerprint(original);
|
||||
expect(revertedFp).toBe(originalFp);
|
||||
expect(await readConfigFingerprint(engine, id)).toBe(originalFp);
|
||||
// ⇒ Gate compares storedFp (==originalFp) to currentFp (==originalFp): no drift, no force-full.
|
||||
});
|
||||
|
||||
test('include_globs drift detected independently', async () => {
|
||||
const before = { include_globs: ['people/**'] };
|
||||
const after = { include_globs: ['people/**', 'companies/**'] };
|
||||
const id = await makeSource('e2e-include-drift', before);
|
||||
await writeConfigFingerprint(engine, id, computeSourceConfigFingerprint(before));
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[JSON.stringify(after), id],
|
||||
);
|
||||
const stored = await readConfigFingerprint(engine, id);
|
||||
const current = computeSourceConfigFingerprint(after);
|
||||
expect(stored).not.toBe(current);
|
||||
});
|
||||
|
||||
test('strategy drift detected', async () => {
|
||||
const before = { strategy: 'markdown' };
|
||||
const after = { strategy: 'code' };
|
||||
const id = await makeSource('e2e-strategy-drift', before);
|
||||
await writeConfigFingerprint(engine, id, computeSourceConfigFingerprint(before));
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[JSON.stringify(after), id],
|
||||
);
|
||||
expect(await readConfigFingerprint(engine, id))
|
||||
.not.toBe(computeSourceConfigFingerprint(after));
|
||||
});
|
||||
|
||||
test('mutating unrelated config field (federated) does NOT drift', async () => {
|
||||
// The fingerprint hashes ONLY walk-affecting fields. Federation
|
||||
// changes search visibility, not the walk set — must not invalidate
|
||||
// the checkpoint.
|
||||
const id = await makeSource('e2e-federated-toggle', {
|
||||
federated: true,
|
||||
exclude_globs: ['Templates/**'],
|
||||
});
|
||||
await writeConfigFingerprint(
|
||||
engine,
|
||||
id,
|
||||
computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] }),
|
||||
);
|
||||
await engine.executeRaw(
|
||||
`UPDATE sources SET config = $1::text::jsonb WHERE id = $2`,
|
||||
[
|
||||
JSON.stringify({ federated: false, exclude_globs: ['Templates/**'] }),
|
||||
id,
|
||||
],
|
||||
);
|
||||
const stored = await readConfigFingerprint(engine, id);
|
||||
const current = computeSourceConfigFingerprint({
|
||||
federated: false,
|
||||
exclude_globs: ['Templates/**'],
|
||||
});
|
||||
expect(stored).toBe(current);
|
||||
});
|
||||
|
||||
test('double-encoded JSONB config (the sources-add stringify bug) hashes equivalently to the parsed object', async () => {
|
||||
// `gbrain sources add` writes `JSON.stringify(config)::jsonb`, which
|
||||
// double-encodes the value into a JSON-string scalar (`"{\"x\":1}"`)
|
||||
// rather than a proper JSONB object. The defensive reader in
|
||||
// postgres-engine.ts:1274 + readSourceConfig parses the string back
|
||||
// before the fingerprint sees it, so a double-encoded row and a
|
||||
// properly-shaped row must fingerprint identically.
|
||||
const cfg = { exclude_globs: ['Templates/**'], strategy: 'markdown' };
|
||||
const direct = computeSourceConfigFingerprint(cfg);
|
||||
// The pure compute fn handles a pre-parsed object; the persistence
|
||||
// layer's job is to deliver a parsed object. We assert that the
|
||||
// round-trip a real read would produce (parse the string scalar)
|
||||
// hashes to the same value.
|
||||
const parsed = JSON.parse(JSON.stringify(cfg));
|
||||
expect(computeSourceConfigFingerprint(parsed)).toBe(direct);
|
||||
});
|
||||
});
|
||||
@@ -355,69 +355,6 @@ describe('sync monorepo subdir-source support (#753/#774)', () => {
|
||||
expect(await engine.getPage('wiki/draft-a')).toBeNull();
|
||||
});
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────
|
||||
// --include: allow-list counterpart (#2156). Same scope-relative anchoring
|
||||
// as --exclude; exclude applies after include.
|
||||
// ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
test('--include: only matching files import on full sync', async () => {
|
||||
const { performSync } = await import('../src/commands/sync.ts');
|
||||
const result = await performSync(engine, {
|
||||
repoPath,
|
||||
include: ['wiki/**'],
|
||||
noPull: true,
|
||||
noEmbed: true,
|
||||
full: true,
|
||||
});
|
||||
expect(result.status).toBe('first_sync');
|
||||
expect(result.added).toBe(2); // wiki/page1 + wiki/page2; memory/* miss the allow-list
|
||||
expect(await engine.getPage('wiki/page1')).not.toBeNull();
|
||||
expect(await engine.getPage('memory/note1')).toBeNull();
|
||||
});
|
||||
|
||||
test('--include applies to the incremental path too', async () => {
|
||||
const { performSync } = await import('../src/commands/sync.ts');
|
||||
const first = await performSync(engine, {
|
||||
repoPath,
|
||||
include: ['wiki/**'],
|
||||
noPull: true,
|
||||
noEmbed: true,
|
||||
full: true,
|
||||
});
|
||||
expect(first.status).toBe('first_sync');
|
||||
|
||||
writeFileSync(join(repoPath, 'wiki', 'page3.md'), mdPage('Wiki Page 3'));
|
||||
writeFileSync(join(repoPath, 'memory', 'note3.md'), mdPage('Memory Note 3'));
|
||||
gitCommit(repoPath, 'more pages');
|
||||
|
||||
const second = await performSync(engine, {
|
||||
repoPath,
|
||||
include: ['wiki/**'],
|
||||
noPull: true,
|
||||
noEmbed: true,
|
||||
});
|
||||
expect(second.status).toBe('synced');
|
||||
expect(second.added).toBe(1); // wiki/page3 only; memory/note3 misses the allow-list
|
||||
expect(await engine.getPage('wiki/page3')).not.toBeNull();
|
||||
expect(await engine.getPage('memory/note3')).toBeNull();
|
||||
});
|
||||
|
||||
test('--exclude applies after --include (path in both is rejected)', async () => {
|
||||
const { performSync } = await import('../src/commands/sync.ts');
|
||||
const result = await performSync(engine, {
|
||||
repoPath,
|
||||
include: ['wiki/**'],
|
||||
exclude: ['wiki/page2.md'],
|
||||
noPull: true,
|
||||
noEmbed: true,
|
||||
full: true,
|
||||
});
|
||||
expect(result.status).toBe('first_sync');
|
||||
expect(result.added).toBe(1); // page1 only: page2 included then excluded
|
||||
expect(await engine.getPage('wiki/page1')).not.toBeNull();
|
||||
expect(await engine.getPage('wiki/page2')).toBeNull();
|
||||
});
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────
|
||||
// --exclude '**/*' emits warning (NAV-4)
|
||||
// ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -1,212 +0,0 @@
|
||||
/**
|
||||
* TODO #3 — `parseGlobList` defensive parse.
|
||||
*
|
||||
* `sources.config` is a JSONB column with no schema. The runtime can find
|
||||
* anything in `config.include_globs` / `config.exclude_globs`:
|
||||
* - A user `gbrain sources add` wrote `["people/**"]` (the happy path).
|
||||
* - A stray hand-edit wrote `"people/**"` (string, not array).
|
||||
* - A future migration's null default.
|
||||
* - A test fixture that left the column at `{}`.
|
||||
*
|
||||
* The parse must produce `string[] | undefined` so the downstream
|
||||
* `SyncOpts.include` / `SyncOpts.exclude` are either undefined (no filter)
|
||||
* or a non-empty list of usable globs. Returning `[]` would make
|
||||
* `commands/sync.ts:1454` engage the filter loop with an empty allow-list
|
||||
* that classifies every path as `include-glob-miss`.
|
||||
*/
|
||||
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import { parseGlobList, mergeGlobs, computeSourceConfigFingerprint } from '../src/commands/sync.ts';
|
||||
|
||||
describe('parseGlobList — JSONB-safe coercion to string[] | undefined', () => {
|
||||
test('happy path: array of strings round-trips identically', () => {
|
||||
expect(parseGlobList(['people/**', 'companies/**'])).toEqual(['people/**', 'companies/**']);
|
||||
});
|
||||
|
||||
test('single-element array returned as-is', () => {
|
||||
expect(parseGlobList(['Templates/**'])).toEqual(['Templates/**']);
|
||||
});
|
||||
|
||||
test('non-array values return undefined (string, object, number, null)', () => {
|
||||
expect(parseGlobList('Templates/**')).toBeUndefined();
|
||||
expect(parseGlobList({ globs: ['Templates/**'] })).toBeUndefined();
|
||||
expect(parseGlobList(42)).toBeUndefined();
|
||||
expect(parseGlobList(null)).toBeUndefined();
|
||||
expect(parseGlobList(undefined)).toBeUndefined();
|
||||
});
|
||||
|
||||
test('empty array returns undefined (no engagement of the filter loop)', () => {
|
||||
// Critical: a literal `[]` must not slip through. Empty `include` in
|
||||
// SyncableOptions silently passes everything (good), but empty
|
||||
// `exclude` is fine too — the real motivation is to keep `SyncOpts`
|
||||
// unset so callers can ignore the field entirely. Symmetric with the
|
||||
// `omitted glob flags leave config untouched` regression guard in
|
||||
// sources.test.ts.
|
||||
expect(parseGlobList([])).toBeUndefined();
|
||||
});
|
||||
|
||||
test('mixed array drops non-string entries and keeps the rest', () => {
|
||||
expect(parseGlobList(['people/**', 42, null, 'companies/**'])).toEqual([
|
||||
'people/**',
|
||||
'companies/**',
|
||||
]);
|
||||
});
|
||||
|
||||
test('empty strings dropped (a `""` glob would match every path)', () => {
|
||||
expect(parseGlobList(['', 'people/**', ''])).toEqual(['people/**']);
|
||||
});
|
||||
|
||||
test('array of only empty strings collapses to undefined', () => {
|
||||
expect(parseGlobList(['', '', ''])).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe('mergeGlobs — CLI flags union with persisted source-config globs', () => {
|
||||
test('both sides present: union, deduped, CLI first', () => {
|
||||
expect(mergeGlobs(['a/**', 'b/**'], ['b/**', 'c/**'])).toEqual(['a/**', 'b/**', 'c/**']);
|
||||
});
|
||||
|
||||
test('CLI only', () => {
|
||||
expect(mergeGlobs(['a/**'], undefined)).toEqual(['a/**']);
|
||||
});
|
||||
|
||||
test('persisted only', () => {
|
||||
expect(mergeGlobs([], ['Templates/**'])).toEqual(['Templates/**']);
|
||||
});
|
||||
|
||||
test('neither side: undefined so SyncOpts stays unset', () => {
|
||||
expect(mergeGlobs([], undefined)).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* #2157 follow-on (sources_config_fingerprint, migration v125).
|
||||
*
|
||||
* computeSourceConfigFingerprint hashes the walk-affecting fields of
|
||||
* sources.config (strategy + include_globs + exclude_globs) so the
|
||||
* "Already up to date" gate at performSync's git-HEAD equality check
|
||||
* can detect drift and force a re-walk. These cases pin the contract
|
||||
* the gate depends on:
|
||||
*
|
||||
* - Deterministic over equivalent inputs (order-insensitive,
|
||||
* defensively-coerced via parseGlobList).
|
||||
* - Sensitive to each walk-affecting field separately.
|
||||
* - Insensitive to fields the walker doesn't read (federated,
|
||||
* unrelated keys).
|
||||
* - A toggle-and-revert is a no-op (returns to the original hash).
|
||||
*
|
||||
* Without the canonicalization the gate would fire spuriously on
|
||||
* cosmetic changes (e.g. a user re-ordering their exclude list) and
|
||||
* miss real drift (e.g. an add-then-remove that nets to a different
|
||||
* effective set than the stored fingerprint).
|
||||
*/
|
||||
describe('computeSourceConfigFingerprint — walk-affecting config drift detector', () => {
|
||||
test('empty config produces a stable hash', () => {
|
||||
const a = computeSourceConfigFingerprint({});
|
||||
const b = computeSourceConfigFingerprint({});
|
||||
expect(a).toBe(b);
|
||||
expect(a).toMatch(/^[a-f0-9]{64}$/);
|
||||
});
|
||||
|
||||
test('null / undefined / missing config all hash the same', () => {
|
||||
const empty = computeSourceConfigFingerprint({});
|
||||
expect(computeSourceConfigFingerprint(null)).toBe(empty);
|
||||
expect(computeSourceConfigFingerprint(undefined)).toBe(empty);
|
||||
});
|
||||
|
||||
test('same config → same hash (deterministic)', () => {
|
||||
const cfg = { strategy: 'markdown', exclude_globs: ['Templates/**', 'Photos/**'] };
|
||||
expect(computeSourceConfigFingerprint(cfg)).toBe(computeSourceConfigFingerprint(cfg));
|
||||
});
|
||||
|
||||
test('array order does not affect hash (canonical sort)', () => {
|
||||
const a = computeSourceConfigFingerprint({ exclude_globs: ['a/**', 'b/**', 'c/**'] });
|
||||
const b = computeSourceConfigFingerprint({ exclude_globs: ['c/**', 'a/**', 'b/**'] });
|
||||
expect(a).toBe(b);
|
||||
});
|
||||
|
||||
test('exclude_globs change → different hash', () => {
|
||||
const a = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
|
||||
const b = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**', 'Photos/**'] });
|
||||
expect(a).not.toBe(b);
|
||||
});
|
||||
|
||||
test('include_globs change → different hash', () => {
|
||||
const a = computeSourceConfigFingerprint({ include_globs: ['people/**'] });
|
||||
const b = computeSourceConfigFingerprint({ include_globs: ['people/**', 'companies/**'] });
|
||||
expect(a).not.toBe(b);
|
||||
});
|
||||
|
||||
test('strategy change → different hash', () => {
|
||||
const a = computeSourceConfigFingerprint({ strategy: 'markdown' });
|
||||
const b = computeSourceConfigFingerprint({ strategy: 'code' });
|
||||
expect(a).not.toBe(b);
|
||||
});
|
||||
|
||||
test('strategy unset vs set differ', () => {
|
||||
const unset = computeSourceConfigFingerprint({});
|
||||
const set = computeSourceConfigFingerprint({ strategy: 'markdown' });
|
||||
expect(unset).not.toBe(set);
|
||||
});
|
||||
|
||||
test('add-then-remove returns to original hash (toggle is a no-op)', () => {
|
||||
const original = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
|
||||
const added = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**', 'Photos/**'] });
|
||||
const reverted = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'] });
|
||||
expect(added).not.toBe(original);
|
||||
expect(reverted).toBe(original);
|
||||
});
|
||||
|
||||
test('non-walk-affecting fields are ignored (federated, unrelated keys)', () => {
|
||||
const a = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'], federated: true });
|
||||
const b = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'], federated: false });
|
||||
const c = computeSourceConfigFingerprint({ exclude_globs: ['Templates/**'], some_unrelated_key: 'value' });
|
||||
expect(a).toBe(b);
|
||||
expect(a).toBe(c);
|
||||
});
|
||||
|
||||
test('non-string strategy coerced to null (defensive)', () => {
|
||||
// A hand-edited row could leave `strategy: 42` or `strategy: {}` — both
|
||||
// collapse to the same shape as `strategy: undefined` so the fingerprint
|
||||
// doesn't reflect a value the walker can't honor anyway.
|
||||
const empty = computeSourceConfigFingerprint({});
|
||||
expect(computeSourceConfigFingerprint({ strategy: 42 })).toBe(empty);
|
||||
expect(computeSourceConfigFingerprint({ strategy: {} })).toBe(empty);
|
||||
expect(computeSourceConfigFingerprint({ strategy: null })).toBe(empty);
|
||||
});
|
||||
|
||||
test('defensive parsing: mixed-type glob arrays hash same as cleaned arrays', () => {
|
||||
// parseGlobList drops non-string + empty entries; the fingerprint must
|
||||
// reflect what the walker actually uses, not what the raw row says.
|
||||
const dirty = computeSourceConfigFingerprint({
|
||||
exclude_globs: ['Templates/**', 42, null, '', 'Photos/**'],
|
||||
});
|
||||
const clean = computeSourceConfigFingerprint({
|
||||
exclude_globs: ['Templates/**', 'Photos/**'],
|
||||
});
|
||||
expect(dirty).toBe(clean);
|
||||
});
|
||||
|
||||
test('empty array and missing field hash identically', () => {
|
||||
const missing = computeSourceConfigFingerprint({});
|
||||
const emptyArray = computeSourceConfigFingerprint({ exclude_globs: [] });
|
||||
const emptyAfterClean = computeSourceConfigFingerprint({ exclude_globs: ['', '', ''] });
|
||||
expect(emptyArray).toBe(missing);
|
||||
expect(emptyAfterClean).toBe(missing);
|
||||
});
|
||||
|
||||
test('non-array exclude_globs (string, object) hash same as missing', () => {
|
||||
const missing = computeSourceConfigFingerprint({});
|
||||
expect(computeSourceConfigFingerprint({ exclude_globs: 'Templates/**' })).toBe(missing);
|
||||
expect(computeSourceConfigFingerprint({ exclude_globs: { foo: 'bar' } })).toBe(missing);
|
||||
});
|
||||
|
||||
test('SHA-256 output shape: 64 hex characters', () => {
|
||||
const fp = computeSourceConfigFingerprint({
|
||||
strategy: 'markdown',
|
||||
include_globs: ['people/**'],
|
||||
exclude_globs: ['Templates/**', '.git/**'],
|
||||
});
|
||||
expect(fp).toMatch(/^[a-f0-9]{64}$/);
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user