mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 08:53:22 +00:00
Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7fd0343187 | ||
|
|
d43ed5d8c3 | ||
|
|
fb1b3bfa84 |
@@ -188,7 +188,8 @@ per-release `**vX.Y.Z:**` narration — CI enforces this
|
||||
- `scripts/check-no-double-retry.sh` + `scripts/check-batch-audit-site.sh` — CI lint guards wired into `bun run verify`. The former greps src/ for `withRetry(...engine.{addLinksBatch|addTimelineEntriesBatch|upsertChunks})` patterns and fails the build on hit (prevents 3×3=9 retry amplification on incomplete reverts). The latter extracts every string-literal `auditSite: '...'` from src/ and validates each appears in the `BATCH_AUDIT_SITES` const in `src/core/retry.ts` (typo guard — prevents fragmented doctor output).
|
||||
- `src/core/fail-improve.ts` — Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generation.
|
||||
- `src/core/transcription.ts` — Audio transcription: Groq Whisper (default), OpenAI fallback, ffmpeg segmentation for >25MB.
|
||||
- `src/core/enrichment-service.ts` — Global enrichment service: entity slug generation, tier auto-escalation, batch throttling.
|
||||
- `src/core/enrichment-service.ts` — Global enrichment service: entity slug generation, tier auto-escalation, batch throttling. Write path is trust-gated (issue #160): `enrichEntity` / `enrichEntities` / `extractAndEnrich` take `EnrichmentTrustOptions { trusted?, sourceId? }`; only an explicit `trusted: true` writes authoritative `people/` / `companies/` stubs. Anything else (undefined/false — fail-closed, mirroring `OperationContext.remote`) creates the stub with the extraction quarantine markers from `src/core/extraction-review.ts` and reports `quarantined: true` in `EnrichmentResult`. The ONLY sanctioned op surface is `extract_entities` (operations.ts), which grants `trusted` solely for `ctx.remote === false` callers passing `--trusted-extraction`.
|
||||
- `src/core/extraction-review.ts` — Extraction quarantine lane markers (issue #160), sibling of `src/core/quarantine.ts` / `embed-skip.ts` (frontmatter-key pattern, no schema migration). Auto-extracted stubs from untrusted input carry the PAIR `provenance: 'auto-extracted'` + `status: 'unverified'` (both required — user pages with their own `status`/`provenance` never match). Exports `quarantineMarkers()`, `isUnverifiedExtraction()` (JS predicate) and `unverifiedExtractionFragment(alias)` — the single SQL source of truth consumed by `buildSourceFactorCase` (namespace source-boost guard), both engines' `getUnverifiedExtractionPageIds`, the `extraction_pending` op, and the `unverified_extractions` doctor check, so filter and marker keys can never drift. Consequences: unverified stubs are excluded from the compiled-truth fusion boost + the `people/`/`companies/` source-boost (rank as ordinary content), stamped `unverified: true` in search results (`stampUnverifiedExtractions`, hybrid.ts), listed by `extraction_pending`, promoted (status → `verified`, provenance kept for audit) or rejected (soft-delete) by the owner-only `extraction_review` op. Pinned by `test/extraction-review.test.ts` (PGLite) + `test/e2e/extraction-review-postgres.test.ts` (live Postgres parity).
|
||||
- `src/commands/enrich.ts` + `src/core/enrich/thin.ts` + `src/core/cycle/enrich-thin.ts` — `gbrain enrich --thin`: batch-develops stub (thin) pages via **brain-internal grounded synthesis**. gbrain's model tooling sees only brain-internal context (search / get_page / facts / backlinks), not the web, so enrich consolidates what the brain ALREADY knows about an entity (scattered across meetings, other pages, deals, facts) into one cited page via ONE `gateway.chat` call per page; web research stays the agent-driven `enrich` SKILL's job. `runEnrichCore(engine, opts, signal)` (strict per-source; multi-source iteration is the caller's job) drives `enrichOne` per candidate: `withRefreshingLock('enrich:<src>:<slug>')` → `getPage` → deterministic retrieve (hybridSearch + getBacklinks + facts + raw_data, source-scoped, sanitized via `INJECTION_PATTERNS`) → `assessGrounding` gate (skip < `MIN_CONTEXT_CHARS`, no LLM) → `buildEnrichPrompt` (grounded dossier, `[Source: slug]` citations, SKIP sentinel) → synth → `put_page` handler (`remote:false`, auto-link + write-through) stamping `enriched_at` + `enriched_by:'cli:enrich'`. Candidate selection is the SQL-native `engine.listEnrichCandidates(opts)` (`src/core/engine.ts` interface + `EnrichCandidate`/`EnrichCandidatesOpts`/`ENRICH_ORDER_SQL` in `src/core/types.ts` + pg/pglite impls): thin-filter + per-page source-correct inbound count (`to_page_id = p.id`, `mentions` excluded) + `enriched_at` recency guard + whitelisted ORDER BY + LIMIT, lightweight projection (NO bodies). Resume via `src/core/op-checkpoint.ts` (local `enrichFingerprint`); budget via `BudgetTracker` + `withBudgetTracker` (best-effort under `--workers > 1` — `runSlidingPool` aborts new claims on `BUDGET_EXHAUSTED` but does NOT cancel in-flight `gateway.chat`; pin `--workers 1` for a hard ceiling). `sanitizeContext` (thin.ts) neutralizes the `<context>…</context>` data-envelope delimiters (injection escape, mirrors the `</trajectory>` convention); the `--background` multi-source fan-out idempotency key carries the run fingerprint via exported `backgroundIdempotencyKey(sid, args)` (a bare `enrich:${sid}` would return stale completed jobs); `runEnrichCore` flags `budget_exhausted` post-hoc when `tracker.totalSpent > tracker.cap` even when the gateway swallowed the final-call throw (via read-only `BudgetTracker.cap` getter); `body()` flushes the checkpoint on `BudgetExhausted` before it propagates so resume doesn't re-charge. The opt-in `enrich_thin` cycle phase (default OFF via `cycle.enrich_thin.enabled`) trickles `max_pages_per_tick` (default 3) per source with per-source cost cap enforced as `min(per_source_cap, brain_wide_remaining)` + brain-wide total + walltime caps. Wired into `cycle.ts` (`CyclePhase`/`ALL_PHASES` between `conversation_facts_backfill` and `skillopt`/`embed`; `PHASE_SCOPE='source'`; `NEEDS_LOCK`; dispatch), `cli.ts` (`CLI_ONLY` + `CLI_ONLY_SELF_HELP` + `THIN_CLIENT_REFUSED_COMMANDS` + dispatch), `jobs.ts` (Minion `enrich` handler, strict per-source, NOT in `PROTECTED_JOB_NAMES`). DI seam `opts.synthesizeFn` keeps tests hermetic (no API key, no mock.module). Pinned by `test/enrich/thin.test.ts`, `test/enrich/idempotency.test.ts`, `test/enrich-cycle-phase.test.ts`, `test/e2e/enrich-pglite.test.ts` (grew-cited, skip, ordering, multi-source, recency, resume, budget abort + checkpoint flush, final-call overage, lock-skip, provenance), `test/e2e/engine-parity.test.ts` (`listEnrichCandidates` pg↔pglite parity).
|
||||
- `src/core/data-research.ts` — Recipe validation, field extraction (MRR/ARR regex), dedup, tracker parsing, HTML stripping.
|
||||
- `src/commands/embed.ts` — `gbrain embed [--stale|--all] [--slugs ...]`. `--stale` starts with `engine.countStaleChunks()` (single SELECT count(*) WHERE embedding IS NULL, ~50 bytes wire) so a fully-embedded brain short-circuits with no further reads. When stale chunks exist, `engine.listStaleChunks()` returns just the chunks needing embeddings (slug + chunk_index + chunk_text + metadata, no `vector(1536)` payload); caller groups by slug, embeds, re-upserts via `upsertChunks`. All `console.log`/`console.error` call sites use `slog`/`serr` from `src/core/console-prefix.ts` so when `runEmbedCore` runs inside a per-source `withSourcePrefix` scope (installed by the `gbrain sync --all` worker pool) every line carries the `[<source-id>] ` prefix; standalone callers see identical output because slog/serr fall through to bare console fns outside the wrap. Every embed-write path stamps `pages.embedding_signature` via `engine.setPageEmbeddingSignature(slug, {sourceId, signature: currentEmbeddingSignature()})` so a later model/dims swap is detectable as stale. The per-slug path (`embedPage`, used by `gbrain embed <slug>` AND sync's post-import embed step) and the full-re-embed path (`embedAll`) stamp unconditionally per page. The stale path (`embedAllStale`) first calls `invalidateStaleSignatureEmbeddings` on a live run so signature-drifted pages flow through the NULL cursor, then stamps each page — but ONLY when EVERY chunk was stale this pass (a partially-stale page keeps preserved chunks of unknown provenance, so it stays unstamped rather than falsely marked current; `embed --all` fully re-embeds + stamps those). dry-run never mutates: it counts signature-drift via the widened `countStaleChunks({signature})` predicate without NULLing anything.
|
||||
|
||||
@@ -87,6 +87,15 @@ embedding proximity. Four layers, added after the incident in
|
||||
deciding "is this page already here, safe to NOT write a duplicate?" keys off
|
||||
`create_safety`, not a raw blended score.
|
||||
|
||||
**Extraction quarantine lane (issue #160):** pages carrying the unverified
|
||||
auto-extracted markers (frontmatter `provenance: auto-extracted` +
|
||||
`status: unverified`, see `src/core/extraction-review.ts`) rank as ordinary
|
||||
content — they are skipped by the compiled-truth fusion boost and by the
|
||||
`people/`/`companies/` namespace source-boost, and every search result from
|
||||
such a page carries `unverified: true` so agents can label the provenance.
|
||||
Promote or reject them via `gbrain extraction-pending` / `gbrain
|
||||
extraction-review`.
|
||||
|
||||
The `search` MCP/CLI op is **cheap-hybrid** (vector + keyword + RRF + pool +
|
||||
title + alias, expansion off); `query` is the full-control variant. NamedThingBench
|
||||
(`gbrain eval retrieval-quality`) gates these families on every PR. Diagnose a
|
||||
|
||||
@@ -53,6 +53,7 @@ import { isUndefinedColumnError } from '../core/utils.ts';
|
||||
// drift from what search actually filters.
|
||||
import { resolveHardExcludes, DEFAULT_HARD_EXCLUDES } from '../core/search/source-boost.ts';
|
||||
import { escapeLikePattern, buildVisibilityClause } from '../core/search/sql-ranking.ts';
|
||||
import { unverifiedExtractionFragment } from '../core/extraction-review.ts';
|
||||
import { hnswIndexExpected, hnswMaxDimsForType } from '../core/vector-index.ts';
|
||||
|
||||
export interface Check {
|
||||
@@ -3620,6 +3621,52 @@ export async function checkLinksExtractionLag(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* issue #160 — unverified_extractions doctor check.
|
||||
*
|
||||
* The extraction quarantine lane parks auto-extracted entity stubs
|
||||
* (frontmatter `provenance: 'auto-extracted'` + `status: 'unverified'`)
|
||||
* until the owner promotes or rejects them. A queue nobody reviews decays
|
||||
* into invisible clutter, so this check counts stubs older than N days
|
||||
* (default 7) and nudges toward the review surface. Exported for direct
|
||||
* testing (mirrors checkLinksExtractionLag).
|
||||
*/
|
||||
export async function checkUnverifiedExtractions(
|
||||
engine: BrainEngine,
|
||||
opts?: { sourceId?: string; days?: number },
|
||||
): Promise<Check> {
|
||||
const name = 'unverified_extractions';
|
||||
const days = opts?.days ?? 7;
|
||||
const sourceId = opts?.sourceId;
|
||||
try {
|
||||
const params: unknown[] = [String(days)];
|
||||
let srcClause = '';
|
||||
if (sourceId) {
|
||||
params.push(sourceId);
|
||||
srcClause = 'AND p.source_id = $2';
|
||||
}
|
||||
const rows = await engine.executeRaw<{ n: string | number }>(
|
||||
`SELECT COUNT(*)::int AS n FROM pages p
|
||||
WHERE p.deleted_at IS NULL
|
||||
AND ${unverifiedExtractionFragment('p')}
|
||||
AND p.created_at < now() - ($1 || ' days')::interval
|
||||
${srcClause}`,
|
||||
params,
|
||||
);
|
||||
const n = Number(rows[0]?.n ?? 0);
|
||||
return {
|
||||
name,
|
||||
status: n > 0 ? 'warn' : 'ok',
|
||||
message: n > 0
|
||||
? `${n} unverified auto-extracted entity stub(s) older than ${days} days awaiting review. List with 'gbrain extraction-pending'; promote/reject with 'gbrain extraction-review <promote|reject> --slugs <slug,...>'.`
|
||||
: 'No stale unverified extraction stubs',
|
||||
details: { count: n, days, source_id: sourceId ?? null },
|
||||
};
|
||||
} catch (e) {
|
||||
return { name, status: 'warn', message: `Could not check unverified_extractions: ${(e as Error).message}` };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* issue #1678 — extract_atoms_backlog doctor check.
|
||||
*
|
||||
@@ -6891,6 +6938,10 @@ export async function buildChecks(
|
||||
checks.push({ name: 'flagged_pages', status: 'ok', message: `Skipped (${msg})` });
|
||||
}
|
||||
|
||||
// issue #160: extraction quarantine lane review nudge.
|
||||
progress.heartbeat('unverified_extractions');
|
||||
checks.push(await checkUnverifiedExtractions(engine, { sourceId: orphanRatioSourceId }));
|
||||
|
||||
// 11a. Frontmatter integrity (v0.22.4, hardened in v0.38.2.0).
|
||||
// scanBrainSources walks every registered source's local_path on disk
|
||||
// (not from the DB), invoking parseMarkdown(..., {validate:true}) per
|
||||
|
||||
@@ -112,6 +112,7 @@ export const BRAIN_CHECK_NAMES: ReadonlySet<string> = new Set([
|
||||
'takes_weight_grid',
|
||||
'timeline_coverage',
|
||||
'unified_multimodal_coverage',
|
||||
'unverified_extractions',
|
||||
'voice_gate_health',
|
||||
]);
|
||||
|
||||
|
||||
@@ -1327,6 +1327,16 @@ export interface BrainEngine {
|
||||
getContentFlagsByPageIds(
|
||||
pageIds: number[],
|
||||
): Promise<Map<number, { reason: string; detail: string }>>;
|
||||
/**
|
||||
* Extraction quarantine lane (issue #160): for a list of page_ids, return
|
||||
* the subset that are unverified auto-extracted entity stubs (frontmatter
|
||||
* `provenance: 'auto-extracted'` + `status: 'unverified'`). Used by hybrid
|
||||
* search to stamp `SearchResult.unverified` pre-fusion so the fusion-level
|
||||
* compiled-truth boost skips them. Single SQL query, not N+1. Empty input
|
||||
* → empty set (no query). SQL predicate is the shared
|
||||
* `unverifiedExtractionFragment` (src/core/extraction-review.ts).
|
||||
*/
|
||||
getUnverifiedExtractionPageIds(pageIds: number[]): Promise<Set<number>>;
|
||||
/**
|
||||
* v0.27.0: for a list of slugs, return their updated_at timestamps (or created_at fallback).
|
||||
* Used by hybrid search recency boost. Single SQL query, not N+1.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
import type { BrainEngine } from './engine.ts';
|
||||
import { waitForCapacity } from './backoff.ts';
|
||||
import { quarantineMarkers } from './extraction-review.ts';
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Types
|
||||
@@ -28,9 +29,32 @@ export interface EnrichmentRequest {
|
||||
tier?: 1 | 2 | 3;
|
||||
}
|
||||
|
||||
/**
|
||||
* Trust options for the enrichment write path (issue #160).
|
||||
*
|
||||
* `trusted: true` — the input text comes from the machine owner via the
|
||||
* trusted local CLI (ctx.remote === false) AND the caller passed an explicit
|
||||
* opt-in flag. Stubs write direct as authoritative entity pages.
|
||||
*
|
||||
* Anything else (undefined, false, absent) is UNTRUSTED — fail-closed,
|
||||
* mirroring the OperationContext.remote invariant ("anything not strictly
|
||||
* false is remote"). Created stubs land in the quarantine lane: frontmatter
|
||||
* `provenance: 'auto-extracted'` + `status: 'unverified'`. They are excluded
|
||||
* from authoritative retrieval boosts and wait in the review queue
|
||||
* (`extraction_pending` / `extraction_review` ops) until the owner promotes
|
||||
* or rejects them.
|
||||
*/
|
||||
export interface EnrichmentTrustOptions {
|
||||
trusted?: boolean;
|
||||
/** Source to read/write in (multi-source brains). Omitted → engine default. */
|
||||
sourceId?: string;
|
||||
}
|
||||
|
||||
export interface EnrichmentResult {
|
||||
slug: string;
|
||||
action: 'created' | 'updated' | 'skipped';
|
||||
/** True when the created stub landed in the quarantine lane (issue #160). */
|
||||
quarantined?: boolean;
|
||||
tier: 1 | 2 | 3;
|
||||
backlinkCreated: boolean;
|
||||
timelineAdded: boolean;
|
||||
@@ -72,11 +96,15 @@ export function entityPagePath(name: string, type: 'person' | 'company'): string
|
||||
export async function enrichEntity(
|
||||
engine: BrainEngine,
|
||||
request: EnrichmentRequest,
|
||||
opts?: EnrichmentTrustOptions,
|
||||
): Promise<EnrichmentResult> {
|
||||
const slug = slugifyEntity(request.entityName, request.entityType);
|
||||
// Fail-closed: only an explicit `trusted: true` writes authoritative pages.
|
||||
const trusted = opts?.trusted === true;
|
||||
const scope = opts?.sourceId ? { sourceId: opts.sourceId } : undefined;
|
||||
|
||||
// 1. Count existing mentions for tier auto-escalation
|
||||
const { mentionCount, mentionSources } = await countMentions(engine, request.entityName);
|
||||
const { mentionCount, mentionSources } = await countMentions(engine, request.entityName, opts?.sourceId);
|
||||
|
||||
// 2. Determine tier (auto-escalate based on mentions)
|
||||
const suggestedTier = suggestTier(mentionCount, mentionSources, request.context);
|
||||
@@ -84,7 +112,7 @@ export async function enrichEntity(
|
||||
const tierEscalated = suggestedTier < (request.tier || 3); // lower tier number = higher importance
|
||||
|
||||
// 3. Check if entity page exists
|
||||
const existingPage = await engine.getPage(slug);
|
||||
const existingPage = await engine.getPage(slug, scope);
|
||||
let action: 'created' | 'updated' | 'skipped';
|
||||
|
||||
if (existingPage) {
|
||||
@@ -104,8 +132,11 @@ export async function enrichEntity(
|
||||
created: new Date().toISOString().split('T')[0],
|
||||
source: request.sourceSlug,
|
||||
tier,
|
||||
// issue #160 quarantine lane: stubs extracted from untrusted input
|
||||
// carry provenance + unverified markers until the owner reviews them.
|
||||
...(trusted ? {} : quarantineMarkers()),
|
||||
},
|
||||
});
|
||||
}, scope);
|
||||
action = 'created';
|
||||
}
|
||||
|
||||
@@ -116,7 +147,7 @@ export async function enrichEntity(
|
||||
date: new Date().toISOString().split('T')[0] ?? '',
|
||||
summary: `Referenced in [${request.sourceSlug}](${request.sourceSlug}) — ${request.context}`,
|
||||
source: request.sourceSlug,
|
||||
});
|
||||
}, scope);
|
||||
timelineAdded = true;
|
||||
} catch {
|
||||
// Timeline add failed (page might not support it)
|
||||
@@ -125,7 +156,7 @@ export async function enrichEntity(
|
||||
// 5. Add backlink from entity to source
|
||||
let backlinkCreated = false;
|
||||
try {
|
||||
await engine.addLink(slug, request.sourceSlug, `Entity mention from ${request.sourceSlug}`); // gbrain-allow-direct-insert: auto-link reconciliation triggered by entity reference in source markdown
|
||||
await engine.addLink(slug, request.sourceSlug, `Entity mention from ${request.sourceSlug}`, undefined, undefined, undefined, undefined, opts?.sourceId ? { fromSourceId: opts.sourceId, toSourceId: opts.sourceId } : undefined); // gbrain-allow-direct-insert: auto-link reconciliation triggered by entity reference in source markdown
|
||||
backlinkCreated = true;
|
||||
} catch {
|
||||
// Link might already exist
|
||||
@@ -134,6 +165,7 @@ export async function enrichEntity(
|
||||
return {
|
||||
slug,
|
||||
action,
|
||||
...(action === 'created' && !trusted ? { quarantined: true } : {}),
|
||||
tier,
|
||||
backlinkCreated,
|
||||
timelineAdded,
|
||||
@@ -152,14 +184,14 @@ export async function enrichEntity(
|
||||
export async function enrichEntities(
|
||||
engine: BrainEngine,
|
||||
requests: EnrichmentRequest[],
|
||||
config?: { throttle?: boolean; onProgress?: (done: number, total: number, name: string) => void },
|
||||
config?: { throttle?: boolean; onProgress?: (done: number, total: number, name: string) => void } & EnrichmentTrustOptions,
|
||||
): Promise<EnrichmentResult[]> {
|
||||
const results: EnrichmentResult[] = [];
|
||||
for (const req of requests) {
|
||||
if (config?.throttle !== false) {
|
||||
await waitForCapacity({ maxAttempts: 5 }); // shorter timeout for batch items
|
||||
}
|
||||
const result = await enrichEntity(engine, req);
|
||||
const result = await enrichEntity(engine, req, { trusted: config?.trusted, sourceId: config?.sourceId });
|
||||
results.push(result);
|
||||
config?.onProgress?.(results.length, requests.length, req.entityName);
|
||||
}
|
||||
@@ -175,8 +207,11 @@ export async function extractAndEnrich(
|
||||
engine: BrainEngine,
|
||||
text: string,
|
||||
sourceSlug: string,
|
||||
opts?: EnrichmentTrustOptions & { throttle?: boolean; maxEntities?: number },
|
||||
): Promise<EnrichmentResult[]> {
|
||||
const entities = extractEntities(text);
|
||||
// Bounded by default (#160 hardening): the greedy regex on a large paste
|
||||
// can produce thousands of hits; each enrichment is several DB round-trips.
|
||||
const entities = extractEntities(text).slice(0, opts?.maxEntities ?? 200);
|
||||
if (entities.length === 0) return [];
|
||||
|
||||
const requests: EnrichmentRequest[] = entities.map(e => ({
|
||||
@@ -186,7 +221,7 @@ export async function extractAndEnrich(
|
||||
sourceSlug,
|
||||
}));
|
||||
|
||||
return enrichEntities(engine, requests);
|
||||
return enrichEntities(engine, requests, { trusted: opts?.trusted, sourceId: opts?.sourceId, throttle: opts?.throttle });
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -197,9 +232,10 @@ export async function extractAndEnrich(
|
||||
async function countMentions(
|
||||
engine: BrainEngine,
|
||||
entityName: string,
|
||||
sourceId?: string,
|
||||
): Promise<{ mentionCount: number; mentionSources: string[] }> {
|
||||
try {
|
||||
const results = await engine.searchKeyword(entityName, { limit: 100 });
|
||||
const results = await engine.searchKeyword(entityName, { limit: 100, ...(sourceId ? { sourceId } : {}) });
|
||||
// Derive sources from slug prefixes since SearchResult has no metadata.skill
|
||||
const sources = new Set<string>();
|
||||
for (const r of results) {
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
/**
|
||||
* Extraction quarantine lane (issue #160).
|
||||
*
|
||||
* `extractAndEnrich` regex-extracts entity names from arbitrary ingested text
|
||||
* and creates `people/{slug}` / `companies/{slug}` stub pages. When the input
|
||||
* text comes from an untrusted channel (anything that is not the trusted local
|
||||
* CLI with an explicit opt-in), those stubs must NOT enter the brain as
|
||||
* authoritative entity pages. Instead they land in the quarantine lane:
|
||||
* ordinary pages carrying two frontmatter markers —
|
||||
*
|
||||
* provenance: 'auto-extracted' — HOW the page came to exist
|
||||
* status: 'unverified' — the owner has not reviewed it yet
|
||||
*
|
||||
* Consequences of the markers (each enforced at its own site):
|
||||
* - Search: unverified stubs are excluded from the compiled-truth authority
|
||||
* boost (they rank as ordinary content) and results carry
|
||||
* `unverified: true` so agents can label the provenance.
|
||||
* - Review: `extraction_pending` lists them; `extraction_review` promotes
|
||||
* (status → 'verified', provenance kept for audit) or rejects
|
||||
* (soft-delete) in batch. Promotion is local-owner-only.
|
||||
* - Doctor: counts unverified stubs older than N days as a review nudge.
|
||||
*
|
||||
* Fail-closed trust rule (mirrors OperationContext.remote): only an explicit
|
||||
* `trusted: true` writes direct; undefined/false/anything-else quarantines.
|
||||
*
|
||||
* Known scope (deliberate, documented — not gaps discovered later):
|
||||
* - CREATE-path only. The enrichment UPDATE path (timeline append + edge
|
||||
* onto an EXISTING page when a slug collides) is the separately-tracked
|
||||
* slug-collision finding referenced in issue #160; this lane does not
|
||||
* gate it.
|
||||
* - The markers are ordinary frontmatter keys, not put_page-strip-listed
|
||||
* (#1699). A caller holding generic remote put_page write scope can
|
||||
* rewrite a stub without them — but that caller can author an unmarked
|
||||
* people/ page directly anyway, so stripping here adds no privilege.
|
||||
* The promotion OP surface (extraction_review) is what stays owner-only.
|
||||
*
|
||||
* Sibling of `src/core/quarantine.ts` / `src/core/embed-skip.ts` — same
|
||||
* marker-as-frontmatter-JSONB pattern, same "SQL fragment lives next to the
|
||||
* marker key so they can never drift" rule. No schema migration needed.
|
||||
*/
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Marker keys + values (stable contract)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export const EXTRACTION_PROVENANCE_KEY = 'provenance';
|
||||
export const EXTRACTION_STATUS_KEY = 'status';
|
||||
|
||||
export const PROVENANCE_AUTO_EXTRACTED = 'auto-extracted';
|
||||
export const STATUS_UNVERIFIED = 'unverified';
|
||||
export const STATUS_VERIFIED = 'verified';
|
||||
|
||||
/** Frontmatter markers to spread onto a quarantined stub at create time. */
|
||||
export function quarantineMarkers(): Record<string, string> {
|
||||
return {
|
||||
[EXTRACTION_PROVENANCE_KEY]: PROVENANCE_AUTO_EXTRACTED,
|
||||
[EXTRACTION_STATUS_KEY]: STATUS_UNVERIFIED,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* JS-side predicate: true only when BOTH markers match. Requiring the pair
|
||||
* means user pages that happen to carry their own `status` or `provenance`
|
||||
* frontmatter are never captured by the review lane.
|
||||
*/
|
||||
export function isUnverifiedExtraction(
|
||||
frontmatter: Record<string, unknown> | null | undefined,
|
||||
): boolean {
|
||||
if (!frontmatter) return false;
|
||||
return (
|
||||
frontmatter[EXTRACTION_PROVENANCE_KEY] === PROVENANCE_AUTO_EXTRACTED &&
|
||||
frontmatter[EXTRACTION_STATUS_KEY] === STATUS_UNVERIFIED
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* SQL fragment matching unverified auto-extracted stubs, parameterized on the
|
||||
* page-table alias. Single source of truth for every SQL-side consumer
|
||||
* (extraction_pending list, doctor count) so the filter and the marker keys
|
||||
* can never drift. `pageAlias` is engine-supplied (never user input).
|
||||
* JSONB `->>` works identically on Postgres and PGLite (PostgreSQL-in-WASM).
|
||||
*/
|
||||
export function unverifiedExtractionFragment(pageAlias: string): string {
|
||||
return (
|
||||
`(COALESCE(${pageAlias}.frontmatter, '{}'::jsonb) ->> '${EXTRACTION_PROVENANCE_KEY}') = '${PROVENANCE_AUTO_EXTRACTED}'` +
|
||||
` AND (COALESCE(${pageAlias}.frontmatter, '{}'::jsonb) ->> '${EXTRACTION_STATUS_KEY}') = '${STATUS_UNVERIFIED}'`
|
||||
);
|
||||
}
|
||||
+211
-1
@@ -11,7 +11,7 @@ import type { GBrainConfig } from './config.ts';
|
||||
import type { PageType } from './types.ts';
|
||||
import { importFromContent } from './import-file.ts';
|
||||
import { writePageThrough } from './write-through.ts';
|
||||
import { hybridSearch, hybridSearchCached, stampContentFlags } from './search/hybrid.ts';
|
||||
import { hybridSearch, hybridSearchCached, stampContentFlags, stampUnverifiedExtractions } from './search/hybrid.ts';
|
||||
import { expandQuery } from './search/expansion.ts';
|
||||
import { dedupResults } from './search/dedup.ts';
|
||||
import { captureEvalCandidate, isEvalCaptureEnabled, isEvalScrubEnabled } from './eval-capture.ts';
|
||||
@@ -21,6 +21,8 @@ import { isFactsBackstopEligible } from './facts/eligibility.ts';
|
||||
import { stripTakesFence } from './takes-fence.ts';
|
||||
import { stripFactsFence } from './facts-fence.ts';
|
||||
import { getContentFlag } from './quarantine.ts';
|
||||
import { unverifiedExtractionFragment, isUnverifiedExtraction, EXTRACTION_STATUS_KEY, STATUS_VERIFIED } from './extraction-review.ts';
|
||||
import { buildVisibilityClause } from './search/sql-ranking.ts';
|
||||
import { bumpLastRetrievedAt } from './last-retrieved.ts';
|
||||
import { isSearchMode } from './search/mode.ts';
|
||||
import { stampEvidence } from './search/evidence.ts';
|
||||
@@ -1625,6 +1627,10 @@ const search: Operation = {
|
||||
// agent-warning channel (hybridSearch stamps it; this branch bypasses
|
||||
// hybridSearch, so stamp explicitly). Fail-open inside the helper.
|
||||
await stampContentFlags(ctx.engine, results);
|
||||
// #160: same for the unverified auto-extracted stub marker (no boost
|
||||
// to cancel on this path — keyword-only never applies the compiled-
|
||||
// truth boost — but the provenance marker must still surface).
|
||||
await stampUnverifiedExtractions(ctx.engine, results);
|
||||
bumpLastRetrievedAt(ctx.engine, results.map((r) => r.page_id));
|
||||
maybeCaptureSearch(ctx, queryText, results, Date.now() - startedAt, false);
|
||||
return results;
|
||||
@@ -5587,6 +5593,208 @@ const chronicle_backfill: Operation = {
|
||||
cliHints: { name: 'chronicle-backfill' },
|
||||
};
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Extraction quarantine lane (issue #160)
|
||||
//
|
||||
// `extractAndEnrich` regex-extracts entity names from arbitrary text and
|
||||
// creates people/ + companies/ stub pages. These three ops are its ONLY
|
||||
// sanctioned surface:
|
||||
// - extract_entities — run extraction. Direct authoritative writes need
|
||||
// BOTH the trusted local CLI (ctx.remote === false)
|
||||
// AND the explicit --trusted-extraction flag;
|
||||
// everything else lands in the quarantine lane
|
||||
// (frontmatter provenance/status markers).
|
||||
// - extraction_pending — list unverified stubs awaiting review.
|
||||
// - extraction_review — promote (status → verified) or reject
|
||||
// (soft-delete) in batch. Owner-only (fail-closed
|
||||
// on ctx.remote): THIS surface never lets a remote
|
||||
// caller flip the status markers. Scope note: the
|
||||
// markers are ordinary frontmatter, so a caller who
|
||||
// already holds generic remote put_page write scope
|
||||
// can rewrite the page (markers included) — that
|
||||
// caller could equally author an unmarked people/
|
||||
// page directly, so the lane adds no privilege
|
||||
// there; put_page authz is its own boundary.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// Resource guards for extract_entities (#160 hardening): bound the work a
|
||||
// single remote write-scope call can trigger. ponytail: flat caps; make them
|
||||
// config knobs only if a real workload hits them.
|
||||
const MAX_EXTRACT_TEXT_CHARS = 200_000;
|
||||
const MAX_EXTRACT_ENTITIES = 200;
|
||||
|
||||
const extract_entities: Operation = {
|
||||
name: 'extract_entities',
|
||||
description: 'Extract entity names (people, companies) from text and create/update their brain stub pages. Stubs from untrusted input land in the quarantine lane (frontmatter `provenance: auto-extracted` + `status: unverified`) — excluded from authoritative retrieval boosts until reviewed. Direct authoritative writes require the trusted local CLI AND --trusted-extraction.',
|
||||
params: {
|
||||
text: { type: 'string', required: true, description: 'The text to extract entities from (email, transcript, pasted content, …). Max 200k characters — split larger inputs.' },
|
||||
source_slug: { type: 'string', required: true, description: 'Slug of the source page the text came from (used for backlinks + timeline attribution).' },
|
||||
trusted_extraction: { type: 'boolean', required: false, description: 'Local CLI only: write stubs directly as authoritative pages, skipping the quarantine lane. Ignored (always quarantined) for remote callers.' },
|
||||
},
|
||||
mutating: true,
|
||||
scope: 'write',
|
||||
handler: async (ctx, p) => {
|
||||
// Trust rule (#160, fail-closed like the CV6 provenance gate above):
|
||||
// `ctx.remote === false` is the ONLY truthy condition that can admit a
|
||||
// direct authoritative write, and even then the caller must opt in
|
||||
// explicitly. Remote/unset trust → quarantine lane, flag ignored.
|
||||
const trusted = ctx.remote === false && p.trusted_extraction === true;
|
||||
const text = p.text as string;
|
||||
// Resource guards: the greedy name regex on a huge paste can yield tens
|
||||
// of thousands of "entities", each costing several DB round-trips. Cap
|
||||
// input size loudly and entity count softly (surfaced as `truncated`).
|
||||
if (text.length > MAX_EXTRACT_TEXT_CHARS) {
|
||||
throw new OperationError(
|
||||
'invalid_params',
|
||||
`extract_entities: text is ${text.length} chars (max ${MAX_EXTRACT_TEXT_CHARS}).`,
|
||||
'Split the input and call extract_entities per section.',
|
||||
);
|
||||
}
|
||||
if (ctx.dryRun) return { dry_run: true, action: 'extract_entities', trusted };
|
||||
const { extractEntities, enrichEntities } = await import('./enrichment-service.ts');
|
||||
const found = extractEntities(text);
|
||||
const capped = found.slice(0, MAX_EXTRACT_ENTITIES);
|
||||
const results = await enrichEntities(
|
||||
ctx.engine,
|
||||
capped.map((e) => ({ entityName: e.name, entityType: e.type, context: e.context, sourceSlug: p.source_slug as string })),
|
||||
{
|
||||
trusted,
|
||||
...(ctx.sourceId ? { sourceId: ctx.sourceId } : {}),
|
||||
// Pure local DB writes — no external API call to pace, so the
|
||||
// system-load capacity gate would only stall the caller.
|
||||
throttle: false,
|
||||
},
|
||||
);
|
||||
return {
|
||||
status: 'ok',
|
||||
trusted,
|
||||
quarantined: results.filter((r) => r.quarantined === true).length,
|
||||
count: results.length,
|
||||
entities_found: found.length,
|
||||
truncated: found.length > capped.length,
|
||||
entities: results,
|
||||
};
|
||||
},
|
||||
cliHints: { name: 'extract-entities' },
|
||||
};
|
||||
|
||||
const extraction_pending: Operation = {
|
||||
name: 'extraction_pending',
|
||||
description: 'List unverified auto-extracted entity stubs awaiting owner review (the quarantine lane from extract_entities). Promote or reject them with extraction_review.',
|
||||
params: {
|
||||
limit: { type: 'number', required: false, description: 'Max rows (default 100, cap 500).' },
|
||||
offset: { type: 'number', required: false, description: 'Pagination offset.' },
|
||||
},
|
||||
scope: 'read',
|
||||
handler: async (ctx, p) => {
|
||||
const limit = Math.min(Math.max(Number(p.limit ?? 100) || 100, 1), 500);
|
||||
const offset = Math.max(Number(p.offset ?? 0) || 0, 0);
|
||||
// Read-side source isolation: route through sourceScopeOpts (federated
|
||||
// array > scalar > nothing), applied in SQL below.
|
||||
const scope = sourceScopeOpts(ctx);
|
||||
const params: unknown[] = [];
|
||||
let srcClause = '';
|
||||
if (scope.sourceIds && scope.sourceIds.length > 0) {
|
||||
params.push(scope.sourceIds);
|
||||
srcClause = `AND p.source_id = ANY($${params.length}::text[])`;
|
||||
} else if (scope.sourceId) {
|
||||
params.push(scope.sourceId);
|
||||
srcClause = `AND p.source_id = $${params.length}`;
|
||||
}
|
||||
params.push(limit, offset);
|
||||
const rows = await ctx.engine.executeRaw<{
|
||||
slug: string; title: string; type: string; source_id: string;
|
||||
extracted_from: string | null; created_at: string;
|
||||
}>(
|
||||
`SELECT p.slug, p.title, p.type, p.source_id,
|
||||
p.frontmatter ->> 'source' AS extracted_from,
|
||||
p.created_at::text AS created_at
|
||||
FROM pages p
|
||||
JOIN sources s ON s.id = p.source_id
|
||||
WHERE ${unverifiedExtractionFragment('p')}
|
||||
${buildVisibilityClause('p', 's')}
|
||||
${srcClause}
|
||||
ORDER BY p.created_at DESC
|
||||
LIMIT $${params.length - 1} OFFSET $${params.length}`,
|
||||
params,
|
||||
);
|
||||
return { count: rows.length, pending: rows };
|
||||
},
|
||||
cliHints: { name: 'extraction-pending' },
|
||||
};
|
||||
|
||||
const extraction_review: Operation = {
|
||||
name: 'extraction_review',
|
||||
description: 'Promote or reject unverified auto-extracted entity stubs (batch). Promote flips `status` to verified (provenance kept for audit); reject soft-deletes the stub. Owner-only: this op is refused for any non-local caller. (The markers are ordinary frontmatter — the boundary against rewriting them wholesale is put_page write authz, same as for any page.)',
|
||||
params: {
|
||||
action: { type: 'string', required: true, description: "'promote' or 'reject'." },
|
||||
slugs: { type: 'array', required: true, items: { type: 'string' }, description: 'Stub slugs to act on (batch).' },
|
||||
},
|
||||
mutating: true,
|
||||
scope: 'write',
|
||||
localOnly: true,
|
||||
handler: async (ctx, p) => {
|
||||
// The review decision IS the trust gate — if a remote caller could
|
||||
// promote, injected content could self-promote and the quarantine lane
|
||||
// would be decorative. Fail-closed: only strictly-local callers pass.
|
||||
if (ctx.remote !== false) {
|
||||
throw new OperationError(
|
||||
'permission_denied',
|
||||
'extraction_review is owner-only: promote/reject decisions must come from the trusted local CLI.',
|
||||
'Run `gbrain extraction-review <promote|reject> --slugs ...` on the host machine.',
|
||||
);
|
||||
}
|
||||
const action = p.action as string;
|
||||
if (action !== 'promote' && action !== 'reject') {
|
||||
throw new OperationError('invalid_params', `extraction_review: action must be 'promote' or 'reject'; got '${action}'.`);
|
||||
}
|
||||
// CLI passes `--slugs a,b,c` as one string; MCP passes a real array.
|
||||
const slugs = Array.isArray(p.slugs)
|
||||
? (p.slugs as string[])
|
||||
: typeof p.slugs === 'string'
|
||||
? p.slugs.split(',').map((s) => s.trim()).filter(Boolean)
|
||||
: [];
|
||||
if (slugs.length === 0) {
|
||||
throw new OperationError('invalid_params', 'extraction_review: slugs must be a non-empty array (CLI: --slugs slug1,slug2).');
|
||||
}
|
||||
if (ctx.dryRun) return { dry_run: true, action: `extraction_review:${action}`, slugs };
|
||||
const results: Array<{ slug: string; status: string }> = [];
|
||||
for (const slug of slugs) {
|
||||
const page = await ctx.engine.getPage(slug, ctx.sourceId ? { sourceId: ctx.sourceId } : undefined);
|
||||
if (!page) {
|
||||
results.push({ slug, status: 'not_found' });
|
||||
continue;
|
||||
}
|
||||
if (!isUnverifiedExtraction(page.frontmatter)) {
|
||||
results.push({ slug, status: 'not_unverified' });
|
||||
continue;
|
||||
}
|
||||
if (action === 'promote') {
|
||||
// Frontmatter-only flip via a targeted JSONB merge — NOT putPage,
|
||||
// whose upsert would reset non-carried columns (page_kind →
|
||||
// 'markdown', content_hash, …) for a change that only touches one
|
||||
// frontmatter key. provenance stays 'auto-extracted' as the audit
|
||||
// trail of HOW the page came to exist; status → 'verified' records
|
||||
// the owner's call. jsonb_build_object binds as text (no
|
||||
// JSON.stringify-into-::jsonb hazard); identical on both engines.
|
||||
await ctx.engine.executeRaw(
|
||||
`UPDATE pages
|
||||
SET frontmatter = COALESCE(frontmatter, '{}'::jsonb) || jsonb_build_object($1::text, $2::text),
|
||||
updated_at = now()
|
||||
WHERE slug = $3 AND source_id = $4`,
|
||||
[EXTRACTION_STATUS_KEY, STATUS_VERIFIED, slug, page.source_id],
|
||||
);
|
||||
results.push({ slug, status: 'promoted' });
|
||||
} else {
|
||||
await ctx.engine.softDeletePage(slug, { sourceId: page.source_id });
|
||||
results.push({ slug, status: 'rejected' });
|
||||
}
|
||||
}
|
||||
return { status: 'ok', action, results };
|
||||
},
|
||||
cliHints: { name: 'extraction-review', positional: ['action'] },
|
||||
};
|
||||
|
||||
export const operations: Operation[] = [
|
||||
// Page CRUD
|
||||
get_page, put_page, delete_page, list_pages,
|
||||
@@ -5643,6 +5851,8 @@ export const operations: Operation[] = [
|
||||
volunteer_chronicle, chronicle_backfill,
|
||||
// v0.43 (#2095): push-based context
|
||||
volunteer_context,
|
||||
// Extraction quarantine lane (#160): gated entity extraction + review queue
|
||||
extract_entities, extraction_pending, extraction_review,
|
||||
// v0.31: hot memory (facts table)
|
||||
extract_facts, recall, forget_fact,
|
||||
// v0.32.6: contradiction probe MCP surface (M3)
|
||||
|
||||
@@ -58,6 +58,7 @@ import { finalizeLastSeen } from './chronicle/last-seen.ts';
|
||||
import { computeAnomaliesFromBuckets } from './cycle/anomaly.ts';
|
||||
import { resolveBoostMap, resolveHardExcludes } from './search/source-boost.ts';
|
||||
import { buildSourceFactorCase, buildHardExcludeClause, buildVisibilityClause, buildRecencyComponentSql, buildBestPerPagePoolCte, buildOrFallbackWebsearchQuery } from './search/sql-ranking.ts';
|
||||
import { unverifiedExtractionFragment } from './extraction-review.ts';
|
||||
import { shouldExcludeFromOrphanReporting, loadOrphanPolicyOverrides } from './orphan-policy.ts';
|
||||
import { LINK_EXTRACTOR_VERSION_TS } from './link-extraction.ts';
|
||||
import {
|
||||
@@ -2072,7 +2073,10 @@ export class PGLiteEngine implements BrainEngine {
|
||||
// Built on the bare `slug` output column: applied inside the `scored` CTE
|
||||
// whose FROM is the single relation `hnsw_candidates`, so unqualified
|
||||
// `slug` resolves cleanly (T1 per-page pool restructure).
|
||||
const sourceFactorCaseOnSlug = buildSourceFactorCase('slug', boostMap, opts?.detail);
|
||||
// issue #160: guard predicate projected as `unverified_stub` in
|
||||
// hnsw_candidates (parity with postgres-engine) so unverified stubs get
|
||||
// factor 1.0, not the people/ 1.2x, inside the pre-LIMIT re-rank.
|
||||
const sourceFactorCaseOnSlug = buildSourceFactorCase('slug', boostMap, opts?.detail, 'unverified_stub');
|
||||
const hardExcludePrefixes = resolveHardExcludes(opts?.exclude_slug_prefixes, opts?.include_slug_prefixes);
|
||||
const hardExcludeClause = buildHardExcludeClause('p.slug', hardExcludePrefixes);
|
||||
const innerLimit = offset + Math.max(limit * 5, 100);
|
||||
@@ -2148,6 +2152,7 @@ export class PGLiteEngine implements BrainEngine {
|
||||
CASE WHEN NULLIF(regexp_replace(p.frontmatter->>'message_id', '^[[:space:]]+|[[:space:]]+$', '', 'g'), '') IS NOT NULL
|
||||
THEN NULLIF(p.frontmatter->>'subject', '') END AS source_subject,
|
||||
cc.id as chunk_id, cc.chunk_index, cc.chunk_text, cc.chunk_source,
|
||||
(${unverifiedExtractionFragment('p')}) AS unverified_stub,
|
||||
1 - (cc.${col} <=> ${castSql}) AS raw_score
|
||||
FROM content_chunks cc
|
||||
JOIN pages p ON p.id = cc.page_id
|
||||
@@ -3407,6 +3412,20 @@ export class PGLiteEngine implements BrainEngine {
|
||||
return result;
|
||||
}
|
||||
|
||||
async getUnverifiedExtractionPageIds(pageIds: number[]): Promise<Set<number>> {
|
||||
if (pageIds.length === 0) return new Set();
|
||||
// Parity with PostgresEngine.getUnverifiedExtractionPageIds (issue #160).
|
||||
// Predicate is the shared unverifiedExtractionFragment so this query and
|
||||
// the SQL-side source-boost guard can never drift.
|
||||
const { rows } = await this.db.query(
|
||||
`SELECT id FROM pages
|
||||
WHERE id = ANY($1::int[])
|
||||
AND ${unverifiedExtractionFragment('pages')}`,
|
||||
[pageIds]
|
||||
);
|
||||
return new Set((rows as { id: number }[]).map((r) => Number(r.id)));
|
||||
}
|
||||
|
||||
async getPageTimestamps(slugs: string[]): Promise<Map<string, Date>> {
|
||||
if (slugs.length === 0) return new Map();
|
||||
const { rows } = await this.db.query(
|
||||
|
||||
@@ -65,6 +65,7 @@ import { logConnectionEvent } from './connection-audit.ts';
|
||||
import { validateSlug, contentHash, rowToPage, rowToStalePage, rowToChunk, rowToSearchResult, parseEmbedding, tryParseEmbedding, takeRowToTake, takeHitRowToHit, isUndefinedTableError, warnOncePerProcess } from './utils.ts';
|
||||
import { resolveBoostMap, resolveHardExcludes } from './search/source-boost.ts';
|
||||
import { buildSourceFactorCase, buildHardExcludeClause, buildVisibilityClause, buildRecencyComponentSql, buildBestPerPagePoolCte, buildOrFallbackWebsearchQuery } from './search/sql-ranking.ts';
|
||||
import { unverifiedExtractionFragment } from './extraction-review.ts';
|
||||
import { DEFAULT_EMBEDDING_MODEL, DEFAULT_EMBEDDING_DIMENSIONS } from './ai/defaults.ts';
|
||||
import { DELETE_BATCH_SIZE } from './engine-constants.ts';
|
||||
import { SOURCE_CONFIG_OBJECT_SQL } from './source-config-sql.ts';
|
||||
@@ -2120,7 +2121,10 @@ export class PostgresEngine implements BrainEngine {
|
||||
// innerLimit scales with offset to preserve the pagination contract:
|
||||
// a fixed cap of 100 would silently empty offset > 100.
|
||||
const boostMap = resolveBoostMap();
|
||||
const sourceFactorCaseOnSlug = buildSourceFactorCase('slug', boostMap, opts?.detail);
|
||||
// issue #160: the guard predicate is projected as `unverified_stub` in
|
||||
// hnsw_candidates (frontmatter isn't otherwise available at re-rank), so
|
||||
// unverified auto-extracted stubs get factor 1.0, not the people/ 1.2x.
|
||||
const sourceFactorCaseOnSlug = buildSourceFactorCase('slug', boostMap, opts?.detail, 'unverified_stub');
|
||||
const hardExcludePrefixes = resolveHardExcludes(opts?.exclude_slug_prefixes, opts?.include_slug_prefixes);
|
||||
const hardExcludeClause = buildHardExcludeClause('p.slug', hardExcludePrefixes);
|
||||
const innerLimit = offset + Math.max(limit * 5, 100);
|
||||
@@ -2220,6 +2224,7 @@ export class PostgresEngine implements BrainEngine {
|
||||
CASE WHEN NULLIF(regexp_replace(p.frontmatter->>'message_id', '^[[:space:]]+|[[:space:]]+$', '', 'g'), '') IS NOT NULL
|
||||
THEN NULLIF(p.frontmatter->>'subject', '') END AS source_subject,
|
||||
cc.id as chunk_id, cc.chunk_index, cc.chunk_text, cc.chunk_source,
|
||||
(${unverifiedExtractionFragment('p')}) AS unverified_stub,
|
||||
1 - (cc.${col} <=> ${castSql}) AS raw_score
|
||||
FROM content_chunks cc
|
||||
JOIN pages p ON p.id = cc.page_id
|
||||
@@ -3559,6 +3564,20 @@ export class PostgresEngine implements BrainEngine {
|
||||
return result;
|
||||
}
|
||||
|
||||
async getUnverifiedExtractionPageIds(pageIds: number[]): Promise<Set<number>> {
|
||||
if (pageIds.length === 0) return new Set();
|
||||
const sql = this.sql;
|
||||
// Predicate is the shared unverifiedExtractionFragment (issue #160) so
|
||||
// this query and the SQL-side source-boost guard can never drift.
|
||||
const rows = await sql.unsafe(
|
||||
`SELECT id FROM pages
|
||||
WHERE id = ANY($1::int[])
|
||||
AND ${unverifiedExtractionFragment('pages')}`,
|
||||
[pageIds] as never,
|
||||
);
|
||||
return new Set((rows as unknown as { id: number }[]).map((r) => Number(r.id)));
|
||||
}
|
||||
|
||||
async getPageTimestamps(slugs: string[]): Promise<Map<string, Date>> {
|
||||
if (slugs.length === 0) return new Map();
|
||||
const sql = this.sql;
|
||||
|
||||
@@ -76,6 +76,37 @@ export async function stampContentFlags(engine: BrainEngine, results: SearchResu
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Extraction quarantine lane (issue #160). Stamps `SearchResult.unverified`
|
||||
* for any result whose page is an unverified auto-extracted entity stub
|
||||
* (frontmatter `provenance: 'auto-extracted'` + `status: 'unverified'`).
|
||||
* MUST run PRE-fusion: rrfFusion/rrfFusionWeighted read the flag to skip the
|
||||
* COMPILED_TRUTH_BOOST for these pages, so a stub fabricated by hostile
|
||||
* ingested text ranks as ordinary content, never with entity authority.
|
||||
* One batched query over the candidate arms' page_ids. Fail-open on the
|
||||
* fetch (a marker-fetch failure must not break retrieval) — the boost then
|
||||
* applies, but the SQL-side source-boost guard still holds.
|
||||
*/
|
||||
export async function stampUnverifiedExtractions(
|
||||
engine: BrainEngine,
|
||||
results: SearchResult[],
|
||||
): Promise<void> {
|
||||
if (results.length === 0) return;
|
||||
try {
|
||||
const ids = [...new Set(
|
||||
results.map((r) => r.page_id).filter((n): n is number => typeof n === 'number' && Number.isFinite(n)),
|
||||
)];
|
||||
if (ids.length === 0) return;
|
||||
const unverified = await engine.getUnverifiedExtractionPageIds(ids);
|
||||
if (unverified.size === 0) return;
|
||||
for (const r of results) {
|
||||
if (unverified.has(r.page_id)) r.unverified = true;
|
||||
}
|
||||
} catch {
|
||||
// best-effort: never break retrieval.
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* v0.42.20.0 — bounded drain (was an unbounded `Promise.allSettled`, codex
|
||||
* confirmed; TODOS retrofit). Mirrors `awaitPendingLastRetrievedWrites`: races
|
||||
@@ -1107,6 +1138,9 @@ export async function hybridSearch(
|
||||
// valuable exactly when vector is unavailable). The title arm fuses here
|
||||
// too — an exact-title lookup on a keyless install is precisely where
|
||||
// chunk-grain keyword FTS alone fails (D1).
|
||||
// issue #160: stamp unverified stubs BEFORE fusion so the compiled-truth
|
||||
// boost skips them (flag survives fusion's result spread).
|
||||
await stampUnverifiedExtractions(engine, [...keywordResults, ...titleResults, ...relationalList]);
|
||||
let noEmbedResults = keywordResults;
|
||||
if (relationalList.length > 0 || titleResults.length > 0) {
|
||||
const fk = opts?.rrfK ?? RRF_K;
|
||||
@@ -1342,6 +1376,9 @@ export async function hybridSearch(
|
||||
// v0.43: fuse the relational arm with keyword via RRF so typed-edge
|
||||
// answers survive even when vector is unavailable. The title arm fuses
|
||||
// here too (same rationale as the no-embedding-provider path — D1).
|
||||
// issue #160: stamp unverified stubs BEFORE fusion (see the
|
||||
// no-embedding-provider path for rationale).
|
||||
await stampUnverifiedExtractions(engine, [...keywordResults, ...titleResults, ...relationalList]);
|
||||
let fallbackResults = keywordResults;
|
||||
if (relationalList.length > 0 || titleResults.length > 0) {
|
||||
const fk = opts?.rrfK ?? RRF_K;
|
||||
@@ -1431,6 +1468,10 @@ export async function hybridSearch(
|
||||
allLists.push({ list: relationalList, k: baseRrfK });
|
||||
}
|
||||
|
||||
// issue #160: stamp unverified auto-extracted stubs across ALL candidate
|
||||
// arms BEFORE fusion so the compiled-truth authority boost skips them.
|
||||
await stampUnverifiedExtractions(engine, allLists.flatMap((l) => l.list));
|
||||
|
||||
let fused = rrfFusionWeighted(allLists, detail !== 'high');
|
||||
|
||||
// Cosine re-scoring before dedup so semantically better chunks survive.
|
||||
@@ -1997,7 +2038,9 @@ export function rrfFusionWeighted(
|
||||
if (maxScore > 0) {
|
||||
for (const e of entries) {
|
||||
e.score = e.score / maxScore;
|
||||
const boost = applyBoost && e.result.chunk_source === 'compiled_truth' ? COMPILED_TRUTH_BOOST : 1.0;
|
||||
// issue #160: unverified auto-extracted stubs (stamped pre-fusion by
|
||||
// stampUnverifiedExtractions) never get the compiled-truth authority boost.
|
||||
const boost = applyBoost && e.result.chunk_source === 'compiled_truth' && e.result.unverified !== true ? COMPILED_TRUTH_BOOST : 1.0;
|
||||
e.score *= boost;
|
||||
}
|
||||
}
|
||||
@@ -2040,8 +2083,9 @@ export function rrfFusion(lists: SearchResult[][], k: number, applyBoost = true)
|
||||
const rawScore = e.score;
|
||||
e.score = e.score / maxScore;
|
||||
|
||||
// Apply compiled truth boost after normalization (skip for detail=high)
|
||||
const boost = applyBoost && e.result.chunk_source === 'compiled_truth' ? COMPILED_TRUTH_BOOST : 1.0;
|
||||
// Apply compiled truth boost after normalization (skip for detail=high;
|
||||
// skip for unverified auto-extracted stubs — issue #160)
|
||||
const boost = applyBoost && e.result.chunk_source === 'compiled_truth' && e.result.unverified !== true ? COMPILED_TRUTH_BOOST : 1.0;
|
||||
e.score *= boost;
|
||||
|
||||
if (DEBUG) {
|
||||
|
||||
@@ -18,6 +18,7 @@
|
||||
*/
|
||||
|
||||
import { quarantineFilterFragment } from '../quarantine.ts';
|
||||
import { unverifiedExtractionFragment } from '../extraction-review.ts';
|
||||
|
||||
/**
|
||||
* Escape `%`, `_`, and `\` so a string can be used as a LIKE prefix literal.
|
||||
@@ -63,6 +64,7 @@ export function buildSourceFactorCase(
|
||||
slugColumn: string,
|
||||
boostMap: Record<string, number>,
|
||||
detail: 'low' | 'medium' | 'high' | undefined,
|
||||
unverifiedGuardColumn?: string,
|
||||
): string {
|
||||
// Loose-string guard: agents passing `"HIGH"` or `"high "` over MCP/JSON
|
||||
// should still hit the temporal-bypass path. TypeScript narrows `detail`
|
||||
@@ -80,7 +82,26 @@ export function buildSourceFactorCase(
|
||||
`WHEN ${slugColumn} LIKE ${buildLikePrefixLiteral(prefix)} THEN ${factor}`
|
||||
).join(' ');
|
||||
|
||||
return `(CASE ${whens} ELSE 1.0 END)`;
|
||||
// Extraction quarantine lane (issue #160): unverified auto-extracted stubs
|
||||
// never receive the namespace-authority factor (people/ / companies/ 1.2x)
|
||||
// — they rank as ordinary content until promoted. Two forms:
|
||||
// - table-qualified slug column ('p.slug'): reference the sibling
|
||||
// `frontmatter` column inline via unverifiedExtractionFragment.
|
||||
// - bare column + `unverifiedGuardColumn`: the vector arm's re-rank CTE
|
||||
// has no frontmatter column, so its inner hnsw_candidates CTE projects
|
||||
// the predicate as a boolean (`... AS unverified_stub`) and passes the
|
||||
// column name here. Without this the 1.2x would apply INSIDE the
|
||||
// scored/best_per_page pipeline pre-LIMIT — an unverified stub could
|
||||
// outrank AND evict a legitimate page from the candidate pool, which
|
||||
// nothing downstream can restore.
|
||||
const alias = slugColumn.includes('.') ? slugColumn.split('.')[0] : null;
|
||||
const unverifiedGuard = unverifiedGuardColumn
|
||||
? `WHEN ${unverifiedGuardColumn} THEN 1.0 `
|
||||
: alias
|
||||
? `WHEN ${unverifiedExtractionFragment(alias)} THEN 1.0 `
|
||||
: '';
|
||||
|
||||
return `(CASE ${unverifiedGuard}${whens} ELSE 1.0 END)`;
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -699,6 +699,17 @@ export interface SearchResult {
|
||||
* Absent when the page is clean.
|
||||
*/
|
||||
content_flag?: { reason: string; detail: string };
|
||||
/**
|
||||
* Extraction quarantine lane (issue #160): true when the result's page is
|
||||
* an unverified auto-extracted entity stub (frontmatter
|
||||
* `provenance: 'auto-extracted'` + `status: 'unverified'`). Such pages are
|
||||
* excluded from the compiled-truth authority boost and the namespace
|
||||
* source-boost — they rank as ordinary content — and this marker tells the
|
||||
* agent the page has NOT been reviewed by the owner. Stamped pre-fusion by
|
||||
* `stampUnverifiedExtractions` (hybrid.ts). Absent for reviewed/ordinary
|
||||
* pages.
|
||||
*/
|
||||
unverified?: boolean;
|
||||
/**
|
||||
* v0.36 (cross-modal wave): the chunk's modality discriminator from
|
||||
* content_chunks.modality. 'text' for the existing text-embedding rows,
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
/**
|
||||
* Extraction quarantine lane (issue #160) — LIVE Postgres parity.
|
||||
*
|
||||
* The PGLite coverage lives in test/extraction-review.test.ts; this file
|
||||
* re-runs the SQL-touching pieces on a real Postgres so the shared
|
||||
* `unverifiedExtractionFragment` predicate, the engine method
|
||||
* `getUnverifiedExtractionPageIds`, the source-boost guard inside
|
||||
* `buildSourceFactorCase`, and the review-op raw SQL are proven on both
|
||||
* engines (PGLite can hide postgres.js-specific behavior).
|
||||
*
|
||||
* Gated by DATABASE_URL via hasDatabase(); skips cleanly when unset.
|
||||
*
|
||||
* Run: DATABASE_URL=... bun test test/e2e/extraction-review-postgres.test.ts
|
||||
*/
|
||||
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
||||
import type { PostgresEngine } from '../../src/core/postgres-engine.ts';
|
||||
import { hasDatabase, setupDB, teardownDB } from './helpers.ts';
|
||||
import { enrichEntity } from '../../src/core/enrichment-service.ts';
|
||||
import { isUnverifiedExtraction, STATUS_VERIFIED, EXTRACTION_STATUS_KEY } from '../../src/core/extraction-review.ts';
|
||||
import { operationsByName, type OperationContext } from '../../src/core/operations.ts';
|
||||
|
||||
const RUN = hasDatabase();
|
||||
const d = RUN ? describe : describe.skip;
|
||||
|
||||
let engine: PostgresEngine;
|
||||
|
||||
function ctx(over: Partial<OperationContext> = {}): OperationContext {
|
||||
return {
|
||||
engine,
|
||||
config: {} as OperationContext['config'],
|
||||
logger: { info() {}, warn() {}, error() {}, debug() {} } as unknown as OperationContext['logger'],
|
||||
dryRun: false,
|
||||
remote: true,
|
||||
sourceId: 'default',
|
||||
...over,
|
||||
} as OperationContext;
|
||||
}
|
||||
|
||||
d('extraction quarantine lane (live Postgres)', () => {
|
||||
beforeAll(async () => {
|
||||
engine = await setupDB();
|
||||
}, 60_000);
|
||||
|
||||
afterAll(async () => {
|
||||
await teardownDB();
|
||||
}, 60_000);
|
||||
|
||||
test('untrusted enrichEntity → markers; getUnverifiedExtractionPageIds sees them', async () => {
|
||||
await enrichEntity(engine, { entityName: 'Pg Fake', entityType: 'person', context: 'c', sourceSlug: 's' });
|
||||
await enrichEntity(engine, { entityName: 'Pg Real', entityType: 'person', context: 'c', sourceSlug: 's' }, { trusted: true });
|
||||
const fake = await engine.getPage('people/pg-fake');
|
||||
const real = await engine.getPage('people/pg-real');
|
||||
expect(isUnverifiedExtraction(fake!.frontmatter)).toBe(true);
|
||||
expect(isUnverifiedExtraction(real!.frontmatter)).toBe(false);
|
||||
const set = await engine.getUnverifiedExtractionPageIds([fake!.id, real!.id]);
|
||||
expect(set.has(fake!.id)).toBe(true);
|
||||
expect(set.has(real!.id)).toBe(false);
|
||||
});
|
||||
|
||||
test('SQL source-boost guard: unverified stub loses the people/ 1.2x in searchKeyword', async () => {
|
||||
await engine.upsertChunks('people/pg-fake', [{ chunk_index: 0, chunk_text: 'flurbo synergy report alpha', chunk_source: 'compiled_truth', token_count: 4 }]);
|
||||
await engine.upsertChunks('people/pg-real', [{ chunk_index: 0, chunk_text: 'flurbo synergy report bravo', chunk_source: 'compiled_truth', token_count: 4 }]);
|
||||
const rows = await engine.searchKeyword('flurbo', { limit: 10 });
|
||||
const fake = rows.find((r) => r.slug === 'people/pg-fake')!;
|
||||
const real = rows.find((r) => r.slug === 'people/pg-real')!;
|
||||
expect(fake).toBeDefined();
|
||||
expect(real).toBeDefined();
|
||||
// Same base ts_rank; only the verified page carries the 1.2 factor.
|
||||
expect(real.score / fake.score).toBeCloseTo(1.2, 5);
|
||||
});
|
||||
|
||||
test('vector arm: unverified stub gets source factor 1.0 in searchVector re-rank', async () => {
|
||||
// The 1.2x people/ factor multiplies raw_score inside the scored CTE,
|
||||
// pre-LIMIT — the guard column projected in hnsw_candidates must zero it
|
||||
// out for unverified stubs. Identical basis embeddings → identical
|
||||
// cosine → the score ratio IS the factor. (1536-dim basis vectors match
|
||||
// the shared e2e schema, same as test/e2e/engine-parity.test.ts.)
|
||||
const basis = new Float32Array(1536);
|
||||
basis[7] = 1.0;
|
||||
await engine.upsertChunks('people/pg-fake', [{ chunk_index: 1, chunk_text: 'vec alpha', chunk_source: 'compiled_truth', embedding: basis, token_count: 2 }]);
|
||||
await engine.upsertChunks('people/pg-real', [{ chunk_index: 1, chunk_text: 'vec bravo', chunk_source: 'compiled_truth', embedding: basis, token_count: 2 }]);
|
||||
const rows = await engine.searchVector(basis, { limit: 10 });
|
||||
const fake = rows.find((r) => r.slug === 'people/pg-fake')!;
|
||||
const real = rows.find((r) => r.slug === 'people/pg-real')!;
|
||||
expect(fake).toBeDefined();
|
||||
expect(real).toBeDefined();
|
||||
expect(real.score / fake.score).toBeCloseTo(1.2, 5);
|
||||
});
|
||||
|
||||
test('extraction_pending + extraction_review promote/reject run on Postgres', async () => {
|
||||
const pending = (await operationsByName['extraction_pending']!.handler(ctx(), {})) as {
|
||||
pending: Array<{ slug: string }>;
|
||||
};
|
||||
expect(pending.pending.map((r) => r.slug)).toContain('people/pg-fake');
|
||||
|
||||
const out = (await operationsByName['extraction_review']!.handler(ctx({ remote: false }), {
|
||||
action: 'promote', slugs: ['people/pg-fake'],
|
||||
})) as { results: Array<{ slug: string; status: string }> };
|
||||
expect(out.results[0].status).toBe('promoted');
|
||||
const promoted = await engine.getPage('people/pg-fake');
|
||||
expect(promoted!.frontmatter[EXTRACTION_STATUS_KEY]).toBe(STATUS_VERIFIED);
|
||||
|
||||
await enrichEntity(engine, { entityName: 'Pg Reject', entityType: 'person', context: 'c', sourceSlug: 's' });
|
||||
const rej = (await operationsByName['extraction_review']!.handler(ctx({ remote: false }), {
|
||||
action: 'reject', slugs: ['people/pg-reject'],
|
||||
})) as { results: Array<{ slug: string; status: string }> };
|
||||
expect(rej.results[0].status).toBe('rejected');
|
||||
expect(await engine.getPage('people/pg-reject')).toBeNull();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,446 @@
|
||||
/**
|
||||
* Extraction quarantine lane (issue #160).
|
||||
*
|
||||
* `extractAndEnrich` regex-extracts entity names from arbitrary ingested text
|
||||
* and creates people/ + companies/ stub pages. These tests pin the lane
|
||||
* end-to-end:
|
||||
* - fail-closed trust: only an explicit `trusted: true` (which the op layer
|
||||
* only grants for ctx.remote === false AND --trusted-extraction) writes
|
||||
* authoritative pages; undefined/false/remote → quarantine markers.
|
||||
* - unverified stubs are excluded from authoritative retrieval boosts
|
||||
* (compiled-truth fusion boost + the SQL namespace source-boost) and
|
||||
* carry `unverified: true` in search-result metadata.
|
||||
* - review queue: extraction_pending lists; extraction_review promotes
|
||||
* (status → verified) / rejects (soft-delete) in batch, owner-only.
|
||||
* - doctor nudge: unverified_extractions counts stale stubs.
|
||||
*
|
||||
* Hermetic via PGLite (both engines share the SQL through
|
||||
* unverifiedExtractionFragment + the same literal method SQL; postgres runs
|
||||
* via the DATABASE_URL-gated e2e lane).
|
||||
*/
|
||||
|
||||
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
|
||||
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
||||
import { configureGateway, resetGateway } from '../src/core/ai/gateway.ts';
|
||||
import {
|
||||
quarantineMarkers,
|
||||
isUnverifiedExtraction,
|
||||
unverifiedExtractionFragment,
|
||||
EXTRACTION_STATUS_KEY,
|
||||
STATUS_UNVERIFIED,
|
||||
STATUS_VERIFIED,
|
||||
PROVENANCE_AUTO_EXTRACTED,
|
||||
} from '../src/core/extraction-review.ts';
|
||||
import { enrichEntity, extractAndEnrich } from '../src/core/enrichment-service.ts';
|
||||
import { rrfFusion, hybridSearch } from '../src/core/search/hybrid.ts';
|
||||
import { buildSourceFactorCase } from '../src/core/search/sql-ranking.ts';
|
||||
import { operationsByName, OperationError, type OperationContext } from '../src/core/operations.ts';
|
||||
import { checkUnverifiedExtractions } from '../src/commands/doctor.ts';
|
||||
import { categorizeCheck } from '../src/core/doctor-categories.ts';
|
||||
import type { SearchResult } from '../src/core/types.ts';
|
||||
|
||||
let engine: PGLiteEngine;
|
||||
|
||||
function basisEmbedding(idx: number, dim = 1536): Float32Array {
|
||||
const emb = new Float32Array(dim);
|
||||
emb[idx % dim] = 1.0;
|
||||
return emb;
|
||||
}
|
||||
|
||||
beforeAll(async () => {
|
||||
// Deterministic no-embedding-provider path: configure the gateway with NO
|
||||
// auth env so hybridSearch never attempts a real embedding call, even on a
|
||||
// dev machine with provider keys in process.env. Pins the vector dim too
|
||||
// (shard-order defense, same class as doctor-hidden-by-search-policy).
|
||||
configureGateway({
|
||||
embedding_model: 'openai:text-embedding-3-large',
|
||||
embedding_dimensions: 1536,
|
||||
env: {},
|
||||
});
|
||||
engine = new PGLiteEngine();
|
||||
await engine.connect({});
|
||||
await engine.initSchema();
|
||||
}, 60_000);
|
||||
|
||||
afterAll(async () => {
|
||||
await engine.disconnect();
|
||||
resetGateway();
|
||||
});
|
||||
|
||||
beforeEach(async () => {
|
||||
await engine.executeRaw('DELETE FROM content_chunks');
|
||||
await engine.executeRaw('DELETE FROM links');
|
||||
await engine.executeRaw('DELETE FROM timeline_entries');
|
||||
await engine.executeRaw('DELETE FROM pages');
|
||||
});
|
||||
|
||||
function ctx(over: Partial<OperationContext> = {}): OperationContext {
|
||||
return {
|
||||
engine,
|
||||
config: {} as OperationContext['config'],
|
||||
logger: { info() {}, warn() {}, error() {}, debug() {} } as unknown as OperationContext['logger'],
|
||||
dryRun: false,
|
||||
remote: true,
|
||||
sourceId: 'default',
|
||||
...over,
|
||||
} as OperationContext;
|
||||
}
|
||||
|
||||
/** ctx with `remote` deleted entirely — the type-bypass case the fail-closed
|
||||
* invariant exists for ("anything not strictly false is untrusted"). */
|
||||
function ctxNoRemote(): OperationContext {
|
||||
const c = ctx() as unknown as Record<string, unknown>;
|
||||
delete c.remote;
|
||||
return c as unknown as OperationContext;
|
||||
}
|
||||
|
||||
const extract_entities = operationsByName['extract_entities']!;
|
||||
const extraction_pending = operationsByName['extraction_pending']!;
|
||||
const extraction_review = operationsByName['extraction_review']!;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Marker module (pure)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe('extraction-review markers', () => {
|
||||
test('quarantineMarkers → provenance + status pair', () => {
|
||||
expect(quarantineMarkers()).toEqual({ provenance: PROVENANCE_AUTO_EXTRACTED, status: STATUS_UNVERIFIED });
|
||||
});
|
||||
|
||||
test('isUnverifiedExtraction requires BOTH markers', () => {
|
||||
expect(isUnverifiedExtraction(quarantineMarkers())).toBe(true);
|
||||
expect(isUnverifiedExtraction({ status: 'unverified' })).toBe(false);
|
||||
expect(isUnverifiedExtraction({ provenance: 'auto-extracted' })).toBe(false);
|
||||
expect(isUnverifiedExtraction({ provenance: 'auto-extracted', status: 'verified' })).toBe(false);
|
||||
expect(isUnverifiedExtraction({ status: 'unverified', provenance: 'user' })).toBe(false);
|
||||
expect(isUnverifiedExtraction(null)).toBe(false);
|
||||
expect(isUnverifiedExtraction(undefined)).toBe(false);
|
||||
});
|
||||
|
||||
test('SQL fragment references both keys on the given alias', () => {
|
||||
const frag = unverifiedExtractionFragment('p');
|
||||
expect(frag).toContain("p.frontmatter");
|
||||
expect(frag).toContain(PROVENANCE_AUTO_EXTRACTED);
|
||||
expect(frag).toContain(STATUS_UNVERIFIED);
|
||||
});
|
||||
|
||||
test('buildSourceFactorCase guards unverified stubs in both forms', () => {
|
||||
const qualified = buildSourceFactorCase('p.slug', { 'people/': 1.2 }, 'low');
|
||||
expect(qualified).toContain(unverifiedExtractionFragment('p'));
|
||||
expect(qualified.indexOf(unverifiedExtractionFragment('p'))).toBeLessThan(qualified.indexOf('people/'));
|
||||
// Vector re-rank form: bare slug column + pre-computed guard column
|
||||
// (projected in hnsw_candidates) — the guard WHEN must come first.
|
||||
const guarded = buildSourceFactorCase('slug', { 'people/': 1.2 }, 'low', 'unverified_stub');
|
||||
expect(guarded).toContain('CASE WHEN unverified_stub THEN 1.0');
|
||||
expect(guarded.indexOf('unverified_stub')).toBeLessThan(guarded.indexOf('people/'));
|
||||
});
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Fusion boost skip (pure)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe('rrfFusion compiled-truth boost skip', () => {
|
||||
function result(slug: string, over: Partial<SearchResult> = {}): SearchResult {
|
||||
return {
|
||||
slug,
|
||||
page_id: over.page_id ?? 1,
|
||||
title: slug,
|
||||
type: 'person',
|
||||
chunk_text: 'x',
|
||||
chunk_source: 'compiled_truth',
|
||||
chunk_id: over.chunk_id ?? 1,
|
||||
chunk_index: 0,
|
||||
score: 1,
|
||||
stale: false,
|
||||
...over,
|
||||
} as SearchResult;
|
||||
}
|
||||
|
||||
test('unverified compiled_truth chunk does NOT get the 2x boost', () => {
|
||||
const verified = result('people/real', { page_id: 1, chunk_id: 1 });
|
||||
const unverified = result('people/fake', { page_id: 2, chunk_id: 2, unverified: true });
|
||||
// Two single-result lists at the same rank → identical raw RRF scores.
|
||||
const fused = rrfFusion([[verified], [unverified]], 60, true);
|
||||
const v = fused.find((r) => r.slug === 'people/real')!;
|
||||
const u = fused.find((r) => r.slug === 'people/fake')!;
|
||||
expect(u.unverified).toBe(true);
|
||||
// Same normalized base; verified gets 2.0x, unverified stays 1.0x.
|
||||
expect(v.score).toBeCloseTo(u.score * 2.0, 10);
|
||||
});
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Enrichment write path (PGLite)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe('enrichEntity trust lane', () => {
|
||||
test('default (opts omitted) → fail-closed quarantine markers', async () => {
|
||||
const r = await enrichEntity(engine, {
|
||||
entityName: 'Mallory Fake',
|
||||
entityType: 'person',
|
||||
context: 'injected sentence',
|
||||
sourceSlug: 'inbox/hostile-email',
|
||||
});
|
||||
expect(r.action).toBe('created');
|
||||
expect(r.quarantined).toBe(true);
|
||||
const page = await engine.getPage('people/mallory-fake');
|
||||
expect(isUnverifiedExtraction(page!.frontmatter)).toBe(true);
|
||||
expect(page!.frontmatter[EXTRACTION_STATUS_KEY]).toBe(STATUS_UNVERIFIED);
|
||||
});
|
||||
|
||||
test('trusted: true → direct authoritative write, no markers', async () => {
|
||||
const r = await enrichEntity(engine, {
|
||||
entityName: 'Alice Example',
|
||||
entityType: 'person',
|
||||
context: 'my own notes',
|
||||
sourceSlug: 'notes/daily',
|
||||
}, { trusted: true });
|
||||
expect(r.action).toBe('created');
|
||||
expect(r.quarantined).toBeUndefined();
|
||||
const page = await engine.getPage('people/alice-example');
|
||||
expect(isUnverifiedExtraction(page!.frontmatter)).toBe(false);
|
||||
expect(page!.frontmatter[EXTRACTION_STATUS_KEY]).toBeUndefined();
|
||||
});
|
||||
|
||||
test('trusted: false explicitly → quarantine markers', async () => {
|
||||
await enrichEntity(engine, {
|
||||
entityName: 'Widget Co Corp',
|
||||
entityType: 'company',
|
||||
context: 'ctx',
|
||||
sourceSlug: 'inbox/x',
|
||||
}, { trusted: false });
|
||||
const page = await engine.getPage('companies/widget-co-corp');
|
||||
expect(isUnverifiedExtraction(page!.frontmatter)).toBe(true);
|
||||
});
|
||||
|
||||
test('vector arm: unverified stub gets source factor 1.0, not the people/ 1.2x', async () => {
|
||||
// The 1.2x namespace factor is applied INSIDE searchVector's re-rank SQL,
|
||||
// pre-LIMIT — an unguarded stub would outrank AND could evict legitimate
|
||||
// pages from the candidate pool before fusion ever sees them. Identical
|
||||
// basis embeddings → identical cosine → the score ratio IS the factor.
|
||||
await enrichEntity(engine, { entityName: 'Vec Fake', entityType: 'person', context: 'c', sourceSlug: 's' });
|
||||
await enrichEntity(engine, { entityName: 'Vec Real', entityType: 'person', context: 'c', sourceSlug: 's' }, { trusted: true });
|
||||
const e = basisEmbedding(7);
|
||||
await engine.upsertChunks('people/vec-fake', [{ chunk_index: 0, chunk_text: 'vector text alpha', chunk_source: 'compiled_truth', embedding: e, token_count: 3 }]);
|
||||
await engine.upsertChunks('people/vec-real', [{ chunk_index: 0, chunk_text: 'vector text bravo', chunk_source: 'compiled_truth', embedding: e, token_count: 3 }]);
|
||||
const rows = await engine.searchVector(e, { limit: 10 });
|
||||
const fake = rows.find((r) => r.slug === 'people/vec-fake')!;
|
||||
const real = rows.find((r) => r.slug === 'people/vec-real')!;
|
||||
expect(fake).toBeDefined();
|
||||
expect(real).toBeDefined();
|
||||
expect(real.score / fake.score).toBeCloseTo(1.2, 5);
|
||||
});
|
||||
|
||||
test('getUnverifiedExtractionPageIds returns only marked pages', async () => {
|
||||
await enrichEntity(engine, { entityName: 'Fake Guy', entityType: 'person', context: 'c', sourceSlug: 's' });
|
||||
await enrichEntity(engine, { entityName: 'Real Guy', entityType: 'person', context: 'c', sourceSlug: 's' }, { trusted: true });
|
||||
const fake = await engine.getPage('people/fake-guy');
|
||||
const real = await engine.getPage('people/real-guy');
|
||||
const set = await engine.getUnverifiedExtractionPageIds([fake!.id, real!.id]);
|
||||
expect(set.has(fake!.id)).toBe(true);
|
||||
expect(set.has(real!.id)).toBe(false);
|
||||
expect((await engine.getUnverifiedExtractionPageIds([])).size).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Ops: trust-boundary matrix + review queue
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe('extract_entities op trust boundary', () => {
|
||||
const TEXT = 'I had lunch with Bobby Injected today. He said Evil Widgets Inc is pivoting.';
|
||||
|
||||
test('remote: true → quarantined even WITH trusted_extraction flag', async () => {
|
||||
const out = (await extract_entities.handler(ctx({ remote: true }), {
|
||||
text: TEXT, source_slug: 'inbox/mail', trusted_extraction: true,
|
||||
})) as { trusted: boolean; quarantined: number; count: number };
|
||||
expect(out.trusted).toBe(false);
|
||||
expect(out.count).toBeGreaterThan(0);
|
||||
expect(out.quarantined).toBe(out.count);
|
||||
const page = await engine.getPage('people/bobby-injected');
|
||||
expect(isUnverifiedExtraction(page!.frontmatter)).toBe(true);
|
||||
});
|
||||
|
||||
test('remote UNSET (type bypass) → fail-closed quarantine', async () => {
|
||||
const out = (await extract_entities.handler(ctxNoRemote(), {
|
||||
text: TEXT, source_slug: 'inbox/mail', trusted_extraction: true,
|
||||
})) as { trusted: boolean; quarantined: number };
|
||||
expect(out.trusted).toBe(false);
|
||||
expect(out.quarantined).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
test('remote: false WITHOUT flag → still quarantined (explicit opt-in required)', async () => {
|
||||
const out = (await extract_entities.handler(ctx({ remote: false }), {
|
||||
text: TEXT, source_slug: 'inbox/mail',
|
||||
})) as { trusted: boolean; quarantined: number };
|
||||
expect(out.trusted).toBe(false);
|
||||
expect(out.quarantined).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
test('resource guards: oversize text rejected; entity flood capped + surfaced', async () => {
|
||||
// Oversize input → loud invalid_params, nothing written.
|
||||
await expect(extract_entities.handler(ctx(), {
|
||||
text: 'A'.repeat(200_001), source_slug: 'inbox/big',
|
||||
})).rejects.toBeInstanceOf(OperationError);
|
||||
// 300 distinct name-shaped tokens → capped at 200, truncated surfaced.
|
||||
// 300 distinct two-word names (letters only — the extractor regex is
|
||||
// [A-Z][a-z]+ per word, digits would break the match).
|
||||
const flood = Array.from({ length: 300 }, (_, i) =>
|
||||
`Flood Name${String.fromCharCode(97 + (i % 26))}${String.fromCharCode(97 + Math.floor(i / 26))}`,
|
||||
).join('. ');
|
||||
const out = (await extract_entities.handler(ctx(), { text: flood, source_slug: 'inbox/flood' })) as {
|
||||
count: number; entities_found: number; truncated: boolean;
|
||||
};
|
||||
expect(out.entities_found).toBeGreaterThan(200);
|
||||
expect(out.count).toBe(200);
|
||||
expect(out.truncated).toBe(true);
|
||||
}, 120_000);
|
||||
|
||||
test('remote: false WITH --trusted-extraction → direct authoritative write', async () => {
|
||||
const out = (await extract_entities.handler(ctx({ remote: false }), {
|
||||
text: TEXT, source_slug: 'notes/mine', trusted_extraction: true,
|
||||
})) as { trusted: boolean; quarantined: number; count: number };
|
||||
expect(out.trusted).toBe(true);
|
||||
expect(out.quarantined).toBe(0);
|
||||
const page = await engine.getPage('people/bobby-injected');
|
||||
expect(isUnverifiedExtraction(page!.frontmatter)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('extraction_pending + extraction_review', () => {
|
||||
async function seedStub(name: string): Promise<string> {
|
||||
const r = await enrichEntity(engine, { entityName: name, entityType: 'person', context: 'c', sourceSlug: 'inbox/x' });
|
||||
return r.slug;
|
||||
}
|
||||
|
||||
test('pending lists unverified stubs; promoted/rejected drop out', async () => {
|
||||
const a = await seedStub('Fake Aa');
|
||||
const b = await seedStub('Fake Bb');
|
||||
await enrichEntity(engine, { entityName: 'Real Cc', entityType: 'person', context: 'c', sourceSlug: 's' }, { trusted: true });
|
||||
|
||||
const before = (await extraction_pending.handler(ctx(), {})) as { count: number; pending: Array<{ slug: string }> };
|
||||
expect(before.pending.map((r) => r.slug).sort()).toEqual([a, b].sort());
|
||||
|
||||
const out = (await extraction_review.handler(ctx({ remote: false }), {
|
||||
action: 'promote', slugs: [a],
|
||||
})) as { results: Array<{ slug: string; status: string }> };
|
||||
expect(out.results).toEqual([{ slug: a, status: 'promoted' }]);
|
||||
|
||||
const promoted = await engine.getPage(a);
|
||||
expect(promoted!.frontmatter[EXTRACTION_STATUS_KEY]).toBe(STATUS_VERIFIED);
|
||||
// provenance survives as the audit trail.
|
||||
expect(promoted!.frontmatter.provenance).toBe(PROVENANCE_AUTO_EXTRACTED);
|
||||
expect(isUnverifiedExtraction(promoted!.frontmatter)).toBe(false);
|
||||
|
||||
const rej = (await extraction_review.handler(ctx({ remote: false }), {
|
||||
action: 'reject', slugs: b, // CLI string form
|
||||
})) as { results: Array<{ slug: string; status: string }> };
|
||||
expect(rej.results).toEqual([{ slug: b, status: 'rejected' }]);
|
||||
expect(await engine.getPage(b)).toBeNull(); // soft-deleted → hidden
|
||||
|
||||
const after = (await extraction_pending.handler(ctx(), {})) as { count: number };
|
||||
expect(after.count).toBe(0);
|
||||
});
|
||||
|
||||
test('batch promote is batch-friendly and reports per-slug statuses', async () => {
|
||||
const a = await seedStub('Fake Dd');
|
||||
const b = await seedStub('Fake Ee');
|
||||
await enrichEntity(engine, { entityName: 'Real Ff', entityType: 'person', context: 'c', sourceSlug: 's' }, { trusted: true });
|
||||
const out = (await extraction_review.handler(ctx({ remote: false }), {
|
||||
action: 'promote', slugs: [a, b, 'people/real-ff', 'people/missing'],
|
||||
})) as { results: Array<{ slug: string; status: string }> };
|
||||
expect(out.results.map((r) => r.status)).toEqual(['promoted', 'promoted', 'not_unverified', 'not_found']);
|
||||
});
|
||||
|
||||
test('extraction_review is owner-only: remote and unset-trust callers are refused', async () => {
|
||||
const a = await seedStub('Fake Gg');
|
||||
await expect(extraction_review.handler(ctx({ remote: true }), { action: 'promote', slugs: [a] }))
|
||||
.rejects.toBeInstanceOf(OperationError);
|
||||
await expect(extraction_review.handler(ctxNoRemote(), { action: 'promote', slugs: [a] }))
|
||||
.rejects.toBeInstanceOf(OperationError);
|
||||
// and it is not exposed over HTTP MCP at all
|
||||
expect(extraction_review.localOnly).toBe(true);
|
||||
// stub untouched
|
||||
expect(isUnverifiedExtraction((await engine.getPage(a))!.frontmatter)).toBe(true);
|
||||
});
|
||||
|
||||
test('invalid action / empty slugs → invalid_params', async () => {
|
||||
await expect(extraction_review.handler(ctx({ remote: false }), { action: 'bless', slugs: ['x'] }))
|
||||
.rejects.toBeInstanceOf(OperationError);
|
||||
await expect(extraction_review.handler(ctx({ remote: false }), { action: 'promote', slugs: [] }))
|
||||
.rejects.toBeInstanceOf(OperationError);
|
||||
});
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Doctor nudge
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe('unverified_extractions doctor check', () => {
|
||||
test('fresh stubs → ok; stale stubs → warn with review commands', async () => {
|
||||
await enrichEntity(engine, { entityName: 'Fake Hh', entityType: 'person', context: 'c', sourceSlug: 's' });
|
||||
const fresh = await checkUnverifiedExtractions(engine);
|
||||
expect(fresh.status).toBe('ok');
|
||||
|
||||
await engine.executeRaw(`UPDATE pages SET created_at = now() - interval '30 days' WHERE slug = 'people/fake-hh'`);
|
||||
const stale = await checkUnverifiedExtractions(engine, { days: 7 });
|
||||
expect(stale.status).toBe('warn');
|
||||
expect(stale.message).toContain('extraction-pending');
|
||||
expect(stale.message).toContain('extraction-review');
|
||||
expect((stale.details as { count: number }).count).toBe(1);
|
||||
});
|
||||
|
||||
test('categorized as a brain check', () => {
|
||||
expect(categorizeCheck('unverified_extractions')).toBe('brain');
|
||||
});
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// End-to-end: hostile transcript → quarantined stubs, NOT boosted in search
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
describe('e2e: hostile transcript', () => {
|
||||
test('fake entities land quarantined and rank without entity authority', async () => {
|
||||
// 1. Hostile transcript arrives through an agent-facing (remote) caller.
|
||||
const transcript =
|
||||
'Meeting notes. I had lunch with Zorbulon Fakeperson today. ' +
|
||||
'He mentioned the zorbulon pivot is confirmed.';
|
||||
const out = (await extract_entities.handler(ctx({ remote: true }), {
|
||||
text: transcript, source_slug: 'meetings/2026-04-03', trusted_extraction: true,
|
||||
})) as { trusted: boolean; quarantined: number };
|
||||
expect(out.trusted).toBe(false);
|
||||
expect(out.quarantined).toBeGreaterThan(0);
|
||||
|
||||
const stub = await engine.getPage('people/zorbulon-fakeperson');
|
||||
expect(stub).not.toBeNull();
|
||||
expect(isUnverifiedExtraction(stub!.frontmatter)).toBe(true);
|
||||
|
||||
// 2. Owner-authored control page with the same lexical relevance.
|
||||
await engine.putPage('people/zorbulon-realperson', {
|
||||
type: 'person', title: 'Zorbulon Realperson', compiled_truth: 'zorbulon notes', timeline: '', frontmatter: {},
|
||||
});
|
||||
const real = await engine.getPage('people/zorbulon-realperson');
|
||||
|
||||
// 3. Chunk both with equal lexical relevance (distinct texts — identical
|
||||
// ones would be Jaccard-deduped). Enrichment stubs are chunked by the
|
||||
// normal reindex/import pipeline later; seed what it would write.
|
||||
await engine.upsertChunks(stub!.slug, [{ chunk_index: 0, chunk_text: 'zorbulon pivot details from the injected meeting', chunk_source: 'compiled_truth', token_count: 7 }]);
|
||||
await engine.upsertChunks(real!.slug, [{ chunk_index: 0, chunk_text: 'zorbulon launch update in my own written notes', chunk_source: 'compiled_truth', token_count: 7 }]);
|
||||
|
||||
// 4. Search. No embedding provider configured → keyword(+title) fusion path.
|
||||
const results = await hybridSearch(engine, 'zorbulon', { limit: 10 });
|
||||
const fake = results.find((r) => r.slug === stub!.slug);
|
||||
const legit = results.find((r) => r.slug === real!.slug);
|
||||
expect(fake).toBeDefined();
|
||||
expect(legit).toBeDefined();
|
||||
|
||||
// Clearly marked in search-result metadata…
|
||||
expect(fake!.unverified).toBe(true);
|
||||
expect(legit!.unverified).toBeUndefined();
|
||||
// …and stripped of entity authority: the verified page outranks the
|
||||
// injected stub despite identical chunk text (2x compiled-truth boost +
|
||||
// people/ source-boost apply only to the verified page).
|
||||
expect(legit!.score).toBeGreaterThan(fake!.score);
|
||||
});
|
||||
});
|
||||
@@ -6,6 +6,7 @@ import {
|
||||
escapeLikePattern as topLevelEscapeLikePattern,
|
||||
__test__,
|
||||
} from '../src/core/search/sql-ranking.ts';
|
||||
import { unverifiedExtractionFragment } from '../src/core/extraction-review.ts';
|
||||
import {
|
||||
DEFAULT_SOURCE_BOOSTS,
|
||||
DEFAULT_HARD_EXCLUDES,
|
||||
@@ -87,9 +88,11 @@ describe('buildSourceFactorCase', () => {
|
||||
expect(buildSourceFactorCase('p.slug', {}, 'medium')).toBe('1.0');
|
||||
});
|
||||
|
||||
test('emits a CASE expression for non-high detail', () => {
|
||||
test('emits a CASE expression for non-high detail (unverified guard first — issue #160)', () => {
|
||||
const result = buildSourceFactorCase('p.slug', { 'originals/': 1.5 }, 'medium');
|
||||
expect(result).toBe("(CASE WHEN p.slug LIKE 'originals/%' THEN 1.5 ELSE 1.0 END)");
|
||||
expect(result).toBe(
|
||||
`(CASE WHEN ${unverifiedExtractionFragment('p')} THEN 1.0 WHEN p.slug LIKE 'originals/%' THEN 1.5 ELSE 1.0 END)`,
|
||||
);
|
||||
});
|
||||
|
||||
test('sorts prefixes by length descending so longest-match wins', () => {
|
||||
@@ -119,7 +122,9 @@ describe('buildSourceFactorCase', () => {
|
||||
{ 'good/': 1.5, 'nan/': NaN, 'neg/': -1, 'inf/': Infinity },
|
||||
'medium',
|
||||
);
|
||||
expect(result).toBe("(CASE WHEN p.slug LIKE 'good/%' THEN 1.5 ELSE 1.0 END)");
|
||||
expect(result).toBe(
|
||||
`(CASE WHEN ${unverifiedExtractionFragment('p')} THEN 1.0 WHEN p.slug LIKE 'good/%' THEN 1.5 ELSE 1.0 END)`,
|
||||
);
|
||||
});
|
||||
|
||||
test('uses the supplied slug column reference', () => {
|
||||
|
||||
Reference in New Issue
Block a user