mirror of
https://github.com/garrytan/gbrain.git
synced 2026-08-14 08:53:22 +00:00
* feat: search quality boost — compiled truth ranking, detail parameter, cosine re-scoring Compiled truth chunks now rank 2x higher in hybrid search via RRF normalization + source boost. New --detail flag (low/medium/high) controls timeline inclusion. Cosine re-scoring blends query-chunk similarity before dedup for query-specific ranking. Also: remove DISTINCT ON from keyword search (dedup handles per-page capping), add chunk_id + chunk_index to SearchResult, add getEmbeddingsByChunkIds to BrainEngine interface. Inspired by Ramp Labs' "Latent Briefing" paper (April 2026). * feat: RRF normalization, source-aware dedup, detail param in operations RRF scores normalized to 0-1 before 2.0x compiled truth boost. Source-aware dedup guarantees compiled truth chunk per page. Detail parameter added to query operation, dedupResults added to bare search operation. Debug logging via GBRAIN_SEARCH_DEBUG=1. * chore: bump version and changelog (v0.8.1) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: CJK word count in query expansion CJK text is not space-delimited. A query like "向量搜索优化" was counted as 1 word and silently skipped expansion. Now counts characters for CJK queries instead of space-separated tokens. Co-Authored-By: YIING99 <yiing99@users.noreply.github.com> * feat: retrieval evaluation harness — P@k, R@k, MRR, nDCG@k + gbrain eval Full IR evaluation framework: precisionAtK, recallAtK, mrr, ndcgAtK metrics with runEval() orchestrator. gbrain eval CLI with single-run table and A/B comparison mode (--config-a / --config-b) for parameter tuning. HybridSearchOpts now accepts rrfK and dedupOpts overrides. Co-Authored-By: 4shut0sh <4shut0sh@users.noreply.github.com> * test: search quality tests — RRF boost, dedup guarantee, cosine similarity, E2E benchmark 42 new tests across 3 files: - test/search.test.ts: RRF normalization, compiled truth 2x boost, dedup key collision prevention, cosine similarity edge cases, CJK word count detection - test/dedup.test.ts: source-aware compiled truth guarantee, layer interactions, custom maxPerPage, empty/single result edge cases - test/e2e/search-quality.test.ts: full pipeline against PGLite with basis vector embeddings — chunk_id/chunk_index fields, detail parameter filtering, getEmbeddingsByChunkIds, keyword multi-chunk, vector ordering Also: export rrfFusion + cosineSimilarity for unit testing, fix PGLite getEmbeddingsByChunkIds to parse string vectors from pgvector. * test: search quality benchmark with A/B comparison (baseline vs PR#64) Benchmark measures P@1, MRR, nDCG@5, and source accuracy across 8 queries against 5 seeded pages. Key finding: boost helps entity lookups but over-corrects temporal queries. Validates the --detail parameter as the right control mechanism. Output at docs/benchmarks/2026-04-13.md. * feat: query intent classifier — auto-selects detail level, 100% source accuracy Zero-latency heuristic classifier detects query intent from text patterns: - "Who is Pedro?" → entity → detail=low (compiled truth only) - "When did we last meet?" → temporal → detail=high (no boost, natural ranking) - "Variant fund announcement" → event → detail=high - General queries → detail=medium (default with boost) The key insight: skip the 2.0x compiled truth boost for detail=high queries. Temporal/event queries want natural ranking where timeline entries can win. Benchmark results (source accuracy = does the top chunk match expected type): - Baseline: 100% (already good, no boost needed) - Boost only: 71.4% (boost over-corrects temporal queries) - Boost + intent classifier: 100% (best of both worlds) 35 unit tests for the classifier. 590 total tests pass. * feat: query intent classifier — auto-selects detail level, 100% source accuracy Heuristic classifier detects query intent from text patterns (zero latency, no LLM call). Maps temporal queries ("when did we last meet") to detail=high, entity queries ("who is X") to detail=low, events to detail=high. Benchmark results (29 pages, 20 queries, graded relevance): - Baseline: P@1=0.947, MRR=0.974, source accuracy=89.5% - Boost only: P@1=0.895, MRR=0.939, source accuracy=63.2% (over-correction) - Boost + intent: P@1=0.947, MRR=0.974, source accuracy=89.5% (fully recovered) The intent classifier eliminates the boost's over-correction on temporal queries while preserving its benefits for entity lookups. 35 unit tests for the classifier. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test: search quality benchmark with A/B comparison (baseline vs PR#64) Rich benchmark: 29 pages, 58 chunks, 20 queries with graded relevance. Now measures CHUNK-LEVEL quality, not just page-level retrieval. Key findings (C. Boost+Intent vs A. Baseline): - Unique pages in top-10: 7.2 → 8.7 (+21% broader coverage) - Compiled truth ratio: 51.6% → 66.8% (+15pp more signal) - CT-first rate: 100% (compiled truth leads for entity queries) - Timeline accessible: 100% (temporal queries still find dates) - Source accuracy: 89.5% maintained (intent classifier prevents regression) The boost alone (B) causes -26pp source accuracy regression. Intent classifier (C) recovers it fully. * docs: clean benchmark report — ELI10 search quality analysis for PR#64 Replaces two drafts with one clean report. Explains what changed, why it matters, and what the numbers mean. All fictional data, no private info. Key findings: 21% more page coverage per query, 29% more compiled truth in results. Intent classifier prevents boost from burying timeline for temporal queries. Full per-query breakdown with before/after comparison. * chore: remove auto-generated benchmark file (clean version is 2026-04-14-search-quality.md) * docs: update project documentation for search quality boost CLAUDE.md: added search/intent.ts, search/eval.ts, commands/eval.ts to key files. Added 5 new test files (search, dedup, intent, eval, e2e/search-quality). Updated test count from 23+4 to 28+5. Added docs/benchmarks/ to key files. README.md: updated search pipeline diagram with intent classifier, RRF normalization, compiled truth boost, cosine re-scoring, and 5-layer dedup. Added --detail flag explanation and benchmark instructions. CHANGELOG.md: added search quality entries to v0.9.3 (intent classifier, --detail flag, gbrain eval, CJK fix). Credited @4shut0sh and @YIING99. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * docs: headline benchmark gains in changelog * docs: add community attribution rule to CHANGELOG voice section --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: YIING99 <yiing99@users.noreply.github.com> Co-authored-by: 4shut0sh <4shut0sh@users.noreply.github.com>
183 lines
5.1 KiB
TypeScript
183 lines
5.1 KiB
TypeScript
/**
|
|
* 4-Layer Dedup Pipeline + Compiled Truth Guarantee
|
|
* Ported from production Ruby implementation (content_chunk.rb)
|
|
*
|
|
* 1. By source: top 3 chunks per page by score
|
|
* 2. By text similarity: remove chunks >0.85 Jaccard-similar to kept results
|
|
* 3. By type: no page type exceeds 60% of results
|
|
* 4. By page: max N chunks per page (default 2)
|
|
* 5. Compiled truth guarantee: ensure at least 1 compiled_truth chunk per page
|
|
*/
|
|
|
|
import type { SearchResult } from '../types.ts';
|
|
|
|
const COSINE_DEDUP_THRESHOLD = 0.85;
|
|
const MAX_TYPE_RATIO = 0.6;
|
|
const MAX_PER_PAGE = 2;
|
|
|
|
export function dedupResults(
|
|
results: SearchResult[],
|
|
opts?: {
|
|
cosineThreshold?: number;
|
|
maxTypeRatio?: number;
|
|
maxPerPage?: number;
|
|
},
|
|
): SearchResult[] {
|
|
const threshold = opts?.cosineThreshold ?? COSINE_DEDUP_THRESHOLD;
|
|
const maxRatio = opts?.maxTypeRatio ?? MAX_TYPE_RATIO;
|
|
const maxPerPage = opts?.maxPerPage ?? MAX_PER_PAGE;
|
|
|
|
// Preserve pre-dedup input for compiled truth guarantee
|
|
const preDedup = results;
|
|
|
|
let deduped = results;
|
|
|
|
// Layer 1: Top 3 chunks per page by score
|
|
deduped = dedupBySource(deduped);
|
|
|
|
// Layer 2: Text similarity dedup (Jaccard on word sets)
|
|
deduped = dedupByTextSimilarity(deduped, threshold);
|
|
|
|
// Layer 3: Type diversity (no page type exceeds 60%)
|
|
deduped = enforceTypeDiversity(deduped, maxRatio);
|
|
|
|
// Layer 4: Cap chunks per page
|
|
deduped = capPerPage(deduped, maxPerPage);
|
|
|
|
// Final pass: guarantee compiled_truth representation
|
|
deduped = guaranteeCompiledTruth(deduped, preDedup);
|
|
|
|
return deduped;
|
|
}
|
|
|
|
/**
|
|
* Layer 1: Keep top 3 chunks per page.
|
|
* Later layers (text similarity, cap per page) handle further reduction.
|
|
*/
|
|
function dedupBySource(results: SearchResult[]): SearchResult[] {
|
|
const byPage = new Map<string, SearchResult[]>();
|
|
|
|
for (const r of results) {
|
|
const existing = byPage.get(r.slug) || [];
|
|
existing.push(r);
|
|
byPage.set(r.slug, existing);
|
|
}
|
|
|
|
const kept: SearchResult[] = [];
|
|
for (const chunks of byPage.values()) {
|
|
chunks.sort((a, b) => b.score - a.score);
|
|
kept.push(...chunks.slice(0, 3));
|
|
}
|
|
|
|
return kept.sort((a, b) => b.score - a.score);
|
|
}
|
|
|
|
/**
|
|
* Layer 2: Remove chunks that are too similar to already-kept results.
|
|
* Uses Jaccard similarity on word sets as a proxy for cosine similarity.
|
|
*/
|
|
function dedupByTextSimilarity(results: SearchResult[], threshold: number): SearchResult[] {
|
|
const kept: SearchResult[] = [];
|
|
|
|
for (const r of results) {
|
|
const rWords = new Set(r.chunk_text.toLowerCase().split(/\s+/));
|
|
let tooSimilar = false;
|
|
|
|
for (const k of kept) {
|
|
const kWords = new Set(k.chunk_text.toLowerCase().split(/\s+/));
|
|
const intersection = new Set([...rWords].filter(w => kWords.has(w)));
|
|
const union = new Set([...rWords, ...kWords]);
|
|
const jaccard = intersection.size / union.size;
|
|
|
|
if (jaccard > threshold) {
|
|
tooSimilar = true;
|
|
break;
|
|
}
|
|
}
|
|
|
|
if (!tooSimilar) {
|
|
kept.push(r);
|
|
}
|
|
}
|
|
|
|
return kept;
|
|
}
|
|
|
|
/**
|
|
* Layer 3: No page type exceeds maxRatio of total results.
|
|
*/
|
|
function enforceTypeDiversity(results: SearchResult[], maxRatio: number): SearchResult[] {
|
|
const maxPerType = Math.max(1, Math.ceil(results.length * maxRatio));
|
|
const typeCounts = new Map<string, number>();
|
|
const kept: SearchResult[] = [];
|
|
|
|
for (const r of results) {
|
|
const count = typeCounts.get(r.type) || 0;
|
|
if (count < maxPerType) {
|
|
kept.push(r);
|
|
typeCounts.set(r.type, count + 1);
|
|
}
|
|
}
|
|
|
|
return kept;
|
|
}
|
|
|
|
/**
|
|
* Layer 4: Cap chunks per page.
|
|
*/
|
|
function capPerPage(results: SearchResult[], maxPerPage: number): SearchResult[] {
|
|
const pageCounts = new Map<string, number>();
|
|
const kept: SearchResult[] = [];
|
|
|
|
for (const r of results) {
|
|
const count = pageCounts.get(r.slug) || 0;
|
|
if (count < maxPerPage) {
|
|
kept.push(r);
|
|
pageCounts.set(r.slug, count + 1);
|
|
}
|
|
}
|
|
|
|
return kept;
|
|
}
|
|
|
|
/**
|
|
* Final pass: for each page in results that has no compiled_truth chunk,
|
|
* swap in the best compiled_truth chunk from the pre-dedup set (if one exists).
|
|
*/
|
|
function guaranteeCompiledTruth(results: SearchResult[], preDedup: SearchResult[]): SearchResult[] {
|
|
// Group results by page
|
|
const byPage = new Map<string, SearchResult[]>();
|
|
for (const r of results) {
|
|
const existing = byPage.get(r.slug) || [];
|
|
existing.push(r);
|
|
byPage.set(r.slug, existing);
|
|
}
|
|
|
|
const output = [...results];
|
|
|
|
for (const [slug, pageChunks] of byPage) {
|
|
const hasCompiledTruth = pageChunks.some(c => c.chunk_source === 'compiled_truth');
|
|
if (hasCompiledTruth) continue;
|
|
|
|
// Find the best compiled_truth chunk from pre-dedup input for this page
|
|
const candidate = preDedup
|
|
.filter(r => r.slug === slug && r.chunk_source === 'compiled_truth')
|
|
.sort((a, b) => b.score - a.score)[0];
|
|
|
|
if (!candidate) continue;
|
|
|
|
// Swap: replace the lowest-scored chunk from this page
|
|
const lowestIdx = output.reduce((minIdx, r, idx) => {
|
|
if (r.slug !== slug) return minIdx;
|
|
if (minIdx === -1) return idx;
|
|
return r.score < output[minIdx].score ? idx : minIdx;
|
|
}, -1);
|
|
|
|
if (lowestIdx !== -1) {
|
|
output[lowestIdx] = candidate;
|
|
}
|
|
}
|
|
|
|
return output;
|
|
}
|