Adversarial review: survived a hostile reviewer plus two independent refuters, each told to assume the PR was broken and to default to refuting when uncertain. `GBRAIN_FTS_LANGUAGE` was absent from the query-cache key, so a language switch served stale pre-switch rows. The hash now folds it (v14→15) with all five pin sites updated; reverting the fix fails 4 of 15 tests at the exact claimed step. Landing first in the knobs_hash cluster — the constant is single-writer, so #3617 rebases onto this and takes 16. Verified before merge: the PR's own tests fail when the production change is reverted (11 of the previous 32 PRs failed exactly there — one had 7 of 8 new tests passing on master); typecheck clean; MERGEABLE/CLEAN with 22/22 checks green on the current base, not a stale one.
3.8 KiB
Multi-language full-text search
GBrain's keyword search arm uses Postgres full-text search (tsvector/tsquery).
The tokenizer language is configurable via the GBRAIN_FTS_LANGUAGE
environment variable. Default: english.
How it works
Postgres text-search configurations control stemming and stop-word removal.
GBRAIN_FTS_LANGUAGE is read by src/core/fts-language.ts and applied on
both sides of the search:
- Query side —
websearch_to_tsquery('<lang>', $query)in both engines (Postgres and PGLite). - Write side — the
update_page_search_vectorandupdate_chunk_search_vectortrigger functions that populatepages.search_vectorandcontent_chunks.search_vector.
The value is validated against /^[a-z][a-z0-9_]*$/ before it is ever
interpolated into SQL (tsvector functions don't accept parameterized config
names). Invalid values fall back to english with a warning.
Built-in languages
Set the env var to any configuration your Postgres instance ships:
export GBRAIN_FTS_LANGUAGE=portuguese
export GBRAIN_FTS_LANGUAGE=spanish
export GBRAIN_FTS_LANGUAGE=german
List what's available:
SELECT cfgname FROM pg_ts_config;
PGLite (the embedded default engine) ships the same built-in snowball configurations as stock Postgres.
First install vs. changing language later
On first install (or upgrade), the configurable_fts_language schema
migration reads GBRAIN_FTS_LANGUAGE and stamps the trigger functions with
that language. After the migration has run, changing the env var alone does
NOT retokenize existing rows — the migration shows as applied and is skipped.
Use the explicit command:
export GBRAIN_FTS_LANGUAGE=portuguese
gbrain reindex-search-vector --dry-run # preview: language + row counts
gbrain reindex-search-vector --yes # recreate triggers + backfill
The command recreates both trigger functions under the new language and
backfills every existing pages and content_chunks row in batches,
streaming progress to stderr. It is idempotent: re-running with the same
language produces identical vectors. --json prints a machine-readable
result envelope but still requires --yes (or an interactive confirm).
No cache purge is needed. The resolved language is part of the query-cache
key, so rows written under the previous language are unreachable after the
switch — searches read the retokenized index immediately instead of being
served pre-switch results for up to search.cache.ttl_seconds. Switching
back reaches the original rows rather than rebuilding them.
Recipe: accent-insensitive Portuguese (pt_br)
Brazilian Portuguese content often mixes accented and unaccented spellings
("São Paulo" vs "Sao Paulo"). Build a custom config that folds accents via
the unaccent extension, then stems with the portuguese snowball dictionary:
CREATE EXTENSION IF NOT EXISTS unaccent;
CREATE TEXT SEARCH CONFIGURATION pt_br (COPY = portuguese);
ALTER TEXT SEARCH CONFIGURATION pt_br
ALTER MAPPING FOR hword, hword_part, word
WITH unaccent, portuguese_stem;
Then point GBrain at it:
export GBRAIN_FTS_LANGUAGE=pt_br
gbrain reindex-search-vector --yes
Note: custom configurations require a real Postgres instance (e.g. the
Supabase engine). The config must exist BEFORE the migration or the reindex
command runs, or Postgres will reject the trigger recreation with
text search configuration "pt_br" does not exist.
Caveats
- One language per brain: the setting is global to the database, not per-source. Mixed-language brains should pick the dominant language (the vector-search arm is language-agnostic and covers the rest).
- Keep
GBRAIN_FTS_LANGUAGEset consistently in every environment that writes to the brain (CLI shells, MCP server, cron jobs) — a writer without the env var tokenizes new rows inenglishuntil the next reindex.