diff --git a/docs/architecture/KEY_FILES.md b/docs/architecture/KEY_FILES.md index 893726577..3711004e4 100644 --- a/docs/architecture/KEY_FILES.md +++ b/docs/architecture/KEY_FILES.md @@ -190,7 +190,7 @@ per-release `**vX.Y.Z:**` narration — CI enforces this - `src/commands/enrich.ts` + `src/core/enrich/thin.ts` + `src/core/cycle/enrich-thin.ts` — `gbrain enrich --thin`: batch-develops stub (thin) pages via **brain-internal grounded synthesis**. gbrain's model tooling sees only brain-internal context (search / get_page / facts / backlinks), not the web, so enrich consolidates what the brain ALREADY knows about an entity (scattered across meetings, other pages, deals, facts) into one cited page via ONE `gateway.chat` call per page; web research stays the agent-driven `enrich` SKILL's job. `runEnrichCore(engine, opts, signal)` (strict per-source; multi-source iteration is the caller's job) drives `enrichOne` per candidate: `withRefreshingLock('enrich::')` → `getPage` → deterministic retrieve (hybridSearch + getBacklinks + facts + raw_data, source-scoped, sanitized via `INJECTION_PATTERNS`) → `assessGrounding` gate (skip < `MIN_CONTEXT_CHARS`, no LLM) → `buildEnrichPrompt` (grounded dossier, `[Source: slug]` citations, SKIP sentinel) → synth → `put_page` handler (`remote:false`, auto-link + write-through) stamping `enriched_at` + `enriched_by:'cli:enrich'`. Candidate selection is the SQL-native `engine.listEnrichCandidates(opts)` (`src/core/engine.ts` interface + `EnrichCandidate`/`EnrichCandidatesOpts`/`ENRICH_ORDER_SQL` in `src/core/types.ts` + pg/pglite impls): thin-filter + per-page source-correct inbound count (`to_page_id = p.id`, `mentions` excluded) + `enriched_at` recency guard + whitelisted ORDER BY + LIMIT, lightweight projection (NO bodies). Resume via `src/core/op-checkpoint.ts` (local `enrichFingerprint`); budget via `BudgetTracker` + `withBudgetTracker` (best-effort under `--workers > 1` — `runSlidingPool` aborts new claims on `BUDGET_EXHAUSTED` but does NOT cancel in-flight `gateway.chat`; pin `--workers 1` for a hard ceiling). `sanitizeContext` (thin.ts) neutralizes the `` data-envelope delimiters (injection escape, mirrors the `` convention); the `--background` multi-source fan-out idempotency key carries the run fingerprint via exported `backgroundIdempotencyKey(sid, args)` (a bare `enrich:${sid}` would return stale completed jobs); `runEnrichCore` flags `budget_exhausted` post-hoc when `tracker.totalSpent > tracker.cap` even when the gateway swallowed the final-call throw (via read-only `BudgetTracker.cap` getter); `body()` flushes the checkpoint on `BudgetExhausted` before it propagates so resume doesn't re-charge. The opt-in `enrich_thin` cycle phase (default OFF via `cycle.enrich_thin.enabled`) trickles `max_pages_per_tick` (default 3) per source with per-source cost cap enforced as `min(per_source_cap, brain_wide_remaining)` + brain-wide total + walltime caps. Wired into `cycle.ts` (`CyclePhase`/`ALL_PHASES` between `conversation_facts_backfill` and `skillopt`/`embed`; `PHASE_SCOPE='source'`; `NEEDS_LOCK`; dispatch), `cli.ts` (`CLI_ONLY` + `CLI_ONLY_SELF_HELP` + `THIN_CLIENT_REFUSED_COMMANDS` + dispatch), `jobs.ts` (Minion `enrich` handler, strict per-source, NOT in `PROTECTED_JOB_NAMES`). DI seam `opts.synthesizeFn` keeps tests hermetic (no API key, no mock.module). Pinned by `test/enrich/thin.test.ts`, `test/enrich/idempotency.test.ts`, `test/enrich-cycle-phase.test.ts`, `test/e2e/enrich-pglite.test.ts` (grew-cited, skip, ordering, multi-source, recency, resume, budget abort + checkpoint flush, final-call overage, lock-skip, provenance), `test/e2e/engine-parity.test.ts` (`listEnrichCandidates` pg↔pglite parity). - `src/core/data-research.ts` — Recipe validation, field extraction (MRR/ARR regex), dedup, tracker parsing, HTML stripping. - `src/commands/embed.ts` — `gbrain embed [--stale|--all] [--slugs ...]`. `--stale` starts with `engine.countStaleChunks()` (single SELECT count(*) WHERE embedding IS NULL, ~50 bytes wire) so a fully-embedded brain short-circuits with no further reads. When stale chunks exist, `engine.listStaleChunks()` returns just the chunks needing embeddings (slug + chunk_index + chunk_text + metadata, no `vector(1536)` payload); caller groups by slug, embeds, re-upserts via `upsertChunks`. All `console.log`/`console.error` call sites use `slog`/`serr` from `src/core/console-prefix.ts` so when `runEmbedCore` runs inside a per-source `withSourcePrefix` scope (installed by the `gbrain sync --all` worker pool) every line carries the `[] ` prefix; standalone callers see identical output because slog/serr fall through to bare console fns outside the wrap. Every embed-write path stamps `pages.embedding_signature` via `engine.setPageEmbeddingSignature(slug, {sourceId, signature: currentEmbeddingSignature()})` so a later model/dims swap is detectable as stale. The per-slug path (`embedPage`, used by `gbrain embed ` AND sync's post-import embed step) and the full-re-embed path (`embedAll`) stamp unconditionally per page. The stale path (`embedAllStale`) first calls `invalidateStaleSignatureEmbeddings` on a live run so signature-drifted pages flow through the NULL cursor, then stamps each page — but ONLY when EVERY chunk was stale this pass (a partially-stale page keeps preserved chunks of unknown provenance, so it stays unstamped rather than falsely marked current; `embed --all` fully re-embeds + stamps those). dry-run never mutates: it counts signature-drift via the widened `countStaleChunks({signature})` predicate without NULLing anything. -- `src/core/conversation-parser/` — 15-pattern built-in chat-format registry + opt-in LLM polish/fallback. Modules: `types.ts` (PatternEntry + ParseResult + DateContext + CaptureMap + TimezonePolicy), `builtins.ts` (15 hand-vetted patterns sourced from public format docs — iMessage/Slack, Telegram bracket + text-export, bold-paren-time, bold-name-no-time, Discord classic + export, WhatsApp ISO + US, Signal, Matrix/Element, IRC classic + weechat, Teams export; module-load validation runs every `test_positive[]` + `test_negative[]` sample at startup so a typo in any built-in regex makes gbrain refuse to start; `DEFAULT_SPEAKER_CLEAN` exported as a module-level default), `parse.ts` (orchestrator with pattern-priority scoring across the first 10 lines + date derivation chain `explicit > frontmatter.date > effective_date > '1970-01-01'` + multi-line continuation + timezone warning), `llm-base.ts` (shared `runLlmCall` with content-hash cache in-process + DB-persistent via migration v97 + 4-strategy JSON repair + Anthropic-key probe; polish and fallback are thin wrappers), `llm-polish.ts` (opt-IN; headroom guard skips when tracker within $0.10 of cap; pure `applyPolish` for merge/drop/edit ops), `llm-fallback.ts` (opt-IN; NO regex inference + NO persistence), `eval.ts` (`scoreFixture` + `aggregateScores` + `parseFixtureJsonl` for the fixture-corpus CI gate), `nightly-probe.ts` (DI-stubbed; mode-gated default tokenmax=ON, conservative/balanced opt-in; adversarial false-positive detection). Pattern `bold-name-no-time` (regex `/^\*\*(?!\[)(.+?):\*\*\s*(.*)$/`, index 3 after `bold-paren-time`) parses `**Speaker:** text` with NO per-line timestamp (Circleback/Granola/Zoom), anchoring every message at `T00:00:00Z` of the frontmatter date (line order preserves sequence, same no-time convention as `irc-classic`); the `(?!\[)` lookahead rejects telegram-bracket `**[18:37] Name:**`; non-shadow is the colon-INSIDE-bold regex (NOT declaration order — `parse.ts` scores every candidate independently, order is only the tie-break). Because `**Label:** text` is a common prose idiom, the pattern sets optional `PatternEntry.score_full_body: true` so `parse.ts` recomputes the winner's acceptance score over the FULL body before the `SCORING_MIN_ACCEPTANCE` floor, keeping a bold-label notes page at `no_match`. Pattern `bold-paren-time` parses `**Speaker** (HH:MM): text` and `(HH:MM:SS)` (date_source: frontmatter). Fallback gates: `SCORING_HEAD_TRIGGER_THRESHOLD = 0.3` triggers a full-body re-score when the head pass scores below that; `SCORING_MIN_ACCEPTANCE = 0.05` blocks essay false-positives. Exported `scorePatternFull(body, entry)`; private `getNonBlankLines(body, headCap?)` + `scoreFromLines(lines, entry)` DRY the quick_reject+regex loop. CLI surfaces at `src/commands/eval-conversation-parser.ts` (`gbrain eval conversation-parser ` exit 0/1/2, wired into `bun run verify` via `check:conversation-parser`) and `src/commands/conversation-parser.ts` (`scan ` debug, `list-builtins`, `validate `). Doctor checks: `conversation_format_coverage`, `progressive_batch_audit_health`, `conversation_parser_probe_health`. Pinned by `test/conversation-parser/{parse,llm-base,llm-fallback,llm-polish,nightly-probe}.test.ts` + the 27-case baseline at `test/extract-conversation-facts.test.ts` (back-compat invariant). Migration v97 (`conversation_parser_llm_cache_table`). Fixtures at `test/fixtures/conversation-formats/{imessage,imessage-time-only-12h,telegram-bracket,whatsapp-iso,whatsapp-us,signal-export,irc-classic,irc-weechat,matrix-element,teams-export,all,adversarial,bold-name-no-time}.jsonl` with `scripts/check-fixture-privacy.sh` banning real-name leaks. +- `src/core/conversation-parser/` — 17-pattern built-in chat-format registry + opt-in LLM polish/fallback. Modules: `types.ts` (PatternEntry + ParseResult + DateContext + CaptureMap + TimezonePolicy), `builtins.ts` (17 hand-vetted patterns sourced from public format docs — iMessage/Slack, Telegram bracket + text-export, bold-paren-time, bold-time-dash, bold-name-no-time, Discord classic + export, WhatsApp ISO + US, Signal, Matrix/Element, IRC classic + weechat, Teams export; module-load validation runs every `test_positive[]` + `test_negative[]` sample at startup so a typo in any built-in regex makes gbrain refuse to start; `DEFAULT_SPEAKER_CLEAN` exported as a module-level default), `parse.ts` (orchestrator with pattern-priority scoring across the first 10 lines + date derivation chain `explicit > frontmatter.date > effective_date > '1970-01-01'` + multi-line continuation + timezone warning), `llm-base.ts` (shared `runLlmCall` with content-hash cache in-process + DB-persistent via migration v97 + 4-strategy JSON repair + Anthropic-key probe; polish and fallback are thin wrappers), `llm-polish.ts` (opt-IN; headroom guard skips when tracker within $0.10 of cap; pure `applyPolish` for merge/drop/edit ops), `llm-fallback.ts` (opt-IN; NO regex inference + NO persistence), `eval.ts` (`scoreFixture` + `aggregateScores` + `parseFixtureJsonl` for the fixture-corpus CI gate), `nightly-probe.ts` (DI-stubbed; mode-gated default tokenmax=ON, conservative/balanced opt-in; adversarial false-positive detection). Pattern `bold-name-no-time` (regex `/^\*\*(?!\[)(.+?):\*\*\s*(.*)$/`, ordered after the time-bearing bold patterns) parses `**Speaker:** text` with NO per-line timestamp (Circleback/Granola/Zoom), anchoring every message at `T00:00:00Z` of the frontmatter date (line order preserves sequence, same no-time convention as `irc-classic`); the `(?!\[)` lookahead rejects telegram-bracket `**[18:37] Name:**`; non-shadow is the colon-INSIDE-bold regex (NOT declaration order — `parse.ts` scores every candidate independently, order is only the tie-break). Because `**Label:** text` is a common prose idiom, the pattern sets optional `PatternEntry.score_full_body: true` so `parse.ts` recomputes the winner's acceptance score over the FULL body before the `SCORING_MIN_ACCEPTANCE` floor, keeping a bold-label notes page at `no_match`. Pattern `bold-paren-time` parses `**Speaker** (HH:MM): text` and `(HH:MM:SS)` (date_source: frontmatter). Fallback gates: `SCORING_HEAD_TRIGGER_THRESHOLD = 0.3` triggers a full-body re-score when the head pass scores below that; `SCORING_MIN_ACCEPTANCE = 0.05` blocks essay false-positives. Exported `scorePatternFull(body, entry)`; private `getNonBlankLines(body, headCap?)` + `scoreFromLines(lines, entry)` DRY the quick_reject+regex loop. CLI surfaces at `src/commands/eval-conversation-parser.ts` (`gbrain eval conversation-parser ` exit 0/1/2, wired into `bun run verify` via `check:conversation-parser`) and `src/commands/conversation-parser.ts` (`scan ` debug, `list-builtins`, `validate `). Doctor checks: `conversation_format_coverage`, `progressive_batch_audit_health`, `conversation_parser_probe_health`. Pinned by `test/conversation-parser/{parse,llm-base,llm-fallback,llm-polish,nightly-probe}.test.ts` + the 27-case baseline at `test/extract-conversation-facts.test.ts` (back-compat invariant). Migration v97 (`conversation_parser_llm_cache_table`). Fixtures at `test/fixtures/conversation-formats/{imessage,imessage-time-only-12h,telegram-bracket,whatsapp-iso,whatsapp-us,signal-export,irc-classic,irc-weechat,matrix-element,teams-export,all,adversarial,bold-name-no-time,bold-time-dash}.jsonl` with `scripts/check-fixture-privacy.sh` banning real-name leaks. Maintainer guidance: [conversation parser patterns](conversation-parser-patterns.md). - `src/core/progressive-batch/` — shared ramp-up + cost-cap + verification primitive (trial 10 → ramp 100 → ramp 500 → full, with verification at each stage), with verifier+policy injection (callers describe HOW TO MEASURE SUCCESS, not WHEN TO WAIT FOR CTRL-C). Modules: `types.ts` (Stage, StageVerdict, AbortReason, discriminated `Verifier` union `OutputCountVerifier | IdempotentMutationVerifier | NoopVerifier`, Policy, StageReport), `orchestrator.ts` (`runProgressiveBatch(items, verifier, policy, runner)` — reads `getCurrentBudgetTracker()` ahead of `Policy.maxCostUsd` fail-closed; null both ways triggers `abort_cost_cap reason='no_budget_safety_net'`), `audit.ts` (ISO-week JSONL at `~/.gbrain/audit/progressive-batch-YYYY-Www.jsonl` via the shared `audit-writer` primitive), `stage-report.ts` (ASCII formatter for the default `Policy.onStageReport`). Env knobs: `GBRAIN_PROGRESSIVE_BATCH_DISABLED=1`, `GBRAIN_PROGRESSIVE_BATCH_AUTO=1` (skip Ctrl-C grace), `GBRAIN_PROGRESSIVE_BATCH_STAGES=10,100,500`. Sites that "jump straight to full" stay that way by default; ramp is opt-in per-site via `Policy.interactiveAbortMs > 0`. Pinned by `test/progressive-batch/orchestrator.test.ts` (35 cases, every verdict path). - `src/commands/extract-conversation-facts.ts` + `src/core/cycle/conversation-facts-backfill.ts` — bulk fact extraction for long-form conversation pages. Walks conversation/meeting/slack/email/imessage/imessage-daily pages, splits into time-windowed segments (30-min gap or 30-msg cap), prepends a topical/temporal header, and runs through `extractFactsFromTurn()` so anchor-rich facts surface in `gbrain search`. Invariants: **strict per-source core** (`runExtractConversationFactsCore({sourceId, ...})` always takes one sourceId; CLI + cycle phase each do their own multi-source iteration because PHASE_SCOPE='source' is taxonomy-only); **two-phase enumeration** (paginated `listPages({type, sourceId, limit:10})`; per-page body cap MAX_PAGE_BODY_BYTES=25MB with `pages_skipped_too_large` counter surfaced in doctor); **page-global row_num accumulator** (facts unique index is `(source_id, source_markdown_slug, row_num)` per migration v51 — per-segment row_num would collide); **page-level TERMINAL audit row** to facts table after all segments commit (source='cli:extract-conversation-facts:terminal'; doctor's NOT EXISTS matches the terminal row so partial-extraction pages stay in backlog); **optional `opts.budgetTracker?`** (when present, used as-is — nested `withBudgetTracker` REPLACES; when absent, core auto-wraps with `BudgetTracker({maxCostUsd})`); **body read covers compiled_truth + timeline**; **honors `facts.extraction_enabled` kill-switch** with `--override-disabled` escape; **--types LIST allowlist** (`conversation,meeting,slack,email,imessage,imessage-daily`) with CLI default reading `cycle.conversation_facts_backfill.types`; **fingerprint on sourceId only**; **string-encoded op-checkpoint** entries `"||"` for resume (durable audit is the facts terminal row); **`--background` via maybeBackground** (Minion handler `extract-conversation-facts` re-creates BudgetTracker from `data.max_cost_usd`; on `BudgetExhausted` mid-job catches + persists + marks `completed` with `result.budget_exhausted=true`). The companion cycle phase `conversation_facts_backfill` (default OFF) iterates `listSources(engine)`, creates ONE brain-wide tracker per tick + wraps the loop in `withBudgetTracker` + passes the tracker into every per-source call. Two-layer cost AND walltime caps: per-source (`max_cost_usd=$1`, `max_walltime_min=20`) AND brain-wide (`max_total_cost_usd=$5`, `max_total_walltime_min=30`). Pinned by `test/extract-conversation-facts.test.ts` (27 cases). Migration v94 adds partial index `idx_facts_extract_conversation_session ON facts(source_id, source_session) WHERE source LIKE 'cli:extract-conversation-facts%'` (`transaction:false` + invalid-index pre-drop on Postgres; plain CREATE INDEX on PGLite). `src/commands/doctor.ts:computeConversationFactsBacklogCheck` is 3-state (SKIPPED when disabled; OK when caught up; WARN when >10 pages lack the terminal row, with paste-ready `gbrain doctor --remediate` step). `src/commands/sources.ts:runAudit` adds `facts_backfill_estimate: {pages, est_segments, est_cost_usd, types}`. Schema-pack `gbrain-base.yaml` promotes `conversation` (temporal, extractable) + `atom` (annotation, NOT extractable) into the base seed; backstop uses hardcoded `ELIGIBLE_TYPES` in `src/core/facts/eligibility.ts:51` not pack extractable. `ALL_PAGE_TYPES` in `src/core/types.ts` extended with the two new types. - `src/core/link-extraction.ts` — shared library for the graph layer. `extractEntityRefs` (canonical) matches `[Name](people/slug)` markdown links and Obsidian `[[people/slug|Name]]` wikilinks; `extractPageLinks`, `inferLinkType` heuristics (attended/works_at/invested_in/founded/advises/source/mentions), `parseTimelineEntries`, `isAutoLinkEnabled`. `DIR_PATTERN` covers `people`, `companies`, `deals`, `topics`, `concepts`, `projects`, `entities`, `tech`, `finance`, `personal`, `openclaw`. Used by extract.ts, operations.ts auto-link post-hook, and backlinks.ts. **Opt-in global-basename wikilink resolution** (issue #972, default off): `WIKILINK_GENERIC_RE` catches bare `[[name]]` wikilinks outside `DIR_PATTERN` (third pass `2c` in `extractEntityRefs`); `EntityRef.needsResolution: true` tags refs from this pass (the ref's `slug` is the wikilink TARGET, `name` the optional display alias). `SlugResolver` gains optional `resolveBasenameMatches(name): Promise` (multi-match by design — emits one edge per matching page). The single shared basename matcher is `buildBasenameIndex(slugs)` + `queryBasenameIndex(index, name)` + `normalizeBasename` (keys raw/lower/slugified tail, stable-sorted shorter-first then lexical), used by `makeResolver`, the FS `resolveBasenameMatchesFromSlugs`, AND the doctor check so they cannot drift. `makeResolver(engine, {mode, sourceId})` builds the index lazily via `engine.getAllSlugs({sourceId})` — source-scoped so a bare `[[name]]` never resolves to a same-tail page in a different source. `extractPageLinks` gains `opts.globalBasename` (routes `needsResolution` refs through `resolveBasenameMatches` keyed on `ref.slug`, emits candidates tagged `linkType: 'wikilink_basename'` + `linkSource: 'wikilink-resolved'`, skips self-loops) and `opts.skipFrontmatter` (replaces the old `nullResolver` ternary). All three surfaces (FS extract, DB extract, `put_page` auto-link) tag provenance with `link_source='wikilink-resolved'`; `put_page` includes it in its reconcilable-edge set so stale basename edges are removed when the wikilink or the flag goes away. Exports `WIKILINK_BASENAME_LINK_TYPE` + `isGlobalBasenameEnabled(engine)` (resolution order: env `GBRAIN_LINK_RESOLUTION_GLOBAL_BASENAME` → DB config `link_resolution.global_basename` → default false). `gbrain doctor`'s `link_resolution_opportunity` check surfaces a paste-ready enable hint when ≥5 bare wikilinks would resolve AND ≥20% match. Migration v113 widens `links_link_source_check` to admit `'wikilink-resolved'`; v114 (#1941) then opens it to any kebab-case provenance (`^[a-z][a-z0-9]*(-[a-z0-9]+)*$`, ≤64 chars) so external derivers register their own tag (e.g. `citation-graph`) without a migration. `LINK_EXTRACTOR_VERSION_TS` also lives here (bump like `CHUNKER_VERSION` to invalidate prior extract-stale stamps). Pinned by `test/link-extraction.test.ts`, `test/extract-fs.test.ts`, `test/doctor.test.ts`, `test/e2e/global-basename-pglite.test.ts`. diff --git a/docs/architecture/conversation-parser-patterns.md b/docs/architecture/conversation-parser-patterns.md new file mode 100644 index 000000000..005372d3d --- /dev/null +++ b/docs/architecture/conversation-parser-patterns.md @@ -0,0 +1,146 @@ +# Conversation parser patterns + +The conversation parser turns exported chat and meeting transcripts into a +common message stream without requiring an LLM call for known formats. This +document describes the built-in pattern contract and the checks required when +adding or changing a format. + +## Data flow + +`parseConversation` uses this sequence: + +1. Resolve the page date and timezone context. +2. Score every enabled built-in and user pattern against the first ten + non-blank lines. +3. Re-score the full body when the head score is inconclusive, or when a broad + pattern explicitly requires full-body scoring. +4. Reject the winner when its acceptance score is below the false-positive + floor. +5. Apply the winning pattern to every line and attach continuation lines to the + preceding message. +6. Optionally run LLM polish or fallback when those features are enabled. + +Pattern order is only a tie-breaker. A new regex must be structurally distinct +from neighboring formats; moving it earlier in the registry is not a valid +non-shadowing strategy. + +## Built-in pattern contract + +Every `PatternEntry` in `builtins.ts` declares: + +- A stable, kebab-case `id`. +- A hand-vetted line regex and explicit capture-group indexes. +- Where the date comes from and how the time is represented. +- A timezone policy. +- Whether the format supports multi-line message bodies. +- Positive and negative samples that run during module initialization. +- A documentation pointer describing the source format. + +The registry refuses to load when a positive sample stops matching, a negative +sample starts matching, or a capture map becomes invalid. This catches local +regex mistakes before extraction can silently produce empty conversations. + +### Date and timezone rules + +Formats with an inline date should capture it from each message. Time-only +formats use an explicit caller fallback first, then the page frontmatter date, +then the page effective date. If none is available, the parser uses +`1970-01-01` so the missing date remains visible instead of inventing a current +date. + +Time-only formats normally use `utc_assumed_with_warn`. The parser constructs a +UTC timestamp and returns a timezone warning when the page does not provide a +timezone. A new pattern should not imply local-time precision that the source +format does not contain. + +### Multi-line messages + +An anchor regex identifies the first line of a message. Subsequent non-anchor +lines are appended to that message until another anchor appears. Set +`multi_line: true` when continuation content is part of the documented format, +such as Markdown bullets, blockquotes, or an exported message body on the next +line. + +Tests for a multi-line format should assert the complete message text, including +newlines. A message-count assertion alone will not detect lost bullets or a +continuation attached to the wrong speaker. + +### Scoring and false positives + +The score compares matched anchors with the pattern's relevant candidate lines. +The first pass uses the head of the page for speed. Low-confidence pages are +re-scored across the full body before the parser accepts a winner. + +Multi-line formats may opt into `score_continuations_as_body` when their anchor +grammar is distinctive. Candidate-only scoring activates only after two anchors +match, or when the first non-blank line is an anchor. This evidence threshold +lets a single long message keep its continuation body without turning one stray +anchor in a prose page into a conversation. Candidate anchor lines that fail the +full regex still lower the score. Other patterns continue to use all non-blank +lines in their density score. + +Use `score_full_body: true` for a broad grammar that also occurs in ordinary +prose. For example, `**Label:** text` can be either a transcript line or a bold +label in meeting notes. Narrow formats with a timestamp and a distinctive +separator generally do not need this override. + +`quick_reject` is a performance hint, not an acceptance rule. It should cheaply +exclude obviously unrelated lines while admitting every string accepted by the +main regex. + +## Normalized Slack Markdown + +The `bold-time-dash` pattern parses message anchors shaped like: + +```text +**Alice Example** 09:15 — first message +- supporting detail +**Bob Example** 09:18 — second message +``` + +Its grammar is: + +```text +**speaker** H:MM text +``` + +where: + +- `H:MM` is a valid 24-hour time from `0:00` through `23:59`. +- `` may be an em dash (`—`), en dash (`–`), or ASCII hyphen (`-`). +- The date comes from the resolved page date context. +- Continuation lines belong to the preceding message. +- The captured clock value is emitted with `Z`. Timezone metadata suppresses + the missing-timezone warning but is not currently used for IANA conversion. + +The required time and dash distinguish it from all existing bold-speaker +formats: + +- `**Speaker** (09:15): text` uses `bold-paren-time`. +- `**Speaker** (9:15 AM): text` uses `bold-paren-time-12h`. +- `**Speaker:** text` uses `bold-name-no-time`. +- `**Speaker** (2026-04-09 9:15 AM): text` uses `imessage-slack`. + +Keeping these examples in both `test_negative` and parser regression tests makes +the non-shadowing contract executable. + +## Adding a built-in format + +1. Collect multiple anonymized examples, including separator and timestamp + variants that occur in the same export family. +2. Choose the narrowest grammar that represents the format. Constrain numeric + fields such as hours and minutes when possible. +3. Add at least two positive module-load samples and negative samples for every + neighboring pattern that could plausibly overlap. +4. Add parser tests that verify speakers, timestamps, text, continuation + handling, and non-shadowing behavior. +5. Add a dedicated JSONL fixture and include the same cases in + `test/fixtures/conversation-formats/all.jsonl`. +6. Run the focused parser tests and the fixture evaluator. +7. Run the repository verification and full test suites before submission. +8. Update `docs/architecture/KEY_FILES.md` when the registry count or supported + format inventory changes. + +Use generic fixture identities such as `Alice Example`, `Bob Example`, and +`Summary Bot`. Never copy real transcript names or private content into source, +tests, documentation, commits, or pull-request descriptions. diff --git a/src/core/conversation-parser/builtins.ts b/src/core/conversation-parser/builtins.ts index 4928cbe26..969edb036 100644 --- a/src/core/conversation-parser/builtins.ts +++ b/src/core/conversation-parser/builtins.ts @@ -1,7 +1,7 @@ /** * v0.41.16.0 — Built-in conversation parser pattern registry. * - * Fifteen hand-vetted patterns covering the chat-export formats this + * Seventeen hand-vetted patterns covering the chat-export formats this * codebase is most likely to encounter. Each pattern's regex was * derived from a public format reference (source_doc field) so future * maintainers can verify against the wild shape. @@ -50,7 +50,7 @@ export function cleanSpeaker(raw: string, override?: RegExp): string { return stripped || raw.trim(); } -/** The 15 hand-vetted built-in patterns. */ +/** The 17 hand-vetted built-in patterns. */ export const BUILTIN_PATTERNS: readonly PatternEntry[] = [ // ------------------------------------------------------------------- // INLINE-DATE patterns (date in every line; less ambiguous; tried first). @@ -213,6 +213,67 @@ export const BUILTIN_PATTERNS: readonly PatternEntry[] = [ 'Time-only 12h AM/PM iMessage export shape: `**Speaker** (H:MM AM): text`', }, + { + // Some Slack-to-Markdown normalizers render one message anchor as: + // + // **Speaker Name** 09:15 — message text + // + // The date lives in page frontmatter while each line supplies a 24-hour + // wall-clock time. The separator varies by renderer: Unicode em dash, + // Unicode en dash, and ASCII hyphen all appear in otherwise identical + // exports. Treating all three as the same deterministic grammar avoids + // sending long, regular transcripts through the bounded LLM fallback. + // + // CONTINUATION SEMANTICS: normalized messages can contain Markdown lists, + // quoted blocks, or generated summaries below the anchor line. multi_line + // is therefore true; applyPattern appends every non-anchor line to the + // preceding message until the next matching anchor. + // + // DATE/TIME SEMANTICS: date_source='frontmatter' combines the resolved page + // date with the captured hour and minute. timezone_policy intentionally + // matches the other time-only Markdown formats: the captured clock value + // is emitted with `Z`; timezone metadata controls the warning but does not + // currently convert the wall-clock value. + // + // NON-SHADOW GUARANTEE: this grammar requires the closing bold marker, + // whitespace, a valid 24-hour time, and a dash. It cannot match the + // parenthesized bold formats (`**Name** (09:15): text`), the no-time bold + // format (`**Name:** text`), or the inline-date iMessage format. Parser + // declaration order is only a score tie-breaker, so these distinctions + // must remain structural in the regex. + id: 'bold-time-dash', + origin: 'builtin', + regex: + /^\*\*(.+?)\*\*\s+([01]?\d|2[0-3]):([0-5]\d)\s+[-\u2013\u2014]\s*(.*)$/, + captures: { + speaker_group: 1, + hour_group: 2, + minute_group: 3, + text_group: 4, + }, + date_source: 'frontmatter', + time_format: '24h', + timezone_policy: 'utc_assumed_with_warn', + multi_line: true, + score_continuations_as_body: true, + quick_reject: /^\*\*/, + test_positive: [ + '**Alice Example** 09:15 — hello world', + '**Summary Bot** 23:04 – nightly summary follows', + '**Bob Example** 7:05 - ASCII dash export', + ], + test_negative: [ + '**Alice Example** (09:15): parenthesized meeting shape', + '**Alice Example** (9:15 AM): parenthesized 12-hour shape', + '**Alice Example:** no-time transcript shape', + '**Alice Example** (2024-03-15 9:00 AM): inline-date shape', + '**Alice Example** 24:00 — invalid 24-hour time', + '**Alice Example** 09:60 — invalid minute', + ], + source_doc: + 'Normalized Slack Markdown: `**Speaker** HH:MM — text`, with the date in page frontmatter', + }, + { // Fathom/phone-call raw transcripts in this workspace use a plain // `Speaker A: ...` / `Speaker B: ...` shape with no per-line time. @@ -647,17 +708,28 @@ export function validatePatternEntry(entry: PatternEntry): void { if (entry.test_positive.length > 0) { const m = entry.regex.exec(entry.test_positive[0]); if (m === null) return; // already thrown above - const requiredGroups = [ - entry.captures.speaker_group, - entry.captures.date_group, - entry.captures.hour_group, - entry.captures.minute_group, - entry.captures.ampm_group, - ].filter((g): g is number => typeof g === 'number'); - for (const g of requiredGroups) { - if (g >= m.length) { + const captureGroups: Array<[ + name: string, + group: number | undefined, + minimum: number, + ]> = [ + ['speaker_group', entry.captures.speaker_group, 1], + ['text_group', entry.captures.text_group, 0], + ['date_group', entry.captures.date_group, 1], + ['hour_group', entry.captures.hour_group, 1], + ['minute_group', entry.captures.minute_group, 1], + ['ampm_group', entry.captures.ampm_group, 1], + ]; + for (const [name, group, minimum] of captureGroups) { + if (group === undefined) continue; + if (!Number.isInteger(group) || group < minimum) { throw new Error( - `[conversation-parser] PatternEntry '${entry.id}' captures group ${g} but regex only emits ${m.length - 1} groups`, + `[conversation-parser] PatternEntry '${entry.id}' ${name} must be an integer >= ${minimum}; got ${group}`, + ); + } + if (group > 0 && group >= m.length) { + throw new Error( + `[conversation-parser] PatternEntry '${entry.id}' captures group ${group} but regex only emits ${m.length - 1} groups`, ); } } diff --git a/src/core/conversation-parser/parse.ts b/src/core/conversation-parser/parse.ts index 98689bdb1..66aa76745 100644 --- a/src/core/conversation-parser/parse.ts +++ b/src/core/conversation-parser/parse.ts @@ -391,7 +391,7 @@ function getNonBlankLines(body: string, headCap?: number): string[] { * window) and `scorePatternFull` (whole body) delegate here so the * quick_reject + regex loop lives in one place. Reused by * `parseConversation`'s fallback path which pre-splits ONCE and - * passes the array to all 15 candidates (saves 14 redundant body + * passes the array to all 17 candidates (saves 16 redundant body * splits per fallback pass). */ function scoreFromLines( @@ -400,9 +400,28 @@ function scoreFromLines( ): number { if (lines.length === 0) return 0; let anchored = 0; - for (const line of lines) { - if (entry.quick_reject && !entry.quick_reject.test(line)) continue; - if (entry.regex.test(line)) anchored++; + let anchorCandidates = 0; + let firstLineAnchored = false; + for (let index = 0; index < lines.length; index++) { + const line = lines[index]; + if (entry.quick_reject && !entry.quick_reject.test(line)) { + continue; + } + anchorCandidates++; + if (entry.regex.test(line)) { + anchored++; + if (index === 0) firstLineAnchored = true; + } + } + + if ( + entry.score_continuations_as_body && + entry.multi_line && + entry.quick_reject && + anchorCandidates > 0 && + (anchored >= 2 || firstLineAnchored) + ) { + return anchored / anchorCandidates; } return anchored / lines.length; } @@ -411,8 +430,10 @@ function scoreFromLines( * Score how well a pattern matches the first N lines of a body (D18). * Returns 0..1 ratio of matched lines. Higher = more confident. * - * Quick_reject is honored (lines that don't pass quick_reject still - * count as "could be continuation"; not penalized). + * Quick_reject is honored. Patterns that opt into + * `score_continuations_as_body` may exclude continuation lines from the + * denominator only after the scorer sees two anchors, or an anchor on the + * first non-blank line. Otherwise the ordinary full-body density applies. * * Exported for tests. */ diff --git a/src/core/conversation-parser/types.ts b/src/core/conversation-parser/types.ts index f439f410a..417ffaa3e 100644 --- a/src/core/conversation-parser/types.ts +++ b/src/core/conversation-parser/types.ts @@ -159,6 +159,14 @@ export interface PatternEntry { * message; continuation logic still applies for orphan lines. */ multi_line: boolean; + /** + * When true, scoring may treat lines that fail `quick_reject` as message + * continuation rather than independent evidence. To preserve the global + * false-positive floor, the candidate-only score is used only after two + * anchors match, or when the first non-blank line is itself an anchor. + * Requires `multi_line: true` and a `quick_reject`. + */ + score_continuations_as_body?: boolean; /** * D11: optional cheap O(1) prefix check. If set, orchestrator runs * this FIRST per line; only tries `regex` if quick_reject matches. diff --git a/test/conversation-parser/parse.test.ts b/test/conversation-parser/parse.test.ts index a43e96692..c9cc02648 100644 --- a/test/conversation-parser/parse.test.ts +++ b/test/conversation-parser/parse.test.ts @@ -24,7 +24,10 @@ import { scorePattern, scorePatternFull, } from '../../src/core/conversation-parser/parse.ts'; -import { BUILTIN_PATTERNS } from '../../src/core/conversation-parser/builtins.ts'; +import { + BUILTIN_PATTERNS, + validatePatternEntry, +} from '../../src/core/conversation-parser/builtins.ts'; import type { Page } from '../../src/core/types.ts'; // Helper to construct a minimal Page for date-derivation tests. @@ -139,6 +142,35 @@ describe('parseConversation — every built-in matches its test_positive sample' } }); +test('validatePatternEntry rejects invalid capture indexes', () => { + const base = BUILTIN_PATTERNS[0]; + const aboveRange = { + ...base, + id: 'invalid-text-capture', + captures: { ...base.captures, text_group: 99 }, + }; + const zeroSpeaker = { + ...base, + id: 'invalid-speaker-capture', + captures: { ...base.captures, speaker_group: 0 }, + }; + const negativeText = { + ...base, + id: 'negative-text-capture', + captures: { ...base.captures, text_group: -1 }, + }; + + expect(() => validatePatternEntry(aboveRange)).toThrow( + "captures group 99 but regex only emits", + ); + expect(() => validatePatternEntry(zeroSpeaker)).toThrow( + 'speaker_group must be an integer >= 1', + ); + expect(() => validatePatternEntry(negativeText)).toThrow( + 'text_group must be an integer >= 0', + ); +}); + // --------------------------------------------------------------------------- // Date derivation precedence (D8) // --------------------------------------------------------------------------- @@ -512,6 +544,152 @@ describe('bold-paren-time pattern (Circleback meeting transcripts)', () => { }); }); +// --------------------------------------------------------------------------- +// bold-time-dash pattern (normalized Slack Markdown) +// --------------------------------------------------------------------------- + +describe('bold-time-dash pattern (normalized Slack Markdown)', () => { + test('parses anchors, dash variants, and multi-line continuation text', () => { + const body = [ + '# Team channel — 2026-04-09', + '**Alice Example** 09:15 — first line', + '- detailed bullet one', + '- detailed bullet two', + '**Summary Bot** 09:18 – second message', + '> continuation of second message', + '**Bob Example** 10:01 - final message', + ].join('\n'); + const r = parseConversation(body, { fallbackDate: '2026-04-09' }); + + expect(r.phase).toBe('regex_match'); + expect(r.matched_pattern_id).toBe('bold-time-dash'); + expect(r.messages).toHaveLength(3); + expect(r.messages[0]).toEqual({ + speaker: 'Alice Example', + timestamp: '2026-04-09T09:15:00Z', + text: 'first line\n- detailed bullet one\n- detailed bullet two', + }); + expect(r.messages[1]).toEqual({ + speaker: 'Summary Bot', + timestamp: '2026-04-09T09:18:00Z', + text: 'second message\n> continuation of second message', + }); + expect(r.messages[2]).toEqual({ + speaker: 'Bob Example', + timestamp: '2026-04-09T10:01:00Z', + text: 'final message', + }); + }); + + test('parses one anchor with a long Markdown continuation body', () => { + const continuation = Array.from( + { length: 30 }, + (_, index) => `- supporting detail ${index + 1}`, + ); + const body = [ + '**Alice Example** 09:15 — summary', + ...continuation, + ].join('\n'); + const r = parseConversation(body, { fallbackDate: '2026-04-09' }); + + expect(r.matched_pattern_id).toBe('bold-time-dash'); + expect(r.messages).toHaveLength(1); + expect(r.messages[0].text.split('\n')).toHaveLength(31); + expect(r.messages[0].text.endsWith('- supporting detail 30')).toBe(true); + }); + + test('does not treat one stray anchor in long prose as a conversation', () => { + const before = Array.from( + { length: 150 }, + (_, index) => `Prose paragraph before ${index + 1}.`, + ); + const after = Array.from( + { length: 150 }, + (_, index) => `Prose paragraph after ${index + 1}.`, + ); + const body = [ + ...before, + '**Deadline** 09:15 — quoted schedule entry', + ...after, + ].join('\n'); + const r = parseConversation(body, { fallbackDate: '2026-04-09' }); + + expect(r.phase).toBe('no_match'); + expect(r.messages).toEqual([]); + }); + + test('uses date headings to advance the frontmatter date anchor', () => { + const body = [ + '## 2026-04-09', + '**Alice Example** 23:59 — day one', + '## 2026-04-10', + '**Bob Example** 00:01 — day two', + ].join('\n'); + const r = parseConversation(body, { fallbackDate: '2026-04-09' }); + + expect(r.matched_pattern_id).toBe('bold-time-dash'); + expect(r.messages.map((message) => message.timestamp)).toEqual([ + '2026-04-09T23:59:00Z', + '2026-04-10T00:01:00Z', + ]); + }); + + test('uses page date and preserves the time-only timezone policy', () => { + const body = '**Alice Example** 09:15 — hello'; + const withoutTimezone = parseConversation(body, { + page: makePage({ date: '2026-04-09' }), + }); + const withTimezone = parseConversation(body, { + page: makePage({ + date: '2026-04-09', + timezone: 'America/Los_Angeles', + }), + }); + + expect(withoutTimezone.messages[0].timestamp).toBe( + '2026-04-09T09:15:00Z', + ); + expect(withoutTimezone.timezone_warning).toContain('bold-time-dash'); + // Current time-only policy records the captured wall-clock fields with Z; + // timezone metadata suppresses the warning but does not convert the time. + expect(withTimezone.messages[0].timestamp).toBe('2026-04-09T09:15:00Z'); + expect(withTimezone.timezone_warning).toBeUndefined(); + }); + + test('does not shadow existing bold transcript formats', () => { + const opts = { fallbackDate: '2026-04-09' }; + + expect( + parseConversation('**Alice Example** (00:00): hello', opts) + .matched_pattern_id, + ).toBe('bold-paren-time'); + expect( + parseConversation('**Alice Example** (9:15 AM): hello', opts) + .matched_pattern_id, + ).toBe('bold-paren-time-12h'); + expect( + parseConversation('**Alice Example:** hello', opts).matched_pattern_id, + ).toBe('bold-name-no-time'); + expect( + parseConversation( + '**Alice Example** (2026-04-09 9:15 AM): hello', + opts, + ).matched_pattern_id, + ).toBe('imessage-slack'); + }); + + test('rejects invalid 24-hour times', () => { + const body = [ + '**Alice Example** 24:00 — invalid hour', + '**Bob Example** 09:60 — invalid minute', + ].join('\n'); + const r = parseConversation(body, { fallbackDate: '2026-04-09' }); + + expect(r.phase).toBe('no_match'); + expect(r.messages).toEqual([]); + }); +}); + // --------------------------------------------------------------------------- // bold-name-no-time pattern (Circleback / Granola / Zoom transcripts with NO // per-line timestamp — `**Speaker:** text`). Additive pattern; the colon diff --git a/test/fixtures/conversation-formats/adversarial.jsonl b/test/fixtures/conversation-formats/adversarial.jsonl index 2be6cc67a..19d81666c 100644 --- a/test/fixtures/conversation-formats/adversarial.jsonl +++ b/test/fixtures/conversation-formats/adversarial.jsonl @@ -4,3 +4,4 @@ {"fixture_id":"adversarial-lyrics","pattern":null,"frontmatter":{"date":"2024-03-15"},"body":"Twinkle twinkle little star\nHow I wonder what you are\nUp above the world so high\nLike a diamond in the sky","expected_messages":0,"expected_participants":[]} {"fixture_id":"adversarial-json","pattern":null,"frontmatter":{"date":"2024-03-15"},"body":"{\n \"key\": \"value\",\n \"num\": 42,\n \"list\": [1, 2, 3]\n}","expected_messages":0,"expected_participants":[]} {"fixture_id":"adversarial-bold-labels-clustered-head","pattern":null,"frontmatter":{"date":"2026-05-28"},"body":"**Attendees:** Alice Example, Bob Example, Participant 2\n**Date:** 2026-05-28\n**Goal:** decide on the Q3 roadmap and unblock the vendor migration\nThis is an ordinary prose sentence number 0 describing the meeting in detail.\nThis is an ordinary prose sentence number 1 describing the meeting in detail.\nThis is an ordinary prose sentence number 2 describing the meeting in detail.\nThis is an ordinary prose sentence number 3 describing the meeting in detail.\nThis is an ordinary prose sentence number 4 describing the meeting in detail.\nThis is an ordinary prose sentence number 5 describing the meeting in detail.\nThis is an ordinary prose sentence number 6 describing the meeting in detail.\nThis is an ordinary prose sentence number 7 describing the meeting in detail.\nThis is an ordinary prose sentence number 8 describing the meeting in detail.\nThis is an ordinary prose sentence number 9 describing the meeting in detail.\nThis is an ordinary prose sentence number 10 describing the meeting in detail.\nThis is an ordinary prose sentence number 11 describing the meeting in detail.\nThis is an ordinary prose sentence number 12 describing the meeting in detail.\nThis is an ordinary prose sentence number 13 describing the meeting in detail.\nThis is an ordinary prose sentence number 14 describing the meeting in detail.\nThis is an ordinary prose sentence number 15 describing the meeting in detail.\nThis is an ordinary prose sentence number 16 describing the meeting in detail.\nThis is an ordinary prose sentence number 17 describing the meeting in detail.\nThis is an ordinary prose sentence number 18 describing the meeting in detail.\nThis is an ordinary prose sentence number 19 describing the meeting in detail.\nThis is an ordinary prose sentence number 20 describing the meeting in detail.\nThis is an ordinary prose sentence number 21 describing the meeting in detail.\nThis is an ordinary prose sentence number 22 describing the meeting in detail.\nThis is an ordinary prose sentence number 23 describing the meeting in detail.\nThis is an ordinary prose sentence number 24 describing the meeting in detail.\nThis is an ordinary prose sentence number 25 describing the meeting in detail.\nThis is an ordinary prose sentence number 26 describing the meeting in detail.\nThis is an ordinary prose sentence number 27 describing the meeting in detail.\nThis is an ordinary prose sentence number 28 describing the meeting in detail.\nThis is an ordinary prose sentence number 29 describing the meeting in detail.\nThis is an ordinary prose sentence number 30 describing the meeting in detail.\nThis is an ordinary prose sentence number 31 describing the meeting in detail.\nThis is an ordinary prose sentence number 32 describing the meeting in detail.\nThis is an ordinary prose sentence number 33 describing the meeting in detail.\nThis is an ordinary prose sentence number 34 describing the meeting in detail.\nThis is an ordinary prose sentence number 35 describing the meeting in detail.\nThis is an ordinary prose sentence number 36 describing the meeting in detail.\nThis is an ordinary prose sentence number 37 describing the meeting in detail.\nThis is an ordinary prose sentence number 38 describing the meeting in detail.\nThis is an ordinary prose sentence number 39 describing the meeting in detail.\nThis is an ordinary prose sentence number 40 describing the meeting in detail.\nThis is an ordinary prose sentence number 41 describing the meeting in detail.\nThis is an ordinary prose sentence number 42 describing the meeting in detail.\nThis is an ordinary prose sentence number 43 describing the meeting in detail.\nThis is an ordinary prose sentence number 44 describing the meeting in detail.\nThis is an ordinary prose sentence number 45 describing the meeting in detail.\nThis is an ordinary prose sentence number 46 describing the meeting in detail.\nThis is an ordinary prose sentence number 47 describing the meeting in detail.\nThis is an ordinary prose sentence number 48 describing the meeting in detail.\nThis is an ordinary prose sentence number 49 describing the meeting in detail.\nThis is an ordinary prose sentence number 50 describing the meeting in detail.\nThis is an ordinary prose sentence number 51 describing the meeting in detail.\nThis is an ordinary prose sentence number 52 describing the meeting in detail.\nThis is an ordinary prose sentence number 53 describing the meeting in detail.\nThis is an ordinary prose sentence number 54 describing the meeting in detail.\nThis is an ordinary prose sentence number 55 describing the meeting in detail.\nThis is an ordinary prose sentence number 56 describing the meeting in detail.\nThis is an ordinary prose sentence number 57 describing the meeting in detail.\nThis is an ordinary prose sentence number 58 describing the meeting in detail.\nThis is an ordinary prose sentence number 59 describing the meeting in detail.\nThis is an ordinary prose sentence number 60 describing the meeting in detail.\nThis is an ordinary prose sentence number 61 describing the meeting in detail.\nThis is an ordinary prose sentence number 62 describing the meeting in detail.\nThis is an ordinary prose sentence number 63 describing the meeting in detail.\nThis is an ordinary prose sentence number 64 describing the meeting in detail.\nThis is an ordinary prose sentence number 65 describing the meeting in detail.\nThis is an ordinary prose sentence number 66 describing the meeting in detail.\nThis is an ordinary prose sentence number 67 describing the meeting in detail.\nThis is an ordinary prose sentence number 68 describing the meeting in detail.\nThis is an ordinary prose sentence number 69 describing the meeting in detail.\nThis is an ordinary prose sentence number 70 describing the meeting in detail.\nThis is an ordinary prose sentence number 71 describing the meeting in detail.\nThis is an ordinary prose sentence number 72 describing the meeting in detail.\nThis is an ordinary prose sentence number 73 describing the meeting in detail.\nThis is an ordinary prose sentence number 74 describing the meeting in detail.\nThis is an ordinary prose sentence number 75 describing the meeting in detail.\nThis is an ordinary prose sentence number 76 describing the meeting in detail.\nThis is an ordinary prose sentence number 77 describing the meeting in detail.\nThis is an ordinary prose sentence number 78 describing the meeting in detail.\nThis is an ordinary prose sentence number 79 describing the meeting in detail.","expected_messages":0,"expected_participants":[]} +{"fixture_id":"adversarial-bold-time-dash-stray","pattern":null,"frontmatter":{"date":"2026-04-09"},"body":"Context paragraph before 1.\nContext paragraph before 2.\nContext paragraph before 3.\nContext paragraph before 4.\nContext paragraph before 5.\nContext paragraph before 6.\nContext paragraph before 7.\nContext paragraph before 8.\nContext paragraph before 9.\nContext paragraph before 10.\n**Deadline** 09:15 — quoted schedule entry\nContext paragraph after 1.\nContext paragraph after 2.\nContext paragraph after 3.\nContext paragraph after 4.\nContext paragraph after 5.\nContext paragraph after 6.\nContext paragraph after 7.\nContext paragraph after 8.\nContext paragraph after 9.\nContext paragraph after 10.","expected_messages":0,"expected_participants":[]} diff --git a/test/fixtures/conversation-formats/all.jsonl b/test/fixtures/conversation-formats/all.jsonl index 5f19b8024..68143b30f 100644 --- a/test/fixtures/conversation-formats/all.jsonl +++ b/test/fixtures/conversation-formats/all.jsonl @@ -14,4 +14,7 @@ {"fixture_id":"teams-export-001","pattern":"teams-export","frontmatter":{"date":"2024-03-15"},"body":"Alice Example, 3/15/2024 6:37 PM: hello team\nBob Example, 3/15/2024 6:38 PM: hey\nAlice Example, 3/15/2024 6:39 PM: meeting at 4?\nBob Example, 3/15/2024 6:40 PM: works for me","expected_messages":4,"expected_participants":["Alice Example","Bob Example"]} {"fixture_id":"bold-name-no-time-001","pattern":"bold-name-no-time","frontmatter":{"date":"2026-05-28"},"body":"**Alice Example:** Okay, start on. And then weirdly like zoom doesn’t work.\n**Bob Example:** he tried to reset it remotely the other night. Let me ask him.\n**Alice Example:** I mean it’s really just like we need to get zoom to fix this.\n**Bob Example:** Okay, let me.","expected_messages":4,"expected_participants":["Alice Example","Bob Example"]} {"fixture_id":"bold-name-no-time-002","pattern":"bold-name-no-time","frontmatter":{"date":"2026-05-28"},"body":"**Alice Example:** can you hear me now\n**Participant 2:** yeah loud and clear\n**Alice Example:** great, let us start the review","expected_messages":3,"expected_participants":["Alice Example","Participant 2"]} +{"fixture_id":"bold-time-dash-001","pattern":"bold-time-dash","frontmatter":{"date":"2026-04-09"},"body":"**Alice Example** 09:15 — first message\n- supporting detail\n**Bob Example** 09:18 — second message","expected_messages":2,"expected_participants":["Alice Example","Bob Example"]} +{"fixture_id":"bold-time-dash-002","pattern":"bold-time-dash","frontmatter":{"date":"2026-04-10"},"body":"**Summary Bot** 7:05 - ASCII separator\n**Alice Example** 12:30 – Unicode en dash\n**Summary Bot** 23:59 — Unicode em dash","expected_messages":3,"expected_participants":["Summary Bot","Alice Example"]} +{"fixture_id":"adversarial-bold-time-dash-stray","pattern":null,"frontmatter":{"date":"2026-04-09"},"body":"Context paragraph before 1.\nContext paragraph before 2.\nContext paragraph before 3.\nContext paragraph before 4.\nContext paragraph before 5.\nContext paragraph before 6.\nContext paragraph before 7.\nContext paragraph before 8.\nContext paragraph before 9.\nContext paragraph before 10.\n**Deadline** 09:15 — quoted schedule entry\nContext paragraph after 1.\nContext paragraph after 2.\nContext paragraph after 3.\nContext paragraph after 4.\nContext paragraph after 5.\nContext paragraph after 6.\nContext paragraph after 7.\nContext paragraph after 8.\nContext paragraph after 9.\nContext paragraph after 10.","expected_messages":0,"expected_participants":[]} {"fixture_id":"adversarial-bold-labels-clustered-head","pattern":null,"frontmatter":{"date":"2026-05-28"},"body":"**Attendees:** Alice Example, Bob Example, Participant 2\n**Date:** 2026-05-28\n**Goal:** decide on the Q3 roadmap and unblock the vendor migration\nThis is an ordinary prose sentence number 0 describing the meeting in detail.\nThis is an ordinary prose sentence number 1 describing the meeting in detail.\nThis is an ordinary prose sentence number 2 describing the meeting in detail.\nThis is an ordinary prose sentence number 3 describing the meeting in detail.\nThis is an ordinary prose sentence number 4 describing the meeting in detail.\nThis is an ordinary prose sentence number 5 describing the meeting in detail.\nThis is an ordinary prose sentence number 6 describing the meeting in detail.\nThis is an ordinary prose sentence number 7 describing the meeting in detail.\nThis is an ordinary prose sentence number 8 describing the meeting in detail.\nThis is an ordinary prose sentence number 9 describing the meeting in detail.\nThis is an ordinary prose sentence number 10 describing the meeting in detail.\nThis is an ordinary prose sentence number 11 describing the meeting in detail.\nThis is an ordinary prose sentence number 12 describing the meeting in detail.\nThis is an ordinary prose sentence number 13 describing the meeting in detail.\nThis is an ordinary prose sentence number 14 describing the meeting in detail.\nThis is an ordinary prose sentence number 15 describing the meeting in detail.\nThis is an ordinary prose sentence number 16 describing the meeting in detail.\nThis is an ordinary prose sentence number 17 describing the meeting in detail.\nThis is an ordinary prose sentence number 18 describing the meeting in detail.\nThis is an ordinary prose sentence number 19 describing the meeting in detail.\nThis is an ordinary prose sentence number 20 describing the meeting in detail.\nThis is an ordinary prose sentence number 21 describing the meeting in detail.\nThis is an ordinary prose sentence number 22 describing the meeting in detail.\nThis is an ordinary prose sentence number 23 describing the meeting in detail.\nThis is an ordinary prose sentence number 24 describing the meeting in detail.\nThis is an ordinary prose sentence number 25 describing the meeting in detail.\nThis is an ordinary prose sentence number 26 describing the meeting in detail.\nThis is an ordinary prose sentence number 27 describing the meeting in detail.\nThis is an ordinary prose sentence number 28 describing the meeting in detail.\nThis is an ordinary prose sentence number 29 describing the meeting in detail.\nThis is an ordinary prose sentence number 30 describing the meeting in detail.\nThis is an ordinary prose sentence number 31 describing the meeting in detail.\nThis is an ordinary prose sentence number 32 describing the meeting in detail.\nThis is an ordinary prose sentence number 33 describing the meeting in detail.\nThis is an ordinary prose sentence number 34 describing the meeting in detail.\nThis is an ordinary prose sentence number 35 describing the meeting in detail.\nThis is an ordinary prose sentence number 36 describing the meeting in detail.\nThis is an ordinary prose sentence number 37 describing the meeting in detail.\nThis is an ordinary prose sentence number 38 describing the meeting in detail.\nThis is an ordinary prose sentence number 39 describing the meeting in detail.\nThis is an ordinary prose sentence number 40 describing the meeting in detail.\nThis is an ordinary prose sentence number 41 describing the meeting in detail.\nThis is an ordinary prose sentence number 42 describing the meeting in detail.\nThis is an ordinary prose sentence number 43 describing the meeting in detail.\nThis is an ordinary prose sentence number 44 describing the meeting in detail.\nThis is an ordinary prose sentence number 45 describing the meeting in detail.\nThis is an ordinary prose sentence number 46 describing the meeting in detail.\nThis is an ordinary prose sentence number 47 describing the meeting in detail.\nThis is an ordinary prose sentence number 48 describing the meeting in detail.\nThis is an ordinary prose sentence number 49 describing the meeting in detail.\nThis is an ordinary prose sentence number 50 describing the meeting in detail.\nThis is an ordinary prose sentence number 51 describing the meeting in detail.\nThis is an ordinary prose sentence number 52 describing the meeting in detail.\nThis is an ordinary prose sentence number 53 describing the meeting in detail.\nThis is an ordinary prose sentence number 54 describing the meeting in detail.\nThis is an ordinary prose sentence number 55 describing the meeting in detail.\nThis is an ordinary prose sentence number 56 describing the meeting in detail.\nThis is an ordinary prose sentence number 57 describing the meeting in detail.\nThis is an ordinary prose sentence number 58 describing the meeting in detail.\nThis is an ordinary prose sentence number 59 describing the meeting in detail.\nThis is an ordinary prose sentence number 60 describing the meeting in detail.\nThis is an ordinary prose sentence number 61 describing the meeting in detail.\nThis is an ordinary prose sentence number 62 describing the meeting in detail.\nThis is an ordinary prose sentence number 63 describing the meeting in detail.\nThis is an ordinary prose sentence number 64 describing the meeting in detail.\nThis is an ordinary prose sentence number 65 describing the meeting in detail.\nThis is an ordinary prose sentence number 66 describing the meeting in detail.\nThis is an ordinary prose sentence number 67 describing the meeting in detail.\nThis is an ordinary prose sentence number 68 describing the meeting in detail.\nThis is an ordinary prose sentence number 69 describing the meeting in detail.\nThis is an ordinary prose sentence number 70 describing the meeting in detail.\nThis is an ordinary prose sentence number 71 describing the meeting in detail.\nThis is an ordinary prose sentence number 72 describing the meeting in detail.\nThis is an ordinary prose sentence number 73 describing the meeting in detail.\nThis is an ordinary prose sentence number 74 describing the meeting in detail.\nThis is an ordinary prose sentence number 75 describing the meeting in detail.\nThis is an ordinary prose sentence number 76 describing the meeting in detail.\nThis is an ordinary prose sentence number 77 describing the meeting in detail.\nThis is an ordinary prose sentence number 78 describing the meeting in detail.\nThis is an ordinary prose sentence number 79 describing the meeting in detail.","expected_messages":0,"expected_participants":[]} diff --git a/test/fixtures/conversation-formats/bold-time-dash.jsonl b/test/fixtures/conversation-formats/bold-time-dash.jsonl new file mode 100644 index 000000000..b30a8f732 --- /dev/null +++ b/test/fixtures/conversation-formats/bold-time-dash.jsonl @@ -0,0 +1,2 @@ +{"fixture_id":"bold-time-dash-001","pattern":"bold-time-dash","frontmatter":{"date":"2026-04-09"},"body":"**Alice Example** 09:15 — first message\n- supporting detail\n**Bob Example** 09:18 — second message","expected_messages":2,"expected_participants":["Alice Example","Bob Example"]} +{"fixture_id":"bold-time-dash-002","pattern":"bold-time-dash","frontmatter":{"date":"2026-04-10"},"body":"**Summary Bot** 7:05 - ASCII separator\n**Alice Example** 12:30 – Unicode en dash\n**Summary Bot** 23:59 — Unicode em dash","expected_messages":3,"expected_participants":["Summary Bot","Alice Example"]}