Four findings from a second independent blind review, all reproduced against
HEAD before the fix and all pinned by a mutation-tested case.
- sanitizeModelText escapes &, < and > after the existing comment/mention/
block-marker stripping, so raw HTML from any dynamic string (model output,
mechanical flag details built from PR filenames, neutralReason, the
policy-exempt note) renders as literal text. A <details open><summary>MERGE
LANE — approved</summary> string is a forged verdict inside a close-lane
comment; it now renders as <details.
- hasScreenshot requires a markdown image whose URL looks like a URL or path,
an <img> carrying a non-empty src=, or a bare paste URL, and strips HTML
comments before scanning. ``, `<img alt=proof>` and an image
hidden inside `<!-- -->` no longer clear it. Documented in the code as a
FLOOR against zero-effort submissions, not proof.
- stripCodeFences is a line scanner following the CommonMark rule that a
closing fence must be the same character and at least as long as the
opening one. The old backreference read ```` as a NEW opening fence and
stripped to EOF, so a compliant body documenting fence syntax lost its
intent paragraph and was closed. The scanner is also linear, retiring the
superlinear-backtracking hazard the 16KB cap was sized against.
- Labels are reconciled BEFORE the sticky comment carrying the cached state.
Written the other way, one transient label-API 500 left stale/missing/
duplicate labels forever: the rerun short-circuited on the persisted state
and returned success without repairing them.
Also:
- hashInputs covers what the run consumes — the truncated model body plus the
mechanical policy outcome — so an edit past the 6KB model cap no longer
mints a new hash and buys an identical paid call, while a policy fix landing
past that cap still invalidates the cached verdict.
- The test's YAML block-scalar scanner accepts indentation indicators in both
legal orders (`>2-` as well as `|-2`), with a guard-the-guard case proving
it sees a `run: >2-` block interpolating attacker-controlled text.
- A block at the top of the script states plainly that this gate is a triage
signal and a reviewer checklist, not an authorization boundary; one line of
the sticky comment footer says the same to the contributor.
145 pass / 0 fail (was 126), typecheck clean, actionlint clean, verify 34/34.
Blind review round 2 rejected the branch. Five findings, all fixed.
BLOCKING 1 — the verdict was forgeable through a PR filename.
renderComment wrote mechanical red-flag details raw, and two of them
interpolate filenames (adds_recipe, deletes_tests). git allows a newline
inside a filename and JS `[^/]` matches one, so RECIPE_RE's anchors were
decorative: a PR adding `src/core/ai/recipes/x\n## PR Gate — ...\ncc
@octocat\n<!-- ...state... -->\nz.ts` put that text verbatim into the bot's
comment — forged heading, live third-party mention, and on the NEUTRAL
render (which writes no state block of its own) parseState() returned the
ATTACKER's block, so the next run hit the spend guard and silently skipped
the verdict with no label and exit 0. Three layers:
(a) flag details go through sanitizeList, exactly like the model's
strings. Audited every other interpolation into the comment; the rest
are literals in the file, Number()-coerced, or already sanitized.
(b) every path regex spells its segment class [^/\n], not [^/] —
RECIPE_RE, SOURCE_EXT_RE and the test-path check.
(c) parseState reads the state block only off line 2 of a marker-leading
comment (where renderComment writes it) and STATE_RE is whole-line
anchored. A block anywhere else is somebody else's text.
BLOCKING 2 — no exemption, so every release PR was close-lane.
Measured: 40 of the last 40 merged PRs would be close-lane on
missing_screenshot, including every /ship release PR. A check that is red
on every release gets switched off within a week, and then it filters
nothing. The #3745 policy exists to filter INCOMING OUTSIDE CONTRIBUTIONS;
release automation cannot take a screenshot of itself. It is now waived for
OWNER/MEMBER/COLLABORATOR, bot authors and drafts — the usefulness verdict,
the title rule and every mechanical red flag still run, and the sticky
comment says the check was skipped. author_association / draft / user.type
are read from the pr.json the workflow already fetches: no new API call, one
source of truth. `draft` is the one author-settable input, so
ready_for_review joins the trigger list and the exemption is folded into the
spend-guard hash — the draft-era verdict cannot be reused after the flip.
Stated as a deliberate decision in both the workflow header and the script.
3 — DOWNGRADE_FLAG_IDS omitted deletes_tests, adds_symlink and
adds_node_modules, so a PR deleting test/e2e/engine-parity.test.ts kept
merge-lane and a green check on the strength of its prose. All three added.
The old test used deletes_tests as its example of a NON-downgrading flag;
rewritten to pin the stronger invariant instead — every id detectRedFlags
can emit is a downgrade trigger (derived from the detector, so a new flag
fails until it is classified on purpose), and the set is still an allowlist
(an unrecognized id changes nothing).
4 — FENCE_RE backtracks superlinearly on a hostile body: 65KB of backticks
(GitHub's max body length) measured 8.2s across the two policy scans on a
pull_request_target runner. stripCodeFences now caps the scan at 16KB —
same input, 0.40s. Tradeoff documented at the constant: the intent
paragraph and the screenshot both sit near the top in practice (the PR
template puts them in the first two sections, and the model payload already
caps the same body at 6KB), so a real contributor is not judged on a
truncated tail.
Verified: test/pr-gate-workflow.test.ts 126 pass / 0 fail (was 95),
typecheck clean, actionlint clean, check-privacy /
check-no-tracked-symlinks / check-progress-to-stdout /
check-bun-test-timeout / check-key-files-current-state all exit 0. Each new
pin was mutation-tested against the pre-fix behavior: all six mutants fail
the suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three fixes on the strict PR usefulness gate (#3698), plus the merge that
brings in the policy it enforces.
1. The mechanical policy check now outlives the model. The
ANTHROPIC_API_KEY guard used to sit above detectPolicyMisses, so a PR
with no intent paragraph and no screenshot got a green NEUTRAL skip
whenever the key was absent or Anthropic was down — "wait for a 500"
was a documented way past the one hard requirement. The #3745 branch
now sits above the key guard and the spend guard: a policy miss is
close-lane + the friendly fix-it comment + exit 1 with no API
dependency at all. A compliant PR that hits a missing key or a dead
API keeps the round-1 NEUTRAL behavior unchanged (loud comment,
::warning::, exit 0, stale gate:* labels cleared) — and the NEUTRAL
comment now says plainly that the *usefulness verdict* did not run,
while still reporting the title check and mechanical red flags it was
able to compute without a model.
2. Merged origin/master, which carries #3745's CONTRIBUTING.md section
and .github/pull_request_template.md. No conflicts: this branch never
touched VERSION / package.json / CHANGELOG.md, so master's 0.42.72.1
carried through untouched — the feature branch adds no version bump.
The test's inlined pull_request_template fallback (only needed while
the branch predated the merge) is gone; it now reads the real file, so
growing the template's own prose past the 40-word bar fails here
instead of silently letting an untouched template through. The
CONTRIBUTING_URL deep link is pinned against a GitHub-style slug of
every heading in the merged CONTRIBUTING.md, with the slugger itself
pinned so it cannot "pass" against an anchor GitHub never generates.
3. hashInputs joined its three fields with literal NUL bytes, which made
grep treat the whole of scripts/pr-gate.mjs as binary — any future
grep-based CI guard over that file would have matched nothing and
passed silently. Replaced with JSON.stringify of the tuple: still
unforgeable (each field is quoted and escaped), still stable by
construction, and printable. `grep -c hashInputs scripts/pr-gate.mjs`
now returns 2 instead of nothing. Existing sticky-comment state hashes
are invalidated once, costing one re-verdict per open PR.
Tests: 95 pass / 0 fail in test/pr-gate-workflow.test.ts. The no-API-key
policy-miss case was verified to fail against the pre-fix ordering.
Verified live against the Anthropic API: HTTP 200, strict JSON, all seven
required keys, merge-lane on a compliant fixture.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CONTRIBUTING.md (#3745) requires every PR to carry a paragraph the author
wrote themselves and a screenshot of gbrain actually in use. The gate now
checks both, mechanically, before it spends anything on review.
What it checks (no LLM, both exported for testing):
- hasScreenshot: markdown ``, a bare user-images.githubusercontent
or github.com/user-attachments/assets URL, or an <img> tag. Anything inside
a fenced code block does not count — pasting the syntax is not attaching
the picture.
- hasIntentParagraph: >= 40 words of prose left after stripping fenced code,
blockquotes, list items, headings, HTML comments, links and the PR
template's own boilerplate. Per-character scripts are tokenized per
character, so a paragraph written in Chinese counts as one.
The two lane consequences:
- missing_screenshot OR missing_intent forces close-lane from any recommended
lane (exit 1) and skips the model call entirely — closed without review is
the documented consequence, so there is nothing to spend a review on. The
sticky comment leads with what is missing, how to fix it, and the reopen
path; both misses are also recorded in the existing "Mechanical downgrades
applied" section.
- The model's new advisory intent_authenticity verdict forces
needs-maintainer (exit 0) when it reads "ai_generated", and never
close-lane on that signal alone. The comment says only that a maintainer
will read the paragraph personally; the model's reasoning is consumed and
never published at the contributor.
Rubric + strict-JSON schema gain intent_authenticity and
intent_authenticity_reason, with explicit instructions that rough grammar,
terseness and non-native English are evidence of a HUMAN and that "unclear"
is the answer whenever the evidence is not clear-cut.
test/pr-gate-workflow.test.ts: 63 -> 88 tests, all green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two independent blind reviews of #3698 (one APPROVE, one REJECT). Every
confirmed finding from the reject side, plus the cheap hardening both
reviews flagged.
BLOCKING
1. Sticky-comment hijack. upsertStickyComment adopted ANY comment containing
the marker, so a contributor could pre-post `<!-- gbrain-pr-gate -->`, have
the gate PATCH it, then edit it into a fake green verdict. isOwnComment()
now requires user.type === 'Bot' AND login === 'github-actions[bot]' AND
the body to START WITH the marker; anything else gets a fresh comment.
2. LLM output injected as raw Markdown. Model-produced reasons[] and
reviewer_checklist[] reached the comment unescaped — PR-body-driven
injection could forge headings, a second marker, and live @mentions.
sanitizeModelText()/sanitizeList() are now the single choke point in
renderComment(): HTML comments stripped, mentions zero-width-broken,
leading block markers removed, newlines collapsed, 300 chars per string,
8 entries per list, both caps self-marking.
3. Lane was purely model-decided. A well-written feature pitch could talk
itself into merge-lane. The model now RECOMMENDS; applyMechanicalDowngrades
forces merge-lane -> needs-maintainer on any of: workflow edits, a new
package.json dependency, a new src/core/ai/recipes/ provider file, new
KNOWN_CONFIG_KEYS entries, >40 changed files, >400 net source lines outside
test/, or a src/ change with no test file touched (#3665). The sticky
comment reports them under "Mechanical downgrades applied".
4. The test file did not pin what it claimed. Added: no gh pr checkout /
git fetch / refs/pull / pull/*/head in any spelling; exact permissions
key->value map plus a single-permissions-block assertion so no job-level
grant re-widens contents; the ${{ }}-in-run scanner now covers folded
(`run: >`) and chomped blocks, with a guard-the-guard test; and mocked
end-to-end runGate() runs for close-lane exit 1, marker hijack, sanitizer,
truncation, refusal routing, NEUTRAL label clearing, label swap, and the
spend guard.
5. Version-first title regex rejected the documented suffix form.
`v0.31.1.1-fixwave fix: ...` now passes. VERSION_AT_END_RE no longer
false-positives on `chore: bump zod (3.25.76)`: it fires only on a
v-prefixed or 4-segment trailing version, i.e. this project's own shape.
6. Refusal fail-open. stop_reason=refusal exhausted retries into a green
NEUTRAL — a deterministic way to dodge the red X. Refusal and
schema-invalid output now route to needs-maintainer with an explicit
note; only transport failure stays NEUTRAL. Refusal also short-circuits
the retry loop, since retrying a deterministic refusal only burns spend.
ALSO
7. persist-credentials: false on the checkout step.
8. Dropped pull-requests:write. Everything the script calls is the issues
API (comments, label create, label add/remove), so issues:write is the
only grant that is actually needed.
9. Spend guard for the edited/synchronize amplification: the LLM call is
skipped when sha256(title+body+head_sha) matches the hash recorded in the
previous sticky comment's state block, and the stored lane's exit code is
reused.
10. NEUTRAL runs now clear every gate:* label instead of leaving a stale
verdict behind.
Verified: bun test test/pr-gate-workflow.test.ts 63 pass / 0 fail,
bun run typecheck clean, actionlint clean, check-privacy /
check-no-tracked-symlinks / check-progress-to-stdout /
check-bun-test-timeout clean, plus a live Anthropic smoke of the exact
request shape (HTTP 200, valid strict JSON, injection attempt rejected).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Runs on pull_request_target against master. PR code is never checked out
or executed: metadata + a 120KB-capped diff come from the GitHub API only;
the base-repo checkout provides the rubric script. All interpolations are
env-bound; actions SHA-pinned; permissions limited to contents:read +
pull-requests:write + issues:write; per-PR concurrency with cancel.
scripts/pr-gate.mjs (no deps, fetch only) classifies via the strict
usefulness rubric (claude-sonnet-5, strict JSON schema, thinking disabled;
sampling params omitted — rejected on this model), posts one sticky
comment (<!-- gbrain-pr-gate -->), applies exactly one gate:* label, and
exits 1 only on close-lane. Missing ANTHROPIC_API_KEY or API failure after
2 retries NEUTRAL-skips loudly (comment + warning, exit 0). Mechanical
(no-LLM) checks: version-first title rule and diff red flags (>40 files,
node_modules, symlinks, workflow edits force needs-maintainer, new
package.json deps, deleted tests).
test/pr-gate-workflow.test.ts pins the security invariants and unit-tests
the exported title rule + red-flag detector (import side-effect guarded).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
runImport printed its informational lines ('Found N markdown files' and four
siblings) through console.log unconditionally, so `gbrain import <dir> --json`
emitted them ahead of the payload and stdout did not parse as JSON. A consumer
parsing stdout reads that as zero imports while its own bookkeeping records the
files as ingested, so the next run skips them permanently.
Route those five lines through an `info` helper that switches to stderr when
--json is set. The lines are relocated, not removed: human mode is byte-for-byte
unchanged, and the JSON payload itself is untouched. This is the contract
CLAUDE.md already states ('Stdout stays clean for data output (--json
payloads)') and that import.ts's own progress comment repeats.
scripts/check-progress-to-stdout.sh only greps for process.stdout.write('\\r…),
so this class was never guarded; the new spawn-level test pins it.
installHelper's "already current" fast path ran chmodSync(helperPath,
0o755) unconditionally, before the dryRun check further down — so
`gbrain sources harden --dry-run` mutated the helper's permissions even
though dry-run is documented as a pure preview. This is the 5th
instance of the #3594 class (see #3692 for the sibling fix to the
step-1 pull, same function).
Gate the chmod with `if (!dryRun)`, mirroring the guard pattern already
used in installLocalHook's own "already current" branch. Non-dry-run
behavior is unchanged.
Adds a regression pair to test/brain-repo-durability.serial.test.ts:
one asserting dry-run leaves a drifted-permission helper untouched,
and a control asserting the real run still restores the exec bit.
Two cycle phases still carried the label-vs-actual defect that
propose_takes shed in v0.42.62 (and that the #2805 review noted was worth
tracking so it isn't lost):
- grade_takes hardcoded 'claude-sonnet-4-6' into judge_model_id, the
evidence signature, and budget metering, while the default judge call
passed NO model hint and rode the gateway's chat_model. On any brain
with a non-default chat_model, the verdict cache and telemetry recorded
a model that never ran.
- calibration_profile persisted TIER_DEFAULTS.reasoning to model_id but
never passed a model to the patterns generator at all, so the recorded
model and the executed one were unrelated.
Fix, matching the propose_takes convention: each phase resolves ONE
string via explicit override > getChatModel(), and that string drives the
chat call, the cache key, and the stored id.
- grade_takes: the judge hint gets the FULL provider-prefixed string; the
stored judge_model_id and evidence signature keep the historical bare
tail. Stock installs are unchanged: getChatModel() defaults to
'anthropic:claude-sonnet-4-6', whose tail equals the old hardcoded
value, so no verdict cache invalidates. A genuinely different
chat_model invalidates, which is correct — the judge really changed.
- calibration_profile: getChatModel() is provider-prefixed, preserving
the #2451 contract, and its default IS the old TIER_DEFAULTS.reasoning
value — stock behavior unchanged. The generator now receives the same
resolved string that is persisted to model_id.
Tests: regression per phase pinning configured-chat-model routing (full
string to the call, bare tail to the cache key for grade_takes, full
string persisted for calibration); all existing suites pass unchanged,
pinning the no-change-on-stock property. 112 tests across the four
affected suites; typecheck clean.
Co-authored-by: Paolo Belcastro <p3ob7o@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The comment listed `gbrain integrity --dry-run` as a subcommand, but
runIntegrity dispatches on `check`/`auto` and parses --dry-run inside
the auto path, so that form exits 1 with "Unknown subcommand". The
runtime help already shows --dry-run indented under `auto [options]`;
only the source comment disagreed.
Addresses #3738.
The dead-link branch in cmdAuto called logSkip() (writing to
integrity.log.jsonl) but incremented bucketReview, the same counter fed by
the bare-tweet path's appendReview() calls (which write to
integrity-review.md). The printed "Review queue" summary line therefore
included findings that were never written to the review file, so
`gbrain integrity review` silently disagreed with `integrity auto`'s own
summary.
Give dead-link findings their own counter (bucketDeadLink) and print it as
a separate line. No change to bare-tweet bucketing or CLI surface.
retype mapping_rules' existing path_filter matches pages.source_path,
which is only populated for pages synced from a git repo. Pages ingested
via the put_page MCP tool (or any write path that doesn't go through
sync) have source_path = NULL, so path_filter can never disambiguate
them — a same-from_type retype rule targeting a slug prefix has no way
to address this class of page at all.
slug_filter adds an independent, orthogonal LIKE filter on pages.slug
(the field that's always populated), combinable with path_filter via AND
when both are given. Wired through the full call path: schema validation
(manifest-v1.ts), the RetypeRule interface + probeRule/applyRetypeRule
(retype.ts), and the pack-manifest → RetypeRule conversion in the
unify-types job handler (unify-types-handler.ts) — the last one matters
because a field only present in the zod schema but not carried through
that conversion would validate fine yet silently no-op at execution time.
- manifest-v1.ts: add optional slug_filter to RetypeMappingRuleSchema
- retype.ts: slug_filter on RetypeRule; probeRule + applyRetypeRule both
apply `AND slug LIKE $N` when present, independent of path_filter
- unify-types-handler.ts: carry rule.slug_filter through the
pack-mapping-rule → RetypeRule conversion
- tests: skips-outside-slug_filter, matches-despite-NULL-source_path
(the motivating case), path_filter+slug_filter combined via AND
320/320 existing schema-pack tests pass; 3 new tests added; tsc --noEmit
clean.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Time Attakc <89218912+time-attack@users.noreply.github.com>
Effective 2026-08-02. Every issue and PR must carry a paragraph the author
wrote themselves explaining why they are opening it, and a screenshot showing
gbrain actually in use in that situation. AI-generated or AI-polished intent
text is not accepted; AI assistance for the code is still fine. Missing either
one means closed without review, reopenable once added.
Stated in CONTRIBUTING.md and pre-filled in both issue templates plus a new
pull request template so the fields are in front of the author.
Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Registering an OAuth client with --bound-slug-prefixes now makes the write boundary real: writes outside the bound prefixes are refused on every op that can name a page, and ops that write by something other than a slug are refused outright rather than left unfenced. Deny-by-default at dispatch, so a write op added later is refused to bound clients until it is explicitly fenced.
Adds docs/integrations/qm-harness.md (gbrain as the company brain for a qm deployment) with a roster-driven provisioning script and deployment templates, plus a Known limitations section stating plainly that this is a write boundary and not a privacy boundary.
Five review rounds, including three clean-room passes by codex gpt-5.6-sol and Claude Fable 5 against an instruction-stripped tree.
The first Release run (30698650484) failed before building anything: the
build job re-ran the entire unit suite serially on the release runner —
a different environment from the sharded Test workflow that had already
gated the exact SHA at merge — and died on ambient-env tests. The review
of #3573 predicted this ('a flaky test now blocks the release'). The
build job now compiles and smoke-tests the binary (--version must match
the VERSION file), which validates the actual artifact — something the
test suite never did. The Test workflow remains the code gate.
Review hardening from the #3573 gatekeeper pass: a ${{ }} inside run: is
shell injection by construction; bind through env instead. Not
attacker-reachable today (VERSION comes from master), correct anyway.
CI caught what the composite review's watermark 'fix' actually did: the
stamp path clamps links_extracted_at up to the watermark
(GREATEST(updated_at, versionTs), extract.ts:1835), so a future watermark
stamps every page at 2026-08-02 and masks any concurrent edit until then —
the exact race D4 guards — and makes the doctor lag check warn on fresh
brains. The limitation the future date tried to cover (stamps written by
pre-wave code after the watermark) is inherent: no fixed watermark covers
code that keeps running past it. Documented instead.
Two cross-PR findings from the wave-i hostile review (Codex xhigh):
1. The #3560-ungated bare-path pass scanned inside [[...]] spans, so a
dir-qualified wikilink's lowercase prefix (its parent page) became a
spurious 'markdown' edge whenever the parent existed. Wikilink spans
are now masked with equal-length blanks before pass 2; discriminating
test added (fails without the mask).
2. LINK_EXTRACTOR_VERSION_TS was midnight today with a strict-< staleness
predicate, so same-day stamps from pre-wave code read as fresh and
never re-extracted. Bumped to 2026-08-02T00:00:00Z.
Plus two comment corrections from both reviewers: upgrade.ts's X1 hook
rationale (stale after #3085) and PGLiteEngine.transaction's tx-engine
db-proxy hazard for #3613's searchVector wrapper.
In-wave interaction with #3560: extractPageLinks now also emits raw-literal and bare-path candidates that downstream existence checks drop; the tests' whole-list toEqual predates that. Filter to linkSource='wikilink-resolved' — the PR's resolution + no-cross-dir-leak claims are still fully pinned.
Master's #1072 test asserted bare qwen3-embedding@1024 threads dimensions:1024;
wave PR #3699 suppresses the param when the request equals the native width
(1024 for the bare id), so the same input now correctly returns undefined —
pinned by dims-qwen3-native.test.ts. Moved this test to 512 to keep its actual
intent (bare id recognized as Matryoshka-capable) without contradicting the
suppression. Third composition defect of the wave; caught by CI shard 8.
#3574 flipped the unify-types worker default to dry-run (jobs.ts:2221,
apply: data.apply ?? false) and updated the architecture docs, but three
agent-facing surfaces still presented the bare submit as the Apply step:
skills/schema-unify/SKILL.md 'Phase 3: Apply', skills/conventions/
schema-evolution.md, and README.md. Because #3545 also edited SKILL.md in
this wave, each PR looked self-consistent in isolation — only the composed
branch shipped a playbook whose apply step silently retypes nothing and
never flips the active pack. Skills distribute downstream via the skillpack,
so this would have propagated. Found by an independent cross-PR review pass.
#3691's regression test used llama-server as an 'unknown provider' example.
#3541 (same wave) prices ollama/llama-server at $0 via FREE_LOCAL_CHAT_PROVIDERS,
so that example is now priceable and the assertion inverted. Swapped in groq —
the paid-but-unpriced case #3691's own description cites — and added the
positive assertion that free local providers keep their cap enforced.