diff --git a/README.md b/README.md index 8c78fd6..0a4f5e0 100644 --- a/README.md +++ b/README.md @@ -2,7 +2,7 @@ Autonomous experiment loop for any optimization target. -Faithful OpenClaw port of [`davebcn87/pi-autoresearch`](https://github.com/davebcn87/pi-autoresearch). +Faithful OpenClaw port of [`davebcn87/pi-autoresearch`](https://github.com/davebcn87/pi-autoresearch), including upstream statistical confidence scoring. ## How it works @@ -12,13 +12,13 @@ Three tools drive the loop: | Tool | What it does | |---|---| -| `init_experiment` | Configures the session: name, primary metric, unit, direction (lower/higher). Re-calling starts a new segment. | +| `init_experiment` | Configures the session: name, primary metric, unit, direction (lower/higher). Once runs exist, starting a new segment requires `reset: true`, and the prior segment's best result is carried forward into checkpoint context. | | `run_experiment` | Executes a shell command, times it, captures stdout/stderr, parses `METRIC name=number` lines, and opens a pending experiment window that must be logged before another run can start. | -| `log_experiment` | Records the pending run. `keep` auto-commits to git. `discard`/`crash` log without committing. If the prior `run_experiment` captured the primary metric, `log_experiment` can infer `commit` and `metric` automatically. | +| `log_experiment` | Records the pending run. The first logged run in a segment is tagged as the baseline automatically. `keep` auto-commits to git. `discard`/`crash` log without committing, and `discard` now requires an `idea` note that is appended to `autoresearch.ideas.md`. If the prior `run_experiment` captured the primary metric, `log_experiment` can infer `commit` and `metric` automatically. After 3+ runs in a segment, it also reports a confidence score for the best improvement versus noise. | Each tool also accepts an optional `cwd` so callers can target a nested repo explicitly instead of relying on the current session working directory. -All state lives in five repo-root files: +All state lives in six repo-root files: | File | Purpose | |---|---| @@ -27,6 +27,7 @@ All state lives in five repo-root files: | `autoresearch.jsonl` | Structured log: config headers + experiment entries (metric, status, timestamp, segment, commit hash). | | `autoresearch.ideas.md` | Backlog of promising ideas not yet tried. Optional. | | `autoresearch.checkpoint.json` | Plugin-managed checkpoint: latest logged state, recent runs, and any pending unlogged run. | +| `autoresearch.lock` | Session lock with PID + timestamp so another agent can detect an active or stale loop before forking a second session. | The design is file-first: any agent can pick up the repo-root files and continue the loop without prior context. @@ -77,9 +78,13 @@ Prefer the explicit `/autoresearch` command surface in OpenClaw. The auto-genera - `run_experiment` refuses to start a second run until the previous one is logged. - `run_experiment` parses `METRIC name=number` lines and stores a pending run so `log_experiment` can default from the actual benchmark output. +- `init_experiment` refuses to reset a live history unless `reset: true` is passed explicitly. +- The first `log_experiment` in a segment is tagged as the baseline automatically, even if it is later discarded. +- `discard` logs must include an `idea` note, and that note is appended to `autoresearch.ideas.md`. - During active autoresearch mode, raw benchmark execution through OpenClaw `exec`/`bash` is blocked. Use `run_experiment` instead. -- `autoresearch_status` warns when a pending run is unlogged or git history has moved ahead of the last logged experiment. -- The plugin updates `autoresearch.checkpoint.json` and refreshes plugin-managed sections in `autoresearch.md` after init, run, and log transitions. +- `autoresearch_status` warns when a pending run is unlogged, when the canonical branch has drifted, when a stale/live lock exists, or when git history has moved ahead of the last logged experiment. On `autoresearch/*` branches it explicitly warns not to push unlogged commits. +- After 3+ positive-metric runs in a segment, `log_experiment`, `autoresearch_status`, and the synced session doc report a MAD-based confidence score so the agent can distinguish likely wins from noise. +- The plugin updates `autoresearch.checkpoint.json`, `autoresearch.lock`, and the plugin-managed sections in `autoresearch.md` after init, run, and log transitions. ## Use @@ -89,7 +94,7 @@ In the repo you want to optimize: 2. Run `/autoresearch` or `/autoresearch setup `. 3. Send a normal message with the goal, command, metric (+ direction), files in scope, and constraints. 4. If you need the raw skill invocation, use `/skill autoresearch-create`. -5. The agent writes `autoresearch.md` and `autoresearch.sh`, runs a baseline with `run_experiment`, then records it with `log_experiment`. +5. The agent writes `autoresearch.md` and `autoresearch.sh`, captures or reuses the canonical `autoresearch/*` branch, runs a baseline with `run_experiment`, then records it with `log_experiment`. 6. Use `/autoresearch` or `/autoresearch status` to re-prime context on a later turn. To resume an existing session, a new agent reads the repo-root files and continues from where the last one stopped. @@ -100,7 +105,7 @@ Messages sent while an experiment is running are queued and surfaced after the n ### Ideas backlog -When the agent discovers promising but complex ideas mid-loop, it appends them to `autoresearch.ideas.md`. On resume, the agent reads the backlog, prunes stale entries, and uses the remaining ideas as experiment paths. +When the agent discovers promising but complex ideas mid-loop, it appends them to `autoresearch.ideas.md`. Discarded experiments now require an `idea` note, so failed paths leave behind concrete follow-up suggestions instead of disappearing. On resume, the agent reads the backlog, prunes stale entries, and uses the remaining ideas as experiment paths. ## Upstream reference @@ -108,6 +113,7 @@ This port preserves upstream semantics, names, and file contracts while adapting - upstream repo: `https://github.com/davebcn87/pi-autoresearch` - pinned upstream commit: `2227029fa5712944a36938b5fe59f709cb30ed22` (`2227029f`) +- later upstream parity cherry-pick: confidence scoring from `cf1bbf03debca8f3fb2cca2c3e799b9e23320f87` (`cf1bbf0`, March 19, 2026) ## Validation diff --git a/docs/non-parity.md b/docs/non-parity.md index 988c91d..e335ae2 100644 --- a/docs/non-parity.md +++ b/docs/non-parity.md @@ -6,6 +6,7 @@ Pinned upstream reference: - Repo: `https://github.com/davebcn87/pi-autoresearch` - Commit: `2227029fa5712944a36938b5fe59f709cb30ed22` (`2227029f`) +- Later upstream semantics also ported: confidence scoring from `cf1bbf03debca8f3fb2cca2c3e799b9e23320f87` (`cf1bbf0`) ## Principle diff --git a/extensions/openclaw-autoresearch/src/checkpoint.ts b/extensions/openclaw-autoresearch/src/checkpoint.ts index 5e88e60..59d3caf 100644 --- a/extensions/openclaw-autoresearch/src/checkpoint.ts +++ b/extensions/openclaw-autoresearch/src/checkpoint.ts @@ -6,10 +6,19 @@ import type { } from "./state.js"; import type { PendingExperimentRun } from "./runtime-state.js"; +export type AutoresearchCarryForwardContext = { + readonly metricName: string; + readonly metricUnit: string; + readonly bestDirection: "lower" | "higher"; + readonly run: AutoresearchRunSnapshot; +}; + export type AutoresearchCheckpoint = { readonly version: 1; readonly updatedAt: number; readonly sessionStartCommit: string | null; + readonly canonicalBranch?: string | null; + readonly carryForwardContext?: AutoresearchCarryForwardContext | null; readonly session: { readonly name: string | null; readonly metricName: string; @@ -20,6 +29,7 @@ export type AutoresearchCheckpoint = { readonly totalRunCount: number; readonly currentBaselineMetric: number | null; readonly currentBestMetric: number | null; + readonly confidence: number | null; }; readonly lastLoggedRun: AutoresearchRunSnapshot | null; readonly recentLoggedRuns: readonly AutoresearchRunSnapshot[]; @@ -44,6 +54,8 @@ export function writeAutoresearchCheckpoint(options: { cwd: string; state: AutoresearchStateSnapshot; sessionStartCommit: string | null; + canonicalBranch: string | null; + carryForwardContext: AutoresearchCarryForwardContext | null; recentLoggedRuns: readonly AutoresearchRunSnapshot[]; pendingRun: PendingExperimentRun | null; }): AutoresearchCheckpoint { @@ -51,6 +63,8 @@ export function writeAutoresearchCheckpoint(options: { version: 1, updatedAt: Date.now(), sessionStartCommit: options.sessionStartCommit, + canonicalBranch: options.canonicalBranch, + carryForwardContext: options.carryForwardContext, session: { name: options.state.name, metricName: options.state.metricName, @@ -61,6 +75,7 @@ export function writeAutoresearchCheckpoint(options: { totalRunCount: options.state.totalRunCount, currentBaselineMetric: options.state.currentBaselineMetric, currentBestMetric: options.state.currentBestMetric, + confidence: options.state.confidence, }, lastLoggedRun: options.state.lastRun, recentLoggedRuns: options.recentLoggedRuns, diff --git a/extensions/openclaw-autoresearch/src/commands/autoresearch.ts b/extensions/openclaw-autoresearch/src/commands/autoresearch.ts index 85e16f0..e4847d2 100644 --- a/extensions/openclaw-autoresearch/src/commands/autoresearch.ts +++ b/extensions/openclaw-autoresearch/src/commands/autoresearch.ts @@ -14,6 +14,11 @@ import { setAutoresearchRunInFlight, setAutoresearchRuntimeMode, } from "../runtime-state.js"; +import { + acquireAutoresearchSessionLock, + getAutoresearchSessionLockStatus, + removeAutoresearchSessionLock, +} from "../session-lock.js"; type CommandContext = { args?: string; @@ -59,6 +64,7 @@ export function registerAutoresearchCommand(api: OpenClawPluginApi): void { setAutoresearchPendingCommand(cwd, null); clearAutoresearchSteers(cwd); setAutoresearchRunInFlight(cwd, false); + removeAutoresearchSessionLock(cwd); return { text: [ "Autoresearch mode OFF.", @@ -86,7 +92,11 @@ export function buildAutoresearchCommandText( ): string { const runtimeState = getAutoresearchRuntimeState(cwd); const presentFiles = getPresentCanonicalFiles(cwd); - const hasSession = presentFiles.length > 0; + const presentSessionFiles = presentFiles.filter( + (file) => file !== AUTORESEARCH_ROOT_FILES.sessionLock, + ); + const hasSession = presentSessionFiles.length > 0; + const lockStatus = getAutoresearchSessionLockStatus(cwd); if (!hasSession) { return [ @@ -95,6 +105,7 @@ export function buildAutoresearchCommandText( `Expected canonical files: ${Object.values(AUTORESEARCH_ROOT_FILES).join(", ")}`, "Recommended OpenClaw entrypoint: `/autoresearch` or `/autoresearch setup `.", "Direct skill fallback: `/skill autoresearch-create`.", + `Session lock: ${formatLockStatus(lockStatus)}`, ].join("\n"); } @@ -105,7 +116,7 @@ export function buildAutoresearchCommandText( ]; if (mode === "status") { - lines.push("", formatAutoresearchStatusText(state, runtimeState)); + lines.push("", formatAutoresearchStatusText(state, runtimeState), `Session lock: ${formatLockStatus(lockStatus)}`); } else if (state.mode === "active" || state.hasSessionDoc) { lines.push( "Use `/autoresearch` or `/autoresearch on` to enable mode for the next agent turn, then continue the upstream loop with `init_experiment`, `run_experiment`, and `log_experiment` as needed.", @@ -120,9 +131,18 @@ export function buildAutoresearchCommandText( } function enableAutoresearchMode(cwd: string, args: string | null): string { + const lockStatus = acquireAutoresearchSessionLock(cwd); + if (lockStatus.state === "active" && !lockStatus.ownedByCurrentProcess) { + return [ + "Autoresearch mode NOT enabled.", + `Another live autoresearch loop holds autoresearch.lock (PID ${lockStatus.pid}, started ${new Date(lockStatus.timestamp ?? 0).toISOString()}).`, + "Resume that loop instead of creating a parallel session.", + ].join("\n"); + } + setAutoresearchRuntimeMode(cwd, "on"); const presentFiles = getPresentCanonicalFiles(cwd); - const hasSession = presentFiles.length > 0; + const hasSession = presentFiles.some((file) => file !== AUTORESEARCH_ROOT_FILES.sessionLock); if (!hasSession) { setAutoresearchPendingCommand(cwd, { @@ -150,6 +170,15 @@ function enableAutoresearchMode(cwd: string, args: string | null): string { } function primeAutoresearchSetup(cwd: string, args: string | null): string { + const lockStatus = acquireAutoresearchSessionLock(cwd); + if (lockStatus.state === "active" && !lockStatus.ownedByCurrentProcess) { + return [ + "Autoresearch setup NOT primed.", + `Another live autoresearch loop holds autoresearch.lock (PID ${lockStatus.pid}, started ${new Date(lockStatus.timestamp ?? 0).toISOString()}).`, + "Resume that loop instead of starting a parallel setup flow.", + ].join("\n"); + } + setAutoresearchRuntimeMode(cwd, "on"); setAutoresearchPendingCommand(cwd, { kind: "setup", @@ -180,3 +209,14 @@ function getPresentCanonicalFiles(cwd: string): string[] { } return present; } + +function formatLockStatus(lockStatus: ReturnType): string { + if (lockStatus.state === "missing") { + return "missing"; + } + + const timestamp = lockStatus.timestamp + ? new Date(lockStatus.timestamp).toISOString() + : "unknown time"; + return `${lockStatus.state} (pid ${lockStatus.pid}, started ${timestamp})`; +} diff --git a/extensions/openclaw-autoresearch/src/confidence.ts b/extensions/openclaw-autoresearch/src/confidence.ts new file mode 100644 index 0000000..8c124c9 --- /dev/null +++ b/extensions/openclaw-autoresearch/src/confidence.ts @@ -0,0 +1,82 @@ +export type ConfidenceRun = { + readonly metric: number; + readonly status: string; +}; + +export function computeConfidence( + runs: readonly ConfidenceRun[], + direction: "lower" | "higher", +): number | null { + const usableRuns = runs.filter((run) => Number.isFinite(run.metric) && run.metric > 0); + if (usableRuns.length < 3) { + return null; + } + + const baseline = runs.find((run) => Number.isFinite(run.metric)); + if (!baseline) { + return null; + } + + const values = usableRuns.map((run) => run.metric); + const median = sortedMedian(values); + const deviations = values.map((value) => Math.abs(value - median)); + const mad = sortedMedian(deviations); + if (mad === 0) { + return null; + } + + let bestKept: number | null = null; + for (const run of usableRuns) { + if (run.status !== "keep") { + continue; + } + + if (bestKept === null || isBetter(run.metric, bestKept, direction)) { + bestKept = run.metric; + } + } + + if (bestKept === null || bestKept === baseline.metric) { + return null; + } + + return Math.abs(bestKept - baseline.metric) / mad; +} + +export function formatConfidenceLine( + confidence: number | null, + label = "Confidence", +): string { + return confidence === null ? `${label}: n/a` : `${label}: ${describeConfidence(confidence)}`; +} + +export function describeConfidence(confidence: number): string { + const rendered = confidence.toFixed(1); + if (confidence >= 2.0) { + return `${rendered}x noise floor - improvement is likely real`; + } + if (confidence >= 1.0) { + return `${rendered}x noise floor - improvement is above noise but marginal`; + } + return `${rendered}x noise floor - improvement is within noise. Consider re-running to confirm before keeping`; +} + +function sortedMedian(values: readonly number[]): number { + if (values.length === 0) { + return 0; + } + + const sorted = [...values].sort((left, right) => left - right); + const mid = Math.floor(sorted.length / 2); + return sorted.length % 2 === 0 + ? (sorted[mid - 1] + sorted[mid]) / 2 + : sorted[mid]; +} + +function isBetter( + current: number, + best: number, + direction: "lower" | "higher", +): boolean { + return direction === "lower" ? current < best : current > best; +} diff --git a/extensions/openclaw-autoresearch/src/files.ts b/extensions/openclaw-autoresearch/src/files.ts index 65a4d41..30d60bf 100644 --- a/extensions/openclaw-autoresearch/src/files.ts +++ b/extensions/openclaw-autoresearch/src/files.ts @@ -6,6 +6,7 @@ export const AUTORESEARCH_ROOT_FILES = { resultsLog: "autoresearch.jsonl", ideasBacklog: "autoresearch.ideas.md", checkpoint: "autoresearch.checkpoint.json", + sessionLock: "autoresearch.lock", } as const; export type AutoresearchRootFileKey = keyof typeof AUTORESEARCH_ROOT_FILES; @@ -29,10 +30,6 @@ export function readAutoresearchRootFile( return fs.readFileSync(filePath, "utf8"); } -/** - * PR 2 skeleton only. - * This module will own canonical root-level file IO helpers in later PRs. - */ export function describeCanonicalFiles(): typeof AUTORESEARCH_ROOT_FILES { return AUTORESEARCH_ROOT_FILES; } diff --git a/extensions/openclaw-autoresearch/src/git.ts b/extensions/openclaw-autoresearch/src/git.ts index ae7a3ad..57cb41e 100644 --- a/extensions/openclaw-autoresearch/src/git.ts +++ b/extensions/openclaw-autoresearch/src/git.ts @@ -54,6 +54,14 @@ export async function readShortHeadCommit(options: GitRuntimeOptions): Promise 0 ? result.stdout.trim() : null; } +export async function readCurrentBranch(options: GitRuntimeOptions): Promise { + const result = await runGitCommand(options.runCommandWithTimeout, options.cwd, [ + "branch", + "--show-current", + ]); + return result.code === 0 && result.stdout.trim().length > 0 ? result.stdout.trim() : null; +} + export async function countCommitsSince( options: GitRuntimeOptions & { sinceCommit: string }, ): Promise { diff --git a/extensions/openclaw-autoresearch/src/hooks.ts b/extensions/openclaw-autoresearch/src/hooks.ts index 75e760a..8e18cde 100644 --- a/extensions/openclaw-autoresearch/src/hooks.ts +++ b/extensions/openclaw-autoresearch/src/hooks.ts @@ -1,6 +1,8 @@ import type { OpenClawPluginApi } from "openclaw/plugin-sdk"; import { AUTORESEARCH_ROOT_FILES } from "./files.js"; import { reconstructStateFromJsonl } from "./state.js"; +import { readAutoresearchCheckpoint } from "./checkpoint.js"; +import { formatConfidenceLine } from "./confidence.js"; import { clearAutoresearchRuntimeState, consumeAutoresearchContinuationReminder, @@ -9,6 +11,7 @@ import { queueAutoresearchSteer, setAutoresearchContinuationReminder, } from "./runtime-state.js"; +import { removeAutoresearchSessionLock } from "./session-lock.js"; type BeforeAgentStartEvent = { systemPrompt?: string; @@ -106,6 +109,7 @@ export function registerAutoresearchHooks(api: OpenClawPluginApi): void { return; } + removeAutoresearchSessionLock(cwd); clearAutoresearchRuntimeState(cwd); }); return; @@ -136,6 +140,7 @@ export function registerAutoresearchHooks(api: OpenClawPluginApi): void { export function buildBeforePromptBuildContext(cwd: string): string | null { const state = reconstructStateFromJsonl(cwd); const runtimeState = getAutoresearchRuntimeState(cwd); + const checkpoint = readAutoresearchCheckpoint(cwd); if (!shouldEnforceAutoresearchMode(cwd, state, runtimeState)) { return null; } @@ -166,7 +171,24 @@ export function buildBeforePromptBuildContext(cwd: string): string | null { "Use init_experiment, run_experiment, and log_experiment for experiment state changes. Never stop unless the user explicitly interrupts the loop.", "Never run benchmark or test commands through raw exec/bash during autoresearch mode. Use run_experiment so the plugin can capture metrics, enforce logging, and preserve resumable state.", "After every run_experiment, call log_experiment before starting another run. If METRIC lines were captured, log_experiment can infer commit and metric from the pending run.", + "Once runs have been logged, init_experiment requires reset: true before starting a new segment.", + `Discarded experiments must include an idea note in log_experiment so it can be appended to ${AUTORESEARCH_ROOT_FILES.ideasBacklog}.`, ); + if (state.confidence !== null) { + lines.push( + `${formatConfidenceLine(state.confidence, "Current confidence")}. Treat low-confidence wins as provisional and re-run before keeping when the score is below 1.0x.`, + ); + } + if (checkpoint?.canonicalBranch) { + lines.push( + `Canonical branch: ${checkpoint.canonicalBranch}. Continue on this branch instead of creating a new autoresearch/* lineage.`, + ); + } + if (checkpoint?.carryForwardContext) { + lines.push( + `Carry-forward best context from the previous segment: ${checkpoint.carryForwardContext.metricName} ${checkpoint.carryForwardContext.run.metric}${checkpoint.carryForwardContext.metricUnit} on ${checkpoint.carryForwardContext.run.commit}.`, + ); + } if (pendingCommand?.args) { lines.push(`Additional resume instruction from /autoresearch: ${pendingCommand.args}`); } diff --git a/extensions/openclaw-autoresearch/src/ideas.ts b/extensions/openclaw-autoresearch/src/ideas.ts new file mode 100644 index 0000000..1b8934f --- /dev/null +++ b/extensions/openclaw-autoresearch/src/ideas.ts @@ -0,0 +1,14 @@ +import * as fs from "node:fs"; +import { getAutoresearchRootFilePath } from "./files.js"; + +export function appendIdeaBacklogEntry(cwd: string, idea: string): void { + const normalized = idea.trim(); + if (!normalized) { + return; + } + + const ideasPath = getAutoresearchRootFilePath(cwd, "ideasBacklog"); + const prefix = + fs.existsSync(ideasPath) && fs.readFileSync(ideasPath, "utf8").trim().length > 0 ? "\n" : ""; + fs.appendFileSync(ideasPath, `${prefix}- ${normalized}\n`); +} diff --git a/extensions/openclaw-autoresearch/src/logging.ts b/extensions/openclaw-autoresearch/src/logging.ts index fe49ca1..00f30b0 100644 --- a/extensions/openclaw-autoresearch/src/logging.ts +++ b/extensions/openclaw-autoresearch/src/logging.ts @@ -44,9 +44,11 @@ export type AutoresearchResultEntry = { readonly metric: number; readonly metrics: Record; readonly status: "keep" | "discard" | "crash"; + readonly baseline?: boolean; readonly description: string; readonly timestamp: number; readonly segment: number; + readonly confidence: number | null; }; export function appendResultEntry(cwd: string, entry: AutoresearchResultEntry): void { diff --git a/extensions/openclaw-autoresearch/src/session-doc.ts b/extensions/openclaw-autoresearch/src/session-doc.ts index 449b4d1..8ac5b63 100644 --- a/extensions/openclaw-autoresearch/src/session-doc.ts +++ b/extensions/openclaw-autoresearch/src/session-doc.ts @@ -1,6 +1,7 @@ import * as fs from "node:fs"; import { AUTORESEARCH_ROOT_FILES, getAutoresearchRootFilePath } from "./files.js"; import type { AutoresearchCheckpoint } from "./checkpoint.js"; +import { formatConfidenceLine } from "./confidence.js"; export function syncAutoresearchSessionDoc( cwd: string, @@ -65,7 +66,8 @@ function buildTriedSection(checkpoint: AutoresearchCheckpoint): string { const metricUnit = checkpoint.session.metricUnit; const renderedMetric = metricUnit && metricUnit.length > 0 ? `${run.metric}${metricUnit}` : `${run.metric}`; - return `- #${run.run} ${run.status} ${renderedMetric} ${run.commit} — ${run.description}`; + const baselineLabel = run.baseline ? " baseline" : ""; + return `- #${run.run}${baselineLabel} ${run.status} ${renderedMetric} ${run.commit} — ${run.description}`; }) .join("\n"); } @@ -76,8 +78,19 @@ function buildCheckpointSection(checkpoint: AutoresearchCheckpoint): string { `- Runs tracked: ${checkpoint.session.currentRunCount} current / ${checkpoint.session.totalRunCount} total`, `- Baseline: ${formatMetric(checkpoint.session.currentBaselineMetric, checkpoint.session.metricUnit)}`, `- Best kept: ${formatMetric(checkpoint.session.currentBestMetric, checkpoint.session.metricUnit)}`, + `- ${formatConfidenceLine(checkpoint.session.confidence)}`, ]; + if (checkpoint.canonicalBranch) { + lines.push(`- Canonical branch: ${checkpoint.canonicalBranch}`); + } + + if (checkpoint.carryForwardContext) { + lines.push( + `- Carry-forward best: ${checkpoint.carryForwardContext.metricName} ${formatMetric(checkpoint.carryForwardContext.run.metric, checkpoint.carryForwardContext.metricUnit)} from #${checkpoint.carryForwardContext.run.run} ${checkpoint.carryForwardContext.run.commit} — ${checkpoint.carryForwardContext.run.description}`, + ); + } + if (checkpoint.lastLoggedRun) { lines.push( `- Last logged run: #${checkpoint.lastLoggedRun.run} ${checkpoint.lastLoggedRun.status} ${checkpoint.lastLoggedRun.commit} — ${checkpoint.lastLoggedRun.description}`, diff --git a/extensions/openclaw-autoresearch/src/session-lock.ts b/extensions/openclaw-autoresearch/src/session-lock.ts new file mode 100644 index 0000000..04c0135 --- /dev/null +++ b/extensions/openclaw-autoresearch/src/session-lock.ts @@ -0,0 +1,96 @@ +import * as fs from "node:fs"; +import { getAutoresearchRootFilePath } from "./files.js"; + +export type AutoresearchSessionLock = { + readonly pid: number; + readonly timestamp: number; +}; + +export type AutoresearchSessionLockStatus = { + readonly state: "missing" | "active" | "stale"; + readonly pid: number | null; + readonly timestamp: number | null; + readonly ownedByCurrentProcess: boolean; +}; + +export function readAutoresearchSessionLock(cwd: string): AutoresearchSessionLock | null { + const lockPath = getAutoresearchRootFilePath(cwd, "sessionLock"); + if (!fs.existsSync(lockPath)) { + return null; + } + + try { + const parsed = JSON.parse(fs.readFileSync(lockPath, "utf8")) as Partial; + if (typeof parsed.pid !== "number" || typeof parsed.timestamp !== "number") { + return null; + } + return { + pid: parsed.pid, + timestamp: parsed.timestamp, + }; + } catch { + return null; + } +} + +export function getAutoresearchSessionLockStatus(cwd: string): AutoresearchSessionLockStatus { + const lock = readAutoresearchSessionLock(cwd); + if (!lock) { + return { + state: "missing", + pid: null, + timestamp: null, + ownedByCurrentProcess: false, + }; + } + + const active = isProcessAlive(lock.pid); + return { + state: active ? "active" : "stale", + pid: lock.pid, + timestamp: lock.timestamp, + ownedByCurrentProcess: lock.pid === process.pid, + }; +} + +export function acquireAutoresearchSessionLock(cwd: string): AutoresearchSessionLockStatus { + const status = getAutoresearchSessionLockStatus(cwd); + if (status.state === "active" && !status.ownedByCurrentProcess) { + return status; + } + + const nextLock: AutoresearchSessionLock = { + pid: process.pid, + timestamp: Date.now(), + }; + const lockPath = getAutoresearchRootFilePath(cwd, "sessionLock"); + fs.writeFileSync(lockPath, `${JSON.stringify(nextLock, null, 2)}\n`); + + return { + state: "active", + pid: nextLock.pid, + timestamp: nextLock.timestamp, + ownedByCurrentProcess: true, + }; +} + +export function removeAutoresearchSessionLock(cwd: string): void { + const lockPath = getAutoresearchRootFilePath(cwd, "sessionLock"); + if (fs.existsSync(lockPath)) { + fs.unlinkSync(lockPath); + } +} + +function isProcessAlive(pid: number): boolean { + if (!Number.isInteger(pid) || pid <= 0) { + return false; + } + + try { + process.kill(pid, 0); + return true; + } catch (error) { + const code = error && typeof error === "object" ? (error as NodeJS.ErrnoException).code : null; + return code === "EPERM"; + } +} diff --git a/extensions/openclaw-autoresearch/src/state.ts b/extensions/openclaw-autoresearch/src/state.ts index d31e958..a022707 100644 --- a/extensions/openclaw-autoresearch/src/state.ts +++ b/extensions/openclaw-autoresearch/src/state.ts @@ -1,4 +1,5 @@ import { readAutoresearchRootFile } from "./files.js"; +import { computeConfidence, type ConfidenceRun } from "./confidence.js"; export type SecondaryMetricDef = { readonly name: string; @@ -19,9 +20,11 @@ export type AutoresearchRunSnapshot = { readonly metric: number; readonly metrics: Record; readonly status: "keep" | "discard" | "crash"; + readonly baseline: boolean; readonly description: string; readonly timestamp: number; readonly segment: number; + readonly confidence: number | null; }; export type AutoresearchStateSnapshot = { @@ -35,6 +38,7 @@ export type AutoresearchStateSnapshot = { readonly totalRunCount: number; readonly currentBaselineMetric: number | null; readonly currentBestMetric: number | null; + readonly confidence: number | null; readonly lastRun: AutoresearchRunSnapshot | null; readonly mode: AutoresearchMode; readonly hasSessionDoc: boolean; @@ -51,6 +55,7 @@ type MutableStateSnapshot = { totalRunCount: number; currentBaselineMetric: number | null; currentBestMetric: number | null; + confidence: number | null; lastRun: AutoresearchRunSnapshot | null; mode: AutoresearchMode; hasSessionDoc: boolean; @@ -68,9 +73,11 @@ type JsonlEntry = { readonly metric?: number; readonly metrics?: Record; readonly status?: "keep" | "discard" | "crash"; + readonly baseline?: boolean; readonly description?: string; readonly timestamp?: number; readonly segment?: number; + readonly confidence?: number | null; }; export function createEmptyStateSnapshot(): AutoresearchStateSnapshot { @@ -85,6 +92,7 @@ export function createEmptyStateSnapshot(): AutoresearchStateSnapshot { totalRunCount: 0, currentBaselineMetric: null, currentBestMetric: null, + confidence: null, lastRun: null, mode: "inactive", hasSessionDoc: false, @@ -120,6 +128,7 @@ export function reconstructStateFromJsonl(cwd: string): AutoresearchStateSnapsho } const currentSecondaryMetrics = new Map(); + let currentSegmentRuns: ConfidenceRun[] = []; let currentRunIndex = 0; let hasSeenAnyRun = false; @@ -155,7 +164,9 @@ export function reconstructStateFromJsonl(cwd: string): AutoresearchStateSnapsho state.currentRunCount = 0; state.currentBaselineMetric = null; state.currentBestMetric = null; + state.confidence = null; currentRunIndex = 0; + currentSegmentRuns = []; currentSecondaryMetrics.clear(); continue; } @@ -168,6 +179,7 @@ export function reconstructStateFromJsonl(cwd: string): AutoresearchStateSnapsho currentRunIndex += 1; state.currentRunCount = currentRunIndex; state.totalRunCount += 1; + const isBaseline = entry.baseline === true || currentRunIndex === 1; const run: AutoresearchRunSnapshot = { run: typeof entry.run === "number" ? entry.run : currentRunIndex, @@ -175,9 +187,11 @@ export function reconstructStateFromJsonl(cwd: string): AutoresearchStateSnapsho metric: entry.metric, metrics: normalizeMetrics(entry.metrics), status: entry.status ?? "keep", + baseline: isBaseline, description: entry.description ?? "", timestamp: typeof entry.timestamp === "number" ? entry.timestamp : 0, segment: typeof entry.segment === "number" ? entry.segment : state.currentSegment, + confidence: typeof entry.confidence === "number" ? entry.confidence : null, }; if (state.currentBaselineMetric === null) { @@ -202,11 +216,16 @@ export function reconstructStateFromJsonl(cwd: string): AutoresearchStateSnapsho } } + currentSegmentRuns.push({ + metric: run.metric, + status: run.status, + }); state.lastRun = run; } return { ...state, + confidence: computeConfidence(currentSegmentRuns, state.bestDirection), secondaryMetrics: [...currentSecondaryMetrics.values()], }; } @@ -215,8 +234,36 @@ export function readRecentLoggedRuns( cwd: string, limit: number, ): readonly AutoresearchRunSnapshot[] { + const runs = readAllLoggedRuns(cwd); + return limit <= 0 ? [] : runs.slice(-limit); +} + +export function readBestLoggedRun( + cwd: string, + direction: "lower" | "higher", + segment?: number, +): AutoresearchRunSnapshot | null { + const runs = readAllLoggedRuns(cwd).filter((run) => + typeof segment === "number" ? run.segment === segment : true, + ); + if (runs.length === 0) { + return null; + } + + const keepRuns = runs.filter((run) => run.status === "keep"); + const candidates = keepRuns.length > 0 ? keepRuns : runs; + let best = candidates[0] ?? null; + for (const run of candidates.slice(1)) { + if (!best || isBetter(run.metric, best.metric, direction)) { + best = run; + } + } + return best; +} + +function readAllLoggedRuns(cwd: string): readonly AutoresearchRunSnapshot[] { const jsonl = readAutoresearchRootFile(cwd, "resultsLog"); - if (jsonl === null || limit <= 0) { + if (jsonl === null) { return []; } @@ -252,19 +299,22 @@ export function readRecentLoggedRuns( hasSeenAnyRun = true; currentRunIndex += 1; + const isBaseline = entry.baseline === true || currentRunIndex === 1; runs.push({ run: typeof entry.run === "number" ? entry.run : currentRunIndex, commit: entry.commit ?? "", metric: entry.metric, metrics: normalizeMetrics(entry.metrics), status: entry.status ?? "keep", + baseline: isBaseline, description: entry.description ?? "", timestamp: typeof entry.timestamp === "number" ? entry.timestamp : 0, segment: typeof entry.segment === "number" ? entry.segment : currentSegment, + confidence: typeof entry.confidence === "number" ? entry.confidence : null, }); } - return runs.slice(-limit); + return runs; } function normalizeMetrics(metrics: Record | undefined): Record { diff --git a/extensions/openclaw-autoresearch/src/tools/autoresearch-status.ts b/extensions/openclaw-autoresearch/src/tools/autoresearch-status.ts index 6792bbd..5dcd3fd 100644 --- a/extensions/openclaw-autoresearch/src/tools/autoresearch-status.ts +++ b/extensions/openclaw-autoresearch/src/tools/autoresearch-status.ts @@ -10,12 +10,19 @@ import { readAutoresearchCheckpoint, type AutoresearchCheckpoint, } from "../checkpoint.js"; -import { countCommitsSince, readShortHeadCommit } from "../git.js"; +import { countCommitsSince, readCurrentBranch, readShortHeadCommit } from "../git.js"; +import { formatConfidenceLine } from "../confidence.js"; +import { + getAutoresearchSessionLockStatus, + type AutoresearchSessionLockStatus, +} from "../session-lock.js"; export type AutoresearchStatusDiagnostics = { readonly warnings: readonly string[]; readonly checkpoint: AutoresearchCheckpoint | null; readonly gitHead: string | null; + readonly gitBranch: string | null; + readonly lock: AutoresearchSessionLockStatus; }; const AutoresearchStatusParams = Type.Object( @@ -83,6 +90,7 @@ export function formatAutoresearchStatusText( `Runs: ${state.currentRunCount} current / ${state.totalRunCount} total`, `Baseline: ${formatMetric(state.currentBaselineMetric, state.metricUnit)}`, `Best kept: ${formatMetric(state.currentBestMetric, state.metricUnit)}`, + formatConfidenceLine(state.confidence), ]; if (state.name) { @@ -114,6 +122,24 @@ export function formatAutoresearchStatusText( lines.push(`Git HEAD: ${diagnostics.gitHead}`); } + if (diagnostics?.gitBranch) { + lines.push(`Current branch: ${diagnostics.gitBranch}`); + } + + if (diagnostics?.checkpoint?.canonicalBranch) { + lines.push(`Canonical branch: ${diagnostics.checkpoint.canonicalBranch}`); + } + + if (diagnostics?.checkpoint?.carryForwardContext) { + lines.push( + `Carry-forward best: ${diagnostics.checkpoint.carryForwardContext.metricName} ${formatMetric(diagnostics.checkpoint.carryForwardContext.run.metric, diagnostics.checkpoint.carryForwardContext.metricUnit)} from ${diagnostics.checkpoint.carryForwardContext.run.commit}`, + ); + } + + if (diagnostics?.lock) { + lines.push(`Session lock: ${formatLockStatus(diagnostics.lock)}`); + } + if (diagnostics && diagnostics.warnings.length > 0) { lines.push("", "Warnings:"); for (const warning of diagnostics.warnings) { @@ -143,6 +169,11 @@ async function buildAutoresearchStatusDiagnostics( runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, cwd, }); + const gitBranch = await readCurrentBranch({ + runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, + cwd, + }); + const lock = getAutoresearchSessionLockStatus(cwd); const warnings: string[] = []; if (checkpoint?.pendingRun) { @@ -151,6 +182,26 @@ async function buildAutoresearchStatusDiagnostics( ); } + if (lock.state === "stale") { + warnings.push( + `Stale autoresearch.lock detected for PID ${lock.pid}. Remove it after confirming that process is gone.`, + ); + } else if (lock.state === "active" && !lock.ownedByCurrentProcess) { + warnings.push( + `Another live autoresearch loop appears active (PID ${lock.pid}, started ${new Date(lock.timestamp ?? 0).toISOString()}). Resume that loop instead of branching off a new one.`, + ); + } + + if ( + checkpoint?.canonicalBranch && + gitBranch && + checkpoint.canonicalBranch !== gitBranch + ) { + warnings.push( + `Branch drift: current branch is ${gitBranch}, but the canonical autoresearch branch is ${checkpoint.canonicalBranch}. Switch back before continuing the loop.`, + ); + } + const driftBase = state.lastRun?.commit && state.lastRun.commit.length > 0 ? state.lastRun.commit @@ -163,10 +214,17 @@ async function buildAutoresearchStatusDiagnostics( }); if (commitsAhead !== null && commitsAhead > 0) { + const branchLabel = gitBranch ? `Branch ${gitBranch}` : "Current HEAD"; + const shouldBlockPush = + (gitBranch ?? checkpoint?.canonicalBranch ?? "").startsWith("autoresearch/"); warnings.push( - state.lastRun - ? `${commitsAhead} commit${commitsAhead === 1 ? "" : "s"} since the last logged experiment (${state.lastRun.commit}).` - : `${commitsAhead} commit${commitsAhead === 1 ? "" : "s"} since init_experiment, but no experiment has been logged yet.`, + shouldBlockPush + ? state.lastRun + ? `UNLOGGED COMMITS: ${branchLabel} is ${commitsAhead} commit${commitsAhead === 1 ? "" : "s"} ahead of the last logged experiment (${state.lastRun.commit}). Do not push this branch until each commit is captured by run_experiment -> log_experiment.` + : `UNLOGGED COMMITS: ${branchLabel} is ${commitsAhead} commit${commitsAhead === 1 ? "" : "s"} ahead of init_experiment with no logged experiment yet. Do not push this branch until the baseline and follow-up runs are logged.` + : state.lastRun + ? `${commitsAhead} commit${commitsAhead === 1 ? "" : "s"} since the last logged experiment (${state.lastRun.commit}).` + : `${commitsAhead} commit${commitsAhead === 1 ? "" : "s"} since init_experiment, but no experiment has been logged yet.`, ); } } @@ -175,5 +233,22 @@ async function buildAutoresearchStatusDiagnostics( warnings, checkpoint, gitHead, + gitBranch, + lock, }; } + +function formatLockStatus(lock: AutoresearchSessionLockStatus): string { + if (lock.state === "missing") { + return "missing"; + } + + const owner = `pid ${lock.pid}`; + const timestamp = lock.timestamp ? new Date(lock.timestamp).toISOString() : "unknown time"; + if (lock.state === "stale") { + return `stale (${owner}, started ${timestamp})`; + } + return lock.ownedByCurrentProcess + ? `active (current process, ${owner}, started ${timestamp})` + : `active (${owner}, started ${timestamp})`; +} diff --git a/extensions/openclaw-autoresearch/src/tools/init-experiment.ts b/extensions/openclaw-autoresearch/src/tools/init-experiment.ts index 99592d1..dd631a8 100644 --- a/extensions/openclaw-autoresearch/src/tools/init-experiment.ts +++ b/extensions/openclaw-autoresearch/src/tools/init-experiment.ts @@ -3,15 +3,17 @@ import { InitExperimentParams } from "./schemas.js"; import { createConfigHeader, writeConfigHeader } from "../logging.js"; import { createEmptyStateSnapshot, + readBestLoggedRun, readRecentLoggedRuns, reconstructStateFromJsonl, type AutoresearchStateSnapshot, } from "../state.js"; import { readAutoresearchCheckpoint, writeAutoresearchCheckpoint } from "../checkpoint.js"; import { syncAutoresearchSessionDoc } from "../session-doc.js"; -import { readShortHeadCommit } from "../git.js"; +import { readCurrentBranch, readShortHeadCommit } from "../git.js"; import { setAutoresearchPendingRun, setAutoresearchRunInFlight } from "../runtime-state.js"; import { resolveToolCwd } from "./tool-cwd.js"; +import { acquireAutoresearchSessionLock } from "../session-lock.js"; export function createInitExperimentTool(api: OpenClawPluginApi) { return { @@ -28,14 +30,72 @@ export function createInitExperimentTool(api: OpenClawPluginApi) { metric_name: string; metric_unit?: string; direction?: "lower" | "higher"; + reset?: boolean; }, _signal: AbortSignal, _onUpdate: unknown, ) { const cwd = resolveToolCwd(api, params.cwd); + const lockStatus = acquireAutoresearchSessionLock(cwd); + if (lockStatus.state === "active" && !lockStatus.ownedByCurrentProcess) { + return { + content: [ + { + type: "text" as const, + text: + `Another autoresearch loop already holds ${"autoresearch.lock"}.\n` + + `Lock owner PID: ${lockStatus.pid}\n` + + `Started: ${new Date(lockStatus.timestamp ?? 0).toISOString()}\n` + + "Resume that loop instead of starting a parallel session, or remove the stale lock after confirming the process is gone.", + }, + ], + details: { + status: "error", + phase: "lock", + }, + }; + } + const previousState = reconstructStateFromJsonl(cwd); const previousCheckpoint = readAutoresearchCheckpoint(cwd); - const isReinit = previousState.currentRunCount > 0; + const hasLoggedRuns = previousState.totalRunCount > 0; + const isReinit = hasLoggedRuns && (params.reset ?? false); + + if (previousCheckpoint?.pendingRun) { + return { + content: [ + { + type: "text" as const, + text: + "A run_experiment result is still pending log_experiment.\n" + + `Pending command: ${previousCheckpoint.pendingRun.command}\n` + + "Log or clear that run before re-initializing the session.", + }, + ], + details: { + status: "error", + phase: "pending_log", + }, + }; + } + + if (hasLoggedRuns && !params.reset) { + return { + content: [ + { + type: "text" as const, + text: + `This session already has ${previousState.totalRunCount} logged experiment${previousState.totalRunCount === 1 ? "" : "s"} across ${previousState.currentSegment + 1} segment${previousState.currentSegment === 0 ? "" : "s"}.\n` + + "init_experiment now requires reset: true before starting a new segment so the reset is explicit and prior results stay comparable.", + }, + ], + details: { + status: "error", + phase: "reset_required", + }, + }; + } + const nextState: AutoresearchStateSnapshot = { ...createEmptyStateSnapshot(), name: params.name, @@ -80,10 +140,28 @@ export function createInitExperimentTool(api: OpenClawPluginApi) { runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, cwd, }); + const currentBranch = await readCurrentBranch({ + runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, + cwd, + }); + const previousBestRun = + isReinit + ? readBestLoggedRun(cwd, previousState.bestDirection, previousState.currentSegment) + : null; const checkpoint = writeAutoresearchCheckpoint({ cwd, state: nextPersistentState, sessionStartCommit: sessionStartCommit ?? previousCheckpoint?.sessionStartCommit ?? null, + canonicalBranch: previousCheckpoint?.canonicalBranch ?? currentBranch, + carryForwardContext: + previousBestRun && isReinit + ? { + metricName: previousState.metricName, + metricUnit: previousState.metricUnit, + bestDirection: previousState.bestDirection, + run: previousBestRun, + } + : previousCheckpoint?.carryForwardContext ?? null, recentLoggedRuns: readRecentLoggedRuns(cwd, 8), pendingRun: null, }); @@ -92,6 +170,10 @@ export function createInitExperimentTool(api: OpenClawPluginApi) { const reinitNote = isReinit ? " (re-initialized - previous results archived, new baseline needed)" : ""; + const carryForwardNote = + checkpoint.carryForwardContext + ? `\nCarry-forward context: best prior result was ${checkpoint.carryForwardContext.metricName}=${checkpoint.carryForwardContext.run.metric}${checkpoint.carryForwardContext.metricUnit} on ${checkpoint.carryForwardContext.run.commit}.` + : ""; return { content: [ @@ -100,7 +182,8 @@ export function createInitExperimentTool(api: OpenClawPluginApi) { text: `Experiment initialized: "${nextState.name}"${reinitNote}\n` + `Metric: ${nextState.metricName} (${nextState.metricUnit || "unitless"}, ${nextState.bestDirection} is better)\n` + - "Config written to autoresearch.jsonl. Now run the baseline with run_experiment, then log it before starting another run.", + "Config written to autoresearch.jsonl. Now run the baseline with run_experiment, then log it before starting another run." + + carryForwardNote, }, ], details: { diff --git a/extensions/openclaw-autoresearch/src/tools/log-experiment.ts b/extensions/openclaw-autoresearch/src/tools/log-experiment.ts index 54d0e03..568d598 100644 --- a/extensions/openclaw-autoresearch/src/tools/log-experiment.ts +++ b/extensions/openclaw-autoresearch/src/tools/log-experiment.ts @@ -1,9 +1,10 @@ import * as fs from "node:fs"; import type { OpenClawPluginApi } from "openclaw/plugin-sdk"; import { LogExperimentParams } from "./schemas.js"; -import { commitKeptExperiment, readShortHeadCommit } from "../git.js"; +import { commitKeptExperiment, readCurrentBranch, readShortHeadCommit } from "../git.js"; import { appendResultEntry, type AutoresearchResultEntry } from "../logging.js"; import { getAutoresearchRootFilePath } from "../files.js"; +import { appendIdeaBacklogEntry } from "../ideas.js"; import { readRecentLoggedRuns, reconstructStateFromJsonl, @@ -19,6 +20,8 @@ import { import { resolveToolCwd } from "./tool-cwd.js"; import { readAutoresearchCheckpoint, writeAutoresearchCheckpoint } from "../checkpoint.js"; import { syncAutoresearchSessionDoc } from "../session-doc.js"; +import { computeConfidence, formatConfidenceLine } from "../confidence.js"; +import { acquireAutoresearchSessionLock } from "../session-lock.js"; export function createLogExperimentTool(api: OpenClawPluginApi) { return { @@ -35,6 +38,7 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { metric?: number; status: "keep" | "discard" | "crash"; description: string; + idea?: string; metrics?: Record; force?: boolean; }, @@ -42,6 +46,26 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { _onUpdate: unknown, ) { const cwd = resolveToolCwd(api, params.cwd); + const lockStatus = acquireAutoresearchSessionLock(cwd); + if (lockStatus.state === "active" && !lockStatus.ownedByCurrentProcess) { + return { + content: [ + { + type: "text" as const, + text: + "Another autoresearch loop is already active for this repo.\n" + + `Lock owner PID: ${lockStatus.pid}\n` + + `Started: ${new Date(lockStatus.timestamp ?? 0).toISOString()}\n` + + "Resume that loop instead of logging results from a parallel session.", + }, + ], + details: { + status: "error", + phase: "lock", + }, + }; + } + const checkpoint = readAutoresearchCheckpoint(cwd); const pendingRun = getAutoresearchPendingRun(cwd) ?? checkpoint?.pendingRun ?? null; const state = reconstructStateFromJsonl(cwd); @@ -74,6 +98,23 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { }; } + if (params.status === "discard" && !(params.idea ?? "").trim()) { + return { + content: [ + { + type: "text" as const, + text: + "Discarded experiments must record what was learned.\n" + + 'Provide idea with a short note like "cache key was too coarse; retry with per-file invalidation". It will be appended to autoresearch.ideas.md automatically.', + }, + ], + details: { + status: "error", + phase: "idea_required", + }, + }; + } + if (state.secondaryMetrics.length > 0) { const validationError = validateSecondaryMetrics( state.secondaryMetrics, @@ -93,15 +134,18 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { const knownSecondaryMetrics = mergeSecondaryMetrics(state.secondaryMetrics, secondaryMetrics); const currentResults = readCurrentSegmentResults(cwd, state.currentSegment); + const isBaselineRun = currentResults.length === 0; const experiment: AutoresearchResultEntry = { run: state.currentRunCount + 1, commit: inferredCommit.slice(0, 7), metric: inferredMetric, metrics: secondaryMetrics, status: params.status, + baseline: isBaselineRun, description: params.description, timestamp: Date.now(), segment: state.currentSegment, + confidence: null, }; let finalExperiment = experiment; @@ -145,6 +189,21 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { }; } + finalExperiment = { + ...finalExperiment, + confidence: computeConfidence( + [ + ...currentResults, + { + metric: finalExperiment.metric, + metrics: finalExperiment.metrics, + status: finalExperiment.status, + }, + ], + state.bestDirection, + ), + }; + try { appendResultEntry(cwd, finalExperiment); } catch (error) { @@ -164,6 +223,12 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { }; } + let ideaAppended = false; + if ((params.idea ?? "").trim()) { + appendIdeaBacklogEntry(cwd, params.idea ?? ""); + ideaAppended = true; + } + const baselineMetric = currentResults.length > 0 ? currentResults[0].metric : experiment.metric; const baselineSecondaryMetrics = findBaselineSecondaryMetrics( @@ -174,10 +239,16 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { const queuedSteers = consumeAutoresearchSteers(cwd); consumeAutoresearchPendingRun(cwd); setAutoresearchRunInFlight(cwd, false); + const currentBranch = await readCurrentBranch({ + runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, + cwd, + }); const nextCheckpoint = writeAutoresearchCheckpoint({ cwd, state: nextState, sessionStartCommit: checkpoint?.sessionStartCommit ?? experiment.commit, + canonicalBranch: checkpoint?.canonicalBranch ?? currentBranch, + carryForwardContext: checkpoint?.carryForwardContext ?? null, recentLoggedRuns: readRecentLoggedRuns(cwd, 8), pendingRun: null, }); @@ -197,6 +268,9 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { knownSecondaryMetrics, queuedSteers, usedPendingRun: pendingRun !== null, + baseline: isBaselineRun, + ideaAppended, + confidence: finalExperiment.confidence, }), }, ], @@ -214,6 +288,7 @@ export function createLogExperimentTool(api: OpenClawPluginApi) { type CurrentSegmentResult = { readonly metric: number; readonly metrics: Record; + readonly status: "keep" | "discard" | "crash"; }; function validateSecondaryMetrics( @@ -303,6 +378,10 @@ function readCurrentSegmentResults(cwd: string, segment: number): CurrentSegment entry.metrics && typeof entry.metrics === "object" ? (entry.metrics as Record) : {}, + status: + entry.status === "keep" || entry.status === "discard" || entry.status === "crash" + ? entry.status + : "keep", }); } @@ -349,8 +428,13 @@ function buildResultText(options: { knownSecondaryMetrics: readonly SecondaryMetricDef[]; queuedSteers: readonly string[]; usedPendingRun: boolean; + baseline: boolean; + ideaAppended: boolean; + confidence: number | null; }): string { - let text = `Logged #${options.experiment.run}: ${options.experiment.status} - ${options.experiment.description}`; + let text = options.baseline + ? `Logged #${options.experiment.run}: baseline - ${options.experiment.description}` + : `Logged #${options.experiment.run}: ${options.experiment.status} - ${options.experiment.description}`; text += `\nBaseline ${options.state.metricName}: ${formatMetric(options.baselineMetric, options.state.metricUnit)}`; if (options.experiment.run > 1 && options.experiment.status === "keep" && options.experiment.metric > 0) { @@ -382,10 +466,20 @@ function buildResultText(options: { text += `\nSecondary: ${parts.join(" ")}`; } + if (options.confidence !== null) { + text += `\n${formatConfidenceLine(options.confidence)}`; + } + text += `\n(${options.totalRunCount} experiments in current segment)`; if (options.usedPendingRun) { text += "\nUsed the pending run_experiment result as the source of truth for commit/metric defaults."; } + if (options.baseline) { + text += "\nThis run established the baseline for the current segment."; + } + if (options.ideaAppended) { + text += "\nAdded a follow-up idea to autoresearch.ideas.md."; + } text += `\n${options.gitSummary}`; if (options.queuedSteers.length > 0) { diff --git a/extensions/openclaw-autoresearch/src/tools/run-experiment.ts b/extensions/openclaw-autoresearch/src/tools/run-experiment.ts index 377c96c..399cff7 100644 --- a/extensions/openclaw-autoresearch/src/tools/run-experiment.ts +++ b/extensions/openclaw-autoresearch/src/tools/run-experiment.ts @@ -8,10 +8,11 @@ import { } from "../runtime-state.js"; import { resolveToolCwd } from "./tool-cwd.js"; import { parseMetricLines } from "../metrics.js"; -import { readShortHeadCommit } from "../git.js"; +import { readCurrentBranch, readShortHeadCommit } from "../git.js"; import { readAutoresearchCheckpoint, writeAutoresearchCheckpoint } from "../checkpoint.js"; import { readRecentLoggedRuns, reconstructStateFromJsonl } from "../state.js"; import { syncAutoresearchSessionDoc } from "../session-doc.js"; +import { acquireAutoresearchSessionLock } from "../session-lock.js"; export function createRunExperimentTool(api: OpenClawPluginApi) { return { @@ -31,6 +32,26 @@ export function createRunExperimentTool(api: OpenClawPluginApi) { onUpdate: ((update: unknown) => void | Promise) | undefined, ) { const cwd = resolveToolCwd(api, params.cwd); + const lockStatus = acquireAutoresearchSessionLock(cwd); + if (lockStatus.state === "active" && !lockStatus.ownedByCurrentProcess) { + return { + content: [ + { + type: "text" as const, + text: + "Another autoresearch loop is already active for this repo.\n" + + `Lock owner PID: ${lockStatus.pid}\n` + + `Started: ${new Date(lockStatus.timestamp ?? 0).toISOString()}\n` + + "Resume that loop instead of running a parallel experiment session.", + }, + ], + details: { + status: "error", + phase: "lock", + }, + }; + } + const checkpoint = readAutoresearchCheckpoint(cwd); const existingPendingRun = getAutoresearchPendingRun(cwd) ?? checkpoint?.pendingRun ?? null; @@ -96,6 +117,10 @@ export function createRunExperimentTool(api: OpenClawPluginApi) { runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, cwd, }); + const currentBranch = await readCurrentBranch({ + runCommandWithTimeout: api.runtime.system.runCommandWithTimeout, + cwd, + }); const pendingRun = { command: params.command, commit: currentCommit, @@ -113,6 +138,8 @@ export function createRunExperimentTool(api: OpenClawPluginApi) { cwd, state, sessionStartCommit: checkpoint?.sessionStartCommit ?? currentCommit, + canonicalBranch: checkpoint?.canonicalBranch ?? currentBranch, + carryForwardContext: checkpoint?.carryForwardContext ?? null, recentLoggedRuns: readRecentLoggedRuns(cwd, 8), pendingRun, }); diff --git a/extensions/openclaw-autoresearch/src/tools/schemas.ts b/extensions/openclaw-autoresearch/src/tools/schemas.ts index fbdc2a9..bbcbe1d 100644 --- a/extensions/openclaw-autoresearch/src/tools/schemas.ts +++ b/extensions/openclaw-autoresearch/src/tools/schemas.ts @@ -29,6 +29,12 @@ export const InitExperimentParams = Type.Object({ enum: ["lower", "higher"], }), ), + reset: Type.Optional( + Type.Boolean({ + description: + "Required when starting a new segment after experiments have already been logged. Preserves prior history but makes the reset explicit.", + }), + ), }); export const RunExperimentParams = Type.Object({ @@ -64,6 +70,12 @@ export const LogExperimentParams = Type.Object({ description: Type.String({ description: "Short description of what this experiment tried.", }), + idea: Type.Optional( + Type.String({ + description: + "Required when status is discard. Summarize what you learned and what you would try differently; this is appended to autoresearch.ideas.md.", + }), + ), metrics: Type.Optional( Type.Record(Type.String(), Type.Number(), { description: diff --git a/openclaw.plugin.json b/openclaw.plugin.json index 8ce77d9..dc8c1d2 100644 --- a/openclaw.plugin.json +++ b/openclaw.plugin.json @@ -5,7 +5,7 @@ "skills": [ "./skills" ], - "version": "1.0.3", + "version": "1.0.5", "configSchema": { "type": "object", "additionalProperties": false, diff --git a/package-lock.json b/package-lock.json index fbdf935..b319e02 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "@gianfrancopiana/openclaw-autoresearch", - "version": "1.0.3", + "version": "1.0.5", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "@gianfrancopiana/openclaw-autoresearch", - "version": "1.0.3", + "version": "1.0.5", "license": "MIT", "dependencies": { "@sinclair/typebox": "0.34.48" diff --git a/package.json b/package.json index ae1d026..99829c6 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@gianfrancopiana/openclaw-autoresearch", - "version": "1.0.3", + "version": "1.0.5", "description": "Faithful OpenClaw port of pi-autoresearch.", "type": "module", "main": "./index.ts", diff --git a/skills/autoresearch-create/SKILL.md b/skills/autoresearch-create/SKILL.md index 54bf444..47f4e89 100644 --- a/skills/autoresearch-create/SKILL.md +++ b/skills/autoresearch-create/SKILL.md @@ -9,14 +9,14 @@ Autonomous experiment loop: try ideas, keep what works, discard what doesn't, ne ## Tools -- **`init_experiment`** — configure session (name, metric, unit, direction). Call again to re-initialize with a new baseline when the optimization target changes. +- **`init_experiment`** — configure session (name, metric, unit, direction). Once runs exist, calling it again requires `reset: true` to start a new segment explicitly. - **`run_experiment`** — runs the benchmark command, times it, captures output, parses `METRIC name=number` lines, and opens a pending run that must be logged before another run can start. -- **`log_experiment`** — records the pending run. `keep` auto-commits. `discard`/`crash` → `git checkout -- .` to revert. If the previous `run_experiment` captured the primary metric, `commit` and `metric` can be omitted and will default from the pending run. +- **`log_experiment`** — records the pending run. The first logged run in a segment becomes the baseline automatically. `keep` auto-commits. `discard`/`crash` → `git checkout -- .` to revert. `discard` also requires an `idea` note so the failed path gets appended to `autoresearch.ideas.md`. If the previous `run_experiment` captured the primary metric, `commit` and `metric` can be omitted and will default from the pending run. ## Setup 1. Ask (or infer): **Goal**, **Command**, **Metric** (+ direction), **Files in scope**, **Constraints**. -2. `git checkout -b autoresearch/-` +2. If `autoresearch.checkpoint.json` already names a canonical branch, reuse it. Otherwise create one with `git checkout -b autoresearch/-` and keep using that branch for the whole loop. 3. Read the source files. Understand the workload deeply before writing anything. 4. Write `autoresearch.md` and `autoresearch.sh` (see below). Commit both. 5. `init_experiment` → `run_experiment` baseline → `log_experiment` → start looping immediately. @@ -68,7 +68,10 @@ Bash script (`set -euo pipefail`) that: pre-checks fast (syntax errors in <1s), - **Crashes:** fix if trivial, otherwise log and move on. Don't over-invest. - **Think longer when stuck.** Re-read source files, study the profiling data, reason about what the CPU is actually doing. The best ideas come from deep understanding, not from trying random variations. - **Resuming:** if `autoresearch.md` exists, read it plus `autoresearch.checkpoint.json`, then continue looping. +- **Respect the canonical branch.** If the checkpoint names a canonical branch, switch back to it before resuming. Do not create a fresh autoresearch branch for every session. +- **Respect the lock file.** If `autoresearch.lock` exists and the PID is alive, another loop is active. Resume that loop instead of forking a second session. - **No raw benchmark exec:** during active autoresearch mode, benchmark/test commands should go through `run_experiment`, not raw `exec`/`bash`. +- **Reset explicitly.** If you need a new segment, call `init_experiment` with `reset: true`; do not silently wipe the current baseline. **NEVER STOP.** The user may be away for hours. Keep going until interrupted. diff --git a/test/command.test.ts b/test/command.test.ts index 13dcef8..79c6ed0 100644 --- a/test/command.test.ts +++ b/test/command.test.ts @@ -61,6 +61,7 @@ describe("buildAutoresearchCommandText", () => { expect(text).toContain("Mode: active"); expect(text).toContain("Runtime mode: auto"); expect(text).toContain("Baseline: 101ms"); + expect(text).toContain("Confidence: n/a"); }); it("registers a mode-aware /autoresearch command that primes resume instructions", () => { diff --git a/test/lifecycle.test.ts b/test/lifecycle.test.ts index 19ce7ff..e0a4b8f 100644 --- a/test/lifecycle.test.ts +++ b/test/lifecycle.test.ts @@ -49,6 +49,7 @@ async function initExperiment( metric_name: string; metric_unit?: string; direction?: "lower" | "higher"; + reset?: boolean; }, ) { return await createInitExperimentTool(createApi(cwd) as never).execute( @@ -82,6 +83,7 @@ async function logExperiment( metric?: number; status: "keep" | "discard" | "crash"; description: string; + idea?: string; metrics?: Record; force?: boolean; }, @@ -139,6 +141,44 @@ describe("experiment lifecycle tools", () => { ]); }); + it("requires reset: true before starting a new segment after logged runs exist", async () => { + const cwd = createTempDir("autoresearch-reset-required-"); + + await initExperiment(cwd, { + name: "Parser optimization", + metric_name: "total_ms", + metric_unit: "ms", + direction: "lower", + }); + fs.appendFileSync( + getAutoresearchRootFilePath(cwd, "resultsLog"), + `${JSON.stringify({ + run: 1, + commit: "abc1234", + metric: 120, + metrics: {}, + status: "keep", + baseline: true, + description: "baseline", + timestamp: 1700000000000, + segment: 0, + })}\n`, + ); + + const result = await initExperiment(cwd, { + name: "Parser optimization v2", + metric_name: "total_ms", + metric_unit: "ms", + direction: "lower", + }); + + expect(result.details).toMatchObject({ + status: "error", + phase: "reset_required", + }); + expect((result.content[0] as { text: string }).text).toContain("reset: true"); + }); + it("appends a new config header on re-init instead of overwriting prior history", async () => { const cwd = createTempDir("autoresearch-reinit-"); @@ -156,6 +196,7 @@ describe("experiment lifecycle tools", () => { metric: 120, metrics: {}, status: "keep", + baseline: true, description: "baseline", timestamp: 1700000000000, segment: 0, @@ -167,6 +208,7 @@ describe("experiment lifecycle tools", () => { metric_name: "total_ms", metric_unit: "ms", direction: "lower", + reset: true, }); expect(result.content[0]?.type).toBe("text"); @@ -185,6 +227,7 @@ describe("experiment lifecycle tools", () => { metric: 120, metrics: {}, status: "keep", + baseline: true, description: "baseline", timestamp: 1700000000000, segment: 0, @@ -204,6 +247,9 @@ describe("experiment lifecycle tools", () => { currentRunCount: 0, }, }); + expect(fs.readFileSync(path.join(cwd, "autoresearch.checkpoint.json"), "utf8")).toContain( + '"carryForwardContext"', + ); }); it("reports a successful run and emits a running update", async () => { @@ -311,6 +357,7 @@ describe("experiment lifecycle tools", () => { metric: 120, status: "discard", description: "baseline", + idea: "baseline is noisy; retry with warm cache isolation", metrics: { compile_ms: 15, }, @@ -326,8 +373,10 @@ describe("experiment lifecycle tools", () => { compile_ms: 15, }, status: "discard", + baseline: true, description: "baseline", segment: 0, + confidence: null, }); expect(result.details).toMatchObject({ status: "ok", @@ -345,9 +394,15 @@ describe("experiment lifecycle tools", () => { attempted: false, }, }); + expect((result.content[0] as { text: string }).text).toContain( + "This run established the baseline for the current segment.", + ); expect((result.content[0] as { text: string }).text).toContain( "Git: skipped commit (discard) - revert tracked changes yourself with git checkout -- .", ); + expect(fs.readFileSync(path.join(cwd, "autoresearch.ideas.md"), "utf8")).toContain( + "baseline is noisy; retry with warm cache isolation", + ); }); it("validates missing and newly added secondary metrics unless forced", async () => { @@ -358,6 +413,7 @@ describe("experiment lifecycle tools", () => { metric: 120, status: "discard", description: "baseline", + idea: "record compile_ms on every run before comparing changes", metrics: { compile_ms: 15, }, @@ -369,6 +425,7 @@ describe("experiment lifecycle tools", () => { metric: 110, status: "discard", description: "missing secondary metric", + idea: "keep compile_ms in the log for comparability", metrics: {}, }); @@ -385,6 +442,7 @@ describe("experiment lifecycle tools", () => { metric: 110, status: "discard", description: "add bundle size", + idea: "bundle_kb looks useful enough to track on future runs", metrics: { compile_ms: 12, bundle_kb: 7, @@ -404,6 +462,7 @@ describe("experiment lifecycle tools", () => { metric: 110, status: "discard", description: "add bundle size", + idea: "bundle_kb looks useful enough to track on future runs", metrics: { compile_ms: 12, bundle_kb: 7, @@ -464,6 +523,7 @@ describe("experiment lifecycle tools", () => { metric: 130, status: "discard", description: "discard regression", + idea: "regression suggests parser cache invalidation is too broad", }); expect(discardResult.details).toMatchObject({ @@ -508,6 +568,56 @@ describe("experiment lifecycle tools", () => { expect(statuses).toEqual(["keep", "discard", "crash"]); }); + it("computes and persists the upstream-style confidence score after enough runs", async () => { + const cwd = createTempDir("autoresearch-confidence-"); + await seedExperiment(cwd); + + await logExperiment(cwd, { + commit: "abc1234", + metric: 100, + status: "keep", + description: "baseline", + }); + await logExperiment(cwd, { + commit: "def5678", + metric: 99, + status: "discard", + description: "noise sample", + idea: "rerun without the optimization to confirm the noise floor", + }); + + const result = await logExperiment(cwd, { + commit: "fedcba9", + metric: 95, + status: "keep", + description: "confirmed improvement", + }); + + expect(result.details).toMatchObject({ + status: "ok", + experiment: { + confidence: 5, + }, + state: { + confidence: 5, + }, + }); + expect((result.content[0] as { text: string }).text).toContain( + "Confidence: 5.0x noise floor - improvement is likely real", + ); + + const rows = readJsonl(cwd).filter((entry) => entry.type !== "config"); + expect(rows[2]).toMatchObject({ + confidence: 5, + }); + expect(fs.readFileSync(path.join(cwd, "autoresearch.checkpoint.json"), "utf8")).toContain( + '"confidence": 5', + ); + expect(fs.readFileSync(path.join(cwd, "autoresearch.md"), "utf8")).toContain( + "Confidence: 5.0x noise floor - improvement is likely real", + ); + }); + it("keeps the experiment window open across run_experiment and surfaces queued steers in log_experiment", async () => { const cwd = createTempDir("autoresearch-queued-steers-"); await seedExperiment(cwd); @@ -528,6 +638,7 @@ describe("experiment lifecycle tools", () => { metric: 120, status: "discard", description: "baseline", + idea: "parser cache might still help if keyed on file contents", }); expect((result.content[0] as { text: string }).text).toContain( @@ -581,6 +692,7 @@ describe("experiment lifecycle tools", () => { const result = await logExperiment(cwd, { status: "discard", description: "baseline", + idea: "compile_ms should stay in the log for future comparisons", }); expect(result.details).toMatchObject({ diff --git a/test/state.test.ts b/test/state.test.ts index 6d92702..0631813 100644 --- a/test/state.test.ts +++ b/test/state.test.ts @@ -16,6 +16,7 @@ describe("reconstructStateFromJsonl", () => { expect(state.totalRunCount).toBe(4); expect(state.currentBaselineMetric).toBe(130); expect(state.currentBestMetric).toBe(118); + expect(state.confidence).toBeNull(); expect(state.lastRun).toMatchObject({ run: 2, commit: "7654321", @@ -23,6 +24,7 @@ describe("reconstructStateFromJsonl", () => { status: "keep", description: "keep winner", segment: 1, + confidence: null, }); expect(state.secondaryMetrics).toEqual([ { name: "compile_ms", unit: "ms" }, diff --git a/test/status.test.ts b/test/status.test.ts index 2904c30..f93b931 100644 --- a/test/status.test.ts +++ b/test/status.test.ts @@ -15,6 +15,13 @@ describe("formatAutoresearchStatusText", () => { warnings: ["2 commits since the last logged experiment (7654321)."], checkpoint: null, gitHead: "89abcde", + gitBranch: "autoresearch/parser-cache", + lock: { + state: "missing", + pid: null, + timestamp: null, + ownedByCurrentProcess: false, + }, }; const text = formatAutoresearchStatusText( reconstructStateFromJsonl(activeSessionFixture), @@ -31,8 +38,11 @@ describe("formatAutoresearchStatusText", () => { expect(text).toContain("Metric: total_ms (ms, lower is better)"); expect(text).toContain("Runs: 2 current / 4 total"); expect(text).toContain("Best kept: 118ms"); + expect(text).toContain("Confidence: n/a"); expect(text).toContain("Last run: #2 keep 118ms 7654321 keep winner"); expect(text).toContain("Git HEAD: 89abcde"); + expect(text).toContain("Current branch: autoresearch/parser-cache"); + expect(text).toContain("Session lock: missing"); expect(text).toContain("Warnings:"); expect(text).toContain("2 commits since the last logged experiment"); expect(text).toContain( diff --git a/test/tool-execute.test.ts b/test/tool-execute.test.ts index da71ea4..b664aad 100644 --- a/test/tool-execute.test.ts +++ b/test/tool-execute.test.ts @@ -115,6 +115,7 @@ describe("autoresearch tools", () => { metric: 5, status: "discard", description: "baseline", + idea: "keep a discard note so follow-up experiments are not lost", }, new AbortController().signal, undefined, @@ -122,7 +123,7 @@ describe("autoresearch tools", () => { expect(api.resolvePath).toHaveBeenCalledWith("."); expect(result.details).toMatchObject({ status: "ok" }); - expect(result.content[0]?.text).toContain("Logged #1: discard - baseline"); + expect(result.content[0]?.text).toContain("Logged #1: baseline - baseline"); }); it("supports explicit cwd overrides for nested repo tool execution", async () => { @@ -179,12 +180,13 @@ describe("autoresearch tools", () => { metric: 5, status: "discard", description: "baseline", + idea: "keep a discard note so follow-up experiments are not lost", }, new AbortController().signal, undefined, ); expect(logResult.details).toMatchObject({ status: "ok" }); - expect(logResult.content[0]?.text).toContain("Logged #1: discard - baseline"); + expect(logResult.content[0]?.text).toContain("Logged #1: baseline - baseline"); }); });