Repository navigation
fix: observers and benchmark harnesses survive transient failures and report real paths - #1194
Conversation
…e launches The adapter ended a benchmark attempt on the first `ultrafuzz status` call that exited non-zero, timed out at the five-minute call limit, or printed no JSON, although the run it watches is detached and keeps going. It also aborted with "unknown status verdict" on `launch-incomplete`, which `status` reports for a run whose launch preparation has not finished publishing. A failed status call is now retried after 2, 4, 8 and 16 poll intervals; the fifth consecutive failure ends the attempt with the last failure's message. This applies to both status reads: the one that decides whether an existing run needs `resume`, and the poll loop. A contract-invalid envelope still fails at once, because retrying cannot fix a CLI and adapter that disagree. `launch-incomplete` is now a verdict the adapter waits on. Co-Authored-By: Claude Opus 5.5 <[email protected]>
`ultrafuzz dashboard` picks the newest run by reading every run
directory's run.json, and one unreadable entry threw out of that scan. A
launch that fails before run.json is written leaves exactly such a
directory, so the dashboard could not start at all ("cannot open regular
file .../run.json: ENOENT") until someone deleted it by hand.
The scan now skips an entry whose name is not a run ID or whose run.json
cannot be read, and picks the newest readable run. `listRuns` already
tolerates the same directories (it lists them as unreadable).
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…s failing Two ways one eval row could cost the whole suite: - A row whose run evidence could not be read threw out of the watch: for example an invalid graph.json, read when the row is recorded (recovery classification and expansion), or a foreign state.json. The rejection went up through Promise.all, so runEvalSuite rejected before run-summary.json was written, and the other rows' watches kept running unawaited. A failed watch is now caught per row: the row is appended to runs.jsonl with an EVAL_ROW_WATCH_FAILED error and stays launched with no observed outcome, so it counts as incomplete, and the summary is written for every row. `eval run` still exits non-zero because the row is incomplete. - A sync that failed identically on every poll (a deterministic WORKFLOW_CONTROL_EVIDENCE_INVALID, say) was retried until the six-hour watch deadline. The watch now stops after ten identical consecutive failures and records EVAL_ROW_SYNC_ABANDONED next to the existing EVAL_ROW_SYNC_FAILED warning. A different failure, or a success, restarts the count. The detached run is not cancelled. The watch-deadline suite test depended on the default syncRun failing the same way against its fixture until the one-second deadline; it now injects a sync that succeeds, so it still tests the deadline. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ydration Hydration swaps each pinned submodule root through a transaction directory inside the task worktree. A controller killed mid-hydration (SIGKILL, OOM, reboot) left that directory behind, and every later hydration and verification of the worktree then failed with "stale pinned submodule transaction is present". Only deleting the directory by hand recovered it, and a test pinned that dead end. Hydration now rolls a stale transaction back first: each root it had moved aside under backup/ goes back into place, the entry is removed, and the hydration proceeds from the sealed snapshot as usual. The entry lives where the agent can write, so a backup is put back only when it is a physical directory under the worktree; anything else, including an entry that is a file rather than a directory, is removed without being restored, since hydration replaces every root from the sealed snapshot anyway. The same path also completes a transaction that an earlier hydration kept on purpose after its own rollback failed. The post-agent verification is unchanged: a transaction entry found after the agent ran still fails the attempt, and the retry's hydration clears it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Runner text shown to an operator (process diagnostics, status reasons, `why`, events, attempt errors, inspection failures) had every "smithers" replaced with "workflow runner", including inside paths. A diagnostic about <target>/.smithers/workflows/ultrafuzz-<id>.tsx therefore pointed at <target>/.workflow runner/workflows/..., a directory that does not exist. The five copies of that replacement (smithers.ts, lifecycle-inspection.ts, state-export.ts twice, workflow-sync.ts) now share one scrubWorkflowRunnerText. It leaves alone any whitespace- or quote-delimited word that contains a path separator, file:// URLs included; http(s) URLs and prose are scrubbed as before. Lifecycle inspection's rename of the runner's own `smithers <command>` suggestions no longer fires after a path separator or a dot, so ".../bin/smithers ENOENT" stays as written. Co-Authored-By: Claude Opus 5.5 <[email protected]>
- evmbench: stop waiting on `launch-incomplete`. The poll loop cannot see it: `run` exits 0 only after startRun has sealed the run and set it `running`, and a successful `resume` sets `running` too. A half-launched existing run still fails at `resume`, as before, instead of stalling until the workflow timeout. The retry test keeps the fail/success/fail sequence that checks a success restarts the backoff. - pinned submodules: the next hydration now just removes a stale transaction entry. Restoring its backup bought nothing, since hydration replaces every root from the sealed snapshot and already handles a missing root, and it moved agent-writable bytes back into the worktree. The Aave-shaped test also leaves a symlink to an outside directory and checks that only the link goes. - scrub: a word holding a dot before a letter or digit is kept whole, like a word holding a path separator, and web URLs are no longer special cased. `.smithers`, `smithers.db` and the runner's docs link stay real instead of becoming `.workflow runner`, `workflow runner.db` and a dead `https://workflow runner.sh/...` link. A sentence's closing dot does not count, so prose is still scrubbed. - evals: state the identical-failure rule as the policy it is. Co-Authored-By: Claude Opus 5.5 <[email protected]>
| const reason = error instanceof Error ? error.message : String(error); | ||
| throw new Error(`Ultrafuzz status failed ${String(failures)} consecutive times: ${reason}`, { cause: error }); | ||
| } | ||
| await wait(pollIntervalMs * 2 ** failures); |
There was a problem hiding this comment.
Retries overrun workflow deadline When
status starts failing near the workflow deadline, these waits continue without checking it, and each status call can take up to five minutes. The adapter can run past the configured timeout and even accept a run that finishes late if a later status call succeeds. Check the deadline during retries, not only after a successful poll.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/evmbench/src/adapter.ts
Line: 268
Comment:
**Retries overrun workflow deadline** When `status` starts failing near the workflow deadline, these waits continue without checking it, and each status call can take up to five minutes. The adapter can run past the configured timeout and even accept a run that finishes late if a later status call succeeds. Check the deadline during retries, not only after a successful poll.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| return [{ runId, createdAt: metadata.created_at, mtimeMs: fs.statSync(root).mtimeMs }]; | ||
| } catch { | ||
| return []; |
There was a problem hiding this comment.
Skipped runs go unexplained This catch silently skips a corrupt or temporarily unreadable
run.json as well as a directory left by a failed launch. Because the dashboard caches its startup choice, an operator can remain on an older run for the whole session without knowing a newer one was skipped. Report the skipped-run reason while keeping the fallback.
Knowledge Base Used: Dashboard and observability
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/dashboard/src/index.ts
Line: 576-578
Comment:
**Skipped runs go unexplained** This catch silently skips a corrupt or temporarily unreadable `run.json` as well as a directory left by a failed launch. Because the dashboard caches its startup choice, an operator can remain on an older run for the whole session without knowing a newer one was skipped. Report the skipped-run reason while keeping the fallback.
**Knowledge Base Used:** [Dashboard and observability](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/dashboard-observability.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| return scrubWorkflowRunnerText( | ||
| redactSecretsInText(value).replace( | ||
| // Not after a path separator or dot: `.../bin/smithers ENOENT` names a file, not a runner command. | ||
| /(`?)\b(?<![\\/.])smithers\s+([a-z][a-z-]*)(`?)/giu, |
There was a problem hiding this comment.
Command rewrite corrupts paths The lookbehind protects
smithers only when /, \, or . is immediately before it. In a diagnostic such as spawn /opt/runner/pre-smithers run ENOENT, the match starts after - and changes the executable path to /opt/runner/pre-ultrafuzz run`` before the path-preserving scrub runs. This leaves an operator with a path that does not exist; restrict command rewriting to standalone command tokens.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/lifecycle-inspection.ts
Line: 1940
Comment:
**Command rewrite corrupts paths** The lookbehind protects `smithers` only when `/`, `\`, or `.` is immediately before it. In a diagnostic such as `spawn /opt/runner/pre-smithers run ENOENT`, the match starts after `-` and changes the executable path to `/opt/runner/pre-`ultrafuzz run`` before the path-preserving scrub runs. This leaves an operator with a path that does not exist; restrict command rewriting to standalone command tokens.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.…gnostics publicWorkflowText renames the runner's own `smithers <command>` suggestions before the path-preserving scrub runs. #1194 stopped the rename from firing after a path separator or a dot, but it still fired after any other non-word character. A path whose last segment ends in "smithers" after a hyphen (or @, +, ~, ...) and is followed by a space and a word was rewritten before the scrub could keep it whole: "spawn /opt/runner/pre-smithers ENOENT" became "spawn /opt/runner/pre-workflow runner ENOENT", and "Command failed: /opt/tools/pinned-smithers why --run-id r1" became "Command failed: /opt/tools/pinned-`ultrafuzz why` --run-id r1". Exposure is narrow. On Linux the runner is executed through /proc/<pid>/fd paths, so a failed runner query does not name the runner's own path (a probe with a failing runner at <dir>/pinned-smithers reported "Command failed: /proc/<pid>/fd/29 /proc/<pid>/fd/28 why ..."), and a missing runner is reported as "ENOENT: ... open '<path>'", which the rename leaves alone. The rewrite still applies to any other text that reaches publicWorkflowText: runner stderr, `why` prose, event labels and details, and node errors. The rename now fires only where "smithers" starts a word: at the start of the text, or after whitespace, a quote, a backtick, or an opening parenthesis or bracket. The scrub still replaces the name in prose. Over the 1,064 string literals that mention "smithers" in packages/*/src and packages/*/test, main's and the new publicWorkflowText differ only on this change's own new pre-smithers strings. Over the 1,664 distinct lines that mention `smithers <word>` in the pinned Smithers 0.35.0 packages they differ on one, an HTML <title> template in @smthrs/cli's runReport.js. The diagnoseRun fixture now also carries a hint after whitespace ("then run smithers inspect"), the form missing from the tests, so narrowing the lookbehind to the start of the text and backticks fails. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…gnostics publicWorkflowText renames the runner's own `smithers <command>` suggestions before the path-preserving scrub runs. #1194 stopped the rename from firing after a path separator or a dot, but it still fired after any other non-word character. A path whose last segment ends in "smithers" after a hyphen (or @, +, ~, ...) and is followed by a space and a word was rewritten before the scrub could keep it whole: "spawn /opt/runner/pre-smithers ENOENT" became "spawn /opt/runner/pre-workflow runner ENOENT", and "Command failed: /opt/tools/pinned-smithers why --run-id r1" became "Command failed: /opt/tools/pinned-`ultrafuzz why` --run-id r1". Exposure is narrow. On Linux the runner is executed through /proc/<pid>/fd paths, so a failed runner query does not name the runner's own path (a probe with a failing runner at <dir>/pinned-smithers reported "Command failed: /proc/<pid>/fd/29 /proc/<pid>/fd/28 why ..."), and a missing runner is reported as "ENOENT: ... open '<path>'", which the rename leaves alone. The rewrite still applies to any other text that reaches publicWorkflowText: runner stderr, `why` prose, event labels and details, and node errors. The rename now fires only where "smithers" starts a word: at the start of the text, or after whitespace, a quote, a backtick, or an opening parenthesis or bracket. The scrub still replaces the name in prose. Over the 1,064 string literals that mention "smithers" in packages/*/src and packages/*/test, main's and the new publicWorkflowText differ only on this change's own new pre-smithers strings. Over the 1,664 distinct lines that mention `smithers <word>` in the pinned Smithers 0.35.0 packages they differ on one, an HTML <title> template in @smthrs/cli's runReport.js. The diagnoseRun fixture now also carries a hint after whitespace ("then run smithers inspect"), the form missing from the tests, so narrowing the lookbehind to the start of the text and backticks fails. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem
In five places, an observer or harness stopped, or sent an operator to a path that does not exist, for reasons unrelated to the health of the run it watches. All five reproduce on
origin/integration/wave1(a48ad9c3). Item 4 was reproduced from the on-disk state such a crash leaves, not by killing a controller.packages/evmbench/src/adapter.ts). The firstultrafuzz statuscall that exited non-zero, hit the five-minute call limit, or printed no JSON ended the benchmark attempt. The Ultrafuzz run it watches is detached and keeps going.packages/dashboard/src/index.ts,latestRunId). A launch that fails before writingrun.jsonleaves its run directory behind. After that,ultrafuzz dashboardcould not start:cannot open regular file …/run.json: ENOENT.packages/evals/src/runner.ts).graph.json, a foreignstate.json),runEvalSuiterejected beforerun-summary.jsonwas written. This is the item refactor(evals): delete the dead reporter/telemetry pipeline and publication leftovers #1174 left open. Because the rows run underPromise.all, the other rows' watches also kept running, unawaited (by code reading).packages/runtime/src/pinned-submodules.ts). A controller killed during submodule hydration (SIGKILL, OOM, reboot) left a.ultrafuzz-submodule-transaction-*directory in the task worktree. Every later hydration and verification of that worktree then failed withstale pinned submodule transaction is present, until someone deleted the directory by hand. A test pinned that dead end.smithersin operator-facing runner text becameworkflow runner, including inside paths, file names and links. A diagnostic about<target>/.smithers/workflows/ultrafuzz-<id>.tsxtherefore pointed at<target>/.workflow runner/workflows/…, which does not exist. The runner's ownSee https://smithers.sh/reference/errorsbecame the dead linkhttps://workflow runner.sh/reference/errors.Root cause
commandData("status", execute(…))directly, so any throw fromexecuteended the attempt.latestRunIdread every run directory with a strictreadRunMetadataDocumentand no per-entry guard.listRunsalready guards the same read.runEvalSuiteawaitedwatchEvalRowwithout a per-row catch.evalRunExpansion→readStaticNodeIds, andclassifyRecoveryEquivalence. One row's bad evidence therefore rejected the whole suite.assertNoStaleTransactions), and nothing ever removed an entry left by a crash..replace(/smithers/giu, "workflow runner")treated the whole text as prose:smithers.ts,lifecycle-inspection.ts,state-export.ts(twice) andworkflow-sync.ts.Change
There is one commit per item, then one commit that applies the review. Source is +160/−57 across 8 files.
fix(evmbench)statuscall is retried after 2, 4, 8 and 16 poll intervals: 30 s, 1, 2 and 4 min at the smoke profile's 15 s interval.Ultrafuzz status failed 5 consecutive times: <last failure>.resume, and the poll loop.fix(dashboard):latestRunIdskips any entry whose name is not a run ID or whoserun.jsoncannot be read, and picks the newest readable run.fix(evals)runs.jsonlwith anEVAL_ROW_WATCH_FAILEDerror;status: "launched"with no observed outcome, so it counts as incomplete;run-summary.jsonis written for every row, andeval runstill exits non-zero (by code reading ofcommands/eval/run.ts).EVAL_ROW_SYNC_ABANDONEDnext to the existingEVAL_ROW_SYNC_FAILEDwarning.docs/reference/evals.mddescribes both.fix(runtime), pinned submodules. Hydration first removes every stale transaction entry, whatever its shape. Nothing in the entry is read or restored. Hydration then runs from the sealed snapshot, unchanged: it replaces every root, and a root that the crash left missing is hydrated like any other (replaceTaskRootsTransactionallyalready handles a missing root).Post-agent verification is also unchanged. An entry found after the agent ran still fails the attempt, and the retry's hydration clears it.
fix(runtime), scrub.scrubWorkflowRunnerTextreplaces the five copies./or\, or a dot before a letter or digit. That covers paths,file://and web URLs, and names such as.smithersandsmithers.db.restart smithers.is still scrubbed. Prose is scrubbed as before.publicWorkflowTextrenames the runner's ownsmithers <command>suggestions. That rename no longer fires after/,\or., so…/bin/smithers ENOENTstays as written.Deliberately not built
launch-incomplete. The poll loop cannot see that verdict. It needs run statependingplus a missing control seal or link journal, andultrafuzz runexits 0 only afterstartRunhas written both and set the runrunning. A successfulresumealso setsrunning.resume, as on the base:resumeRunon a planned, unsealed run returnsWORKFLOW_LIFECYCLE_FAILED: ENOENT … lstat '<runs>/<id>/smithers'. Nothing else would finish that launch, so waiting would only turn the error into a stall untilworkflow_timeout_seconds.status: "failed". That field means the launch failed, and downstream a failed row reads as a run that never reached a model, i.e. a relaunch candidate.final_status: "failed"either, because nobody observed that outcome.readStaticNodeIds,classifyRecoveryEquivalence) are untouched. Only the damage is contained, to the one row.listRunsalready lists such entries as unreadable.preserveTransaction,PinnedSubmoduleRollbackError). The backup can be inspected there until the next hydration of that worktree removes it. Dropping that too would change no run outcome, and is left for a follow-up.publicHealthReason(smithers why,smithers supervise -r …) are unchanged.Verification
Every behavioural test below fails on
origin/integration/wave1and passes on this branch (heade6d6f50f). For base runs, the base versions of the changed source files were checked out under the branch's tests, and the tree was confirmed clean after restoring them.adapter.test.ts"keeps waiting through failed status calls". A success between two failure streaks restarts the backoff at two intervals.Ultrafuzz command status failed with exit code 1at the first failed poll.failed to invoke Ultrafuzz: spawnSync node ETIMEDOUTafter one call.dashboard.test.ts"starts on the latest readable run when a half-launched run directory has no run.json"cannot open regular file …/half-launched/run.json: ENOENTrunner-publish.test.ts"publishes the run summary when a watched row's run evidence cannot be read"runEvalSuiterejects:planned graph is schema-invalid: / must have required property 'graph_version'; …expected 40 to be 20: the watch kept syncing until the fixture's 40th call made the run terminal.pinned-submodules.test.ts"the next hydration clears a transaction an interrupted one left behind" (replaces "a failed immediate submodule restore preserves the transaction backup")stale pinned submodule transaction is present.stale pinned submodule transaction is present. A cleanup that follows the symlink fails withENOENT … outside-the-task-worktree/kept.txt.smithers-diagnostic.test.ts"workflow process diagnostics drop the runner name from prose but keep paths, file names and URLs"…/.workflow runner/workflows/…,file:///work/target/.workflow runner/agents/index.ts,workflow runner.db,.workflow runnerandhttps://workflow runner.sh/reference/errors. A rule that keeps every word holding a dot also fails it, onrestart smithers..lifecycle-inspection.test.ts"diagnoseRun keeps a runner path intact in public text"/work/target/.workflow runner/workflows/…. Without the lookbehind:spawn /opt/runner/bin/workflow runner ENOENT.One existing test changed: "counts a watched row as incomplete when it misses the watch deadline". Against its runner-less fixture, the default
syncRunfails identically on every poll, so the watch would now stop after ten failures, before the one-second deadline the test is about. It now injects a sync that succeeds.Suites run at this head (
e6d6f50f):runner-publish.test.ts32/32. The full suite, 386/386 in 20 files, ran atd6dcc66c; this head changes only a comment inpackages/evals.pinned-submodules.test.ts: 5/5.smithers-diagnostic.test.ts: 3/3.dashboard.test.ts: 30/30. The dashboard is unchanged sinced6dcc66c, where all three of its files ran 39/39.lifecycle-inspection.test.ts: 52/52, over two runs. A 30-minute cap stopped the first after 24 tests at load ~33, and the other 28 ran in a second. This includes every test that asserts public inspection text carries no runner name (assertNoEngineBranding).runtime.test.ts, the eight tests that assert on runner-derived text: 8/8.doesNotMatch(/smithers/iu);workflow execution file dependencies/packages/000126/license changed while reading, before any runner text was produced. That reader compares link count and ctime, which change when other worktrees on the machine hard-link the same pnpm store files. It passed alone on the same build.cli.test.ts: 8/8. These are "init and validate …", "run, ps, status, inspect, report, materialize, clean, and lifecycle commands expose product workflow evidence", "status surfaces a terminal product and live workflow lifecycle divergence", "runtime command failures …", "run rejects an OpenRouter override …", both "report bundle …" tests, and "product surface checks …".lifecycle-commands.test.ts: 12/12, including "doctor reports install posture in human and JSON output".https://workflow runner.sh/…link, and now asserts the real one. The tests above that assert operator output never names the runner all pass.Gates, at this head:
npx prettier --checkon all 15 changed files.pnpm -w lint.CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/integration/wave1 pnpm -w lint:strict:ci.pnpm --filter @ultrafuzz/{evmbench,evals,runtime} typecheck, and the dashboard's twotsc --noEmitpasses.synchronizeLinkedWorkflowRun), so the ceiling stays.pnpm -w knip.pnpm -w docs:check.Risk / compatibility
statuscall that keeps failing now ends the attempt only after 30 poll intervals of backoff, not at once: 7.5 min on the smoke profile and 15 min on full. If every call also hits the five-minute limit, that adds up to 25 min more. The retries do not check the workflow deadline, so they can overrun it by that much.run.jsonis corrupt, the dashboard now opens the newest readable run, or the preview, instead of refusing to start, and does not say it skipped one.EVAL_ROW_WATCH_FAILEDandEVAL_ROW_SYNC_ABANDONED. The schema is unchanged, since durable diagnostics accept any non-empty code.fs.rmSyncunlinks a symlinked entry, and any symlink inside a directory entry, without following it. The Aave-shaped test covers the symlinked entry under Node. A probe covers both cases under Bun 1.3.14, which runs the controller.backup/is deleted, not restored. It was either a fresh worktree's empty submodule directory or an earlier attempt's hydrated tree. The sealed snapshot replaces either one.smithersinside paths, file names and URLs, for example.smithers/workflows/…,node_modules/.bin/smithers,smithers.dborhttps://smithers.sh/reference/errors. Anything that greps operator output for the runner name will now find it there.SMITHERS_BIN=/usr/bin/smitherskeeps both names, and a package-like word such assmithers/clicounts as a path. So does a word with an inner dot, such as a version spec like[email protected].smithers.ts: the scrub helper only;workflow-sync.ts: one import and one line;state-export.ts: one import and two lines;lifecycle-inspection.ts: one import andpublicWorkflowText.Changelog entry
Observers and benchmark harnesses no longer stop on transient failures: the EVMBench adapter retries failed
ultrafuzz statuscalls (giving up after five in a row),ultrafuzz dashboardstarts even when a run directory has no readablerun.json, andeval runwritesrun-summary.jsoneven when one row's evidence cannot be read and stops watching a row after ten identical sync failures instead of polling for six hours. A crash during pinned-submodule hydration no longer bricks the task worktree. Operator diagnostics keep real.smithers/…paths, file names and links instead of rewriting them to.workflow runner/….🤖 Generated with Claude Code
The PR should not merge until status retries respect the benchmark workflow deadline.
Fix with agent prompt
Summary
The PR makes benchmark polling, dashboard startup, and evaluation watches more tolerant of failures; clears abandoned submodule transactions; and preserves real paths in runner diagnostics.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR Run[Detached run] --> Status[EVMBench status polling] Status --> Retry[Retry failed calls] Retry --> Deadline[Workflow deadline] Run --> Evidence[Saved run evidence] Evidence --> Dashboard[Dashboard selects readable run] Evidence --> Eval[Eval watches rows and writes summary] Worktree[Task worktree] --> Hydrate[Clear abandoned transaction and hydrate] Runner[Runner diagnostics] --> Scrub[Preserve paths in operator text]Reviews (1) · Last reviewed commit: "fix: apply review of the observer and ha..."