Skip to content

fix: observers and benchmark harnesses survive transient failures and report real paths - #1194

Merged
aviggiano merged 6 commits into
mainfrom
claude/v06-observer-harness-robustness
Sep 29, 2026
Merged

aviggiano merged 6 commits into
mainfrom
claude/v06-observer-harness-robustness

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

In five places, an observer or harness stopped, or sent an operator to a path that does not exist, for reasons unrelated to the health of the run it watches. All five reproduce on origin/integration/wave1 (a48ad9c3). Item 4 was reproduced from the on-disk state such a crash leaves, not by killing a controller.

  1. EVMBench adapter (packages/evmbench/src/adapter.ts). The first ultrafuzz status call that exited non-zero, hit the five-minute call limit, or printed no JSON ended the benchmark attempt. The Ultrafuzz run it watches is detached and keeps going.
  2. Dashboard (packages/dashboard/src/index.ts, latestRunId). A launch that fails before writing run.json leaves its run directory behind. After that, ultrafuzz dashboard could not start: cannot open regular file …/run.json: ENOENT.
  3. Evals (packages/evals/src/runner.ts).
    • When a watched row's run evidence could not be read (an invalid graph.json, a foreign state.json), runEvalSuite rejected before run-summary.json was written. This is the item refactor(evals): delete the dead reporter/telemetry pipeline and publication leftovers #1174 left open. Because the rows run under Promise.all, the other rows' watches also kept running, unawaited (by code reading).
    • A sync that failed the same way on every poll was retried until the six-hour watch deadline.
  4. Pinned submodules (packages/runtime/src/pinned-submodules.ts). A controller killed during submodule hydration (SIGKILL, OOM, reboot) left a .ultrafuzz-submodule-transaction-* directory in the task worktree. Every later hydration and verification of that worktree then failed with stale pinned submodule transaction is present, until someone deleted the directory by hand. A test pinned that dead end.
  5. Runner-name scrub. Every smithers in operator-facing runner text became workflow runner, including inside paths, file names and links. A diagnostic about <target>/.smithers/workflows/ultrafuzz-<id>.tsx therefore pointed at <target>/.workflow runner/workflows/…, which does not exist. The runner's own See https://smithers.sh/reference/errors became the dead link https://workflow runner.sh/reference/errors.

Root cause

  1. EVMBench. Both status reads called commandData("status", execute(…)) directly, so any throw from execute ended the attempt.
  2. Dashboard. latestRunId read every run directory with a strict readRunMetadataDocument and no per-entry guard. listRuns already guards the same read.
  3. Evals.
    • runEvalSuite awaited watchEvalRow without a per-row catch.
    • When a row is recorded, the watch runs fail-closed reads: evalRunExpansion → readStaticNodeIds, and classifyRecoveryEquivalence. One row's bad evidence therefore rejected the whole suite.
    • The watch loop counted sync failures, but its only exits were a terminal state or the deadline.
  4. Pinned submodules. Hydration refused to start while any transaction entry existed (assertNoStaleTransactions), and nothing ever removed an entry left by a crash.
  5. Scrub. Five copies of .replace(/smithers/giu, "workflow runner") treated the whole text as prose: smithers.ts, lifecycle-inspection.ts, state-export.ts (twice) and workflow-sync.ts.

Change

There is one commit per item, then one commit that applies the review. Source is +160/−57 across 8 files.

  • fix(evmbench)

    • A failed status call is retried after 2, 4, 8 and 16 poll intervals: 30 s, 1, 2 and 4 min at the smoke profile's 15 s interval.
    • The fifth consecutive failure ends the attempt with Ultrafuzz status failed 5 consecutive times: <last failure>.
    • Both status reads retry: the check that decides whether an existing run needs resume, and the poll loop.
    • An envelope that fails the CLI result contract still fails at once, because retrying cannot fix a CLI and adapter that disagree.
  • fix(dashboard): latestRunId skips any entry whose name is not a run ID or whose run.json cannot be read, and picks the newest readable run.

  • fix(evals)

    • A watch that throws is caught per row:
      • the row is appended to runs.jsonl with an EVAL_ROW_WATCH_FAILED error;
      • it keeps status: "launched" with no observed outcome, so it counts as incomplete;
      • run-summary.json is written for every row, and eval run still exits non-zero (by code reading of commands/eval/run.ts).
    • The watch stops after ten identical consecutive sync failures, meaning the same message each time.
      • It records EVAL_ROW_SYNC_ABANDONED next to the existing EVAL_ROW_SYNC_FAILED warning.
      • A success or a different failure restarts the count.
      • The run is not cancelled.
    • docs/reference/evals.md describes both.
  • fix(runtime), pinned submodules. Hydration first removes every stale transaction entry, whatever its shape. Nothing in the entry is read or restored. Hydration then runs from the sealed snapshot, unchanged: it replaces every root, and a root that the crash left missing is hydrated like any other (replaceTaskRootsTransactionally already handles a missing root).

    Post-agent verification is also unchanged. An entry found after the agent ran still fails the attempt, and the retry's hydration clears it.

  • fix(runtime), scrub.

    • One exported scrubWorkflowRunnerText replaces the five copies.
    • It leaves whole any whitespace- or quote-delimited word that contains / or \, or a dot before a letter or digit. That covers paths, file:// and web URLs, and names such as .smithers and smithers.db.
    • A sentence's closing dot does not count, so restart smithers. is still scrubbed. Prose is scrubbed as before.
    • publicWorkflowText renames the runner's own smithers <command> suggestions. That rename no longer fires after /, \ or ., so …/bin/smithers ENOENT stays as written.

Deliberately not built

  • EVMBench
    • No retry or backoff settings in the profile schema. Five tries with a doubling wait are constants.
    • No wait on launch-incomplete. The poll loop cannot see that verdict. It needs run state pending plus a missing control seal or link journal, and ultrafuzz run exits 0 only after startRun has written both and set the run running. A successful resume also sets running.
    • An existing half-launched run therefore still fails at resume, as on the base: resumeRun on a planned, unsealed run returns WORKFLOW_LIFECYCLE_FAILED: ENOENT … lstat '<runs>/<id>/smithers'. Nothing else would finish that launch, so waiting would only turn the error into a stall until workflow_timeout_seconds.
  • Evals
    • A row whose watch failed is not marked status: "failed". That field means the launch failed, and downstream a failed row reads as a run that never reached a model, i.e. a relaunch candidate.
    • It is not marked final_status: "failed" either, because nobody observed that outcome.
    • The fail-closed evidence readers (readStaticNodeIds, classifyRecoveryEquivalence) are untouched. Only the damage is contained, to the one row.
    • The abandon rule compares consecutive messages for equality. It does not classify failures and does not match error text.
  • Dashboard reports no diagnostic for a skipped entry. listRuns already lists such entries as unreadable.
  • Pinned submodules. A hydration whose own rollback fails still keeps its transaction (preserveTransaction, PinnedSubmoduleRollbackError). The backup can be inspected there until the next hydration of that worktree removes it. Dropping that too would change no run outcome, and is left for a follow-up.
  • Scrub: the phrase rewrites in publicHealthReason (smithers why, smithers supervise -r …) are unchanged.

Verification

Every behavioural test below fails on origin/integration/wave1 and passes on this branch (head e6d6f50f). For base runs, the base versions of the changed source files were checked out under the branch's tests, and the tree was confirmed clean after restoring them.

Item Test On the base
EVMBench adapter.test.ts "keeps waiting through failed status calls". A success between two failure streaks restarts the backoff at two intervals. Throws Ultrafuzz command status failed with exit code 1 at the first failed poll.
EVMBench "ends the attempt after five consecutive failed status calls" Rejects with the raw failed to invoke Ultrafuzz: spawnSync node ETIMEDOUT after one call.
Dashboard dashboard.test.ts "starts on the latest readable run when a half-launched run directory has no run.json" cannot open regular file …/half-launched/run.json: ENOENT
Evals runner-publish.test.ts "publishes the run summary when a watched row's run evidence cannot be read" runEvalSuite rejects: planned graph is schema-invalid: / must have required property 'graph_version'; …
Evals "stops watching after ten identical consecutive synchronization failures" expected 40 to be 20: the watch kept syncing until the fixture's 40th call made the run terminal.
Submodules pinned-submodules.test.ts "the next hydration clears a transaction an interrupted one left behind" (replaces "a failed immediate submodule restore preserves the transaction backup") stale pinned submodule transaction is present.
Submodules Aave-shaped test, stale-transaction block. It used to assert that hydration throws forever. It now asserts that hydration clears a leftover directory, a leftover file, and a symlink to a directory outside the worktree, whose contents stay. stale pinned submodule transaction is present. A cleanup that follows the symlink fails with ENOENT … outside-the-task-worktree/kept.txt.
Scrub smithers-diagnostic.test.ts "workflow process diagnostics drop the runner name from prose but keep paths, file names and URLs" …/.workflow runner/workflows/…, file:///work/target/.workflow runner/agents/index.ts, workflow runner.db, .workflow runner and https://workflow runner.sh/reference/errors. A rule that keeps every word holding a dot also fails it, on restart smithers..
Scrub lifecycle-inspection.test.ts "diagnoseRun keeps a runner path intact in public text" /work/target/.workflow runner/workflows/…. Without the lookbehind: spawn /opt/runner/bin/workflow runner ENOENT.

One existing test changed: "counts a watched row as incomplete when it misses the watch deadline". Against its runner-less fixture, the default syncRun fails identically on every poll, so the watch would now stop after ten failures, before the one-second deadline the test is about. It now injects a sync that succeeds.

Suites run at this head (e6d6f50f):

  • evmbench: 67/67.
  • evals: runner-publish.test.ts 32/32. The full suite, 386/386 in 20 files, ran at d6dcc66c; this head changes only a comment in packages/evals.
  • pinned-submodules.test.ts: 5/5.
  • smithers-diagnostic.test.ts: 3/3.
  • dashboard.test.ts: 30/30. The dashboard is unchanged since d6dcc66c, where all three of its files ran 39/39.
  • lifecycle-inspection.test.ts: 52/52, over two runs. A 30-minute cap stopped the first after 24 tests at load ~33, and the other 28 ran in a second. This includes every test that asserts public inspection text carries no runner name (assertNoEngineBranding).
  • runtime.test.ts, the eight tests that assert on runner-derived text: 8/8.
    • "getRunHealth adapts the workflow health summary …", whose check is doesNotMatch(/smithers/iu);
    • "getRunHealth reports a live product status while workflow health is terminal";
    • "getRunHealth accepts strict 0.35 orphan, cancel-pending, quota, and operation metadata shapes";
    • "listRuns, observers and status re-read live run documents replaced while they were read";
    • the two "startRun rejects … project-local …" tests;
    • "a launch that fails after creating its run directory …";
    • "startRun includes bounded workflow runner stdio when submission fails". It failed once in the batch at load ~36 with workflow execution file dependencies/packages/000126/license changed while reading, before any runner text was produced. That reader compares link count and ctime, which change when other worktrees on the machine hard-link the same pnpm store files. It passed alone on the same build.
  • CLI, the tests whose operator output is checked for the runner name:
    • cli.test.ts: 8/8. These are "init and validate …", "run, ps, status, inspect, report, materialize, clean, and lifecycle commands expose product workflow evidence", "status surfaces a terminal product and live workflow lifecycle divergence", "runtime command failures …", "run rejects an OpenRouter override …", both "report bundle …" tests, and "product surface checks …".
    • lifecycle-commands.test.ts: 12/12, including "doctor reports install posture in human and JSON output".
  • The spec asked me to update tests that pin the mangled text. None on the base asserted a mangled path. The branch's own scrub test first asserted the mangled https://workflow runner.sh/… link, and now asserts the real one. The tests above that assert operator output never names the runner all pass.

Gates, at this head:

  • npx prettier --check on all 15 changed files.
  • pnpm -w lint.
  • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/integration/wave1 pnpm -w lint:strict:ci.
  • pnpm --filter @ultrafuzz/{evmbench,evals,runtime} typecheck, and the dashboard's two tsc --noEmit passes.
  • The complexity maximum is unchanged at 90 (synchronizeLinkedWorkflowRun), so the ceiling stays.
  • pnpm -w knip.
  • pnpm -w docs:check.

Risk / compatibility

  • EVMBench. A status call that keeps failing now ends the attempt only after 30 poll intervals of backoff, not at once: 7.5 min on the smoke profile and 15 min on full. If every call also hits the five-minute limit, that adds up to 25 min more. The retries do not check the workflow deadline, so they can overrun it by that much.
  • Dashboard. If the newest run's run.json is corrupt, the dashboard now opens the newest readable run, or the preview, instead of refusing to start, and does not say it skipped one.
  • Evals.
    • An identical sync failure that would have cleared after its tenth repeat now ends that row's watch early: ten attempts, 15 s apart by default. The row is recorded incomplete.
    • Failures whose message changes from poll to poll (timestamps, counts) never trip the rule and still run to the deadline.
    • Two new diagnostic codes: EVAL_ROW_WATCH_FAILED and EVAL_ROW_SYNC_ABANDONED. The schema is unchanged, since durable diagnostics accept any non-empty code.
  • Pinned submodules.
    • The stale entry sits in the agent-writable worktree. Nothing in it is read or moved back. It is only removed, and fs.rmSync unlinks a symlinked entry, and any symlink inside a directory entry, without following it. The Aave-shaped test covers the symlinked entry under Node. A probe covers both cases under Bun 1.3.14, which runs the controller.
    • The pre-hydration root that an interrupted hydration moved into backup/ is deleted, not restored. It was either a fresh worktree's empty submodule directory or an earlier attempt's hydrated tree. The sealed snapshot replaces either one.
  • Scrub.
    • Operator text can now contain smithers inside paths, file names and URLs, for example .smithers/workflows/…, node_modules/.bin/smithers, smithers.db or https://smithers.sh/reference/errors. Anything that greps operator output for the runner name will now find it there.
    • A word that contains a path is kept whole, so SMITHERS_BIN=/usr/bin/smithers keeps both names, and a package-like word such as smithers/cli counts as a path. So does a word with an inner dot, such as a version spec like [email protected].
    • Prose and command suggestions are scrubbed as before.
  • Shared files. The hunks are small and should merge cleanly with v01, v02, v03, v07, v08 and v10:
    • smithers.ts: the scrub helper only;
    • workflow-sync.ts: one import and one line;
    • state-export.ts: one import and two lines;
    • lifecycle-inspection.ts: one import and publicWorkflowText.

Changelog entry

Observers and benchmark harnesses no longer stop on transient failures: the EVMBench adapter retries failed ultrafuzz status calls (giving up after five in a row), ultrafuzz dashboard starts even when a run directory has no readable run.json, and eval run writes run-summary.json even when one row's evidence cannot be read and stops watching a row after ten identical sync failures instead of polling for six hours. A crash during pinned-submodule hydration no longer bricks the task worktree. Operator diagnostics keep real .smithers/… paths, file names and links instead of rewriting them to .workflow runner/….

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR should not merge until status retries respect the benchmark workflow deadline.

Fix All in Claude CodeFindings

  1. P1 Retries overrun workflow deadline ▶
  2. P2 Skipped runs go unexplained ▶
  3. P2 Command rewrite corrupts paths ▶
Fix with agent prompt
### Issue 1
packages/evmbench/src/adapter.ts:268
When `status` starts failing near the workflow deadline, these waits continue without checking it, and each status call can take up to five minutes. The adapter can run past the configured timeout and even accept a run that finishes late if a later status call succeeds. Check the deadline during retries, not only after a successful poll.

### Issue 2
packages/dashboard/src/index.ts:576-578
This catch silently skips a corrupt or temporarily unreadable `run.json` as well as a directory left by a failed launch. Because the dashboard caches its startup choice, an operator can remain on an older run for the whole session without knowing a newer one was skipped. Report the skipped-run reason while keeping the fallback.

### Issue 3
packages/runtime/src/lifecycle-inspection.ts:1940
The lookbehind protects `smithers` only when `/`, `\`, or `.` is immediately before it. In a diagnostic such as `spawn /opt/runner/pre-smithers run ENOENT`, the match starts after `-` and changes the executable path to `/opt/runner/pre-`ultrafuzz run`` before the path-preserving scrub runs. This leaves an operator with a path that does not exist; restrict command rewriting to standalone command tokens.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR makes benchmark polling, dashboard startup, and evaluation watches more tolerant of failures; clears abandoned submodule transactions; and preserves real paths in runner diagnostics.

  • The benchmark retry can overrun its workflow deadline.
  • Dashboard fallback hides why a newer run was skipped, and one diagnostic rewrite can still corrupt a path.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Run[Detached run] --> Status[EVMBench status polling]
  Status --> Retry[Retry failed calls]
  Retry --> Deadline[Workflow deadline]
  Run --> Evidence[Saved run evidence]
  Evidence --> Dashboard[Dashboard selects readable run]
  Evidence --> Eval[Eval watches rows and writes summary]
  Worktree[Task worktree] --> Hydrate[Clear abandoned transaction and hydrate]
  Runner[Runner diagnostics] --> Scrub[Preserve paths in operator text]
Loading

Reviews (1) · Last reviewed commit: "fix: apply review of the observer and ha..."

aviggiano and others added 6 commits September 29, 2026 09:35
…e launches

The adapter ended a benchmark attempt on the first `ultrafuzz status` call
that exited non-zero, timed out at the five-minute call limit, or printed
no JSON, although the run it watches is detached and keeps going. It also
aborted with "unknown status verdict" on `launch-incomplete`, which
`status` reports for a run whose launch preparation has not finished
publishing.

A failed status call is now retried after 2, 4, 8 and 16 poll intervals;
the fifth consecutive failure ends the attempt with the last failure's
message. This applies to both status reads: the one that decides whether
an existing run needs `resume`, and the poll loop. A contract-invalid
envelope still fails at once, because retrying cannot fix a CLI and
adapter that disagree. `launch-incomplete` is now a verdict the adapter
waits on.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
`ultrafuzz dashboard` picks the newest run by reading every run
directory's run.json, and one unreadable entry threw out of that scan. A
launch that fails before run.json is written leaves exactly such a
directory, so the dashboard could not start at all ("cannot open regular
file .../run.json: ENOENT") until someone deleted it by hand.

The scan now skips an entry whose name is not a run ID or whose run.json
cannot be read, and picks the newest readable run. `listRuns` already
tolerates the same directories (it lists them as unreadable).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…s failing

Two ways one eval row could cost the whole suite:

- A row whose run evidence could not be read threw out of the watch: for
  example an invalid graph.json, read when the row is recorded (recovery
  classification and expansion), or a foreign state.json. The rejection
  went up through Promise.all, so runEvalSuite rejected before
  run-summary.json was written, and the other rows' watches kept running
  unawaited. A failed watch is now caught per row: the row is appended to
  runs.jsonl with an EVAL_ROW_WATCH_FAILED error and stays launched with
  no observed outcome, so it counts as incomplete, and the summary is
  written for every row. `eval run` still exits non-zero because the row
  is incomplete.

- A sync that failed identically on every poll (a deterministic
  WORKFLOW_CONTROL_EVIDENCE_INVALID, say) was retried until the six-hour
  watch deadline. The watch now stops after ten identical consecutive
  failures and records EVAL_ROW_SYNC_ABANDONED next to the existing
  EVAL_ROW_SYNC_FAILED warning. A different failure, or a success,
  restarts the count. The detached run is not cancelled.

The watch-deadline suite test depended on the default syncRun failing the
same way against its fixture until the one-second deadline; it now
injects a sync that succeeds, so it still tests the deadline.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ydration

Hydration swaps each pinned submodule root through a transaction
directory inside the task worktree. A controller killed mid-hydration
(SIGKILL, OOM, reboot) left that directory behind, and every later
hydration and verification of the worktree then failed with "stale
pinned submodule transaction is present". Only deleting the directory by
hand recovered it, and a test pinned that dead end.

Hydration now rolls a stale transaction back first: each root it had
moved aside under backup/ goes back into place, the entry is removed,
and the hydration proceeds from the sealed snapshot as usual. The entry
lives where the agent can write, so a backup is put back only when it is
a physical directory under the worktree; anything else, including an
entry that is a file rather than a directory, is removed without being
restored, since hydration replaces every root from the sealed snapshot
anyway. The same path also completes a transaction that an earlier
hydration kept on purpose after its own rollback failed.

The post-agent verification is unchanged: a transaction entry found
after the agent ran still fails the attempt, and the retry's hydration
clears it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Runner text shown to an operator (process diagnostics, status reasons,
`why`, events, attempt errors, inspection failures) had every "smithers"
replaced with "workflow runner", including inside paths. A diagnostic
about <target>/.smithers/workflows/ultrafuzz-<id>.tsx therefore pointed
at <target>/.workflow runner/workflows/..., a directory that does not
exist.

The five copies of that replacement (smithers.ts, lifecycle-inspection.ts,
state-export.ts twice, workflow-sync.ts) now share one
scrubWorkflowRunnerText. It leaves alone any whitespace- or
quote-delimited word that contains a path separator, file:// URLs
included; http(s) URLs and prose are scrubbed as before. Lifecycle
inspection's rename of the runner's own `smithers <command>` suggestions
no longer fires after a path separator or a dot, so
".../bin/smithers ENOENT" stays as written.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
- evmbench: stop waiting on `launch-incomplete`. The poll loop cannot see
  it: `run` exits 0 only after startRun has sealed the run and set it
  `running`, and a successful `resume` sets `running` too. A half-launched
  existing run still fails at `resume`, as before, instead of stalling
  until the workflow timeout. The retry test keeps the fail/success/fail
  sequence that checks a success restarts the backoff.
- pinned submodules: the next hydration now just removes a stale
  transaction entry. Restoring its backup bought nothing, since hydration
  replaces every root from the sealed snapshot and already handles a
  missing root, and it moved agent-writable bytes back into the worktree.
  The Aave-shaped test also leaves a symlink to an outside directory and
  checks that only the link goes.
- scrub: a word holding a dot before a letter or digit is kept whole, like
  a word holding a path separator, and web URLs are no longer special
  cased. `.smithers`, `smithers.db` and the runner's docs link stay real
  instead of becoming `.workflow runner`, `workflow runner.db` and a dead
  `https://workflow runner.sh/...` link. A sentence's closing dot does not
  count, so prose is still scrubbed.
- evals: state the identical-failure rule as the policy it is.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 09:36
const reason = error instanceof Error ? error.message : String(error);
throw new Error(`Ultrafuzz status failed ${String(failures)} consecutive times: ${reason}`, { cause: error });
}
await wait(pollIntervalMs * 2 ** failures);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Retries overrun workflow deadline When status starts failing near the workflow deadline, these waits continue without checking it, and each status call can take up to five minutes. The adapter can run past the configured timeout and even accept a run that finishes late if a later status call succeeds. Check the deadline during retries, not only after a successful poll.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/evmbench/src/adapter.ts
Line: 268

Comment:
**Retries overrun workflow deadline** When `status` starts failing near the workflow deadline, these waits continue without checking it, and each status call can take up to five minutes. The adapter can run past the configured timeout and even accept a run that finishes late if a later status call succeeds. Check the deadline during retries, not only after a successful poll.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

Comment on lines +576 to +578
return [{ runId, createdAt: metadata.created_at, mtimeMs: fs.statSync(root).mtimeMs }];
} catch {
return [];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Skipped runs go unexplained This catch silently skips a corrupt or temporarily unreadable run.json as well as a directory left by a failed launch. Because the dashboard caches its startup choice, an operator can remain on an older run for the whole session without knowing a newer one was skipped. Report the skipped-run reason while keeping the fallback.

Knowledge Base Used: Dashboard and observability

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/dashboard/src/index.ts
Line: 576-578

Comment:
**Skipped runs go unexplained** This catch silently skips a corrupt or temporarily unreadable `run.json` as well as a directory left by a failed launch. Because the dashboard caches its startup choice, an operator can remain on an older run for the whole session without knowing a newer one was skipped. Report the skipped-run reason while keeping the fallback.

**Knowledge Base Used:** [Dashboard and observability](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/dashboard-observability.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

return scrubWorkflowRunnerText(
redactSecretsInText(value).replace(
// Not after a path separator or dot: `.../bin/smithers ENOENT` names a file, not a runner command.
/(`?)\b(?<![\\/.])smithers\s+([a-z][a-z-]*)(`?)/giu,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Command rewrite corrupts paths The lookbehind protects smithers only when /, \, or . is immediately before it. In a diagnostic such as spawn /opt/runner/pre-smithers run ENOENT, the match starts after - and changes the executable path to /opt/runner/pre-ultrafuzz run`` before the path-preserving scrub runs. This leaves an operator with a path that does not exist; restrict command rewriting to standalone command tokens.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/lifecycle-inspection.ts
Line: 1940

Comment:
**Command rewrite corrupts paths** The lookbehind protects `smithers` only when `/`, `\`, or `.` is immediately before it. In a diagnostic such as `spawn /opt/runner/pre-smithers run ENOENT`, the match starts after `-` and changes the executable path to `/opt/runner/pre-`ultrafuzz run`` before the path-preserving scrub runs. This leaves an operator with a path that does not exist; restrict command rewriting to standalone command tokens.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

@aviggiano
aviggiano merged commit 777cb73 into main Sep 29, 2026
14 of 16 checks passed
@aviggiano
aviggiano deleted the claude/v06-observer-harness-robustness branch September 29, 2026 10:01
aviggiano added a commit that referenced this pull request Sep 29, 2026
…gnostics

publicWorkflowText renames the runner's own `smithers <command>`
suggestions before the path-preserving scrub runs. #1194 stopped the
rename from firing after a path separator or a dot, but it still fired
after any other non-word character. A path whose last segment ends in
"smithers" after a hyphen (or @, +, ~, ...) and is followed by a space
and a word was rewritten before the scrub could keep it whole:
"spawn /opt/runner/pre-smithers ENOENT" became
"spawn /opt/runner/pre-workflow runner ENOENT", and
"Command failed: /opt/tools/pinned-smithers why --run-id r1" became
"Command failed: /opt/tools/pinned-`ultrafuzz why` --run-id r1".

Exposure is narrow. On Linux the runner is executed through
/proc/<pid>/fd paths, so a failed runner query does not name the
runner's own path (a probe with a failing runner at
<dir>/pinned-smithers reported "Command failed: /proc/<pid>/fd/29
/proc/<pid>/fd/28 why ..."), and a missing runner is reported as
"ENOENT: ... open '<path>'", which the rename leaves alone. The rewrite
still applies to any other text that reaches publicWorkflowText: runner
stderr, `why` prose, event labels and details, and node errors.

The rename now fires only where "smithers" starts a word: at the start
of the text, or after whitespace, a quote, a backtick, or an opening
parenthesis or bracket. The scrub still replaces the name in prose.
Over the 1,064 string literals that mention "smithers" in
packages/*/src and packages/*/test, main's and the new
publicWorkflowText differ only on this change's own new pre-smithers
strings. Over the 1,664 distinct lines that mention `smithers <word>`
in the pinned Smithers 0.35.0 packages they differ on one, an HTML
<title> template in @smthrs/cli's runReport.js.

The diagnoseRun fixture now also carries a hint after whitespace
("then run smithers inspect"), the form missing from the tests, so
narrowing the lookbehind to the start of the text and backticks fails.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…gnostics

publicWorkflowText renames the runner's own `smithers <command>`
suggestions before the path-preserving scrub runs. #1194 stopped the
rename from firing after a path separator or a dot, but it still fired
after any other non-word character. A path whose last segment ends in
"smithers" after a hyphen (or @, +, ~, ...) and is followed by a space
and a word was rewritten before the scrub could keep it whole:
"spawn /opt/runner/pre-smithers ENOENT" became
"spawn /opt/runner/pre-workflow runner ENOENT", and
"Command failed: /opt/tools/pinned-smithers why --run-id r1" became
"Command failed: /opt/tools/pinned-`ultrafuzz why` --run-id r1".

Exposure is narrow. On Linux the runner is executed through
/proc/<pid>/fd paths, so a failed runner query does not name the
runner's own path (a probe with a failing runner at
<dir>/pinned-smithers reported "Command failed: /proc/<pid>/fd/29
/proc/<pid>/fd/28 why ..."), and a missing runner is reported as
"ENOENT: ... open '<path>'", which the rename leaves alone. The rewrite
still applies to any other text that reaches publicWorkflowText: runner
stderr, `why` prose, event labels and details, and node errors.

The rename now fires only where "smithers" starts a word: at the start
of the text, or after whitespace, a quote, a backtick, or an opening
parenthesis or bracket. The scrub still replaces the name in prose.
Over the 1,064 string literals that mention "smithers" in
packages/*/src and packages/*/test, main's and the new
publicWorkflowText differ only on this change's own new pre-smithers
strings. Over the 1,664 distinct lines that mention `smithers <word>`
in the pinned Smithers 0.35.0 packages they differ on one, an HTML
<title> template in @smthrs/cli's runReport.js.

The diagnoseRun fixture now also carries a hint after whitespace
("then run smithers inspect"), the form missing from the tests, so
narrowing the lookbehind to the start of the text and backticks fails.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant