fix(cli): recovery hints name ultrafuzz commands, not raw runner invocations - #1207
Merged
aviggiano merged 6 commits intoSep 29, 2026
Merged
Conversation
…y suggestion `ultrafuzz why` passed the workflow runner's recovery suggestions through a scrub that only renamed the command. An unblocker such as `smithers up <target>/.smithers/workflows/ultrafuzz-<id>.tsx --run-id ultrafuzz-<id> --resume true` reached the operator as `workflow runner up ...`, and the failed-run checkpoint note as "`ultrafuzz replay` <path> --run-id ultrafuzz-<id> --frame 29`": a runner invocation, the runner's run ID, and a `--frame` flag `ultrafuzz replay` does not have. Each suggestion is now rebuilt as the ultrafuzz command for the run, keyed by the Ultrafuzz run ID: an unblocker as a whole command, and commands the runner quotes in backticks in the notes and summary. - resume (`up`) and task retry (`retry-task`) on a failed run become `ultrafuzz resume <run-id> --retry-failed`: a plain resume leaves a failed node failed, and --retry-failed also retries a failed artifact verifier from its producer; - on a live run, `up` becomes `ultrafuzz resume <run-id>` and `retry-task --node-id <node>` becomes `ultrafuzz resume <run-id> --reset-node <node>`; - replay from checkpoint frame N becomes `ultrafuzz fork <run-id> --frame N`; - `logs` becomes `ultrafuzz events <run-id> --watch`, and `inspect` becomes `ultrafuzz inspect <run-id>`. A rebuilt command is not scrubbed, so a run ID containing "smithers" stays intact. Suggestions ultrafuzz has no command for (approve, deny, signal, which generated workflows never wait on) keep the neutral rename. Co-Authored-By: Claude Opus 5.5 <[email protected]>
`ultrafuzz node <run> final-report` failed with "WORKFLOW_NODE_FAILED:
Command failed: /proc/<pid>/fd/24 --config=... node final-report --run-id
ultrafuzz-<run> --format json --full-output". The runner had answered
with `{"ok":false,"error":{"code":"NODE_NOT_FOUND","message":"Node not
found: final-report"}}` on stdout and exit 4, but a failed inspection
query kept the process error as its message, which only says the
command exited and names the whole runner invocation.
A failed query's `error` is now the message from the runner's JSON
envelope when it printed one. `why`, `node`, `timeline` and `snapshots`
report it, and so do synchronization and status diagnostics when the
runner wrote nothing to stderr, which they prefer. A timeout, or a
failure with no envelope, reports as before.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
The pinned runner's status reason for a paused run ends with "resume with `smithers up --resume <runId>`". publicHealthReason rewrote only the runner's `why` and `supervise -r` suggestions, so `ultrafuzz status` printed "resume with `workflow runner up --resume <runId>`". The paused suggestion now becomes `ultrafuzz resume <run-id>`, like the orphaned one. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every backticked `smithers up ...` was rebuilt as a resume and lost its flags. The runner's concurrency-saturation remediation, `smithers up --max-concurrency 8`, starts a run with a higher limit; it became `ultrafuzz resume <id>` (or `--retry-failed` on a failed run), dropping the limit that was the whole remediation. An `up` is now rebuilt only when it carries `--resume`; any other `up` keeps the neutral rename it had before. The reference now lists which suggestions are rebuilt, and says that approve and signal suggestions stay in the runner's words. Co-Authored-By: Claude Opus 5.5 <[email protected]>
… invented one Reporting the runner's own error message (0739a4d) also carried the runner's advice into `why`, `node`, `status` and sync warnings. For a project whose store has no run history the pinned runner says "No Smithers run history found at <db>. Run 'smithers up <workflow>' to start a run first. See https://smithers.sh/reference/errors", and `why` printed "No `ultrafuzz run` history found at <db>. Run 'workflow runner up <workflow>' to start a run first. ...": an invented ultrafuzz command, and a runner command that is the wrong advice for a run that exists. - runnerReportedError drops each sentence that quotes a runner command. - publicWorkflowText renames only lowercase `smithers <command>`, the form the runner's command suggestions take. The capitalized name in runner prose is no longer read as a command; the scrub turns it into "workflow runner". Co-Authored-By: Claude Opus 5.5 <[email protected]>
| case "up": | ||
| case "retry-task": { | ||
| if (command === "up" && !args.includes("--resume")) return undefined; | ||
| if (runFailed) return `ultrafuzz resume ${runId} --retry-failed`; |
There was a problem hiding this comment.
On a failed run, a runner suggestion to retry one task becomes ultrafuzz resume <run-id> --retry-failed, dropping its --node-id. That command resets every failed task, so an operator following a single-node hint can also rerun unrelated failed work, including costly producers. The hint should preserve the target where possible or make its broader effect clear.
Knowledge Base Used: Runtime orchestration
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/lifecycle-inspection.ts
Line: 1985
Comment:
**Targeted retry becomes broad**
On a failed run, a runner suggestion to retry one task becomes `ultrafuzz resume <run-id> --retry-failed`, dropping its `--node-id`. That command resets every failed task, so an operator following a single-node hint can also rerun unrelated failed work, including costly producers. The hint should preserve the target where possible or make its broader effect clear.
**Knowledge Base Used:** [Runtime orchestration](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/runtime-orchestration.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
aviggiano
force-pushed
the
claude/x04-operator-hints-use-ultrafuzz-commands
branch
from
September 29, 2026 14:01
fd6015b to
2716620
Compare
…istory A runner 'logs' suggestion became 'ultrafuzz events <run> --watch', which streams only new events. On a stalled or ended run that shows none of the events that explain the blocker. Add --history so the hint replays existing events first, as the runner's logs command does (Greptile, on #1207). Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
deleted the
claude/x04-operator-hints-use-ultrafuzz-commands
branch
September 29, 2026 15:28
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #1145
Problem
For the failed
smoke-main-2smoke run,ultrafuzz whyon main (2cacf4c) printed recovery steps an operator cannot run:workflow runneris not a command.ultrafuzz-smoke-main-2is the runner's run ID, not the Ultrafuzz one.ultrafuzz replaytakes neither a workflow path nor--frame. Even a well-formedultrafuzz resume smoke-main-2would not help, because a plain resume leaves the rejected verifier failed. #1194 already keeps the path intact. This PR fixes the hint text.ultrafuzz statuson a paused run has the same fault. Its reason reads "run was gracefully paused; resume withworkflow runner up --resume <runId>".ultrafuzz node smoke-main-2 final-reportfailed withWORKFLOW_NODE_FAILED: Command failed: /proc/<pid>/fd/24 --config=/proc/<pid>/fd/22/controls/bunfig.toml ... /proc/<pid>/fd/22/dependencies/packages/000004/src/bin/smithers.js node final-report --run-id ultrafuzz-smoke-main-2 --format json --full-output. Other failed runner queries, apart from timeouts, were reported the same way. For example,whyon a project whose store holds no run history showed aCommand failed: <argv>line too.Root cause
Runner suggestions were renamed, not rebuilt.
publicWorkflowText(lifecycle-inspection.ts) rewrote a runner suggestion by its command name only. If Ultrafuzz had a command with the same name,smithers <cmd>became`ultrafuzz <cmd>`, and the runner's arguments stayed after the closing backtick. Otherwise it becameworkflow runner <cmd>. A rename cannot produce the right command here:replay --frame Ncorresponds toultrafuzz fork --frame N.up --resumehas to becomeultrafuzz resume --retry-failed.up,retry-taskandlogshave no Ultrafuzz command of the same name.These suggestions reach an operator in two places:
why(summary, unblockers and notes) and thestatushealth reason. fix(runtime): make run synchronization single-writer, idempotent, bounded and non-fatal for observers #1180'spublicHealthReasonrewrote the status reason'ssmithers whyandsmithers supervise -r <id>. It did not rewrite the paused reason'ssmithers up --resume <runId>(pinned runner, run-status.js:484). Run-level errors already leave out the resume commands the runner appends to them, becausecurrentSmithersRunErrorkeeps thecauseorsummary.node final-reporthad two causes:final-reportis the topology node ID. The workflow's task IDs areprepare:final-report,node:final-reportandverify:final-report, andultrafuzz node smoke-main-2 verify:final-reportworks. So the runner answered{"ok":false,"error":{"code":"NODE_NOT_FOUND","message":"Node not found: final-report"}}on stdout, with exit 4. I captured this with a probe against a copy of the target.runSmithersInspectionCommandstored Node's exec error (Command failed: <argv>) as the snapshot'serror. Thenodediagnostic printserrorfirst, so the operator saw the bun command line instead of the runner's reason.The runner's own message carries its advice and its name. It follows the reason with its own next step. For a store with no run table it says "No Smithers run history found at . Run 'smithers up ' to start a run first. See https://smithers.sh/reference/errors" (openSmithersStore.js:262). For a missing store it says "No smithers.db found at ." followed by the same advice (line 115). The
publicWorkflowTextrename was case-insensitive, so it read the capitalized product name in "Smithers run history" as the commandsmithers run.Change
whyrebuilds each runner recovery suggestion as theultrafuzzcommand for the run, using the Ultrafuzz run ID. An unblocker is rewritten as one whole command. In the notes and the summary, only the commands the runner quotes in backticks are rewritten:whynow saysup … --resume true,retry-task … --node-id Non a failed runultrafuzz resume <run-id> --retry-failedup … --resume …on a live or paused runultrafuzz resume <run-id>retry-task … --node-id Non a live runultrafuzz resume <run-id> --reset-node Nreplay … --frame Nultrafuzz fork <run-id> --frame Nlogs <id>ultrafuzz events <run-id> --watchinspect <id>ultrafuzz inspect <run-id>up --max-concurrency N(no--resume)`workflow runner up --max-concurrency N`A rebuilt command skips the runner-name scrub, so a run ID that contains "smithers" stays intact.
approve,denyandsignalhave no Ultrafuzz command and keep the existing neutral rename. The runner raises them only for approval and signal waits, and generated workflows contain no approval, signal or timer components.statusreason for a paused run.publicHealthReasonnow also turnssmithers up --resume <runId>intoultrafuzz resume <run-id>, as it already did for the orphaned run'ssupervise -r. This extends the pattern in fix(runtime): make run synchronization single-writer, idempotent, bounded and non-fatal for observers #1180's function by one alternative.Failed runner queries. When a query fails and the runner printed a JSON error envelope, the snapshot's
erroris the envelope's message instead ofCommand failed: <argv>. Any sentence in it that quotes a runner command is dropped. Timeouts, and failures with no envelope, report as before.The rename is lowercase only.
publicWorkflowTextrenames only lowercasesmithers <command>, which is how the runner writes its command suggestions. "Smithers run history" now scrubs to "workflow runner run history" instead of becoming`ultrafuzz run`.Docs.
docs/reference/cli.mdgets one paragraph listing which suggestionswhyrebuilds, and saying that approve and signal suggestions stay in the runner's words.After the change, on the same copied project:
whywith the copy'ssmithers.dbreplaced by an empty file, which makes the pinned runner report missing run history:With
smithers.dbmoved away, both lines read "No smithers.db found at /smithers.db. See https://smithers.sh/reference/errors".Deliberately not built
retries={0}, so they fail rather than stall. That means--reset-nodeon a live run never reruns a verifier on output it already rejected.statuskeeps fix(runtime): make run synchronization single-writer, idempotent, bounded and non-fatal for observers #1180'spublicHealthReason, extended by one alternative, andwhyhasultrafuzzRecoveryCommand. Merging them would move code between state-export.ts and lifecycle-inspection.ts, and neither diff here needs that. If the runner changes how it words its suggestions, both need updating.ultrafuzzform ofup --max-concurrency N. That suggestion starts a run with a higher limit.ultrafuzz resume <run-id> --max-concurrency Ndoes nothing on a live run, so no singleultrafuzzcommand is right for it, and the hint keeps main's wording. For a run with a pinned limit, the engine records this note only when a descendant run waits for a slot (engine.js:9595-9607). Submit always pins the limit (smithers.ts:4474), so the note should be rare on generated runs.--forcefor an orphaned run. The runner suggestsup … --force truefor stale-heartbeat, but plainultrafuzz resumeis enough. The runner classifies such a run asorphaned(orstalewhen the owner cannot be checked), and neither is an active state, so ultrafuzz goes on toup --resume. The runner refuses that only for a fresh heartbeat or a live owner. This matches the fix(runtime): make run synchronization single-writer, idempotent, bounded and non-fatal for observers #1180 orphan hint instatus.ultrafuzz node.final-reportcould mean any of three tasks. Insmoke-main-2the failing one isverify:final-report, so guessingnode:final-reportwould show a finished node. The error now names the ID the runner did not find.ultrafuzz replayis left alone, but it is broken on the pinned runner. It runssmithers replay <workflow> --run-id <id> --format json --full-outputwithout--frame. Smithers 0.35.0 requires that option: running the pinned runner'sreplaywithout it returnsVALIDATION_ERROR: Missing required option --framewith exit 4. The hints therefore nameultrafuzz fork <run-id> --frame N, which does the runner's fork and resume. Adding--frametoreplayor retiring it is a CLI decision, and a parallel change in this batch reworks the replay and fork result types.why --jsonstill rejects five blocker kinds (found, not fixed). The CLI result schema's blockerkindenum listsbinding,side-effect-boundaryandother. It does not liststale-task-heartbeat,bound-stale,binding-missing,approval-decided-resume-requiredorside-effect-boundary-crossed.validateCliResultEnveloperejects awhyresult carrying any of them. Sinceenvelope()throws on that,why --jsonfails for a hung task (stale-task-heartbeat) or after a forced crossing. Humanwhyis unaffected. This needs a separate schema fix.Run ultrafuzz-smoke-main-2 is failed), whichwhy --jsonalso reports asworkflow_run_id. It is not a recovery hint.Verification
lifecycle-inspection.test.ts› "diagnoseRun names the ultrafuzz command for each recovery the runner suggests". It uses the pinned runner's exact suggestion strings for a failed, a live and a paused diagnosis, under run IDsmithers-probe.workflow runner up /work/target/.smithers/workflows/ultrafuzz-smithers-probe.tsx --run-id ultrafuzz-workflow runner-probe --resume true.Remediation: `ultrafuzz resume smithers-probe`.instead of`workflow runner up --max-concurrency 8`.runtime.test.ts› "getRunHealth accepts strict 0.35 orphan, cancel-pending, quota, and operation metadata shapes". It gains a paused case with the runner's exact reason. On both main and 0739a4d the reason wasresume with `workflow runner up --resume <runId>`.cli/lifecycle-commands.test.ts› "why reports a runner failure in the runner's words, without its advice". The fake runner answerswhywith the pinned runner's missing-history envelope and exit 1, and the test asserts the exacterror:line.error: WORKFLOW_DIAGNOSIS_FAILED: Command failed: /proc/<pid>/fd/29 /proc/<pid>/fd/28 why ultrafuzz-lifecycle-cli-run --format json --full-output.No `ultrafuzz run` history found at <db>. Run 'workflow runner up <workflow>' to start a run first. See ….`ultrafuzz run`, and the lowercase-only rename alone leaves the advice sentence.smithers-executable-capability.test.ts› "a failed runner query reports the runner's own reason, not the command line it ran". On main,errorwasCommand failed: /proc/<pid>/fd/22 /proc/<pid>/fd/21 node final-report --run-id ultrafuzz-run --format json --full-output\n.cli/lifecycle-commands.test.ts› "why reports the diagnosis in human and JSON output". The fixture uses the runner's realretry-tasksuggestion. On main, the rendered line isworkflow runner retry-task /tmp/…/.smithers/workflows/ultrafuzz-lifecycle-cli-run.tsx --run-id … --iteration 0.diagnoseRun|getWorkflowNode|watchWorkflowNode|getRunTimeline|listRunSnapshots|event queries|queryWorkflowEvents|watchWorkflowEvents|truncated event stream|cancelRun|lifecycle commands rejectgetRunHealth|controller refresh|parseCurrentSmithersInspect|status reports a runner query|status keeps reporting runner health--test-concurrency=1): every test exceptdoctorwhy smoke-main-2,node smoke-main-2 final-report,status smoke-main-2, andwhywith the store emptied or moved away). I ran them against an rsync copy of/home/ubuntu/targets/tiny-vault, so the real run directory was not touched. The outputs are above. I did not reproduce a pausedstatusend to end; the runtime test covers it with the runner's exact reason string.pnpm -w lint,CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci,pnpm -w knip,pnpm -w docs:check, andtypecheckfor runtime and cli all pass. The complexity ceiling (83,verifyCoverageProductionInventory) is unchanged.Risk / compatibility
unblocker,notes,summaryand the pausedstatusreason change for runner suggestions only. The JSON shape and types are unchanged. Prose unblockers and non-command notes go through the same scrub as before.error. A failed inspection query'serroris now the runner's reason, without the runner's advice, instead ofCommand failed: <argv>. Nothing in the repository matches on that text.smithersSnapshotReportsMissingRunreads the raw envelope, noterror. Synchronization and status diagnostics still prefer stderr when the runner wrote any.'smithers …or`smithers …). If a runner message puts its reason and its advice in one sentence, that sentence is dropped. If nothing is left, the error falls back toCommand failed: <argv>as before.Smithers <command>in runner text is no longer rewritten to anultrafuzzcommand; it scrubs toworkflow runner <word>. In the pinned runner, command suggestions are written in lowercase, and the capitalized name appears in prose such as "No Smithers run history".ultrafuzz resume <run-id> --retry-failedis now the headline hint for a verifier rejection likesmoke-main-2. I did not run it against that run. Making--retry-failedrecover a verifier rejection is a separate fix in this batch ("resume --retry-failed can recover a run whose verifier rejected output"). This PR only makes the hint name the command.--projectin hints. Rebuilt hints omit--project, like fix(runtime): make run synchronization single-writer, idempotent, bounded and non-fatal for observers #1180's status hint, so they assume the operator's working directory is the target project.packages/runtime/test/lifecycle-inspection.test.ts. This PR adds arunId = "inspect-run"parameter tolaunchedProject, and test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203 restructures that function. Whichever lands second adds the parameter to the restructuredlaunchedProject.git merge-treeshows no other conflict with the open PRs. feat(runtime): let agents record roadblocks in a run-local friction log #1162's conflict in smithers.ts already exists against main.Changelog entry
ultrafuzz whynow gives each recovery step as a runnableultrafuzzcommand for the run, such asultrafuzz resume <run-id> --retry-failedfor a failed node orultrafuzz fork <run-id> --frame <n>for the last good checkpoint, instead of a workflow runner invocation with its workflow path and run ID.ultrafuzz statuson a paused run now suggestsultrafuzz resume <run-id>. A failed runner query, such asultrafuzz nodewith an unknown node ID, now reports the runner's reason ("Node not found: final-report") without the runner's own advice, instead of the command line it ran.Greptile follow-up
logssuggestion now becomesultrafuzz events <run-id> --watch --history.--watchalone streams only new events, so on a stalled or ended run it showed none of the events that explain the blocker (comment).retry-task --node-id <task>suggestion still becomesultrafuzz resume <run-id> --retry-failed(comment). Its node ID is a runner task ID, which may name a verifier.--retry-failedis the recovery path that re-runs a failed verifier from its producer (fix(runtime): resume --retry-failed can recover a run whose verifier rejected output #1205). On a failed run, the failed nodes are the work the campaign needs to finish. On a run that has not failed, the hint keeps the target (--reset-node <task>), as before.Also rebased onto
mainafter #1203. The rebase resolves thelaunchedProjecthelper conflict by keeping #1203's shared-run structure and adding this PR'srunIdparameter.lifecycle-inspection.test.ts: 53/53 on the rebased tree.🤖 Generated with Claude Code
The PR appears safe to merge, although the existing targeted-retry wording concern remains non-blocking.
Fix with agent prompt
Summary
The PR rebuilds runner recovery suggestions as Ultrafuzz commands, updates the paused-status hint, and reports runner error messages instead of failed command lines. The change since the previous review adds
--historyto the events hint so it includes existing events.Reviews (3) · Last reviewed commit: "fix(runtime): the events hint for a runn..."