fix(cli): plain status, inspect and why print a successful poll's warnings, and the e2e covers a sealed engine relaunch - #1209
Merged
Conversation
…ssful poll commandFromRuntime rendered a result's diagnostics only when the result failed. getRunHealth returns ok unless a diagnostic is an error, and the sync and deadline problems it reports are warnings (WORKFLOW_DEADLINE_CANCEL_FAILED, WORKFLOW_STATE_SYNC_SKIPPED, WORKFLOW_CONTROL_EVIDENCE_DIVERGED and the transient sync codes), so plain `ultrafuzz status` exited 0 without them. inspect and why render through the same helper. docs/reference/configuration.md tells operators to run `ultrafuzz status` from cron and act on its warnings. A successful result now prints its diagnostics after the rendered value, and a failure still leads with them. resume and init appended successful diagnostics themselves; those copies are gone. JSON output is unchanged. The status reference now says where the warnings appear. The new CLI test sets a past workflow_deadline_at and has the fake runner answer cancel outside its status contract. On origin/main plain status prints no warning line; with this change it prints `warning: WORKFLOW_DEADLINE_CANCEL_FAILED: ...`. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…aled engine re-runs the interrupted node No test checked that a fresh `ultrafuzz run`, whose engine runs from the sealed execution snapshot, recovers when only its engine dies. The real-engine relaunch test runs a plain-path engine, the fd-transfer harness runs a toy script, and this e2e killed the engine together with its supervisor, then recovered through `ultrafuzz resume`. The e2e now first SIGKILLs only the detached engine (the `up` process, not `supervise`) while summarize is held, and waits, with no ultrafuzz command, for the relaunched engine to hold summarize again in a new process. It then kills the whole controller, which is now the relaunched engine and its supervisor, and resumes, so summarize starts three times and the event log has three RunStarted and a RunAutoResumed. The relaunched engine does not finish the run: it is killed while it holds summarize, and the engine `ultrafuzz resume` starts finishes it. The resume phase therefore starts from a supervisor-relaunched engine instead of the one `ultrafuzz run` started. Resume takes the workflow path from Ultrafuzz's run metadata, so what differs is the state the relaunched engine wrote to the Smithers database. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…agnostics A successful status poll prints its warnings on stdout without changing the exit status, and a message can continue on lines without a severity prefix. A cron job that discards stdout, or greps for `warning:` lines, can therefore miss them or see only the first line of each. The status reference now says so, and the deadline advice in the configuration reference points scripts at the `diagnostics` of `ultrafuzz status <run-id> --json`. Co-Authored-By: Claude Opus 5.5 <[email protected]>
| text: result.value | ||
| ? `${result.ok ? "" : diagnosticsText(result.diagnostics)}${text(result.value)}` | ||
| ? result.ok | ||
| ? `${text(result.value)}${diagnosticsText(result.diagnostics)}` |
There was a problem hiding this comment.
Warnings depend on renderer newlines The helper appends diagnostics without adding a separator. All current renderers end with a newline, so output is correct today, but a future renderer that does not would attach
warning: to its final line and make the warning harder to recognize. Adding the separator here would keep that formatting requirement in one place.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/cli/src/command-shared.ts
Line: 137
Comment:
**Warnings depend on renderer newlines** The helper appends diagnostics without adding a separator. All current renderers end with a newline, so output is correct today, but a future renderer that does not would attach `warning:` to its final line and make the warning harder to recognize. Adding the separator here would keep that formatting requirement in one place.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
aviggiano
added a commit
that referenced
this pull request
Sep 29, 2026
…ptor once (#1215) `packages/modal/test/deterministic-archive.test.ts` › "rejects a caller-visible output parent replaced after descriptor normalization" fails intermittently. It failed in #1209's package gates lane: ``` AssertionError: expected [Function] to throw error including 'deterministic archive output changed while it was written' but got 'EBADF: bad file descriptor, close' ``` Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Plain
statushid the warnings the deadline docs tell operators to act on. Greptile on #1157. docs: state that workflow_deadline_seconds is only enforced on sync #1157 (docs: state that workflow_deadline_seconds is only enforced on sync) added this advice todocs/reference/configuration.md: "runultrafuzz status <run-id>periodically (for example from cron) and act on its warnings". ButcommandFromRuntime(packages/cli/src/command-shared.ts) rendered a result's diagnostics only when the result failed.getRunHealthreturns ok unless a diagnostic is an error, and these deadline and sync diagnostics are warnings:WORKFLOW_DEADLINE_CANCEL_FAILED,WORKFLOW_STATE_SYNC_SKIPPED,WORKFLOW_CONTROL_EVIDENCE_DIVERGEDand the transient synchronization codes. So a cron job polling a run whose deadline cancel keeps failing, or whose sync is skipped because its control evidence diverged (the No supported way to re-seal workflow control evidence: one divergent file permanently disables product-level status and reporting #674 shape), saw exit 0 andStatus: running-healthy (running)on every poll, while nothing cancelled the run.inspectandwhyrender through the same helper. Onlyresumeandinitappended successful diagnostics themselves.No test checked that a sealed engine killed on its own gets relaunched by its supervisor. Greptile on #1199. fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199 (fix(runtime): the Smithers supervisor can relaunch a crashed engine) fixed the relaunch of both plain-path and sealed engines. A sealed engine is what every fresh
ultrafuzz runstarts. But:ultrafuzz resume.The sealed relaunch was checked only by hand, in fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199's manual experiment. A later change to the sealed launch path, such as Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921 or feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201, could leave every engine-only crash of a fresh run orphaned until an operator resumes it. No test takes a sealed engine through a supervisor relaunch, so none would fail.
Change
commandFromRuntimeprints a successful result's diagnostics after the rendered value. Each one prints on stdout as<severity>: <code>: <message>, and the exit status stays 0. Failures still lead with them. The duplicate appenders incommands/resume.tsandcommands/init.tsare deleted. Source is +5/−9 lines. The JSON envelope is unchanged. In the docs:statussection ofdocs/reference/cli.mdsays where these lines appear and that a message can span several lines. It also says that warnings do not change the exit status, so scripts should read the--jsondiagnosticsarray.docs/reference/configuration.mdnow tells cron jobs to pollultrafuzz status <run-id> --jsonand act on the warnings in itsdiagnostics. It notes that plainstatusprints them on stdout.packages/cli/test/e2e/campaign-resume.test.ts) first kills only the engine. Test code only, no production change:summarizeis held, it SIGKILLs only the detached engine: the process whose argv hasupand notsupervise.ultrafuzzcommand, it waits for the relaunched engine to holdsummarizeagain in a new process. A newheldCallhelper serves both waits.ultrafuzz resumenow follows a relaunched engine, not the engineultrafuzz runstarted (see Risk).summarizestarts 3 times, the event log has 3RunStarted, and it must contain aRunAutoResumed. The test title now names both kills.Deliberately not fixed
summarize.RunFinishedcomes from the engineultrafuzz resumestarts, and that engine runs the project-relative workflow file recorded inrun.json, not the sealed snapshot.codexincluded, under@smthrs/agents/src/BaseCliAgent/parentDeathWatchdog.js(runCommandEffect→parentDeathCommand). The watchdog checks the engine's pid every 100 ms and SIGKILLs the agent's process group once the engine is gone. So killing the engine kills the held stub. origin/main's existing wait for "the controller and its agent to exit" already relies on this.smithers-supervisor-relaunch.integration.test.ts. Draft feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 rewrites that file. The e2e drives the CLI black-box, so it keeps guarding the sealed path through feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 and the Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921 rewrites.statusprints theWORKFLOW_CONTROL_SEAL_PENDINGwarning, and its message is also theReason:line.statusthen prints bothLifecycle divergence: Ultrafuzz is terminal timed-out, but the workflow runner is running; ...andwarning: RUN_WORKFLOW_STATUS_DIVERGED: .... From reading the code, a poll that fails but still returns a health value (a non-transient sync error) already printed both on origin/main, diagnostics first. A follow-up could deletelifecycleDivergenceLinesincommands/status.ts(about −7 lines), because the same condition now also prints the warning, and pointcli.test.ts's divergence assertion at the warning line.RunAutoResumedmust be present, but its count is not checked. A relaunch that needs a second supervisor attempt under load still finishes the campaign, and this test is not meant to measure activation latency.RunStartedis still required to be exactly 3.Verification
packages/cli/test/lifecycle-commands.test.ts). It sets a pastworkflow_deadline_at, and the fake runner answerscanceloutside its status contract.AssertionError [ERR_ASSERTION]: The input did not match the regular expression /^warning: WORKFLOW_DEADLINE_CANCEL_FAILED: /mu. The input is the plain status text, which ends atPace: 0 finished in the last 10mwith no warning line. Both reviews reproduced this failure with main's text logic inpackages/cli/dist. oclif loads commands from there, so rebuilding onlydist-testdoes not test main.lifecycle-commands.test.ts(13/13);cli.test.tstests selected by name, most of which run plain-text commands: the product-evidence test, plain init, ps text, status launch-incomplete, status lifecycle divergence, status quota parking, resume already-active, references status and symlinked--project.cli.test.tsplusaudit-profile-commands(88/91):report bundle --require-verifiedtests whoseultrafuzz run --jsonfailed withWORKFLOW_SUBMISSION_FAILED(... changed while reading) on a dependency file. That is the launch-snapshot race that x03 fixes.runCliwith a fake runner. Every scenario exits 0, and the JSON is unchanged. On origin/main the same script printed no warning line in any scenario.statusnow ends withwarning: WORKFLOW_DEADLINE_CANCEL_FAILED: workflow runner command failed (exit 1), followed by the message's ownstderr: runner busyline. Plainwhyprints the same warning.warning: WORKFLOW_CONTROL_EVIDENCE_DIVERGED: ...andwarning: WORKFLOW_STATE_SYNC_SKIPPED: ..., and plaineventsprints the first.statusprints theLifecycle divergence:line and theRUN_WORKFLOW_STATUS_DIVERGEDwarning described above.summarize45 s after the engine SIGKILL;/proc/self/fd/3/...paths,--log-dir <run>/smithers/logsand--resume;RunAutoResumed×1,RunStarted×3,summarizeNodeStarted×3, thenRunFinished;supervisor_resume_rootpatch to upstream behaviour (return direct;) and rebuilt the runtime.AssertionError [ERR_ASSERTION]: timed out after 15 minutes waiting for the relaunched engine to rerun summarize. The supervisor log repeatsCannot resume run ...: detached log is unavailable at <run>/smithers/execution-snapshots/<gen>/.smithers/workflows/.smithers/logs/<id>.log, the pre-fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199 sealed failure.resume_snapshot_transferpatch, by pointing a sealed relaunch at a missing--config=/proc/self/fd/3/controls/ablated-bunfig.toml. The plain-path relaunch was left alone.the workflow stopped while waiting for the relaunched engine to rerun summarize. The review reports that the supervisor log showed ENOENT for that config on attempts 1–3, after which the supervisor gave up.npx prettier --checkandnpx eslinton the changed files;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci. It also passes with the merge base as the diff commit, which is what CI's pull-request diff sees; origin/main has two later commits that touch no file here;pnpm -w lint;pnpm --filter @ultrafuzz/cli typecheck;pnpm -w knip;pnpm -w docs:check, which includesnode scripts/docs-check.mjs.Risk / compatibility
commandFromRuntime. A successful result with diagnostics now prints them on stdout after its text, whichresumeandinitalready did. The affected commands arestatus(each--watchpoll too),inspect,why,events,node,timeline,snapshots,validate,doctor,run,ps,pause,cancel,replay,fork,clean,materializeandreferences status|sync|update. The output ofresumeandinitis unchanged.statusstdout will now seewarning:lines.docs/cli.mdalready says human text is not the automation contract,--jsonis unchanged, and the status and deadline docs now point scripts at itsdiagnostics.doctorrepeats what its check lines summarise. It now also lists the validation, temporary-directory and registry warnings behind those lines. This is from reading the code.ultrafuzz runstarted, so the e2e no longer covers a host crash before any relaunch. Resume takes the workflow path from Ultrafuzz's run metadata (start-run.ts), not from the engine that ran last. What differs is the state the relaunched engine wrote to the Smithers database.cli-e2erelease-validation lane. CI runs that lane on every pull request and on pushes to main, with a 60-minute lane timeout (scripts/ci/release-validation-lanes.mjs). A broken relaunch therefore fails the lane after the 15-minute wait instead of passing.commands/resume.tsjust above the deleted block.whyhints, and plainwhynow also prints its sync warnings. x04 and x05 also editdocs/reference/cli.md.configuration.md.Changelog entry
Plain-text
ultrafuzz status(including each--watchpoll),inspect,why,events,node,timeline,snapshots,run,ps,pause,cancel,replay,fork,clean,materialize,validate,doctorandreferences status|sync|updatenow print the warnings of a successful result on stdout after their output, asinitandresumealready did. Examples areWORKFLOW_DEADLINE_CANCEL_FAILEDandWORKFLOW_STATE_SYNC_SKIPPED, which before only--jsonshowed; the exit status is unchanged. The end-to-end campaign test now also checks that the supervisor relaunches a SIGKILLed engine of a fresh run, and that the relaunched engine re-runs the interrupted node.Greptile follow-up
commandFromRuntimeends its text with a newline, so the output is correct. A separator helper would guard only a hypothetical renderer.🤖 Generated with Claude Code