fix(runtime): resume --retry-failed can recover a run whose verifier rejected output - #1205
Merged
aviggiano merged 4 commits intoSep 29, 2026
Merged
Conversation
…rejected output Once a synchronization observed a verifier rejection that the host reproduced, the node was sealed with the task-output-validation-failure disposition, and immutableTerminalFinalization kept that seal for the life of the task. resume --retry-failed and --reset-node rerun the producer and its verifier, but no later synchronization finalized the rerun: the node and the run stayed failed even when the verifier accepted the rerun's output, and the attempt ledger recorded the accepted attempt as failed. A reset restarts Smithers' attempt numbering, so the rerun carries the same task ID and attempt number as the sealed record. Tell the two apart by time instead: evidence of an occurrence that started after the record finished is a later occurrence. It lifts an invalid-output seal (a successful publication stays immutable), and workflowEvidenceSupersedesPrevious treats it as new even when status, task and attempt match, so a rerun that is rejected again is finalized with its own error and disposition. The occurrence's start is compared, not its finish, because a later terminal event of the same occurrence, such as a trailing NodeCancelled for an attempt that already ended, moves the finish but not the start. Files written after a rejection therefore still cannot change its verdict. Lifting the seal when the rerun starts also stops a sync during the rerun from recording the rerun's producer attempt with the old failed status. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…occurrence Smithers stamps a run cancellation's NodeCancelled with an instant taken before its transaction, so a verifier that started meanwhile is recorded as finishing before it started. Comparing a later start with the recorded finish alone then took the recorded occurrence for a newer one: a sealed output-validation failure was re-finalized from files written after its verdict, and an unsealed failure was re-finalized on every sync, each time with an error-severity ARTIFACT_VERIFIER_FAILED. A later occurrence now has to start after both the recorded start and the recorded finish. Co-Authored-By: Claude Opus 5.5 <[email protected]>
After a reset, Smithers re-inserts the verifier's node row as pending, so until the rerun's verifier starts, the only start that identifies the rerun is its producer's. A sync between the rerun producer's NodeFinished and its verifier's NodeStarted therefore kept the seal: the node read failed, and the accepted producer attempt was appended to the attempt ledger as failed, which no later sync corrects. The producer's start now counts alongside the verifier's. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…dation failure The test named for keeping an immutable output-validation failure now takes the opposite path: its replacement verifier starts after the hand-sealed finish, so the replacement is judged on its own output and sealed again. It passed only because that judgment matches the old one. Name it for what it checks and assert the re-derived finish time and disposition. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 29, 2026
Collaborator
Author
|
Real-run check (main 73e70f3 + #1204 + #1205, tiny-vault smoke campaign Before this, the campaign had ended
The resume also ran on code newer than the run's launch build, which #1188 allows. |
aviggiano
deleted the
claude/x02-retry-failed-recovers-verifier-failures
branch
September 29, 2026 15:18
aviggiano
added a commit
that referenced
this pull request
Sep 29, 2026
… through slow CI setup (#1214) `model-fanout dynamic nodes stay pending until every generated attempt has evidence` (`packages/runtime/test/dynamic-lifecycle.test.ts`) fails intermittently in the `runtime-supporting` lane. It failed on `main` at 142ba80 and on #1205: ``` AssertionError [ERR_ASSERTION]: Expected values to be strictly equal: + 'controller-loss' - 'dependency' ``` Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 29, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When a task's verifier rejects its output and the host reproduces the rejection, the first synchronization that observes it seals the node. That can be any
status,inspect,whyorstatscall, or the eval poller. The node becomesfailedwith terminal dispositiontask-output-validation-failure. The documented recovery isultrafuzz resume --retry-failed, which resets every failed node, "retrying a failed artifact verifier from its agent producer" (docs/how-to/restart-continue.md). It reruns the producer and the verifier, but the run could never endsucceeded.On the pinned engine (ad hoc probe,
origin/main2cacf4c). The probe is adapted frompackages/cli/test/e2e/campaign-resume.test.ts: everyultrafuzzcommand is its own process, and the agent is a stubcodex.summarize'ssummary.txt, and its verifier rejects the output.ultrafuzz statusseals the node.ultrafuzz resume <run> --retry-failedrerunssummarize, this time with its output. The engine finishes the run (RunFinished), and the final-report agent runs.ultrafuzz statusstill reports the runfailed, withsummarizefailedunder the old disposition.final-reportfails too, withARTIFACT_VERIFICATION_AUTHORITY_INVALID, and the report isunavailable. So the recovered campaign ends without its report.attempts.jsonlrecords the accepted rerun ofsummarizeasfailed.With the fake runner (ad hoc, same base):
findings.jsonis missing. OnesyncRunseals the node.resumeRun({ retryFailed: true })issuestimetravel … --node-id node:project-discovery.syncRunstill returns runfailedand nodefailed, with the old disposition and no artifact manifest.attempts.jsonlrecords the rerun's accepted producer attempt asfailed.--reset-node node:project-discoveryends the same way. That holds after a verifier rejection, and also after the host rejects output that the verifier accepted.failed. It also writes the rerun's producer attempt to the append-only ledger asfailed, before any verdict.Reach.
succeededwhen no sync ran in between.final-reporthas optional dependencies, so the rejection was not sealed. Its retry was synchronized normally: the node wentfailed→running→failed, andattempts.jsonlholds two failed attempts. The rerun failed the same report semantic gate again, which is a different problem.Root cause
immutableTerminalFinalization(workflow-sync.ts) kept an invalid-output disposition for the life of the task. The seal only has to stop one thing: a rejected occurrence being finalized again from files written after its verdict.--retry-failedand--reset-noderestart Smithers' attempt numbering (docs/reference/artifacts-reports.md). The pinned Smithers time-travel package says so too (resetCancelMarker.js: "a deliberate reset restarts at attempt 1"). So the rerun has the same task ID and attempt number as the sealed record.workflowEvidenceSupersedesPrevioushas the same blind spot, for every terminal record. It treats an equal status, task and attempt as the same occurrence. So a rerun that ends with the recorded status, with no sync while it ran, was never finalized, and the node kept the earlier error.Change
workflow-sync.ts: +43/−10 lines; 17 of the added lines are comments.startedAfterRecordedOccurrence(previous, attemptEvidence). The evidence belongs to a later occurrence when its verifier's or its producer'sNodeStartedtime is after both the record'sstarted_atand itsfinished_at.completionEvidenceForTaskrecords it asagentStartedAt, next to the producer's attempt number.immutableTerminalFinalizationno longer keeps an invalid-output disposition against a later occurrence. Successful publications stay immutable, as before.workflowEvidenceSupersedesPrevioustreats a later occurrence as new evidence even when status, task and attempt match. The rerun is finalized, so the node gets the rerun's error and disposition.finishedAtbut keeps its start, and Smithers emits such events: aNodeCancelledfor an attempt that already ended (Attempt and cancellation ledgers are incomplete after recovery #1139).pending(pinned@smthrs/time-travel,timetravel.js,buildPendingNode). Until the rerun's verifier starts, the producer's start is the only sign of the rerun. Without it, a sync between the rerun producer'sNodeFinishedand its verifier'sNodeStartedkept the seal, and appended the accepted producer attempt to the ledger asfailed, which no later sync corrects.NodeCancelledwith an instant taken before its transaction. So a verifier that started meanwhile is recorded as finishing before it started;terminalWorkflowAttemptsalready documents this. Compared with the recorded finish alone, that record's own occurrence looked newer:ARTIFACT_VERIFIER_FAILED. By code reading, that makesgetRunHealth, and so everyultrafuzz status, report not ok.runningwhile a task of the rerun runs, andpendingbetween the producer finishing and the verifier starting. Between the reset and the producer's start it still reads the old verdict. The ledger records the rerun's producer attempt only with the rerun's verdict.Tests (
runtime.test.ts).resume --retry-failed recovers a node whose verifier rejected its outputdrives the fake runner throughstartRun,syncRunandresumeRun. Every occurrence reuses attempt 1, as a Smithers reset does.NodeCancelledfor the same attempt appears. The node record is unchanged (deepEqual), and no manifest is written.--retry-failed. The rerun is rejected again, with no sync in between. The node now carries the rerun's error and a re-derived disposition.--retry-failed. A sync after the rerun's producer finished, with its verifier stillpending, shows the nodependingand leaves the ledger at two rows. The verifier then accepts. Run and node endsucceeded, the manifest is written,retry_countis 2, and the ledger readsfailed,failed,succeeded.No pass reports an attempt-bookkeeping diagnostic.
syncRun keeps a verifier occurrence whose cancellation is stamped before its start(new). The verifier'sNodeStartedis at +400 ms, and its run-cancellationNodeCancelledis stamped at +300 ms. It runs two cases:findings.jsonis written;findings.jsonis removed.In both, the next sync leaves the record unchanged and reports no error-severity diagnostic.
syncRun judges a replacement that started after a sealed output-validation failure on its own output(renamed fromsyncRun keeps an immutable output-validation failure when its successful occurrence is superseded). Its replacement verifier starts at +800 ms, after the hand-sealed finish at +350 ms. With this change the replacement is judged and sealed again, at +900 ms. The old test still passed only because that judgment matches the hand-sealed one. It now asserts the re-derivedfinished_atand disposition, and its comment no longer says the seal carries no finalization diagnostic.Deliberately not built
resume(the other candidate fix).--retry-failedresets a failed verifier's producer together with all its dependents.--retry-failedand--reset-nodealike.finishedAt, as the audit's 8-line prototype did. A trailingNodeCancelledfor the same attempt then looks like a newer occurrence. In the matrix below, the rule keyed on the finish finalizes the rejected occurrence again, as "workflow task was cancelled".failed.--reset-nodeof a node that succeeded. This is unchanged. Republishing a success rewrites the manifest that its dependents' prerequisite digests point at, which is a separate change.--retry-failedrerun a node whose output the host rejected after its verifier accepted. Smithers reports that task finished, so--retry-faileddoes not reset it.--reset-nodedoes, and with this change its rerun counts (probe below).Verification
Matrix. Each
workflow-sync.tsvariant was compiled into its owndist-testcopy, with the committed test file. Base isorigin/main2cacf4c.workflow-sync.tsorigin/mainlast_erroris still the first rejectionfinished_atstays at +350 msfailed, expectedpendingfailed, expectedpendingevidence.finishedAtlast_erroris still the first rejectionfailed, expectedpendingorigin/main, which never lifts a seal. It pins the same-occurrence behavior that the previous head broke.origin/mainand fails on the previous head and on the variant without the bound: there the unsealed record is sealed from the removed file.Pinned-engine probe (ad hoc, described under Problem). It ran once on this revision, under
eatmydataon a loaded host, in 348 s.summarizeat attempt 1. Its verifier started at 13:15:16.134Z, after the sealed record'sfinished_at(13:12:53.388Z).statusreports runsucceeded, verdictdone, and reportavailable,completeandverified.summarizeissucceededwithretry_count1. Its ledger rows are attempt 1failed(sequence 51) and attempt 1succeeded(sequence 79).statsshows all three nodessucceeded.origin/main(run for the previous revision): runfailedand reportunavailable(report-agent-output-unavailable).summarizeandfinal-reportarefailed, and the ledger holds bothsummarizeattempts asfailed.--reset-nodeprobes (ad hoc, fake runner), on this revision:succeeded, the ledger readsfailed,succeeded, and there are no warnings. Onorigin/main(run for the previous revision), both endfailed, with the ledger atfailed,failed.NodeCancelledbefore the reset: the sealed record stays unchanged until--reset-node.Reviewers' window and skew probes (ad hoc, fake runner), on this revision:
NodeFinishedand its verifier'sNodeStartedleaves the nodependingand the ledger unchanged. After the verifier accepts, run and node aresucceeded, and the ledger readsfailed,succeeded.Targeted
runtime.test.tssweep on this revision, in 8 shards. It selects every test whose name starts withsyncRun,resume,ordinary resume,native continuation,statusorgetRunHealth, or contains immutable, disposition, output-validation, retry, reset, rerun, supersed, reused, occurrence, verifier, re-finaliz or "invalid output". 145 tests were selected: 141 pass, none fail, and 4 Bun-lane adapter contracts are skipped under Node. They include the three tests above and the existing immutability, supersession, reset and #1099 attempt-reuse tests.Supporting test files:
dynamic-lifecycleandworkflow-control. 33/33 pass.Static checks, all passing:
npx prettier --checkandnpx eslinton both changed files;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;pnpm -w lint,pnpm --filter @ultrafuzz/runtime typecheckandpnpm -w knip.Complexity.
synchronizeTasksstays at 66.completionEvidenceForTaskgoes from 11 to 12,immutableTerminalFinalizationfrom 6 to 7 andworkflowEvidenceSupersedesPreviousfrom 4 to 5; the new helper measures 4. The global ceiling (83) is unchanged.Not run: the rest of
runtime.test.ts, the other test files and packages,--reset-nodeon the pinned engine, and a campaign with a real model.Risk / compatibility
NodeStartedtimes against the recordedstarted_atandfinished_at. If the runner's clock steps back by more than the time between the verdict and the rerun's start, the old seal stays, which is the behavior before this change.WORKFLOW_EVENTS_TRUNCATEDwhen it hits the limit. If a rerun's start falls past the limit, the old seal stays. If only its terminal event does, the record keeps the earlier finish time, so it shows a finish before its start; the recorded start keeps later syncs from finalizing it again.synchronizeTaskschanges in two call lines. The other source hunks are one field ofAttemptWorkflowEvidence, one statement incompletionEvidenceForTask, and the helpers near the bottom of the file.resume --retry-failed after a pre-agent failure keeps synchronizing the reused attempt.git merge-treefinds no conflict with the other branches of this wave (x01, x03, x04, x05), or with the open PR branches fix(runtime): record final-report producer selections in the run instead of querying smithers #1183, refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197, feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198, feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201, fix(modal): keep the model-work flag when a model node has no run-state record #1202 and test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203.origin/mainin five files, and this branch adds none.Changelog entry
ultrafuzz resume --retry-failedand--reset-nodenow recover a node whose verifier rejected its output, and--reset-nodealso recovers one whose output the host rejected after its verifier accepted it: the rerun is judged on its own output, so the node and the run can endsucceededand publish their report. Before, once any status call had observed the rejection, the node kept its old verdict even after the engine finished the rerun, so the run endedfailedand a final report that depended on the node was not published. Output written after a rejection, without a rerun, still does not change the verdict.Refs #1141.
🤖 Generated with Claude Code
The PR appears safe to merge; no actionable regression was established.
Summary
The PR uses producer and verifier start times to distinguish a rerun from a sealed output-validation failure, allowing the rerun to receive its own verdict.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Verifier output rejected] --> B[Sync seals failed occurrence] B --> C[Producer starts after recorded occurrence] C --> D[Sync follows rerun] D --> E{Rerun verdict} E -->|Accepted| F[Publish artifacts and succeed] E -->|Rejected| G[Record new failure]Reviews (1) · Last reviewed commit: "test(runtime): assert the re-judged reco..."