Skip to content

fix(runtime): resume --retry-failed can recover a run whose verifier rejected output - #1205

Merged
aviggiano merged 4 commits into
mainfrom
claude/x02-retry-failed-recovers-verifier-failures
Sep 29, 2026
Merged

aviggiano merged 4 commits into
mainfrom
claude/x02-retry-failed-recovers-verifier-failures

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

When a task's verifier rejects its output and the host reproduces the rejection, the first synchronization that observes it seals the node. That can be any status, inspect, why or stats call, or the eval poller. The node becomes failed with terminal disposition task-output-validation-failure. The documented recovery is ultrafuzz resume --retry-failed, which resets every failed node, "retrying a failed artifact verifier from its agent producer" (docs/how-to/restart-continue.md). It reruns the producer and the verifier, but the run could never end succeeded.

On the pinned engine (ad hoc probe, origin/main 2cacf4c). The probe is adapted from packages/cli/test/e2e/campaign-resume.test.ts: every ultrafuzz command is its own process, and the agent is a stub codex.

  • The stub leaves out summarize's summary.txt, and its verifier rejects the output. ultrafuzz status seals the node.
  • ultrafuzz resume <run> --retry-failed reruns summarize, this time with its output. The engine finishes the run (RunFinished), and the final-report agent runs.
  • ultrafuzz status still reports the run failed, with summarize failed under the old disposition.
  • final-report fails too, with ARTIFACT_VERIFICATION_AUTHORITY_INVALID, and the report is unavailable. So the recovered campaign ends without its report.
  • attempts.jsonl records the accepted rerun of summarize as failed.

With the fake runner (ad hoc, same base):

  • The verifier rejects the output: the required findings.json is missing. One syncRun seals the node.
  • resumeRun({ retryFailed: true }) issues timetravel … --node-id node:project-discovery.
  • The rerun writes valid output, and its verifier accepts it.
  • The next syncRun still returns run failed and node failed, with the old disposition and no artifact manifest. attempts.jsonl records the rerun's accepted producer attempt as failed.
  • --reset-node node:project-discovery ends the same way. That holds after a verifier rejection, and also after the host rejects output that the verifier accepted.
  • A sync after the rerun's producer finishes, before its verifier's verdict, shows the node failed. It also writes the rerun's producer attempt to the append-only ledger as failed, before any verdict.

Reach.

  • A rejection is sealed only when the host re-derives it. That is either a failed verifier whose failure the host's artifact gate reproduces, or a verifier success that the host rejects. The host re-runs the gate after a verifier failure only for nodes without optional dependencies.
  • A sync has to observe the rejection before the retry. The audit that found this bug (sync-state map, failure mode "A successful retry cannot clear an output-validation failure once it is sealed") showed on its base that the same evidence ends succeeded when no sync ran in between.
  • smoke-main-2 is not an instance. Its final-report has optional dependencies, so the rejection was not sealed. Its retry was synchronized normally: the node went failed → running → failed, and attempts.jsonl holds two failed attempts. The rerun failed the same report semantic gate again, which is a different problem.

Root cause

  • The seal outlived the verdict. immutableTerminalFinalization (workflow-sync.ts) kept an invalid-output disposition for the life of the task. The seal only has to stop one thing: a rejected occurrence being finalized again from files written after its verdict.
  • Synchronization could not tell a rerun from the sealed occurrence. --retry-failed and --reset-node restart Smithers' attempt numbering (docs/reference/artifacts-reports.md). The pinned Smithers time-travel package says so too (resetCancelMarker.js: "a deliberate reset restarts at attempt 1"). So the rerun has the same task ID and attempt number as the sealed record.
  • workflowEvidenceSupersedesPrevious has the same blind spot, for every terminal record. It treats an equal status, task and attempt as the same occurrence. So a rerun that ends with the recorded status, with no sync while it ran, was never finalized, and the node kept the earlier error.

Change

workflow-sync.ts: +43/−10 lines; 17 of the added lines are comments.

  • New helper, startedAfterRecordedOccurrence(previous, attemptEvidence). The evidence belongs to a later occurrence when its verifier's or its producer's NodeStarted time is after both the record's started_at and its finished_at.
  • The producer's start travels with the evidence. completionEvidenceForTask records it as agentStartedAt, next to the producer's attempt number.
  • immutableTerminalFinalization no longer keeps an invalid-output disposition against a later occurrence. Successful publications stay immutable, as before.
  • workflowEvidenceSupersedesPrevious treats a later occurrence as new evidence even when status, task and attempt match. The rerun is finalized, so the node gets the rerun's error and disposition.
  • Why starts, not the finish. A later terminal event of the same occurrence moves finishedAt but keeps its start, and Smithers emits such events: a NodeCancelled for an attempt that already ended (Attempt and cancellation ledgers are incomplete after recovery #1139).
  • Why the producer's start counts. A reset re-inserts the verifier's node row as pending (pinned @smthrs/time-travel, timetravel.js, buildPendingNode). Until the rerun's verifier starts, the producer's start is the only sign of the rerun. Without it, a sync between the rerun producer's NodeFinished and its verifier's NodeStarted kept the seal, and appended the accepted producer attempt to the ledger as failed, which no later sync corrects.
  • Why the bound includes the recorded start. Smithers stamps a run cancellation's NodeCancelled with an instant taken before its transaction. So a verifier that started meanwhile is recorded as finishing before it started; terminalWorkflowAttempts already documents this. Compared with the recorded finish alone, that record's own occurrence looked newer:
    • a sealed record was finalized again from files written after its verdict;
    • an unsealed one was finalized again on every sync, each time with an error-severity ARTIFACT_VERIFIER_FAILED. By code reading, that makes getRunHealth, and so every ultrafuzz status, report not ok.
  • When the seal lifts. When the rerun's producer starts. From then on the node reads what an unsealed node reads: running while a task of the rerun runs, and pending between the producer finishing and the verifier starting. Between the reset and the producer's start it still reads the old verdict. The ledger records the rerun's producer attempt only with the rerun's verdict.
  • Comment. The immutable-path comment now says the disposition holds until the task runs again.

Tests (runtime.test.ts).

resume --retry-failed recovers a node whose verifier rejected its output drives the fake runner through startRun, syncRun and resumeRun. Every occurrence reuses attempt 1, as a Smithers reset does.

  1. The verifier rejects, and a sync seals the node.
  2. Valid files and a fresh verifier marker are written, and a trailing NodeCancelled for the same attempt appears. The node record is unchanged (deepEqual), and no manifest is written.
  3. --retry-failed. The rerun is rejected again, with no sync in between. The node now carries the rerun's error and a re-derived disposition.
  4. --retry-failed. A sync after the rerun's producer finished, with its verifier still pending, shows the node pending and leaves the ledger at two rows. The verifier then accepts. Run and node end succeeded, the manifest is written, retry_count is 2, and the ledger reads failed, failed, succeeded.

No pass reports an attempt-bookkeeping diagnostic.

syncRun keeps a verifier occurrence whose cancellation is stamped before its start (new). The verifier's NodeStarted is at +400 ms, and its run-cancellation NodeCancelled is stamped at +300 ms. It runs two cases:

  • output rejected, so the node is sealed, and then findings.json is written;
  • output accepted, so the node is failed without a seal, and then findings.json is removed.

In both, the next sync leaves the record unchanged and reports no error-severity diagnostic.

syncRun judges a replacement that started after a sealed output-validation failure on its own output (renamed from syncRun keeps an immutable output-validation failure when its successful occurrence is superseded). Its replacement verifier starts at +800 ms, after the hand-sealed finish at +350 ms. With this change the replacement is judged and sealed again, at +900 ms. The old test still passed only because that judgment matches the hand-sealed one. It now asserts the re-derived finished_at and disposition, and its comment no longer says the seal carries no finalization diagnostic.

Deliberately not built

  • Resetting Ultrafuzz node records to pending from resume (the other candidate fix).
    • It needs its own copy of Smithers' reset scope: --retry-failed resets a failed verifier's producer together with all its dependents.
    • It adds a second writer of node records on the lifecycle path. Synchronizations do not take the lifecycle lock.
    • By code reading, a pending record that meets Smithers evidence still showing the old rejection finalizes that rejection again from the current files. That happens when a reset fails part-way, or when a sync runs before it. Preventing that is exactly what the seal is for.
    • The sync-side rule needs no bookkeeping, and it covers --retry-failed and --reset-node alike.
  • Comparing finishedAt, as the audit's 8-line prototype did. A trailing NodeCancelled for the same attempt then looks like a newer occurrence. In the matrix below, the rule keyed on the finish finalizes the rejected occurrence again, as "workflow task was cancelled".
  • Lifting the seal only on the rerun's terminal event. A sync after the rerun's producer finished, before its verifier's verdict, would then keep the seal and write that producer attempt to the append-only ledger as failed.
  • Letting a later occurrence replace a successful publication, for example after --reset-node of a node that succeeded. This is unchanged. Republishing a success rewrites the manifest that its dependents' prerequisite digests point at, which is a separate change.
  • Making --retry-failed rerun a node whose output the host rejected after its verifier accepted. Smithers reports that task finished, so --retry-failed does not reset it. --reset-node does, and with this change its rerun counts (probe below).

Verification

Matrix. Each workflow-sync.ts variant was compiled into its own dist-test copy, with the committed test file. Base is origin/main 2cacf4c.

workflow-sync.ts retry test cancellation-skew test renamed test
origin/main fails at step 3: last_error is still the first rejection passes fails: finished_at stays at +350 ms
previous head of this branch (2b15ca4) fails at step 4: node failed, expected pending fails: the sealed case loses its disposition passes
this change without the producer's start fails at step 4: node failed, expected pending passes passes
this change without the recorded-start bound passes fails: the sealed case loses its disposition passes
this change keyed on evidence.finishedAt fails at step 2: finalized again as "workflow task was cancelled" passes passes
this change without the supersession clause fails at step 3: last_error is still the first rejection passes passes
this change, lifting the seal only on terminal evidence fails at step 4: node failed, expected pending passes passes
this change passes passes passes
  • The cancellation-skew test passes on origin/main, which never lifts a seal. It pins the same-occurrence behavior that the previous head broke.
  • Its unsealed case was also run alone, by patching the compiled loop. It passes on origin/main and fails on the previous head and on the variant without the bound: there the unsealed record is sealed from the removed file.

Pinned-engine probe (ad hoc, described under Problem). It ran once on this revision, under eatmydata on a loaded host, in 348 s.

  • The engine reran summarize at attempt 1. Its verifier started at 13:15:16.134Z, after the sealed record's finished_at (13:12:53.388Z).
  • status reports run succeeded, verdict done, and report available, complete and verified.
  • summarize is succeeded with retry_count 1. Its ledger rows are attempt 1 failed (sequence 51) and attempt 1 succeeded (sequence 79). stats shows all three nodes succeeded.
  • On origin/main (run for the previous revision): run failed and report unavailable (report-agent-output-unavailable). summarize and final-report are failed, and the ledger holds both summarize attempts as failed.

--reset-node probes (ad hoc, fake runner), on this revision:

  • After a verifier rejection, and after the host rejected output the verifier had accepted: run and node end succeeded, the ledger reads failed, succeeded, and there are no warnings. On origin/main (run for the previous revision), both end failed, with the ledger at failed, failed.
  • A reviewer's probe of the host-rejection case also adds valid files and a trailing NodeCancelled before the reset: the sealed record stays unchanged until --reset-node.

Reviewers' window and skew probes (ad hoc, fake runner), on this revision:

  • A sync between the rerun producer's NodeFinished and its verifier's NodeStarted leaves the node pending and the ledger unchanged. After the verifier accepts, run and node are succeeded, and the ledger reads failed, succeeded.
  • A sealed record whose cancellation is stamped before its start stays unchanged after its files change, and the second and third syncs report no diagnostics.

Targeted runtime.test.ts sweep on this revision, in 8 shards. It selects every test whose name starts with syncRun, resume, ordinary resume, native continuation, status or getRunHealth, or contains immutable, disposition, output-validation, retry, reset, rerun, supersed, reused, occurrence, verifier, re-finaliz or "invalid output". 145 tests were selected: 141 pass, none fail, and 4 Bun-lane adapter contracts are skipped under Node. They include the three tests above and the existing immutability, supersession, reset and #1099 attempt-reuse tests.

Supporting test files: dynamic-lifecycle and workflow-control. 33/33 pass.

Static checks, all passing:

  • npx prettier --check and npx eslint on both changed files;
  • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;
  • pnpm -w lint, pnpm --filter @ultrafuzz/runtime typecheck and pnpm -w knip.

Complexity. synchronizeTasks stays at 66. completionEvidenceForTask goes from 11 to 12, immutableTerminalFinalization from 6 to 7 and workflowEvidenceSupersedesPrevious from 4 to 5; the new helper measures 4. The global ceiling (83) is unchanged.

Not run: the rest of runtime.test.ts, the other test files and packages, --reset-node on the pinned engine, and a campaign with a real model.

Risk / compatibility

Changelog entry

ultrafuzz resume --retry-failed and --reset-node now recover a node whose verifier rejected its output, and --reset-node also recovers one whose output the host rejected after its verifier accepted it: the rerun is judged on its own output, so the node and the run can end succeeded and publish their report. Before, once any status call had observed the rejection, the node kept its old verdict even after the engine finished the rerun, so the run ended failed and a final report that depended on the node was not published. Output written after a rejection, without a rerun, still does not change the verdict.

Refs #1141.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no actionable regression was established.

Summary

The PR uses producer and verifier start times to distinguish a rerun from a sealed output-validation failure, allowing the rerun to receive its own verdict.

  • Adds recovery and cancellation-timestamp regression tests.
  • Preserves the seal when files change without a rerun.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Verifier output rejected] --> B[Sync seals failed occurrence]
  B --> C[Producer starts after recorded occurrence]
  C --> D[Sync follows rerun]
  D --> E{Rerun verdict}
  E -->|Accepted| F[Publish artifacts and succeed]
  E -->|Rejected| G[Record new failure]
Loading

Reviews (1) · Last reviewed commit: "test(runtime): assert the re-judged reco..."

aviggiano and others added 4 commits September 29, 2026 11:06
…rejected output

Once a synchronization observed a verifier rejection that the host
reproduced, the node was sealed with the task-output-validation-failure
disposition, and immutableTerminalFinalization kept that seal for the life
of the task. resume --retry-failed and --reset-node rerun the producer and
its verifier, but no later synchronization finalized the rerun: the node
and the run stayed failed even when the verifier accepted the rerun's
output, and the attempt ledger recorded the accepted attempt as failed.

A reset restarts Smithers' attempt numbering, so the rerun carries the same
task ID and attempt number as the sealed record. Tell the two apart by
time instead: evidence of an occurrence that started after the record
finished is a later occurrence. It lifts an invalid-output seal (a
successful publication stays immutable), and workflowEvidenceSupersedesPrevious
treats it as new even when status, task and attempt match, so a rerun that
is rejected again is finalized with its own error and disposition.

The occurrence's start is compared, not its finish, because a later
terminal event of the same occurrence, such as a trailing NodeCancelled for
an attempt that already ended, moves the finish but not the start. Files
written after a rejection therefore still cannot change its verdict.
Lifting the seal when the rerun starts also stops a sync during the rerun
from recording the rerun's producer attempt with the old failed status.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…occurrence

Smithers stamps a run cancellation's NodeCancelled with an instant taken
before its transaction, so a verifier that started meanwhile is recorded as
finishing before it started. Comparing a later start with the recorded
finish alone then took the recorded occurrence for a newer one: a sealed
output-validation failure was re-finalized from files written after its
verdict, and an unsealed failure was re-finalized on every sync, each time
with an error-severity ARTIFACT_VERIFIER_FAILED.

A later occurrence now has to start after both the recorded start and the
recorded finish.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
After a reset, Smithers re-inserts the verifier's node row as pending, so
until the rerun's verifier starts, the only start that identifies the rerun
is its producer's. A sync between the rerun producer's NodeFinished and its
verifier's NodeStarted therefore kept the seal: the node read failed, and
the accepted producer attempt was appended to the attempt ledger as failed,
which no later sync corrects.

The producer's start now counts alongside the verifier's.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…dation failure

The test named for keeping an immutable output-validation failure now takes
the opposite path: its replacement verifier starts after the hand-sealed
finish, so the replacement is judged on its own output and sealed again.
It passed only because that judgment matches the old one. Name it for what
it checks and assert the re-derived finish time and disposition.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano

Copy link
Copy Markdown
Collaborator Author

Real-run check (main 73e70f3 + #1204 + #1205, tiny-vault smoke campaign smoke-main-2)

Before this, the campaign had ended failed twice, both times with Report: unavailable. Its verify:final-report rejected the agent's report because the report reworded one finding's proof_of_concept.

  1. ultrafuzz status with this build, before any retry: Report: available, completion PARTIAL, verification not-checked. The rejected report is published unchecked from its workspace copy (fix: a reworded finding no longer discards the final report #1204).
  2. ultrafuzz resume smoke-main-2 --retry-failed re-ran only the final report: its producer and verifier (fix(runtime): resume --retry-failed can recover a run whose verifier rejected output #1205). Four minutes later the run was done (succeeded), with report completion COMPLETE and verification verified.
  3. The fresh report agent again reworded issue 1's proof of concept. This time it is recorded as a warning, not a lost report:
    warning: ARTIFACT_SEMANTIC_GATE_WARNING: Bounded report did not preserve dedupe field "proof_of_concept" at artifacts/final-report/report.json#$.issues[1].proof_of_concept

The resume also ran on code newer than the run's launch build, which #1188 allows.

@aviggiano
aviggiano merged commit 0786563 into main Sep 29, 2026
28 of 30 checks passed
@aviggiano
aviggiano deleted the claude/x02-retry-failed-recovers-verifier-failures branch September 29, 2026 15:18
aviggiano added a commit that referenced this pull request Sep 29, 2026
… through slow CI setup (#1214)

`model-fanout dynamic nodes stay pending until every generated attempt has evidence` (`packages/runtime/test/dynamic-lifecycle.test.ts`) fails intermittently in the `runtime-supporting` lane. It failed on `main` at 142ba80 and on #1205:
```
AssertionError [ERR_ASSERTION]: Expected values to be strictly equal:
+ 'controller-loss'
- 'dependency'
```

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant