Skip to content

Fence recovery handoffs and preserve artifact generations - #1154

Closed
aviggiano wants to merge 6 commits into
mainfrom
codex/1141-recovery-authority
Closed

aviggiano wants to merge 6 commits into
mainfrom
codex/1141-recovery-authority

Conversation

@aviggiano

Copy link
Copy Markdown
Collaborator

Recovered runs now reevaluate dependency-blocked descendants, keep immutable verified artifact generations addressable across replacement attempts, and fence pause/cancel publication by the current runtime owner. This gives continuation a safer path without weakening artifact verification.

What changed

  • Restore only completed Smithers tasks so previously skipped descendants reevaluate their dependency predicates on resume.
  • Publish immutable, byte-verified artifact generations and retain a current marker only while its generation remains complete and authentic.
  • Pin each admitted consumer to the exact generation it authenticated, so a concurrent replacement cannot revoke or silently replace that authority.
  • Bind generation authority to the marker's complete publication set and enforce file-count, per-file, and aggregate read limits.
  • Archive dynamic artifact-generation roots during explicit source retry.
  • Owner-fence cancellation and paused-state publication, including finalization of a pending cancellation.
  • Add regressions for resume hydration, concurrent A/B generation replacement, duplicate or extra generation entries, marker retention and tamper rejection, dynamic retry archival, and controller handoff fencing.

Validation

  • 243 recovery and integration tests passed.
  • 142 generated-workflow verifier tests passed.
  • pnpm --filter @ultrafuzz/artifacts test
  • pnpm -r --workspace-concurrency=1 typecheck
  • Strict runtime ESLint and full Prettier checks passed.

Mitigates #1141
Mitigates #1142
Mitigates #1153

aviggiano added a commit that referenced this pull request Sep 28, 2026
…d a recovered producer

The resume hydration patches restored durable `skipped` rows as terminal.
A skip is a verdict on prerequisites, and a reset can overturn it: after
`resume --retry-failed` (timetravel --no-deps) re-ran a failed agent, its
verifier and every descendant stayed skipped, so the recovered output was
never verified or consumed (#1141).

Restore only `finished` rows and let the workflow re-derive skips. That
alone (the resume half of #1154) regresses a plain resume: the first
render of a resumed session runs before hydration, so every skip
predicate sees an empty state map and a still-failed agent's verifier
and descendants get dispatched. The scheduler therefore re-renders once
hydration has run, before its first decision.

It also re-renders when a decision pass re-enters after skipping nodes.
Upstream recurses there with the previous render's predicates, so the
direct consumer of a failed agent, and each level below it, ran its
preparation into a failure instead of being skipped: the "retried per
descendant" cascade #1141 also describes. Both re-renders are one patch
at the top of the scheduler's decide().

Resetting the skipped descendants from `--retry-failed` instead is not
possible with Smithers' timetravel: it needs an attempt row for its
target and selects dependents by attempt start time, and a skipped node
has no attempts.

The new integration test drives the pinned Smithers 0.35.0 engine with
the scheduler and hydration patches applied to a private copy of the
Smithers packages: a failed agent's skip cascade executes nothing, a
plain resume executes nothing, and a --retry-failed style reset reopens
the verifier and descendants without re-running finished work. Both
tests fail on main. The string test that pinned the old "restore
finished and skipped" rule is deleted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…d a recovered producer

The resume hydration patches restored durable `skipped` rows as terminal.
A skip is a verdict on prerequisites, and a reset can overturn it: after
`resume --retry-failed` (timetravel --no-deps) re-ran a failed agent, its
verifier and every descendant stayed skipped, so the recovered output was
never verified or consumed (#1141).

Restore only `finished` rows and let the workflow re-derive skips. That
alone (the resume half of #1154) regresses a plain resume: the first
render of a resumed session runs before hydration, so every skip
predicate sees an empty state map and a still-failed agent's verifier
and descendants get dispatched. The scheduler therefore re-renders once
hydration has run, before its first decision.

It also re-renders when a decision pass re-enters after skipping nodes.
Upstream recurses there with the previous render's predicates, so the
direct consumer of a failed agent, and each level below it, ran its
preparation into a failure instead of being skipped: the "retried per
descendant" cascade #1141 also describes. Both re-renders are one patch
at the top of the scheduler's decide().

Resetting the skipped descendants from `--retry-failed` instead is not
possible with Smithers' timetravel: it needs an attempt row for its
target and selects dependents by attempt start time, and a skipped node
has no attempts.

The new integration test drives the pinned Smithers 0.35.0 engine with
the scheduler and hydration patches applied to a private copy of the
Smithers packages: a failed agent's skip cascade executes nothing, a
plain resume executes nothing, and a --retry-failed style reset reopens
the verifier and descendants without re-running finished work. Both
tests fail on main. The string test that pinned the old "restore
finished and skipped" rule is deleted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…d a recovered producer (#1169)

* fix(runtime): resume --retry-failed reopens descendants skipped behind a recovered producer

The resume hydration patches restored durable `skipped` rows as terminal.
A skip is a verdict on prerequisites, and a reset can overturn it: after
`resume --retry-failed` (timetravel --no-deps) re-ran a failed agent, its
verifier and every descendant stayed skipped, so the recovered output was
never verified or consumed (#1141).

Restore only `finished` rows and let the workflow re-derive skips. That
alone (the resume half of #1154) regresses a plain resume: the first
render of a resumed session runs before hydration, so every skip
predicate sees an empty state map and a still-failed agent's verifier
and descendants get dispatched. The scheduler therefore re-renders once
hydration has run, before its first decision.

It also re-renders when a decision pass re-enters after skipping nodes.
Upstream recurses there with the previous render's predicates, so the
direct consumer of a failed agent, and each level below it, ran its
preparation into a failure instead of being skipped: the "retried per
descendant" cascade #1141 also describes. Both re-renders are one patch
at the top of the scheduler's decide().

Resetting the skipped descendants from `--retry-failed` instead is not
possible with Smithers' timetravel: it needs an attempt row for its
target and selects dependents by attempt start time, and a skipped node
has no attempts.

The new integration test drives the pinned Smithers 0.35.0 engine with
the scheduler and hydration patches applied to a private copy of the
Smithers packages: a failed agent's skip cascade executes nothing, a
plain resume executes nothing, and a --retry-failed style reset reopens
the verifier and descendants without re-running finished work. Both
tests fail on main. The string test that pinned the old "restore
finished and skipped" rule is deleted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): re-render after every skip, not only after a pass that only skipped

The skip re-render fired on `depth > 0`, that is only when a decision
pass had changed node states without dispatching anything. When a skip
shared its pass with dispatched work, decide() returned Execute, and the
next decision made without a render (after a retryable failure, or at a
retry deadline) scheduled the skip's dependents from predicates rendered
before the skip. A failed agent's consumer then still ran its
preparation into a failure: on a fresh run, and on a plain resume with
the producer still failed. Both #1169 reviews reproduced this on the
pinned engine.

Mark the predicates stale where the skip happens instead. Upstream's
skipIf branch now sets the same flag resume hydration sets, and the top
of decide() re-renders whenever it is set; the `depth > 0` test is gone.
The cost is at most one redundant render per pass that both skipped and
dispatched, when a completion's own re-render already saw the skip.

The flag is renamed `skipPredicatesStale`, the new anchor is registered
as `skip_marks_predicates_stale`, and `skip_predicate_rerender` no
longer claims to retire with `terminal_state_restore`, because it now
serves both flag setters.

A third integration test adds a task that becomes runnable in the pass
that skips the failed agent's verifier and fails once retryably. It
fails on main and on the previous patch (`prepare:consumer` ends
`failed`) and passes with this one.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano

Copy link
Copy Markdown
Collaborator Author

Superseded. Two independent reviews (both reproduced with the real pinned @smthrs 0.35.0 scheduler) found the #1141 part correct but incomplete — restoring only finished on resume also broke plain resume unless skip predicates are re-evaluated after hydration — and found the immutable artifact-generation store checked a copy that no consumer reads while weakening the existing fail-closed check. The root fix landed in #1169 (restore only finished tasks plus a one-shot re-render after hydration, with a real-engine test). #1153's handoff defect turned out to be in the resume wrapper, fixed in #1177; the engine already fences superseded owners, so the four owner-fence patches were not needed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant