Skip to content

fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code - #1171

Merged
aviggiano merged 5 commits into
mainfrom
claude/w08-retry-policy
Sep 29, 2026
Merged

aviggiano merged 5 commits into
mainfrom
claude/w08-retry-policy

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Root cause

  1. Retry wait. compileTask sets retryPolicy: { backoff: "exponential", initialDelayMs: 1_000 }, so retries waited 1 s, then 2 s. Claude Code's refresh lock can take about a minute to clear.
  2. Stall verdict. The agent Task inherited Smithers' default maxIdenticalFailures of 3. That verdict overrides Ultrafuzz's explicit chain: identical failures ended the chain before later same-agent attempts (exhaustive plans five) or any [retry].agents fallback profile ran.
  3. Admission failures looked retryable. They re-read the same producer bytes but were thrown as plain Errors, so Smithers retried them. preparationStep also discarded details when it rewrapped an error.
  4. Timeouts inferred from text. errorLooksLikeTimeout ran /timeout|timed out|heartbeat/ over every string in the NodeFailed error, including stack and causes. That is a failure taxonomy built from free text, which Add bounded error-agnostic agent retries with backoff and optional model fallback #572 rules out. Since fix(runtime): let the smoke lane reach a scoreable result #1027, every JSON-validator preflight failure message mentions "timeout". The agent-error normalizer also stripped Smithers' PROCESS_TIMEOUT and PROCESS_IDLE_TIMEOUT codes, which left that regex as the only thing labelling real agent CLI deadlines.

Change

  1. Retry wait and full chain (smithers.ts, workflow template)
    • Agent retries now wait 60 s, 120 s and 240 s, then Smithers' 300 s cap.
    • The agent Task gets maxIdenticalFailures: 0, so identical failures no longer end the planned chain early.
    • maxIdenticalFailures is set only in the template. The compiled manifest, cloud handoff schema and sealed task documents keep their exact shape.
    • The real smithers graph smoke test never ran: it needed a workspace-root .smithers install that no checkout has, and it could not resolve the sealed module paths. It now uses the runtime package's own Smithers and asserts the retryPolicy Smithers receives.
  2. Classify deterministic failures correctly and stop retry cascades #1144 (workflow template, workflow-sync.ts)
    • A tiny nonRetryableFailure helper sets details.failureRetryable = false. Smithers already honors that flag.
    • The admission entry points throw through it: preparation's assertTaskDependencyInputs, and the agent's pre- and post-generation assertDependencyArtifactAdmissionCurrent rechecks. Both wrap their unchanged bodies.
    • preparationStep keeps the flag when it rewraps. Failures outside admission stay retryable, including the Concurrent prepare:* worktree tasks fail nondeterministically with a Bun-erased TypeError: undefined is not an object (evaluating 'get') #672 engine-boundary TypeError.
    • errorLooksLikeTimeout is replaced by a check of the top-level typed code (TASK_TIMEOUT, TASK_HEARTBEAT_TIMEOUT, PROCESS_TIMEOUT, PROCESS_IDLE_TIMEOUT). The existing TaskHeartbeatTimeout event rule stays.
    • The normalizer allowlist now keeps PROCESS_TIMEOUT and PROCESS_IDLE_TIMEOUT.
  3. Docs from review. docs/config.md documents the retry timing, what can still end a chain early (a failure Smithers classifies as non-retryable, or a quota pause), and that admission failures are not retried, including file-system and validator errors during admission. docs/reference/artifacts-reports.md states the timeout-labelling rule next to the node statuses.
  4. CI. <!-- markdownlint-disable-file MD013 --> at the end of CHANGELOG.md, byte-identical to perf(runtime): halve launch fsync work, and doctor stops failing on unused agent profiles #1179's. Without it, super-linter fails every PR that touches CHANGELOG.md, release-gates fails, and every full release-validation lane is skipped. That happened on this PR's first push.

Deliberately not built (and why)

Verification

Each test below fails with origin/main (b6dd1da9) source and passes at this PR's head. I re-ran the node tests against main after the review changes. I copied this PR's test files into a fresh origin/main worktree, compiled them, and ran them there. The same results were reproduced independently by the implementer and both reviews.

Test On origin/main source
runtime.test.ts › compiled Smithers workflow passes a real non-executing graph smoke actual {backoff:"exponential", initialDelayMs:1000}, expected {…, initialDelayMs:60000, maxIdenticalFailures:0}
runtime.test.ts › compileSmithersWorkflow exhausts same-profile retries before ordered fallback initialDelayMs 1000 vs 60000
runtime.test.ts › syncRun maps failed workflow nodes into durable failed run state, with realistic payloads: a TASK_HEARTBEAT_TIMEOUT code on attempt 1 and the #1144 timeout-worded failure on attempt 2 node status timed-out, expected failed
generated-workflow-verifier › optional admission rejects a present malformed marker…: real admission code, marked non-retryable details undefined, expected { failureRetryable: false }
generated-workflow-verifier › dependency admission retains one exact snapshot epoch…: the agent recheck is marked non-retryable details undefined, expected { failureRetryable: false }
generated-workflow-verifier › agent failure normalization preserves only validated Smithers recovery controls: CLI deadline codes survive code undefined, expected PROCESS_TIMEOUT
smithers-dependency-skip.integration › a dependency admission failure fails its preparation once while a transient failure is retried. Real Smithers 0.35.0 runs the template's own preparationStep and helper. Fails only because the helper is missing. Against this PR's template, with preparationStep dropping the flag or with the helper marking nothing, prepare:consumer ends stalled instead of failed. Unmodified, it passes.

Experiment, not committed. The implementer ran this and one review reproduced it. In real Smithers 0.35.0, a [fail, fail, fail, fallback] agent chain with retries=3 ended stalled after attempt 3 under the default policy and never ran the fallback. With maxIdenticalFailures: 0, the fallback ran at attempt 4 and the run finished.

Ran at this head, all passing:

  • The whole generated-workflow-verifier file (140), plus smithers-dependency-skip.integration (2) and agent-adapter-boundaries (10). The adapter pins are main's again, because claude.tsx is unchanged.
  • 31 targeted runtime.test.ts tests: the 29 whose bodies feed NodeFailed, TaskHeartbeatTimeout, timed-out, typed deadline codes or retryPolicy through compile, start or sync, plus the init test and resume retries stalled nodes alongside failed ones. 29 passed in one batch. The other two, syncRun keeps redacted failure state… and syncRun records failed primaries and the actual fallback producer, failed with WORKFLOW_SUBMISSION_FAILED … changed while reading on a dependency file while a concurrent pnpm install ran on this host. Both passed when re-run alone.
  • prettier --check and eslint on the changed files, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, pnpm --filter @ultrafuzz/runtime typecheck, pnpm -w knip and node scripts/docs-check.mjs.

Not run locally: the full runtime and CLI suites, and the live OAuth refresh race against a real credential (a real refresh would rotate a token that other sessions on this host share). CI ran the full runtime suite instead; its lane selector did not pick the CLI lane for this diff. On this head, External static analysis passes, and so do the five full release-validation lanes (runtime shards 1–4, and runtime-supporting with the Bun adapter contracts) and release-gates. The now-unconditional graph smoke test ran in shard 4 and passed. On the first push those lanes were skipped.

Risk / compatibility

  • Longer time to exhaust deterministic failures. Waits add up to 3 minutes for the default three attempts and 12 minutes for exhaustive's five. Each fallback rung adds up to 5 more. That is small next to the 1–2 h node timeouts. With the stall verdict gone, repeated identical agent failures now use the full planned chain.
  • Environmental admission failures are not retried. Suppose admission is reading a producer's files and hits a validator-isolate deadline (5 s default), a worker error, or a file-system error other than ENOENT. The task now fails at once. On main, preparation retried immediately, and the agent-side recheck retried under the agent policy. One review measured a first validation at about 200 ms at load 46 on 32 cores, so the deadline has a wide margin. Nobody measured behavior under memory pressure.
  • Some failures change label. Provider timeouts that an agent CLI reports only as text, with no typed code, are now failed/agent-failure rather than timed-out/provider-interruption. That changes report categories and eval dispositions.
    • Modal cloud deadlines. They carry no code: the provider's own cloud node execution timed out, and an inner agent timeout reported through the worker's error document. Both are now failed unless a typed Smithers deadline fires first. By my reading, the provider's deadline usually wins, because its clock starts before the provider's only in-flight heartbeat (at launch). In Modal eval disposition such a deadline moves from incomplete to operational-failure. I traced this by reading node-provider.ts and terminal-disposition.ts, not by running a cloud node.
  • Runs recorded before the upgrade.
    • Some attempts got their timed-out outcome only from the old text rule, for example an agent CLI deadline whose code the old normalizer stripped. Those attempts now re-derive as failed, while the immutable attempt ledger holds timed-out. reconcileNodeAttemptLedgerEntry has no match for that difference.
    • So every sync of such a run reports an error-severity NODE_ATTEMPT_LEDGER_WRITE_FAILED and appends no later ledger entries for that task. That includes ultrafuzz status on a run that completed before the upgrade.
    • Node state and run completion are unaffected: the error is caught and sync still returns ok: true. fix(runtime): attempt and usage ledgers never block run synchronization #1186 (w02) deletes that reconcile path, so merging it first or together removes the diagnostic. I confirmed this by reading the code; I did not re-run the upgrade experiment.
    • Runs launched before the upgrade keep initialDelayMs: 1000 in their sealed task manifests. The template's maxIdenticalFailures: 0 reaches them only through resume --refresh-controller, which re-renders the current template. Ordinary resume continues the persisted workflow.
  • No adapter change. Existing projects do not need ultrafuzz init --force for this PR.

Refs #1084
Refs #1144
Refs #676

🤖 Generated with Claude Code

RetriggerConfidence Score: 2/5

The PR is not ready to merge: changelog linting blocks CI, and two previously reported runtime issues remain outstanding.

Fix All in Claude CodeFindings

  1. P1 Upgraded runs block ledger updates ▶
  2. P1 Transient admission errors cannot retry ▶
Fix with agent prompt
### Issue 1
packages/runtime/src/workflow-sync.ts:undefined-5686
If an in-flight run already recorded an attempt as timed out, but its error has no typed code, this change reclassifies the same attempt as failed on the next sync. The attempt ledger treats the recorded outcome as immutable, so reconciliation fails with `NODE_ATTEMPT_LEDGER_WRITE_FAILED` and later entries for that task cannot be appended. Preserve the recorded outcome or provide a migration path.

### Issue 2
packages/runtime/src/templates/smithers/workflows/workflow.tsx:5584-5589
This catch marks every admission error as non-retryable, including a temporary failure while checking or reading a dependency file. Preparation normally has retries, but Smithers will now fail the consumer after one attempt even if the filesystem operation would succeed on retry. Mark only deterministic admission failures as non-retryable.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR lengthens agent retry waits, preserves the planned retry chain, makes dependency-admission failures non-retryable, and classifies workflow timeouts by typed code. It also expands tests and documentation. The latest change removes a changelog lint suppression needed by the current CI gate.

Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."

Agent retries used an exponential backoff from 1s, so a three-attempt
budget was spent in about 25 seconds, inside the minute a contended
Claude Code OAuth refresh can need to clear (#1084). The agent Task also
inherited Smithers' default identical-failure stall verdict (3), which
ended a chain before later same-agent attempts (exhaustive plans five) or
any [retry].agents fallback profile ran. Against real Smithers 0.35.0, a
[fail, fail, fail, fallback] chain ended `stalled` after attempt 3 and
never ran the fallback; with maxIdenticalFailures: 0 the fallback ran at
attempt 4 and the run finished.

- compileTask: initialDelayMs 60_000, so retries wait 60s, 120s, 240s,
  then Smithers' 300s cap.
- Agent Task: maxIdenticalFailures: 0, so the planned chain is the
  budget. It is set in the template only; the compiled manifest, cloud
  handoff schema and sealed task documents keep their exact shape.
- The real `smithers graph` smoke test never ran: it required a
  workspace-root .smithers install that no checkout has, and could not
  resolve the sealed module paths. It now uses the runtime package's own
  Smithers and asserts the retryPolicy Smithers receives.

Refs #1084

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 28, 2026 23:41
case "NodeFailed": {
const failureMessage = errorText(event.payload.error);
return errorLooksLikeTimeout(event.payload.error)
return workflowErrorIsTimeout(event.payload.error)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Upgraded runs block ledger updates

If an in-flight run already recorded an attempt as timed out, but its error has no typed code, this change reclassifies the same attempt as failed on the next sync. The attempt ledger treats the recorded outcome as immutable, so reconciliation fails with NODE_ATTEMPT_LEDGER_WRITE_FAILED and later entries for that task cannot be appended. Preserve the recorded outcome or provide a migration path.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/workflow-sync.ts
Line: 5686

Comment:
**Upgraded runs block ledger updates**

If an in-flight run already recorded an attempt as timed out, but its error has no typed code, this change reclassifies the same attempt as failed on the next sync. The attempt ledger treats the recorded outcome as immutable, so reconciliation fails with `NODE_ATTEMPT_LEDGER_WRITE_FAILED` and later entries for that task cannot be appended. Preserve the recorded outcome or provide a migration path.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

Comment on lines 5584 to +5589
function assertTaskDependencyInputs(task: (typeof taskSpecs)[number]): void {
try {
admitTaskDependencyInputs(task);
} catch (error) {
throw nonRetryableFailure(error);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Transient admission errors cannot retry

This catch marks every admission error as non-retryable, including a temporary failure while checking or reading a dependency file. Preparation normally has retries, but Smithers will now fail the consumer after one attempt even if the filesystem operation would succeed on retry. Mark only deterministic admission failures as non-retryable.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/templates/smithers/workflows/workflow.tsx
Line: 5584-5589

Comment:
**Transient admission errors cannot retry**

This catch marks every admission error as non-retryable, including a temporary failure while checking or reading a dependency file. Preparation normally has retries, but Smithers will now fail the consumer after one attempt even if the filesystem operation would succeed on retry. Mark only deterministic admission failures as non-retryable.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

aviggiano and others added 3 commits September 29, 2026 03:00
…l timeouts by code

Two defects from #1144.

Retries. A consumer's dependency-admission failure re-reads the same
producer bytes, so every retry fails the same way; each consumer still
spent its whole retry budget and agent fallback chain. The admission
entry points (preparation's assertTaskDependencyInputs and the agent's
assertDependencyArtifactAdmissionCurrent rechecks) now throw with
details.failureRetryable=false, which Smithers honors, and
preparationStep keeps that flag when it rewraps the error. Everything
else stays retryable, including the #672 engine-boundary TypeError. A
real Smithers run of the template's preparationStep and helper pins
this: the admission failure runs once and its agent is skipped, while a
transient TypeError is retried and finishes. With the flag removed the
same preparation runs three times and ends `stalled`.

Timeout labels. errorLooksLikeTimeout ran /timeout|timed out|heartbeat/
over every string in the NodeFailed error, including the stack and
causes. That is a failure taxonomy from free text, which #572 rules
out, and since #1027 every JSON-validator preflight failure mentions
"timeout", so a 2ms `spawnSync ultrafuzz ENOENT` was recorded as a
timed-out provider interruption. A node now times out only on Smithers'
typed deadline codes (TASK_TIMEOUT, TASK_HEARTBEAT_TIMEOUT,
PROCESS_TIMEOUT, PROCESS_IDLE_TIMEOUT) or a TaskHeartbeatTimeout event.
The agent-failure normalizer used to strip PROCESS_TIMEOUT and
PROCESS_IDLE_TIMEOUT, which left the regex as the only label for real
agent CLI deadlines, so it now keeps them.

The #1144 analysis also proposed guarding generate()'s catch-path
recheck. It is not included: resetTaskArtifactsForRetry admits
dependencies before that try block on every first generation in a
process, so the catch path never runs without an admission.

Refs #1144

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ines are labelled failed

Review of this change found the docs, CHANGELOG and one comment stronger
than the code:

- docs/config.md said the planned chain is the whole budget. Smithers
  still stops a chain at a failure it classifies as non-retryable, such
  as a CLI auth or configuration error, and pauses the run on a quota
  limit.
- The admission wrappers mark every failure inside admission
  non-retryable, including a file-system or validator error while
  reading the producer files, not only missing or changed bytes.
- Timeout labelling also keeps the TaskHeartbeatTimeout event rule, and
  a deadline reported only as text, such as the Modal provider's
  cloud-node deadline, is now labelled failed.
  docs/reference/artifacts-reports.md states the rule next to the node
  statuses.
- The workflow-sync comment read as if it listed every Smithers
  deadline code.

Refs #1144

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every CHANGELOG entry is one long line, so super-linter's markdownlint
fails MD013 (line length 400) on any pull request that touches the file.
release-gates then fails and every full release-validation lane is
skipped, so the runtime suite never runs on the pull request. This is
the same directive #1179 adds, byte for byte, so whichever lands first
merges cleanly with the other.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/w08-retry-policy branch from 25ee12c to 668ad91 Compare September 29, 2026 03:13
@aviggiano aviggiano changed the title fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and report what Claude Code said fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code Sep 29, 2026
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant