Skip to content

fix(runtime): attempts abandoned by a crash are recorded, so stats and status agree - #1200

Merged
aviggiano merged 4 commits into
mainfrom
claude/v08-record-abandoned-attempts
Sep 29, 2026
Merged

aviggiano merged 4 commits into
mainfrom
claude/v08-record-abandoned-attempts

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

The campaign-resume e2e test (packages/cli/test/e2e/campaign-resume.test.ts) SIGKILLs the detached controller while summarize runs, then resumes the run. Afterwards ultrafuzz stats counted 3 agent attempts, while status (model_mix) and the stub agent counted 4. The test recorded this as a todo subtest. Smithers' own attempt rows, which ultrafuzz node node:summarize --attempts lists, hold two attempts for the node, the first one cancelled (observed in a run with this change; Ultrafuzz does not write those rows). attempts.jsonl held only the second.

Separately (#1087), stats counted a node interrupted by ultrafuzz cancel as failed, while status counted the same task as other.

Root cause

  • Missing attempt. terminalWorkflowAttempts (packages/runtime/src/workflow-sync.ts) pairs each NodeStarted with a terminal NodeFinished, NodeFailed or NodeCancelled event. At each RunStarted it cleared the starts that were still open, and nothing recorded them. When a run resumes, Smithers 0.35.0 marks the dead activation's in-progress attempt rows cancelled in its resume-cancel-stale-attempt transaction but emits no event for them. So the attempt the kill interrupted never reached the ledger, and fix(runtime): attempt and usage ledgers never block run synchronization #1186's NodeCancelled handling had no event to handle.
  • stats counts operator-cancelled nodes as failed while status counts them as other #1087. Run state has no cancelled status: statusFromWorkflowState maps a Smithers cancelled task to failed, and stats counted run state. status counts Smithers' own task states, where cancelled is other.

Change

Source and schema +88/−39, tests +260/−15, docs +22/−7.

  1. fix(runtime): record attempts a stopped controller abandoned as canceled. A start still open at a RunStarted is kept as abandoned. It is recorded when its task has its next event, normally the replacement attempt's NodeStarted, with:

    • outcome and category canceled;
    • the message abandoned: the controller stopped during this attempt, and the resumed run cancelled it;
    • that event's sequence and timestamp as its terminal.

    It needs its own terminal event because the ledger identity is (workflow_run_id, source_event_sequence) and one RunStarted can abandon several attempts. The rest follows the existing rules:

    • Agent provenance comes from the smithers node attempt row that synchronization already fetches, so a pre-agent attempt is still skipped.
    • An abandoned attempt whose number a reset reuses is recorded without the agent block, like other superseded occurrences.
    • Nothing new can throw. An abandoned start whose task has no later event stays unrecorded, and a terminal stamped before its start is still skipped.

    Normal pairing and the abandoned path share one end() helper. The e2e todo is now a hard assertion. Docs: the attempt-ledger section of docs/reference/artifacts-reports.md, and the e2e description in docs/reference/development.md.

  2. fix(cli): stats counts a node whose task Smithers cancelled as canceled (stats counts operator-cancelled nodes as failed while status counts them as other #1087). A node that run state records as failed is reported, and counted in totals.status_counts, as canceled when every failed task of it has Smithers state cancelled in its recorded provenance.workflow.state. That is the Smithers task state that status counts. The CLI result schema gains canceled in the stats node status enum and in statsStatusCounts. The envelope stays ultrafuzz.cli.result.v2 and stats stays ultrafuzz.stats.v1. Docs: docs/reference/cli.md.

  3. fix(cli): stats takes a node's latest attempt in event order. aggregateAttemptOutcome took the last ledger row per strategy attempt in file order. After change 1, the next sync of a run that an earlier version already synchronized appends the abandoned row after the replacement's row. Without this change, that run's outcome column would read canceled for a node whose last attempt succeeded. Within one workflow run the higher source_event_sequence now wins. Rows from different workflow runs keep file order.

Deliberately not built

  • No upstream Smithers event and no new compatibility patch. The Attempt and cancellation ledgers are incomplete after recovery #1139 analysis proposed asking Smithers to emit NodeCancelled{reason: "resumed"} for stale attempts. The events Ultrafuzz already reads are enough.
  • No reconciliation from smithers node attempt rows. The spec offered it as an alternative source. Those rows carry no event sequence to key the ledger identity on, and using them would add a second pairing path.
  • No cancelled status in run state (NODE_STATE_STATUSES). stats derives canceled from the recorded Smithers state instead, so state.json and event schemas are unchanged.
  • stats counts operator-cancelled nodes as failed while status counts them as other #1087 does not use the attempt ledger. A task cancelled before Smithers selected an agent has no ledger row. A cancellation while the verifier runs is recorded as a failed agent attempt (the finalization overlay). A ledger-based rule would keep disagreeing with status in both cases.
  • The attempt-source-event-join semantic gate is unchanged. It expects NodeFailed for every non-success outcome, so it would reject these rows, as it already rejects fix(runtime): attempt and usage ledgers never block run synchronization #1186's NodeCancelled-terminated rows. No production code runs it. Nothing supplies its eventLog context, and no production caller runs any of the node-attempt-ledger.schema.json gates (the ledger reader has its own checks). Only packages/artifacts/test/semantic-gates.test.ts runs it. Deleting this one gate would leave the other ledger-schema gates, which no production code runs either, so deleting them together is a separate cleanup.
  • Only four events end an abandoned start. An abandoned start ends at the task's next NodeStarted, NodeFinished, NodeFailed or NodeCancelled. If the task's next event is some other type (NodeSkipped, NodePending, NodeRetrying) or never comes, the attempt stays unrecorded, as before.

Verification

Discriminating tests, each run against origin/integration/wave1 sources by restoring the base file, running, then re-applying the change:

  • syncRun records each attempt that a killed controller abandoned as canceled (new, packages/runtime/test/runtime.test.ts). Two tasks are left in progress by one RunStarted and restarted as attempt 2. Two syncs must record one canceled row per task, with agent provenance, no diagnostics and retry_count 1. Base: the ledger is []. Fixed: passes.
  • syncRun records an attempt abandoned at a run activation boundary even when a reset reuses its number. This replaces syncRun abandons an unterminated occurrence at a later run activation boundary, which asserted only ok and running. Base: []. Fixed: one canceled row without the agent block.
  • stats counts a failed node whose task Smithers cancelled as canceled (new, packages/cli/test/run-statistics.test.ts). It covers a single-task node and a fan-out node. The fan-out node's canonical state aggregates two strategy tasks and has no Smithers state of its own. The test also validates the envelope against the closed CLI result schema. Base: [ 'failed', 1, undefined ]. Fixed: [ 'canceled', 0, 1 ]. With the fixed run-statistics.ts and the base schema it also fails; the schema is the only difference. Two simpler rules fail it too:
    • Requiring every failed state, the aggregate included, to carry Smithers state cancelled: the fan-out node whose tasks were both cancelled gets [ 'failed', 1, 0 ].
    • Accepting any cancelled task instead of every one: a fan-out node with one cancelled and one failed task gets [ 'canceled', 0, 1 ].
  • stats reports a node's latest attempt outcome in event order, not ledger row order (new). With the base aggregateAttemptOutcome, the case where the abandoned row comes after its replacement gets 'canceled' instead of 'succeeded'. Fixed: passes.
  • syncRun records an abandoned attempt of an already reported run and publishes the report again (new, packages/runtime/test/runtime.test.ts). The first synchronization sees the event log without the abandoned NodeStarted. That reproduces what an earlier version recorded for a finished, reported run: ledger [[2, 3, 4, 'succeeded']] (attempt, start, terminal, outcome), retry_count 0 and a verified report. The second synchronization sees the full log. It must append [1, 1, 3, 'canceled'] after the replacement's row and raise retry_count to 1. The report publication status that status reads must stay available and verified, and loadCurrentFinalReportSnapshot must still return verified-runtime-report. Base: the ledger keeps one row and retry_count stays 0. Two mutants fail it with terminal report receipt does not match current controller evidence: one skips the terminal-report republish once a publication status exists, and one stops an immutable node's retry_count from following the ledger.
  • e2e campaign-resume.test.ts. Base: a probe copy of this test with the same flow printed per-node attempt_count 1/1/1 (3) against 4 status attempts, failure_categories [] for summarize, and the todo failing with 3 !== 4. Fixed: the test passes with the former todo as a hard assertion, on e9ee6dd (351 s) and on e64c6db (384 s), the last commit that changes source; the later commit adds tests only. The fixed run's ledger shows summarize attempt 1 canceled from sequence 39 (its NodeStarted) to 45 (attempt 2's NodeStarted), with agent chain index 0; attempt 2 succeeded 45 to 57. stats gave summarize attempt_count 2, retry_count 1 and failure_categories ["canceled"]. ultrafuzz node listed attempt 1 as cancelled, finishing 0.25 s before the ledger's finished_at.

Existing tests:

  • The runtime tests matching attempt|ledger|occurrence|supersed|activation|abandon|cancel ran on the final HEAD: 48 tests, including the new one; 46 pass and 2 are skipped (Bun-only adapter contracts).
  • CLI: the run-statistics, stats-command and cli-contracts test files pass (34 tests).
  • Gates: prettier --check on the changed files, eslint, pnpm -w lint, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/integration/wave1 pnpm -w lint:strict:ci, pnpm --filter @ultrafuzz/runtime --filter @ultrafuzz/cli typecheck, pnpm -w knip and node scripts/docs-check.mjs pass. The complexity ceiling is unchanged. The repository maximum is still 90 (synchronizeLinkedWorkflowRun, not touched), and the changed functions measure 2 to 37.

Risk / compatibility

  • More ledger rows. attempts.jsonl gains a canceled row for each crash-abandoned attempt. Pre-agent attempts are skipped as before, unless a reset supersedes them. Those rows count toward retry_count, stats attempt_count, executed-attempt counts, ultrafuzz inspect attempt summaries and the eval recovery metrics (observed_node_attempts, repeated_model_backed_node_executions). All of these undercounted before.
  • Durations include downtime. An abandoned attempt ends when the resumed run restarts its task, so its finished_at, and stats duration_ms for the node, include the time the controller was down. Smithers' own attempt row ends at the resume too: 0.25 s earlier in the e2e run.
  • Runs synchronized before this change get their abandoned rows appended on the next sync, after later rows in the file. Change 3 keeps the outcome column correct. The other ledger readers I found (retry and attempt counts, inspect summaries, eval recovery metrics) count rows and do not depend on order.
  • Finished runs with an abandoned attempt republish their report once. On such a run, the first synchronization after the upgrade (status, stats, inspect and why each run one) also raises retry_count in state.json for the node that lost an attempt. This happens even when the node's successful publication is immutable, because the retry count follows the ledger. The raised retry_count, and the workflow-synced event that the same synchronization appends, both make the published terminal report stale. The report publication status is keyed to a state fingerprint that includes retry_count, so the same synchronization publishes the report again: a new review/runtime-report/<generation>, a rewritten current.json, and a rewritten review/report-publication.json. The report stays verified and status keeps reporting it available. The retry_count change is what triggers that republish; the new runtime test fails without it.
  • Schema. The stats JSON gains a canceled node status and a canceled key in totals.status_counts. A consumer that validates stats output against the previous closed schema will reject the new key.
  • Scope of stats counts operator-cancelled nodes as failed while status counts them as other #1087. A failed node is reported canceled only when every failed task of it has recorded Smithers state cancelled. A failed node whose provenance has no Smithers state (the field is optional) keeps reporting failed.

Changelog entry

Refs #1139, #1087

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR should not merge until stats preserves failures when a fan-out task lacks cancellation evidence.

Fix All in Claude CodeFindings

  1. P1 Failed tasks can disappear ▶
Fix with agent prompt
### Issue 1
packages/cli/src/run-statistics.ts:491-495
When a fan-out node has one task marked `cancelled` and another failed task without a recorded workflow state, this check drops the second task. It then reports the whole node as `canceled`, hiding the failure in `totals.status_counts`. A failed task needs cancellation evidence before it can be counted as canceled.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR records controller-abandoned attempts on resume, aligns stats cancellation counts with Smithers task state, and selects latest attempt outcomes by event sequence.

  • Adds runtime and CLI regression coverage, strengthens the resume end-to-end assertion, and updates the result schema and documentation.
  • The cancellation aggregation still needs to account for failed tasks without recorded workflow state.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[RunStarted after controller stop] --> B[Retain open attempt by task]
  B --> C[Next task event]
  C --> D[Append canceled attempt to ledger]
  D --> E[stats attempt counts and outcome]
  F[Stored task workflow states] --> G[stats node status]
  G --> H[status counts]
Loading

Reviews (1) · Last reviewed commit: "test: cover a cancelled fan-out node and..."

aviggiano and others added 4 commits September 29, 2026 09:55
When a controller dies mid-attempt, the resumed Smithers run marks the
in-progress attempt row cancelled but emits no event for it (the
engine's resume-cancel-stale-attempt path). terminalWorkflowAttempts
dropped any start still open at the next RunStarted, so attempts.jsonl
never recorded that attempt. After a controller SIGKILL and resume,
`ultrafuzz stats` counted 3 agent attempts while `status` and the stub
agent counted 4, and `ultrafuzz node` listed the interrupted node's
first attempt as cancelled (#1139).

A start still open at a RunStarted is now kept as abandoned. It is
recorded as outcome and category canceled, with the message "abandoned:
the controller stopped during this attempt, and the resumed run
cancelled it", when its task has its next event, normally the
replacement attempt's NodeStarted. That event is the terminal because
one RunStarted can abandon several attempts, and each needs its own
(workflow_run_id, source_event_sequence) identity. Agent provenance
comes from the `smithers node` attempt row that synchronization already
reads, so a pre-agent attempt is still skipped. An abandoned attempt
whose number a reset reuses is recorded without the agent block, like
other superseded occurrences. Nothing new can throw: an abandoned start
whose task has no later event stays unrecorded, and a terminal stamped
before its start is still skipped.

The campaign-resume e2e todo is now a hard assertion that stats and
status count the same agent attempts.

Refs #1139

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Run state has no cancelled status: synchronization maps a Smithers
`cancelled` task to failed, so `stats` counted nodes interrupted by
`ultrafuzz cancel` as failed, while `status` counted the same tasks as
other (#1087).

stats now reports a failed node as canceled when every failed task of
it has Smithers state `cancelled` in its recorded workflow provenance.
That is the state `status` reads. The attempt ledger is not used for
this: a task cancelled before Smithers selected an agent has no ledger
row, and a cancellation while the verifier runs is recorded as a
failed agent attempt. This adds `canceled` to the stats node status
enum and to `totals.status_counts` in the CLI result schema; the
envelope stays ultrafuzz.cli.result.v2 and stats stays
ultrafuzz.stats.v1.

Refs #1087

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aggregateAttemptOutcome took the last attempts.jsonl row of each
strategy attempt in file order. Rows are appended when they are
recorded, so once a version records attempts an earlier one skipped,
such as the crash-abandoned attempts recorded by the previous commit,
the next sync of an already synchronized run appends the earlier
attempt after its replacement. `stats` would then report the node's
outcome as canceled although its last attempt succeeded.

Within one workflow run the row with the higher source_event_sequence
now wins; rows from different workflow runs keep file order.

Refs #1139

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…n upgrade

- stats: the #1087 test also covers a fan-out node, whose canonical
  state aggregates its strategy tasks and has no Smithers state. It now
  fails if the rule requires the aggregate state to be cancelled, or
  accepts any cancelled task instead of every one.
- runtime: a finished, reported run that an earlier version synchronized
  without its abandoned attempt gets the canceled row appended on the
  next sync. That raises retry_count on the immutable node, and the same
  pass publishes the terminal report again, so status keeps reporting a
  verified report. The test fails on the base sync code and when the
  republish is skipped.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 09:55
Comment on lines +491 to +495
if (nodeState.status !== "failed" || provenance === undefined || !("workflow" in provenance)) return [];
const workflow = provenance.workflow;
return workflow !== undefined && "state" in workflow ? [workflow.state] : [];
});
return failedTaskStates.length > 0 && failedTaskStates.every((state) => state === "cancelled");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Failed tasks can disappear When a fan-out node has one task marked cancelled and another failed task without a recorded workflow state, this check drops the second task. It then reports the whole node as canceled, hiding the failure in totals.status_counts. A failed task needs cancellation evidence before it can be counted as canceled.

Knowledge Base Used: CLI workflows

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/cli/src/run-statistics.ts
Line: 491-495

Comment:
**Failed tasks can disappear** When a fan-out node has one task marked `cancelled` and another failed task without a recorded workflow state, this check drops the second task. It then reports the whole node as `canceled`, hiding the failure in `totals.status_counts`. A failed task needs cancellation evidence before it can be counted as canceled.

**Knowledge Base Used:** [CLI workflows](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/cli-workflows.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

@aviggiano
aviggiano merged commit 2cacf4c into main Sep 29, 2026
7 of 16 checks passed
@aviggiano
aviggiano deleted the claude/v08-record-abandoned-attempts branch September 29, 2026 10:01
aviggiano added a commit that referenced this pull request Sep 29, 2026
…1213)

`stats` reports a failed node as `canceled` when Smithers cancelled every failed task of it (#1087, #1200). `workflowCancelled` in `packages/cli/src/run-statistics.ts` used a missing Smithers `state` as the mark of a fan-out aggregate record. So it also skipped a failed *task* record without a state. For a fan-out node with one cancelled task and one failed task without a state, it saw only `cancelled` and reported the node `canceled`. The failure disappeared from `totals.status_counts`, which contradicts `docs/reference/cli.md` ("When Smithers cancelled every failed task of a node").
A triage of the finding found no writer on `main`, or in past releases, that records a failed task without a state. So this fixes the rule, not an observed run.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant