fix(runtime): make run synchronization single-writer, idempotent, bounded and non-fatal for observers - #1180
Conversation
… on a failed deadline cancel The workflow deadline is checked only when something synchronizes the run (#1110). An operator-paused run executes nothing, and resuming it records a new deadline, so cancelling it at the first observation past its deadline bounded nothing. The control projection now leaves paused runs out of deadlineExceeded. A failed deadline cancel was an error diagnostic, so `status` returned ok:false and `status --watch` stopped polling, and nothing asked again. It is now a warning: the run stays active and the next synchronization requests cancellation again. Refs #1110 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…runner queries Every control projection renews the controller lease and advances the concurrency clock, and any difference counted as a change, so each `status` poll rewrote state.json and a finished run kept an "active" lease renewed by whoever looked at it (#1145). The projection now also reports whether those clocks are its only change; an observe-only pass (status) and any pass over a terminal run skip writing such a projection. Explicit passes over a live run (syncRun, inspect, why, stats) still renew its lease, which the eval and Modal pumps read as liveness. Read-only runner queries had no timeout unless the caller passed one, so a wedged runner process blocked status and the synchronization pumps indefinitely. runSmithersInspectionCommand now bounds the runner process (never executable preparation) at 120 s by default; ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS overrides it, capped at 600 s. A timeout becomes the existing failed-query diagnostic, with a message naming the bound. The runner's orphan reason recommends `smithers supervise -r <id>`, which the public rewrite turned into `workflow runner supervise -r <id>`, a command nobody can run. It now recommends `ultrafuzz resume <run-id>`. Refs #1145 Co-Authored-By: Claude Opus 5.5 <[email protected]>
… redundant artifact events status, inspect, why, stats, the dashboard and the eval runner all run the full mutating synchronization, and nothing serialized them (#1152). Two passes over one newly finished node both finalize it and race on its manifest temp file, state.json and the journals; the issue analysis reproduced duplicate node-synced events, a false failure that the immutable attempt ledger then kept, and passes that died midway through a write group. synchronizeLinkedWorkflowRun now runs each pass under a non-blocking proper-lockfile lock on <run>/.workflow-sync. A pass that finds it held returns ok with an info WORKFLOW_SYNC_IN_PROGRESS diagnostic and writes nothing, so status never waits or fails on contention. Acquisition never waits, so it cannot deadlock with the control lock, and its target is not the run root because proper-lockfile keys its in-process registry by target and the control lock already locks that path. The lock lives in a small wrapper that re-enters the unchanged pass body. node-artifacts-verified and node-artifacts-missing duplicated node provenance, had no reader, and reported `missing: []` for every schema, semantic or authority failure. They are no longer written; their schema variants stay so existing journals replay. A failed node's durable last_error now prefixes each error with its run-relative path and JSON pointer, never a host path. The eventProvenanceForTask plumbing is deleted: createEventRecord never persisted it. Refs #1152 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…sh fails status ran its observe-only synchronization strictly. A failed or malformed runner events query made it return ok:false, exiting 1 and stopping --watch, although the runner's own health was available, and any error thrown by the refresh (for example a torn usage.jsonl append) escaped getRunHealth. stats already tolerated the transient class. The observe-only pass now sets tolerateInvalidEventStreams and downgrades the same transient code set as stats; workflow-sync exports it once as TRANSIENT_SYNC_DIAGNOSTIC_CODES for both. An error thrown by the refresh becomes a WORKFLOW_STATE_SYNC_FAILED warning (the snapshot race keeps WORKFLOW_STATE_SYNC_RACED), and status goes on to query the runner. getRunHealth also passes the evidence it already verified into the pass instead of reading and verifying the sealed snapshot a second time. Refs #1145 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…nd deadline changes configuration.md gains the ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS row and a paragraph on skipped passes, status warnings and when state.json is left untouched. The workflow_deadline_seconds paragraph from #1157 now says paused runs are not cancelled and a failed cancel is a retried warning, and the supervisor paragraph no longer claims that reporting checks the deadline. Refs #1110, #1145, #1152 Co-Authored-By: Claude Opus 5.5 <[email protected]>
| stale: 300_000, | ||
| update: 60_000, | ||
| // The default handler throws from a timer and would kill an observer or the eval runner. | ||
| onCompromised: () => undefined |
There was a problem hiding this comment.
Lost lock permits overlapping writes If a synchronization process is suspended long enough for its five-minute lock to become stale, another process can acquire the lock while the first pass is still active. This no-op handler lets the first pass continue writing state and journals, so both passes can finalize the same node concurrently. Stop the pass when it loses lock ownership without throwing from the timer.
Knowledge Base Used: Runtime orchestration
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/workflow-sync.ts
Line: 1366
Comment:
**Lost lock permits overlapping writes** If a synchronization process is suspended long enough for its five-minute lock to become stale, another process can acquire the lock while the first pass is still active. This no-op handler lets the first pass continue writing state and journals, so both passes can finalize the same node concurrently. Stop the pass when it loses lock ownership without throwing from the timer.
**Knowledge Base Used:** [Runtime orchestration](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/runtime-orchestration.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| const evidence = | ||
| control.evidence ?? | ||
| (await readLinkedWorkflowEvidence(projectRoot, input.runId, { | ||
| ...(control.observeOnly === true ? { observeOnly: true } : {}) | ||
| })); |
There was a problem hiding this comment.
Stale workflow evidence gets written
getRunHealth reads workflow evidence before acquiring the synchronization lock. If a concurrent resume or replay switches this run to a new workflow link in between, this pass uses the old evidence without checking the current link. It can then write the old workflow’s state into the new link’s state.json and journals. Revalidate the link under the lock before writing.
Knowledge Base Used: Runtime orchestration
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/workflow-sync.ts
Line: 1415-1419
Comment:
**Stale workflow evidence gets written** `getRunHealth` reads workflow evidence before acquiring the synchronization lock. If a concurrent resume or replay switches this run to a new workflow link in between, this pass uses the old evidence without checking the current link. It can then write the old workflow’s state into the new link’s `state.json` and journals. Revalidate the link under the lock before writing.
**Knowledge Base Used:** [Runtime orchestration](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/runtime-orchestration.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.…instead of throwing The single-writer lock rethrew every acquisition error other than contention, so `inspect`, `syncRun` and the dashboard's inspect action threw on a run root they could not write to (read-only, full or removed), where origin/main returned a diagnostic. The Modal worker stops its `inspect` watch at the first command that fails. A failed acquisition now fails the pass with WORKFLOW_SYNC_LOCK_FAILED and writes nothing. `inspect` already reports a failed pass as warnings, and the code joins the transient set that `status` and `stats` report as a warning, because a lock the pass could not take says nothing about the run's evidence. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…say which errors it downgrades The observe-only pass reused the evidence getRunHealth had verified on every retry. A retry follows a snapshot race, and the writer it raced may be a lifecycle command (resume, replay, fork) that switched the run's workflow link, so a retry now reads the evidence again, as each attempt did on origin/main. Only the first attempt reuses it. The comment claimed that every synchronization error becomes a warning in `status`. Only the transient codes, an exhausted race budget and thrown errors do; any other error the refresh returns keeps its severity, which is how #707 deliberately made `status` report refresh errors. The comment now says so, and the torn-ledger scenario asserts the WORKFLOW_STATE_SYNC_FAILED warning of the catch-all rather than only a healthy result, so it fails if that trigger stops reaching the catch-all. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…phan resume advice publicHealthReason joined `ultrafuzz resume <run-id>` into the runner's orphan reason and then renamed every remaining "smithers" to "workflow runner", so a chosen run ID such as `smithers-upgrade-check` came out as `ultrafuzz resume workflow runner-upgrade-check`, a command that does not exist. The generic rename now runs on the text around the runner's `smithers supervise -r <id>` command, and the resume command is joined in afterwards. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ive lease The idempotence test asserted that a finished run's controller lease is still `active` before comparing bytes. That is the #1145 defect, not the behaviour under test, so a later change that releases the lease at the terminal transition would have failed it for no reason. The run status check stays. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e mode The CHANGELOG said synchronizations of one run "no longer overlap", but a lock whose holder stops refreshing it for five minutes is taken over, and the last_error location applies only to nodes failed by host-side output validation. Both claims now say so, and the entry names WORKFLOW_SYNC_LOCK_FAILED. The configuration reference names the lock file, its five-minute takeover and WORKFLOW_SYNC_LOCK_FAILED, adds a lock the pass cannot take to the refresh failures `status` reports as warnings, and states that `0`, `off` or any other invalid ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS value means the default, unlike the neighbouring observation timeout where `0` and `off` disable it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
| observeOnly: true, | ||
| tolerateInvalidEventStreams: true, | ||
| deadlineMs, | ||
| ...(reuseEvidence ? { evidence } : {}) |
There was a problem hiding this comment.
Health can mix workflow links If a concurrent resume, replay, or fork replaces run state during synchronization, this retry can reload and synchronize the new workflow link.
getRunHealth then queries the original link for its verdict and counts, so one response can combine new-link local state with old-link health. Keep the synchronization and health query bound to the same link.
Knowledge Base Used: Runtime orchestration
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/state-export.ts
Line: 386
Comment:
**Health can mix workflow links** If a concurrent resume, replay, or fork replaces run state during synchronization, this retry can reload and synchronize the new workflow link. `getRunHealth` then queries the original link for its verdict and counts, so one response can combine new-link local state with old-link health. Keep the synchronization and health query bound to the same link.
**Knowledge Base Used:** [Runtime orchestration](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/runtime-orchestration.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Every pull request in this batch inserts its entry at the same place in CHANGELOG.md, so each merge would conflict with the next. The entries are collected into one changelog update instead. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem
Run state (
state.json, the event/attempt/usage journals,run.jsonaccounting) converges only when some command synchronizes the run:status,inspect,why,stats, the dashboard's inspect action, and the eval runner's 15 s poll. That pump had four stability defects:state.jsonand the journals. The analysis reproduced duplicatenode-syncedevents, a node recordedfailedin the immutable attempt ledger while it succeeded, and passes crashing midway through a write group.statuspoll rewrotestate.jsonand a finished run kept an "active" lease renewed by whoever looked at it. Read-only runner queries had no default timeout, so one wedged runner process blockedstatusand the pumps indefinitely. An orphaned run's status told operators to runworkflow runner supervise -r <id>, which is not a command.eventsquery madestatusreturnok:false(exit 1,--watchstops) although the runner's health was available, and any error thrown by the refresh (for example a tornusage.jsonlappend) escapedgetRunHealthentirely.WORKFLOW_DEADLINE_CANCEL_FAILEDwas an error, sostatus --watchstopped and nothing asked again.Also:
node-artifacts-missingwas emitted for every gate failure withmissing: []for schema, semantic and authority failures, and a failed node'slast_errorjoined messages without saying which artifact failed.Root cause
The synchronization pass is written as if it were the only writer and as if every observation were a change. Neither holds: every observer is a writer, and most observations change nothing but the clock.
Change
workflow-sync.ts).synchronizeLinkedWorkflowRunruns each pass under a non-blocking proper-lockfile lock on<run>/.workflow-sync(lockfile<run>/.workflow-sync.lock). A pass that finds it held returnsokwith an infoWORKFLOW_SYNC_IN_PROGRESSdiagnostic and writes nothing, sostatusnever waits or fails on contention. A pass that cannot create the lock (a read-only, full or removed run root) fails withWORKFLOW_SYNC_LOCK_FAILEDand writes nothing; it no longer throws out ofinspect,syncRunor the dashboard. The target is not the run root, because proper-lockfile keys its in-process registry by target and the control lock already locks that path.onCompromisedis a no-op, because the default throws from a timer and would kill an observer or the eval runner.state.jsonand the journals untouched (workflow-control.ts,workflow-sync.ts). The projection also reportsobservationOnly, which is true when its only changes are the leaserenewed_at/expires_atand the concurrencyobserved_atand durations. An observe-only pass (status) and any pass over a terminal run skip writing such a projection. Each pass still creates and removes<run>/.workflow-sync.lock. Explicit passes over a live run (syncRun,inspect,why,stats) still renew the lease, so the eval runner and the Modal worker, which pump throughsyncRunandinspect, keep it current.eval statusreads it. Durations stay exact: a skipped write means nothing they integrate over changed.smithers.ts).runSmithersInspectionCommandbounds the runner process, never executable preparation, at 120 s by default.ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MSoverrides it, capped at 600 s;0,offor any other value that is not a positive integer means the default, so the bound cannot be disabled. A timeout becomes the existing failed-query diagnostic, with a message naming the bound.state-export.ts,stats.ts). The observe-only pass setstolerateInvalidEventStreamsand downgrades the transient code setstatsalready tolerated, now exported once fromworkflow-syncasTRANSIENT_SYNC_DIAGNOSTIC_CODESfor both and extended withWORKFLOW_SYNC_LOCK_FAILED. Any error thrown by the refresh becomes aWORKFLOW_STATE_SYNC_FAILEDwarning (the snapshot race keepsWORKFLOW_STATE_SYNC_RACED). Other errors the refresh returns keep their severity (see below).getRunHealthpasses the evidence it already verified into the first attempt, which skips one full sealed-snapshot verification perstatus; a retry after a snapshot race reads the evidence again, as every attempt did before.workflow-control.ts,workflow-sync.ts). Paused runs are left out ofdeadlineExceeded.WORKFLOW_DEADLINE_CANCEL_FAILEDis now a warning, and the next synchronization requests cancellation again.workflow-sync.ts,events.ts).node-artifacts-verifiedandnode-artifacts-missingare no longer written. They duplicated node provenance and had no reader. Their schema variants stay, so existing journals replay. A node failed by host-side output validation now gets alast_errorthat prefixes each error with its run-relative path and JSON pointer, for exampleartifacts/project-discovery/findings.json#/0: must have required property 'status', and never with a host path. Runner-reported failures keep the runner's text, and a diagnostic with no in-run path keeps its bare message. The deadeventProvenanceForTaskplumbing is deleted;createEventRecordnever persisted it.state-export.ts). The runner'ssmithers supervise -r <id>text is rewritten toultrafuzz resume <run-id>. Resume continues an orphaned run, becauseorphanedis not an active runner state. The generic "smithers" → "workflow runner" rename runs on the text around that command, so a chosen run ID such assmithers-upgrade-checkstays intact.docs/reference/configuration.mdgains theULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MSrow and a paragraph on skipped passes, the lock file and its five-minute takeover,WORKFLOW_SYNC_LOCK_FAILED,statuswarnings and whenstate.jsonis left untouched. Theworkflow_deadline_secondsparagraph from docs: state that workflow_deadline_seconds is only enforced on sync #1157 now covers paused runs and the retried warning, and the supervisor paragraph no longer claims that reporting checks the deadline.The lock sits in a small wrapper,
synchronizeExclusively, that re-enters the unchanged 450-line pass body. The wrapper recognizes its own call through a module-privateWeakSetof control objects. This keeps every caller on the locked path under the same public name. It also leaves the body's lines untouched apart from the few this PR changes. That limits textual conflicts with other in-flight edits of the body, and the diff-limited strict lint does not re-flag it (a review confirmed that a plain rename of the body failslint:strict:cion complexity and length). Once #1184's lint baseline lands, this can become an exported wrapper plus a private body. The lock is taken before the evidence read, not after it as the work-item spec and analysis proposed. That is safe because acquisition never waits, so it cannot form a cycle with the control lock, and a skipped pass then does no evidence work at all.Deliberately not built (and why)
status. A skipped pass leaves the run to the pass in progress, so there is nothing to wait for.statusstaying a pump is the documented way state converges (workflow_deadline_seconds is only enforced while an operator runs a CLI command; unattended runs have no wall clock #1110 depends on it).statusdoes not downgrade every refresh error, althoughwhydoes. A review suggestedwhy's rule. fix(runtime): recover terminal pending resumes #707 deliberately did the opposite forstatus: it removed a blanket downgrade of failed refreshes and madeokfalse on any refresh error, and its testgetRunHealth reports a terminal product status while workflow health is liverequiresWORKFLOW_TERMINAL_WITHOUT_FAILED_NODEto stay an error. I ran that test with the blanket downgrade and it fails (true !== falseonhealth.ok). So only the transient set, the lock failure and thrown errors are downgraded, as scoped for this item. A failed runnernodequery (WORKFLOW_ATTEMPT_INSPECT_FAILED) therefore still failsstatusand stops--watchhere; from its diff, fix(runtime): attempt and usage ledgers never block run synchronization #1186 (w02), which owns that path, makes it a warning where it is raised.resume,replay,fork) do not take the sync lock, so an observe-only pass could already read evidence and write after a link switch on origin/main. PassinggetRunHealth's evidence into the first attempt adds only the time between that read and the start of the pass, which is a local divergence check. Retries now read the evidence again.ultrafuzz eval scorestep that resolves the report when it runs (scoring.ts), so a holder that finishes publishes in time. The residual needs the holder to die inside that window.eval scorethen fails withEVAL_TERMINAL_REPORT_INVALIDrather than silently, and any later synchronization of the run publishes the report. On origin/main the same outcome needed the eval runner's own pass to end just before another pass wrote the terminal status. The runner's final catch-up block is also being deleted by refactor(evals): delete the dead reporter/telemetry pipeline and publication leftovers #1174 (w20a), so a hunk there would conflict.minItems: 1onmissing) would make existing journals unreplayable, and a schema change rotates the validator build identity for in-flight runs. The analysis also found that the issue's "and N more" truncation loses nothing durable: the retained artifact, whichultrafuzz artifact validatere-checks, the sealed graph andtasks.json, andstate.jsonkeep the full data.pricing_catalog.fetched_at, sorun.jsonis rewritten on everystatusof such a run. Work item w02 owns accounting. This is one reason this PR only refers to Make run status authoritative, bounded, and read-only #1145.preserveFailedWorkflowAttemptsBeforeReset) does not take the lock. That area belongs to w02. Lifecycle writers (cancel,pause,resume) also stay outside it. This lock covers sync-versus-sync writes only.cancelcommand is not bounded. Only read-only queries got a timeout, andcancelis shared withultrafuzz cancel.ULTRAFUZZ_OBSERVATION_SYNC_TIMEOUT_MScheckpoints. They are now partly redundant with the per-query bound, but removing them is a separate cut.Verification
Discriminating evidence for the first round. I copied the new and changed tests onto an
origin/mainworktree (fbbcf6c5, same sync code asb6dd1da9), compiled them against main'ssrc, and ran them there and on the branch. Two reviewers reproduced this table independently:a paused run is not cancelled by an observation past its workflow deadline(unit)re-projecting an unchanged run only advances its observation clock(unit)observationOnlyabsent)repeated observations leave an unchanged finished or orphaned run byte-identicalstatus reports a runner query that stops answering instead of waiting on ita synchronization that finds another in progress skips it without waiting or writingnode-synceda node whose outputs fail validation names each failing artifact in its last errorlast_errorhas no location; with the assertions reordered,node-artifacts-missingwith"missing":[]is emitteda failed deadline cancel is a warning and the next status requests it againWORKFLOW_DEADLINE_CANCEL_FAILEDhas severityerroran explicit synchronization keeps renewing a live run's controller leaseReview round. Each new or changed test ran against
origin/main(b6dd1da9) and against the previous head (689e92e0), both in separate worktrees with the final test file compiled against that commit'ssrc:a synchronization that cannot take its lock is reported instead of thrown(new; run rootchmod 0555, registered only when not running as root)WORKFLOW_SYNC_LOCK_FAILED; main has no sync lock, andinspectreturns onlyWORKFLOW_CONTROL_EVIDENCE_INVALIDwarningsgetRunStatusrejects withEACCES: permission denied, mkdir '<run>/.workflow-sync.lock'inspectandstatusareokwith aWORKFLOW_SYNC_LOCK_FAILEDwarning, andsyncRunreturnsok:falsewith itgetRunHealth accepts strict 0.35 orphan, …(run ID nowsmithers-034-shapes)superviseultrafuzz resume workflow runner-034-shapes"status keeps reporting runner health when run-state synchronization fails(now also asserts theWORKFLOW_STATE_SYNC_FAILEDwarning for the torn ledger)WORKFLOW_EVENTS_FAILEDis an errorrepeated observations …(precondition pinning a finished run'sactivelease removed)The retry change (a retry reads the evidence again) has no dedicated test; making a lifecycle command switch the link between two attempts of one pass is not something the fake runner can stage. The existing snapshot-race test covers the retry path and passes.
What else I ran on the final tree:
runtime.test.tstests on the final commit, all passing: the nine new or changed ones above, every othergetRunHealth*andgetRunStatus*test (including fix(runtime): recover terminal pending resumes #707'sgetRunHealth reports a terminal product status while workflow health is live), the observer snapshot-race test (listRuns, observers and status re-read live run documents replaced while they were read), the four divergence-tolerance status tests, and both existing deadline tests (syncRun cancels a nonterminal workflow at its durable workflow deadline,syncRun honors cancellation and an overall deadline before terminal synchronization).workflow-control.test.js: all 14 pass.status --watch --json keeps a failing poll on one NDJSON line,status surfaces a terminal product and live workflow lifecycle divergence,status --watch stops immediately on a degraded verdict even while product state is nonterminal,stats falls back to unchanged local evidence when workflow event output is malformed, andstatus recommends ultrafuzz why instead of the engine commandpass. In the first run, two of them failed at theirultrafuzz runlaunch step (exit 1, before any status or stats code ran; I did not capture its output) while the machine was heavily loaded; both passed when rerun. Separately, two runtime tests once failed insidestartRunwithWORKFLOW_SUBMISSION_FAILED("workflow execution file dependencies/… changed while reading", the host flake a reviewer also hit) and passed when rerun.npx prettier --checkandnpx eslinton the changed files;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;pnpm --filter @ultrafuzz/{runtime,cli,artifacts} typecheck;pnpm -w knip;node scripts/docs-check.mjs. All clean. The runtimesrcand tests also typecheck at each of this round's commits.Run in the first round and not re-run on this head: the five finalization tests (
syncRun accepts canonical findings…,…rejects a valid output swap…,…persists a schema-valid failed node when controller manifest publication fails,…marks task-output validation failures for terminal disposition,…rejects workspace-mirrored outputs…),lifecycle-inspection.test.tstests 1–32,observation-deadline.test.tsandobservation-snapshot.test.ts(12/12),dynamic-lifecycle.test.tsdynamic child failure, skip, and timeout keep strict joins blocked with durable terminal state, and artifactsevent-record.test.ts(3/3, oldnode-artifacts-*fixtures still validate).What I did not run: the full runtime, CLI, evals or Modal suites; the Bun adapter contracts (untouched); and a multi-process soak of concurrent syncs. The in-process gate test is deterministic; the analysis's 8/8 multi-process soak was on a prototype, not on this code.
CI:
External static analysis(andrelease-gates, which depends on it) fails on this PR because Super-Linter's MD013 flags every long line inCHANGELOG.md, including lines that predate this PR. Every PR that editsCHANGELOG.mdhits it; #1181 and #1184 address it, so this PR adds no third exemption and should be rebased after one of them lands.Risk / compatibility
statusnow exits 0 with warnings when its refresh hits a transient runner failure, a lock it cannot take, or a thrown bookkeeping error. A paused run is no longer cancelled at its deadline. A runner query slower than 120 s now fails unlessULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MSis raised.events.jsonlstops receivingnode-artifacts-*records. I found no reader of those records in the repo, and old journals still replay.statsnow toleratesWORKFLOW_SYNC_LOCK_FAILEDlike its other transient codes, so on a run root it cannot write to it should fall back to local evidence with a warning (by reading the code; not run).WORKFLOW_SYNC_IN_PROGRESS, andstatusstill reports runner health. A normal exit, SIGINT or SIGTERM removes it through proper-lockfile's exit hook.statusno longer renews the lease of a live run. Explicit passes and the eval/Modal pumps still do. Nothing in the runtime decides on the lease timestamps;evals/status.tsreads them for display heuristics, fed by the eval runner's explicit pass.last_errorformat applies only to new finalizations; a node already finalized as an immutable output failure is not re-finalized.preserveFailedWorkflowAttemptsBeforeResetnext to wheresynchronizeExclusivelyis inserted, a textual conflict to resolve by keeping the wrapper. From its diff (not run together with this branch), fix(runtime): attempt and usage ledgers never block run synchronization #1186 also catches accounting errors as aWORKFLOW_ACCOUNTING_FAILEDwarning. After both land, the torn-ledger scenario above would then probably stop reaching the thrown-error catch-all, and its newWORKFLOW_STATE_SYNC_FAILEDassertion would fail loudly rather than pass without testing anything; whichever PR lands second needs to point that scenario at another thrown error.Refs #1145
Closes #1152
Refs #1110
🤖 Generated with Claude Code
The PR does not appear safe to merge while the three outstanding synchronization and workflow-link findings remain unresolved.
Fix with agent prompt
Summary
The PR serializes run-state synchronization, avoids observation-only state writes, bounds read-only runner queries, and keeps status reporting runner health when synchronization fails. It also changes deadline handling and improves failure and orphan-remediation text. Since the previous review, the author removed the changelog entry; the runtime changes are unchanged.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR O[Observer or explicit sync] --> L{Run sync lock available?} L -->|Held elsewhere| S[Skip pass; report in progress] L -->|Acquired| E[Read workflow evidence] E --> P[Project and persist eligible changes] P --> U[Release lock]Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."