fix(runtime): record final-report producer selections in the run instead of querying smithers - #1183
Conversation
cef55b1 to
594d6b8
Compare
| const previous = | ||
| cached ?? | ||
| (local && attempt > 1 ? readFinalReportSelections(task).filter((selection) => selection.attempt < attempt) : []); | ||
| const selections = [...previous, { attempt, chainIndex }]; |
There was a problem hiding this comment.
Stale failed-attempt history If a reset reuses an attempt number and that attempt fails in preflight, it never updates the record. The next attempt keeps the old entry through this filter, so
agent_execution.failed_attempts can name a model that did not run in the current round. This is a non-blocking provenance inaccuracy that can mislead readers comparing model runs.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/templates/smithers/workflows/workflow.tsx
Line: 3185-3188
Comment:
**Stale failed-attempt history** If a reset reuses an attempt number and that attempt fails in preflight, it never updates the record. The next attempt keeps the old entry through this filter, so `agent_execution.failed_attempts` can name a model that did not run in the current round. This is a non-blocking provenance inaccuracy that can mislead readers comparing model runs.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.3a1d343 to
c9b278a
Compare
| if (cloudCleanup !== undefined) { | ||
| return runtimeFailure([cloudCleanup]); | ||
| } | ||
| try { |
There was a problem hiding this comment.
Cloud resources survive cleanup When an operator cleans a run launched with per-node Modal execution,
cleanRun now deletes the local run directory without terminating its sandboxes or deleting its volume. Those remote resources can remain billable, while the deleted plan.json held the Modal app information needed to locate them for manual cleanup.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/clean.ts
Line: 85
Comment:
**Cloud resources survive cleanup** When an operator cleans a run launched with per-node Modal execution, `cleanRun` now deletes the local run directory without terminating its sandboxes or deleting its volume. Those remote resources can remain billable, while the deleted `plan.json` held the Modal app information needed to locate them for manual cleanup.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.…ead of querying smithers After a controller restart (resume, quota park, supervisor relaunch) the generated workflow's process-local record of which chain rung each report-producer attempt used is gone. The producer's retry prompt and the verifier both rebuilt it by running a bare `smithers node ... --full-output` from inside the workflow: a runner subprocess on the finalizer path that reads up to 64 MiB under a 180 s budget. Wherever the controller PATH resolved no `smithers`, every such attempt failed with "Smithers report-producer authority is unavailable" (cause: ENOENT) until the report node ran out of retries (#1143). The only real-Smithers test of this path gave the detached engine a `smithers` on its PATH. Record each executed selection {attempt, chainIndex} with writeFileDurable in <run>/smithers/final-report-selections/<attempt-id>.json before the agent runs, and read it back when the process-local cache is empty. A re-dispatched attempt number replaces its own entry; finalReportAgentExecution keeps validating attempt order and chain bounds. This deletes readFinalReportSmithersAuthority, priorFinalReportAgentSelections, the 180 s SMITHERS_REPORT_PRODUCER_AUTHORITY_TIMEOUT_MS budget and the workflow's use of the Smithers attempt reconcilers. The runtime exports stay: host-side workflow-sync uses inspectSmithersAttemptAgentSelection, and workflows rendered before this change import both. Per-node cloud execution is gone (#1197) and rendered task specs no longer carry `execution`, so every producer attempt writes the record and the verifier has no single-rung cloud shortcut. The history and integration fixtures carry no `execution` either, so a leftover `task.execution` read in the template fails both. This reverses #585's "a non-Codex fallback cannot forge final-report producer authority through the run filesystem" test. That property was not a real boundary: an unsandboxed agent running as the same user can edit the Smithers database as easily as a run-directory file, and trusted-cli.ts already states the host is not a same-UID sandbox boundary. Both the database and the new record sit outside the agent's worktree and declared artifact directories. Tests: the real-Smithers restart test now runs the detached engine with a PATH that resolves no `smithers` (and calls pause by absolute path); with the previous template both variants fail with ENOENT, with this change both pass. report-retry-history.test.ts is rewritten to drive the extracted helpers across simulated restarts against a temporary run directory. The tests that pinned the CLI query (fake execFileSync, read counts, the budget, source regexes) are deleted. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The artifacts reference said the verifier compares agent_execution with "independently persisted Smithers attempt authority ... without trusting a model-writable file". It now compares with controller memory or, after a restart, the run directory's selection record; say where that record lives and that, like the Smithers database, it is host evidence rather than a boundary against an unsandboxed same-user agent. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…on a bare smithers Review follow-ups for the final-report selection record tests: - The first history test now also checks that a verifier in the producer's own controller returns what that controller observed after the record is rewritten. The #585 forge test that covered controller-memory precedence was deleted with the Smithers query; without this, a verifier that always re-read the record passed every test. - The re-dispatch test now dispatches attempt 2 again on a different rung, so it tells replacing the recorded entry apart from keeping it. A real Smithers timetravel of a failed report verifier's producer re-dispatches the latest attempt number, which is the case this models. - The real-Smithers integration test no longer filters PATH entries that hold a `smithers`. That filter also dropped `bun` on hosts where a global `smithers` sits beside it, and the node_modules/.bin shim then failed with `exec: bun: not found`. A stub that exits 97 now goes first on the detached engine's PATH instead, so a bare `smithers` call fails even on a host that has one. Co-Authored-By: Claude Opus 5.5 <[email protected]>
A run launched before the selection record existed keeps its persisted workflow on plain `resume`; only `resume --refresh-controller` renders the recording template. If that refresh lands after the report producer succeeded, the verifier finds no record and fails. The error now says to run `ultrafuzz resume <run-id> --refresh-controller --retry-failed`, which resets the verifier's producer and reruns it under the current controller, so the rerun writes the record. A comment notes that after a reset reuses attempt numbers, an attempt that fails before recording leaves the old round's entry for its number in place. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ng change New entries now go straight into CHANGELOG.md under Unreleased, so the entry this branch had moved to consolidated release notes comes back. It sits under Breaking changes rather than Other changes: a run whose report producer succeeded under an earlier release's workflow and that is resumed with --refresh-controller before its verifier runs now fails that verifier until `resume --refresh-controller --retry-failed` reruns the producer, and the reference docs retire #585's claim that the run filesystem cannot be used to forge the producer. Main already files changes without a `!` whose only break is to older runs there (#1176, #1193). The entry also says that plain `resume` keeps the workflow persisted at launch, and that the recovered producer's failed_attempts omits the attempts from before the refresh. It describes what the replaced query cost (a runner subprocess with a 64 MiB read under a 180 s budget) rather than a missing `smithers` on PATH, since the run's `trusted-bin/smithers` shim now gives the controller PATH a runner on every host, and it says every producer attempt writes the record, since per-node cloud workers are gone. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ort record The verifier's "never recorded" error told operators to run `resume --refresh-controller --retry-failed`, but `--retry-failed` resets every failed or stalled task, so in a best-effort campaign it would also rerun each failed strategy task. Both record errors now name `resume <run-id> --refresh-controller --reset-node <producer node>`, which resets the producer's latest attempt and its dependents, here its verifier. A scratch run against the pinned Smithers confirmed that this timetravel resets exactly node:report and verify:report and that the re-dispatched attempt rewrites the record. A malformed record blocked every re-dispatched attempt above 1 without naming a way out; its error now names the file to delete. The history test now asserts both errors' exact text and pins that a first attempt replaces an unreadable record instead of failing on it. The docs paragraph states the two cases where failed_attempts is inexact and that a Codex workspace-write agent can reach the record when the project lives under /tmp or $TMPDIR. The CHANGELOG entry names the mixed plain and refreshed resume failure with its recovery. The integration test's poison-stub comment no longer describes the controller PATH, which the run's trusted-bin/smithers shim now gives a runner. Co-Authored-By: Claude Opus 5.5 <[email protected]>
c9b278a to
1fde8c7
Compare
…shim Problem: a Breaking-changes entry announced that launch, resume, replay and fork no longer write <run>/trusted-bin/smithers. No release wrote it: #1201 added it after v0.1.2, and this branch already rewrote #1201's entry without it. Measured against v0.1.2, nothing the entry lists changes: replay and fork never refused an unpatched install, and a run launched by an earlier release already found its persisted workflow's bare `smithers` only on the operator's PATH. Change: remove the entry. Its one piece of advice moves to the #1183 entry, which already explains that plain resume keeps running the persisted workflow: such a run still runs `smithers node` in a restarted controller, so continue it with `resume --refresh-controller`. The #1183 entry already gives the --reset-node form for a report producer that has succeeded. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: #1201 wrote <run>/trusted-bin/smithers at launch, resume, replay and fork so that the generated workflow's bare `smithers node` call (#1143) found a runner. #1183 removed that call: final-report producer retries and verifiers now read the selections each producer attempt records in the run. The current template spawns only git, bash and the run's `ultrafuzz` launcher, and every command the runtime spawns runs an executable it bound by path, never `smithers` from PATH. smithers-report-retry.integration.test.ts puts a failing `smithers` first on the engine PATH around the production report code and still passes. The shim therefore served only workflows persisted by earlier releases, and it put an engine CLI first on every task's PATH. It also made replay and fork bind the installed runner, which they otherwise never run: they run the run's sealed engine. Change: delete writeTrustedSmithersShim and its three callers. When resume cannot re-verify the trusted CLI, it again keeps an existing run launcher first on PATH itself, which the shim had done as a side effect; the existing test for that path covers it. Launch still binds the installed runner before creating a run, because resume runs it; the comments and the doctor summaries now say launch and resume. The capability and native-continuation tests drop their shim assertions, and a resume test now asserts that trusted-bin holds only the ultrafuzz launcher. The docs no longer say that the shim serves the workflow's smithers calls or that tasks can drive their run through it. The CHANGELOG entry for #1201, which has not been released, drops its shim claims, and a new entry records the removal. BREAKING CHANGE: tasks no longer find a `smithers` CLI that Ultrafuzz provides on PATH. A run launched by an earlier release and continued by plain `resume` still runs its persisted workflow, whose bare `smithers` call in a restarted controller finds a runner only on the operator's own PATH, as before #1201; continue such a run with `resume --refresh-controller`. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…shim Problem: a Breaking-changes entry announced that launch, resume, replay and fork no longer write <run>/trusted-bin/smithers. No release wrote it: #1201 added it after v0.1.2, and this branch already rewrote #1201's entry without it. Measured against v0.1.2, nothing the entry lists changes: replay and fork never refused an unpatched install, and a run launched by an earlier release already found its persisted workflow's bare `smithers` only on the operator's PATH. Change: remove the entry. Its one piece of advice moves to the #1183 entry, which already explains that plain resume keeps running the persisted workflow: such a run still runs `smithers node` in a restarted controller, so continue it with `resume --refresh-controller`. The #1183 entry already gives the --reset-node form for a report producer that has succeeded. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The owner approved this PR's security-posture change: after a controller restart the verifier's reference for
agent_executioncomes from a run-directory record instead of the Smithers database, which changes the source but not the boundary (a same-UID agent could forge either) and retires #585's claim, with no migration for runs whose producer ran under an older template.Problem
On current main, a final-report producer retry or verifier that runs in a restarted controller (resume, quota park, supervisor relaunch, crash) cannot rebuild
report.json#run_metadata.agent_executionunless asmithersexecutable happens to be on the operator's ownPATH. When none is:unverified-report-inputs.tsaccepts a node that its verifier failed (read in the code, not exercised here), and the run ends FAILED.This reproduces with the pinned Smithers (0.35.0): with main's
workflow.tsx(previous round) and with the stacked base's (this round; #1201 keeps the shell-out), both variants ofsmithers-report-retry.integration.test.tsfail once a stubsmithersthat exits 97 is first on the detached engine'sPATH(table under Verification).command -v smithersis empty on this host.Overlap with #1201. This branch is stacked on #1201, which merges first. #1201 writes a
<run>/trusted-bin/smithersshim so the template's baresmithersresolves, and cites #1143 point 1, so on top of it the ENOENT no longer reproduces in production. What this PR adds there is that the report path no longer spawns a runner subprocess (a--full-outputread of up to 64 MiB under a 180 s budget, which #1026's contention forced). Instead it reads a file of about 30 bytes per attempt that the workflow wrote itself. See Merge notes.Root cause
Checked on main in the previous round (
a46a4960; #1229, which main has since added, does not touch these paths). On this stack, #1201 changes the third and fifth points: its shim puts asmithersfirst on the engine'sPATH, and the integration test used that shim.execFileSync("smithers", ["node", …, "--full-output"])with a bare command name. This is the template's only baresmitherscall.PATHcomes fromcomposeSmithersCommandPath(smithers.ts). It is the run'strusted-bin, which on main holds only theultrafuzzlauncher, followed by the operator'sPATHwith target-local directories (such as.smithers/node_modules/.bin) removed. The controller launches its runner by absolute path, and nothing adds a directory that contains asmithersexecutable toPATH.node_modules/.binto the engine'sPATH, andpnpm testprepends it too. So CI never ran the query without asmithersonPATH.Change
In
packages/runtime/src/templates/smithers/workflows/workflow.tsx(+54/−98):[{attempt, chainIndex}, …]withwriteFileDurableto<realpath(runRoot)>/smithers/final-report-selections/<attempt-id>.json.resume --retry-failedand--reset-nodeboth do exactly that for a failed verifier: they issuetimetravel --node-id <producer>without--attempt(smithers.ts), which resets the producer's latest attempt. The engine'snextAttemptNumbercounts only attempts that are not marked reset-cancelled, so that attempt number is dispatched again. A failed preflight never reachesgenerate, so it is never recorded; the old code skipped those attempts too.report producer selection was never recorded; run `ultrafuzz resume <run-id> --refresh-controller --reset-node <node>` to rerun the producer, with the producer's workflow node ID filled in (node:final-reportin the packaged topologies).--reset-nodeon the producer resets its latest attempt and its dependents. Smithers'timetravelresets dependents unless given--no-deps, and the--reset-nodepath passes none. Smithers counts as dependents the nodes whose attempts started after the target attempt. In every packaged topology, every node except__finish__(which depends on it) is upstream offinal-report, so for the report producer that is its verifier. The first version of this PR named--retry-failed, which resets every failed or stalled task (smithers.tsloops oversmithersFailedTasks), so in a best-effort campaign with failed strategy tasks it would also have rerun each of them.--refresh-controllerstays in the hint: without it, a run launched before this release goes back to its persisted workflow, which does not write the record.smithers/final-report-selections/<attempt-id>.jsonin the run directory) and the same rerun command. Only an outside edit can produce such a file, sincewriteFileDurablewrites a temp file and renames it. Without the deletion, every re-dispatched attempt above 1 reads it again and fails.finalReportAgentExecutionstill checks attempt order and chain bounds, so the same rules are not written twice.readFinalReportSmithersAuthority,priorFinalReportAgentSelections, the 180 sSMITHERS_REPORT_PRODUCER_AUTHORITY_TIMEOUT_MSbudget, and the workflow's imports ofinspectSmithersAttemptAgentSelectionandreconcileSmithersAttemptAgentSelection.workflow-sync.tsusesinspectSmithersAttemptAgentSelection.resumeof an in-flight run keeps its old code path.reconcileSmithersAttemptAgentSelectionhas no caller in the repository except its own tests (see Risk).Docs and CHANGELOG:
docs/reference/artifacts-reports.md: one paragraph. It says where the record lives and that each producer attempt writes it before starting its agent. It names the two cases wherefailed_attemptsis inexact: attempts that ran before a refresh replaced an earlier release's workflow are omitted, and after a reset that reuses attempt numbers a pre-reset attempt can be listed. It also says what the record is not: a boundary against an unsandboxed same-UID agent, or against a Codexworkspace-writeagent when the project lives under/tmpor$TMPDIR.CHANGELOG.md: one entry at the top of## Unreleased→### Breaking changes, not under Other changes. The reasons are the in-flight cases under Risk and the retired feat(runtime): add error-agnostic agent retries #585 claim. Main files such consequences of PRs without a!there too (fix(runtime): stop host artifact gates rejecting valid campaign output #1176, refactor(runtime): delete dead controller-generation, sealed-refresh, re-finalization and recovery-authority code #1193). The entry names the recovery for both in-flight failure modes. It describes the replaced query by its cost, not by a missingsmithersonPATH, which feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201's shim fixes.This PR reverses #585's test "a non-Codex fallback cannot forge final-report producer authority through the run filesystem" and deletes it. That property was never a real boundary:
permissionMode: "bypassPermissions"and OpenCode setsyolo: true, so an unsandboxed agent can edit the Smithers database (<target>/smithers.db) orstate.json.echo), not the boundary.trusted-cli.tsalready says the host is "reproducible execution evidence, not a same-UID sandbox boundary".<run>/workspaces/<attempt-id>) and its declared artifact roots. The Codex template setssandbox: "workspace-write"and the workflow passesaddDir: [task.artifactDir, ...dependencyArtifactDirs]. From reading the code, a sandboxed Codex agent is therefore not given either location as a writable root, unless the project lives under/tmpor$TMPDIR. Codex'sworkspace-writeleaves both writable by default, andcodex.tsxsets no exclusion.run.output_dirmust be project-local, so the record and the database are always exposed together. I did not verify the sandbox experimentally.The other half of that deleted test, that controller memory outranks the on-disk source, still holds for edits made after the controller read the record, and
report-retry-history.test.tschecks it. A producer attempt dispatched after a restart fills memory from the record. So an earlier attempt's agent that rewrote the record before the restart (for example to erase its own failure before a quota park) has that history carried into the prompt authority, and the in-process verifier accepts it. A security review confirmed this with a probe against the extracted helpers; it is the same #585 property.Deliberately not built (and why)
agent_executionand this PR's fail-closed record errors are tamper checks too. They are candidates for downgrading to warnings in fix(runtime): a renderer or projection change no longer strands runs whose dynamic groups expanded #1216's rethink, not in this PR, which replaces main's equivalent stop instead of adding one.smithers, and no reliance on feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201's shim for this path. Either one keeps a runner subprocess in the finalizer path, with a 64 MiB--full-outputread and a 180 s budget, just to read data the workflow already had when the attempt started.report producer is outside the sealed agent chainfromfinalReportAgentExecution, which also validates in-memory selections. Naming a record-specific recovery there would mean writing its checks again in the reader. The malformed-record recovery (delete the file, then reset the producer) applies unchanged.trusted-binon main, fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code #1171 fixed ENOENT being reported as a timeout, and feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 handles the TMPDIR controller. HenceRefs, notCloses.Verification
command -v smithersis empty on this host. This round ran on the stacked headc9b278a7afterpnpm install --frozen-lockfile,pnpm -w buildand a cleandist-testrebuild. Everything after it is from earlier rounds, before this stacked rebase, and says which head it ran on where the earlier description did. Since then the report path changed only in dropping the local-mode guard and the single-rung cloud verifier shortcut, which #1197's removal ofexecutionfrom task specs required.This round
Discriminating evidence. I wrote the stacked base's
workflow.tsx(#1201's5689f7f3, which still shells out) into the worktree, ran this PR's two test files against it, and restored it. The previous round got the same failures with origin/main'sworkflow.tsx.workflow.tsx(#1201)smithers-report-retry.integration.test.tsfallback variant (real Smithers 0.35.0; a stubsmithersthat exits 97 is first onPATH)Smithers report-producer authority is unavailable after 1ms against a 180000ms budget, caused byCommand failed: smithers node node:report …with the stub'sbare smithers resolved from PATHand"status":97report-retry-history.test.ts: a restarted producer and a restarted verifier rebuildagent_executionfrom the run, and controller memory outranks the recordSmithers report-producer authority is unavailable after 0ms …, caused byReferenceError: execFileSync is not defined(the harness injects noexecFileSync)Smithers report-producer authority is unavailable after 0ms …, with the same causeMutation checks, each applied to
workflow.tsx, run, and then reverted:task.execution.modereads that refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197 made unsafe (thelocalguard on the read and the write, and the single-rung cloud shortcut in the verifier) fails all five history tests withTypeError: Cannot read properties of undefined (reading 'mode'), because the fixtures no longer carryexecution.node:report failed 3 consecutive attempts with an identical error: undefined is not an object (evaluating 'task.execution.mode').execution: { mode: "local" }, so such a leftover would have passed them and only thecli-e2elane would have caught it.What I ran on the stacked head:
report-retry-history.test.js: 5/5. The cloud test is deleted (see Rebase notes).smithers-report-retry.integration.test.js: 2/2.generated-workflow-verifier,generated-workflow-render,generated-workflow-footprintinvariant-suite-{ancestor-order,handoff-durability,enumeration-overflow}workspace-patch-{supersede,replay},workspace-preparation-lifecycle,stale-workspace-cleanup-overflowdynamic-lifecycle,dynamic-workflow,pinned-submodulesterminal-report-projection,task-workflow-identity,workflow-dependency-policy,smithers-attempt-authoritycloud-worker-handoff, which the previous round also ran, is gone with refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197.smithers-{preparation-race,dependency-skip,resume-reopen}.integration: 9/9.runtime.test.tswith--test-name-patternfor the six resume tests named under the targeted recovery below: 6/6.packages/cli/test/e2e/campaign-resume.test.ts, the wholecli-e2elane (pnpm --filter @ultrafuzz/cli test:e2eruns only this file): 1/1 on this head (273 s). Itsfinal-reportnode declaresreport@3andnonempty-markdown@1and succeeds on attempt 1, so the run goes through the now unconditional record write and the in-memory verifier, and the report endsavailable,completeandverified. It does not reach the restart read, which the integration test covers.workflow.tsxtranspiled with Bun'stsxloader, the loader the engine uses.pnpm -w format:checkpnpm -w lintCI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/claude/v10-install-based-controller pnpm -w lint:strict:cipnpm -w knipafterpnpm -w build, and again with everydistremovedpnpm --filter @ultrafuzz/runtime typechecknode scripts/docs-check.mjspnpm typecheckcovers onlysrc/**/*.ts, so it does not type-checkworkflow.tsx. ESLint checks the template for undefined names, and the template's runtime coverage comes from the Bun-run integration tests and the e2e run.Not run this round: the rest of
runtime.test.ts, the CLI unit suite, the scratch recovery and scratch e2e variants below, and a live campaign with real model agents.Earlier rounds
The targeted recovery against real Smithers 0.35.0 (previous round, on
3a1d3430). This is a scratch variant ofsmithers-report-retry.integration.test.tsand is not committed.[{"attempt":1,"chainIndex":0},{"attempt":2,"chainIndex":1}]. I deleted it and resumed.report producer selection was never recorded; run `ultrafuzz resume <run-id> --refresh-controller --reset-node node:report` to rerun the producer.timetravelthatresume --reset-node node:reportsends for a producer that is not itself failed:--node-id node:report --no-vcs --force, with no--iteration. It reset["node:report","verify:report"]and nothing else.[{"attempt":2,"chainIndex":1}]. After one more resume the verifier passed, and the runfinished. The four phases that log (three producer attempts and the passing verifier) ran in four distinct controller pids. The rerun'sagent_executionlists no failed attempts: attempt 1 went with the deleted record, the kind of gap the docs paragraph now names.The ultrafuzz side of
--refresh-controller --reset-nodecomes fromruntime.test.ts.an unobserved failed occurrence stays in the attempt ledger across a stopped refreshpassesresetNodeandrefreshControllertogether, and it passes along with these tests:resetandretryvariants;resume --reset-node does not repeat a committed reset after a failed continuation;resume --retry-failed recovers a node whose verifier rejected its output;a refresh resume reuses its own ownership inspection instead of inspecting twice.start-run.tspasses the path rendered by--refresh-controllerinto the samerunSmithersLifecycleCommandcall that carriesresetNode. The round before that ran the same recovery through thetimetravelthat--retry-failedissues (--iteration 0, no--no-deps), with the same result.Mutation checks, each applied to
workflow.tsx, run, and then reverted:3a1d3430, each of these fails only the malformed-record test:attempt > 1guard, after which a first attempt fails on an unreadable record instead of replacing it;task.attemptIdinstead oftask.smithersNodeId;--retry-failed;controller memory outranks the run record. A re-dispatched attempt keeps its old recorded rung, and only the re-dispatch test fails.End to end through the ultrafuzz CLI, with a producer retry (two rounds back; since then, the code it ran changed only in error text and, in this round, in dropping the local-mode guard). This is a scratch variant of
packages/cli/test/e2e/campaign-resume.test.tsand is not committed. The only differences are that the stubcodexholdsfinal-reportinstead ofsummarize, plus checks on the record and on the report. Each command runs as its own process on the pinned engine:init,run, a SIGKILL of the engine, the supervisor's relaunch, a SIGKILL of the whole controller, thenresume,status,reportandstats.[{"attempt":1,"chainIndex":0},{"attempt":2,"chainIndex":1}], and the agent started. On main, that attempt would shell out to the baresmithersbefore its agent starts. That comes from the code; I did not run this variant against main.resumestarted a fresh controller, which dispatched attempt 3 on rung 2, again with an empty cache, and the attempt completed.[{"attempt":1,"chainIndex":0},{"attempt":2,"chainIndex":1},{"attempt":3,"chainIndex":2}].agent_executionlists failed attempts 1 and 2 and producer attempt 3.statusreports the runsucceeded, with the reportavailable,completeandverified, and every check in the original test passed (406 s).The integration test keeps its existing assertions:
failed_attemptslists attempt[1], and the producer is attempt 2.[0,1]for the fallback variant and[0,0]for the quota variant.The detached engine's
PATHstarts with a stubsmithersthat prints to stderr and exits 97, and the fixture's ownpausecall uses the runner's absolute path.Deleted tests that pinned the removed implementation:
loadFinalReportAgentExecutionAuthorityharness with a fakeexecFileSyncreport-retry-history.test.ts)smithersReads() === 0testreconcileSmithersAttemptAgentSelection(task, authority.authorityDetailcall, thedetail.ok === true && …envelope unwrap, and the absence of anagent-execution/execution.jsonrecordWhat I ran on
3a1d3430(onmaina46a4960):report-retry-history.test.js: 6/6smithers-report-retry.integration.test.js: 2/2generated-workflow-verifier(116) andcloud-worker-handoff(14)invariant-suite-{ancestor-order,handoff-durability,enumeration-overflow}workspace-patch-{supersede,replay},workspace-preparation-lifecycle,stale-workspace-cleanup-overflowdynamic-lifecycle(19),dynamic-workflow,generated-workflow-footprint,pinned-submodulesterminal-report-projection,task-workflow-identity,workflow-dependency-policy,smithers-attempt-authoritysmithers-{preparation-race,dependency-skip,resume-reopen}.integration: 9/9runtime.test.tswith--test-name-patternfor the six resume tests above: 6/6workflow.tsxtranspiled with Bun'stsxloader, the loader the engine uses.npx prettier --checkon the changed files, andpnpm -w format:checkpnpm -w lintpnpm -w knipbefore the build (everydistremoved) and afterpnpm -w buildCI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:cipnpm --filter @ultrafuzz/runtime typechecknode scripts/docs-check.mjsNot run on
3a1d3430:runtime.test.tsand the CLI unit suite;cli-e2elane on3a1d3430. A review rancampaign-resume.test.tson594d6b87: 1/1, 309 s. Its report producer succeeds on attempt 1, so it does not reach the lines that round changed.Risk / compatibility
agent_executioncomes from a run-directory file, not the Smithers database. A same-UID unsandboxed agent could forge either one, so this changes the source, not the boundary, although the file takes less effort to forge. It retires feat(runtime): add error-agnostic agent retries #585's claim that the run filesystem cannot be used to forge producer provenance. Controller memory outranks edits made after that controller read the record; a producer attempt dispatched after a restart starts from the record.resumeruns the workflow persisted at launch (start-run.tsreadsrun.json#workflow.path), and nothing on the resume path rewrites that pointer.--refresh-controllerrenders into.smithers/continuations/<uuid>/and passes that path only to its own lifecycle call. So existing runs keep the old code, including the shell-out. Onlyresume --refresh-controllerrenders this template, and a later plainresumegoes back to the launch workflow.--refresh-controller --reset-node <node>on purpose. Without the refresh, the reset would go back to the persisted workflow, and its re-dispatched producer attempt would shell out again, since that attempt is above 1 whenever the producer needed a retry.agent_execution.failed_attemptsomits the attempts that ran before the refresh.failed_attemptsomits the attempts that ran before the refresh. Before this change, on main, those attempts failed outright unlesssmitherswas on the operator'sPATH; on top of feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 they kept full provenance.run_metadata.agent_execution differs from the controller-observed producer, which names no recovery.--refresh-controller --reset-noderecovery fixes it, because the re-dispatched attempt rewrites the record. The CHANGELOG names this case and the recovery.smitherson the operator'sPATHlose anything.resumestill works: feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 writes the shim on every resume (start-run.ts:630), and the persisted workflow's shell-out resolves through it.failed_attempts. This affects provenance only; the run does not stop. The docs paragraph now states this gap and the refresh gap. Greptile raised the same gap; it is declined here (see Rebase notes).reconcileSmithersAttemptAgentSelectionhas no caller in the repository after this PR; main'sworkflow-sync.tsno longer uses it. It stays exported because plainresumeruns the persisted pre-change workflow against the installed runtime (ULTRAFUZZ_RUNTIME_MODULE: import.meta.resolve("@ultrafuzz/runtime"),start-run.ts:583); only launch uses the sealed snapshot copy. Then delete it with its tests insmithers-attempt-authority.test.tsandtask-workflow-identity.test.ts.smithersRunIdis dead in the same way. feat(runtime): add error-agnostic agent retries #585 added it to the serialized task specs (smithers.ts:8023) forsmithers node -r task.smithersRunId, and fix(runtime): preserve reconstructed task workflow identity #1104 copied it onto generated tasks (workflow.tsx:418). This PR deletes its only reader. Delete the field from both places, with the asserts intask-workflow-identity.test.ts(thegenerated tasks inherit the sealed workflow run identitytest exists only for it) anddynamic-workflow.test.ts:149, in the same follow-up.Merge notes
claude/v10-install-based-controller(5689f7f3), which sits on refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197's. The approved merge order is refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197, feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201, this PR, then feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198. Until refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197 and feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 merge, GitHub's diff againstmainincludes their commits; this PR's own six are5689f7f3..c9b278a7.cli-e2e. Green on this head:campaign-resume.test.ts, the lane's only test, passed 1/1 (see Verification).smitherscall, so feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201's<run>/trusted-bin/smithersshim serves only workflows persisted by earlier releases, which plainresumekeeps running.writeTrustedSmithersShimdocstring (smithers.ts:6280),docs/reference/cli.md:359, feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201's CHANGELOG entry ("which gives the workflow's ownsmitherscalls a runner"), and a comment insmithers-preparation-race.integration.test.ts:256.trusted-bincomes first on the controllerPATH, and agents inherit thatPATH(agents/environment.tsx), so until then every agent sees a working engine CLI on itsPATH. A same-UID agent can reach the runner by its path anyway, so this is least privilege, not a boundary.resume, and the--refresh-controller --reset-noderecovery is their way out. Per owner policy there is no fallback.docs/security.mdamendment, so thattrusted-binagain holds only the validator launcher.Rebase notes
Stacked rebase onto #1201 (
5689f7f3). This PR's six commits moved frommaina46a4960ontoorigin/claude/v10-install-based-controller5689f7f3: #1201's five commits on #1197's fourteen, onmain538b6188, which adds #1229. The pre-rebase head was3a1d3430, and the new head isc9b278a7. This PR changes no lockfile, patch or manifest, sopnpm install --frozen-lockfilewas enough. Commits 2 and 3 applied unchanged (git range-diffshows=). The conflicts, and how each was resolved:workflow.tsx(commit 1, with refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197, infinalReportAgentSelectionsForAttemptandauthoritativeFinalReportAgentExecution; commit 4 again in the same comment).const local = task.execution.mode !== "cloud", thelocal &&guard on the read, theif (local)guard on the write, and the verifier's single-rung cloud shortcut.taskSpecsFromCompiledno longer emitsexecution, and nothing type-checks the template, so any of these would throw aTypeErroron every report-producer attempt before its agent starts. The template now readstask.executionnowhere.generated-workflow-verifier.test.ts(commit 1). Both sides deleted the single-rung cloud restart test. I kept that deletion and this PR's deletions of the feat(runtime): add error-agnostic agent retries #585 forge test and the 180 s budget test. The rebased deletions equal the original ones minus the cloud test.smithers-report-retry.integration.test.ts(commit 1, with feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201). feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 had replaced the samecli()PATHwithcontrollerPath(root), which puts the run'strusted-binshim first.pause.controllerPathand thewriteTrustedSmithersShimimport. With the shim first, the stub would never be reached, and the stub is what proves that the report path makes nosmitherscall.CHANGELOG.md(commits 5 and 6, with refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197's new top entry). This PR's entry goes above refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197's, which is unchanged.Changes made during the rebase that no conflict forced:
executionfrom the history test'sTasktype andtaskFixture(with itsmodeoption), and theexecution: { mode: "local" }line from the integration fixture'stask. A leftovertask.executionread now fails both files (mutation check under Verification).smithers node, which Ultrafuzz never put on the controllerPATH" became "instead of runningsmithers nodeas a subprocess (a--full-outputread of up to 64 MiB under a 180 s budget)". The sentence that started "Unless the operator's ownPATHhad asmithers" is gone, since feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 fixes that failure and claims it.PATHwithout a runner. Commit 4's subject lost "and qualify upgrade claims", because what is left of that commit changes only the template.start-run.ts:583and:630,smithers.ts:8023,workflow.tsx:418.Greptile (review of
3a1d3430):workflow.tsx:3188(now line 2236), "Stale failed-attempt history": declined. The observation is right. After a reset that reuses attempt numbers, a new-round attempt that fails before recording (in preflight, say) leaves the old round's entry for its number, and later attempts list it inagent_execution.failed_attempts. It affects provenance only: the producer's prompt and the verifier read the same record, so the run does not stop. The code comment, the docs paragraph and Risk ("Resets and provenance") already state it. A fix needs a host-side hook into the reset path to clear the record, which "Deliberately not built" explains.Earlier rebase. Rebased onto origin/main
a46a4960, 52 commits past the old baseb6dd1da9. The pre-rebase head wascef55b10and the rebased head is594d6b87. Four of its five commits were kept, and one CHANGELOG commit was added. The review fixes are one commit on top of594d6b87.packages/runtime/test/generated-workflow-verifier.test.ts: this PR deletes three tests: the feat(runtime): add error-agnostic agent retries #585 forge test, the single-rung cloud restart test and the 180 s budget test. test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203 deleted the neighbouring source-text testgenerated retries do not inspect or inject previous failure text. I kept both deletions and dropped the closing brace the two hunks shared. TheloadFinalReportAgentExecutionAuthorityharness and the three source-regex asserts applied cleanly.CHANGELOG.md: the entry that the docs commit added under Other changes conflicted with main's new entries. The branch's last commit (chore: move the changelog entry to the consolidated release notes) had removed that entry again. I kept main's file through the rebase, dropped that commit, and added one entry at the top of### Breaking changesin a new commit.workflow.tsx,report-retry-history.test.ts,smithers-report-retry.integration.test.tsanddocs/reference/artifacts-reports.mdapplied cleanly.594d6b87, every changed line in these files, and in the verifier test, was identical to the pre-rebase branch.pnpm install --frozen-lockfilewas enough.Refs #1143, #585
🤖 Generated with Claude Code
The PR does not appear safe to merge while the previously reported cloud cleanup defect remains outstanding.
Fix with agent prompt
Summary
The PR records final-report producer selections in the run directory so a restarted controller can rebuild
agent_executionwithout spawningsmithers. It adds recovery guidance and tests, and documents the changed provenance and security posture. No changes were made since the previous review.Diagram
sequenceDiagram participant P as Report producer participant R as Run-directory record participant V as Verifier P->>R: Write selected attempt and chain rung P->>P: Run agent and remember selection alt Same controller V->>V: Compare report with remembered selection else Restarted controller V->>R: Read recorded selections V->>V: Compare report with reconstructed execution endReviews (3) · Last reviewed commit: "fix(runtime): name a targeted recovery f..."