refactor!: remove per-node cloud execution (execution.mode = "cloud") - #1197
Merged
Merged
Conversation
aviggiano
force-pushed
the
claude/v04-remove-cloud-node-execution
branch
from
September 29, 2026 10:55
f212184 to
52cb250
Compare
This was referenced Sep 29, 2026
Merged
aviggiano
force-pushed
the
claude/v04-remove-cloud-node-execution
branch
from
September 30, 2026 08:24
52cb250 to
57daecf
Compare
aviggiano
marked this pull request as ready for review
September 30, 2026 09:56
This was referenced Sep 30, 2026
Merged
…ION_CLOUD_REMOVED Per-node cloud execution cannot complete a node on main or in any release (v0.1.0 to v0.1.2). Two checks reject every cloud node before inference: - The node worker requires realpath(data root) to equal the lexical path. Modal's filesystem-v2 /data mount is a provider symlink (#712 live probe), so initializeDurableNodeWorkspace throws "durable data root is unsafe" for any data root below it. - The worker exports its /proc/<pid>/fd/<n>/... descriptor path as ULTRAFUZZ_WORKFLOW_PERSISTED_PATH. The generated workflow evaluates that path at module load and throws "persisted workflow path must stay inside the cloud handoff project" (#713). The fixes for both landed only on release/v0.1.0, a branch that no longer exists; all three release tags carry the unfixed worker. This is the decision commit of the removal. validateExecutionConfig now reports CONFIG_EXECUTION_CLOUD_REMOVED for execution.mode = "cloud", execution.provider and [execution.providers.modal], one diagnostic per path, so planning fails before a run directory exists. The gate sits in resolve, not in the TOML parser, because artifact gates re-parse each run's persisted config.resolved.toml mid-run. The rest of [execution] stays accepted and inert: ultrafuzz init scaffolds it and every persisted run config carries it. Removing it would change the config schemas, so it is left for later. Tests: config.test.ts's three cloud tests and four cloud diagnostic cases become one removal test (which also round-trips the table init writes) and one named-diagnostic case. The six runtime tests that fed cloud TOML through the loader are deleted; one of them becomes "planRun rejects removed cloud execution before creating the run directory". Refs #134, #712, #713 Co-Authored-By: Claude Opus 5.5 <[email protected]>
… workflow and lifecycle
With execution.mode = "cloud" rejected at resolve time, the cloud branch
of every runtime layer is unreachable. Delete it:
- Compiler (smithers.ts): compileTask always emits the inert local
execution block {mode: "local", resources, agentCredentialEnv: []}, the
exact value it already produced for local runs, so tasks.json and the
artifact schemas are unchanged. The cloud timeout reservation, the cloud
credential classifier and its re-check, the relative-path rewriting for
relocated workers, the modal module entry and the reset-node
cloud-execution-generation.json write are gone. The config comments
on resourceTimeoutOrigin now describe its one remaining use
(serialization) rather than the removed cloud timeout cap.
- Generated workflow: the cloud-worker input envelope, the selected-task
handoff DTO, the Modal Sandbox branch, the top-level dynamic import of
@ultrafuzz/modal and the react Fragment import are gone. Every attempt renders
as the existing Worktree prepare/agent/verify triplet.
- Lifecycle: doctor and run probe required commands on the local PATH
only (no Modal app creation), clean no longer imports @ultrafuzz/modal
to terminate a run's purpose=ultrafuzz-node sandboxes and delete its
ultrafuzz-node-* volume, and data governance, validate, retry-chain and
start-run lose their cloud-only branches.
- The runtime cloud-execution-generation contract and schema are deleted.
Only CLI `json validate` reads the runtime schema bundle digest.
- Secret scans: the generated verifier's pre-publication scan uses
sensitiveEnvironmentValues(process.env) alone, the value it already used
for local tasks, whose agentCredentialEnv is always empty. The scan of a
verifier-rejected final report (unverified-report-inputs.ts, #1204) uses
the same value, and currentReportAttempt no longer returns the task's
execution block.
The runtime no longer references @ultrafuzz/modal, so the five runtime
release-validation lanes stop building the modal closure.
workflow-sync.ts keeps its execution.mode === "local" filters: every task
compiled from now on passes them, and parallel changes are editing that
file.
Tests: cloud-worker-handoff.test.ts and its harness, the only tests that
rendered the whole generated program in process, are deleted with the
remaining cloud cases. generated-workflow-render.test.ts replaces that
coverage for local runs: it imports a compiled two-node workflow in
process and checks the rendered Worktree/Task tree and its dependency
edges. A new clean test removes a run whose plan.json recorded cloud
execution; before this change clean failed on it with
CLEAN_CLOUD_STORAGE_FAILED. Template-slice markers and counts that
pointed into the deleted code are updated. The tests that named
credentials through a cloud task go: the verifier test for configured
credential names (with its task `execution` field and the harness's
scanProcessEnvironment option) and the unverified-report test that named
them through a cloud task manifest. The remaining secret-gate tests cover
the environment-name heuristic and the vendor-format patterns that still
run.
Refs #134
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…schemas and SDK patch With the runtime no longer dispatching cloud nodes, nothing imports the controller or worker half of per-node Modal execution. Delete it: - packages/modal: node-provider.ts (sandbox provider, handoff archive, result polling, probeModalCommands, cleanupModalNodeRun), node-worker.ts, and their archive helpers modal-download.ts, deterministic-archive.ts and safe-archive.ts, with their tests. - The seven node schemas (node input, result, checkpoint, checkpoint index, restore, worker error, execution-dependency manifest), their modal-common definitions, contracts, registry entries and semantic gates. The Modal schema registry drops from 21 to 14 schemas. - packages/artifacts: the controller-to-worker selected-task DTO (cloud-selected-task.ts) and its export. - packages/config: the MODAL_* sandbox lifetime constants. - patches/[email protected] and its patchedDependencies entry. The patch only rewrote SandboxFilesystem.copyToLocal, which only modal-download.ts called. [email protected] stays and is now unpatched. packages/modal also drops tar-stream; packages/cli keeps its own tar-stream dependency. - scripts/ci: the modal-sdk-download and safe-archive Bun tests. The ultrafuzz-modal eval runner, which runs a whole local campaign inside one sandbox, is unchanged. It imports none of these modules. Schema identities: VALIDATOR_BUILD_IDENTITY and the artifact and config schema bundle digests are byte-identical to the base. The Modal bundle digest changes. It is computed at validation time only (Modal document parsing and `json validate`); no persisted document records it. Tests: the CLI json-validate Modal case and the Modal document tests now use modal-smoke-checkpoint instead of the deleted node-input schema. The Modal test fixture currentTaskOutputBinding is no longer exported: node-provider.test.ts was its only importer. Refs #134, #712, #713 Co-Authored-By: Claude Opus 5.5 <[email protected]>
scripts/validate-release.mjs imports packages/artifacts/dist when it loads. The five runtime lanes only got that build as a side effect of build_modal_dependencies, which the runtime dispatch deletion dropped. On a fresh checkout, `node scripts/validate-release.mjs --gates runtime-1` then fails with ERR_MODULE_NOT_FOUND before it selects a gate, so all five lanes would fail on every pull request. Every lane runs validate-release.mjs, so the reporter build is now unconditional and the build_release_reporter lane flag is deleted. A new lane-table test fails when a package that validate-release.mjs imports from packages/*/dist is not built by an unconditional step before `validate:release`; it fails on the previous workflow. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
The earlier commits deleted the Modal node provider, the only production caller of two exports: - runtime: verifyCommittedControllerGenerationAuthority and its CommittedControllerGenerationAuthority result. #1193 deleted every other export of workflow-controller-generation.ts and kept this one only for the node provider, so the whole module goes: the journal and manifest readers it keeps are private to that verifier. The runtime index stops re-exporting it. - artifacts: referenceArtifactManifestAuthorityForArtifactDir. No generated-workflow template in v0.1.0, v0.1.1, v0.1.2, main or integration/wave1 references either name, so resuming an older run is unaffected. The artifacts test that called its export now asserts the same fact through a live path: the parsed manifest keeps its reference authority. The runtime verifier's only test was the dynamic controller-refresh test, which #1193 deleted with the refresh code. In the generated workflow, dynamicExecutionPath(task, value, label) read `task` only for the removed cloud branch and had become a copy of currentProjectPath(value, label). It is deleted, its 12 call sites call currentProjectPath, and the seven path.resolve(process.cwd(), ...) wrappers around its already absolute result are gone. The verifier test's one-element local/cloud loop is inlined. Stale cloud wording goes from six code comments, the artifact-contract migration reference (the node schemas are deleted, not retained) and the harness research page. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
The deleted cloud config tests were the only coverage for [execution] behaviour that still runs for local projects: validate and run call validateExecutionNodeOverrides (CONFIG_EXECUTION_NODE_UNKNOWN), every compiled task records resolveExecutionResources in tasks.json, and the resolved TOML serializes [execution.nodes.*.resources]. A local-mode test now checks the node-override merge, an unknown node, the serialization round trip and a cpu = 0 rejection. Swapping the override precedence or inverting the unknown-node filter fails it. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
Adds the breaking entry for this change under Unreleased. Two earlier Unreleased entries described cloud-mode behaviour this change deletes, so their cloud clauses go: the review-timeout entry (#1150) no longer tells cloud users to add per-node timeout overrides to avoid CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED, an error that no longer exists, and the Codex openai_base_url route entry no longer says cloud planning rejects that config, since cloud planning is gone. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…tion A run planned with `[execution] mode = "cloud"` before the removal still records `execution.mode = "cloud"` for every task in its `smithers/tasks.json`, and its persisted `resolved-config.json` still parses, because the JSON schema accepts the value and the new CONFIG_EXECUTION_CLOUD_REMOVED gate lives only in `resolveConfig`. `resume --refresh-controller` therefore rendered the current controller, which has no Sandbox branch and ignores `task.execution`, and continued the run with every remaining agent attempt running as a local worktree task on the operator's host, with no diagnostic. A native resume ran the historical workflow instead, which review found fails with an unrelated module-not-found error for the deleted Modal module. `submitSmithersContinuation` now reads the task manifest it already parses and, when any task was planned for cloud execution, returns WORKFLOW_CLOUD_EXECUTION_REMOVED before it inspects, renders or starts anything. That covers both native and refreshed resume. Replay and fork are unchanged: they run the historical sealed workflow rather than the current controller, so they cannot move a run onto this host. A new runtime test launches a local run, rewrites its task manifest to what the pre-removal compiler wrote for a cloud config, and checks that both resume forms fail with only that diagnostic and issue no Smithers command. With the guard removed, the native resume succeeds and the test fails. The changelog entry now says such runs cannot be continued, and tightens two instructions reviewers flagged: setting `mode = "local"` alone does not clear the error (provider and the providers table must go too), and the leftover `ultrafuzz-node-*` volumes live in the Modal app the run's `plan.json` records until `ultrafuzz clean` deletes it. The configuration reference says the same, and its `[execution]` row no longer calls the table "retained for compatibility": it is validated but has no effect. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…onment secret scan The rebase deleted #1203's test that the generated verifier refuses to publish a configured credential's value, together with the harness's `scanProcessEnvironment` option, because that test covered the cloud credential-name lists this PR removes. That option was the only path that passed the real `sensitiveEnvironmentValues` into `verifyArtifacts`, so the pre-publication scan of environment values, which still runs for every local task, lost its only test: replacing `sensitiveEnvironmentValues(process.env)` with `[]` in the template's `verifyArtifacts` left all 115 verifier tests passing. The only remaining verifier secret test uses a vendor-format token, which pattern matching catches without any environment value. The harness option is back, and a local test replaces the deleted one. It writes a value with no vendor format into `result.json`. Held only by ULTRAFUZZ_TEST_GATEWAY_LABEL, a name the credential heuristic does not match, the value is published; once ULTRAFUZZ_TEST_GATEWAY_API_KEY also holds it, `verifyArtifacts` throws "contains sensitive data" and publishes nothing and writes no marker. With the template's scan emptied as above, the new test fails. Two agent-failure redaction tests no longer configure any credential name (the task-level lists were cloud-only), so their names now say they redact credential-named environment values, which is what they check. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
MODAL_CAMPAIGN_ENVELOPE_TOO_SHORT, which the eval runner raises when the public full-benchmark row budget cannot hold the exhaustive campaign, still told operators to "use local or per-node Modal execution", and the Modal eval how-to described a 1,800-second per-node lifecycle reserve and an explicit resource cap that this PR deletes. An operator following that advice now gets CONFIG_EXECUTION_CLOUD_REMOVED, whose message links to this same how-to page. Both now say to run the four-hour profile locally until the full benchmark's envelopes are enlarged, and the how-to drops the reserve and cap sentences. The tests match only the diagnostic's prefix, so none change. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…message A config that sets `mode = "cloud"`, `provider` and `[execution.providers.modal]` gets one CONFIG_EXECUTION_CLOUD_REMOVED diagnostic per setting, but all three carried the same message, and plain-text `ultrafuzz validate` and `doctor` print messages without their paths. The operator saw one ~330-character line three times with nothing to tell them apart. Each message now starts with its setting, for example `execution.provider is no longer supported: per-node cloud execution was removed; ...`. The JSON paths are unchanged. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…does not resolve `diagnoseProject` skipped the required-command probe whenever the project config failed to resolve, and then reported every required command as missing. The gate existed because the probe needed the config to choose between the local PATH and a Modal sandbox; after this PR, `probeCommandsForExecution` takes no config, so the gate only produced a false DOCTOR_TOOLCHAIN_MISSING for git, node and forge next to every upgraded cloud user's CONFIG_EXECUTION_CLOUD_REMOVED. Doctor now always probes the local PATH. A new test gives doctor a `mode = "cloud"` config and an all-available probe, and checks that git, node and forge are probed and the toolchain check is ok; with the gate restored it fails. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
Review found code and wording the earlier cleanup commit missed because
neither TypeScript nor knip flags them:
- `compileTask`'s `env` input: its only reader was the deleted
`cloudAgentCredentialEnv`. `startRun` still built
`{ ...process.env, ...input.env }` and passed it through
`compileSmithersWorkflow`, where nothing read it. The field goes from
`SmithersCompileInput` and `compileTask`, with both call-site props and
the `startRun` argument.
- `workflowModuleEntryUrls()` and the empty-string filter on its result
existed only for the conditional `modal: ""` entry. Both URLs are now
always present, so the two `import.meta.resolve` calls are inlined.
- The fake runners' `SMITHERS_FAKE_CLOUD_ENV_LOG` hooks, which printed
MODAL_TOKEN_ID and MODAL_TOKEN_SECRET, and the
`SMITHERS_FAKE_CLOUD_ENV_LOG`/`SMITHERS_FAKE_CLOUD_SELECTOR_LOG` entries
in the Smithers test environment allowlist: only the deleted cloud
credential-forwarding tests set them.
- The Effect-pinning comments in smithers-package.ts, runtime.test.ts and
the workspace override CI test justified the pin with two cloud
containers; the reason holds for a run's launch and a later resume,
which install at different times. The pnpm-workspace.yaml copy is left
to #1201, which rewrites that block.
- A stateful-profile test named "preserves cloud timeout inheritance" now
only checks that the resource timeout origin survives TOML snapshots.
Refs #134
Co-Authored-By: Claude Opus 5.5 <[email protected]>
The resume guard for removed per-node cloud execution read the mode from
`smithers/tasks.json`. Native resume treats that manifest as optional:
when it is missing or does not parse, the guard was skipped and the
historical workflow of a pre-removal cloud run went to Smithers instead
of failing with WORKFLOW_CLOUD_EXECUTION_REMOVED (Greptile P1 on
start-run.ts).
The pre-removal compiler gave every task `config.execution.mode`, so the
run's `smithers/resolved-config.json` is the record of that mode.
`submitSmithersContinuation` now reads that config first and returns
WORKFLOW_CLOUD_EXECUTION_REMOVED when it says cloud, before it reads the
task manifest. A config that is missing or does not parse fails the
resume with WORKFLOW_LIFECYCLE_FAILED ("run <id> cannot be resumed
without its resolved config, ...") for a plain resume too. Before, a
plain resume continued such a run with default concurrency and lease
and no ULTRAFUZZ_CONFIG_PATH, so its planned execution mode was never
checked; `--refresh-controller` already failed in both cases.
With the config always present, the continuation's fallbacks for a run
without one are dead and are deleted: the forge-guard and
ULTRAFUZZ_CONFIG_PATH branches, the `config?.` defaults for concurrency,
workspaces and lease, and recordNativeContinuationState's legacy
deadline branch. submitSmithersContinuation's complexity drops from 51
to 35.
Tests: the cloud-refusal test now rewrites the resolved config as well
as the task manifest, as a pre-removal cloud launch wrote both, and
checks that both resume forms are still refused once the manifest is
malformed and then missing; on the previous guard the plain resume of
the malformed-manifest run succeeds. A new test damages a local run's
config (malformed, then missing), checks that both resume forms fail
before any Smithers command, then resumes the run once the config is
restored; on the previous code the plain resume with a malformed config
succeeds. Three fixtures that hand-build a run directory (the unsealed
legacy launch test and the two real-engine continuation tests) now
publish a default local resolved config through a shared test helper.
The CLI reference and the changelog entry describe the new refusal.
Refs #134
Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/v04-remove-cloud-node-execution
branch
from
September 30, 2026 12:58
533dbfd to
b8a2bc5
Compare
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…ead of querying smithers After a controller restart (resume, quota park, supervisor relaunch) the generated workflow's process-local record of which chain rung each report-producer attempt used is gone. The producer's retry prompt and the verifier both rebuilt it by running a bare `smithers node ... --full-output` from inside the workflow: a runner subprocess on the finalizer path that reads up to 64 MiB under a 180 s budget. Wherever the controller PATH resolved no `smithers`, every such attempt failed with "Smithers report-producer authority is unavailable" (cause: ENOENT) until the report node ran out of retries (#1143). The only real-Smithers test of this path gave the detached engine a `smithers` on its PATH. Record each executed selection {attempt, chainIndex} with writeFileDurable in <run>/smithers/final-report-selections/<attempt-id>.json before the agent runs, and read it back when the process-local cache is empty. A re-dispatched attempt number replaces its own entry; finalReportAgentExecution keeps validating attempt order and chain bounds. This deletes readFinalReportSmithersAuthority, priorFinalReportAgentSelections, the 180 s SMITHERS_REPORT_PRODUCER_AUTHORITY_TIMEOUT_MS budget and the workflow's use of the Smithers attempt reconcilers. The runtime exports stay: host-side workflow-sync uses inspectSmithersAttemptAgentSelection, and workflows rendered before this change import both. Per-node cloud execution is gone (#1197) and rendered task specs no longer carry `execution`, so every producer attempt writes the record and the verifier has no single-rung cloud shortcut. The history and integration fixtures carry no `execution` either, so a leftover `task.execution` read in the template fails both. This reverses #585's "a non-Codex fallback cannot forge final-report producer authority through the run filesystem" test. That property was not a real boundary: an unsandboxed agent running as the same user can edit the Smithers database as easily as a run-directory file, and trusted-cli.ts already states the host is not a same-UID sandbox boundary. Both the database and the new record sit outside the agent's worktree and declared artifact directories. Tests: the real-Smithers restart test now runs the detached engine with a PATH that resolves no `smithers` (and calls pause by absolute path); with the previous template both variants fail with ENOENT, with this change both pass. report-retry-history.test.ts is rewritten to drive the extracted helpers across simulated restarts against a temporary run directory. The tests that pinned the CLI query (fake execFileSync, read counts, the budget, source regexes) are deleted. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…ead of querying smithers After a controller restart (resume, quota park, supervisor relaunch) the generated workflow's process-local record of which chain rung each report-producer attempt used is gone. The producer's retry prompt and the verifier both rebuilt it by running a bare `smithers node ... --full-output` from inside the workflow: a runner subprocess on the finalizer path that reads up to 64 MiB under a 180 s budget. Wherever the controller PATH resolved no `smithers`, every such attempt failed with "Smithers report-producer authority is unavailable" (cause: ENOENT) until the report node ran out of retries (#1143). The only real-Smithers test of this path gave the detached engine a `smithers` on its PATH. Record each executed selection {attempt, chainIndex} with writeFileDurable in <run>/smithers/final-report-selections/<attempt-id>.json before the agent runs, and read it back when the process-local cache is empty. A re-dispatched attempt number replaces its own entry; finalReportAgentExecution keeps validating attempt order and chain bounds. This deletes readFinalReportSmithersAuthority, priorFinalReportAgentSelections, the 180 s SMITHERS_REPORT_PRODUCER_AUTHORITY_TIMEOUT_MS budget and the workflow's use of the Smithers attempt reconcilers. The runtime exports stay: host-side workflow-sync uses inspectSmithersAttemptAgentSelection, and workflows rendered before this change import both. Per-node cloud execution is gone (#1197) and rendered task specs no longer carry `execution`, so every producer attempt writes the record and the verifier has no single-rung cloud shortcut. The history and integration fixtures carry no `execution` either, so a leftover `task.execution` read in the template fails both. This reverses #585's "a non-Codex fallback cannot forge final-report producer authority through the run filesystem" test. That property was not a real boundary: an unsandboxed agent running as the same user can edit the Smithers database as easily as a run-directory file, and trusted-cli.ts already states the host is not a same-UID sandbox boundary. Both the database and the new record sit outside the agent's worktree and declared artifact directories. Tests: the real-Smithers restart test now runs the detached engine with a PATH that resolves no `smithers` (and calls pause by absolute path); with the previous template both variants fail with ENOENT, with this change both pass. report-retry-history.test.ts is rewritten to drive the extracted helpers across simulated restarts against a temporary run directory. The tests that pinned the CLI query (fake execFileSync, read counts, the budget, source regexes) are deleted. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 30, 2026
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…loud run Problem: before #1197 removed per-node cloud execution, `ultrafuzz clean` of a run planned with `[execution] mode = "cloud"` terminated its sandboxes (tagged purpose=ultrafuzz-node) and deleted its Modal volume before deleting the run directory. #1197 dropped that step and kept the deletion, so cleaning such a run now leaves its Modal volume and any running sandbox in place, still billed, and deletes the plan.json that records the Modal app they live in, without saying so. Change: before anything is removed, clean reads the mode and Modal app from each selected run's plan.json, which copies the execution block of the run's resolved config and is what the removed cleanup read. For a cloud run it returns a CLEAN_CLOUD_STORAGE_RETAINED warning naming the app, the volume and the sandboxes' run tag, and the `modal volume delete` command. `--dry-run` reports it without deleting, and every later result, including a failed removal, carries it. The deletion itself is unchanged. The volume is not `ultrafuzz-node-<run-id>`: the removed provider named it `ultrafuzz-node-` plus a bounded identity of the Smithers run ID `ultrafuzz-<run-id>` (at most 32 normalized characters, then 12 hex digits of its SHA-256). The same identity was the sandboxes' `run` tag. clean reproduces that function, and the test pins the name the removed code computes for a fixed run ID. An unreadable plan.json yields no warning. It also makes the source revision step fail before anything is removed, as it did before, so no run that clean deletes skips the check. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
Problem: #1197 made compilation write every task's execution block as `{ mode: "local", resources, agentCredentialEnv: [] }` with no `modal` entry, but synchronization still added task.execution.agentCredentialEnv and task.execution.modal?.credentialEnv to its exact-value redaction list. For every run compiled since #1197 both are empty. The generated agent environment also still blanked ULTRAFUZZ_MODAL_MODULE, which only the removed cloud provider set. And the doctor tests still wrote a fake engine into the target's .smithers, although since #1201 doctor inspects only the runner Ultrafuzz's own install provides. Change: synchronization derives its redaction values from the environment alone, the agent environment no longer lists ULTRAFUZZ_MODAL_MODULE, and seven doctor tests lose the dead fake-engine setup. One test used it only to create the .smithers/node_modules/.bin directory it writes into, and now creates that directory itself. The doctor test that installs a 0.29.0 project-local engine keeps it: it asserts that doctor ignores that engine. The task manifest keeps agentCredentialEnv, because its sealed schema requires the field. Effect on runs compiled before #1197 in cloud mode, whose lists were not empty: sensitiveEnvironmentValues still redacts, by name, every variable that looks like a credential (*_API_KEY, *_TOKEN, *_SECRET, ...), which includes MODAL_TOKEN_ID and MODAL_TOKEN_SECRET. What these lists added beyond that were route variables such as *_BASE_URL and allowlisted variables. Those carry no secret, or have a value the secret patterns already match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
… about old cloud runs' Modal storage in clean, and drop the smithers shim (#1235) ### 1. Forge guard dropped from the engine PATH, while run.json said it was active (bug on `main`) Under a umask that leaves new directories group writable, such as Ubuntu's default `0002`, `ensureSafeDirectory` created `<run>/safe-bin` with mode `0775`. `composeSmithersCommandPath` admits a PATH entry inside the target only through `isPreparedForgeGuardBin`, which refuses a directory where `mode & 0o022` is set. The workflow engine's PATH therefore dropped the wrapper, and agents ran the real `forge` without `forge_vmem_limit_kb` or `forge_rayon_threads`, while `run.json` recorded `forge_guard.active: true`. - To reproduce on `main`, run it under this host's umask `0002`. `startRun injects the configured Forge guard into the workflow environment and metadata` fails because the fake engine resolves `/tmp/ufz-runtime-…-fake-bin/forge` instead of `<run>/safe-bin/forge`. - A second cause, found while fixing the first: with any custom `run.output_dir`, `run.json` makes the same false claim under umask `022` too, because `isPreparedForgeGuardBin` only admits the wrapper from `<project>/.ultrafuzz/runs/<run-id>/safe-bin`. A probe on `main` printed `.ultrafuzz/runs active: true safe-bin on engine PATH: true` and `audit-runs active: true safe-bin on engine PATH: false`. - **The OpenRouter adapter-contract failure was a test-fixture bug, not a product bug.** `generated OpenRouter adapter preserves opaque model IDs…` failed on `main` under `0002` with `provider-home ancestors cannot be group/world writable`. The ancestor the check refused was `<project>/.ultrafuzz`, created by `ultrafuzz init` with the process umask (a probe showed `0775`). The fixture had rooted the operator-owned provider homes under the target. The adapter itself creates every provider-home component with `mkdirSync(…, { mode: 0o700 })`, and the umask cannot widen that mode. In production the default root is `$XDG_STATE_HOME` or `~/.local/state/ultrafuzz/provider-homes`, which is not under the target. ### 2. `ultrafuzz clean` of a run planned for per-node Modal execution (Greptile P1 on #1197) Before #1197, `clean` terminated such a run's sandboxes (tagged `purpose=ultrafuzz-node`) and deleted its Modal volume, then deleted the run directory. Now it deletes only the directory. The volume and any running sandboxes stay in place and may still be billed, and the deletion removes the `plan.json` that names their Modal app. ### 3. Leftovers of #1197 and #1201 Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #134, #921, #712, #713
The owner approved removing per-node cloud execution outright: a config that sets
execution.mode = "cloud",execution.provideror[execution.providers.modal]now fails withCONFIG_EXECUTION_CLOUD_REMOVED, and the Modal eval runner, which runs a whole campaign in one sandbox, stays.Problem
Per-node cloud execution (
[execution] mode = "cloud": one Modal sandbox per agentic attempt, #134) cannot complete a node onmainor in any release. Two defects in the node worker stop every cloud node before inference:realpath(data root)to equal the lexical path (assertDurableDirectory, node-worker.ts). The Modal filesystem-v2 mount alias makes cloud-node durable checkpoint parents fail safety checks #712 live probe shows Modal's filesystem-v2/datamount is a provider symlink (realpath_equal=false). I reproduced the check with the base's own node-provider test fixture.initializeDurableNodeWorkspaceinitializes under a plain directory. Under a symlinked mount it throwsdurable data root is unsafe./proc/<pid>/fd/<n>/…descriptor path asULTRAFUZZ_WORKFLOW_PERSISTED_PATH. The generated workflow passes that path throughcloudSnapshotRelativePathat module load. I ran the base template's own function against a descriptor path to the same file. The canonical path is accepted. The descriptor path throwspersisted workflow path must stay inside the cloud handoff project, so the worker cannot load the workflow.Tags
v0.1.0,v0.1.1andv0.1.2are ancestors ofmain, and each still has the worker'srealpathcheck, its descriptor-path export and the template'scloudSnapshotRelativePathcheck;node-worker.tsandnode-provider.tsare byte-identical fromv0.1.0throughmain.originhas norelease/*branch. The prior analysis says the fixes (#715, #735, #751, #753–756, #761) merged only into the now-deletedrelease/v0.1.0; I did not re-check each of those PRs.The feature is also the largest reason the runtime ships relocatable whole-tree snapshots (#921). It adds a hidden runtime-to-modal import cycle: variable-name dynamic imports of a package the runtime does not declare. And it keeps a 4,448-line provider plus a 1,794-line worker under test on every PR.
Root cause
Both defects sit in the controller/worker handoff layer: the tarred sealed snapshot, the durable FS-v2 checkpoints, and a cloud-worker re-run of the generated workflow. That layer exists only for per-node cloud execution. Making it work again would mean re-landing the deleted release-branch fixes, then the further failures filed as #718, #740, #741, #743 and #745–747, then validating live on Modal. Removal is the smaller, root-level change.
Change
Fourteen commits. The first three do the removal, in dependency order. The next three follow up on review, the seventh records the change in the changelog, commits 8–13 address the promotion review, and commit 14 addresses its Greptile review:
refactor(config)!:validateExecutionConfigreportsCONFIG_EXECUTION_CLOUD_REMOVEDforexecution.mode = "cloud",execution.providerand[execution.providers.modal], one diagnostic per path.validate,doctor,runandeval runfail before a run directory exists. Each message names its setting (commit 11), then says: per-node cloud execution was removed; set [execution] mode = "local" (or delete the table) and remove execution.provider and [execution.providers.*]. Every agentic attempt now runs locally; to use Modal, run the whole campaign inside one sandbox (the ultrafuzz-modal eval runner does this for benchmark rows, see docs/how-to/run-evals-on-modal.md). The gate sits in resolve, not in the TOML parser, because artifact gates re-parse each run's persistedconfig.resolved.tomlmid-run. The rest of[execution]stays accepted and validated but no longer affects execution:ultrafuzz initscaffolds it and every persisted run config carries it. Docs: deletesdocs/reference/cloud-execution.mdand removes cloud clauses from the configuration, CLI, topology and backends pages.refactor(runtime)!: deletes cloud dispatch.compileTasknow always emits{mode: "local", resources, agentCredentialEnv: []}. That is the value it already produced for local runs, sotasks.jsonand the artifact schemas are unchanged.<Sandbox>branch, the top-level dynamic import of@ultrafuzz/modaland thereactimport (−1,078/+33 lines).doctorandrunprobe required commands on the localPATHonly, so no Modal app is created.cleanno longer calls Modal. Data governance, validate, retry-chain and start-run lose their cloud branches.artifactAwareAgent, the scan of a verifier-rejected final report (unverified-report-inputs.ts, added by fix: a reworded finding no longer discards the final report #1204) andstartRun's launch-failure redaction. Each was empty for local tasks onmain:cloudAgentCredentialEnvreturned[]outside cloud, and local mode with[execution.providers.modal]failedCONFIG_EXECUTION_LOCAL_PROVIDER_SETTINGS. The first three now usesensitiveEnvironmentValues(process.env)alone; the launch-failure redaction keeps its agent API-key names and dropsproviders.modal.credentialEnv.currentReportAttemptstops returning the task'sexecutionblock.cloud-execution-generationcontract and schema are deleted, so--reset-nodestops writing that file.@ultrafuzz/modaloutside two code comments. The five runtime release-validation lanes dropbuild_modal_dependencies, andpackage-gatesstill builds modal. That flag was also how those lanes built@ultrafuzz/artifacts; commit 4 restores that build.@ultrafuzz/modal, which is still needed forpnpm exec ultrafuzz-modal, andreactleaves the runtime ignore list. Both edits are required: on the rebased tree, reverting the first fails knip withUnused devDependencies: @ultrafuzz/modal, and restoringreactfails it withreact … Remove from ignoreDependencies.refactor(modal)!:node-provider.ts,node-worker.tsand the archive helpers (modal-download.ts,deterministic-archive.ts,safe-archive.ts).modal-commondefinitions, with their contracts, registry entries and gates. The Modal registry goes from 21 to 14 schemas.cloud-selected-task.tsDTO and the configMODAL_*constants.patches/[email protected]. It only rewroteSandboxFilesystem.copyToLocal, and onlymodal-download.tscalled that, so[email protected]stays, unpatched.packages/modaldropstar-stream(packages/clikeeps its own), and the two relatedscripts/ciBun tests go.ultrafuzz-modaleval runner (one sandbox per eval row, running a whole local campaign) is unchanged and imports none of the deleted modules.ci: every release-validation lane builds@ultrafuzz/artifacts....scripts/validate-release.mjsimportspackages/artifacts/distwhen it loads. The runtime lanes got that build only throughbuild_modal_dependencies. After commit 2,node scripts/validate-release.mjs --gates runtime-1on a fresh checkout failed withERR_MODULE_NOT_FOUNDbefore it selected a gate, so all five runtime lanes would have failed on every PR.build_release_reporterlane flag is deleted.package-gatesnow builds the artifacts closure twice.validate-release.mjsimports frompackages/*/distis not built by an unconditional step beforevalidate:release.refactor: deletes the cloud-only code that commits 1–3 left without a production caller.verifyCommittedControllerGenerationAuthorityand itsCommittedControllerGenerationAuthorityresult; artifactsreferenceArtifactManifestAuthorityForArtifactDir. refactor(runtime): delete dead controller-generation, sealed-refresh, re-finalization and recovery-authority code #1193 deleted every other export ofworkflow-controller-generation.tsand kept the verifier only because the node provider called it. Its journal and manifest readers served only the verifier, so the whole module is deleted and the runtime index stops re-exporting it. No generated-workflow template in v0.1.0, v0.1.1, v0.1.2,mainorintegration/wave1references any of these names. The artifacts test that called its export now checks the same result through live code: the parsed manifest keeps its reference authority. The verifier's only test was the dynamic controller-refresh test, which refactor(runtime): delete dead controller-generation, sealed-refresh, re-finalization and recovery-authority code #1193 deleted along with the refresh code.dynamicExecutionPath(task, value, label)readtaskonly for the deleted cloud branch, and had become a copy ofcurrentProjectPath(value, label). It is deleted: its 12 call sites callcurrentProjectPath, and the sevenpath.resolve(process.cwd(), …)wrappers around its already absolute result go.test(config): restores coverage for the[execution]validation that still runs for local projects. The test covers the node-override merge thattasks.jsonrecords,CONFIG_EXECUTION_NODE_UNKNOWNfromvalidateandrun, the[execution.nodes.*.resources]serialization round trip and the rejection ofcpu = 0.docs(changelog): adds the breaking entry under Unreleased (quoted at the end). Two earlier Unreleased entries lose the clauses that described cloud behaviour this PR deletes: the Checkpoint long fan-in stages and recover according to failure type #1150 review-timeout entry no longer tells cloud users to add per-node timeout overrides to avoidCLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED(an error this PR deletes), and the Codexopenai_base_urlroute entry (fix(runtime): Codex route checks follow openai_base_url the way Codex does, and -smithers paths stay intact in diagnostics #1210) no longer says cloud planning rejects that config.fix(runtime):resumerefuses a run planned for cloud execution. Such a run'stasks.jsonrecordsexecution.mode = "cloud", and itsresolved-config.jsonstill parses because the JSON schema accepts the value; the removal gate lives only inresolveConfig.resume --refresh-controllertherefore rendered the current controller, which has no Sandbox branch, and ran the remaining attempts as local worktree tasks on the operator's host, with no diagnostic.submitSmithersContinuationnow returnsWORKFLOW_CLOUD_EXECUTION_REMOVEDbefore it inspects, renders or starts anything when any task in the manifest it already parses is a cloud task, for native and refreshed resume alike. (Commit 14 takes the mode from the run's resolved config instead, because native resume treats the task manifest as optional.) Replay and fork are unchanged: they run the historical sealed workflow, not the current controller. The changelog entry and the configuration reference say so, and the changelog's fix instructions are tightened (below).test(runtime): restores a discriminating test for the verifier's pre-publication scan of environment values. The rebase had deleted test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203's test and the harness'sscanProcessEnvironmentoption, the only path that passed the realsensitiveEnvironmentValuesintoverifyArtifacts. The option is back, and the new test writes a value with no vendor format intoresult.json: underULTRAFUZZ_TEST_GATEWAY_LABELalone it is published, and onceULTRAFUZZ_TEST_GATEWAY_API_KEYalso holds it the verifier throwscontains sensitive dataand publishes nothing. Two agent-failure redaction tests that no longer configure credential names are renamed to say they redact credential-named environment values.fix(modal):MODAL_CAMPAIGN_ENVELOPE_TOO_SHORTand the Modal eval how-to stop recommending per-node Modal execution. Both now say to run the four-hour profile locally until the full benchmark's envelopes are enlarged, and the how-to drops the deleted 1,800-second lifecycle reserve and explicit-cap sentences.CONFIG_EXECUTION_CLOUD_REMOVEDlinks to that page.fix(config): eachCONFIG_EXECUTION_CLOUD_REMOVEDmessage starts with its setting (execution.mode = "cloud",execution.provideror[execution.providers.modal]). Plain-textvalidateanddoctorprint messages without paths, so a config with all three printed the same ~330-character line three times.fix(runtime):doctorprobes the local toolchain even when the config does not resolve. The skip existed because the probe once needed the config to choose between the localPATHand a Modal sandbox; after commit 2 it only reported git, node and forge missing next to every upgraded cloud user'sCONFIG_EXECUTION_CLOUD_REMOVED.refactor: deletes leftovers commit 5 missed because neither TypeScript nor knip flags them:compileTask'senvinput (its only reader was the deletedcloudAgentCredentialEnv;startRunstill built and passed it), theworkflowModuleEntryUrls()wrapper and its empty-string filter (it existed for the conditionalmodal: ""entry), the fake runners'SMITHERS_FAKE_CLOUD_ENV_LOGhooks that printedMODAL_TOKEN_ID/MODAL_TOKEN_SECRETand their two allowlist entries, the "two cloud containers" rationale in three Effect-pinning comments (thepnpm-workspace.yamlcopy is left to feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201, which rewrites that block), and one stateful-profile test name.fix(runtime):resumedecides whether a run was planned for cloud execution from itssmithers/resolved-config.json, not fromsmithers/tasks.json(Greptile P1). Native resume treats the task manifest as optional, so with a missing or malformed manifest commit 8's guard was skipped and a pre-removal cloud run's historical workflow went to Smithers. The pre-removal compiler gave every taskconfig.execution.mode, so the config is the record of that mode.submitSmithersContinuationreads the config before the manifest and returnsWORKFLOW_CLOUD_EXECUTION_REMOVEDwhen it says cloud.WORKFLOW_LIFECYCLE_FAILED(run <id> cannot be resumed without its resolved config, which records whether it was planned for the removed per-node cloud execution: …). Before, a plain resume continued such a run with default concurrency and lease and noULTRAFUZZ_CONFIG_PATH;--refresh-controlleralready failed.ULTRAFUZZ_CONFIG_PATHbranches, theconfig?.defaults for concurrency, workspaces and lease, andrecordNativeContinuationState's legacy-deadline branch.test/local-resolved-config.tshelper.ULTRAFUZZ_CONFIG_PATHfor a run whose config does not parse) and the changelog entry describe the refusal.Size: 107 files, +979/−21,513 (net −20,534):
About 180 cloud-only test declarations are deleted: 163 in 8 deleted test files (including the 119-declaration node-provider suite and the 14-declaration cloud-worker handoff suite) and 17 in edited suites. Ten tests are added (four by commits 8, 9, 12 and 14), and eight are renamed to drop "cloud" or configured-credential wording from their names.
Deliberately not built
[execution]data stays. That coversResolvedConfig.executionwith its loader, defaults, serializer and zod/JSON schema (includingproviders.modalparsing andresourceTimeoutOrigin), theexecutionblocks inplan.jsonandtasks.json, validate'sexecution_mode, and the CLI result'sexecution_mode/execution_providerenum. Removing it changes the artifact and config schema bundle digests, which strands in-flight runs on resume. No issue tracks it yet (Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921, the sealed-snapshot cost/benefit issue, does not list these items). The owner's direction on fix(runtime): a renderer or projection change no longer strands runs whose dynamic groups expanded #1216, to rethink tamper detection and possibly drop the sealing, would remove that digest reason, so it belongs in one follow-up issue with the items below.workflow-sync.tskeeps itsexecution.mode === "local"filters and its scan of each task's recorded credential names. Every newly compiled task passes the filters and records no names, and feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 edits that file.environment.tsxstill listsULTRAFUZZ_MODAL_MODULE. The adapters are digest-enforced, so changing them for one unused name would make every project rerunultrafuzz initbefore its next run.hydrateTaskSpec'spath.resolve(process.cwd(), …)wrappers, now no-ops because the renderer emits absolute paths (source-regex tests pin them, and fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 edits the template); andexecution: { mode: "local" }fields in test fixtures of functions that no longer read them (runtime.test.ts,smithers-report-retry.integration.test.ts; the two ingenerated-workflow-verifier.test.tssit in tests fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 deletes).CONFIG_EXECUTION_NODE_UNKNOWNstays an error. An[execution.nodes.<id>]override for a node the topology lacks still blocksvalidateandrun, as onmain, although the override no longer changes execution. Downgrading it to a warning is a separate behaviour change for the follow-up; commit 6's test pins the current behaviour.Verification
Everything below ran on the rebased tree,
origin/main(a46a496) plus the first seven commits (head57daecfa), unless it says otherwise. What ran for commits 8–13 is under Review follow-ups, and what ran for the stacked rebase ontoorigin/main538b6188and commit 14 is under Stacked rebase and commit 14.Sealed identities. I printed them from the built trees of this head and of
origin/main(a46a496):VALIDATOR_BUILD_IDENTITY(…929e741c…), the artifact schema bundle (6a05a7a6…) and the config schema bundle (b799a29b…) are byte-identical toorigin/main. (VALIDATOR_BUILD_IDENTITYwas…028be325…on the previous base;mainchanged it, not this branch.)c61b8387…→17b5f648…, 19 → 18 schemas) and the Modal bundle (9a626ddb…→6f0fd0af…, 21 → 14) change. Both are computed only at validation time, in CLIjson validateand Modal document parsing. No persisted document records either one: the trusted-CLI identity binds the artifact bundle, which is unchanged.VALIDATOR_BUILD_IDENTITYhashes, and538b6188(fix(security): merge per-range registry advisories and patch brace-expansion and undici #1229) touches nothing underpackages/artifacts,packages/config,packages/clior a schema directory, so all five values hold at the final head (b8a2bc5b).Discriminating tests. I ran each of these against
origin/main(a46a496) by copying the test file into amainworktree, and each fails there as described:rejects removed per-node cloud execution and still loads the local [execution] table init scaffoldsand the named-diagnostic caseremoved cloud executionboth fail. The first also round-trips the exact[execution]tableultrafuzz initwrites.planRun rejects removed cloud execution before creating the run directoryfails:mainreturns noCONFIG_EXECUTION_CLOUD_REMOVEDdiagnostics (actual: []).cleanRun removes a run planned for removed cloud execution without provider credentialsfails withCLEAN_CLOUD_STORAGE_FAILED, becausemain'scleanimports@ultrafuzz/modaland needs Modal credentials.build every package validate-release.mjs imports before any lane runs itfails (Received: "matrix.build_release_reporter == true").Replacement coverage (not discriminating tests):
generated-workflow-render.test.tstakes over from the deleted cloud-worker harness, the only thing that rendered the whole generated program in process. It compiles a two-node local workflow, imports it with Smithers components stubbed, and renders it against the built runtime and artifacts. It checks the Worktree prepare/agent/verify tree, the dependency edges and the prompt body. Before this rebase I checked it with two mutations: it fails when a module-scope call to a deleted helper is added back (ReferenceError: readCloudExecutionGeneration is not defined), and when a verifier'sdependsOnits agent is dropped. The test file and the template are byte-identical to that pre-rebase head, andmainhas not changed the template since the previous base.merges local [execution] node resource overrides, rejects unknown nodes and bad bounds, and round-trips themrestores the coverage that the deleted cloud tests gave the live local path. It also passes onmain. Before this rebase, swapping the override precedence inresolveExecutionResourcesfailed it, and so did inverting its unknown-node filter.Resume compatibility:
main's template imports 7 names that commits 1–3 retire (6 from artifacts,CLOUD_EXECUTION_GENERATION_JSON_SCHEMA_IDfrom runtime). They are destructured at module scope and used only inside cloud-only functions; the one module-scope caller,readCloudExecutionGeneration(), returns"base"before touching its schema ID when no task is a cloud task.main's template is unchanged since the previous base, so the earlier AST scan still applies. The names commit 5 deletes appear in no template of v0.1.0, v0.1.1, v0.1.2,mainorintegration/wave1.codex. This is a one-off harness adapted from thecli-e2ecampaign-resume test and is not committed.origin/main's CLI (a46a496) runsinitandrunon the test's three-node campaign; the engine is SIGKILLed so the supervisor relaunches it, then the whole controller is SIGKILLed mid-node. This branch's CLI runs every later command (status,resume,events,report,stats). I ran it twice on the rebased tree, once with a plainresumeand once withresume --refresh-controller, and both pass every assertion of the e2e test: the resumed run endsRunFinishedwith a verified report, only the interrupted node runs again, no finished task starts again, andstatsrecords the interrupted attempt ascanceledand counts the same attempts asstatus.Suites on the rebased tree:
runtime.test.tstests, listed belowjson-validateinit and validate emit schema-versioned launch JSON, therun … clean …lifecycle test anddoctor reports install posturecli-e2ecampaign resumeorigin/main's CLI, plain and--refresh-controllerpackaged-topologiesworkflow.tsx(dynamic-lifecycle, invariant-suite-ancestor-order, invariant-suite-enumeration-overflow, invariant-suite-handoff-durability, report-retry-history, smithers-preparation-race, smithers-report-retry, stale-workspace-cleanup-overflow, terminal-report-projection, workspace-patch-replay, workspace-patch-supersede, workspace-preparation-lifecycle), and dynamic-workflow, generated-workflow-footprint, workflow-task-metrics and trusted-cli.runtime.test.tsselection: every test whose code this PR edits (including the four that share the edited invariant-budget fixture), every test that fix: a reworded finding no longer discards the final report #1204–fix(cli): recovery hints name ultrafuzz commands, not raw runner invocations #1207 added or changed there (among them the threesyncRun distinguishes available agent reports after stopped failuresvariants; the failed-verifier and changed-output variants read a verifier-rejected report through the simplified secret scan), all controller-refresh, native-continuation, current-controller-rendering and reset-node tests, and the command-probe and credential-forwarding tests of the start-run and required-commands code this PR edits.eatmydata, as CI does. Without it,ultrafuzz runfsyncs every file of the sealed execution snapshot, and my first pre-rebasecli-e2eand cross-version runs timed out.Gates: all pass on the rebased tree.
pnpm install --frozen-lockfile(a real install: this PR changes the lockfile and removes a patch) andpnpm -w build, which also passes at each of the seven commitspnpm -w lint, andCI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:cipnpm -w knip, both on a fresh unbuilt worktree and afterpnpm -w buildnpx prettier --checkon the 71 added or modified files,pnpm -w format:check, andpnpm -w docs:check(which runsscripts/docs-check.mjs)pnpm --filtertypecheck of artifacts, cli, config, modal, runtime and topology, andpnpm -w typecheck(all 12 workspace projects)The complexity ceiling stays at 83. ESLint's per-function complexity in every source file this PR edits (the template included) is never higher than on
origin/main; the maximum in those files is 76.Review follow-ups (commits 8–13). Run on commit 13's head (
533dbfdf, ona46a4960), built withpnpm -w build:resume refuses a run planned for removed per-node cloud execution before invoking Smitherslaunches a local run, rewrites itstasks.jsonto what the pre-removal compiler wrote for a cloud config, and checks that plain and--refresh-controllerresume fail with onlyWORKFLOW_CLOUD_EXECUTION_REMOVEDand issue no Smithers command. With the guard removed, the plain resume succeeds and the test fails (actual: true).generated Smithers verifier refuses to publish the value of a credential-named environment variablefails when the template'sassertArtifactPublicationsContainNoSecrets(publications, sensitiveEnvironmentValues(process.env))becomes(publications, []); before commit 9, all 115 verifier tests passed under that mutation.diagnoseProject still probes the local toolchain when the project config does not resolvefails with doctor's oldresolved.config === undefined ? [] :skip restored.config.test.tsandstateful-profiles.test.ts52/52; modalconfig.test.ts23/23 andrunner.test.ts79/79; runtimegenerated-workflow-verifierandunverified-report137/137;lifecycle-inspection,data-governance,clean,trusted-clianddynamic-lifecycle107/107; thesmithers-terminal-resume,smithers-preparation-raceandsmithers-resume-reopenreal-engine suites 8/8; the sevencompileSmithersWorkflowcaller files (generated-workflow-render,task-workflow-identity,generated-workflow-footprint,source-revision,pinned-submodules,workflow-dependency-policy,dynamic-workflow) 21/21; 22 selectedruntime.test.tstests (the new resume test, both controller-refresh tests, the 7 native-continuation tests, the 5startRun forwards …tests,planRun rejects removed cloud execution, the command-probe test, the 4 tests that use the edited fake npm installer, and the runner-pin test whose comment commit 13 rewords), all passing, with 2 Bun-only variants skipped under node; CLIdoctor reports install posture in human and JSON outputandinit and validate emit schema-versioned launch JSON2/2; Bunworkspace-engine-overrides3/3.ultrafuzz.tomlsets all three removed settings:validateanddoctorprint threeCONFIG_EXECUTION_CLOUD_REMOVEDlines whose messages each start with their setting, anddoctornow lists git, node and forge with their paths and versions, reporting only the topology's genuinely absent commands missing.pnpm -w lint;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;pnpm -w knipafterpnpm -w buildand on a fresh unbuilt worktree (pnpm install --frozen-lockfile --offline);npx prettier --checkon the 77 added or modified files;pnpm -w format:check;node scripts/docs-check.mjs;pnpm --filtertypecheck of config, modal and runtime; and the runtime and CLI test builds (tsc -p tsconfig.test.json).submitSmithersContinuationstays at 51 (51 onmain) with the guard, anddiagnoseProjectdrops to 39 (40 onmain).42a1bc85), fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 (594d6b87) and feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 (b3808c24) onto the old head and commit 13's head conflict in the same files with the same number of conflict hunks, so commits 8–13 add no merge work.Stacked rebase and commit 14. Run on the final head (
b8a2bc5b, onorigin/main538b6188) afterpnpm install --frozen-lockfileandpnpm -w build:start-run.tsbuilt into the same tree:resume refuses a run planned for removed per-node cloud execution before invoking Smithersnow rewrites the resolved config as well as the task manifest, as a pre-removal cloud launch wrote both, and checks both resume forms again with the manifest malformed and then deleted. Against commit 13 it fails atmalformed task manifest, refreshController=false(actual: true: the plain resume reached Smithers).resume refuses a run whose resolved config cannot be read instead of continuing it locallydamages a local run's config (malformed, then deleted), checks that both resume forms fail withWORKFLOW_LIFECYCLE_FAILEDand the config message before any Smithers command, then resumes the run once the config is restored. Against commit 13 it fails atmalformed config, refreshController=false(actual: true).runtime.test.ts: every test that callsresumeRun, the two above included (40/40), and the 32 tests whose code this PR edits or that use its edited fixtures (32/32).dynamic-lifecycle,lifecycle-inspectionand the real-enginesmithers-terminal-resume,smithers-preparation-raceandsmithers-resume-reopensuites: 81/81. The two hand-built continuation fixtures first failed with the new refusal; that is how they were found.config.test.tsandrunner.test.ts102/102.json-validate16/16, plusinit and validate emit schema-versioned launch JSON, therun, ps, status, … lifecycle commands …test (which resumes),resume of an already-active run …anddoctor reports install posture in human and JSON output: 4/4.release-validation-lanes,workspace-engine-overridesanddependency-advisories: 28/28.pnpm -w format:check;pnpm -w lint;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;pnpm -w knipafterpnpm -w buildand on a fresh unbuilt worktree (pnpm install --frozen-lockfile --offline);pnpm --filter @ultrafuzz/runtime typecheck;node scripts/docs-check.mjs; andnode scripts/check-production-dependency-advisories.mjs, because the lockfile was the conflicted file (0 High/Critical production advisories, 0 active exceptions).submitSmithersContinuationdrops from 51 to 35 andrecordNativeContinuationStatefrom 12 to 5.Not run:
validate:packand super-linterRisk / compatibility
ultrafuzz-modaldoes for eval rows. No release ever completed a cloud node, and the feature allowed only one attempt with no subscription auth, so no working capability is lost in practice. The plan is what goes. With this merged, Run each node on the cloud #134 would be closed as not planned.mode = "cloud",execution.provideror[execution.providers.modal]failsvalidate,doctor,runandeval runwithCONFIG_EXECUTION_CLOUD_REMOVEDuntil those keys are removed. While such a config fails to resolve, the commands that only locate runs (status,resume,events, …) fall back to.ultrafuzz/runs, as they already do for any invalid config. A project that also sets a customrun.output_dirtherefore has to fix its config first.ultrafuzz cleanno longer terminates sandboxes taggedpurpose=ultrafuzz-nodeor deletes theultrafuzz-node-*volume of an earlier cloud attempt. Operators must remove those by hand in the Modal app the run used; the run'splan.jsonrecords it underexecution.providers.modaluntilultrafuzz cleandeletes the run.resume, plain or with--refresh-controller, refuses one withWORKFLOW_CLOUD_EXECUTION_REMOVED(commits 8 and 14) instead of re-rendering it as local tasks on the operator's machine. The mode comes from the run's resolved config, so a missing or malformed task manifest does not let such a run through.replayandforkrun the historical sealed workflow and are unchanged.pause,canceland the read-only commands are not affected by the guard.smithers/resolved-config.jsonis missing or does not parse can no longer be resumed: it fails withWORKFLOW_LIFECYCLE_FAILEDbefore Smithers starts. A plain resume used to continue it with default concurrency, lease and deadline handling. Every launch since v0.1.0 writes that file. v0.1.2's resolved-config schema is byte-identical to this branch's, and v0.1.1's differs only by a lowerworkflowDeadlineSecondsmaximum, so their runs parse. What is affected is a run whose file was deleted or damaged, and possibly a v0.1.0 run; I did not check whether v0.1.0's configs still parse.dist. The names retired by commits 1–3 appear only in cloud-only functions, commit 5 deletes no export that an old template uses, and the cross-version runs under Verification resumed a run thatorigin/main's CLI created.resume --refresh-controller. It renders the new template, whose input schema drops the 5 nullable cloud-worker columns and whose task specs no longer carryexecution. The engine's input load selects only declared columns, and the cross-version refresh run under Verification continued a pre-upgrade run to completion.ULTRAFUZZ_MODAL_MODULEis no longer forwarded to the workflow engine or listed as controller-only in the runtime. Nothing reads it for local runs, and the stock adapter still strips it from agent environments.startRun's launch-failure redaction no longer add configured cloud credential names to the values they refuse. Each of those lists was empty for local tasks onmain, so results for local runs are unchanged; commit 9 restores a test that the pre-publication scan still refuses credential-named environment values.git merge-tree) of each PR's current head onto this one (b8a2bc5b) conflicts in: feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 (a811973d) —start-run.ts,lifecycle-inspection.test.ts,pnpm-lock.yamlandpnpm-workspace.yaml(this PR removes the Modal patch next to the Smithers patches feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 adds); fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 (3a1d3430) —CHANGELOG.md,workflow.tsxandgenerated-workflow-verifier.test.ts; feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 (eb783f17) —CHANGELOG.mdonly. Commit 14 adds only thestart-run.tsconflict, one hunk in the continuation'srunSmithersLifecycleCommandcall: keep this PR'skeepWorkspaces: config.run.keepWorkspacesandcontrollerLeaseSeconds: config.run.controllerLeaseSecondswith feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201'senv: lifecycleEnvironment. None of the three adds a test that resumes a hand-built run directory, which would now needwriteLocalResolvedConfig. fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 also needs a semantic fix when it is rebased: its new final-report selection code readstask.execution.mode !== "cloud"inworkflow.tsx, and after this PR the rendered task specs no longer carryexecution, so that expression would throw; every task is local now, so the branch should go. Source-slice tests indexworkflow.tsxby text markers; the markers and occurrence counts that pointed into or counted deleted template code were updated here.Rebase notes
Stacked rebase onto
538b6188(base of the #1197 → #1201 → #1183 → #1198 stack). The 13 commits moved froma46a4960ontoorigin/main538b6188, which adds one commit (#1229). One conflict:pnpm-lock.yaml(content, in commit 3, with fix(security): merge per-range registry advisories and patch brace-expansion and undici #1229): fix(security): merge per-range registry advisories and patch brace-expansion and undici #1229 changed the[email protected]patch hash (beeb1f98…→cb8c7558…) on the line next to the[email protected]patch entry this PR deletes. I kept fix(security): merge per-range registry advisories and patch brace-expansion and undici #1229's npm hash and dropped the Modal entry; the rest of commit 3's lockfile change (unpatched[email protected], notar-stream) applied cleanly.pnpm install --frozen-lockfileaccepts the result, and the production dependency advisory gate passes on it.CHANGELOG.mdmerged cleanly: fix(security): merge per-range registry advisories and patch brace-expansion and undici #1229's entry is under Other changes, this PR's under Breaking changes.git range-diffshows the other twelve commits unchanged; commit 3 differs only in that lockfile context. Commit 14 is new.Greptile (promotion review):
packages/runtime/src/start-run.ts:546, "Missing manifest bypasses cloud refusal": fixed in commit 14. The finding is valid: native resume swallowed a missing or malformedtasks.json, which skipped the guard. The mode now comes from the run'ssmithers/resolved-config.json, which the continuation already parsed. A config that cannot be read fails the resume closed instead of continuing the run as local. Both new test branches fail on commit 13's code.Earlier rebase. Rebased from
52cb2506ontoorigin/maina46a4960(19 new commits onmain). Conflicts and how each was resolved:packages/runtime/src/clean.ts(content, with ci: fail on unused exports, types and duplicates #1219): ci: fail on unused exports, types and duplicates #1219 deleted thecleanGeneratedalias next to thereadPersistedModalExecutionhelper this PR deletes. Both are gone; theModalExecutionProviderConfigandisRecordimports go with them.packages/runtime/test/cloud-worker-harness.ts(modify/delete, with ci: fail on unused exports, types and duplicates #1219): took the deletion.packages/runtime/test/generated-workflow-verifier.test.ts(content, with test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203): test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203 deletedgenerated Smithers verifier publishes the complete validated set before task successandgenerated Smithers preparation requires a successful dependency artifact verification, two tests this PR edited; the deletion stands. Git also droppedimport { z } from "zod/v4"cleanly, but test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203's new finalizer test passes that module-scopezinto anew Function, so the import is restored. test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203'sgenerated Smithers verifier refuses to publish the value of any credential the task is configured withcovered the cloud credential lists this PR removes from the template's pre-publication scan, so it is deleted with itsVerifyArtifactsTask.executionfield. The rebase also deleted the harness'sscanProcessEnvironmentoption; commit 9 restores it for a replacement test that covers the scan that still runs.packages/modal/src/deterministic-archive.tsand its test (modify/delete, with fix(modal): a rejected deterministic archive closes its output descriptor once #1215): took the deletion; the node provider was the archive's only production caller. The fix(modal): a rejected deterministic archive closes its output descriptor once #1215 changelog entry is left alone: it records a fix that shipped in v0.1.2.packages/modal/test/current-artifact-fixtures.ts: withnode-provider.test.tsdeleted,currentTaskOutputBindinghas no importer, and knip now fails on unused exports (ci: fail on unused exports, types and duplicates #1219), so it is no longer exported.packages/runtime/src/unverified-report-inputs.ts(no textual conflict, fix: a reworded finding no longer discards the final report #1204): its new rejected-report secret scan read the task manifest'sagentCredentialEnvandmodal.credentialEnv. It now scanssensitiveEnvironmentValues(process.env), the same value as the verifier's pre-publication scan, andcurrentReportAttemptstops returning the execution block. fix: a reworded finding no longer discards the final report #1204'sa report its verifier rejected stays unavailable when it holds a credential its task namestest and its cloud task-manifest fixture are deleted; the remaining secret-gate tests (vendor token, JSON-escaped credentials named like credentials, line-broken and cross-field recovery phrases) still run against the simplified scan, andruntime.test.tscovers the manifest path end to end.CHANGELOG.md: new entry under Unreleased > Breaking changes, plus the two clause removals described in commit 7.Changelog
The entry is in
CHANGELOG.mdunder Unreleased > Breaking changes:🤖 Generated with Claude Code
The PR appears safe to merge based on the issues established in this review.
Summary
The PR removes per-node cloud execution while retaining the whole-campaign Modal eval runner.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Resume request] --> B[Read persisted resolved config] B -->|Missing or invalid| C[Lifecycle failure] B -->|Cloud mode| D[Cloud execution removed] B -->|Local mode| E[Continue Smithers workflow]Reviews (2) · Last reviewed commit: "fix(runtime): decide a resumed run's clo..."