Skip to content

refactor!: remove per-node cloud execution (execution.mode = "cloud") - #1197

Merged
aviggiano merged 14 commits into
mainfrom
claude/v04-remove-cloud-node-execution
Sep 30, 2026
Merged

aviggiano merged 14 commits into
mainfrom
claude/v04-remove-cloud-node-execution

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #134, #921, #712, #713

The owner approved removing per-node cloud execution outright: a config that sets execution.mode = "cloud", execution.provider or [execution.providers.modal] now fails with CONFIG_EXECUTION_CLOUD_REMOVED, and the Modal eval runner, which runs a whole campaign in one sandbox, stays.

Problem

Per-node cloud execution ([execution] mode = "cloud": one Modal sandbox per agentic attempt, #134) cannot complete a node on main or in any release. Two defects in the node worker stop every cloud node before inference:

Tags v0.1.0, v0.1.1 and v0.1.2 are ancestors of main, and each still has the worker's realpath check, its descriptor-path export and the template's cloudSnapshotRelativePath check; node-worker.ts and node-provider.ts are byte-identical from v0.1.0 through main. origin has no release/* branch. The prior analysis says the fixes (#715, #735, #751, #753–756, #761) merged only into the now-deleted release/v0.1.0; I did not re-check each of those PRs.

The feature is also the largest reason the runtime ships relocatable whole-tree snapshots (#921). It adds a hidden runtime-to-modal import cycle: variable-name dynamic imports of a package the runtime does not declare. And it keeps a 4,448-line provider plus a 1,794-line worker under test on every PR.

Root cause

Both defects sit in the controller/worker handoff layer: the tarred sealed snapshot, the durable FS-v2 checkpoints, and a cloud-worker re-run of the generated workflow. That layer exists only for per-node cloud execution. Making it work again would mean re-landing the deleted release-branch fixes, then the further failures filed as #718, #740, #741, #743 and #745–747, then validating live on Modal. Removal is the smaller, root-level change.

Change

Fourteen commits. The first three do the removal, in dependency order. The next three follow up on review, the seventh records the change in the changelog, commits 8–13 address the promotion review, and commit 14 addresses its Greptile review:

  1. refactor(config)!: validateExecutionConfig reports CONFIG_EXECUTION_CLOUD_REMOVED for execution.mode = "cloud", execution.provider and [execution.providers.modal], one diagnostic per path. validate, doctor, run and eval run fail before a run directory exists. Each message names its setting (commit 11), then says: per-node cloud execution was removed; set [execution] mode = "local" (or delete the table) and remove execution.provider and [execution.providers.*]. Every agentic attempt now runs locally; to use Modal, run the whole campaign inside one sandbox (the ultrafuzz-modal eval runner does this for benchmark rows, see docs/how-to/run-evals-on-modal.md). The gate sits in resolve, not in the TOML parser, because artifact gates re-parse each run's persisted config.resolved.toml mid-run. The rest of [execution] stays accepted and validated but no longer affects execution: ultrafuzz init scaffolds it and every persisted run config carries it. Docs: deletes docs/reference/cloud-execution.md and removes cloud clauses from the configuration, CLI, topology and backends pages.
  2. refactor(runtime)!: deletes cloud dispatch.
    • compileTask now always emits {mode: "local", resources, agentCredentialEnv: []}. That is the value it already produced for local runs, so tasks.json and the artifact schemas are unchanged.
    • The generated workflow loses the cloud-worker input envelope, the selected-task handoff, the Modal <Sandbox> branch, the top-level dynamic import of @ultrafuzz/modal and the react import (−1,078/+33 lines).
    • doctor and run probe required commands on the local PATH only, so no Modal app is created. clean no longer calls Modal. Data governance, validate, retry-chain and start-run lose their cloud branches.
    • Secret scans: four secret-value lists no longer add cloud credential names: the generated verifier's pre-publication scan, the agent-failure redaction in artifactAwareAgent, the scan of a verifier-rejected final report (unverified-report-inputs.ts, added by fix: a reworded finding no longer discards the final report #1204) and startRun's launch-failure redaction. Each was empty for local tasks on main: cloudAgentCredentialEnv returned [] outside cloud, and local mode with [execution.providers.modal] failed CONFIG_EXECUTION_LOCAL_PROVIDER_SETTINGS. The first three now use sensitiveEnvironmentValues(process.env) alone; the launch-failure redaction keeps its agent API-key names and drops providers.modal.credentialEnv. currentReportAttempt stops returning the task's execution block.
    • The runtime cloud-execution-generation contract and schema are deleted, so --reset-node stops writing that file.
    • The runtime now has no reference to @ultrafuzz/modal outside two code comments. The five runtime release-validation lanes drop build_modal_dependencies, and package-gates still builds modal. That flag was also how those lanes built @ultrafuzz/artifacts; commit 4 restores that build.
    • knip: the root workspace ignores @ultrafuzz/modal, which is still needed for pnpm exec ultrafuzz-modal, and react leaves the runtime ignore list. Both edits are required: on the rebased tree, reverting the first fails knip with Unused devDependencies: @ultrafuzz/modal, and restoring react fails it with react … Remove from ignoreDependencies.
  3. refactor(modal)!:
    • Deletes node-provider.ts, node-worker.ts and the archive helpers (modal-download.ts, deterministic-archive.ts, safe-archive.ts).
    • Deletes 7 node schemas and 7 modal-common definitions, with their contracts, registry entries and gates. The Modal registry goes from 21 to 14 schemas.
    • Deletes the artifacts cloud-selected-task.ts DTO and the config MODAL_* constants.
    • Deletes patches/[email protected]. It only rewrote SandboxFilesystem.copyToLocal, and only modal-download.ts called that, so [email protected] stays, unpatched.
    • packages/modal drops tar-stream (packages/cli keeps its own), and the two related scripts/ci Bun tests go.
    • The ultrafuzz-modal eval runner (one sandbox per eval row, running a whole local campaign) is unchanged and imports none of the deleted modules.
  4. ci: every release-validation lane builds @ultrafuzz/artifacts....
    • scripts/validate-release.mjs imports packages/artifacts/dist when it loads. The runtime lanes got that build only through build_modal_dependencies. After commit 2, node scripts/validate-release.mjs --gates runtime-1 on a fresh checkout failed with ERR_MODULE_NOT_FOUND before it selected a gate, so all five runtime lanes would have failed on every PR.
    • Every lane runs that script, so the reporter build step is now unconditional and the build_release_reporter lane flag is deleted. package-gates now builds the artifacts closure twice.
    • A new lane-table test fails when a package that validate-release.mjs imports from packages/*/dist is not built by an unconditional step before validate:release.
  5. refactor: deletes the cloud-only code that commits 1–3 left without a production caller.
    • Runtime verifyCommittedControllerGenerationAuthority and its CommittedControllerGenerationAuthority result; artifacts referenceArtifactManifestAuthorityForArtifactDir. refactor(runtime): delete dead controller-generation, sealed-refresh, re-finalization and recovery-authority code #1193 deleted every other export of workflow-controller-generation.ts and kept the verifier only because the node provider called it. Its journal and manifest readers served only the verifier, so the whole module is deleted and the runtime index stops re-exporting it. No generated-workflow template in v0.1.0, v0.1.1, v0.1.2, main or integration/wave1 references any of these names. The artifacts test that called its export now checks the same result through live code: the parsed manifest keeps its reference authority. The verifier's only test was the dynamic controller-refresh test, which refactor(runtime): delete dead controller-generation, sealed-refresh, re-finalization and recovery-authority code #1193 deleted along with the refresh code.
    • In the generated workflow, dynamicExecutionPath(task, value, label) read task only for the deleted cloud branch, and had become a copy of currentProjectPath(value, label). It is deleted: its 12 call sites call currentProjectPath, and the seven path.resolve(process.cwd(), …) wrappers around its already absolute result go.
    • Stale cloud wording is removed from six code comments, from the artifact-contract migration reference (the node schemas were deleted, not retained) and from the harness research page.
  6. test(config): restores coverage for the [execution] validation that still runs for local projects. The test covers the node-override merge that tasks.json records, CONFIG_EXECUTION_NODE_UNKNOWN from validate and run, the [execution.nodes.*.resources] serialization round trip and the rejection of cpu = 0.
  7. docs(changelog): adds the breaking entry under Unreleased (quoted at the end). Two earlier Unreleased entries lose the clauses that described cloud behaviour this PR deletes: the Checkpoint long fan-in stages and recover according to failure type #1150 review-timeout entry no longer tells cloud users to add per-node timeout overrides to avoid CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED (an error this PR deletes), and the Codex openai_base_url route entry (fix(runtime): Codex route checks follow openai_base_url the way Codex does, and -smithers paths stay intact in diagnostics #1210) no longer says cloud planning rejects that config.
  8. fix(runtime): resume refuses a run planned for cloud execution. Such a run's tasks.json records execution.mode = "cloud", and its resolved-config.json still parses because the JSON schema accepts the value; the removal gate lives only in resolveConfig. resume --refresh-controller therefore rendered the current controller, which has no Sandbox branch, and ran the remaining attempts as local worktree tasks on the operator's host, with no diagnostic. submitSmithersContinuation now returns WORKFLOW_CLOUD_EXECUTION_REMOVED before it inspects, renders or starts anything when any task in the manifest it already parses is a cloud task, for native and refreshed resume alike. (Commit 14 takes the mode from the run's resolved config instead, because native resume treats the task manifest as optional.) Replay and fork are unchanged: they run the historical sealed workflow, not the current controller. The changelog entry and the configuration reference say so, and the changelog's fix instructions are tightened (below).
  9. test(runtime): restores a discriminating test for the verifier's pre-publication scan of environment values. The rebase had deleted test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203's test and the harness's scanProcessEnvironment option, the only path that passed the real sensitiveEnvironmentValues into verifyArtifacts. The option is back, and the new test writes a value with no vendor format into result.json: under ULTRAFUZZ_TEST_GATEWAY_LABEL alone it is published, and once ULTRAFUZZ_TEST_GATEWAY_API_KEY also holds it the verifier throws contains sensitive data and publishes nothing. Two agent-failure redaction tests that no longer configure credential names are renamed to say they redact credential-named environment values.
  10. fix(modal): MODAL_CAMPAIGN_ENVELOPE_TOO_SHORT and the Modal eval how-to stop recommending per-node Modal execution. Both now say to run the four-hour profile locally until the full benchmark's envelopes are enlarged, and the how-to drops the deleted 1,800-second lifecycle reserve and explicit-cap sentences. CONFIG_EXECUTION_CLOUD_REMOVED links to that page.
  11. fix(config): each CONFIG_EXECUTION_CLOUD_REMOVED message starts with its setting (execution.mode = "cloud", execution.provider or [execution.providers.modal]). Plain-text validate and doctor print messages without paths, so a config with all three printed the same ~330-character line three times.
  12. fix(runtime): doctor probes the local toolchain even when the config does not resolve. The skip existed because the probe once needed the config to choose between the local PATH and a Modal sandbox; after commit 2 it only reported git, node and forge missing next to every upgraded cloud user's CONFIG_EXECUTION_CLOUD_REMOVED.
  13. refactor: deletes leftovers commit 5 missed because neither TypeScript nor knip flags them: compileTask's env input (its only reader was the deleted cloudAgentCredentialEnv; startRun still built and passed it), the workflowModuleEntryUrls() wrapper and its empty-string filter (it existed for the conditional modal: "" entry), the fake runners' SMITHERS_FAKE_CLOUD_ENV_LOG hooks that printed MODAL_TOKEN_ID/MODAL_TOKEN_SECRET and their two allowlist entries, the "two cloud containers" rationale in three Effect-pinning comments (the pnpm-workspace.yaml copy is left to feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201, which rewrites that block), and one stateful-profile test name.
  14. fix(runtime): resume decides whether a run was planned for cloud execution from its smithers/resolved-config.json, not from smithers/tasks.json (Greptile P1). Native resume treats the task manifest as optional, so with a missing or malformed manifest commit 8's guard was skipped and a pre-removal cloud run's historical workflow went to Smithers. The pre-removal compiler gave every task config.execution.mode, so the config is the record of that mode.
  • submitSmithersContinuation reads the config before the manifest and returns WORKFLOW_CLOUD_EXECUTION_REMOVED when it says cloud.
  • It fails closed: a config that is missing or does not parse now fails a plain resume too, with WORKFLOW_LIFECYCLE_FAILED (run <id> cannot be resumed without its resolved config, which records whether it was planned for the removed per-node cloud execution: …). Before, a plain resume continued such a run with default concurrency and lease and no ULTRAFUZZ_CONFIG_PATH; --refresh-controller already failed.
  • With the config always present, the continuation's no-config fallbacks are deleted: the forge-guard and ULTRAFUZZ_CONFIG_PATH branches, the config?. defaults for concurrency, workspaces and lease, and recordNativeContinuationState's legacy-deadline branch.
  • Three fixtures that hand-build a run directory (the unsealed-legacy-launch test and the two real-engine continuation tests) now publish a default local resolved config through a new test/local-resolved-config.ts helper.
  • The CLI reference's resume paragraph (it said resume sets no ULTRAFUZZ_CONFIG_PATH for a run whose config does not parse) and the changelog entry describe the refusal.

Size: 107 files, +979/−21,513 (net −20,534):

Area Files Added Removed
Source 38 +209 −10,635
Tests 38 +691 −10,193
Schemas 9 0 −369
Docs and changelog 15 +58 −237
Build, CI and dependencies 7 +21 −79

About 180 cloud-only test declarations are deleted: 163 in 8 deleted test files (including the 119-declaration node-provider suite and the 14-declaration cloud-worker handoff suite) and 17 in edited suites. Ten tests are added (four by commits 8, 9, 12 and 14), and eight are renamed to drop "cloud" or configured-credential wording from their names.

Deliberately not built

  • No replacement for per-node remote execution. There is no other provider, no "bigger VM per heavy node", and no node-boundary machine rotation.
  • The inert [execution] data stays. That covers ResolvedConfig.execution with its loader, defaults, serializer and zod/JSON schema (including providers.modal parsing and resourceTimeoutOrigin), the execution blocks in plan.json and tasks.json, validate's execution_mode, and the CLI result's execution_mode/execution_provider enum. Removing it changes the artifact and config schema bundle digests, which strands in-flight runs on resume. No issue tracks it yet (Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921, the sealed-snapshot cost/benefit issue, does not list these items). The owner's direction on fix(runtime): a renderer or projection change no longer strands runs whose dynamic groups expanded #1216, to rethink tamper detection and possibly drop the sealing, would remove that digest reason, so it belongs in one follow-up issue with the items below.
  • workflow-sync.ts keeps its execution.mode === "local" filters and its scan of each task's recorded credential names. Every newly compiled task passes the filters and records no names, and feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 edits that file.
  • Stock adapters are untouched. environment.tsx still lists ULTRAFUZZ_MODAL_MODULE. The adapters are digest-enforced, so changing them for one unused name would make every project rerun ultrafuzz init before its next run.
  • Other deferred leftovers, for the same follow-up: hydrateTaskSpec's path.resolve(process.cwd(), …) wrappers, now no-ops because the renderer emits absolute paths (source-regex tests pin them, and fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 edits the template); and execution: { mode: "local" } fields in test fixtures of functions that no longer read them (runtime.test.ts, smithers-report-retry.integration.test.ts; the two in generated-workflow-verifier.test.ts sit in tests fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 deletes).
  • CONFIG_EXECUTION_NODE_UNKNOWN stays an error. An [execution.nodes.<id>] override for a node the topology lacks still blocks validate and run, as on main, although the override no longer changes execution. Downgrading it to a warning is a separate behaviour change for the follow-up; commit 6's test pins the current behaviour.
  • No committed resume-compatibility test for old generated workflows. A one-off check is described under Verification.

Verification

Everything below ran on the rebased tree, origin/main (a46a496) plus the first seven commits (head 57daecfa), unless it says otherwise. What ran for commits 8–13 is under Review follow-ups, and what ran for the stacked rebase onto origin/main 538b6188 and commit 14 is under Stacked rebase and commit 14.

Sealed identities. I printed them from the built trees of this head and of origin/main (a46a496):

  • VALIDATOR_BUILD_IDENTITY (…929e741c…), the artifact schema bundle (6a05a7a6…) and the config schema bundle (b799a29b…) are byte-identical to origin/main. (VALIDATOR_BUILD_IDENTITY was …028be325… on the previous base; main changed it, not this branch.)
  • The runtime bundle (c61b8387… → 17b5f648…, 19 → 18 schemas) and the Modal bundle (9a626ddb… → 6f0fd0af…, 21 → 14) change. Both are computed only at validation time, in CLI json validate and Modal document parsing. No persisted document records either one: the trusted-CLI identity binds the artifact bundle, which is unchanged.
  • Commits 8–14 change no schema file and none of the artifacts modules VALIDATOR_BUILD_IDENTITY hashes, and 538b6188 (fix(security): merge per-range registry advisories and patch brace-expansion and undici #1229) touches nothing under packages/artifacts, packages/config, packages/cli or a schema directory, so all five values hold at the final head (b8a2bc5b).

Discriminating tests. I ran each of these against origin/main (a46a496) by copying the test file into a main worktree, and each fails there as described:

  • config: rejects removed per-node cloud execution and still loads the local [execution] table init scaffolds and the named-diagnostic case removed cloud execution both fail. The first also round-trips the exact [execution] table ultrafuzz init writes.
  • runtime: planRun rejects removed cloud execution before creating the run directory fails: main returns no CONFIG_EXECUTION_CLOUD_REMOVED diagnostics (actual: []).
  • runtime: cleanRun removes a run planned for removed cloud execution without provider credentials fails with CLEAN_CLOUD_STORAGE_FAILED, because main's clean imports @ultrafuzz/modal and needs Modal credentials.
  • CI: build every package validate-release.mjs imports before any lane runs it fails (Received: "matrix.build_release_reporter == true").

Replacement coverage (not discriminating tests):

  • generated-workflow-render.test.ts takes over from the deleted cloud-worker harness, the only thing that rendered the whole generated program in process. It compiles a two-node local workflow, imports it with Smithers components stubbed, and renders it against the built runtime and artifacts. It checks the Worktree prepare/agent/verify tree, the dependency edges and the prompt body. Before this rebase I checked it with two mutations: it fails when a module-scope call to a deleted helper is added back (ReferenceError: readCloudExecutionGeneration is not defined), and when a verifier's dependsOn its agent is dropped. The test file and the template are byte-identical to that pre-rebase head, and main has not changed the template since the previous base.
  • config: merges local [execution] node resource overrides, rejects unknown nodes and bad bounds, and round-trips them restores the coverage that the deleted cloud tests gave the live local path. It also passes on main. Before this rebase, swapping the override precedence in resolveExecutionResources failed it, and so did inverting its unknown-node filter.

Resume compatibility:

  • main's template imports 7 names that commits 1–3 retire (6 from artifacts, CLOUD_EXECUTION_GENERATION_JSON_SCHEMA_ID from runtime). They are destructured at module scope and used only inside cloud-only functions; the one module-scope caller, readCloudExecutionGeneration(), returns "base" before touching its schema ID when no task is a cloud task. main's template is unchanged since the previous base, so the earlier AST scan still applies. The names commit 5 deletes appear in no template of v0.1.0, v0.1.1, v0.1.2, main or integration/wave1.
  • Cross-version, on the real engine, with a stub codex. This is a one-off harness adapted from the cli-e2e campaign-resume test and is not committed. origin/main's CLI (a46a496) runs init and run on the test's three-node campaign; the engine is SIGKILLed so the supervisor relaunches it, then the whole controller is SIGKILLed mid-node. This branch's CLI runs every later command (status, resume, events, report, stats). I ran it twice on the rebased tree, once with a plain resume and once with resume --refresh-controller, and both pass every assertion of the e2e test: the resumed run ends RunFinished with a verified report, only the interrupted node runs again, no finished task starts again, and stats records the interrupted attempt as canceled and counts the same attempts as status.

Suites on the rebased tree:

Suite Result
config (vitest) 90/90
artifacts 324/324
modal (vitest, 30 files) 522/522
29 runtime test files, listed below 391/391
Selected runtime.test.ts tests, listed below 50/50
CLI json-validate 16/16
CLI init and validate emit schema-versioned launch JSON, the run … clean … lifecycle test and doctor reports install posture 3/3
CLI cli-e2e campaign resume 1/1
Cross-version resume from origin/main's CLI, plain and --refresh-controller 2/2
topology packaged-topologies 14/14
CI policy scripts (Bun) 77/77
  • The 29 runtime files: the 13 this PR edits or whose fixture it edits (audit-contracts, clean, data-governance, generated-workflow-render, generated-workflow-verifier, lifecycle-inspection, pinned-submodules, runtime-document-contracts, smithers-dependency-skip, smithers-resume-reopen, task-workflow-identity, unverified-report, workflow-dependency-policy), the 12 other files that slice workflow.tsx (dynamic-lifecycle, invariant-suite-ancestor-order, invariant-suite-enumeration-overflow, invariant-suite-handoff-durability, report-retry-history, smithers-preparation-race, smithers-report-retry, stale-workspace-cleanup-overflow, terminal-report-projection, workspace-patch-replay, workspace-patch-supersede, workspace-preparation-lifecycle), and dynamic-workflow, generated-workflow-footprint, workflow-task-metrics and trusted-cli.
  • The runtime.test.ts selection: every test whose code this PR edits (including the four that share the edited invariant-budget fixture), every test that fix: a reworded finding no longer discards the final report #1204–fix(cli): recovery hints name ultrafuzz commands, not raw runner invocations #1207 added or changed there (among them the three syncRun distinguishes available agent reports after stopped failures variants; the failed-verifier and changed-output variants read a verifier-rejected report through the simplified secret scan), all controller-refresh, native-continuation, current-controller-rendering and reset-node tests, and the command-probe and credential-forwarding tests of the start-run and required-commands code this PR edits.
  • eatmydata: every real-engine test above ran under eatmydata, as CI does. Without it, ultrafuzz run fsyncs every file of the sealed execution snapshot, and my first pre-rebase cli-e2e and cross-version runs timed out.

Gates: all pass on the rebased tree.

  • pnpm install --frozen-lockfile (a real install: this PR changes the lockfile and removes a patch) and pnpm -w build, which also passes at each of the seven commits
  • pnpm -w lint, and CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci
  • pnpm -w knip, both on a fresh unbuilt worktree and after pnpm -w build
  • npx prettier --check on the 71 added or modified files, pnpm -w format:check, and pnpm -w docs:check (which runs scripts/docs-check.mjs)
  • pnpm --filter typecheck of artifacts, cli, config, modal, runtime and topology, and pnpm -w typecheck (all 12 workspace projects)

The complexity ceiling stays at 83. ESLint's per-function complexity in every source file this PR edits (the template included) is never higher than on origin/main; the maximum in those files is 76.

Review follow-ups (commits 8–13). Run on commit 13's head (533dbfdf, on a46a4960), built with pnpm -w build:

  • Discriminating tests, each checked by mutation:
    • resume refuses a run planned for removed per-node cloud execution before invoking Smithers launches a local run, rewrites its tasks.json to what the pre-removal compiler wrote for a cloud config, and checks that plain and --refresh-controller resume fail with only WORKFLOW_CLOUD_EXECUTION_REMOVED and issue no Smithers command. With the guard removed, the plain resume succeeds and the test fails (actual: true).
    • generated Smithers verifier refuses to publish the value of a credential-named environment variable fails when the template's assertArtifactPublicationsContainNoSecrets(publications, sensitiveEnvironmentValues(process.env)) becomes (publications, []); before commit 9, all 115 verifier tests passed under that mutation.
    • diagnoseProject still probes the local toolchain when the project config does not resolve fails with doctor's old resolved.config === undefined ? [] : skip restored.
  • Suites: config config.test.ts and stateful-profiles.test.ts 52/52; modal config.test.ts 23/23 and runner.test.ts 79/79; runtime generated-workflow-verifier and unverified-report 137/137; lifecycle-inspection, data-governance, clean, trusted-cli and dynamic-lifecycle 107/107; the smithers-terminal-resume, smithers-preparation-race and smithers-resume-reopen real-engine suites 8/8; the seven compileSmithersWorkflow caller files (generated-workflow-render, task-workflow-identity, generated-workflow-footprint, source-revision, pinned-submodules, workflow-dependency-policy, dynamic-workflow) 21/21; 22 selected runtime.test.ts tests (the new resume test, both controller-refresh tests, the 7 native-continuation tests, the 5 startRun forwards … tests, planRun rejects removed cloud execution, the command-probe test, the 4 tests that use the edited fake npm installer, and the runner-pin test whose comment commit 13 rewords), all passing, with 2 Bun-only variants skipped under node; CLI doctor reports install posture in human and JSON output and init and validate emit schema-versioned launch JSON 2/2; Bun workspace-engine-overrides 3/3.
  • Built-CLI probe on a project whose ultrafuzz.toml sets all three removed settings: validate and doctor print three CONFIG_EXECUTION_CLOUD_REMOVED lines whose messages each start with their setting, and doctor now lists git, node and forge with their paths and versions, reporting only the topology's genuinely absent commands missing.
  • Gates: pnpm -w lint; CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci; pnpm -w knip after pnpm -w build and on a fresh unbuilt worktree (pnpm install --frozen-lockfile --offline); npx prettier --check on the 77 added or modified files; pnpm -w format:check; node scripts/docs-check.mjs; pnpm --filter typecheck of config, modal and runtime; and the runtime and CLI test builds (tsc -p tsconfig.test.json).
  • Complexity: submitSmithersContinuation stays at 51 (51 on main) with the guard, and diagnoseProject drops to 39 (40 on main).
  • Trial merges of feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 (42a1bc85), fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 (594d6b87) and feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 (b3808c24) onto the old head and commit 13's head conflict in the same files with the same number of conflict hunks, so commits 8–13 add no merge work.

Stacked rebase and commit 14. Run on the final head (b8a2bc5b, on origin/main 538b6188) after pnpm install --frozen-lockfile and pnpm -w build:

  • Discriminating tests, each also run against commit 13's start-run.ts built into the same tree:
    • resume refuses a run planned for removed per-node cloud execution before invoking Smithers now rewrites the resolved config as well as the task manifest, as a pre-removal cloud launch wrote both, and checks both resume forms again with the manifest malformed and then deleted. Against commit 13 it fails at malformed task manifest, refreshController=false (actual: true: the plain resume reached Smithers).
    • resume refuses a run whose resolved config cannot be read instead of continuing it locally damages a local run's config (malformed, then deleted), checks that both resume forms fail with WORKFLOW_LIFECYCLE_FAILED and the config message before any Smithers command, then resumes the run once the config is restored. Against commit 13 it fails at malformed config, refreshController=false (actual: true).
  • Suites:
    • runtime.test.ts: every test that calls resumeRun, the two above included (40/40), and the 32 tests whose code this PR edits or that use its edited fixtures (32/32).
    • dynamic-lifecycle, lifecycle-inspection and the real-engine smithers-terminal-resume, smithers-preparation-race and smithers-resume-reopen suites: 81/81. The two hand-built continuation fixtures first failed with the new refusal; that is how they were found.
    • 16 other runtime files this PR edits or that slice its code (audit-contracts, clean, data-governance, dynamic-workflow, generated-workflow-footprint, generated-workflow-render, generated-workflow-verifier, operator-npm, pinned-submodules, runtime-document-contracts, smithers-dependency-skip, task-workflow-identity, trusted-cli, unverified-report, workflow-dependency-policy, workflow-task-metrics): 206/206.
    • config vitest 90/90; modal config.test.ts and runner.test.ts 102/102.
    • CLI json-validate 16/16, plus init and validate emit schema-versioned launch JSON, the run, ps, status, … lifecycle commands … test (which resumes), resume of an already-active run … and doctor reports install posture in human and JSON output: 4/4.
    • Bun release-validation-lanes, workspace-engine-overrides and dependency-advisories: 28/28.
  • Gates: pnpm -w format:check; pnpm -w lint; CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci; pnpm -w knip after pnpm -w build and on a fresh unbuilt worktree (pnpm install --frozen-lockfile --offline); pnpm --filter @ultrafuzz/runtime typecheck; node scripts/docs-check.mjs; and node scripts/check-production-dependency-advisories.mjs, because the lockfile was the conflicted file (0 High/Critical production advisories, 0 active exceptions).
  • Complexity: submitSmithersContinuation drops from 51 to 35 and recordNativeContinuationState from 12 to 5.

Not run:

  • the full runtime shards 1–4, the full CLI lane, validate:pack and super-linter
  • anything against live Modal
  • a resume of a real pre-upgrade cloud run: the resume tests rewrite a local run's resolved config and task manifest instead
  • a resume of a run launched by v0.1.0 (see Resume without a readable resolved config below)

Risk / compatibility

  • Product loss (Run each node on the cloud #134). This drops the planned ability to give each heavy node its own sized cloud VM, to fan independent branches out across machines, and to outlive a single sandbox's 24-hour lifetime. What remains is running the whole campaign on a bigger host, or inside one sandbox as ultrafuzz-modal does for eval rows. No release ever completed a cloud node, and the feature allowed only one attempt with no subscription auth, so no working capability is lost in practice. The plan is what goes. With this merged, Run each node on the cloud #134 would be closed as not planned.
  • Config break. A project with mode = "cloud", execution.provider or [execution.providers.modal] fails validate, doctor, run and eval run with CONFIG_EXECUTION_CLOUD_REMOVED until those keys are removed. While such a config fails to resolve, the commands that only locate runs (status, resume, events, …) fall back to .ultrafuzz/runs, as they already do for any invalid config. A project that also sets a custom run.output_dir therefore has to fix its config first.
  • Modal leftovers. ultrafuzz clean no longer terminates sandboxes tagged purpose=ultrafuzz-node or deletes the ultrafuzz-node-* volume of an earlier cloud attempt. Operators must remove those by hand in the Modal app the run used; the run's plan.json records it under execution.providers.modal until ultrafuzz clean deletes the run.
  • In-flight cloud runs. They cannot be continued; they could never complete a node anyway. resume, plain or with --refresh-controller, refuses one with WORKFLOW_CLOUD_EXECUTION_REMOVED (commits 8 and 14) instead of re-rendering it as local tasks on the operator's machine. The mode comes from the run's resolved config, so a missing or malformed task manifest does not let such a run through. replay and fork run the historical sealed workflow and are unchanged. pause, cancel and the read-only commands are not affected by the guard.
  • Resume without a readable resolved config (commit 14). A run whose smithers/resolved-config.json is missing or does not parse can no longer be resumed: it fails with WORKFLOW_LIFECYCLE_FAILED before Smithers starts. A plain resume used to continue it with default concurrency, lease and deadline handling. Every launch since v0.1.0 writes that file. v0.1.2's resolved-config schema is byte-identical to this branch's, and v0.1.1's differs only by a lower workflowDeadlineSeconds maximum, so their runs parse. What is affected is a run whose file was deleted or damaged, and possibly a v0.1.0 run; I did not check whether v0.1.0's configs still parse.
  • Native resume of older local runs. These run their old generated workflow against the new dist. The names retired by commits 1–3 appear only in cloud-only functions, commit 5 deletes no export that an old template uses, and the cross-version runs under Verification resumed a run that origin/main's CLI created.
  • resume --refresh-controller. It renders the new template, whose input schema drops the 5 nullable cloud-worker columns and whose task specs no longer carry execution. The engine's input load selects only declared columns, and the cross-version refresh run under Verification continued a pre-upgrade run to completion.
  • Environment variables. ULTRAFUZZ_MODAL_MODULE is no longer forwarded to the workflow engine or listed as controller-only in the runtime. Nothing reads it for local runs, and the stock adapter still strips it from agent environments.
  • Secret scans. The verifier's pre-publication scan, the agent-failure redaction, the verifier-rejected-report scan and startRun's launch-failure redaction no longer add configured cloud credential names to the values they refuse. Each of those lists was empty for local tasks on main, so results for local runs are unchanged; commit 9 restores a test that the pre-publication scan still refuses credential-named environment values.
  • PRs merging after this one. A trial merge (git merge-tree) of each PR's current head onto this one (b8a2bc5b) conflicts in: feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 (a811973d) — start-run.ts, lifecycle-inspection.test.ts, pnpm-lock.yaml and pnpm-workspace.yaml (this PR removes the Modal patch next to the Smithers patches feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 adds); fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 (3a1d3430) — CHANGELOG.md, workflow.tsx and generated-workflow-verifier.test.ts; feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 (eb783f17) — CHANGELOG.md only. Commit 14 adds only the start-run.ts conflict, one hunk in the continuation's runSmithersLifecycleCommand call: keep this PR's keepWorkspaces: config.run.keepWorkspaces and controllerLeaseSeconds: config.run.controllerLeaseSeconds with feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201's env: lifecycleEnvironment. None of the three adds a test that resumes a hand-built run directory, which would now need writeLocalResolvedConfig. fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 also needs a semantic fix when it is rebased: its new final-report selection code reads task.execution.mode !== "cloud" in workflow.tsx, and after this PR the rendered task specs no longer carry execution, so that expression would throw; every task is local now, so the branch should go. Source-slice tests index workflow.tsx by text markers; the markers and occurrence counts that pointed into or counted deleted template code were updated here.

Rebase notes

Stacked rebase onto 538b6188 (base of the #1197 → #1201 → #1183 → #1198 stack). The 13 commits moved from a46a4960 onto origin/main 538b6188, which adds one commit (#1229). One conflict:

Greptile (promotion review):

  • P1, packages/runtime/src/start-run.ts:546, "Missing manifest bypasses cloud refusal": fixed in commit 14. The finding is valid: native resume swallowed a missing or malformed tasks.json, which skipped the guard. The mode now comes from the run's smithers/resolved-config.json, which the continuation already parsed. A config that cannot be read fails the resume closed instead of continuing the run as local. Both new test branches fail on commit 13's code.

Earlier rebase. Rebased from 52cb2506 onto origin/main a46a4960 (19 new commits on main). Conflicts and how each was resolved:

Changelog

The entry is in CHANGELOG.md under Unreleased > Breaking changes:

[config] [runtime] [modal] [docs] Removes per-node cloud execution (#134), which never completed a node in any release (#712, #713). validate, doctor, run and eval run reject [execution] mode = "cloud", execution.provider and [execution.providers.modal] with CONFIG_EXECUTION_CLOUD_REMOVED before a run directory exists; remove execution.provider and [execution.providers.modal], and set mode = "local" or delete it. The rest of [execution] is still accepted and validated but no longer affects execution: every agentic attempt runs locally. To use Modal, run the whole campaign inside one sandbox, as the unchanged ultrafuzz-modal eval runner does. A run planned with mode = "cloud" cannot be continued: resume, with or without --refresh-controller, reads the mode from the run's smithers/resolved-config.json and fails with WORKFLOW_CLOUD_EXECUTION_REMOVED before starting Smithers, so start a new run. It also fails, with WORKFLOW_LIFECYCLE_FAILED, for a run whose resolved config is missing or does not parse, which a plain resume used to continue without it. ultrafuzz clean no longer terminates Modal sandboxes tagged purpose=ultrafuzz-node or deletes the ultrafuzz-node-* volumes of earlier cloud attempts; remove those by hand in the Modal app the run used, which its plan.json records under execution.providers.modal until ultrafuzz clean deletes the run. The runtime cloud-execution-generation schema and the seven Modal node schemas are deleted, so json validate no longer recognizes them.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the issues established in this review.

Summary

The PR removes per-node cloud execution while retaining the whole-campaign Modal eval runner.

  • Cloud execution settings now fail configuration validation, and historical cloud runs are refused on resume.
  • Runtime and Modal node-worker code, schemas, and tests are removed; release-validation lanes build their required artifacts dependency.
  • Since the previous review, resume reads the persisted resolved config independently of the optional task manifest, and the bundled operator npm and dependency-advisory handling were updated.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Resume request] --> B[Read persisted resolved config]
  B -->|Missing or invalid| C[Lifecycle failure]
  B -->|Cloud mode| D[Cloud execution removed]
  B -->|Local mode| E[Continue Smithers workflow]
Loading

Reviews (2) · Last reviewed commit: "fix(runtime): decide a resumed run's clo..."

@aviggiano
aviggiano force-pushed the claude/v04-remove-cloud-node-execution branch from f212184 to 52cb250 Compare September 29, 2026 10:55
@aviggiano
aviggiano force-pushed the claude/v04-remove-cloud-node-execution branch from 52cb250 to 57daecf Compare September 30, 2026 08:24
@aviggiano
aviggiano marked this pull request as ready for review September 30, 2026 09:56
@aviggiano
aviggiano requested a review from a team as a code owner September 30, 2026 09:56
Comment thread packages/runtime/src/start-run.ts Outdated
aviggiano and others added 12 commits September 30, 2026 12:00
…ION_CLOUD_REMOVED

Per-node cloud execution cannot complete a node on main or in any
release (v0.1.0 to v0.1.2). Two checks reject every cloud node before
inference:

- The node worker requires realpath(data root) to equal the lexical path.
  Modal's filesystem-v2 /data mount is a provider symlink (#712 live
  probe), so initializeDurableNodeWorkspace throws "durable data root is
  unsafe" for any data root below it.
- The worker exports its /proc/<pid>/fd/<n>/... descriptor path as
  ULTRAFUZZ_WORKFLOW_PERSISTED_PATH. The generated workflow evaluates that
  path at module load and throws "persisted workflow path must stay inside
  the cloud handoff project" (#713).

The fixes for both landed only on release/v0.1.0, a branch that no longer
exists; all three release tags carry the unfixed worker.

This is the decision commit of the removal. validateExecutionConfig now
reports CONFIG_EXECUTION_CLOUD_REMOVED for execution.mode = "cloud",
execution.provider and [execution.providers.modal], one diagnostic per
path, so planning fails before a run directory exists. The gate sits in
resolve, not in the TOML parser, because artifact gates re-parse each
run's persisted config.resolved.toml mid-run.

The rest of [execution] stays accepted and inert: ultrafuzz init
scaffolds it and every persisted run config carries it. Removing it
would change the config schemas, so it is left for later.

Tests: config.test.ts's three cloud tests and four cloud diagnostic cases
become one removal test (which also round-trips the table init writes)
and one named-diagnostic case. The six runtime tests that fed cloud TOML
through the loader are deleted; one of them becomes "planRun rejects
removed cloud execution before creating the run directory".

Refs #134, #712, #713

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… workflow and lifecycle

With execution.mode = "cloud" rejected at resolve time, the cloud branch
of every runtime layer is unreachable. Delete it:

- Compiler (smithers.ts): compileTask always emits the inert local
  execution block {mode: "local", resources, agentCredentialEnv: []}, the
  exact value it already produced for local runs, so tasks.json and the
  artifact schemas are unchanged. The cloud timeout reservation, the cloud
  credential classifier and its re-check, the relative-path rewriting for
  relocated workers, the modal module entry and the reset-node
  cloud-execution-generation.json write are gone. The config comments
  on resourceTimeoutOrigin now describe its one remaining use
  (serialization) rather than the removed cloud timeout cap.
- Generated workflow: the cloud-worker input envelope, the selected-task
  handoff DTO, the Modal Sandbox branch, the top-level dynamic import of
  @ultrafuzz/modal and the react Fragment import are gone. Every attempt renders
  as the existing Worktree prepare/agent/verify triplet.
- Lifecycle: doctor and run probe required commands on the local PATH
  only (no Modal app creation), clean no longer imports @ultrafuzz/modal
  to terminate a run's purpose=ultrafuzz-node sandboxes and delete its
  ultrafuzz-node-* volume, and data governance, validate, retry-chain and
  start-run lose their cloud-only branches.
- The runtime cloud-execution-generation contract and schema are deleted.
  Only CLI `json validate` reads the runtime schema bundle digest.
- Secret scans: the generated verifier's pre-publication scan uses
  sensitiveEnvironmentValues(process.env) alone, the value it already used
  for local tasks, whose agentCredentialEnv is always empty. The scan of a
  verifier-rejected final report (unverified-report-inputs.ts, #1204) uses
  the same value, and currentReportAttempt no longer returns the task's
  execution block.

The runtime no longer references @ultrafuzz/modal, so the five runtime
release-validation lanes stop building the modal closure.

workflow-sync.ts keeps its execution.mode === "local" filters: every task
compiled from now on passes them, and parallel changes are editing that
file.

Tests: cloud-worker-handoff.test.ts and its harness, the only tests that
rendered the whole generated program in process, are deleted with the
remaining cloud cases. generated-workflow-render.test.ts replaces that
coverage for local runs: it imports a compiled two-node workflow in
process and checks the rendered Worktree/Task tree and its dependency
edges. A new clean test removes a run whose plan.json recorded cloud
execution; before this change clean failed on it with
CLEAN_CLOUD_STORAGE_FAILED. Template-slice markers and counts that
pointed into the deleted code are updated. The tests that named
credentials through a cloud task go: the verifier test for configured
credential names (with its task `execution` field and the harness's
scanProcessEnvironment option) and the unverified-report test that named
them through a cloud task manifest. The remaining secret-gate tests cover
the environment-name heuristic and the vendor-format patterns that still
run.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…schemas and SDK patch

With the runtime no longer dispatching cloud nodes, nothing imports the
controller or worker half of per-node Modal execution. Delete it:

- packages/modal: node-provider.ts (sandbox provider, handoff archive,
  result polling, probeModalCommands, cleanupModalNodeRun), node-worker.ts,
  and their archive helpers modal-download.ts, deterministic-archive.ts
  and safe-archive.ts, with their tests.
- The seven node schemas (node input, result, checkpoint, checkpoint
  index, restore, worker error, execution-dependency manifest), their
  modal-common definitions, contracts, registry entries and semantic
  gates. The Modal schema registry drops from 21 to 14 schemas.
- packages/artifacts: the controller-to-worker selected-task DTO
  (cloud-selected-task.ts) and its export.
- packages/config: the MODAL_* sandbox lifetime constants.
- patches/[email protected] and its patchedDependencies entry. The patch
  only rewrote SandboxFilesystem.copyToLocal, which only modal-download.ts
  called. [email protected] stays and is now unpatched. packages/modal also
  drops tar-stream; packages/cli keeps its own tar-stream dependency.
- scripts/ci: the modal-sdk-download and safe-archive Bun tests.

The ultrafuzz-modal eval runner, which runs a whole local campaign
inside one sandbox, is unchanged. It imports none of these modules.

Schema identities: VALIDATOR_BUILD_IDENTITY and the artifact and config
schema bundle digests are byte-identical to the base. The Modal bundle
digest changes. It is computed at validation time only (Modal document
parsing and `json validate`); no persisted document records it.

Tests: the CLI json-validate Modal case and the Modal document tests now
use modal-smoke-checkpoint instead of the deleted node-input schema. The
Modal test fixture currentTaskOutputBinding is no longer exported:
node-provider.test.ts was its only importer.

Refs #134, #712, #713

Co-Authored-By: Claude Opus 5.5 <[email protected]>
scripts/validate-release.mjs imports packages/artifacts/dist when it loads.
The five runtime lanes only got that build as a side effect of
build_modal_dependencies, which the runtime dispatch deletion dropped. On a
fresh checkout, `node scripts/validate-release.mjs --gates runtime-1` then
fails with ERR_MODULE_NOT_FOUND before it selects a gate, so all five lanes
would fail on every pull request.

Every lane runs validate-release.mjs, so the reporter build is now
unconditional and the build_release_reporter lane flag is deleted. A new
lane-table test fails when a package that validate-release.mjs imports from
packages/*/dist is not built by an unconditional step before
`validate:release`; it fails on the previous workflow.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The earlier commits deleted the Modal node provider, the only production
caller of two exports:

- runtime: verifyCommittedControllerGenerationAuthority and its
  CommittedControllerGenerationAuthority result. #1193 deleted every
  other export of workflow-controller-generation.ts and kept this one
  only for the node provider, so the whole module goes: the journal and
  manifest readers it keeps are private to that verifier. The runtime
  index stops re-exporting it.
- artifacts: referenceArtifactManifestAuthorityForArtifactDir.

No generated-workflow template in v0.1.0, v0.1.1, v0.1.2, main or
integration/wave1 references either name, so resuming an older run is
unaffected. The artifacts test that called its export now asserts the
same fact through a live path: the parsed manifest keeps its reference
authority. The runtime verifier's only test was the dynamic
controller-refresh test, which #1193 deleted with the refresh code.

In the generated workflow, dynamicExecutionPath(task, value, label) read
`task` only for the removed cloud branch and had become a copy of
currentProjectPath(value, label). It is deleted, its 12 call sites call
currentProjectPath, and the seven path.resolve(process.cwd(), ...) wrappers
around its already absolute result are gone. The verifier test's
one-element local/cloud loop is inlined.

Stale cloud wording goes from six code comments, the artifact-contract
migration reference (the node schemas are deleted, not retained) and the
harness research page.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The deleted cloud config tests were the only coverage for [execution]
behaviour that still runs for local projects: validate and run call
validateExecutionNodeOverrides (CONFIG_EXECUTION_NODE_UNKNOWN), every
compiled task records resolveExecutionResources in tasks.json, and the
resolved TOML serializes [execution.nodes.*.resources].

A local-mode test now checks the node-override merge, an unknown node,
the serialization round trip and a cpu = 0 rejection. Swapping the
override precedence or inverting the unknown-node filter fails it.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Adds the breaking entry for this change under Unreleased. Two earlier
Unreleased entries described cloud-mode behaviour this change deletes, so
their cloud clauses go: the review-timeout entry (#1150) no longer tells
cloud users to add per-node timeout overrides to avoid
CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED, an error that no longer exists, and the
Codex openai_base_url route entry no longer says cloud planning rejects that
config, since cloud planning is gone.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…tion

A run planned with `[execution] mode = "cloud"` before the removal still
records `execution.mode = "cloud"` for every task in its
`smithers/tasks.json`, and its persisted `resolved-config.json` still
parses, because the JSON schema accepts the value and the new
CONFIG_EXECUTION_CLOUD_REMOVED gate lives only in `resolveConfig`.
`resume --refresh-controller` therefore rendered the current controller,
which has no Sandbox branch and ignores `task.execution`, and continued the
run with every remaining agent attempt running as a local worktree task on
the operator's host, with no diagnostic. A native resume ran the historical
workflow instead, which review found fails with an unrelated
module-not-found error for the deleted Modal module.

`submitSmithersContinuation` now reads the task manifest it already parses
and, when any task was planned for cloud execution, returns
WORKFLOW_CLOUD_EXECUTION_REMOVED before it inspects, renders or starts
anything. That covers both native and refreshed resume. Replay and fork
are unchanged: they run the historical sealed workflow rather than the
current controller, so they cannot move a run onto this host.

A new runtime test launches a local run, rewrites its task manifest to what
the pre-removal compiler wrote for a cloud config, and checks that both
resume forms fail with only that diagnostic and issue no Smithers command.
With the guard removed, the native resume succeeds and the test fails.

The changelog entry now says such runs cannot be continued, and tightens
two instructions reviewers flagged: setting `mode = "local"` alone does not
clear the error (provider and the providers table must go too), and the
leftover `ultrafuzz-node-*` volumes live in the Modal app the run's
`plan.json` records until `ultrafuzz clean` deletes it. The configuration
reference says the same, and its `[execution]` row no longer calls the
table "retained for compatibility": it is validated but has no effect.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…onment secret scan

The rebase deleted #1203's test that the generated verifier refuses to
publish a configured credential's value, together with the harness's
`scanProcessEnvironment` option, because that test covered the cloud
credential-name lists this PR removes. That option was the only path that
passed the real `sensitiveEnvironmentValues` into `verifyArtifacts`, so the
pre-publication scan of environment values, which still runs for every
local task, lost its only test: replacing
`sensitiveEnvironmentValues(process.env)` with `[]` in the template's
`verifyArtifacts` left all 115 verifier tests passing. The only remaining
verifier secret test uses a vendor-format token, which pattern matching
catches without any environment value.

The harness option is back, and a local test replaces the deleted one. It
writes a value with no vendor format into `result.json`. Held only by
ULTRAFUZZ_TEST_GATEWAY_LABEL, a name the credential heuristic does not
match, the value is published; once ULTRAFUZZ_TEST_GATEWAY_API_KEY also
holds it, `verifyArtifacts` throws "contains sensitive data" and publishes
nothing and writes no marker. With the template's scan emptied as above,
the new test fails.

Two agent-failure redaction tests no longer configure any credential name
(the task-level lists were cloud-only), so their names now say they redact
credential-named environment values, which is what they check.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
MODAL_CAMPAIGN_ENVELOPE_TOO_SHORT, which the eval runner raises when the
public full-benchmark row budget cannot hold the exhaustive campaign,
still told operators to "use local or per-node Modal execution", and the
Modal eval how-to described a 1,800-second per-node lifecycle reserve and
an explicit resource cap that this PR deletes. An operator following that
advice now gets CONFIG_EXECUTION_CLOUD_REMOVED, whose message links to
this same how-to page.

Both now say to run the four-hour profile locally until the full
benchmark's envelopes are enlarged, and the how-to drops the reserve and
cap sentences. The tests match only the diagnostic's prefix, so none
change.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…message

A config that sets `mode = "cloud"`, `provider` and
`[execution.providers.modal]` gets one CONFIG_EXECUTION_CLOUD_REMOVED
diagnostic per setting, but all three carried the same message, and
plain-text `ultrafuzz validate` and `doctor` print messages without their
paths. The operator saw one ~330-character line three times with nothing
to tell them apart.

Each message now starts with its setting, for example
`execution.provider is no longer supported: per-node cloud execution was
removed; ...`. The JSON paths are unchanged.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…does not resolve

`diagnoseProject` skipped the required-command probe whenever the project
config failed to resolve, and then reported every required command as
missing. The gate existed because the probe needed the config to choose
between the local PATH and a Modal sandbox; after this PR,
`probeCommandsForExecution` takes no config, so the gate only produced a
false DOCTOR_TOOLCHAIN_MISSING for git, node and forge next to every
upgraded cloud user's CONFIG_EXECUTION_CLOUD_REMOVED.

Doctor now always probes the local PATH. A new test gives doctor a
`mode = "cloud"` config and an all-available probe, and checks that git,
node and forge are probed and the toolchain check is ok; with the gate
restored it fails.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 2 commits September 30, 2026 12:00
Review found code and wording the earlier cleanup commit missed because
neither TypeScript nor knip flags them:

- `compileTask`'s `env` input: its only reader was the deleted
  `cloudAgentCredentialEnv`. `startRun` still built
  `{ ...process.env, ...input.env }` and passed it through
  `compileSmithersWorkflow`, where nothing read it. The field goes from
  `SmithersCompileInput` and `compileTask`, with both call-site props and
  the `startRun` argument.
- `workflowModuleEntryUrls()` and the empty-string filter on its result
  existed only for the conditional `modal: ""` entry. Both URLs are now
  always present, so the two `import.meta.resolve` calls are inlined.
- The fake runners' `SMITHERS_FAKE_CLOUD_ENV_LOG` hooks, which printed
  MODAL_TOKEN_ID and MODAL_TOKEN_SECRET, and the
  `SMITHERS_FAKE_CLOUD_ENV_LOG`/`SMITHERS_FAKE_CLOUD_SELECTOR_LOG` entries
  in the Smithers test environment allowlist: only the deleted cloud
  credential-forwarding tests set them.
- The Effect-pinning comments in smithers-package.ts, runtime.test.ts and
  the workspace override CI test justified the pin with two cloud
  containers; the reason holds for a run's launch and a later resume,
  which install at different times. The pnpm-workspace.yaml copy is left
  to #1201, which rewrites that block.
- A stateful-profile test named "preserves cloud timeout inheritance" now
  only checks that the resource timeout origin survives TOML snapshots.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The resume guard for removed per-node cloud execution read the mode from
`smithers/tasks.json`. Native resume treats that manifest as optional:
when it is missing or does not parse, the guard was skipped and the
historical workflow of a pre-removal cloud run went to Smithers instead
of failing with WORKFLOW_CLOUD_EXECUTION_REMOVED (Greptile P1 on
start-run.ts).

The pre-removal compiler gave every task `config.execution.mode`, so the
run's `smithers/resolved-config.json` is the record of that mode.
`submitSmithersContinuation` now reads that config first and returns
WORKFLOW_CLOUD_EXECUTION_REMOVED when it says cloud, before it reads the
task manifest. A config that is missing or does not parse fails the
resume with WORKFLOW_LIFECYCLE_FAILED ("run <id> cannot be resumed
without its resolved config, ...") for a plain resume too. Before, a
plain resume continued such a run with default concurrency and lease
and no ULTRAFUZZ_CONFIG_PATH, so its planned execution mode was never
checked; `--refresh-controller` already failed in both cases.

With the config always present, the continuation's fallbacks for a run
without one are dead and are deleted: the forge-guard and
ULTRAFUZZ_CONFIG_PATH branches, the `config?.` defaults for concurrency,
workspaces and lease, and recordNativeContinuationState's legacy
deadline branch. submitSmithersContinuation's complexity drops from 51
to 35.

Tests: the cloud-refusal test now rewrites the resolved config as well
as the task manifest, as a pre-removal cloud launch wrote both, and
checks that both resume forms are still refused once the manifest is
malformed and then missing; on the previous guard the plain resume of
the malformed-manifest run succeeds. A new test damages a local run's
config (malformed, then missing), checks that both resume forms fail
before any Smithers command, then resumes the run once the config is
restored; on the previous code the plain resume with a malformed config
succeeds. Three fixtures that hand-build a run directory (the unsealed
legacy launch test and the two real-engine continuation tests) now
publish a default local resolved config through a shared test helper.

The CLI reference and the changelog entry describe the new refusal.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/v04-remove-cloud-node-execution branch from 533dbfd to b8a2bc5 Compare September 30, 2026 12:58
aviggiano added a commit that referenced this pull request Sep 30, 2026
…ead of querying smithers

After a controller restart (resume, quota park, supervisor relaunch) the
generated workflow's process-local record of which chain rung each
report-producer attempt used is gone. The producer's retry prompt and the
verifier both rebuilt it by running a bare `smithers node ... --full-output`
from inside the workflow: a runner subprocess on the finalizer path that
reads up to 64 MiB under a 180 s budget. Wherever the controller PATH
resolved no `smithers`, every such attempt failed with "Smithers
report-producer authority is unavailable" (cause: ENOENT) until the report
node ran out of retries (#1143). The only real-Smithers test of this path
gave the detached engine a `smithers` on its PATH.

Record each executed selection {attempt, chainIndex} with writeFileDurable
in <run>/smithers/final-report-selections/<attempt-id>.json before the agent
runs, and read it back when the process-local cache is empty. A
re-dispatched attempt number replaces its own entry;
finalReportAgentExecution keeps validating attempt order and chain bounds.
This deletes readFinalReportSmithersAuthority,
priorFinalReportAgentSelections, the 180 s
SMITHERS_REPORT_PRODUCER_AUTHORITY_TIMEOUT_MS budget and the workflow's use
of the Smithers attempt reconcilers. The runtime exports stay: host-side
workflow-sync uses inspectSmithersAttemptAgentSelection, and workflows
rendered before this change import both.

Per-node cloud execution is gone (#1197) and rendered task specs no longer
carry `execution`, so every producer attempt writes the record and the
verifier has no single-rung cloud shortcut. The history and integration
fixtures carry no `execution` either, so a leftover `task.execution` read
in the template fails both.

This reverses #585's "a non-Codex fallback cannot forge final-report
producer authority through the run filesystem" test. That property was not
a real boundary: an unsandboxed agent running as the same user can edit the
Smithers database as easily as a run-directory file, and trusted-cli.ts
already states the host is not a same-UID sandbox boundary. Both the
database and the new record sit outside the agent's worktree and declared
artifact directories.

Tests: the real-Smithers restart test now runs the detached engine with a
PATH that resolves no `smithers` (and calls pause by absolute path); with
the previous template both variants fail with ENOENT, with this change both
pass. report-retry-history.test.ts is rewritten to drive the extracted
helpers across simulated restarts against a temporary run directory. The
tests that pinned the CLI query (fake execFileSync, read counts, the budget,
source regexes) are deleted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit d2b5cce into main Sep 30, 2026
17 checks passed
@aviggiano
aviggiano deleted the claude/v04-remove-cloud-node-execution branch September 30, 2026 15:29
aviggiano added a commit that referenced this pull request Sep 30, 2026
…ead of querying smithers

After a controller restart (resume, quota park, supervisor relaunch) the
generated workflow's process-local record of which chain rung each
report-producer attempt used is gone. The producer's retry prompt and the
verifier both rebuilt it by running a bare `smithers node ... --full-output`
from inside the workflow: a runner subprocess on the finalizer path that
reads up to 64 MiB under a 180 s budget. Wherever the controller PATH
resolved no `smithers`, every such attempt failed with "Smithers
report-producer authority is unavailable" (cause: ENOENT) until the report
node ran out of retries (#1143). The only real-Smithers test of this path
gave the detached engine a `smithers` on its PATH.

Record each executed selection {attempt, chainIndex} with writeFileDurable
in <run>/smithers/final-report-selections/<attempt-id>.json before the agent
runs, and read it back when the process-local cache is empty. A
re-dispatched attempt number replaces its own entry;
finalReportAgentExecution keeps validating attempt order and chain bounds.
This deletes readFinalReportSmithersAuthority,
priorFinalReportAgentSelections, the 180 s
SMITHERS_REPORT_PRODUCER_AUTHORITY_TIMEOUT_MS budget and the workflow's use
of the Smithers attempt reconcilers. The runtime exports stay: host-side
workflow-sync uses inspectSmithersAttemptAgentSelection, and workflows
rendered before this change import both.

Per-node cloud execution is gone (#1197) and rendered task specs no longer
carry `execution`, so every producer attempt writes the record and the
verifier has no single-rung cloud shortcut. The history and integration
fixtures carry no `execution` either, so a leftover `task.execution` read
in the template fails both.

This reverses #585's "a non-Codex fallback cannot forge final-report
producer authority through the run filesystem" test. That property was not
a real boundary: an unsandboxed agent running as the same user can edit the
Smithers database as easily as a run-directory file, and trusted-cli.ts
already states the host is not a same-UID sandbox boundary. Both the
database and the new record sit outside the agent's worktree and declared
artifact directories.

Tests: the real-Smithers restart test now runs the detached engine with a
PATH that resolves no `smithers` (and calls pause by absolute path); with
the previous template both variants fail with ENOENT, with this change both
pass. report-retry-history.test.ts is rewritten to drive the extracted
helpers across simulated restarts against a temporary run directory. The
tests that pinned the CLI query (fake execFileSync, read counts, the budget,
source regexes) are deleted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…loud run

Problem: before #1197 removed per-node cloud execution, `ultrafuzz clean`
of a run planned with `[execution] mode = "cloud"` terminated its sandboxes
(tagged purpose=ultrafuzz-node) and deleted its Modal volume before
deleting the run directory. #1197 dropped that step and kept the deletion,
so cleaning such a run now leaves its Modal volume and any running sandbox
in place, still billed, and deletes the plan.json that records the Modal
app they live in, without saying so.

Change: before anything is removed, clean reads the mode and Modal app
from each selected run's plan.json, which copies the execution block of
the run's resolved config and is what the removed cleanup read. For a
cloud run it returns a CLEAN_CLOUD_STORAGE_RETAINED warning naming the
app, the volume and the sandboxes' run tag, and the `modal volume delete`
command. `--dry-run` reports it without deleting, and every later result,
including a failed removal, carries it. The deletion itself is unchanged.

The volume is not `ultrafuzz-node-<run-id>`: the removed provider named it
`ultrafuzz-node-` plus a bounded identity of the Smithers run ID
`ultrafuzz-<run-id>` (at most 32 normalized characters, then 12 hex digits
of its SHA-256). The same identity was the sandboxes' `run` tag. clean
reproduces that function, and the test pins the name the removed code
computes for a fixed run ID.

An unreadable plan.json yields no warning. It also makes the source
revision step fail before anything is removed, as it did before, so no run
that clean deletes skips the check.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
Problem: #1197 made compilation write every task's execution block as
`{ mode: "local", resources, agentCredentialEnv: [] }` with no `modal`
entry, but synchronization still added task.execution.agentCredentialEnv
and task.execution.modal?.credentialEnv to its exact-value redaction list.
For every run compiled since #1197 both are empty. The generated agent
environment also still blanked ULTRAFUZZ_MODAL_MODULE, which only the
removed cloud provider set. And the doctor tests still wrote a fake engine
into the target's .smithers, although since #1201 doctor inspects only the
runner Ultrafuzz's own install provides.

Change: synchronization derives its redaction values from the environment
alone, the agent environment no longer lists ULTRAFUZZ_MODAL_MODULE, and
seven doctor tests lose the dead fake-engine setup. One test used it only
to create the .smithers/node_modules/.bin directory it writes into, and
now creates that directory itself. The doctor test that installs a 0.29.0
project-local engine keeps it: it asserts that doctor ignores that engine.
The task manifest keeps agentCredentialEnv, because its sealed schema
requires the field.

Effect on runs compiled before #1197 in cloud mode, whose lists were not
empty: sensitiveEnvironmentValues still redacts, by name, every variable
that looks like a credential (*_API_KEY, *_TOKEN, *_SECRET, ...), which
includes MODAL_TOKEN_ID and MODAL_TOKEN_SECRET. What these lists added
beyond that were route variables such as *_BASE_URL and allowlisted
variables. Those carry no secret, or have a value the secret patterns
already match.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
… about old cloud runs' Modal storage in clean, and drop the smithers shim (#1235)

### 1. Forge guard dropped from the engine PATH, while run.json said it was active (bug on `main`)
Under a umask that leaves new directories group writable, such as Ubuntu's default `0002`, `ensureSafeDirectory` created `<run>/safe-bin` with mode `0775`. `composeSmithersCommandPath` admits a PATH entry inside the target only through `isPreparedForgeGuardBin`, which refuses a directory where `mode & 0o022` is set. The workflow engine's PATH therefore dropped the wrapper, and agents ran the real `forge` without `forge_vmem_limit_kb` or `forge_rayon_threads`, while `run.json` recorded `forge_guard.active: true`.
- To reproduce on `main`, run it under this host's umask `0002`. `startRun injects the configured Forge guard into the workflow environment and metadata` fails because the fake engine resolves `/tmp/ufz-runtime-…-fake-bin/forge` instead of `<run>/safe-bin/forge`.
- A second cause, found while fixing the first: with any custom `run.output_dir`, `run.json` makes the same false claim under umask `022` too, because `isPreparedForgeGuardBin` only admits the wrapper from `<project>/.ultrafuzz/runs/<run-id>/safe-bin`. A probe on `main` printed `.ultrafuzz/runs active: true safe-bin on engine PATH: true` and `audit-runs active: true safe-bin on engine PATH: false`.
- **The OpenRouter adapter-contract failure was a test-fixture bug, not a product bug.** `generated OpenRouter adapter preserves opaque model IDs…` failed on `main` under `0002` with `provider-home ancestors cannot be group/world writable`. The ancestor the check refused was `<project>/.ultrafuzz`, created by `ultrafuzz init` with the process umask (a probe showed `0775`). The fixture had rooted the operator-owned provider homes under the target. The adapter itself creates every provider-home component with `mkdirSync(…, { mode: 0o700 })`, and the umask cannot widen that mode. In production the default root is `$XDG_STATE_HOME` or `~/.local/state/ultrafuzz/provider-homes`, which is not under the target.
### 2. `ultrafuzz clean` of a run planned for per-node Modal execution (Greptile P1 on #1197)
Before #1197, `clean` terminated such a run's sandboxes (tagged `purpose=ultrafuzz-node`) and deleted its Modal volume, then deleted the run directory. Now it deletes only the directory. The volume and any running sandboxes stay in place and may still be billed, and the deletion removes the `plan.json` that names their Modal app.
### 3. Leftovers of #1197 and #1201

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant