Skip to content

feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) - #1201

Open
aviggiano wants to merge 3 commits into
mainfrom
claude/v10-install-based-controller
Open

aviggiano wants to merge 3 commits into
mainfrom
claude/v10-install-based-controller

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Step 1 of the #921 plan (the judged plan's PR-1 and PR-7). It changes which engine bytes a resumed run executes.

Owner decision: a pnpm checkout with the Smithers patchedDependencies is the only supported install, so packed and plain-npm installs are refused (#921 question 3), and long campaigns run from a dedicated checkout or worktree (#921 question 4), which docs/reference/cli.md now tells operators.

Problem

Any Ultrafuzz process that ran a Smithers command without a sealed runner installed its own copy of the engine first. That covers ultrafuzz resume (including an attach to a live run), its pre-resume inspect, ultrafuzz ps and the dashboard's run listing, the --refresh-controller ownership inspection, and commands on runs without a sealed snapshot. The install was a registry npm install of the pinned closure (the vendored npm 11.19.0 with the --before cutoff) into mkdtemp(os.tmpdir())/ultrafuzz-controller-*, followed by the 69-entry string-patch registry.

On this host, measured on main a46a496:

  • Slow first command. The first Smithers command of a fresh process, such as a resume, took 42.8 s and 43.8 s (load 14–15).
  • Registry dependency. Resume needs registry.npmjs.org. With every proxy and npm_config_registry pointed at a listener that drops connections, the native-continuation integration test's resume opened 6 registry connections, and the test took 213 s instead of 9 s.
  • A TMPDIR root left per resume. The resumed detached engine and supervisor run from that install, so resume marks it retained and nothing ever deletes it. The newest retained roots on this host are 595 MB each, and 142 ultrafuzz-controller-* roots have piled up in /tmp. A tmp cleaner that reaps a root also breaks the live engine's later supervisor relaunches (Pin required executables and runtime dependencies for detached runs #1143, point 5).
  • Bare smithers unresolvable (Pin required executables and runtime dependencies for detached runs #1143 point 1). The generated workflow calls a bare smithers node to rebuild final-report producer authority after a restart. The controller PATH carries no runner, so on a production-like PATH every restarted report attempt fails with Executable not found in $PATH: "smithers".
  • Tests ran unpatched bytes. The repository's own smthrs install was unpatched, so the real-Smithers tests ran different bytes from production, and three of them patched their own copy of the runner.

Root cause

The Smithers compatibility patches existed only as a TypeScript registry applied at runtime to a freshly npm-installed tree. So every process that needed a runner had to create and patch one, and the repository install never had them.

Change

Commit 1: build(runtime), ship the patches as pnpm patchedDependencies

  • Generated patch files. scripts/smithers-patches.mjs generates patches/[email protected] and patches/@smthrs__{agents,cli,db,engine,scheduler}@0.35.0.patch (2,166 lines). It applies SMITHERS_COMPATIBILITY_PATCHES to pristine copies of the pinned packages with applySmithersCompatibilityPatches, the same function launch uses, and diffs them the way pnpm patch-commit does.
  • Workspace wiring. pnpm-workspace.yaml gets the patchedDependencies entries, and the lockfile is refreshed.
  • Drift check. A new CI build-gate step runs node scripts/smithers-patches.mjs --check. It fails when a committed patch file or its workspace entry no longer matches the registry. The registry stays the source of truth until a later step deletes it.
  • Tests now see production bytes:
    • The upstream-shape tests (patch anchors, envelope contracts, delegation refusal) undo the installed replacements, asserting that each is installed exactly once and that re-applying the registry reproduces the installed file.
    • The usage, cost and path-acceptance tests import the installed patched modules instead of patching a copy.
    • The controller-engine and resume-reopen integration tests drop their private patchers.
    • fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199's supervisor-relaunch integration test runs the installed runner instead of a patched copy, so the shared patched-smithers-runner.ts test helper is deleted. Its leftover-process cleanup now matches the run ID, because the runner no longer lives below the test root.
    • The Bun adapter contracts' OpenRouter helper checks the installed ordered-stdout patch instead of applying it.
    • The DeepSeek usage contract asserts what the patched agents report: inputTokens 527 = 120 fresh + 400 cache read + 7 cache write, with each component in inputTokenDetails. It used to expect upstream's 120, which production never reported.

Commit 2: feat(runtime)!, commands after launch run the installed runner

The installed runner

  • Every command that used to install a controller now binds installedWorkflowRunner() instead. That is the smthrs package @ultrafuzz/runtime resolves, which pnpm has already patched. These commands are the ones listed under Problem.
  • Runner guard. The runner is refused unless all 69 registry patches and both engine anchors classify as applied. Commit 3 changes the refusal text and adds the outside-target check.
  • Bun flags. Its command process runs with BUN_TARGET_CONFIGURATION_GUARD_ARGS (--config=/dev/null --no-env-file --no-install --no-addons), the constant fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199 added for the unsealed native continuation, so it does not load the target's bunfig.toml or .env. The flags do not confine what the runner imports; commit 3 closes the one target import Smithers makes. the installed Bun runner ignores target startup files and still runs its attested bytes covers the flags. fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199's narrower test (native operator continuations run Bun without the target repository's bunfig.toml or .env) repeated that test's setup and assertions, so it is deleted.
  • Spawned children. fix(runtime): the Smithers supervisor can relaunch a crashed engine #1199's spawn patches already pass the same flags to the detached engine and supervisor this process spawns, and to the supervisor's relaunch of a dead engine. The regenerated CLI patch file carries them, and docs/security.md now says so.
  • NODE_PATH. A continued target workflow resolves its bare imports from the runner's own pnpm dependency directory.
  • Resolution. The package is located with import.meta.resolve, because once Bun has loaded smthrs, its require.resolve("smthrs") returns the bare specifier. The Bun contract suite hit this.

Deleted

  • The native-continuation environment marker.
  • The .ultrafuzz-native-continuation retention of the TMPDIR root, and its exit-hook exemption.
  • The per-command controller-closure seal check.
  • The TMPDIR controller's Bun-interpreter assertion.

<run>/trusted-bin/smithers shim

smthrs becomes a runtime dependency of @ultrafuzz/runtime

  • The production advisory gate now audits the engine tree (910 packages instead of 332). It found GHSA-p95v-992w-h6c3 (high: prototype pollution when decoding untrusted TOON) in @toon-format/toon 2.3.0, which @smthrs/cli 0.35.0 pins exactly.
  • A scoped override, "@smthrs/cli>@toon-format/toon": "2.3.1", takes the patch release. With it, the engine tree adds no High/Critical advisory: the registry reports the same advisories for main's production tree as for this branch's (see Verification).
  • Only the Smithers pack and share flows decode TOON (packs.js, manifest.js). Ultrafuzz runs neither.
  • The Effect-overrides CI test now compares only the Effect entries.
  • Both closure walkers skip smthrs. Each walker follows every declared dependency of the @ultrafuzz/* packages: launch's execution-snapshot walker and the trusted-CLI closure walker. Following smthrs would copy a second engine closure into every run's snapshot and into every trusted-CLI closure, which is re-hashed on each ultrafuzz call. So both skip a first-party package's smthrs edge, and the sealed engine keeps using the snapshot's own root runner. This is one line in each walker, and both walkers are deleted in later steps (plan PR-8 and PR-10).

validate:pack

  • It copies the Smithers patch files, the root overrides and the full allowBuilds map into the consumer workspace.
  • It then requires ultrafuzz doctor there to report both workflow-engine checks as ok.

doctor

  • It reports the installed runner's version, bin target, layout and per-patch posture as ok/error checks. Before, these checks were always unknown and described a project-local tree launch never used.
  • Human output drops the runner-path line, because the path contains the engine package name.

Docs and changelog

  • docs/reference/cli.md (resume, doctor) and docs/security.md (workflow engine install).
  • The resume section says to run long campaigns from a dedicated checkout or worktree (commit 3 corrects what triggers the problem).
  • CHANGELOG.md gets the entry under ## Unreleased → Breaking changes.

Commit 3: fix(runtime), review fixes

Pin the SQLite store for every engine process (the major review finding)

  • Smithers 0.35.0 picks its store backend in resolveSmithersBackendPreference (smthrs/src/resolveSmithersBackendChoice.js). Unless --backend or SMITHERS_BACKEND is set, that function runs await import("<workspace>/.smithers/smithers.config.ts"). Ultrafuzz set neither, and the store-opening commands it runs (ps, inspect, node) all go through it.
  • The installed runner has no module confinement to refuse that import. On 42a1bc8 a planted <target>/.smithers/smithers.config.ts wrote its marker file from runSmithersInspectionCommand(["ps", …]), the code behind ultrafuzz ps and the dashboard's listing, and from <run>/trusted-bin/smithers ps. The shim's run had ANTHROPIC_API_KEY in its environment. Until fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 lands, the finalizer calls smithers node through that shim with the engine's provider credentials.
  • smithersCommandEnv now sets SMITHERS_BACKEND=sqlite for every engine process, and the shim exports it. Ultrafuzz's store is always <target>/smithers.db. With the pin, Smithers' .smithers/backend.json and migrated.json markers no longer choose the backend, and it skips its PGlite probe.
  • The detached engine, the supervisor and the supervisor's relaunch all spawn with ...process.env, so they inherit the pin.
  • A sealed runner benefits too. Its confinement refuses the config import, and before the pin that refusal failed the command.
  • New test: installed-runner commands and the run's smithers shim never import the target's Smithers config, on the real installed runner. It fails if either pin is removed.

Launch refuses an unusable installed runner before creating the run

  • The launch preflight (beforeMaterialize) calls the new bindInstalledWorkflowRunner. This is the same binding the shim writer and every installed-runner command use: the patch guard, the outside-target check, the runner digest and Bun resolution.
  • A failure is RUN_PREFLIGHT_FAILED. It happens before the run directory, the controller install or any registry access exists.
  • Resume, replay and fork keep their late check. Before their shim write they only take the lifecycle lock, read evidence, and prepare the forge guard and trusted CLI, so a refusal there changes no run state.

The installed runner is refused inside the target project, as an explicit SMITHERS_BIN already was. The check is assertExecutableOutsideRoot in bindOperatorSmithersExecutableCapability, and capability binding rejects target-contained runner and interpreter paths covers it.

Refusal text

  • Before: …lacks Ultrafuzz's compatibility patches (<2.7 KB of id: posture pairs>); reinstall Ultrafuzz with pnpm install --frozen-lockfile.
  • That remedy fixed neither of two cases:
    • A stale build. The registry that classifies the installed files is compiled into Ultrafuzz, so a pull and an install without a build is refused the same way.
    • A consumer project, where pnpm install --frozen-lockfile applies no patches.
  • Now: installed workflow runner <path> lacks N of 71 Ultrafuzz compatibility patches (ultrafuzz doctor lists them); reinstall and rebuild Ultrafuzz in its repository checkout: pnpm install --frozen-lockfile && pnpm -w build.
  • Doctor's two workflow-engine-* checks share the hint (WORKFLOW_RUNNER_REINSTALL_HINT) and the filter (unappliedCompatibilityPatches). The patches check still lists the patch ids.

A long-lived process whose runner directory is gone is told to restart

  • The resolved runner is cached per process. After a pnpm install moves and prunes the engine, a running dashboard or eval run still binds the deleted path.
  • The suggested fix, re-resolving, cannot follow the move, because Node 24.21.0 and Bun 1.3.14 both cache import.meta.resolve. I checked with a throwaway package: I re-pointed its node_modules symlink and removed the old directory. import.meta.resolve in the same process still returned the old path, and realpath failed with ENOENT.
  • installedWorkflowRunner() now throws installed workflow runner <path> no longer exists, most likely because a pnpm install in Ultrafuzz's checkout replaced it; restart this Ultrafuzz process, instead of an lstat ENOENT. docs/reference/cli.md tells operators to restart the dashboard and eval run after such an install.

Dead code the per-command path left behind

  • Launch's controller install no longer writes controls/{bun-module-confinement.js,bun-empty.env,bunfig.toml} into the TMPDIR controller, and no longer hashes them into its in-process seal. Sealed snapshots take theirs from <run>/smithers/bun-startup-controls. writeCurrentBunStartupControls is deleted.
  • Deleted: operatorControllerProjectRoot's control parameter, which no caller passed, and the signal plumbing in ensureSmithersDependencies.
  • Resume's trusted-CLI fallback no longer sets ULTRAFUZZ_TRUSTED_BIN just before the shim writer sets it to the same directory.

Tests

  • The launch test asserts that the execution snapshot's module:@ultrafuzz/runtime issuer has no smthrs edge. Before, disabling that walker exclusion only slowed the test (36 s instead of 7 s), because every snapshot then copied the engine closure.
  • The agent-usage helper reads the installed agents and rewrites only BaseCliAgent.js, the one file it changes.

Docs

  • The dedicated-checkout guidance, in the docs and the changelog, now names the real trigger: any install that changes the engine's resolved dependency tree, not only a patch change. pnpm names smthrs's directory from its full dependency path, which includes about 25 peers such as [email protected] and [email protected].
  • It also says:
    • leave the build alone too, because a resumed workflow imports the checkout's built @ultrafuzz/runtime;
    • the running engine can fail on a lazy import before any relaunch;
    • every run's trusted-bin/smithers points into the install.
  • docs/security.md says:
    • which commands run the unsealed, unconfined installed engine, and what is still anchored;
    • that every engine process gets the SQLite pin;
    • that agents can drive their own run through the shim.
  • The validate:pack comment and doctor's ok summary no longer describe the guard as resume-only.

Unchanged

  • Launch, apart from binding the installed runner in its preflight and writing the shim, still installs, patches and seals its own npm controller into the run's execution snapshot, and the sealed engine runs from there.
  • Sealed-run lifecycle commands. pause, cancel, status, replay and fork on a sealed run still execute the run's sealed runner through the snapshot. They never installed a controller.

Deliberately not built

  • Moving sealed-run pause/cancel/status/replay/fork to the installed runner.
    • These commands run through the snapshot's descriptor anchors and Bun module confinement, which a runner outside the snapshot cannot pass.
    • replay and fork load the sealed workflow, whose bare imports resolve to the snapshot's own smthrs, React and Effect copies.
    • This needs the plan's PR-6 (plain-path execution) first.
  • A digest manifest of the patched files (patches/smithers-patched-files.json in the plan).
    • The guard classifies every registry entry against the installed files instead. That catches unpatched (packed or npm) installs and missing, reverted or partially applied patches, with no second generated artifact to keep in sync.
    • It does not detect other edits to those files.
  • A per-run engine closure (the plan's content-addressed engine mirror). The owner chose the dedicated-checkout guidance instead (Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921 question 4).
  • Later plan steps:
    • deleting the TS registry, applier, vendored npm and launch controller install (PR-8);
    • sealing or hashing the installed closure (the plan drops it).
  • An early refusal in resume, replay and fork. Nothing changes run state before their shim write (see Commit 3).
  • Re-resolving the engine in long-lived processes. Node and Bun cache the resolution, so this would need a hand-written node_modules walk. The documented answer is to restart after such an install.
  • Dropping the shim, or keeping it off agents' PATH. The owner kept the shim; fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 removes the finalizer's own call. docs/security.md says agents can reach it.
  • Deleting the dead writeFakeInstalledEngine(project, …) setup in seven doctor tests. Doctor now reads only Ultrafuzz's install. lifecycle-inspection.test.ts conflicts with refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197 and test(runtime): delete vacuous and source-text tests, keep behavioural coverage #1203, so this waits until after the merge train.
  • A Windows .cmd shim.

Verification

Review-fix head (a811973, on main a46a496)

Load 4–7. Every result below ran on this head's code, except for three later edits. One was a test-only lint fix, assert.ok(run.value) in place of a non-null assertion, and the launch test was re-run after it. The other two were a code comment and changelog wording. The static checks were re-run after all three.

Suite Result
smithers-executable-capability.test.ts and trusted-cli.test.ts 36/36 (8 s), including the new planted-config test (1.5 s)
lifecycle-inspection.test.ts, diagnoseProject|installation inspection|diagnoseRun 21/21 (2 min 3 s)
runtime.test.ts with ^startRun|operator controller locks|usage compatibility|resume|replay|fork|continuation|refresh|compatibility patch|patch anchors|envelope contracts|manifest parsing|package-manager-owned 99 tests: 86 pass, 12 skipped (Bun-only contracts under Node), 1 fail (10 min 29 s). See the note below the table.
All 8 real-Smithers integration files 17/17 (36 s)
CLI doctor reports install posture in human and JSON output 1/1 (27 s)
Bun adapter contracts (bun test … '^Bun adapter contract:', umask 022) 51 pass, 1 skip, 0 fail (78 s)
  • The one runtime failure is this host's umask. The failing test is startRun injects the configured Forge guard into the workflow environment and metadata. With umask 0002 the guard's safe-bin is created group-writable, and isPreparedForgeGuardBin refuses a group-writable directory, so the guard's forge never reaches PATH. With umask 022 the test passes; I re-ran it together with the launch test (2/2). A reviewer saw it fail the same way on main.
  • Earlier Bun failure explained. The umask also explains the Bun contract failure recorded below for 42a1bc8 (provider-home ancestors cannot be group/world writable). Under umask 022 that test passes.

Mutation checks, run on the compiled test tree, which was recompiled afterwards:

Removed Result
The SMITHERS_BACKEND pin in smithersCommandEnv the planted-config test fails at its ps assertion
The shim's export SMITHERS_BACKEND=sqlite the same test fails at its shim assertion
assertExecutableOutsideRoot in the operator binding the target-contained test fails (Missing expected exception)
The snapshot walker's smthrs skip the launch test fails on the new issuer assertion, after 36 s instead of 7 s

Launch against an unpatched installed runner (a manual check; a test process cannot swap Ultrafuzz's own install)

  • Setup: I reverted the installed smithers.js to upstream's delegation line, then restored it byte-identically and confirmed with sha256 and smithers-patches.mjs --check.
  • With the reverted runner, startRun returned RUN_PREFLIGHT_FAILED in 0.9 s: installed workflow runner … lacks 1 of 71 Ultrafuzz compatibility patches (ultrafuzz doctor lists them); reinstall and rebuild Ultrafuzz in its repository checkout: …. It left no run directory, and the fresh TMPDIR held no ultrafuzz-controller-* root.
  • Control, with the patched runner: the preflight passed, and the launch went on to its controller install (an ultrafuzz-controller-* root appeared).
  • On 42a1bc8 the same refusal came only after the controller install. Reviewers measured that install at about 43 s, and the failed run stayed behind.

Planted .smithers/smithers.config.ts

  • On 42a1bc8, both runSmithersInspectionCommand(ps) and <run>/trusted-bin/smithers ps ran it.
  • On this head, both leave no marker. With SMITHERS_BACKEND=sqlite set by hand on the old shim, inspect and node did not import it either.

Static checks, all passing

  • prettier --check on the changed files, and format:check;
  • lint (complexity ceiling 83), and lint:strict:ci against origin/main;
  • knip with every dist/dist-test moved out of the worktree, and again after pnpm -w build;
  • pnpm --filter @ultrafuzz/runtime --filter @ultrafuzz/cli typecheck;
  • docs:check and node scripts/docs-check.mjs;
  • node scripts/smithers-patches.mjs --check;
  • pnpm install --frozen-lockfile --offline, after deleting node_modules/.pnpm-workspace-state-v1.json. Commit 3 changes no manifest, lockfile or patch file.

Not re-run on this head: validate:pack, the CLI e2e lane, bun test scripts/ci/*.test.ts, the full runtime shards, and the full CLI and modal suites. Commit 3 changes no packaging, CI script or e2e code; its validate-packed-install.mjs edit is a comment.

Before the review fixes (42a1bc8)

All on the rebased head (42a1bc8, on main a46a496), at load 3–17 unless noted. Every test ran on this head's code; the only edits made after the test runs are to docs, the changelog, comments and the smithers-patches.mjs failure hint, and the static checks and --check were re-run after them.

Patch files and install

  • main changed no registry entry since this branch's previous base (2cacf4c), and node scripts/smithers-patches.mjs regenerates all six patch files byte-identically.
  • node scripts/smithers-patches.mjs --check is clean.
  • pnpm install --frozen-lockfile --prefer-offline installs the patched packages. pnpm install --frozen-lockfile --offline passes, including after deleting node_modules/.pnpm-workspace-state-v1.json.
  • In the installed tree, all 71 postures (69 patches, 2 engine anchors) classify as applied.

When pnpm deletes the directory a resumed engine runs from (a throwaway worktree of this branch; the docs' dedicated-checkout guidance rests on this)

  • A pnpm install after editing patches/[email protected] adds a new [email protected]_patch_hash=… directory and keeps the old one as an orphan. pnpm skipped that install until node_modules/.pnpm-workspace-state-v1.json was deleted.
  • A second install still keeps the orphan. pnpm deletes orphans only once prunedAt in node_modules/.modules.yaml is older than modules-cache-max-age (7 days by default).
  • With prunedAt set 8 days back, the patch-changing install itself deleted the old directory. So the deletion happens in that install or a later one, not always immediately.

Discriminating test against main

  • native continuation keeps a finished producer and runs only a newly rendered downstream task resumes via resumeRun. Every proxy variable and npm_config_registry point at a local listener that counts and drops each connection, and TMPDIR is fresh.
  • On this branch it passes in 9.3 s:
    • zero listener connections;
    • TMPDIR stays empty;
    • NODE_PATH is the installed runner's dependency directory;
    • PATH starts with <run>/trusted-bin;
    • trusted-bin/smithers inspect answers finished.
  • The same compiled test on main fails after 213 s with resume reached the package registry (6 connections).

First Smithers command of a fresh process

  • Harness: runSmithersInspectionCommand with no runner override, running inspect of a missing run with a fresh TMPDIR, at load 14–15.
Runs TMPDIR during the command
This branch 1.2 s, 1.0 s, 0.9 s empty
main a46a496 42.8 s, 43.8 s one ultrafuzz-controller-* root per run. An inspect deletes it at exit; a resume keeps it.

Suites

Suite Result
All 8 real-Smithers integration files 17/17 (36 s)
trusted-cli.test.ts and smithers-executable-capability.test.ts 35/35 (8 s)
runtime.test.ts, patterns for: compatibility patch, usage and cost compatibility, patched engine/runner admission, patch anchors and envelope contracts, manifest parsing, #1199's fd-transfer and module-confinement tests, native continuation, native resume, controller refresh, trusted CLI, resume/replay/fork delegation, startRun runner install and target-local runners 49 tests: 47 pass, 2 skipped (Bun-only contracts), 0 fail (8 min 14 s)
lifecycle-inspection.test.ts, diagnoseProject|installation inspection|diagnoseRun 21/21 (1 min 59 s)
CLI doctor reports install posture in human and JSON output 1/1 (53 s)
Bun adapter contracts (bun test … '^Bun adapter contract:') 50 pass, 1 skip, 1 fail. The failing test, generated OpenRouter adapter preserves opaque model IDs and enables the authenticated provider catalogue, fails the same way on main a46a496: provider-home ancestors cannot be group/world writable. That is the host's umask 0002; it passes under umask 022 (see the review-fix head above).
bun test scripts/ci/*.test.ts 82/82
validate:pack passes (1 min 10 s). Its doctor now scans the host's own TMPDIR, which holds 142 controller roots.
CLI e2e lane (campaign-resume.test.ts): SIGKILL the engine, let the supervisor relaunch it, SIGKILL the controller, then ultrafuzz resume 1/1 (5 min 52 s). The run is submitted at 146 s, the controller killed at 189 s, resume submitted at 246 s, and the run ended at 263 s. From controller kill to resume submission, which includes waiting for the dead engine's heartbeat lease to lapse, took 57 s here and 167 s in main's CI run of a46a496; the hosts differ, so this is indicative only.
  • The retained Bun-flags test discriminates. With the guard flags dropped from the installed runner's argument prefix, it fails on its argument assertion. With that assertion also removed, it fails on .env loading (dotenv: 'hostile').
  • Static checks: format:check, prettier --check on the changed files, lint (under main's complexity ceiling of 83), lint:strict:ci against origin/main, knip before and after pnpm -w build (now also checking unused exports, types and duplicates), docs:check, pnpm --filter @ultrafuzz/runtime --filter @ultrafuzz/cli typecheck, and the runtime and CLI test compiles all pass.

security:dependency-advisories fails, on main a46a496 as well

  • Both fail with the same error, approved registry advisory brace-expansion[2] duplicates GHSA-6j4f-fj2g-mc7p. The registry now returns that advisory twice, once per vulnerable range, and the gate's strict parser rejects duplicates. main's own CI run for a46a496 fails this step too.
  • Behind that error, the registry reports the same High advisories for both production trees:
  • This branch adds no advisory. Its larger tree resolves @toon-format/toon to 2.3.1.
  • Fixing the gate is not part of this PR.

Not run: the full runtime shards and the full CLI and modal suites.

Earlier results that still apply (measured before this rebase, on origin/integration/wave1 a48ad9c):

  • The a real Smithers {fallback,quota} report retry survives producer and verifier restarts tests give the engine a PATH whose only smithers is the shim. Both pass on this branch (in the integration run above). On the base, both failed with Executable not found in $PATH: "smithers" / "Smithers report-producer authority is unavailable" (Pin required executables and runtime dependencies for detached runs #1143), because the base had no shim writer.
  • Walker exclusions: with smthrs moved to dependencies but the exclusion removed, the trusted-CLI closure lists smthrs and the snapshot seals @smthrs/engine.
  • validate:pack negative control: a consumer without the patch files fails as intended, with doctor reporting workflow-engine-patches: error … lacks compatibility patches.
  • security:dependency-advisories failed on toon 2.3.0 before the override and passed after it.

Rebase notes

Rebased from 2cacf4c onto main a46a496, which adds #1202–#1215, #1219–#1222 and #1228. Commit 1 applied cleanly. Commit 2 conflicted in two files:

Merged without a textual conflict and checked by hand:

Changes made while promoting:

  • Added the dedicated-checkout guidance to docs/reference/cli.md and the changelog entry. The guidance states when pnpm deletes the old engine directory as measured above; the earlier claim that any patch-changing install deletes it at once was only true when the last prune was over seven days old.
  • Deleted the duplicated Bun-flags test (see Commit 2, Bun flags).
  • Dropped the empty TMPDIR that validate:pack gave doctor. Its only purpose was to keep doctor from sizing a host's leftover controller roots, which stalled one run for over 18 minutes. refactor(runtime): give resume, replay and fork explicit result types #1208 now stops that sizing after about one second.
  • Updated the CLI e2e test's header comment, which said resume installs the engine from npm.
  • Fixed the regeneration recipe in scripts/smithers-patches.mjs (its header and its --check failure hint). After an earlier install, pnpm 11's optimistic repeat-install check ignores patch-file contents: a plain pnpm install printed Already up to date and left the lockfile's patch hash stale, and a frozen install then failed with ERR_PNPM_LOCKFILE_CONFIG_MISMATCH. With --config.optimistic-repeat-install=false, the install refreshed the hash (tested in a throwaway worktree).
  • Dropped the "found while verifying" risk about launch failing with … changed while reading when pnpm hardlinks store files: fix(runtime): a pnpm install elsewhere on the host no longer fails a launch #1206 fixed it.

Risk / compatibility

  • Breaking: installs without the patches are refused.
    • Launch, resume, replay and fork refuse an install without the Smithers patches, because each writes the shim through the runner guard. That includes a packed or plain npm install without patchedDependencies. validate:pack shows the supported packed path.
    • They also refuse an installed engine inside the target project, such as an Ultrafuzz checkout under --project, as they already refused an explicit SMITHERS_BIN there. Before this PR that layout worked, because the engine was installed in TMPDIR.
    • These four commands also need bun on PATH to write the shim, even when SMITHERS_BIN points elsewhere.
    • Launch refuses all of these in its preflight, before it creates the run directory. doctor reports the patch and layout problems too.
  • A run changes dependency trees across a resume. It launches on the npm-cutoff tree and resumes on the pnpm-lock tree. The design analysis counted 76 differing package versions; toon 2.3.0 → 2.3.1 is one more, and the count was not redone for this rebase. The Smithers packages, patches and DB schema are the same.
  • The resumed engine runs from the operator's checkout.
    • Any pnpm install there that changes the engine's resolved dependency tree moves the engine: a Smithers patch, or a version change anywhere in its dependency graph. pnpm deletes the old directory in that install or a later one (see Verification).
    • After that, a live resumed engine can fail on its next lazy import, its supervisor cannot relaunch it, and every run's trusted-bin/smithers fails.
    • The docs say to run long campaigns from a dedicated checkout or worktree and to leave its install and build alone.
    • Recovery is ultrafuzz resume, which also rewrites the shim.
    • Before this PR, tmp cleaners could remove the TMPDIR root instead (Pin required executables and runtime dependencies for detached runs #1143).
  • Long-lived processes cache the engine path.
    • eval run calls startRun in-process for every row, and the dashboard also stays up; both resolve the engine once.
    • After such an install they keep binding the old directory until pnpm prunes it, so runs they launch in that window get shims into it.
    • Once the directory is gone, their engine commands fail with … restart this Ultrafuzz process. The docs say to restart them after such an install.
  • Older runs. Runs resumed by an older release keep running from their retained TMPDIR roots. This release creates no new ones and never deletes old ones; doctor still reports them.
  • Security:
    • The installed engine is not sealed. The commands that run it no longer check a controller-closure seal. Its runner and interpreter files are still digest-anchored for every command, its package closure is never hashed, and its patch posture is checked once per process. Launch still seals its own npm controller into the run's snapshot.
    • Commands that lose Bun module confinement. These ran the TMPDIR controller under confinement on main and now run the installed runner without it:
      • ps and the dashboard's listing;
      • the --refresh-controller ownership inspection;
      • commands on runs without a sealed snapshot.
    • Native continuation already ran unconfined. Two guards remain: BUN_TARGET_CONFIGURATION_GUARD_ARGS, so Bun loads no target bunfig.toml or .env, and the SMITHERS_BACKEND=sqlite pin, so Smithers imports no target smithers.config.ts.
    • The finalizer's shim call has the provider credentials. The shim runs the finalizer's smithers node call with the engine's full provider-credential environment, until fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 removes that call.
    • Agents can reach the shim. Agents inherit the engine PATH (trusted-bin first), so besides the validator launcher they can run smithers through the shim against their own run (ps, cancel, signal, …). Same-UID agents could already run the runner by absolute path. docs/security.md says so.
  • The launch controller's npm tree still resolves toon 2.3.0. The production advisory gate never audits that tree; the plan's PR-8 deletes it.
  • Merge coordination:

Changelog entry

Added to CHANGELOG.md under ## Unreleased → Breaking changes:

[runtime] [cli] [ci] [docs] resume, and every other command that installed its own copy of the workflow engine (the pre-resume inspect, ps, the --refresh-controller ownership inspection, and commands on a run without a sealed runner), now runs the engine from Ultrafuzz's pnpm install. pnpm applies the Smithers compatibility patches, committed under patches/ as patchedDependencies, at install time, so these commands make no registry install and leave no ultrafuzz-controller-* directory in TMPDIR; on the test host the first engine command of a resume took about 1 s instead of about 43 s. That engine is not sealed: its package closure is not hashed per command, and it runs without the Bun module confinement the TMPDIR controller gave every command except a native resume. Every engine process now gets SMITHERS_BACKEND=sqlite, so none imports the target's .smithers/smithers.config.ts to choose a store. A pnpm checkout with those patches is now the only supported install: launch, resume, replay and fork refuse an install without them, such as a packed or plain npm install, and an engine inside the target project, telling the operator to run pnpm install --frozen-lockfile && pnpm -w build in the Ultrafuzz repository checkout; launch refuses before it creates the run directory. They also need bun on PATH to write <run>/trusted-bin/smithers, which gives the workflow's own smithers calls a runner (#1143). doctor reports that engine's layout and patch posture as ok or error instead of unknown. smthrs is now a production dependency of @ultrafuzz/runtime, with @toon-format/toon pinned to 2.3.1 under @smthrs/cli (GHSA-p95v-992w-h6c3). Run long campaigns from a dedicated checkout or worktree and leave its install and build alone while they run: a pnpm install there that changes the engine's resolved dependency tree (a Smithers patch, or a version change anywhere in its dependency graph, such as TypeScript or React) moves the engine, and pnpm deletes the old directory in that install or a later one. From then on the resumed engine can fail, its supervisor cannot relaunch it, and every run's trusted-bin/smithers fails, until ultrafuzz resume continues the run from the new install; restart long-lived processes such as the dashboard or eval run after such an install. CI fails when a patch file drifts from the registry; regenerate them with pnpm -w build && node scripts/smithers-patches.mjs && pnpm install --config.optimistic-repeat-install=false (a plain pnpm install after an earlier install leaves the lockfile's patch hashes stale).

Refs #921, #1143, #1147

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR appears safe to merge, with non-blocking gaps in doctor’s diagnosis and shim invocation checks.

Fix All in Claude CodeFindings

  1. P2 Doctor misses target-contained runner ▶
  2. P2 Shim skips runner revalidation ▶
Fix with agent prompt
### Issue 1
packages/runtime/src/doctor.ts:69
When Ultrafuzz’s checkout is inside the target project, doctor reports the installed engine as healthy, but launch and resume reject it because the runner is inside the target. This leaves operators without a diagnosis of why those commands fail. Have doctor check the runner against the project root too.

### Issue 2
packages/runtime/src/smithers.ts:6332
The shim records the installed runner’s pathname when it is written, then executes that pathname without checking it again. If the runner’s bytes change later, a workflow’s `smithers node` call runs the changed bytes, unlike direct Ultrafuzz commands, which recheck the executable before each invocation. Revalidate the runner when the shim runs, or avoid describing shim calls as checked on every command.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

This PR moves post-launch Smithers commands to Ultrafuzz’s pnpm-patched installation while launch continues to seal a separate controller. It adds an installed-runner guard, a per-run smithers shim, SQLite backend pinning, generated pnpm patches, and install guidance. Doctor does not yet diagnose one runner location that commands refuse, and shim invocations do not repeat direct commands’ executable checks.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Launch] --> B[Sealed controller snapshot]
  C[Resume and inspection] --> D[Installed patched runner]
  B --> E[Workflow]
  D --> E
  E --> F[trusted-bin/smithers shim]
  F --> D
Loading

Reviews (1) · Last reviewed commit: "fix(runtime): pin the SQLite store for e..."

aviggiano and others added 2 commits September 30, 2026 08:18
…edDependencies

The repository install now carries every Smithers compatibility patch: pnpm
applies patches/[email protected] and
patches/@smthrs__{agents,cli,db,engine,scheduler}@0.35.0.patch at install
time. scripts/smithers-patches.mjs generates those files: it applies
SMITHERS_COMPATIBILITY_PATCHES to pristine copies of the pinned packages with
applySmithersCompatibilityPatches, the function launch uses, and diffs them
the way `pnpm patch-commit` does. A new CI build-gate step runs it with
--check and fails when a committed patch file or its pnpm-workspace.yaml entry
no longer matches the registry. The registry stays the source of truth, and
launch still installs and patches its own TMPDIR controller, so production
behaviour is unchanged by this commit.

Regenerate the files and the lockfile's patch hashes with:

  pnpm -w build && node scripts/smithers-patches.mjs &&
    pnpm install --config.optimistic-repeat-install=false

pnpm's repeat-install check ignores patch-file contents, so after an earlier
install a plain `pnpm install` prints "Already up to date", leaves the hashes
stale, and CI's frozen install then fails with
ERR_PNPM_LOCKFILE_CONFIG_MISMATCH.

Tests that read the pinned runner now see the bytes production runs:
- the upstream-shape tests (patch anchors, envelope contract, delegation
  refusal) undo the registry's replacements, checking that each is installed
  exactly once and that re-applying the registry reproduces the installed file;
- the usage, cost and path-acceptance tests use the installed patched modules
  instead of patching a copy;
- the controller-engine, resume-reopen and supervisor-relaunch integration
  tests drop their own load-time patch plugin and patched store copies, and
  the shared patched-smithers-runner test helper goes (the relaunch test's
  leftover-process cleanup matches its run ID instead of the copy's path);
- the Bun adapter contracts' OpenRouter helper checks the installed
  ordered-stdout patch instead of applying it, and the DeepSeek usage
  contract asserts what the patched agents report: inputTokens counts cache
  reads and writes (527 = 120 fresh + 400 read + 7 write), with each component
  in inputTokenDetails.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ling one

Every Smithers command whose environment carries no sealed runner used to
npm-install the pinned engine closure into mkdtemp(os.tmpdir()) and patch it
at runtime: native resume and its pre-resume inspect, the
--refresh-controller ownership inspection, and commands on runs without a
sealed snapshot. Resume marked that root retained for its detached engine, so
each resume left one behind in TMPDIR and depended on the registry.

Those commands now bind installedWorkflowRunner(): the smthrs package
@ultrafuzz/runtime resolves, which pnpm has patched at install time. It is
refused, with "reinstall Ultrafuzz with pnpm install --frozen-lockfile",
unless every registered compatibility patch and engine anchor classifies as
applied. Like the unsealed native continuation it replaces, its command
process runs with BUN_TARGET_CONFIGURATION_GUARD_ARGS (--config=/dev/null
--no-env-file --no-install --no-addons), and a continued target workflow
resolves its bare imports from the runner's own dependency directory
(NODE_PATH).

- Delete the native-continuation marker, the retention of its TMPDIR root,
  the per-command controller-closure seal and the TMPDIR controller's Bun
  interpreter assertion.
- Launch, resume, replay and fork write <run>/trusted-bin/smithers, a shim
  that runs the installed runner with the same flags, so the generated
  workflow's bare `smithers` calls resolve (#1143 point 1).
- smthrs becomes a runtime dependency. The production advisory gate now
  audits the engine tree and found GHSA-p95v-992w-h6c3 in @toon-format/toon
  2.3.0, which @smthrs/cli pins exactly; a scoped override takes 2.3.1.
  The execution-snapshot and trusted-CLI closure walkers skip a first-party
  package's smthrs edge, so launch does not copy a second engine closure
  into every run (both walkers go away in later #921 steps).
- validate:pack copies the patch files, root overrides and build allowlist
  into the consumer and requires doctor to report both engine checks ok.
- doctor reports the installed runner's layout and patch posture as ok/error
  checks instead of an ignored project-local tree.
- docs/reference/cli.md tells operators to run long campaigns from a
  dedicated checkout or worktree. The resumed engine and its supervisor run
  from that install, and after a pnpm install that changes a Smithers patch,
  pnpm deletes the package directory the supervisor relaunches the engine
  from, in that install or a later one (by default it clears orphaned package
  directories once seven days have passed since it last did). `ultrafuzz
  resume` continues the run from the new install. CHANGELOG.md records the
  change under Unreleased.

Launch otherwise works as before: it still installs, patches and seals its
own controller into the run's execution snapshot. pause, cancel, status,
replay and fork on a sealed run still execute the sealed runner, which never
installed one.

Tests: the real-engine native-continuation test resumes with every proxy and
npm_config_registry pointed at a listener that counts connections, and
asserts zero of them, an empty TMPDIR, NODE_PATH at the installed runner's
dependency directory and a working trusted-bin/smithers. The report-retry
tests give the engine a PATH with no runner but the shim. Unit tests cover the
guard, doctor and the closure exclusion.

BREAKING CHANGE: a pnpm checkout with the Smithers patchedDependencies is the
only supported install. An install without the patches (a packed or plain npm
install) is refused at launch, resume, replay and fork, which also need bun on
PATH to write the shim.

Refs #921, #1143, #1147

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/v10-install-based-controller branch from 84b7dac to 42a1bc8 Compare September 30, 2026 08:19
aviggiano added a commit that referenced this pull request Sep 30, 2026
…ort record

The verifier's "never recorded" error told operators to run
`resume --refresh-controller --retry-failed`, but `--retry-failed` resets
every failed or stalled task, so in a best-effort campaign it would also
rerun each failed strategy task. Both record errors now name
`resume <run-id> --refresh-controller --reset-node <producer node>`, which
resets the producer's latest attempt and its dependents, here its verifier.
A scratch run against the pinned Smithers confirmed that this timetravel
resets exactly node:report and verify:report and that the re-dispatched
attempt rewrites the record.

A malformed record blocked every re-dispatched attempt above 1 without
naming a way out; its error now names the file to delete. The history test
now asserts both errors' exact text and pins that a first attempt replaces
an unreadable record instead of failing on it.

The docs paragraph states the two cases where failed_attempts is inexact
and that a Codex workspace-write agent can reach the record when the
project lives under /tmp or $TMPDIR. The CHANGELOG entry drops its
self-contradicting PATH clause and names the mixed plain and refreshed
resume failure with its recovery. The integration test's poison-stub
comment is reworded so it stays true after #1201.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
Review found code and wording the earlier cleanup commit missed because
neither TypeScript nor knip flags them:

- `compileTask`'s `env` input: its only reader was the deleted
  `cloudAgentCredentialEnv`. `startRun` still built
  `{ ...process.env, ...input.env }` and passed it through
  `compileSmithersWorkflow`, where nothing read it. The field goes from
  `SmithersCompileInput` and `compileTask`, with both call-site props and
  the `startRun` argument.
- `workflowModuleEntryUrls()` and the empty-string filter on its result
  existed only for the conditional `modal: ""` entry. Both URLs are now
  always present, so the two `import.meta.resolve` calls are inlined.
- The fake runners' `SMITHERS_FAKE_CLOUD_ENV_LOG` hooks, which printed
  MODAL_TOKEN_ID and MODAL_TOKEN_SECRET, and the
  `SMITHERS_FAKE_CLOUD_ENV_LOG`/`SMITHERS_FAKE_CLOUD_SELECTOR_LOG` entries
  in the Smithers test environment allowlist: only the deleted cloud
  credential-forwarding tests set them.
- The Effect-pinning comments in smithers-package.ts, runtime.test.ts and
  the workspace override CI test justified the pin with two cloud
  containers; the reason holds for a run's launch and a later resume,
  which install at different times. The pnpm-workspace.yaml copy is left
  to #1201, which rewrites that block.
- A stateful-profile test named "preserves cloud timeout inheritance" now
  only checks that the resource timeout origin survives TOML snapshots.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…usable installed runners at launch

Review of the installed-runner change found that it let target code run
inside read-only and controller commands. Smithers 0.35.0 imports
<workspace>/.smithers/smithers.config.ts to choose its store backend unless
--backend or SMITHERS_BACKEND is set, and the installed runner has no module
confinement to refuse that import. A planted config therefore ran inside
`ultrafuzz ps`, the dashboard listing, the refresh inspection and the run's
trusted-bin/smithers shim, which the finalizer calls with the engine's
provider credentials. smithersCommandEnv now gives every engine process
SMITHERS_BACKEND=sqlite, the shim exports it, and the detached engine,
supervisor and relaunch inherit it. For a sealed runner the pin also removes
a failure: its confinement refused the import and failed the command.

Launch now binds the installed runner in its preflight, so an unpatched
install, a missing Bun or an engine inside the target project fails with
RUN_PREFLIGHT_FAILED before the run directory, the controller install or any
registry access. The installed runner is refused inside the target project,
as an explicit SMITHERS_BIN already was. The refusal gives a count instead of
2.7 KB of patch ids and says to reinstall and rebuild in the repository
checkout, because a stale build is refused the same way; doctor shares that
hint and the unapplied-patch filter. A long-lived process whose cached runner
directory a pnpm install deleted is told to restart: re-resolving cannot
follow the move, because Node and Bun cache import.meta.resolve.

Also deletes launch-controller code only the removed per-command path used
(the TMPDIR controller's Bun startup controls and their seal entries,
writeCurrentBunStartupControls, a never-passed control parameter and its
signal plumbing) and a dead ULTRAFUZZ_TRUSTED_BIN store in resume. The launch
test now asserts that the snapshot never follows @ultrafuzz/runtime's smthrs
edge. The dedicated-checkout guidance names the real trigger: any install
that changes the engine's resolved dependency tree, not only a patch change,
moves the engine.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
const resolved = await loadResolvedProject({ projectRoot, env });
const references = referencesStatus({ projectRoot });
const installation = inspectSmithersInstallation(projectRoot);
const installation = inspectSmithersInstallation();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Doctor misses target-contained runner When Ultrafuzz’s checkout is inside the target project, doctor reports the installed engine as healthy, but launch and resume reject it because the runner is inside the target. This leaves operators without a diagnosis of why those commands fail. Have doctor check the runner against the project root too.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/doctor.ts
Line: 69

Comment:
**Doctor misses target-contained runner** When Ultrafuzz’s checkout is inside the target project, doctor reports the installed engine as healthy, but launch and resume reject it because the runner is inside the target. This leaves operators without a diagnosis of why those commands fail. Have doctor check the runner against the project root too.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

const trustedBin = ensureSafeDirectory(runRoot, "trusted-bin");
writeFileDurable(
path.join(trustedBin, "smithers"),
`#!/bin/sh\nexport SMITHERS_BACKEND=sqlite\nexec ${[capability.interpreter.path, ...BUN_TARGET_CONFIGURATION_GUARD_ARGS, capability.runner.path].map(quote).join(" ")} "$@"\n`,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Shim skips runner revalidation The shim records the installed runner’s pathname when it is written, then executes that pathname without checking it again. If the runner’s bytes change later, a workflow’s smithers node call runs the changed bytes, unlike direct Ultrafuzz commands, which recheck the executable before each invocation. Revalidate the runner when the shim runs, or avoid describing shim calls as checked on every command.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/smithers.ts
Line: 6332

Comment:
**Shim skips runner revalidation** The shim records the installed runner’s pathname when it is written, then executes that pathname without checking it again. If the runner’s bytes change later, a workflow’s `smithers node` call runs the changed bytes, unlike direct Ultrafuzz commands, which recheck the executable before each invocation. Revalidate the runner when the shim runs, or avoid describing shim calls as checked on every command.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant