test: model-free end-to-end campaign with controller kill and resume on the real pinned engine - #1187
Merged
Merged
Conversation
Every runtime and CLI test drove a fake `smithers` shell script or ran the engine on a hand-written workflow, so a break between the generated workflow, the pinned Smithers engine, the sealed execution snapshot, and the CLI only surfaced in a real campaign. The new test runs `init`, `run`, `status`, `resume`, `stats`, `report`, and `events` as separate CLI processes against the pinned engine that `run` and `resume` install and start under Bun. A stub `codex` on PATH writes the artifacts each prompt's output contract names, builds the final report from the host-injected authorities, and renders it with the prompt's `ultrafuzz report render` command. The stub holds the second node open while the test SIGKILLs the detached engine and supervisor; once status reports the run orphaned, the test resumes it. It asserts that the run succeeds with a verified report, that only the interrupted node's agent ran twice, that no engine task started again after it finished, and that status and stats describe the same complete run. A todo subtest records a gap it found: stats counts three agent attempts where status and the stub count four, because attempts.jsonl is built from NodeFinished/NodeFailed events and resume cancels the interrupted attempt without one. The file lives in test/e2e/ with its own `test:e2e` script, so the CLI suite glob does not run it twice. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Adds a `cli-e2e` release-validation gate that runs `pnpm --filter @ultrafuzz/cli test:e2e`, and a lane for it that is required on pull requests. The lane has a 60-minute budget and the test itself a 45-minute timeout. It runs beside the runtime lanes rather than inside the push-only CLI lane. The budget test in release-validation-lanes.test.ts keeps its 120-minute expectation for the complete runtime and CLI suites only. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Super-linter lints every changed file in full, so any pull request that adds a CHANGELOG entry fails MD013 on the file's existing entries, which are single lines of up to 1,800 characters. This is the same file-level directive #1179 adds, byte for byte, so the two merge without conflict. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/w22-hermetic-e2e
branch
from
September 29, 2026 02:28
bcceb51 to
58a19e5
Compare
`ultrafuzz status` synchronizes the run before it answers, unless the control evidence has diverged. `events` streams the engine's event log without synchronizing. The test now polls `events` until a terminal run event, calls `status` once, and reuses those events for the no-restart assertion. Locally, resume to run end fell from 165 s to 122 s. The wait for the held node also fails at once if the workflow stops before that node starts, instead of after the 15-minute bound. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/w22-hermetic-e2e
branch
from
September 29, 2026 02:49
58a19e5 to
75cb1fc
Compare
…on interrupt Review fixes for the end-to-end campaign test. - The no-rerun checks now run straight after the events poll, before any command synchronizes the run. Those checks cover the stub start counts, the RunStarted count, the untruncated event page and no NodeStarted after NodeFinished. First the test checks that the engine's first run-terminal event is RunFinished. A resume that re-runs a finished node now fails on the named count. Before, it failed with an opaque ARTIFACT_VERIFICATION_AUTHORITY_INVALID from the status call that came first. - The stub now fails, and logs why, in three cases: it finds no output contract, a named authority file is missing, or the report render line is not all `--flag 'value'` pairs. Before, a missed contract match exited 0 with a successful Codex turn and wrote nothing. The test's "workflow stopped" and RunFinished failures include the stub's call log. - SIGINT and SIGTERM handlers, and the exit hook, now SIGKILL the campaign's processes and delete the fixture; the signal handlers then re-raise. A Ctrl-C'd run used to leave the detached engine, the supervisor and the ~1 GB fixture behind. - The wait for the held agent to exit counts a zombie as exited. kill(pid, 0) succeeds on a zombie; its /proc cmdline is empty. - Drops the stats `attempts_complete === true` pin. That field only says the attempt ledger exists, and pinning it would break a fix that reports the undercounted attempts as partial evidence. - The todo reason now names the cause that outlasts #1186. Smithers emits no terminal event for the attempt it abandons at resume. The reason cites #1187. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The CHANGELOG entry said the CLI commands ran on the pinned Smithers engine under Bun. They run as Node CLI processes; only the generated workflow runs on the engine under Bun. The development guide gets the same precise wording, plus one sentence on what the test does when it is interrupted. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in CHANGELOG.md, so each merge would conflict with the next. The entries are collected into one changelog update instead. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
No test ran the generated workflow on the real engine through a report and a resume. The runtime and CLI suites drive a fake
smithersshell script (fakeSmithersEnvinruntime.test.tsandcli.test.ts, the lifecycle-inspection fixtures). The four*.integration.test.tsfiles do run a realsmithers up, but only on hand-written mini workflows. So a break between the generatedworkflow.tsx, the pinned Smithers engine, the sealed execution snapshot and the CLI only showed up in a real campaign. #973 is one example: native continuation could not resolvereact/smthrs/zodoutside the target. #974 fixed it.Root cause
There was no way to run a campaign without a model. The product spawns a real
codex.init's default Codex auth mode needsOPENAI_API_KEY, and according to the code its preflight probesapi.openai.com/v1/models. A private campaign needs disclosure acknowledgements, and the pricing catalog is fetched from models.dev.Change
packages/cli/test/e2e/campaign-resume.test.ts: one test that goes through the shipped product path.ultrafuzz init,run,status,resume,stats,reportandeventseach run as their ownnode packages/cli/dist/index.jsprocess. The generated workflow runs on the pinned Smithers 0.35.0 engine thatrunandresumeinstall into their operator controller, under Bun. The pieces:project-discovery→summarize→final-report, with the realreview/final-report.mdprompt.init'sultrafuzz.toml, with the[agents.CodexAgent]table switched to subscription auth.CODEX_HOME/auth.jsonholds a fake ChatGPT token, so by the adapter code the engine's agent preflight makes no network call.ULTRAFUZZ_DATA_GOVERNANCE_POLICYis a public policy andULTRAFUZZ_PRICING_CATALOG_URL=off.TMPDIRpoints inside the fixture, so the per-command controller installs are deleted with it.codexexecutable onPATH. It writes thetext@1andreport@3outputs that the prompt's output contract names. Forreport@3it copies the two host-injected authority files the prompt names, then runs theultrafuzz report rendercommand the prompt gives, which writesreport.md. It prints Codex JSONL with usage. When the prompt lacks what it parses, it fails and logs why. That covers a missing output contract, a missing authority file, and a render command that is not all--flag 'value'pairs. The test's failure messages include that log. The stub is a typed function in the test, serialized into the binary.summarizeopen. The test then SIGKILLs every process whose argv names the Smithers run ID, which are the detached engine and its supervisor. It waits until they and the held agent are gone; a zombie counts as gone.statusto reportorphaned, then runsultrafuzz resume.finallyblock, anexithook, and SIGINT/SIGTERM handlers all SIGKILL the campaign's processes and delete the fixture. The signal handlers then re-raise the signal. At the two interrupt points I tried (see Verification), a Ctrl-C'd run left no campaign process and no fixture behind. The fixture is about 1 GB.resumereportssubmitted: true.RunFinished.project-discoveryandfinal-reportonce each andsummarizetwice (killed, then resumed).eventsis not truncated and shows twoRunStartedactivations.NodeStartedafter its ownNodeFinished.eventsresult. They run before any command synchronizes the run, so a regression is reported as the invariant it breaks.succeeded, with the reportavailable/complete/verified.reportreturnsverified-runtime-reportfor this run ID.statusprogress isfinished == total, with nothing running, pending, failed or skipped.statsshows exactly the three topology nodes, allsucceeded.statuscounts 4 agent attempts, matching the stub's 4 invocations.packages/cli/package.json:test:e2escript. The file sits intest/e2e/, so theclilane'sdist-test/test/*.test.jsglob does not run it a second time.scripts/validate-release.mjs,scripts/ci/release-validation-lanes.mjs: a newcli-e2egate and lane, required on pull requests (60 min lane budget, 45 min test timeout).scripts/ci/release-validation-lanes.test.ts: expects the new PR lane, and keeps the 120-minute budget check scoped to the complete runtime and CLI suites.docs/reference/development.md,CHANGELOG.md.Product bug this exposes (left as a
todosubtest)statsundercounts attempts after a controller crash:ultrafuzz node <run> node:summarize --attemptsshows attempt 1cancelled(the killed one) and attempt 2finished.statusmodel_mixcounts 4 agent attempts, which matches the stub's 4 invocations.statsreportssummarizewithattempt_count: 1, retry_count: 0, 3 in total.The cause, from reading
terminalWorkflowAttemptsinpackages/runtime/src/workflow-sync.ts: its own comment says Smithers cancels stale in-progress rows before a resumed activation without emittingNodeCancelled. With no terminal event for it, the abandoned attempt, and its time, never reachattempts.jsonl, which is the ledgerstatsreads. A review run confirmed this in the lifecycle log: 10NodeStarted, 9NodeFinished, and noNodeCancelled.The subtest
stats counts the agent attempt the controller crash interruptedasserts the agreement and is markedtodo, so it reports3 !== 4without failing the lane. Its todo string cites this PR. #1186 (w02) reworks this ledger but explicitly leaves crash-abandoned attempts unrecorded until Smithers emits an event for them (an upstreamNodeCancelled{reason:"resumed"}). So no open change fixes this gap. Node does not flag atodosubtest that starts passing, so whoever closes the gap should turn it into a plain assertion.Other things I saw while building this (not asserted)
statusgives the reason... the run is orphaned — resume it with \workflow runner supervise -r ultrafuzz-`. That is Smithers' own remediation with "smithers" rewritten, not a command an operator can type. It should sayultrafuzz resume `.resumefinds the run "running" and returnssubmitted: false, while the text output still printsSubmitted resume: …. The test waits fororphanedfor this reason. A parallel change covers the no-op attach path inresume, so I did not touch it.ultrafuzz runtook 7–14 min.node-pty, two copies each ofeffectandtypescript, and@smthrs/jj-linux-x64.ultrafuzz statustook 50–77 s wall and 48–68 s user CPU. I did not profile where that CPU goes.<target>/.workflow runner/workflows/…, a directory that does not exist; the real path is<target>/.smithers/workflows/….statsafterstatus) reportedNODE_ATTEMPT_LEDGER_WRITE_FAILED: node attempt ["ultrafuzz-e2e-3",29] was already recorded with different immutable data. That is not on this test's path. It is attempt-ledger code, like the undercount above.resumeRunto forceresetNode: "node:project-discovery".--reset-nodeis documented for failed nodes, so this is off-label use. With that change, the engine re-ran the finished node and reachedRunFinished. Every laterultrafuzz statusthen failed withARTIFACT_VERIFICATION_AUTHORITY_INVALIDforproject-discoveryandsummarize. I reproduced the first half below (the rerun andRunFinished). The test now stops beforestatus, so I did not re-check thestatusfailure.Deliberately not built (and why)
todosubtest instead of asserting today's buggy count. I did not file an issue from this PR; the todo cites this description.ULTRAFUZZ_PRICING_CATALOG_URL=offskips models.dev. That comes from the configuration and the code; I did not capture network traffic. Butrunandresumeinstall the pinned engine from registry.npmjs.org, exactly as a campaign does, so the lane needs registry access. The title used to say "hermetic" and now says "model-free".resume --forceto skip the lease wait. The test uses the command an operator would use, afterstatussays the run is orphaned./proc, and the test skips on other platforms.clilane. That lane is push-only and takes 37–76 min, so the test gets its own PR lane instead.attempts_completeassertion. An earlier revision pinnedstats.totals.attempts_complete === true. That field only means the ledger file exists (packages/cli/src/run-statistics.ts), and it sat in the same run as the undercount. A fix that reports the evidence as partial would have broken it, so it is gone.Verification
Review-fix round (head dce193d). The test code is the same as in bf9d36c; the last commit only changes docs.
Local, under eatmydata, load average 5–13:
node --test dist-test/test/e2e/campaign-resume.test.jspassed in 5 min 55 s (1 pass, 0 fail, 1 todo,3 !== 4). The phase timings were: run submitted after 89 s, controller killed after 99 s, resume submitted after 196 s, run ended after 232 s, checks done after 349 s. It left no fixture and no process.Rerun mutant: this is a reviewer's mutant. Behind an env var, the built
resumeRunforcesresetNode: "node:project-discovery", so resume re-runs a finished node.[["project-discovery",2],["summarize",2],["final-report",1]], expectedproject-discovery1.RunFinishedcheck before it passed.status, withARTIFACT_VERIFICATION_AUTHORITY_INVALID(the reviewer's runs).cmp.Interrupt cleanup: I ran
setsid node --test …and sent SIGINT to its process group.project-discovery(therunCLI and two enginebunprocesses alive): this revision exited 8 s later. No process named the fixture or run ID, and the fixture was gone. The previous revision (75cb1fc), interrupted at the same point, exited at once and left bothbunprocesses (engine and supervisor) and the fixture behind. I killed and deleted those by hand.Stub strictness: I ran the compiled stub by itself.
* Path:, the previous stub exited 0 with a successful Codex turn and wrote nothing. This one exits 1 withstub codex found no output contract in the promptand logs afailedcall.stub codex cannot parse the report render command: ….Zombie: for a zombie child (state
Z),process.kill(pid, 0)succeeds and/proc/<pid>/cmdlineis empty. I did not run the test under a PID 1 that never reaps.Gates: all of these pass:
npx prettier --checkandnpx eslinton every changed file;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;pnpm --filter @ultrafuzz/cli typecheckandnpx tsc -p packages/cli/tsconfig.test.json;pnpm -w knipandnode scripts/docs-check.mjs;bun test scripts/ci/release-validation-lanes.test.ts scripts/ci/validate-release.test.ts(12 pass).I did not re-run
pnpm -w test:ci-scripts, because this round did not touchscripts/ci.CI on dce193d (run 36523181500): the
cli-e2elane passed. The test took 677 s (1 pass, 1 todo,3 !== 4) and the job 14 min 11 s. The phase timings were: run submitted after 177 s, controller killed after 201 s, resume submitted after 378 s, run ended after 448 s, checks done after 667 s. "Run ended" now marks the end of theeventspoll. Thestatus,reportandstatschecks after it took 219 s on CI.Earlier revisions, before the review fixes. These runs predate this round, and I did not repeat them:
Local, under eatmydata, on origin/main (b6dd1da):
3 !== 4). No/tmp/ufz-e2e-*directory and no run process was left behind.statusattempt-count check passed in 8 min 44 s at load average about 10.events-polling revision passed in 7 min 54 s.fix(runtime): resolve native continuation dependencies #974 revert: I made
nativeOperatorSmithersNodePathreturnundefinedin the built runtime, which reintroduces Native continuation cannot resolve workflow dependencies outside the target #973. The test then fails atultrafuzz resumewithDETACHED_PREFLIGHT_FAILED: ResolveMessage: Cannot find package 'react' from '<target>/.workflow runner/workflows/ultrafuzz-<run>.tsx': the resumed engine cannot load the persisted workflow. That is the production failure Native continuation cannot resolve workflow dependencies outside the target #973 described, and no fake-smitherstest can see it. The run took 12 min 15 s and left no fixture or process behind. The file was restored afterwards.CI,
cli-e2elane:Each run was 1 pass and 1 todo. That is above the 5–8 min target. Most of it is product cost:
runinstalls the engine and copies the 772 MB snapshot,resumeinstalls the engine again, and everystatus/statscall synchronizes the run.Risk / compatibility
The product is unchanged. The new required lane adds one job per PR, 12.5–14.2 minutes on the four CI runs so far. It runs alongside the 52-minute runtime-supporting lane, so PR wall time should not grow.
The lane depends on registry.npmjs.org, as
pnpm installalready does.The stub, not a real agent, depends on three exact prompt strings. A model reads prose and would cope with a rewording. The three strings are:
- Path:/Contract:);workspace-relative file "…"authority sentences;ultrafuzz report renderline inreview/final-report.md.A change to any of them makes the stub fail with a named error. The test's failure message then includes the stub's call log, which records that error.
Conflicts with ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling #1184 (w21). ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling #1184 deletes
pull_request,PULL_REQUEST_REQUIRED_GATESandselectReleaseValidationLanes, rewrites the lanes test and the samedevelopment.mdparagraph, and runs every lane under eatmydata. Whichever PR merges second should:cli-e2elane object withoutpull_request;release-validation-lanes.test.ts;ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling #1184 also makes the per-file
markdownlint-disable-file MD013directive inCHANGELOG.mdredundant.🤖 Generated with Claude Code
The PR appears safe to merge based on the changes reviewed.
Summary
This PR adds a model-free end-to-end CLI campaign test that runs a generated workflow on the pinned engine, kills its controller mid-node, resumes it, and checks completion and report verification. It adds a dedicated pull-request-required
cli-e2erelease-validation lane and documents how to run the test.Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR CI[PR validation] --> Lane[cli-e2e lane] Lane --> Init[CLI init and run] Init --> Engine[Pinned engine] Engine --> Kill[Kill controller mid-node] Kill --> Resume[CLI resume] Resume --> Verify[Events, status, stats and verified report]Reviews (6) · Last reviewed commit: "chore: move the changelog entry to the c..."