Skip to content

test: model-free end-to-end campaign with controller kill and resume on the real pinned engine - #1187

Merged
aviggiano merged 7 commits into
mainfrom
claude/w22-hermetic-e2e
Sep 29, 2026
Merged

aviggiano merged 7 commits into
mainfrom
claude/w22-hermetic-e2e

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

No test ran the generated workflow on the real engine through a report and a resume. The runtime and CLI suites drive a fake smithers shell script (fakeSmithersEnv in runtime.test.ts and cli.test.ts, the lifecycle-inspection fixtures). The four *.integration.test.ts files do run a real smithers up, but only on hand-written mini workflows. So a break between the generated workflow.tsx, the pinned Smithers engine, the sealed execution snapshot and the CLI only showed up in a real campaign. #973 is one example: native continuation could not resolve react/smthrs/zod outside the target. #974 fixed it.

Root cause

There was no way to run a campaign without a model. The product spawns a real codex. init's default Codex auth mode needs OPENAI_API_KEY, and according to the code its preflight probes api.openai.com/v1/models. A private campaign needs disclosure acknowledgements, and the pricing catalog is fetched from models.dev.

Change

  • packages/cli/test/e2e/campaign-resume.test.ts: one test that goes through the shipped product path. ultrafuzz init, run, status, resume, stats, report and events each run as their own node packages/cli/dist/index.js process. The generated workflow runs on the pinned Smithers 0.35.0 engine that run and resume install into their operator controller, under Bun. The pieces:
    • Target: a one-file git repo; no Foundry.
    • Topology: three agentic nodes, project-discovery → summarize → final-report, with the real review/final-report.md prompt.
    • Config: init's ultrafuzz.toml, with the [agents.CodexAgent] table switched to subscription auth. CODEX_HOME/auth.json holds a fake ChatGPT token, so by the adapter code the engine's agent preflight makes no network call.
    • Environment: ULTRAFUZZ_DATA_GOVERNANCE_POLICY is a public policy and ULTRAFUZZ_PRICING_CATALOG_URL=off. TMPDIR points inside the fixture, so the per-command controller installs are deleted with it.
    • Stub agent: a codex executable on PATH. It writes the text@1 and report@3 outputs that the prompt's output contract names. For report@3 it copies the two host-injected authority files the prompt names, then runs the ultrafuzz report render command the prompt gives, which writes report.md. It prints Codex JSONL with usage. When the prompt lacks what it parses, it fails and logs why. That covers a missing output contract, a missing authority file, and a render command that is not all --flag 'value' pairs. The test's failure messages include that log. The stub is a typed function in the test, serialized into the binary.
    • Kill: the stub holds summarize open. The test then SIGKILLs every process whose argv names the Smithers run ID, which are the detached engine and its supervisor. It waits until they and the held agent are gone; a zombie counts as gone.
    • Resume: the test waits for status to report orphaned, then runs ultrafuzz resume.
    • Cleanup: a finally block, an exit hook, and SIGINT/SIGTERM handlers all SIGKILL the campaign's processes and delete the fixture. The signal handlers then re-raise the signal. At the two interrupt points I tried (see Verification), a Ctrl-C'd run left no campaign process and no fixture behind. The fixture is about 1 GB.
  • It asserts, in this order:
    • resume reports submitted: true.
    • The engine event log's first run-terminal event is RunFinished.
    • The stub ran project-discovery and final-report once each and summarize twice (killed, then resumed).
    • events is not truncated and shows two RunStarted activations.
    • No engine task has a NodeStarted after its own NodeFinished.
    • These no-rerun checks use only the stub's log and the events result. They run before any command synchronizes the run, so a regression is reported as the invariant it breaks.
    • The run ends succeeded, with the report available / complete / verified. report returns verified-runtime-report for this run ID.
    • status progress is finished == total, with nothing running, pending, failed or skipped.
    • stats shows exactly the three topology nodes, all succeeded.
    • status counts 4 agent attempts, matching the stub's 4 invocations.
  • packages/cli/package.json: test:e2e script. The file sits in test/e2e/, so the cli lane's dist-test/test/*.test.js glob does not run it a second time.
  • scripts/validate-release.mjs, scripts/ci/release-validation-lanes.mjs: a new cli-e2e gate and lane, required on pull requests (60 min lane budget, 45 min test timeout).
  • scripts/ci/release-validation-lanes.test.ts: expects the new PR lane, and keeps the 120-minute budget check scoped to the complete runtime and CLI suites.
  • docs/reference/development.md, CHANGELOG.md.

Product bug this exposes (left as a todo subtest)

stats undercounts attempts after a controller crash:

  • ultrafuzz node <run> node:summarize --attempts shows attempt 1 cancelled (the killed one) and attempt 2 finished.
  • status model_mix counts 4 agent attempts, which matches the stub's 4 invocations.
  • stats reports summarize with attempt_count: 1, retry_count: 0, 3 in total.

The cause, from reading terminalWorkflowAttempts in packages/runtime/src/workflow-sync.ts: its own comment says Smithers cancels stale in-progress rows before a resumed activation without emitting NodeCancelled. With no terminal event for it, the abandoned attempt, and its time, never reach attempts.jsonl, which is the ledger stats reads. A review run confirmed this in the lifecycle log: 10 NodeStarted, 9 NodeFinished, and no NodeCancelled.

The subtest stats counts the agent attempt the controller crash interrupted asserts the agreement and is marked todo, so it reports 3 !== 4 without failing the lane. Its todo string cites this PR. #1186 (w02) reworks this ledger but explicitly leaves crash-abandoned attempts unrecorded until Smithers emits an event for them (an upstream NodeCancelled{reason:"resumed"}). So no open change fixes this gap. Node does not flag a todo subtest that starts passing, so whoever closes the gap should turn it into a plain assertion.

Other things I saw while building this (not asserted)

  • For an orphaned run, status gives the reason ... the run is orphaned — resume it with \workflow runner supervise -r ultrafuzz-`. That is Smithers' own remediation with "smithers" rewritten, not a command an operator can type. It should say ultrafuzz resume `.
  • From reading code, not by running it: for about 30 s after a crash, the engine's heartbeat is still fresh. During that window resume finds the run "running" and returns submitted: false, while the text output still prints Submitted resume: …. The test waits for orphaned for this reason. A parallel change covers the no-op attach path in resume, so I did not touch it.
  • Launch cost, measured on this shared 32-core host at load average 40–48:
    • ultrafuzz run took 7–14 min.
    • The sealed execution snapshot it copies is 772 MB in 51,418 files. Nearly all of that is the Smithers dependency tree, including node-pty, two copies each of effect and typescript, and @smthrs/jj-linux-x64.
    • Each ultrafuzz status took 50–77 s wall and 48–68 s user CPU. I did not profile where that CPU goes.
  • Diagnostic sanitizing rewrites paths as well as names. The fix(runtime): resolve native continuation dependencies #974-revert failure below reported <target>/.workflow runner/workflows/…, a directory that does not exist; the real path is <target>/.smithers/workflows/….
  • In one exploratory run where the stub wrote no artifacts (a failing verify), the second synchronization (stats after status) reported NODE_ATTEMPT_LEDGER_WRITE_FAILED: node attempt ["ultrafuzz-e2e-3",29] was already recorded with different immutable data. That is not on this test's path. It is attempt-ledger code, like the undercount above.
  • One reviewer changed the built resumeRun to force resetNode: "node:project-discovery". --reset-node is documented for failed nodes, so this is off-label use. With that change, the engine re-ran the finished node and reached RunFinished. Every later ultrafuzz status then failed with ARTIFACT_VERIFICATION_AUTHORITY_INVALID for project-discovery and summarize. I reproduced the first half below (the rerun and RunFinished). The test now stops before status, so I did not re-check the status failure.

Deliberately not built (and why)

  • No fix for the undercount. fix(runtime): attempt and usage ledgers never block run synchronization #1186 owns the attempt ledger and defers this gap to upstream Smithers (see above). The test records the undercount as a todo subtest instead of asserting today's buggy count. I did not file an issue from this PR; the todo cites this description.
  • No npm mirror or cache, so not hermetic. The test is model-free and needs no real credential. Subscription auth with a fake token keeps the agent preflight offline, and ULTRAFUZZ_PRICING_CATALOG_URL=off skips models.dev. That comes from the configuration and the code; I did not capture network traffic. But run and resume install the pinned engine from registry.npmjs.org, exactly as a campaign does, so the lane needs registry access. The title used to say "hermetic" and now says "model-free".
  • No eatmydata or fsync bypass. The lane pays the real launch cost. Nothing asserts on that cost; the phase timings are diagnostics only. A parallel change is reducing that cost.
  • No resume --force to skip the lease wait. The test uses the command an operator would use, after status says the run is orphaned.
  • No per-command timing assertions, and no macOS path: the kill step reads /proc, and the test skips on other platforms.
  • Not added to the cli lane. That lane is push-only and takes 37–76 min, so the test gets its own PR lane instead.
  • No attempts_complete assertion. An earlier revision pinned stats.totals.attempts_complete === true. That field only means the ledger file exists (packages/cli/src/run-statistics.ts), and it sat in the same run as the undercount. A fix that reports the evidence as partial would have broken it, so it is gone.

Verification

Review-fix round (head dce193d). The test code is the same as in bf9d36c; the last commit only changes docs.

  • Local, under eatmydata, load average 5–13: node --test dist-test/test/e2e/campaign-resume.test.js passed in 5 min 55 s (1 pass, 0 fail, 1 todo, 3 !== 4). The phase timings were: run submitted after 89 s, controller killed after 99 s, resume submitted after 196 s, run ended after 232 s, checks done after 349 s. It left no fixture and no process.

  • Rerun mutant: this is a reviewer's mutant. Behind an env var, the built resumeRun forces resetNode: "node:project-discovery", so resume re-runs a finished node.

    • The test now fails after 4 min 05 s, at the stub start-count assertion: actual [["project-discovery",2],["summarize",2],["final-report",1]], expected project-discovery 1.
    • The RunFinished check before it passed.
    • Before this change, the same mutant failed after 967–988 s at status, with ARTIFACT_VERIFICATION_AUTHORITY_INVALID (the reviewer's runs).
    • I restored the built file afterwards and checked it byte for byte with cmp.
  • Interrupt cleanup: I ran setsid node --test … and sent SIGINT to its process group.

    • While the detached campaign was running project-discovery (the run CLI and two engine bun processes alive): this revision exited 8 s later. No process named the fixture or run ID, and the fixture was gone. The previous revision (75cb1fc), interrupted at the same point, exited at once and left both bun processes (engine and supervisor) and the fixture behind. I killed and deleted those by hand.
    • Right after the test's own kill step (two runs): no process was left, and the fixture was deleted.
    • I did not try other interrupt points, or processes that name neither the fixture nor the run ID.
  • Stub strictness: I ran the compiled stub by itself.

    • With a prompt using * Path:, the previous stub exited 0 with a successful Codex turn and wrote nothing. This one exits 1 with stub codex found no output contract in the prompt and logs a failed call.
    • A render line with a double-quoted value fails with stub codex cannot parse the report render command: ….
  • Zombie: for a zombie child (state Z), process.kill(pid, 0) succeeds and /proc/<pid>/cmdline is empty. I did not run the test under a PID 1 that never reaps.

  • Gates: all of these pass:

    • npx prettier --check and npx eslint on every changed file;
    • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci;
    • pnpm --filter @ultrafuzz/cli typecheck and npx tsc -p packages/cli/tsconfig.test.json;
    • pnpm -w knip and node scripts/docs-check.mjs;
    • bun test scripts/ci/release-validation-lanes.test.ts scripts/ci/validate-release.test.ts (12 pass).

    I did not re-run pnpm -w test:ci-scripts, because this round did not touch scripts/ci.

  • CI on dce193d (run 36523181500): the cli-e2e lane passed. The test took 677 s (1 pass, 1 todo, 3 !== 4) and the job 14 min 11 s. The phase timings were: run submitted after 177 s, controller killed after 201 s, resume submitted after 378 s, run ended after 448 s, checks done after 667 s. "Run ended" now marks the end of the events poll. The status, report and stats checks after it took 219 s on CI.

Earlier revisions, before the review fixes. These runs predate this round, and I did not repeat them:

  • Local, under eatmydata, on origin/main (b6dd1da):

    • The first two runs passed in 27 min 07 s and 15 min 48 s on the loaded shared host (1 pass, 0 fail, 1 todo, 3 !== 4). No /tmp/ufz-e2e-* directory and no run process was left behind.
    • The revision that added phase timings and the hard status attempt-count check passed in 8 min 44 s at load average about 10.
    • The events-polling revision passed in 7 min 54 s.
  • fix(runtime): resolve native continuation dependencies #974 revert: I made nativeOperatorSmithersNodePath return undefined in the built runtime, which reintroduces Native continuation cannot resolve workflow dependencies outside the target #973. The test then fails at ultrafuzz resume with DETACHED_PREFLIGHT_FAILED: ResolveMessage: Cannot find package 'react' from '<target>/.workflow runner/workflows/ultrafuzz-<run>.tsx': the resumed engine cannot load the persisted workflow. That is the production failure Native continuation cannot resolve workflow dependencies outside the target #973 described, and no fake-smithers test can see it. The run took 12 min 15 s and left no fixture or process behind. The file was restored afterwards.

  • CI, cli-e2e lane:

    Run Test Job
    36510198801 610 s 13 min 43 s
    36512728303 654 s 13 min 40 s
    36514402758 (head 75cb1fc) 524 s 12 min 33 s
    36523181500 (head dce193d, this round) 677 s 14 min 11 s

    Each run was 1 pass and 1 todo. That is above the 5–8 min target. Most of it is product cost: run installs the engine and copies the 772 MB snapshot, resume installs the engine again, and every status/stats call synchronizes the run.

Risk / compatibility

  • The product is unchanged. The new required lane adds one job per PR, 12.5–14.2 minutes on the four CI runs so far. It runs alongside the 52-minute runtime-supporting lane, so PR wall time should not grow.

  • The lane depends on registry.npmjs.org, as pnpm install already does.

  • The stub, not a real agent, depends on three exact prompt strings. A model reads prose and would cope with a rewording. The three strings are:

    • the output-contract lines (- Path: / Contract:);
    • the injected workspace-relative file "…" authority sentences;
    • the ultrafuzz report render line in review/final-report.md.

    A change to any of them makes the stub fail with a named error. The test's failure message then includes the stub's call log, which records that error.

  • Conflicts with ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling #1184 (w21). ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling #1184 deletes pull_request, PULL_REQUEST_REQUIRED_GATES and selectReleaseValidationLanes, rewrites the lanes test and the same development.md paragraph, and runs every lane under eatmydata. Whichever PR merges second should:

    • keep the cli-e2e lane object without pull_request;
    • drop this PR's hunks in release-validation-lanes.test.ts;
    • say nine lanes, at most eight in parallel, in the merged docs.

    ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling #1184 also makes the per-file markdownlint-disable-file MD013 directive in CHANGELOG.md redundant.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the changes reviewed.

Summary

This PR adds a model-free end-to-end CLI campaign test that runs a generated workflow on the pinned engine, kills its controller mid-node, resumes it, and checks completion and report verification. It adds a dedicated pull-request-required cli-e2e release-validation lane and documents how to run the test.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  CI[PR validation] --> Lane[cli-e2e lane]
  Lane --> Init[CLI init and run]
  Init --> Engine[Pinned engine]
  Engine --> Kill[Kill controller mid-node]
  Kill --> Resume[CLI resume]
  Resume --> Verify[Events, status, stats and verified report]
Loading

Reviews (6) · Last reviewed commit: "chore: move the changelog entry to the c..."

aviggiano and others added 2 commits September 29, 2026 01:42
Every runtime and CLI test drove a fake `smithers` shell script or ran the
engine on a hand-written workflow, so a break between the generated
workflow, the pinned Smithers engine, the sealed execution snapshot, and the
CLI only surfaced in a real campaign.

The new test runs `init`, `run`, `status`, `resume`, `stats`, `report`, and
`events` as separate CLI processes against the pinned engine that `run` and
`resume` install and start under Bun. A stub `codex` on PATH writes the
artifacts each prompt's output contract names, builds the final report from
the host-injected authorities, and renders it with the prompt's `ultrafuzz
report render` command. The stub holds the second node open while the test
SIGKILLs the detached engine and supervisor; once status reports the run
orphaned, the test resumes it. It asserts that the run succeeds with a
verified report, that only the interrupted node's agent ran twice, that no
engine task started again after it finished, and that status and stats
describe the same complete run.

A todo subtest records a gap it found: stats counts three agent attempts
where status and the stub count four, because attempts.jsonl is built from
NodeFinished/NodeFailed events and resume cancels the interrupted attempt
without one.

The file lives in test/e2e/ with its own `test:e2e` script, so the CLI suite
glob does not run it twice.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Adds a `cli-e2e` release-validation gate that runs
`pnpm --filter @ultrafuzz/cli test:e2e`, and a lane for it that is
required on pull requests. The lane has a 60-minute budget and the test
itself a 45-minute timeout. It runs beside the runtime lanes rather than
inside the push-only CLI lane.

The budget test in release-validation-lanes.test.ts keeps its 120-minute
expectation for the complete runtime and CLI suites only.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 01:43
Super-linter lints every changed file in full, so any pull request that
adds a CHANGELOG entry fails MD013 on the file's existing entries, which
are single lines of up to 1,800 characters. This is the same file-level
directive #1179 adds, byte for byte, so the two merge without conflict.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/w22-hermetic-e2e branch from bcceb51 to 58a19e5 Compare September 29, 2026 02:28
`ultrafuzz status` synchronizes the run before it answers, unless the
control evidence has diverged. `events` streams the engine's event log
without synchronizing. The test now polls `events` until a terminal run
event, calls `status` once, and reuses those events for the no-restart
assertion. Locally, resume to run end fell from 165 s to 122 s.

The wait for the held node also fails at once if the workflow stops
before that node starts, instead of after the 15-minute bound.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 2 commits September 29, 2026 04:46
…on interrupt

Review fixes for the end-to-end campaign test.

- The no-rerun checks now run straight after the events poll, before any
  command synchronizes the run. Those checks cover the stub start counts,
  the RunStarted count, the untruncated event page and no NodeStarted after
  NodeFinished. First the test checks that the engine's first run-terminal
  event is RunFinished. A resume that re-runs a finished node now fails on
  the named count. Before, it failed with an opaque
  ARTIFACT_VERIFICATION_AUTHORITY_INVALID from the status call that came
  first.
- The stub now fails, and logs why, in three cases: it finds no output
  contract, a named authority file is missing, or the report render line
  is not all `--flag 'value'` pairs. Before, a missed contract match
  exited 0 with a successful Codex turn and wrote nothing. The test's
  "workflow stopped" and RunFinished failures include the stub's call log.
- SIGINT and SIGTERM handlers, and the exit hook, now SIGKILL the
  campaign's processes and delete the fixture; the signal handlers then
  re-raise. A Ctrl-C'd run used to leave the detached engine, the
  supervisor and the ~1 GB fixture behind.
- The wait for the held agent to exit counts a zombie as exited. kill(pid,
  0) succeeds on a zombie; its /proc cmdline is empty.
- Drops the stats `attempts_complete === true` pin. That field only
  says the attempt ledger exists, and pinning it would break a fix that
  reports the undercounted attempts as partial evidence.
- The todo reason now names the cause that outlasts #1186. Smithers
  emits no terminal event for the attempt it abandons at resume. The
  reason cites #1187.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The CHANGELOG entry said the CLI commands ran on the pinned Smithers
engine under Bun. They run as Node CLI processes; only the generated
workflow runs on the engine under Bun. The development guide gets the
same precise wording, plus one sentence on what the test does when it
is interrupted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano aviggiano changed the title test: hermetic end-to-end campaign with controller kill and resume on the real pinned engine test: model-free end-to-end campaign with controller kill and resume on the real pinned engine Sep 29, 2026
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 629118f into main Sep 29, 2026
14 checks passed
@aviggiano
aviggiano deleted the claude/w22-hermetic-e2e branch September 29, 2026 06:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant