Skip to content

test(runtime): delete vacuous and source-text tests, keep behavioural coverage - #1203

Merged
aviggiano merged 4 commits into
mainfrom
claude/v09a-runtime-test-quality
Sep 29, 2026
Merged

aviggiano merged 4 commits into
mainfrom
claude/v09a-runtime-test-quality

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #1000

Problem

Two kinds of runtime test cost buy no protection.

  • Tests that do not test behaviour. A large part of generated-workflow-verifier.test.ts reads workflow.tsx as text and matches regexes, indexOf positions or identifier names. These tests fail when a function moves or is reformatted, and they pass when the code they point at is broken. A few other assertions compare a value with its own definition, repeat the assertion before them, or sit in a branch that never runs.
  • Launches nothing reads. lifecycle-inspection.test.ts launched a full run (startRun, which seals a complete execution snapshot) for 38 of its 52 tests. 29 of them only issue read-only inspection commands against the fake runner, and the 5 diagnoseProject tests never read the run at all. The file runs in the runtime-supporting release lane, whose comment says runner contention can push that lane past 75 minutes.

Root cause

  • The generated workflow is a template, not an importable module, so tests pinned its source text instead of running it. Where a harness does run template functions, several of its stubs silently dropped the behaviour a source-text test was standing in for. For example, clearArtifactVerificationMarker and verifyGeneratedTestFiles were no-ops in the verifyArtifacts harness, and the review-context stub ignored directOnly.
  • lifecycle-inspection.test.ts baked each test's fixture answers into the fake runner script it wrote before launching, so every test needed its own launched run.

Change

Test-only; no file under src/ changes. Four commits: the three below, and a review follow-up (section 4) that restores the guards review found lost.

1. lifecycle-inspection.test.ts: one shared run for read-only commands

  • The fake runner now reads every answer (why, timeline, snapshots, node, events, cancel) from files. writeInspectionFixtures rewrites those files before each test, and clears the control files and the command log.
  • The 29 read-only tests share one launched run through inspectedRun. That includes "diagnoseRun keeps a runner path intact in public text", which fix: observers and benchmark harnesses survive transient failures and report real paths #1194 added with its own launch for a single why call.
  • The 5 diagnoseProject tests use a project with no run (projectWithFakeRunner).
  • The 4 tests that record a cancel or rewrite run.json keep their own run (launchedProject).
  • Launches go from 38 to 5, and no assertion changes.
  • Each inspectedRun call first checks that run.json, state.json, events.jsonl and smithers/workflow-run-link-journal.json still hold the bytes they had right after the shared launch. A future test that changes the shared run therefore fails the next shared test, with a message saying such a test must launch its own run. A shared launch that fails is reported by every shared test as that launch failing, not as its own regression.
  • To check that the shared tests leave the run alone, a scratch copy of the file digested the shared run's project twice: right after the launch, and after all 52 tests. That is 14,631 files, directories and links, excluding the runner's fixture files. No file was added, removed or changed. Review repeated the digest before every shared test, in three test orders, and found the same. Only two directory mtimes move: the project root's (fixture rewrites) and the run root's (the control-lock directory that each evidence read not marked observeOnly creates and removes).

2. Assertions that cannot catch a regression

Test What it asserted Why it could not catch a regression
agent-adapter-boundaries: "OpenRouter retains the manually reviewed argv responsibility inherited from Codex" (deleted) the file's own adapterPolicies table lists argv-construction for openrouter.tsx it reads test data defined in the same file, so only an edit to that data can fail it
prompt-artifact-authority: promptArtifactAuthorityJsonSchema.$id === PROMPT_ARTIFACT_AUTHORITY_JSON_SCHEMA_ID (assertion deleted) the schema's $id equals a constant the $id is defined as that constant
final-report-markdown: second half of "directive conformance requires both coverage headings…" (assertion deleted, test renamed) calling the checker with an extra false under @ts-expect-error still returns true JavaScript ignores the extra argument, so the call repeats the one above it. At compile time it only pins the parameter list of isDirectiveConformingFinalReportMarkdown, which no production code calls
runtime-test-shard: [1,2,3,4].filter(… === shard(name)) equals [first] (replaced in section 4) a name maps to exactly one shard Corrected by review: it does not follow from the range and determinism assertions. A fractional shard passes both, and the shard wrapper would then skip every test while the release lanes stay green. Section 4 replaces it with [1, 2, 3, 4].includes(first)
dynamic-workflow: the smithers graph branch (branch deleted, test renamed) the compiled workflow runs under smithers graph it ran only if .smithers/node_modules/smithers-orchestrator existed. That is the Smithers 0.32 package name, which the dependency migration removed, so the branch never ran
runtime.test.ts: "reported CLI cost compatibility patches validate and prefer the adapter estimate" (deleted) regexes over the two cost patch strings "reported CLI cost compatibility preserves zero-token adapter estimates" applies the same patches to the pinned engine. It fails if either patched behaviour is removed: the finite/non-negative guard, the guard in the patched constant alone, or the early return

3. Source-text tests of workflow.tsx: 30 deleted, 10 kept, coverage added first

How each candidate was judged.

  • Forty candidate tests mainly assert the template's text.
  • I wrote 93 mutations of workflow.tsx, each breaking one behaviour that one of those tests pins.
  • I deleted all 40 candidates and ran each mutation against the 18 runtime test files that read workflow.tsx and run code from it. The 13 fast ones ran first, then 5 lifecycle and integration files if nothing had failed.
  • runtime.test.ts, lifecycle-inspection and dynamic-lifecycle only render the template through startRun, so they were not run against mutations.
  • 68 mutations were caught.

The 25 survivors were then run against the full base generated-workflow-verifier.test.ts, which shows which candidates catch each one:

  • 18 are caught only by candidate tests. 10 candidates stay, and together they catch all 18.
  • 3 are equivalent to the original code, as noted in the table.
  • 4 are caught by no test, candidates included. These gaps predate this PR; they are listed below, and section 4 closes one of them.

Of the 40 candidates, 10 stay, 29 are deleted, and the relocatable-prompt-path test is rewritten in place as a behavioural test. One test that is not a source-text test is also deleted: it ran git ls-files with its own arguments, so it re-implemented the call under test. That makes 30 deletions.

Review found this sample too thin for the broad tests. Review applied 41 more mutations aimed at what the deleted tests pinned. The deleted tests caught 39 of them and the branch 13; about 13 of the rest were real losses. Review also found one of this PR's recorded catches (submodules-no-expectation) to be spurious. Section 4 restores each loss with a behavioural test, and the per-test rows below say which test now covers what review found.

Behavioural coverage added before deleting. Where a survivor pointed at a behaviour no test checked, and an existing harness could run the real template function, the harness now observes it. Each addition fails on the mutations it targets and passes on the unmodified template.

Test Covers (mutations)
new: "artifact-aware agents recheck every task-local authority after the model returns, succeeding or failing" (harness stubs now report which authority they check) dependency admission before and after generate; prompt, report run-metadata and report-prompt authority after generate; the same on the failure path (6)
new: "retry cleanup clears every task-owned root before it re-prepares the attempt and its authorities" (invariant-suite-handoff-durability) canonical, mirror and generated-test resets, and prompt authority after re-preparation (5)
new: "the verifier publishes nothing for an unsuccessful agent and verifies pinned submodules without restoring them" the agent-success gate, clearing a stale marker, pinnedSubmodules: "verify" (3)
new: "a task worktree must still sit at the recorded launch commit, through its ref as well as HEAD" (real git repository) HEAD and ref checks (2)
new: "the verifier publishes every generated-test companion the manifest authenticates" (the harness now takes the companion list; it was hard-wired to []) companion publication (1)
new, replacing the one-regex test "generated Smithers workflow prefers its relocatable task prompt path": "a task prompt comes from its relocatable prompt path before the dispatch input's path" prompt-path preference (1)
new, replacing "invariant git discovery includes tracked, untracked, and ignored sources", which ran git ls-files with its own arguments: "invariant source discovery lists tracked, untracked and gitignored sources under every supported root", run through the template's invariantWorkspaceSourcePaths --exclude-standard added, --others dropped (2)
case added to "generated Smithers verifier rejects invalid UTF-8 and duplicate JSON keys…" (the harness records marker clears; the stub was a no-op) a failed verification clears the stale marker (1)
case added to "generated task-local prompt authority is minimized, tamper-evident, and restored for retries" (injectable writer) the byte comparison after writing the authority (1)
case added to "generated Smithers accepts direct lenses, excludes transitive lenses, and projects a renamed ledger" one lens per producer (1)
stub in "generated semantic context projects every required review authority field" now answers only a direct-only triage lookup severity binds its direct triage (1)
case added to "#213 the protected baseline outranks a sidecar the agent rewrote" a rewritten protected baseline is refused (1)
cases added to "#211 invariant suite provenance is restricted to supported source roots" (src/.envrc, tests/.npmrc) the nested .envrc rule (1)

Per-test evidence.

  • The mutation ids are those in the verification scripts.
  • "Caught" means a test outside the 40 candidates fails with the mutation applied.
  • Where a candidate's survivor is caught only by a kept candidate, the row names that test.
Deleted test What it asserted Evidence
generated Smithers verifier rejects zero-byte generated-test companions regexes that verifyGeneratedTestFiles and its snapshot reader contain empty-file and size/digest checks 1 of 2 mutations caught (companions-size-mismatch-accepted), e.g. by "generated Smithers authenticates the final generated-test publication snapshot". companions-accept-empty is equivalent: the schema requires size_bytes >= 1, so an empty companion still fails the size check
generated severity verification authenticates the triaged finding preservation context regexes for a direct-only triaged-findings lookup and its run-relative path 2 of 2 mutations caught (severity-triaged-transitive, severity-triaged-path-dropped), e.g. by "generated semantic context projects every required review authority field"
generated review verification uses declared immutable review-stage authority regexes over the review-stage context helpers 2 of 2 mutations caught (review-context-missing, review-severity-authority-missing), e.g. by "1091: generated verifier publishes 41 detections with six family omissions and durable warnings"; "generated semantic context projects every required review authority field" (+1 more)
generated differential verification uses exact declared siblings and ancestors regexes that the differential helper names both lookups and all eight lane schemas differential-no-siblings survives with this test present too: its identifier regexes still pass when siblingDifferentialBindings returns no bindings. Review: its schema-name regexes were the only check that each lane schema reaches its context builder; renaming a case label makes that schema's reconciliation gate fail every node that produces it. Both are now caught by the new "generated differential verification gives each lane schema its exact declared siblings and ancestors"
generated Smithers restores sealed submodules before inputs and verifies them only in the finalizer regexes and indexOf order of the submodule steps 2 of 3 mutations caught: submodules-never-verify by "the verifier publishes nothing for an unsuccessful agent and verifies pinned submodules without restoring them", inputs-asserted-before-submodules by "#1081 complete preparation regenerates every pre-agent store after a dependency replacement" (+3 more). Corrected by review: the recorded catch of submodules-no-expectation was spurious (the failing test's fixture has no pinned submodules), and this test missed it too, because its regex matches the hydrate call. It is now caught by the extended "the post-agent verify pass checks the task's pinned submodules…". Moving preservePinnedSourceProof ahead of the submodule step is caught only by this test's order check; that is left as a residual
generated Smithers workflow binds every planned output to the preflighted schema bundle before agent work regexes and order of the schema-binding and preflight steps 1 of 2 mutations caught (schema-binding-ignores-bundle), e.g. by "generated task preparation binds output schema content but not the validator build". no-schema-binding-check is caught by the kept "generated Smithers preparation names its failing step and carries a retry budget". This test misses it: its order check indexOf(call) < indexOf(preflight) still passes when the call is removed, because indexOf returns -1
generated Smithers workflow does not precreate runtime-owned workspace patch outputs doesNotMatch of patch file names inside prepareArtifactMirror 1 of 1 mutations caught (prepare-precreates-outputs), e.g. by "#1081 complete preparation regenerates every pre-agent store after a dependency replacement"; "#1115 an interrupted workspace replacement still binds property-only dependency markers"
generated Smithers workflow guards runtime-owned workspace patch publication regexes over writeWorkspacePatchArtifact 2 of 2 mutations caught (patch-publication-always-replaces, patch-publication-default-replace), e.g. by "#357 a capture against a different pinned commit is still rejected"; "#357 a foreign schema version is not treated as a superseded capture" (+4 more)
generated prompt authority is derived from sealed controls immediately before each first generation regexes over the authority derivation, the retry reset and the agent wrapper 6 of 6 mutations caught (prompt-authority-no-admission, prompt-authority-no-reread-check, reset-no-prompt-authority, agent-post-prompt-authority, agent-no-repeat-generation-checks, agent-catch-authority), e.g. by "generated task-local prompt authority is minimized, tamper-evident, and restored for retries"; "retry cleanup clears every task-owned root before it re-prepares the attempt and its authorities"
generated Smithers worktrees fail closed on any source other than the pinned benchmark ref regexes for the git commands and messages of the source checks 3 of 4 mutations caught (source-revision-accepts-any-head, source-revision-ignores-ref, agent-no-source-revision-check), e.g. by "a task worktree must still sit at the recorded launch commit, through its ref as well as HEAD"; "agent retries are error-agnostic fresh generations with Smithers' effective prompt" (+1 more). prepare-skips-source-revision is caught by the kept "generated Smithers preparation names its failing step and carries a retry budget". Review: none of the four touched preservePinnedSourceProof, which is the only check that a pinned worktree has no remote, and whose early return for unpinned runs no unit or integration test exercised. Both are now in the extended "generated Smithers pinned source proof rejects any previously published byte drift". Dropping the worktree base branch is caught by runtime.test.ts's rendered-text test
generated retries do not inspect or inject previous failure text regexes over the retry arguments and the post-generation checks 6 of 6 mutations caught (retry-keeps-resume-session, retry-attempt1-unscrubbed, retry-keeps-messages, retry-uses-original-prompt, retry-injects-failure-text, agent-post-admission), e.g. by "a resumed activation's first dispatch carries no stale continuation pointer"; "agent retries are error-agnostic fresh generations with Smithers' effective prompt" (+1 more)
generated Smithers resets exact task-owned artifact contents before every selected attempt regexes over resetTaskArtifactsForRetry and resetTaskArtifactContents 4 of 5 mutations caught (reset-no-canonical, reset-no-mirror, reset-deletes-preserved-prompt, agent-no-reset-before-generation), e.g. by "agent retries are error-agnostic fresh generations with Smithers' effective prompt"; "final-report prompt authority is bounded, tamper-evident, and constant-size across large projections" (+4 more). reset-no-basename-guard is equivalent for an existing root, because the anchoredRoot !== path.join(parent, attemptId) check rejects the same root; for a missing root nothing is deleted either way. Review: none of the five touched the canonical preserve list or that anchor check, and this test was their only guard. Without workspace-patch-baseline.json in the list, every patch-publishing task fails before its model runs; without the anchor check, a reset root replaced by a symlink has its target emptied. Both are now in the extended "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt"
post-agent preparation preserves newly added invariant sources for workspace-patch capture regexes for the preserveCurrentSources flag 1 of 1 mutations caught (post-agent-deletes-new-sources), e.g. by "#1081 complete preparation regenerates every pre-agent store after a dependency replacement". Review: the opposite direction, the flag set on every pass, survived everything else. A pre-agent pass restores the workspace preparation tree first, so the flag shows only on sources inherited after that tree was captured: with it set, a preparation retry or a durable resume leaves the inherited invariant-suite source out of the worktree. Now caught by the new "a pre-agent preparation pass restores inherited invariant sources…"
generated Smithers agent boundary performs only task-local authority checks after completion one regex for the post-generation checks, doesNotMatch of repair helpers, and retries={0} on the verifier 6 of 7 mutations caught (agent-post-admission, agent-post-run-metadata-authority, agent-post-report-prompt-authority, agent-pre-generate-admission, agent-output-schema-leak, agent-task-runtime-leak), e.g. by "artifact-aware agents own process completion without constraining terminal responses"; "artifact-aware agents recheck every task-local authority after the model returns, succeeding or failing" (+2 more). local-verifier-retries survives with this test present too: its retries={0} regex matches the first verifier in the file, the cloud one, so a local verifier with retries={2} passes it
generated Smithers workflow contains no output repair or legacy normalization helpers doesNotMatch of helper names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers agent never promotes dedupe findings into a final report doesNotMatch of helper names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers agent never synthesizes final-report Markdown doesNotMatch of helper names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers agent rejects v2 and legacy generated-test lists without conversion doesNotMatch of helper names and v2 field names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers agent does not strip finding path suffixes doesNotMatch of helper names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers agent does not convert legacy finding field shapes doesNotMatch of helper names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers agent does not convert legacy report provenance doesNotMatch of helper names removed in #558 it checks names only, so the behaviour could return under any other name
generated Smithers never infers or repairs missing generated-test companions doesNotMatch of removed helper names and a regex over a code comment it checks names only, so the behaviour could return under any other name
generated Smithers retries clear every task-owned generated-test work directory regexes over the generated-test reset loop 2 of 2 mutations caught (reset-no-generated-test, reset-only-logical-generated-test), e.g. by "#212 retry cleanup resets generated tests under the repository's plural tests/ root"
generated Smithers workflow preserves the complete invariant suite across worktree handoffs about 60 regexes that identifiers and messages occur in the template, plus the file order of four functions 3 of 4 mutations caught (invariant-dependency-change-accepted, invariant-baseline-tamper-accepted, invariant-ancestor-conflict-accepted), e.g. by "#213 the protected baseline outranks a sidecar the agent rewrote"; "#219 the durable handoff record fails closed when a recorded dependency no longer matches" (+4 more). invariant-ancestor-order-flipped is equivalent since #315: sources are resolved per path over the set of publishers, not by visit order. Review: 4 mutations were too few for about 60 regexes. This test was the only guard of the fail-fast for a producer that claims implemented sources but published no suite, and of the hard-link refusal in test-tree discovery; both now have tests in invariant-suite-handoff-durability. Its other two review survivors are not losses: the symlink-destination check (the canonical-path check after it still throws) and the manifest self-validation (consumers re-validate)
generated Smithers invariant discovery uses a Git-compatible ls-files invocation regexes for the literal ls-files argument lists 2 of 2 mutations caught (ls-files-excludes-ignored, ls-files-tracked-only), e.g. by "#323 a listing past Node's 1 MB default is enumerated in full under the template's own bound"; "#323 an enumeration that outgrows its capture buffer names the subcommand, the bound and the roots" (+3 more). Review: those two mutations cover only invariantWorkspaceSourcePaths. Adding --exclude-standard to changedInvariantSourcePaths, or to the test tree's git fallback, or dropping that fallback's --others, survived; gitignored agent-authored sources would then drop out of the published suite. All three are now caught by the new "changed-source discovery includes untracked and gitignored sources under every supported root"
generated Smithers invariant discovery bounds every git enumeration it captures regexes that six functions call the bounded git helper 1 of 1 mutations caught (enumeration-unbounded), e.g. by "#323 a listing past Node's 1 MB default is enumerated in full under the template's own bound"; "#323 an enumeration that outgrows its capture buffer names the subcommand, the bound and the roots" (+4 more)
generated Smithers invariant provenance accepts only supported source roots regexes over assertSafeInvariantSuitePath 2 of 2 mutations caught (invariant-any-root, invariant-envrc-allowed), e.g. by "#211 invariant suite provenance is restricted to supported source roots"; "#214 an unreferenced file under invariant-suite/ fails the publication closed"
generated Smithers verifier publishes the complete validated set before task success regexes and indexOf order over verifyArtifacts 5 of 5 mutations caught (verifier-no-secret-scan, verifier-no-epoch-check, verifier-marker-before-publication, verifier-no-companion-publication, verifier-no-semantic-gates), e.g. by "1091: generated verifier publishes 41 detections with six family omissions and durable warnings"; "generated Smithers binds generated-test manifests to the current run and logical producer" (+13 more). Review: none of the five touched the scan's configured credential names. Dropping agentCredentialEnv or modal.credentialEnv survived everything else, so an opaque credential under a name the heuristic does not recognise would be published. Both are now caught by the new "generated Smithers verifier refuses to publish the value of any credential the task is configured with"
generated Smithers preparation requires a successful dependency artifact verification regexes that dependency-verification identifiers occur, and the marker order in verifyArtifacts 1 of 1 mutations caught (verifier-keeps-stale-marker), e.g. by "generated Smithers verifier rejects invalid UTF-8 and duplicate JSON keys from captured bytes"
invariant git discovery includes tracked, untracked, and ignored sources ran git ls-files with its own argument list, never the template's discovery 2 of 2 mutations caught (ls-files-excludes-ignored, ls-files-tracked-only), e.g. by "#323 a listing past Node's 1 MB default is enumerated in full under the template's own bound"; "#323 an enumeration that outgrows its capture buffer names the subcommand, the bound and the roots" (+3 more). It re-implemented the call under test
Kept test Mutation that only this test catches
generated Smithers selects property and discovery inputs by their declared contracts companions-select-by-path, expectations-select-by-path
generated Smithers workflow prepares output directories without creating agent-owned files prepare-no-parent-dirs, agent-not-after-preparation
generated Smithers workflow quarantines optional tasks and reads only verified optional ancestors agent-gets-unadmitted-dirs, optional-never-unavailable, ancestor-outputs-not-filtered-by-admission
generated verifiers require the runtime-owned process marker before publishing artifacts verifier-drops-agent-output
generated Smithers verification publishes its exact workspace patch baseline snapshot baseline-tree-unchecked, baseline-not-published
generated Smithers workspace handoff enforces declared production source roots handoff-ignores-declared-roots
generated Smithers verifier treats the final-report projector only as a non-mutating oracle projection-report-not-compared
generated Smithers retry snapshots are durable and restore through canonical parents snapshot-run-root-not-canonical, stale-cleanup-before-reset, stale-cleanup-deletes-runtime-roots
generated Smithers preserves setup-patch baselines across post-agent preparation preparation-evidence-not-required
generated Smithers preparation names its failing step and carries a retry budget no-schema-binding-check, prepare-skips-source-revision

Coverage gaps found, not fixed here. These mutations survive the 18 files and the full base generated-workflow-verifier.test.ts, on the base and on this branch:

  • local-verifier-retries: the local verifier task gets retries={2} instead of 0.
  • coverage-by-default-path: implemented-property coverage is refused unless it uses the default file name.
  • verifier-no-final-report-projection: verifyArtifacts skips verifyFinalReportCanonicalProjection.

Closing them needs harnesses around the rendered task tree, contract-based input selection and the final-report branch of verifyArtifacts. Section 4 closes the fourth gap listed here before review, differential-no-siblings, and two that no test caught before this PR either (the canonical reset deleting invariant-suite-baseline.json or workspace-patch-preparation.json).

Known residuals from review. These also survive every branch test and were caught only by deleted tests, but review judged them low-exposure or near-equivalent, so they are not restored:

  • relocatedRunRoot without realpathSync: breaks only cloud tasks, whose run root is relative; v04 removes cloud execution.
  • The validator preflight envelope parsed as plain JSON: the CLI's exit status is still enforced.
  • The workspace-patch writer reading through a symlink: the durable writer replaces symlinks, and capture rejects non-regular files.
  • The companion's UTF-8 re-decode dropped: the generated-test file-integrity gate rejects the file earlier.
  • The repeat-generation authority pre-checks dropped: they are bracketed by the previous generation's post-checks.
  • preservePinnedSourceProof moved ahead of the submodule step.

4. Review follow-up: guards the deletions lost

Review found that four deleted tests were the only guard of behaviour this PR's mutations never touched, plus smaller losses elsewhere. Each is restored as a behavioural test rather than by bringing back the source-text test. Every template row was checked the same way: the listed mutations fail the named test, and the unmodified template passes it. The shard row was checked with review's fractional mapping applied to the compiled helper, and the last row with a probe test that appends to the shared run's event log (the next shared test fails). The mutation ids are review's (rv-…) unless marked as this PR's; the files and runner are attached in this comment.

Test Covers (mutations)
extended: "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt" (the three runtime-owned file names are now read from the template, not restated) the canonical reset keeps invariant-suite-baseline.json, workspace-patch-baseline.json and workspace-patch-preparation.json; a reset root replaced by a symlink is refused and its target is left intact (4: rv-reset-deletes-workspace-patch-baseline, rv-reset-no-anchor-check, and rv-reset-deletes-invariant-baseline / rv-reset-deletes-patch-preparation, which no test caught before)
extended: "retry cleanup clears every task-owned root before it re-prepares the attempt and its authorities" generated tests are reset under both test/ and tests/ when both exist (1: rv-reset-generated-tests-first-root-only)
new: "generated differential verification gives each lane schema its exact declared siblings and ancestors" (runs the template's semantic context and its real declaration lookups over the differential chain of config/topologies/exhaustive.yml) each of the eight lane schemas gets exactly its declared sibling and ancestor artifacts, and those satisfy every contextual gate the schema has (4: the three case-label renames, and this PR's differential-no-siblings)
new: "generated Smithers verifier refuses to publish the value of any credential the task is configured with" (the verifyArtifacts harness can opt into the real environment scan) an opaque value under a name the heuristic ignores is published when the task does not name it, and refused, with nothing published and no marker, when agentCredentialEnv or modal.credentialEnv names it (2: rv-secret-scan-drops-agent-credentials, rv-secret-scan-drops-modal-credentials)
new: "invariant-suite expectations fail closed when a producer claims implemented sources but published no suite" (1: rv-invariant-missing-suite-root-accepted)
new: "changed-source discovery includes untracked and gitignored sources under every supported root" (real git repository) changedInvariantSourcePaths, and the test tree's git fallback (3: rv-changed-sources-exclude-ignored, rv-changed-tests-exclude-ignored, rv-changed-tests-no-untracked)
new: "test-tree discovery refuses a hard-linked source, whose bytes another path can still change" (1: rv-invariant-hard-link-accepted)
extended: "generated Smithers pinned source proof rejects any previously published byte drift" a pinned worktree with a remote is refused before any proof is written, and an unpinned run writes none (2: rv-proof-ignores-remotes, rv-proof-runs-for-unpinned-sources)
extended and renamed: "the post-agent verify pass checks the task's pinned submodules and does not preflight the agent-facing validator CLI" (the preparation harness records what each submodule step receives) the verify pass compares against the task's pinned-submodule expectation (1: this PR's submodules-no-expectation)
new: "a pre-agent preparation pass restores inherited invariant sources, and only the post-agent pass keeps the worktree's" a preparation retry and a durable resume restore an inherited invariant-suite source from the snapshot; the post-agent pass keeps the attempt's edit (2: rv-preserve-current-sources-always, and this PR's post-agent-deletes-new-sources)
restored: "every runtime test name belongs to exactly one deterministic shard" now asserts [1, 2, 3, 4].includes(first) a fractional shard, which passes the range and determinism assertions (review's (u32 / 2**32) * (total - 1) + 1)
lifecycle-inspection: inspectedRun checks the shared run is unchanged (section 1) a shared test that changes the run fails the next shared test instead of silently changing what later tests see

The configured credential lists are non-empty only for cloud tasks: cloudAgentCredentialEnv returns [] for local execution, and modal.credentialEnv is the Modal provider's configuration. That is where they matter: a cloud task's list also names its agents' route and allowlisted variables, which the name heuristic need not recognise. Stock local agents must use canonical credential names, which it does recognise. v04 removes cloud execution and these lists with it (see Merge interactions).

Deliberately not built

  • No product code changes. Where a deletion leaves product code with no user, it is noted here, not removed:

    • promptArtifactAuthorityJsonSchema is no longer referenced by any test. It stays exported through src/index.ts.
    • isDirectiveConformingFinalReportMarkdown has no production caller.
  • Real-engine execution of a compiled dynamic workflow, which the deleted smithers graph branch claimed to check. It needs the sealed execution snapshot that only startRun builds; a probe without one fails with Cannot find module …/modules/@ultrafuzz/runtime/dist/index.js. That path is changing in v10.

  • Consolidating the multi-launch tests in runtime.test.ts: "syncRun accepts the pinned 0.35.0 usage payload…", "syncRun seals reproducible verifier output failures…" and "ordinary resume tolerates link-journal gaps…". Sharing one run there needs either an exported usage parser (a product change) or resetting run-root documents between cases, which ties the test to the run layout.

  • Fixtures that encode behaviour the analysis calls a bug. The spec lists these for removal, but I kept them:

    • the provider-home 0755 refusal;
    • strict extra-field rejection in lifecycle-inspection;
    • the crashed submodule transaction;
    • the capacity-wait stall label;
    • unresumable runs.

    Each is the only test of what the product does today, so deleting it would remove coverage without fixing anything. Each should change in the PR that changes the behaviour.

  • Regex-over-patch tests with no behavioural twin stay: event_probe_index, engine_agent_event_ownership, and the engine_agent_usage_progress ordering. So do three whole-file scans that behave like lint rules and catch a real class of bug: #691 (every git capture states maxBuffer), the compatibility-patch helper declarations, and the shard-wrapper import check.

Verification

  • Mutation evidence for section 3 was gathered as described there, in two scratch worktrees at integration/wave1 (a48ad9c). The scripts are not committed; the 93 mutations and their runner are attached in this comment, and the review mutations with the targeted runner used for section 4 in this one.

  • Test files, with eatmydata node --test after tsc -p tsconfig.test.json, on this branch:

    File Result
    lifecycle-inspection 52/52. At d3e171c one own-run test's launch hit "…changed while reading" (see below) in a full run; it passed when re-run
    generated-workflow-verifier 119/119 (base 143: 30 deleted, 6 added, 1 rewritten in place)
    invariant-suite-handoff-durability 41/41 (base 37, 4 added)
    workspace-preparation-lifecycle 8/8 (base 7, 1 added)
    invariant-suite-enumeration-overflow 6/6 (base 5, 1 added)
    dynamic-workflow 2/2
    agent-adapter-boundaries 7/7 (base 8, 1 deleted)
    final-report-markdown 40/40
    prompt-artifact-authority 8/8
    runtime-test-shard 5/5, including its Bun off-shard check
    runtime.test.ts, by pattern 3/3: the remaining cost test, the patcher test, and "every runner compatibility patch still anchors in the pinned Smithers release"

    The full runtime.test.ts suite was not run locally; it takes hours.

  • Cost-patch deletion (section 2): I mutated the compiled smithers.js three ways: without the finite/non-negative guard at both anchors, without it in the patched constant alone, and without the early return. The remaining behavioural cost test failed each time.

  • lifecycle-inspection time (the runtime-supporting lane), measured locally with eatmydata on a shared 32-core host:

    Measurement Base This branch
    startRun launches 38 5
    On origin/main 2cacf4c, run one after the other (base at load 11–28, branch at 28–18): wall 20m03s 12m23s
    Same runs, CPU (user+sys) 771 s 527 s
    Same runs, sum of per-test durations 1,108.8 s 716.3 s
    Earlier, on integration/wave1 (37 launches on the base), both started together at load ~40, so contended: wall 35m47s 7m58s
    Same runs, CPU (user+sys) 634 s 356 s
    Earlier, on integration/wave1, one after the other at load 5–40: wall 459.6 s 265.0 s
    • Host load dominates these numbers, so read them as a range. The branch's wall time is about 40% shorter (38% and 42% in the two pairs run one after the other) and it spends 32–44% less CPU. The concurrent pair's 78% is not a fair wall-time comparison: the two runs competed with each other, and a base launch failed.
    • At d3e171c, with the shared-run check, the branch took 9m09s wall and 610 s CPU (user+sys) at load 6–14.
    • In two base runs, one launch failed with "…changed while reading". That is the launch snapshot's hardlink check, which fails when a pnpm install elsewhere on the host touches the shared store. It makes those base times underestimates.
    • The lane runs every supporting file in one concurrent node --test call. This file's wall-time saving shortens the lane only while it is the slowest file; the CPU saving applies either way.
    • generated-workflow-verifier takes about 5 s before and after; its change is about quality, not time.
  • Gates:

    • npx prettier --check and npx eslint on the changed files.
    • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=2cacf4ca pnpm -w lint:strict:ci (at d3e171c too, with origin/main at 2cacf4c).
    • pnpm -w lint, pnpm --filter @ultrafuzz/runtime typecheck, pnpm -w knip, before and after the review follow-up.
    • The complexity ceiling (83) is unchanged. The repository maximum is verifyCoverageProductionInventory, which is product code, and the most complex function in the changed files scores 64.
  • Base. The branch was rebased from integration/wave1 onto origin/main at 2cacf4c (fix(runtime): attempts abandoned by a crash are recorded, so stats and status agree #1200). The only change to the change set is converting fix: observers and benchmark harnesses survive transient failures and report real paths #1194's new test (section 1). The changed test files were re-run after the rebase.

Risk / compatibility

  • Test-only. No file under src/, no schema, config, docs or CHANGELOG.
  • Shared run in lifecycle-inspection. The tests that use inspectedRun share one run, so a new test that changes that run's files would affect later tests. The helper's comment says which tests must launch their own run: those that record a cancel or rewrite run metadata. inspectedRun enforces that for the run's metadata, state, event log and link journal, so a violation fails the next shared test. Fixtures, control files and the command log are reset before each test.
  • Kept source-text tests still fail on unrelated refactors. They stay because each one is the only test that catches a mutation.
  • Merge interactions. These were checked with git merge-tree of each open branch's own commits against this branch, and re-checked at d3e171c: the review follow-up adds no textual conflict with any of them. v01, v02, v03, v06, v07, v08 and v09b are already in the base, and fix(modal): keep the model-work flag when a model node has no run-state record #1202 merges cleanly.
    • v13 (feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198, draft) merges cleanly and does not touch workflow.tsx.
    • v05 (no PR yet) merges cleanly. Its stricter knip run (exports,types,duplicates) passes with this change applied, because promptArtifactAuthorityJsonSchema is still exported through src/index.ts.
    • v04 (refactor!: remove per-node cloud execution (execution.mode = "cloud") #1197, draft) has two interactions:
      • A textual conflict in generated-workflow-verifier.test.ts. v04 edits "…publishes the complete validated set before task success" and "…preparation requires a successful dependency artifact verification", which this PR deletes. Resolve it by keeping the deletion.
      • A conflict git does not flag. v04 deletes import { z } from "zod/v4" from the same file along with its last user, and this PR's new finalizer test uses z. The merged file then fails to compile (TS2304: Cannot find name 'z'), so keep the import.
      • A behaviour v04 removes. v04 drops cloud execution, and with it the configured credential names from the pre-publication scan (sensitiveEnvironmentValues(process.env)). Section 4's "generated Smithers verifier refuses to publish the value of any credential the task is configured with" therefore fails on the merge; delete it in v04 together with the lists. With the import kept, the merge of d3e171c and v04 passes 115 of 116 in generated-workflow-verifier, the one failure being that test. It passes 41/41 in invariant-suite-handoff-durability, 8/8 in workspace-preparation-lifecycle, 6/6 in invariant-suite-enumeration-overflow and 5/5 in runtime-test-shard.
    • v10 (feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201, draft) conflicts in four diagnoseProject hunks of lifecycle-inspection.test.ts. Resolve by taking v10's side of each hunk. v10's versions no longer use the replaced launch lines, and the one test that still launches ("…reports the installed runner that commands after launch execute") needs its run.

Changelog entry

Runtime tests: lifecycle-inspection shares one launched run across its read-only tests, cutting its launches from 38 to 5, and checks that those tests leave the run unchanged. 30 tests that pinned the workflow template's source text (or re-ran its git call themselves) and five redundant, tautological or unreachable assertions are removed. Mutation testing of the template decided which source-text tests stay (10), and 13 new and 9 extended behavioural tests cover what the removed ones pinned, including what review's further mutations found only the removed tests had guarded.

🤖 Generated with Claude Code

aviggiano and others added 3 commits September 29, 2026 10:21
…ection tests

lifecycle-inspection.test.ts launched a fresh run for 38 of its 52 tests,
and each launch seals a full execution snapshot. 29 of those tests only run
read-only inspection commands (why, timeline, snapshots, events, node and a
cancel the runner rejects) against the fake runner, and the five
diagnoseProject tests never read the run they launched.

The fake runner now reads every answer from files that
writeInspectionFixtures resets before each test, so one launched run serves
the read-only tests. The diagnoseProject tests use a project with no run,
and the four tests that record a cancel or rewrite run metadata keep their
own run. Launches go from 38 to 5 and no assertion changes. A digest of
every file of the shared run, taken after launch and again after the whole
file ran, found no change.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
- agent-adapter-boundaries: "OpenRouter retains the manually reviewed argv
  responsibility inherited from Codex" read a value from this file's own
  adapterPolicies table, so only an edit to that test data could fail it.
- prompt-artifact-authority: compared promptArtifactAuthorityJsonSchema.$id
  with the constant the $id is defined from.
- final-report-markdown: the second call passed an extra argument under
  @ts-expect-error. JavaScript ignores it, so the call repeated the first
  one; at compile time it only pinned the parameter list of
  isDirectiveConformingFinalReportMarkdown, which no production code calls.
- runtime-test-shard: the [1, 2, 3, 4].filter assertion follows from the
  range and determinism assertions before it.
- dynamic-workflow: the `smithers graph` branch ran only when
  .smithers/node_modules/smithers-orchestrator existed. That is the Smithers
  0.32 package name, which the dependency migration removes, so the branch
  never ran. The test is renamed to what it checks.
- runtime.test.ts: "reported CLI cost compatibility patches validate and
  prefer the adapter estimate" matched both patch strings with regexes.
  "reported CLI cost compatibility preserves zero-token adapter estimates"
  applies the same patches to the pinned engine and fails when either
  patched behaviour is removed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ioural ones

generated-workflow-verifier.test.ts pinned workflow.tsx with regexes,
indexOf positions and identifier names. To find which of those tests were
the only guard of real behaviour, the template was mutated 93 ways. With
all 40 candidate tests deleted, the 18 test files that run template
functions then ran against each mutation, and 68 mutations were caught.
runtime.test.ts, lifecycle-inspection and dynamic-lifecycle, which only
render the template through startRun, were not run.

Where a survivor showed a behaviour no test checked and an existing harness
could observe it, the harness now does:

- New: artifact-aware agents recheck all four task-local authorities
  after the model returns, whether it succeeds or fails.
- New: retry cleanup clears the canonical, mirror and generated-test roots
  before it re-prepares the attempt and its prompt authority.
- New: the finalizer publishes nothing for an unsuccessful agent, clears a
  stale marker, and verifies pinned submodules without restoring them.
- New: a task worktree must sit at the recorded launch commit, through HEAD
  and through its ref.
- New, replacing a regex over one line: a task prompt prefers its
  relocatable prompt path.
- New, replacing a test that ran git itself: invariant discovery lists
  tracked, untracked and gitignored sources through the template's own
  function.
- The verifyArtifacts harness records marker clears and takes the companion
  list, so failing verification must clear the marker and companions must
  be published.
- One case each on the prompt-authority, property-lens, review-context,
  protected-baseline and source-root tests.

30 tests are deleted: every mutation of what they pinned is now caught,
equivalent, or not caught by the test itself either. The 10 tests that
catch a mutation surviving the rest of the suite stay; one of them, the
preparation step-name test, is the only guard of two preparation steps
pinned by tests deleted here.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 11:03
…tions lost

Review applied 41 more mutations to workflow.tsx, aimed at what the deleted
source-text tests pinned. Four of those tests turned out to be the only guard
of behaviour this change's own 93 mutations never touched. Each loss is
restored as a behavioural test rather than by bringing a source-text test back:

- the canonical retry reset keeps its runtime-owned evidence files and refuses
  a reset root replaced by a symlink;
- each differential lane schema gets exactly its declared sibling and ancestor
  artifacts, and they satisfy every contextual gate it has (this also closes
  the differential-no-siblings gap);
- the pre-publication secret scan covers a cloud task's configured credential
  names;
- the invariant-suite handoff fails closed for a producer that claims sources
  but published none, changed-source discovery keeps untracked and gitignored
  sources, and test-tree discovery refuses a hard-linked source;
- a pinned worktree with a remote is refused before any source proof is kept;
- the post-agent verify pass checks the task's pinned submodules (the recorded
  catch of that mutation was spurious);
- a pre-agent preparation pass restores inherited invariant sources from the
  workspace snapshot.

The shard test again asserts membership, which the range and determinism
checks do not imply. lifecycle-inspection now checks that shared tests leave
the shared run's documents unchanged, and reports a failed shared launch as
that launch failing.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano

Copy link
Copy Markdown
Collaborator Author

Mutation evidence for section 3: the 93 mutations and their runner

These are the 93 mutations of packages/runtime/src/templates/smithers/workflows/workflow.tsx behind section 3's tables, and the runner that produced them. They are attached here, not committed, so the kept and deleted decisions can be re-checked after the template changes.

  • Each entry has an id, the candidate test it targets, and edits: exact find → replace pairs with the expected occurrence count. The runner refuses a mutation whose find no longer matches, so a template change shows up as an error, not a silent pass.
  • All 93 still apply to workflow.tsx at 2cacf4c.
  • To run: build a worktree of this branch (pnpm install && pnpm -w build), run npx tsc -p tsconfig.test.json in packages/runtime, save both files at the repository root, and run node mutate.mjs mutations.json out.jsonl [ids...]. Stage 1 runs the 13 fast template-executing test files. Stage 2 runs the 5 slow ones, and only when stage 1 caught nothing. The template is restored after each mutation, and on SIGINT or SIGTERM.
  • "Caught" means any test failed with the mutation applied. Review showed that is not enough on its own: the recorded catch of submodules-no-expectation failed a test whose fixture has no pinned submodules. When you rely on a catch, check that the failing test can observe the mutated code.
mutations.json (93 mutations)
[
{"id":"agent-post-admission","candidate":"prompt authority / agent boundary","edits":[{"find":"const result = await executionAgent.generate(unstructuredArgs);\n        assertDependencyArtifactAdmissionCurrent(task);\n","replace":"const result = await executionAgent.generate(unstructuredArgs);\n","count":1}],"batch":1},
{"id":"agent-post-prompt-authority","candidate":"prompt authority / agent boundary","edits":[{"find":"const result = await executionAgent.generate(unstructuredArgs);\n        assertDependencyArtifactAdmissionCurrent(task);\n        assertPromptArtifactAuthorityUnchanged(task);\n","replace":"const result = await executionAgent.generate(unstructuredArgs);\n        assertDependencyArtifactAdmissionCurrent(task);\n","count":1}],"batch":1},
{"id":"agent-post-run-metadata-authority","candidate":"prompt authority / agent boundary","edits":[{"find":"        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportRunMetadataAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n        return {","replace":"        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n        return {","count":1}],"batch":1},
{"id":"agent-post-report-prompt-authority","candidate":"prompt authority / agent boundary","edits":[{"find":"        assertFinalReportRunMetadataAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n        return {","replace":"        assertFinalReportRunMetadataAuthorityUnchanged(task);\n        return {","count":1}],"batch":1},
{"id":"agent-catch-authority","candidate":"prompt authority / agent boundary","edits":[{"find":"        try {\n          assertDependencyArtifactAdmissionCurrent(task);\n          assertPromptArtifactAuthorityUnchanged(task);\n          assertFinalReportRunMetadataAuthorityUnchanged(task);\n          assertFinalReportPromptAuthorityUnchanged(task);\n        } catch (authorityError) {","replace":"        try {\n        } catch (authorityError) {","count":1}],"batch":1},
{"id":"agent-output-schema-leak","candidate":"prompt authority / agent boundary","edits":[{"find":"        Reflect.deleteProperty(unstructuredArgs, \"outputSchema\");\n","replace":"","count":1}],"batch":1},
{"id":"agent-task-runtime-leak","candidate":"prompt authority / agent boundary","edits":[{"find":"        Reflect.deleteProperty(unstructuredArgs, \"ultrafuzzTaskRuntime\");\n","replace":"","count":1}],"batch":1},
{"id":"agent-no-reset-before-generation","candidate":"prompt authority / resets before every selected attempt","edits":[{"find":"        assertWorkspaceSourceRevision(task);\n        await resetTaskArtifactsForRetry(task, finalReportTaskRuntimeFromAgentArgs(task, args));\n","replace":"        assertWorkspaceSourceRevision(task);\n","count":1}],"batch":1},
{"id":"agent-no-source-revision-check","candidate":"worktrees fail closed on other sources","edits":[{"find":"      if (firstGenerationForAttempt) {\n        assertWorkspaceSourceRevision(task);\n","replace":"      if (firstGenerationForAttempt) {\n","count":1}],"batch":1},
{"id":"agent-no-repeat-generation-checks","candidate":"prompt authority","edits":[{"find":"      } else {\n        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportRunMetadataAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n      }\n","replace":"      }\n","count":1}],"batch":1},
{"id":"agent-pre-generate-admission","candidate":"agent boundary","edits":[{"find":"        executionAgent ??= admittedAgent();\n        assertDependencyArtifactAdmissionCurrent(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n","replace":"        executionAgent ??= admittedAgent();\n        assertFinalReportPromptAuthorityUnchanged(task);\n","count":1}],"batch":1},
{"id":"retry-keeps-resume-session","candidate":"retries do not inspect previous failure","edits":[{"find":"          resumeSession: undefined,\n          continueSession: false,\n          lastHeartbeat: undefined\n","replace":"          continueSession: false,\n          lastHeartbeat: undefined\n","count":1}],"batch":1},
{"id":"retry-attempt1-unscrubbed","candidate":"retries do not inspect previous failure","edits":[{"find":"        if (smithersAttempt <= 1) return continuationFreeArgs;","replace":"        if (smithersAttempt <= 1) return args;","count":1}],"batch":1},
{"id":"retry-keeps-messages","candidate":"retries do not inspect previous failure","edits":[{"find":"        Reflect.deleteProperty(continuationFreeArgs, \"messages\");\n","replace":"","count":1}],"batch":1},
{"id":"retry-uses-original-prompt","candidate":"retries do not inspect previous failure","edits":[{"find":"          prompt: typeof args?.prompt === \"string\" ? args.prompt : originalPrompt","replace":"          prompt: originalPrompt","count":1}],"batch":1},
{"id":"retry-injects-failure-text","candidate":"retries do not inspect previous failure","edits":[{"find":"          prompt: typeof args?.prompt === \"string\" ? args.prompt : originalPrompt","replace":"          prompt: `${typeof args?.prompt === \"string\" ? args.prompt : originalPrompt}\\nPrevious attempt failed: ${String((args as { error?: unknown } | undefined)?.error ?? \"unknown\")}`","count":1}],"batch":1},
{"id":"reset-no-canonical","candidate":"resets exact task-owned artifact contents","edits":[{"find":"  resetTaskArtifactContents(task.metadata.artifacts.dir, task.attemptId, \"canonical\", promptPath);\n","replace":"  void promptPath;\n","count":1}],"batch":1},
{"id":"reset-no-mirror","candidate":"resets exact task-owned artifact contents","edits":[{"find":"  resetTaskArtifactContents(path.join(artifactsParent, task.attemptId), task.attemptId, \"mirror\");\n","replace":"","count":1}],"batch":1},
{"id":"reset-no-generated-test","candidate":"retries clear every generated-test work directory","edits":[{"find":"        resetTaskArtifactContents(path.join(foundryParent, nodeId), nodeId, \"generated-test\");\n","replace":"        void nodeId;\n","count":1}],"batch":1},
{"id":"reset-only-logical-generated-test","candidate":"retries clear every generated-test work directory","edits":[{"find":"new Set([task.metadata.node.logicalNodeId, task.metadata.node.concreteNodeId])","replace":"new Set([task.metadata.node.logicalNodeId])","count":1}],"batch":1},
{"id":"reset-deletes-preserved-prompt","candidate":"resets exact task-owned artifact contents","edits":[{"find":"      candidate === preservedInput ||\n","replace":"      false ||\n","count":1}],"batch":1},
{"id":"reset-no-basename-guard","candidate":"resets exact task-owned artifact contents","edits":[{"find":"  if (path.basename(candidate) !== attemptId) {\n    throw new Error(`artifact-contract failure: unsafe ${label} task artifact root ${attemptId}`);\n  }\n","replace":"","count":1}],"batch":1},
{"id":"reset-no-prompt-authority","candidate":"prompt authority derived before each first generation","edits":[{"find":"  prepareArtifactMirror(task, { replayWorkspacePatches: false, evidenceMode: \"require\" });\n  materializePromptArtifactAuthority(task);\n","replace":"  prepareArtifactMirror(task, { replayWorkspacePatches: false, evidenceMode: \"require\" });\n","count":1}],"batch":1},
{"id":"prompt-authority-no-admission","candidate":"prompt authority derived from sealed controls","edits":[{"find":"    admittedDependencyArtifactDirs: admission.directories,\n","replace":"    admittedDependencyArtifactDirs: task.dependencyArtifactDirs,\n","count":1}],"batch":1},
{"id":"prompt-authority-no-reread-check","candidate":"prompt authority derived from sealed controls","edits":[{"find":"  if (!captured.bytes.equals(expected)) {\n    throw new Error(\n      `artifact-contract failure: prompt artifact authority changed while materialized ${task.attemptId}`\n    );\n  }\n","replace":"","count":1}],"batch":1},
{"id":"agent-gets-unadmitted-dirs","candidate":"quarantines optional tasks","edits":[{"find":"        const admitted = baseAgentForProfile(task, profile, admittedDependencyArtifactDirs(task));","replace":"        const admitted = baseAgentForProfile(task, profile, task.dependencyArtifactDirs);","count":1}],"batch":1},
{"id":"agent-gets-no-dependency-dirs","candidate":"quarantines optional tasks","edits":[{"find":"    addDir: [task.artifactDir, ...dependencyArtifactDirs]","replace":"    addDir: [task.artifactDir]","count":1}],"batch":1},
{"id":"optional-never-unavailable","candidate":"quarantines optional tasks","edits":[{"find":"  return marker === undefined || !pathEntryExists(marker.path);\n}","replace":"  return false;\n}","count":1}],"batch":1},
{"id":"optional-always-unavailable","candidate":"quarantines optional tasks","edits":[{"find":"  return marker === undefined || !pathEntryExists(marker.path);\n}","replace":"  return true;\n}","count":1}],"batch":1},
{"id":"ancestor-outputs-not-filtered-by-admission","candidate":"quarantines optional tasks","edits":[{"find":"  return outputs.filter((output) => admittedDirectories.has(path.resolve(output.artifactDir)));","replace":"  return outputs;","count":1}],"batch":1},
{"id":"lenses-transitive","candidate":"selects property inputs by declared contracts","edits":[{"find":"  const lensOutputs = declaredAncestorContractOutputs(task, \"ultrafuzz/property-lens@2\", {\n    directOnly: true\n  });","replace":"  const lensOutputs = declaredAncestorContractOutputs(task, \"ultrafuzz/property-lens@2\", {\n    directOnly: false\n  });","count":1}],"batch":1},
{"id":"lenses-ledger-projection-required","candidate":"selects property inputs by declared contracts","edits":[{"find":"      sourceNodeId: output.logicalNodeId,\n      projectionRequired: false,","replace":"      sourceNodeId: output.logicalNodeId,\n      projectionRequired: true,","count":1}],"batch":1},
{"id":"lenses-ambiguity-accepted","candidate":"selects property inputs by declared contracts","edits":[{"find":"    if (seenLensProducers.has(output.attemptId)) {\n      throw new Error(`artifact-contract failure: property-lens producer is ambiguous ${output.attemptId}`);\n    }\n","replace":"","count":1}],"batch":1},
{"id":"submodules-never-verify","candidate":"restores sealed submodules and verifies in finalizer","edits":[{"find":"    evidenceMode: \"require\",\n    pinnedSubmodules: \"verify\"\n  });","replace":"    evidenceMode: \"require\"\n  });","count":1}],"batch":1},
{"id":"submodules-no-expectation","candidate":"restores sealed submodules and verifies in finalizer","edits":[{"find":"        expectation: task.pinnedSubmodules ?? undefined\n      })\n    );\n  } else {","replace":"        expectation: undefined\n      })\n    );\n  } else {","count":1}],"batch":1},
{"id":"finalizer-ignores-agent-process","candidate":"require runtime-owned process marker","edits":[{"find":"  if (!agentProcessOutput.safeParse(agentProcess).success) {\n    throw new Error(`artifact-contract failure: agent task did not succeed ${task.attemptId}`);\n  }\n","replace":"","count":1}],"batch":1},
{"id":"finalizer-keeps-stale-marker","candidate":"require runtime-owned process marker","edits":[{"find":"  clearArtifactVerificationMarker(task);\n  if (!agentProcessOutput.safeParse(agentProcess).success) {","replace":"  if (!agentProcessOutput.safeParse(agentProcess).success) {","count":1}],"batch":1},
{"id":"no-schema-binding-check","candidate":"binds planned output to preflighted schema bundle","edits":[{"find":"  preparationStep(task.attemptId, \"assert-task-output-schema-bindings\", () => assertTaskOutputSchemaBindings(task));\n","replace":"","count":1}],"batch":1},
{"id":"schema-binding-ignores-bundle","candidate":"binds planned output to preflighted schema bundle","edits":[{"find":"      binding?.schema_sha256 !== output.schemaSha256 ||\n      binding?.schema_bundle_sha256 !== output.schemaBundleSha256\n","replace":"      binding?.schema_sha256 !== output.schemaSha256\n","count":1}],"batch":1},
{"id":"verifier-no-secret-scan","candidate":"publishes complete validated set before task success","edits":[{"find":"    assertArtifactPublicationsContainNoSecrets(\n      publications,\n      sensitiveEnvironmentValues(process.env, [\n        ...(task.execution?.agentCredentialEnv ?? []),\n        ...(task.execution?.modal?.credentialEnv ?? [])\n      ])\n    );\n","replace":"","count":1}],"batch":1},
{"id":"verifier-no-epoch-check","candidate":"publishes complete validated set before task success","edits":[{"find":"    assertVerifiedDependencySnapshotEpochRemainedCurrent(task, dependencySnapshotEpoch);\n","replace":"","count":1}],"batch":1},
{"id":"verifier-marker-before-publication","candidate":"publishes complete validated set before task success","edits":[{"find":"    publishVerifiedArtifacts(artifactDir, publications);\n    assertVerifiedDependencySnapshotEpochRemainedCurrent(task, dependencySnapshotEpoch);\n    const verificationMarker = writeArtifactVerificationMarker(task, artifacts, publications, validationWarnings);\n","replace":"    assertVerifiedDependencySnapshotEpochRemainedCurrent(task, dependencySnapshotEpoch);\n    const verificationMarker = writeArtifactVerificationMarker(task, artifacts, publications, validationWarnings);\n    publishVerifiedArtifacts(artifactDir, publications);\n","count":1}],"batch":1},
{"id":"verifier-no-companion-publication","candidate":"publishes complete validated set before task success","edits":[{"find":"        for (const companion of verifyGeneratedTestFiles(verified.artifactRoot, verified.value)) {\n          rememberVerifiedPublication(publications, companion.path, companion.contents);\n        }\n","replace":"        verifyGeneratedTestFiles(verified.artifactRoot, verified.value);\n","count":1}],"batch":1},
{"id":"verifier-no-semantic-gates","candidate":"publishes complete validated set before task success","edits":[{"find":"    const validationWarnings = verifyOutputSemanticGates(task, verifiedOutputs, campaignEvidence);\n","replace":"    const validationWarnings: never[] = [];\n","count":1}],"batch":1},
{"id":"verifier-keeps-stale-marker","candidate":"preparation requires successful dependency verification","edits":[{"find":"  // A model-controlled workspace can pre-create arbitrary sidecars. Remove\n  // any stale marker before validating so only this verifier can publish the\n  // success boundary consumed by downstream preparation tasks.\n  clearArtifactVerificationMarker(task);\n","replace":"","count":1}],"batch":1},
{"id":"verifier-no-final-report-projection","candidate":"verifier treats final-report projector as oracle","edits":[{"find":"    verifyFinalReportCanonicalProjection(task, verifiedOutputs);\n","replace":"","count":1}],"batch":1},
{"id":"companions-accept-empty","candidate":"zero-byte generated-test companions","edits":[{"find":"    if (companion.stats.size === 0) {\n      throw new Error(`artifact-contract failure: generated-test bundle file is empty ${companion.relativePath}`);\n    }\n","replace":"","count":1},{"find":"  if (requireNonEmpty && bytes.length === 0) {\n    throw new Error(`${failureMessage}: file is empty`);\n  }\n","replace":"","count":1}],"batch":1},
{"id":"companions-size-mismatch-accepted","candidate":"zero-byte generated-test companions","edits":[{"find":"    if (snapshot.bytes.length !== entry.size_bytes) {\n      throw new Error(`artifact-contract failure: generated-test bundle file size does not match ${relativePath}`);\n    }\n","replace":"","count":1}],"batch":1},
{"id":"source-revision-accepts-any-head","candidate":"worktrees fail closed on other sources","edits":[{"find":"  if (head !== task.sourceRevision || sourceRef !== task.sourceRevision) {","replace":"  if (false && (head !== task.sourceRevision || sourceRef !== task.sourceRevision)) {","count":1}],"batch":1},
{"id":"source-revision-ignores-ref","candidate":"worktrees fail closed on other sources","edits":[{"find":"  if (head !== task.sourceRevision || sourceRef !== task.sourceRevision) {","replace":"  if (head !== task.sourceRevision) {","count":1}],"batch":1},
{"id":"prepare-skips-source-revision","candidate":"worktrees fail closed on other sources","edits":[{"find":"    preparationStep(task.attemptId, \"assert-workspace-source-revision\", () => assertWorkspaceSourceRevision(task));\n","replace":"","count":1}],"batch":1},
{"id":"post-agent-deletes-new-sources","candidate":"post-agent preparation preserves new invariant sources","edits":[{"find":"      preserveCurrentSources: options.replayWorkspacePatches === false\n","replace":"      preserveCurrentSources: false\n","count":1}],"batch":1},
{"id":"preparation-step-unnamed","candidate":"preparation names failing step","edits":[{"find":"      `prepare:${attemptId} failed at step ${step}: ${error instanceof Error ? error.message : String(error)}`,","replace":"      `prepare:${attemptId} failed: ${error instanceof Error ? error.message : String(error)}`,","count":1}],"batch":1},
{"id":"preparation-step-retryable","candidate":"preparation names failing step and carries retry budget","edits":[{"find":"    throw isNonRetryableFailure(error) ? nonRetryableFailure(wrapped) : wrapped;","replace":"    throw wrapped;","count":1}],"batch":1},
{"id":"severity-triaged-transitive","candidate":"severity triaged context","edits":[{"find":"      \"ultrafuzz/triaged-findings@1\",\n      \"triaged findings\",\n      { directOnly: true }\n    );","replace":"      \"ultrafuzz/triaged-findings@1\",\n      \"triaged findings\",\n      { directOnly: false }\n    );","count":1}],"batch":2},
{"id":"severity-triaged-path-dropped","candidate":"severity triaged context","edits":[{"find":"            triagedFindings: triagedFindings.value,\n            triagedFindingsArtifactPath: triagedFindings.runRelativePath\n","replace":"            triagedFindings: triagedFindings.value\n","count":1}],"batch":2},
{"id":"review-context-missing","candidate":"review-stage authority","edits":[{"find":"    context.artifactSet = {\n      reviewStage,\n","replace":"    context.artifactSet = {\n","count":1}],"batch":2},
{"id":"review-severity-authority-missing","candidate":"review-stage authority","edits":[{"find":"...finalSeverityAuthority","replace":"...{}","count":1}],"batch":2},
{"id":"differential-no-siblings","candidate":"differential siblings/ancestors","edits":[{"find":"  return declaredSiblingOutputsByContract(current, contract).map(","replace":"  return [].map(","count":1}],"batch":2},
{"id":"prepare-precreates-outputs","candidate":"output directories without agent-owned files","edits":[{"find":"      mkdirSync(parentPath, { recursive: true });\n","replace":"      mkdirSync(parentPath, { recursive: true });\n      if (!existsSync(artifactPath)) writeFileDurable(artifactPath, \"\");\n","count":1}],"batch":2},
{"id":"prepare-no-parent-dirs","candidate":"output directories without agent-owned files","edits":[{"find":"      mkdirSync(parentPath, { recursive: true });\n","replace":"","count":1}],"batch":2},
{"id":"patch-publication-always-replaces","candidate":"guards runtime-owned patch publication","edits":[{"find":"      if (!replaceSuperseded) {\n        throw new Error(`artifact-contract failure: workspace patch artifact was modified ${relativePath}`);","replace":"      if (false) {\n        throw new Error(`artifact-contract failure: workspace patch artifact was modified ${relativePath}`);","count":1}],"batch":2},
{"id":"patch-publication-default-replace","candidate":"guards runtime-owned patch publication","edits":[{"find":"  replaceSuperseded = false\n): void {","replace":"  replaceSuperseded = true\n): void {","count":1}],"batch":2},
{"id":"baseline-tree-unchecked","candidate":"workspace patch baseline snapshot","edits":[{"find":"  if (parsed.attempt_id !== task.attemptId || expectedTree === undefined || parsed.baseline_tree !== expectedTree) {","replace":"  if (parsed.attempt_id !== task.attemptId) {","count":1}],"batch":2},
{"id":"baseline-not-published","candidate":"workspace patch baseline snapshot","edits":[{"find":"    if (taskPublishesWorkspacePatch(task)) {\n      rememberVerifiedPublication(","replace":"    if (false) {\n      rememberVerifiedPublication(","count":1}],"batch":2},
{"id":"handoff-ignores-declared-roots","candidate":"declared production source roots","edits":[{"find":"  const captured = captureWorkspacePatch(workspaceRoot, baselineTree, task.productionSourceRoots);","replace":"  const captured = captureWorkspacePatch(workspaceRoot, baselineTree, [\"src\", \"contracts\"]);","count":1}],"batch":2},
{"id":"prompt-path-prefers-input","candidate":"relocatable task prompt path","edits":[{"find":"    const promptPath = task.promptPath ?? inputTask?.prompt_path;","replace":"    const promptPath = inputTask?.prompt_path ?? task.promptPath;","count":1}],"batch":2},
{"id":"invariant-ancestor-order-flipped","candidate":"invariant suite across handoffs","edits":[{"find":"    if (leftDirect !== rightDirect) return leftDirect ? 1 : -1;","replace":"    if (leftDirect !== rightDirect) return leftDirect ? -1 : 1;","count":1}],"batch":2},
{"id":"invariant-dependency-change-accepted","candidate":"invariant suite across handoffs","edits":[{"find":"    if (bytes.length !== entry.size || createHash(\"sha256\").update(bytes).digest(\"hex\") !== entry.sha256) {\n      throw new Error(`artifact-contract failure: invariant suite dependency changed ${relativePath}`);\n    }\n","replace":"","count":1}],"batch":2},
{"id":"invariant-baseline-tamper-accepted","candidate":"invariant suite across handoffs","edits":[{"find":"    if (snapshot !== undefined && snapshot.sha256 !== digest) {\n      throw new Error(\"artifact-contract failure: protected invariant suite baseline was modified\");\n    }\n","replace":"","count":1}],"batch":2},
{"id":"invariant-ancestor-conflict-accepted","candidate":"invariant suite across handoffs","edits":[{"find":"      if (disagreeing !== undefined) {\n        throw new Error(","replace":"      if (false) {\n        throw new Error(","count":1}],"batch":2},
{"id":"ls-files-excludes-ignored","candidate":"Git-compatible ls-files invocation","edits":[{"find":"  const values = invariantSuiteGitPaths(workspaceRoot, [\n    \"ls-files\",\n    \"--cached\",\n    \"--others\",\n    \"--\",\n    \"src\",","replace":"  const values = invariantSuiteGitPaths(workspaceRoot, [\n    \"ls-files\",\n    \"--cached\",\n    \"--others\",\n    \"--exclude-standard\",\n    \"--\",\n    \"src\",","count":1}],"batch":2},
{"id":"ls-files-tracked-only","candidate":"Git-compatible ls-files invocation","edits":[{"find":"  const values = invariantSuiteGitPaths(workspaceRoot, [\n    \"ls-files\",\n    \"--cached\",\n    \"--others\",\n    \"--\",\n    \"src\",","replace":"  const values = invariantSuiteGitPaths(workspaceRoot, [\n    \"ls-files\",\n    \"--cached\",\n    \"--\",\n    \"src\",","count":1}],"batch":2},
{"id":"enumeration-unbounded","candidate":"bounds every git enumeration","edits":[{"find":"    return execFileSync(\"git\", [...args], {\n      cwd: workspaceRoot,\n      maxBuffer: MAX_INVARIANT_SUITE_ENUMERATION_BYTES\n    }).toString(\"utf8\");","replace":"    return execFileSync(\"git\", [...args], {\n      cwd: workspaceRoot\n    }).toString(\"utf8\");","count":1}],"batch":2},
{"id":"invariant-any-root","candidate":"supported source roots","edits":[{"find":"  if (!INVARIANT_SUITE_ALLOWED_ROOTS.some((prefix) => value.startsWith(`${prefix}/`))) {","replace":"  if (false) {","count":1}],"batch":2},
{"id":"invariant-envrc-allowed","candidate":"supported source roots","edits":[{"find":"segment === \".envrc\"","replace":"segment === \".envrc-disabled\"","count":1}],"batch":2},
{"id":"snapshot-run-root-not-canonical","candidate":"retry snapshots through canonical parents","edits":[{"find":"    runRootStat.isSymbolicLink() ||\n    realpathSync(runRootCandidate) !== runRootCandidate\n","replace":"    runRootStat.isSymbolicLink()\n","count":1}],"batch":2},
{"id":"stale-cleanup-before-reset","candidate":"retry snapshots through canonical parents","edits":[{"find":"  restoreWorkspaceTreeWithIndexLockRecovery(workspaceRoot, preparationTree);\n  removeStaleWorkspaceFiles(workspaceRoot, preparationTree);\n","replace":"  removeStaleWorkspaceFiles(workspaceRoot, preparationTree);\n  restoreWorkspaceTreeWithIndexLockRecovery(workspaceRoot, preparationTree);\n","count":1}],"batch":2},
{"id":"stale-cleanup-deletes-runtime-roots","candidate":"retry snapshots through canonical parents","edits":[{"find":"    if (expected.has(relativePath) || isWorkspaceRuntimePath(relativePath)) continue;","replace":"    if (expected.has(relativePath)) continue;","count":1}],"batch":2},
{"id":"replay-skip-not-chained","candidate":"setup-patch baselines / replay","edits":[{"find":"return chained ? index + 1 : 0;","replace":"return index + 1;","count":1}],"batch":2},
{"id":"replay-skip-on-post-agent","candidate":"setup-patch baselines / replay","edits":[{"find":"    replayWorkspacePatches && captures.length > 0\n","replace":"    captures.length > 0\n","count":1}],"batch":2},
{"id":"replay-skipped-capture-unvalidated","candidate":"setup-patch baselines / replay","edits":[{"find":"for (const capture of captures) validateWorkspacePatchCapture(workspaceRoot, capture, task.productionSourceRoots);","replace":"void captures;","count":1}],"batch":2},
{"id":"preparation-evidence-not-required","candidate":"setup-patch baselines / replay","edits":[{"find":"persistedPreparation === undefined && evidenceMode === \"require\"","replace":"false","count":1}],"batch":2},
{"id":"local-verifier-retries","candidate":"agent boundary / verifier retries={0}","edits":[{"find":"                depsOptional\n                continueOnFail={task.continueOnFail}\n                retries={0}\n                metadata={{\n                  category: \"artifact-contract\",","replace":"                depsOptional\n                continueOnFail={task.continueOnFail}\n                retries={2}\n                metadata={{\n                  category: \"artifact-contract\",","count":1}],"batch":3},
{"id":"projection-report-not-compared","candidate":"final-report projector is a non-mutating oracle","edits":[{"find":"isDeepStrictEqual(projection.report, report.value)","replace":"true","count":1}],"batch":3},
{"id":"projection-markdown-not-compared","candidate":"final-report projector is a non-mutating oracle","edits":[{"find":"markdown.file.bytes.equals(Buffer.from(projection.markdown, \"utf8\"))","replace":"true","count":1}],"batch":3},
{"id":"unplanned-track-reason-dropped","candidate":"final-report projector / not-planned reason","edits":[{"find":"reason: \"property-implementation-track-not-declared\"","replace":"reason: \"unknown\"","count":1}],"batch":3},
{"id":"agent-not-after-preparation","candidate":"prepares output directories / preparation dependency","edits":[{"find":"dependsOn={[task.preparationId]}","replace":"dependsOn={task.dependsOn}","count":1}],"batch":3},
{"id":"coverage-by-default-path","candidate":"selects inputs by declared contracts","edits":[{"find":"  const implementation = verifiedSingletonAncestorJsonArtifact(\n    task,\n    \"ultrafuzz/implemented-properties@3\",\n    \"implemented property coverage\"\n  );","replace":"  const implementation = verifiedSingletonAncestorJsonArtifact(\n    task,\n    \"ultrafuzz/implemented-properties@3\",\n    \"implemented property coverage\"\n  ) ?? undefined;\n  if (implementation !== undefined && implementation.path !== \"implemented-properties.json\") throw new Error(\"artifact-contract failure: implemented properties must use the default path\");","count":1}],"batch":4},
{"id":"companions-select-by-path","candidate":"selects inputs by declared contracts","edits":[{"find":"  const implementationOutputs = task.outputs.filter(\n    (output) => output.contract === \"ultrafuzz/implemented-properties@3\"\n  );","replace":"  const implementationOutputs = task.outputs.filter(\n    (output) => output.path === \"implemented-properties.json\"\n  );","count":1}],"batch":4},
{"id":"expectations-select-by-path","candidate":"selects inputs by declared contracts","edits":[{"find":"    const implementationOutputs = producer.outputs.filter(\n      (output) => output.contract === \"ultrafuzz/implemented-properties@3\"\n    );","replace":"    const implementationOutputs = producer.outputs.filter(\n      (output) => output.path === \"implemented-properties.json\"\n    );","count":1}],"batch":4},
{"id":"verifier-drops-agent-output","candidate":"runtime-owned process marker / render wiring","edits":[{"find":"{(deps) => finalizeAndVerifyArtifacts(task, deps.agent)}","replace":"{() => finalizeAndVerifyArtifacts(task, undefined)}","count":2}],"batch":5},
{"id":"inputs-asserted-before-submodules","candidate":"restores sealed submodules before inputs","edits":[{"find":"  preparationStep(task.attemptId, \"assert-task-inputs\", () => assertTaskInputs(task, workspaceRoot));\n","replace":"","count":1},{"find":"  if (options.pinnedSubmodules === \"verify\") {\n","replace":"  preparationStep(task.attemptId, \"assert-task-inputs\", () => assertTaskInputs(task, workspaceRoot));\n  if (options.pinnedSubmodules === \"verify\") {\n","count":1}],"batch":5}
]
mutate.mjs
// Runs each mutation of the workflow template against the behavioural suite.
// Usage (from the repository root, after a build and `npx tsc -p tsconfig.test.json` in packages/runtime):
//   node mutate.mjs <mutations.json> <out.jsonl> [ids...]
// Stage 1 runs the fast template-executing files; stage 2 (only for survivors) runs the slow ones.
import fs from "node:fs";
import path from "node:path";
import { spawnSync } from "node:child_process";
const RUNTIME = path.resolve(process.env.ULTRAFUZZ_RUNTIME ?? "packages/runtime");
const TEMPLATE = path.join(RUNTIME, "src/templates/smithers/workflows/workflow.tsx");
const files = (names) => names.map((name) => `dist-test/test/${name}.test.js`);
const STAGES = [
  files([
    "generated-workflow-verifier", "invariant-suite-ancestor-order", "invariant-suite-enumeration-overflow",
    "invariant-suite-handoff-durability", "stale-workspace-cleanup-overflow", "workflow-dependency-policy",
    "terminal-report-projection", "report-retry-history", "workspace-patch-supersede", "task-workflow-identity",
    "workspace-patch-replay", "workspace-preparation-replacement", "smithers-dependency-skip.integration"
  ]),
  files([
    "smithers-preparation-race.integration", "smithers-report-retry.integration", "smithers-resume-reopen.integration",
    "workspace-preparation-lifecycle", "workspace-handoff"
  ])
];
const [mutationsFile, outFile, ...only] = process.argv.slice(2);
const mutations = JSON.parse(fs.readFileSync(mutationsFile, "utf8")).filter((m) => only.length === 0 || only.includes(m.id));
const pristine = fs.readFileSync(TEMPLATE, "utf8");
const restore = () => fs.writeFileSync(TEMPLATE, pristine);
for (const signal of ["SIGINT", "SIGTERM", "SIGHUP"]) process.on(signal, () => { restore(); process.exit(1); });
function run(stageFiles) {
  const result = spawnSync("eatmydata", ["node", "--test", "--test-concurrency=6", "--test-reporter=tap", ...stageFiles], {
    cwd: RUNTIME, encoding: "utf8", maxBuffer: 256 * 1024 * 1024, timeout: 1_200_000
  });
  const failed = [...result.stdout.matchAll(/^\s*not ok \d+ - (.*)$/gmu)].map((m) => m[1].trim())
    .filter((name) => !/\.test\.js$/u.test(name) && !name.startsWith("/"));
  return { status: result.status, failed: [...new Set(failed)] };
}
try {
  for (const mutation of mutations) {
    let mutated = pristine;
    let error;
    for (const edit of mutation.edits) {
      const occurrences = mutated.split(edit.find).length - 1;
      const expected = edit.count ?? 1;
      if (occurrences !== expected) { error = `expected ${expected} occurrence(s), found ${occurrences}: ${edit.find.slice(0, 120)}`; break; }
      mutated = mutated.split(edit.find).join(edit.replace);
    }
    if (error) {
      fs.appendFileSync(outFile, `${JSON.stringify({ id: mutation.id, error })}\n`);
      console.log(mutation.id, "ERROR", error);
      continue;
    }
    fs.writeFileSync(TEMPLATE, mutated);
    const started = Date.now();
    let outcome = { status: 0, failed: [] };
    let stage = 0;
    for (const stageFiles of STAGES) {
      stage += 1;
      outcome = run(stageFiles);
      if (outcome.failed.length > 0 || outcome.status !== 0) break;
    }
    restore();
    const record = { id: mutation.id, candidate: mutation.candidate, caught: outcome.failed.length > 0 || outcome.status !== 0, stage, failed: outcome.failed, status: outcome.status, seconds: Math.round((Date.now() - started) / 1000) };
    fs.appendFileSync(outFile, `${JSON.stringify(record)}\n`);
    console.log(mutation.id, record.caught ? `CAUGHT(stage ${stage})` : "SURVIVED", outcome.failed.slice(0, 4).join(" | "));
  }
} finally {
  restore();
}

@aviggiano

aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Mutation checks for the review follow-up (d3e171c)

Review applied 42 more mutations to the template, aimed at what the deleted source-text tests pinned: 41 new ones, plus a re-run of this PR's submodules-no-expectation. The follow-up commit adds or extends tests for the losses review found.

Each mutation below was run against the one test meant to catch it. That test fails with the mutation applied, and passes on the unmodified template: 21 of 21 caught, 10 of 10 controls pass. Separately, the shard test fails on review's fractional mapping (u32 / 2**32) * (total - 1) + 1, applied to the compiled runtime-test-shard.js.

Mutation Test it fails (file)
rv-reset-deletes-workspace-patch-baseline "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt" (generated-workflow-verifier)
rv-reset-no-anchor-check "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt" (generated-workflow-verifier)
rv-reset-deletes-invariant-baseline "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt" (generated-workflow-verifier)
rv-reset-deletes-patch-preparation "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt" (generated-workflow-verifier)
rv-reset-generated-tests-first-root-only "retry cleanup clears every task-owned root before it re-prepares the attempt and its authorities" (invariant-suite-handoff-durability)
rv-differential-gap-review-case-dropped "generated differential verification gives each lane schema its exact declared siblings and ancestors" (generated-workflow-verifier)
rv-differential-reference-harness-case-dropped "generated differential verification gives each lane schema its exact declared siblings and ancestors" (generated-workflow-verifier)
rv-differential-red-triage-case-dropped "generated differential verification gives each lane schema its exact declared siblings and ancestors" (generated-workflow-verifier)
differential-no-siblings (this PR's) "generated differential verification gives each lane schema its exact declared siblings and ancestors" (generated-workflow-verifier)
rv-secret-scan-drops-agent-credentials "generated Smithers verifier refuses to publish the value of any credential the task is configured with" (generated-workflow-verifier)
rv-secret-scan-drops-modal-credentials "generated Smithers verifier refuses to publish the value of any credential the task is configured with" (generated-workflow-verifier)
rv-invariant-missing-suite-root-accepted "invariant-suite expectations fail closed when a producer claims implemented sources but published no suite" (invariant-suite-handoff-durability)
rv-changed-sources-exclude-ignored "changed-source discovery includes untracked and gitignored sources under every supported root" (invariant-suite-handoff-durability)
rv-changed-tests-exclude-ignored "changed-source discovery includes untracked and gitignored sources under every supported root" (invariant-suite-handoff-durability)
rv-changed-tests-no-untracked "changed-source discovery includes untracked and gitignored sources under every supported root" (invariant-suite-handoff-durability)
rv-invariant-hard-link-accepted "test-tree discovery refuses a hard-linked source, whose bytes another path can still change" (invariant-suite-handoff-durability)
rv-proof-ignores-remotes "generated Smithers pinned source proof rejects any previously published byte drift" (generated-workflow-verifier)
rv-proof-runs-for-unpinned-sources "generated Smithers pinned source proof rejects any previously published byte drift" (generated-workflow-verifier)
submodules-no-expectation (this PR's) "the post-agent verify pass checks the task's pinned submodules and does not preflight the agent-facing validator CLI" (workspace-preparation-lifecycle)
rv-preserve-current-sources-always "a pre-agent preparation pass restores inherited invariant sources, and only the post-agent pass keeps the worktree's" (workspace-preparation-lifecycle)
post-agent-deletes-new-sources (this PR's) "a pre-agent preparation pass restores inherited invariant sources, and only the post-agent pass keeps the worktree's" (workspace-preparation-lifecycle)

The review mutations this commit does not restore are listed under "Known residuals from review" in the PR description, with the reason each is low-exposure.

To re-run: save the two files below next to mutations.json from the comment above, build as described there, and run node mutcheck.mjs out.jsonl [ids...] from the repository root.

review-mutations.json (review's 42 mutations)
[
{"id":"rv-secret-scan-drops-agent-credentials","target":"verifier publishes complete validated set","edits":[{"find":"      sensitiveEnvironmentValues(process.env, [\n        ...(task.execution?.agentCredentialEnv ?? []),\n        ...(task.execution?.modal?.credentialEnv ?? [])\n      ])\n    );\n    publishVerifiedArtifacts","replace":"      sensitiveEnvironmentValues(process.env, [\n        ...(task.execution?.modal?.credentialEnv ?? [])\n      ])\n    );\n    publishVerifiedArtifacts","count":1}]},
{"id":"rv-secret-scan-drops-modal-credentials","target":"verifier publishes complete validated set","edits":[{"find":"        ...(task.execution?.agentCredentialEnv ?? []),\n        ...(task.execution?.modal?.credentialEnv ?? [])\n      ])\n    );\n    publishVerifiedArtifacts","replace":"        ...(task.execution?.agentCredentialEnv ?? [])\n      ])\n    );\n    publishVerifiedArtifacts","count":1}]},
{"id":"rv-secret-scan-after-publication","target":"verifier publishes complete validated set","edits":[{"find":"    assertArtifactPublicationsContainNoSecrets(\n      publications,\n      sensitiveEnvironmentValues(process.env, [\n        ...(task.execution?.agentCredentialEnv ?? []),\n        ...(task.execution?.modal?.credentialEnv ?? [])\n      ])\n    );\n    publishVerifiedArtifacts(artifactDir, publications);\n","replace":"    publishVerifiedArtifacts(artifactDir, publications);\n    assertArtifactPublicationsContainNoSecrets(\n      publications,\n      sensitiveEnvironmentValues(process.env, [\n        ...(task.execution?.agentCredentialEnv ?? []),\n        ...(task.execution?.modal?.credentialEnv ?? [])\n      ])\n    );\n","count":1}]},
{"id":"rv-epoch-check-after-marker","target":"verifier publishes complete validated set","edits":[{"find":"    publishVerifiedArtifacts(artifactDir, publications);\n    assertVerifiedDependencySnapshotEpochRemainedCurrent(task, dependencySnapshotEpoch);\n    const verificationMarker = writeArtifactVerificationMarker(task, artifacts, publications, validationWarnings);\n","replace":"    publishVerifiedArtifacts(artifactDir, publications);\n    const verificationMarker = writeArtifactVerificationMarker(task, artifacts, publications, validationWarnings);\n    assertVerifiedDependencySnapshotEpochRemainedCurrent(task, dependencySnapshotEpoch);\n","count":1}]},
{"id":"rv-semantic-gates-after-publication","target":"verifier publishes complete validated set","edits":[{"find":"    const validationWarnings = verifyOutputSemanticGates(task, verifiedOutputs, campaignEvidence);\n","replace":"","count":1},{"find":"    publishVerifiedArtifacts(artifactDir, publications);\n    assertVerifiedDependencySnapshotEpochRemainedCurrent","replace":"    publishVerifiedArtifacts(artifactDir, publications);\n    const validationWarnings = verifyOutputSemanticGates(task, verifiedOutputs, campaignEvidence);\n    assertVerifiedDependencySnapshotEpochRemainedCurrent","count":1}]},
{"id":"rv-marker-cleared-after-validation","target":"preparation requires dependency verification (marker order)","edits":[{"find":"  clearArtifactVerificationMarker(task);\n  const artifactRoots = taskArtifactRoots(task, artifactDir);\n  const capturedOutputs = capturedTaskOutputs ?? captureTaskOutputs(task);\n  const { artifacts, verifiedOutputs } = validateCapturedTaskOutputs(task, capturedOutputs);\n","replace":"  const artifactRoots = taskArtifactRoots(task, artifactDir);\n  const capturedOutputs = capturedTaskOutputs ?? captureTaskOutputs(task);\n  const { artifacts, verifiedOutputs } = validateCapturedTaskOutputs(task, capturedOutputs);\n  clearArtifactVerificationMarker(task);\n","count":1}]},
{"id":"rv-reset-deletes-workspace-patch-baseline","target":"resets exact task-owned artifact contents","edits":[{"find":"          entry === WORKSPACE_PATCH_BASELINE_FILE ||\n","replace":"","count":1}]},
{"id":"rv-reset-no-anchor-check","target":"resets exact task-owned artifact contents","edits":[{"find":"  const anchoredRoot = realpathSync(candidate);\n  if (anchoredRoot !== path.join(parent, attemptId)) {\n    throw new Error(`artifact-contract failure: unsafe ${label} task artifact root ${attemptId}`);\n  }\n","replace":"  const anchoredRoot = realpathSync(candidate);\n","count":1}]},
{"id":"rv-reset-generated-tests-first-root-only","target":"resets exact contents / generated-test dirs","edits":[{"find":"    for (const testRoot of invariantTestRoots(workspaceRoot)) {\n      const foundryParentCandidate","replace":"    for (const testRoot of invariantTestRoots(workspaceRoot).slice(0, 1)) {\n      const foundryParentCandidate","count":1}]},
{"id":"rv-retry-keeps-continue-session","target":"retries do not inspect previous failure text","edits":[{"find":"          resumeSession: undefined,\n          continueSession: false,\n          lastHeartbeat: undefined\n        };","replace":"          resumeSession: undefined,\n          lastHeartbeat: undefined\n        };","count":1}]},
{"id":"rv-retry-keeps-last-heartbeat","target":"retries do not inspect previous failure text","edits":[{"find":"          continueSession: false,\n          lastHeartbeat: undefined\n        };","replace":"          continueSession: false\n        };","count":1}]},
{"id":"rv-authority-assert-skips-identity","target":"prompt authority derived from sealed controls","edits":[{"find":"    !sameImmutableFileIdentity(captured.identity, expected.identity) ||\n    !captured.bytes.equals(expected.bytes)\n  ) {\n    throw new Error(`artifact-contract failure: prompt artifact authority was modified","replace":"    !captured.bytes.equals(expected.bytes)\n  ) {\n    throw new Error(`artifact-contract failure: prompt artifact authority was modified","count":1}]},
{"id":"rv-authority-relocated-run-root-unresolved","target":"prompt authority derived from sealed controls","edits":[{"find":"    relocatedRunRoot: realpathSync(task.runRoot),\n    admittedDependencyArtifactDirs: admission.directories,","replace":"    relocatedRunRoot: task.runRoot,\n    admittedDependencyArtifactDirs: admission.directories,","count":1}]},
{"id":"rv-repeat-generation-skips-run-metadata","target":"prompt authority (repeat-generation checks)","edits":[{"find":"      } else {\n        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportRunMetadataAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n      }","replace":"      } else {\n        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n      }","count":1}]},
{"id":"rv-repeat-generation-skips-report-prompt","target":"prompt authority (repeat-generation checks)","edits":[{"find":"      } else {\n        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportRunMetadataAuthorityUnchanged(task);\n        assertFinalReportPromptAuthorityUnchanged(task);\n      }","replace":"      } else {\n        assertPromptArtifactAuthorityUnchanged(task);\n        assertFinalReportRunMetadataAuthorityUnchanged(task);\n      }","count":1}]},
{"id":"rv-proof-ignores-remotes","target":"worktrees fail closed (source-isolation proof)","edits":[{"find":"  const remotes = git([\"remote\"]).split(\"\\n\").filter(Boolean);","replace":"  const remotes = [];","count":1}]},
{"id":"rv-proof-not-run-in-preparation","target":"worktrees fail closed (source-isolation proof)","edits":[{"find":"  preparationStep(task.attemptId, \"preserve-pinned-source-proof\", () => preservePinnedSourceProof(task));\n","replace":"","count":1}]},
{"id":"rv-proof-runs-for-unpinned-sources","target":"worktrees fail closed (source-isolation proof)","edits":[{"find":"  if (!usesPinnedSource) return;\n  const workspaceRoot = realpathSync(task.workspacePath);","replace":"  const workspaceRoot = realpathSync(task.workspacePath);","count":1}]},
{"id":"rv-worktree-base-branch-dropped","target":"worktrees fail closed (render: worktree base)","edits":[{"find":"              {...(baseBranch === undefined ? {} : { baseBranch })}\n","replace":"","count":1}]},
{"id":"rv-preserve-current-sources-always","target":"post-agent preparation preserves new sources","edits":[{"find":"      preserveCurrentSources: options.replayWorkspacePatches === false\n","replace":"      preserveCurrentSources: true\n","count":1}]},
{"id":"rv-post-agent-restores-baseline-bytes","target":"post-agent preparation preserves new sources","edits":[{"find":"  if (preserveCurrentSources) return;\n  for (const [relativePath, bytes] of snapshot) {","replace":"  for (const [relativePath, bytes] of snapshot) {","count":1}]},
{"id":"rv-patch-writer-follows-symlink","target":"guards runtime-owned workspace patch publication","edits":[{"find":"    const existing = resolveRegularArtifactFile(\n      root,\n      target,\n      \"artifact-contract failure: workspace patch artifact is unsafe\"\n    );\n    const existingContents = readFileSync(existing, \"utf8\");","replace":"    const existingContents = readFileSync(target, \"utf8\");","count":1}]},
{"id":"rv-companion-digest-unchecked","target":"rejects zero-byte generated-test companions (size/digest)","edits":[{"find":"    if (createHash(\"sha256\").update(snapshot.bytes).digest(\"hex\") !== entry.sha256) {\n      throw new Error(`artifact-contract failure: generated-test bundle file digest does not match ${relativePath}`);\n    }\n","replace":"","count":1}]},
{"id":"rv-companion-utf8-unchecked","target":"rejects zero-byte generated-test companions (utf8)","edits":[{"find":"    decodeStrictUtf8Snapshot(snapshot, `artifact-contract failure: generated-test bundle file ${relativePath}`);\n","replace":"","count":1}]},
{"id":"rv-ledger-gets-no-review-context","target":"review-stage authority","edits":[{"find":"    output.schemaFile === \"finding-lifecycle-ledger.schema.json\" ||\n    output.schemaFile === \"strategy-detections.schema.json\"\n  ) {","replace":"    output.schemaFile === \"strategy-detections.schema.json\"\n  ) {","count":1}]},
{"id":"rv-strategy-detections-gets-no-review-context","target":"review-stage authority","edits":[{"find":"    output.schemaFile === \"finding-lifecycle-ledger.schema.json\" ||\n    output.schemaFile === \"strategy-detections.schema.json\"\n  ) {","replace":"    output.schemaFile === \"finding-lifecycle-ledger.schema.json\"\n  ) {","count":1}]},
{"id":"rv-triage-gets-no-deduped-context","target":"review-stage authority","edits":[{"find":"  } else if (output.schemaFile === \"triaged-findings.schema.json\") {","replace":"  } else if (output.schemaFile === \"triaged-findings-renamed.schema.json\") {","count":1}]},
{"id":"rv-differential-gap-review-case-dropped","target":"differential siblings/ancestors","edits":[{"find":"    case \"differential-gap-review.schema.json\":\n","replace":"    case \"differential-gap-review-renamed.schema.json\":\n","count":1}]},
{"id":"rv-differential-reference-harness-case-dropped","target":"differential siblings/ancestors","edits":[{"find":"    case \"reference-harness.schema.json\":\n","replace":"    case \"reference-harness-renamed.schema.json\":\n","count":1}]},
{"id":"rv-differential-red-triage-case-dropped","target":"differential siblings/ancestors","edits":[{"find":"    case \"differential-red-triage.schema.json\":\n","replace":"    case \"differential-red-triage-renamed.schema.json\":\n","count":1}]},
{"id":"rv-invariant-hard-link-accepted","target":"invariant suite across handoffs","edits":[{"find":"        if (sourceStat.nlink !== 1) {\n          throw new Error(`artifact-contract failure: invariant suite source is hard-linked ${relativePath}`);\n        }\n","replace":"","count":1}]},
{"id":"rv-invariant-symlink-destination-accepted","target":"invariant suite across handoffs","edits":[{"find":"    if (destinationStat.isSymbolicLink()) {\n      throw new Error(`artifact-contract failure: invariant suite destination is a symlink ${relativePath}`);\n    }\n","replace":"","count":1}]},
{"id":"rv-invariant-missing-suite-root-accepted","target":"invariant suite across handoffs","edits":[{"find":"    if (expectedPaths.size > 0 && !existsSync(suiteRoot)) {\n      throw new Error(\n        `artifact handoff is missing invariant-suite sources for ${task.metadata.node.logicalNodeId}: ${dependency}`\n      );\n    }\n","replace":"","count":1}]},
{"id":"rv-invariant-manifest-not-validated","target":"invariant suite across handoffs","edits":[{"find":"  assertValidInvariantSuiteManifest(manifest);\n","replace":"","count":1}]},
{"id":"rv-preflight-envelope-not-parsed","target":"binds planned output to preflighted schema bundle","edits":[{"find":"    parseJsonValidatorPreflightSuccessEnvelope(Buffer.from(stdout, \"utf8\"));","replace":"    JSON.parse(stdout);","count":1}]},
{"id":"rv-reset-deletes-invariant-baseline","target":"resets exact contents (sibling, not pinned by deleted test)","edits":[{"find":"        (entry === INVARIANT_SUITE_BASELINE_FILE ||\n","replace":"        (false ||\n","count":1}]},
{"id":"rv-reset-deletes-patch-preparation","target":"resets exact contents (sibling, not pinned by deleted test)","edits":[{"find":"          entry === WORKSPACE_PATCH_BASELINE_FILE ||\n          entry === WORKSPACE_PATCH_PREPARATION_FILE))\n","replace":"          entry === WORKSPACE_PATCH_BASELINE_FILE))\n","count":1}]},
{"id":"rv-proof-before-submodule-step","target":"restores sealed submodules before inputs (proof after submodule step)","edits":[{"find":"    );\n  }\n  preparationStep(task.attemptId, \"preserve-pinned-source-proof\", () => preservePinnedSourceProof(task));\n","replace":"    );\n  }\n","count":1},{"find":"  if (options.pinnedSubmodules === \"verify\") {\n    preparationStep(task.attemptId, \"verify-pinned-submodules\", () =>","replace":"  preparationStep(task.attemptId, \"preserve-pinned-source-proof\", () => preservePinnedSourceProof(task));\n  if (options.pinnedSubmodules === \"verify\") {\n    preparationStep(task.attemptId, \"verify-pinned-submodules\", () =>","count":1}]},
{"id":"submodules-no-expectation","candidate":"restores sealed submodules and verifies in finalizer","edits":[{"find":"        expectation: task.pinnedSubmodules ?? undefined\n      })\n    );\n  } else {","replace":"        expectation: undefined\n      })\n    );\n  } else {","count":1}],"batch":1},
{"id":"rv-changed-tests-exclude-ignored","target":"Git-compatible ls-files invocation (changedTestTreePaths)","edits":[{"find":"      [\"ls-files\", \"--others\", \"--\", \"test\", \"tests\"]\n","replace":"      [\"ls-files\", \"--others\", \"--exclude-standard\", \"--\", \"test\", \"tests\"]\n","count":1}]},
{"id":"rv-changed-sources-exclude-ignored","target":"Git-compatible ls-files invocation (changedInvariantSourcePaths)","edits":[{"find":"      [\"ls-files\", \"--others\", \"--\", \"src\", \"contracts\"]\n","replace":"      [\"ls-files\", \"--others\", \"--exclude-standard\", \"--\", \"src\", \"contracts\"]\n","count":1}]},
{"id":"rv-changed-tests-no-untracked","target":"Git-compatible ls-files invocation (changedTestTreePaths)","edits":[{"find":"      [\"diff\", \"--name-only\", \"HEAD\", \"--\", \"test\", \"tests\"],\n      [\"ls-files\", \"--others\", \"--\", \"test\", \"tests\"]\n","replace":"      [\"diff\", \"--name-only\", \"HEAD\", \"--\", \"test\", \"tests\"]\n","count":1}]}
]
mutcheck.mjs (targeted runner: one test per mutation, plus unmutated controls)
// Targeted mutation check for the PR #1203 fix-up: each mutation is run only against the test(s) meant
// to catch it, and every (file, pattern) pair is also run unmutated as a control.
// Usage (from the repository root, with review-mutations.json and mutations.json beside it, after a build
// and `npx tsc -p tsconfig.test.json` in packages/runtime): node mutcheck.mjs <out.jsonl> [ids...]
import fs from "node:fs";
import path from "node:path";
import { spawnSync } from "node:child_process";
const RUNTIME = path.resolve(process.env.ULTRAFUZZ_RUNTIME ?? "packages/runtime");
const TEMPLATE = path.join(RUNTIME, "src/templates/smithers/workflows/workflow.tsx");
const mutations = new Map(
  [
    ...JSON.parse(fs.readFileSync("review-mutations.json", "utf8")),
    ...JSON.parse(fs.readFileSync("mutations.json", "utf8"))
  ].map((m) => [m.id, m])
);
const GWV = "generated-workflow-verifier";
const HANDOFF = "invariant-suite-handoff-durability";
const PREP = "workspace-preparation-lifecycle";
const CHECKS = [
  ["rv-reset-deletes-workspace-patch-baseline", GWV, "retry cleanup preserves only a task-owned prompt"],
  ["rv-reset-no-anchor-check", GWV, "retry cleanup preserves only a task-owned prompt"],
  ["rv-reset-deletes-invariant-baseline", GWV, "retry cleanup preserves only a task-owned prompt"],
  ["rv-reset-deletes-patch-preparation", GWV, "retry cleanup preserves only a task-owned prompt"],
  ["rv-reset-generated-tests-first-root-only", HANDOFF, "retry cleanup clears every task-owned root"],
  ["rv-differential-gap-review-case-dropped", GWV, "exact declared siblings and ancestors"],
  ["rv-differential-reference-harness-case-dropped", GWV, "exact declared siblings and ancestors"],
  ["rv-differential-red-triage-case-dropped", GWV, "exact declared siblings and ancestors"],
  ["differential-no-siblings", GWV, "exact declared siblings and ancestors"],
  ["rv-secret-scan-drops-agent-credentials", GWV, "refuses to publish the value of any credential"],
  ["rv-secret-scan-drops-modal-credentials", GWV, "refuses to publish the value of any credential"],
  ["rv-invariant-missing-suite-root-accepted", HANDOFF, "producer claims implemented sources but published no suite"],
  ["rv-changed-sources-exclude-ignored", HANDOFF, "changed-source discovery includes untracked and gitignored"],
  ["rv-changed-tests-exclude-ignored", HANDOFF, "changed-source discovery includes untracked and gitignored"],
  ["rv-changed-tests-no-untracked", HANDOFF, "changed-source discovery includes untracked and gitignored"],
  ["rv-invariant-hard-link-accepted", HANDOFF, "refuses a hard-linked source"],
  ["rv-proof-ignores-remotes", GWV, "pinned source proof rejects any previously published byte drift"],
  ["rv-proof-runs-for-unpinned-sources", GWV, "pinned source proof rejects any previously published byte drift"],
  ["submodules-no-expectation", PREP, "post-agent verify pass checks the task's pinned submodules"],
  ["rv-preserve-current-sources-always", PREP, "a pre-agent preparation pass restores inherited invariant sources"],
  ["post-agent-deletes-new-sources", PREP, "a pre-agent preparation pass restores inherited invariant sources"]
];
const only = process.argv.slice(3);
const outFile = process.argv[2];
const escape = (value) => value.replace(/[.*+?^${}()|[\]\\#]/gu, "\\$&");
const pristine = fs.readFileSync(TEMPLATE, "utf8");
const restore = () => fs.writeFileSync(TEMPLATE, pristine);
for (const signal of ["SIGINT", "SIGTERM", "SIGHUP"])
  process.on(signal, () => {
    restore();
    process.exit(1);
  });
function run(file, pattern) {
  const result = spawnSync(
    "eatmydata",
    ["node", "--test", "--test-reporter=tap", `--test-name-pattern=${escape(pattern)}`, `dist-test/test/${file}.test.js`],
    { cwd: RUNTIME, encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 900_000 }
  );
  const pass = Number(/^# pass (\d+)/mu.exec(result.stdout)?.[1] ?? -1);
  const fail = Number(/^# fail (\d+)/mu.exec(result.stdout)?.[1] ?? -1);
  const firstError = /^\s+error: (.*)$/mu.exec(result.stdout)?.[1];
  return { status: result.status, pass, fail, firstError };
}
try {
  const checks = CHECKS.filter(([id]) => only.length === 0 || only.includes(id));
  const controls = [...new Set(checks.map(([, file, pattern]) => JSON.stringify([file, pattern])))];
  for (const control of controls) {
    const [file, pattern] = JSON.parse(control);
    const result = run(file, pattern);
    const record = { control: true, file, pattern, ...result, ok: result.status === 0 && result.pass >= 1 && result.fail === 0 };
    fs.appendFileSync(outFile, `${JSON.stringify(record)}\n`);
    console.log("CONTROL", record.ok ? "PASS" : "FAIL", file, "|", pattern, JSON.stringify(result));
  }
  for (const [id, file, pattern] of checks) {
    const mutation = mutations.get(id);
    let mutated = pristine;
    let error;
    for (const edit of mutation.edits) {
      const occurrences = mutated.split(edit.find).length - 1;
      if (occurrences !== (edit.count ?? 1)) {
        error = `expected ${edit.count ?? 1} occurrence(s), found ${occurrences}`;
        break;
      }
      mutated = mutated.split(edit.find).join(edit.replace);
    }
    if (error) {
      fs.appendFileSync(outFile, `${JSON.stringify({ id, file, pattern, error })}\n`);
      console.log(id, "ERROR", error);
      continue;
    }
    fs.writeFileSync(TEMPLATE, mutated);
    const result = run(file, pattern);
    restore();
    const record = { id, file, pattern, ...result, caught: result.status !== 0 && result.fail >= 1 };
    fs.appendFileSync(outFile, `${JSON.stringify(record)}\n`);
    console.log(id, record.caught ? "CAUGHT" : "SURVIVED", "|", file, "|", (result.firstError ?? "").slice(0, 160));
  }
} finally {
  restore();
}

@aviggiano
aviggiano merged commit 142ba80 into main Sep 29, 2026
17 checks passed
@aviggiano
aviggiano deleted the claude/v09a-runtime-test-quality branch September 29, 2026 13:49
aviggiano added a commit that referenced this pull request Sep 29, 2026
Brings in #1202 (modal) and #1203 (runtime tests). Neither touches the
files this branch changes, and the merge had no conflicts.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…onment secret scan

The rebase deleted #1203's test that the generated verifier refuses to
publish a configured credential's value, together with the harness's
`scanProcessEnvironment` option, because that test covered the cloud
credential-name lists this PR removes. That option was the only path that
passed the real `sensitiveEnvironmentValues` into `verifyArtifacts`, so the
pre-publication scan of environment values, which still runs for every
local task, lost its only test: replacing
`sensitiveEnvironmentValues(process.env)` with `[]` in the template's
`verifyArtifacts` left all 115 verifier tests passing. The only remaining
verifier secret test uses a vendor-format token, which pattern matching
catches without any environment value.

The harness option is back, and a local test replaces the deleted one. It
writes a value with no vendor format into `result.json`. Held only by
ULTRAFUZZ_TEST_GATEWAY_LABEL, a name the credential heuristic does not
match, the value is published; once ULTRAFUZZ_TEST_GATEWAY_API_KEY also
holds it, `verifyArtifacts` throws "contains sensitive data" and publishes
nothing and writes no marker. With the template's scan emptied as above,
the new test fails.

Two agent-failure redaction tests no longer configure any credential name
(the task-level lists were cloud-only), so their names now say they redact
credential-named environment values, which is what they check.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…onment secret scan

The rebase deleted #1203's test that the generated verifier refuses to
publish a configured credential's value, together with the harness's
`scanProcessEnvironment` option, because that test covered the cloud
credential-name lists this PR removes. That option was the only path that
passed the real `sensitiveEnvironmentValues` into `verifyArtifacts`, so the
pre-publication scan of environment values, which still runs for every
local task, lost its only test: replacing
`sensitiveEnvironmentValues(process.env)` with `[]` in the template's
`verifyArtifacts` left all 115 verifier tests passing. The only remaining
verifier secret test uses a vendor-format token, which pattern matching
catches without any environment value.

The harness option is back, and a local test replaces the deleted one. It
writes a value with no vendor format into `result.json`. Held only by
ULTRAFUZZ_TEST_GATEWAY_LABEL, a name the credential heuristic does not
match, the value is published; once ULTRAFUZZ_TEST_GATEWAY_API_KEY also
holds it, `verifyArtifacts` throws "contains sensitive data" and publishes
nothing and writes no marker. With the template's scan emptied as above,
the new test fails.

Two agent-failure redaction tests no longer configure any credential name
(the task-level lists were cloud-only), so their names now say they redact
credential-named environment values, which is what they check.

Refs #134

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant