test(runtime): delete vacuous and source-text tests, keep behavioural coverage - #1203
Conversation
…ection tests lifecycle-inspection.test.ts launched a fresh run for 38 of its 52 tests, and each launch seals a full execution snapshot. 29 of those tests only run read-only inspection commands (why, timeline, snapshots, events, node and a cancel the runner rejects) against the fake runner, and the five diagnoseProject tests never read the run they launched. The fake runner now reads every answer from files that writeInspectionFixtures resets before each test, so one launched run serves the read-only tests. The diagnoseProject tests use a project with no run, and the four tests that record a cancel or rewrite run metadata keep their own run. Launches go from 38 to 5 and no assertion changes. A digest of every file of the shared run, taken after launch and again after the whole file ran, found no change. Co-Authored-By: Claude Opus 5.5 <[email protected]>
- agent-adapter-boundaries: "OpenRouter retains the manually reviewed argv responsibility inherited from Codex" read a value from this file's own adapterPolicies table, so only an edit to that test data could fail it. - prompt-artifact-authority: compared promptArtifactAuthorityJsonSchema.$id with the constant the $id is defined from. - final-report-markdown: the second call passed an extra argument under @ts-expect-error. JavaScript ignores it, so the call repeated the first one; at compile time it only pinned the parameter list of isDirectiveConformingFinalReportMarkdown, which no production code calls. - runtime-test-shard: the [1, 2, 3, 4].filter assertion follows from the range and determinism assertions before it. - dynamic-workflow: the `smithers graph` branch ran only when .smithers/node_modules/smithers-orchestrator existed. That is the Smithers 0.32 package name, which the dependency migration removes, so the branch never ran. The test is renamed to what it checks. - runtime.test.ts: "reported CLI cost compatibility patches validate and prefer the adapter estimate" matched both patch strings with regexes. "reported CLI cost compatibility preserves zero-token adapter estimates" applies the same patches to the pinned engine and fails when either patched behaviour is removed. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ioural ones generated-workflow-verifier.test.ts pinned workflow.tsx with regexes, indexOf positions and identifier names. To find which of those tests were the only guard of real behaviour, the template was mutated 93 ways. With all 40 candidate tests deleted, the 18 test files that run template functions then ran against each mutation, and 68 mutations were caught. runtime.test.ts, lifecycle-inspection and dynamic-lifecycle, which only render the template through startRun, were not run. Where a survivor showed a behaviour no test checked and an existing harness could observe it, the harness now does: - New: artifact-aware agents recheck all four task-local authorities after the model returns, whether it succeeds or fails. - New: retry cleanup clears the canonical, mirror and generated-test roots before it re-prepares the attempt and its prompt authority. - New: the finalizer publishes nothing for an unsuccessful agent, clears a stale marker, and verifies pinned submodules without restoring them. - New: a task worktree must sit at the recorded launch commit, through HEAD and through its ref. - New, replacing a regex over one line: a task prompt prefers its relocatable prompt path. - New, replacing a test that ran git itself: invariant discovery lists tracked, untracked and gitignored sources through the template's own function. - The verifyArtifacts harness records marker clears and takes the companion list, so failing verification must clear the marker and companions must be published. - One case each on the prompt-authority, property-lens, review-context, protected-baseline and source-root tests. 30 tests are deleted: every mutation of what they pinned is now caught, equivalent, or not caught by the test itself either. The 10 tests that catch a mutation surviving the rest of the suite stay; one of them, the preparation step-name test, is the only guard of two preparation steps pinned by tests deleted here. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…tions lost Review applied 41 more mutations to workflow.tsx, aimed at what the deleted source-text tests pinned. Four of those tests turned out to be the only guard of behaviour this change's own 93 mutations never touched. Each loss is restored as a behavioural test rather than by bringing a source-text test back: - the canonical retry reset keeps its runtime-owned evidence files and refuses a reset root replaced by a symlink; - each differential lane schema gets exactly its declared sibling and ancestor artifacts, and they satisfy every contextual gate it has (this also closes the differential-no-siblings gap); - the pre-publication secret scan covers a cloud task's configured credential names; - the invariant-suite handoff fails closed for a producer that claims sources but published none, changed-source discovery keeps untracked and gitignored sources, and test-tree discovery refuses a hard-linked source; - a pinned worktree with a remote is refused before any source proof is kept; - the post-agent verify pass checks the task's pinned submodules (the recorded catch of that mutation was spurious); - a pre-agent preparation pass restores inherited invariant sources from the workspace snapshot. The shard test again asserts membership, which the range and determinism checks do not imply. lifecycle-inspection now checks that shared tests leave the shared run's documents unchanged, and reports a failed shared launch as that launch failing. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Mutation evidence for section 3: the 93 mutations and their runnerThese are the 93 mutations of
|
Mutation checks for the review follow-up (d3e171c)Review applied 42 more mutations to the template, aimed at what the deleted source-text tests pinned: 41 new ones, plus a re-run of this PR's Each mutation below was run against the one test meant to catch it. That test fails with the mutation applied, and passes on the unmodified template: 21 of 21 caught, 10 of 10 controls pass. Separately, the shard test fails on review's fractional mapping
The review mutations this commit does not restore are listed under "Known residuals from review" in the PR description, with the reason each is low-exposure. To re-run: save the two files below next to
|
Brings in #1202 (modal) and #1203 (runtime tests). Neither touches the files this branch changes, and the merge had no conflicts. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…onment secret scan The rebase deleted #1203's test that the generated verifier refuses to publish a configured credential's value, together with the harness's `scanProcessEnvironment` option, because that test covered the cloud credential-name lists this PR removes. That option was the only path that passed the real `sensitiveEnvironmentValues` into `verifyArtifacts`, so the pre-publication scan of environment values, which still runs for every local task, lost its only test: replacing `sensitiveEnvironmentValues(process.env)` with `[]` in the template's `verifyArtifacts` left all 115 verifier tests passing. The only remaining verifier secret test uses a vendor-format token, which pattern matching catches without any environment value. The harness option is back, and a local test replaces the deleted one. It writes a value with no vendor format into `result.json`. Held only by ULTRAFUZZ_TEST_GATEWAY_LABEL, a name the credential heuristic does not match, the value is published; once ULTRAFUZZ_TEST_GATEWAY_API_KEY also holds it, `verifyArtifacts` throws "contains sensitive data" and publishes nothing and writes no marker. With the template's scan emptied as above, the new test fails. Two agent-failure redaction tests no longer configure any credential name (the task-level lists were cloud-only), so their names now say they redact credential-named environment values, which is what they check. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…onment secret scan The rebase deleted #1203's test that the generated verifier refuses to publish a configured credential's value, together with the harness's `scanProcessEnvironment` option, because that test covered the cloud credential-name lists this PR removes. That option was the only path that passed the real `sensitiveEnvironmentValues` into `verifyArtifacts`, so the pre-publication scan of environment values, which still runs for every local task, lost its only test: replacing `sensitiveEnvironmentValues(process.env)` with `[]` in the template's `verifyArtifacts` left all 115 verifier tests passing. The only remaining verifier secret test uses a vendor-format token, which pattern matching catches without any environment value. The harness option is back, and a local test replaces the deleted one. It writes a value with no vendor format into `result.json`. Held only by ULTRAFUZZ_TEST_GATEWAY_LABEL, a name the credential heuristic does not match, the value is published; once ULTRAFUZZ_TEST_GATEWAY_API_KEY also holds it, `verifyArtifacts` throws "contains sensitive data" and publishes nothing and writes no marker. With the template's scan emptied as above, the new test fails. Two agent-failure redaction tests no longer configure any credential name (the task-level lists were cloud-only), so their names now say they redact credential-named environment values, which is what they check. Refs #134 Co-Authored-By: Claude Opus 5.5 <[email protected]>
Refs #1000
Problem
Two kinds of runtime test cost buy no protection.
generated-workflow-verifier.test.tsreadsworkflow.tsxas text and matches regexes,indexOfpositions or identifier names. These tests fail when a function moves or is reformatted, and they pass when the code they point at is broken. A few other assertions compare a value with its own definition, repeat the assertion before them, or sit in a branch that never runs.lifecycle-inspection.test.tslaunched a full run (startRun, which seals a complete execution snapshot) for 38 of its 52 tests. 29 of them only issue read-only inspection commands against the fake runner, and the 5diagnoseProjecttests never read the run at all. The file runs in theruntime-supportingrelease lane, whose comment says runner contention can push that lane past 75 minutes.Root cause
clearArtifactVerificationMarkerandverifyGeneratedTestFileswere no-ops in theverifyArtifactsharness, and the review-context stub ignoreddirectOnly.lifecycle-inspection.test.tsbaked each test's fixture answers into the fake runner script it wrote before launching, so every test needed its own launched run.Change
Test-only; no file under
src/changes. Four commits: the three below, and a review follow-up (section 4) that restores the guards review found lost.1.
lifecycle-inspection.test.ts: one shared run for read-only commandswhy,timeline,snapshots,node,events,cancel) from files.writeInspectionFixturesrewrites those files before each test, and clears the control files and the command log.inspectedRun. That includes "diagnoseRun keeps a runner path intact in public text", which fix: observers and benchmark harnesses survive transient failures and report real paths #1194 added with its own launch for a singlewhycall.diagnoseProjecttests use a project with no run (projectWithFakeRunner).run.jsonkeep their own run (launchedProject).inspectedRuncall first checks thatrun.json,state.json,events.jsonlandsmithers/workflow-run-link-journal.jsonstill hold the bytes they had right after the shared launch. A future test that changes the shared run therefore fails the next shared test, with a message saying such a test must launch its own run. A shared launch that fails is reported by every shared test as that launch failing, not as its own regression.observeOnlycreates and removes).2. Assertions that cannot catch a regression
adapterPoliciestable listsargv-constructionforopenrouter.tsxpromptArtifactAuthorityJsonSchema.$id === PROMPT_ARTIFACT_AUTHORITY_JSON_SCHEMA_ID(assertion deleted)$idequals a constant$idis defined as that constantfalseunder@ts-expect-errorstill returns trueisDirectiveConformingFinalReportMarkdown, which no production code calls[1,2,3,4].filter(… === shard(name))equals[first](replaced in section 4)[1, 2, 3, 4].includes(first)smithers graphbranch (branch deleted, test renamed)smithers graph.smithers/node_modules/smithers-orchestratorexisted. That is the Smithers 0.32 package name, which the dependency migration removed, so the branch never ran3. Source-text tests of
workflow.tsx: 30 deleted, 10 kept, coverage added firstHow each candidate was judged.
workflow.tsx, each breaking one behaviour that one of those tests pins.workflow.tsxand run code from it. The 13 fast ones ran first, then 5 lifecycle and integration files if nothing had failed.runtime.test.ts,lifecycle-inspectionanddynamic-lifecycleonly render the template throughstartRun, so they were not run against mutations.The 25 survivors were then run against the full base
generated-workflow-verifier.test.ts, which shows which candidates catch each one:Of the 40 candidates, 10 stay, 29 are deleted, and the relocatable-prompt-path test is rewritten in place as a behavioural test. One test that is not a source-text test is also deleted: it ran
git ls-fileswith its own arguments, so it re-implemented the call under test. That makes 30 deletions.Review found this sample too thin for the broad tests. Review applied 41 more mutations aimed at what the deleted tests pinned. The deleted tests caught 39 of them and the branch 13; about 13 of the rest were real losses. Review also found one of this PR's recorded catches (
submodules-no-expectation) to be spurious. Section 4 restores each loss with a behavioural test, and the per-test rows below say which test now covers what review found.Behavioural coverage added before deleting. Where a survivor pointed at a behaviour no test checked, and an existing harness could run the real template function, the harness now observes it. Each addition fails on the mutations it targets and passes on the unmodified template.
generate; prompt, report run-metadata and report-prompt authority aftergenerate; the same on the failure path (6)invariant-suite-handoff-durability)pinnedSubmodules: "verify"(3)[])git ls-fileswith its own arguments: "invariant source discovery lists tracked, untracked and gitignored sources under every supported root", run through the template'sinvariantWorkspaceSourcePaths--exclude-standardadded,--othersdropped (2)src/.envrc,tests/.npmrc).envrcrule (1)Per-test evidence.
verifyGeneratedTestFilesand its snapshot reader contain empty-file and size/digest checkscompanions-size-mismatch-accepted), e.g. by "generated Smithers authenticates the final generated-test publication snapshot".companions-accept-emptyis equivalent: the schema requiressize_bytes >= 1, so an empty companion still fails the size checkseverity-triaged-transitive,severity-triaged-path-dropped), e.g. by "generated semantic context projects every required review authority field"review-context-missing,review-severity-authority-missing), e.g. by "1091: generated verifier publishes 41 detections with six family omissions and durable warnings"; "generated semantic context projects every required review authority field" (+1 more)differential-no-siblingssurvives with this test present too: its identifier regexes still pass whensiblingDifferentialBindingsreturns no bindings. Review: its schema-name regexes were the only check that each lane schema reaches its context builder; renaming acaselabel makes that schema's reconciliation gate fail every node that produces it. Both are now caught by the new "generated differential verification gives each lane schema its exact declared siblings and ancestors"indexOforder of the submodule stepssubmodules-never-verifyby "the verifier publishes nothing for an unsuccessful agent and verifies pinned submodules without restoring them",inputs-asserted-before-submodulesby "#1081 complete preparation regenerates every pre-agent store after a dependency replacement" (+3 more). Corrected by review: the recorded catch ofsubmodules-no-expectationwas spurious (the failing test's fixture has no pinned submodules), and this test missed it too, because its regex matches thehydratecall. It is now caught by the extended "the post-agent verify pass checks the task's pinned submodules…". MovingpreservePinnedSourceProofahead of the submodule step is caught only by this test's order check; that is left as a residualschema-binding-ignores-bundle), e.g. by "generated task preparation binds output schema content but not the validator build".no-schema-binding-checkis caught by the kept "generated Smithers preparation names its failing step and carries a retry budget". This test misses it: its order checkindexOf(call) < indexOf(preflight)still passes when the call is removed, becauseindexOfreturns -1doesNotMatchof patch file names insideprepareArtifactMirrorprepare-precreates-outputs), e.g. by "#1081 complete preparation regenerates every pre-agent store after a dependency replacement"; "#1115 an interrupted workspace replacement still binds property-only dependency markers"writeWorkspacePatchArtifactpatch-publication-always-replaces,patch-publication-default-replace), e.g. by "#357 a capture against a different pinned commit is still rejected"; "#357 a foreign schema version is not treated as a superseded capture" (+4 more)prompt-authority-no-admission,prompt-authority-no-reread-check,reset-no-prompt-authority,agent-post-prompt-authority,agent-no-repeat-generation-checks,agent-catch-authority), e.g. by "generated task-local prompt authority is minimized, tamper-evident, and restored for retries"; "retry cleanup clears every task-owned root before it re-prepares the attempt and its authorities"source-revision-accepts-any-head,source-revision-ignores-ref,agent-no-source-revision-check), e.g. by "a task worktree must still sit at the recorded launch commit, through its ref as well as HEAD"; "agent retries are error-agnostic fresh generations with Smithers' effective prompt" (+1 more).prepare-skips-source-revisionis caught by the kept "generated Smithers preparation names its failing step and carries a retry budget". Review: none of the four touchedpreservePinnedSourceProof, which is the only check that a pinned worktree has no remote, and whose early return for unpinned runs no unit or integration test exercised. Both are now in the extended "generated Smithers pinned source proof rejects any previously published byte drift". Dropping the worktree base branch is caught byruntime.test.ts's rendered-text testretry-keeps-resume-session,retry-attempt1-unscrubbed,retry-keeps-messages,retry-uses-original-prompt,retry-injects-failure-text,agent-post-admission), e.g. by "a resumed activation's first dispatch carries no stale continuation pointer"; "agent retries are error-agnostic fresh generations with Smithers' effective prompt" (+1 more)resetTaskArtifactsForRetryandresetTaskArtifactContentsreset-no-canonical,reset-no-mirror,reset-deletes-preserved-prompt,agent-no-reset-before-generation), e.g. by "agent retries are error-agnostic fresh generations with Smithers' effective prompt"; "final-report prompt authority is bounded, tamper-evident, and constant-size across large projections" (+4 more).reset-no-basename-guardis equivalent for an existing root, because theanchoredRoot !== path.join(parent, attemptId)check rejects the same root; for a missing root nothing is deleted either way. Review: none of the five touched the canonical preserve list or that anchor check, and this test was their only guard. Withoutworkspace-patch-baseline.jsonin the list, every patch-publishing task fails before its model runs; without the anchor check, a reset root replaced by a symlink has its target emptied. Both are now in the extended "retry cleanup preserves only a task-owned prompt and accepts a sealed snapshot prompt"preserveCurrentSourcesflagpost-agent-deletes-new-sources), e.g. by "#1081 complete preparation regenerates every pre-agent store after a dependency replacement". Review: the opposite direction, the flag set on every pass, survived everything else. A pre-agent pass restores the workspace preparation tree first, so the flag shows only on sources inherited after that tree was captured: with it set, a preparation retry or a durable resume leaves the inherited invariant-suite source out of the worktree. Now caught by the new "a pre-agent preparation pass restores inherited invariant sources…"doesNotMatchof repair helpers, andretries={0}on the verifieragent-post-admission,agent-post-run-metadata-authority,agent-post-report-prompt-authority,agent-pre-generate-admission,agent-output-schema-leak,agent-task-runtime-leak), e.g. by "artifact-aware agents own process completion without constraining terminal responses"; "artifact-aware agents recheck every task-local authority after the model returns, succeeding or failing" (+2 more).local-verifier-retriessurvives with this test present too: itsretries={0}regex matches the first verifier in the file, the cloud one, so a local verifier withretries={2}passes itdoesNotMatchof helper names removed in #558doesNotMatchof helper names removed in #558doesNotMatchof helper names removed in #558doesNotMatchof helper names and v2 field names removed in #558doesNotMatchof helper names removed in #558doesNotMatchof helper names removed in #558doesNotMatchof helper names removed in #558doesNotMatchof removed helper names and a regex over a code commentreset-no-generated-test,reset-only-logical-generated-test), e.g. by "#212 retry cleanup resets generated tests under the repository's plural tests/ root"invariant-dependency-change-accepted,invariant-baseline-tamper-accepted,invariant-ancestor-conflict-accepted), e.g. by "#213 the protected baseline outranks a sidecar the agent rewrote"; "#219 the durable handoff record fails closed when a recorded dependency no longer matches" (+4 more).invariant-ancestor-order-flippedis equivalent since #315: sources are resolved per path over the set of publishers, not by visit order. Review: 4 mutations were too few for about 60 regexes. This test was the only guard of the fail-fast for a producer that claims implemented sources but published no suite, and of the hard-link refusal in test-tree discovery; both now have tests ininvariant-suite-handoff-durability. Its other two review survivors are not losses: the symlink-destination check (the canonical-path check after it still throws) and the manifest self-validation (consumers re-validate)ls-filesargument listsls-files-excludes-ignored,ls-files-tracked-only), e.g. by "#323 a listing past Node's 1 MB default is enumerated in full under the template's own bound"; "#323 an enumeration that outgrows its capture buffer names the subcommand, the bound and the roots" (+3 more). Review: those two mutations cover onlyinvariantWorkspaceSourcePaths. Adding--exclude-standardtochangedInvariantSourcePaths, or to the test tree's git fallback, or dropping that fallback's--others, survived; gitignored agent-authored sources would then drop out of the published suite. All three are now caught by the new "changed-source discovery includes untracked and gitignored sources under every supported root"enumeration-unbounded), e.g. by "#323 a listing past Node's 1 MB default is enumerated in full under the template's own bound"; "#323 an enumeration that outgrows its capture buffer names the subcommand, the bound and the roots" (+4 more)assertSafeInvariantSuitePathinvariant-any-root,invariant-envrc-allowed), e.g. by "#211 invariant suite provenance is restricted to supported source roots"; "#214 an unreferenced file under invariant-suite/ fails the publication closed"indexOforder oververifyArtifactsverifier-no-secret-scan,verifier-no-epoch-check,verifier-marker-before-publication,verifier-no-companion-publication,verifier-no-semantic-gates), e.g. by "1091: generated verifier publishes 41 detections with six family omissions and durable warnings"; "generated Smithers binds generated-test manifests to the current run and logical producer" (+13 more). Review: none of the five touched the scan's configured credential names. DroppingagentCredentialEnvormodal.credentialEnvsurvived everything else, so an opaque credential under a name the heuristic does not recognise would be published. Both are now caught by the new "generated Smithers verifier refuses to publish the value of any credential the task is configured with"verifyArtifactsverifier-keeps-stale-marker), e.g. by "generated Smithers verifier rejects invalid UTF-8 and duplicate JSON keys from captured bytes"git ls-fileswith its own argument list, never the template's discoveryls-files-excludes-ignored,ls-files-tracked-only), e.g. by "#323 a listing past Node's 1 MB default is enumerated in full under the template's own bound"; "#323 an enumeration that outgrows its capture buffer names the subcommand, the bound and the roots" (+3 more). It re-implemented the call under testcompanions-select-by-path,expectations-select-by-pathprepare-no-parent-dirs,agent-not-after-preparationagent-gets-unadmitted-dirs,optional-never-unavailable,ancestor-outputs-not-filtered-by-admissionverifier-drops-agent-outputbaseline-tree-unchecked,baseline-not-publishedhandoff-ignores-declared-rootsprojection-report-not-comparedsnapshot-run-root-not-canonical,stale-cleanup-before-reset,stale-cleanup-deletes-runtime-rootspreparation-evidence-not-requiredno-schema-binding-check,prepare-skips-source-revisionCoverage gaps found, not fixed here. These mutations survive the 18 files and the full base
generated-workflow-verifier.test.ts, on the base and on this branch:local-verifier-retries: the local verifier task getsretries={2}instead of0.coverage-by-default-path: implemented-property coverage is refused unless it uses the default file name.verifier-no-final-report-projection:verifyArtifactsskipsverifyFinalReportCanonicalProjection.Closing them needs harnesses around the rendered task tree, contract-based input selection and the final-report branch of
verifyArtifacts. Section 4 closes the fourth gap listed here before review,differential-no-siblings, and two that no test caught before this PR either (the canonical reset deletinginvariant-suite-baseline.jsonorworkspace-patch-preparation.json).Known residuals from review. These also survive every branch test and were caught only by deleted tests, but review judged them low-exposure or near-equivalent, so they are not restored:
relocatedRunRootwithoutrealpathSync: breaks only cloud tasks, whose run root is relative; v04 removes cloud execution.preservePinnedSourceProofmoved ahead of the submodule step.4. Review follow-up: guards the deletions lost
Review found that four deleted tests were the only guard of behaviour this PR's mutations never touched, plus smaller losses elsewhere. Each is restored as a behavioural test rather than by bringing back the source-text test. Every template row was checked the same way: the listed mutations fail the named test, and the unmodified template passes it. The shard row was checked with review's fractional mapping applied to the compiled helper, and the last row with a probe test that appends to the shared run's event log (the next shared test fails). The mutation ids are review's (
rv-…) unless marked as this PR's; the files and runner are attached in this comment.invariant-suite-baseline.json,workspace-patch-baseline.jsonandworkspace-patch-preparation.json; a reset root replaced by a symlink is refused and its target is left intact (4:rv-reset-deletes-workspace-patch-baseline,rv-reset-no-anchor-check, andrv-reset-deletes-invariant-baseline/rv-reset-deletes-patch-preparation, which no test caught before)test/andtests/when both exist (1:rv-reset-generated-tests-first-root-only)config/topologies/exhaustive.yml)case-label renames, and this PR'sdifferential-no-siblings)verifyArtifactsharness can opt into the real environment scan)agentCredentialEnvormodal.credentialEnvnames it (2:rv-secret-scan-drops-agent-credentials,rv-secret-scan-drops-modal-credentials)rv-invariant-missing-suite-root-accepted)changedInvariantSourcePaths, and the test tree's git fallback (3:rv-changed-sources-exclude-ignored,rv-changed-tests-exclude-ignored,rv-changed-tests-no-untracked)rv-invariant-hard-link-accepted)rv-proof-ignores-remotes,rv-proof-runs-for-unpinned-sources)submodules-no-expectation)rv-preserve-current-sources-always, and this PR'spost-agent-deletes-new-sources)[1, 2, 3, 4].includes(first)(u32 / 2**32) * (total - 1) + 1)lifecycle-inspection:inspectedRunchecks the shared run is unchanged (section 1)The configured credential lists are non-empty only for cloud tasks:
cloudAgentCredentialEnvreturns[]for local execution, andmodal.credentialEnvis the Modal provider's configuration. That is where they matter: a cloud task's list also names its agents' route and allowlisted variables, which the name heuristic need not recognise. Stock local agents must use canonical credential names, which it does recognise. v04 removes cloud execution and these lists with it (see Merge interactions).Deliberately not built
No product code changes. Where a deletion leaves product code with no user, it is noted here, not removed:
promptArtifactAuthorityJsonSchemais no longer referenced by any test. It stays exported throughsrc/index.ts.isDirectiveConformingFinalReportMarkdownhas no production caller.Real-engine execution of a compiled dynamic workflow, which the deleted
smithers graphbranch claimed to check. It needs the sealed execution snapshot that onlystartRunbuilds; a probe without one fails withCannot find module …/modules/@ultrafuzz/runtime/dist/index.js. That path is changing in v10.Consolidating the multi-launch tests in
runtime.test.ts: "syncRun accepts the pinned 0.35.0 usage payload…", "syncRun seals reproducible verifier output failures…" and "ordinary resume tolerates link-journal gaps…". Sharing one run there needs either an exported usage parser (a product change) or resetting run-root documents between cases, which ties the test to the run layout.Fixtures that encode behaviour the analysis calls a bug. The spec lists these for removal, but I kept them:
lifecycle-inspection;Each is the only test of what the product does today, so deleting it would remove coverage without fixing anything. Each should change in the PR that changes the behaviour.
Regex-over-patch tests with no behavioural twin stay:
event_probe_index,engine_agent_event_ownership, and theengine_agent_usage_progressordering. So do three whole-file scans that behave like lint rules and catch a real class of bug:#691(every git capture statesmaxBuffer), the compatibility-patch helper declarations, and the shard-wrapper import check.Verification
Mutation evidence for section 3 was gathered as described there, in two scratch worktrees at
integration/wave1(a48ad9c). The scripts are not committed; the 93 mutations and their runner are attached in this comment, and the review mutations with the targeted runner used for section 4 in this one.workflow.tsxon main has lost only four lines of dead code (retryFailureTemplate, refactor(runtime): delete dead controller-generation, sealed-refresh, re-finalization and recovery-authority code #1193), and all 93 mutations still apply to it unchanged.Test files, with
eatmydata node --testaftertsc -p tsconfig.test.json, on this branch:lifecycle-inspectiongenerated-workflow-verifierinvariant-suite-handoff-durabilityworkspace-preparation-lifecycleinvariant-suite-enumeration-overflowdynamic-workflowagent-adapter-boundariesfinal-report-markdownprompt-artifact-authorityruntime-test-shardruntime.test.ts, by patternThe full
runtime.test.tssuite was not run locally; it takes hours.Cost-patch deletion (section 2): I mutated the compiled
smithers.jsthree ways: without the finite/non-negative guard at both anchors, without it in the patched constant alone, and without the early return. The remaining behavioural cost test failed each time.lifecycle-inspectiontime (theruntime-supportinglane), measured locally witheatmydataon a shared 32-core host:startRunlaunchesorigin/main2cacf4c, run one after the other (base at load 11–28, branch at 28–18): wallintegration/wave1(37 launches on the base), both started together at load ~40, so contended: wallintegration/wave1, one after the other at load 5–40: wallnode --testcall. This file's wall-time saving shortens the lane only while it is the slowest file; the CPU saving applies either way.generated-workflow-verifiertakes about 5 s before and after; its change is about quality, not time.Gates:
npx prettier --checkandnpx eslinton the changed files.CI=1 ESLINT_PLUGIN_DIFF_COMMIT=2cacf4ca pnpm -w lint:strict:ci(at d3e171c too, withorigin/mainat 2cacf4c).pnpm -w lint,pnpm --filter @ultrafuzz/runtime typecheck,pnpm -w knip, before and after the review follow-up.verifyCoverageProductionInventory, which is product code, and the most complex function in the changed files scores 64.Base. The branch was rebased from
integration/wave1ontoorigin/mainat 2cacf4c (fix(runtime): attempts abandoned by a crash are recorded, so stats and status agree #1200). The only change to the change set is converting fix: observers and benchmark harnesses survive transient failures and report real paths #1194's new test (section 1). The changed test files were re-run after the rebase.Risk / compatibility
src/, no schema, config, docs or CHANGELOG.lifecycle-inspection. The tests that useinspectedRunshare one run, so a new test that changes that run's files would affect later tests. The helper's comment says which tests must launch their own run: those that record a cancel or rewrite run metadata.inspectedRunenforces that for the run's metadata, state, event log and link journal, so a violation fails the next shared test. Fixtures, control files and the command log are reset before each test.git merge-treeof each open branch's own commits against this branch, and re-checked at d3e171c: the review follow-up adds no textual conflict with any of them. v01, v02, v03, v06, v07, v08 and v09b are already in the base, and fix(modal): keep the model-work flag when a model node has no run-state record #1202 merges cleanly.workflow.tsx.exports,types,duplicates) passes with this change applied, becausepromptArtifactAuthorityJsonSchemais still exported throughsrc/index.ts.generated-workflow-verifier.test.ts. v04 edits "…publishes the complete validated set before task success" and "…preparation requires a successful dependency artifact verification", which this PR deletes. Resolve it by keeping the deletion.import { z } from "zod/v4"from the same file along with its last user, and this PR's new finalizer test usesz. The merged file then fails to compile (TS2304: Cannot find name 'z'), so keep the import.sensitiveEnvironmentValues(process.env)). Section 4's "generated Smithers verifier refuses to publish the value of any credential the task is configured with" therefore fails on the merge; delete it in v04 together with the lists. With the import kept, the merge of d3e171c and v04 passes 115 of 116 ingenerated-workflow-verifier, the one failure being that test. It passes 41/41 ininvariant-suite-handoff-durability, 8/8 inworkspace-preparation-lifecycle, 6/6 ininvariant-suite-enumeration-overflowand 5/5 inruntime-test-shard.diagnoseProjecthunks oflifecycle-inspection.test.ts. Resolve by taking v10's side of each hunk. v10's versions no longer use the replaced launch lines, and the one test that still launches ("…reports the installed runner that commands after launch execute") needs its run.Changelog entry
Runtime tests:
lifecycle-inspectionshares one launched run across its read-only tests, cutting its launches from 38 to 5, and checks that those tests leave the run unchanged. 30 tests that pinned the workflow template's source text (or re-ran its git call themselves) and five redundant, tautological or unreachable assertions are removed. Mutation testing of the template decided which source-text tests stay (10), and 13 new and 9 extended behavioural tests cover what the removed ones pinned, including what review's further mutations found only the removed tests had guarded.🤖 Generated with Claude Code