Skip to content

fix: a validator rebuild no longer strands in-flight runs (validator build becomes provenance) - #1188

Merged
aviggiano merged 7 commits into
mainfrom
claude/w24-validator-build-provenance
Sep 29, 2026
Merged

aviggiano merged 7 commits into
mainfrom
claude/w24-validator-build-provenance

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

VALIDATOR_BUILD_IDENTITY is a SHA-256 over five compiled validator modules (json-file-validator, json-schema-validator, json-validation-worker, schema-registry, strict-json) plus the ajv/ajv-formats versions. Planning records it, with the contract digest and the schema binding, in every planned output. Readers then required those recorded values to equal what the reading build computes. So a rebuild that changes any hashed byte (a comment, a TypeScript or ajv bump) stranded every in-flight run. That is cost 3 in #921.

On origin/main (b6dd1da), measured by the new test below: a run launched by one build, then operated on by a process whose strict-json.js has one appended comment line:

operation (rebuilt validator) origin/main this branch
strict linked evidence (pause/cancel/replay/fork read it) WORKFLOW_CONTROL_EVIDENCE_INVALID: … planned graph output schema binding changed for "findings.json" ok
status (getRunHealth) ok, but two WORKFLOW_CONTROL_EVIDENCE_DIVERGED warnings and WORKFLOW_STATE_SYNC_SKIPPED ok, no divergence
host artifact gate, valid findings.json throws at the graph read (planned graph output schema binding changed) passes
host artifact gate, {} in findings.json throws at the graph read rejected with JSON_SCHEMA_VIOLATION
validator preflight of the run's trusted CLI (what every task preparation runs) JSON validator preflight success envelope has a mismatched identity ok
pause WORKFLOW_CONTROL_EVIDENCE_INVALID pause-requested
cancel WORKFLOW_CONTROL_EVIDENCE_INVALID cancel-requested
resume (native) ok ok

resume --refresh-controller by the rebuilt build failed too, measured by a second new test on a dynamic run. The refreshed workflow's first materialization republished its task plan. Strict run evidence, which every gated command reads, then failed with WORKFLOW_CONTROL_EVIDENCE_INVALID: published dynamic runtime controls no longer re-derive from their sealed base: persisted dynamic runtime task plan does not match its sealed templates and manifests. That happened on origin/main and on this branch before d6677fa.

The same comparison also sat in verified reads (current JSON Schema binding changed), the terminal report (terminal report found schema/control drift), trusted-CLI rotation on --refresh-controller, and the rendered workflow's task preparation and dependency admission.

Root cause

Provenance was being used as an equality gate against the reading build. The planned schema_sha256, plus the run's sealed schema snapshot when the bundle digest differs, already pins what an artifact has to satisfy. validator_build and the description-derived contract_digest add nothing to content validation. What they did add was a failure on every rebuild.

--refresh-controller added a second form of the same problem. It rebound each declared output's contract digest and validator build to the refreshing build (#982). The refreshed workflow copies those declarations into its verification markers, which host synchronization compares deep-equal with graph.json. A dynamic run's first tick also republishes them in the runtime task plan that every gated command re-derives from the sealed base. graph.json and the sealed base still carry the launch build's values.

Change

  • packages/artifacts/src/planned-graph.ts: assertPlannedGraphSemantics checks shape and internal consistency only. It no longer looks outputs up in the reading build's contract registry. assertSealedPlannedGraph is now the same check as assertPlannedGraph and stays as a thin wrapper for its existing callers. The function is split into per-node helpers so the diff-limited strict lint passes on the touched function. The only behaviour change is the removed comparisons.
  • packages/runtime/src/artifact-gates.ts (verifyRequiredArtifactSchemaBinding): validates against the schema content the planned output names. It uses the installed schemas when their bundle digest is the planned one, and otherwise the run's sealed snapshot at the planned schema_file. After validation it checks the schema ID, SHA-256 and bundle digest, and no longer the validator build. It is now exported so verified reads can share it.
  • packages/artifacts/src/sealed-schema-registry.ts: loads a sealed bundle as it was sealed. It no longer requires the installed build's file set and $ids, which a later build that adds, removes or re-versions a schema file breaks. That has happened: usage-ledger 1→2 on 2026-08-31, run-plan 2→3, and four files added in the v0.1.0 release. The gate's bundle-digest comparison decides whether a bundle is the planned one.
  • packages/runtime/src/verified-output.ts: deletes assertCurrentContractBinding. Verified outputs are validated with the same planned-content check as the gate.
  • packages/runtime/src/terminal-report.ts: deletes assertCurrentContractBindings.
  • Validator preflight (semantic-gates.ts, json-validator-preflight.ts, trusted-cli.ts): the identity gate compares schema ID, schema SHA-256, bundle digest and fixture SHA-256, and no longer the validator build. The workflow's per-task preflight accepts a rebuilt CLI of the same schemas, and trusted-CLI rotation on --refresh-controller adopts one.
  • packages/runtime/src/smithers.ts (taskWithCurrentArtifactSchemas, new in this round): a refresh still rebinds each output's schema file, ID, digest and bundle digest to the installed bundle (refresh-controller cannot retry tasks across schema bundle upgrades #982). It now keeps the recorded contract digest and validator build. Nothing in the rendered workflow compares those two with the running build any more.
  • workflow.tsx template:
    • assertTaskOutputSchemaBindings drops validatorBuild. It still requires the schema file, ID, digest and bundle digest to match the schemas the workflow validates with.
    • Dependency admission keeps the marker's sha256 == bytes check and the marker's schema file, ID and SHA-256, and keeps re-validating the bytes.
    • Dependency admission drops contract-digest, bundle-digest and validator-build equality, both against the declaration and against the running build.
    • Its now-unused artifactContractDefinition import is removed. The export is kept.
  • start-run.ts and workflow-integrity.ts: call-site updates for the removed option, and the observer doc comment is rewritten to match.
  • Docs (SPECS.md, schemas.md, reference/topology-yaml.md, reference/cli.md, explanation/topology-prompts-artifacts.md) and CHANGELOG now describe validator_build as provenance. The SPECS requirement names the readers that stop comparing it, and says that task preparation and the preflight still require the planned schema bundle.
  • CHANGELOG.md gains the <!-- markdownlint-disable MD013 --> line other branches add. Super-Linter lints the whole changed file, and 43 lines already on main exceed 400 characters.

The five validator modules, the schema files and contract descriptions are untouched. VALIDATOR_BUILD_IDENTITY is still …028be325… after a rebuild, so taking this change does not itself rotate any identity.

Deliberately not built (and why)

  • Resuming across a schema edit. Task preparation still copies the installed schemas into the workspace and requires them to equal the task's binding. The preflight also still compares schema and bundle identity with the schemas the workflow validates with. So a natively resumed task stops in preparation when any schema byte changed, as it did before this change.
    • A refresh across a schema edit rebinds the declared schema fields (refresh-controller cannot retry tasks across schema bundle upgrades #982), so preparation passes. But every schema-backed node verified after that refresh then records a different schema_sha256 or schema_bundle_sha256 than graph.json, and synchronization compares the two deep-equal. On a dynamic run the republished task plan differs from its sealed base in the same way. The rest of this bullet is from code reading; the same comparison demonstrably fails a node when only validator_build differs (the scratch run below).
    • On origin/main such a refresh could not finish either. Tasks whose dependencies were verified before the refresh also failed earlier there, at dependency admission, before any agent work. This PR removes the bundle comparison from dependency admission, so those tasks now run their agent before failing at synchronization. I left it removed: it only fires in a run that already cannot finish, and it is the fail-closed layer this change is removing.
    • The real fix is to validate against the planned bundle inside the workflow, or to re-seal the plan on refresh. Both are larger changes.
  • Runs whose workflow was rendered before this change keep the old template's checks in .smithers/workflows/ultrafuzz-<id>.tsx. A later validator rebuild can still stop their task preparation, and a refresh re-renders them with this version's template. There is no honest way to change already-rendered code short of faking the identity.
  • resolveStateDeclaredPropertyLens still compares the validator build and contract digest with the reading build. It is only reached when a gate is called without attempt authority, and every production caller passes it. The parallel "delete host re-implementations of registry artifact gates" change (refactor(runtime): delete host re-implementations of registry artifact gates #1178) deletes it, so I left its hunk alone to avoid a conflict.
  • The planned-graph-contract-identity semantic gate (plannedContractIdentityIssues) still re-derives each planned output's contract digest and binding, including validator_build, against the reading build. No production code dispatches planned-graph gates. The parallel artifacts dead-code change (refactor(artifacts): delete never-dispatched gates, production-dead code, and the unread event index #1181) deletes the gate and its registration, so deleting it here would only conflict.
  • Sealed-bundle validation cost. A validation against a sealed bundle reads state.json, reloads the bundle and starts a dedicated validation worker. A reviewer measured about 176 ms per call against about 9 ms on the installed path. Verified reads, and so finalized dependency reads, now take that path when a run's planned bundle differs from the installed one; before, they validated against the installed schemas. The per-call worker is deliberate in json-file-validator.ts, which this change must not edit because it is hashed into VALIDATOR_BUILD_IDENTITY. Memoizing the registry load would not remove the worker start. It costs time only in upgraded runs and does not stall them.
  • contract_digest still hashes the description prose. It is no longer compared on reads. Changing the digest itself would rotate every identity.
  • A real Smithers engine resuming a rendered workflow across a rebuild. Task preparation and dependency admission are covered by the extracted-function harness. The refresh test replays the refreshed workflow's own materialization call and writes the marker its verifier would write, rather than running the engine. The hermetic end-to-end harness is its own work item (test: model-free end-to-end campaign with controller kill and resume on the real pinned engine #1187).

Verification

Discriminating tests. Each fails on origin/main b6dd1da, with main's src and this branch's test files, and passes here. On main, runtime.test.ts also needs a one-line import shim, because main does not export verifyRequiredArtifactSchemaBinding:

  • runtime.test.ts: "a validator rebuild after launch leaves lifecycle commands, status and schema gates working". A child process preloads a one-line fs.readFileSync wrapper that appends a comment to the operator's artifacts/dist/strict-json.js, which gives it a different VALIDATOR_BUILD_IDENTITY. It runs strict evidence, getRunHealth, the host gate on a valid and an invalid findings.json, the run's real trusted-launcher preflight, resumeRun, pauseRun and cancelRun against a run this process launched. On main every step except resume fails; the table above is copied from main's output.
  • dynamic-lifecycle.test.ts (new in this round): "a controller refresh by a rebuilt validator keeps the run's recorded output bindings". It launches a dynamic run and renders the refreshed controller in a rebuilt-validator child. It asserts that the child's identity differs, then replays the refreshed workflow's first-tick materializeDynamicRuntime with its literals. It then checks strict evidence, writes a generated attempt's finding and marker from the refreshed declaration, and syncs.
    • It fails on origin/main and on 7683771 with the WORKFLOW_CONTROL_EVIDENCE_INVALID quoted above.
    • With d6677fa the evidence is admitted and the attempt synchronizes as succeeded.
    • The preload now lives in test/rebuilt-validator.ts, shared with the test above.
  • runtime.test.ts: "current-controller rendering preserves prompts idempotently…" (fix(runtime): refresh task schema bundles #983's test) now asserts that the refreshed spec keeps the recorded contract digest and validator build while its bundle digest is rebound. Without d6677fa's smithers.ts change it fails with actual: '8123…1081', expected: '0000…0000'.
  • Scratch run, not committed, for the synchronization half of the refresh problem: a generated attempt's marker whose validator_build differs from graph.json. syncRun reports ARTIFACT_VERIFICATION_AUTHORITY_INVALID: artifact verification marker outputs do not exactly match planned attempt … and records the attempt as failed. Before d6677fa, every schema-backed node verified after such a refresh wrote a marker like that.
  • runtime.test.ts: "artifact gates validate a historical bundle through its active sealed schema snapshot". The sealed snapshot now also differs in the output's own schema bytes, carries a schema file this build does not ship, and the planned validator_build is foreign. On main it fails with schema registry mismatch; unregistered: retired.schema.json. Here the valid artifact passes and {} is rejected with JSON_SCHEMA_VIOLATION. I removed the test's hard-coded VALIDATOR_BUILD_IDENTITY pin; it only locked in the stranding.
  • artifact-gates.test.ts: "artifact validation binds schema content and treats the validator build as provenance". On main: ARTIFACT_SCHEMA_BINDING_MISMATCH … does not match validator build.
  • generated-workflow-verifier.test.ts:
    • "generated task preparation binds output schema content but not the validator build". On main: planned schema binding changed for findings.json. In this round it also checks that a change to each of the schema file, ID, digest and bundle digest is still refused. That replaces a loop that regex-matched the template source for those field names.
    • "generated Smithers dependency verification fails closed before descendant preparation" now also admits a marker with a different contract digest, bundle digest or validator build. It still rejects a different schema SHA-256, a forged path or contract, and tampered bytes. On main: verification marker artifact is not a declared output.
  • trusted-cli.test.ts: refresh adopts a rebuilt CLI whose validator build differs and still refuses one whose schema bundle differs; recorded-identity provenance. On main both fail with mismatched identity.
  • artifacts/test/schema.test.ts: a planned graph with another build's contract digest, schema SHA-256, bundle and validator build loads, while the shape still requires bindings exactly for schema-backed contracts (main: planned graph output contract digest changed). A sealed bundle with a dropped file, an extra file and a re-versioned $id loads with its own IDs (main: schema registry mismatch).
  • artifacts/test/json-validator-preflight.test.ts: an envelope from another validator build of the same schemas parses, and a different schema SHA-256 is still rejected (main: mismatched identity).
  • verified-output.test.ts: "verified reads and the terminal report accept outputs planned by another validator build". The planned outputs carry a foreign validator_build and contract_digest. The report is still read through verifier/controller authority and published as the terminal report. On main it fails at the graph read (planned graph output contract digest changed). A reviewer confirmed that main's assertCurrentContractBinding and assertCurrentContractBindings each fail it on their own ("current contract digest changed", "terminal report found schema/control drift").

Coverage test, not a regression test: runtime.test.ts "a sealed planned graph that fails graph semantics leaves status readable" (new in this round). This branch deleted the only test of parseSealedPlannedGraphForObserver's divergence branch. A rewritten graph.json with two primary outputs reaches that branch again: strict evidence refuses, and the observer and getRunHealth report the divergence as a warning. It would fail on main only because the divergence text changed.

Suites and checks, at 7683771:

  • @ultrafuzz/artifacts: whole suite 354/354.
  • @ultrafuzz/runtime:
    • artifact-gates 161/161.
    • verified-output, terminal-report-completion, terminal-report-projection, trusted-cli and generated-workflow-verifier 236/236.
    • dynamic-lifecycle plus workflow-control 29/31 in one run. The two failures were WORKFLOW_SUBMISSION_FAILED: workflow execution file dependencies/packages/0000NN/{LICENSE,README.md} changed while reading during fixture launch. I think those files are hard links into the pnpm store that other worktrees on this shared host were installing into, but I did not confirm the cause. Both tests pass when rerun alone on this branch.
  • @ultrafuzz/cli: json-validate 16/16, including host-gate diagnostic parity.
  • @ultrafuzz/modal: node-worker-preflight 6/6.

Suites and checks, in this round at the head:

  • @ultrafuzz/runtime:
    • generated-workflow-verifier, verified-output and trusted-cli 219/219.
    • dynamic-lifecycle, the whole file: 19/20 in one run, alongside other test processes on this shared host. The failure was "dynamic state materialization checks the synchronization deadline before publication": its diagnostics lacked WORKFLOW_SYNC_DEADLINE_EXCEEDED. It passes when rerun alone at this head (45 s) and on origin/main (67 s). This PR does not touch the synchronization clock it counts, and I did not identify the cause.
    • runtime.test.ts: the rotation, fix(runtime): refresh task schema bundles #983-rendering, historical-bundle and graph-semantics tests plus "controller refresh selects current source without rewriting historical evidence" (5/5); every test matching controller refresh|refresh resume|current-controller (13/13); and, after the last edits, the graph-semantics and rotation tests again (2/2).
  • prettier --check and eslint on the changed files, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, typecheck for runtime and artifacts, pnpm -w knip, and node scripts/docs-check.mjs.

Not run locally: the full runtime.test.ts and CLI suites (too heavy for the shared host these ran on), the evals and dashboard suites, and a real Smithers engine resuming a run across a rebuild. CI's full release validation did not run on the earlier push, because External static analysis failed on MD013 in CHANGELOG.md and the validation job depends on it. On ba6bb92 (CI run 36520380854) every check passed: External static analysis, the PR build and Node.js 24 runtime smoke, all four runtime integration shards, the runtime support tests with the Bun adapter contracts, and release-gates.

Risk / compatibility

  • Exports.
    • assertSealedPlannedGraph is kept.
    • assertPlannedGraphSemantics loses its options parameter; start-run was the only caller.
    • JsonValidatorPreflightExpectedIdentity and SemanticValidatorPreflightContext lose validatorBuild. Old JS callers that still pass it are unaffected.
    • New runtime export: verifyRequiredArtifactSchemaBinding.
    • The generated workflow still imports the same names from @ultrafuzz/artifacts/@ultrafuzz/runtime; it just no longer destructures artifactContractDefinition.
  • Diagnostics. The text and details of ARTIFACT_SCHEMA_BINDING_MISMATCH and ARTIFACT_VALIDATOR_IDENTITY_MISMATCH change, as does the observer divergence text for a graph that fails semantics. Codes are unchanged.
  • Posture. A different validator build of the same schema bytes is no longer treated as a reason to refuse. The host still validates every artifact with its own validator against the planned schema bytes, and the preflight still pins schema ID, digest and bundle. What is lost is detecting that ajv or a validator rebuild changed a verdict for identical schemas. That is the cost of Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921's requirement that an identity rotation never strand an in-flight run.
  • Refresh. A refreshed workflow now records the launch build's contract digest and validator build in its markers and runtime task plan, as the sealed plan already does. The schema fields are still rebound as in refresh-controller cannot retry tasks across schema bundle upgrades #982.
  • Test removed. This PR deletes runtime.test.ts "a planned graph that stops matching this build's contracts leaves status readable". It emulated a rebuild by rewriting graph.json after sealing and asserted the re-derivation divergence this PR removes. The two-build rotation test replaces it for the rebuild, and the new graph-semantics test covers the observer branch it used to reach. The parallel "cancel and pause work even when run control evidence has diverged" change (fix(runtime): cancel and pause work even when run control evidence has diverged #1170) edits the same test, so whichever lands second resolves the conflict by dropping it.
  • Merge surface. A trial git merge-tree of this head against each of the other 25 claude/w* branches conflicts in CHANGELOG.md for nine of them (every change adds at the same line) and, for fix(runtime): cancel and pause work even when run control evidence has diverged #1170, in runtime.test.ts at the deleted test above. Nothing else conflicts. The hunks are small and local:
    • start-run.ts 1094–1122
    • workflow-integrity.ts 672 and 2085
    • artifact-gates.ts 3691–3797
    • smithers.ts 12 and 8507–8541
    • workflow.tsx 38, 4066–4080 and 5856–5885
    • semantic-gates.ts 227–232, 428–432 and 7808–7811

Refs #921

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the reviewed changes.

Summary

The PR treats recorded validator-build and contract identities as provenance rather than requiring them to match a later reading build. It validates artifacts against their planned schema content, permits historical sealed schema bundles, and retains recorded output provenance when refreshing a controller.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Plan[Planned output binding] --> Gate[Host gate or verified read]
  Gate --> Bundle{Installed bundle matches plan?}
  Bundle -->|Yes| Installed[Validate with installed schemas]
  Bundle -->|No| Sealed[Validate with sealed snapshot]
  Installed --> Result[Check planned schema identity and artifact]
  Sealed --> Result
Loading

Reviews (4) · Last reviewed commit: "Merge origin/main into claude/w24-valida..."

aviggiano and others added 2 commits September 29, 2026 01:46
VALIDATOR_BUILD_IDENTITY hashes five compiled validator modules and the
ajv/ajv-formats versions. Planning records it, with the contract digest
and schema binding, in every planned output, and readers required those
values to equal what the reading build computes. A rebuild that changes
any hashed byte therefore stranded in-flight runs (#921): strict run
evidence failed with "planned graph output schema binding changed", so
pause, cancel, replay and fork refused the run; status reported a
divergence and skipped synchronization; host artifact gates, verified
reads, the terminal report and the validator preflight refused it after
resume.

The planned schema content already pins what an artifact must satisfy,
so the recorded validator build and contract digest are kept as
provenance and no longer compared with the reading build:

- planned-graph: sealed and fresh graphs are checked for shape and
  internal consistency only. The function is split into per-node helpers
  so the touched code passes the diff-limited strict lint.
- artifact gates and verified reads validate each artifact against the
  schema content its planned output names: the installed schemas when
  the bundle digest matches, otherwise the bundle sealed in the run's
  execution snapshot, at the planned schema file.
- the sealed schema loader loads a bundle as it was sealed, so a later
  build that adds, removes or re-versions a schema file no longer makes
  it throw; the gate's bundle digest check decides.
- the terminal report no longer refuses on contract or binding drift.
- the validator preflight identity gate compares schema id, digest,
  bundle digest and fixture digest, not the validator build, which also
  lets controller refresh adopt a rebuilt trusted CLI of the same schemas.
- the workflow template stops comparing validatorBuild during task
  preparation, and dependency admission keeps the marker's sha256 and
  schema content binding but stops comparing contract digest, bundle
  digest and validator build.

Tests that pinned VALIDATOR_BUILD_IDENTITY or asserted the stranding are
rewritten; the new rotation test runs the operator's commands in a child
process whose rebuilt strict-json.js yields a different identity.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The specification, schema reference, topology reference, CLI reference
and artifacts explanation said the host and preflight verify the
validator build. They now say the host validates against the planned
schema content (the run's sealed snapshot when the installed bundle
differs) and records the validator build without comparing it. Adds the
CHANGELOG entry for the change.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 01:47
Comment thread packages/runtime/test/generated-workflow-verifier.test.ts
aviggiano and others added 3 commits September 29, 2026 04:06
…schemas

`resume --refresh-controller` rebinds every declared output to the
refreshing build (#982), including its contract digest and validator
build. The refreshed workflow copies those declarations into its
verification markers, which host synchronization deep-compares with
graph.json, and a dynamic run's first tick republishes them in the
runtime task plan that every gated command re-derives from the sealed
base. After a rebuild that changed only VALIDATOR_BUILD_IDENTITY, both
comparisons failed: gated commands with WORKFLOW_CONTROL_EVIDENCE_INVALID
("published dynamic runtime controls no longer re-derive from their
sealed base"), and a node verified after the refresh with
ARTIFACT_VERIFICATION_AUTHORITY_INVALID ("artifact verification marker
outputs do not exactly match planned attempt").

Nothing in the rendered workflow compares these two fields with the
running build any more, so the refresh keeps their recorded values and
rebinds only the schema file, ID, digest and bundle digest.

The rebuilt-validator preload moves into a shared test helper, so the
new dynamic-lifecycle test can render the refreshed controller in a
child process whose validator build really differs.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… a source regex

This branch deleted the only test of the divergence branch in
parseSealedPlannedGraphForObserver. A rewritten graph with two primary
outputs reaches it again: strict evidence refuses the run, while
observers and `status` report the divergence as a warning.

The task-preparation test now checks that a change to each schema field,
including the bundle digest, is still refused. That replaces a loop that
regex-matched the template source for the same field names. A test
comment now quotes the failures origin/main actually reports.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The CHANGELOG entry said task preparation stopped comparing bundle
digests; it still requires the schema bundle it validates with to be
the planned one, so resuming across a schema edit is not covered. The
entry and the new SPECS requirement now name the readers that stop
comparing the recorded validator build and contract digest, and the
entry describes the controller-refresh fix.

CHANGELOG.md gains the markdownlint MD013 opt-out other branches add,
because Super-Linter lints the whole file and 43 existing lines exceed
400 characters.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano aviggiano changed the title fix: a rebuild or upgrade no longer strands in-flight runs (validator build becomes provenance) fix: a validator rebuild no longer strands in-flight runs (validator build becomes provenance) Sep 29, 2026
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit a8625ff into main Sep 29, 2026
1 of 5 checks passed
@aviggiano
aviggiano deleted the claude/w24-validator-build-provenance branch September 29, 2026 06:37
aviggiano added a commit that referenced this pull request Sep 29, 2026
…cord it in the changelog (#1191)

Test and lint fixes for five interactions between the 25 merged PRs #1163-#1188, found by the combined CI run in #1190, plus the batch's consolidated changelog entries.
aviggiano added a commit that referenced this pull request Sep 29, 2026
…whose dynamic groups expanded

Every render and every lifecycle admission re-derives the published
dynamic runtime controls from their sealed base with the current code,
including re-rendering each runtime-rendered prompt, and refused a
published prompt.rendered.md whose bytes differed from the fresh render
("runtime rendered prompt changed for <attempt>"). So any renderer or
projection change between the build that published a prompt and the
reading build stranded the run: why, stats, resume, replay, fork and
verified reads failed with WORKFLOW_CONTROL_EVIDENCE_INVALID, status and
inspect stopped synchronizing it, and a resumed controller threw on its
first render. #1176 changed the coverage projection that the stock final
report embeds, and the final report renders as soon as the goal groups
expand, so current main can neither synchronize nor resume a
packaged-topology run that v0.1.1 took past goal-plan. #1195's selector
order does the same for custom prompts with multi-path selectors.

A published prompt is what its task was, or will be, handed, so it is
now kept as published whatever this build renders: renderReadyRuntimePrompts
returns the drifted attempt IDs instead of throwing. The controller
adopts them silently. verifyWorkflowControlSnapshot carries them on the
verified snapshot (and only reuses a remembered snapshot whose drift list
matches), and synchronization adds one WORKFLOW_PUBLISHED_PROMPT_DRIFT
warning per pass naming the count and up to three attempts, so status,
inspect, why and stats return it. Every other re-derivation check is
unchanged: a missing prompt, a changed group template or expansion
manifest, and a task plan or graph that no longer re-derives still fail
closed. A prompt not yet published is rendered and published exactly as
before.

The cost: a prompt.rendered.md that a same-UID process writes or rewrites
for a later task, even before the controller first renders it, is now
adopted and reported as drift rather than refused.

Refs #921, #1188, #1195, #1176

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task

A native resume runs the launch-rendered workflow against the installed
@ultrafuzz/artifacts, and three task-preparation steps required the
installed schema bundle to equal the one the run was launched with. One
`$comment` added to an installed schema failed every resumed task, and
neither --retry-failed nor --refresh-controller recovered it (#921, the
prompt-sealing investigation's map 4):

- materialize-prompt-schemas refused a workspace copy that differed from
  the installed bundle ("prompt schema destination differs from
  checked-in source");
- assert-task-output-schema-bindings compared this build's binding with
  the plan ("planned schema binding changed for <path>");
- preflight-json-validator parsed the run's own validator envelope
  against this build's findings binding by default, while that
  validator, the launch closure, reports the launch bundle ("mismatched
  identity").

The run's plan already records which schema each output must satisfy, and
the host gates validate against it since #1188. The workflow now does the
same:

- plannedArtifactSchemaBundle(runRoot, sha256) resolves the bundle a plan
  names: this build's schemas when their digest is the planned one,
  otherwise the copy sealed in the run's execution snapshot. The host
  gate uses it too, so the host and the workflow share one selection
  rule and one snapshot lookup.
- Task preparation copies that bundle into .ultrafuzz/schemas and
  rewrites a copy that differs instead of failing: the copy is only the
  agent's view. The bindings step checks that the planned bundle holds
  each planned schema, and the preflight expects the planned bundle's
  identity. The verifier, dependency admission and the goal-search
  count validate against the planned schema
  (validateArtifactContractBytes takes it as an optional argument, and no
  longer compares the reported validator build).
- --refresh-controller no longer rebinds declared outputs to the
  installed bundle (#982), so the refreshed verifier's markers keep
  matching the plan. The replacePromptSchemas flag and its
  __ULTRAFUZZ_REPLACE_PROMPT_SCHEMAS__ placeholder are deleted.

Breaking: parseJsonValidatorPreflightSuccessEnvelope requires the expected
identity instead of defaulting to this build's, and
materializePromptSchemas takes the bundle and drops replaceExisting. A
workflow rendered before this change calls both the old way, so a run
launched before it must be continued with --refresh-controller. A run no
longer adopts an upgrade's schemas.

The new runtime test launches a run, loads its project workflow the way a
native resume does against an upgraded copy of @ultrafuzz/artifacts, and
runs task preparation and verification in process. On the base branch it
fails at preflight-json-validator (and at materialize-prompt-schemas or
assert-task-output-schema-bindings when those run first).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task

A native resume runs the launch-rendered workflow against the installed
@ultrafuzz/artifacts, and three task-preparation steps required the
installed schema bundle to equal the one the run was launched with. One
`$comment` added to an installed schema failed every resumed task, and
neither --retry-failed nor --refresh-controller recovered it (#921, the
prompt-sealing investigation's map 4):

- materialize-prompt-schemas refused a workspace copy that differed from
  the installed bundle ("prompt schema destination differs from
  checked-in source");
- assert-task-output-schema-bindings compared this build's binding with
  the plan ("planned schema binding changed for <path>");
- preflight-json-validator parsed the run's own validator envelope
  against this build's findings binding by default, while that
  validator, the launch closure, reports the launch bundle ("mismatched
  identity").

The run's plan already records which schema each output must satisfy, and
the host gates validate against it since #1188. The workflow now does the
same:

- plannedArtifactSchemaBundle(runRoot, sha256) resolves the bundle a plan
  names: this build's schemas when their digest is the planned one,
  otherwise the copy sealed in the run's execution snapshot. The host
  gate uses it too, so the host and the workflow share one selection
  rule and one snapshot lookup.
- Task preparation copies that bundle into .ultrafuzz/schemas and
  rewrites a copy that differs instead of failing: the copy is only the
  agent's view. The bindings step checks that the planned bundle holds
  each planned schema, and the preflight expects the planned bundle's
  identity. The verifier, dependency admission and the goal-search
  count validate against the planned schema
  (validateArtifactContractBytes takes it as an optional argument, and no
  longer compares the reported validator build).
- --refresh-controller no longer rebinds declared outputs to the
  installed bundle (#982), so the refreshed verifier's markers keep
  matching the plan. The replacePromptSchemas flag and its
  __ULTRAFUZZ_REPLACE_PROMPT_SCHEMAS__ placeholder are deleted.

Breaking: parseJsonValidatorPreflightSuccessEnvelope requires the expected
identity instead of defaulting to this build's, and
materializePromptSchemas takes the bundle and drops replaceExisting. A
workflow rendered before this change calls both the old way, so a run
launched before it must be continued with --refresh-controller. A run no
longer adopts an upgrade's schemas.

The new runtime test launches a run, loads its project workflow the way a
native resume does against an upgraded copy of @ultrafuzz/artifacts, and
runs task preparation and verification in process. On the base branch it
fails at preflight-json-validator (and at materialize-prompt-schemas or
assert-task-output-schema-bindings when those run first).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task

A native resume runs the launch-rendered workflow against the installed
@ultrafuzz/artifacts, and three task-preparation steps required the
installed schema bundle to equal the one the run was launched with. One
`$comment` added to an installed schema failed every resumed task, and
neither --retry-failed nor --refresh-controller recovered it (#921, the
prompt-sealing investigation's map 4):

- materialize-prompt-schemas refused a workspace copy that differed from
  the installed bundle ("prompt schema destination differs from
  checked-in source");
- assert-task-output-schema-bindings compared this build's binding with
  the plan ("planned schema binding changed for <path>");
- preflight-json-validator parsed the run's own validator envelope
  against this build's findings binding by default, while that
  validator, the launch closure, reports the launch bundle ("mismatched
  identity").

The run's plan already records which schema each output must satisfy, and
the host gates validate against it since #1188. The workflow now does the
same:

- plannedArtifactSchemaBundle(runRoot, sha256) resolves the bundle a plan
  names: this build's schemas when their digest is the planned one,
  otherwise the copy sealed in the run's execution snapshot. The host
  gate uses it too, so the host and the workflow share one selection
  rule and one snapshot lookup.
- Task preparation copies that bundle into .ultrafuzz/schemas and
  rewrites a copy that differs instead of failing: the copy is only the
  agent's view. The bindings step checks that the planned bundle holds
  each planned schema, and the preflight expects the planned bundle's
  identity. The verifier, dependency admission and the goal-search
  count validate against the planned schema
  (validateArtifactContractBytes takes it as an optional argument, and no
  longer compares the reported validator build).
- --refresh-controller no longer rebinds declared outputs to the
  installed bundle (#982), so the refreshed verifier's markers keep
  matching the plan. The replacePromptSchemas flag and its
  __ULTRAFUZZ_REPLACE_PROMPT_SCHEMAS__ placeholder are deleted.

Breaking: parseJsonValidatorPreflightSuccessEnvelope requires the expected
identity instead of defaulting to this build's, and
materializePromptSchemas takes the bundle and drops replaceExisting. A
workflow rendered before this change calls both the old way, so a run
launched before it must be continued with --refresh-controller. A run no
longer adopts an upgrade's schemas.

The new runtime test launches a run, loads its project workflow the way a
native resume does against an upgraded copy of @ultrafuzz/artifacts, and
runs task preparation and verification in process. On the base branch it
fails at preflight-json-validator (and at materialize-prompt-schemas or
assert-task-output-schema-bindings when those run first).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task

A native resume runs the launch-rendered workflow against the installed
@ultrafuzz/artifacts, and three task-preparation steps required the
installed schema bundle to equal the one the run was launched with. One
`$comment` added to an installed schema failed every resumed task, and
neither --retry-failed nor --refresh-controller recovered it (#921, the
prompt-sealing investigation's map 4):

- materialize-prompt-schemas refused a workspace copy that differed from
  the installed bundle ("prompt schema destination differs from
  checked-in source");
- assert-task-output-schema-bindings compared this build's binding with
  the plan ("planned schema binding changed for <path>");
- preflight-json-validator parsed the run's own validator envelope
  against this build's findings binding by default, while that
  validator, the launch closure, reports the launch bundle ("mismatched
  identity").

The run's plan already records which schema each output must satisfy, and
the host gates validate against it since #1188. The workflow now does the
same:

- plannedArtifactSchemaBundle(runRoot, sha256) resolves the bundle a plan
  names: this build's schemas when their digest is the planned one,
  otherwise the copy sealed in the run's execution snapshot. The host
  gate uses it too, so the host and the workflow share one selection
  rule and one snapshot lookup.
- Task preparation copies that bundle into .ultrafuzz/schemas and
  rewrites a copy that differs instead of failing: the copy is only the
  agent's view. The bindings step checks that the planned bundle holds
  each planned schema, and the preflight expects the planned bundle's
  identity. The verifier, dependency admission and the goal-search
  count validate against the planned schema
  (validateArtifactContractBytes takes it as an optional argument, and no
  longer compares the reported validator build).
- --refresh-controller no longer rebinds declared outputs to the
  installed bundle (#982), so the refreshed verifier's markers keep
  matching the plan. The replacePromptSchemas flag and its
  __ULTRAFUZZ_REPLACE_PROMPT_SCHEMAS__ placeholder are deleted.

Breaking: parseJsonValidatorPreflightSuccessEnvelope requires the expected
identity instead of defaulting to this build's, and
materializePromptSchemas takes the bundle and drops replaceExisting. A
workflow rendered before this change calls both the old way, so a run
launched before it must be continued with --refresh-controller. A run no
longer adopts an upgrade's schemas.

The new runtime test launches a run, loads its project workflow the way a
native resume does against an upgraded copy of @ultrafuzz/artifacts, and
runs task preparation and verification in process. On the base branch it
fails at preflight-json-validator (and at materialize-prompt-schemas or
assert-task-output-schema-bindings when those run first).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task

A native resume runs the launch-rendered workflow against the installed
@ultrafuzz/artifacts, and three task-preparation steps required the
installed schema bundle to equal the one the run was launched with. One
`$comment` added to an installed schema failed every resumed task, and
neither --retry-failed nor --refresh-controller recovered it (#921, the
prompt-sealing investigation's map 4):

- materialize-prompt-schemas refused a workspace copy that differed from
  the installed bundle ("prompt schema destination differs from
  checked-in source");
- assert-task-output-schema-bindings compared this build's binding with
  the plan ("planned schema binding changed for <path>");
- preflight-json-validator parsed the run's own validator envelope
  against this build's findings binding by default, while that
  validator, the launch closure, reports the launch bundle ("mismatched
  identity").

The run's plan already records which schema each output must satisfy, and
the host gates validate against it since #1188. The workflow now does the
same:

- plannedArtifactSchemaBundle(runRoot, sha256) resolves the bundle a plan
  names: this build's schemas when their digest is the planned one,
  otherwise the copy sealed in the run's execution snapshot. The host
  gate uses it too, so the host and the workflow share one selection
  rule and one snapshot lookup.
- Task preparation copies that bundle into .ultrafuzz/schemas and
  rewrites a copy that differs instead of failing: the copy is only the
  agent's view. The bindings step checks that the planned bundle holds
  each planned schema, and the preflight expects the planned bundle's
  identity. The verifier, dependency admission and the goal-search
  count validate against the planned schema
  (validateArtifactContractBytes takes it as an optional argument, and no
  longer compares the reported validator build).
- --refresh-controller no longer rebinds declared outputs to the
  installed bundle (#982), so the refreshed verifier's markers keep
  matching the plan. The replacePromptSchemas flag and its
  __ULTRAFUZZ_REPLACE_PROMPT_SCHEMAS__ placeholder are deleted.

Breaking: parseJsonValidatorPreflightSuccessEnvelope requires the expected
identity instead of defaulting to this build's, and
materializePromptSchemas takes the bundle and drops replaceExisting. A
workflow rendered before this change calls both the old way, so a run
launched before it must be continued with --refresh-controller. A run no
longer adopts an upgrade's schemas.

The new runtime test launches a run, loads its project workflow the way a
native resume does against an upgraded copy of @ultrafuzz/artifacts, and
runs task preparation and verification in process. On the base branch it
fails at preflight-json-validator (and at materialize-prompt-schemas or
assert-task-output-schema-bindings when those run first).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task

A native resume runs the launch-rendered workflow against the installed
@ultrafuzz/artifacts, and three task-preparation steps required the
installed schema bundle to equal the one the run was launched with. One
`$comment` added to an installed schema failed every resumed task, and
neither --retry-failed nor --refresh-controller recovered it (#921, the
prompt-sealing investigation's map 4):

- materialize-prompt-schemas refused a workspace copy that differed from
  the installed bundle ("prompt schema destination differs from
  checked-in source");
- assert-task-output-schema-bindings compared this build's binding with
  the plan ("planned schema binding changed for <path>");
- preflight-json-validator parsed the run's own validator envelope
  against this build's findings binding by default, while that
  validator, the launch closure, reports the launch bundle ("mismatched
  identity").

The run's plan already records which schema each output must satisfy, and
the host gates validate against it since #1188. The workflow now does the
same:

- plannedArtifactSchemaBundle(runRoot, sha256) resolves the bundle a plan
  names: this build's schemas when their digest is the planned one,
  otherwise the copy sealed in the run's execution snapshot. The host
  gate uses it too, so the host and the workflow share one selection
  rule and one snapshot lookup.
- Task preparation copies that bundle into .ultrafuzz/schemas and
  rewrites a copy that differs instead of failing: the copy is only the
  agent's view. The bindings step checks that the planned bundle holds
  each planned schema, and the preflight expects the planned bundle's
  identity. The verifier, dependency admission and the goal-search
  count validate against the planned schema
  (validateArtifactContractBytes takes it as an optional argument, and no
  longer compares the reported validator build).
- --refresh-controller no longer rebinds declared outputs to the
  installed bundle (#982), so the refreshed verifier's markers keep
  matching the plan. The replacePromptSchemas flag and its
  __ULTRAFUZZ_REPLACE_PROMPT_SCHEMAS__ placeholder are deleted.

Breaking: parseJsonValidatorPreflightSuccessEnvelope requires the expected
identity instead of defaulting to this build's, and
materializePromptSchemas takes the bundle and drops replaceExisting. A
workflow rendered before this change calls both the old way, so a run
launched before it must be continued with --refresh-controller. A run no
longer adopts an upgrade's schemas.

The new runtime test launches a run, loads its project workflow the way a
native resume does against an upgraded copy of @ultrafuzz/artifacts, and
runs task preparation and verification in process. On the base branch it
fails at preflight-json-validator (and at materialize-prompt-schemas or
assert-task-output-schema-bindings when those run first).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 30, 2026
…sumed task (#1238)

A native `resume` runs the run's launch-rendered workflow against the **installed** `@ultrafuzz/artifacts`, and three task-preparation steps required the installed schema bundle to equal the one the run was launched with. Map 4 of the prompt-sealing investigation measured that one `"$comment"` added to an installed schema file (no rebuild is needed; the registry reads schema files at runtime) fails **every** resumed task, including a text-only one, and that neither `--retry-failed` nor `--refresh-controller --retry-failed` recovers it. The #1230 implementer hit the same two failures.
| Gate | Where | Failure |
| --- | --- | --- |
| C1 | `materializePromptSchemas` refused a workspace copy that differed from the installed bundle (only a refresh could replace it) | `prepare:<attempt> failed at step materialize-prompt-schemas: prompt schema destination differs from checked-in source: …` |
| C2 | `assertTaskOutputSchemaBindings` compared this build's binding with the plan | `…step assert-task-output-schema-bindings: artifact-contract failure: planned schema binding changed for <path>` |
| C3 | `parseJsonValidatorPreflightSuccessEnvelope` defaulted its expected identity to this build's findings binding, while the run's own validator (the launch closure) reports the launch bundle | `…step preflight-json-validator: … invalid success envelope` (`mismatched identity`) |
| C4 | `--refresh-controller` rebound every declared output to the installed bundle (#982) | the refreshed verifier's markers stop matching `graph.json` after a schema change |
Behind those gates, the workflow's verifier and dependency admission also validated with the installed schemas, and so did synchronization's findings count, so after a substantive schema change the engine and the host would check artifacts against a schema the run was never planned with; C1–C3 stopped every task before that could be reached. The host artifact gates had already moved to the planned schema in #1188.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant