fix(runtime): a renderer or projection change no longer strands runs whose dynamic groups expanded - #1216
fix(runtime): a renderer or projection change no longer strands runs whose dynamic groups expanded#1216aviggiano wants to merge 3 commits into
Conversation
…whose dynamic groups expanded
Every render and every lifecycle admission re-derives the published
dynamic runtime controls from their sealed base with the current code,
including re-rendering each runtime-rendered prompt, and refused a
published prompt.rendered.md whose bytes differed from the fresh render
("runtime rendered prompt changed for <attempt>"). So any renderer or
projection change between the build that published a prompt and the
reading build stranded the run: why, stats, resume, replay, fork and
verified reads failed with WORKFLOW_CONTROL_EVIDENCE_INVALID, status and
inspect stopped synchronizing it, and a resumed controller threw on its
first render. #1176 changed the coverage projection that the stock final
report embeds, and the final report renders as soon as the goal groups
expand, so current main can neither synchronize nor resume a
packaged-topology run that v0.1.1 took past goal-plan. #1195's selector
order does the same for custom prompts with multi-path selectors.
A published prompt is what its task was, or will be, handed, so it is
now kept as published whatever this build renders: renderReadyRuntimePrompts
returns the drifted attempt IDs instead of throwing. The controller
adopts them silently. verifyWorkflowControlSnapshot carries them on the
verified snapshot (and only reuses a remembered snapshot whose drift list
matches), and synchronization adds one WORKFLOW_PUBLISHED_PROMPT_DRIFT
warning per pass naming the count and up to three attempts, so status,
inspect, why and stats return it. Every other re-derivation check is
unchanged: a missing prompt, a changed group template or expansion
manifest, and a task plan or graph that no longer re-derives still fail
closed. A prompt not yet published is rendered and published exactly as
before.
The cost: a prompt.rendered.md that a same-UID process writes or rewrites
for a later task, even before the controller first renders it, is now
adopted and reported as drift rather than refused.
Refs #921, #1188, #1195, #1176
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…withdrawn expansion
`resume --retry-failed` on a dynamic source whose verifier failed withdraws
the published expansion generation into `dynamic-expansion-history/`, so the
group expands again from the source's new output. The archive moved only the
generation-owned attempts. It left the published `prompt.rendered.md` of every
planned task whose prompt waits on the group, and that prompt was rendered
from the withdrawn children.
When the re-run source planned different items, the next render of such a
task, for example a custom join whose prompt uses `{{artifact_path:<group>}}`,
differed from the file on disk. On main that threw `runtime rendered prompt
changed for <attempt>` and stranded the run. With published prompts now kept
as published, the join would instead have been handed the old prompt: it named
a child whose artifact directory is now in the history and omitted the new one.
The stock topology is not affected, because its deferred prompts reach the goal
groups only through authority selectors.
The archive now also moves the published prompt of each sealed base task whose
`deferredPromptGroups` names a withdrawn group, validated with the rest of the
attempt state before the Smithers reset. The next expansion renders it afresh
against the new generation, and `retry.json` lists it with the other archived
paths.
The re-derivation test now gives the fixture's join a deferred prompt that
names the group's children, and checks that the retry archives it and that
the re-expansion renders it from the new child with no drift. The lifecycle
retry test checks the same through a real `resume --retry-failed` of a
launched run, whose compiled base tasks mark the join's prompt as deferred:
the prompt is in the archive after the resume and is rendered again by the
re-expansion. Both fail with origin/main's retry module and on the previous
head: the join's prompt is still in place after the archive.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…kept too The previous commit kept a published runtime prompt that the current build renders differently. It still rendered every published prompt first, and a render or validator-command check that threw for one failed both the controller's render and every lifecycle admission. So a later build whose renderer refuses what an earlier build rendered, for example over a removed variable or a stricter validator-command rule, would still strand every run whose dynamic groups had expanded. That fresh render of a published prompt is only ever compared and then discarded, so it has no reason to fail a run. `renderReadyRuntimePrompts` still builds the render input for every ready task, but runs `renderPrompt` and the validator-command check in a closure. A published prompt whose render throws is recorded as drift and kept like one that renders differently. A prompt that is not published yet is rendered, checked and published exactly as before, and a render that throws for it still fails. Building the input reads the template and checks the vulnerability-catalog digest, so both still fail closed for every ready task. The warning text now says the prompts "do not match what this build renders", which covers both cases. Review wording fixes: - docs/reference/artifacts-reports.md names the commands that synchronize a run (`status`, `inspect`, `why`, `stats`; there is no `sync` command) and says the warning is in their `--json` diagnostics, printed only by `stats`. It says a planned node renders from its template snapshot, which is not sealed, and that a missing prompt fails admission while a render republishes it. - The smithers.ts comment says "template snapshot" instead of "sealed template". The drift test now also makes the join's published prompt unrenderable and checks that it is kept and reported, and that the same prompt, once removed, still fails to publish. It fails on origin/main and on the previous head with `unknown prompt template variable`. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…withdrawn expansion
`resume --retry-failed` on a dynamic source whose verifier failed withdraws
the published expansion generation into `dynamic-expansion-history/`, so the
group expands again from the source's new output. The archive moved only the
generation-owned attempts. It left the published `prompt.rendered.md` of every
planned task whose prompt waits on the group, and that prompt was rendered
from the withdrawn children.
When the re-run source planned different items, the next render of such a
task, for example a custom join whose prompt uses `{{artifact_path:<group>}}`,
differed from the file on disk. The re-expansion then threw `runtime rendered
prompt changed for <attempt>`, and so did every later render and lifecycle
admission check, so the run stranded until the file was deleted by hand. The
stock topologies are not affected, because their deferred prompts reach the
goal groups only through authority selectors, which render the same bytes
whatever the children are.
The archive now also moves the published prompt of each sealed base task whose
`deferredPromptGroups` names a withdrawn group. It is validated with the rest
of the attempt state before the Smithers reset, the next expansion renders it
afresh against the new generation, and `retry.json` lists it with the other
archived paths. A prompt that was never rendered is skipped.
The re-derivation test gives the fixture's join a deferred prompt that names
the group's children and re-runs the source with a different item after the
retry. It checks that the re-expansion renders the join's prompt from the new
child, that the archive holds the withdrawn prompt, and that `retry.json`
lists it. With origin/main's retry module it fails at the re-expansion with
`runtime rendered prompt changed for join`. The lifecycle retry test checks
the same through a real `resume --retry-failed` of a launched run, whose
compiled base tasks mark the join's prompt as deferred: the prompt is in the
archive after the resume and is rendered again by the re-expansion. With
origin/main's retry module it fails because the join's prompt is still in
place after the resume.
Split out of draft #1216, whose change to how drifted published prompts are
handled this fix does not depend on.
Refs #1141
Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
The second commit's fix, |
|
Owner decision: tamper detection on prompts is not worth what it costs. The main cost is that a run cannot be restarted after a minor prompt change, and stability matters more than exact reproducibility. So this PR will not merge as is. A follow-up will remove prompt tamper-sealing outright and keep only a record of the prompt each attempt used, following the #1188 precedent (record, don't enforce). This PR closes when that replacement opens. |
Problem
Every lifecycle read that trusts a dynamic run's runtime controls re-derives them from the sealed base with the reading build's code (
readLinkedWorkflowEvidence→verifyWorkflowControlSnapshot→verifyDynamicRuntimeMaterialization). The re-derivation re-renders every runtime-rendered prompt: a generated child's, and a planned node's whose prompt waits on a dynamic group. It then byte-compares each one with the publishedartifacts/<attempt>/prompt.rendered.mdand throwsruntime rendered prompt changed for <attempt>on any difference. The controller'smaterializeDynamicRuntimemakes the same comparison on every render. That includes the first render of a resumed run, which loads the installed runtime (start-run.tssetsULTRAFUZZ_RUNTIME_MODULEtoimport.meta.resolve("@ultrafuzz/runtime")).So any change to renderer code, or to a projection the code generates, between the build that published a prompt and the build reading it strands the run:
{{coverage_evidence_markdown_projection}}text, whichrenderPromptreads from the installed build's bundled_templates/output-contract/coverage-evidence-markdown.mdx.review/final-report.mdembeds it.final-reporthasthreat-goalsandclass-goalsas dynamic ancestors, so its prompt is rendered and published as soon as both groups expand, right aftergoal-plan. So every packaged-topology run launched on v0.1.1 that is pastgoal-planholds the v0.1.1 text. Current main (the upcoming v0.1.2) cannot synchronize or resume it. Neither fix(runtime): stop host artifact gates rejecting valid campaign output #1176 nor fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195 is in a tag yet.{{ancestor_artifact_path_authority:…}}selector whose paths sort differently under host collation. The published prompt embeds the old selector ID.Measured with the new lifecycle test's fixture: a launched run past its planner whose two runtime-rendered prompts differ from the current render. Main is origin/main 142ba80, before the rebase. The later main commits up to 7728e7b do not touch
dynamic-runtime.ts,dynamic-expansion.ts,dynamic-expansion-retry.tsor the prompts package, and the unit tests below fail the same way on 7728e7b.resume,replay,forkand verified reads take)WORKFLOW_CONTROL_EVIDENCE_INVALID: published dynamic runtime controls no longer re-derive from their sealed base: runtime rendered prompt changed for strict-joinsyncRun, the synchronization pass behindstatus,inspect,whyandstatsWORKFLOW_PUBLISHED_PROMPT_DRIFTwarningstatusWORKFLOW_CONTROL_EVIDENCE_DIVERGEDandWORKFLOW_STATE_SYNC_SKIPPED, so run state is never synchronized againwhywhyenvelope, which is a harness artefact)materializeDynamicRuntime, first tick of a resume)runtime rendered prompt changed for …Read from the code, not measured, because the fixture's fake runner cannot serve their queries: on main
inspectdowngrades the failed synchronization to a warning and never synchronizes, andstatsfails withlinked workflow authority is invalid. On this branch both return the warning.pauseandcancelalready tolerate the divergence (#1170).This is the same class of problem as #1188 fixed for validator rebuilds: something an earlier build recorded was used as an equality gate against the reading build. It is #921's cost 3, and ask 3 there says an identity rotation should never strand a run already in progress.
Root cause
renderReadyRuntimePrompts(packages/runtime/src/dynamic-runtime.ts) treated "this build renders different bytes" as tampering. Runtime-rendered prompts are produced after the control seal, so that byte comparison was their only integrity check, and it cannot tell an upgrade from an edit.Change
dynamic-runtime.ts: when a runtime-rendered prompt is already published, the published bytes are kept and never rewritten. That holds when the current build renders different bytes, and also when it cannot render the prompt at all, because the render or its validator-command check throws. The attempt ID is returned in a newDynamicRuntimeMaterialization.promptDriftAttemptIds.renderPromptand the validator-command check run in a closure, and a failure there counts as drift.dynamic-expansion-retry.ts:resume --retry-failedon a dynamic source whose verifier failed withdraws the published expansion so the group expands again from the source's new output. It moved the manifests and the generated children's attempt state, but not the published prompt of a planned task whose prompt waits on the group, which was rendered from the withdrawn children. The archive now moves that prompt too, validated with the rest of the attempt state before the Smithers reset, and lists it inretry.json. Without this, adoption would have handed such a task the old prompt when the re-run source planned different items. An example is a custom join whose prompt uses{{artifact_path:<group>}}: the old prompt names a child whose directory is now in the history and omits the new one. On main the same case threwruntime rendered prompt changedand stranded the run. The stock topology is not affected, because its deferred prompts reach the goal groups only through authority selectors.workflow-integrity.ts:verifyWorkflowControlSnapshotcarries the list asVerifiedWorkflowControlSnapshot.promptDriftAttemptIds. It is not a divergence, so strict callers do not throw andstatusdoes not skip synchronization. A remembered snapshot is now reused only when its drift list is equal too (this is the WeakRef reuse next to the ultrafuzz status is unusable on every run: manifest re-derivation throws past tolerateControlDivergence (#674 incomplete) #866 comment). Without that condition, a snapshot remembered before the drift appeared hides it. The lifecycle test covers this case.workflow-sync.ts:synchronizeLinkedWorkflowRunadds oneWORKFLOW_PUBLISHED_PROMPT_DRIFTwarning per pass. It gives the count and up to three attempt IDs, for example2 published runtime prompt(s) do not match what this build renders and were kept as published: strict-join, dynamic-fanout-45b2….status,inspect,whyandstatsreturn the pass's diagnostics.docs/reference/artifacts-reports.mdsays what is still checked, what is adopted and where the warning appears.docs/reference/topology-yaml.mdsays what a retry withdraws. The comment insmithers.ts(currentControllerPromptBindings) is updated too.Every other re-derivation check is unchanged:
No new files, events, seals or digests.
The warning comes from synchronization, not from the verifier's own diagnostics. The only channel the verifier has is
divergences: a divergence fails strict callers and makesstatusskip synchronization, which is the opposite of what is needed here. A second list on the verified snapshot plus one diagnostic in the synchronizer is the least plumbing that reachesstatus,inspect,whyandstats.Trade-off for the owner
This is a threat-model decision, which is why the PR is a draft.
artifacts/<attempt>/prompt.rendered.md, including before the controller first renders it, and that file is adopted as published. Examples are an agent in a skip-permissions workspace, or anything else running as the operator. On main, a file that differed from the fresh render failed every later render and every gated lifecycle command. Here it is only reported as drift:WORKFLOW_PUBLISHED_PROMPT_DRIFTwarning in the--jsondiagnostics ofstatus,inspect,whyandstats.statsalso prints it; the other three print no warnings on success.dynamic-prompt-templates/after its prompt was published, including one that makes the template unrenderable. Main caught that only indirectly: through this byte comparison, or because the re-render threw.plan.jsondigests and the sealed execution snapshot. Dynamic-group manifests and group templates, the runtime graph and the task plan still fail closed. Main never checked two things before a prompt was first published, and they stay unchecked: an edit to a planned node's template snapshot underdynamic-prompt-templates/(digest-checked only at compile time), and an edit to the inputs the task reads.goal-plan.statusdegrades to divergence warnings and never synchronizes state. Keeping the check means treating every renderer or projection change as a run-stranding break, or shipping old renderers alongside new ones.dynamic-runtime.tshas no such change. Main keeps a compatibility test for manifests published before the sequence field. No real v0.1.1 run was resumed end to end.Deliberately not built
eventsobservers (queryWorkflowEvents,watchWorkflowEvents). They report control divergences but do not synchronize.status,inspectandwhyoutput. LikeWORKFLOW_CONTROL_EVIDENCE_DIVERGED, the warning is in the result's diagnostics (--json).commandFromRuntimeprints diagnostics only for a failed result, so printing warnings on success would change every command's output.Verification
Discriminating tests. Each fails with this branch's test files on origin/main's
srcand passes here. For the twodynamic-expansion.test.tstests that is origin/main 7728e7b. For the lifecycle retry assertions it is origin/main'sdynamic-expansion-retry.tsswapped into this branch. The lifecycle drift test was checked on 142ba80. On maintscalso reports the missingpromptDriftAttemptIdsfield; the emitted JS is what fails.dynamic-expansion.test.ts, "a published prompt this build renders differently or cannot render is kept for renders and admission":verifyDynamicRuntimeMaterializationandmaterializeDynamicRuntimeboth return[join, <generated>]inpromptDriftAttemptIdsand leave both files' bytes unchanged. Once the join's prompt is removed, publishing it still fails withunknown prompt template variable. A tampered manifest item still fails withDYNAMIC_MANIFEST_INVALID, and an edited group template withDYNAMIC_TEMPLATE_CHANGED.unknown prompt template variable: variable_a_later_build_removed, because the join renders first. Before the second part was added, the test failed on main withError: runtime rendered prompt changed for dynamic-fanout-4650387e2fa7c4800d4aeacfd625d35d. The first commit alone also fails withunknown prompt template variable.dynamic-expansion.test.ts, "explicit source retry re-derives the base runtime controls after archiving an expansion":Join {{artifact_path:fanout}}., like the stock final report.dynamic-lifecycle.test.ts, "explicit source retry prunes the withdrawn generation from run state and keeps observers admitted", extended:resume --retry-failedof a launched run. The compiler marks its join's prompt as deferred, so this does not rely on the unit fixture's hand-built task.dynamic-expansion-retry.ts:expected false, actual trueon the prompt still being in place.dynamic-lifecycle.test.ts, "runtime prompts an earlier build published keep the run synchronizable and are reported":syncRunandgetRunHealthsucceed. Each reports exactly oneWORKFLOW_PUBLISHED_PROMPT_DRIFTwarning, and no control-evidence or sync-skipped diagnostic. The prompt bytes are unchanged.WORKFLOW_CONTROL_EVIDENCE_INVALID: published dynamic runtime controls no longer re-derive from their sealed base: runtime rendered prompt changed for strict-join.These existing tamper tests still pass unchanged:
Repro scripts, run against both builds:
repro-f1.mjs(the fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195 selector case with a join template) andrepro-f1-cov.mjswithCOV=1. The latter rewrites the coverage projection to the v0.1.1 text, which I checked againstv0.1.1:.ultrafuzz/prompts/_templates/output-contract/coverage-evidence-markdown.mdx. Main: both the verify and the publish path throwruntime rendered prompt changed for join. Here: both return, and the reviewer'scov.mjsvariant withCOV=1prints drift[join, dynamic-fanout-…]for both. The scripts do not check the bytes; the unit tests above do.retry-stale.mjs: main throwsruntime rendered prompt changed for join. The first commit alone kept the stale join prompt as drift. Here the join's prompt is archived, re-rendered with the new child only, and there is no drift.render-throws.mjs: main and the first commit alone throwunknown prompt template variable. Here both paths return with drift[join], and the bytes are unchanged.Suites,
@ultrafuzz/runtime, at the head (rebased on origin/main 7728e7b):dynamic-expansion17/17 anddynamic-workflow2/2dynamic-lifecycle, whole file: 20/20cloud-worker-handoffplusverified-output: 64/64lifecycle-inspection: 55/55runtime.test.ts, the fiveresumeretry tests: "resume derives reset identities from the canonical nodes of a failed workflow", "resume retries a failed artifact verifier from its agent producer and dependent closure" (the retry archive throughresumeRun), "resume retries stalled nodes alongside failed ones", "resume retries failed tasks reported inside a successful terminal workflow" and "resume --retry-failed after a pre-agent failure keeps synchronizing the reused attempt": 5/5Checks:
npx prettier --checkandnpx eslinton the changed files,CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci,pnpm -w lint,pnpm --filter @ultrafuzz/runtime typecheck,pnpm -w knipandnode scripts/docs-check.mjs.Not run:
runtime.test.tsand the other packages' suites (the host is shared)materializeDynamicRuntimecall the controller makes.Risk / compatibility
DynamicRuntimeMaterializationandVerifiedWorkflowControlSnapshotgain a requiredpromptDriftAttemptIdsfield.verifyWorkflowControlSnapshotis the only producer of the snapshot, andstart-run.ts's spread keeps the field. Every workspace package is private and no in-tree code builds either object by hand, so this is not marked as a breaking change.runtime rendered prompt changed for …disappears;runtime rendered prompt is missing for …stays.retry.json'sarchived_attempt_pathscan now also listartifacts/<attempt>/prompt.rendered.mdfor planned tasks. A symlinked or non-regular prompt file there refuses the retry before the Smithers reset, like the other attempt state.ULTRAFUZZ_RUNTIME_MODULEpoints at the installed runtime, so the fix reaches v0.1.1 runs resumed by this build. It needs no--refresh-controller.renderReadyRuntimePrompts/deriveDynamicRuntimeindynamic-runtime.ts(therenderPromptinput object keeps its lines; only the call and the validator-command check move into a closure),collectAttemptStateMovesindynamic-expansion-retry.ts,verifyWorkflowControlSnapshotinworkflow-integrity.ts, one line plus a helper inworkflow-sync.ts, and a comment insmithers.ts.Changelog entry
A dynamic run past its expansion, such as a packaged-topology run past
goal-plan, no longer strands after an upgrade that changes how a prompt renders (#1176, #1195); before,statusandinspectstopped synchronizing it andwhy,statsandresumefailed. A runtime-rendered prompt that is already published is now kept as published, andstatus,inspect,whyandstatsreturn aWORKFLOW_PUBLISHED_PROMPT_DRIFTwarning (in--json;statsalso prints it), which is also the only sign left when another process running as the same user writes such a prompt. Aresume --retry-failedthat re-expands a dynamic group now renders the later prompts that wait on it again, instead of failing on the ones rendered from the withdrawn expansion.Refs #921, #1188, #1195, #1176
🤖 Generated with Claude Code