Skip to content

fix(runtime): a renderer or projection change no longer strands runs whose dynamic groups expanded - #1216

Closed
aviggiano wants to merge 3 commits into
mainfrom
claude/x06-published-prompts-survive-upgrades
Closed

aviggiano wants to merge 3 commits into
mainfrom
claude/x06-published-prompts-survive-upgrades

Conversation

@aviggiano

Copy link
Copy Markdown
Collaborator

Draft for v0.1.3, owner decision. The owner's policy is that 0.x releases may break in-flight runs without fallbacks. So this PR's main motivation, keeping v0.1.1 runs resumable after upgrading to v0.1.2, does not apply to v0.1.2. It remains useful only if the owner wants a renderer or prompt-asset change to stop stranding in-flight runs on later upgrades, at the cost of the tamper check described under "Trade-off for the owner".

The second commit (a source retry withdraws the prompts rendered from the withdrawn expansion) also fixes a bug on main that is independent of upgrades. It is being split into its own PR for v0.1.2. Once that lands, this branch will be rebased onto it.

Problem

Every lifecycle read that trusts a dynamic run's runtime controls re-derives them from the sealed base with the reading build's code (readLinkedWorkflowEvidence → verifyWorkflowControlSnapshot → verifyDynamicRuntimeMaterialization). The re-derivation re-renders every runtime-rendered prompt: a generated child's, and a planned node's whose prompt waits on a dynamic group. It then byte-compares each one with the published artifacts/<attempt>/prompt.rendered.md and throws runtime rendered prompt changed for <attempt> on any difference. The controller's materializeDynamicRuntime makes the same comparison on every render. That includes the first render of a resumed run, which loads the installed runtime (start-run.ts sets ULTRAFUZZ_RUNTIME_MODULE to import.meta.resolve("@ultrafuzz/runtime")).

So any change to renderer code, or to a projection the code generates, between the build that published a prompt and the build reading it strands the run:

Measured with the new lifecycle test's fixture: a launched run past its planner whose two runtime-rendered prompts differ from the current render. Main is origin/main 142ba80, before the rebase. The later main commits up to 7728e7b do not touch dynamic-runtime.ts, dynamic-expansion.ts, dynamic-expansion-retry.ts or the prompts package, and the unit tests below fail the same way on 7728e7b.

operation origin/main this branch
strict linked evidence (what resume, replay, fork and verified reads take) WORKFLOW_CONTROL_EVIDENCE_INVALID: published dynamic runtime controls no longer re-derive from their sealed base: runtime rendered prompt changed for strict-join ok
syncRun, the synchronization pass behind status, inspect, why and stats fails with the same error ok, one WORKFLOW_PUBLISHED_PROMPT_DRIFT warning
status ok, but WORKFLOW_CONTROL_EVIDENCE_DIVERGED and WORKFLOW_STATE_SYNC_SKIPPED, so run state is never synchronized again ok and synchronized, with the same warning
why fails with the same error the same warning (the fixture's fake runner then returns an invalid why envelope, which is a harness artefact)
controller render (materializeDynamicRuntime, first tick of a resume) throws runtime rendered prompt changed for … returns and lists the drift

Read from the code, not measured, because the fixture's fake runner cannot serve their queries: on main inspect downgrades the failed synchronization to a warning and never synchronizes, and stats fails with linked workflow authority is invalid. On this branch both return the warning.

pause and cancel already tolerate the divergence (#1170).

This is the same class of problem as #1188 fixed for validator rebuilds: something an earlier build recorded was used as an equality gate against the reading build. It is #921's cost 3, and ask 3 there says an identity rotation should never strand a run already in progress.

Root cause

renderReadyRuntimePrompts (packages/runtime/src/dynamic-runtime.ts) treated "this build renders different bytes" as tampering. Runtime-rendered prompts are produced after the control seal, so that byte comparison was their only integrity check, and it cannot tell an upgrade from an edit.

Change

  • dynamic-runtime.ts: when a runtime-rendered prompt is already published, the published bytes are kept and never rewritten. That holds when the current build renders different bytes, and also when it cannot render the prompt at all, because the render or its validator-command check throws. The attempt ID is returned in a new DynamicRuntimeMaterialization.promptDriftAttemptIds.
    • The fresh render of a published prompt is only compared, never used. So renderPrompt and the validator-command check run in a closure, and a failure there counts as drift.
    • A prompt that is not yet published is rendered, checked and published exactly as before, and a render that throws for it still fails. A missing published prompt still throws in verify mode.
    • The render input is still built for every ready task. Building it reads the template and checks the vulnerability-catalog digest, so both still fail closed.
    • The controller ignores the list: it adopts silently, which the code comment at the site documents. It has no diagnostics channel and renders on every tick, and synchronization reports the same list.
  • dynamic-expansion-retry.ts: resume --retry-failed on a dynamic source whose verifier failed withdraws the published expansion so the group expands again from the source's new output. It moved the manifests and the generated children's attempt state, but not the published prompt of a planned task whose prompt waits on the group, which was rendered from the withdrawn children. The archive now moves that prompt too, validated with the rest of the attempt state before the Smithers reset, and lists it in retry.json. Without this, adoption would have handed such a task the old prompt when the re-run source planned different items. An example is a custom join whose prompt uses {{artifact_path:<group>}}: the old prompt names a child whose directory is now in the history and omits the new one. On main the same case threw runtime rendered prompt changed and stranded the run. The stock topology is not affected, because its deferred prompts reach the goal groups only through authority selectors.
  • workflow-integrity.ts: verifyWorkflowControlSnapshot carries the list as VerifiedWorkflowControlSnapshot.promptDriftAttemptIds. It is not a divergence, so strict callers do not throw and status does not skip synchronization. A remembered snapshot is now reused only when its drift list is equal too (this is the WeakRef reuse next to the ultrafuzz status is unusable on every run: manifest re-derivation throws past tolerateControlDivergence (#674 incomplete) #866 comment). Without that condition, a snapshot remembered before the drift appeared hides it. The lifecycle test covers this case.
  • workflow-sync.ts: synchronizeLinkedWorkflowRun adds one WORKFLOW_PUBLISHED_PROMPT_DRIFT warning per pass. It gives the count and up to three attempt IDs, for example 2 published runtime prompt(s) do not match what this build renders and were kept as published: strict-join, dynamic-fanout-45b2…. status, inspect, why and stats return the pass's diagnostics.
  • Docs: docs/reference/artifacts-reports.md says what is still checked, what is adopted and where the warning appears. docs/reference/topology-yaml.md says what a retry withdraws. The comment in smithers.ts (currentControllerPromptBindings) is updated too.

Every other re-derivation check is unchanged:

  • The runtime graph and task plan must still re-derive to the same whole-document fingerprint.
  • Each expansion manifest must still validate and match its group: template digest, fingerprint, source and key paths, node-ID template and limits.
  • The group template digest, path confinement and symlink checks still apply, as do the vulnerability-catalog digest and, for a prompt not yet published, the render and its validator-command check.
  • A missing published prompt still fails admission (a render republishes it).

No new files, events, seals or digests.

The warning comes from synchronization, not from the verifier's own diagnostics. The only channel the verifier has is divergences: a divergence fails strict callers and makes status skip synchronization, which is the opposite of what is needed here. A second list on the verified snapshot plus one diagnostic in the synchronizer is the least plumbing that reaches status, inspect, why and stats.

Trade-off for the owner

This is a threat-model decision, which is why the PR is a draft.

  • What is lost. A same-UID process can write or rewrite a later task's artifacts/<attempt>/prompt.rendered.md, including before the controller first renders it, and that file is adopted as published. Examples are an agent in a skip-permissions workspace, or anything else running as the operator. On main, a file that differed from the fresh render failed every later render and every gated lifecycle command. Here it is only reported as drift:
    • The report is one WORKFLOW_PUBLISHED_PROMPT_DRIFT warning in the --json diagnostics of status, inspect, why and stats. stats also prints it; the other three print no warnings on success.
    • The same attempt IDs are listed on every synchronization for the rest of the run. Once an upgrade has caused drift, a later edit to one of those attempts adds nothing new to the warning.
    • Validator commands are included. For a published prompt, the validator-command check runs on the discarded fresh render, not on the bytes the agent receives, so an edited validator command in a published prompt is no longer caught. Host verification still runs after the agent finishes.
    • The same applies to an in-place edit of a planned node's template snapshot under dynamic-prompt-templates/ after its prompt was published, including one that makes the template unrenderable. Main caught that only indirectly: through this byte comparison, or because the re-render threw.
    • Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921 ask 1 is exactly this question: whether in-flight tamper detection is in the threat model. Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921 also notes that no tamper event has been reported. The failures this check has caused came from upgrades.
  • What is not weakened. Plan-time prompts stay bound by their plan.json digests and the sealed execution snapshot. Dynamic-group manifests and group templates, the runtime graph and the task plan still fail closed. Main never checked two things before a prompt was first published, and they stay unchecked: an edit to a planned node's template snapshot under dynamic-prompt-templates/ (digest-checked only at compile time), and an edit to the inputs the task reads.
  • If this is not merged. v0.1.2 can neither synchronize nor resume any v0.1.1 run past dynamic expansion, which means every packaged-topology run past goal-plan. status degrades to divergence warnings and never synchronizes state. Keeping the check means treating every renderer or projection change as a run-stranding break, or shipping old renderers alongside new ones.
  • What can still strand a run across an upgrade. A later build whose renderer rejects a published prompt no longer does. What remains are the other re-derivations: a build that derives a different task plan or graph from the same sealed base and manifests, or that rejects an older manifest, still fails admission. The v0.1.1 to main diff of dynamic-runtime.ts has no such change. Main keeps a compatibility test for manifests published before the sequence field. No real v0.1.1 run was resumed end to end.

Deliberately not built

  • Telling an upgrade from an edit. One option is to record each published prompt's digest and the build that rendered it, and adopt only when the build changed. That is a new digest or seal layer, and the same same-UID process can rewrite that record too.
  • A warning from the controller. It renders reactively on every tick and has no diagnostics channel. Synchronization reports the same list.
  • The warning in the events observers (queryWorkflowEvents, watchWorkflowEvents). They report control divergences but do not synchronize.
  • Human-readable status, inspect and why output. Like WORKFLOW_CONTROL_EVIDENCE_DIVERGED, the warning is in the result's diagnostics (--json). commandFromRuntime prints diagnostics only for a failed result, so printing warnings on success would change every command's output.
  • Reporting each drift only once. That needs a record of what was already reported, which is new run state.
  • A digest check on deferred template snapshots at render time. That is the pre-existing gap noted above, and closing it would add fail-closed machinery.

Verification

Discriminating tests. Each fails with this branch's test files on origin/main's src and passes here. For the two dynamic-expansion.test.ts tests that is origin/main 7728e7b. For the lifecycle retry assertions it is origin/main's dynamic-expansion-retry.ts swapped into this branch. The lifecycle drift test was checked on 142ba80. On main tsc also reports the missing promptDriftAttemptIds field; the emitted JS is what fails.

  • dynamic-expansion.test.ts, "a published prompt this build renders differently or cannot render is kept for renders and admission":
    • Setup, first part: a generated prompt whose multi-path selector ID is rewritten to the pre-fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195 host-collation ID.
    • Setup, second part: the join's prompt is published, and then its template gains a variable the renderer does not know. This stands in for a later build that refuses what an earlier one rendered.
    • Checks: verifyDynamicRuntimeMaterialization and materializeDynamicRuntime both return [join, <generated>] in promptDriftAttemptIds and leave both files' bytes unchanged. Once the join's prompt is removed, publishing it still fails with unknown prompt template variable. A tampered manifest item still fails with DYNAMIC_MANIFEST_INVALID, and an edited group template with DYNAMIC_TEMPLATE_CHANGED.
    • On main: unknown prompt template variable: variable_a_later_build_removed, because the join renders first. Before the second part was added, the test failed on main with Error: runtime rendered prompt changed for dynamic-fanout-4650387e2fa7c4800d4aeacfd625d35d. The first commit alone also fails with unknown prompt template variable.
  • dynamic-expansion.test.ts, "explicit source retry re-derives the base runtime controls after archiving an expansion":
    • Setup: the fixture's join now has a deferred prompt, Join {{artifact_path:fanout}}., like the stock final report.
    • Checks: the retry archive moves the join's prompt into the history. After the source plans a different item, the re-expansion renders the join from the new child, not the withdrawn one, with no drift, and verification agrees.
    • On main and on the first commit alone: the join's prompt is still in place after the archive.
  • dynamic-lifecycle.test.ts, "explicit source retry prunes the withdrawn generation from run state and keeps observers admitted", extended:
    • Setup: a real resume --retry-failed of a launched run. The compiler marks its join's prompt as deferred, so this does not rely on the unit fixture's hand-built task.
    • Checks: after the resume, the join's prompt is in the archive and no longer in place, and the re-expansion renders it again.
    • With origin/main's dynamic-expansion-retry.ts: expected false, actual true on the prompt still being in place.
  • dynamic-lifecycle.test.ts, "runtime prompts an earlier build published keep the run synchronizable and are reported":
    • Setup: a launched dynamic run past its planner whose join prompt and generated prompt are rewritten. The join is deferred the same way as the stock final report.
    • Checks: strict linked evidence is admitted, and syncRun and getRunHealth succeed. Each reports exactly one WORKFLOW_PUBLISHED_PROMPT_DRIFT warning, and no control-evidence or sync-skipped diagnostic. The prompt bytes are unchanged.
    • The test reads and holds evidence once before the rewrite, so the snapshot reuse path runs. With the new reuse condition removed, the test fails: expected 1 drift warning, got 0.
    • On main: WORKFLOW_CONTROL_EVIDENCE_INVALID: published dynamic runtime controls no longer re-derive from their sealed base: runtime rendered prompt changed for strict-join.

These existing tamper tests still pass unchanged:

  • "dynamic lifecycle admission rejects graph and task extensions not derived from sealed controls" (edits to the task plan and graph)
  • "a half-published dynamic expansion stays readable while execution stays closed" (a missing prompt)
  • "persisted expansion rejects tampering, transplantation, and symlink manifests"
  • "persisted expansion rejects prompt-template, topology-contract, and dynamic-limit changes"

Repro scripts, run against both builds:

  • The triage agent's repro-f1.mjs (the fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195 selector case with a join template) and repro-f1-cov.mjs with COV=1. The latter rewrites the coverage projection to the v0.1.1 text, which I checked against v0.1.1:.ultrafuzz/prompts/_templates/output-contract/coverage-evidence-markdown.mdx. Main: both the verify and the publish path throw runtime rendered prompt changed for join. Here: both return, and the reviewer's cov.mjs variant with COV=1 prints drift [join, dynamic-fanout-…] for both. The scripts do not check the bytes; the unit tests above do.
  • The reviewer's retry-stale.mjs: main throws runtime rendered prompt changed for join. The first commit alone kept the stale join prompt as drift. Here the join's prompt is archived, re-rendered with the new child only, and there is no drift.
  • The reviewer's render-throws.mjs: main and the first commit alone throw unknown prompt template variable. Here both paths return with drift [join], and the bytes are unchanged.

Suites, @ultrafuzz/runtime, at the head (rebased on origin/main 7728e7b):

  • dynamic-expansion 17/17 and dynamic-workflow 2/2
  • dynamic-lifecycle, whole file: 20/20
  • cloud-worker-handoff plus verified-output: 64/64
  • lifecycle-inspection: 55/55
  • runtime.test.ts, the five resume retry tests: "resume derives reset identities from the canonical nodes of a failed workflow", "resume retries a failed artifact verifier from its agent producer and dependent closure" (the retry archive through resumeRun), "resume retries stalled nodes alongside failed ones", "resume retries failed tasks reported inside a successful terminal workflow" and "resume --retry-failed after a pre-agent failure keeps synchronizing the reused attempt": 5/5

Checks: npx prettier --check and npx eslint on the changed files, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, pnpm -w lint, pnpm --filter @ultrafuzz/runtime typecheck, pnpm -w knip and node scripts/docs-check.mjs.

Not run:

  • the full runtime.test.ts and the other packages' suites (the host is shared)
  • a real Smithers engine resuming an actual v0.1.1 run. The controller path is covered at the materializeDynamicRuntime call the controller makes.

Risk / compatibility

  • API. DynamicRuntimeMaterialization and VerifiedWorkflowControlSnapshot gain a required promptDriftAttemptIds field. verifyWorkflowControlSnapshot is the only producer of the snapshot, and start-run.ts's spread keeps the field. Every workspace package is private and no in-tree code builds either object by hand, so this is not marked as a breaking change.
  • Behaviour. A published runtime prompt that differs from the current render, or that the current build cannot render, is no longer an admission failure anywhere. The text runtime rendered prompt changed for … disappears; runtime rendered prompt is missing for … stays.
  • Retry archive. retry.json's archived_attempt_paths can now also list artifacts/<attempt>/prompt.rendered.md for planned tasks. A symlinked or non-regular prompt file there refuses the retry before the Smithers reset, like the other attempt state.
  • Runs rendered by an earlier build. On resume, ULTRAFUZZ_RUNTIME_MODULE points at the installed runtime, so the fix reaches v0.1.1 runs resumed by this build. It needs no --refresh-controller.
  • Cost. The verifier already rendered every prompt. The only additions are comparing one short list when deciding whether to reuse a snapshot, and one existence check per planned deferred task during a retry archive.
  • Merge surface. Small, local hunks: renderReadyRuntimePrompts/deriveDynamicRuntime in dynamic-runtime.ts (the renderPrompt input object keeps its lines; only the call and the validator-command check move into a closure), collectAttemptStateMoves in dynamic-expansion-retry.ts, verifyWorkflowControlSnapshot in workflow-integrity.ts, one line plus a helper in workflow-sync.ts, and a comment in smithers.ts.

Changelog entry

A dynamic run past its expansion, such as a packaged-topology run past goal-plan, no longer strands after an upgrade that changes how a prompt renders (#1176, #1195); before, status and inspect stopped synchronizing it and why, stats and resume failed. A runtime-rendered prompt that is already published is now kept as published, and status, inspect, why and stats return a WORKFLOW_PUBLISHED_PROMPT_DRIFT warning (in --json; stats also prints it), which is also the only sign left when another process running as the same user writes such a prompt. A resume --retry-failed that re-expands a dynamic group now renders the later prompts that wait on it again, instead of failing on the ones rendered from the withdrawn expansion.

Refs #921, #1188, #1195, #1176

🤖 Generated with Claude Code

aviggiano and others added 3 commits September 29, 2026 15:30
…whose dynamic groups expanded

Every render and every lifecycle admission re-derives the published
dynamic runtime controls from their sealed base with the current code,
including re-rendering each runtime-rendered prompt, and refused a
published prompt.rendered.md whose bytes differed from the fresh render
("runtime rendered prompt changed for <attempt>"). So any renderer or
projection change between the build that published a prompt and the
reading build stranded the run: why, stats, resume, replay, fork and
verified reads failed with WORKFLOW_CONTROL_EVIDENCE_INVALID, status and
inspect stopped synchronizing it, and a resumed controller threw on its
first render. #1176 changed the coverage projection that the stock final
report embeds, and the final report renders as soon as the goal groups
expand, so current main can neither synchronize nor resume a
packaged-topology run that v0.1.1 took past goal-plan. #1195's selector
order does the same for custom prompts with multi-path selectors.

A published prompt is what its task was, or will be, handed, so it is
now kept as published whatever this build renders: renderReadyRuntimePrompts
returns the drifted attempt IDs instead of throwing. The controller
adopts them silently. verifyWorkflowControlSnapshot carries them on the
verified snapshot (and only reuses a remembered snapshot whose drift list
matches), and synchronization adds one WORKFLOW_PUBLISHED_PROMPT_DRIFT
warning per pass naming the count and up to three attempts, so status,
inspect, why and stats return it. Every other re-derivation check is
unchanged: a missing prompt, a changed group template or expansion
manifest, and a task plan or graph that no longer re-derives still fail
closed. A prompt not yet published is rendered and published exactly as
before.

The cost: a prompt.rendered.md that a same-UID process writes or rewrites
for a later task, even before the controller first renders it, is now
adopted and reported as drift rather than refused.

Refs #921, #1188, #1195, #1176

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…withdrawn expansion

`resume --retry-failed` on a dynamic source whose verifier failed withdraws
the published expansion generation into `dynamic-expansion-history/`, so the
group expands again from the source's new output. The archive moved only the
generation-owned attempts. It left the published `prompt.rendered.md` of every
planned task whose prompt waits on the group, and that prompt was rendered
from the withdrawn children.

When the re-run source planned different items, the next render of such a
task, for example a custom join whose prompt uses `{{artifact_path:<group>}}`,
differed from the file on disk. On main that threw `runtime rendered prompt
changed for <attempt>` and stranded the run. With published prompts now kept
as published, the join would instead have been handed the old prompt: it named
a child whose artifact directory is now in the history and omitted the new one.
The stock topology is not affected, because its deferred prompts reach the goal
groups only through authority selectors.

The archive now also moves the published prompt of each sealed base task whose
`deferredPromptGroups` names a withdrawn group, validated with the rest of the
attempt state before the Smithers reset. The next expansion renders it afresh
against the new generation, and `retry.json` lists it with the other archived
paths.

The re-derivation test now gives the fixture's join a deferred prompt that
names the group's children, and checks that the retry archives it and that
the re-expansion renders it from the new child with no drift. The lifecycle
retry test checks the same through a real `resume --retry-failed` of a
launched run, whose compiled base tasks mark the join's prompt as deferred:
the prompt is in the archive after the resume and is rendered again by the
re-expansion. Both fail with origin/main's retry module and on the previous
head: the join's prompt is still in place after the archive.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…kept too

The previous commit kept a published runtime prompt that the current build
renders differently. It still rendered every published prompt first, and a
render or validator-command check that threw for one failed both the
controller's render and every lifecycle admission. So a later build whose
renderer refuses what an earlier build rendered, for example over a removed
variable or a stricter validator-command rule, would still strand every run
whose dynamic groups had expanded. That fresh render of a published prompt is
only ever compared and then discarded, so it has no reason to fail a run.

`renderReadyRuntimePrompts` still builds the render input for every ready
task, but runs `renderPrompt` and the validator-command check in a closure. A
published prompt whose render throws is recorded as drift and kept like one
that renders differently. A prompt that is not published yet is rendered,
checked and published exactly as before, and a render that throws for it
still fails. Building the input reads the template and checks the
vulnerability-catalog digest, so both still fail closed for every ready task.
The warning text now says the prompts "do not match what this build renders",
which covers both cases.

Review wording fixes:
- docs/reference/artifacts-reports.md names the commands that synchronize a
  run (`status`, `inspect`, `why`, `stats`; there is no `sync` command) and
  says the warning is in their `--json` diagnostics, printed only by `stats`.
  It says a planned node renders from its template snapshot, which is not
  sealed, and that a missing prompt fails admission while a render
  republishes it.
- The smithers.ts comment says "template snapshot" instead of "sealed
  template".

The drift test now also makes the join's published prompt unrenderable
and checks that it is kept and reported, and that the same prompt, once
removed, still fails to publish. It fails on origin/main and on the previous
head with `unknown prompt template variable`.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…withdrawn expansion

`resume --retry-failed` on a dynamic source whose verifier failed withdraws
the published expansion generation into `dynamic-expansion-history/`, so the
group expands again from the source's new output. The archive moved only the
generation-owned attempts. It left the published `prompt.rendered.md` of every
planned task whose prompt waits on the group, and that prompt was rendered
from the withdrawn children.

When the re-run source planned different items, the next render of such a
task, for example a custom join whose prompt uses `{{artifact_path:<group>}}`,
differed from the file on disk. The re-expansion then threw `runtime rendered
prompt changed for <attempt>`, and so did every later render and lifecycle
admission check, so the run stranded until the file was deleted by hand. The
stock topologies are not affected, because their deferred prompts reach the
goal groups only through authority selectors, which render the same bytes
whatever the children are.

The archive now also moves the published prompt of each sealed base task whose
`deferredPromptGroups` names a withdrawn group. It is validated with the rest
of the attempt state before the Smithers reset, the next expansion renders it
afresh against the new generation, and `retry.json` lists it with the other
archived paths. A prompt that was never rendered is skipped.

The re-derivation test gives the fixture's join a deferred prompt that names
the group's children and re-runs the source with a different item after the
retry. It checks that the re-expansion renders the join's prompt from the new
child, that the archive holds the withdrawn prompt, and that `retry.json`
lists it. With origin/main's retry module it fails at the re-expansion with
`runtime rendered prompt changed for join`. The lifecycle retry test checks
the same through a real `resume --retry-failed` of a launched run, whose
compiled base tasks mark the join's prompt as deferred: the prompt is in the
archive after the resume and is rendered again by the re-expansion. With
origin/main's retry module it fails because the join's prompt is still in
place after the resume.

Split out of draft #1216, whose change to how drifted published prompts are
handled this fix does not depend on.

Refs #1141

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano

Copy link
Copy Markdown
Collaborator Author

The second commit's fix, a source retry withdraws the prompts rendered from the withdrawn expansion, is split into #1220 so it can ship in v0.1.2. #1220 adapts the tests to main's throw-on-drift behaviour. Once #1220 merges, this draft should be rebased onto it; the duplicate hunks in dynamic-expansion-retry.ts then drop out.

@aviggiano

Copy link
Copy Markdown
Collaborator Author

Owner decision: tamper detection on prompts is not worth what it costs. The main cost is that a run cannot be restarted after a minor prompt change, and stability matters more than exact reproducibility. So this PR will not merge as is. A follow-up will remove prompt tamper-sealing outright and keep only a record of the prompt each attempt used, following the #1188 precedent (record, don't enforce). This PR closes when that replacement opens.

@aviggiano

Copy link
Copy Markdown
Collaborator Author

Superseded by #1230. That PR removes the prompt tamper-sealing outright instead of warning on drift, and it ports this PR's lifecycle test without the drift warning. Its follow-up, resume applying edited project prompts by default with an ultrafuzz.toml switch to turn it off, is stacked on #1230.

@aviggiano aviggiano closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant