refactor(runtime)!: a run's prompt files are records, not seals - #1230
Merged
Merged
Conversation
Base automatically changed from
claude/v04-remove-cloud-node-execution
to
main
September 30, 2026 15:29
An attempt's artifacts/<attempt>/prompt.rendered.md is now the prompt it receives on every verb, and nothing hashes, compares or seals it after launch. Prompt tamper detection cost more than it bought: an operator's edit of a task that had not run, or an upgrade that renders templates differently (#1176, #1195), stranded the run, and a static prompt had three different sources depending on the verb. Deleted: - the runtime byte comparison of published prompts, and the verify-mode "missing" throw (renderReadyRuntimePrompts renders only a missing file and otherwise adopts it; admission renders nothing); - the group-template digest gate (DYNAMIC_TEMPLATE_CHANGED) and its now-unused templatePath input; the manifest compatibility rows, which compare launch values with launch values, stay; - the prompt entries of the control seal and execution snapshot (controls/rendered-prompts/, controls/prompt-snapshots/), the workflow's sealed-copy read, and the --refresh-controller retained binding with its digest checks; - the digest gate on restoring a missing static prompt; - the dead dynamicRuntimePromptDigest export. Added: - a prompt problem fails only its task: a runtime prompt that cannot be rendered is returned as promptRenderFailures instead of thrown, a missing prompt file reads as empty in the render, and assert-task-inputs fails that task with the renderer's message or a "is missing; ultrafuzz resume restores ..." message. Every other task, status and sync keep working; - a missing static prompt is restored from prompt-snapshots/ before every engine start (resume, replay and fork), never over an existing file; - the #1220 retry archive moves the withdrawn generation's attempt state before it renames the manifests, so an interrupted archive leaves the published generation in place instead of a withdrawn item's prompt beside a new manifest. The docs say what is no longer checked and how to change a prompt of a running campaign; the CHANGELOG entry replaces the #1176/#1195 upgrade note, whose failure this removes. Runs launched before this change keep their launch engine's behaviour until they are resumed. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Only a missing file at a prompt path was scoped to its task. Any other entry there, such as a directory, a symlink or an unreadable file, made promptForTask throw on every render and stopped the whole run. Now that every engine reads the run's agent-writable prompt file instead of a sealed copy, the task's own agent or an agent of a dependent task can cause that. - The workflow reads each prompt with readRegularFileSnapshot, which follows no symlink and never blocks on a FIFO. A prompt that cannot be read reads as empty and its cause is kept per attempt, so assert-task-inputs fails that task with "rendered prompt for <attempt> could not be read: <cause>". - The post-agent verify pass no longer checks the prompt. The agent has already received it, so a later edit, deletion or render failure can no longer fail completed work. Preparation and the first-generation reset, which run before the agent, still check it. - A published runtime prompt that is not a regular file is reported as its task's render failure instead of thrown from every publishing render, and admission no longer opens runtime prompt entries at all. A symlinked artifact directory still stops the render. - Restore skips any entry at a static prompt path (lstat), so a dangling symlink there no longer makes resume, replay and fork refuse to start. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…thdrawal Moving the withdrawn generation's attempt state before the manifest rename closed the stale-prompt hole, but it left an interrupted withdrawal with no way to finish. Smithers has already reset the retried source by then, so a later `resume --retry-failed` finds no failed task and plans nothing, and the published manifest keeps the group at its old items. The retried source's new output was then ignored for the rest of the run. archiveDynamicExpansionsForRetry now records the withdrawal in `dynamic-expansions/.retry-withdrawal.json` (the source node IDs, the history directory and the planned moves) before anything moves. The record lives in the manifest directory, so the manifest rename takes it along, and while it is still there the withdrawal is unfinished. Before resume, replay or fork starts an engine, finishInterruptedDynamicExpansionRetry re-validates the manifests as a new plan would and completes the recorded withdrawal into the same history entry. An entry that a render recreated after it had moved belongs to the withdrawn generation too, and is archived beside the first copy. Co-Authored-By: Claude Opus 5.5 <[email protected]>
… and records Review follow-ups to the prompt-records change. - restart-continue.md: a template copy is shared by every task launched with the same text, so an edit to it reaches all of them that have not rendered yet. `--retry-failed` of a dynamic source withdraws edits to the prompts that wait on the group, which then render again from their template copies. `--reset-node` and `--retry-failed` restart the engine's attempt numbers, and each rerun attempt replaces the engine's row of the earlier attempt with the same number, prompt included: run `ultrafuzz status` first and keep a copy of the prompt file. A run launched by an earlier release gets task-scoped prompt failures only with `resume --refresh-controller`, because a plain resume continues its launch workflow. - security.md also names the template copies under dynamic-prompt-templates/ and the launch copies under prompt-snapshots/ as unchecked, says the catalog digest is now checked only when a prompt is rendered, and no longer follows "a prompt file edited after launch is not acknowledged again" with "Any change makes an acknowledgement stale". - backends-safety.md and prompt-variables.md match security.md and SPECS.md, and dashboard.md says the node view shows the prompt the task's next attempt receives. cli.md and artifacts-reports.md name the unreadable and non-regular prompt files that now fail only their task. - The CHANGELOG entry names an agent in another task as the realistic actor, the `--refresh-controller` route for earlier runs, and what the dashboard's rendered prompt is. Co-Authored-By: Claude Opus 5.5 <[email protected]>
dynamic-lifecycle's "--refresh-controller keeps a hand-edited static prompt" checked what the runtime.test.ts test of the same name checks end to end through resumeRun, which also edits the launch copy. Since renderCurrentSmithersController reads no prompt bytes, and two neighbouring dynamic-lifecycle tests already assert that the task spec binds the run's own prompt file, the copy caught no regression the others miss. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…t helpers dynamic-runtime.ts is already over the strict lint's 500-line budget on main (668 counted lines), and ESLint reports max-lines only on the 501st counted line. lint:strict:ci keeps that report only when its line is one the branch changed. Before the rebase it fell on resolvedConfigForRuntimeRoot, 11 lines past this PR's hunk. #1198 adds lines above it (optional-input lowering), so on current main it falls inside renderRuntimePrompt, which this PR extracts, and the gate fails over a file size this PR did not create. Move renderRuntimePrompt, byte for byte, below resolvedConfigForRuntimeRoot and promptGraphContext, the helpers it calls or whose return type it takes. The report now falls inside promptGraphContext, which this PR does not change. This is the fix #1062 used for lifecycle-inspection.ts; no eslint-disable is added, since #1184 chose not to keep a suppression baseline. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/prompt-records-not-seals
branch
from
September 30, 2026 16:18
0e6025d to
94123ad
Compare
…rness workspace-preparation-lifecycle.test.ts builds prepareArtifactMirror from workflow.tsx by evaluating every function declaration, but only the module-level variables whose names its allowlist regex matches. The first review commit made assert-task-inputs call assertTaskPromptInput, which reads the new module-level runtimePromptRenderFailures and taskPromptReadFailures maps. Neither name matched, so every preparation in the harness threw "prepare:worker failed at step assert-task-inputs: runtimePromptRenderFailures is not defined", and all 8 tests of the file failed in the runtime-supporting CI lane (on 0e6025d and on 94123ad; main passes it). Add both names to the allowlist. The production template is unchanged; the harness now evaluates the same empty maps the workflow starts with. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…fests moved The withdrawal record lived in dynamic-expansions/, so the manifest rename took it into the archive, where the next line deleted it. A failure or a kill in any step after the rename (re-deriving the runtime graph and task plan, pruning state.json, publishing retry.json) left nothing for finishInterruptedDynamicExpansionRetry to find. The task plan then no longer re-derived from the manifests, so replay, fork and synchronization refused the run until an engine render republished it, and nothing ever removed the withdrawn children's state records: a regenerated child that reuses a storage ID started from the archived record, which is what the prune is for. This is Greptile's P1 on #1230, reproduced by making smithers/ read-only so that the re-derivation fails with EACCES. Keep the record at dynamic-expansion-history/.retry-withdrawal.json, beside the entry it fills, from before the first move until after the last step. The manifest rename is the phase boundary: once <archive>/manifests exists, the completion skips the moves and the rename and repeats the steps after it, each safe to repeat (recreate the manifest directory, re-derive the runtime controls, publish retry.json unless an earlier completion did, prune state.json, remove the record). retry.json is now published before the prune, so a completion that finds the records already pruned still lists them; a copy of an entry that a render recreated is recorded before the rename, so a completion after it lists the copy too. This differs from the review's suggestion of a phase marker written into the recreated manifest directory: a kill between the rename and that marker's write would strand the withdrawal the same way, while a record outside the renamed directory leaves no such window. Tests: a new dynamic-expansion test stops the withdrawal before the task plan (read-only smithers/), before the manifest directory was recreated, and before the record's removal (an existing retry.json is kept). The dynamic-lifecycle interrupted-withdrawal test now also fails a post-rename step through resumeRun (a directory at graph.json), and checks the state prune and that strict admission refuses the run until the next resume. Its body moved into a loop over the two cases; git diff -w shows the change. On the previous commit both fail: the completion returns undefined and retry.json is never written. The new unit test sits at the end of dynamic-expansion.test.ts: placed beside the interrupted-archive test, it moved the file's max-lines report (1000 counted lines) onto a line this PR changed, and lint:strict:ci failed. Docs: CHANGELOG and topology-yaml.md name the new location, and say that replay, fork and synchronization refuse a run whose withdrawal stopped between the manifest rename and the task-plan re-derivation until a resume completes it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…e, and only files a later render reads Review of the prompt refresh found records rewritten that the rest of the resume could still refuse or withdraw, and a record of the refresh that a failure could leave out. Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them and just before the first timetravel; the completion of an earlier command's withdrawal runs first, so the refresh sees the run as the resets will find it. Its resets come from the same list the reset loop issues; continuationResets, a copy of that logic, is gone. A timetravel that fails after the refresh is the one window left, and the docs say so. Withdrawn generations. A `--retry-failed` of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them to dynamic-expansion-history/, so the archive no longer held what they ran with. The hook now names the groups the withdrawal plan covers; the refresh treats them as unexpanded, leaves their generation's prompts as they are, and refreshes the template copies the next expansion renders from. Template copies. Every copy was rewritten, even one no render would read again, which also broke replay and fork of runs launched before #1230 for groups that had finished. A copy is now rewritten only when a later render reads it: its group has not expanded or is withdrawn, or an unfinished task that renders from it has no published prompt. Prompts that share a copy must still agree on it, counting every prompt bound to the copy, and one pass settles that because a prompt has one copy. Record. refresh.json is published before the first file changes, so after a failure or a kill every file it lists holds its old or its new bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced instead of claiming replacements that did not happen. Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals. It now compares the effective topology file with plan.json's launch digest, and the topology's node set with the run's, which an eval run's exclusions change. A comment-only topology edit now skips the refresh too. Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order, the derivation sorts them), which skipped every refresh. The refresh uses a new derive-only mode instead. Also: one shared test helper for the toml key; the hook's context type is exported and reused; an unread failure field is dropped; the native resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message says the attach applied none of the project's current prompts. Docs: restart-continue.md starts with pause, wait, edit, resume, and covers the new rules, the reset window, catalog-wide errors, INCOMPLETE, the launch digests and replay and fork of runs launched before #1230. cli.md and artifacts-reports.md point there; the CHANGELOG entry, #1230's edit sentence, security.md and SPECS.md match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…e, and only files a later render reads Review of the prompt refresh found records rewritten that the rest of the resume could still refuse or withdraw, and a record of the refresh that a failure could leave out. Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them and just before the first timetravel; the completion of an earlier command's withdrawal runs first, so the refresh sees the run as the resets will find it. Its resets come from the same list the reset loop issues; continuationResets, a copy of that logic, is gone. A timetravel that fails after the refresh is the one window left, and the docs say so. Withdrawn generations. A `--retry-failed` of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them to dynamic-expansion-history/, so the archive no longer held what they ran with. The hook now names the groups the withdrawal plan covers; the refresh treats them as unexpanded, leaves their generation's prompts as they are, and refreshes the template copies the next expansion renders from. Template copies. Every copy was rewritten, even one no render would read again, which also broke replay and fork of runs launched before #1230 for groups that had finished. A copy is now rewritten only when a later render reads it: its group has not expanded or is withdrawn, or an unfinished task that renders from it has no published prompt. Prompts that share a copy must still agree on it, counting every prompt bound to the copy, and one pass settles that because a prompt has one copy. Record. refresh.json is published before the first file changes, so after a failure or a kill every file it lists holds its old or its new bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced instead of claiming replacements that did not happen. Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals. It now compares the effective topology file with plan.json's launch digest, and the topology's node set with the run's, which an eval run's exclusions change. A comment-only topology edit now skips the refresh too. Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order, the derivation sorts them), which skipped every refresh. The refresh uses a new derive-only mode instead. Also: one shared test helper for the toml key; the hook's context type is exported and reused; an unread failure field is dropped; the native resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message says the attach applied none of the project's current prompts. Docs: restart-continue.md starts with pause, wait, edit, resume, and covers the new rules, the reset window, catalog-wide errors, INCOMPLETE, the launch digests and replay and fork of runs launched before #1230. cli.md and artifacts-reports.md point there; the CHANGELOG entry, #1230's edit sentence, security.md and SPECS.md match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 30, 2026
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…e, and only files a later render reads Review of the prompt refresh found records rewritten that the rest of the resume could still refuse or withdraw, and a record of the refresh that a failure could leave out. Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them and just before the first timetravel; the completion of an earlier command's withdrawal runs first, so the refresh sees the run as the resets will find it. Its resets come from the same list the reset loop issues; continuationResets, a copy of that logic, is gone. A timetravel that fails after the refresh is the one window left, and the docs say so. Withdrawn generations. A `--retry-failed` of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them to dynamic-expansion-history/, so the archive no longer held what they ran with. The hook now names the groups the withdrawal plan covers; the refresh treats them as unexpanded, leaves their generation's prompts as they are, and refreshes the template copies the next expansion renders from. Template copies. Every copy was rewritten, even one no render would read again, which also broke replay and fork of runs launched before #1230 for groups that had finished. A copy is now rewritten only when a later render reads it: its group has not expanded or is withdrawn, or an unfinished task that renders from it has no published prompt. Prompts that share a copy must still agree on it, counting every prompt bound to the copy, and one pass settles that because a prompt has one copy. Record. refresh.json is published before the first file changes, so after a failure or a kill every file it lists holds its old or its new bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced instead of claiming replacements that did not happen. Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals. It now compares the effective topology file with plan.json's launch digest, and the topology's node set with the run's, which an eval run's exclusions change. A comment-only topology edit now skips the refresh too. Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order, the derivation sorts them), which skipped every refresh. The refresh uses a new derive-only mode instead. Also: one shared test helper for the toml key; the hook's context type is exported and reused; an unread failure field is dropped; the native resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message says the attach applied none of the project's current prompts. Docs: restart-continue.md starts with pause, wait, edit, resume, and covers the new rules, the reset window, catalog-wide errors, INCOMPLETE, the launch digests and replay and fork of runs launched before #1230. cli.md and artifacts-reports.md point there; the CHANGELOG entry, #1230's edit sentence, security.md and SPECS.md match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…e, and only files a later render reads Review of the prompt refresh found records rewritten that the rest of the resume could still refuse or withdraw, and a record of the refresh that a failure could leave out. Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them and just before the first timetravel; the completion of an earlier command's withdrawal runs first, so the refresh sees the run as the resets will find it. Its resets come from the same list the reset loop issues; continuationResets, a copy of that logic, is gone. A timetravel that fails after the refresh is the one window left, and the docs say so. Withdrawn generations. A `--retry-failed` of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them to dynamic-expansion-history/, so the archive no longer held what they ran with. The hook now names the groups the withdrawal plan covers; the refresh treats them as unexpanded, leaves their generation's prompts as they are, and refreshes the template copies the next expansion renders from. Template copies. Every copy was rewritten, even one no render would read again, which also broke replay and fork of runs launched before #1230 for groups that had finished. A copy is now rewritten only when a later render reads it: its group has not expanded or is withdrawn, or an unfinished task that renders from it has no published prompt. Prompts that share a copy must still agree on it, counting every prompt bound to the copy, and one pass settles that because a prompt has one copy. Record. refresh.json is published before the first file changes, so after a failure or a kill every file it lists holds its old or its new bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced instead of claiming replacements that did not happen. Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals. It now compares the effective topology file with plan.json's launch digest, and the topology's node set with the run's, which an eval run's exclusions change. A comment-only topology edit now skips the refresh too. Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order, the derivation sorts them), which skipped every refresh. The refresh uses a new derive-only mode instead. Also: one shared test helper for the toml key; the hook's context type is exported and reused; an unread failure field is dropped; the native resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message says the attach applied none of the project's current prompts. Docs: restart-continue.md starts with pause, wait, edit, resume, and covers the new rules, the reset window, catalog-wide errors, INCOMPLETE, the launch digests and replay and fork of runs launched before #1230. cli.md and artifacts-reports.md point there; the CHANGELOG entry, #1230's edit sentence, security.md and SPECS.md match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…e, and only files a later render reads Review of the prompt refresh found records rewritten that the rest of the resume could still refuse or withdraw, and a record of the refresh that a failure could leave out. Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them and just before the first timetravel; the completion of an earlier command's withdrawal runs first, so the refresh sees the run as the resets will find it. Its resets come from the same list the reset loop issues; continuationResets, a copy of that logic, is gone. A timetravel that fails after the refresh is the one window left, and the docs say so. Withdrawn generations. A `--retry-failed` of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them to dynamic-expansion-history/, so the archive no longer held what they ran with. The hook now names the groups the withdrawal plan covers; the refresh treats them as unexpanded, leaves their generation's prompts as they are, and refreshes the template copies the next expansion renders from. Template copies. Every copy was rewritten, even one no render would read again, which also broke replay and fork of runs launched before #1230 for groups that had finished. A copy is now rewritten only when a later render reads it: its group has not expanded or is withdrawn, or an unfinished task that renders from it has no published prompt. Prompts that share a copy must still agree on it, counting every prompt bound to the copy, and one pass settles that because a prompt has one copy. Record. refresh.json is published before the first file changes, so after a failure or a kill every file it lists holds its old or its new bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced instead of claiming replacements that did not happen. Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals. It now compares the effective topology file with plan.json's launch digest, and the topology's node set with the run's, which an eval run's exclusions change. A comment-only topology edit now skips the refresh too. Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order, the derivation sorts them), which skipped every refresh. The refresh uses a new derive-only mode instead. Also: one shared test helper for the toml key; the hook's context type is exported and reused; an unread failure field is dropped; the native resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message says the attach applied none of the project's current prompts. Docs: restart-continue.md starts with pause, wait, edit, resume, and covers the new rules, the reset window, catalog-wide errors, INCOMPLETE, the launch digests and replay and fork of runs launched before #1230. cli.md and artifacts-reports.md point there; the CHANGELOG entry, #1230's edit sentence, security.md and SPECS.md match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…asks (#1237) The owner's decision on this follow-up: > We should do this 'enabled by default', so having a toml config to forcefully disable is the best way. It follows the direction that stability matters more than reproducibility and that an operator must be able to restart a run after a minor prompt change. After #1230, a run's prompt files are no longer sealed, but a run still reads `.ultrafuzz/prompts/**` only when it launches. To change what an in-flight task receives, the operator has to hand-edit the run's own files: - There is one `artifacts/<attempt>/prompt.rendered.md` per attempt, with attempt-specific paths inside it. - A prompt that is not rendered yet lives in a template copy under `dynamic-prompt-templates/`, named by the digest of its launch text. - An upgrade that changes a packaged prompt never reaches a run in flight. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…sumed task (#1238) A native `resume` runs the run's launch-rendered workflow against the **installed** `@ultrafuzz/artifacts`, and three task-preparation steps required the installed schema bundle to equal the one the run was launched with. Map 4 of the prompt-sealing investigation measured that one `"$comment"` added to an installed schema file (no rebuild is needed; the registry reads schema files at runtime) fails **every** resumed task, including a text-only one, and that neither `--retry-failed` nor `--refresh-controller --retry-failed` recovers it. The #1230 implementer hit the same two failures. | Gate | Where | Failure | | --- | --- | --- | | C1 | `materializePromptSchemas` refused a workspace copy that differed from the installed bundle (only a refresh could replace it) | `prepare:<attempt> failed at step materialize-prompt-schemas: prompt schema destination differs from checked-in source: …` | | C2 | `assertTaskOutputSchemaBindings` compared this build's binding with the plan | `…step assert-task-output-schema-bindings: artifact-contract failure: planned schema binding changed for <path>` | | C3 | `parseJsonValidatorPreflightSuccessEnvelope` defaulted its expected identity to this build's findings binding, while the run's own validator (the launch closure) reports the launch bundle | `…step preflight-json-validator: … invalid success envelope` (`mismatched identity`) | | C4 | `--refresh-controller` rebound every declared output to the installed bundle (#982) | the refreshed verifier's markers stop matching `graph.json` after a schema change | Behind those gates, the workflow's verifier and dependency admission also validated with the installed schemas, and so did synchronization's findings count, so after a substantive schema change the engine and the host would check artifacts against a schema the run was never planned with; C1–C3 stopped every task before that could be reached. The host artifact gates had already moved to the planned schema in #1188. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #1216. Based on
mainatc1bd7361, which has #1197, #1201, #1183 and #1198; see "Rebase notes" at the end. It does not touchworkflow-sync.tsorpackages/artifacts.Problem
The owner's direction:
On the base, prompt sealing is what stops runs:
Every workflow render, and every lifecycle admission (
status, sync,why,stats, verified reads,replay,fork), re-renders each published runtime prompt and byte-compares it with the file. An operator's edit, or an upgrade that renders a template differently (fix(runtime): stop host artifact gates rejecting valid campaign output #1176, fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195), fails withruntime rendered prompt changed for …. The engine then dispatches nothing, andstatusstops synchronizing.A template copy under
dynamic-prompt-templates/is re-hashed on every render (DYNAMIC_TEMPLATE_CHANGED), with the same effect.A static prompt has three sources, depending on the verb:
replayandforkread the sealed copy in the execution snapshot;resumereads the run's file;--refresh-controllerreads the digest-gated copy inprompt-snapshots/.So an edit is used, silently reverted, or refused, depending on the command. A changed sealed copy even blocks
status,pauseandcancel.The launch engine's attempt reset deletes the run's static prompt files, because the prompt it binds lives in the snapshot. An engine that reads the run file instead, as native resume does, fails every render with
ENOENTwhile one is missing (reset-node can delete prompt required by continuation render #1048), which is why resume restores them first.Change
The rule: an attempt's
artifacts/<attempt>/prompt.rendered.mdis its prompt on every verb (launch,resumewith or without--refresh-controller,--retry-failedor--reset-node,replay,fork). It is rendered once. After launch, nothing hashes it, compares it or seals it. This implements PR 1 of the design; the owner already decided on #1216 to "remove prompt tamper-sealing outright and keep only a record of the prompt each attempt used". The per-attempt record is not in this PR (see "Dropped from this PR").refactor(runtime)!: a run's prompt files are records, not sealsDeleted:
renderReadyRuntimePromptsrenders only a missing file and otherwise adopts it; admission renders nothing);DYNAMIC_TEMPLATE_CHANGED, and its now-unusedtemplatePathinput;controls/rendered-prompts/,controls/prompt-snapshots/), and the tautological plan-row check beside them;sealedTaskPromptPath,promptExecutionSnapshotRoot);--refresh-controllerretained binding and its digest checks (currentControllerPromptBindings,retainedPromptPaths);dynamicRuntimePromptDigestexport.Added:
promptRenderFailures, never thrown, and the task plan and graph are still published.assert-task-inputsthen fails that task alone, either withrendered prompt for <a> could not be rendered: <renderer message>or withrendered prompt for <a> is missing; ultrafuzz resume restores static prompts from prompt-snapshots/. Every other task,statusand sync keep working. Preparation retries, so a template fixed while the engine runs is picked up.resume,replayandforkstart an engine, instead of two resume-only calls. It never overwrites an existing file.fix(runtime): a prompt the render cannot read fails only its taskThe first commit scoped only a missing file (ENOENT, ENOTDIR) to its task. A directory, a symlink or an unreadable file at a prompt path still made
promptForTaskthrow on every render (render frame: EISDIR,ELOOP,EACCES), and after the first commit every engine reads the agent-writable run file, so the task's own agent or a dependent task's agent can cause it.readRegularFileSnapshot(O_NOFOLLOW | O_NONBLOCK, regular-file check). A prompt that cannot be read reads as empty and its cause is kept per attempt;assert-task-inputsfails the task withrendered prompt for <a> could not be read: <cause>. Preparation and the agent's first generation both run that check, so the empty text does not reach a model.pinnedSubmodules: "verify") no longer checks the prompt: the agent already received it, so a later edit, deletion or render failure cannot fail completed work. Preparation and the first-generation reset, which both run before the agent, still check it.lstat), so a dangling symlink there no longer makesresume,replayandforkrefuse to start.fix(runtime): the next engine start completes an interrupted retry withdrawalThe first commit's reorder closed the stale-prompt hole but left an interrupted withdrawal with no way to finish: Smithers has already reset the retried source, so a later
--retry-failedfinds no failed task and plans nothing, and the published manifest keeps the group at its old items for the rest of the run.archiveDynamicExpansionsForRetrynow records the withdrawal indynamic-expansions/.retry-withdrawal.json(source node IDs, history directory, planned moves) before anything moves. The manifest rename takes the record along, so while it is still there the withdrawal is unfinished. Beforeresume,replayorforkstarts an engine,finishInterruptedDynamicExpansionRetryre-validates the manifests as a new plan would and completes the withdrawal into the same history entry. An entry that a render recreated after it had moved is archived beside the first copy.docs: say what editing a run's prompts does to shared copies, retries and recordsrestart-continue.md:--retry-failedof a dynamic source withdraws edits to prompts that wait on the group;--reset-nodeand--retry-failedreplace the engine's attempt row, prompt included, so runultrafuzz statusfirst and keep a copy of the prompt file;--refresh-controllerto get task-scoped prompt failures.security.mdalso names the template copies and launch copies as unchecked, says the catalog digest is checked only when a prompt is rendered, and fixes the governance sentence that contradicted itself.backends-safety.md,prompt-variables.mdanddashboard.mdmatch, andcli.mdandartifacts-reports.mdname the unreadable and non-regular prompt files that now fail only their task.test(runtime): keep one owner of the refresh-controller prompt-edit testdeletes the dynamic-lifecycle copy of a test thatruntime.test.tsowns end to end.Dropped from this PR:
agent.prompt_sha256inattempts.jsonlThe earlier head carried a second commit (
0559e74b,feat(artifacts)!: attempts.jsonl records the prompt each attempt received). It is removed, with its CHANGELOG entry, docs lines and tests, because it broke every in-flight run on upgrade:node-attempt-ledger.schema.jsonis covered by the artifact schema-bundle digest, and the field moved it from6a05a7a6…(v0.1.2 and the base) toe1df8284….materializePromptSchemas, and the JSON-validator preflight's identity check), so every run launched by v0.1.x would fail each task it resumes. That includes the fix(runtime): stop host artifact gates rejecting valid campaign output #1176/fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195 runs the first commit recovers.packages/artifactsis untouched and the bundle digest on this branch is6a05a7a672e0…9598, the v0.1.2 value.It re-lands in its own PR, after the schema "record, don't enforce" follow-up (the design's K1). That PR also needs
resume --reset-node/--retry-failed(andfork --reset-node) to sync the attempt ledger before they reset. The reason is Smithers' reset: it marks the node's attemptsresetCancelled, restarts attempt numbering at 1 (engine.jsnextAttemptNumber), andinsertAttemptupserts on(run_id, node_id, iteration, attempt). So each rerun attempt overwrites the replaced attempt's row,meta.promptincluded.0559e74bis the commit to re-land; this PR's force-push entry links it.The earlier head justified dropping the commit with "Smithers already keeps the full prompt of every attempt in
meta.prompt". That is not true across those resets. Until the follow-up lands, no Ultrafuzz record says which prompt a replaced attempt received. The how-to therefore tells operators to runultrafuzz statusbefore such a reset and to keep a copy of the prompt file.What is kept and why
prompt-snapshots/and the restore. A deleted static prompt has no other source. v0.1.x launch engines have already deleted the static prompt files of the tasks they started: 3 of the 4 tiny-vault smoke runs have none left. Keeping the restore adds no code.plan-run.ts,snapshotPromptTemplate). They run once, in the launch process, right after the bytes are written.--refresh-controllernever recompiles, so they cannot strand a run.template.prompt_sha256andfingerprint. Both sides are compiled launch values, so an edit or an upgrade never trips them. A tampered or transplanted manifest still fails. The docs now say the manifest digest is the launch template's.promptDigestsin the graph fingerprint,rendered_prompt_digestandprompt_digest. They are now documented as launch provenance, and nothing compares them with a prompt file.Deliberately not built
resume --refresh-prompts(the design's PR 2). The owner has not decided on it (the design's Q1; its recommendation is "not now"). Edits to.ultrafuzz/prompts/**still reach only new runs, as today, and the docs now say so..ultrafuzz/authorities/<attempt>.jsontheir prose names is still never written for deferred and generated tasks, as on the base. The fix belongs in its own PR.security.mdstates current behaviour: acknowledgements bind the prompt catalog at launch, and a prompt file edited in a run is not acknowledged again.Verification
Tests that fail on base and pass here. Each was proven by copying this branch's test files into a detached, built worktree of the base and running them there. The only other change to that copy was to put back the
templatePathinput the base'sloadOrCreateDynamicExpansionstill requires (and, for the new withdrawal test, a stub for the missing export), so that only behaviour differs. The first block was proven on533dbfdf, #1197's head before its force-push; the reviewers re-ran it there. The second block covers the review fixes and was proven onb8a2bc5b, #1197's final head (squash-merged asd2b5ccee, the same tree). Neither block was re-proven onc1bd7361: the PRs merged since then do not change prompt sealing.533dbfdf)PromptError: unknown prompt template variable: variable_a_later_build_removed(the base re-renders the published prompt)DynamicExpansionError: Dynamic group fanout prompt template changedDynamicExpansionError: Dynamic group fanout prompt template changedruntime rendered prompt is missing for joinPromptError: unknown prompt template variable: artifact_pth:fanoutWORKFLOW_CONTROL_EVIDENCE_INVALID: published dynamic runtime controls no longer re-derive from their sealed base: runtime rendered prompt changed for strict-joinDynamicExpansionError: Dynamic group fanout prompt template changedWORKFLOW_CONTROL_EVIDENCE_INVALID: … runtime rendered prompt is missing for dynamic-fanout-…prompt-snapshots/<digest>.mdWORKFLOW_LIFECYCLE_FAILED: retained rendered prompt snapshot digest does not match task project-discoveryprompt-snapshots/<digest>.mdWORKFLOW_LIFECYCLE_FAILED: retained rendered prompt snapshot does not match task node:project-discoverycontrols/rendered-prompts/ReferenceError: sealedTaskPromptPath is not defined. A probe of the base template shows the snapshot-loaded engine binding…/controls/rendered-prompts/project-discovery.mdcli-e2e, real pinned Smithers engine, stubcodex): after the engine and then the whole controller are SIGKILLed mid-summarize,artifacts/summarize/prompt.rendered.mdstill exists; a line appended to it beforeresumereaches only the resumed attempt's stdinAssertionError: …/artifacts/summarize/prompt.rendered.md did not survive the interrupted attempts. The base launch engine's attempt reset deleted it as soon assummarizestarted, because that engine binds the sealedcontrols/rendered-prompts/copyb8a2bc5b)assertTaskPromptInputand whosepromptForTaskthrows on any read error. Against the first commit's template,promptForTaskthrowsEISDIR: illegal operation on a directory, readassertTaskPromptInput. Against the first commit's template, the verify pass throwsrendered prompt for summarize is missingArtifactPathError: runtime rendered prompt for dynamic-fanout-… cannot be a symlinkfrom the whole render. The first commit throws the samedynamic-expansions/is[]right after the failed move (the base withdraws the manifests first)[]. With only the completion call removed,dynamic-expansions/still holds[".retry-withdrawal.json", "fanout.json"]after the second resume, so the retried source's new output would be ignoredsmithers graph, missing, directory, symlink and unreadable)render frame: ENOENT: … prompt.rendered.md(#1048's signature). Against the first commit's template,render frame: EISDIRfor the directoryENOENT … artifacts/project-discovery/prompt.rendered.md(nothing restored it). Against the first commit,WORKFLOW_LIFECYCLE_FAILED: published artifact cannot be a symlinkThe owner's case is covered at three levels:
status, sync andresume, and the resumed workflow's first render keeps the edited bytes.94123ada) in 4.5 minutes, as it did before the rebase (0e6025d4, 5 minutes). After the engine and then the whole controller were SIGKILLed mid-summarize,artifacts/summarize/prompt.rendered.mdstill existed. The line appended to it beforeresumereached only the resumed attempt.Conflicts. The rebase met the conflicts that the earlier trial merges against #1201, #1198 and #1183 predicted in
docs/reference/cli.md,docs/how-to/restart-continue.mdandgenerated-workflow-verifier.test.ts, all in the first commit.CHANGELOG.md, which the trial merge with #1198 also flagged, merged cleanly. No production file conflicted. "Rebase notes" says how each was resolved.Gates on the rebased head (
94123ada), each exit code checked directly:pnpm -w format:check;pnpm -w lint;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci. It failed after the plain rebase, and the new commit in "Rebase notes" fixes it;pnpm -w knip, on an unbuilt checkout and afterpnpm -w build;pnpm --filter @ultrafuzz/runtime --filter @ultrafuzz/cli typecheck;node scripts/docs-check.mjs.All exit 0.
Tests run on the rebased tree:
generated-workflow-verifier,dynamic-expansion(23),task-workflow-identity,smithers-attempt-authorityandpinned-submodules, whole files: 153/153. Before the rebase it was 155/155; fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 deletes two tests from the verifier file;dynamic-lifecycle, whole file: 20/20;runtime.test.ts, every test whose name matchesprompt|restore|resume|replay|fork|reset|graph|retry|continuation|refresh|snapshot|seal: 129 tests, 111 pass, 0 fail, and 18 Bun-only tests skipped under Node;campaign-resume, throughpnpm -w validate:release -- --gates cli-e2eunderumask 022: 1/1 in 4.5 minutes.The
dynamic-lifecycleandruntime.test.tsruns used the five rebased commits. TherenderRuntimePromptmove came after them, and the whole-file group and the e2e were then re-run on94123ada. The move only relocates a function declaration inside its module.Risk / compatibility
What is given up, stated plainly:
Ultrafuzz no longer detects prompt tampering. A process running as the operator, including an agent in another task (its artifact directory is writable), can rewrite the prompt or template copy of a later task, and nothing reports it:
status;On the base this was already possible in three ways: through native resume's unchecked run file, through a deferred template before its first render, and through the project workflow file that native resume runs.
security.mdnow says so. It names the prompt files, template copies and launch copies, and lists prompt files you edit as part of the review boundary.replayandforkrender the run's current prompt files, not the launch bytes.One run can mix prompt versions, for example after an upgrade that changes templates. Until the per-attempt digest lands, only Smithers'
meta.promptsays which prompt an attempt ran with, and a--reset-nodeor--retry-failedrerun overwrites it.An edited prompt is not checked for the validator-command lines. That check now runs only on fresh renders. The engine's and the host's output verifiers still reject bad outputs, and the docs say to keep the output-contract block.
Behaviour to know:
Runs launched before this release:
resume. Until then, a static edit is lost, and a runtime edit stops the engine with the old comparison.replayandforkkeep that behaviour.status, sync andresumeuse the installed code, so they are fixed.resumecontinues their launch workflow, which ignorespromptRenderFailuresand reads prompts strictly. A runtime prompt that no longer renders therefore still stops that engine, now reported asrender frame: ENOENT …/prompt.rendered.md; this case also stopped the engine on the base.resume --refresh-controllerrenders the current template and gives such a run the full new behaviour. The docs and the CHANGELOG now say to use it.status, sync andresumeloading the installed runtime, and from this branch leaving the schema bundle at the v0.1.2 digest. A reviewer checked that the 4 v0.1.2 tiny-vault runs'plan.jsonparse with this branch's reader and that all 24 missing static prompts have launch copies. No real v0.1.2 run has been resumed on this branch.Pre-PR runs keep one sealed check: the installed verifier still digest-checks the sealed
controls/rendered-prompts/copies in their execution snapshots. The design accepts this.Changelog
## Unreleased>### Breaking changes, with the entry "Prompts are no longer sealed or compared after launch. …". The review commits extend that one entry: unreadable and non-regular prompt entries, the verify pass, the completed withdrawal, the agent in another task, the--refresh-controllerroute, and what the dashboard's prompt shows. It lists every removed error, what is given up, and the rule for runs launched earlier.attempts.jsonlentry: the commit that needed one is dropped.Rebase notes
This PR was rebased from
b8a2bc5b, #1197's final head, ontomainatc1bd7361, which has #1197 (squash-merged asd2b5ccee, the same tree), #1201, #1183 and #1198. The command wasgit rebase --onto origin/main b8a2bc5b.git range-diffshows every code hunk of the five commits unchanged; beyond the resolutions below, only context lines differ.Conflicts, all in the first commit:
docs/reference/cli.md: feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201 appends two paragraphs at the spot where this PR appends "Prompt files are used as they are…". One covers the patched installed engine andtrusted-bin/smithers; the other covers running long campaigns from a dedicated checkout. All three are kept, feat(runtime)!: run lifecycle commands from the pnpm-patched install instead of per-command npm installs (#921 step 1) #1201's first.docs/how-to/restart-continue.md: feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 appends its paragraph on retrying a node of afailure_policy: continuegroup where this PR's new "Change A Prompt Of A Running Campaign" section begins. Both are kept, and feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198's paragraph stays at the end of its own section, before the new heading.packages/runtime/test/generated-workflow-verifier.test.ts: fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 deletes the two final-report producer-authority tests right above the retry-cleanup test this PR renames. The deletion and the rename are both kept, so the file has 115 tests:main's 112 plus the 3 this PR adds.CHANGELOG.mdmerged without a conflict. This PR's entry still replaces the fix(runtime): stop host artifact gates rejecting valid campaign output #1176/fix: symlinked project roots, locale-independent ordering, clock-skew-tolerant audit journals and other small verified bugs #1195 upgrade note in place, now below the new feat(topology)!: a failed property lens no longer skips the rest of the campaign #1198 and fix(runtime): record final-report producer selections in the run instead of querying smithers #1183 entries.New commit:
refactor(runtime): define renderRuntimePrompt after the prompt-context helpers.dynamic-runtime.tsis already over the strict lint's 500-line budget onmain(668 counted lines). ESLint reportsmax-linesonly on the 501st counted line, andlint:strict:cikeeps that report only when its line is one the branch changed. Before the rebase, that line was inresolvedConfigForRuntimeRoot, 11 lines past this PR's hunk. #1198 adds lines above it, which moved the line intorenderRuntimePrompt, a function this PR extracts, so the gate failed. The commit movesrenderRuntimePrompt, byte for byte, belowresolvedConfigForRuntimeRootandpromptGraphContext. The report now falls insidepromptGraphContext, which this PR does not change. #1062 fixed the same failure this way; noeslint-disableis added.Interplay checked beyond the conflicts:
resume,replayandforkstill reachrunSmithersLifecycleCommandwithrelaunchPathsafter they writetrusted-bin/smithers. The static-prompt restore and the interrupted-withdrawal completion therefore still run before any engine starts.readTaskPromptFileis still the only one. No code onmainreferences the identifiers this PR deletes:sealedTaskPromptPath,promptExecutionSnapshotRoot,currentControllerPromptBindings,retainedPromptPaths,dynamicRuntimePromptDigestandDYNAMIC_TEMPLATE_CHANGED.dynamic-runtime.tschanges do not touch prompt rendering. They cover optional inputs for generated children andreconcilesPartialResultstaking the producer's group.Validation of the rebased head is in "Verification" above: the gates, the test counts and the e2e.
🤖 Generated with Claude Code
The PR appears safe to merge on the reviewed changes; no outstanding blocking finding remains.
Summary
The PR makes a run’s rendered prompt files editable records rather than sealed execution inputs. It also scopes prompt render and read failures to their tasks, restores missing static prompts before engine starts, and adds recovery for interrupted dynamic-source retry withdrawals.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Record withdrawal in history] --> B[Archive attempt state] B --> C[Rename manifests] C --> D[Rebuild graph and task plan] D --> E[Prune withdrawn state] E --> F[Remove withdrawal record] B -. interrupted .-> R[Next engine start completes withdrawal] C -. interrupted .-> R D -. interrupted .-> R R --> FReviews (3) · Last reviewed commit: "fix(runtime): complete a retry withdrawa..."