Conversation
Comment on lines
+370
to
+372
| tryPropose(state, [target.prompt], attemptId, () => { | ||
| if ("error" in render) throw new Error(render.error); | ||
| return { relativePath: promptFilePath(attemptId), prompt: target.prompt, attemptId, next: render.markdown }; |
There was a problem hiding this comment.
Runtime prompts bypass authority checks
When an edited prompt for an unfinished generated child or deferred task adds a valid ancestor-artifact authority selector, this path publishes the rendered prompt without checking that selector against the task’s compiled selectors. The prompt can instruct the agent to use an authority absent from its sealed authority document, instead of being rejected as the static refresh path rejects it. Return the rendered artifact references from the runtime renderer and check them against the compiled selectors before publishing.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/prompt-refresh.ts
Line: 370-372
Comment:
**Runtime prompts bypass authority checks**
When an edited prompt for an unfinished generated child or deferred task adds a valid ancestor-artifact authority selector, this path publishes the rendered prompt without checking that selector against the task’s compiled selectors. The prompt can instruct the agent to use an authority absent from its sealed authority document, instead of being rejected as the static refresh path rejects it. Return the rendered artifact references from the runtime renderer and check them against the compiled selectors before publishing.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.…ritable umask Problem: under a umask that leaves new directories group writable, such as Ubuntu's default 0002, ensureSafeDirectory created <run>/safe-bin 0775. composeSmithersCommandPath admits a target-local PATH entry only through isPreparedForgeGuardBin, which refuses a group-writable directory, so the workflow engine's PATH dropped the wrapper and agents ran the real forge without forge_vmem_limit_kb or forge_rayon_threads while run.json recorded forge_guard.active: true. A custom run.output_dir produces the same false claim under any umask, because that check admits the wrapper only from <project>/.ultrafuzz/runs/<run-id>/safe-bin. Change: prepareForgeGuardEnvironment sets safe-bin to 0700 after creating it (chmod does not apply the umask, and it also repairs the directory of a run launched earlier), then applies the engine's own admission check to it. It reports the guard active only when that check passes. Otherwise it leaves PATH and the guard variables unset and returns a FORGE_GUARD_INACTIVE warning, which launch, resume, replay and fork return with their result. The admission check itself is unchanged. resume now also records the guard of the controller it starts in run.json, best effort like its state projection, so run.json no longer keeps a launch's claim after a resume that dropped the wrapper. The OpenRouter adapter contract test failed on such hosts for a different reason: its fixture rooted the operator-owned provider homes under the target's .ultrafuzz, which `ultrafuzz init` creates with the process umask. The adapter creates every provider-home component itself with mode 0700, so the fixture now uses its own private root; the provider-home ancestor check is unchanged. The new tests pin umask 0002 through a shared helper, so they reproduce on CI runners, whose umask is 022. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…loud run Problem: before #1197 removed per-node cloud execution, `ultrafuzz clean` of a run planned with `[execution] mode = "cloud"` terminated its sandboxes (tagged purpose=ultrafuzz-node) and deleted its Modal volume before deleting the run directory. #1197 dropped that step and kept the deletion, so cleaning such a run now leaves its Modal volume and any running sandbox in place, still billed, and deletes the plan.json that records the Modal app they live in, without saying so. Change: before anything is removed, clean reads the mode and Modal app from each selected run's plan.json, which copies the execution block of the run's resolved config and is what the removed cleanup read. For a cloud run it returns a CLEAN_CLOUD_STORAGE_RETAINED warning naming the app, the volume and the sandboxes' run tag, and the `modal volume delete` command. `--dry-run` reports it without deleting, and every later result, including a failed removal, carries it. The deletion itself is unchanged. The volume is not `ultrafuzz-node-<run-id>`: the removed provider named it `ultrafuzz-node-` plus a bounded identity of the Smithers run ID `ultrafuzz-<run-id>` (at most 32 normalized characters, then 12 hex digits of its SHA-256). The same identity was the sandboxes' `run` tag. clean reproduces that function, and the test pins the name the removed code computes for a fixed run ID. An unreadable plan.json yields no warning. It also makes the source revision step fail before anything is removed, as it did before, so no run that clean deletes skips the check. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: #1197 made compilation write every task's execution block as `{ mode: "local", resources, agentCredentialEnv: [] }` with no `modal` entry, but synchronization still added task.execution.agentCredentialEnv and task.execution.modal?.credentialEnv to its exact-value redaction list. For every run compiled since #1197 both are empty. The generated agent environment also still blanked ULTRAFUZZ_MODAL_MODULE, which only the removed cloud provider set. And the doctor tests still wrote a fake engine into the target's .smithers, although since #1201 doctor inspects only the runner Ultrafuzz's own install provides. Change: synchronization derives its redaction values from the environment alone, the agent environment no longer lists ULTRAFUZZ_MODAL_MODULE, and seven doctor tests lose the dead fake-engine setup. One test used it only to create the .smithers/node_modules/.bin directory it writes into, and now creates that directory itself. The doctor test that installs a 0.29.0 project-local engine keeps it: it asserts that doctor ignores that engine. The task manifest keeps agentCredentialEnv, because its sealed schema requires the field. Effect on runs compiled before #1197 in cloud mode, whose lists were not empty: sensitiveEnvironmentValues still redacts, by name, every variable that looks like a credential (*_API_KEY, *_TOKEN, *_SECRET, ...), which includes MODAL_TOKEN_ID and MODAL_TOKEN_SECRET. What these lists added beyond that were route variables such as *_BASE_URL and allowlisted variables. Those carry no secret, or have a value the secret patterns already match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: #1201 wrote <run>/trusted-bin/smithers at launch, resume, replay and fork so that the generated workflow's bare `smithers node` call (#1143) found a runner. #1183 removed that call: final-report producer retries and verifiers now read the selections each producer attempt records in the run. The current template spawns only git, bash and the run's `ultrafuzz` launcher, and every command the runtime spawns runs an executable it bound by path, never `smithers` from PATH. smithers-report-retry.integration.test.ts puts a failing `smithers` first on the engine PATH around the production report code and still passes. The shim therefore served only workflows persisted by earlier releases, and it put an engine CLI first on every task's PATH. It also made replay and fork bind the installed runner, which they otherwise never run: they run the run's sealed engine. Change: delete writeTrustedSmithersShim and its three callers. When resume cannot re-verify the trusted CLI, it again keeps an existing run launcher first on PATH itself, which the shim had done as a side effect; the existing test for that path covers it. Launch still binds the installed runner before creating a run, because resume runs it; the comments and the doctor summaries now say launch and resume. The capability and native-continuation tests drop their shim assertions, and a resume test now asserts that trusted-bin holds only the ultrafuzz launcher. The docs no longer say that the shim serves the workflow's smithers calls or that tasks can drive their run through it. The CHANGELOG entry for #1201, which has not been released, drops its shim claims, and a new entry records the removal. BREAKING CHANGE: tasks no longer find a `smithers` CLI that Ultrafuzz provides on PATH. A run launched by an earlier release and continued by plain `resume` still runs its persisted workflow, whose bare `smithers` call in a restarted controller finds a runner only on the operator's own PATH, as before #1201; continue such a run with `resume --refresh-controller`. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/prompts-refresh-on-resume
branch
from
September 30, 2026 21:48
a7a86f3 to
29f03e4
Compare
…e run Problem: for a run an earlier release planned for per-node Modal execution, the CLEAN_CLOUD_STORAGE_RETAINED warning names the Modal volume, app and sandbox tag that `clean` leaves behind, and the run's plan.json, which the removal deletes, is the only record of the app. cleanRun computed the warning before removing anything but returned it only with its result, and the CLI prints a successful result's diagnostics after "Removed: ...". The operator therefore saw it only once the plan was gone, and never when clean did not return: the audit append after the removal had no try/catch, so a journal lock timeout, a full disk or a read-only .ultrafuzz rejected cleanRun after the run was deleted and dropped the warning with it. The CHANGELOG and the how-to also gave the volume as ultrafuzz-node-ultrafuzz-<run-id>-<hash>, which is wrong for every generated run ID. The removed provider kept at most 32 characters of ultrafuzz-<run-id>, and a generated ID such as run-20260930t161149123z-ab12cd34 makes that 42, so its volume is ultrafuzz-node-ultrafuzz-run-20260930t161149123-d99772718708. The runtime message was right; the test pinned only a short run ID. Change: - CleanGeneratedInput takes an optional onRetainedStorage callback. cleanRun calls it with the warnings just before the first deletion, after the source-ref step that can still stop the removal; a dry run does not call it. Every result still carries the warnings. - The CLI's text output prints them from that callback and leaves them out of its final text, so they appear once, before anything is removed. JSON output and the dashboard keep reading them from the result. - A failed audit append is now a CLEAN_AUDIT_FAILED failure that says whether anything was removed and still carries the warnings. - The docs describe the volume name as derived, say to copy it from the warning or find it with `modal volume list`, and the runtime test pins the truncated name for a generated run ID. On the branch before this commit, the new tests fail as intended: the callback is never called, the audit test rejects with EACCES on .ultrafuzz/clean-audit.jsonl.lock, and the CLI test sees the warning first written after plan.json was deleted. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: four ways the Forge guard follow-up could still leave run.json or the operator with a wrong picture. - A synchronization pass (status --watch, the dashboard, the eval runner) reads run.json, then awaits usage replay and possibly a pricing fetch, and wrote back that copy plus its accounting. It holds only .workflow-sync.lock, and resume records its controller's guard holding only the lifecycle lock, so a pass that read before a resume's write and wrote after it put the launch's forge_guard back, and nothing rewrote it until the next lifecycle command. A test whose pricing fetch rewrites the guard mid-pass reproduces this on the branch: run.json ends with the launch's active: true. replay and fork, which record the guard the same way, had this on main. - The engine PATH admits safe-bin only while it holds the wrapper alone. The .forge.tmp-* file an interrupted writeFileDurable leaves, or any other entry, therefore dropped the guard for every later launch, resume, replay and fork of the run, and nothing removed it. - writeFileDurable created the wrapper 0600 and chmod made it 0700 after the rename. resume and replay rewrite the wrapper while tasks may be running, and a PATH lookup in that window skips a file without its execute bit and runs the real Forge. - A resume that only attached to a running controller still returned FORGE_GUARD_INACTIVE about the controller it did not start. The warning, config.md and configuration.md also left out the admission's no-symbolic-link condition, so a project reached through a symlinked parent directory was told the engine admits only the directory the wrapper was in. Change: - @ultrafuzz/artifacts adds updateRunMetadataDocument, which re-reads run.json and writes the update while holding run.json.lock, taken with the existing journal lock. Recording the Forge guard and synchronization's accounting write both use it. The accounting now goes into the document as it is at the write, which must still be bound to the workflow it was computed for; otherwise the pass reports WORKFLOW_ACCOUNTING_FAILED, as it already did for that mismatch at its start, and writes nothing. - prepareForgeGuardEnvironment removes every entry but forge from safe-bin before writing the wrapper, best effort: an entry it cannot remove still fails the admission, which it reports. It writes the wrapper with mode 0700, so the rename replaces an executable file with an executable file. - resume returns the Forge guard diagnostics only when it starts a controller, the same condition under which it records the guard. - The warning and the docs state the symbolic-link condition. The unit test's stray-file case, which expected the guard inactive, now expects it cleared and active; the inactive cases are a custom output_dir and a symlinked parent. The resume test makes the guard inactive with an entry resume cannot remove, and is skipped under root. On the branch before this commit, the new tests fail as intended, including the attach case, which returned the warning. Not covered: replay and fork rebind run.json's workflow under the control lock, not run.json.lock. The accounting write now refuses a document bound to another workflow instead of writing the old binding back, which narrows that race on main to the lock-held re-read and write, but does not close it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…shim Problem: a Breaking-changes entry announced that launch, resume, replay and fork no longer write <run>/trusted-bin/smithers. No release wrote it: #1201 added it after v0.1.2, and this branch already rewrote #1201's entry without it. Measured against v0.1.2, nothing the entry lists changes: replay and fork never refused an unpatched install, and a run launched by an earlier release already found its persisted workflow's bare `smithers` only on the operator's PATH. Change: remove the entry. Its one piece of advice moves to the #1183 entry, which already explains that plain resume keeps running the persisted workflow: such a run still runs `smithers node` in a restarted controller, so continue it with `resume --refresh-controller`. The #1183 entry already gives the --reset-node form for a report producer that has succeeded. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…safe-bin Launch holds the workflow control lock and resume, replay and fork the lifecycle lock, so a launch and a resume can prepare the same run's Forge guard at once. The wrapper was written through a temporary file inside safe-bin, so the stray-entry cleanup of one command could delete the other's temporary file before its rename, failing that command with ENOENT. An engine PATH composed meanwhile also saw two entries in safe-bin and dropped the wrapper, although run.json recorded the guard active. writeFileDurable now takes an optional temporaryDirectory, and the wrapper's temporary file goes in the run root, so safe-bin only ever holds the wrapper. The cleanup still removes temporary files that interrupted writes left in safe-bin before this change. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/prompts-refresh-on-resume
branch
from
September 30, 2026 22:34
29f03e4 to
4bd75ad
Compare
…ugh testWhen
The test "resume records the Forge guard of the controller it starts"
passed { skip: ... } to test() to stay out of root runs. Bun's node:test
shim ignores that option, so the monolithic suite bans it (#709), and "the
monolithic runtime suite registers exclusively through the shard wrapper"
failed on the stray skip:.
Register the test through testWhen(process.getuid?.() !== 0), which
selects test.skip for root in both the Node and Bun lanes, as the lock
synchronization test does. A comment keeps the reason the skip option
carried.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…asks Before it resets, archives or submits anything, every `ultrafuzz resume`, with or without --retry-failed, --reset-node or --refresh-controller, now applies the project's current prompts (.ultrafuzz/prompts/** and the packaged built-ins, selected exactly as `run` selects them) to every task of the run that has not finished. It re-renders, with planning's own renderer and the run's frozen config, the prompt of every static task whose agent the engine does not report `finished`, plus the tasks a reset reopens and their dependents; the published prompts of unfinished generated children and of nodes that waited on a dynamic group; and the template copies that prompts not rendered yet are rendered from. Finished tasks keep their prompt files as records. The refresh never fails a resume. A topology that no longer expands to the run's graph, an unreadable ultrafuzz.toml or a run the engine still reports active skips it with PROMPT_REFRESH_SKIPPED. Each prompt is validated as `run` validates it and applied to all of its tasks or to none: one that does not validate or render, that names an artifact authority a task was not compiled with, or that shares a template copy with a prompt whose new text differs keeps its tasks' files and is reported as PROMPT_REFRESH_REJECTED. Every replaced file is copied first to prompt-history/<time>-<uuid>/, next to a refresh.json listing each rewritten file with its old and new SHA-256, and each file is replaced atomically; resume reports PROMPTS_REFRESHED. The record is refresh.json rather than an events.jsonl event: a new event type changes event-record.schema.json, which the artifact schema-bundle digest covers, and the engine still requires a run's launch bundle, so every in-flight run would fail each task it resumes. The bundle digest is unchanged. The new [run] refresh_prompts_on_resume key (default true) turns it off. Resume reads it from the project's current ultrafuzz.toml; applyRunConfig drops it, so it never enters a run's resolved config, whose strict schema older runs' sealed configs must keep matching. Supporting changes: planning's static renderer can render in memory without creating directories and collect per-attempt failures (renderRunStaticPrompts); the runtime renderer takes a template body (renderRuntimePromptsFromTemplates); effectiveTopologyPath selects the run's topology file without effectiveAuditPolicy's prompt-validating load; the lifecycle command gains a beforeContinuation hook with the inspected node states and the resets it will issue. BREAKING CHANGE: a hand edit of an unfinished task's prompt.rendered.md or of a template copy is replaced by the next resume (the old bytes stay in prompt-history/) unless run.refresh_prompts_on_resume = false, and edits to .ultrafuzz/prompts/** now reach runs already in flight. Co-Authored-By: Claude Opus 5.5 <[email protected]>
One file that does not parse, such as invalid frontmatter, fails the
whole prompt catalog, and the error did not say which file it was: with
about 40 scaffolded prompts, `run`, `validate` and the new prompt
refresh on resume reported only "invalid prompt YAML frontmatter: ...".
Each per-file parse or variable error now starts with the prompt's
source and relative path ("project prompt setup/variant.md: ..."), and
a duplicate project prompt id names both files instead of the second.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e, and only files a later render reads Review of the prompt refresh found records rewritten that the rest of the resume could still refuse or withdraw, and a record of the refresh that a failure could leave out. Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them and just before the first timetravel; the completion of an earlier command's withdrawal runs first, so the refresh sees the run as the resets will find it. Its resets come from the same list the reset loop issues; continuationResets, a copy of that logic, is gone. A timetravel that fails after the refresh is the one window left, and the docs say so. Withdrawn generations. A `--retry-failed` of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them to dynamic-expansion-history/, so the archive no longer held what they ran with. The hook now names the groups the withdrawal plan covers; the refresh treats them as unexpanded, leaves their generation's prompts as they are, and refreshes the template copies the next expansion renders from. Template copies. Every copy was rewritten, even one no render would read again, which also broke replay and fork of runs launched before #1230 for groups that had finished. A copy is now rewritten only when a later render reads it: its group has not expanded or is withdrawn, or an unfinished task that renders from it has no published prompt. Prompts that share a copy must still agree on it, counting every prompt bound to the copy, and one pass settles that because a prompt has one copy. Record. refresh.json is published before the first file changes, so after a failure or a kill every file it lists holds its old or its new bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced instead of claiming replacements that did not happen. Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals. It now compares the effective topology file with plan.json's launch digest, and the topology's node set with the run's, which an eval run's exclusions change. A comment-only topology edit now skips the refresh too. Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order, the derivation sorts them), which skipped every refresh. The refresh uses a new derive-only mode instead. Also: one shared test helper for the toml key; the hook's context type is exported and reused; an unread failure field is dropped; the native resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message says the attach applied none of the project's current prompts. Docs: restart-continue.md starts with pause, wait, edit, resume, and covers the new rules, the reset window, catalog-wide errors, INCOMPLETE, the launch digests and replay and fork of runs launched before #1230. cli.md and artifacts-reports.md point there; the CHANGELOG entry, #1230's edit sentence, security.md and SPECS.md match. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…orities refreshes on resume Greptile flagged that the refresh publishes runtime prompts without the compiled-selector check that static prompts get. Declined: every task whose prompt is rendered at runtime (generated children and tasks that wait on a dynamic group) is compiled with no artifact authority selectors, on main and since v0.1.0 (#1234), and the engine's own render names whatever its template names. The stock review prompts (dedupe-findings, aggregate-test-files and final-report) already name authorities at launch, so that check would reject every edit of them once their groups expand. For the flagged case, the refreshed prompt is byte for byte what a launch of the edited prompt renders. The new test publishes the stock review prompts the way the engine's first render after goal-plan does. It checks that an unedited resume reports nothing and that an edit of review/final-report.md is applied. It fails under the proposed check in both forms: run on every render, the unedited resume gets three PROMPT_REFRESH_REJECTED warnings; run only on changed bytes, the edit is rejected. None of the existing refresh tests fail under either form. The docs and the doc comments now say the authority check covers static tasks, which is what it does, and why runtime prompts are not checked. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…reference cache "resume of a stock run renders its prompts as launch did and refreshes an unexpanded group's template copy" and "resume applies an edited stock review prompt that names artifact authorities once its groups expanded" launch the stock project, whose plan verifies and materializes every reference the shipped catalog pins. Both read the host's reference cache, so they passed only where ~/.cache/ultrafuzz/references was synced and failed in CI with MISSING_CACHE for vulnerability-database.owasp-scs. A withShippedReferenceCache helper seeds an XDG cache in the project with writeShippedDocumentReferenceCaches and writeShippedVulnerabilityDatabaseCache, as the clean-scaffold plan tests do, and points XDG_CACHE_HOME at it for the launch. Launch is the only step that reads the cache: with XDG_CACHE_HOME at an empty directory both tests failed with MISSING_CACHE before this change and pass after it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/prompts-refresh-on-resume
branch
from
September 30, 2026 22:57
4bd75ad to
16bc398
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1235 (
claude/post-train-followups,66dc5434) for the merge train #1235 → #1237 → #1238: merge #1235 first; until then this PR also lists its eight commits, and this PR's own are the top four. #1230, named below, has merged (e45076a4).Stacked on #1230 (
claude/prompt-records-not-seals, "a run's prompt files are records, not seals"), which sits onmainafter the train (#1197, #1201, #1183, #1198) merged. This branch is rebased onto #1230's current head,9b42ae0b, and merges after it.Problem
The owner's decision on this follow-up:
It follows the direction that stability matters more than reproducibility and that an operator must be able to restart a run after a minor prompt change.
After #1230, a run's prompt files are no longer sealed, but a run still reads
.ultrafuzz/prompts/**only when it launches. To change what an in-flight task receives, the operator has to hand-edit the run's own files:artifacts/<attempt>/prompt.rendered.mdper attempt, with attempt-specific paths inside it.dynamic-prompt-templates/, named by the digest of its launch text.Change
Four commits. The fourth,
test(runtime): an edited stock review prompt that names artifact authorities refreshes on resume, is described under "Greptile follow-up" below.feat(runtime)!: resume applies edited project prompts to unfinished tasksEvery
ultrafuzz resume <run>(plain,--retry-failed,--reset-nodeor--refresh-controller) now runs a prompt refresh. It runs after the active-run attach, the completion of an interrupted retry withdrawal, the retry-archive plan and the unsupported-run-ID refusal, each of which can end the resume before it resets or archives anything, and before the first reset, the retry archive and the submission.What it renders.
rundoes, from.ultrafuzz/prompts/**and the packaged built-ins.smithers/expanded-graph.jsonand its frozensmithers/resolved-config.json;Which tasks it covers. A task counts as unfinished unless Smithers reports its agent node (
node:<attempt>)finishedand this resume does not reset it. For resets:--no-deps), so is every task that depends on it.The refresh then rewrites:
prompt.rendered.mdof every unfinished static task;dynamic-prompt-templates/that a later render reads: that of a group that has not expanded, and that of an unfinished task whose prompt is not published yet.A finished task keeps its file, which is the record of the prompt it ran with. When
--retry-failedreruns a dynamic source, the groups the withdrawal plan covers count as unexpanded: their generated children and the prompts that wait on them keep their files, which the withdrawal moves todynamic-expansion-history/as the record of what ran, and their template copies take the edits. A file is rewritten only when its bytes change: an unchanged project renders byte-identical prompts, and the refresh writes nothing (tested, also on the stock default topology).It never fails the resume.
It skips entirely with
PROMPT_REFRESH_SKIPPED(warning), and changes nothing, when:plan.jsonrecorded at launch (any edit, even to a comment), or the topology's nodes are not the run's (an eval run launched with exclusions);run, such as invalid frontmatter or a duplicateid; the warning names the file;ultrafuzz.tomlcannot be read;--forceor--reset-nodeon an active run).Each prompt is applied to every file rendered from it, or to none. It is validated on its own, the way
runvalidates it:It is then rendered for each of its unfinished tasks, which also runs the validator-command check.
A prompt that fails any of these keeps every file rendered from it, and resume reports
PROMPT_REFRESH_REJECTED(warning), naming the prompt and the error. So does a prompt that meets either of these:The other prompts still apply. A rejected prompt that no unfinished task or template copy uses changes nothing, so it is not reported.
Before it changes any file, the refresh publishes
prompt-history/<time>-<uuid>/refresh.json, which lists each file it will rewrite with its attempt, its prompt, and its old and new SHA-256. It then copies each file into that entry, at its path in the run, and replaces it atomically withwriteFileDurable. Resume reportsPROMPTS_REFRESHED(info), which names the tasks and the entry.If a step fails part way, resume reports
PROMPT_REFRESH_INCOMPLETEwith how many of the planned files it replaced. Each filerefresh.jsonlists holds either its old or its new bytes, and a resume after the cause is fixed applies the rest.Turning it off. The new
[run] refresh_prompts_on_resume = falseturns the refresh off. Resume reads the key from the project's currentultrafuzz.tomlevery time, so it also applies to runs already in flight. The config loader accepts the key, andapplyRunConfigdrops it, so it never enters a run's resolved config. That matters: the resolved-config schema is strict and requires everyrunfield, so adding the key there would make every older run's sealed config fail to parse on resume.Where the code is.
prompt-refresh.ts(new) plans the refresh.prompt-history.ts(new) records and applies it.plan-run.ts: the plan-time static renderer is split so thatrenderRunStaticPromptsrenders in memory, creates no artifact directories, and collects per-attempt failures. Launch'srenderPromptsForPlannow writes that renderer's results, so both use the same render code.dynamic-runtime.ts:renderRuntimePrompttakes an optional template body;renderRuntimePromptsFromTemplates, at the end of the file, renders listed tasks in memory;deriveDynamicRuntimeMaterializationderives the runtime without publishing or comparing anything.audit-profile-policy.ts:effectiveTopologyPathselects the run's topology file aseffectiveAuditPolicydoes. It skips that function's topology load, which validates the project's prompt files and would let one bad prompt skip the whole refresh.smithers.ts: abeforeContinuationhook receives aSmithersContinuationContext: the inspected node states, the resets the command will issue (from the same list the reset loop iterates), the groups the retry withdrawal covers, and whether the run is active.start-run.tspasses the hook.dynamic-expansion-retry.tsexportsreadDynamicRuntimeBase.Docs.
restart-continue.md: "Change A Prompt Of A Running Campaign" is rewritten around "pause, wait forpaused, edit the project prompt, resume". It says what the refresh covers and reports, when it skips, what a failed reset leaves, and how to turn it off to edit a run's own files. The refactor(runtime)!: a run's prompt files are records, not seals #1230 hand-edit guidance moves under "Edit A Run's Own Prompt Files", and the rule for runs launched earlier gets its own subsection.configuration.mddocuments the key and the one exception to "configuration is frozen at launch".SPECS.md,security.md,edit-prompts-topology.md,prompt-variables.mdanddashboard.mdare updated to match;cli.mdandartifacts-reports.mdsay it in a sentence and link to the how-to.Size. About +1180 / −200 production lines across the three commits, including doc comments, and +775 test lines. The review fixes account for +336 / −310 of the production lines: they add the derive-only mode, the withdrawal and live-copy rules and the record-first apply, and delete the topology re-expansion,
continuationResetsand the fixed-point loop.fix(prompts): a prompt catalog error names the file that caused itOne file that does not parse fails the whole catalog, for
run,validateand the refresh, and the error did not say which file it was. Each per-file parse or variable error now starts with the prompt's source and relative path (project prompt setup/variant.md: invalid prompt YAML frontmatter: ...), and a duplicate project promptidnames both files.fix(runtime): refresh prompts only after a resume can no longer refuse, and only files a later render readsThe review fixes; the first commit's description above already describes the result. What changed, and why:
timetravel.continuationResets, which repeated the reset loop's logic, is gone: the hook and the reset loop read one list.--retry-failedof a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them, sodynamic-expansion-history/no longer held what they ran with.refresh.jsonwas written after the last file, so a failure or a kill left no record, andPROMPT_REFRESH_INCOMPLETEclaimed replacements that had not happened.references.ymlchange, would have skipped every in-flight run's refresh. It now compares the file withaudit_profile.topology_digestand the node sets.dynamic_groupsin compile order and the derivation sorts them), which skipped every refresh. The refresh now derives without comparing.Review findings and what was done
--retry-failedchanges no prompt file; a successful one archives the children's and the join's launch bytes. The one window left, atimetravelthat fails after the refresh, is documented.id); the skip is documented. Not isolated per prompt:runselects prompts byidacross every file, so a file that does not parse could carry any prompt'sid, andrunrefuses the same catalog.refresh.jsonwritten last; misleadingPROMPT_REFRESH_INCOMPLETEartifacts-reports.mdandsecurity.mdgapscontinuationResetsduplicates the reset loop--retry-failedtest added.settleTemplateCopieswithout a testGreptile follow-up
prompt-refresh.ts:372): declined. Instead,a7a86f3dadds a guard test and makes the docs precise.mainand has been since v0.1.0: plan time renders only static tasks, and only those rows seal selectors (Runtime-rendered prompts may name prompt-artifact authority files that are never written #1234). On a stock launch,dedupe-findings,aggregate-test-files,final-reportand both goal-group templates seal none, while their launch prompts already name 3, 1 and 2 authorities.resume applies an edited stock review prompt that names artifact authorities once its groups expanded, fails under both forms of the rule. The 16 existing refresh tests pass under both, so nothing guarded this before.a7a86f3d. The 12 refresh tests inruntime.test.tspass.pnpm -w format:check,pnpm -w lint,lint:strict:ciagainstorigin/main,pnpm -w knip, the runtime typecheck andnode scripts/docs-check.mjsall exit 0.Deliberately not built
An
events.jsonlevent. The task asked for one event that lists the tasks whose prompts changed, but that event cannot be added safely:event-record.schema.json;The record is
prompt-history/<entry>/refresh.jsontogether with thePROMPTS_REFRESHEDdiagnostic. The bundle digest is unchanged, andpackages/artifactsis untouched.A CLI flag. The owner chose a TOML key only.
A refresh on
replayandfork. They keep rendering the run's current files, as in refactor(runtime)!: a run's prompt files are records, not seals #1230.Exact prediction of Smithers' reset set. Smithers resets by attempt start time:
timetravelalso resets every node whose attempt started after the target's (resolveResetNodes), sibling branches included. The refresh runs before the reset, andinspectcarries no start times, so it counts only the reset task and the tasks that depend on it. A finished task that merely started later reruns with its old prompt. The docs say so.Applying after the reset. The design's §9 sketch applied after the resets, which would also close the failed-
timetravelwindow. The brief puts the refresh before any timetravel, and a split plan and apply adds a second hook for a window that ends in a failed resume, whose filesrefresh.jsonnames.Per-prompt isolation of a broken catalog (see the table).
Re-acknowledging data disclosure after a refresh (the design's Q2). Acknowledgements stay launch-scoped, and
security.mdsays so.Refusing
--refresh-controllertogether with the refresh (the design's §9 and T10). The refresh renders with the run's planned schema bindings, exactly as launch did. The refreshed controller's rebinding to the installed schemas is the same skew as before (the K2 follow-up).Treating runs launched before refactor(runtime)!: a run's prompt files are records, not seals #1230 specially. On a plain resume, their launch workflow still stops the whole run when a prompt no longer renders, and their
replayandforkrun the launch snapshot, which still compares runtime prompts and template copies. The docs and the CHANGELOG say to resume them with--refresh-controller, and to turn the refresh off before resuming such a run that must still be replayed or forked.Updating launch provenance.
prompt_digest, the manifesttemplate.prompt_sha256andrendered_prompt_digeststay what launch recorded; the CHANGELOG andartifacts-reports.mdsay so.refresh.jsonis the record of what changed.Verification
Discriminating tests. Each was proven by copying this branch's
runtime.test.ts,dynamic-lifecycle.test.ts,prompt-refresh-config.tsandscaffold-catalog.test.tsinto a detached, freshly built worktree and running them there, with nothing else changed: on #1230's current head9b42ae0b, and onb32874f4, this branch's head before the review fixes (on #1230's previous head94123ada), which shows what the review fixes change. The first round proved the same tests, andconfig.test.ts, on0e6025d4and94123ada.9b42ae0bb32874f4PROMPTS_REFRESHEDrun.refresh_prompts_on_resume = falsekeeps the run's prompts until it is turned back on--reset-nodeapplies edited prompts to the finished task it reruns and to its dependents--retry-failedapplies edited prompts to the finished producer of a failed verifier and to its dependentsPROMPT_REFRESH_SKIPPED, the tamper checkrefresh.jsoncheckcatalog.tsis the base's)native resume delegates the persisted workflow after mutable project sources are replacednow turns the refresh off and asserts no prompt diagnostics; it passes everywhere, as a pinned test should.Assertions inside them cover each required case:
--retry-failedof a dynamic source leaves the withdrawn children's and the join's launch bytes indynamic-expansion-history/, and a refused retry changes no prompt file.prompt-history/is written.syncRunandgetRunHealthstay clean.refresh.jsonlists every rewritten file, the archived copies hold the old bytes, and a failed apply has already published it.Real engine.
packages/cli/test/e2e/campaign-resume.test.ts(CI lanecli-e2e) edits the project prompt instead of the run file; the first round ran it (1/1, about 10 minutes) on this branch and saw it fail on94123ada. It was not re-run after the review fixes: they change only the order of steps it does not reach (no retry, no reset) and which template copies are rewritten, and its topology has none.Tests run after the rebase onto
9b42ae0b. The long runs usede2c57e34, which differs from the pushed head2f215b53only in one skip message's wording, a type annotation, the CLI attach sentence and docs; after those edits the refresh tests, the dynamic prompt tests and the two CLI tests were run again on2f215b53and passed.runtime.test.ts, every test whose name matchesprompt|restore|resume|replay|fork|reset|retry|continuation|refresh|snapshot|render|selector|authority: 124 tests: 107 pass, 0 fail, and 17 Bun-only adapter contracts skipped under Node.dynamic-lifecycle.test.ts, whole file: 23/23 (on2f215b53, its four prompt tests again: 4/4).dynamic-expansion.test.tsandworkspace-preparation-lifecycle.test.ts, which refactor(runtime)!: a run's prompt files are records, not seals #1230's two new commits change: 32/32.packages/prompts:scaffold-catalog,frontmatterandrender: 52/52.runtime.test.ts, the 14 refresh and prompt-edit tests again on2f215b53: 14/14.packages/cli, on2f215b53:resume of an already-active run says no controller was started instead of claiming a submissionandlifecycle commands expose product workflow evidence: 2/2.Gates on
2f215b53, each exit code checked directly:npx prettier --checkon the changed files;pnpm -w lint;CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/claude/prompt-records-not-seals pnpm -w lint:strict:ci, against9b42ae0b;pnpm -w knip, on an unbuilt checkout of this commit and after the build;pnpm --filter @ultrafuzz/config --filter @ultrafuzz/prompts --filter @ultrafuzz/runtime --filter @ultrafuzz/cli typecheck;node scripts/docs-check.mjs;pnpm -w format:check.All exit 0. The two new dynamic-lifecycle tests sit at the end of the file: the file is over its 1000-line budget, and placed beside the other refresh test they moved the
max-linesreport onto a line this branch changed.Rebase. Rebased onto
9b42ae0b, which moved the retry-withdrawal record todynamic-expansion-history/.retry-withdrawal.json. The refresh does not read that record:finishInterruptedDynamicExpansionRetryruns before it and completes any withdrawal it names. OnlyCHANGELOG.mdconflicted, twice, because #1230's entry changed: I took its new text and re-applied this branch's one qualifier to it. Everything above was run after the rebase.Risk / compatibility
Breaking, stated plainly:
resumenow rewrites the prompts of unfinished tasks from the project's current prompts. A hand edit to an unfinished task'sprompt.rendered.md, or to a template copy a later render reads, is replaced by the next resume. The old bytes stay inprompt-history/. Setrun.refresh_prompts_on_resume = falseto keep hand edits..ultrafuzz/prompts/**, or an upgrade that changes a packaged prompt, now reaches runs in flight at their next resume, not only new runs. No disclosure acknowledgement is asked for again.failed,stalled,cancelledorskippedcounts as unfinished, because Smithers can rerun it on a plain resume: retry state is restored from its attempts. So its file is replaced even when its retries are exhausted and it will not rerun. The old bytes are archived.replayandforkstop at their first render once a resume has rewritten one of its runtime prompts or template copies.Behaviour to know:
ultrafuzz.toml, derives the dynamic runtime, and renders every unfinished prompt in memory. That is cheap per prompt, but it was not measured on a large campaign.references.ymlor to the installed build's expansion no longer do.refresh.jsonlists it and its old bytes are beside it.assert-task-inputs(from refactor(runtime)!: a run's prompt files are records, not seals #1230). On a plain resume of a run launched before refactor(runtime)!: a run's prompt files are records, not seals #1230, it stops the whole run instead.Changelog
One
### Breaking changesbullet under## Unreleased, right after #1230's "Prompts are no longer sealed" entry:prompt-history/andrefresh.json, published first;[run] refresh_prompts_on_resume = false, which is read from the currentultrafuzz.tomland never frozen into a run;--refresh-controlleradvice for runs launched earlier, and their replay and fork.#1230's entry gains one qualifier: editing a run's prompt file changes what the next attempt receives "unless a later
resumefirst renders the project's prompt over it, which[run] refresh_prompts_on_resume = falseturns off".🤖 Generated with Claude Code
The PR appears safe to merge; the changes since the previous review introduce no actionable issue.
Fix with agent prompt
Summary
The PR makes resume refresh unfinished tasks’ prompts from current project or packaged prompts by default, with a current-project TOML opt-out. It also records replaced prompt files before applying them, improves prompt-catalog errors, and updates Forge-guard and legacy cloud-cleanup diagnostics. The changes since the previous review are confined to test setup and test registration.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart TD A[Resume passes refusal checks] --> B{Refresh enabled and run eligible?} B -- No --> C[Keep run prompts] B -- Yes --> D[Validate and render unfinished prompts] D --> E[Record planned replacements in refresh.json] E --> F[Archive old bytes and replace changed files] F --> G[Reset or submit controller] C --> GReviews (5) · Last reviewed commit: "test(runtime): launch the stock prompt-r..."