Skip to content

feat(runtime)!: resume applies edited project prompts to unfinished tasks - #1237

Open
aviggiano wants to merge 14 commits into
mainfrom
claude/prompts-refresh-on-resume
Open

aviggiano wants to merge 14 commits into
mainfrom
claude/prompts-refresh-on-resume

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1235 (claude/post-train-followups, 66dc5434) for the merge train #1235 → #1237 → #1238: merge #1235 first; until then this PR also lists its eight commits, and this PR's own are the top four. #1230, named below, has merged (e45076a4).

Stacked on #1230 (claude/prompt-records-not-seals, "a run's prompt files are records, not seals"), which sits on main after the train (#1197, #1201, #1183, #1198) merged. This branch is rebased onto #1230's current head, 9b42ae0b, and merges after it.

Problem

The owner's decision on this follow-up:

We should do this 'enabled by default', so having a toml config to forcefully disable is the best way.

It follows the direction that stability matters more than reproducibility and that an operator must be able to restart a run after a minor prompt change.

After #1230, a run's prompt files are no longer sealed, but a run still reads .ultrafuzz/prompts/** only when it launches. To change what an in-flight task receives, the operator has to hand-edit the run's own files:

  • There is one artifacts/<attempt>/prompt.rendered.md per attempt, with attempt-specific paths inside it.
  • A prompt that is not rendered yet lives in a template copy under dynamic-prompt-templates/, named by the digest of its launch text.
  • An upgrade that changes a packaged prompt never reaches a run in flight.

Change

Four commits. The fourth, test(runtime): an edited stock review prompt that names artifact authorities refreshes on resume, is described under "Greptile follow-up" below.

feat(runtime)!: resume applies edited project prompts to unfinished tasks

Every ultrafuzz resume <run> (plain, --retry-failed, --reset-node or --refresh-controller) now runs a prompt refresh. It runs after the active-run attach, the completion of an interrupted retry withdrawal, the retry-archive plan and the unsupported-run-ID refusal, each of which can end the resume before it resets or archives anything, and before the first reset, the retry archive and the submission.

What it renders.

  • It selects each prompt exactly as run does, from .ultrafuzz/prompts/** and the packaged built-ins.
  • It renders the prompt the way the run's launch did:
    • static prompts with planning's own renderer, against the run's sealed smithers/expanded-graph.json and its frozen smithers/resolved-config.json;
    • runtime prompts with the runtime renderer, against the dynamic runtime derived from the sealed base controls and the published manifests exactly as the next render derives it.

Which tasks it covers. A task counts as unfinished unless Smithers reports its agent node (node:<attempt>) finished and this resume does not reset it. For resets:

  • the reset task is unfinished, except that a reset verifier does not rerun its agent;
  • when the reset reopens its dependents (every reset without --no-deps), so is every task that depends on it.

The refresh then rewrites:

  • the prompt.rendered.md of every unfinished static task;
  • the published prompts of unfinished generated children, and of nodes that waited on a dynamic group, such as the final report;
  • the template copies under dynamic-prompt-templates/ that a later render reads: that of a group that has not expanded, and that of an unfinished task whose prompt is not published yet.

A finished task keeps its file, which is the record of the prompt it ran with. When --retry-failed reruns a dynamic source, the groups the withdrawal plan covers count as unexpanded: their generated children and the prompts that wait on them keep their files, which the withdrawal moves to dynamic-expansion-history/ as the record of what ran, and their template copies take the edits. A file is rewritten only when its bytes change: an unchanged project renders byte-identical prompts, and the refresh writes nothing (tested, also on the stock default topology).

It never fails the resume.

  • It skips entirely with PROMPT_REFRESH_SKIPPED (warning), and changes nothing, when:

    • the effective topology file's bytes differ from the digest plan.json recorded at launch (any edit, even to a comment), or the topology's nodes are not the run's (an eval run launched with exclusions);
    • a prompt file breaks the whole catalog as it would break run, such as invalid frontmatter or a duplicate id; the warning names the file;
    • ultrafuzz.toml cannot be read;
    • the run's documents cannot be read or its dynamic runtime cannot be derived;
    • the engine still reports the run active (--force or --reset-node on an active run).
  • Each prompt is applied to every file rendered from it, or to none. It is validated on its own, the way run validates it:

    • the project copy must exist as a regular file;
    • its variables must parse;
    • its artifact references must name ancestors with those outputs.

    It is then rendered for each of its unfinished tasks, which also runs the validator-command check.

    A prompt that fails any of these keeps every file rendered from it, and resume reports PROMPT_REFRESH_REJECTED (warning), naming the prompt and the error. So does a prompt that meets either of these:

    • it names an ancestor-artifact authority that a static task was not compiled with. Dropping one is accepted, following the stability critique;
    • it shares a template copy that a later render reads with another prompt whose new text differs. Every prompt bound to the copy counts.

    The other prompts still apply. A rejected prompt that no unfinished task or template copy uses changes nothing, so it is not reported.

  • Before it changes any file, the refresh publishes prompt-history/<time>-<uuid>/refresh.json, which lists each file it will rewrite with its attempt, its prompt, and its old and new SHA-256. It then copies each file into that entry, at its path in the run, and replaces it atomically with writeFileDurable. Resume reports PROMPTS_REFRESHED (info), which names the tasks and the entry.

  • If a step fails part way, resume reports PROMPT_REFRESH_INCOMPLETE with how many of the planned files it replaced. Each file refresh.json lists holds either its old or its new bytes, and a resume after the cause is fixed applies the rest.

Turning it off. The new [run] refresh_prompts_on_resume = false turns the refresh off. Resume reads the key from the project's current ultrafuzz.toml every time, so it also applies to runs already in flight. The config loader accepts the key, and applyRunConfig drops it, so it never enters a run's resolved config. That matters: the resolved-config schema is strict and requires every run field, so adding the key there would make every older run's sealed config fail to parse on resume.

Where the code is.

  • prompt-refresh.ts (new) plans the refresh.
  • prompt-history.ts (new) records and applies it.
  • plan-run.ts: the plan-time static renderer is split so that renderRunStaticPrompts renders in memory, creates no artifact directories, and collects per-attempt failures. Launch's renderPromptsForPlan now writes that renderer's results, so both use the same render code.
  • dynamic-runtime.ts: renderRuntimePrompt takes an optional template body; renderRuntimePromptsFromTemplates, at the end of the file, renders listed tasks in memory; deriveDynamicRuntimeMaterialization derives the runtime without publishing or comparing anything.
  • audit-profile-policy.ts: effectiveTopologyPath selects the run's topology file as effectiveAuditPolicy does. It skips that function's topology load, which validates the project's prompt files and would let one bad prompt skip the whole refresh.
  • smithers.ts: a beforeContinuation hook receives a SmithersContinuationContext: the inspected node states, the resets the command will issue (from the same list the reset loop iterates), the groups the retry withdrawal covers, and whether the run is active.
  • start-run.ts passes the hook.
  • dynamic-expansion-retry.ts exports readDynamicRuntimeBase.

Docs.

  • restart-continue.md: "Change A Prompt Of A Running Campaign" is rewritten around "pause, wait for paused, edit the project prompt, resume". It says what the refresh covers and reports, when it skips, what a failed reset leaves, and how to turn it off to edit a run's own files. The refactor(runtime)!: a run's prompt files are records, not seals #1230 hand-edit guidance moves under "Edit A Run's Own Prompt Files", and the rule for runs launched earlier gets its own subsection.
  • configuration.md documents the key and the one exception to "configuration is frozen at launch".
  • SPECS.md, security.md, edit-prompts-topology.md, prompt-variables.md and dashboard.md are updated to match; cli.md and artifacts-reports.md say it in a sentence and link to the how-to.

Size. About +1180 / −200 production lines across the three commits, including doc comments, and +775 test lines. The review fixes account for +336 / −310 of the production lines: they add the derive-only mode, the withdrawal and live-copy rules and the record-first apply, and delete the topology re-expansion, continuationResets and the fixed-point loop.

fix(prompts): a prompt catalog error names the file that caused it

One file that does not parse fails the whole catalog, for run, validate and the refresh, and the error did not say which file it was. Each per-file parse or variable error now starts with the prompt's source and relative path (project prompt setup/variant.md: invalid prompt YAML frontmatter: ...), and a duplicate project prompt id names both files.

fix(runtime): refresh prompts only after a resume can no longer refuse, and only files a later render reads

The review fixes; the first commit's description above already describes the result. What changed, and why:

  • Order. The refresh ran before the interrupted-withdrawal completion, the retry-archive plan and the unsupported-run-ID refusal, so a resume that one of them refused had already rewritten the prompts of finished tasks it only predicted it would reset. It now runs after all of them, just before the first timetravel. continuationResets, which repeated the reset loop's logic, is gone: the hook and the reset loop read one list.
  • Withdrawn generations. A --retry-failed of a dynamic source rewrote the finished children's and the join's prompts just before the withdrawal moved them, so dynamic-expansion-history/ no longer held what they ran with.
  • Template copies. Every copy was rewritten, even one no render would read again. That also broke replay and fork of runs launched before refactor(runtime)!: a run's prompt files are records, not seals #1230 for groups that had finished. The shared-copy rule now settles in one pass: a prompt has exactly one copy, named by its launch text.
  • Record. refresh.json was written after the last file, so a failure or a kill left no record, and PROMPT_REFRESH_INCOMPLETE claimed replacements that had not happened.
  • Topology check. It re-expanded the topology with the installed build and diffed the graph through copies of planning internals, so a build that adds a field to expanded nodes, or a references.yml change, would have skipped every in-flight run's refresh. It now compares the file with audit_profile.topology_digest and the node sets.
  • Dynamic runtime. Admission's tamper check threw on a stock run before its first render (launch writes dynamic_groups in compile order and the derivation sorts them), which skipped every refresh. The refresh now derives without comparing.
  • Smaller. One shared test helper for the TOML key; the hook's context type is exported and reused; an unread failure field is dropped; the native-resume test turns the refresh off instead of relying on the fake run being active; the CLI attach message now ends "This resume applied none of the project's current prompts; only a resume that starts a controller does."

Review findings and what was done

Finding Disposition
Stability, major: finished tasks rewritten before steps that can refuse the resume or withdraw them Fixed (order, withdrawn groups). Tested: a refused --retry-failed changes no prompt file; a successful one archives the children's and the join's launch bytes. The one window left, a timetravel that fails after the refresh, is documented.
Stability, minor: runs launched before #1230 lose replay and fork after a copy rewrite Narrowed (only copies a later render reads) and documented in the CHANGELOG and "Runs Launched By An Earlier Release", with the remedy. Not special-cased (0.x policy).
Stability and scope, minor: a catalog-wide error skips every edit and names no file The error now names the file (both files for a duplicate id); the skip is documented. Not isolated per prompt: run selects prompts by id across every file, so a file that does not parse could carry any prompt's id, and run refuses the same catalog.
Stability, nit, and scope, minor: refresh.json written last; misleading PROMPT_REFRESH_INCOMPLETE Fixed; two tests.
Stability, nit: native-resume test depends on the fake run being active Fixed: the key is off in that test, which asserts no prompt diagnostics.
Scope, minor: a running campaign's resume attaches and applies nothing, silently Docs lead with pause, wait, edit, resume; the CLI attach message says so.
Scope, nits: every copy is a target; CHANGELOG, artifacts-reports.md and security.md gaps Fixed; #1230's "editing that file" sentence is qualified.
Simplicity, minor: continuationResets duplicates the reset loop Fixed; --retry-failed test added.
Simplicity, minor: topology re-expansion Replaced by the digest and node-set check. A comment-only topology edit now skips the refresh; the docs say so.
Simplicity, minor: the tamper check skips the refresh on a stock run Derive-only mode; stock-topology test.
Simplicity, minor: fixed-point settleTemplateCopies without a test One pass; shared-copy test.
Simplicity, nits: unread field, duplicate helpers and types, docs repeated in seven places Fixed, except that the two runtime render loops stay separate: merging them adds a mode to the engine's per-render loop to save about 20 lines.

Greptile follow-up

  • P1 "Runtime prompts bypass authority checks" (prompt-refresh.ts:372): declined. Instead, a7a86f3d adds a guard test and makes the docs precise.
    • The observation is correct. Runtime prompts are published without the compiled-selector check that static prompts get.
    • That check has no sealed set to compare against. No task whose prompt is rendered at runtime, meaning generated children and tasks that wait on a dynamic group, is compiled with any authority selector. This is also true on main and has been since v0.1.0: plan time renders only static tasks, and only those rows seal selectors (Runtime-rendered prompts may name prompt-artifact authority files that are never written #1234). On a stock launch, dedupe-findings, aggregate-test-files, final-report and both goal-group templates seal none, while their launch prompts already name 3, 1 and 2 authorities.
    • The refresh adds nothing launch cannot produce. In the flagged case, an edit that adds a selector, the refreshed prompt is byte for byte what a launch of the edited prompt renders (compared with run paths normalized).
    • The proposed rule would regress the stock prompts. Applied here, the static rule rejects every edit of those three stock review prompts once their groups expand. If it runs on every render, it also warns three times on every resume of a stock run past goal-plan.
    • Test. The new test, resume applies an edited stock review prompt that names artifact authorities once its groups expanded, fails under both forms of the rule. The 16 existing refresh tests pass under both, so nothing guarded this before.
    • Docs. The how-to, the CHANGELOG entry and the doc comments now say the check covers static tasks, and the comment at the flagged function says why.
    • What is left. Serving authorities to runtime-rendered tasks is Runtime-rendered prompts may name prompt-artifact authority files that are never written #1234. That fix should also decide what the refresh checks them against.
    • Checks on a7a86f3d. The 12 refresh tests in runtime.test.ts pass. pnpm -w format:check, pnpm -w lint, lint:strict:ci against origin/main, pnpm -w knip, the runtime typecheck and node scripts/docs-check.mjs all exit 0.

Deliberately not built

  • An events.jsonl event. The task asked for one event that lists the tasks whose prompts changed, but that event cannot be added safely:

    • every event type is a variant of the closed event-record.schema.json;
    • that file is covered by the artifact schema-bundle digest;
    • the engine still requires a run's launch bundle, so a new variant would fail every task of every in-flight run on resume. This is the failure refactor(runtime)!: a run's prompt files are records, not seals #1230 dropped its ledger commit for.

    The record is prompt-history/<entry>/refresh.json together with the PROMPTS_REFRESHED diagnostic. The bundle digest is unchanged, and packages/artifacts is untouched.

  • A CLI flag. The owner chose a TOML key only.

  • A refresh on replay and fork. They keep rendering the run's current files, as in refactor(runtime)!: a run's prompt files are records, not seals #1230.

  • Exact prediction of Smithers' reset set. Smithers resets by attempt start time: timetravel also resets every node whose attempt started after the target's (resolveResetNodes), sibling branches included. The refresh runs before the reset, and inspect carries no start times, so it counts only the reset task and the tasks that depend on it. A finished task that merely started later reruns with its old prompt. The docs say so.

  • Applying after the reset. The design's §9 sketch applied after the resets, which would also close the failed-timetravel window. The brief puts the refresh before any timetravel, and a split plan and apply adds a second hook for a window that ends in a failed resume, whose files refresh.json names.

  • Per-prompt isolation of a broken catalog (see the table).

  • Re-acknowledging data disclosure after a refresh (the design's Q2). Acknowledgements stay launch-scoped, and security.md says so.

  • Refusing --refresh-controller together with the refresh (the design's §9 and T10). The refresh renders with the run's planned schema bindings, exactly as launch did. The refreshed controller's rebinding to the installed schemas is the same skew as before (the K2 follow-up).

  • Treating runs launched before refactor(runtime)!: a run's prompt files are records, not seals #1230 specially. On a plain resume, their launch workflow still stops the whole run when a prompt no longer renders, and their replay and fork run the launch snapshot, which still compares runtime prompts and template copies. The docs and the CHANGELOG say to resume them with --refresh-controller, and to turn the refresh off before resuming such a run that must still be replayed or forked.

  • Updating launch provenance. prompt_digest, the manifest template.prompt_sha256 and rendered_prompt_digest stay what launch recorded; the CHANGELOG and artifacts-reports.md say so. refresh.json is the record of what changed.

Verification

Discriminating tests. Each was proven by copying this branch's runtime.test.ts, dynamic-lifecycle.test.ts, prompt-refresh-config.ts and scaffold-catalog.test.ts into a detached, freshly built worktree and running them there, with nothing else changed: on #1230's current head 9b42ae0b, and on b32874f4, this branch's head before the review fixes (on #1230's previous head 94123ada), which shows what the review fixes change. The first round proved the same tests, and config.test.ts, on 0e6025d4 and 94123ada.

Test Base 9b42ae0b Before the review fixes b32874f4 Here
runtime: resume applies an edited project prompt to a task that has not run and keeps a finished task's prompt fails: no PROMPTS_REFRESHED passes passes
runtime: run.refresh_prompts_on_resume = false keeps the run's prompts until it is turned back on fails passes passes
runtime: an edited prompt that does not validate keeps its tasks' prompts, and the resume and other edits go ahead fails passes passes
runtime: an edited prompt that names an artifact authority its task was not compiled with keeps the run's prompt fails passes passes
runtime: a topology changed since launch skips the prompt refresh with a warning fails fails: the old check's message passes
runtime: resume --reset-node applies edited prompts to the finished task it reruns and to its dependents fails passes passes
runtime (new): resume --retry-failed applies edited prompts to the finished producer of a failed verifier and to its dependents fails passes passes
runtime (new): resume of a stock run renders its prompts as launch did and refreshes an unexpanded group's template copy fails fails: PROMPT_REFRESH_SKIPPED, the tamper check passes
runtime (new): a prompt file that breaks the prompt catalog skips the refresh with a warning naming the file fails fails: no file named passes
runtime (new): a refresh that cannot create its history entry changes no prompt file and says so fails fails: claims replaced files passes
runtime (new): a refresh that fails part way has already recorded every file it planned fails fails at the old message, before its refresh.json check passes
dynamic-lifecycle (changed): resume applies edited project prompts to unfinished dynamic children and to prompts not rendered yet fails fails: rewrites the group's template copy, which no render reads passes
dynamic-lifecycle (new): a source retry withdraws its generation's prompts as they ran and refreshes the prompts it reruns fails fails: the refused resume had already rewritten a finished child's prompt passes
dynamic-lifecycle (new): prompts that share a template copy apply their own edits until a later render reads the copy fails fails: rejects both while no render reads the copy passes
prompts (new): names the prompt file that breaks the catalog, and both files of a duplicate id fails not run (its catalog.ts is the base's) passes

native resume delegates the persisted workflow after mutable project sources are replaced now turns the refresh off and asserts no prompt diagnostics; it passes everywhere, as a pinned test should.

Assertions inside them cover each required case:

  • A not-yet-run static task receives the edit. actors-flows' file contains the edit and still has its output contract.
  • A finished task's file is unchanged, in the static tests, the dynamic tests (the planner) and the e2e (project-discovery). A --retry-failed of a dynamic source leaves the withdrawn children's and the join's launch bytes in dynamic-expansion-history/, and a refused retry changes no prompt file.
  • An invalid edit warns, keeps the old prompt, and the resume still submits. The rejected prompt is named, and a valid edit to the other task still applies. A catalog-breaking file names itself.
  • A topology change skips the refresh with a warning, and no prompt-history/ is written.
  • An edited group template reaches unfinished dynamic children, through their published prompts, and through the refreshed copy once a render reads it: a child rendered after the refresh renders from it, and so do the children a retried source regenerates. Admission, syncRun and getRunHealth stay clean.
  • The record is complete. refresh.json lists every rewritten file, the archived copies hold the old bytes, and a failed apply has already published it.
  • Unchanged prompts produce no rewrite and no history, on the test topology and on the stock default one, and a second resume after a refresh changes nothing.

Real engine. packages/cli/test/e2e/campaign-resume.test.ts (CI lane cli-e2e) edits the project prompt instead of the run file; the first round ran it (1/1, about 10 minutes) on this branch and saw it fail on 94123ada. It was not re-run after the review fixes: they change only the order of steps it does not reach (no retry, no reset) and which template copies are rewritten, and its topology has none.

Tests run after the rebase onto 9b42ae0b. The long runs used e2c57e34, which differs from the pushed head 2f215b53 only in one skip message's wording, a type annotation, the CLI attach sentence and docs; after those edits the refresh tests, the dynamic prompt tests and the two CLI tests were run again on 2f215b53 and passed.

  • runtime.test.ts, every test whose name matches prompt|restore|resume|replay|fork|reset|retry|continuation|refresh|snapshot|render|selector|authority: 124 tests: 107 pass, 0 fail, and 17 Bun-only adapter contracts skipped under Node.
  • dynamic-lifecycle.test.ts, whole file: 23/23 (on 2f215b53, its four prompt tests again: 4/4).
  • dynamic-expansion.test.ts and workspace-preparation-lifecycle.test.ts, which refactor(runtime)!: a run's prompt files are records, not seals #1230's two new commits change: 32/32.
  • packages/prompts: scaffold-catalog, frontmatter and render: 52/52.
  • runtime.test.ts, the 14 refresh and prompt-edit tests again on 2f215b53: 14/14.
  • packages/cli, on 2f215b53: resume of an already-active run says no controller was started instead of claiming a submission and lifecycle commands expose product workflow evidence: 2/2.

Gates on 2f215b53, each exit code checked directly:

  • npx prettier --check on the changed files;
  • pnpm -w lint;
  • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/claude/prompt-records-not-seals pnpm -w lint:strict:ci, against 9b42ae0b;
  • pnpm -w knip, on an unbuilt checkout of this commit and after the build;
  • pnpm --filter @ultrafuzz/config --filter @ultrafuzz/prompts --filter @ultrafuzz/runtime --filter @ultrafuzz/cli typecheck;
  • node scripts/docs-check.mjs;
  • pnpm -w format:check.

All exit 0. The two new dynamic-lifecycle tests sit at the end of the file: the file is over its 1000-line budget, and placed beside the other refresh test they moved the max-lines report onto a line this branch changed.

Rebase. Rebased onto 9b42ae0b, which moved the retry-withdrawal record to dynamic-expansion-history/.retry-withdrawal.json. The refresh does not read that record: finishInterruptedDynamicExpansionRetry runs before it and completes any withdrawal it names. Only CHANGELOG.md conflicted, twice, because #1230's entry changed: I took its new text and re-applied this branch's one qualifier to it. Everything above was run after the rebase.

Risk / compatibility

Breaking, stated plainly:

  • By default, every resume now rewrites the prompts of unfinished tasks from the project's current prompts. A hand edit to an unfinished task's prompt.rendered.md, or to a template copy a later render reads, is replaced by the next resume. The old bytes stay in prompt-history/. Set run.refresh_prompts_on_resume = false to keep hand edits.
  • An edit to .ultrafuzz/prompts/**, or an upgrade that changes a packaged prompt, now reaches runs in flight at their next resume, not only new runs. No disclosure acknowledgement is asked for again.
  • A task in state failed, stalled, cancelled or skipped counts as unfinished, because Smithers can rerun it on a plain resume: retry state is restored from its attempts. So its file is replaced even when its retries are exhausted and it will not rerun. The old bytes are archived.
  • For a run launched before refactor(runtime)!: a run's prompt files are records, not seals #1230, replay and fork stop at their first render once a resume has rewritten one of its runtime prompts or template copies.

Behaviour to know:

  • Resume now reads the project's topology file, prompts and ultrafuzz.toml, derives the dynamic runtime, and renders every unfinished prompt in memory. That is cheap per prompt, but it was not measured on a large campaign.
  • Any edit to the effective topology file, even to a comment, skips the refresh until the file is restored. Edits to references.yml or to the installed build's expansion no longer do.
  • If the engine fails to reset a task after the refresh, the resume fails with that task's file already rewritten; refresh.json lists it and its old bytes are beside it.
  • A resume of a run the engine still reports active attaches and applies nothing.
  • A prompt that validates and renders now can still fail later: an item-variable typo fails when its group expands. That fails only the child's task at assert-task-inputs (from refactor(runtime)!: a run's prompt files are records, not seals #1230). On a plain resume of a run launched before refactor(runtime)!: a run's prompt files are records, not seals #1230, it stops the whole run instead.
  • A published runtime prompt whose template copy is not rewritten, and which is then deleted, is rendered again from the copy's older text.

Changelog

One ### Breaking changes bullet under ## Unreleased, right after #1230's "Prompts are no longer sealed" entry:

  • resume now applies the project's current prompts to unfinished tasks by default, when, and what it covers;
  • prompt-history/ and refresh.json, published first;
  • the four diagnostics, when the refresh is skipped and when a prompt is rejected, and that catalog errors now name their file;
  • that hand edits are replaced unless [run] refresh_prompts_on_resume = false, which is read from the current ultrafuzz.toml and never frozen into a run;
  • that project and packaged prompt edits now reach runs in flight, and that an attach applies nothing;
  • that launch digests keep their values;
  • the reset gap and the failed-reset window;
  • the --refresh-controller advice for runs launched earlier, and their replay and fork.

#1230's entry gains one qualifier: editing a run's prompt file changes what the next attempt receives "unless a later resume first renders the project's prompt over it, which [run] refresh_prompts_on_resume = false turns off".

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge; the changes since the previous review introduce no actionable issue.

Fix All in Claude CodeFindings

  1. P1 Runtime prompts bypass authority checks ▶
Fix with agent prompt
### Issue 1
packages/runtime/src/prompt-refresh.ts:373-375
When an edited prompt for an unfinished generated child or deferred task adds a valid ancestor-artifact authority selector, this path publishes the rendered prompt without checking that selector against the task’s compiled selectors. The prompt can instruct the agent to use an authority absent from its sealed authority document, instead of being rejected as the static refresh path rejects it. Return the rendered artifact references from the runtime renderer and check them against the compiled selectors before publishing.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR makes resume refresh unfinished tasks’ prompts from current project or packaged prompts by default, with a current-project TOML opt-out. It also records replaced prompt files before applying them, improves prompt-catalog errors, and updates Forge-guard and legacy cloud-cleanup diagnostics. The changes since the previous review are confined to test setup and test registration.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Resume passes refusal checks] --> B{Refresh enabled and run eligible?}
  B -- No --> C[Keep run prompts]
  B -- Yes --> D[Validate and render unfinished prompts]
  D --> E[Record planned replacements in refresh.json]
  E --> F[Archive old bytes and replace changed files]
  F --> G[Reset or submit controller]
  C --> G
Loading

Reviews (5) · Last reviewed commit: "test(runtime): launch the stock prompt-r..."

@aviggiano
aviggiano requested a review from a team as a code owner September 30, 2026 19:27
Comment on lines +370 to +372
tryPropose(state, [target.prompt], attemptId, () => {
if ("error" in render) throw new Error(render.error);
return { relativePath: promptFilePath(attemptId), prompt: target.prompt, attemptId, next: render.markdown };

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Runtime prompts bypass authority checks

When an edited prompt for an unfinished generated child or deferred task adds a valid ancestor-artifact authority selector, this path publishes the rendered prompt without checking that selector against the task’s compiled selectors. The prompt can instruct the agent to use an authority absent from its sealed authority document, instead of being rejected as the static refresh path rejects it. Return the rendered artifact references from the runtime renderer and check them against the compiled selectors before publishing.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/prompt-refresh.ts
Line: 370-372

Comment:
**Runtime prompts bypass authority checks**

When an edited prompt for an unfinished generated child or deferred task adds a valid ancestor-artifact authority selector, this path publishes the rendered prompt without checking that selector against the task’s compiled selectors. The prompt can instruct the agent to use an authority absent from its sealed authority document, instead of being rejected as the static refresh path rejects it. Return the rendered artifact references from the runtime renderer and check them against the compiled selectors before publishing.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

aviggiano and others added 4 commits September 30, 2026 20:56
…ritable umask

Problem: under a umask that leaves new directories group writable, such as
Ubuntu's default 0002, ensureSafeDirectory created <run>/safe-bin 0775.
composeSmithersCommandPath admits a target-local PATH entry only through
isPreparedForgeGuardBin, which refuses a group-writable directory, so the
workflow engine's PATH dropped the wrapper and agents ran the real forge
without forge_vmem_limit_kb or forge_rayon_threads while run.json recorded
forge_guard.active: true. A custom run.output_dir produces the same false
claim under any umask, because that check admits the wrapper only from
<project>/.ultrafuzz/runs/<run-id>/safe-bin.

Change: prepareForgeGuardEnvironment sets safe-bin to 0700 after creating
it (chmod does not apply the umask, and it also repairs the directory of a
run launched earlier), then applies the engine's own admission check to it.
It reports the guard active only when that check passes. Otherwise it
leaves PATH and the guard variables unset and returns a FORGE_GUARD_INACTIVE
warning, which launch, resume, replay and fork return with their result.
The admission check itself is unchanged. resume now also records the guard
of the controller it starts in run.json, best effort like its state
projection, so run.json no longer keeps a launch's claim after a resume
that dropped the wrapper.

The OpenRouter adapter contract test failed on such hosts for a different
reason: its fixture rooted the operator-owned provider homes under the
target's .ultrafuzz, which `ultrafuzz init` creates with the process
umask. The adapter creates every provider-home component itself with mode
0700, so the fixture now uses its own private root; the provider-home
ancestor check is unchanged.

The new tests pin umask 0002 through a shared helper, so they reproduce on
CI runners, whose umask is 022.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…loud run

Problem: before #1197 removed per-node cloud execution, `ultrafuzz clean`
of a run planned with `[execution] mode = "cloud"` terminated its sandboxes
(tagged purpose=ultrafuzz-node) and deleted its Modal volume before
deleting the run directory. #1197 dropped that step and kept the deletion,
so cleaning such a run now leaves its Modal volume and any running sandbox
in place, still billed, and deletes the plan.json that records the Modal
app they live in, without saying so.

Change: before anything is removed, clean reads the mode and Modal app
from each selected run's plan.json, which copies the execution block of
the run's resolved config and is what the removed cleanup read. For a
cloud run it returns a CLEAN_CLOUD_STORAGE_RETAINED warning naming the
app, the volume and the sandboxes' run tag, and the `modal volume delete`
command. `--dry-run` reports it without deleting, and every later result,
including a failed removal, carries it. The deletion itself is unchanged.

The volume is not `ultrafuzz-node-<run-id>`: the removed provider named it
`ultrafuzz-node-` plus a bounded identity of the Smithers run ID
`ultrafuzz-<run-id>` (at most 32 normalized characters, then 12 hex digits
of its SHA-256). The same identity was the sandboxes' `run` tag. clean
reproduces that function, and the test pins the name the removed code
computes for a fixed run ID.

An unreadable plan.json yields no warning. It also makes the source
revision step fail before anything is removed, as it did before, so no run
that clean deletes skips the check.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: #1197 made compilation write every task's execution block as
`{ mode: "local", resources, agentCredentialEnv: [] }` with no `modal`
entry, but synchronization still added task.execution.agentCredentialEnv
and task.execution.modal?.credentialEnv to its exact-value redaction list.
For every run compiled since #1197 both are empty. The generated agent
environment also still blanked ULTRAFUZZ_MODAL_MODULE, which only the
removed cloud provider set. And the doctor tests still wrote a fake engine
into the target's .smithers, although since #1201 doctor inspects only the
runner Ultrafuzz's own install provides.

Change: synchronization derives its redaction values from the environment
alone, the agent environment no longer lists ULTRAFUZZ_MODAL_MODULE, and
seven doctor tests lose the dead fake-engine setup. One test used it only
to create the .smithers/node_modules/.bin directory it writes into, and
now creates that directory itself. The doctor test that installs a 0.29.0
project-local engine keeps it: it asserts that doctor ignores that engine.
The task manifest keeps agentCredentialEnv, because its sealed schema
requires the field.

Effect on runs compiled before #1197 in cloud mode, whose lists were not
empty: sensitiveEnvironmentValues still redacts, by name, every variable
that looks like a credential (*_API_KEY, *_TOKEN, *_SECRET, ...), which
includes MODAL_TOKEN_ID and MODAL_TOKEN_SECRET. What these lists added
beyond that were route variables such as *_BASE_URL and allowlisted
variables. Those carry no secret, or have a value the secret patterns
already match.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: #1201 wrote <run>/trusted-bin/smithers at launch, resume, replay
and fork so that the generated workflow's bare `smithers node` call
(#1143) found a runner. #1183 removed that call: final-report producer
retries and verifiers now read the selections each producer attempt
records in the run. The current template spawns only git, bash and the
run's `ultrafuzz` launcher, and every command the runtime spawns runs an
executable it bound by path, never `smithers` from PATH.
smithers-report-retry.integration.test.ts puts a failing `smithers` first
on the engine PATH around the production report code and still passes.
The shim therefore served only workflows persisted by earlier releases,
and it put an engine CLI first on every task's PATH. It also made replay
and fork bind the installed runner, which they otherwise never run: they
run the run's sealed engine.

Change: delete writeTrustedSmithersShim and its three callers. When resume
cannot re-verify the trusted CLI, it again keeps an existing run launcher
first on PATH itself, which the shim had done as a side effect; the
existing test for that path covers it. Launch still binds the installed
runner before creating a run, because resume runs it; the comments and
the doctor summaries now say launch and resume. The capability and
native-continuation tests drop their shim assertions, and a resume test
now asserts that trusted-bin holds only the ultrafuzz launcher. The docs
no longer say that the shim serves the workflow's smithers calls or that
tasks can drive their run through it. The CHANGELOG entry for #1201,
which has not been released, drops its shim claims, and a new entry
records the removal.

BREAKING CHANGE: tasks no longer find a `smithers` CLI that Ultrafuzz
provides on PATH. A run launched by an earlier release and continued by
plain `resume` still runs its persisted workflow, whose bare `smithers`
call in a restarted controller finds a runner only on the operator's own
PATH, as before #1201; continue such a run with
`resume --refresh-controller`.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 4 commits September 30, 2026 22:28
…e run

Problem: for a run an earlier release planned for per-node Modal execution,
the CLEAN_CLOUD_STORAGE_RETAINED warning names the Modal volume, app and
sandbox tag that `clean` leaves behind, and the run's plan.json, which the
removal deletes, is the only record of the app. cleanRun computed the
warning before removing anything but returned it only with its result, and
the CLI prints a successful result's diagnostics after "Removed: ...". The
operator therefore saw it only once the plan was gone, and never when clean
did not return: the audit append after the removal had no try/catch, so a
journal lock timeout, a full disk or a read-only .ultrafuzz rejected
cleanRun after the run was deleted and dropped the warning with it.

The CHANGELOG and the how-to also gave the volume as
ultrafuzz-node-ultrafuzz-<run-id>-<hash>, which is wrong for every generated
run ID. The removed provider kept at most 32 characters of
ultrafuzz-<run-id>, and a generated ID such as
run-20260930t161149123z-ab12cd34 makes that 42, so its volume is
ultrafuzz-node-ultrafuzz-run-20260930t161149123-d99772718708. The runtime
message was right; the test pinned only a short run ID.

Change:
- CleanGeneratedInput takes an optional onRetainedStorage callback. cleanRun
  calls it with the warnings just before the first deletion, after the
  source-ref step that can still stop the removal; a dry run does not call
  it. Every result still carries the warnings.
- The CLI's text output prints them from that callback and leaves them out
  of its final text, so they appear once, before anything is removed. JSON
  output and the dashboard keep reading them from the result.
- A failed audit append is now a CLEAN_AUDIT_FAILED failure that says
  whether anything was removed and still carries the warnings.
- The docs describe the volume name as derived, say to copy it from the
  warning or find it with `modal volume list`, and the runtime test pins
  the truncated name for a generated run ID.

On the branch before this commit, the new tests fail as intended: the
callback is never called, the audit test rejects with EACCES on
.ultrafuzz/clean-audit.jsonl.lock, and the CLI test sees the warning first
written after plan.json was deleted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem: four ways the Forge guard follow-up could still leave run.json or
the operator with a wrong picture.
- A synchronization pass (status --watch, the dashboard, the eval runner)
  reads run.json, then awaits usage replay and possibly a pricing fetch, and
  wrote back that copy plus its accounting. It holds only
  .workflow-sync.lock, and resume records its controller's guard holding
  only the lifecycle lock, so a pass that read before a resume's write and
  wrote after it put the launch's forge_guard back, and nothing rewrote it
  until the next lifecycle command. A test whose pricing fetch rewrites the
  guard mid-pass reproduces this on the branch: run.json ends with the
  launch's active: true. replay and fork, which record the guard the same
  way, had this on main.
- The engine PATH admits safe-bin only while it holds the wrapper alone.
  The .forge.tmp-* file an interrupted writeFileDurable leaves, or any other
  entry, therefore dropped the guard for every later launch, resume, replay
  and fork of the run, and nothing removed it.
- writeFileDurable created the wrapper 0600 and chmod made it 0700 after the
  rename. resume and replay rewrite the wrapper while tasks may be running,
  and a PATH lookup in that window skips a file without its execute bit and
  runs the real Forge.
- A resume that only attached to a running controller still returned
  FORGE_GUARD_INACTIVE about the controller it did not start. The warning,
  config.md and configuration.md also left out the admission's
  no-symbolic-link condition, so a project reached through a symlinked
  parent directory was told the engine admits only the directory the
  wrapper was in.

Change:
- @ultrafuzz/artifacts adds updateRunMetadataDocument, which re-reads
  run.json and writes the update while holding run.json.lock, taken with the
  existing journal lock. Recording the Forge guard and synchronization's
  accounting write both use it. The accounting now goes into the document as
  it is at the write, which must still be bound to the workflow it was
  computed for; otherwise the pass reports WORKFLOW_ACCOUNTING_FAILED, as it
  already did for that mismatch at its start, and writes nothing.
- prepareForgeGuardEnvironment removes every entry but forge from safe-bin
  before writing the wrapper, best effort: an entry it cannot remove still
  fails the admission, which it reports. It writes the wrapper with mode
  0700, so the rename replaces an executable file with an executable file.
- resume returns the Forge guard diagnostics only when it starts a
  controller, the same condition under which it records the guard.
- The warning and the docs state the symbolic-link condition.

The unit test's stray-file case, which expected the guard inactive, now
expects it cleared and active; the inactive cases are a custom output_dir
and a symlinked parent. The resume test makes the guard inactive with an
entry resume cannot remove, and is skipped under root. On the branch before
this commit, the new tests fail as intended, including the attach case,
which returned the warning.

Not covered: replay and fork rebind run.json's workflow under the control
lock, not run.json.lock. The accounting write now refuses a document bound
to another workflow instead of writing the old binding back, which narrows
that race on main to the lock-held re-read and write, but does not close it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…shim

Problem: a Breaking-changes entry announced that launch, resume, replay and
fork no longer write <run>/trusted-bin/smithers. No release wrote it: #1201
added it after v0.1.2, and this branch already rewrote #1201's entry without
it. Measured against v0.1.2, nothing the entry lists changes: replay and fork
never refused an unpatched install, and a run launched by an earlier release
already found its persisted workflow's bare `smithers` only on the
operator's PATH.

Change: remove the entry. Its one piece of advice moves to the #1183 entry,
which already explains that plain resume keeps running the persisted
workflow: such a run still runs `smithers node` in a restarted controller,
so continue it with `resume --refresh-controller`. The #1183 entry already
gives the --reset-node form for a report producer that has succeeded.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…safe-bin

Launch holds the workflow control lock and resume, replay and fork the
lifecycle lock, so a launch and a resume can prepare the same run's Forge
guard at once. The wrapper was written through a temporary file inside
safe-bin, so the stray-entry cleanup of one command could delete the
other's temporary file before its rename, failing that command with
ENOENT. An engine PATH composed meanwhile also saw two entries in safe-bin
and dropped the wrapper, although run.json recorded the guard active.

writeFileDurable now takes an optional temporaryDirectory, and the
wrapper's temporary file goes in the run root, so safe-bin only ever holds
the wrapper. The cleanup still removes temporary files that interrupted
writes left in safe-bin before this change.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/prompts-refresh-on-resume branch from 29f03e4 to 4bd75ad Compare September 30, 2026 22:34
aviggiano and others added 6 commits September 30, 2026 22:47
…ugh testWhen

The test "resume records the Forge guard of the controller it starts"
passed { skip: ... } to test() to stay out of root runs. Bun's node:test
shim ignores that option, so the monolithic suite bans it (#709), and "the
monolithic runtime suite registers exclusively through the shard wrapper"
failed on the stray skip:.

Register the test through testWhen(process.getuid?.() !== 0), which
selects test.skip for root in both the Node and Bun lanes, as the lock
synchronization test does. A comment keeps the reason the skip option
carried.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…asks

Before it resets, archives or submits anything, every `ultrafuzz resume`,
with or without --retry-failed, --reset-node or --refresh-controller, now
applies the project's current prompts (.ultrafuzz/prompts/** and the
packaged built-ins, selected exactly as `run` selects them) to every task
of the run that has not finished. It re-renders, with planning's own
renderer and the run's frozen config, the prompt of every static task whose
agent the engine does not report `finished`, plus the tasks a reset reopens
and their dependents; the published prompts of unfinished generated
children and of nodes that waited on a dynamic group; and the template
copies that prompts not rendered yet are rendered from. Finished tasks keep
their prompt files as records.

The refresh never fails a resume. A topology that no longer expands to the
run's graph, an unreadable ultrafuzz.toml or a run the engine still reports
active skips it with PROMPT_REFRESH_SKIPPED. Each prompt is validated as
`run` validates it and applied to all of its tasks or to none: one that does
not validate or render, that names an artifact authority a task was not
compiled with, or that shares a template copy with a prompt whose new text
differs keeps its tasks' files and is reported as PROMPT_REFRESH_REJECTED.
Every replaced file is copied first to prompt-history/<time>-<uuid>/, next
to a refresh.json listing each rewritten file with its old and new
SHA-256, and each file is replaced atomically; resume reports
PROMPTS_REFRESHED.

The record is refresh.json rather than an events.jsonl event: a new event
type changes event-record.schema.json, which the artifact schema-bundle
digest covers, and the engine still requires a run's launch bundle, so
every in-flight run would fail each task it resumes. The bundle digest is
unchanged.

The new [run] refresh_prompts_on_resume key (default true) turns it off.
Resume reads it from the project's current ultrafuzz.toml; applyRunConfig
drops it, so it never enters a run's resolved config, whose strict schema
older runs' sealed configs must keep matching.

Supporting changes: planning's static renderer can render in memory without
creating directories and collect per-attempt failures
(renderRunStaticPrompts); the runtime renderer takes a template body
(renderRuntimePromptsFromTemplates); effectiveTopologyPath selects the run's
topology file without effectiveAuditPolicy's prompt-validating load; the
lifecycle command gains a beforeContinuation hook with the inspected node
states and the resets it will issue.

BREAKING CHANGE: a hand edit of an unfinished task's prompt.rendered.md or of
a template copy is replaced by the next resume (the old bytes stay in
prompt-history/) unless run.refresh_prompts_on_resume = false, and edits to
.ultrafuzz/prompts/** now reach runs already in flight.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
One file that does not parse, such as invalid frontmatter, fails the
whole prompt catalog, and the error did not say which file it was: with
about 40 scaffolded prompts, `run`, `validate` and the new prompt
refresh on resume reported only "invalid prompt YAML frontmatter: ...".
Each per-file parse or variable error now starts with the prompt's
source and relative path ("project prompt setup/variant.md: ..."), and
a duplicate project prompt id names both files instead of the second.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e, and only files a later render reads

Review of the prompt refresh found records rewritten that the rest of
the resume could still refuse or withdraw, and a record of the refresh
that a failure could leave out.

Order. The refresh ran before the interrupted-withdrawal completion, the
retry-archive plan and the unsupported-run-ID refusal, so a resume that
one of them refused had already rewritten the prompts of finished tasks
it only predicted it would reset. It now runs after all of them and just
before the first timetravel; the completion of an earlier command's
withdrawal runs first, so the refresh sees the run as the resets will
find it. Its resets come from the same list the reset loop issues;
continuationResets, a copy of that logic, is gone. A timetravel that
fails after the refresh is the one window left, and the docs say so.

Withdrawn generations. A `--retry-failed` of a dynamic source rewrote
the finished children's and the join's prompts just before the
withdrawal moved them to dynamic-expansion-history/, so the archive no
longer held what they ran with. The hook now names the groups the
withdrawal plan covers; the refresh treats them as unexpanded, leaves
their generation's prompts as they are, and refreshes the template
copies the next expansion renders from.

Template copies. Every copy was rewritten, even one no render would read
again, which also broke replay and fork of runs launched before #1230
for groups that had finished. A copy is now rewritten only when a later
render reads it: its group has not expanded or is withdrawn, or an
unfinished task that renders from it has no published prompt. Prompts
that share a copy must still agree on it, counting every prompt bound to
the copy, and one pass settles that because a prompt has one copy.

Record. refresh.json is published before the first file changes, so
after a failure or a kill every file it lists holds its old or its new
bytes. PROMPT_REFRESH_INCOMPLETE says how many files were replaced
instead of claiming replacements that did not happen.

Topology check. It re-expanded the topology with the installed build and
diffed the graph through copies of planning internals. It now compares
the effective topology file with plan.json's launch digest, and the
topology's node set with the run's, which an eval run's exclusions
change. A comment-only topology edit now skips the refresh too.

Dynamic runtime. Admission's tamper check threw on a stock run before
its first render (launch writes dynamic_groups in compile order, the
derivation sorts them), which skipped every refresh. The refresh uses a
new derive-only mode instead.

Also: one shared test helper for the toml key; the hook's context type
is exported and reused; an unread failure field is dropped; the native
resume test turns the refresh off instead of relying on the fake run
being active; the CLI attach message says the attach applied none of the
project's current prompts.

Docs: restart-continue.md starts with pause, wait, edit, resume, and
covers the new rules, the reset window, catalog-wide errors, INCOMPLETE,
the launch digests and replay and fork of runs launched before #1230.
cli.md and artifacts-reports.md point there; the CHANGELOG entry,
#1230's edit sentence, security.md and SPECS.md match.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…orities refreshes on resume

Greptile flagged that the refresh publishes runtime prompts without the
compiled-selector check that static prompts get. Declined: every task whose
prompt is rendered at runtime (generated children and tasks that wait on a
dynamic group) is compiled with no artifact authority selectors, on main and
since v0.1.0 (#1234), and the engine's own render names whatever its template
names. The stock review prompts (dedupe-findings, aggregate-test-files and
final-report) already name authorities at launch, so that check would reject
every edit of them once their groups expand. For the flagged case, the
refreshed prompt is byte for byte what a launch of the edited prompt renders.

The new test publishes the stock review prompts the way the engine's first
render after goal-plan does. It checks that an unedited resume reports
nothing and that an edit of review/final-report.md is applied. It fails under
the proposed check in both forms: run on every render, the unedited resume
gets three PROMPT_REFRESH_REJECTED warnings; run only on changed bytes, the
edit is rejected. None of the existing refresh tests fail under either form.

The docs and the doc comments now say the authority check covers static
tasks, which is what it does, and why runtime prompts are not checked.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…reference cache

"resume of a stock run renders its prompts as launch did and refreshes an
unexpanded group's template copy" and "resume applies an edited stock
review prompt that names artifact authorities once its groups expanded"
launch the stock project, whose plan verifies and materializes every
reference the shipped catalog pins. Both read the host's reference cache,
so they passed only where ~/.cache/ultrafuzz/references was synced and
failed in CI with MISSING_CACHE for vulnerability-database.owasp-scs.

A withShippedReferenceCache helper seeds an XDG cache in the project with
writeShippedDocumentReferenceCaches and
writeShippedVulnerabilityDatabaseCache, as the clean-scaffold plan tests
do, and points XDG_CACHE_HOME at it for the launch. Launch is the only
step that reads the cache: with XDG_CACHE_HOME at an empty directory both
tests failed with MISSING_CACHE before this change and pass after it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/prompts-refresh-on-resume branch from 4bd75ad to 16bc398 Compare September 30, 2026 22:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant