Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion CHANGELOG.md

Large diffs are not rendered by default.

16 changes: 12 additions & 4 deletions docs/SPECS.md
Original file line number Diff line number Diff line change
Expand Up @@ -277,10 +277,17 @@ A static prompt MUST be rendered before workflow launch. A prompt that waits on
a dynamic group MUST be rendered when the group expands, from the run's own
template copy. Each rendered prompt MUST be stored as its attempt's
`artifacts/<attempt-id>/prompt.rendered.md`, and that file is the prompt every
later attempt of the task receives: it MUST be rendered only while it is
missing and MUST NOT be re-rendered, compared, or sealed afterwards. A prompt
that cannot be rendered at runtime MUST fail only its own task. Unknown
template variables MUST fail validation.
later attempt of the task receives: it MUST NOT be compared or sealed, and it
MUST be rendered again only while it is missing or when `resume` applies the
project's current prompts. `resume` MUST apply them, unless
`run.refresh_prompts_on_resume` is `false` in the project's current
`ultrafuzz.toml`, to every task that has not finished and to the template
copies a later render reads, after the checks that can refuse the resume and
before it resets or submits anything, and MUST NOT fail because of it: a prompt
that does not validate or render keeps its tasks' files, and a topology that
differs from the one the run launched with skips the refresh. A prompt that
cannot be rendered at runtime MUST fail only its own task. Unknown template
variables MUST fail validation.

The prompt variable set includes:

Expand Down Expand Up @@ -480,6 +487,7 @@ Before or at launch, each run MUST persist:
- `plan.json`
- `trusted-cli.json`, a run-owned trusted launcher, and content-addressed trusted CLI closures for schema-backed producers
- launch copies of static rendered prompts under `prompt-snapshots/`, used only to restore a missing static prompt
- the prompt files `resume` replaced when it applied the project's current prompts, with a `refresh.json` record, under `prompt-history/`
- per-node artifacts under `artifacts/`
- review artifacts under `review/`
- workspace metadata under `workspaces/`
Expand Down
4 changes: 2 additions & 2 deletions docs/how-to/edit-prompts-topology.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,8 +17,8 @@ workflow prompt and its role before choosing what to customize.
Runs use these project copies, and `ultrafuzz init` keeps them unless you pass
`--force`, which also overwrites `ultrafuzz.toml` and the topology. After an
upgrade a copy therefore keeps the text of the release that scaffolded it. A run
reads them only when it launches; to change a prompt of a run that has already
launched, see
renders them when it launches, and every `resume` applies their current text to
the tasks of the run that have not finished; see
[Change A Prompt Of A Running Campaign](restart-continue.md#change-a-prompt-of-a-running-campaign).
`ultrafuzz validate` warns about every project prompt that differs from the
built-in prompt at the same path. To take the built-in version of a prompt,
Expand Down
151 changes: 127 additions & 24 deletions docs/how-to/restart-continue.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,10 @@ only unfinished or newly rendered downstream tasks. Ultrafuzz control seals,
link journals, controller generations, graph fingerprints, schema bindings,
and metadata projections do not authorize continuation, so runs created before
those records existed can still reach Smithers. Historical artifact bytes and
embedded run IDs are never rewritten.
embedded run IDs are never rewritten. Before it resets anything or starts the
engine, resume applies the project's current prompts to the tasks that have not
finished; see
[Change A Prompt Of A Running Campaign](#change-a-prompt-of-a-running-campaign).

Use `--refresh-controller` when continuation should render the currently
installed controller and stock adapters. The new source is retained beside the
Expand Down Expand Up @@ -95,20 +98,111 @@ those nodes run again.

## Change A Prompt Of A Running Campaign

A run keeps its own prompts. Edits to `.ultrafuzz/prompts/**` apply to new runs
only. To change what a task of a launched run receives, edit that task's file
in the run and resume:
Edit the project's prompt under `.ultrafuzz/prompts/`, then resume the run. A
run that is still running must be paused first: while the workflow engine
reports the run active, `resume` only attaches to it and applies nothing. Wait
until `status` reports the run paused before you resume it:

```bash
ultrafuzz pause <run-id> --project /path/to/target-protocol
ultrafuzz status <run-id> --project /path/to/target-protocol
ultrafuzz resume <run-id> --project /path/to/target-protocol
```

Every `resume`, with or without `--refresh-controller`, `--retry-failed` or
`--reset-node`, applies the project's current prompts to every task of the run
that has not finished. It does so after the checks that can refuse the resume,
and before the resume resets anything or starts the engine. It selects each
prompt exactly as `ultrafuzz run` does, from `.ultrafuzz/prompts/**` and the
packaged built-ins, and renders it the way the run's launch did, with the run's
own config. A task has finished when the workflow engine reports its agent
`finished` and the resume does not reset it. The refresh reaches:

- static tasks that have not run, that failed or were interrupted, and those
that `--retry-failed` or `--reset-node` reruns, with the tasks that depend on
them;
- the rendered prompts of generated children and of later nodes such as the
final report, once their dynamic group has expanded;
- the template copies under `dynamic-prompt-templates/` that a later render
reads: the copy of a group that has not expanded, from which its children are
rendered, and the copy of a task whose prompt is not rendered yet, such as
the final report's before its groups expand.

A finished task keeps its `prompt.rendered.md`, the record of the prompt it
ran with. To rerun a finished task with the edited prompt, use
`resume --reset-node node:<attempt-id>`. A reset also reruns finished tasks
that merely started after the reset task; those keep their prompt. If the
workflow engine then fails to reset a task, the resume fails after the refresh,
so that task's file holds an edited prompt it did not run with; `refresh.json`
lists the file, and the bytes it ran with are beside it. An attempt that is
running keeps the prompt it started with.

When `resume --retry-failed` reruns the source of a dynamic group, the group's
generated children and the rendered prompts that wait on the group move to
`dynamic-expansion-history/` with the prompts they ran with. The refresh applies
the edits to the group's template copies instead, and the next expansion
renders from them.

Resume reports what the refresh did, and none of it fails the resume:

- `PROMPTS_REFRESHED` (info) names the tasks whose prompts it rewrote and
counts the template copies. Before it changes a file, it writes
`.ultrafuzz/runs/<run-id>/prompt-history/<time>-<id>/refresh.json`, which
lists every file it rewrites with its old and new SHA-256. It then copies
each file there, at the file's path in the run, and replaces it atomically.
- `PROMPT_REFRESH_REJECTED` (warning) names a prompt it did not apply, and why:
`ultrafuzz run` would reject it, it does not render for one of its tasks, it
names an artifact authority that one of its static tasks was not compiled
with, or it shares a template copy that a later render reads with another
prompt whose new text differs. A prompt is applied to every task rendered
from it or to none, and the other prompts still apply. Fix it and resume
again.
- `PROMPT_REFRESH_SKIPPED` (warning) means it applied nothing, and says why:
the effective topology file differs from the one the run launched with (any
edit counts, even to a comment; change the topology only for a new run), a
file under `.ultrafuzz/prompts/` breaks the whole prompt catalog as it would
break `ultrafuzz run`, such as invalid frontmatter or two files with the same
`id` (the warning names the file), `ultrafuzz.toml` cannot be read, or the
engine still reports the run active (`--force` or `--reset-node` on a running
run).
- `PROMPT_REFRESH_INCOMPLETE` (warning) means it stopped part way, for example
at a directory it cannot write, and says how many files it replaced. Every
file that `refresh.json` lists holds either its old bytes or its new ones.
Fix the cause and resume again to apply the rest.

A refresh changes none of the run's launch provenance: `prompt_digest` and the
other prompt digests keep their launch values, and `refresh.json` is the only
record of it.

A prompt that renders now can still fail later: a typo in an item variable,
such as `{{item.goal_promt}}`, fails only when the group expands, and then only
the child's task, at the `assert-task-inputs` step. Fix the project prompt and
run `resume --retry-failed`.

### Edit A Run's Own Prompt Files

To edit a run's prompt files by hand instead, turn the refresh off first, or
the next resume renders the project's prompt over your edit:

```toml
[run]
refresh_prompts_on_resume = false
```

Resume reads this key from the project's current `ultrafuzz.toml`, so it also
applies to runs already in flight. With the refresh off, every engine hands the
agent the run's own file as it is: `resume` with or without
`--refresh-controller`, `--retry-failed` or `--reset-node`, and `replay` and
`fork`, which never refresh prompts. Nothing re-renders it or compares it with
the launch render, so the edit neither strands the run nor stops `status` from
synchronizing.

```text
.ultrafuzz/runs/<run-id>/artifacts/<attempt-id>/prompt.rendered.md
```

Every engine hands the agent that file as it is: `resume` with or without
`--refresh-controller`, `--retry-failed` or `--reset-node`, and `replay` and
`fork`. Nothing re-renders it or compares it with the launch render, so the
edit neither strands the run nor stops `status` from synchronizing. Edit the
file in place and keep it a regular file: a symlink or a directory at that path
fails the task.
Edit the file in place and keep it a regular file: a symlink or a directory at
that path fails the task.

- A prompt that waits on a dynamic group, a generated child's or a later
node's such as the final report, has no file until the group expands. Before
Expand All @@ -129,13 +223,12 @@ fails the task.
lines. The agent needs them, and they are checked only when a prompt is
rendered, not when you edit it.
- A running attempt keeps the prompt it started with; the edit reaches the
task's next attempt. To rerun a finished task with an edited prompt, use
`resume --reset-node node:<attempt-id>`. This reset, like `--retry-failed`,
restarts the task's attempt numbers in the workflow engine, and each rerun
attempt replaces the engine's record of the earlier attempt with the same
number, including the prompt that attempt received. Run `ultrafuzz status`
before the reset so that `attempts.jsonl` records the finished attempt, and
keep a copy of the prompt file if you need to know what it received.
task's next attempt. `--reset-node` and `--retry-failed` restart the task's
attempt numbers in the workflow engine, and each rerun attempt replaces the
engine's record of the earlier attempt with the same number, including the
prompt that attempt received. Run `ultrafuzz status` before the reset so that
`attempts.jsonl` records the finished attempt, and keep a copy of the prompt
file if you need to know what it received.
- A deleted static prompt is restored from its launch copy in
`prompt-snapshots/` before the next engine starts, so it comes back without
your edit. A deleted runtime prompt is rendered again from its template copy.
Expand All @@ -144,19 +237,29 @@ fails the task.
own task, at the `assert-task-inputs` preparation step, with the cause. Fix
the file and run `resume --retry-failed`.

### Runs Launched By An Earlier Release

For a run launched by a release before this one, edit prompts only while the
run is stopped (`pause` it first), then resume it with this release, using
`resume --refresh-controller`. Until then, its original engine still reads
sealed copies of static prompts, so a static edit is lost, and still compares
runtime prompts, so a runtime edit stops the run. A plain `resume` continues
the run's launch workflow, which stops the whole run, instead of one task, when
a runtime prompt no longer renders; with `--refresh-controller` the run
continues on this release's workflow instead. `replay` and `fork` of such a run
keep running its launch snapshot, and so behave like its original engine. That
engine has also removed the prompt file of every static task it started: copy
the launch copy that `plan.json` names for the attempt
(`rendered_prompts[].rendered_prompt_snapshot_path`) to the attempt's
`prompt.rendered.md`, then edit it.
a runtime prompt no longer renders, including one rendered from a template copy
that the refresh rewrote; with `--refresh-controller` the run continues on this
release's workflow instead. `replay` and `fork` of such a run keep running its
launch snapshot, and so behave like its original engine: once a resume has
rewritten one of the run's runtime prompts or template copies, they stop at
their first render, with `runtime rendered prompt changed` or
`DYNAMIC_TEMPLATE_CHANGED`. To keep replay and fork of such a run working, set
`refresh_prompts_on_resume = false` before you resume it; to recover them, copy
each file that a `prompt-history/` entry lists back from the earliest entry.
That engine has also removed the prompt file of every static task it started. Resume renders the
file again for such a task that has not finished; for a finished one it
restores the launch copy that `plan.json` names for the attempt
(`rendered_prompts[].rendered_prompt_snapshot_path`). To edit one by hand with
the refresh off, copy that launch copy to the attempt's `prompt.rendered.md`,
then edit it.

## Replay A Linked Run

Expand Down
25 changes: 16 additions & 9 deletions docs/reference/artifacts-reports.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ usage.jsonl
attempts.jsonl
plan.json
prompt-snapshots/
prompt-history/
trusted-cli.json
trusted-bin/
artifacts/
Expand Down Expand Up @@ -141,13 +142,16 @@ every engine hands the agent that file: launch, `resume` (with or without
`fork`. A static prompt is rendered at plan time. A prompt that waits on a
dynamic group, a generated child's or a later node's such as the final report,
is rendered from the run's template copy under `dynamic-prompt-templates/` when
the group expands, and only while its file is missing. After that, no prompt
file is re-rendered, compared with a recorded digest, or sealed: an edited file
is what the task's next attempt receives, and an upgrade that renders templates
differently leaves published prompts as they are. A runtime prompt that cannot
be rendered, or a prompt file that is missing, unreadable or not a regular file,
fails only its task, at the `assert-task-inputs` preparation step, with the
cause.
the group expands, and only while its file is missing. No prompt file is
compared with a recorded digest or sealed, and an edited file is what the
task's next attempt receives, except that `resume` renders the files of
unfinished tasks again from the project's current prompts, unless
`run.refresh_prompts_on_resume = false`, and keeps every file it replaces under
`prompt-history/` (see
[Change A Prompt Of A Running Campaign](../how-to/restart-continue.md#change-a-prompt-of-a-running-campaign)).
A runtime prompt that cannot be rendered, or a prompt file that is missing,
unreadable or not a regular file, fails only its task, at the
`assert-task-inputs` preparation step, with the cause.

`plan.json` records the run plan, graph/config fingerprints, topology summary,
the launch render of each static prompt (its path and digest) and the path of
Expand All @@ -157,8 +161,11 @@ restores a missing static prompt from its copy, as it is; it never replaces a
prompt file that exists. That recovery concerns runtime-owned task input only;
it never reconstructs an agent-owned output. `prompt_digest` in `run.json` and
`plan.json`, like the final report's `run_metadata.prompt_digest`, is the
digest of the prompt catalog the run launched with; editing a run's prompt
files does not change it.
digest of the prompt catalog the run launched with. Neither editing a run's
prompt files nor a `resume` that applies the project's current prompts changes
it, an expansion manifest's `template.prompt_sha256` or `plan.json`'s
`rendered_prompts[].rendered_prompt_digest`: `prompt-history/*/refresh.json` is
the only record of such a refresh.

`trusted-cli.json` binds the run-owned launcher in `trusted-bin/` to the exact
CLI entrypoint bytes, validator build, and artifact schema-bundle digest. The
Expand Down
18 changes: 9 additions & 9 deletions docs/reference/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -382,15 +382,15 @@ Restart long-lived Ultrafuzz processes, such as the dashboard or
once that directory is gone their engine commands fail with
`restart this Ultrafuzz process`.

Prompt files are used as they are. Every task of a resumed run, with or without
`--refresh-controller`, `--retry-failed` or `--reset-node`, and of a `replay` or
`fork`, receives the run's own `artifacts/<attempt-id>/prompt.rendered.md`, so
an edit to that file reaches the task's next attempt. Before any of these
commands starts an engine, it restores a missing static prompt from its launch
copy in `prompt-snapshots/`; it never replaces anything at a prompt path. A
prompt that is still missing or cannot be read, or a runtime prompt that no
longer renders, fails only its task. Edits to `.ultrafuzz/prompts/**` apply to
new runs. See
Every task of a resumed run, with or without `--refresh-controller`,
`--retry-failed` or `--reset-node`, and of a `replay` or `fork`, receives the
run's own `artifacts/<attempt-id>/prompt.rendered.md`. Unless
`run.refresh_prompts_on_resume = false`, `resume` first renders that file again
from the project's current prompts for every task that has not finished;
`replay` and `fork` never do. Before any of these commands starts an engine, it
restores a missing static prompt from its launch copy in `prompt-snapshots/`,
never over an existing file. A prompt that is still missing or cannot be read,
or a runtime prompt that no longer renders, fails only its task. See
[Change a prompt of a running campaign](../how-to/restart-continue.md#change-a-prompt-of-a-running-campaign).

A run that ends `failed` with no failed durable node was stopped by something
Expand Down
13 changes: 13 additions & 0 deletions docs/reference/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ workspace_mode = "git-worktree"
default_timeout_seconds = 3600
workflow_deadline_seconds = 86400
controller_lease_seconds = 30
refresh_prompts_on_resume = true

[execution]
mode = "local"
Expand Down Expand Up @@ -124,9 +125,21 @@ empty path components, and dot components fail validation.
| `default_timeout_seconds` | integer | Default node timeout in seconds. The generated default is 3,600 (one hour). |
| `workflow_deadline_seconds` | integer | Workflow wall-time limit, checked only when a command syncs (see below). |
| `controller_lease_seconds` | integer | Lost-controller threshold used by the scoped renewable recovery supervisor. |
| `refresh_prompts_on_resume` | boolean | Re-render unfinished tasks' prompts on `resume`. Defaults to `true`. |

Other workspace modes are outside the product contract.

A run's configuration is frozen when it launches, with one exception:
`refresh_prompts_on_resume` is never part of it. Every `resume` reads the key
from the project's current `ultrafuzz.toml`, so setting it to `false` also
keeps the prompts of a run already in flight, and it appears in neither the
run's `config.resolved.toml` nor its `smithers/resolved-config.json`. When
`ultrafuzz.toml` cannot be read, resume keeps the run's prompts and warns. With
the key `true` or absent, resume re-renders the prompts of the run's unfinished
tasks from `.ultrafuzz/prompts/**` and the packaged built-ins before it starts
the engine; see
[Change A Prompt Of A Running Campaign](../how-to/restart-continue.md#change-a-prompt-of-a-running-campaign).

`max_parallel_agents` is the only concurrency limit the runtime enforces. It
bounds every task the workflow submits, not only agent tasks.

Expand Down
Loading
Loading