Skip to content

fix(runtime): name the runner error when a workflow fails with no failed node, and pin same-id recovery - #1172

Merged
aviggiano merged 6 commits into
mainfrom
claude/w15a-terminal-without-failed-node
Sep 29, 2026
Merged

aviggiano merged 6 commits into
mainfrom
claude/w15a-terminal-without-failed-node

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

A Smithers run can end failed without failing any task. The runner raises a run-level error, in practice WORKFLOW_RENDER_FAILED when the generated workflow throws while rendering, and durable work is left pending. For that run, Ultrafuzz's WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE diagnostic could only say failing workflow task(s): unreported (or name a wrapper task). It never said why the run stopped. Separately, the only regression for resuming this shape (resume continues a run-level render failure in place without a no-op rewind) used a fake runner, so nothing checked that a same-ID resume actually recovers it (#272).

Root cause

Change

  • smithers.ts: parseCurrentSmithersInspect returns an optional runError: { code?, message }, read loosely from data.run.error:
    • The message is cause.message, then the runner's summary, then message. The runner stores what was thrown as cause. summary is the runner's own message before it appends a docs link (a SmithersError default) and, for a failed run, raw smithers up … --resume true / smithers replay … commands (@smthrs/engine run-failure-recovery.js). Those commands bypass the attested operator controller, so they should not reach this diagnostic.
    • Both fields are redacted with redactSecretsInText, runner-name scrubbed like the module's other diagnostic text, and capped: code at 100 characters, message at 1,000.
    • Only a prefix eight times the kept length is redacted. The runner stores error text untruncated, and the redactor's cost grows with the square of one long token.
    • Any other shape yields undefined. Parsing this field never throws.
  • workflow-sync.ts: the WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE message appends ; workflow run error <code>: <message>, or ; workflow run error: <message> when the runner gives no code.
    • details is untouched, so the durable workflow-failure-unattributed event stays ids-only and needs no event schema change.
    • I also reworded the comment above it. It used to say the next resume "re-finalizes identically", which is not what the runner does.
  • New packages/runtime/test/smithers-terminal-resume.integration.test.ts. It uses the real pinned runner plus Ultrafuzz's own resumeRun({ force: true, retryFailed: true }), the ultrafuzz resume --force --retry-failed entry point, on the attested operator controller that production resumes use. The scenario:
    1. A producer finishes, then the render throws. The run is failed with WORKFLOW_RENDER_FAILED, the producer is finished@1 and the dependent is pending. The new reader returns exactly { code: "WORKFLOW_RENDER_FAILED", message: "synthetic run-level render failure" } from the runner's real error JSON.
    2. While the cause persists, resume is refused with WORKFLOW_LIFECYCLE_FAILED carrying the render error, and the parsed inspection is deep-equal to before.
    3. After the cause is removed, resume continues the same workflow run ID and the run succeeds. The producer is still at attempt 1 and was executed once; the dependent was executed once.
  • Deleted smithersSnapshotUnverifiedDependencies, an R43-era regex helper with no caller in src, and its test (separate commit).
  • docs/reference/cli.md: a short note on the two shapes this diagnostic covers (a run-level runner error, or a failed workflow task outside the durable graph) and how to recover from a run-level error. CHANGELOG entry added, plus a markdownlint-disable MD013 directive at the top of CHANGELOG.md so the super-linter check can pass (see Verification).

Deliberately not built (and why)

  • No rewind, fork or replacement-lineage recovery. feat: add strict producer-visible JSON contracts with release-lane fix #558 removed it and fix(runtime): recover terminal pending workflows on the real runner #861 was closed. Each re-renders the same program with the same inputs, so none can remove a render-time cause, and a same-ID resume already continues the run.
  • No "pending ready nodes but terminal failed" detector. Every accepted continuation already sets product status running. wait_reason: ready is Ultrafuzz's own projection, not runner evidence.
  • No error text in the durable event. That would need an event schema change and would break the existing ids-only rule for that record. The message is enough for an operator.
  • No stripping of the runner's resume advice from message. summary avoids it without text matching. Only a run error with neither cause nor summary, such as a plain Error, still carries it; removing it there would need an error-text matcher.
  • The runner-name scrub stays on this text. Operator output never names the runner, and tests pin that: cli.test.ts:1525 (status), lifecycle-commands.test.ts:53/397, lifecycle-inspection.test.ts:323, runtime.test.ts:13061. As a consequence, a thrown error that names a .smithers/... path shows .workflow runner/.... Changing that is a product decision for all diagnostic text, not this field alone.
  • Not added to PR smoke. The test takes 66–424 s here depending on host load. Most of that is the resume path itself: operatorControllerProjectSeal re-hashes the whole controller closure, about 40 s of CPU across two resumes (profiled), plus the controller's npm install.

Verification

Discriminating runs for the original change used an origin/main worktree (fbbcf6c) with this branch's tests copied in and compiled against main's src:

Test origin/main this branch
parseCurrentSmithersInspect reads the run-level error loosely and never fails on its shape (new) fails: actual: undefined, expected { code: 'WORKFLOW_RENDER_FAILED', message: 'boom' } passes
syncRun reports a typed diagnostic when a terminal workflow failure has no failed durable node (extended; fixture now uses the runner's coded SESSION_ERROR shape) fails: message is workflow run ended failed with no failed durable node; failing workflow task(s): ultrafuzz-agent-tasks passes
a run-level render failure resumes under the same id once its cause is gone, without re-running finished work (new, real runner) fails at the runError assertion on real runner output (actual: undefined) passes

The same integration test on main with only that one runError assertion removed passes (76 s). So the recovery half characterizes existing runner and resume behaviour, by design: it pins what #272 asked to be proven live. It doesn't claim a behaviour change.

Review follow-ups (8c832d8), each case run on its own against this PR's previous head (da3653e) compiled src and against the fix:

Case da3653e fix
no cause, runner summary present (DUPLICATE_ID) message ends See https://workflow runner.sh/reference/errors Resume with: workflow runner up w.tsx --run-id r --resume true Duplicate Task id detected: dependent
secret-shaped code returned verbatim <redacted>
500-character code 500 characters 100
200,000-character hex token as the message 7,023 ms for one parse 13 ms

The updated unit test fails on da3653e at the first new assertion (the DUPLICATE_ID summary) and passes on the fix. Its timing bound is 2 s, so it has about 150x headroom over the fixed parse and fails the unbounded one.

Why eight times and not four: with a 4,000-character prefix, a probe that put a 4096-bit RSA PEM key (3,272 characters) at offset 800 or 950 of a longer message leaked two of the key's base64 body lines into the kept 1,000 characters, because the key's END marker fell past the cut. With 8,000 the whole block is redacted; I checked offsets 0, 500, 800, 950 and 999 through the compiled reader.

Also ran, on the fix, with node --test and name patterns, all passing:

  • both parseCurrentSmithersInspect tests and pinned runner state and envelope contracts match Ultrafuzz's mirrors;
  • syncRun reports a typed diagnostic … (53 s) and syncRun records an unattributed terminal workflow failure durably and only once (39 s);
  • the fake-runner resume continues a run-level render failure in place without a no-op rewind (19 s), which remains the guard against no-op rewind subprocesses;
  • the real-runner integration test: 169 s at load average ~27 before a behaviour-neutral rename, and 123 s on the final source.

Run on the first commit and not rerun after the follow-ups: the other syncRun … unattributed … tests, getRunHealth reports a terminal product status while workflow health is live, the seven stopped reset … tests, and the CLI test status surfaces a terminal product and live workflow lifecycle divergence. Their fixtures carry no run error summary, no code and no long error text, so the follow-up does not change what they exercise.

Correction to an earlier claim in this description: the stopped-reset tests do not exercise runError inside assertStoppedResetAuthorityUnchanged, because their fake inspect has no run.error. That comparison stays sound by construction: both reads are the same inspect --format json --full-output of a stopped run whose error row does not change, parsed by pure functions. #1186 deletes that comparison.

Gates on the final source: prettier and eslint on the changed files; CI=1 ESLINT_PLUGIN_DIFF_COMMIT=fbbcf6c5 pnpm -w lint:strict:ci, diffed against this PR's base as CI does (against the current origin/main it also flags packages/cli/src/commands/report/bundle.ts, only because this branch predates #1161); pnpm --filter @ultrafuzz/runtime typecheck; pnpm -w knip; node scripts/docs-check.mjs.

Not run: the full runtime or CLI suites and the Bun adapter lane.

CI: on da3653e, "External static analysis" failed on super-linter MD013 (400-character lines) across CHANGELOG.md. Super-linter lints a changed file whole, and dozens of existing entries are longer than that. That failed release-gates, which skipped "Full release validation", so none of the runtime lanes, including the new test, ran in CI. 6af84e4 adds the same <!-- markdownlint-disable MD013 --> hunk as #1169, which passed that check with it. Locally, markdownlint-cli with only MD013 enabled, at 400, fails the previous CHANGELOG.md and passes the new one and docs/reference/cli.md. A trial merge with #1169's branch applies the directive hunk cleanly; the two entries under "Other changes" still conflict, as any two new entries there do.

Risk / compatibility

  • The diagnostic's code, severity and details are unchanged. Its message can now be up to about 1,100 characters longer and can contain multi-line text from the thrown error. The only existing consumers I found match the code or a message prefix (cli.test.ts, cli-contracts.test.ts).
  • CurrentSmithersInspect gains an optional field. It is part of the stopped-reset deep comparison of two parsed inspections (see the correction above).
  • A parse of a failed run now redacts at most 8,800 characters of error text (8,000 of the message, 800 of the code). In the worst case measured here, one unbroken token, that costs tens of milliseconds.
  • No schema, validator build identity or contract description is touched. The generated workflow.tsx imports nothing that changed.
  • The new test lands in runtime-supporting, which is pull_request: true in scripts/ci/release-validation-lanes.mjs, so it runs on every PR. It provisions the operator controller with a real npm install from registry.npmjs.org (5-minute timeout), as the sibling native-continuation test in that lane already does, so the lane gains one more networked install. One local run on this heavily loaded host failed that way earlier (the error names the npm install step); a rerun passed.
  • Each run of the test leaves one retained ~595 MB operator controller in os.tmpdir(). That is production behaviour: the controller is retained for the detached runner, and the OS tmp policy is the stated reclamation boundary. The sibling test does the same.

Closing #272

I recommend this closes #272. The 08-25 reopen describes a resume that is accepted and then re-fails at once, and asks for rewind or replacement-lineage recovery. The new test pins the refused variant, not the accepted one. The analysis behind this PR reproduced the accepted variant only with a cause that renders see once live task state is present, and that cause still persisted. A rewind or replacement lineage re-renders the same program with the same inputs, so neither can remove such a cause. What recovers the run is removing the cause and resuming under the same ID, which the test now proves on the real runner, and this PR makes the cause visible in status. If you want the accepted variant pinned before closing, change this to Refs #272.

Closes #272

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no outstanding finding or new blocking issue remains.

Summary

The PR adds a bounded, redacted run-level error to the diagnostic for a failed workflow with no failed durable node, and tests same-ID recovery against the pinned runner.

  • The durable event payload remains unchanged.
  • The previous finding about unredacted error codes is fixed.

Reviews (3) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

aviggiano and others added 2 commits September 28, 2026 23:12
…encies helper

The helper regex-scanned a Smithers inspect snapshot for the generated
workflow's "artifact dependency has not passed verification" message so a
resume could name the dependency behind a failed `prepare:` wrapper (the R43
shape behind #272). Nothing in src calls it: prepare-wrapper failures are now
attributed to their durable node (#288), and the lens sanitizer that caused
R43 was deleted (#558). Its only caller was its own unit test, which goes too.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…led node, and pin same-id recovery

A Smithers run can end `failed` without failing any task. That is a run-level
error such as WORKFLOW_RENDER_FAILED, which the runner raises when the
generated workflow throws while rendering, and it leaves durable work pending.
WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE could then only say "failing workflow
task(s): unreported", because parseCurrentSmithersInspect accepted
data.run.error and dropped it, and that error is the only record of why the
run stopped.

parseCurrentSmithersInspect now returns an optional runError read loosely from
data.run.error. It takes the cause's message (the runner records what was
thrown as `cause` under its own summary and `smithers up` recovery advice),
falling back to the message, redacted and scrubbed like other runner text and
capped at 1,000 characters. An unexpected shape yields nothing and never fails
the parse. The WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE message appends
"; workflow run error <code>: <message>". Its details, and so the durable
workflow-failure-unattributed event payload, are unchanged: that record stays
ids-only and needs no schema change.

A new integration test drives the pinned runner and Ultrafuzz's own resumeRun
(the `resume --force --retry-failed` entry) through this shape. While the
render-time cause persists, the resume is refused with
WORKFLOW_LIFECYCLE_FAILED and the run is left exactly as it was. Once the
cause is removed, the same run ID resumes and runs only the pending task, and
the finished producer is not run again. The test also checks the new reader
against the runner's real error JSON. No rewind, fork or replacement lineage
is added: each would re-render the same program and inputs and hit the same
cause (#272).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 28, 2026 23:42
Comment thread packages/runtime/src/smithers.ts Outdated
aviggiano and others added 2 commits September 29, 2026 02:37
…action

Review follow-ups on the run-level error reader added for #272.

- Prefer the runner's structured `summary` over `message` when the run
  error has no `cause`. The runner appends a docs link to every
  SmithersError message and, for a failed run, raw `smithers up ...
  --resume true` / `smithers replay ...` commands. Those commands bypass
  the attested operator controller and contradict the documented
  "fix the cause, then resume" recovery. `summary` carries neither, so
  no text matching is needed. A plain Error without a cause (no
  `summary`) still falls back to `message`.
- Redact only a prefix eight times the kept length. The runner stores
  error text untruncated and the redactor's cost grows with the square
  of one long token, so a 200,000-character hex token took about 7 s per
  parse, paid by every sync, status and resume of a failed run. It now
  takes about 13 ms. Eight times (not four) keeps a 4096-bit RSA PEM
  block that starts in the kept 1,000 characters whole: with a
  4,000-character prefix a probe leaked two of its body lines.
- Apply the same redaction, runner-name scrub and a 100-character cap
  to `code`, which was copied verbatim.
- Reword the reader comment (it only guarantees the parse never fails)
  and the CLI reference, which now names both shapes the diagnostic
  covers and scopes the resume advice to run-level errors.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Super-linter lints a changed Markdown file whole, and dozens of existing
CHANGELOG.md entries are single lines longer than its 400-character
MD013 limit. So a PR that adds an entry fails "External static
analysis", which fails `release-gates` and skips every release
validation lane, including the runtime lanes this PR's new test runs in.

This is the same two-line hunk #1169 adds, so whichever merges second
applies it cleanly.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
Since the external static-analysis gate was adopted (#1005), super-linter
lints every changed Markdown file with MD013 at 400 characters.
CHANGELOG.md keeps one entry per line, and markdownlint reports 42 existing
entries on main as longer than that, so any change to the file fails
`External static analysis`.
That also fails `release-gates` and skips the release validation lanes on
the pull request. No commit has changed CHANGELOG.md since the gate
landed. Disable MD013 below the title of this file only, using the same
directive line that #1169, #1172 and #1181 add, so the branches do not
end up with two spellings of it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
#1181, #1172 and #1169 add `<!-- markdownlint-disable MD013 -->` right
after the changelog title. Adding the identical hunk here instead of a
disable-file comment at the end of the file means a merge of any two of
these PRs keeps one directive rather than two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 2 commits September 29, 2026 03:55
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 1a6d716 into main Sep 29, 2026
13 checks passed
@aviggiano
aviggiano deleted the claude/w15a-terminal-without-failed-node branch September 29, 2026 06:33
aviggiano added a commit that referenced this pull request Sep 29, 2026
…r-render cost drops (#1168)

* fix(runtime): literal braces in prompt text no longer fail the workflow render

renderAgentPrompt substituted the operator and task prompts into the
trusted agent-prompt template and then regex-scanned the whole result for
leftover `{{word}}` placeholders. Inserted text that merely contained such
a sequence threw inside the Smithers render: an escaped `\{{word}}` example
in a prompt template (the prompt renderer emits it as literal `{{word}}`),
the same escape inside a dynamic item value (kept verbatim), or an
operator `--prompt` note. Every render builds every task's prompt, so the
run failed, and failed again on every resume.

Check only the template's own placeholders, inside the replace callback,
as renderAgentPreambleTemplate already does. A template placeholder with
no value still throws, and now names the placeholder.

The dynamic prompt renderer also resolves `{{...}}` inside goal-plan
replacement values, and an unbound name there throws in the same render.
The goal-plan contract now rejects `{{` in replacement values, which are
plain-text labels, so that failure lands on goal-plan's own verify instead
of on every render. The check is a zod refinement; goal-plan.schema.json
is unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): run the JSON validator preflight once per engine process

prepareArtifactMirror spawned `ultrafuzz json validate` against the smoke
fixture in every prepare, in every agent-attempt reset, when an agent is
built after a controller restart, and again in the zero-retry verify task.
The trusted launcher's cold start was measured at ~35 s under contention
(#1026), so each spawn was another chance to fail an attempt, including
one whose agent work had already finished.

What the spawn proves -- that this process can launch the agent-facing
validator -- does not depend on the task: materializePromptSchemas has
just digest-checked the workspace's schema copy, and start-run already
runs the trusted-launcher preflight at launch and resume. Remember the
first success in the engine process. A failure is not remembered, so the
next caller spawns again.

Co-Authored-By: Claude Opus 5.5 <[email protected]>


* perf(runtime): cut per-process and per-render cost of the generated workflow

agent-registry.ts and smithers.ts imported the TypeScript compiler at
module load, and both are reachable from @ultrafuzz/runtime's index. Only
registry analysis (validate and init) and the controller-refresh helper
that only tests call ever use it. Require it at first use instead.

The saving lands in processes that import the runtime without smthrs:
the ultrafuzz CLI, including every agent `ultrafuzz json validate` call.
Smithers engine processes still load the compiler: the generated
workflow imports smthrs, whose index re-exports @smthrs/scorers, and two
scorer modules import typescript at module load. Importing the built
runtime index alone on this host went from 528-547 ms / 248-250 MB RSS to
409-437 ms / 196-198 MB under Node 24, and from 449-483 ms / 252-261 MB
to 336-351 ms / 209-213 MB under Bun 1.3.14 (5 runs each).

Changing the parameter type on topLevelNameIsBound's signature makes the
strict diff lint report that function's existing complexity, so its
import-clause check moves into a helper with the same conditions.

materializeDynamicRuntime runs on every render and rewrote tasks.json and
graph.json, with fsyncs, even when it re-derived the bytes already on
disk. Skip the write in that case.

taskSpecsFromCompiled looked every task up with serializedTaskSpecs.find;
index the specs by id once per call instead. At the default topology's 65
static tasks the difference is not measurable.

Add a test that compiles the packaged default topology and holds the
generated workflow under 2.3 MB, measured with a fixed-length project root
(about 2.06-2.07 MB today, depending on the checkout path), so the size
#1146 describes cannot grow silently.

Refs #1146

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: changelog for render robustness and runtime footprint

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: exempt CHANGELOG.md from the markdown line-length rule

Since the external static-analysis gate was adopted (#1005), super-linter
lints every changed Markdown file with MD013 at 400 characters.
CHANGELOG.md keeps one entry per line, and markdownlint reports 42 existing
entries on main as longer than that, so any change to the file fails
`External static analysis`.
That also fails `release-gates` and skips the release validation lanes on
the pull request. No commit has changed CHANGELOG.md since the gate
landed. Disable MD013 below the title of this file only, using the same
directive line that #1169, #1172 and #1181 add, so the branches do not
end up with two spellings of it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): the post-agent verify pass no longer preflights the validator CLI

finalizeAndVerifyArtifacts calls prepareArtifactMirror with
pinnedSubmodules "verify", and that pass still ran the
`ultrafuzz json validate` preflight. Verify checks the finished agent's
outputs in-process and never runs the agent-facing CLI, and its task has
no retry. The per-process memo only helps later callers: after an engine
restart, the first preflight in the new process can be the verify of an
attempt whose agent had already finished, and a CLI cold start that timed
out there (the trusted launcher measured ~35 s under contention in #1026)
failed verification and discarded that work.

Skip the preflight in the verify pass. prepare:* and the reset before each
agent attempt, which run ahead of an agent that uses the CLI, keep it.

The new lifecycle test runs the production preparation body with the
options finalizeAndVerifyArtifacts passes. Without this change it counts
three preflights instead of two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: state where the TypeScript saving and the goal-plan brace rule apply

The comments, the footprint test and the changelog said Smithers engine
processes stop loading the TypeScript compiler. They do not: the
generated workflow imports smthrs, whose index re-exports
@smthrs/scorers, and its sideEffectAnalysis and workflowUiCompliance
modules import typescript at module load. The saving lands in ultrafuzz
CLI processes, including every agent `ultrafuzz json validate` call.

The prompt-variables reference and the changelog also said the
goal-plan@1 contract rejects `{{` in replacement values. Only goal-plan's
own verification (and the evals goal-plan reader) run that zod rule;
`ultrafuzz json validate` and `ultrafuzz artifact validate` accept such a
plan. Say so, replace the exact generated-workflow byte count, which does
not reproduce across environments, with a range, and describe the
preflight as running until its first success.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* test(runtime): cover imported Object bindings in registry shadowing

importClauseBindsName, split out of topLevelNameIsBound on this branch,
had no test: the registry-shadowing test only declared Object locally.
Add default, namespace, named, aliased and default-plus-named imports of
Object, each of which must stop Object.freeze from being read as the
global, and an unrelated named import that must not. Removing any one of
the helper's three checks, or making it match every named import, fails
the test.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…nused agent profiles (#1179)

* perf(runtime): write execution-snapshot files with their final mode and flush each once

Snapshot publication wrote every file with mode 0400 and fsynced it, then
the permission seal reopened every file, fchmodded it (to 0500 for
executables, otherwise to the 0400 it already had) and fsynced it again.
A launch of the test fixture's 12,552-file snapshot therefore issued
27,876 fsyncs, two per file plus two per directory, and launch time is
fsync-bound.

writeSnapshotFile now creates each file with its final mode and sets that
mode through the creating descriptor (the umask can clear creation bits)
before its single flush, so the bytes and the mode are durable together,
as the second flush previously made them. The permission seal now only
makes directories read-only; the publication verification that already
follows it still checks every entry's type, mode, link count and bytes.
The same fixture now issues 15,324 fsyncs (one per file plus two per
directory).

The new test republishes a launched run's snapshot and bounds its fsync
calls at one per file plus two per directory; origin/main fails it with
27,876 flushes for 12,552 files in 1,386 directories.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): doctor requires only the agent CLIs a run can dispatch to

doctor required the executable of every configured model profile, so a
fresh `ultrafuzz init` on a Codex-only host reported "needs attention"
because the scaffold's opt-in Claude, Kimi and Pi profiles were not
installed, although no node would run them. It now requires an agent's
CLI only when the selected topology or its [retry] agents fallback chain
can dispatch to that agent (the same selection the OpenRouter credential
check already used). Other configured profiles' CLIs are still probed and
listed, as not required, and the human output marks them "(not
required)". Agents that share a CLI (Codex and OpenRouter, Claude and
DeepSeek) require it if any of them is selected.

doctor also built DOCTOR_WORKFLOW_ENGINE_* diagnostics and check statuses
for the project-local engine and then discarded them, reporting both
checks as "unknown". The builders are deleted; layout_status and the
per-patch posture are still reported, and the reference docs no longer
list the five codes that were never emitted.

Launch and resume install the workflow engine controller under the OS
temporary directory, and a native resume keeps its install there for the
detached engine (#921 cost 1: a tmpfs /tmp exhausted a host's memory). A
new temporary-directory check warns, without failing the verdict, when
that directory is a tmpfs or has under 2 GiB free, and reports how many
ultrafuzz-controller-* directories it holds and their total size. It
never removes them, because a live engine may still use them.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): keep refreshed controllers out of the governed target identity

`resume --refresh-controller` renders its current controller into
<project>/.smithers/continuations/<uuid>/, but that directory was missing
from controllerOwnedGovernancePaths. Unless the project ignores .smithers,
the rendered files are untracked target files: the next launch in that
project records the target as dirty, which a private campaign (the
default policy) and every cloud launch reject, and the changed worktree
digest invalidates disclosure acknowledgements computed for the target.

The directory is now controller-owned, like .smithers/workflows. The new
test renders files in that layout and checks the target identity is
unchanged; origin/main reports the target dirty.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): parse the control seal with an item limit that covers its schema

Runtime documents are parsed with a 100,000-item limit that the strict
JSON parser counts across the whole document. The control seal's schema
admits 100,000 execution files plus three identity arrays of up to
100,000 entries each, and launch writes the seal after schema validation
alone. A closure near the execution-file bound plus the run's task
bindings would therefore produce a seal that every later command fails
to parse. (An earlier measurement put the production closure at about
51,000 files; it was not re-measured for this change.)

The seal is now parsed with a 400,000-item limit, the schema's total; its
property total (four per execution file plus 35 fixed) already fits the
default limit. Byte limits are unchanged. The new test parses a seal with
the schema's maximum execution files and the fixture's bindings;
origin/main rejects it with "JSON exceeds the item limit of 100000".

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: changelog entry for launch I/O and doctor fixes

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: disable only MD013 for the one-entry-per-line changelog

Super-linter's markdownlint (MD013, 400 columns) validates every changed
Markdown file in full. The gate landed (#1005) after CHANGELOG.md was last
edited, and nearly every entry is one line longer than 400 columns, so
any pull request that adds a changelog entry fails "External static
analysis" on the file's existing lines. Line length is the only rule the
file breaks under super-linter's configuration, so disable just MD013 for
this file and keep every other Markdown rule.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): count only required commands in doctor's toolchain summary

Doctor now lists the CLIs of configured but unselected agents with
required: false, so the toolchain summary's "N required commands
available" counted them too. On a fresh init with only Codex installed it
reported 7 required commands when 4 are required.

The selection test also gains the case the reviewers found missing: a
selected Codex profile beside an unused OpenRouter profile keeps codex
required. Because configured agent refs are sorted, the existing
OpenRouter-only case could not tell "any requirer wins" from "last
writer wins"; this case fails under the latter.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* refactor(runtime): parse every runtime document with one 400,000-item limit

The previous commit gave the control seal its own limits constant, a
schema-ID import and a ternary. Raising the single shared item limit is
the same fix in one number: runtime documents stay bounded by the
unchanged 128 MiB parse cap (and the seal by the 64 MiB control-file read
cap), and 400,000 is what a seal whose arrays are all at their schema
bounds holds (100,000 execution files plus three identity arrays of
100,000), with 400,035 properties under the unchanged 500,000 limit.

The seal test's comment now claims only what it fills: execution_files at
its bound. Filling the three identity arrays too would pin the 400,000
figure, but their schema validation grows quadratically (uniqueItems on
$ref items): with 30,000 identities per array it alone took 6.1 s, and a
test with every array at its bound ran for 85 s, so the test does not.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: use the changelog MD013 directive sibling PRs add

#1181, #1172 and #1169 add `<!-- markdownlint-disable MD013 -->` right
after the changelog title. Adding the identical hunk here instead of a
disable-file comment at the end of the file means a merge of any two of
these PRs keeps one directive rather than two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* test(runtime): make the snapshot tests exercise the phases they name

- "publishing an execution snapshot flushes each file once" now counts
  only fsyncs of regular-file descriptors and asserts that count equals
  the published file count (12,552 on the fixture; main flushes 25,104).
  The old total bound also encoded how directories are flushed.
- The leaf-swap test's hook now fires on the write-time fchmod, since the
  seal no longer opens files, so it is renamed after what it checks.
- The parent-swap test's hook fires only on a directory fchmod, so the
  swap lands in the directory seal after every file is written, as it did
  on main. Without that, no test swapped anything while directories were
  being sealed.
- The seal's doc comment now says what publication verification checks
  for each entry type instead of "every entry's type, mode, link count,
  and bytes".

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Resume terminal failed workflows with pending ready nodes

1 participant