Skip to content

perf(runtime): halve launch fsync work, and doctor stops failing on unused agent profiles - #1179

Merged
aviggiano merged 11 commits into
mainfrom
claude/w16-launch-io-hygiene
Sep 29, 2026
Merged

aviggiano merged 11 commits into
mainfrom
claude/w16-launch-io-hygiene

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Four launch and lifecycle hygiene problems that came up while costing #921:

  1. Launch time goes to fsync. Publishing the sealed execution snapshot fsyncs every file twice. On the runtime test fixture (12,552 files in 1,386 directories), one publication issues 27,876 fsync calls, 25,104 of them on files. Where fsync is slow, as on the loaded host used here, that time dominates (see Verification). An earlier measurement put a production closure at about 51,000 files; I did not re-measure it.
  2. doctor fails on agent profiles no node uses. It requires the CLI of every configured model profile. On a fresh ultrafuzz init on a Codex-only host, the verdict is "needs attention" only because the scaffold's opt-in Claude, Kimi and Pi profiles are not installed. doctor also builds DOCTOR_WORKFLOW_ENGINE_* diagnostics and then throws them away. It says nothing about the temporary directory, even though launch and resume install the workflow engine controller there and native resumes leave theirs behind (Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921 cost 1: a tmpfs /tmp exhausted a host's memory).
  3. resume --refresh-controller dirties the governed target. It renders into <project>/.smithers/continuations/<uuid>/, and that path is not controller-owned for governance. Unless the project ignores .smithers, the next launch sees untracked target files, and a private campaign (the default policy) or any cloud launch rejects a dirty target.
  4. The control seal can be written but not read back. Runtime documents are parsed with a 100,000-item limit counted across the whole document. The seal's schema admits 100,000 execution files plus three identity arrays of up to 100,000 entries each, and launch writes the seal after schema validation alone.

Root cause

  1. writeSnapshotFile created files 0400 and fsynced them. sealSnapshotDirectoryPermissions then reopened every file, fchmoded it (to 0500 for executables, and to the 0400 it already had for everything else) and fsynced it again.
  2. diagnoseProject built the command list from configuredAgentRefs(models.profiles), which is every configured profile, and ignored the topology selection it already computed for the OpenRouter credential check. workflowEngineCheck and compatibilityPatchCheck ran, but only their check names were used.
  3. controllerOwnedGovernancePaths lists .smithers/workflows and .smithers/node_modules but not .smithers/continuations.
  4. parseRuntimeDocumentBytes parsed every runtime document with a 100,000-item limit, below the 400,000 items a seal at its schema's array bounds holds.

Change

  • Snapshot publication (workflow-integrity.ts, net −26 lines). Each file is created with its final mode (0500 for executables, otherwise 0400). The mode is also set through the creating descriptor, because the umask can clear creation bits, before the file's single flush, so bytes and mode become durable together. The permission seal now only makes directories read-only. The publication verification that already runs right after it still checks each file's mode, link count and bytes, each link's target, that every directory is read-only, and that no entry is missing or unexpected. This removes the second open, fchmod and fsync per file: per-file flushes are halved exactly, and total fsyncs per publication drop by 45%.
  • doctor (doctor.ts, CLI renderer):
    • An agent's CLI is required only if the selected topology or its [retry] agents chain can dispatch to that agent (activeTopologyAgentRefs, the same selection the credential check uses). Other configured profiles' CLIs are still probed and listed with required: false, and the human output marks them (not required). Agents that share a CLI (Codex/OpenRouter, Claude/DeepSeek) require it when any of them is selected. The toolchain summary counts only the required commands.
    • The two discarded check builders are deleted. layout_status and per-patch posture are still reported, and the reference docs drop the five codes that were never emitted.
    • A new temporary-directory check warns (it never fails the verdict) when os.tmpdir() is a tmpfs or has under 2 GiB free. It reports the free space and how many ultrafuzz-controller-* directories the temporary directory holds, with their total size, and never deletes them. The warning's diagnostic code is DOCTOR_TEMPORARY_DIRECTORY_CONSTRAINED.
  • Governance (data-governance.ts). <project>/.smithers/continuations is now controller-owned.
  • Runtime document parsing (runtime-document-codec.ts, one number). parseRuntimeDocumentBytes now parses every runtime document with a 400,000-item limit instead of 100,000: the item count of a control seal whose arrays are all at their schema bounds. The seal's property total at those bounds (400,035) was already under the unchanged 500,000 limit, and byte limits are unchanged. The first version gave the seal its own limits; review pointed out that one shared number is the same fix with less code, and it also covers the execution-dependency map. Launch writes that map the same way, after assertRuntimeDocument alone (smithers.ts), and reads it back through this parser, while its schema admits 320,001 items. A map with more than 100,000 items in total (for example 100,001 issuers, within the schema's 110,001) was writable but unreadable under the old limit.
  • Docs: docs/reference/cli.md (Doctor section), docs/config.md (OpenCode note) and docs/tutorials/first-campaign.md now describe the selection rule and the temporary-directory check. There is also a CHANGELOG entry.
  • Changelog lint. Super-linter's MD013 (400 columns) checks every changed Markdown file in full. The gate landed (chore: adopt external static-analysis gates #1005) after the changelog was last edited, and nearly every entry is a single line longer than 400 columns, so any PR that adds a changelog entry fails "External static analysis" (this PR's first push failed with 43 MD013 errors, almost all on lines it did not touch). CHANGELOG.md now has <!-- markdownlint-disable MD013 --> right after its title. That is the same hunk refactor(artifacts): delete never-dispatched gates, production-dead code, and the unread event index #1181, fix(runtime): name the runner error when a workflow fails with no failed node, and pin same-id recovery #1172 and fix(runtime): resume --retry-failed reopens descendants skipped behind a recovered producer #1169 add, so merging any of them together leaves one directive. The implementer checked locally with super-linter's v8.7.0 markdownlint config that line length is the only rule the file breaks, and "External static analysis" passes on this head with the directive in its new place.

Deliberately not built (and why)

  • Per-file fsync is not removed entirely. The earlier analysis measured that as the larger win, but without it a crash mid-publication needs a "republish on verification failure" path to stay safe. This change keeps publication's durability exactly as it was.
  • Snapshots are not retired, and node_modules is still copied. That is Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921's open design decision, not a hygiene fix.
  • doctor does not clean up controller roots. A native resume keeps its root for the detached engine, so deleting roots safely would need liveness tracking.
  • The controller-root size scan is not bounded. It walks each ultrafuzz-controller-* directory. On this host the reviewers measured 1.3 s (warm cache) for 20 retained roots (495,000 entries, 5.0 GiB) and 1.6 s for 23 roots (585,000 entries, 5.9 GiB). The scan grows with the roots it reports, and I left it as is.
  • No write-side size check for the seal. The seal file is still read under the 64 MiB workflow-control file cap. On the test fixture, the seal for 12,551 execution files measured 5.3 MB in the implementer's run and 5.5 MB in a reviewer's (426 to 440 bytes per entry, with long worktree source paths). Extrapolated to the 100,000-entry schema bound, that is about 43 to 44 MB, so only unusually long paths would exceed the cap. I have not seen that happen, and a guard would be one more fail-closed layer.
  • The seal test does not pin 400,000. It fills execution_files to its bound (100,005 items in all), which fails on main but would pass with any limit of at least 100,005. Filling the three identity arrays as well would pin the figure, but that test ran for 85 s (see the validation-cost risk below), so the 400,000 figure rests on the schema arithmetic in the code comment.
  • The execution-dependency map gets the shared limit and nothing more. Its schema's array total (10,000 + 100,000 + 110,001 + 100,000 = 320,001 items) now fits. Its property total is not bounded by the schema, because issuer dependencies objects have no maxProperties. I did not measure a real map, and I did not audit the other runtime schemas' totals.
  • Continuations are still rendered where they were. Moving them under the run root would change paths Smithers records and the adapters' relative imports. Excluding the directory is a one-line change.
  • The fsync test keeps its own launch instead of sharing the adjacent recovery test's, so each test checks one property. It costs one more startRun (the test took 13 s here under eatmydata).
  • The five doctor tests that start a run first are not converted to init-only fixtures. They do not need a launch, and converting them would save CI time, but that is out of scope.

Verification

Each new test below fails on origin/main (b6dd1da9) and passes on this branch. For the main runs I copied this branch's test files into a main worktree and compiled them against main's sources.

Test origin/main this branch
publishing an execution snapshot flushes each file once (runtime.test.ts; counts fsyncs of regular-file descriptors) fails: 25104 file flushes for 12552 files (27876 flushes in all) passes: 12552 file flushes for 12552 files (15324 flushes in all)
a control seal at its schema's execution-file bound parses back fails: StrictJsonError: JSON exceeds the item limit of 100000 passes (~0.5 s)
a refreshed controller rendered into the target leaves its governed identity unchanged fails: dirty: true, and the worktree digest differs passes
diagnoseProject requires only the CLIs of agents the selected topology can dispatch to fails: toolchain 'error' !== 'ok' passes; also checks a [retry] agents fallback, the toolchain summary count, and codex staying required in both directions (Codex selected beside an unused OpenRouter profile, and OpenRouter selected beside an unused Codex profile)
diagnoseProject reports controller roots in the temporary directory and leaves them in place fails: no temporary-directory check (empty summary) passes
diagnoseProject warns, without failing, when the temporary directory is RAM-backed (real tmpfs at /dev/shm; skipped where that is not tmpfs) fails: check status undefined, expected warning (ran, not skipped) passes (ran, not skipped)
diagnoseProject warns when the temporary directory has little free space (fs.statfsSync mocked) fails: check status undefined, expected warning passes
CLI doctor reports install posture in human and JSON output (adds a (not required) assertion for kimi) fails: no - kimi: missing from execution environment (not required) line passes

Mutation checks on this branch's compiled doctor.js, each run against the selection test:

  • Restoring the old summary line (${toolchain.length} required commands) fails it: '7 required commands available in the configured execution environment' does not match /^4 required commands available /.
  • Making the shared-CLI merge last-writer-wins (requiredByName.set(name, required)) fails it at the new "Codex beside an unused OpenRouter profile" assertion. Review found that the earlier version of the test passed under this mutant.

Also run on this branch, all passing:

  • All 10 snapshot tests (the snapshot recovery, materialization, publication and writes tests plus the fsync test). The leaf-swap test is renamed to snapshot writes reject a swapped leaf without touching the outside file, since its hook now fires on the write-time fchmod. The parent-swap test's hook now fires only on a directory fchmod, so the swap lands in the directory seal after every file is written, as it did on main. I confirmed that from its diagnostic, an ENOENT from lstat of the temporary snapshot's .smithers/agents directory, which the seal checks right after its fchmod. Both swap tests also pass on main.
  • All 16 diagnoseProject tests, including the five that start a run.
  • The target identity tests in data-governance.test.ts, and all 5 runtime document contract tests.
  • prettier --check and eslint on the changed files, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, pnpm --filter typecheck for runtime and cli, pnpm -w knip, and node scripts/docs-check.mjs.

Test runs used eatmydata except where noted.

Wall-clock, from the implementer's runs without eatmydata on this shared host (about 25 other agents' test runs, load average 22 to 41): republishing the fixture snapshot took 2,008 s on origin/main for 27,876 fsyncs and 855 s on this branch for 15,324. The startRun launches inside those two tests took about 21 minutes each, but under different host load, so they are not a controlled comparison. A reviewer's micro-benchmark under eatmydata (6,000 files, 8 alternating samples each) had medians of about 480 ms on main and 320 ms on this branch. On GitHub runners, CI shard 1 logged the fsync test's republish at 17,269 ms on the earlier push and at 10,557 ms on this one (12552 file flushes for 12552 files (15324 flushes in all, 10557 ms)). The release-validation lanes took 23 to 31 minutes on the earlier push and 24.5 to 34 minutes on this one, about the same as the lanes of #1161 and #1157 (23 to 30 minutes). I found no measurable CI speedup on GitHub's runners and do not claim one.

CI on d1cc424c is all green: the build and Node 24 smoke, External static analysis, the five release-validation lanes (runtime.test.ts shards 1 to 4 plus the supporting test files and the Bun adapter contracts) and release-gates.

Not run: the full CLI suite (CI runs it only on pushes to main); locally I ran only the CLI doctor test.

Risk / compatibility

  • Snapshots. Only new publications change. Already-published generations and their verification are untouched. A umask that masks creation bits is handled by the descriptor fchmod. The seal no longer reopens files, and the verification pass that follows still rejects a swapped or unexpected entry.
  • doctor JSON gains a temporary-directory entry in checks. The CLI result schema already allows any check name, and a warning does not change ok. A project that was "needs attention" only because unselected agents' CLIs were missing now reports healthy, which is the intent. doctor still does not know about a one-off ultrafuzz run --agent X.
  • Governance. Files under .smithers/continuations no longer count toward the target identity. Only resume --refresh-controller writes there, and the exclusion is read only at launch. No disclosure acknowledgement needs recomputing: acknowledgements are required only for private campaigns, a private campaign already rejected a target dirtied by continuation files, and a target whose only untracked files were continuations now gets its clean-target digest back.
  • Runtime document parsing accepts up to 400,000 items in any runtime document instead of 100,000. This does not let an invalid document through: schema validation still applies each schema's own maxItems. It changes how much work happens before an oversized document is rejected. Memory stays bounded by the unchanged 128 MiB parse cap, but the validator runs with allErrors: true, so the quadratic uniqueItems check below still runs on an array that is over its maxItems. A corrupt document with up to 400,000 items in such an array therefore costs more CPU to reject than one cut off at 100,000. The seal needs the 400,000 limit under either design, and its identity arrays are such arrays.
  • Seal validation cost (pre-existing, not changed here). While trying to pin the 400,000 figure I found that schema validation of the seal's three identity arrays grows quadratically. With 30,000 identities per array, the runtime's validation took 6.1 s (0.76 s at 10,000). In a plain Ajv 2020 check of the same document, removing uniqueItems from identityArray took it from 6.6 s to 4 ms. The semantic gate already checks those arrays for uniqueness and order in linear time, as it does for execution_files, whose schema dropped uniqueItems in fix(runtime): keep execution-file validation linear #959 ("validate execution file uniqueness linearly"). Changing the schema changes its digest, so I left it. These arrays grow with the workflow's node count (the fixture's one node contributes one state ID, one attempt ID and three task node IDs); I did not measure a real campaign's seal.

Refs #921

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR does not appear ready to merge because removing the changelog exception is expected to fail Markdown validation.

Fix All in Claude CodeFindings

  1. P2 Unbounded controller size scan ▶
  2. P2 Non-seal item limit increased ▶
Fix with agent prompt
### Issue 1
packages/runtime/src/doctor.ts:undefined-245
Every `doctor` invocation walks all `ultrafuzz-controller-*` directories to total their file sizes, including the installed dependencies in retained controller roots. After several native resumes, this synchronous scan can make a readiness check increasingly slow even when temporary storage has ample space. Consider bounding the scan or making size collection optional.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

### Issue 2
packages/runtime/src/runtime-document-codec.ts:undefined-32
The control seal needs the higher limit, but this shared parser now allows up to 400,000 items for every runtime document instead of 100,000. For example, a dependency map with more than 100,000 issuers can pass parsing and require more work from later validation. Keep the higher limit specific to control seals so other documents retain their previous bound.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR reduces execution-snapshot file flushes, narrows doctor’s required agent CLIs to selected agents, adds temporary-directory diagnostics, excludes refreshed controllers from governed target identity, and raises the runtime-document parse limit.

  • The latest changes remove both the changelog’s MD013 exception and this PR’s release note.

Reviews (5) · Last reviewed commit: "chore: move the changelog entry to the c..."

aviggiano and others added 5 commits September 28, 2026 23:30
…nd flush each once

Snapshot publication wrote every file with mode 0400 and fsynced it, then
the permission seal reopened every file, fchmodded it (to 0500 for
executables, otherwise to the 0400 it already had) and fsynced it again.
A launch of the test fixture's 12,552-file snapshot therefore issued
27,876 fsyncs, two per file plus two per directory, and launch time is
fsync-bound.

writeSnapshotFile now creates each file with its final mode and sets that
mode through the creating descriptor (the umask can clear creation bits)
before its single flush, so the bytes and the mode are durable together,
as the second flush previously made them. The permission seal now only
makes directories read-only; the publication verification that already
follows it still checks every entry's type, mode, link count and bytes.
The same fixture now issues 15,324 fsyncs (one per file plus two per
directory).

The new test republishes a launched run's snapshot and bounds its fsync
calls at one per file plus two per directory; origin/main fails it with
27,876 flushes for 12,552 files in 1,386 directories.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
doctor required the executable of every configured model profile, so a
fresh `ultrafuzz init` on a Codex-only host reported "needs attention"
because the scaffold's opt-in Claude, Kimi and Pi profiles were not
installed, although no node would run them. It now requires an agent's
CLI only when the selected topology or its [retry] agents fallback chain
can dispatch to that agent (the same selection the OpenRouter credential
check already used). Other configured profiles' CLIs are still probed and
listed, as not required, and the human output marks them "(not
required)". Agents that share a CLI (Codex and OpenRouter, Claude and
DeepSeek) require it if any of them is selected.

doctor also built DOCTOR_WORKFLOW_ENGINE_* diagnostics and check statuses
for the project-local engine and then discarded them, reporting both
checks as "unknown". The builders are deleted; layout_status and the
per-patch posture are still reported, and the reference docs no longer
list the five codes that were never emitted.

Launch and resume install the workflow engine controller under the OS
temporary directory, and a native resume keeps its install there for the
detached engine (#921 cost 1: a tmpfs /tmp exhausted a host's memory). A
new temporary-directory check warns, without failing the verdict, when
that directory is a tmpfs or has under 2 GiB free, and reports how many
ultrafuzz-controller-* directories it holds and their total size. It
never removes them, because a live engine may still use them.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…dentity

`resume --refresh-controller` renders its current controller into
<project>/.smithers/continuations/<uuid>/, but that directory was missing
from controllerOwnedGovernancePaths. Unless the project ignores .smithers,
the rendered files are untracked target files: the next launch in that
project records the target as dirty, which a private campaign (the
default policy) and every cloud launch reject, and the changed worktree
digest invalidates disclosure acknowledgements computed for the target.

The directory is now controller-owned, like .smithers/workflows. The new
test renders files in that layout and checks the target identity is
unchanged; origin/main reports the target dirty.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ts schema

Runtime documents are parsed with a 100,000-item limit that the strict
JSON parser counts across the whole document. The control seal's schema
admits 100,000 execution files plus three identity arrays of up to
100,000 entries each, and launch writes the seal after schema validation
alone. A closure near the execution-file bound plus the run's task
bindings would therefore produce a seal that every later command fails
to parse. (An earlier measurement put the production closure at about
51,000 files; it was not re-measured for this change.)

The seal is now parsed with a 400,000-item limit, the schema's total; its
property total (four per execution file plus 35 fixed) already fits the
default limit. Byte limits are unchanged. The new test parses a seal with
the schema's maximum execution files and the fixture's bindings;
origin/main rejects it with "JSON exceeds the item limit of 100000".

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 00:06
diagnostics: []
};
const free = stats.bavail * stats.bsize;
const rootBytes = roots.reduce((total, root) => total + regularFileBytes(root), 0);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Unbounded controller size scan Every doctor invocation walks all ultrafuzz-controller-* directories to total their file sizes, including the installed dependencies in retained controller roots. After several native resumes, this synchronous scan can make a readiness check increasingly slow even when temporary storage has ample space. Consider bounding the scan or making size collection optional.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/doctor.ts
Line: 245

Comment:
**Unbounded controller size scan** Every `doctor` invocation walks all `ultrafuzz-controller-*` directories to total their file sizes, including the installed dependencies in retained controller roots. After several native resumes, this synchronous scan can make a readiness check increasingly slow even when temporary storage has ample space. Consider bounding the scan or making size collection optional.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Claude Code

Comment thread .markdownlintignore Outdated
Super-linter's markdownlint (MD013, 400 columns) validates every changed
Markdown file in full. The gate landed (#1005) after CHANGELOG.md was last
edited, and nearly every entry is one line longer than 400 columns, so
any pull request that adds a changelog entry fails "External static
analysis" on the file's existing lines. Line length is the only rule the
file breaks under super-linter's configuration, so disable just MD013 for
this file and keep every other Markdown rule.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/w16-launch-io-hygiene branch from 6fce903 to 87fa795 Compare September 29, 2026 00:22
aviggiano added a commit that referenced this pull request Sep 29, 2026
Super-linter lints every changed file in full, so any pull request that
adds a CHANGELOG entry fails MD013 on the file's existing entries, which
are single lines of up to 1,800 characters. This is the same file-level
directive #1179 adds, byte for byte, so the two merge without conflict.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
Review follow-up for the OpenRouter guardrails section.

- The section said the OpenCode openrouter/ model came from "the shipped
  profile". No OpenCode model profile ships: packages/config/defaults.toml
  and ultrafuzz.toml carry [agents.OpenCodeAgent] but no [models.opencode].
  Drop the parenthetical, and correct the OpenCode agent section's "The
  default root config includes an opt-in OpenCode profile", the claim it
  repeated. The doctor paragraph's copy of the claim is left to #1179,
  which rewrites that paragraph.
- OpenRouter's scan_scope for the prompt-injection builtin defaults to
  all_messages and can be set to user_only, so say that every message is
  scanned by default rather than always.
- Say that the workspace default and member guardrails also cover other
  keys, that a workspace used only for the Ultrafuzz key confines the
  workspace-default change to that key, and that in an organization
  account only an organization admin can change guardrails.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
Every CHANGELOG entry is one long line, so super-linter's markdownlint
fails MD013 (line length 400) on any pull request that touches the file.
release-gates then fails and every full release-validation lane is
skipped, so the runtime suite never runs on the pull request. This is
the same directive #1179 adds, byte for byte, so whichever lands first
merges cleanly with the other.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 4 commits September 29, 2026 03:37
Doctor now lists the CLIs of configured but unselected agents with
required: false, so the toolchain summary's "N required commands
available" counted them too. On a fresh init with only Codex installed it
reported 7 required commands when 4 are required.

The selection test also gains the case the reviewers found missing: a
selected Codex profile beside an unused OpenRouter profile keeps codex
required. Because configured agent refs are sorted, the existing
OpenRouter-only case could not tell "any requirer wins" from "last
writer wins"; this case fails under the latter.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… limit

The previous commit gave the control seal its own limits constant, a
schema-ID import and a ternary. Raising the single shared item limit is
the same fix in one number: runtime documents stay bounded by the
unchanged 128 MiB parse cap (and the seal by the 64 MiB control-file read
cap), and 400,000 is what a seal whose arrays are all at their schema
bounds holds (100,000 execution files plus three identity arrays of
100,000), with 400,035 properties under the unchanged 500,000 limit.

The seal test's comment now claims only what it fills: execution_files at
its bound. Filling the three identity arrays too would pin the 400,000
figure, but their schema validation grows quadratically (uniqueItems on
$ref items): with 30,000 identities per array it alone took 6.1 s, and a
test with every array at its bound ran for 85 s, so the test does not.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
#1181, #1172 and #1169 add `<!-- markdownlint-disable MD013 -->` right
after the changelog title. Adding the identical hunk here instead of a
disable-file comment at the end of the file means a merge of any two of
these PRs keeps one directive rather than two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
- "publishing an execution snapshot flushes each file once" now counts
  only fsyncs of regular-file descriptors and asserts that count equals
  the published file count (12,552 on the fixture; main flushes 25,104).
  The old total bound also encoded how directories are flushed.
- The leaf-swap test's hook now fires on the write-time fchmod, since the
  seal no longer opens files, so it is renamed after what it checks.
- The parent-swap test's hook fires only on a directory fchmod, so the
  swap lands in the directory seal after every file is written, as it did
  on main. Without that, no test swapped anything while directories were
  being sealed.
- The seal's doc comment now says what publication verification checks
  for each entry type instead of "every entry's type, mode, link count,
  and bytes".

Co-Authored-By: Claude Opus 5.5 <[email protected]>
// after schema validation alone. At its schema's array bounds a seal holds 400_000
// items (100_000 execution files and three identity arrays of 100_000 entries) and
// 400_035 properties.
maxItems: 400_000,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Non-seal item limit increased The control seal needs the higher limit, but this shared parser now allows up to 400,000 items for every runtime document instead of 100,000. For example, a dependency map with more than 100,000 issuers can pass parsing and require more work from later validation. Keep the higher limit specific to control seals so other documents retain their previous bound.

Knowledge Base Used: Execution runtime

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/runtime-document-codec.ts
Line: 32

Comment:
**Non-seal item limit increased** The control seal needs the higher limit, but this shared parser now allows up to 400,000 items for every runtime document instead of 100,000. For example, a dependency map with more than 100,000 issuers can pass parsing and require more work from later validation. Keep the higher limit specific to control seals so other documents retain their previous bound.

**Knowledge Base Used:** [Execution runtime](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/execution-runtime.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Claude Code

Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
* docs: explain OpenRouter prompt-injection guardrail rejections

OpenRouterAgent, PiAgent, and OpenCodeAgent with an openrouter/ model
all use the key in OPENROUTER_API_KEY. When a guardrail that covers that
key sets prompt-injection detection to Block, OpenRouter rejects each
matching request with HTTP 403 "Request blocked: prompt injection
patterns detected" before it reaches a model.

#1149 blamed transcript-like examples in Ultrafuzz's prompts. OpenRouter's
documented exact regexes match none of Ultrafuzz's prompt sources or the
rendered prompts of a local run. They do match text Ultrafuzz does not
write: OpenCode 1.18.18's default system prompt, used for models without
a model-specific prompt (DeepSeek, Qwen, GLM), matches
role_delimiter_injection, and so does ordinary Vyper or YAML source.

Document the operator fix in docs/config.md. Set prompt-injection
detection to Flag, or turn it off, on every guardrail that covers the
key, because OpenRouter applies the most restrictive action across the
workspace default and member or key guardrails. Do not use Redact, which
forwards the request with each match replaced. Runtime behavior is
unchanged: the rejection is an ordinary agent failure under the [retry]
policy, and the doc points there instead of restating it.

Closes #1149

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: correct the OpenCode profile claim and scope the guardrail advice

Review follow-up for the OpenRouter guardrails section.

- The section said the OpenCode openrouter/ model came from "the shipped
  profile". No OpenCode model profile ships: packages/config/defaults.toml
  and ultrafuzz.toml carry [agents.OpenCodeAgent] but no [models.opencode].
  Drop the parenthetical, and correct the OpenCode agent section's "The
  default root config includes an opt-in OpenCode profile", the claim it
  repeated. The doctor paragraph's copy of the claim is left to #1179,
  which rewrites that paragraph.
- OpenRouter's scan_scope for the prompt-injection builtin defaults to
  all_messages and can be set to user_only, so say that every message is
  scanned by default rather than always.
- Say that the workspace default and member guardrails also cover other
  keys, that a workspace used only for the Ultrafuzz key confines the
  workspace-default change to that key, and that in an organization
  account only an organization admin can change guardrails.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…stic failures, and label timeouts by code (#1171)

* fix(runtime): give agent retries a real wait and the whole planned chain

Agent retries used an exponential backoff from 1s, so a three-attempt
budget was spent in about 25 seconds, inside the minute a contended
Claude Code OAuth refresh can need to clear (#1084). The agent Task also
inherited Smithers' default identical-failure stall verdict (3), which
ended a chain before later same-agent attempts (exhaustive plans five) or
any [retry].agents fallback profile ran. Against real Smithers 0.35.0, a
[fail, fail, fail, fallback] chain ended `stalled` after attempt 3 and
never ran the fallback; with maxIdenticalFailures: 0 the fallback ran at
attempt 4 and the run finished.

- compileTask: initialDelayMs 60_000, so retries wait 60s, 120s, 240s,
  then Smithers' 300s cap.
- Agent Task: maxIdenticalFailures: 0, so the planned chain is the
  budget. It is set in the template only; the compiled manifest, cloud
  handoff schema and sealed task documents keep their exact shape.
- The real `smithers graph` smoke test never ran: it required a
  workspace-root .smithers install that no checkout has, and could not
  resolve the sealed module paths. It now uses the runtime package's own
  Smithers and asserts the retryPolicy Smithers receives.

Refs #1084

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): stop retrying deterministic admission failures and label timeouts by code

Two defects from #1144.

Retries. A consumer's dependency-admission failure re-reads the same
producer bytes, so every retry fails the same way; each consumer still
spent its whole retry budget and agent fallback chain. The admission
entry points (preparation's assertTaskDependencyInputs and the agent's
assertDependencyArtifactAdmissionCurrent rechecks) now throw with
details.failureRetryable=false, which Smithers honors, and
preparationStep keeps that flag when it rewraps the error. Everything
else stays retryable, including the #672 engine-boundary TypeError. A
real Smithers run of the template's preparationStep and helper pins
this: the admission failure runs once and its agent is skipped, while a
transient TypeError is retried and finishes. With the flag removed the
same preparation runs three times and ends `stalled`.

Timeout labels. errorLooksLikeTimeout ran /timeout|timed out|heartbeat/
over every string in the NodeFailed error, including the stack and
causes. That is a failure taxonomy from free text, which #572 rules
out, and since #1027 every JSON-validator preflight failure mentions
"timeout", so a 2ms `spawnSync ultrafuzz ENOENT` was recorded as a
timed-out provider interruption. A node now times out only on Smithers'
typed deadline codes (TASK_TIMEOUT, TASK_HEARTBEAT_TIMEOUT,
PROCESS_TIMEOUT, PROCESS_IDLE_TIMEOUT) or a TaskHeartbeatTimeout event.
The agent-failure normalizer used to strip PROCESS_TIMEOUT and
PROCESS_IDLE_TIMEOUT, which left the regex as the only label for real
agent CLI deadlines, so it now keeps them.

The #1144 analysis also proposed guarding generate()'s catch-path
recheck. It is not included: resetTaskArtifactsForRetry admits
dependencies before that try block on every first generation in a
process, so the catch path never runs without an admission.

Refs #1144

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: state what still ends the retry chain, and that text-only deadlines are labelled failed

Review of this change found the docs, CHANGELOG and one comment stronger
than the code:

- docs/config.md said the planned chain is the whole budget. Smithers
  still stops a chain at a failure it classifies as non-retryable, such
  as a CLI auth or configuration error, and pauses the run on a quota
  limit.
- The admission wrappers mark every failure inside admission
  non-retryable, including a file-system or validator error while
  reading the producer files, not only missing or changed bytes.
- Timeout labelling also keeps the TaskHeartbeatTimeout event rule, and
  a deadline reported only as text, such as the Modal provider's
  cloud-node deadline, is now labelled failed.
  docs/reference/artifacts-reports.md states the rule next to the node
  statuses.
- The workflow-sync comment read as if it listed every Smithers
  deadline code.

Refs #1144

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* ci: disable markdownlint MD013 for CHANGELOG.md

Every CHANGELOG entry is one long line, so super-linter's markdownlint
fails MD013 (line length 400) on any pull request that touches the file.
release-gates then fails and every full release-validation lane is
skipped, so the runtime suite never runs on the pull request. This is
the same directive #1179 adds, byte for byte, so whichever lands first
merges cleanly with the other.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 6b88ee4 into main Sep 29, 2026
13 checks passed
@aviggiano
aviggiano deleted the claude/w16-launch-io-hygiene branch September 29, 2026 06:36
aviggiano added a commit that referenced this pull request Sep 29, 2026
…on the real pinned engine (#1187)

* test(cli): run one campaign end to end with a controller kill and resume

Every runtime and CLI test drove a fake `smithers` shell script or ran the
engine on a hand-written workflow, so a break between the generated
workflow, the pinned Smithers engine, the sealed execution snapshot, and the
CLI only surfaced in a real campaign.

The new test runs `init`, `run`, `status`, `resume`, `stats`, `report`, and
`events` as separate CLI processes against the pinned engine that `run` and
`resume` install and start under Bun. A stub `codex` on PATH writes the
artifacts each prompt's output contract names, builds the final report from
the host-injected authorities, and renders it with the prompt's `ultrafuzz
report render` command. The stub holds the second node open while the test
SIGKILLs the detached engine and supervisor; once status reports the run
orphaned, the test resumes it. It asserts that the run succeeds with a
verified report, that only the interrupted node's agent ran twice, that no
engine task started again after it finished, and that status and stats
describe the same complete run.

A todo subtest records a gap it found: stats counts three agent attempts
where status and the stub count four, because attempts.jsonl is built from
NodeFinished/NodeFailed events and resume cancels the interrupted attempt
without one.

The file lives in test/e2e/ with its own `test:e2e` script, so the CLI suite
glob does not run it twice.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* ci: require the end-to-end campaign lane on pull requests

Adds a `cli-e2e` release-validation gate that runs
`pnpm --filter @ultrafuzz/cli test:e2e`, and a lane for it that is
required on pull requests. The lane has a 60-minute budget and the test
itself a 45-minute timeout. It runs beside the runtime lanes rather than
inside the push-only CLI lane.

The budget test in release-validation-lanes.test.ts keeps its 120-minute
expectation for the complete runtime and CLI suites only.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: disable markdownlint line length for CHANGELOG.md

Super-linter lints every changed file in full, so any pull request that
adds a CHANGELOG entry fails MD013 on the file's existing entries, which
are single lines of up to 1,800 characters. This is the same file-level
directive #1179 adds, byte for byte, so the two merge without conflict.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* test(cli): poll the event log for the end of the resumed campaign

`ultrafuzz status` synchronizes the run before it answers, unless the
control evidence has diverged. `events` streams the engine's event log
without synchronizing. The test now polls `events` until a terminal run
event, calls `status` once, and reuses those events for the no-restart
assertion. Locally, resume to run end fell from 165 s to 122 s.

The wait for the held node also fails at once if the workflow stops
before that node starts, instead of after the 15-minute bound.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* test(cli): report e2e reruns and stub failures by name, and clean up on interrupt

Review fixes for the end-to-end campaign test.

- The no-rerun checks now run straight after the events poll, before any
  command synchronizes the run. Those checks cover the stub start counts,
  the RunStarted count, the untruncated event page and no NodeStarted after
  NodeFinished. First the test checks that the engine's first run-terminal
  event is RunFinished. A resume that re-runs a finished node now fails on
  the named count. Before, it failed with an opaque
  ARTIFACT_VERIFICATION_AUTHORITY_INVALID from the status call that came
  first.
- The stub now fails, and logs why, in three cases: it finds no output
  contract, a named authority file is missing, or the report render line
  is not all `--flag 'value'` pairs. Before, a missed contract match
  exited 0 with a successful Codex turn and wrote nothing. The test's
  "workflow stopped" and RunFinished failures include the stub's call log.
- SIGINT and SIGTERM handlers, and the exit hook, now SIGKILL the
  campaign's processes and delete the fixture; the signal handlers then
  re-raise. A Ctrl-C'd run used to leave the detached engine, the
  supervisor and the ~1 GB fixture behind.
- The wait for the held agent to exit counts a zombie as exited. kill(pid,
  0) succeeds on a zombie; its /proc cmdline is empty.
- Drops the stats `attempts_complete === true` pin. That field only
  says the attempt ledger exists, and pinning it would break a fix that
  reports the undercounted attempts as partial evidence.
- The todo reason now names the cause that outlasts #1186. Smithers
  emits no terminal event for the attempt it abandons at resume. The
  reason cites #1187.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: say the e2e CLI runs under Node and the workflow under Bun

The CHANGELOG entry said the CLI commands ran on the pinned Smithers
engine under Bun. They run as Node CLI processes; only the generated
workflow runs on the engine under Bun. The development guide gets the
same precise wording, plus one sentence on what the test does when it
is interrupted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant