Skip to content

refactor(artifacts): delete never-dispatched gates, production-dead code, and the unread event index - #1181

Merged
aviggiano merged 7 commits into
mainfrom
claude/w19-artifacts-dead-code
Sep 29, 2026
Merged

aviggiano merged 7 commits into
mainfrom
claude/w19-artifacts-dead-code

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

packages/artifacts ships a lot of code that production never runs, and part of it slows down or can stop a campaign:

  • Semantic gates. 48 of the 173 registered gates are never dispatched. The 12 planned-graph gates duplicate assertPlannedGraphSemantics, which the runtime does call, and at least one copy had drifted from it: planned-graph-artifact-dir-identity required artifacts/<id> for every node, while assertPlannedGraphSemantics accepts the storage/attempt directory of dynamically generated nodes.

  • Dead code. Several modules and functions have no production caller, and some have large test suites: finding-provenance.ts, the generated-test manifest writer/reader, appendLineDurable, truncateDurable, buildRunArtifactIndex, loadOrCreateRunState, validateWorkflowContract, resolvePathInside and about 60 declaration-only exports.

  • Event journal. appendEvent validated the whole journal several times per append:

    • It pre-read events.jsonl plus up to five events.index/* fan-out files.
    • Each append then read its file again.
    • appendStrictJsonlRecords read the file a third time after writing.

    So append cost grew with journal size, and a run's total append cost grew quadratically. Separately, the event codec inherited the generic 100,000-record cap. Once a journal reached it, every later append and replay threw, and the sync, cancelRun and other lifecycle paths all append events.

  • Parity throws. The findings and property validators ran the registered JSON Schema, which is documented as authoritative. They then parsed the document again with Zod and threw internal schema parity invariant violated whenever Zod disagreed. These validators back workflow-sync and the artifact gates.

Root cause

  • Gates are registered per schema file, including runtime-state schemas. Production only dispatches:

    • artifact-contract schemas (workflow.tsx, artifact-gates.ts, artifact validate);
    • report.schema.json (modal);
    • the validator preflight, the invariant source proof and the attempt/usage ledgers;
    • a fixed set of gates called by name.

    No caller ever passed the planned-graph, task-manifest, artifact-verification/manifest, agent-source-proof, run-metadata/plan/source-run, config-redactions, invariant-suite-manifest, single-finding, run-state-fingerprint or analysis-bundle-digest schemas.

  • The events.index/ fan-out and its query-inputs.json facade have no reader. queryEvents filters events.jsonl, and the only thing that touches the directory is the report bundle, which copies it.

  • The 100k cap is the strict-JSONL default. The operator journal never overrode it, and neither did the report bundle's private copy of the journal parser.

Change

  1. Delete the 48 never-dispatched gates (4c64f583). This removes:
    • their registrations, exclusive handlers and helpers;
    • their names in ARTIFACT_SCHEMA_METADATA;
    • the context fields only they read (plannedGraph, runtimeState, git.refs) and the one runtime line that built the unused plannedGraph context;
    • the tests and fixtures that only exercised these gates. The single-finding provenance cases now run through the live findings-campaign-provenance-coherence gate.
  2. Delete production-dead code in artifacts and security (6c3c69e0). Each symbol was checked with git grep -w over the whole repository (every package's src and tests, scripts, the workflow templates) and with git log -S over the workflow template's history; none has a production caller. Their tests go with them. The two usage-ledger tests that used appendLineDurable to inject a bad line now use fs.appendFileSync. In security this removes resolvePathInside and its helpers, the UnsafeModeAudit* types, and PolicyResult.audit, which no caller ever set.
  3. Event journal (5ae10beb):
    • appendEvent writes only events.jsonl. It checks the new record against the final record plus any trailing records with the same timestamp, using appendStrictJsonlRecordsAfterTail. The event ID hashes the timestamp and timestamps never decrease, so a record outside that window can repeat the ID only through a 96-bit hash collision.
    • replayEvents, queryEvents and the new parseEventJournalBytes still validate the whole journal. The byte length read before the write still fences it.
    • appendStrictJsonlRecords returns the snapshot it validated plus the fenced append instead of re-reading the file.
    • The event codec has no record cap. Its 64 MiB byte limit still applies.
    • createRunLayout no longer creates events.index/ or query-inputs.json, and RunLayout.eventsIndexDir is gone.
    • report bundle validates its journal snapshot with parseEventJournalBytes. This deletes its ~80-line copy of the parser, which also enforced the 100k cap.
    • Docs (artifacts-reports.md, SPECS.md) are updated.
  4. Return the Ajv verdict (eaa1ed59). The findings and property validators drop the second Zod parse and its throws. The Zod schemas still generate the JSON Schema and the TS types, and contract-fixtures.test.ts still checks that Ajv and Zod agree on every fixture, so drift fails CI instead of a campaign.
  5. CHANGELOG entry (15944d03).
  6. <!-- markdownlint-disable MD013 --> at the top of CHANGELOG.md (638ff30f). Every entry there is one line past markdownlint's 400-character limit, so any PR that touches the file fails External static analysis, and that failure skips the release-validation lanes. fix(runtime): resume --retry-failed reopens descendants skipped behind a recovered producer #1169 adds the same line.

LOC delta: +285 / −4,660 (net −4,375). Source is −2,530, tests −1,851, docs/CHANGELOG +6.

Deliberately not built (and why)

  • writeArtifact stays. It has no production caller, but about 350 test call sites across three suites use it as a fixture helper. Moving it into each suite would add code, and one of those suites (runtime/test/artifact-gates.test.ts) is being rewritten in parallel.
  • The event-query-facade.schema.json schema, its Zod twin and its JSON export stay registered. Deleting a schema file changes the schema-bundle digest, and that strands in-flight runs.
  • The CLI bundle manifest still lists events.index in included_roots. The CLI manifest schema pins that constant list, and changing it is a schema change. The bundle still copies the directory when an older run has one.
  • The 64 MiB journal byte limit is unchanged. Readers load the whole journal into memory, so raising it is a separate decision. At about 454 B per workflow-synced record it holds roughly 148k records.
  • No positioned tail read. An append still reads the journal's bytes. Only the parse is windowed, and the byte read is what remains of the per-append growth in the numbers below. A positioned read would duplicate the hardened reader in schema-registry.ts, which must not be edited without rotating the validator identity.
  • Other readers not changed.
    • evals/src/node-telemetry.ts keeps its own 100k cap on run journals. It is eval-only, and the evals dead-code work item covers that file.
    • analysis-bundle.ts has its own Ajv/Zod disagreement path, which returns issues rather than throwing.
    • Security's validateSafeId(label, id) still shares a name with the artifacts version. It is not dead code (materialize-policy uses it).
  • One stale test title. The planned graph v4 validates whole documents and executes every registered document semantic gate test keeps its title and gate-style rule labels, although planned-graph now has no registered gates; the test itself exercises assertPlannedGraphSemantics. Retitling touches the first line of a 176-line test function, so the diff-limited strict lint reports max-lines-per-function on it, and the validator-build work item is editing that test.
  • Ledgers unchanged. Attempt and usage ledgers still validate their whole history on append. They need it for idempotent identity replay, and that code belongs to the ledger work item.

Verification

Discriminating tests

  • event appends and replays keep working past 100,000 records (artifacts.test.ts): fails on main with event journal exceeds the record limit, and passes here. I ran it by copying the test file into a main worktree.
  • an event append refuses a repeated or out-of-order event without changing the journal passes on both main and this branch. It fails with Missing expected exception when I temporarily shrank the window to the final record only (() => false). So it guards the same-timestamp window, not the old behaviour.
  • Part 4 has no discriminating test. I know of no document on which the two validators disagree: the contract fixtures pin their agreement, and I probed timestamp, integer and string edge cases without finding one. On such documents the change only removes the second parse.

Identity unchanged. I compared artifactSchemaBundleDigest(), VALIDATOR_BUILD_IDENTITY, all 44 contract digests and all 70 schema SHA-256s between a main build and this branch: they are byte-identical. pnpm --filter @ultrafuzz/artifacts schema:check passes, and no schema/*.json, validator-identity module or contract description changed.

Dispatch analysis. I computed the gate set from the built registry: registered names minus the gates of every schema file a dispatch site can pass, minus the gates called by name. The result is exactly 48 gates, and none of the 48 names appears outside the artifacts gate module, metadata and tests.

Benchmark. CPU per appendEvent, measured with fsync disabled (eatmydata) on a 32-core host at load average about 38 (noisy):

existing events main branch
250 17.8 ms 4.8 ms
10,000 478 ms 12.5 ms
50,000 2,183 ms 39 ms

Suites and checks run

  • artifacts full suite: 322/322.
  • security full suite: 21/21.
  • runtime:
    • getRunStatus counts only a canonical event-v2 journal and fails closed on invalid presence: pass. This is the corrupt-journal status test.
    • aggregation-semantic-context, audit-contracts, verified-output, workflow-dependency-policy, semantic-artifact-context, forge-guard, artifact-gates and generated-workflow-verifier: 404/404 pass.
  • cli: every report bundle and stats test (--test-name-pattern='report bundle|^stats', --test-concurrency=1): 54 of 55 passed. report bundle --require-verified rejects manifest bytes injected only into the recursive archive read failed in its setup. There, ultrafuzz run returned WORKFLOW_SUBMISSION_FAILED: workflow execution file dependencies/packages/000006/LICENSE changed while reading while I was rebuilding the workspace in parallel. Rerun alone, it passes on this branch and on main. report bundle --require-verified fails closed on every present invalid event journal passes; it exercises the parser the bundle now shares.
  • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, eslint and prettier --check on the changed files: pass.
  • pnpm --filter @ultrafuzz/{security,artifacts,runtime,cli} typecheck: pass. tsc --noEmit for every other package against the new builds: pass.
  • pnpm -w knip and node scripts/docs-check.mjs: pass.

PR CI on 638ff30f: every check passes. That includes the four runtime integration shards and the runtime support/Bun adapter-contract lane, which together run the whole runtime suite.

Not run: the whole cli.test.ts suite (only the patterns above; the CLI lane is push-only), and the modal suite. The evals and dashboard suites ran only partially: evals recovery-equivalence.test.ts 42/42, and the dashboard event|journal|audit tests 4/4. None of those packages imported any removed export, and all of them typecheck.

Risk / compatibility

  • Removed public exports. @ultrafuzz/artifacts and @ultrafuzz/security lose exports: the gate names in SemanticGateName, the functions listed above, appendEventRecord, the facade reader/validators, and RunLayout.eventsIndexDir. No package in the repo uses them, and none has ever appeared in the workflow template (checked with git log -S). Generated workflows load @ultrafuzz/artifacts from their sealed execution snapshot, and none of their imports changed.
  • Existing runs. Stale events.index/ directories stay on disk. Nothing reads them, and report bundles still copy them.
  • Append validation is narrower. A journal already corrupted before the window, for example by a hand edit, is no longer rejected at append time. It is still rejected by the next read, and every sync replays the journal. The append applies the readers' rules to the new record and its window.
  • Merge overlap.
    • semantic-gates.ts: the open Reconcile invocation usage and attempt gaps #1155 and the parallel timeout-evidence change edit it. The hunks here are the deleted handlers and registrations, away from propertyCampaignTimeoutEvidenceIssues.
    • artifacts.test.ts: the parallel ledger and secret-gate changes edit the attempt-ledger tests and the secret-gate test. This PR deletes other blocks and touches the import list.

Refs #462

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the changes reviewed since the previous review.

Summary

The PR removes unused artifact gates and production-dead code, drops the unread event index, changes event appends to validate a trailing window, and uses the registered JSON Schema verdict for findings and property validation.

  • Event replay and report bundling continue to validate complete journal snapshots.
  • No new actionable issue was established in the changes since the previous review.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[appendEvent] --> B[Validate trailing timestamp window]
  B --> C[Append to events.jsonl]
  C --> D[replayEvents and queryEvents]
  C --> E[Report bundle]
  D --> F[Validate complete journal]
  E --> F
Loading

Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."

aviggiano and others added 5 commits September 28, 2026 23:53
Of the 173 registered artifact semantic gates, 48 were only reachable
through executeSchemaSemanticGates for runtime documents that no caller
passes (planned-graph, smithers-task-manifest, artifact-verification,
artifact-manifest, agent-source-proof, run-metadata, run-plan, source-run,
config-redactions, invariant-suite-manifest, the single-finding subschema,
the run-state fingerprint and the analysis-bundle file digest). Production
dispatches gates only for artifact-contract schemas (workflow.tsx,
artifact-gates.ts, `artifact validate`), report.schema.json (modal), the
validator preflight, the invariant source proof, the attempt/usage ledgers,
and a fixed list of by-name calls; none of those names any of the 48.

The runtime validates those documents with their JSON Schemas and the
TypeScript assertions it does call (assertPlannedGraph and
assertSealedPlannedGraph, assertSmithersTaskManifestMatchesPlannedGraph,
assertRunMetadataDocument, assertArtifactVerificationMarkerSemantics, ...).
The planned-graph gates duplicated assertPlannedGraphSemantics, whose own
test covers each of the 12 rules, and at least one copy had drifted:
planned-graph-artifact-dir-identity required artifacts/<id> for every node,
while assertPlannedGraphSemantics accepts the storage/attempt directory of
dynamically generated nodes.

Delete the registrations, their exclusive handlers and helpers, their names
in ARTIFACT_SCHEMA_METADATA, the context fields only they read
(plannedGraph, runtimeState, git.refs) and the runtime line that built the
unused plannedGraph context. Tests that only exercised the deleted gates go
with them; the single-finding provenance case now runs through the live
findings@2 gate.

Schema bytes, the schema-bundle digest, every contract digest and
VALIDATOR_BUILD_IDENTITY are byte-identical to main (compared from both
builds), so the schema bindings sealed into existing runs still match.

Refs #462

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Each symbol below had no caller in any package source, script or workflow
template (git grep -w over the whole repository), and none ever appeared
in the workflow template's history:

- finding-provenance.ts (normalizeFindings, buildFindingSourceExpectations,
  readFindings, ...): only threat-goal-artifacts tests called it.
- writeGeneratedTestManifest/readGeneratedTestManifest and their seven
  private helpers: agents write generated-tests.json and the host only
  validates it, so only artifacts.test.ts exercised the writer.
- appendLineDurable and truncateDurable (safe-paths.ts), and
  buildRunArtifactIndex with its two types (manifests.ts).
- loadOrCreateRunState, validateWorkflowContract, and 29 other workflow-contracts
  exports that occurred only at their declaration.
- Declaration-only exports in findings.ts, findings-schema.ts,
  finding-note-vocabulary.ts, goal-plan.ts, invariant-ledger.ts,
  invariant-source-proof.ts, property-provenance.ts, state-schema.ts,
  runtime-schemas.ts, report-observation.ts, threat-model.ts,
  usage-ledger.ts and artifact-contract-ids.ts.
- security: resolvePathInside and its helpers/types, the UnsafeModeAudit*
  types, and PolicyResult.audit (no caller ever passed it; it was always []).

Tests of the deleted code go with it. The two usage-ledger tests that used
appendLineDurable to inject a bad line now use fs.appendFileSync, and the
stale-size/hard-link test keeps its appendBytesDurableAt assertions.

writeArtifact stays: it has no production caller either, but ~350 test call
sites across three suites use it as a fixture helper, and moving it into
each suite would add code rather than remove it.

Schema bytes, contract digests and VALIDATOR_BUILD_IDENTITY are unchanged.

Refs #462

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… an unread index

appendEvent validated the whole journal several times per append: it
pre-read events.jsonl and five events.index/* fan-out files, then each
append re-read its file in full, and appendStrictJsonlRecords re-read the
file once more after writing. Nothing reads events.index/* or its
query-inputs.json facade (queryEvents filters events.jsonl; the report
bundle only copies the directory). The 100,000-record cap on the event
codec also failed every later append and replay once reached, and the
sync, lifecycle and cancel paths all append events.

- appendEvent writes only events.jsonl. It checks the new record against
  the final record and the trailing records that share its timestamp. The
  event ID hashes the timestamp and timestamps never decrease, so a record
  outside that window can repeat the ID only through a 96-bit hash
  collision. Readers still validate the whole journal, and the byte length
  read before the write still fences it.
- appendStrictJsonlRecords returns the snapshot it validated plus the
  fenced append instead of re-reading the file.
- The event codec has no record cap; its 64 MiB byte limit still applies.
- createRunLayout no longer creates events.index/ or query-inputs.json,
  and RunLayout drops eventsIndexDir. The facade schema stays registered
  because removing a schema file would change the schema-bundle digest.
- The report bundle validates its journal snapshot with
  parseEventJournalBytes instead of its own copy of the parser, which
  still enforced the 100,000-record cap.

Measured CPU per append with fsync disabled (eatmydata, loaded 32-core
host): 17.8 -> 4.8 ms at 250 existing events, 478 -> 12.5 ms at 10,000,
2,183 -> 39 ms at 50,000. The remaining growth is the byte read.

"event appends and replays keep working past 100,000 records" fails on
main ("event journal exceeds the record limit") and passes here. The
same-millisecond duplicate test passes on both, and fails if the window
is shrunk to the final record only.

Refs #462

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…verdict

validateFindingSchema, validateFindingsSchema and the five property
validators ran the registered JSON Schema (documented as authoritative),
then parsed the same document again with its Zod schema and threw an
internal "schema parity invariant violated" error if Zod rejected or
transformed it. The Zod result was discarded either way. Those validators
back workflow-sync and several artifact gates, so a disagreement would
have aborted sync or verification with an internal error instead of
returning a validation result.

Return the Ajv result directly. The Zod schemas still generate the JSON
Schema and the TypeScript types, and contract-fixtures.test.ts still checks
that Ajv and Zod agree on every fixture, so a drift fails CI instead of a
campaign.

No test here discriminates: no document is known on which the two
validators disagree. The contract fixtures pin their agreement, and a probe
of timestamp, integer and string edge cases found none. On such documents
the change only removes the second parse.

Refs #462

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 00:13
Super-Linter lints only the files a pull request changes, and every
CHANGELOG.md entry is a single line far past markdownlint's 400-character
MD013 limit, so any PR that adds an entry fails external static analysis.
That failure also skips the release-validation lanes that depend on it.
Other open PRs add the same directive.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
Since the external static-analysis gate was adopted (#1005), super-linter
lints every changed Markdown file with MD013 at 400 characters.
CHANGELOG.md keeps one entry per line, and markdownlint reports 42 existing
entries on main as longer than that, so any change to the file fails
`External static analysis`.
That also fails `release-gates` and skips the release validation lanes on
the pull request. No commit has changed CHANGELOG.md since the gate
landed. Disable MD013 below the title of this file only, using the same
directive line that #1169, #1172 and #1181 add, so the branches do not
end up with two spellings of it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
#1181, #1172 and #1169 add `<!-- markdownlint-disable MD013 -->` right
after the changelog title. Adding the identical hunk here instead of a
disable-file comment at the end of the file means a merge of any two of
these PRs keeps one directive rather than two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 4b2cabd into main Sep 29, 2026
13 checks passed
@aviggiano
aviggiano deleted the claude/w19-artifacts-dead-code branch September 29, 2026 06:32
aviggiano added a commit that referenced this pull request Sep 29, 2026
…r-render cost drops (#1168)

* fix(runtime): literal braces in prompt text no longer fail the workflow render

renderAgentPrompt substituted the operator and task prompts into the
trusted agent-prompt template and then regex-scanned the whole result for
leftover `{{word}}` placeholders. Inserted text that merely contained such
a sequence threw inside the Smithers render: an escaped `\{{word}}` example
in a prompt template (the prompt renderer emits it as literal `{{word}}`),
the same escape inside a dynamic item value (kept verbatim), or an
operator `--prompt` note. Every render builds every task's prompt, so the
run failed, and failed again on every resume.

Check only the template's own placeholders, inside the replace callback,
as renderAgentPreambleTemplate already does. A template placeholder with
no value still throws, and now names the placeholder.

The dynamic prompt renderer also resolves `{{...}}` inside goal-plan
replacement values, and an unbound name there throws in the same render.
The goal-plan contract now rejects `{{` in replacement values, which are
plain-text labels, so that failure lands on goal-plan's own verify instead
of on every render. The check is a zod refinement; goal-plan.schema.json
is unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): run the JSON validator preflight once per engine process

prepareArtifactMirror spawned `ultrafuzz json validate` against the smoke
fixture in every prepare, in every agent-attempt reset, when an agent is
built after a controller restart, and again in the zero-retry verify task.
The trusted launcher's cold start was measured at ~35 s under contention
(#1026), so each spawn was another chance to fail an attempt, including
one whose agent work had already finished.

What the spawn proves -- that this process can launch the agent-facing
validator -- does not depend on the task: materializePromptSchemas has
just digest-checked the workspace's schema copy, and start-run already
runs the trusted-launcher preflight at launch and resume. Remember the
first success in the engine process. A failure is not remembered, so the
next caller spawns again.

Co-Authored-By: Claude Opus 5.5 <[email protected]>


* perf(runtime): cut per-process and per-render cost of the generated workflow

agent-registry.ts and smithers.ts imported the TypeScript compiler at
module load, and both are reachable from @ultrafuzz/runtime's index. Only
registry analysis (validate and init) and the controller-refresh helper
that only tests call ever use it. Require it at first use instead.

The saving lands in processes that import the runtime without smthrs:
the ultrafuzz CLI, including every agent `ultrafuzz json validate` call.
Smithers engine processes still load the compiler: the generated
workflow imports smthrs, whose index re-exports @smthrs/scorers, and two
scorer modules import typescript at module load. Importing the built
runtime index alone on this host went from 528-547 ms / 248-250 MB RSS to
409-437 ms / 196-198 MB under Node 24, and from 449-483 ms / 252-261 MB
to 336-351 ms / 209-213 MB under Bun 1.3.14 (5 runs each).

Changing the parameter type on topLevelNameIsBound's signature makes the
strict diff lint report that function's existing complexity, so its
import-clause check moves into a helper with the same conditions.

materializeDynamicRuntime runs on every render and rewrote tasks.json and
graph.json, with fsyncs, even when it re-derived the bytes already on
disk. Skip the write in that case.

taskSpecsFromCompiled looked every task up with serializedTaskSpecs.find;
index the specs by id once per call instead. At the default topology's 65
static tasks the difference is not measurable.

Add a test that compiles the packaged default topology and holds the
generated workflow under 2.3 MB, measured with a fixed-length project root
(about 2.06-2.07 MB today, depending on the checkout path), so the size
#1146 describes cannot grow silently.

Refs #1146

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: changelog for render robustness and runtime footprint

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: exempt CHANGELOG.md from the markdown line-length rule

Since the external static-analysis gate was adopted (#1005), super-linter
lints every changed Markdown file with MD013 at 400 characters.
CHANGELOG.md keeps one entry per line, and markdownlint reports 42 existing
entries on main as longer than that, so any change to the file fails
`External static analysis`.
That also fails `release-gates` and skips the release validation lanes on
the pull request. No commit has changed CHANGELOG.md since the gate
landed. Disable MD013 below the title of this file only, using the same
directive line that #1169, #1172 and #1181 add, so the branches do not
end up with two spellings of it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): the post-agent verify pass no longer preflights the validator CLI

finalizeAndVerifyArtifacts calls prepareArtifactMirror with
pinnedSubmodules "verify", and that pass still ran the
`ultrafuzz json validate` preflight. Verify checks the finished agent's
outputs in-process and never runs the agent-facing CLI, and its task has
no retry. The per-process memo only helps later callers: after an engine
restart, the first preflight in the new process can be the verify of an
attempt whose agent had already finished, and a CLI cold start that timed
out there (the trusted launcher measured ~35 s under contention in #1026)
failed verification and discarded that work.

Skip the preflight in the verify pass. prepare:* and the reset before each
agent attempt, which run ahead of an agent that uses the CLI, keep it.

The new lifecycle test runs the production preparation body with the
options finalizeAndVerifyArtifacts passes. Without this change it counts
three preflights instead of two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: state where the TypeScript saving and the goal-plan brace rule apply

The comments, the footprint test and the changelog said Smithers engine
processes stop loading the TypeScript compiler. They do not: the
generated workflow imports smthrs, whose index re-exports
@smthrs/scorers, and its sideEffectAnalysis and workflowUiCompliance
modules import typescript at module load. The saving lands in ultrafuzz
CLI processes, including every agent `ultrafuzz json validate` call.

The prompt-variables reference and the changelog also said the
goal-plan@1 contract rejects `{{` in replacement values. Only goal-plan's
own verification (and the evals goal-plan reader) run that zod rule;
`ultrafuzz json validate` and `ultrafuzz artifact validate` accept such a
plan. Say so, replace the exact generated-workflow byte count, which does
not reproduce across environments, with a range, and describe the
preflight as running until its first success.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* test(runtime): cover imported Object bindings in registry shadowing

importClauseBindsName, split out of topLevelNameIsBound on this branch,
had no test: the registry-shadowing test only declared Object locally.
Add default, namespace, named, aliased and default-plus-named imports of
Object, each of which must stop Object.freeze from being read as the
global, and an unrelated named import that must not. Removing any one of
the helper's three checks, or making it match every named import, fails
the test.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…nused agent profiles (#1179)

* perf(runtime): write execution-snapshot files with their final mode and flush each once

Snapshot publication wrote every file with mode 0400 and fsynced it, then
the permission seal reopened every file, fchmodded it (to 0500 for
executables, otherwise to the 0400 it already had) and fsynced it again.
A launch of the test fixture's 12,552-file snapshot therefore issued
27,876 fsyncs, two per file plus two per directory, and launch time is
fsync-bound.

writeSnapshotFile now creates each file with its final mode and sets that
mode through the creating descriptor (the umask can clear creation bits)
before its single flush, so the bytes and the mode are durable together,
as the second flush previously made them. The permission seal now only
makes directories read-only; the publication verification that already
follows it still checks every entry's type, mode, link count and bytes.
The same fixture now issues 15,324 fsyncs (one per file plus two per
directory).

The new test republishes a launched run's snapshot and bounds its fsync
calls at one per file plus two per directory; origin/main fails it with
27,876 flushes for 12,552 files in 1,386 directories.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): doctor requires only the agent CLIs a run can dispatch to

doctor required the executable of every configured model profile, so a
fresh `ultrafuzz init` on a Codex-only host reported "needs attention"
because the scaffold's opt-in Claude, Kimi and Pi profiles were not
installed, although no node would run them. It now requires an agent's
CLI only when the selected topology or its [retry] agents fallback chain
can dispatch to that agent (the same selection the OpenRouter credential
check already used). Other configured profiles' CLIs are still probed and
listed, as not required, and the human output marks them "(not
required)". Agents that share a CLI (Codex and OpenRouter, Claude and
DeepSeek) require it if any of them is selected.

doctor also built DOCTOR_WORKFLOW_ENGINE_* diagnostics and check statuses
for the project-local engine and then discarded them, reporting both
checks as "unknown". The builders are deleted; layout_status and the
per-patch posture are still reported, and the reference docs no longer
list the five codes that were never emitted.

Launch and resume install the workflow engine controller under the OS
temporary directory, and a native resume keeps its install there for the
detached engine (#921 cost 1: a tmpfs /tmp exhausted a host's memory). A
new temporary-directory check warns, without failing the verdict, when
that directory is a tmpfs or has under 2 GiB free, and reports how many
ultrafuzz-controller-* directories it holds and their total size. It
never removes them, because a live engine may still use them.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): keep refreshed controllers out of the governed target identity

`resume --refresh-controller` renders its current controller into
<project>/.smithers/continuations/<uuid>/, but that directory was missing
from controllerOwnedGovernancePaths. Unless the project ignores .smithers,
the rendered files are untracked target files: the next launch in that
project records the target as dirty, which a private campaign (the
default policy) and every cloud launch reject, and the changed worktree
digest invalidates disclosure acknowledgements computed for the target.

The directory is now controller-owned, like .smithers/workflows. The new
test renders files in that layout and checks the target identity is
unchanged; origin/main reports the target dirty.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): parse the control seal with an item limit that covers its schema

Runtime documents are parsed with a 100,000-item limit that the strict
JSON parser counts across the whole document. The control seal's schema
admits 100,000 execution files plus three identity arrays of up to
100,000 entries each, and launch writes the seal after schema validation
alone. A closure near the execution-file bound plus the run's task
bindings would therefore produce a seal that every later command fails
to parse. (An earlier measurement put the production closure at about
51,000 files; it was not re-measured for this change.)

The seal is now parsed with a 400,000-item limit, the schema's total; its
property total (four per execution file plus 35 fixed) already fits the
default limit. Byte limits are unchanged. The new test parses a seal with
the schema's maximum execution files and the fixture's bindings;
origin/main rejects it with "JSON exceeds the item limit of 100000".

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: changelog entry for launch I/O and doctor fixes

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: disable only MD013 for the one-entry-per-line changelog

Super-linter's markdownlint (MD013, 400 columns) validates every changed
Markdown file in full. The gate landed (#1005) after CHANGELOG.md was last
edited, and nearly every entry is one line longer than 400 columns, so
any pull request that adds a changelog entry fails "External static
analysis" on the file's existing lines. Line length is the only rule the
file breaks under super-linter's configuration, so disable just MD013 for
this file and keep every other Markdown rule.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): count only required commands in doctor's toolchain summary

Doctor now lists the CLIs of configured but unselected agents with
required: false, so the toolchain summary's "N required commands
available" counted them too. On a fresh init with only Codex installed it
reported 7 required commands when 4 are required.

The selection test also gains the case the reviewers found missing: a
selected Codex profile beside an unused OpenRouter profile keeps codex
required. Because configured agent refs are sorted, the existing
OpenRouter-only case could not tell "any requirer wins" from "last
writer wins"; this case fails under the latter.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* refactor(runtime): parse every runtime document with one 400,000-item limit

The previous commit gave the control seal its own limits constant, a
schema-ID import and a ternary. Raising the single shared item limit is
the same fix in one number: runtime documents stay bounded by the
unchanged 128 MiB parse cap (and the seal by the 64 MiB control-file read
cap), and 400,000 is what a seal whose arrays are all at their schema
bounds holds (100,000 execution files plus three identity arrays of
100,000), with 400,035 properties under the unchanged 500,000 limit.

The seal test's comment now claims only what it fills: execution_files at
its bound. Filling the three identity arrays too would pin the 400,000
figure, but their schema validation grows quadratically (uniqueItems on
$ref items): with 30,000 identities per array it alone took 6.1 s, and a
test with every array at its bound ran for 85 s, so the test does not.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: use the changelog MD013 directive sibling PRs add

#1181, #1172 and #1169 add `<!-- markdownlint-disable MD013 -->` right
after the changelog title. Adding the identical hunk here instead of a
disable-file comment at the end of the file means a merge of any two of
these PRs keeps one directive rather than two.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* test(runtime): make the snapshot tests exercise the phases they name

- "publishing an execution snapshot flushes each file once" now counts
  only fsyncs of regular-file descriptors and asserts that count equals
  the published file count (12,552 on the fixture; main flushes 25,104).
  The old total bound also encoded how directories are flushed.
- The leaf-swap test's hook now fires on the write-time fchmod, since the
  seal no longer opens files, so it is renamed after what it checks.
- The parent-swap test's hook fires only on a directory fchmod, so the
  swap lands in the directory seal after every file is written, as it did
  on main. Without that, no test swapped anything while directories were
  being sealed.
- The seal's doc comment now says what publication verification checks
  for each entry type instead of "every entry's type, mode, link count,
  and bytes".

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant