Skip to content

integration: wave-1 combined CI check (do not merge) - #1190

Closed
aviggiano wants to merge 234 commits into
mainfrom
integration/wave1
Closed

aviggiano wants to merge 234 commits into
mainfrom
integration/wave1

Conversation

@aviggiano

Copy link
Copy Markdown
Collaborator

Do not merge. This draft exists only to run the full CI on the combination of the 25 wave-1 pull requests (#1163–#1188, excluding the draft #1183) merged together, with the conflict resolutions that the merge train will replay. Each PR is merged individually into main; this PR is closed afterwards.

Merge order: #1165, #1163, #1166, #1169, #1177, #1170, #1176, #1178, #1181, #1172, #1167, #1164, #1174, #1175, #1182, #1180, #1185, #1186, #1171, #1173, #1168, #1179, #1187, #1188, #1184.

🤖 Generated with Claude Code

aviggiano and others added 30 commits September 28, 2026 22:41
…y for fan-out membership

A published dynamic-expansions/<group>.json is the durable record of which
generated nodes a group has, but loadOrCreateDynamicExpansion re-proved it on
every workflow render and every strict lifecycle admission: it re-read and
re-hashed the live source artifact and required it to match
source.output_sha256, all under an O_EXCL lock that was never reclaimed.

Both are self-inflicted stops:

- A reset that re-runs the source (a --reset-node, or a Smithers timetravel
  that sweeps the source into its dependent set) empties the source's
  canonical artifact directory when its agent attempt starts. Every render
  then threw "dynamic source artifact does not exist" and, once the source
  wrote a different plan, DYNAMIC_EXPANSION_CHANGED, with no way back. The
  same check made strict admission refuse cancel, pause, why, fork and replay.
- A controller killed (SIGKILL, OOM) or an observer interrupted while holding
  .expansion.lock left every later render and admission busy-waiting 5 s and
  then throwing DYNAMIC_EXPANSION_LOCKED.

An existing manifest is now validated (the manifest set plus the group's
static identity: run, group, source node, attempt and path, JSON and key
paths, node-ID template, prompt template digest and fingerprint, limit) and
returned without touching the source. The source is read only to create the
manifest, and its digest is still recorded as provenance. The lock is
deleted: reads need none, creation happens in the workflow render, which
Smithers runs for one live driver per run, publishFileDurableExclusive never
replaces a published file, and the existing re-read after publication refuses
an inconsistent set.

Semantic change: after a reset re-runs a source, the fan-out keeps its
originally published items instead of failing. resume --retry-failed of a
failed source verifier still archives the manifests (#1064) and re-expands.

The tests that encoded the old behaviour (a changed source must throw, the
lock must never be reclaimed, admission must fail while the source is absent)
are replaced by tests that a re-run source keeps the published fan-out for
renders and admission, and that a leftover lock does not block expansion.

Refs #1142

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ource retry

planDynamicExpansionRetryArchive refuses a manifest directory that holds
anything but the published manifests, so the .expansion.lock an older build
left behind after a crash (the failure the previous commit removes) made
`resume --retry-failed` of the dynamic source fail with
DYNAMIC_RETRY_EXPANSION_INVALID until an operator deleted it by hand. The
temporary file of an interrupted publishFileDurableExclusive has the same
effect.

Dot entries are never manifests (readExpansionManifests already skips them)
and the directory rename archives them along with everything else, so the
refusal now ignores them. Any other unrecognized entry is still refused; the
existing test for that case now uses a non-dot file.

Refs #1142

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The publication secret gate scans agent output in positive-only mode and
fails the node on any hit. Its supplemental JWT rule matched any three
dotted runs of base64url characters (20/10/10 minimum lengths), which also
describes ordinary qualified identifiers: a report citing
`ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck`, or a
JSON path such as
`auditProfileResolution.settingOrigins.property_priority_threshold`.
The gate rejected such a file after the agent had finished its work.

A JWT's header and payload are base64url-encoded JSON objects. In the
compact form that JWT libraries emit, each opens with '{"' and a letter,
which encodes to "eyJ". The rule now adds a (?=eyJ) lookahead to the
header and payload segments and keeps the old length minimums, so it
matches exactly the old matches whose first two segments begin with
"eyJ". The existing k07 fixture and an HS256 token are still rejected.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…nic and keys

Anvil prints the "test test test test test test test test test test test
junk" mnemonic and the 10 keys it derives at m/44'/60'/0'/0/0-9 on every
startup. Foundry tests and scripts sign with them: forge-std's own
test_DeriveRememberKey quotes the mnemonic and key (0). The fail-closed
publication gate rejected both. The mnemonic has a valid BIP39 checksum,
and `uint256 privateKey = 0xac09...ff80;` satisfies the context-labeled
private-key rule. So a generated test that derived or signed with a
default Anvil account failed its node's verification.

The labeled-key rule now skips those exact 10 keys (compared
case-insensitively, with or without 0x), and the mnemonic rule skips a
window that is exactly that phrase. The keys and the mnemonic were
checked against `anvil` 1.8.3's startup output and
`cast wallet private-key --mnemonic-index 0..9`. The keys are
public, so exempting them hides nothing. Any other labeled 64-hex key,
including key (0) with one digit changed, and any other valid mnemonic,
including one that shares the first 11 words, is still rejected.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…tions

The final-report agent copies a run summary that the host captures when the
report task starts, and both runtime presentations (the verified terminal
publication and the unchecked fallback report) republished that copy
unchanged. Elapsed time, models, tokens and spend therefore excluded the
report task itself and everything that finished after it started. On a local
smoke run the published report showed 2,363,772 tokens, $5.64 and 20m 17s,
while run.json recorded 4,591,556 tokens and $10.17 and the run finished
after 36m 23s.

projectTerminalReport and captureUnverifiedReport already read the final
run.json and state.json, so withWholeRunSummary() restates elapsed time
(run.json created_at to state.json finished_at) and models, tokens, spend
and partial_pricing from accounting.cumulative. A missing, malformed or
"unavailable" value keeps the agent's copy, and the helper never throws.
The agent's own report.json and report.md are untouched.

Unchecked reports also never received the run-root goal-search census, so
they always said "Goal search coverage is unknown". They now read
goal-search-coverage.json the same way the verified path does; an unreadable
census still renders as unknown coverage instead of failing the report.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…tput

projectCanonicalFinalReport re-validates the Markdown it just rendered with
regex rules written for hand-authored reports. publicProse did not escape
`](`, so upstream prose that the review gates require byte-for-byte (a spec
link, an image, `handlers[id](payload)`, a relative link whose `_` gets
escaped) tripped the link or image rule. The link rule also rejected the
renderer's own index anchors for any issue title with a non-ASCII letter,
and the public projection threw when a redacted path was followed by `(`
(`[redacted-path](line 12)`). The agent cannot repair byte-preserved fields,
so every retry failed the same way and the run ended without a report.

publicProse now escapes `(` after every `]`, so prose cannot form an inline
link or image, and the image and link rules are deleted; the raw-HTML rule
stays. Rendering of prose without `](` is unchanged, so existing reports
keep verifying.

The audit-context renderer (appendAuditContext, safeReportLink,
SAFE_REPORT_RELATIVE_LINK_PATTERN) is deleted: report@3 has no audit_context
field and rejects unknown keys, so it could never run. The prompt's
`## Audit context` instructions, which asked the agent to write a section
that the byte-exact canonical render never contains, are removed.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The artifacts reference claimed report.md links to THREAT_MODEL.md,
threat-model.json and goal-plan.json (only dead code ever did) and that
terminal presentations keep the report-start accounting snapshot. Document
what runtime presentations now restate, that prose link syntax renders as
literal text, and the upgrade effect on existing terminal publications.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Native continuation set ULTRAFUZZ_CONFIG_PATH to
<run>/smithers/resolved-config.json. Every stock agent adapter reads that
path with its TOML reader, which finds no [agents.*] tables in JSON and
returns an empty config, so each adapter fell back to its defaults on
resume: auth, api_key_env and config_dir were silently dropped. With the
default ultrafuzz.toml (CodexAgent auth = "api-key"), every resumed Codex
task ran with subscription auth and a cleared OPENAI_API_KEY.

Launch hands adapters a sealed copy of <run>/smithers/execution-config.toml,
the TOML rendering of the same resolved config. Resume now points them at
that file. It is written at compile time since #635, and resume only sets
the variable when resolved-config.json parses as the current v4 schema
(#1120), so every run that reaches this branch was compiled with it.

The new regression test resumes a run, rebuilds the generated CodexAgent
from the config path the runner received, and checks it still carries the
API key. It fails on main (CODEX_API_KEY is empty) and passes with this
change. "native resume delegates the persisted workflow after mutable
project sources are replaced" asserted the JSON path, i.e. the bug; it now
asserts the TOML path.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…bracket

Escaping `](` as `\](` would break the public projection on two shapes:
after publicProse doubles a prose backslash, `\\\](` matches the UNC
private-path pattern, and `token=REDACTED\](` extends the redacted
assignment value so the fixed-point re-scan redacts it again. Both public
projections succeed with the `]\(` escape; with `\](` this test fails.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ow render

renderAgentPrompt substituted the operator and task prompts into the
trusted agent-prompt template and then regex-scanned the whole result for
leftover `{{word}}` placeholders. Inserted text that merely contained such
a sequence threw inside the Smithers render: an escaped `\{{word}}` example
in a prompt template (the prompt renderer emits it as literal `{{word}}`),
the same escape inside a dynamic item value (kept verbatim), or an
operator `--prompt` note. Every render builds every task's prompt, so the
run failed, and failed again on every resume.

Check only the template's own placeholders, inside the replace callback,
as renderAgentPreambleTemplate already does. A template placeholder with
no value still throws, and now names the placeholder.

The dynamic prompt renderer also resolves `{{...}}` inside goal-plan
replacement values, and an unbound name there throws in the same render.
The goal-plan contract now rejects `{{` in replacement values, which are
plain-text labels, so that failure lands on goal-plan's own verify instead
of on every render. The check is a zod refinement; goal-plan.schema.json
is unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
prepareArtifactMirror spawned `ultrafuzz json validate` against the smoke
fixture in every prepare, in every agent-attempt reset, when an agent is
built after a controller restart, and again in the zero-retry verify task.
The trusted launcher's cold start was measured at ~35 s under contention
(#1026), so each spawn was another chance to fail an attempt, including
one whose agent work had already finished.

What the spawn proves -- that this process can launch the agent-facing
validator -- does not depend on the task: materializePromptSchemas has
just digest-checked the workspace's schema copy, and start-run already
runs the trusted-launcher preflight at launch and resume. Remember the
first success in the engine process. A failure is not remembered, so the
next caller spawns again.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…encies helper

The helper regex-scanned a Smithers inspect snapshot for the generated
workflow's "artifact dependency has not passed verification" message so a
resume could name the dependency behind a failed `prepare:` wrapper (the R43
shape behind #272). Nothing in src calls it: prepare-wrapper failures are now
attributed to their durable node (#288), and the lens sanitizer that caused
R43 was deleted (#558). Its only caller was its own unit test, which goes too.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…led node, and pin same-id recovery

A Smithers run can end `failed` without failing any task. That is a run-level
error such as WORKFLOW_RENDER_FAILED, which the runner raises when the
generated workflow throws while rendering, and it leaves durable work pending.
WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE could then only say "failing workflow
task(s): unreported", because parseCurrentSmithersInspect accepted
data.run.error and dropped it, and that error is the only record of why the
run stopped.

parseCurrentSmithersInspect now returns an optional runError read loosely from
data.run.error. It takes the cause's message (the runner records what was
thrown as `cause` under its own summary and `smithers up` recovery advice),
falling back to the message, redacted and scrubbed like other runner text and
capped at 1,000 characters. An unexpected shape yields nothing and never fails
the parse. The WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE message appends
"; workflow run error <code>: <message>". Its details, and so the durable
workflow-failure-unattributed event payload, are unchanged: that record stays
ids-only and needs no schema change.

A new integration test drives the pinned runner and Ultrafuzz's own resumeRun
(the `resume --force --retry-failed` entry) through this shape. While the
render-time cause persists, the resume is refused with
WORKFLOW_LIFECYCLE_FAILED and the run is left exactly as it was. Once the
cause is removed, the same run ID resumes and runs only the pending task, and
the finished producer is not run again. The test also checks the new reader
against the runner's real error JSON. No rewind, fork or replacement lineage
is added: each would re-render the same program and inputs and hit the same
cause (#272).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Agent retries used an exponential backoff from 1s, so a three-attempt
budget was spent in about 25 seconds, inside the minute a contended
Claude Code OAuth refresh can need to clear (#1084). The agent Task also
inherited Smithers' default identical-failure stall verdict (3), which
ended a chain before later same-agent attempts (exhaustive plans five) or
any [retry].agents fallback profile ran. Against real Smithers 0.35.0, a
[fail, fail, fail, fallback] chain ended `stalled` after attempt 3 and
never ran the fallback; with maxIdenticalFailures: 0 the fallback ran at
attempt 4 and the run finished.

- compileTask: initialDelayMs 60_000, so retries wait 60s, 120s, 240s,
  then Smithers' 300s cap.
- Agent Task: maxIdenticalFailures: 0, so the planned chain is the
  budget. It is set in the template only; the compiled manifest, cloud
  handoff schema and sealed task documents keep their exact shape.
- The real `smithers graph` smoke test never ran: it required a
  workspace-root .smithers install that no checkout has, and could not
  resolve the sealed module paths. It now uses the runtime package's own
  Smithers and asserts the retryPolicy Smithers receives.

Refs #1084

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…s diverged

`cancel` and `pause` are how an operator stops a run, but both read the
run through the strict, execution-grade mode of
`readLinkedWorkflowEvidence`. That mode exists to authorize executing
workflow code: it re-derives every sealed control document and binding
against the current build and throws on the first divergence. So after
a rebuild changed the validator build identity, or after an operator
hot-patched the published workflow, both commands refused with
WORKFLOW_CONTROL_EVIDENCE_INVALID and the run kept executing, while
`status` read the same run and `resume` handed it to Smithers.

Neither command executes workflow code; each only asks the workflow
runner to cancel or pause the linked run. Both now read evidence the
way `status` and `events` already do (observe-only, divergence
tolerant) and invoke the runner through the same published-snapshot
environment as before. A confirmed cancellation still persists
`canceled`. A run whose published execution snapshot files changed is
still refused, as `status` refuses it.

The two tests that asserted `cancelRun` refuses a diverged run now
assert that it cancels (and the contract-binding one also pauses). The
hand-patched-workflow test also pauses and cancels, and the
execution-file test pins that cancel still refuses there without
invoking the runner. The shared fake runner answers `cancel` with the
engine's confirmed-cancellation contract.

Refs #674, #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… on a failed deadline cancel

The workflow deadline is checked only when something synchronizes the run
(#1110). An operator-paused run executes nothing, and resuming it records a
new deadline, so cancelling it at the first observation past its deadline
bounded nothing. The control projection now leaves paused runs out of
deadlineExceeded.

A failed deadline cancel was an error diagnostic, so `status` returned
ok:false and `status --watch` stopped polling, and nothing asked again. It is
now a warning: the run stays active and the next synchronization requests
cancellation again.

Refs #1110

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…runner queries

Every control projection renews the controller lease and advances the
concurrency clock, and any difference counted as a change, so each `status`
poll rewrote state.json and a finished run kept an "active" lease renewed by
whoever looked at it (#1145). The projection now also reports whether those
clocks are its only change; an observe-only pass (status) and any pass over a
terminal run skip writing such a projection. Explicit passes over a live run
(syncRun, inspect, why, stats) still renew its lease, which the eval and
Modal pumps read as liveness.

Read-only runner queries had no timeout unless the caller passed one, so a
wedged runner process blocked status and the synchronization pumps
indefinitely. runSmithersInspectionCommand now bounds the runner process
(never executable preparation) at 120 s by default;
ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS overrides it, capped at 600 s. A timeout
becomes the existing failed-query diagnostic, with a message naming the bound.

The runner's orphan reason recommends `smithers supervise -r <id>`, which the
public rewrite turned into `workflow runner supervise -r <id>`, a command
nobody can run. It now recommends `ultrafuzz resume <run-id>`.

Refs #1145

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… redundant artifact events

status, inspect, why, stats, the dashboard and the eval runner all run the
full mutating synchronization, and nothing serialized them (#1152). Two passes
over one newly finished node both finalize it and race on its manifest temp
file, state.json and the journals; the issue analysis reproduced duplicate
node-synced events, a false failure that the immutable attempt ledger then
kept, and passes that died midway through a write group.

synchronizeLinkedWorkflowRun now runs each pass under a non-blocking
proper-lockfile lock on <run>/.workflow-sync. A pass that finds it held
returns ok with an info WORKFLOW_SYNC_IN_PROGRESS diagnostic and writes
nothing, so status never waits or fails on contention. Acquisition never
waits, so it cannot deadlock with the control lock, and its target is not the
run root because proper-lockfile keys its in-process registry by target and
the control lock already locks that path. The lock lives in a small wrapper
that re-enters the unchanged pass body.

node-artifacts-verified and node-artifacts-missing duplicated node
provenance, had no reader, and reported `missing: []` for every schema,
semantic or authority failure. They are no longer written; their schema
variants stay so existing journals replay. A failed node's durable last_error
now prefixes each error with its run-relative path and JSON pointer, never a
host path. The eventProvenanceForTask plumbing is deleted: createEventRecord
never persisted it.

Refs #1152

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…sh fails

status ran its observe-only synchronization strictly. A failed or malformed
runner events query made it return ok:false, exiting 1 and stopping --watch,
although the runner's own health was available, and any error thrown by the
refresh (for example a torn usage.jsonl append) escaped getRunHealth. stats
already tolerated the transient class.

The observe-only pass now sets tolerateInvalidEventStreams and downgrades the
same transient code set as stats; workflow-sync exports it once as
TRANSIENT_SYNC_DIAGNOSTIC_CODES for both. An error thrown by the refresh
becomes a WORKFLOW_STATE_SYNC_FAILED warning (the snapshot race keeps
WORKFLOW_STATE_SYNC_RACED), and status goes on to query the runner.

getRunHealth also passes the evidence it already verified into the pass
instead of reading and verifying the sealed snapshot a second time.

Refs #1145

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…nd deadline changes

configuration.md gains the ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS row and a
paragraph on skipped passes, status warnings and when state.json is left
untouched. The workflow_deadline_seconds paragraph from #1157 now says paused
runs are not cancelled and a failed cancel is a retried warning, and the
supervisor paragraph no longer claims that reporting checks the deadline.

Refs #1110, #1145, #1152

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…nd flush each once

Snapshot publication wrote every file with mode 0400 and fsynced it, then
the permission seal reopened every file, fchmodded it (to 0500 for
executables, otherwise to the 0400 it already had) and fsynced it again.
A launch of the test fixture's 12,552-file snapshot therefore issued
27,876 fsyncs, two per file plus two per directory, and launch time is
fsync-bound.

writeSnapshotFile now creates each file with its final mode and sets that
mode through the creating descriptor (the umask can clear creation bits)
before its single flush, so the bytes and the mode are durable together,
as the second flush previously made them. The permission seal now only
makes directories read-only; the publication verification that already
follows it still checks every entry's type, mode, link count and bytes.
The same fixture now issues 15,324 fsyncs (one per file plus two per
directory).

The new test republishes a launched run's snapshot and bounds its fsync
calls at one per file plus two per directory; origin/main fails it with
27,876 flushes for 12,552 files in 1,386 directories.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
doctor required the executable of every configured model profile, so a
fresh `ultrafuzz init` on a Codex-only host reported "needs attention"
because the scaffold's opt-in Claude, Kimi and Pi profiles were not
installed, although no node would run them. It now requires an agent's
CLI only when the selected topology or its [retry] agents fallback chain
can dispatch to that agent (the same selection the OpenRouter credential
check already used). Other configured profiles' CLIs are still probed and
listed, as not required, and the human output marks them "(not
required)". Agents that share a CLI (Codex and OpenRouter, Claude and
DeepSeek) require it if any of them is selected.

doctor also built DOCTOR_WORKFLOW_ENGINE_* diagnostics and check statuses
for the project-local engine and then discarded them, reporting both
checks as "unknown". The builders are deleted; layout_status and the
per-patch posture are still reported, and the reference docs no longer
list the five codes that were never emitted.

Launch and resume install the workflow engine controller under the OS
temporary directory, and a native resume keeps its install there for the
detached engine (#921 cost 1: a tmpfs /tmp exhausted a host's memory). A
new temporary-directory check warns, without failing the verdict, when
that directory is a tmpfs or has under 2 GiB free, and reports how many
ultrafuzz-controller-* directories it holds and their total size. It
never removes them, because a live engine may still use them.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…dentity

`resume --refresh-controller` renders its current controller into
<project>/.smithers/continuations/<uuid>/, but that directory was missing
from controllerOwnedGovernancePaths. Unless the project ignores .smithers,
the rendered files are untracked target files: the next launch in that
project records the target as dirty, which a private campaign (the
default policy) and every cloud launch reject, and the changed worktree
digest invalidates disclosure acknowledgements computed for the target.

The directory is now controller-owned, like .smithers/workflows. The new
test renders files in that layout and checks the target identity is
unchanged; origin/main reports the target dirty.

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ts schema

Runtime documents are parsed with a 100,000-item limit that the strict
JSON parser counts across the whole document. The control seal's schema
admits 100,000 execution files plus three identity arrays of up to
100,000 entries each, and launch writes the seal after schema validation
alone. A closure near the execution-file bound plus the run's task
bindings would therefore produce a seal that every later command fails
to parse. (An earlier measurement put the production closure at about
51,000 files; it was not re-measured for this change.)

The seal is now parsed with a 400,000-item limit, the schema's total; its
property total (four per execution file plus 35 fixed) already fits the
default limit. Byte limits are unchanged. The new test parses a seal with
the schema's maximum execution files and the fixture's bindings;
origin/main rejects it with "JSON exceeds the item limit of 100000".

Refs #921

Co-Authored-By: Claude Opus 5.5 <[email protected]>
agent-adapter-boundaries.test.ts pinned every adapter template's SHA-256
twice (a structural policy and a separate responsibility policy) plus
line and syntax-node ceilings. Those assertions check no behaviour: a
comment edit or the one-line missing `import path` fix in deepseek.tsx
failed the gate with "changed from its reviewed source fingerprint", so
every adapter fix needed two hash edits.

Merge the two tables into one policy per source (purpose, declared
responsibilities, upstream links) and delete the fingerprints, the
ceilings, and the tests that only exercised them. The real boundary
checks stay: every source needs a policy, only registered adapters may
own responsibilities, non-adapter helpers may carry no orchestration
signals, and a statically detected responsibility still fails until it
is declared.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…for undefined names

deepseek.tsx called path.join without importing path, so a render
without ULTRAFUZZ_CONFIG_PATH threw "path is not defined" from the
factory. Nothing caught it: runtime templates are copied into projects
and run by Bun, tsconfig only includes src/**/*.ts, and ESLint turns
no-undef off for every .ts/.tsx file.

Enable no-undef for packages/runtime/src/templates/**/*.tsx and declare
the __ULTRAFUZZ_*__ placeholders the compiler substitutes as readonly
globals. Against main the rule reports exactly this bug (deepseek.tsx
200:59 'path' is not defined) and nothing else.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…uts are optional

assertSmithersTaskManifestMatchesPlannedGraph re-derived which dependency
directories the compiler marks optional and required the manifest to
match exactly. That is a second copy of compiler policy: when #1120
narrowed the compiler to review-only opt-in, the copy kept the old rule,
and every launch of the default, low-cost, exhaustive and invariant-only
profiles failed with "optional dependency artifact directories do not
match" before any model ran (#1140). #1160 fixed that by adding a third
copy of the review rule to the gate, which leaves the next policy change
free to break launch, sync, cloud tasks and report completion again,
since all of them run this gate.

Keep only the safety property the gate can check without knowing the
policy: every optional dependency directory must belong to a producer in
a failure_policy: continue group, so a halting producer's output can
never become optional. Which consumers opt in stays compiler policy, now
in one exported helper (reconcilesPartialResults) that the static
compiler and dynamic lowering share. The gate also accepts a review task
that keeps a continuing input required, which is what dynamic lowering
emits for an empty expansion's source.

The #1160 unit test asserted the equality rule itself, so it is
rewritten to pin the safety property. A new runtime test plans and
compiles every packaged audit profile and runs the gate on the result;
against the pre-#1160 gate it rejects default, exhaustive,
invariant-only and low-cost, the #1140 regression.

Refs #1140, #1160

Co-Authored-By: Claude Opus 5.5 <[email protected]>
A launch that failed after its run directory existed was either never
recorded or recorded where no reader looked:

- planRun creates the run directory and then returns later failures
  (reference materialization, prompt rendering, source-ref publication)
  without recording them, and startRun's control-lock failure records
  nothing either. The run stays pending with an empty event journal, and
  status reports "launch-incomplete ... Launcher liveness is unknown ...
  wait for it to finish" indefinitely.
- startRun's catch does mark the run failed and appends a
  workflow-submit-failed event with the original error, but a run without
  a control seal is reported as WORKFLOW_CONTROL_SEAL_MISSING ("may
  predate sealed runs"), so status, events and why hide the error.
- resume of either run takes the lifecycle lock first, whose smithers/
  directory a pre-compile failure never created, and fails with an ENOENT
  lstat error.

recordLaunchFailure now writes the existing record (state failed plus the
workflow-submit-failed event) for planRun failures after the directory
exists, the control-lock failure, and the existing catch, where it
replaces a payload wrapper that could throw inside the catch. The event
contract has one code, so other codes are kept in the message. For a run
with no control seal, the linked-evidence reader and resume report the
recorded error as RUN_LAUNCH_FAILED. No new status, seal or schema.

The vulnerability-database reference check needs only the graph, and the
reference caches can be verified with the existing verifyReferencesCached
(previously unused), so both now run before createRunLayout and those
errors commit no run directory. The topology-summary invariant moves up
with them, since it is a throw that planRun's caller does not catch.

Closes #1140

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…authority

The host re-check of aggregate-test-files built its source bundles from
its own walk of every planned transitive ancestor, keyed by planned node
ID, with its own marker, manifest and prerequisite-chain authentication.
That walk disagreed with the in-workflow verifier in two ways:

- It read every generated-tests producer, including a failure_policy
  continue strategy or goal that failed before its verifier wrote a
  marker. The verifier had admitted the aggregation without that
  optional producer, but the host threw "artifact verification marker
  does not exist", failing aggregate-test-files and, through it, the
  final report.
- It looked a materialized dynamic producer up under its node ID
  ("dynamic:threat:<id>") instead of its storage attempt ID, and threw
  "contains unsafe segment".

Take the producers from finalizedDeclaredContractProducers instead: it
walks the attempt's sealed ancestor closure, skips optional producers
the verifier-persisted admission omitted, maps each attempt to its
planned node, and reads each producer through the canonical
loadFinalizedNodeOutputSnapshot authority. The module keeps only the
bundle projection, attributed to the producer's loop attempt index as
the verifier does. The bespoke authentication chain and its tests go
away; the canonical reader has its own coverage.

The artifact-gates fixture sealer also stops giving deferred dynamic
templates workflow bindings, which real graphs never carry and the task
manifest validator rejects.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 26 commits September 29, 2026 04:28
…e gate test

#1178 made attemptAuthority required for verifyRuntimeRequiredArtifactsForAttempt;
the unscoped-prose test from #1176 still used the three-argument form.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…on interrupt

Review fixes for the end-to-end campaign test.

- The no-rerun checks now run straight after the events poll, before any
  command synchronizes the run. Those checks cover the stub start counts,
  the RunStarted count, the untruncated event page and no NodeStarted after
  NodeFinished. First the test checks that the engine's first run-terminal
  event is RunFinished. A resume that re-runs a finished node now fails on
  the named count. Before, it failed with an opaque
  ARTIFACT_VERIFICATION_AUTHORITY_INVALID from the status call that came
  first.
- The stub now fails, and logs why, in three cases: it finds no output
  contract, a named authority file is missing, or the report render line
  is not all `--flag 'value'` pairs. Before, a missed contract match
  exited 0 with a successful Codex turn and wrote nothing. The test's
  "workflow stopped" and RunFinished failures include the stub's call log.
- SIGINT and SIGTERM handlers, and the exit hook, now SIGKILL the
  campaign's processes and delete the fixture; the signal handlers then
  re-raise. A Ctrl-C'd run used to leave the detached engine, the
  supervisor and the ~1 GB fixture behind.
- The wait for the held agent to exit counts a zombie as exited. kill(pid,
  0) succeeds on a zombie; its /proc cmdline is empty.
- Drops the stats `attempts_complete === true` pin. That field only
  says the attempt ledger exists, and pinning it would break a fix that
  reports the undercounted attempts as partial evidence.
- The todo reason now names the cause that outlasts #1186. Smithers
  emits no terminal event for the attempt it abandons at resume. The
  reason cites #1187.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The CHANGELOG entry said the CLI commands ran on the pinned Smithers
engine under Bun. They run as Node CLI processes; only the generated
workflow runs on the engine under Bun. The development guide gets the
same precise wording, plus one sentence on what the test does when it
is interrupted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…y ceiling

#1005 deliberately added no baseline or suppression registry, and #960 asks for
a complexity ceiling in the style of modem-dev/hunk#861: one global maximum set
at today's worst function, with nothing grandfathered, lowered as hotspots are
simplified. Drop eslint-suppressions.json and the always-on size budgets, keep
the diff-limited strict budgets for changed lines, and fail any function above
complexity 90 (synchronizeLinkedWorkflowRun is at 90 once the pending
simplification PRs land). The changelog entry moves to the consolidated
release notes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano

Copy link
Copy Markdown
Collaborator Author

Served its purpose: the combined CI run found five interaction failures; the 25 PRs are merged and the fixes are in the follow-up PR.

@aviggiano aviggiano closed this Sep 29, 2026
aviggiano added a commit that referenced this pull request Sep 29, 2026
…cord it in the changelog (#1191)

Test and lint fixes for five interactions between the 25 merged PRs #1163-#1188, found by the combined CI run in #1190, plus the batch's consolidated changelog entries.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant