Skip to content

refactor: delete dead prompt rename, config helpers, and CI scripts orphaned by #1131 - #1175

Merged
aviggiano merged 10 commits into
mainfrom
claude/w20b-prompts-config-ci-dead-code
Sep 29, 2026
Merged

aviggiano merged 10 commits into
mainfrom
claude/w20b-prompts-config-ci-dead-code

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Several features and helpers have no production caller but are still built, exported, documented and tested:

  • prompts: the prompt rename feature (rename.ts, renamePromptArtifactReferences, diffPromptIdentity, serializePromptDocument). Output-contract templates also had a second asset resolver that still probed dist/prompts, a layout the build stopped producing in fix(prompts): build bundled assets into the directory the resolver reads #796.
  • config: the prompt-metadata layer, restoreRedactedConfig, assertNoRedactionPlaceholders, applyDefaultProfileOverrides, assertResolvedConfigZod and resolvedConfigSchemaEntry. The prompt-metadata layer did run: resolveConfig applied it on every resolution. But createDefaultPromptMetadataLayer() always returned {}, and only one test passed promptMetadata, so in production it only ever applied an empty layer. docs/config.md described a restore step and a placeholder launch guard, and no code runs either one.
  • ci: outside tests, the only callers of scripts/ci/validate-threat-model-benchmark-gate.mjs and scripts/ci/modal-benchmark-control-window.mjs were eval-benchmarks.yml and the watch-modal-benchmark.sh it ran. Remove paid benchmark CI and repair CI gates #1131 deleted both. The gate's 333-line test still ran in every PR's test:ci-scripts step. The control-window script stayed alive only through a knip entry and a ci-config.test.ts block.
  • modal: exported constants and helpers that nothing references, or that only tests reference.
  • dashboard: commandCapabilities() returned seven flags as the constant false (restartWholeRun, doctor, triage, merge, restartFromNode, rerunSelectedNode, arbitraryShell). The frontend only declares them in a type, and capabilityForCommand never maps a command to any of them.
  • runtime: scripts/run-tests.mjs defines three selectors (materialize, clean, smithers) that nothing passes.

Root cause

These features were removed, or never wired up, but their library code and tests were left behind. CI cannot see this: pnpm -w knip runs with --include files,dependencies,unlisted,unresolved, so it does not report unused exports. And any file that a test imports still counts as used.

Change

One commit per area. Only deletions, apart from the resolver change.

  1. refactor(prompts):
    • Deletes the rename feature and its tests, along with the frontmatter round-trip support only rename needed: the allowUnknownFields option, unknownFrontmatter, rawFrontmatter, the invalid-rename code and TemplateOccurrence.rawName.
    • Output-contract templates now resolve through builtInPromptRoot(), the packaged-first lookup the agent-preamble templates already use.
    • Removes the render test that rebuilt the old dist/prompts layout. Its content assertions move into the packaged-layout test, which copies the real dist under a directory with no repository above it (the sealed-snapshot case).
  2. refactor(config):
    • Deletes the prompt-metadata layer, including PromptMetadataLayer, ResolveConfigInput.promptMetadata and the prompt-metadata diagnostic source.
    • Deletes restoreRedactedConfig and assertNoRedactionPlaceholders, plus their path helpers and the redaction diagnostic source.
    • Deletes applyDefaultProfileOverrides; its test now drives applyModelProfileOverrides, which validate.ts and plan-run.ts call. Also deletes assertResolvedConfigZod and resolvedConfigSchemaEntry.
    • resolvedConfigValidatorsAgree moves into its only test.
    • docs/config.md now states that no command restores values from the redaction manifest.
  3. ci: deletes both scripts, the gate's test, the knip.jsonc entry, and the control-window block (~lines 206-297) of the EVMBench full-mode test in packages/modal/test/ci-config.test.ts. The manifest assertions around that block stay.
  4. refactor(modal):
    • Deleted, no reference anywhere: MAX_MODAL_DOCUMENT_BYTES, MODAL_SMOKE_STOP_PATH, DEFAULT_NODE_TIMEOUT_SECONDS.
    • Deleted together with the tests that only exercised them: createTrackedSourceArchive, persistentWorkspaceRoot, and locateModalResumeWorkspace. The last one's "locates one exact linked durable run" test duplicated the existing tolerant-lookup test.
    • Moved into tests:
      • DEFAULT_BENCHMARK_MODELS becomes test/model-spec-fixtures.ts, typed as a const tuple.
      • modalBenchmarkConfigValidatorsAgree moves into its test.
      • WORKER_RESULT_ALLOWED_KEYS: the test now states the expected key set.
      • publicEvalFailureDiagnosticLogPayload: the same two calls runCommand already makes inline.
      • The two PUBLIC_*_MAX_PARALLEL_EVAL_ROWS aliases.
    • Un-exported five public-worker.ts constants and WORKER_STDERR_TAIL_BYTES; they are used only inside their own modules.
  5. refactor(dashboard): drops the seven flags from the server, dashboard-http.schema.json and the frontend type. The schema edit is required: commandCapabilities is closed (required plus additionalProperties: false).
  6. chore(runtime): drops the unused run-tests.mjs selectors.
  7. CHANGELOG.md entry.
  8. Review follow-ups:
    • refactor(prompts): leave getPrompt to the catalog work item restores packages/prompts/src/catalog.ts to origin/main. The earlier revision deleted getPrompt from it.
    • docs(changelog) tightens the entry's wording (see Verification).

Net: 47 files, +170 / −2164.

Deliberately not built (and why)

  • getPrompt (packages/prompts/src/catalog.ts) has no caller either. But fix(topology): restore the review group time budget and remove prompt/gate contradictions #1164 adds projectPromptsDifferingFromBuiltIns directly above it, and catalog.ts is outside this item's files. Deleting it here made git merge-tree report CONFLICT (content) in catalog.ts against fix(topology): restore the review group time budget and remove prompt/gate contradictions #1164, so it is left for a follow-up after fix(topology): restore the review group time budget and remove prompt/gate contradictions #1164 lands.
  • The retry-failure agent-preamble template is still referenced by smithers.ts (the __ULTRAFUZZ_RETRY_FAILURE_TEMPLATE__ injection) and by workflow.tsx:1666-1667, a few lines from the renderAgentPrompt region another work item is editing. Deleting only the prompts side would break workflow generation, so it stays.
  • createPublicEvalDiagnostics is test-only, but it is the pure seam that 23 calls in public-eval-diagnostics.test.ts use to reach the private createPublicEvalDiagnosticsFromRecords. Moving it into tests would mean either exporting the private function or rewriting those tests over file fixtures. Kept.
  • assertCurrentPersistentWorkerLineage stays exported. Un-exporting it touches a line that the strict diff lint rejects for an unrelated reason: it is async with no await.
  • Modal's triple 24 h sandbox-timeout aliases. Production uses two of the three names, so collapsing them is a rename, not a deletion.
  • hasRedactionPlaceholder (security): its only callers were the config helpers deleted here, so it now has none. packages/security belongs to another work item, so it is left in place.
  • No new knip exports gate. CI gating belongs to another work item. For reference, pnpm dlx [email protected] --include exports,types,duplicates reports 40 unused exports on origin/main and 33 on this branch, and nothing new. The seven that leave the report are all modal:
    • six are un-exported here: the five public-worker.ts constants and WORKER_STDERR_TAIL_BYTES;
    • the seventh, publicEvalFailureEnvelopeDiagnostics, is still exported. It leaves the report only because the rewritten public-worker.test.ts now imports it. Production uses it inside public-worker.ts.
  • Out of this item's scope:
    • the snake_case aliases on PromptConcreteNode/PromptGraphNode
    • config keys that are parsed but never read
    • RuntimeConfigOverrides.triageQuorum/triagePanelSize
    • dead exports in topology and references
    • the automatic-publication path in scripts/ci/prepare-eval-history-publication.mjs, also orphaned by Remove paid benchmark CI and repair CI gates #1131

Verification

Deletions. For every deleted or un-exported symbol, a grep over packages/, scripts/, docs/ and the runtime templates (excluding dist, dist-test and node_modules) finds no remaining reference, apart from the test-local helpers this PR adds.

At a49a626f^, the parent of #1131, the scripts' only invocations outside tests were:

  • .github/workflows/eval-benchmarks.yml lines 335, 408, 518 and 600 (control-window) and 735 (gate);
  • scripts/ci/watch-modal-benchmark.sh:24 (control-window), which eval-benchmarks.yml ran at lines 426 and 620.

#1131 deleted both files. The other two workflows #1131 deleted never invoked either script.

Discriminating evidence. This PR deletes code, and only two behaviours change.

  • Output-contract template lookup. The supported layouts (packaged dist/assets/prompts and the source tree) resolve as before. The packaged-layout test covers the sealed-snapshot case and passes. The deleted installed-layout test was the only thing that built dist/prompts.
  • Dashboard capabilities. Every dashboard test response is validated against the schema. As a mutation check, I re-added doctor: false to commandCapabilities() in the test build of the server, packages/dashboard/dist-test/src/index.js, which is the file the dashboard test imports.
    • With the mutation, serves logical topology flow with expanded attempt details fails with /capabilities additionalProperties: must NOT have additional properties. It passes before the mutation and again after the file is restored. I re-ran this on this revision.
    • The same edit to dist/index.js does not fail the test, because the test does not import that build.

Tests run on e6c395c1 (targeted). The follow-up commits only restore catalog.ts to origin/main and edit CHANGELOG.md, so they cannot affect any suite except prompts, which is re-run below.

  • prompts: full vitest suite, 98/98.
  • config: full vitest suite, 92/92.
  • modal: vitest over config, layout, resume, runner, worker-result, worker-lineage, workspace-config, public-worker, ci-config, modal-documents, smoke and worker-diagnostics.
    • Everything passes except three public-worker tests, which exceed their explicit 30 s budget: recognizes and cleans the legacy persistent public workspace…, rejects a present dangling public bundle… and accepts the bounded full lane….
    • The same three time out the same way on an origin/main worktree on this machine (load average 30-39), so the timeouts are not caused by this PR.
    • After the fixture-tuple change I re-ran config, worker-lineage, workspace-config, worker-result, layout and resume: 115/115.
  • dashboard: contracts 4/4; frontend-wire-contracts + dashboard 34/34.
  • pnpm -w test:ci-scripts: 94 pass, 1 fail.
    • The failure is safe-archive-bun.test.ts > repeatedly extracts a synthetic archive without stalling Bun callbacks, a 15 s timeout. It also times out on origin/main here, and this PR does not touch safe-archive.ts.
    • prepare-eval-history-publication.test.ts passes 17/17. It includes trustedCandidateRuntimePolicyDimensions(process.cwd(), "smoke"), which parses the real public-worker.ts after the un-exports.
  • runtime: dynamic-expansion 15/15. It materializes dynamic children, which calls renderPrompt and so loads the output-contract templates through the new lookup. prompt-artifact-authority + model-override-validation 10/10.

Re-run on this revision (d7704ae7):

  • prompts: full vitest suite 98/98, and tsc --noEmit.
  • npx prettier --check and npx eslint --max-warnings 0 over all 47 changed files.
  • CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, pnpm -w knip, node scripts/docs-check.mjs and node scripts/prompt-catalog-docs.mjs --check.
  • the knip exports comparison above, run on both trees;
  • the dashboard mutation check above.
  • git merge-tree --write-tree against the 25 other open work-item branches. The only conflict is CHANGELOG.md, which every entry touches. The catalog.ts conflict with fix(topology): restore the review group time budget and remove prompt/gate contradictions #1164 is gone.

Static checks on e6c395c1:

  • pnpm --filter @ultrafuzz/{prompts,config,dashboard,modal,runtime} typecheck.
  • tsc --noEmit for topology, references, runtime, evals, evmbench and cli.

Not run: the full runtime and CLI suites, CI itself, and any real campaign.

Risk / compatibility

  • Removed library exports. @ultrafuzz/prompts, @ultrafuzz/config and @ultrafuzz/modal are private workspace packages, and no in-repo caller of a removed export remains. parsePromptFrontmatter no longer takes an options argument; unknown frontmatter keys already failed by default, and no caller passed allowUnknownFields.
  • Resolver error wording. builtInPromptRoot() rethrows non-ENOENT stat errors, where the old output-contract resolver moved on to its next candidate. Rendering fails either way; only the error message differs.
  • In-flight runs. None of the five validator modules changed, so the validator build identity is unchanged. No artifact schema changed. The dashboard HTTP schema feeds neither the validator build identity nor the artifact schema bundle digest that planned outputs and the trusted CLI identity bind to.
  • Execution snapshots. Snapshots copy each package's whole dist. Since fix(prompts): build bundled assets into the directory the resolver reads #796 that contains dist/assets/prompts and no dist/prompts, so dropping that probe matches every current build.
  • Parallel work in shared files.

Refs #462

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no actionable new issue was established.

Summary

This PR removes unused prompt, configuration, Modal, CI, dashboard, and runtime code. Output-contract templates now use the shared packaged-first prompt resolver, and the dashboard contract drops seven always-false capability flags. A follow-up removes this PR’s changelog entry for consolidation elsewhere.

Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."

aviggiano and others added 7 commits September 28, 2026 23:51
…ate template resolver

renamePromptId, renameConcreteNodeId and rewriteArtifactPathForPromptId
(rename.ts), their render.ts helpers (renamePromptArtifactReferences and
three private rewriters) and the frontmatter round-trip support they
needed (serializePromptDocument, diffPromptIdentity, the
allowUnknownFields option, unknownFrontmatter/rawFrontmatter and the
invalid-rename error code) had no caller outside their own tests. getPrompt
had no caller at all. TemplateOccurrence.rawName existed only so rename
could preserve whitespace.

Output-contract templates had their own three-candidate resolver whose
middle candidate, dist/prompts, is the layout the build stopped producing
in #796. They now resolve through builtInPromptRoot(), the packaged-first
lookup the agent preamble templates already use, so both kinds of prompt
asset follow one rule. The packaged-layout render test, which copies the
real dist below a directory with no repository above it, still covers the
sealed-snapshot case; the installed-layout test that rebuilt the old
dist/prompts layout is removed and its content assertions are folded into
the packaged-layout test.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…o production caller

- The prompt-metadata layer: resolveConfig applied
  createDefaultPromptMetadataLayer(), which always returned {}, and an
  input.promptMetadata that only one test supplied. PromptMetadataLayer,
  ResolveConfigInput.promptMetadata and the "prompt-metadata" diagnostic
  source go with it.
- restoreRedactedConfig and assertNoRedactionPlaceholders: nothing in
  runtime, cli, modal or evals calls them. Runs launch from the unredacted
  resolved config; the redacted copy and its manifest are written for
  readers of the run directory, and no command restores values from them.
  docs/config.md no longer describes a restore step or launch guard that
  does not exist. The "redaction" diagnostic source goes with them.
- applyDefaultProfileOverrides: a one-line wrapper only tests called.
  Its test now drives applyModelProfileOverrides, which validate.ts and
  plan-run.ts use.
- assertResolvedConfigZod and resolvedConfigSchemaEntry: no callers.
- resolvedConfigValidatorsAgree: moved into the schema test, its only
  user.

redactResolvedConfig and serializeRedactedResolvedConfigToml, which
plan-run.ts and init.ts use, are unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
#1131 deleted eval-benchmarks.yml, the only workflow that invoked
scripts/ci/modal-benchmark-control-window.mjs (create/deadline) and
scripts/ci/validate-threat-model-benchmark-gate.mjs. Since then the
control-window script was kept alive only by a knip entry and a block of
ci-config.test.ts that exercised it, and the threat-model gate only by its
own test, which still ran in every PR's test:ci-scripts step.

Delete both scripts, the gate's test, the knip entry, and the
control-window block of the EVMBench full-mode manifest test; the
manifest assertions around it stay.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…rs into tests

Deleted, no reference anywhere: MAX_MODAL_DOCUMENT_BYTES,
MODAL_SMOKE_STOP_PATH (smoke-worker.ts builds its own stop path) and
DEFAULT_NODE_TIMEOUT_SECONDS.

Deleted with the tests that only exercised them:
- createTrackedSourceArchive; production archives the exact candidate
  with createExactCandidateSourceArchive.
- persistentWorkspaceRoot.
- locateModalResumeWorkspace, documented as "retained for tests"; worker.ts
  uses findModalResumeWorkspace, and the removed "locates one exact linked
  durable run" test duplicated the existing tolerant-lookup test. The two
  finalize tests unwrap findModalResumeWorkspace through a local helper.

Moved into tests, their only users:
- DEFAULT_BENCHMARK_MODELS, a stale model list no production code reads,
  is now test/model-spec-fixtures.ts, typed as a const tuple so indexed
  fixtures need no non-null assertion.
- modalBenchmarkConfigValidatorsAgree, WORKER_RESULT_ALLOWED_KEYS (the
  test now states the expected persisted key set), and
  publicEvalFailureDiagnosticLogPayload, which composed the same two calls
  runCommand already makes inline.
- PUBLIC_BENCHMARK_MAX_PARALLEL_EVAL_ROWS and its full-lane twin, aliases
  of the evals lane limits that only a test imported.

Un-exported, used only inside their own module: five public-worker
constants and WORKER_STDERR_TAIL_BYTES. The source-constant reader in
prepare-eval-history-publication.mjs matches top-level const
declarations with or without export, and its test against the real
public-worker.ts still passes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
commandCapabilities() returned restartWholeRun, doctor, triage, merge,
restartFromNode, rerunSelectedNode and arbitraryShell as the constant
false. The frontend only declared them in its CommandCapabilities type:
capabilityForCommand never maps a command to any of them, so nothing read
them.

Remove them from the server, the dashboard HTTP schema's
commandCapabilities definition, and the frontend type. The schema edit
is required because that definition is closed (required plus
additionalProperties: false). The dashboard schema feeds neither the
validator build identity, which hashes only the artifacts validator
modules and the ajv versions, nor the artifact schema bundle digest that
planned outputs and the trusted CLI identity bind to.

The flow test's two assertions that doctor and merge are false are
removed. Every dashboard HTTP response in the tests is still validated
against the schema, so a server that emitted one of these keys again
would fail the flow test with "must NOT have additional properties"
(checked by re-adding doctor: false to the compiled server).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
packages/runtime/package.json calls scripts/run-tests.mjs with no
selector or with "supporting"; nothing passes "materialize", "clean" or
"smithers". A single test file can still be selected by name through the
existing fallback.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 28, 2026 23:55
aviggiano and others added 2 commits September 29, 2026 02:50
getPrompt still has no caller, but #1164 adds
projectPromptsDifferingFromBuiltIns directly above it in catalog.ts, so
deleting it here made the two PRs conflict textually
(git merge-tree reported CONFLICT (content) in catalog.ts). catalog.ts is
outside this change's files, so restore it to origin/main and leave the
deletion for a follow-up once #1164 lands.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…idence

Drop getPrompt, which this change no longer deletes. Name the actual
callers of the two CI scripts (eval-benchmarks.yml and the
watch-modal-benchmark.sh it ran, both deleted in #1131) instead of "the
workflows #1131 deleted": the other two workflows #1131 removed never
invoked either script. Stop describing the config prompt-metadata layer as
code no production path calls: resolveConfig ran it on every resolution,
only ever with an empty layer. Mention the run-tests.mjs selectors and the
docs/config.md change, and add the [runtime] and [docs] tags for them.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…vers

Move projectPromptsDifferingFromBuiltIns below builtInPromptRelativePaths
so it no longer sits in the hunk where #1175 deletes getPrompt; with that
placement the two branches merge without a conflict in catalog.ts. No
behaviour change.

The anchor test comment claimed a prompt and its topology "cannot drift
apart again". The test checks only nodes whose built-in prompt promises
an extended timeout (triage and severity-classification) in the packaged
topologies, so say that, and that dedupe-findings is not covered.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano added a commit that referenced this pull request Sep 29, 2026
…/gate contradictions (#1164)

* fix(topology): restore the review group time budget

The triage and severity-classification prompts tell the agent that "the
topology gives this review node an extended timeout", but no shipped
topology did. The review group lost its pin in v0.0.2, so dedupe,
triage, severity classification, test aggregation and the final report
ran on run.default_timeout_seconds (3600) while every goal, strategy and
specialist lane had 7200. The review chain is the widest fan-in, so a
review stage that needs more than the default window times out
deterministically and is retried with the same window; the review group
halts on failure, so every later review stage and the final-report node
are then skipped (#1150).

Pin the review group in the four non-smoke topologies to the same 7200
seconds as the goal, strategy and specialist groups. smoke.yml has no
triage or severity node and is left alone. Like the existing group pins,
this one takes precedence over run.default_timeout_seconds and
model-profile timeouts (#675).

The new packaged-topology test expands every shipped topology and holds
each node whose prompt promises an extended timeout to a window above
the largest shipped default. It fails on the previous topologies with
"default.yml node `triage` promises an extended timeout but resolves to
ultrafuzz.toml `run.default_timeout_seconds`=3600".

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* chore(topology): drop the unused timeout on pinned reference nodes

Every property reference node in the four non-smoke topologies set
timeout_seconds: 300 (36 lines). Reference nodes are materialized by
plan-run before launch and compileSmithersWorkflow only compiles agentic
nodes into Smithers tasks, so the value never bounded anything. It was
only copied into the expanded graph fingerprint and the planned graph,
and displayed by the dashboard. The vulnerability-database reference
node already had no timeout.

Existing project topologies that still carry the lines keep validating;
the reference docs now say the field has no effect on reference nodes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(prompts): stop inviting production interface edits the handoff rejects

The stateful-invariant setup and implement-properties prompts told the
agent not to edit production src/ or contracts/ "except for interfaces
if they are genuinely required by the harness". Both nodes publish a
workspace patch, and captureWorkspacePatch rejects every change under
the configured production source roots (src and contracts by default),
added files included:

  source-snapshot violation: workspace patch modifies protected
  production source: src/interfaces/IVault.sol

So with the default roots, following the exception failed the handoff,
and a retry starts a fresh session from the same prompt. Remove the
exception and tell the agent to declare harness interfaces in the test
tree, which the same capture accepts.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): warn at validate time when a project prompt differs from its built-in

Runs use the project copy of every prompt, and ultrafuzz init without
--force keeps an existing copy. After an upgrade a scaffolded copy
therefore keeps an older release's text while the artifact gates move
on, and nothing said so until a node failed its gate at the end of an
attempt.

projectPromptsDifferingFromBuiltIns lists the project prompts whose bytes
differ from the built-in prompt at the same path. validate reports each
one as a PROMPT_DIFFERS_FROM_BUILT_IN warning, which sets the prompts
posture to warn without failing validation, planning or the dashboard's
prompt save. A freshly scaffolded project and prompts the project adds
at new paths produce no warning.

The new runtime test fails on the previous validate.ts (prompts posture
'pass' where 'warn' is expected) and also checks the remedy the warning
names: deleting the copy and rerunning init restores a passing posture.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: changelog for the review timeout and prompt/gate fixes

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(prompts): tell Vyper setup agents to declare interfaces in the test tree

The setup-foundry and base-test-setup prompts ask for Solidity interfaces
to the Vyper contracts' ABI but never say where to put them. Both nodes
publish a workspace patch, and captureWorkspacePatch rejects any change
under the production source roots (src and contracts by default), so an
interface written under contracts/interfaces/ fails the setup handoff,
and the setup group halts the run on failure. Say to declare new
interfaces in the test tree, as the stateful-invariant prompts now do.

Whether agents actually chose a production location here was not
observed; this closes the same gap the invariant prompts had.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* fix(runtime): doctor no longer summarizes a validation with warnings as a pass

doctor reported the validate check as `warning` but kept the summary
"config, topology, prompts, paths, agents, and trust posture pass". Now
that validate warns about project prompts that differ from their
built-in, every upgraded project with such a copy would see that
contradiction. The summary now says validation passed with warnings and
points to `ultrafuzz validate --json`, since text-mode output does not
print warnings.

The new test fails on the previous doctor.ts with actual 'config,
topology, prompts, paths, agents, and trust posture pass'.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* chore: move the prompt drift helper and state what the anchor test covers

Move projectPromptsDifferingFromBuiltIns below builtInPromptRelativePaths
so it no longer sits in the hunk where #1175 deletes getPrompt; with that
placement the two branches merge without a conflict in catalog.ts. No
behaviour change.

The anchor test comment claimed a prompt and its topology "cannot drift
apart again". The test checks only nodes whose built-in prompt promises
an extended timeout (triage and severity-classification) in the packaged
topologies, so say that, and that dedupe-findings is not covered.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs: record the review pin's precedence break and the topology refresh

A group pin takes precedence over run.default_timeout_seconds and
model-profile timeout_seconds (#675), so a project that raised either
above 7200 now gets 7200 on review nodes of the packaged exhaustive and
invariant-only topologies and of new scaffolds, and a cloud config whose
explicit global resource timeout is below 7200 needs per-node overrides
for the review nodes as well. Record that under Breaking changes with
the escape hatch.

Tighten the Other changes entry: name the three pinned topologies instead
of "the shipped topologies" (smoke stays unpinned), and give the one-command
refresh for an uncustomized project topology, which init keeps. The
how-to now says the same about the topology as it already did about
prompts.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 3fa57d0 into main Sep 29, 2026
13 checks passed
@aviggiano
aviggiano deleted the claude/w20b-prompts-config-ci-dead-code branch September 29, 2026 06:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant