Skip to content

fix(topology): restore the review group time budget and remove prompt/gate contradictions - #1164

Merged
aviggiano merged 10 commits into
mainfrom
claude/w12-topology-prompt-fixes
Sep 29, 2026
Merged

aviggiano merged 10 commits into
mainfrom
claude/w12-topology-prompt-fixes

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #1150
Refs #675

Problem

  • Review stages were under-budgeted. In the four non-smoke topologies (.ultrafuzz/topology.yml and the packaged default, exhaustive and invariant-only), the review group (dedupe-findings → triage → severity-classification → aggregate-test-files → final-report) had no timeout pin. It ran on run.default_timeout_seconds (3,600), while every goal, strategy and specialist lane ran with 7,200. review/triage.md and review/severity-classification.md still tell the agent "The topology gives this review node an extended timeout", which was false. Expanding main's default.yml gives review:3600 ×5, strategies:7200 ×39, goals:7200 ×3, specialists:7200 ×5. Dedupe is the widest fan-in. When a review stage needs more than an hour, it times out every time and is retried with the same budget. The review group halts on failure, so shouldSkipWorkflowTask then skips every later review stage and the final-report node. That is the "one-hour outer timeout" repro in Checkpoint long fan-in stages and recover according to failure type #1150.
  • Stateful prompts contradicted the gate. strategies/invariants/setup.md and implement-properties.md allowed production edits "except for interfaces if they are genuinely required". Both nodes publish workspace.patch, and captureWorkspacePatch rejects every change under the production source roots, added files included. Following the prompt therefore failed the handoff, and a retry starts from the same prompt. The two Vyper setup prompts (setup/prepare-foundry-harness.md, setup/discover-base-test.md) ask for Solidity interfaces without saying where to put them; their nodes publish workspace.patch too, and the setup group halts the run on failure.
  • Prompt drift was invisible. Runs use the project copy of each prompt, and ultrafuzz init keeps existing copies unless run with --force. After an upgrade, stale copies ran against newer gates, and nothing warned about it until a node failed its artifact gate.
  • Dead config. 36 timeout_seconds: 300 lines sat on reference nodes that never run as tasks.

Root cause

  • The history bears out the drift. At v0.0.1, triage, severity-classification, aggregate-test-files and final-report each pinned timeout_seconds: 3600. v0.0.2 removed those pins, and the "extended timeout" sentence first appears in v0.0.3 (02d8927). timeoutSecondsFor (topology expand.ts) reads only a node pin or a group pin. buildSmithersTask then falls back to the model-profile timeout, then to run.default_timeout_seconds.
  • The interface exception predates the source-snapshot guard in workspace-handoff.ts (assertProductionSourcePreserved), and the prompts were never updated to match it.

Change

  1. Review group pin. Add defaults: { timeout_seconds: 7200 } to the review group in .ultrafuzz/topology.yml and packages/config/topologies/{default,exhaustive,invariant-only}.yml. default.yml stays byte-identical to .ultrafuzz/topology.yml. smoke.yml is untouched because it has no triage or severity node. 7,200 is the value the existing two-sided guard ("pins every agentic timeout to the reviewed window…") requires of every agentic group pin.
  2. Reference-node timeouts. Remove the 36 timeout_seconds: 300 lines. Nothing uses them: plan-run materializes reference nodes before launch (materializeReferenceNodesForPlan), and compileSmithersWorkflow compiles only kind === "agentic" nodes. The value was only copied into the graph fingerprint and planned graph, and shown by the dashboard. topology-yaml.md now says the field has no effect on reference nodes. Existing topologies that still carry the lines keep validating.
  3. Prompts. Delete the interface exception from both invariant prompts. They now say the handoff rejects every change under the production source roots, and tell the agent to declare harness interfaces in the test tree instead. The two Vyper setup prompts get the same rule in one sentence each.
  4. Drift warning. A new projectPromptsDifferingFromBuiltIns(catalog) in @ultrafuzz/prompts lists project prompts whose bytes differ from the built-in prompt at the same path. validate emits one PROMPT_DIFFERS_FROM_BUILT_IN warning per such file and sets the prompts posture to warn. Warnings do not fail validate, plan, doctor or the dashboard's prompt save. The text summary gives the count and points to --json for the paths, because text-mode validate prints diagnostics only on failure. doctor's validate check used to keep the summary "config, topology, prompts, paths, agents, and trust posture pass" next to a warning status; it now says "configuration validation passed with warnings; run ultrafuzz validate --json for detail". Documented in docs/reference/cli.md and docs/how-to/edit-prompts-topology.md.
  5. Upgrade notes. CHANGELOG gains a Breaking-changes entry for the goals/strategies group timeout_seconds: 7200 shadows the profile default and kills high-reasoning nodes at 2h with total loss of node work #675 precedence effect (see Risk) with its escape hatch. The Other-changes entry gives ultrafuzz topology copy default .ultrafuzz/topology.yml --force as the refresh for an uncustomized project topology, which init keeps; the how-to says the same.

Deliberately not built (and why)

Verification

Discriminating tests, each run by me on the rebased branch (base b6dd1da) failing without the change and passing with it:

  • packages/topology/test/packaged-topologies.test.ts → "gives every node whose prompt promises an extended timeout more than the default". It expands every packaged topology with the real expandTopology. For each node whose built-in prompt matches /extended\s+timeout/, the resolved timeout must exceed the largest shipped default. It also fails if no node promises one, so it cannot pass vacuously. With origin/main's three packaged YAMLs swapped in, it fails with default.yml node `triage` promises an extended timeout but resolves to ultrafuzz.toml `run.default_timeout_seconds`=3600: expected 3600 to be greater than 3600. With the branch's topologies, all 15 tests in the file pass.
  • packages/runtime/test/prompt-catalog-validation.test.ts (new, 2 tests).
    • "validate warns about a project prompt that differs from the built-in prompt at its path": a fresh scaffold → posture pass and no warning. Edit review/triage.md and add strategies/project-only.md → ok: true, posture warn, exactly one warning, for review/triage.md. Delete the edited file and rerun initProject → pass again, which checks the remedy the warning names. With origin/main's validate.ts it fails with actual: 'pass', expected: 'warn'.
    • "doctor does not summarize a validation that only warned as a pass". With origin/main's doctor.ts it fails with actual: 'config, topology, prompts, paths, agents, and trust posture pass'. With origin/main's validate.ts it fails earlier, with actual: 'ok', expected: 'warning'.

Supporting evidence (not discriminating):

  • Expansion with the built expander. Default, exhaustive and invariant-only: all five review nodes resolve to 7,200. Smoke: dedupe-findings and final-report stay unset. In default, the setup (6) and properties (9) groups stay unset (3,600 default), as before.
  • Upgrade simulation with the built CLI. I ran init on a scratch repo, then put back origin/main's default.yml and the four edited prompts. validate exits 0 with prompts: warn and four PROMPT_DIFFERS_FROM_BUILT_IN diagnostics, and doctor prints validate: warning - configuration validation passed with warnings; run ultrafuzz validate --json for detail. ultrafuzz topology copy default .ultrafuzz/topology.yml --force then makes the file byte-identical to the packaged default, validate still passes, and all five review nodes expand to 7,200.
  • captureWorkspacePatch against a scratch repo using the built runtime (implementer's probe; I did not re-run it): adding src/interfaces/IVault.sol fails with source-snapshot violation: workspace patch modifies protected production source: src/interfaces/IVault.sol, and adding test/recon/interfaces/IVault.sol is accepted.
  • git merge-tree against every other claude/* branch: CHANGELOG.md is the only conflict. Before the helper moved, packages/prompts/src/catalog.ts also conflicted with refactor: delete dead prompt rename, config helpers, and CI scripts orphaned by #1131 #1175.

Ran, all passing:

  • vitest: @ultrafuzz/topology 8 files / 74 tests, @ultrafuzz/prompts 6 / 104 (including semantic-anchors), @ultrafuzz/config 5 / 94.
  • node:test runtime:
    • runtime.test.js subset, 18/18: all validate … tests, init preserves existing project-owned files…, the four a clean scaffold … tests (one plans the full shipped default topology), the shipped default topology expands…, plan materializes pinned reference nodes…, plan provisions a validated trusted expectation catalog… and planned graphs persist topology overrides….
    • lifecycle-inspection.test.js diagnoseProject*, 12/12.
    • prompt-catalog-validation 2/2 and model-override-validation 4/4.
  • node:test CLI: init and validate emit schema-versioned launch JSON and doctor reports install posture in human and JSON output.
  • Checks: npx prettier --check and npx eslint on the changed files, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, pnpm -w knip, typecheck for prompts, runtime and topology, the three docs:check scripts (the audit-profile and prompt-catalog generators produce no diff), and node scripts/validate-audit-profile-package.mjs.

Not run: the full runtime and CLI suites, and the modal, evals, evmbench and dashboard suites.

CI. "External static analysis" fails on this PR for a repo-wide reason. Super-linter's markdownlint checks every changed file in full, and CHANGELOG.md already had many lines over MD013's 400-character limit when super-linter was added (#1005). Any PR that adds a changelog entry therefore fails, release-validation (runtime shards, runtime-supporting, CLI) is skipped, and release-gates fails. #1184 fixes this once with .github/linters/.markdown-lint.yml (MD013 off); its External static analysis passes with a CHANGELOG edit. This PR does not duplicate that change. Until #1184 merges and CI is re-run here, the runtime and CLI suites have not run in CI for this PR. The new runtime test runs in the runtime-supporting lane.

Risk / compatibility

  • Shadowing (goals/strategies group timeout_seconds: 7200 shadows the profile default and kills high-reasoning nodes at 2h with total loss of node work #675), now under Breaking changes. Like the goal, strategy and specialist pins, the review pin wins over run.default_timeout_seconds and model-profile timeout_seconds. A project that raised either above 7,200 now gets 7,200 on review nodes when it runs a packaged topology (exhaustive, invariant-only) or a new scaffold. The CHANGELOG gives the escape hatch: copy the topology, raise the pin, and set topology_path. For everyone on the shipped 3,600 default, review goes from 1h to 2h.
  • Cloud caps, also under Breaking changes. In cloud mode, an explicit global execution.resources.timeout_seconds below 7,200 already fails launch (CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED) on the strategy nodes. A config that works around this with per-node overrides on the pinned nodes now needs them on the review nodes as well.
  • Cost. A review agent that truly hangs now burns up to 2h per attempt instead of 1h.
  • Modal benchmark lanes (by code reading only; no Modal run).
    • The public lanes write run.default_timeout_seconds = node_timeout_seconds (1,800) and do not cap topology pins. threat-model runs the default topology from a fresh init --force scaffold and full runs exhaustive, so each of their five review stages can now take up to 2h per attempt instead of 30 min, inside the 15,000-second row envelope.
    • The private-lane recovery overseer replaces a live owner as owner-stalled when neither a run-state transition nor a node success has happened within max(30 min, node_timeout_seconds + 15 min). With node_timeout_seconds below 7,200 (the documented example uses 7,200), a single review stage that runs past that window could have its owner replaced mid-attempt. The strategy, goal and specialist pins already have the same exposure.
    • evmbench caps every topology timeout_seconds to its node_timeout_seconds, so its review budget is unchanged.
  • Existing projects. A scaffolded .ultrafuzz/topology.yml (default and low-cost profiles) is not rewritten. It keeps the 1h review budget until it is refreshed (ultrafuzz topology copy default .ultrafuzz/topology.yml --force if uncustomized) or defaults: { timeout_seconds: 7200 } is added to its review group. Projects also keep their old copies of the four edited prompts. The new validate warning flags those copies, along with any other stale or customized prompts. Expect warnings, not failures, after upgrading.
  • Digests. The packaged default, exhaustive and invariant-only topology digests change. In-flight runs are unaffected, because resume uses the sealed plan and does not re-expand the topology. An eval run launched before this merges and published after it would fail the public-history "current packaged policy" check, as with any topology change.
  • No schema, validator-build or contract-description changes.
  • History. The branch was rebased onto b6dd1da and force-pushed to reword one commit message ("the other agentic groups" became "the goal, strategy and specialist groups"). The original five commits' content is unchanged, and four review-fix commits follow them.

🤖 Generated with Claude Code

RetriggerConfidence Score: 3/5

The PR does not appear safe to merge while the previously identified review-timeout gaps remain.

Fix All in Claude CodeFindings

  1. P1 Existing projects keep shorter timeouts ▶
  2. P1 Review pin exceeds cloud caps ▶
Fix with agent prompt
### Issue 1
.ultrafuzz/topology.yml:43-45
Default and low-cost runs use each project's editable `.ultrafuzz/topology.yml`, and `ultrafuzz init` preserves that file unless run with `--force`. Projects created before this change therefore do not get the new review pin. Their review stages still use the existing 3,600-second default, so the one-hour timeout this PR aims to fix persists for those runs.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

### Issue 2
packages/config/topologies/default.yml:44-45
If a cloud project has an explicit resource timeout below 7,200 seconds and per-node caps for the already-pinned strategy tasks, its review tasks now exceed that resource timeout. Task compilation throws `CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED`, so a previously valid run cannot launch. The same review pin is added to the exhaustive and invariant-only topologies.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR extends the review group's timeout, aligns harness prompts with the production-source handoff, and warns when project prompts differ from built-ins.

  • It also removes ineffective reference-node timeout pins and adds validation and topology tests.
  • The latest change removes the release notes explaining the review timeout's breaking configuration effects.

Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."

@aviggiano
aviggiano requested a review from a team as a code owner September 28, 2026 22:54
Comment thread .ultrafuzz/topology.yml
Comment on lines 43 to +45
color: "#0f766e"
defaults:
timeout_seconds: 7200

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Existing projects keep shorter timeouts

Default and low-cost runs use each project's editable .ultrafuzz/topology.yml, and ultrafuzz init preserves that file unless run with --force. Projects created before this change therefore do not get the new review pin. Their review stages still use the existing 3,600-second default, so the one-hour timeout this PR aims to fix persists for those runs.

Prompt To Fix With AI
This is a comment left during a code review.
Path: .ultrafuzz/topology.yml
Line: 43-45

Comment:
**Existing projects keep shorter timeouts**

Default and low-cost runs use each project's editable `.ultrafuzz/topology.yml`, and `ultrafuzz init` preserves that file unless run with `--force`. Projects created before this change therefore do not get the new review pin. Their review stages still use the existing 3,600-second default, so the one-hour timeout this PR aims to fix persists for those runs.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Claude Code

Comment on lines +44 to +45
defaults:
timeout_seconds: 7200

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Review pin exceeds cloud caps

If a cloud project has an explicit resource timeout below 7,200 seconds and per-node caps for the already-pinned strategy tasks, its review tasks now exceed that resource timeout. Task compilation throws CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED, so a previously valid run cannot launch. The same review pin is added to the exhaustive and invariant-only topologies.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/config/topologies/default.yml
Line: 44-45

Comment:
**Review pin exceeds cloud caps**

If a cloud project has an explicit resource timeout below 7,200 seconds and per-node caps for the already-pinned strategy tasks, its review tasks now exceed that resource timeout. Task compilation throws `CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED`, so a previously valid run cannot launch. The same review pin is added to the exhaustive and invariant-only topologies.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

aviggiano and others added 9 commits September 29, 2026 02:42
The triage and severity-classification prompts tell the agent that "the
topology gives this review node an extended timeout", but no shipped
topology did. The review group lost its pin in v0.0.2, so dedupe,
triage, severity classification, test aggregation and the final report
ran on run.default_timeout_seconds (3600) while every goal, strategy and
specialist lane had 7200. The review chain is the widest fan-in, so a
review stage that needs more than the default window times out
deterministically and is retried with the same window; the review group
halts on failure, so every later review stage and the final-report node
are then skipped (#1150).

Pin the review group in the four non-smoke topologies to the same 7200
seconds as the goal, strategy and specialist groups. smoke.yml has no
triage or severity node and is left alone. Like the existing group pins,
this one takes precedence over run.default_timeout_seconds and
model-profile timeouts (#675).

The new packaged-topology test expands every shipped topology and holds
each node whose prompt promises an extended timeout to a window above
the largest shipped default. It fails on the previous topologies with
"default.yml node `triage` promises an extended timeout but resolves to
ultrafuzz.toml `run.default_timeout_seconds`=3600".

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every property reference node in the four non-smoke topologies set
timeout_seconds: 300 (36 lines). Reference nodes are materialized by
plan-run before launch and compileSmithersWorkflow only compiles agentic
nodes into Smithers tasks, so the value never bounded anything. It was
only copied into the expanded graph fingerprint and the planned graph,
and displayed by the dashboard. The vulnerability-database reference
node already had no timeout.

Existing project topologies that still carry the lines keep validating;
the reference docs now say the field has no effect on reference nodes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…jects

The stateful-invariant setup and implement-properties prompts told the
agent not to edit production src/ or contracts/ "except for interfaces
if they are genuinely required by the harness". Both nodes publish a
workspace patch, and captureWorkspacePatch rejects every change under
the configured production source roots (src and contracts by default),
added files included:

  source-snapshot violation: workspace patch modifies protected
  production source: src/interfaces/IVault.sol

So with the default roots, following the exception failed the handoff,
and a retry starts a fresh session from the same prompt. Remove the
exception and tell the agent to declare harness interfaces in the test
tree, which the same capture accepts.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…m its built-in

Runs use the project copy of every prompt, and ultrafuzz init without
--force keeps an existing copy. After an upgrade a scaffolded copy
therefore keeps an older release's text while the artifact gates move
on, and nothing said so until a node failed its gate at the end of an
attempt.

projectPromptsDifferingFromBuiltIns lists the project prompts whose bytes
differ from the built-in prompt at the same path. validate reports each
one as a PROMPT_DIFFERS_FROM_BUILT_IN warning, which sets the prompts
posture to warn without failing validation, planning or the dashboard's
prompt save. A freshly scaffolded project and prompts the project adds
at new paths produce no warning.

The new runtime test fails on the previous validate.ts (prompts posture
'pass' where 'warn' is expected) and also checks the remedy the warning
names: deleting the copy and rerunning init restores a passing posture.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…st tree

The setup-foundry and base-test-setup prompts ask for Solidity interfaces
to the Vyper contracts' ABI but never say where to put them. Both nodes
publish a workspace patch, and captureWorkspacePatch rejects any change
under the production source roots (src and contracts by default), so an
interface written under contracts/interfaces/ fails the setup handoff,
and the setup group halts the run on failure. Say to declare new
interfaces in the test tree, as the stateful-invariant prompts now do.

Whether agents actually chose a production location here was not
observed; this closes the same gap the invariant prompts had.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…as a pass

doctor reported the validate check as `warning` but kept the summary
"config, topology, prompts, paths, agents, and trust posture pass". Now
that validate warns about project prompts that differ from their
built-in, every upgraded project with such a copy would see that
contradiction. The summary now says validation passed with warnings and
points to `ultrafuzz validate --json`, since text-mode output does not
print warnings.

The new test fails on the previous doctor.ts with actual 'config,
topology, prompts, paths, agents, and trust posture pass'.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…vers

Move projectPromptsDifferingFromBuiltIns below builtInPromptRelativePaths
so it no longer sits in the hunk where #1175 deletes getPrompt; with that
placement the two branches merge without a conflict in catalog.ts. No
behaviour change.

The anchor test comment claimed a prompt and its topology "cannot drift
apart again". The test checks only nodes whose built-in prompt promises
an extended timeout (triage and severity-classification) in the packaged
topologies, so say that, and that dedupe-findings is not covered.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
A group pin takes precedence over run.default_timeout_seconds and
model-profile timeout_seconds (#675), so a project that raised either
above 7200 now gets 7200 on review nodes of the packaged exhaustive and
invariant-only topologies and of new scaffolds, and a cloud config whose
explicit global resource timeout is below 7200 needs per-node overrides
for the review nodes as well. Record that under Breaking changes with
the escape hatch.

Tighten the Other changes entry: name the three pinned topologies instead
of "the shipped topologies" (smoke stays unpinned), and give the one-command
refresh for an uncustomized project topology, which init keeps. The
how-to now says the same about the topology as it already did about
prompts.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano force-pushed the claude/w12-topology-prompt-fixes branch from aba2cf4 to ba7bd6c Compare September 29, 2026 02:53
aviggiano added a commit that referenced this pull request Sep 29, 2026
getPrompt still has no caller, but #1164 adds
projectPromptsDifferingFromBuiltIns directly above it in catalog.ts, so
deleting it here made the two PRs conflict textually
(git merge-tree reported CONFLICT (content) in catalog.ts). catalog.ts is
outside this change's files, so restore it to origin/main and leave the
deletion for a follow-up once #1164 lands.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 150ce02 into main Sep 29, 2026
13 checks passed
@aviggiano
aviggiano deleted the claude/w12-topology-prompt-fixes branch September 29, 2026 06:33
aviggiano added a commit that referenced this pull request Sep 29, 2026
…rphaned by #1131 (#1175)

* refactor(prompts): delete the unused prompt rename feature and duplicate template resolver

renamePromptId, renameConcreteNodeId and rewriteArtifactPathForPromptId
(rename.ts), their render.ts helpers (renamePromptArtifactReferences and
three private rewriters) and the frontmatter round-trip support they
needed (serializePromptDocument, diffPromptIdentity, the
allowUnknownFields option, unknownFrontmatter/rawFrontmatter and the
invalid-rename error code) had no caller outside their own tests. getPrompt
had no caller at all. TemplateOccurrence.rawName existed only so rename
could preserve whitespace.

Output-contract templates had their own three-candidate resolver whose
middle candidate, dist/prompts, is the layout the build stopped producing
in #796. They now resolve through builtInPromptRoot(), the packaged-first
lookup the agent preamble templates already use, so both kinds of prompt
asset follow one rule. The packaged-layout render test, which copies the
real dist below a directory with no repository above it, still covers the
sealed-snapshot case; the installed-layout test that rebuilt the old
dist/prompts layout is removed and its content assertions are folded into
the packaged-layout test.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* refactor(config): delete the prompt-metadata layer and helpers with no production caller

- The prompt-metadata layer: resolveConfig applied
  createDefaultPromptMetadataLayer(), which always returned {}, and an
  input.promptMetadata that only one test supplied. PromptMetadataLayer,
  ResolveConfigInput.promptMetadata and the "prompt-metadata" diagnostic
  source go with it.
- restoreRedactedConfig and assertNoRedactionPlaceholders: nothing in
  runtime, cli, modal or evals calls them. Runs launch from the unredacted
  resolved config; the redacted copy and its manifest are written for
  readers of the run directory, and no command restores values from them.
  docs/config.md no longer describes a restore step or launch guard that
  does not exist. The "redaction" diagnostic source goes with them.
- applyDefaultProfileOverrides: a one-line wrapper only tests called.
  Its test now drives applyModelProfileOverrides, which validate.ts and
  plan-run.ts use.
- assertResolvedConfigZod and resolvedConfigSchemaEntry: no callers.
- resolvedConfigValidatorsAgree: moved into the schema test, its only
  user.

redactResolvedConfig and serializeRedactedResolvedConfigToml, which
plan-run.ts and init.ts use, are unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* ci: delete benchmark CI scripts orphaned by #1131

#1131 deleted eval-benchmarks.yml, the only workflow that invoked
scripts/ci/modal-benchmark-control-window.mjs (create/deadline) and
scripts/ci/validate-threat-model-benchmark-gate.mjs. Since then the
control-window script was kept alive only by a knip entry and a block of
ci-config.test.ts that exercised it, and the threat-model gate only by its
own test, which still ran in every PR's test:ci-scripts step.

Delete both scripts, the gate's test, the knip entry, and the
control-window block of the EVMBench full-mode manifest test; the
manifest assertions around it stay.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* refactor(modal): delete unreferenced exports and move test-only helpers into tests

Deleted, no reference anywhere: MAX_MODAL_DOCUMENT_BYTES,
MODAL_SMOKE_STOP_PATH (smoke-worker.ts builds its own stop path) and
DEFAULT_NODE_TIMEOUT_SECONDS.

Deleted with the tests that only exercised them:
- createTrackedSourceArchive; production archives the exact candidate
  with createExactCandidateSourceArchive.
- persistentWorkspaceRoot.
- locateModalResumeWorkspace, documented as "retained for tests"; worker.ts
  uses findModalResumeWorkspace, and the removed "locates one exact linked
  durable run" test duplicated the existing tolerant-lookup test. The two
  finalize tests unwrap findModalResumeWorkspace through a local helper.

Moved into tests, their only users:
- DEFAULT_BENCHMARK_MODELS, a stale model list no production code reads,
  is now test/model-spec-fixtures.ts, typed as a const tuple so indexed
  fixtures need no non-null assertion.
- modalBenchmarkConfigValidatorsAgree, WORKER_RESULT_ALLOWED_KEYS (the
  test now states the expected persisted key set), and
  publicEvalFailureDiagnosticLogPayload, which composed the same two calls
  runCommand already makes inline.
- PUBLIC_BENCHMARK_MAX_PARALLEL_EVAL_ROWS and its full-lane twin, aliases
  of the evals lane limits that only a test imported.

Un-exported, used only inside their own module: five public-worker
constants and WORKER_STDERR_TAIL_BYTES. The source-constant reader in
prepare-eval-history-publication.mjs matches top-level const
declarations with or without export, and its test against the real
public-worker.ts still passes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* refactor(dashboard): drop the always-false command capabilities

commandCapabilities() returned restartWholeRun, doctor, triage, merge,
restartFromNode, rerunSelectedNode and arbitraryShell as the constant
false. The frontend only declared them in its CommandCapabilities type:
capabilityForCommand never maps a command to any of them, so nothing read
them.

Remove them from the server, the dashboard HTTP schema's
commandCapabilities definition, and the frontend type. The schema edit
is required because that definition is closed (required plus
additionalProperties: false). The dashboard schema feeds neither the
validator build identity, which hashes only the artifacts validator
modules and the ajv versions, nor the artifact schema bundle digest that
planned outputs and the trusted CLI identity bind to.

The flow test's two assertions that doctor and merge are false are
removed. Every dashboard HTTP response in the tests is still validated
against the schema, so a server that emitted one of these keys again
would fail the flow test with "must NOT have additional properties"
(checked by re-adding doctor: false to the compiled server).

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* chore(runtime): drop the unused run-tests selectors

packages/runtime/package.json calls scripts/run-tests.mjs with no
selector or with "supporting"; nothing passes "materialize", "clean" or
"smithers". A single test file can still be selected by name through the
existing fallback.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs(changelog): record the dead-code deletions

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* refactor(prompts): leave getPrompt to the catalog work item

getPrompt still has no caller, but #1164 adds
projectPromptsDifferingFromBuiltIns directly above it in catalog.ts, so
deleting it here made the two PRs conflict textually
(git merge-tree reported CONFLICT (content) in catalog.ts). catalog.ts is
outside this change's files, so restore it to origin/main and leave the
deletion for a follow-up once #1164 lands.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

* docs(changelog): state the dead-code entry only as strongly as the evidence

Drop getPrompt, which this change no longer deletes. Name the actual
callers of the two CI scripts (eval-benchmarks.yml and the
watch-modal-benchmark.sh it ran, both deleted in #1131) instead of "the
workflows #1131 deleted": the other two workflows #1131 removed never
invoked either script. Stop describing the config prompt-metadata layer as
code no production path calls: resolveConfig ran it on every resolution,
only ever with an empty layer. Mention the run-tests.mjs selectors and the
docs/config.md change, and add the [runtime] and [docs] tags for them.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant