fix(runtime): workflow renders no longer crash on prompt text, and per-render cost drops - #1168
Merged
Merged
Conversation
…ow render
renderAgentPrompt substituted the operator and task prompts into the
trusted agent-prompt template and then regex-scanned the whole result for
leftover `{{word}}` placeholders. Inserted text that merely contained such
a sequence threw inside the Smithers render: an escaped `\{{word}}` example
in a prompt template (the prompt renderer emits it as literal `{{word}}`),
the same escape inside a dynamic item value (kept verbatim), or an
operator `--prompt` note. Every render builds every task's prompt, so the
run failed, and failed again on every resume.
Check only the template's own placeholders, inside the replace callback,
as renderAgentPreambleTemplate already does. A template placeholder with
no value still throws, and now names the placeholder.
The dynamic prompt renderer also resolves `{{...}}` inside goal-plan
replacement values, and an unbound name there throws in the same render.
The goal-plan contract now rejects `{{` in replacement values, which are
plain-text labels, so that failure lands on goal-plan's own verify instead
of on every render. The check is a zod refinement; goal-plan.schema.json
is unchanged.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
prepareArtifactMirror spawned `ultrafuzz json validate` against the smoke fixture in every prepare, in every agent-attempt reset, when an agent is built after a controller restart, and again in the zero-retry verify task. The trusted launcher's cold start was measured at ~35 s under contention (#1026), so each spawn was another chance to fail an attempt, including one whose agent work had already finished. What the spawn proves -- that this process can launch the agent-facing validator -- does not depend on the task: materializePromptSchemas has just digest-checked the workspace's schema copy, and start-run already runs the trusted-launcher preflight at launch and resume. Remember the first success in the engine process. A failure is not remembered, so the next caller spawns again. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…orkflow agent-registry.ts and smithers.ts imported the TypeScript compiler at module load, and both are reachable from @ultrafuzz/runtime's index. Only registry analysis (validate and init) and the controller-refresh helper that only tests call ever use it. Require it at first use instead. The saving lands in processes that import the runtime without smthrs: the ultrafuzz CLI, including every agent `ultrafuzz json validate` call. Smithers engine processes still load the compiler: the generated workflow imports smthrs, whose index re-exports @smthrs/scorers, and two scorer modules import typescript at module load. Importing the built runtime index alone on this host went from 528-547 ms / 248-250 MB RSS to 409-437 ms / 196-198 MB under Node 24, and from 449-483 ms / 252-261 MB to 336-351 ms / 209-213 MB under Bun 1.3.14 (5 runs each). Changing the parameter type on topLevelNameIsBound's signature makes the strict diff lint report that function's existing complexity, so its import-clause check moves into a helper with the same conditions. materializeDynamicRuntime runs on every render and rewrote tasks.json and graph.json, with fsyncs, even when it re-derived the bytes already on disk. Skip the write in that case. taskSpecsFromCompiled looked every task up with serializedTaskSpecs.find; index the specs by id once per call instead. At the default topology's 65 static tasks the difference is not measurable. Add a test that compiles the packaged default topology and holds the generated workflow under 2.3 MB, measured with a fixed-length project root (about 2.06-2.07 MB today, depending on the checkout path), so the size #1146 describes cannot grow silently. Refs #1146 Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Since the external static-analysis gate was adopted (#1005), super-linter lints every changed Markdown file with MD013 at 400 characters. CHANGELOG.md keeps one entry per line, and markdownlint reports 42 existing entries on main as longer than that, so any change to the file fails `External static analysis`. That also fails `release-gates` and skips the release validation lanes on the pull request. No commit has changed CHANGELOG.md since the gate landed. Disable MD013 below the title of this file only, using the same directive line that #1169, #1172 and #1181 add, so the branches do not end up with two spellings of it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…idator CLI finalizeAndVerifyArtifacts calls prepareArtifactMirror with pinnedSubmodules "verify", and that pass still ran the `ultrafuzz json validate` preflight. Verify checks the finished agent's outputs in-process and never runs the agent-facing CLI, and its task has no retry. The per-process memo only helps later callers: after an engine restart, the first preflight in the new process can be the verify of an attempt whose agent had already finished, and a CLI cold start that timed out there (the trusted launcher measured ~35 s under contention in #1026) failed verification and discarded that work. Skip the preflight in the verify pass. prepare:* and the reset before each agent attempt, which run ahead of an agent that uses the CLI, keep it. The new lifecycle test runs the production preparation body with the options finalizeAndVerifyArtifacts passes. Without this change it counts three preflights instead of two. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…apply
The comments, the footprint test and the changelog said Smithers engine
processes stop loading the TypeScript compiler. They do not: the
generated workflow imports smthrs, whose index re-exports
@smthrs/scorers, and its sideEffectAnalysis and workflowUiCompliance
modules import typescript at module load. The saving lands in ultrafuzz
CLI processes, including every agent `ultrafuzz json validate` call.
The prompt-variables reference and the changelog also said the
goal-plan@1 contract rejects `{{` in replacement values. Only goal-plan's
own verification (and the evals goal-plan reader) run that zod rule;
`ultrafuzz json validate` and `ultrafuzz artifact validate` accept such a
plan. Say so, replace the exact generated-workflow byte count, which does
not reproduce across environments, with a range, and describe the
preflight as running until its first success.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
importClauseBindsName, split out of topLevelNameIsBound on this branch, had no test: the registry-shadowing test only declared Object locally. Add default, namespace, named, aliased and default-plus-named imports of Object, each of which must stop Object.freeze from being read as the global, and an unrelated named import that must not. Removing any one of the helper's three checks, or making it match every named import, fails the test. Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/w15b-render-robustness
branch
from
September 29, 2026 03:38
59f7fe2 to
474f4f8
Compare
Every pull request in this batch inserts its entry at the same place in CHANGELOG.md, so each merge would conflict with the next. The entries are collected into one changelog update instead. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Prompt text could fail the whole run. Any literal
{{word}}that reached an agent prompt made the Smithers render throw. Three sources can produce one:\{{word}}example in a prompt template, which the prompt renderer emits as literal{{word}};--promptnote.Every render builds the prompt of every task it renders, so one such prompt fails the render (
WORKFLOW_RENDER_FAILED). The run stops, and every resume fails the same way. Goal-plan replacement values had a second route:Late tick lets {{amount}} round downpassesvalidateGoalPlan, and then the dynamic prompt renderer throwsdynamic item value liquidation:overdue references non-item variable: amountinside the same render. I reproduced both throws at function level onorigin/main.The validator CLI self-test ran over and over.
prepareArtifactMirrorspawnedultrafuzz json validatein everyprepare:, in every agent-attempt reset, when an agent was built after a controller restart, and again in the zero-retryverify:task. The trusted launcher's cold start was measured at ~35 s under contention (Smoke lane reaches agent work but no target can score: 15s validator preflight budget and a self-contradicting final-report prompt #1026), so each spawn was one more way to fail an attempt, including a verify whose agent work had already finished.Per-process and per-render cost.
ultrafuzzCLI process imports@ultrafuzz/runtime, and that includes every agent validator command. That import loaded the TypeScript compiler. (Smithers engine processes import the runtime index too, but they load the compiler throughsmthrsregardless; see the Enforce aggregate memory limits and avoid type-checking huge generated workflows #1146 section.)tasks.jsonandgraph.json, even when nothing had changed.taskSpecsFromCompiledlooked up each task with a linearfind.Root cause
renderAgentPrompt(workflow.tsx, lines 1673-1688 on main) substituted the values first, then regex-scanned the substituted result. So it validated the inserted text rather than the template.ultrafuzz/goal-plan@1zod contract accepted replacement values containing template syntax, which the dynamic renderer resolves and can fail on.preflightJsonValidatorran on every call, although its answer only depends on the process. It also ran in the post-agent verify pass, which never uses the CLI.materializePromptSchemasdigest-checks each workspace's schema copy just before it, andstart-runalready runs the trusted-launcher preflight at launch and resume.agent-registry.tsandsmithers.tsboth had a module-scopeimport * as ts from "typescript", and both are reachable from the runtime index. Only registry analysis (validate/init) andassertRefreshedModuleAuthorityuse it. The latter is reached only throughrefreshedSmithersControllerSnapshot, which has no caller inpackages/*/src.materializeDynamicRuntime's publish mode calledwriteJsonDurableunconditionally.Change
renderAgentPromptchecks the template's own placeholders inside thereplacecallback, the same wayrenderAgentPreambleTemplatedoes. Inserted text is never re-scanned. A template placeholder with no value still throws, and the error now names it.goal-plan.tsadds a zod refinement: replacement values must not contain{{. A bad label now failsgoal-plan's own verify (assertGoalPlan, reached frommaterializeGoalPlanDatabaseArtifacts) instead of every render. The evalsreadGoalPlanapplies the same zod contract. The agent-facingultrafuzz json validate --schema goal-plan.schema.jsonandultrafuzz artifact validate ultrafuzz/goal-plan@1do not run it: both returnok: truefor such a plan (I ran both), the same gap the existing JSON-record refinement already has.z.toJSONSchemadrops refinements, sogoal-plan.schema.jsonand the validator build identity are unchanged. retry-failed replays failed verifier without rerunning artifact producer #925 maderesume --retry-failedreopen a failed verifier's producer; I did not re-test that path here.docs/reference/prompt-variables.mdsays where the rule runs.preflightJsonValidatorremembers its first success for the engine process, and a failure is not remembered. The post-agent verify pass (finalizeAndVerifyArtifactscallsprepareArtifactMirrorwithpinnedSubmodules: "verify") no longer runs it at all.prepare:*and the reset before each agent attempt keep it, because they run ahead of an agent that uses the CLI.createRequired on first use inagent-registry.tsandsmithers.ts. The value import becameimport type, soagent-registry.tsrenames its type annotations toTypeScript.*. That touched the signature line oftopLevelNameIsBound, so the strict diff lint reported the function's existing complexity (25). I moved its import-clause check intoimportClauseBindsName, with the same conditions, and added test cases for it (item 9). The saving lands inultrafuzzCLI processes, including every agent validator call. It does not reach Smithers engine processes.materializeDynamicRuntimeskips the write when the file already holds the exact byteswriteJsonDurablewould write.taskSpecsFromCompiledindexes the serialized specs by id once per call.generated-workflow-footprint.test.ts. It compiles the packaged default topology and asserts the generated workflow stays under 2,300,000 bytes. It measures with a fixed-length project root, because the absolute root appears 4,841 times and would otherwise make the budget depend on the length of$TMPDIR. The normalized size is about 2.06-2.07 MB; the exact figure differs between environments (2,060,237 bytes on this host at the current head, 2,070,382 in the first round's environment). It also asserts that importing the runtime index does not loadtypescript.CHANGELOG.mdgets<!-- markdownlint-disable MD013 -->below its title. The first push failedExternal static analysis. Super-linter lints every changed Markdown file with MD013 at 400 characters, and markdownlint reports 42 existing entries on main as longer than that. So any change to the file fails the check, which failsrelease-gatesand skips the release validation lanes. No commit has touchedCHANGELOG.mdsince that gate landed in chore: adopt external static-analysis gates #1005. The directive line is now spelled exactly as fix(runtime): resume --retry-failed reopens descendants skipped behind a recovered producer #1169, fix(runtime): name the runner error when a workflow fails with no failed node, and pin same-id recovery #1172 and refactor(artifacts): delete never-dispatched gates, production-dead code, and the unread event index #1181 add it, so the file does not end up with two spellings of the policy. With it, the file passes super-linter's pinned markdownlint config ([email protected]locally: 43 MD013 errors without the directive, 0 with it). This is a lint-policy call, kept to one line in its own commit, so it is easy to drop if you would rather reflow the file.Objectlocally. It now also covers default, namespace, named, aliased and default-plus-named imports ofObject, plus an unrelated named import that must not count as shadowing.#1146 investigation: does anything type-check the generated workflow?
No hot path I could find does. Nothing to remove or bound.
packages/*/src: nothing runstsc,ts.createProgram,ts.transpileModule,Bun.buildorBun.Transpilerover the generated workflow. The only TypeScript compiler API calls in production code are:packages/runtime/src/agent-registry.ts:53on main (ts.createSourceFileover the at-most-256 KiB.smithers/agents/index.ts, called fromvalidateandinit);packages/runtime/src/smithers.ts:3760on main (ts.preProcessFileover sealednode_modules.jsfiles inassertRefreshedModuleAuthority, with no production caller).This PR makes both lazy.
Pinned Smithers 0.35.0 only parses the workflow:
@smthrs/engine/src/workflow-hash.js:59-61and@smthrs/time-travel/src/validateWorkflowIdentity.js:15-26callBun.Transpiler#scanImports;@smthrs/cli/src/packs.js:186-188does the same for packs;.tsxwhen the engine imports it.@smthrs/cli/src/workflow-pack.js:165holdstypecheck: "tsc --noEmit"only as a script string for generated packpackage.jsonfiles; ultrafuzz never runs it.Every engine process still loads the TypeScript compiler, and this PR cannot change that. The generated workflow's
import { createSmithers } from "smthrs"(workflow.tsx:23, and the agent templates importsmthrsthe same way) loads[email protected]/src/index.js. Itsexport { … } from "@smthrs/scorers"(lines 506-520) loads the scorers index, which re-exportsworkflowUiCompliance.jsandsideEffectAnalysis.js(@smthrs/scorers/src/index.js:54-55). Both start withimport ts from "typescript"(workflowUiCompliance.js:2,sideEffectAnalysis.js:1).sideEffects: falsein the package manifest only affects bundlers. I checked it with arequire.cacheprobe: afterimport("smthrs")alone,[email protected]/lib/typescript.jsis loaded under both Bun 1.3.14 and Node 24.Measured parse cost over the ~2.07 MB default workflow with Bun 1.3.14 (first round):
scanImports7-9 ms,transformSync14-15 ms, process RSS under 90 MB.In this repo, only tests call
ts.transpileModuleon the template or on generated output. That includesdynamic-workflow.test.ts:153, which transpiles a whole generated workflow.Conclusion. One engine or CLI process does not come near "tens of GB". My inference, which I have not verified, is that such a figure needs an external type-checker (an agent running
tsc, or an editor LSP) over.smithers/workflows/*.tsx. The budget test is the guard added here.Deliberately not built (and why)
Dropping the duplicated task-spec literal. It is 886 KB of the file, next to the 702 KB compiled-tasks literal, and it is the real size lever for Enforce aggregate memory limits and avoid type-checking huge generated workflows #1146. But it changes the generated-program layout and the helpers the template destructures. The budget test only stops growth.
Catching per-task render errors in
renderReadyRuntimePrompts, so that one bad dynamic prompt fails only its own task. That needs a new way to represent a task that failed at render. The goal-plan refinement moves the one model-authored trigger I found in the default topology to its producer instead.Fixing the dynamic renderer instead of adding the goal-plan rule (review suggestion: keep a non-item reference inside a value as literal text, in
resolveDynamicVariable). I ran the builtrenderPromptwith a{{item.goal_prompt}}→{{liquidation:overdue}}chain and four labels:Late tick lets {{amount}} round downthrowsreferences non-item variable: amount, which that change would fix;Late tick {{ never closedthrowsunclosed-template-variable;Tick {{}} roundingthrowsempty-template-variable;Tick {{item.missing}} roundingthrowsmissing dynamic item template variable.The one-branch change fixes one of these four, and the refinement rejects all four. A renderer change that covers them all means not scanning replacement values at all. That changes the documented nested-replacement semantics ("unresolved, cyclic, non-scalar, or non-item references fail") for custom topologies, in
render.ts, which refactor: delete dead prompt rename, config helpers, and CI scripts orphaned by #1131 #1175 is also editing. The refinement's costs are real, and I am leaving them in place:Overdue {{item.node_id}};goal-planis in thesetupgroup, which has nofailure_policy: continue, so a rejected plan stops the run untilresume --retry-failed;Rejecting escaped braces in
goal_prompt. After change 1 they render as literal text and no longer throw.Editing the goal-plan prompt. It already asks for "one line of plain prose", and the verify error names the rule.
Deleting the workflow preflight outright, or running it only in
createmode (review suggestions). After this change it spawns the CLI only until its first success in each engine process, and only fromprepare:*(which has at least one retry) and the reset before an agent attempt. On resume,start-runswallows a failed trusted-launcher preflight (submitSmithersContinuation, start-run.ts:584-599). On a resumed run, the engine check is then the only thing that confirms, before an agent spends an attempt, that theultrafuzzit will call works. Acreate-only gate would skip the pre-attempt reset. In a restarted engine process that reset runs before the lazy agent rebuild, and itsassert-task-inputsstep records the admission that makes the rebuild unnecessary, so the process's first agent attempt could go unchecked.Sharing fixture setup between the new dynamic-expansion test and the existing retry test. fix(runtime): published dynamic expansions are the lock-free authority for fan-out membership #1163 changes about 250 lines of that file, and a shared helper would conflict with it.
Deleting the dead controller-refresh code, the other TypeScript user. It is out of scope, and several parallel items edit
smithers.ts.cgroup or process-tree memory budgets (Enforce aggregate memory limits and avoid type-checking huge generated workflows #1146's main ask).
Verification
Discriminating tests. Each new behavioural test fails on
origin/mainand passes on this branch. Both reviewers re-ran the first five againstb6dd1da9; I ran the sixth there:origin/maingenerated agent prompt inserts literal braces from task and operator prompts verbatim(runtimegenerated-workflow-verifier.test.ts)agent prompt template contains an unresolved variablegenerated validator preflight spawns the CLI once per engine process and never remembers a failure(same file)goal-plan replacement values cannot carry template placeholders into the dynamic render(artifactsthreat-goal-artifacts.test.ts)Late tick lets {{amount}} round downre-rendering an unchanged dynamic runtime does not replace its published task plan or graph(runtimedynamic-expansion.test.ts)importing the runtime package does not load the TypeScript compiler(newgenerated-workflow-footprint.test.ts)typescript.jsinrequire.cacheright after importthe post-agent verify pass does not preflight the agent-facing validator CLI(runtimeworkspace-preparation-lifecycle.test.ts)b6dd1da9, and on this branch with the template change stashed)The last test runs the production
prepareArtifactMirrorbody, with its external CLI and submodule work stubbed, using the three option shapes the template passes. The verify-shaped call is the onefinalizeAndVerifyArtifactsmakes, and an existing verifier test pins that call's options.The byte-budget test is a regression guard, so it passes on both.
The registry import cases (item 9) cover a refactor, so they pass on
origin/maintoo. To check that they actually exercise the helper, I made four one-line mutations to the compiledimportClauseBindsName. Removing the default-import check fails the test onimport Object from …. Removing the namespace check fails it onimport * as Object …. Removing the named-import check fails it onimport { Object } …. Matching every named import fails it on the unrelated-import control.The two template tests run
renderAgentPromptandpreflightJsonValidatorextracted fromworkflow.tsxwith stubbed collaborators, which is how the existing verifier tests work. No test here runs the generated workflow under the real engine.TypeScript load, checked with a
require.cacheprobe on this host:node packages/cli/dist/index.js json validate --schema findings.schema.json --file validator-smoke.valid.json --jsonloadstypescriptonorigin/mainand does not on this branch.smthrsalone loads it under both Bun and Node, on either branch.First-round measurements on this 32-core host, shared with other jobs:
origin/mainjson validatecommand above, runs 2-5 of 5The runtime-index rows do not describe an engine process, which loads
smthrsas well.What I ran on the current head:
pnpm -w build,pnpm --filter @ultrafuzz/runtime typecheckandpnpm --filter @ultrafuzz/artifacts typecheck.generated-workflow-verifier(142/142);workspace-preparation-lifecycle(7/7);generated-workflow-footprint(2/2);cloud-worker-handoff,dynamic-expansion,dynamic-lifecycle,dynamic-workflow,task-workflow-identityandworkflow-dependency-policytogether (58/58,--test-concurrency=2);runtime.test.tsregistry-shadowing test.npx prettier --checkandnpx eslinton every changed file, thenCI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci,pnpm -w knip,node scripts/docs-check.mjs,node scripts/audit-profile-docs.mjs --checkandnode scripts/prompt-catalog-docs.mjs --check.[email protected]with super-linter's config onCHANGELOG.mdanddocs/reference/prompt-variables.md.First round, not repeated on this head. This round's only code change is the verify-pass gate in the template. Everything else is comments, docs and tests, and the artifacts package did not change:
runtime.test.tstests: the registry-analysisvalidate/inittests, the controller-refresh tests that reach the lazyts.preProcessFile,compileSmithersWorkflow*, andstartRun compiles normal Smithers tasks…;threat-goal-artifacts(25/25) andschema.test.js(51/51, including the goal-plan JSON Schema snapshot).PR smoke lane. It passed in CI on the first-round heads and on this head. When I ran it locally in the first round, 2 of its 12 named
runtime.test.tstests failed withWORKFLOW_SUBMISSION_FAILED ... changed while it was readon files under this worktree's sharednode_modules. Both passed when re-run on their own with--test-concurrency=1. A reviewer saw the same class of flake indynamic-lifecycle.CI on
474f4f85: every check passes. That includesExternal static analysis, the PR smoke lane, all four runtime integration shards, the runtime support tests with the Bun 1.3.14 adapter contracts, andrelease-gates. Integration shard 1 took 47 min, against about 30 min on59f7fe25; I did not look into why.Not run anywhere: the CLI suite and any real campaign.
Risk / compatibility
renderAgentPrompt, the preflight memo, the verify-pass skip and the task-spec index therefore take effect for new runs or re-rendered controllers. A native resume keeps the old template code.goal-plan's own verify. EvalsreadGoalPlanapplies it too, and turns a validation failure intogoal-plan-unreadableevidence rather than throwing. So a historical plan with{{in a value would read as unreadable there.CHANGELOG.mddirective changes lint policy for that one file only (item 8).writeJsonDurableproduces (JSON.stringify(value, null, 2) + "\n"), so the on-disk format is unchanged.Refs #1146
🤖 Generated with Claude Code
The PR appears safe to merge; no new actionable issue was established in the changes since the previous review.
Summary
The PR changes prompt rendering so literal braces in inserted text do not fail a workflow render, validates goal-plan replacement labels earlier, and reduces validator startup and repeated-render costs.
Reviews (4) · Last reviewed commit: "chore: move the changelog entry to the c..."